“Let’s just add more servers.”

That sentence has burned more cloud budget than almost any bug. A team sees a slow page, doubles the instance size, the bill doubles, and the page is still slow. Why? Because nobody found out what was slow.

This short guide gives you a repeatable process: measure, find the bottleneck, fix one thing, verify, repeat. Reading time: about 12 minutes.


The big picture

1. Set a goal "p95 under 300ms" not "make it fast" 2. Measure Real traffic, real data, a baseline 3. Find it Profile. Locate the one big bottleneck 4. Fix one thing Smallest change that moves the number 5. Verify Re-measure. Goal met? Stop. Else loop Goal not met: measure again, the bottleneck has moved
Every wasted dollar in performance work comes from skipping step 2 or 3.

Rule #1: never optimize without a number

“Premature optimization is the root of all evil.” — Donald Knuth

The full quote says we should ignore small efficiencies about 97% of the time, and focus on the critical 3%. Your job is to find that 3%.

Write the goal down first. A good goal is specific:

Vague goal Useful goal
“Make the dashboard faster” “Dashboard p95 under 400 ms at 200 concurrent users”
“Reduce server cost” “Cut API cost per 1,000 requests by 30%”
“Fix the slow report” “Monthly report generates in under 10 s”

Use p95/p99, not averages. An average of 200 ms can hide the 5% of users waiting 6 seconds.


Real-world failure #1: the Go rewrite that changed nothing

A team had a slow checkout endpoint (3.2 s). They decided PHP was the problem and spent four months rewriting it in another language. Result: 3.0 s.

Why? When they finally profiled it, 2.9 seconds was one database query waiting on a missing index. The language was never the bottleneck.

Lesson: the slow part is usually I/O (database, network, disk), not your language. Profile before you rewrite anything.


Step 1: Measure the whole request

Before touching code, see where time goes. Think in layers:

Typical slow request: 3,200 ms Database 70% APIs 15% Code What most teams optimize first They polish the blue slice (10%). Even a perfect result saves at most 320 ms. The red slice (70%) is where 2,240 ms is hiding.
Illustrative numbers. Your split will differ, which is exactly why you measure.

Free and cheap tools to start with:

Layer Tool
Browser Chrome DevTools, Lighthouse
Laravel app Telescope, Debugbar, Clockwork
PHP profiling Xdebug, Blackfire, SPX
Database EXPLAIN, slow query log
Production New Relic, Datadog, Sentry Performance
Load testing k6, Apache Bench, Locust

A tiny timer is enough to start:

$start = microtime(true);

$orders = $this->orderService->summaryFor($user); // suspect

Log::info('order summary', [
    'ms' => round((microtime(true) - $start) * 1000, 1),
]);

Step 2: Fix the database first

In most web apps, the database is the biggest slice. It is also the cheapest to fix. Three problems cover most cases.

Problem A: the N+1 query

Real-world failure #2: an admin page listed 100 orders and showed each customer name. It worked in development with 5 rows. In production it ran 101 queries per page load, and under load the database CPU hit 100%. The team’s first reaction was to buy a bigger database.

// Bad: 1 query for orders + 1 query per order = 101 queries
$orders = Order::latest()->take(100)->get();

foreach ($orders as $order) {
    echo $order->customer->name; // hidden query each loop
}
// Good: 2 queries total
$orders = Order::with('customer')->latest()->take(100)->get();

Guard against it permanently, so it fails loudly in development instead of silently in production:

// AppServiceProvider::boot()
Model::preventLazyLoading(! app()->isProduction());

I covered this in depth in Laravel Eloquent: Lazy Loading, Eager Loading and N+1.

Problem B: the missing index

Real-world failure #3: a payments table grew from 10k to 8 million rows. A query filtering by user_id went from 5 ms to 2.5 s because the database had to read every row (a full table scan).

EXPLAIN SELECT * FROM payments WHERE user_id = 42;
-- type: ALL   rows: 8,000,000   <-- full scan, bad
// migration
Schema::table('payments', function (Blueprint $table) {
    $table->index('user_id');
});
-- after the index
-- type: ref   rows: 37          <-- reads only what it needs

Rule of thumb: index columns you use in WHERE, JOIN, and ORDER BY. Do not index everything, because every index slows writes and uses disk.

Problem C: fetching more than you need

// Bad: loads every column and every row into memory
$users = User::all()->filter(fn ($u) => $u->is_active);

// Good: filter in SQL, select only what you use, page the results
$users = User::query()
    ->where('is_active', true)
    ->select('id', 'name', 'email')
    ->paginate(50);

For big batch jobs use chunkById() or lazyById() so memory stays flat. This also prevents the memory problems described in PHP Memory Management and Memory Leaks.

Measured result

Baseline 3200 ms + Eager loading 1400 ms + Index 380 ms + Caching 90 ms Three small changes. No new servers, no rewrite.
Each fix was re-measured before moving to the next one.

Step 3: Cache what is expensive and rarely changes

Caching is powerful, but it is the step that causes the nastiest bugs. Use it after fixing queries, never instead of it. Caching a slow query just hides it until the cache expires.

$stats = Cache::remember('dashboard:stats:'.$user->id, now()->addMinutes(10), function () use ($user) {
    return $this->reportService->buildStats($user); // expensive
});

Real-world failure #4: the cache stampede

A popular homepage widget was cached for 5 minutes. When the key expired, 2,000 concurrent requests all found the cache empty and ran the same 4-second query at once. The database fell over, and the site went down every 5 minutes like clockwork.

Without a lock 2,000 requests Cache: EMPTY DB: 2,000 queries With a lock 2,000 requests 1 rebuilds, rest wait DB: 1 query
Hot keys need a rebuild lock, or the database pays for every request.

The fix in Laravel is an atomic lock:

$stats = Cache::remember('home:widget', 300, function () {
    return Cache::lock('home:widget:lock', 10)->block(5, function () {
        return $this->buildWidget(); // only one process runs this
    });
});

Also remember the second famous problem: stale data. Cache keys must be cleared when the underlying data changes.

// Invalidate on write, not "hope it expires"
protected static function booted(): void
{
    static::saved(fn ($p) => Cache::forget("product:{$p->id}"));
}

Step 4: Move slow work out of the request

If a user does not need the result right now, do not make them wait for it. Sending email, generating PDFs, and calling slow third-party APIs belong in a queue.

// Bad: user waits 6 seconds for the email provider
public function store(Request $request)
{
    $order = Order::create($request->validated());
    Mail::to($order->user)->send(new OrderReceipt($order)); // slow
    return redirect()->route('orders.show', $order);
}
// Good: respond in milliseconds, work happens in the background
public function store(Request $request)
{
    $order = Order::create($request->validated());
    Mail::to($order->user)->queue(new OrderReceipt($order));
    return redirect()->route('orders.show', $order);
}

If a job can run twice (retries happen), make it safe to repeat. See Idempotency in Payment Systems.


Step 5: Only now look at code and infrastructure

If the database, cache, and queue are healthy and you still miss your goal, then look at:

  • Algorithms: an O(n²) loop over 50,000 items beats any server upgrade in cost.
  • Framework config: in production run php artisan config:cache, route:cache, view:cache, and enable OPcache.
  • Payload size: gzip/brotli, smaller images, a CDN for static assets.
  • Scaling: add instances only when a single instance is efficient but traffic is genuinely higher.
// O(n²): scans the whole array for every item
foreach ($orders as $o) {
    if (in_array($o->customer_id, $blockedIds)) { /* ... */ }
}

// O(n): constant-time lookup
$blocked = array_flip($blockedIds);
foreach ($orders as $o) {
    if (isset($blocked[$o->customer_id])) { /* ... */ }
}

Real-world failure #5: scaling a leak

An API’s memory grew steadily until the container was killed every few hours. The team doubled the memory limit, which turned a 3-hour crash into a 6-hour crash. A profile showed a static array that cached every request’s data and was never cleared. The fix was 3 lines. The extra memory had cost money for two months.

Scaling hardware multiplies an inefficiency. It does not remove it.


The priority order (cheapest and highest impact first)

0. Measure and set a goal 1. Database: N+1, indexes, selects 2. Cache hot, rarely-changing data 3. Queue slow work, async external calls 4. Code, algorithms, config, CDN 5. Scale infrastructure (last, most expensive)
Work from the top down. Stop the moment your goal is met.

Common mistakes checklist

  • Optimizing on your laptop with 50 rows instead of production-like data.
  • Changing five things at once, so you cannot tell what helped.
  • Trusting averages instead of p95/p99.
  • Adding cache with no invalidation plan.
  • Buying bigger servers before reading a single query plan.
  • Never stopping. If the goal is met, stop. More speed past the goal costs time and adds complexity.
  • Forgetting to guard the win: add a performance test or alert so it does not regress next sprint.

Key takeaways

  1. Set a numeric goal before you start.
  2. Measure first. Guessing is the most expensive optimization.
  3. Fix the biggest bottleneck only, then measure again.
  4. Database before cache, cache before code, code before servers.
  5. Cache carefully: locks for hot keys, invalidation on writes.
  6. Don’t make users wait for work they don’t need right now.
  7. Stop when the goal is met.

Performance work is not about being clever. It is about being disciplined: one measurement, one fix, one verification at a time. Do that, and you will save both your users’ time and your company’s money.

Thanks for reading. If this helped, share it with a teammate who is about to say “let’s just add more servers.”