“Let’s just add more servers.”
That sentence has burned more cloud budget than almost any bug. A team sees a slow page, doubles the instance size, the bill doubles, and the page is still slow. Why? Because nobody found out what was slow.
This short guide gives you a repeatable process: measure, find the bottleneck, fix one thing, verify, repeat. Reading time: about 12 minutes.
The big picture
Rule #1: never optimize without a number
“Premature optimization is the root of all evil.” — Donald Knuth
The full quote says we should ignore small efficiencies about 97% of the time, and focus on the critical 3%. Your job is to find that 3%.
Write the goal down first. A good goal is specific:
| Vague goal | Useful goal |
|---|---|
| “Make the dashboard faster” | “Dashboard p95 under 400 ms at 200 concurrent users” |
| “Reduce server cost” | “Cut API cost per 1,000 requests by 30%” |
| “Fix the slow report” | “Monthly report generates in under 10 s” |
Use p95/p99, not averages. An average of 200 ms can hide the 5% of users waiting 6 seconds.
Real-world failure #1: the Go rewrite that changed nothing
A team had a slow checkout endpoint (3.2 s). They decided PHP was the problem and spent four months rewriting it in another language. Result: 3.0 s.
Why? When they finally profiled it, 2.9 seconds was one database query waiting on a missing index. The language was never the bottleneck.
Lesson: the slow part is usually I/O (database, network, disk), not your language. Profile before you rewrite anything.
Step 1: Measure the whole request
Before touching code, see where time goes. Think in layers:
Free and cheap tools to start with:
| Layer | Tool |
|---|---|
| Browser | Chrome DevTools, Lighthouse |
| Laravel app | Telescope, Debugbar, Clockwork |
| PHP profiling | Xdebug, Blackfire, SPX |
| Database | EXPLAIN, slow query log |
| Production | New Relic, Datadog, Sentry Performance |
| Load testing | k6, Apache Bench, Locust |
A tiny timer is enough to start:
$start = microtime(true);
$orders = $this->orderService->summaryFor($user); // suspect
Log::info('order summary', [
'ms' => round((microtime(true) - $start) * 1000, 1),
]);
Step 2: Fix the database first
In most web apps, the database is the biggest slice. It is also the cheapest to fix. Three problems cover most cases.
Problem A: the N+1 query
Real-world failure #2: an admin page listed 100 orders and showed each customer name. It worked in development with 5 rows. In production it ran 101 queries per page load, and under load the database CPU hit 100%. The team’s first reaction was to buy a bigger database.
// Bad: 1 query for orders + 1 query per order = 101 queries
$orders = Order::latest()->take(100)->get();
foreach ($orders as $order) {
echo $order->customer->name; // hidden query each loop
}
// Good: 2 queries total
$orders = Order::with('customer')->latest()->take(100)->get();
Guard against it permanently, so it fails loudly in development instead of silently in production:
// AppServiceProvider::boot()
Model::preventLazyLoading(! app()->isProduction());
I covered this in depth in Laravel Eloquent: Lazy Loading, Eager Loading and N+1.
Problem B: the missing index
Real-world failure #3: a payments table grew from 10k to 8 million rows. A query filtering by user_id went from 5 ms to 2.5 s because the database had to read every row (a full table scan).
EXPLAIN SELECT * FROM payments WHERE user_id = 42;
-- type: ALL rows: 8,000,000 <-- full scan, bad
// migration
Schema::table('payments', function (Blueprint $table) {
$table->index('user_id');
});
-- after the index
-- type: ref rows: 37 <-- reads only what it needs
Rule of thumb: index columns you use in WHERE, JOIN, and ORDER BY. Do not index everything, because every index slows writes and uses disk.
Problem C: fetching more than you need
// Bad: loads every column and every row into memory
$users = User::all()->filter(fn ($u) => $u->is_active);
// Good: filter in SQL, select only what you use, page the results
$users = User::query()
->where('is_active', true)
->select('id', 'name', 'email')
->paginate(50);
For big batch jobs use chunkById() or lazyById() so memory stays flat. This also prevents the memory problems described in PHP Memory Management and Memory Leaks.
Measured result
Step 3: Cache what is expensive and rarely changes
Caching is powerful, but it is the step that causes the nastiest bugs. Use it after fixing queries, never instead of it. Caching a slow query just hides it until the cache expires.
$stats = Cache::remember('dashboard:stats:'.$user->id, now()->addMinutes(10), function () use ($user) {
return $this->reportService->buildStats($user); // expensive
});
Real-world failure #4: the cache stampede
A popular homepage widget was cached for 5 minutes. When the key expired, 2,000 concurrent requests all found the cache empty and ran the same 4-second query at once. The database fell over, and the site went down every 5 minutes like clockwork.
The fix in Laravel is an atomic lock:
$stats = Cache::remember('home:widget', 300, function () {
return Cache::lock('home:widget:lock', 10)->block(5, function () {
return $this->buildWidget(); // only one process runs this
});
});
Also remember the second famous problem: stale data. Cache keys must be cleared when the underlying data changes.
// Invalidate on write, not "hope it expires"
protected static function booted(): void
{
static::saved(fn ($p) => Cache::forget("product:{$p->id}"));
}
Step 4: Move slow work out of the request
If a user does not need the result right now, do not make them wait for it. Sending email, generating PDFs, and calling slow third-party APIs belong in a queue.
// Bad: user waits 6 seconds for the email provider
public function store(Request $request)
{
$order = Order::create($request->validated());
Mail::to($order->user)->send(new OrderReceipt($order)); // slow
return redirect()->route('orders.show', $order);
}
// Good: respond in milliseconds, work happens in the background
public function store(Request $request)
{
$order = Order::create($request->validated());
Mail::to($order->user)->queue(new OrderReceipt($order));
return redirect()->route('orders.show', $order);
}
If a job can run twice (retries happen), make it safe to repeat. See Idempotency in Payment Systems.
Step 5: Only now look at code and infrastructure
If the database, cache, and queue are healthy and you still miss your goal, then look at:
- Algorithms: an
O(n²)loop over 50,000 items beats any server upgrade in cost. - Framework config: in production run
php artisan config:cache,route:cache,view:cache, and enable OPcache. - Payload size: gzip/brotli, smaller images, a CDN for static assets.
- Scaling: add instances only when a single instance is efficient but traffic is genuinely higher.
// O(n²): scans the whole array for every item
foreach ($orders as $o) {
if (in_array($o->customer_id, $blockedIds)) { /* ... */ }
}
// O(n): constant-time lookup
$blocked = array_flip($blockedIds);
foreach ($orders as $o) {
if (isset($blocked[$o->customer_id])) { /* ... */ }
}
Real-world failure #5: scaling a leak
An API’s memory grew steadily until the container was killed every few hours. The team doubled the memory limit, which turned a 3-hour crash into a 6-hour crash. A profile showed a static array that cached every request’s data and was never cleared. The fix was 3 lines. The extra memory had cost money for two months.
Scaling hardware multiplies an inefficiency. It does not remove it.
The priority order (cheapest and highest impact first)
Common mistakes checklist
- Optimizing on your laptop with 50 rows instead of production-like data.
- Changing five things at once, so you cannot tell what helped.
- Trusting averages instead of p95/p99.
- Adding cache with no invalidation plan.
- Buying bigger servers before reading a single query plan.
- Never stopping. If the goal is met, stop. More speed past the goal costs time and adds complexity.
- Forgetting to guard the win: add a performance test or alert so it does not regress next sprint.
Key takeaways
- Set a numeric goal before you start.
- Measure first. Guessing is the most expensive optimization.
- Fix the biggest bottleneck only, then measure again.
- Database before cache, cache before code, code before servers.
- Cache carefully: locks for hot keys, invalidation on writes.
- Don’t make users wait for work they don’t need right now.
- Stop when the goal is met.
Performance work is not about being clever. It is about being disciplined: one measurement, one fix, one verification at a time. Do that, and you will save both your users’ time and your company’s money.
Thanks for reading. If this helped, share it with a teammate who is about to say “let’s just add more servers.”