Skip to main content
Ahmed Salama

Architecture Lab

Caching

Do the expensive thing once, then stop doing it.

The shape of the decision

System diagram: Caching: Do the expensive thing once, then stop doing it.Checks the cache before it touches the database.SERVICEApplicationConsulted only on a miss, then the result is written back.DATADatabaseReturns in microseconds when the key is present.CACHECache1 · read2 · on miss3 · populate
Caching: Do the expensive thing once, then stop doing it.
Read this diagram as text
ApplicationService
Checks the cache before it touches the database.
DatabaseData
Consulted only on a miss, then the result is written back.
CacheCache
Returns in microseconds when the key is present.

Connections

  • Application → Cache (1 · read)
  • Cache → Database (2 · on miss)
  • Database → Cache (3 · populate)
The problem it solves
The same expensive computation or query runs over and over for results that barely change. Read-heavy systems spend most of their capacity re-deriving answers they already had.
How it works
Keep the result somewhere fast (in memory, in Redis, at the CDN), keyed by whatever the result depends on. Reads check the cache first. Writes either invalidate the key or let it expire.
What it costs
Every cache introduces a window where the world has changed and your answer has not. You are trading correctness-in-time for speed, and you have to decide how much of it you can afford, key by key rather than globally.
When not to reach for it
The data is written as often as it is read, the query is already fast, or the correct answer changes per user in ways the key cannot capture. Caching a slow query is also how a missing index survives to production.

What to key on

The cache key is not an implementation detail decided after the design is finished. It is the design. A key that is too broad returns a stored answer to a request that never actually asked for it. A key that is too narrow never gets reused, and the cache buys nothing beyond its own storage cost. I treat the key as the actual contract a cache makes with its callers, and everything else, the storage engine, the eviction policy, is secondary to getting that contract right.

Most keys take one of two shapes. A request-derived key is built from the URL, the method and perhaps a query string, and it works when the answer is the same for anyone who could construct that exact request. An identity-derived key folds in who is asking, a user id, a tenant id, a role, and it is required the moment the answer depends on that identity rather than on the request alone.

The CDN entry elsewhere in this lab warns that a personalised response must never share a cache with anyone else's. Worth extending: the failure is rarely a deliberate choice to key by request when a response is personalised. It is forgetting that a route became identity-dependent because nothing about its shape announces it. An account page that reads a session cookie and renders different content per visitor still looks, to a naive cache, like one URL worth one cached copy. Folding identity into the key explicitly, or refusing to cache the route at all, is the only fix that survives someone adding a second data source to that page later.

Cardinality is the cost nobody budgets for up front. A key space with one entry per user, or one per combination of filters on a listing page, produces mostly single-use entries: computed once, read once, evicted before anyone reads them again. That is a cache in name only, paying memory and eviction overhead for a hit rate near zero. I default to keying on the smallest thing that actually repeats, a filter combination rather than a full query string, a tenant rather than a user, and accept a slightly less precise cache in exchange for one that stays warm.

Invalidation without lying

Choosing a TTL is really answering a different question: how wrong can this answer be before somebody notices or is harmed. I treat it as an honesty budget rather than a tuning knob. An hour-long TTL on a price is a stated claim that being an hour out of date is acceptable, and that claim deserves to be written down somewhere a reviewer can challenge it, rather than left as whatever number was in the example when the code was written.

Explicit invalidation removes the wait: the write path deletes or updates the affected key at the moment the underlying data changes, so a read never sees a stale value for longer than the write itself takes. The cost is coverage. Every path that can mutate the data, the normal application code, the admin panel, a bulk import job, a migration run by hand, has to remember to invalidate the same key, and missing one leaves the cache serving a wrong answer without any error until a TTL, if one exists, eventually corrects it.

This is why cache forever, bust on deploy works for static assets and fails for data. An asset is content-addressed, its filename changes when its content does, so the old URL can be cached without limit because it will never represent anything else. A data record keeps the same key across every version of itself, so keeping that one key current is the entire problem, and a deploy is a poor proxy for when the underlying data actually changed. Most of the changes a cache needs to reflect happen between deploys, driven by users, not by releases.

The two mechanisms work best together rather than as alternatives. A short TTL underneath explicit invalidation is a backstop for the invalidation call somebody forgot to add: instead of a wrong answer lasting indefinitely, it lasts at most as long as the TTL. That ceiling is the actual guarantee the system offers, and it is worth setting deliberately rather than leaving it at whatever default the caching library ships with.

Stampedes and dogpiles

A hot key is a single point of failure wearing a cache's clothing. While it is warm, one lookup serves every request that touches it. The moment its TTL lapses under load, every one of those requests misses in the same instant and reaches the same slow origin at once. A query that was fine at one request a second now runs hundreds of times in parallel, and the backend the cache existed to protect takes exactly the spike the cache was meant to prevent.

Jitter breaks the synchronisation that causes this. Adding a small random offset to each TTL, rather than setting them all to the same round number, means keys populated together, at a deploy, at a cache warm, at the start of a traffic spike, stop expiring together. The regeneration work spreads across a window instead of landing in one moment.

Request coalescing handles the case jitter cannot: a single key popular enough that concurrent requests collide on it regardless of timing. On a miss, the first request proceeds to the origin and the rest wait on its result rather than each making an identical, wasted trip. An in-process lock only protects one instance. A lock held against the cache itself protects the whole fleet, at the cost of one more mechanism that needs its own timeout or it can deadlock instead of failing safe.

Serving the stale value while a single background request refreshes it removes the choice between a slow response and a wasted one: everyone keeps getting an answer at cache speed, and only one request pays to update it. What they receive is a little more out of date than the stated TTL promised. That is a real extension of the honesty budget, not a free win, and I would rather size it deliberately than adopt it as a default just because the caching library happens to offer it.

When the cache is hiding the real problem

A cache placed over a missing index survives a load test and dies in production. While the cache stays warm the underlying query never runs often enough to reveal its real cost, so nothing measures it. The first time a widely used key expires under real traffic, or a stampede forces many requests through at once, the raw query time reappears in full, at the worst possible moment, because that is precisely when the cache is under the most pressure to keep working.

The same masking happens with an N+1 query pattern, where one request turns into dozens of database round trips scattered across a loop rather than a single line anyone would flag in review. Caching the assembled response hides the fan-out from request latency, but every cache refresh still pays the full N+1 cost, and a deploy that clears the cache turns every one of those refreshes into a simultaneous N+1 storm across the whole fleet.

A chatty service boundary gets the same treatment when caching absorbs several calls that a badly shaped interface makes for every one it should need. The cache reduces how often the extra calls happen. It does not reduce how many the interface still requires, and a new call site added later inherits the same cost from day one, because the interface itself never improved, only the average frequency of paying for it.

This is the concept's own caution about when to avoid caching, taken one step further: caching over the wrong problem does not just fail to help, it hides the evidence that would have led somebody to fix the real thing. A missing index that is never fixed does not go away. It waits behind the cache for the day the cache is unavailable, cold, or simply overwhelmed, and it reappears at full size with no warning. I measure before I reach for a cache, name the specific operation that is slow, and only cache it once I can say exactly what expensive thing I am choosing not to repeat.

Let’s talk

Have a product, platform or delivery challenge? Let’s talk about turning it into a structured, scalable solution.

Open to technical leadership, product delivery and senior engineering roles, and available for architecture consulting, technical reviews and mentorship. Engagements run as project-based work, contracts, consulting, freelance engagements, remote collaboration and long-term partnerships.

Based in Cairo, Egypt, working remotely with clients across the MENA region and internationally.

Also on LinkedIn (opens in a new tab)