Cache a high-traffic API by measuring first, then putting a cache-aside layer in Redis in front of the slowest, most-read endpoints, with explicit TTLs, invalidation on write and protection against cache stampedes. Use Elasticsearch for search and filtered listings the primary database handles poorly, and never cache user-specific or strictly consistent data without a deliberate design.
Measure P95 latency before you add any cache
Averages hide the requests users complain about. Track P95 and P99 latency per endpoint, along with request volume, so you can see which routes are both slow and heavily used.
Then find where the time goes. A trace or a simple timing breakdown per request usually shows whether the cost is a slow query, repeated queries, an external API call or serialization. Caching a response whose real problem is a missing index only hides that problem until the next cache miss.
- Record P50, P95 and P99 per endpoint, not just one global figure
- Rank endpoints by request volume multiplied by latency to see where caching pays most
- Check the read-to-write ratio, because data read far more often than it changes is the best candidate
- Fix missing indexes and N+1 queries before caching around them
Use cache-aside as the default pattern
In cache-aside, the application checks Redis first. On a hit it returns the cached value. On a miss it reads from the database, writes the result to Redis with a TTL and returns it.
The pattern keeps the database as the source of truth and fails safely. If Redis is slow or unavailable, the application falls back to the database, with a short timeout and a circuit breaker so a struggling cache does not slow every request.
Design keys deliberately. Include the resource type, the identifier, the API version and every parameter that changes the response, such as locale or page. A predictable key scheme is what makes targeted invalidation possible later.
Set TTLs by how stale data can be, then invalidate on change
A TTL is a business decision written as a number. Ask how old the data can be before a user or a downstream system is harmed, and set the TTL from that answer. A live match score and an archived article tolerate very different staleness.
Relying on expiry alone means serving stale data until the TTL runs out. For data your own application changes, delete or overwrite the key after the write commits, ideally from an event emitted after the transaction, so the cache never holds data the database rolled back.
Add a small random jitter to TTLs so keys written at the same time do not all expire at the same moment.
- Short TTL of a few seconds for fast-changing data where slight staleness is acceptable
- Longer TTL plus invalidation on write for data you control that changes rarely
- Versioned keys, where bumping a version number invalidates a whole group of keys at once
Protect the database from cache stampedes
A stampede happens when a popular key expires and many concurrent requests miss at once, all sending the same expensive query to the database. On a high-traffic API this can overload the database at exactly the moment traffic peaks.
Combine two or more of these defenses on your hottest keys.
- Request coalescing: one request rebuilds the key under a short Redis lock set with NX and an expiry, while the others wait briefly or serve the previous value
- Stale-while-revalidate: store a soft expiry inside the value, keep serving the stale copy past it, and refresh in the background
- Early probabilistic refresh: occasionally rebuild a hot key before it expires, with the probability rising as expiry approaches
- Pre-warming: populate known hot keys before a scheduled traffic spike, such as a major live event
Use Elasticsearch for search and listings, not as a general cache
Redis is best for key-value lookups: a single object, a computed fragment, a rate-limit counter. Elasticsearch fits a different job: full-text search, faceted filters and sorted listings that are expensive to compute in a relational database.
Treat the Elasticsearch index as a read model fed from the primary database, through change events or a scheduled sync, and accept that it is eventually consistent. Keep the database authoritative for writes and for anything that must be exact.
On the NorthStar Network sports media platform, which serves 50M+ monthly users, our engineers re-architected critical APIs and redesigned caching across Redis and Elasticsearch, which reduced P95 response times.
Know what not to cache
The costliest caching bug is serving one user's data to another. Review every cached endpoint for identity in the key, and test it with two different accounts before release.
- Responses that depend on the caller's identity or permissions, unless the key includes the user or role
- Data that must be strictly consistent, such as balances, stock at checkout or anything used in authorization decisions
- Write endpoints, one-time tokens and anything with side effects
- Low-traffic endpoints, where a cache adds complexity and a new failure mode for little gain
- Very large payloads that evict many smaller, hotter keys
Make the cache observable, or you cannot trust it
Track hit ratio per key prefix, Redis memory use and evictions, Redis command latency, and database load, next to API P95 on the same dashboard. When P95 moves, you should be able to tell within minutes whether the cache, the database or an upstream service caused it.
Alert on a sudden drop in hit ratio, which often means a deploy changed a key format, and on rising evictions, which mean the cache is too small for its working set.
We take on API performance work as a defined block: measuring, redesigning the caching layer and handing over dashboards and runbooks to the team that operates it.
Key takeaways
- Measure P95 per endpoint and fix query problems before adding a cache.
- Cache-aside in Redis is the safest default because the database stays the source of truth.
- Set TTLs from how stale the data may be, and invalidate on write for data you control.
- Protect hot keys against stampedes with locking, stale-while-revalidate or pre-warming.
- Use Elasticsearch as an eventually consistent read model for search and listings, not as a general cache.
FAQ
Should I use Redis or Elasticsearch to cache API responses?
Use Redis for key-value caching of objects, computed fragments and counters, because lookups by key are fast and simple to invalidate. Use Elasticsearch when the expensive part is search, filtering or sorting across many records, and treat its index as a read model rather than a cache.
What is a good TTL for API responses?
There is no universal value. Set each TTL from how stale that data can be without harming a user, use a few seconds for fast-changing data, and pair longer TTLs with invalidation on write. Add random jitter so related keys do not expire together.
How do I prevent a cache stampede in Redis?
Let only one request rebuild an expired key by taking a short lock with SET and the NX option, and have other requests wait briefly or serve the previous value. Stale-while-revalidate and early refresh of hot keys reduce the number of hard expiries in the first place.