Caching LLM Calls Without Lying to Your Users
The bill arrives and someone says “we should cache the LLM calls”. Everyone agrees, because caching is the least controversial idea in computing. Then three people implement three different things, and one of them starts serving customers the wrong refund policy.
The confusion is worth clearing up before the money is: “LLM caching” names three mechanisms sitting in different places with different guarantees, and only one of them can change what a user is told.
A cache that returns a different question's answer has not made you faster. It has made you wrong, quickly, and at a very good hit rate.
Prefix caching: take this one immediately
Providers cache the computed attention state for a prefix of your input. Send the same first few thousand tokens again — your system prompt, your tool definitions, the document everyone is asking about — and that part does not have to be processed from scratch. You are billed less for those tokens and the first one comes back sooner.
Nothing about the output changes. The model still runs, still generates, still produces something fresh. This is a pure saving with no correctness surface at all, which makes it the one to reach for first and the one most teams get least value from, because of how they order their prompts.
The rule is stable content first, variable content last. A prefix cache matches from the beginning and stops at the first difference, so a timestamp or a user’s name near the top of the prompt invalidates everything after it.
✓ [system rules][tool defs][retrieved docs] │ [chat history][user turn]
← cacheable, unchanged across calls ──┘
✗ [today's date][user name][system rules][tool defs][retrieved docs]…
↑ changes every call, so nothing after it can be reused
Moving two lines is often the whole optimisation.
Exact-match caching: safe, and the key is the hard part
Store the response against a hash of the request, return it verbatim on a repeat. Byte-identical to what you would have computed, so it is safe by construction — which shifts the entire difficulty into what goes in the key.
Everything that could change the answer belongs there. The model and its
version, because -latest aliases move and a cache that survives a silent
model upgrade is serving output from a model you are no longer running. The
system prompt. The tool definitions. Any retrieved context. The sampling
parameters. And the tenant.
If retrieved documents are in the prompt, their ids and versions have to be in the key too, otherwise re-indexing the corpus leaves you serving answers built from documents that no longer say that. In practice this is why exact-match caching earns its keep on system-shaped traffic — classification, extraction, enrichment, the same twelve support questions — and barely registers on open conversation, where nobody phrases anything the same way twice.
TTLs should come from the underlying data, not from a default. An answer derived from a pricing table is stale the moment pricing changes, and that is an invalidation event you can actually subscribe to rather than a duration you guessed.
Semantic caching: a correctness decision, not an infrastructure one
Embed the incoming question, find the nearest stored question, and if the cosine similarity clears a threshold, return that stored answer. The hit rate is dramatically better than exact match, because it catches every rephrasing.
It is also the only one of the three that can hand a user an answer to a question they did not ask.
Consider the pair in the diagram. What is the refund policy for EU orders and what is the refund policy for US orders embed at around 0.94 similarity — above the threshold most tutorials suggest. They are nearly the same sentence. They have completely different answers. No threshold separates them, because the property that distinguishes them is semantic in a way the embedding is explicitly designed to smooth over.
The errors are asymmetric and that asymmetry is the whole decision. A false miss costs a few hundredths of a cent and some milliseconds. A false hit gives someone confident, well-formatted, wrong information — with no error, no anomaly, and nothing in your logs that looks different from a success. You will find out from a customer.
So the honest framing is not “what threshold should I use”. It is: on this traffic, is a near-miss cheap? Sometimes it plainly is — documentation lookups, definitional questions, “what does this error mean”, the top of a support funnel where a slightly-adjacent answer still helps. Sometimes it plainly is not — anything about a specific account, price, entitlement, date or obligation. Split the traffic and cache only the first kind, rather than picking one number for all of it.
If you do run one, three things make it survivable: keep the threshold high enough to feel wasteful, scope the vector index per tenant so a near-miss can never cross a customer boundary, and log every hit with the served question alongside the asked one. That log is the only way you will ever see a false hit — and it is a ready-made source of cases for your eval set.
Measure the thing you are actually buying
Hit rate is the wrong headline number, because a semantic cache can raise it by being more wrong. Track cost per resolved request and latency at p95, and put a correctness check next to both: sample cache hits and grade whether the served answer actually answered the asked question.
That check is the difference between a cache you can defend and a number that went up.
What you have actually built
Two of these are arithmetic. Prefix caching and a well-keyed exact-match cache give back money and milliseconds and take nothing in return, and most teams have not finished collecting either before they reach for the third.
The third is not a cache in the sense the other two are. It is a decision that questions which are close enough deserve the same answer — which is a product judgement about your domain, made once, applied silently to every request. It can be the right call. It just should not be made because the hit rate looked better.
Quick answers
- What is semantic caching for LLMs?
- A semantic cache embeds the incoming prompt and returns a stored answer when an earlier prompt is close enough in vector space. Unlike an exact-match cache it can hit on rephrasings — and unlike an exact-match cache it can serve the answer to a question that was similar but not the same.
- Is prompt caching the same as caching the response?
- No. Prompt or prefix caching is provider-side and reuses the computation for a shared prefix of the input. The model still runs and still generates fresh output — you pay less for the prefix tokens and get first tokens back sooner. The output is unchanged, so there is no correctness risk.
- What should an LLM cache key include?
- Everything that could change the answer: the model, its version, the system prompt, the tool definitions, any retrieved context, and the tenant or user whose data is in scope. Leaving out the tenant is how one customer's answer is served to another.
- When should you not use a semantic cache?
- Anywhere a near-miss is expensive: pricing, policy, medical, legal, account-specific or anything the user will act on. Similar questions with materially different answers are exactly the case a similarity threshold cannot distinguish, and the failure is silent.
References
Related Discoveries
Lumi's weekly note
A short email when we publish something useful. No spam, unsubscribe anytime.