Prompt Caching
Prompt caching lets a provider reuse the processing it already did for a repeated part of a prompt, instead of reprocessing that same text from scratch on every single call. If a system prompt, a large retrieved document, or a set of tool definitions shows up identically at the start of request after request, caching means the model doesn't have to read through it again each time — only the new part at the end actually needs fresh work.
Matching the shared prefix
A provider checks whether the beginning of a new request matches the beginning of a recent one. If it does, it reuses what it already computed for that shared portion and only processes what's new — usually just the user's latest message. This is why order matters: the part that repeats across calls (system instructions, a big document, tool definitions) needs to come first, and the part that actually changes every time (the user's specific question) needs to come last. If the stable content isn't at the very start, or something early in the prompt changes even slightly, the shared portion no longer matches and the cache doesn't help for that call.
What it actually saves
Both cost and time. Cached tokens are typically billed at a steep discount compared to processing them fresh, and skipping that reprocessing also cuts the time before the model starts responding — a real, current example from Anthropic's own pricing puts cached-token cost at roughly a tenth of the normal rate, with a similarly large cut to time-to-first-response for long prompts. Exact numbers vary by provider and change over time; the general shape — a steep discount on the repeated part, faster responses on top of it — is what to expect, not any single specific figure.
Two ways it gets turned on
Some providers cache automatically once a prompt passes a certain length, with no setup needed. Others require marking explicitly which part of the prompt should be cached and for how long it should stay cached before it expires — giving more control, at the cost of having to actually set it up rather than getting it for free.
Where it actually helps
Anywhere the same large block of text gets sent repeatedly: a long system prompt reused on every request to the same application, a big reference document or set of tool definitions that doesn't change between calls, or a long-running conversation that resends its full history on every turn. The more of a prompt is identical across calls, the more caching actually saves. A prompt that's different every time, with nothing stable at the start, gets no benefit from it at all.
A real limitation
A cache doesn't last forever, and it doesn't help infrequent requests — if too much time passes between calls with the same prefix, or the request volume is too low to keep it active, the benefit disappears and every call goes back to full price. It's a real optimization for a system with steady, repeated traffic, not something that helps a one-off request no matter how it's structured.
In this guide
FAQ
If a system prompt is cached, does changing the user's question break the cache for that request?
No — as long as the user's question comes after the stable, cached part rather than being mixed into it. The cache only cares about whether the beginning of the request matches; a different ending is exactly what's expected to change every time.
Does prompt caching require sending less information to the model?
No. The full prompt still gets sent and is still available to the model exactly as before — caching changes how much of it needs to be freshly processed, not how much of it exists. It's a cost and speed optimization, not a way to shrink what the model actually sees.