Why Isn't Your Prompt Cache Hitting?
A prompt cache that stops hitting throws no error. The answer is still right; only the bill goes up. The misses that are not bugs, the ordering mistakes that are, which changes break which part of the cache, and how to catch a broken cache before it reaches your bill, worked through on the Claude API.
This is the last of three posts on prompt caching. The first, How Does Prompt Caching Actually Cut Your Bill?, explains how a cache works and what it costs. The second, How Do You Keep a Growing Conversation Cached?, places cache markers in a chat. You don't need to read either one first. The next section gives you the facts this post uses.
When a prompt cache works, each request reuses work the server already did for an earlier one. You pay less, and the answer starts sooner.
When it stops working, nothing tells you. The request still succeeds, and the answer is still correct. The only signs are a few numbers in each response and a bill that is higher than it should be.
This post is about finding that problem and fixing it. We'll go in five steps:
- Not every miss is a bug. Some misses are expected. Rule them out first.
- The ordering rule. The most common real cause, and the one idea that applies to every provider.
- Three tiers. Which changes break which part of the cache, and the ways around them.
- The silent breakers. Code that changes the prompt when nobody meant it to.
- Debugging. How to find where the match broke, and catch the next break before it ships.
What you need from the earlier posts
Five facts carry this whole post:
- The cache matches exact bytes. A provider stores the model's computed work for a prefix: the stretch of the prompt from its first token up to some point. A later request reuses it only if its prefix is exactly the same, byte for byte. Change one byte, and the match fails from that point on.
- The order is fixed. Your request is laid out as one long line of tokens. On the Claude API, the order is tools, then system, then messages.
- A breakpoint marks where a cached part ends. On the Claude API, it is a
cache_controlfield on one block. The cached part always starts at the first token. - Three numbers report what happened. On the Claude API, every response has a
usageobject.cache_creation_input_tokenscounts tokens written to the cache, which is a miss.cache_read_input_tokenscounts tokens read from it, which is a hit.input_tokenscounts everything else, at full price. Note thatinput_tokensis what's left over, not the total. - A miss costs more than no cache. On the Claude API, a write costs 1.25× the normal input price, and a read costs about 0.1×. Reads should be much larger than writes. If they are about equal, caching is costing you money.
Not every miss is a bug
Before you hunt for a bug, rule out the misses that are expected. Each of these comes straight from how caching works:
- The first request is always a write. Nothing is stored yet. Only a later request can read.
- The prefix is too short. Each model has a minimum cacheable length. On the Claude API it is between 512 and 4,096 tokens, depending on the model. Below it,
cache_controlquietly does nothing. You get no error, no write, and no read. - The gap was longer than the timer. On the Claude API, an entry nobody reads within its time-to-live, or TTL (5 minutes or 1 hour), is deleted. The next request writes it again.
- The requests come from different workspaces. On the Claude API, cache entries are kept separate per workspace. The same prompt in two workspaces makes two entries. (On Amazon Bedrock and Google Cloud, the boundary is per organization instead.)
- The model changed. Stored work only fits the model that made it. A different model, or even a different version of the same model, can't use it.
Two more cases look like failures from the outside. One is a real cost you can avoid. The other isn't a problem at all.
The parallel-request race
A cache entry can only be read once the first response using that prefix has started streaming back.
Now say you send several identical requests at the same moment. This is common in fan-out or multi-agent designs. Every request is still in prefill when the others start. None of them can read what the others are still writing. So all of them pay the full write price.
The same thing happens at any scale. N parallel workers on the same context each write their own entry and read none of the others'.
The fix: if many workers share one fixed prefix, send one request first. Either wait a moment before starting the rest, or run a single warm-up call before fanning out. Then the others read from the cache instead of each paying to write it.
Pre-warming: an empty response that is working
Pre-warming means writing the cache ahead of time, so the first real request doesn't pay for a miss. The idea works on any provider that caches. On the Claude API, you do it with max_tokens: 0.
Send a request with max_tokens set to zero. The server still runs prefill and writes the cache entry at your breakpoint. Then it returns straight away, with:
content: []stop_reason: "max_tokens"- a filled-in
usageblock showing the write
The empty content looks like something went wrong. Nothing did. You pay the normal write price and nothing for output, because there is no output. This is a fine way to warm an entry before traffic you expect. It is not worth it in three cases:
- Your traffic is already steady enough to keep the entry warm. A warm-up call is then just one more write you didn't need.
- The prefix is too short to cache at all.
- You'd be warming lots of different prefixes that mostly won't get read. Each one is a 1.25× write, and that can cost more than the time it saves.
The ordering rule
Exact prefix matching is strict about where a change happens. Change one byte anywhere in the prefix, and every breakpoint after that point stops matching.
That leads to the most useful idea in this whole series. It applies to any provider that caches this way:
Put everything stable first. Put everything that changes at the very end.
On the Claude API, the order is fixed: tools, then system, then messages. So "stable first" means nothing that changes may sit before the last system block you've marked.
The classic version of this bug is one line: Current date: {datetime.now()} at the top of a 12,000-token system prompt.
A few bytes near the very front change on every request. So every breakpoint after it misses, on every call, forever. cache_read_input_tokens stays at zero, and nothing in the response tells you why.
The RAG version of the same mistake is putting retrieved chunks before the system prompt. It feels natural, because chunks are "context" and context seems to belong at the top. But the chunks change on every query. Putting them before the stable system prompt breaks the cache for everything after them. Chunks belong after the breakpoint, not before it.
For the timestamp bug, there is a newer fix that doesn't touch the system prompt at all. Instead of editing the top-level system field, add a message with role: "system" to the messages list:
messages.append({
"role": "system",
"content": "Current date: 2026-09-09",
})
The model treats this message with the same authority as your top-level system prompt. It comes from you, the operator, not from the user. But it doesn't change the bytes the cache is keyed on.
Where it works:
- Available today, with no beta header, on Claude Opus 5, Opus 4.8, Fable 5, Fable 5.1, Mythos 5, and Mythos 5.1.
- Not on Claude Sonnet 5. It returns a 400 error (
role 'system' is not supported on this model).
Where it can go in messages:
- It must come after a
usermessage, or after anassistantmessage that ended in a server-tool call. - It must be the last entry in
messages, or be followed directly by an assistant turn. - So it can never be
messages[0].
Three tiers, not one
It's easy to think of "the cache" as one thing that either survives a change or doesn't. On the Claude API, it is actually three layers, or tiers. They follow the same order as the prompt: tools, then system, then messages.
A change only breaks its own tier and the tiers after it. It never breaks the tiers in front of it.
This three-tier setup is how the Claude API works. It is not a rule every provider follows. A provider with a different order, or with one single cache instead of tiers, won't match the table below.
Read each row as a question: does this tier survive the change on the left?
| Change | Tools cache | System cache | Messages cache |
|---|---|---|---|
| Tool definitions (add / remove / reorder) | ❌ | ❌ | ❌ |
| Model switch | ❌ | ❌ | ❌ |
speed, web-search, citations toggle | ✅ | ❌ | ❌ |
| System prompt content | ✅ | ❌ | ❌ |
tool_choice, images | ✅ | ✅ | ❌ |
thinking or effort change | model-specific | model-specific | ❌ |
| Message content | ✅ | ✅ | ❌ |
Two rows are easy to misread:
- Changing
tool_choiceper request, or adding and removing images between turns, does not touch the tools and system cache. Only the messages tier is lost. - Normal message content never touches the tools and system cache either. That is why caching survives an ordinary back-and-forth chat.
Only two changes rebuild every tier, on every model: changing the tool definitions, and switching models mid-conversation. Everything else in the table costs you at most the messages tier.
Thinking mode and effort need their own note:
- Changing either one always breaks the messages cache.
- On some models, the thinking settings come before tools and system in the prompt. On those models, the change breaks those tiers too.
The safe habit is to fix both settings per route. Use one setting per endpoint or agent role, and don't change it from request to request. Setting a model's default value explicitly costs nothing. It behaves the same as leaving the field out, so you may as well pin it.
The escape hatches
These are Claude API features. Other providers have no named equivalent.
For most of the changes above that break a tier, there is another way to make the same change without losing the cache. Availability is different for each one. They are not a bundle that comes together.
| Change that normally invalidates | Cache-preserving form | Available on |
|---|---|---|
| Tool definitions (add / remove) | tool_addition / tool_removal blocks | Opus 5 onward, beta mid-conversation-tool-changes-2026-07-01 |
| System prompt content | A {"role": "system", ...} message appended to messages[] | Opus 5, Opus 4.8, Fable 5, Fable 5.1, Mythos 5, Mythos 5.1 — available today, no beta header |
| A per-turn reminder you want removed later | A turn-scoped system message with clear_at: "next_user_message", left visible in the transcript | Same six models, beta mid-conversation-system-clear-at-2026-08-21 |
effort change | {"role": "system", "content": [], "output_config": {"effort": ...}} | Fable 5.1, Mythos 5.1, Opus 5, beta mid-conversation-output-config-2026-07-01 |
Switching models has no row here on purpose. There is no escape hatch for it. A cache entry belongs to the model that wrote it. So switching models mid-task always throws away every cache you've built.
That is a hidden cost of routing cheap steps to a small model and hard steps to a big one. Every switch between models is also a switch between caches.
A better approach: keep one model on the main loop. When you need a cheaper model for a side task, call it as a separate subagent. Don't switch the main loop's model.
The silent invalidator list
Most cache breakage doesn't come from a deliberate change like the ones above. It comes from code that builds the prompt and is less stable than it looks. Search your own code for these:
| Pattern | Why it breaks caching |
|---|---|
datetime.now() / Date.now() / time.time() in the system prompt | The prefix changes on every single request |
uuid4() / crypto.randomUUID() / a request ID rendered early in the prompt | Every request becomes unique by construction |
json.dumps(d) without sort_keys=True, or iterating a set | Serialization order isn't guaranteed stable |
| An f-string interpolating a session or user ID into the system prompt | The prefix becomes per-user, so it can never be shared |
A conditional system section (if flag: system += ...) | Every combination of flags is its own distinct prefix |
tools=build_tools(user) where the tool set varies by user | Tools render first, so this invalidates everything behind it too |
One more cause won't show up in a search: forks.
Summarization, server-side compaction, and subagent calls all build a new request from a parent conversation. If that new request rebuilds system, tools, or model with even one small difference, it misses the parent's cache and starts over. A different key order or a slightly different string is enough.
The fix is simple. Copy those three fields from the parent exactly as they are. Add anything fork-specific only after them.
Debugging a cache that isn't hitting
On the Claude API, a beta feature called cache diagnostics helps here. Instead of cutting your prompt in half again and again to find the problem, you ask the server where the match broke. Call it through the beta namespace, with the beta turned on:
response = client.beta.messages.create(
model="claude-opus-5",
max_tokens=1024,
betas=["cache-diagnosis-2026-04-07"],
diagnostics={"previous_message_id": None}, # None on the first turn
system=[...],
messages=[...],
)
On every turn after the first, pass the previous response's ID instead of None. The result comes back on response.diagnostics. It shows exactly where this request's prefix stopped matching the one before it.
That helps when you already suspect a problem. The better habit is to catch it early. Log all three usage fields on every call: input_tokens, cache_creation_input_tokens, and cache_read_input_tokens. Then set an alert for when reads drop compared to writes.
This matters because the failure is silent. Nothing errors. Nothing warns. The request succeeds and the answer is correct. The only sign is a bill that is quietly higher than it should be.
It is also usually a regression, not a first-time mistake. Caching worked when someone built and tested it. Later, an unrelated change to how the prompt is built broke the byte-for-byte match, and nothing said so. Months can pass before it shows up on a cost dashboard.
Here is a check worth keeping in your test suite. Send the same request twice. Assert that the second response shows cache_read_input_tokens > 0. That one check catches this whole class of silent regression before it ships.
Recap
- A broken cache is silent: no error, no warning, a correct answer, and a higher bill. The usage counters are the only place it shows.
- Rule out the expected misses first: the first request, a prefix below the model's minimum, a gap longer than the TTL, a different workspace, a different model. Parallel requests that start together all pay to write.
- Pre-warming with
max_tokens: 0returns emptycontenton purpose. It writes the entry and costs nothing for output. - What usually breaks a cache is ordering: something that changes, placed before a breakpoint. That is true on every provider. Stable content first, changing content last.
- On the Claude API, a change breaks only its own tier and the ones after it. Only tool changes and model switches rebuild everything. Several changes have a cache-preserving form.
- Search your prompt-building code for timestamps, IDs, unsorted serialization, and per-user sections. Forks must copy
system,tools, andmodelfrom the parent exactly. - Log your provider's cache counters on every call (all three on the Claude API). Add a test that sends the same request twice and checks the second one reads from the cache.
A cache that stops hitting never tells you. Keep the stable part first and byte-for-byte the same, and test that the second request reads from the cache.