Skip to content
Agent Month

How to fix: prompt caching not working / cache_read_input_tokens is zero

Last verified: June 2026· Anthropic and other providers with prompt caching

Where this shows up

Anthropic and other providers with prompt caching

The fix

  1. 1Diff the rendered prompt bytes between two requests to find what differs in the prefix.
  2. 2Move volatile content (timestamps, per-request IDs, the varying question) to the end, after the last cache breakpoint.
  3. 3Serialize JSON deterministically (sorted keys) and keep the tool list stable and ordered.
  4. 4Ensure the cached prefix exceeds the model’s minimum cacheable length — short prefixes silently won’t cache.
  5. 5Verify with cache_read_input_tokens in the usage response once fixed.

Prevent it

Freeze the system prompt and tool list, inject dynamic context later in the messages, and audit for silent cache invalidators.

Common variations and related errors

You'll usually hit this same root cause under a few different names. Same fix.

  • "prompt caching not working"
  • "cache miss"
  • "no cached tokens"
  • "cache_control not respected"

Frequently asked questions

What causes “prompt caching not working / cache_read_input_tokens is zero”?

Something in the cached prefix changes every request — a timestamp, UUID, unsorted JSON, or a varying tool set — invalidating the cache.

How do I prevent “prompt caching not working / cache_read_input_tokens is zero” from recurring?

Freeze the system prompt and tool list, inject dynamic context later in the messages, and audit for silent cache invalidators.