Caching
Prompt caching
A repeated prefix — a long system prompt, a document, a tool catalogue — can be read from the provider's cache instead of being processed again, at a fraction of the input rate. Caching is reported on every response so you can verify it is working.
| Field on <code>usage.prompt_tokens_details</code> | What it is |
|---|---|
cached_tokens | Input tokens served from cache, billed at the cache-read rate |
cache_write_tokens | Input tokens written into the cache on this request |
cache_creation.ephemeral_5m_input_tokenscache_creation.ephemeral_1h_input_tokens | The TTL split of that write. The two are billed at different rates, so the split is reported rather than summed. |
How to get a hit. Keep the stable part of your prompt first and byte-identical between calls, and put anything that varies at the end. On Claude, mark the breakpoint explicitly with
cache_control on a system block — the Anthropic SDK passes it through per block. Whether a prefix is cached at all is the provider's decision, which is why the numbers above are reported rather than promised.