ToolingPublished ExplainedAll news from that day
Ongoing case: Anthropic's Claude models 9 storiesHow prompt caching bills a repeated prompt start
A cached prompt start is billed at a fraction of normal input, but only an identical start counts, and writing to the cache costs more than reading from it.
Check our sources · 26 facts from 2 sources
Illustration made with AI for ai notis
Key points
- A hit needs an identical start: Anthropic says changing tool definitions invalidates the entire cache, and prompts under 512 tokens cannot be cached on Sonnet 5.5 or Opus 5.5.
- Anthropic prices a 5-minute cache write at 1.25 times base input and a read at 0.1 times, or 0.05 times on Opus 5.5 and Sonnet 5.5.
- Anthropic's default cache lifetime is 5 minutes, refreshed free on each use; OpenAI says a GPT-5.6 or later entry stays eligible for 30 minutes after its latest write or reuse.
What happened
Prompt caching lets a model provider reuse work it has already done on the start of a prompt. OpenAI says the cache stores key-value tensors, the intermediate states the model builds while reading input, not the tokens themselves. The model still has to process any new input after the reused part.
A hit needs an identical start. Anthropic says cache hits require 100% identical prompt segments up to and including the block marked for caching, and that its cache follows the order tools, then system, then messages.
A change at one level invalidates that level and every level after it, and Anthropic says modifying tool definitions invalidates the entire cache. OpenAI likewise says reuse requires the entire rendered prefix to match, and lists the model, tools, text format and reasoning effort among the settings that can change it.
There is a minimum length. Anthropic lists 512 tokens for Claude Haiku 5.5, Sonnet 5.5 and Opus 5.5, and says shorter prompts cannot be cached even if marked. OpenAI gives 1,024 tokens for GPT-5.6 and later.
In practice you either mark where the reusable part ends or let the provider choose. With Anthropic you add a cache_control field to the request, and with automatic caching the breakpoint moves forward as a conversation grows; you can set up to four breakpoints. OpenAI says caching is on by default for supported models, and GPT-5.6 and later support implicit and explicit caching.
The cache expires. Anthropic's default lifetime is 5 minutes, refreshed free each time the entry is used, with a 1-hour option at extra cost. For GPT-5.6 and later OpenAI says an entry stays eligible for 30 minutes after its latest write or reuse.
The price has three parts. Anthropic lists a 5-minute cache write at 1.25 times the base input price, a 1-hour write at 2 times, and a read at 0.1 times, with per-model exceptions such as 0.05 times on Opus 5.5 and Sonnet 5.5.
For GPT-5.6 and later OpenAI lists writes at 1.25 times and reads at 0.1 times on most models and 0.05 times on GPT-6.1 Sol. OpenAI's own example is that one write and nine full reads cost 2.15 times the ordinary input cost of the prefix, against 10 times without caching.
Both providers report hits in the response. Anthropic's usage fields are cache_creation_input_tokens, cache_read_input_tokens and input_tokens, and OpenAI's are cached_tokens and cache_write_tokens under input_tokens_details. Anthropic says caching has no effect on output generation: the response is identical to one without caching.
What it means for you
Our viewCaching pays when many requests share a long start, such as a system prompt or a tool list. OpenAI's own example is one write and nine full reads costing 2.15 times the ordinary input cost, against 10 times without caching.
Put stable content first and anything that changes last, because a change at one level invalidates everything after it, and avoid editing tool definitions between requests. Anthropic says caching does not change the response.
Before counting on the saving, check your usage fields, cache_read_input_tokens for Anthropic and cached_tokens for OpenAI, to see your real hit rate.
It adds no new facts.
Your reaction
We count reactions per story and day, never who reacted. The counts help us choose what goes in the monthly issue. If you are signed in, your own page shows yours too.
Check our sources
Every sentence above is checked against these 2 sources.
1 Prompt caching
Open the source archived copy-
OpenAI's prompt caching guide says prompt caching reuses work when requests share the same prompt prefix. Quote: "Prompt caching reuses work when requests share the same prompt prefix."
citePrompt caching reuses work when requests share the same prompt prefix.
-
OpenAI says the prompt cache stores key-value (KV) tensors, not the tokens themselves, and describes KV states as intermediate states the model calculates while processing input tokens. Quote: "The prompt cache stores key-value (KV) tensors, not the tokens themselves."
citeWhen the model processes input tokens, it must calculate intermediate states, known as key-value (KV) states. These states let the model refer back to earlier tokens while processing new input and generating output tokens. Prompt caching preserves that state for a reusable prefix : the unchanged tokens at the beginning of a prompt. When a later request has the same prefix and finds a matching cache entry, the model can reuse the saved state instead of processing those tokens again. It still needs to process any new input to generate a new response. The prompt cache stores key-value (KV) tensors, not the tokens themselves.
-
OpenAI says that when a later request has the same prefix and finds a matching cache entry the model can reuse the saved state, but it still needs to process any new input. Quote: "It still needs to process any new input to generate a new response."
citeWhen a later request has the same prefix and finds a matching cache entry, the model can reuse the saved state instead of processing those tokens again. It still needs to process any new input to generate a new response.
-
OpenAI says cache reuse requires the entire rendered prefix to match. Quote: "Cache reuse requires the entire rendered prefix to match."
citeCache reuse requires the entire rendered prefix to match. If content or a relevant setting changes before a breakpoint, the prefix after that change cannot match the existing cache entry.
-
OpenAI's list of settings that affect the cached prefix includes model, tools, text.format and reasoning.effort. Quote: "Which settings affect the cached prefix?"
citeThe main settings to check are: Setting Impact model A different model can use different weights and caching behavior. tools Changes tool names, descriptions, schemas, ordering, or tool-specific instructions. parallel_tool_calls Can change instructions about calling multiple tools in one turn. text.format ( Structured Outputs ) Adds output-format instructions and the requested schema. reasoning.effort Can change model-side reasoning instructions.
-
OpenAI says the minimum cacheable prompt length is 1,024 tokens for GPT-5.6 and later. Quote: "The minimum cacheable prompt length is 1,024 tokens for GPT‑5.6 and later"
citeThe minimum cacheable prompt length is 1,024 tokens for GPT-5.6 and later and varies by request settings for earlier models.
-
OpenAI says prompt caching is enabled by default for supported OpenAI models. Quote: "Prompt caching is enabled by default for supported OpenAI models."
citePrompt caching is enabled by default for supported OpenAI models.
-
OpenAI says that for GPT-5.6 and later both implicit and explicit caching are supported. Quote: "Both implicit and explicit caching are supported, where explicit caching gives you more control over which context is written to cache."
citeBoth implicit and explicit caching are supported, where explicit caching gives you more control over which context is written to cache.
-
OpenAI says that for GPT-5.6 and later a cached prefix remains eligible for reuse for 30 minutes after its most recent write or reuse. Quote: "A cached prefix remains eligible for reuse for 30 minutes after its most recent write or reuse, though OpenAI may retain it longer."
citeA cached prefix remains eligible for reuse for 30 minutes after its most recent write or reuse, though OpenAI may retain it longer.
-
OpenAI says that for GPT-5.6 and later cache writes cost 1.25 times the standard uncached input-token rate and reads cost 0.1 times on most of these models and 0.05 times on GPT-6.1 Sol. Quote: "cache writes cost 1.25× the standard, uncached input-token rate. Subsequent reads cost 0.1× that rate on most of these models and 0.05× on"
citeFor GPT-5.6 and later, cache writes cost 1.25× the standard, uncached input-token rate. Subsequent reads cost 0.1× that rate on most of these models and 0.05× on GPT-6.1 Sol .
-
OpenAI says that across ten requests, one write and nine full reads cost 2.15 times at the 0.1 times read rate, compared with 10 times without caching. Quote: "Across ten requests, one write and nine full reads cost 2.15× at that rate, compared with 10× without caching."
citeAcross ten requests, one write and nine full reads cost 2.15× at that rate, compared with 10× without caching.
-
OpenAI says to track usage.input_tokens_details.cached_tokens and usage.input_tokens_details.cache_write_tokens to measure actual cache performance. Quote: "Track usage.input_tokens_details.cached_tokens, usage.input_tokens_details.cache_write_tokens"
citeTrack usage.input_tokens_details.cached_tokens , usage.input_tokens_details.cache_write_tokens , input-token counts, latency, and realized cost.
2 Prompt caching
Open the source archived copy-
Anthropic says cache hits require 100% identical prompt segments, including all text and images up to and including the block marked with cache control. Quote: "Cache hits require 100% identical prompt segments, including all text and images up to and including the block marked with cache control."
citeExact matching: Cache hits require 100% identical prompt segments, including all text and images up to and including the block marked with cache control.
-
Anthropic says prompt caching references the entire prompt, tools, system and messages in that order, up to and including the block designated with cache_control. Quote: "Prompt caching references the entire prompt:"
citePrompt caching references the entire prompt: tools , system , and messages (in that order), up to and including the block designated with cache_control .
-
Anthropic says changes at each level of the tools, system, messages hierarchy invalidate that level and all subsequent levels. Quote: "Changes at each level invalidate that level and all subsequent levels."
citeAs described in Structuring your prompt , the cache follows the hierarchy: tools → system → messages . Changes at each level invalidate that level and all subsequent levels.
-
Anthropic says modifying tool definitions (names, descriptions, parameters) invalidates the entire cache. Quote: "Modifying tool definitions (names, descriptions, parameters) invalidates the entire cache"
citeTool definitions ✘ ✘ ✘ Modifying tool definitions (names, descriptions, parameters) invalidates the entire cache
-
Anthropic lists a minimum cacheable prompt length of 512 tokens for Claude Fable 5.1, Mythos 5.1, Opus 5.5, Opus 5, Sonnet 5.5, Fable 5, Mythos 5 and Haiku 5.5. Quote: "512 tokens for Claude Fable 5.1"
cite512 tokens for Claude Fable 5.1, Claude Mythos 5.1 , Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Fable 5, Claude Mythos 5 , and Claude Haiku 5.5
-
Anthropic says shorter prompts cannot be cached, even if marked with cache_control. Quote: "Shorter prompts cannot be cached, even if marked with"
citeShorter prompts cannot be cached, even if marked with cache_control . Any requests to cache fewer than this number of tokens will be processed without caching, and no error is returned.
-
Anthropic says automatic caching, a single cache_control field at the top level of the request, applies the cache breakpoint to the last cacheable block and moves it forward as conversations grow. Quote: "The system automatically applies the cache breakpoint to the last cacheable block and moves it forward as conversations grow."
citeAutomatic caching : Add a single cache_control field at the top level of your request. The system automatically applies the cache breakpoint to the last cacheable block and moves it forward as conversations grow.
-
Anthropic says a prompt can define up to 4 cache breakpoints. Quote: "You can define up to 4 cache breakpoints if you want to:"
citeYou can define up to 4 cache breakpoints if you want to:
-
Anthropic says the cache has a 5-minute lifetime by default and is refreshed for no additional cost each time the cached content is used. Quote: "By default, the cache has a 5-minute lifetime. The cache is refreshed for no additional cost each time the cached content is used."
citeBy default, the cache has a 5-minute lifetime. The cache is refreshed for no additional cost each time the cached content is used.
-
Anthropic says it also offers a 1-hour cache duration at additional cost. Quote: "Anthropic also offers a 1-hour cache duration"
citeIf you find that 5 minutes is too short, Anthropic also offers a 1-hour cache duration at additional cost .
-
Anthropic says 5-minute cache write tokens are 1.25 times the base input price, 1-hour cache write tokens are 2 times, and cache read tokens are 0.1 times, with per-model exceptions in the table footnote. Quote: "5-minute cache write tokens are 1.25 times the base input tokens price"
cite5-minute cache write tokens are 1.25 times the base input tokens price 1-hour cache write tokens are 2 times the base input tokens price Cache read tokens are 0.1 times the base input tokens price (see the table footnote for per-model exceptions)
-
Anthropic says cache hits and refreshes on Claude Opus 5.5 and Claude Sonnet 5.5 are priced at 0.05x the base input price. Quote: "Cache hits and refreshes on Claude Opus 5.5 and Claude Sonnet 5.5 are priced at 0.05x the base input price."
citeCache hits and refreshes on Claude Opus 5.5 and Claude Sonnet 5.5 are priced at 0.05x the base input price.
-
Anthropic names the response usage fields cache_creation_input_tokens, cache_read_input_tokens and input_tokens. Quote: "cache_creation_input_tokens: Number of tokens written to the cache when creating a new entry."
citecache_creation_input_tokens : Number of tokens written to the cache when creating a new entry. cache_read_input_tokens : Number of tokens retrieved from the cache for this request. input_tokens : Number of input tokens which were not read from or used to create a cache
-
Anthropic says prompt caching has no effect on output token generation and that the response is identical to what you would get without caching. Quote: "Prompt caching has no effect on output token generation. The response you receive is identical to what you would get if prompt caching were not used."
citeOutput token generation: Prompt caching has no effect on output token generation. The response you receive is identical to what you would get if prompt caching were not used.
Topics
The morning email
On the mornings we publish: the three top stories and up to four short ones. Free.
We email you a link to confirm. An issue may include one sponsor, always labelled Sponsored · Advertisement. Our emails count opens and clicks, not who made them. Unsubscribe in one click. What we keep