Checked fact 8198 Oct 2026Agents and apps
OpenAI says the prompt cache stores key-value (KV) tensors, not the tokens themselves, and describes KV states as intermediate states the model calculates while processing input tokens. Quote: "The prompt cache stores key-value (KV) tensors, not the tokens themselves."
The exact words it rests on
When the model processes input tokens, it must calculate intermediate states, known as key-value (KV) states. These states let the model refer back to earlier tokens while processing new input and generating output tokens. Prompt caching preserves that state for a reusable prefix : the unchanged tokens at the beginning of a prompt. When a later request has the same prefix and finds a matching cache entry, the model can reuse the saved state instead of processing those tokens again. It still needs to process any new input to generate a new response. The prompt cache stores key-value (KV) tensors, not the tokens themselves.
What the source said when we opened it, on 8 Oct 2026.
The source
Prompt caching
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.