How the KV cache shapes AI agent design

7 min read

An agent sends its entire context to the model on every turn, and almost all of it is identical to what it sent the turn before. The "KV cache" is what lets the model skip that repeated part, and whether an agent benefits from it depends on how the agent builds its requests.

What the KV cache holds

At every layer of a language model, each token is turned into three vectors: a "query", a "key" and a "value". To produce the next token, the model compares its query against the keys of every earlier token and blends in the matching values. This is the attention step.

The keys and values for the whole input have to be computed before the first output token can appear. This phase is called "prefill", and on a long input it makes up most of that wait. The model then keeps those keys and values in memory, so that each new token computes only its own. That store is the KV cache.

Every inference engine does this within a single request. "Prompt caching", now offered by the major APIs, extends it across requests: the provider keeps the KV cache from your last request, and if the next one begins with the same tokens, prefill resumes where the stored copy ends.

The stored copy is large. For an open model such as Llama 3 70B at 16-bit precision, the keys and values take about 320 KB per token, so a 100,000-token context occupies roughly 32 GB of accelerator memory. This is why providers keep it for minutes rather than days, and charge to create it.

A token's keys and values are computed from every token before it, so a stored copy is only valid for an exact prefix.

Change one token, and every token after it has to be computed again.

Append only16 of 19 reused
Timestamp in system prompt1 of 19 reused
Tool added3 of 19 reused
Old turn edited8 of 19 reused

Why agents depend on it

A chat request is usually a short question with a long answer. An agent works the other way round. Each turn appends a tool call and its result to a context that already holds the system prompt, the tool definitions, the project instructions and every earlier turn, and the model writes a few hundred tokens in reply.

Cost

Cached input is typically billed at a tenth of the normal rate or less. In a twenty-turn session with a 12,000-token prefix and 3,000 tokens added each turn, that brings the input bill down to about a fifth.

Latency

A hit skips prefill for the matched prefix. Anthropic measured a 79% shorter time to first token on a 100,000-token cached prompt, and an agent pays that wait on every step of its loop.

Throughput

On Anthropic's API, cached reads do not count against input rate limits on most models, so the same limits cover more work.

All three depend on one condition: each request has to begin with exactly the tokens the previous one sent. Whether it does comes down to how the agent builds its requests.

Ordering the request

Because the cache matches from the front, the first decision is order. A request is assembled in a fixed sequence, usually the system prompt, then the tool definitions, then the messages, and the most stable content belongs first and the most volatile last.

One request, front to backstable to volatile
system promptFixed. No dates, IDs or per-user values.
tool definitionsFixed for the life of the agent.
project instructionsFixed for the session.
historyGrows by appending. Never rewritten.
latest turnThe only part that is new.

Everything above the latest turn should be byte-identical to the previous request.

The boundary between the fixed part and the history is also where a cache marker belongs on APIs that need one. Anthropic's API caches only where a request carries a cache_control marker, and for an agent loop the simplest arrangement is one marker at that boundary, plus automatic caching, which moves a second marker to the end of the conversation on every request.

Keeping the prefix stable

Once the order is right, two things remain: everything above the latest turn has to stay byte-identical from one request to the next, and the cached copy has to still be there when the next request arrives. Each of the following rules prevents one specific way of losing the cache, and the last two are about timing.

  1. 1

    Keep the system prompt frozen

    A timestamp or request ID near the top changes the start of every request, and everything after it misses. Put values that change into the latest message.

  2. 2

    Serialize deterministically

    Tool definitions and tool results are usually JSON. If key order or whitespace varies between calls, two requests that mean the same thing are different bytes. Sort keys and fix the formatting.

  3. 3

    Append, never edit

    Treat the history as a log. Rewriting an old tool result or dropping the oldest turns resets the cache from that point onward.

  4. 4

    Keep the tool set fixed

    Tool definitions sit ahead of the whole history, so adding or removing one invalidates everything after them. Leave unavailable tools defined, and restrict them in the latest message or in the tool's handler.

  5. 5

    Stay on one model, and fork from the same prefix

    A cache belongs to the model that built it, so switching models mid-task starts from nothing. A side call, such as a summary, should copy the parent's model, tools and system prompt exactly and add its own instructions at the end.

  6. 6

    Warm the cache before fanning out

    On Anthropic's API a cache entry becomes readable only once the first response starts streaming, so five sub-agents launched together with the same prefix all pay full price. Send one, wait for its first token, then launch the rest.

  7. 7

    Fit pauses inside the cache lifetime

    A cached copy expires five minutes after its last use on Anthropic's API by default. A slow test run or a human approving a tool call can outlast that, and the next turn pays to write the whole context again. Where such pauses are normal, the one-hour lifetime is the cheaper option, even though writing it costs twice the normal input rate rather than 1.25 times.

Three of these rules are broken in the first request below and kept in the second.

Misses every turn

system:
  You are a coding agent.
  Time: 14:32:07
  Request ID: 7f3a9c

tools:
  chosen for this step

messages:
  last 20 turns only
  new tool result

Hits every turn

system:
  You are a coding agent.

tools:
  every tool, by name

messages:
  every turn, unchanged
  new tool result
  Time: 14:32. Run tests.

When to break the cache on purpose

The rule to append and never edit has one deliberate exception. A long session eventually fills with stale tool output, and the context has to shrink. Both of the usual remedies, summarizing the history and dropping old turns, rewrite the prefix, so the question is whether the saving repays the rewrite. On most of Anthropic's models, writing to the cache costs 1.25 times the normal input rate, reading it costs a tenth, and output costs five times.

Take a 150,000-token session whose first 12,000 tokens are the system prompt and the tools. Replacing the other 138,000 with a 4,000-token summary costs the summary call and one fresh write, and every turn after that reads 16,000 tokens instead of 150,000. The saving repays the cost within a few turns, provided the summary call reuses the cached prefix as described above.

A sliding window that drops the oldest turn on every request works the other way. Each drop changes the history from its first message, so every turn writes nearly the whole context at the higher rate. At these ratios that is about nine times what appending costs, on every turn for the rest of the session.

On Anthropic's newest models a sliding window can also fail outright. The "thinking" blocks those models return are bound to the exact history they were produced from, and for accounts created since 31 August 2026, a request that sends them back after older turns have been removed is rejected by default.

Compact rarely and all at once, and never trim a little on every turn.

Measuring the hit rate

The rules above can be checked rather than trusted, because every major API reports how much of each request came from the cache. On Anthropic's API the fields are cache_read_input_tokens and cache_creation_input_tokens; on OpenAI's they are cached_tokens and, from GPT-5.6, cache_write_tokens in the input token details. Dividing cached tokens by total input tokens gives a hit rate for each turn, and it belongs in the agent's logs next to the turn itself. On Anthropic's API, input_tokens counts only the part that was not cached, so the total is the sum of all three fields.

A hit rate that stays at zero from the start has two possible causes, and the write field tells them apart. If nothing is written either, nothing is being cached at all: the request carries no cache marker, or the prefix is shorter than the provider's minimum, which on Anthropic's models ranges from 512 to 4,096 tokens. Neither case produces an error. If a write appears on every turn, the cache is being filled and never read back, which means something near the start of the request changes every time, as a timestamp in the system prompt does.

In a healthy loop the hit rate jumps on the second turn and keeps rising, because each new turn is a smaller share of the whole. A sudden drop on one turn means something in the prefix changed on that turn, or the cached copy expired during a pause before it. Diffing that request against the previous one tells the two apart quickly, because an expired copy leaves the requests identical up to the new turn. The cause is nearly always one of the rules above.

Cache hit rate per turn
TurnPrefix left aloneSystem prompt edited at turn 9
10%0%
283%83%
386%86%
488%88%
589%89%
690%90%
791%91%
892%92%
992%0%
1093%93%
1193%93%
1294%94%
1394%94%
1494%94%
1595%95%
1695%95%
1795%95%
1895%95%
1996%96%
2096%96%
Cache hit rate per turn in an illustrative session with a 12,000-token prefix and 3,000 tokens added each turn. Editing the system prompt changes the very start of the request, so turn 9 reads nothing from the cache.

A falling hit rate is an early and precise signal that a design rule has been broken, and it appears well before the monthly bill does.

How providers differ

The prefix rule holds everywhere, because it comes from how the model computes attention. What differs between providers is the arrangement around it: whether caching has to be switched on, how long a copy lives and what a hit costs.

As of October 2026
Anthropic
Caching
Needs a cache_control marker; minimum prefix 512 to 4,096 tokens, by model
Lifetime
5 minutes after last use, or 1 hour with writes at 2× input
Price
Reads 0.1× input or less, writes 1.25×
Reported as
cache_read_input_tokens
OpenAI
Caching
Automatic; minimum prefix 1,024 tokens on GPT-5.6 and later
Lifetime
30 minutes after last use on GPT-5.6 and later; up to 24 hours on earlier GPT-5 models
Price
Reads 0.1× input or less, writes 1.25× on GPT-5.6 and later
Reported as
cached_tokens
Google Gemini
Caching
Automatic on 2.5 and later, minimum prefix 2,048 to 4,096 tokens; explicit caches optional
Lifetime
Not published for automatic caching; explicit caches last as long as you set
Price
Reads 0.1× input on current models; explicit caches also bill storage by the hour
Reported as
cached_content_token_count
Self-hosted (vLLM, SGLang)
Caching
Automatic and on by default; vLLM matches in 16-token blocks, SGLang token by token
Lifetime
Until evicted to make room, least recently used first
Price
No charge; the saving is GPU time
Reported as
vllm:prefix_cache_hits

Most open-weight chat templates put the system prompt before the tool definitions, while Anthropic's API puts the tools first, and either order works as long as neither block changes. Several of Mistral's templates are the exception to plan around: they attach the tool definitions to the latest user message, so the tool block moves on every turn and is computed again each time.

Prices and cache lifetimes will keep changing between providers and model generations. The constraint will not, because it comes from how attention works: whatever an agent puts first should never need to change.