Dev.to AI 🤖 Ai 👁 0 📖 4 min read

Your prompt-cache fix is worth 0% if your users only send one message

A 30-turn measurement of how the saving amortises — and why the "one-line fix" quietly decays as conversations get longer. Last week I published a benchmark showing that moving a roughly 30-token volatile header out

A 30-turn measurement of how the saving amortises — and why the "one-line fix" quietly decays as conversations get longer.

Last week I published a benchmark showing that moving a roughly 30-token volatile header out of the top of a system prompt cut steady-state inference cost by 96%. The setup, the code and the raw API data are here: https://github.com/chenyu520-ai/rp-cache-lab

The most useful response came from Max Quimby, who runs scheduled multi-agent jobs and pushed back on the framing:

The amortization depends heavily on session length: if most of your sessions are only 2-3 turns, the cold turn-1 miss dominates and the win shrinks fast. Did you look at the hit-rate curve as a function of conversation length? That distribution is what decides whether this is a 95% win or a rounding error for a given app.

He was right, and I hadn't measured it. The first run was 8 turns, which is too short to answer the question. So I reran it at 30 turns.

This post is the answer, plus something I did not expect to find.

What was measured

Three ways of building the same request. The stable content — character card, world book, long-term memory, about 20,000 tokens — is byte-identical in all three. Only a small volatile header moves:

Session context: local time <ISO timestamp>, turn <n>, session <uuid>, mood index <0.00>.
Variant Where the volatile header goes
A — naive Front of the system message
B — minimal fix End of the system message
C — optimized End of the newest user turn; the system message never changes

30 turns per variant, deepseek-flash, thinking disabled, 90 requests total. Every number below comes from the API's own usage fields. The whole run cost $0.096.

The curve

The turn-1 miss is unavoidable — there is no cache to hit yet. That cost gets spread across the rest of the session, so short sessions carry it disproportionately.

Session length A avg/turn B avg/turn C avg/turn C vs A
1 turn $0.002573 $0.002574 $0.002575 ~0%
2 turns $0.002584 $0.001339 $0.001342 −48.0%
3 turns $0.002590 $0.000930 $0.000932 −64.0%
5 turns $0.002602 $0.000609 $0.000599 −77.0%
10 turns $0.002625 $0.000383 $0.000350 −86.7%
20 turns $0.002673 $0.000301 $0.000229 −91.4%
30 turns $0.002723 $0.000301 $0.000189 −93.1%

Read the first row carefully. On a one-turn session there is no saving at all — C is fractionally more expensive, because its prompt is very slightly longer (the header rides in the user turn). There is no history to amortise the cold miss over.

An app where users routinely send one or two messages and leave gets roughly half the headline number. An app with 10+ turn sessions gets the whole thing.

Which means the honest answer to "does this save 96%?" is: not until you know your session-length distribution.

The part I didn't expect

The first run made B and C look equivalent. Over 8 turns their aggregate hit rates were 85.8% and 86.8% — a 1 percentage point gap, well inside noise.

Over 30 turns they are not equivalent at all:

Turn 2 Turn 30 Change
B cache hit rate 98.9% 90.5% decaying
C cache hit rate 98.7% 98.7% flat
B cost per turn $0.000104 $0.000333 3.2×
C cost per turn $0.000110 $0.000110 flat

B starts out nearly as good as C and then gets worse every turn. C does not move.

The mechanism is visible in the raw cache-hit token counts, and it is the most useful single table in this post:

Turn B hit tokens C hit tokens
2 16,896 16,896
5 16,896 17,280
10 16,896 17,664
20 16,896 18,688
30 16,896 19,584

B's hit count never changes. 16,896 is 264 × 64, and 64 appears to be the cache granularity. Because B's system message changes every turn, only the stable block ahead of the header can ever be reused — conversation history is permanently locked out of the cache. As the history grows, the uncached portion grows with it.

C's hit count climbs, because its system message is frozen and its history is append-only, so each turn's cached prefix includes everything before it.

So the "one-line fix" is real but incomplete. It buys you the stable block and nothing else. If your sessions are short, that is most of the win. If they are long, the gap between B and C becomes the thing that matters.

What this means in practice

  • Sessions of 1–2 turns: don't restructure for this. There is nothing to amortise and you will see no meaningful change.
  • Sessions of 3–10 turns: real, in the 64–87% range. Worth doing. Don't quote 96%.
  • Sessions of 10+ turns: the full effect, ~90%+. This is where it pays for itself.
  • Regardless of session length: prefer the append-only architecture (C) over the header move (B). B's advantage decays, and the extra work to freeze the system message is a one-time cost.

Caveats

I'd rather the numbers hold up than look good:

  1. One provider. These are DeepSeek's numbers. The mechanism should generalise to any prefix-based cache, but I have not measured Anthropic's, which uses explicit breakpoints plus a TTL rather than pure prefix matching. That is a real gap — if you have data there, I would like to see it.
  2. Best-effort caching. No provider guarantees a 100% hit rate, and repeated runs will vary.
  3. History compression changes the picture. A real app summarises or truncates history instead of letting it grow. That would cap the divergence between B and C, so the 30-turn numbers are the pessimistic case for C's advantage, not the optimistic one.
  4. The stable block is synthetic, though its structure and scale are realistic. This measures cache behaviour, not model quality.
  5. Off-peak pricing throughout. Peak is twice off-peak.

Reproducing it

The benchmark is small and self-contained:

node bench.js --self-test --dry-run   # proves the three variants differ, costs nothing
node bench.js --turns 30 --budget 1.00

--self-test diffs the serialised request bodies of turn 1 and turn 2 and reports how many bytes are identical. For variant A it is 75 bytes out of 84,665. That is the whole story in one number.

Full per-turn data for both runs, and the session-length analysis, are in the repository:
https://github.com/chenyu520-ai/rp-cache-lab

Thanks to Max Quimby, Tom Veber, IndianInfraNotes and Rulestack for the comments that turned this into a second measurement rather than a single headline. If your sessions are shorter or longer than the ones here, I would like to know what you see.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.