You Might Not Need a Vector Database
A RAG pipeline is a lot of parts: a chunker, an embedding model, a vector database, a retriever, usually a reranker, and an eval harness to tell you when retrieval quietly got worse. Cache-augmented generation (CAG) del
A RAG pipeline is a lot of parts: a chunker, an embedding model, a vector database, a retriever, usually a reranker, and an eval harness to tell you when retrieval quietly got worse.
Cache-augmented generation (CAG) deletes all of it. You put the entire knowledge base in the prompt, cache it at the provider, and ask your question. No retrieval step, so no retrieval mistakes.
That sounds like a stunt until you do the arithmetic. Then it turns into a straightforward rule.
TL;DR
- CAG = the whole corpus in the prompt + prompt caching. With a hosted API, that's the whole implementation.
- The rule: cache the corpus when it's less than ~10× what you'd retrieve. Cached input costs about 0.1× normal input, so 10× the tokens costs roughly the same.
- The real constraint is the cache TTL, not the context window. Entries expire in minutes.
- Sparse traffic makes CAG the most expensive option, not the cheapest.
- Keep RAG for big corpora, fast-changing content, per-user permissions, and provenance.
What CAG actually is
The paper that named it — Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks — preloads the documents into the model's context, computes the KV cache once, and answers queries against that. No retriever at query time.
Running your own model, you hold that cache in GPU memory. On a hosted API you get the same effect from prompt caching: the provider stores the processed prefix and charges a fraction to reuse it.
So CAG isn't a framework you install. It's a prompt layout:
- Put the whole corpus at the front, in a fixed byte-for-byte order.
- Mark the end of it as a cache breakpoint.
- Put the question after the breakpoint, where it changes freely.
Not the same as caching responses. Response caching stores the answer so an identical question skips the model. Prompt caching stores the processed input so a different question over the same documents skips re-reading them. They compose well — response cache in front, prompt cache behind.
The arithmetic
Cached input tokens cost roughly 0.1× normal input tokens — that holds for Anthropic and, on current models, OpenAI. Writing the cache costs about 1.25×, once.
That one number gives you the decision. Let C be your whole corpus in tokens and R the tokens you'd send per query with retrieval:
Caching the corpus costs about the same as retrieving when C ≈ 10 × R.
Below that, CAG is cheaper per query and has no vector database to run.
Worked through at $5 per million input tokens (so cached reads around $0.50 per million):
| Approach | Tokens per query | Cost per query |
|---|---|---|
| No cache, whole 200K corpus every time | 200,000 | $1.00 |
| CAG, 200K corpus cached | 200,000 cached | $0.10 |
| RAG, 6K of retrieved chunks | 6,000 | $0.03 |
| CAG, 20K corpus cached | 20,000 cached | $0.01 |
At 200K tokens, retrieval still wins on tokens. At 20K — a product manual, an API reference, a policy handbook — caching the lot is cheaper than retrieving part of it, and you delete the entire pipeline.
And the table leaves out everything RAG costs besides tokens: a vector database to run, an embedding model called on every write and every query, a chunking strategy to tune, a reranker, and an eval harness to catch silent regressions.
The TTL is the real constraint
Everyone asks whether the corpus fits. Current flagship models take a million tokens; most corpora that matter fit. The question that actually decides it is how often you ask.
| Provider | Cache lifetime |
|---|---|
| Anthropic | 5 minutes by default, 1-hour TTL at 2× the write price. Every read refreshes the timer. |
| OpenAI (GPT-5.6+) | A prefix stays eligible for about 30 minutes after its last use. |
| Gemini | Implicit caching on by default for 2.5 and newer. |
So the break-even is about traffic shape:
| Gap between queries sharing the corpus | What happens |
|---|---|
| Under 5 minutes | Every query refreshes the entry. Pay the write once, read cheaply forever. CAG shines. |
| 5–60 minutes | Use a longer TTL or re-warm on a schedule. The doubled write price needs a few reads to pay off. |
| Hours apart | Every query is a cold write at 1.25–2× full price. CAG is the most expensive option available. |
A trap worth knowing: every provider has a minimum cacheable prefix, and below it caching is skipped with no error — you just quietly pay full price forever. Anthropic's is model-dependent (512–4096 tokens), OpenAI's is 1,024 on GPT-5.6+, Gemini needs 2,048–4,096. Check the usage fields on the response, not the docs.
The layout in Spring Boot
Corpus in the system prompt with a cache breakpoint; question in the user message, after it.
@Service
public class HandbookAnswerService {
private final AnthropicClient client = AnthropicOkHttpClient.fromEnv();
// Built once at startup, in a fixed order. Any byte that changes here
// invalidates the cache for every request that follows.
private final String corpus;
public String answer(String question) {
MessageCreateParams params = MessageCreateParams.builder()
.model("claude-opus-5")
.maxTokens(4096L)
.systemOfTextBlockParams(List.of(
TextBlockParam.builder()
.text(corpus)
.cacheControl(CacheControlEphemeral.builder()
.ttl(CacheControlEphemeral.Ttl.TTL_1H)
.build())
.build()))
.addUserMessage(question) // after the breakpoint - varies freely
.build();
Message response = client.messages().create(params);
// The only proof that caching is working.
log.info("cache write={} read={} uncached={}",
response.usage().cacheCreationInputTokens(),
response.usage().cacheReadInputTokens(),
response.usage().inputTokens());
return text(response);
}
}
Two details do most of the work:
The corpus never varies. Caching is a prefix match on exact bytes. A timestamp, a reordered map, a "user: Feezan" line — any of those changes the prefix and every request after it misses. Sort deterministically, interpolate nothing.
The question sits after the cached block. Putting the breakpoint at the very end is the common mistake: each unique question writes its own entry that nothing ever reads, so you pay the write premium on every call.
Log those usage fields from day one. If the write count is large on every call, something upstream is rewriting your prefix — and this failure is invisible, because requests still succeed and only the bill changes. Assert on it in a test.
When RAG still wins
- The corpus doesn't fit, or barely fits. Above a few hundred thousand tokens you're re-reading everything for every question, and long-context accuracy isn't free either — measure it.
- It changes constantly. One edited document changes the prefix and rewrites the whole cache at the write premium. A handbook revised weekly is ideal; a ticket system is not.
- Different users see different documents. Per-user filtering means a separate cached prefix per permission set.
- You need provenance. RAG hands you the chunks, so you can cite them. With everything in context, you're trusting the model to cite honestly.
- Volume is large and steady. At high volume, retrieval's smaller prompt wins again.
The hybrid is usually the answer at scale: cache the stable material (schemas, policies, API reference, instructions) as a fixed prefix, and retrieve only the volatile part after the breakpoint.
The takeaway
RAG is a solution to a size problem. If your knowledge base isn't actually that big, you can skip the pipeline, cache the whole thing, and spend the time you saved on answer quality instead of retriever tuning.
Count your corpus before you build. If it's under about ten times what your retriever would send per query, try the boring version first.
I write about backend engineering and payments at feezankhattak.com. The longer version of this post has the full gotchas list, and I build free in-browser developer tools — no sign-up, nothing uploaded.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.