Dev.to AI 🤖 Ai 👁 0 📖 9 min read

An epilogue to the game factory — I changed every model and the cost moved

Post 9 of 9 in the game-factory series. I said I was done. Post 8 ended with the factory parked — functional, proven, not something I was going to keep polishing. Then I came back, changed every model in the pipeline, a

Post 9 of 9 in the game-factory series.

I said I was done. Post 8 ended with the factory parked — functional, proven, not something I was going to keep polishing. Then I came back, changed every model in the pipeline, and learned something I'd had backwards.

The thing that pulled me back was a number. One agent, on a single run, sent about 780,000 tokens into the model and got roughly 5,000 back. Input to output, a ratio of 150 to 1. That agent was the Builder — the same one that fought me for months. By this point it barely used the model at all. So why was it burning three-quarters of a million tokens a run, and why was it the most expensive model I owned?

Where the factory landed

The short version, if you didn't read the rest of the series: six agents turn a one-sentence theme into a deployed, playable slot machine. Designer writes a spec, the image agents make icons and a background, the Builder rewrites the code, the Tester plays it in a browser, the Deployer ships it.

The Builder was the one I eventually turned back into a program. After the redesign it was roughly 99% deterministic — mirror the source, then swap colors, fonts, and strings in plain Python. The only thing left for the model was optional cosmetic CSS: nudge a glow, adjust a keyframe. A few small edits per build.

That's the setup for the question I hadn't asked yet. If the model in an agent is doing almost nothing, why is it the same top-tier model I use for the agent that does the most?

Matching the model to the work

So I made a map. For each agent, how much of the work is the model actually thinking, and how much is plumbing?

  • Designer — generates a spec from a conversation. This is the creative one. Taste shows up here.
  • Image-Gen / Background-Gen — the creative work happens inside the image model. The agent wrapped around it just loops over prompts and saves files. Plumbing around a creative core.
  • Builder — 99% deterministic Python, 1% cosmetic CSS. Plumbing.
  • Tester — half deterministic (drive a browser, take screenshots), half judgment (a vision model looking at the result).
  • Deployer — almost entirely deterministic. Plumbing.

The map made the decision obvious. Keep the strong model where the agent thinks. Try a cheaper, faster one where it's plumbing.

These are the assignments I landed on, from what Bedrock offered me in late 2026:

Agent Model Why
Designer Claude Sonnet 5 The creative step. Taste shows up here, so it keeps the strong model.
Image-Gen / Background-Gen Claude Haiku 4.5, driving Stability The judgment is in the image model. The agent around it just drives a short tool loop, so the cheapest Claude does.
Builder Qwen3-Coder-480B A long code-editing tool loop — a coding-specialized model, and a much cheaper one.
Tester (loop) Qwen3-32B Executes a fixed checklist. No creativity required.
Tester (vision) Pixtral Large Scores screenshots against a rubric. Needs to actually see.
Deployer Qwen3-32B Runs SAM and CLI steps in a fixed order.

Worth being precise about one thing, because "I moved the cheap agents to a cheap model" hides two different decisions. The Builder went to a coding model, because it edits code in a tool loop. The Tester and Deployer went to a plain, smaller general model, because they don't edit anything — they run checklists. And the Tester's vision model is the exception to the whole cost story: Pixtral Large is the second most expensive model in the pipeline, because "does this look like ancient Egypt" is a judgment call and that's the one place I'm buying judgment.

One practical note that isn't in any pricing table: when I ran this, the only region my setup could invoke Qwen3-Coder-480B from was us-west-2, so the Builder talks to a different region than every other agent in the pipeline. Regional availability moves around — AWS has since added others — but that's the point. Where a model is callable from is part of choosing it, not a footnote, and it's the part most likely to have changed by the time you read this.

The Builder got fast. Suspiciously fast — the kind of fast that makes you check whether it actually ran. It had. And then I looked at the token bill.

780,000 in, 5,000 out

The output was tiny, and that's the point. 5,000 tokens covers everything the model emitted across the whole run — its plan, its tool calls, and the handful of small CSS patches inside them. That confirmed "99% deterministic": the model was hardly generating anything. The 780,000 was all input. The cost of looking, not the cost of writing.

So where does 780,000 tokens of input come from when the instructions are barely 2,000?

I logged every call's input size across one run. The shape told the whole story:

call   input tokens   what changed
  1        3,017       system prompt + tool specs (baseline)
  3        3,360       still just reading the spec
  4       31,398       +28k: it read one CSS file into the conversation
 ...      ~33,000      each — that same file, re-sent every turn
 17       64,604       +28k: it read a second big file
 ...      ~65,000      each — now both files, re-sent every turn

The jump at call 4 is the model reading a single stylesheet so it could edit it. The file is 66 KB, and it arrived as about 28,000 tokens. CSS tokenizes worse than prose — all that punctuation and all those hex codes land nearer 2.4 characters per token than the 4 you'd assume from English. To patch a rule, the edit tool needs the file's exact current text, and the lazy way to get that is to hand the model the whole file.

Here's the part I'd had backwards. The file was read once. But in an agent loop, every call re-sends the entire conversation so far. That stylesheet landed in the context on call 4 and then rode along on all seventeen calls after it. Read once, paid for eighteen times. That one file is around 500,000 of the 780,000 tokens by itself.

I hadn't burned 780,000 tokens thinking. I'd burned them re-reading the same two files.

The catch I didn't see coming

At this point Qwen looked like a clear win: same near-empty output, a fraction of the price per token. Then I noticed a column in my usage log that was zero for every single call — cached tokens.

Qwen didn't support prompt caching on the Bedrock path I was using. Claude Sonnet 5 did, and my own integration only ever enabled caching for Anthropic model ids — one is_anthropic() check gating the whole code path. Bedrock supports caching on more than just Anthropic models, so this is a fact about my two candidates and my implementation, not a rule about the platform.

And prompt caching is exactly the thing that fixes a loop like this: the big, stable prefix you re-send every turn gets billed at a small fraction after the first time. That repeated stylesheet — the 500,000 tokens — is the ideal shape for a cache, and on Qwen none of it was cached, so every re-send was full freight.

So the comparison I care about isn't "cheap model versus expensive model." It's:

On published rates, and assuming the cache-hit pattern I'd expect from a prefix this stable, Sonnet 5 with caching lands in roughly the same range on this workload as Qwen3-Coder without it.

I want to be clear that this is an estimate and not a measurement. I never ran this build on Sonnet to compare — every number in my logs is a Qwen run with a cache column of zero. Cache hits also aren't free money: they depend on where you place checkpoints, whether the prefix really is byte-stable, and whether you hit the TTL before the next call. The point isn't a precise figure. It's that a price gap that looks like an order of magnitude on the pricing page can mostly close once caching is in play, so the sticker rate isn't the comparison you want.

Why Qwen stayed anyway

Cost was a wash. Three other things weren't.

The first is speed. Qwen3-Coder decodes quickly, and on an agent producing 5,000 tokens of real output, fast is what you feel. Builds that used to make me wait now finish before I've switched windows.

The second is a constraint I walked into sideways. Qwen3-Coder-480B gave me a 128K context window with a 16K cap on output, and on an agent whose context grows by 28,000 tokens every time it opens a file, that ceiling is closer than it looks. I had to bound the conversation history explicitly and reserve room for the reply, or the loop would overflow on input alone and the call would just fail. Context limits are shared between input and output generally — this isn't a Qwen peculiarity — but the headroom I had here was tight enough that it stopped being an accounting detail. The context bloat wasn't only expensive. On this model it was also the thing most likely to break the build.

The third is that the model isn't thinking here. Look at that output again: 5,000 tokens across twenty-one calls. The Builder isn't reasoning about architecture; it's making a few cosmetic edits and calling tools. When an agent is mostly plumbing, you're not giving up much quality by running a plainer model, because there's little quality being exercised in the first place.

The whole decision fits in one table:

Claude Sonnet 5 Qwen3-Coder-480B
Quality on hard reasoning high lower
Cost (with caching in play) competitive competitive
Speed slower fast
Context window larger; headroom to spare here 128K, 16K max output — tight for this loop
Doing hard reasoning here? no no

When the bottom row is "no," the top row stops mattering, and speed wins. So the plumbing agents keep Qwen. The Designer, where the bottom row is "yes," keeps Sonnet.

What I'd actually fix

The neat ending would be "pick the right model per agent." That's true, but it buries the real lesson, which is that the model was never the expensive part. The expensive part was shipping a 66 KB file into the context and re-sending it twenty times.

The model choice saved some money and a lot of wall-clock. Not shipping the whole file would save far more — the Builder touches maybe six rule blocks, a few hundred tokens, not twenty-eight thousand.

The two stylesheets account for about 646,000 of the 779,474 tokens, just by riding along in the context: one across eighteen calls, the other across five. Strip them out and what's left — system prompt, tool specs, spec lookups, tool results, the model's own replies, all of it re-sent every turn — is roughly 133,000. Hand the model extracted rule blocks instead of whole files and you land somewhere near 140,000. That's a five-fold cut, and it has nothing to do with which model is behind the endpoint.

(Not a twentyfold cut, which is what I'd have guessed before doing the arithmetic. The baseline you re-send every turn is itself most of what remains.)

I haven't shipped that one yet. I'm writing it down because it's the more useful lesson: I went looking for savings in the model and found them in my own context management.

What transfers

  • Measure input against output before you optimize a model. A 150-to-1 ratio isn't a smart agent; it's an agent re-reading. The token count measures how much you sent, not how much the model thought.
  • Match the model to what the agent actually does. Strong model where it reasons, plain model where it plumbs. Map your agents honestly — in this pipeline, most of them turned out to be plumbing.
  • Caching can beat a cheaper sticker price. A premium model that caches a big, stable prefix can land in the same range as a budget model that re-sends everything. Check whether your cheap candidate supports caching at all before you price the swap — not every model on a platform does. Compare the workload, not the per-token rate.
  • The cheapest token is the one you don't send twice. Before swapping models, look at what's sitting in your context getting re-billed every turn.

The most reliable agent I built was the one I turned back into a program. The cheapest one will be the same agent once I stop handing it files it doesn't need. Same lesson as the whole series, one layer down: keep taking work away from the model until the only thing left is the thing a model is for.

This is an epilogue to the game-factory series. The most related earlier posts are the Builder agent and the lessons.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.