Prompt Tricks Died. Context Engineering Didn't
In this video: 0:00 Do 1M windows end choosing? No 0:19 Prompt engineering is half right 1:25 Two reflexes: polish or dump 2:34 Personas, stuffing and the experiment 5:15 Selection, order and length 8:15 Tools, schemas
In this video:
0:00 Do 1M windows end choosing? No
0:19 Prompt engineering is half right
1:25 Two reflexes: polish or dump
2:34 Personas, stuffing and the experiment
5:15 Selection, order and length
8:15 Tools, schemas and examples
10:58 Costs, and when stuffing is fine
12:37 Prompt, context and harness layers
Most developers assume a million-token window means you can stop choosing what goes in. That assumption is wrong. Chroma tested eighteen models in 2025, and reliability fell as inputs grew, even on simple retrieval. The phrasing tricks faded, but deciding what enters the window is still the job.
Rewording the prompt changed nothing
Take a summarisation feature on a current frontier model and rewrite the instruction ten different ways: shorter, longer, politer, with a persona, with a step-by-step nudge. Practitioners often report that the output moves far less than expected.
So the claim that prompt engineering is over is half right. Persona and magic-phrase tricks don't reliably improve accuracy. Tone and behaviour still live in prose, though, and vendor guidance stresses precise instructions plus the logic and data the task needs.
Now leave the wording alone and change the material around it. Swap which documents were retrieved, reorder them, add a tool, trim the history. The same feature gives a different answer. That failure is real, and it was never in the wording.
The claim worth defending is that one technique has been confused with the whole job. Phrasing was the technique. The job is controlling what the model sees. Anthropic's engineering team says as much: it calls context engineering the natural progression of prompt engineering. The discipline was extended, not killed. If your feature is still wrong, rewording was never going to fix it.
Two reflexes: polish the words, or dump everything in
When the output is bad, a competent developer reaches for one of two fixes, and both feel like engineering.
The first is to polish the words. You are a world-class expert. Take a deep breath and think step by step. Then come the bribes, the threats, the elaborate politeness. This is not superstition from nowhere. On earlier models, a line like that could nudge the output toward a register: more formal, more careful, closer to what you wanted. You watched it work, so you keep reaching for it.
The second reflex is to stop selecting at all. Claude Sonnet 5.5, released in late September 2026, has a one-million-token window. Gemini 3.8 Flash accepts 1,048,576 input tokens. GPT-6.1 Sol is reported at 1,050,000. Choosing documents is work, and work has bugs. So why choose? If the window holds that much, include everything and let the model sort it out. It looks free.
The two reflexes share something. Both treat the model as the variable to tune and the input as fixed. One coaxes the model with better words. The other hands it more material and trusts it to cope. Neither asks whether the input itself was the problem. And there is a measurable reason to doubt that dumping is free.
What the persona and reasoning evidence says
Start with the persona line, because it has been tested. A report on a December 2025 Wharton study, covering six frontier models on PhD-level questions, found that telling a model it is an expert did not reliably improve factual accuracy. A report on an EMNLP 2024 study, with 162 personas across four model families, found no persona strategy that consistently beat picking one at random.
Both reached me through secondary write-ups, so read the papers before you quote them. The claim is also narrow. Personas can still shape tone, and some writers still recommend them. The finding is that they are unreliable for accuracy, not that they are useless.
Reasoning has moved too. Claude's docs list the old manual extended-thinking setting as unsupported on the Claude 5 models, and adaptive thinking is the control that replaced it. On Sonnet 5.5 it is on by default, at high effort. Thinking is now a setting you configure, not a sentence you write.
A test: wording versus selection
Here is the test, on Claude Sonnet 5.5. Thirty retrieved chunks sit in a fixed pool. In the first run, three chunks stay fixed and the instruction, summarise this, is reworded fifteen ways: shorter, longer, polite, with a persona, with a step-by-step nudge. In the second run, the wording stays fixed and only the choice changes, meaning which three of the thirty go into the window. Every output is scored the same way.
The expected result is that the spread across wordings comes out small next to the spread across chunk selections. But that is an expectation, not a finding. One caveat matters here: I found no primary source that has measured this contrast, so this is one run on one model, not a law.
Where the dump-everything reflex breaks
Chroma's Context Rot study tested eighteen models and found that reliability falls as input grows, even on simple tasks like retrieval and text replication. Lost in the Middle found position effects: models did best when the relevant passage sat at the start or end of the input, and worse in the middle.
Both studies used older models. Chroma's were mid-2025 systems, and I found no replication on the million-token models shipping now. Treat them as a reason to test your own pipeline, not as a verdict. Thirty chunks is also nowhere near a million tokens, so my experiment tests selection, not rot.
Then there is the bill. GPT-6.1 Sol is reported to have a window of 1,050,000 tokens, but prompts above 272,000 input tokens bill at double the input rate for the entire request. That figure comes from a secondary source, so check the pricing page. Even so, the lesson holds: the size of the window is not the size of the range where use is reliable or cheap. A window is capacity. It does not tell you what belongs in it.
The window is an attention budget
Every token you add to the window is paid for in attention. Anthropic's engineering guidance, as summarised in a knowledge-base write-up, calls context an attention budget. The reasoning is that a transformer relates every token to every other one, so n tokens means n squared pairwise relationships. Add tokens and each one gets a thinner share. Context is not storage. It is a budget, and every passage you include competes with the passage that actually answers the question.
Chroma's study, run on 2025-era models, points the same way. Distractors hurt performance, unevenly, and how similar the needle was to the question changed the results. That hints the worst filler is not random noise but the near miss: the chunk that looks relevant and isn't. Treat that as a hypothesis for current models.
Four levers: selection, order, length, timing
The first lever is selection. Choosing the top three of thirty chunks is a ranking problem and a retrieval-precision decision. If your ranker puts a plausible wrong chunk in the top three, no instruction will rescue the answer. That is the bet worth testing: the choice of chunks is the variable to tune, not the wording. No published measurement of that contrast turned up, so run it on your own task and model. In practice that means measuring retrieval precision as a first-class number, and cutting at a hard count instead of a loose similarity threshold.
The second lever is order. In Lost in the Middle, models used evidence at the start or end of the input more reliably than evidence buried in the middle. That was measured on older models, so the rule is a hypothesis for yours. Put the question-critical material at an edge and the question itself near the end. Then test on your model by moving the key passage through the positions and watching the score.
The third lever is length, and the tool for it is compression. Anthropic's multi-agent pattern, as reported in a summary, lets a subagent burn tens of thousands of tokens on a search and hand back a condensed summary of roughly one to two thousand. The parent agent never sees the forty pages of dead ends. It sees the conclusion. The work is isolated, and only its result enters the window you care about.
The fourth lever is when the data enters. Anthropic describes a just-in-time strategy. Instead of loading every document up front, the agent holds lightweight identifiers: file paths, stored queries, links. It pulls the real data in at runtime, through tools. Claude Code does this with large databases, and it looks at the head and the tail of an object rather than loading the whole thing. Anthropic also notes that Claude Code caps tool responses at twenty-five thousand tokens by default.
The agent learns what exists, then opens only what it needs, then opens the next thing. That is progressive disclosure. The window stays small because the reference is cheap and the load is deliberate.
Selection, order, compression and deferral all decide what earns a place. None of them asks how big the window is. A million tokens only raises the ceiling on how much you can get wrong.
Tool descriptions are context too
Documents are only part of what the model sees. Before your first retrieved chunk arrives, the model has already read your tool definitions. Names, descriptions, parameter docs: all of it is loaded into the window, and all of it steers the agent.
Anthropic's engineering team calls prompt-engineering your tool descriptions one of the most effective ways to improve tools, and its advice is plain. Write the description as if you were onboarding a new hire. Make the implicit explicit: the query format you assume, the niche terminology your team uses, the meaning of an ID.
A tool described as "searches tickets" leaves the model guessing. Searches by what? Does it return titles or full threads? Is the date a string or a range? Claude's own docs draw this line. A good description says what the tool does, when to use it, what data it returns, and what each parameter means. The poor one is too brief, and the gaps are exactly those four points. The model fills them with guesses, and you see the guesses as wrong tool calls.
The other direction matters as well. A tool's response is a selection problem too. If a tool returns every matching record, you have rebuilt the dump-everything reflex inside a function call. Anthropic suggests combining pagination, range selection, filtering and truncation, with sensible defaults, so the first response is small and the agent asks for more on purpose. The Claude Code cap mentioned above is the same idea, enforced by the harness.
Output schemas and examples
An output schema constrains the space the model can write into. Fields, types, allowed values: each one removes a whole family of wrong answers before generation starts.
Examples do similar work, with one warning. Anthropic's guidance favours a few diverse, canonical examples over a pile of edge cases. Three examples that cover the typical shapes pin the format. Twenty edge cases teach the model that the edge cases are the norm, and they cost attention every time.
The ground moves, and prose still has a job
A caution about how fast this ground moves: on Claude Sonnet 5.5, forced tool use now returns an error. The switch many of us used to make a model call a particular tool is gone on that model. Steering tool choice now leans on what you wrote in the descriptions and what else is in context. Your prompt didn't change. The harness under it did.
Prose still has a real job in this window, just a different one. OpenAI's prompting guidance for GPT-6 Astra tells the model to infer intent from context and show a bias towards action. A phrase like "can you" or "help me" is to be treated as a request to act, not an invitation to ask a follow-up. That is a behaviour spec. It is not a persona and not a magic phrase. It states a policy the model cannot read off the documents.
So the window is more than documents. It is tools, schema, examples and a few lines of policy, all competing for the same attention budget. Once all of that is in place and the model is updated underneath you, how do you know it still works?
What this costs, and when to just paste it in
Selection is code, and code has bugs. An aggressive retriever or summariser can drop the one fact the answer needs, and the final answer will still read fluently. That is why the selection step needs its own evals, separate from the output. Did the needed chunk survive into the window? Scoring only the final answer hides exactly this failure.
Price is a selection variable too. Anthropic lists Claude Sonnet 5.5 at two dollars per million input tokens, against twenty cents for cache reads. GPT-6.1 Sol's cached input is reported at ten cents per million, though that comes from a secondary source, so check the pricing page.
Either way, the layout has a cost. Use a stable prefix with a volatile tail. Tools, schema and examples go first, where the cache can reuse them. The retrieved chunks that change on every request go last.
The savings come from tokens used, not wording. Anthropic credits Sonnet 5.5 with needing far fewer tokens to do the same work, for up to thirty percent lower cost per task than Sonnet 5.
So when is plain stuffing fine? When the corpus is small. When the analysis is a one-off. When the whole input sits far below the degradation and billing cliffs, and the reported 272,000-token cliff on GPT-6.1 Sol is one of them. In those cases, curation is the wrong investment. Paste it in and move on.
One more limit, on me as much as anyone: the evidence is thin. The long-context degradation studies predate the 2026 models, so no rule of thumb here has been confirmed on the models you are actually running. Measure your own model and your own task. The practical question is what that measurement looks like, and how you keep it running when the model changes underneath you.
Prompt, context and harness: three nested layers
Practitioners now split this work into three layers. The prompt is the wording. The context is everything the model sees. The harness is the tools, memory, constraints and feedback loops around the model. These layers are nested, not competing: the prompt sits inside the context, and the harness assembles the context.
The term context engineering was popularised in June 2025 by Andrej Karpathy and Tobi Lutke. It now appears in Anthropic's official engineering guidance, as the natural progression of prompt engineering.
Look at what agents in production actually run on. Retrieval decides which documents arrive. Tool design shapes what the agent can do. Long history gets compacted. Subagents hand back a summary of a thousand or two thousand tokens instead of forty pages of dead ends. The wording of the top-level instruction is the smallest part of that. The skill used to be writing the question. Now it is assembling everything around it.
Treat that assembly as code. Version it, so you can say which retriever, which tool descriptions and which schema produced which result. Log exactly what reached the window on each request, not what you meant to send. And evaluate it: the selection step and the final answer, on your model, and again whenever the model changes.
Persona and magic-phrase tricks don't reliably improve accuracy, according to secondhand reports of persona studies. Behaviour is still specified in prose, as OpenAI's GPT-6 Astra guide shows. The window still does not fill itself. At a million tokens, deciding what goes in is harder, not easier. And that decision is still your job.
Sources & further reading β Prompt Tricks Died. Context Engineering Didn't
- Claude Sonnet 5.5 - Claude Platform Docs (Anthropic, 2026) β https://platform.claude.com/docs/en/models/sonnet-5-5/overview Sonnet 5.5's one-million-token window, adaptive thinking on by default replacing manual extended thinking, forced tool use returning an error, and the $2 per million input and $0.20 cache-read pricing.
- Introducing Claude Sonnet 5.5 (Anthropic, 2026) β https://www.anthropic.com/claude-sonnet-5-5 Sonnet 5.5's release timing and the claim that it needs fewer tokens for the same work, cutting cost per task by up to 30 percent versus Sonnet 5.
- Gemini 3.8 Flash (Google AI for Developers, 2026) β https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash Gemini 3.8 Flash accepting 1,048,576 input tokens.
- Effective context engineering for AI agents (Anthropic Engineering) β https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents Context engineering as the natural progression of prompt engineering, context as an attention budget, subagents returning condensed summaries, just-in-time retrieval via lightweight identifiers, and Claude Code's 25,000-token tool-response cap.
- Writing effective tools for AI agentsβusing AI agents (Anthropic Engineering, 2025) β https://www.anthropic.com/engineering/writing-tools-for-agents Prompt-engineering tool descriptions like onboarding a new hire, making implicit conventions explicit, and returning small responses via pagination, filtering, range selection and truncation.
- Define tools - Claude Platform Docs (Anthropic) β https://platform.claude.com/docs/en/agents-and-tools/tool-use/define-tools A good tool description says what the tool does, when to use it, what it returns and what each parameter means; too-brief descriptions leave gaps the model guesses at.
- OpenAI shares prompting tips for GPT-6 Astra including a blocklist of slop words (The Decoder, 2026) β https://the-decoder.com/openai-shares-prompting-tips-for-gpt-6-astra-including-a-blocklist-of-slop-words/ OpenAI's GPT-6 Astra prompting guidance (secondary report): infer intent from context and bias toward action, as an example of prose acting as a behaviour spec.
- GPT-6.1 Sol: Features, Benchmarks, Pricing, and Access (DataCamp, 2026) β https://www.datacamp.com/blog/gpt-6-1-sol Reported GPT-6.1 Sol figures (secondary source, to be checked against the pricing page): roughly 1,050,000-token window, double input rate above 272,000 tokens, and cached input pricing.
Disclosure: this article is the written companion to the video above. Its script was drafted with AI assistance and the narration is an AI voice.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.