October 2, 2026
ProductA model guide for the GPTβ6 family
Practical tips for getting the best results from GPTβ6 models while managing time and cost
GPTβ6 is our most advanced suite of models yet, and offers you a choice of models for different kinds of work.
Whether youβre turning an idea into a working prototype, building and testing a feature, or orchestrating multi-step workflows across code repositories, databases, and external APIs, this guide explains how to choose a GPTβ6 model, give it effective instructions, manage long-running work, and prepare for production.

TL;DR
Run effectively in production. Use cachingβ (opens in a new window) and compactionβ (opens in a new window) to manage context and cost. Measure task success and latency, and plan for monitoring and data controls.
Match the model to your workload. Balance capability, cost, and latency by choosing the model, reasoning effortβ (opens in a new window), and speed that fit the task.
Adjust your prompts and skills. Keep prompts, skills, and repository instructions consistent about what the model should deliver, what it can do independently, and what counts as done.
Keep long-running work on track. Use steeringβ (opens in a new window), async toolsβ (opens in a new window), and delegationβ (opens in a new window) to handle updates and independent work. Set clear boundaries for when the model should ask for input.
1. Run effectively in production
Prepare your workflow for production
Before deploying, there are several checks and best practices youβll want to put into place.
Keep efficiency in mind. Cut context the task doesnβt needβ (opens in a new window) while keeping the evidence it does. Where your application supports it, run independent tasks togetherβ (opens in a new window) so one slow step doesnβt hold up unrelated work.
Reuse shared context through prompt cachingβ (opens in a new window) for recurring work. Cached input tokens cost up to 95% lessβ (opens in a new window) than uncached input tokens, depending on the model. Put stable instructions and reference material before changing task details, and keep tool definitions consistent. The caching dashboardβ (opens in a new window) and diagnostics guideβ (opens in a new window) help you see where that reuse breaks down. Include cache writes and any long-context rates when estimating the cost of a complete workflow.
For longer conversations, compactionβ (opens in a new window) reduces context size while preserving the state needed to continue.
Decide how youβll monitor behaviorβ (opens in a new window) and review the data controlsβ (opens in a new window) for your application.
Test before deploying: Run representative tasks and measure task success, latency, and cost per successful task. Check out our API deployment checklistβ (opens in a new window).
Match the model to the workload
Think of the model choice and reasoning level as an intelligence/ price tradeoff.
Model:
GPTβ6 Astraβ (opens in a new window) for the hardest reasoning work where maximum intelligence is needed.
GPTβ6.1 Solβ (opens in a new window) for complex coding, research, and computer use.
GPTβ6 Lunaβ (opens in a new window) for focused tasks at scale and everyday, repeated work with a clear goal, such as extracting invoice fields, classifying requests, or producing structured summaries.
When evaluating the best model for the task, compare pricingβ (opens in a new window) for each model.
Reasoning level: In the API, choose how much effort the model spends on the task.
Low: Routine tasks, such as extracting facts or making small edits.
Medium: Work requiring judgment, such as planning a feature or comparing options.
High: Difficult debugging, deeper analysis, or careful review.
Extra high / Max: Test where supported when High falls short, and keep only if the improvement justifies the added time and cost.
In the API, you can change reasoning effort mid-conversationβ (opens in a new window) without breaking cache.
In Codex, start with the default reasoning level for that model, then lower it for simpler tasks or increase it for deeper analysis.
Speed:
In the API, use Fast modeβ (opens in a new window) when response time matters, such as in chat apps or coding tools. It provides faster, more consistent response times at a higher per-token cost than Standard processing.
In Codex and the API, use Ultrafastβ (opens in a new window) when faster responses are worth the premium, such as rapid coding iterations. It speeds up token generation independently of reasoning effort. Available for GPTβ6 Astraβ (opens in a new window).

2. Adjust your prompts and skills
Give the model a clear assignment
βEric Provencher, Developer Experience at OpenAI
Start with a clear assignment: the result you want, who itβs for, the relevant context and constraints, and what counts as done. Then review these four areas, summarized from Rethinking skills and prompts for GPTβ6 Astraβ (opens in a new window), for a deeper dive into updating your instructions:
Create better skills: Keep descriptions short and explicit about when each skill should run, load supporting details only when needed, and replace rigid recipes with guidance suited to the models your team uses.
Update your AGENTS.md: Explain when particular documents and tests are relevant, and explicitly authorize safe routine workflows, such as running local tests with disposable data and no production access.
Set decision boundaries: State which actions can proceed independently and which require approval, replacing blanket βalways askβ rules with clear boundaries.
Be prescriptive about persistence: Define what βdoneβ includesβimplementing the change, running it, inspecting the result, and fixing failuresβand identify any decisions that require your review.
For additional guidance, see reasoning best practicesβ (opens in a new window).
Define the output you need
Whether youβre working in Codex or building with the API, specify which decisions the model can make, when it should ask for input, and what a useful response looks like.
Give the model enough direction to keep work moving without guessing at decisions that matter. Tell it which choices it can make and when to ask for your inputβ (opens in a new window), for example, it can choose how to organize a summary, but should check with you before changing the projectβs scope. Describe what a useful response looks likeβ (opens in a new window), for example plain language, technical detail suited to your audience, and a short handoff covering what changed, what was checked, and what still needs attention.
3. Optimize long-running tasks
Keep complex work moving
With the GPTβ6 family of models, you can now take on tasks that span hours or days. Use the following features to better manage agents on long-running tasks.
API
In the API, use steering, asynchronous tools, and parallel work to keep long-running tasks moving.
Update instructions during a run: Mid-turn steeringβ (opens in a new window) lets you send a correction through the Responses WebSocket APIβ (opens in a new window) while the model works. Updates are queued; they donβt cancel running tools or undo completed actions.
Keep working while a tool runs: Asynchronous tool callingβ (opens in a new window) lets the model continue independent work while your app runs a slower task, such as tests. Your app returns the result when itβs ready. Wait for that result before starting work that depends on it.
Delegate independent subtasks: GPTβ6.1 Sol supports multi-agent workflows in the Responses APIβ (opens in a new window). It can assign independent work to subagents (such as investigating different parts of a codebase) and combine their findings into a final response. Multi-agent is currently in beta.
Codex
Long-running tasks can uncover decisions you wouldnβt necessarily anticipate in the initial prompt. Use clarification and steering to keep the work on course.
Answer questions as work progresses: With GPTβ6 Astra, Codex can ask for clarification while it worksβ (opens in a new window). Resolve questions that affect the next step, and specify which independent work can continue while you decide. If youβll be away, tell Codex which tasks can continue and when it should pause for your answer.
Redirect work when requirements change: Steer the active taskβ (opens in a new window) with new information, explaining what should change and what should stay the same. This helps avoid spending more time on an approach that no longer meets your needs.

Leverage computer use to do more of the job
Computer useβ (opens in a new window) lets GPTβ6 Astra, GPTβ6.1 Sol, and GPTβ6 Luna interact directly with websites and desktop apps, even applications without an API. For example, you can ask the model to investigate a bug, fix the code, and open your product in a browser to check that the fix works.
Choose the simplest reliable way to do each step:
Use an API or connected tool when it can do the job directly.
Use computer use when the model needs to read a screen, click buttons, or fill in a form for you.
If youβre building computer use into your own app, give the model a tool that can run code to control a browser or desktop. Playwright works with browsersβ (opens in a new window); PyAutoGUI works with desktop apps.

From testing to production: How teams are building with GPTβ6 Astra
Harvey: more context, more useful drafts.β Harvey combines court information, case law, firm documents, and a lawyerβs preferences to tailor its drafts. βWe can give more context to the model and produce better and better structured outputs,β says cofounder Gabe Pereyra.
Cognition: test results engineers can review.β Cognition uses GPTβ6 Astra inside Devin to test software and return evidence. In an iPhone-game example, Devin produced a simulator recording and a report separating checks that passed from areas left untested, making the remaining work easier to see.
Hex: from a business question to an interactive dashboard.β Hex uses GPTβ6 Astra to turn questions about sales-channel performance into written findings and interactive dashboards, including geographic breakdowns. It also asks the model to examine whether the numbers make sense and whether the analysis answers the business question.
Invideo: more control over the final edit.β Invideo uses GPTβ6 Astra to plan timeline edits and create custom effects that editors can refine. The company reports roughly three times the success rate on color-grading and correction tasks. A few editors also created about 50 effects in one day.


