Dev.to AI 🤖 Ai 👁 0 📖 8 min read

Claude Code vs Codex CLI in 2026: Which Subscription Should You Actually Pay For?

Both terminal agents are free to install. Both will happily burn through your entire monthly plan in a single afternoon if you let them. And in 2026, choosing between Claude Code and Codex CLI is less about which model i

Both terminal agents are free to install. Both will happily burn through your entire monthly plan in a single afternoon if you let them. And in 2026, choosing between Claude Code and Codex CLI is less about which model is smarter and more about which meter you would rather live with.

Because the meters are genuinely different now. Anthropic throttles by clock time: usage windows that reset every five hours, plus a weekly cap. OpenAI moved Codex to a credit system in April 2026, and credits are just tokens wearing a costume. One charges you for the hours you spend working. The other charges you for how much code the agent reads and writes.

I have not run a controlled 30-day trial on both. What follows comes from pricing documentation, public usage-limit guides, and developer reports, with the sources linked where a claim depends on them. Treat it as a decision framework, not a benchmark.

The plan landscape at a glance

Both tools ride on their parent chat subscriptions. There is no separate "Claude Code plan" or "Codex plan."

Claude Code (Anthropic side):

  • Pro: $20/month. Roughly 40 to 80 hours of Sonnet-class work per week according to third-party limit trackers, split into five-hour rolling windows.
  • Max 5x: $100/month, five times the Pro allowance.
  • Max 20x: $200/month, twenty times Pro.
  • Weekly caps apply across all models on every tier.

Codex CLI (OpenAI side):

  • Go: $8/month, light use, and no cloud task delegation.
  • Plus: $20/month, the default for most developers.
  • Pro 5x: $100/month, five times Plus usage.
  • Pro 20x: $200/month, twenty times Plus.
  • Free tier exists for kicking the tires.

The symmetry is not an accident, and the price points being identical ($20, $100, $200) makes the comparison cleaner than it has ever been. Same money, very different mechanics.

The core difference: time-based vs token-based

Anthropic limits your active hours. The commonly reported figures put Pro at roughly 40 to 80 hours of Sonnet usage per week, with a five-hour rolling window on top. If you code two hours a day, you will almost never see the wall. If you spend a Saturday on a 12-hour refactor marathon, you will hit it, even if the tasks themselves were small.

OpenAI limits your tokens. Since April 2026, every Codex task burns credits. One credit works out to roughly four cents when you cross-reference the credit math against OpenAI's own API rates. OpenAI's help documentation says a typical task on GPT-5.5 burns 5 to 45 credits, which is a 9x spread between a light fix and a heavy refactor on the same plan.

This distinction changes which plan "feels" generous. The token meter punishes agents that read forty files before writing ten. The time meter punishes you for working a long day, regardless of task size.

Where your money actually goes

Codex's rate card is public enough to do real math with. Three token types feed the meter: input, cached input, and output.

  • GPT-5.3-Codex: 43.75 credits per million input tokens, 350 per million output, an 8x output multiplier.
  • GPT-5.5: 125 credits per million input, 750 per million output, a 6x multiplier.
  • Cached input runs at roughly 10% of fresh input on every model.

That last line is the single biggest lever in the whole system. Reusing context costs about a tenth of reloading it. OpenAI's own cost guidance makes the same point from another angle: a bloated AGENTS.md file loads its tokens on every single task, so trimming it is the cheapest bill reduction available. Same logic applies to disabling MCP servers you are not using in a session.

Anthropic does not publish an equivalent per-token rate card for plan users, because the plan is not billed per token. Your cost lever on Claude Code is different: it is model selection. Running Opus instead of Sonnet burns your weekly allowance several times faster, which is why the limit guides all repeat the same advice: default to Sonnet, save Opus for the tasks that genuinely need it.

The five-hour window is the real decision

If you only remember one thing from this article, make it this: usage windows reset every five hours on both platforms.

A five-hour window plus a weekly cap means Claude Code is tuned for a specific shape of work: focused sessions, a few hours at a time, multiple sessions per day. Developers who report hitting Pro limits are almost always marathon coders. The forty-hour weekend sprint is exactly the profile that runs dry.

Codex's credits do not care about your schedule. They care about task weight. A day of tiny fixes barely dents Plus. A single migration that reads your whole repo can eat a visible chunk of your week's allowance in one task. The GitHub discussions about Codex limits fill up with developers who hit the weekly cap after a handful of full-scale sessions.

So the honest question is not "which is better?" It is: do your days look like many small tasks, or a few enormous ones?

  • Many small tasks (bug fixes, quick features, questions about the codebase): Codex Plus stretches further, because small tasks burn 5 to 15 credits, not the 45-credit ceiling.
  • A few enormous tasks (migrations, large refactors, whole-system reviews): both plans hurt, but Claude's time-based throttle at least lets you do unlimited small thinking around the big task. On Codex, every file the agent reads is on the meter.
  • Long autonomous runs: Codex has the edge here in raw structure, because cloud task delegation lets you hand off background work rather than sitting in an interactive session burning a window. Note that Go does not include delegation; you need Plus or higher.

When the API beats both subscriptions

Both companies publish numbers that let you find the breakeven point, and third-party calculators put it in a surprisingly low place.

Running Codex through your own API key skips credits entirely. At standard token rates, GPT-5.3-Codex costs $1.75 per million input tokens and $14 per million output. A typical CLI session lands around $0.50 to $2.00 at those rates. Do the division and the API tends to beat the $20 Plus plan somewhere under roughly 10 to 40 sessions per month. Under about 50 to 200 sessions, it can even beat the $100 tier. These are directional numbers, not exact ones, because session weight varies wildly.

The trade-offs of the API route are real, though: no automatic GitHub code review, no Slack integration, and new models usually reach ChatGPT subscribers before API users. If you are a light user, the subscription is a bad deal. If you are a daily driver, it is a very good one.

Claude's equivalent calculus is fuzzier because the plans are throttled rather than metered. The practical version: if you regularly exhaust the Pro window before lunch, price out whether the $100 Max tier costs less than the API sessions you would otherwise run. For most solo developers it does.

The security angle most comparisons skip

If you are picking a tool that will edit code autonomously, the sandbox matters as much as the pricing. The two tools made opposite bets.

Codex CLI sandboxes at the kernel level. On macOS it uses Seatbelt, on Linux Landlock plus seccomp, denying syscalls below the application layer. Three explicit modes: read-only, workspace-write, danger-full-access. A hostile agent literally cannot touch filesystem areas you did not allow. The philosophy is stronger boundaries with coarser control.

Claude Code sandboxes at the hooks layer. It offers granular, pattern-based allow and deny lists per tool, plus a sandboxed Bash mode with a network allowlist. More control surface, more programmability (hooks can run arbitrary scripts), and therefore a larger surface for misconfiguration. Weaker default boundaries, finer control.

For solo work on your own machine, either is fine if you leave the defaults sensible. For letting an agent loose on a repo with production credentials in the environment, Codex's kernel-level denial is the more forgiving failure mode.

What I would actually do

Given all of the above, here is the decision path I would hand a friend picking one subscription this month:

  • Trying things out: both free tiers, then Codex Go at $8 if you want a paid taste without committing.
  • Daily driver on a budget: Codex Plus ($20) if your work is many small tasks, Claude Pro ($20) if you work in long interactive sessions and want the stronger in-repo reasoning reputation.
  • Heavy but solo: Claude Max 5x ($100) if you keep hitting five-hour windows, Codex Pro 5x ($100) if your tasks are huge and token-hungry.
  • Light or sporadic use: skip subscriptions entirely. API keys on both, and you will likely spend less than $20 most months.
  • Any team rollout: budget for the meter, not the seat. The plan price never moves; the credits behind it do. This is the gap that catches teams off guard, and finance surveys back it up: in one 2026 survey of 260 finance leaders, 46% called managing AI spend their most stressful responsibility.

One more practical note: both tools coexist in the same repository without conflict. Claude Code reads CLAUDE.md, Codex reads AGENTS.md, and the two files are independent. Plenty of developers keep both installed and pay one subscription at a time, switching by the project.

The bottom line

This comparison is closer than the fan wars suggest, because at $20 the two plans are aimed at two different shapes of work rather than two different quality levels. Claude Code sells you protected time in five-hour blocks. Codex sells you tokens and lets you spend them however you want. Pick the plan that matches your calendar, not the one that matches your favorite model, and check the API math before you assume the subscription is the cheap option.

And whatever you pick: trim that AGENTS.md. Your future bill is hiding inside it.

I write about developer tools, AI infrastructure, and backend engineering every week. Subscribe, it is free, and it saves you from reading the fan wars to get to the pricing math.

Which one is your daily driver right now, and have you actually hit a usage wall on it? I am curious whether the five-hour window or the credit meter is biting people more in practice.

Further reading:

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.