When two models cost exactly the same, the price is the least useful thing you can know about them. GPT-6.1 Sol and Claude Sonnet 5.5 both launched within 24 hours at $2 per million input tokens and $10 per million output — the identical sticker. So the real decision is: pick GPT-6.1 Sol if your team lives in Codex and ChatGPT and wants raw interactive speed; pick Claude Sonnet 5.5 if you live in Claude Code, want lab-published agentic-coding benchmarks, or need Bedrock/Vertex/Azure and data-residency options. And whichever you lean toward, decide on tokens-per-task from your own repo, not the launch-day charts. Here's why.

Here's the whole decision in one screen:

Why the identical price changes the question#

For two years, "which model?" was mostly a cost question, because the frontier models were priced far apart and you traded quality for dollars. That trade just disappeared for this tier. Anthropic shipped Sonnet 5.5 at $2/$10 on Sept 28, holding the same rate as Sonnet 5. A day later, OpenAI shipped GPT-6.1 Sol at the same $2/$10 — pitched as near-frontier "Astra-like" intelligence at roughly a fifth of the flagship's price. Two of the strongest agentic-coding models in the field, released a day apart, on the exact same number.

When the rate is identical, the interesting differences are the ones the rate card hides. There are three that actually matter.

1. The harness moves your velocity more than the model#

Both models ship inside a coding agent, and that agent — not the weights — is what your engineers touch all day. GPT-6.1 Sol runs natively in Codex (and ChatGPT Work); Sonnet 5.5 runs natively in Claude Code, and is also available on the Anthropic API plus AWS Bedrock, Google Vertex and Azure Foundry with a zero-data-retention option. How each harness edits files, runs your tests, reviews a diff, and asks permission before a destructive command will change throughput more than a two-point benchmark gap ever will.

The practical read: if your team already lives in one of these, that's a strong default — the switching cost of retraining habits and rebuilding config usually swamps the model delta. If you're greenfield, the Bedrock/Vertex/Azure availability makes Sonnet the easier adopt inside a hyperscaler with data-residency rules; the Codex-and-ChatGPT surface makes Sol the easier adopt if your org is already standardized there.

2. The real cost lever is below the sticker price#

Here's the number that actually decides your bill. A coding agent doesn't send one prompt — it loops, re-sending the same file tree, system prompt, and tool definitions on every turn. Over a real task that's tens of thousands of repeated input tokens. Which is why cache-read pricing and tokens-per-task swing your cost far more than the base rate that happens to match.

Both labs price cache reads well below the input rate — Anthropic lists cache reads at $0.20 per million, a tenth of the input price, and OpenAI prices cached input lower still. And Anthropic's headline efficiency claim cuts the other way: Sonnet 5.5 "costs up to 30% less per task" not through a lower rate but by using fewer tokens to finish the same work. Read that carefully: it is a token-efficiency win, not a price cut. So your true cost-per-task is (tokens the model burns) × ($2/$10), discounted by however much of your context is a cheap cache read — and that product can differ by multiples between two models sharing one sticker. This is the same logic behind routing each request to the cheapest capable model and behind why cache-read pricing quietly dominates an agent's bill.

OpenAI's counter-lever is speed, not efficiency: an Ultrafast tier that clocks up to ~300 tokens/sec in Codex at a premium rate (it rolled out first for the flagship and is coming to Sol). If your bottleneck is a human waiting on a generation, raw speed may be worth the surcharge; if it's your monthly invoice, token efficiency wins.

3. Don't decide on a launch-day benchmark#

It's tempting to settle this with a leaderboard. Don't. Anthropic published specific, lab-run agentic-coding numbers for Sonnet 5.5 — Terminal-Bench 4.0 at 70.6%, CursorBench 4.0 at 55.5%, OSWorld 2.1 at 80.1% — on its launch page. GPT-6.1 Sol's coding figures come mostly from third parties and use different benchmark versions, so putting them side by side is comparing two rulers with different inches. And the number most people reach for, SWE-bench Verified, was not reported by either lab for these models — so any head-to-head SWE-bench claim you see is invented, not launched.

The honest, useful test is the one you run yourself. This is the same discipline as reading a coding benchmark critically before trusting the headline: a public score is a starting hypothesis, your repository is the experiment.

How to actually pick, this week#

  1. Take 10 representative tickets from your real backlog — the mix you actually ship, not toy problems.
  2. Run each through both Codex (Sol) and Claude Code (Sonnet 5.5).
  3. Record three numbers per model: did it pass, total tokens per task, and cache-read hit rate.
  4. Compute true cost as tokens-per-task × $2/$10, adjusted for cache hits. This is where the identical sticker fractures into a real difference.
  5. Weigh the harness ergonomics your team will live with every day — the permission model, the diff review, the test loop.

Two models, one price. The launch charts want you to pick a winner; your backlog will actually pick one. When the rate card is a tie, the tiebreaker is your own tokens — go measure them.