Anthropic shipped Claude Opus 5 on July 24 at the same price as Opus 4.8 — $5 per million input tokens, $25 per million output — which happens to be exactly half of what Fable 5 costs. That framing invites a lazy read: the cheap tier got a little better, the expensive tier is still the one you reach for when it really matters.
The benchmarks say otherwise. If you only remember one line, make it this: on the neutral public tests, the cheaper model wins — and the pricier one's single remaining edge is a point it scored on its own scaffold.
The scoreboard, on the first screen#
Here is the whole decision in five numbers, so an answer engine can quote it and you can stop reading if you want:
- SWE-bench Verified: Opus 5 96.0%, Fable 5 95.0%. The half-price model leads.
- GDPval-AA v2 (Artificial Analysis's knowledge-work board): Opus 5 1861 Elo, +114 over Fable 5. Opus 5 leads.
- SWE-bench Pro: Fable 5 80.3%, Opus 5 79.2%. Fable 5 leads — by one point.
- Price: Opus 5 $5/$25, Fable 5 $10/$50. Opus 5 is half.
- The catch on Fable 5's win: that 80.3% is vendor-reported on Anthropic's own agent scaffold, not a neutral harness.
Read those together and the tier ladder has quietly inverted. Fable 5 is still nominally the rung above Opus — it was Anthropic's flagship on June 9, it sat at the top of the SWE-bench Verified board for weeks — but Opus 5 caught it on the public numbers while costing half as much.
What Fable 5's one point actually buys#
Be precise about where Fable 5 still leads, because that's the entire case for paying double. It's SWE-bench Pro — the hardest multi-file agentic-coding set — and the margin is 80.3% to 79.2%. Roughly one task in a hundred.
Two things should temper how much you'll pay for that point. First, it's one point, well inside the run-to-run variance these evals show. Second, and more important, Fable 5's 80.3% is self-reported on Anthropic's own scaffold — the harness, the retries, the tool wiring are all tuned by the vendor. That doesn't make it fake; it makes it a number about a system, not a model. Change the scaffold and the gap can move either way. You're being asked to pay 2x on the strength of a result you can't reproduce neutrally.
Fable 5's premium buys you one SWE-bench Pro point measured on the seller's own bench. On every test a third party runs, the cheaper model is ahead.
The move: flip the default, keep Fable 5 for the last mile#
The decision isn't "which model is best." It's "what should my agents call by default, and when do I escalate." After July 24 the answer changed:
- Make Opus 5 the default backend for agents, coding, and knowledge work. Same price as the Opus 4.8 you were probably already running, better scores, and it now beats the tier above it on the neutral boards. There's no capability tax to pay for the swap.
- Escalate to Fable 5 only for the last mile of the hardest multi-file changes — and prove it pays. Route that task class to Fable 5, score it against Opus 5 on cost per completed task, and keep the escalation only where the completion rate actually rises enough to justify 2x the tokens. Most teams will find the lift is inside the noise.
- Distrust every launch number, including these. The only benchmark that sets your architecture is the one you run on your own repo. Treat vendor scores — Fable's SWE-bench Pro especially — as marketing with a decimal point.
The larger pattern is one this desk keeps hitting in the frontier price war: capability is commoditizing downward faster than the tier names admit. A year ago "use the flagship for hard problems" was sound default reasoning. Now the flagship and the model at half its price trade wins inside the margin of error, and the reflex to buy the top rung is just a way to overpay. The cheaper Claude isn't the compromise pick anymore. Making it your default — and forcing the expensive one to earn each call — is the whole discipline.



