OpenAI’s GPT-5.6 Sol claims new state-of-the-art scores on agentic benchmarks, but Anthropic’s Fable 5 keeps a roughly 15-point lead on SWE-Bench Pro. Here is the full comparison.
The contest for frontier AI leadership has tightened into a benchmark-by-benchmark battle. Just over a week after Anthropic’s Fable 5 returned to full availability, OpenAI answered with a new generation of its own: the GPT-5.6 family, made up of Sol, Terra, and Luna. Each side can now point to headline numbers where it wins, and the honest answer to “which model is better” increasingly depends on which test you look at.
A Turbulent Run-Up to the Launch
Anthropic entered the summer in a strong position. After introducing its Mythos model and granting early access to a select group of companies in its Glasswing Alliance, the company released Fable 5, a version of Mythos with additional safeguards around biology, cybersecurity, and AI research, to a much wider audience.
That momentum was interrupted by a dispute with the U.S. government over export controls, during which Anthropic was treated as a supply-chain risk. Access to Fable 5 and Mythos 5 was suspended for 19 days, from June 12 to July 1, 2026, until the Commerce Department lifted the relevant controls and service resumed.
OpenAI timed its counterpunch carefully. Shortly after Fable 5 came back online, it unveiled GPT-5.6 as three separately priced capability tiers, with Sol at the top and Luna at the bottom. The company argues this structure lets each tier improve on its own schedule instead of waiting for a single flagship release.
Where GPT-5.6 Sol Pulls Ahead
OpenAI has been vocal about the areas where Sol edges out Anthropic’s flagship. The most notable claims include:
- A score of 53.6 on Agents’ Last Exam, an evaluation of long-running professional workflows across 55 fields, which OpenAI says is 13.1 points ahead of Fable 5.
- A new high of 80 on the Artificial Analysis Coding Agent Index, roughly 2.8 points above Fable 5, achieved while using less than half the output tokens and time.
- New state-of-the-art results on BrowseComp at 92.2 percent and OSWorld 2.0 at 62.6 percent.
Efficiency is arguably Sol’s clearest distinction. On the broader Artificial Analysis Intelligence Index, it lands within a single point of Fable 5 while finishing tasks in 61 percent less time at roughly half the cost. For teams that measure cost per completed task rather than leaderboard positions, that gap could matter more than any individual benchmark win.
Where Claude Fable 5 Still Dominates
The picture is far from one-directional. Fable 5 holds a wide lead on SWE-Bench Pro, one of the most closely watched agentic-coding benchmarks. At its June launch, Fable 5 scored 80.3 percent, far ahead of the 58.6 percent posted by GPT-5.5, and it remains roughly 15 points above the 64.6 percent OpenAI reports for Sol.
| Benchmark | GPT-5.6 Sol | Claude Fable 5 |
|---|---|---|
| SWE-Bench Pro | 64.6% | 80.3% |
| Artificial Analysis Coding Agent Index | 80 | About 77.2 |
| Agents’ Last Exam | 53.6 | About 40.5 |
| BrowseComp | 92.2% | Below Sol |
| OSWorld 2.0 | 62.6% | Below Sol |
Why Both Companies Can Claim a Win
Independent observers note that Terminal-Bench and SWE-Bench Pro reward different kinds of work, so the split results are less contradictory than they look. A useful way to picture it is as two different job interviews for the same candidate.
SWE-Bench Pro resembles handing someone a half-finished house left behind by a previous contractor and asking them to identify exactly what is wrong and repair only that, without breaking anything else in the process. It rewards careful reading, navigating a large unfamiliar structure, and making a precise, contained fix, much like an engineer resolving a specific bug in a codebase they did not write. This kind of whole-codebase issue resolution still favours Claude.
Terminal-Bench 2.1 is closer to testing general skill around the entire house: setting up the electricity, getting the plumbing running, and recovering when a step fails midway. It measures how well a model operates a computer’s control panel, including installing software, configuring servers, running multi-step jobs, and troubleshooting failures. That style of terminal automation favours Sol.
A model can be a meticulous bug-fixer without being a great all-around operator, and the reverse is equally true. That is broadly why Fable 5 leads on the precise-repair test while Sol leads on the messy-workflow test.
The Harness Question
There is a second wrinkle worth understanding. These evaluations do not test the model in isolation; they also test the toolkit it is given. OpenAI evaluated its models using its in-house Codex toolkit, which is tuned to how its models prefer to operate. It is a bit like letting a contractor bring their own familiar tools rather than renting them a standard set.
A contractor handed unfamiliar equipment usually needs time to adjust before working at full speed, and even then some tacit knowledge takes longer to build. If that adjustment cost were factored in, Sol’s lead over Fable 5 would likely be narrower. Buyers should check whether Sol’s efficiency advantages hold up outside OpenAI’s own tuned environment before committing a production workflow to it.
The Product Fight Beyond the Models
Alongside the GPT-5.6 family, OpenAI announced ChatGPT Work, an agentic mode built on GPT-5.6 that can read emails, summarise texts, create presentations, run scheduled or recurring tasks, and organise loose ideas into charts and dashboards through a new Sites feature. It competes directly with Claude Cowork, which also completes multi-step, tool-using tasks, though the two differ in their plugin ecosystems and in how tightly each connects to its maker’s coding tools, Codex for OpenAI and Claude Code for Anthropic.
The bottom line is that both flagships are strongly capable and land close together overall, with each holding a genuine specialty rather than one being categorically ahead. Sol brings clear speed and cost advantages plus wins on long-horizon agentic tests, while Fable 5 remains the stronger choice for precise, large-codebase software repair. For most teams, the deciding factor will be the shape of their own workload, not the top of any single leaderboard.
Frequently Asked Questions
It depends on the task. Claude Fable 5 leads SWE-Bench Pro by roughly 15 points, making it stronger for precise bug-fixing inside large, unfamiliar codebases. GPT-5.6 Sol scores higher on terminal automation, long-running agentic workflows, and does so faster and at lower cost.
An export-control dispute with the U.S. government led to Anthropic being treated as a supply-chain risk. Access to Fable 5 and Mythos 5 was suspended for 19 days, from June 12 to July 1, 2026, until the Commerce Department lifted the relevant controls and service was restored.
They are three capability tiers within OpenAI's GPT-5.6 family, priced separately with Sol at the top and Luna at the bottom. OpenAI says the tiered structure allows each level to advance on its own tempo instead of waiting for a single flagship release.




