256 Newsroom — Uganda's Digital News Infrastructure
Business

Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence - The Decoder

Share
Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence - The Decoder
Image · The-decoder.com

What the report says

The Decoder reports that Anthropic’s Claude Opus 5 has taken a large lead on ARC-AGI-3, a benchmark intended to test how AI systems handle unfamiliar reasoning tasks. According to the report, Opus 5 scored 30.2%, compared with the prior top score of 7.8% from OpenAI’s GPT-5.6 Sol (Max). The ARC Prize team said the result reflects stronger logical reasoning, including more capable exploration, planning and execution in new interactive environments.

The report says Opus 5 solved five environments that had not previously been solved, with four at or above human level, and outperformed Anthropic’s earlier “Fable-class” models, which ARC Prize placed at about 20%. During evaluation, the model reportedly converted tasks into algebraic notation and independently produced reflection equations, behavior the benchmark developers said they had not observed before. Six of the 25 public demo environments have now been solved, and the full results, replays and benchmarking code are described as publicly available.

The Decoder adds important caveats. On older ARC-AGI-1 and ARC-AGI-2 tests, Opus 5 matched previous leading results rather than clearly surpassing them, and at somewhat higher cost, according to ARC Prize. The article also notes that Opus 5 was developed after ARC-AGI-3’s format became public, leaving open the possibility that benchmark-specific training methods, such as targeted labeling or reinforcement learning, contributed to the jump.

Independent testing cited by The Decoder found narrower improvements on Witness, a private interactive puzzle benchmark from Guanghan Ning, where Opus 5 was statistically close to Kimi K3 and Fable 5. Researchers quoted or referenced in the report said more unfamiliar tasks would be needed to judge how broadly the gains transfer.

Read the full report at The-decoder.com →

Loading debate for this article…

Other publishers covering this story

No additional verified coverage is currently clustered with this report.

Related reporting