Stories about ARC-AGI
1 related stories
Claude Fable 5.1 results on ARC-AGI
AI InsightClaude Fable 5.1 scores 97.5% on ARC-AGI-1 Semi-Private and 90.0% on ARC-AGI-2 Semi-Private at max effort, with per-task costs of $1.40 and $4.49. This sets a new reference point for abstraction reasoning benchmarks and cost-efficiency comparisons.Key TakeawayThe new model publicly reaches 90% on ARC-AGI-2 for the first time.Why It MattersA 90% score on the difficult ARC-AGI-2 quantifies abstraction capability and cost, raising the bar for future benchmark evaluations.Who's Affected- AI ResearchersCan benchmark their own models' abstraction reasoning gap against this score.
- DevelopersCan estimate inference costs and choose models using per-task pricing.
- Model ProvidersNeed to target the new 90% threshold on ARC-AGI-2.
What's NextWatch for ARC-AGI-3 benchmarks and whether competing models exceed the 97.5% ceiling.Importance 70/100