Stories about Terminal-Bench-Science
1 related stories
Claude Fable 5.1 made me a really nice animated pelican
AI InsightAnthropic released Claude Fable 5.1, with Terminal-Bench-Science score jumping from 24.7% (Fable 5) to 52.6%. Unlike previous models where reasoning was optional, Fable 5.1 forces five reasoning intensity levels without an off switch. This indicates Anthropic is making deep reasoning a default foundational capability rather than an add-on toggle.Key TakeawayCompared to previous toggleable reasoning, the new model mandates 5 reasoning levels.Why It MattersShifting reasoning from optional to mandatory defaults will restructure developer invocation habits and token cost expectations.Who's Affected- DevelopersMust adapt to 5-level mandatory reasoning; can no longer turn it off to cut costs or latency.
- AI ResearchersSignificant gains on science benchmarks from forced reasoning provide new data for scaling laws.
What's NextWatch the specific latency/cost impact of mandatory 5-level reasoning on general apps and whether competitors follow suit.Importance 70/100