Stories about MarkTechPost
2 related stories
Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
AI InsightAnthropic released Claude Fable 5.1 and Mythos 5.1. Fable 5.1 scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5's 24.7%, while cache read cost drops 75% to $0.25 per million tokens. This signals a major leap in scientific terminal task automation and a sharp reduction in long-context reasoning costs.Key TakeawayPerformance doubles and cache read cost drops 75%.Why It MattersThe jump from 24.7% to 52.6% on a science benchmark is generational, and the 75% cache price cut directly lowers developers' inference costs, potentially accelerating terminal agent adoption.Who's Affected- DevelopersCache cost drops 75%, but three breaking API changes require migration.
- Enterprises1M context and cheaper cache make science R&D agents more economical.
- AI ResearchersThe score leap on Terminal-Bench-Science offers a new reference for evaluating scientific operation capability.
What's NextWatch for Mythos 5.1 safety evaluation under Project Glasswing and migration impact of the breaking API changes.Importance 70/100Lowest-Latency Inference APIs for Voice and Realtime Agents: A Time to First Token TTFT-First Benchmark
AI InsightThe first proposal of Token Time to First (TTFT) as a benchmark for inference APIs in voice and real-time agents, focusing on latency performance in LLM and speech processing.Key TakeawayIntroducing Token Time to First (TTFT) as a new benchmark for inference API performance.Why It MattersReducing latency is crucial for the performance of voice and real-time agents.Who's Affected- DevelopersHelps developers choose the appropriate inference API.
What's NextFuture may witness more optimizations for latency performance.Importance 60/100