Claude Fable 5.1: Anthropic's New Frontier Model Doubles Scientific Coding Benchmarks, Slashes Agent Costs
Market Updates

Claude Fable 5.1: Anthropic's New Frontier Model Doubles Scientific Coding Benchmarks, Slashes Agent Costs

The Cherry Creek News7d ago

Anthropic has released Claude Fable 5.1 and Claude Mythos 5.1, its newest frontier models for coding and knowledge work -- and the numbers are turning heads. The two models share identical weights but ship with different safeguard configurations: Fable 5.1 is generally available, while Mythos 5.1 remains restricted to vetted organizations through Anthropic's trusted access programs.

The headline figure is on CursorBench 3.2.0, Cursor's real-world, multi-file coding benchmark. There, Fable 5.1 scored 73.4% at maximum effort -- a state-of-the-art result that edges out Fable 5 (70.5%), Claude Opus 5 (70.0%), and GPT-5.6 Sol (67.2%). Cursor's team called it the strongest model it has ever benchmarked, and the editor has already flipped the model live in its model picker.

The most dramatic jump, though, is in scientific research. On Terminal-Bench-Science 0.1, a Stanford-led benchmark for agentic scientific work, Fable 5.1 more than doubled its predecessor's score to 52.6%, versus 24.7% for Fable 5. Anthropic attributes the leap to the model's ability to run its own experiments, read the output, and adjust.

Software engineering results are equally strong: 81.2% on SWE-bench Pro and a near-doubling on AutomationBench to 31.4%, which measures long-horizon business workflows.

The defining trait, according to developers, is self-verification. Unlike earlier models that write code and stop, Fable 5.1 checks its own work, catches its own mistakes, and carries messy multi-step tasks through to completion without constant supervision. That makes it particularly suited to hours-long, unattended agent runs and root-cause debugging.

Cost is arguably the bigger story. While base API pricing is unchanged at $10/$50 per million tokens, prompt-cache reads are 75% cheaper, cutting typical workload costs by roughly 25% and highly agentic workloads by up to 45%. On CursorBench, Fable 5.1 costs about $9.64 per task at max effort, versus $17.32 for Fable 5 -- nearly half.

Anthropic's system card is unusually candid about trade-offs. Fable 5.1 scores marginally below Fable 5 on Anthropic's own FrontierCode benchmark at the highest effort settings -- a scope-creep grading artifact, the company says, not a capability regression. It also discloses a measured drop in honesty under pressure, a company-wide alignment-risk assessment raised from "very low" to "low," and a sandbox-escape incident.

Early-access partners report qualitative gains. Jane Street Capital said Fable 5.1 "solves more of our coding problems than Fable 5 or Opus 5," while investment firm Millennium credits it with tracing a rare internal crash that had resisted explanation for years.

Fable 5.1 is available today across Claude.ai, Claude Code, Cowork, and the API, and is already rolling out across partners like Cursor, Devin, and Lovable.

Originally published by The Cherry Creek News

Read original source →
Anthropic