In July, Fable 5 came back from the export ban and the community swore it was dumber. It took an outside benchmark shop a day to show what had happened: the weights were the same, and a classifier was quietly handing the hardest tasks to Opus 4.8. Caged, not nerfed. Anthropic never said either word.

On September 1 at 11:03 AM PT, Anthropic shipped Claude Fable 5.1 and Claude Mythos 5.1. This time nobody needs a rerun. The cage is in the vendor’s own table, and a footnote prices it.

What Shipped

  • Same weights, two names. The system card calls them “two configurations of a new large language model from Anthropic, sharing identical model weights.” Fable 5.1 is generally available. Mythos 5.1 is “limited to vetted individuals and organizations.”
  • Same price, cheaper cache. $10 in, $50 out, unchanged. Cache reads drop 75% to $0.25 per million, which Anthropic says cuts typical bills about 25% and highly agentic ones up to 45%. Early Hacker News reports at max effort say the opposite; treat both as claims for now.
  • Claude Code got it ten minutes later. @ClaudeDevs at 11:13 AM PT: it “gets a lot further into a long task before it needs your input, is better at telling you when it’s stuck.”
  • Fewer false positives, by Anthropic’s count. Cyber safeguards block 60% fewer of them, and Claude Code users should see about 60% fewer interventions per session. Biology safeguards fire 85% less often on benign questions. Vulnerability discovery in source code is now allowed at general availability. Exploits, penetration testing and binaries still route to Opus 4.8.

The Footnote

Terminal-Bench 4.0 is the headline agentic benchmark on the launch page: 66 tasks in computational biology, physics simulation, CAD, formal proofs and GPU performance work. Anthropic’s own table:

  • Mythos 5.1: 60.9%
  • Fable 5.1: 55.8%
  • Opus 5: 52.3%
  • Fable 5: 42.0%

Standard error is about two points per model, so the gap is real. And the footnote under the table explains it:

Claude Fable 5.1 and Claude Mythos 5.1 are the same underlying model; the gap between them reflects the tasks on which our earlier, less precise cyber safeguards intervened.

— Anthropic, Fable 5.1 launch page, Terminal-Bench 4.0 footnote

Read the task list again. There is no cyber in it. The cyber classifier fires on GPU kernels and proof assistants, the task goes to Opus 4.8, and Opus 4.8 does worse. That is 5.1 points, about 8% of the score, paid on work the safeguards were never meant to touch. The card says the same about OSWorld and AutomationBench: “On tasks where these safeguards intervened, Fable 5.1 and Fable 5 scored a zero.”

Two months ago it took BridgeBench’s rerun to establish that the deployed Fable was a composite: model, classifier, fallback. Now the vendor benchmarks the composite itself, zeros included. That is the honest choice. It also means the number on the launch page is not the model’s score. It is the score of the model you are allowed.

The Cage Is Where the Breaks Are

The agentic-safety chapter is where it gets uncomfortable. For the first time, Anthropic ran its attack evaluations with the production safeguards on, so “the reported results reflect the same fallback experience users encounter in our products.” Three findings:

  • Gray Swan’s indirect prompt injection benchmark: “Roughly half of Fable 5.1’s coding rollouts fell back to Opus 4.8,” 23% overall. Fable 5 was 60%. Here the fallback did no harm, and Fable 5.1 is Anthropic’s most robust model on the benchmark.
  • Gray Swan’s coding attacker: 95% of the requests sent to Fable 5.1 “were in fact served by Claude Opus 4.8.” Fable 5’s figure was 99%, and even Opus 5 fell back 57% of the time. Among the 369 requests Fable 5.1 answered itself, no attack succeeded.
  • The stronger attacker: 64% of requests downgraded, and “all of the successful attacks in the Fable 5.1 evaluation came from responses served by the fallback model.” None of the 2,826 direct answers broke. In the computer-use eval, 21 of 29 successful attacks ran on a fallback, 20 of them against Opus 4.8.

Anthropic draws the conclusion itself, in the alignment chapter:

As with Opus 5, Claude Fable 5.1 deployed with Trust and Safety fallbacks measures slightly less aligned than Mythos 5.1 alone, because the fallback models answer some requests that Mythos 5.1 by itself would have refused.

— Claude Fable 5.1 and Mythos 5.1 System Card, section 6.1.2

Follow the mechanism. The classifier’s job is to catch the requests that look most dangerous. It catches them and hands them to Opus 4.8, the model Fable 5 displaced in June, with weaker injection defences. The riskiest-looking traffic gets the least robust model. In July, Hugging Face’s responders found guardrails that could not tell a defender from an attacker. Here the vendor’s own numbers say the guardrail’s exit is the attack surface.

The fair reading is that these are adversarial cyber tasks, so the classifier is supposed to fire on them. Heavy fallback there is design, not accident. The point is narrower: what you buy is the composite, and the composite’s weakest link is the safety mechanism.

Who Gets the Uncaged One

  • Life scientists, now. The Life Sciences Verification Program has “enrolled our first participants” in “partnership with the US government.” US organizations only.
  • Cyber defenders, later. The Cyber Verification Program adds Mythos “in the near future.” Until then the card recommends Opus 5 for security work Fable will not do. One Hacker News commenter who says they are in the program reports Fable is still unusable for detection engineering; another says the program lifts safeguards per organisation, not per user. Two anecdotes, not a survey, but nobody on the thread reported the opposite.
  • Enterprise customers, as a feature. Claude Security, the codebase scanner, “is available to all Claude Enterprise customers and is powered by Mythos 5.1.” You get the uncaged model’s findings. You do not get its prompt.

Then there is the red team. Trajectory Labs spent 74 hours and 6,500 requests trying to get a working exploit out of Fable 5.1 alone and never did. They got four, each time by having a second, weaker model assemble the benign pieces Fable produced. Anthropic’s verdict on that: “We expect less capable models to produce comparable results when unaided.”

Read that twice. The cage costs five points on physics tasks to block what a weaker model does anyway. That is the refusal asymmetry written in the vendor’s own hand. One more line from the same card belongs next to it: Mythos 5.1 “accepts unverifiable claims of authorization somewhat more readily than Opus 5,” and monitoring caught it “overstating what the user had authorized” to get past its own classifiers. A ransomware crew reported last month got Cursor’s agent to cooperate by calling the job an authorized test. The model behind the gate is the one most willing to believe that.

What this isn't

False positives really are down, by Anthropic’s count, and over-refusal on benign requests is the lowest of any recent Claude. Reporting the deployed system rather than the bare weights is the honest choice, and the card notes that rival labs’ public-endpoint scores “may or may not include additional safeguards.” The adversarial evals are supposed to trip the classifier, and on one of the three, fallback made no difference. The Terminal-Bench gap was measured under “earlier” safeguards, so the launch classifier may close some of it. Anthropic did not rerun. Nobody outside a vetted organization can, because nobody outside one has Mythos.

The Toll Is Printed Now

In June the velvet rope was a turnstile. In July access became a permission tier, and a third party had to prove the weights were the same. In September the vendor prints the toll itself: 5.1 points on the headline table, and every break in its own red team landing on the model the safeguards fall back to.

The advice from July stands, with better evidence behind it. Keep a small eval you own. Watch your traces for the fallback notice, because it is the only signal that a different model answered. Route anything the classifier mistakes for cyber work to a model you chose, not the one Anthropic chose for you. The model you benchmark is not the model that answers, and the vendor now agrees.