← Field notes

AI systems

Fable 5, the Government Shutdown, and What Frontier AI Really Costs

Fable 5 launched benchmarking well above Opus 4.8, got pulled by the US government three days later, came back more restricted, and leaves Claude subscriptions on July 12. Here is what that actually costs, using my own June usage numbers.

A red velvet rope barrier between brass stanchions, representing access to Fable 5 being cut off

Fable 5 launched benchmarking well above Claude Opus 4.8, got pulled by the US government three days later over jailbreak concerns, came back a few weeks after with tighter guardrails, and loses its place in Claude subscriptions on July 12, 2026. After that date, using it costs real, metered money. I ran my own June numbers through it, and the figure was bigger than I expected.

What actually happened to Fable 5?

Anthropic launched Claude Fable 5 on June 9, 2026, alongside a sibling model, Mythos 5, tuned more for scientific and research work. Fable 5 was pitched as the agentic coding and computer-use model. Within days, both were gone, and the story of why is more interesting than the launch itself.

That is the part that should have been the story. It was not, for long.

How much better is Fable 5, actually?

Not evenly better, which is the honest answer. Anthropic's own published benchmarks show a model that dominates on long, autonomous, high-stakes work and only nudges ahead on everything else.

Benchmark comparison table showing Claude Fable 5 scoring highest across SWE-Bench Pro, FrontierCode, GDPval-AA, Blueprint-Bench, AutomationBench, OSWorld-Verified, Legal Agent Benchmark, Humanity's Last Exam, BioMysteryBench, Terminal-Bench 2.1, ExploitBench, and HealthBench Professional, compared against Claude Opus 4.8, GPT 5.5, and Gemini 3.1 Pro
Source: Anthropic, "Introducing Claude Fable 5 and Mythos 5."

Pull out the numbers that actually matter for a small business deciding what to build on:

  • Coding (SWE-Bench Pro): 80.3% for Fable 5, versus 69.2% for Opus 4.8. This is the headline gap, and it holds up.
  • Terminal work (Terminal-Bench 2.1): 88.0% versus 82.7%. A real lead, not a rounding error.
  • Knowledge work (GDPval-AA): 1,932 versus 1,890. Here the gap nearly disappears. If your work is mostly knowledge tasks rather than agentic coding, Opus 4.8 is not meaningfully behind.
  • Spatial reasoning (Blueprint-Bench 2): 38.6% versus 14.5%, close to the roughly 3x improvement Anthropic described publicly.
  • Cybersecurity (ExploitBench): 78.0% versus 40.0%. Nearly double. Keep this number in mind for the next section.

Anthropic also published task-level results, not just benchmark scores. Fable 5 migrated a 50-million-line codebase for Stripe in one day, work that would normally take about two months by hand. In Slay the Spire, a game-playing benchmark that rewards remembering what happened many turns ago, its persistent memory got it to the final act roughly three times more often than Opus 4.8. Mythos 5, the sibling model built for research, was reportedly used internally to accelerate protein design work by around ten times, and scientists preferred its molecular biology hypotheses over a human baseline in blind comparisons about 80% of the time. None of that is Fable 5's own workload, but it tells you the two models launched from the same very capable base.

Here is the FrontierCode chart Anthropic published, plotting accuracy against mean cost per task as you push the model's reasoning effort higher:

Line chart titled FrontierCode: Accuracy vs Cost, showing Claude Fable 5 reaching over 30 percent accuracy at max reasoning effort for roughly 20 dollars per task, well above Claude Opus 4.8 and GPT-5.5 at the same cost
Source: Anthropic, "Introducing Claude Fable 5 and Mythos 5."

The shape of that chart is the whole pricing story in one picture. Opus 4.8 flattens out around $10 a task. Fable 5 keeps climbing well past that, up to roughly $20 a task at maximum reasoning effort. That capability is real. It is also the reason the cost conversation later in this post matters, because Anthropic is not giving away access to a model that scales like that for free.

Why did the US government pull it three days after launch?

On the evening of June 12, 2026, three days after Fable 5 and Mythos 5 went live, Anthropic disabled both models worldwide. The company had received a US export-control directive at 5:21 PM ET that day, citing national security authorities, after reports that someone had found a jailbreak that could turn Fable 5 into a tool for finding software vulnerabilities in a target codebase. The order technically only applied to foreign nationals, but the only way to comply was to switch the models off for everyone.

Look back at that ExploitBench score from the section above: 78.0% for Fable 5, against 40.0% for Opus 4.8, on a benchmark that specifically measures exploit-finding capability. That is not a small capability jump on the exact skill a jailbreak would need. Whatever you think of how the government handled it, the underlying concern was not invented from nothing.

Anthropic pushed back publicly, saying it disagreed "that the finding of a narrow potential jailbreak should be cause for recalling a commercial model deployed to hundreds of millions of people." I will be honest about how this reads from where I sit: whatever the actual security calculus, a government pulling your model three days after launch because it is judged too capable is the kind of headline most companies could never buy. I am not saying that is why it happened. I am saying it did not hurt Fable 5's reputation.

What changed when Fable 5 came back?

The Commerce Department lifted the directive on June 30, and Fable 5 returned to global service on July 1 with a new safety classifier built specifically to block the reported jailbreak. Anthropic reported the classifier triggers a fallback in less than 5% of sessions on average, and that more than 95% of Fable 5 sessions involve no fallback at all. That is a genuinely low false-positive rate for a brand-new classifier, and worth saying plainly because it would have been easy to assume the fix made the model unusable day to day. It did not.

Anthropic was upfront that this still comes with a real tradeoff: the classifier blocks more benign coding and debugging requests than before, and when it does, your request gets rerouted to Opus 4.8 instead of simply refused. Anthropic's own words on the underlying problem were blunt: "it is probably impossible to make any AI model fully robust to jailbreaks." That is not a company overselling its fix. Worth remembering next time a vendor promises a safety patch closes the book for good.

Why is Fable 5 leaving subscriptions on July 12?

Here is the part that matters more for anyone actually building with it. Fable 5 came back included in Pro, Max, Team, and select Enterprise plans, covering up to half of your weekly usage limit. That was always meant to be temporary. The cutoff was first set for July 7, then pushed to July 12 after backlash over the early date. After that, Fable 5 moves off your plan's included usage and onto prepaid usage credits billed at Anthropic's published API rates: $10 per million input tokens and $50 per million output tokens, on top of whatever you already pay for your subscription. For reference, that is less than half the per-token price of the earlier Mythos preview release, so this is already the cheaper end of what frontier-tier metered pricing looks like right now.

Anthropic says this is not meant to be permanent, and that Fable 5 will return to standard subscription access "as soon as capacity allows." I believe that is a genuine intention. I also think it tells you something true about how this industry actually works: a model that costs up to $20 a task to run at full reasoning effort, per that FrontierCode chart above, is not something any lab can subsidise indefinitely at scale.

What would my own Fable 5 usage actually cost at those rates?

In June, my total usage across Claude models came to about 19 million tokens. Opus 4.8 alone accounted for more than 74% of that, roughly 14 million tokens. None of that cost me a line-item cent beyond my subscription, because it sat inside my plan's limits.

Now recast that same 14 million tokens at Fable 5's post-July-12 metered rate.

The same month's usage, priced two ways

  • On a Max subscription (flat): $100/month (5x tier) or $200/month (20x tier), covering this and everything else within your weekly limit.
  • That same ~14 million tokens, recast at Fable 5's metered rate ($10 input / $50 output per million): roughly $140 if it skewed almost entirely input, up to roughly $700 if it skewed heavily output. Realistic coding-agent work sits somewhere in between, so call it several hundred US dollars, for one model, in one month, for one person.

Scale that across a small team, or a year, or add in whatever Fable 5-specific usage sits on top, and "thousands" stops being an exaggeration. It is closer to a floor.

That gap, a flat $100 to $200 a month versus several hundred dollars for a slice of one model's usage, is the entire story of why subscriptions exist in the first place. They are a bet by the lab that your usage will average out below the metered price, and a bet by you that it will.

Is a subscription still worth it once Fable 5 goes metered?

For me, yes, with a condition. I will still reach for Fable 5 on the handful of tasks where its lead over Opus 4.8 actually shows up: a long refactor, a genuinely hard debugging session, something that runs for a while without me steering it, the kind of work that lines up with that 80.3% SWE-Bench Pro score rather than the near-tied GDPval-AA number. I am not going to route my everyday work through it once it is metered, because the everyday work does not need the extra capability, and a small business cannot casually absorb frontier-model API pricing as a background cost. That is not a knock on Fable 5. It is just the honest math of running a small operation instead of a funded lab.

The broader point I keep coming back to: inference is not free to run, no matter how generous the plan looks today. Every lab currently subsidising your usage through a subscription is making a bet that will get repriced eventually, the same way Fable 5 just did. Build like that is true, because it is.

What should a small business actually do about this?

  • Know which of your workflows actually need frontier-tier capability, and which ones just feel better because the model is newer. The GDPval-AA gap above shows plenty of knowledge work does not need it.
  • Track your own token usage by model, monthly. You cannot make this call with a vague sense that "AI costs are going up." You need your own number, the way I needed my 19 million.
  • Treat any "included in your subscription" frontier access as a temporary discount, not a permanent price. Plan your workflows so a repricing does not break them.
  • If your work genuinely lives in the long-horizon, high-capability zone (research, security work, complex multi-step builds), budget for metered API cost on purpose, rather than getting surprised by it.

Key takeaways

  • Fable 5 launched meaningfully above Opus 4.8 on long, agentic, autonomous work (80.3% vs 69.2% on SWE-Bench Pro), but nearly tied on knowledge work (GDPval-AA).
  • The US government forced a three-day-old model offline worldwide over a reported jailbreak, then lifted the order after 18 days, after Anthropic's own ExploitBench numbers showed Fable 5 nearly doubling Opus 4.8's exploit-finding score.
  • Fable 5 returned with a stricter safety classifier that Anthropic says triggers a fallback in under 5% of sessions.
  • From July 12, Fable 5 leaves standard subscription usage and moves to metered API pricing: $10 input / $50 output per million tokens.
  • Recasting one month of my own real usage at that rate turned a $0 line item into several hundred dollars, for one model alone.

FAQ

Is Fable 5 gone for good after July 12?

No. It stays fully available on the Claude API and consumption-based Enterprise plans. It is leaving the included usage on Pro, Max, Team, and select Enterprise plans, moving to metered usage credits instead.

Does the Pro plan get Fable 5 too, or only Max?

Both, along with Team and select Enterprise plans, each got Fable 5 included for up to half their weekly usage limit during the temporary window. It was never Max-only.

Why did the government actually shut it down?

An export-control directive citing national security authorities, after reports of a jailbreak that could let Fable 5 be used to find software vulnerabilities in a target codebase. Anthropic's own benchmarks show Fable 5 scoring 78.0% on ExploitBench against Opus 4.8's 40.0%, which is the kind of gap that explains the concern even if you disagree with how it was handled.

Is Opus 4.8 a fine replacement for everyday work?

For most day-to-day tasks, yes. The benchmark gap between Fable 5 and Opus 4.8 is small to nonexistent on knowledge work and quick tasks, and only opens up wide on long, complex, autonomous work and security-adjacent tasks.

What happened to Mythos 5, the sibling model?

Suspended and restored on the same timeline as Fable 5, since both were named in the same export-control order. Mythos 5 is tuned for research rather than coding, and Anthropic has reported it accelerating internal protein-design work roughly tenfold, though that is a different workload from anything covered in this post.

Is GPT-5.6, which launched around the same time, a similar story?

Different story, same moment. I have covered GPT-5.6's tiers and pricing separately, since it is a straight three-tier lineup rather than a launch-shutdown-return arc. Worth reading if you are deciding what to build on next.

If you want to try Claude Code yourself before deciding any of this, I have a referral link that gets you a free week. No cost to you either way, and it is the same tool I used to pull these numbers.

Extra Resources

Work with me
WhatsApp us