
TLDR: Fable 5 is the first Mythos-class model the public can touch, and one question follows it everywhere: is this AGI? I don't have the answer. Anthropic never uses the word, and the one team with a standardized AGI eval couldn't run it. What I do have is two days of evidence: it runs longer between errors than any model I've used, it checks its own work unprompted, and it nearly drained my token allowance doing it. A real upgrade. Whether it's more than that is now a question that costs money to ask.
What Fable 5 Is (and Where Mythos 5 Fits)
A quick orientation, from Anthropic's announcement:
It's the first Mythos-class model released to the public. Anthropic calls it "a Mythos-class model that we've made safe for general use." Mythos 5 is the identical model with safeguards lifted in some areas. The names are the same word in two languages: fable comes from the Latin fabula, "that which is told," akin to the Greek mythos. The safeguards are the only difference between them.
Mythos 5 stays restricted. It deploys through Project Glasswing (United States cyberdefenders and infrastructure providers), with a trusted-access program planned for biology researchers.
Built for long-horizon autonomy. It works on its own for longer than any previous Claude model. Stripe ran a codebase migration through it in one day that they'd estimated at two months of manual work.
Memory compounds. It stays focused across millions of tokens, and giving it persistent file-based memory improved its performance three times more than the same setup improved Opus 4.8.
Vision took a real jump. It rebuilds web app source code from screenshots and beat Pokémon FireRed with a vision-only setup, where earlier models needed helper tools.
Science is the headline for Mythos. Scientists preferred its research hypotheses in roughly 80% of blinded comparisons, and Anthropic's protein design experts reported around ten times faster drug-design work.
The safety design is a router. Separate classifier systems watch for cybersecurity, biology and chemistry, and model-distillation requests. Flagged requests get answered by Opus 4.8 instead. More than 95% of sessions involve no fallback, and for those sessions Fable 5 performs the same as Mythos 5.
Pricing: $10 per million input tokens, $50 per million output. Included in paid Claude plans through June 22, then it moves to usage credits while Anthropic builds capacity.
That's the launch story. Here's what two days with it actually looked like.
The Handoff
I had a curriculum review hanging over me. Every lesson needed a read-through and update based on a new format: flow, accuracy, latest instructions that actually work on a student's machine. The kind of job I kept deferring because it needs sustained attention across dozens of files, and sustained attention is the thing I never have two days of. I tried executing it with Opus 4.8, ran into quality issues and parked it for a later time.
So i resumed the same session in claude code and switched the model to Fable 5. My actual first message was casual, messy, not very organized. But I knew it could find the context within my files and catch up quick.

What happened next is the part worth writing about. It created a working log and proposed a phased review plan. It asked which decisions were mine to make before touching anything.


Then it ran: reading every lesson in sequence, keeping its own work log in a markdown file, spawning four parallel agents to verify claims against official Anthropic and OpenAI documentation. I never asked it to do that verification. It decided the lessons made factual claims, and factual claims get checked.
Some lessons came back rewritten from scratch. Some came back with surgical edits. Terminology, numbering, and cross-references stayed consistent across all 31 files. I went through the output expecting the usual cleanup pass. I found virtually nothing to fix.
Two Days, 21 Messages
I pulled the session transcripts afterward, because I wanted numbers instead of vibes.
Across the whole review:
665,000 output tokens, longest running loop of 29 minutes, 137 file operations, 6 subagents, and exactly 21 messages from me.
Reading those 21 back, almost every one is a decision it had escalated: certification grading mechanism, renumbering of lessons, updating capstone mechanics etc. Zero of them are me interrupting it mid-action or steering it back on course. Its longest unattended stretch was 29 minutes of continuous execution off a single "proceed."
Same week, same model, different task: a content strategy session where I sent 56 messages in an afternoon. That one was co-thinking, and co-thinking is chatty by nature. The point is the model handled both modes and seemed to know which one it was in.
Anthropic's launch claim is that Fable's advantage grows with task length. My version is small and says the same thing:
The longer the run, the more the model's habits paid off, because the habits were continuity, verification, and writing things down.
Fewer Mistakes Is the Feature
The benchmark gap is real: 80.3% on SWE-Bench Pro against Opus 4.8's 69.2%, and the testers at Every scored it 91/100 on their senior-engineer benchmark, where Opus got 63. Simon Willison called it "a beast" after it bundled a project he'd been stuck on for months.

The number that changed my week was three. Only 3 tool errors across 137 file operations, each one self-corrected on the next call. Opus 4.8 is a strong model and I run my whole system on it. It still produces the occasional confident wrong turn that I catch two steps later. Fable's runs came back clean, and where it wasn't sure, it flagged the uncertainty instead of papering over it. Rakuten's launch quote matches what I saw: at highest effort, the model checks and validates its own work.
“At highest effort, the model checks and validates its own work.”
Fewer errors compounds exactly like task length does. A model that's wrong 5% of the time needs supervision every few minutes. A model that's wrong almost never, and says so when it might be, can hold and execute a two-day job.
Anthropanic Returns
Now the bill. I wrote a dispatch about Anthropanic earlier this year: the anxiety of watching token limits burn during a workflow you can't pause. The SpaceX data center deal in May fixed it, and I stopped thinking about tokens.
Fable brought the feeling back in one session. The curriculum review pushed me to 80% of my daily allowance on the 5x Max plan. One review. Simon Willison reported $110 in a single day of API testing. The model is verbose by design, because verification is output, and every one of those tokens is billed at 2x Opus rates.

This was mid session, 16 minutes into the Curriculum review task
For long running tasks, the math pays off. Two days of curriculum review would have cost me a week of evenings. Against that, the burn is cheap. Against a quick bug fix, it's absurd.
The Speed Tax
It is slow. Noticeably. The multi-stage checking that makes the output trustworthy also makes every response feel deliberate, and reviewers are converging on the same practical guidance: it's a poor fit for interactive chat. When I want speed, I switch back to Opus, or Sonnet for quick edits, and lose little.
My working rule after two days: Fable for anything long, complex, and well-bounded, where I can describe the outcome and walk away. Opus as the daily driver. Sonnet when latency matters more than depth.
So, Is It AGI Yet?
Honest answer: I don't know, and right now i don’t think anybody does.
Anthropic never uses the word. The launch post and the system card describe a "Mythos-class model made safe for general use" and spend their boldest language on risk, not intelligence. The loudest AGI claims are coming from TikTok demos. Meanwhile the ARC Prize team, who run the closest thing the field has to a standardized AGI test, had early access and couldn't run their verified evals: the new data-retention terms for Mythos-class models ruled it out. The one group holding a measuring stick wasn't able to use it.
So the public evidence is anecdotal, mine included. My two days showed longer runs, fewer errors, self-checking, and judgment calls I never asked for. An upgrade, a very clear one. But my runs didn’t show anything categorically different. It carries no goals from one job to the next unless I hand them over. The autonomy is rented and scoped, task by task.
And here's the part I think is genuinely new: I can't afford to find its limits. Probing the frontier now is gated by economics. Every serious test is a 90% day on my plan and the people who can afford bigger experiments are bound by terms that keep the results private. Although, who knows. Maybe AGI arrives looking like a model that quietly keeps a work log.
“Is it AGI" has become an experiment most of us can't run.”
When to Use Fable 5
When should I reach for Fable 5? Long, complex, bounded work: a multi-file refactor, a full content audit, a research synthesis. If the job would take you days of sustained attention, it's a Fable job.
When should I stay on Opus or Sonnet? Anything interactive or quick should be Sonnet - Chat, short edits, brainstorming, fast iteration loops. The speed difference is real and the quality gap on short tasks is small.
What does it cost? Double Opus 4.8 per token, and verbose on top of that. Budget per finished task: a $50 run that replaces a week of evenings is a bargain.
Will it eat my subscription limits? Yes. One long review took me to 90% of a daily 5x Max allowance. If you’re on the Pro plan, you will run out of daily limits very quickly. Plan your heavy runs; don't start one an hour before you need the tokens for something else.
How should I hand work to it? Bound the scope, tell it the outcome, ask it to keep a work log, and let it escalate decisions instead of guessing. Then leave it alone. Interrupting it mid-run spends the advantage you're paying for.
What about the guardrails? Cybersecurity and biology topics get routed to Opus 4.8 automatically. Anthropic says under 5% of sessions; independent measurement says closer to 8. For everyday building and writing work, I never hit them.
Open Questions
Does the error rate hold at month three, or am I in the honeymoon window every new model gets?
What happens to the economics after June 22, when metered credits land?
If the standard evals can't run and the probes cost $100 a day, who gets to answer the title question, and would we believe them?
P.S: I am hosting a free Webinar on Maven on building Claude Skills. Jun 16th, 12 PM EDT. Register here
Spot the patterns in your work worth bottling into skills
Explore Anthropic Skills & Plugins
Build a skill the Orchestrator Method™ way
Run workflows on command or schedule them
Written with Fable 5, naturally. It mined its own session transcripts for the numbers in this piece.
That’s it Folks
Thanks for reading through.
I’d love to know how you felt about today’s newsletter. This will help me make the newsletter better.


