Image created with Gemini. Image prompt: An overhead swimming pool composed of a grid of slightly mismatched, overlapping square photo panels with hand-drawn white ripple squiggles on flat turquoise water, small flat bar charts and simple stopwatch shapes in terracotta pink, lawn green and hot yellow arranged along the pool’s tiled edge. Flat saturated acrylic style with hard-edged shapes, high noon light, no shadows, no text, generous white space around the pool.
Anthropic has released Claude Fable 5, the first publicly available Mythos-class model that ranks #1 in our agentic real-world knowledge work benchmark GDPval-AA Claude Fable 5 shares the same underlying model as Claude Mythos 5, with added security guardrails for potentially
https://x.com/ArtificialAnlys/status/2064414308289937869
Exciting news: Claude Fable 5 ranks #1 on the new Agent Arena leaderboard! Fable 5 leads by the widest margin ever over Opus-4.8 and GPT-5.5 on two key signals: confirmed task success rate and praise vs. complaint, despite weaker steerability. If Fable can do something, it will
https://x.com/arena/status/2064807170714358193
Just used claude fable (aka mythos) to create this city block simulator complete with multi-agent traffic, live detection boxes + tracks, and day to night cycle. And it just one shotted it. This is gonna be fun — the gap between idea and execution just keeps collapsing.
https://x.com/bilawalsidhu/status/2064524211914223867
This chart from Anthropic is useful, since Agent Teams and Workflows are both very new and very powerful (and token hungry). On the other hand, maybe it doesn’t matter as a lot of the decisions about which approach to use is from the AI itself & it often uses them in combination
https://x.com/emollick/status/2063073955968123062
With multi-agent orchestration in Claude Managed Agents, you can use Fable to delegate work to dedicated agents using smaller models.
https://x.com/ClaudeDevs/status/2064394928948703406
Anthropic has a coding MOAT
https://x.com/scaling01/status/2064399642603802676
Anthropic’s new Claude Fable 5 is #1 on Terminal-Bench 2.1 at 88.0%, beating GPT 5.5 by 4.6%. Under the hood it’s their Mythos model with safeguards that route dangerous requests to Opus 4.8, which Anthropic claims triggers in under 5% of sessions. Available in Cline now!
https://x.com/cline/status/2064427461212045546
As of May 2026, more than 80% of the code we merge into Anthropic’s codebase was authored by Claude.”” Matches independent measures. There really is no sign this is slowing down (which doesn’t mean there aren’t organizational challenges to absorbing this much productivity gain)
https://x.com/emollick/status/2062580212768936344
Claude Fable 5 / Mythos 5 wins everywhere. I thought Fable 5 was just a nerfed Mythos Preview, but it’s literally better. SWE-Bench Pro: Fable 5: 80.3%, GPT-5.5: 58.6%. And the price is only 2x Opus 4.8: $10/input MTok, $50/output MTok. I don’t think GPT 5.6 can beat this…
https://x.com/Yuchenj_UW/status/2064396097075003739
Claude Fable 5 beats Pokémon FireRed only using vision – YouTube
https://www.youtube.com/watch?v=Ty_50J84fMY
claude fable 5 has solved CAD I asked it to make a model of a V8 engine It came back to me with a fully working model in under 10 minutes
https://x.com/aaronli/status/2064876123109089742
Claude Fable 5 is our first generally available Mythos-class model. It ships with new safety classifiers that may flag certain prompts in dual-use domains like cyber and bio. We’ve added fallbacks: a refused request retries on Claude Opus 4.8 instead of dead-ending.
https://x.com/ClaudeDevs/status/2064428347678220691
Claude Mythos 5 scores 161.29 on Anthropic’s ECI
https://x.com/scaling01/status/2064392088003756431
Claude Mythos 5 speeds up kernels by 430x – a 300x speedup equals 40 human expert hours of work
https://x.com/scaling01/status/2064392386520780945
Claude Mythos is next level. h/t @Lentils80 Look at this MacOS output. One shotted.
https://x.com/kimmonismus/status/2062843119864021404
Claude mythos will be on a completely different level. These outputs are insane
https://x.com/kimmonismus/status/2062805570982203820
Fable 5 is the biggest step up I’ve felt in our models since Opus 4.5 back in November. After 4.5 came out I uninstalled my IDE when I realized that I’d been doing 100% of my coding in a terminal for a few weeks. With Fable, it’s felt like Claude has stepped up from being a
https://x.com/bcherny/status/2064431111154053187
It’s state of the art on nearly every benchmark we tested and the lead grows the longer the task. Made safe for general release: cyber & bio requests fall back transparently to Opus 4.8 and 95%+ of sessions never see one. $10/$50 on the API, in paid Claude plans today.
https://x.com/mikeyk/status/2064392996288901392
The capabilities of Claude Code and Codex have expanded a lot in recent months, they added many ways to approach work (subagents, skills, goal, workflows, plugins, etc). Given the AI labs can use their own AI to help documentation, a surprising amount is effectively undocumented
https://x.com/emollick/status/2062510975513747606
This is a super exciting release – Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it’s SOTA on everything by a margin but I’ll add that *qualitatively* also, this is a major-version-bump-deserving step change forward
https://x.com/karpathy/status/2064409694761054332
Want to be very upfront about the subscription rollout (Pro, Max, Team, seat-based Enterprise). We’re giving users Fable 5 within subscription limits for 2 weeks, and then we’re taking it away. Here’s exactly what’s happening and why:
https://x.com/TheAmolAvasare/status/2064393574431764928
Today I’m publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast–much faster than the policy process was built to handle. The essay lays out where I think the technology is now, and the action needed to close the gap:
https://x.com/DarioAmodei/status/2064781775247950326
New #1 on PostTrainBench: Opus 4.8 (max reasoning) hits 37.23% — up from 28.56% for 4.7, the largest single improvement we’ve seen. Fable 5 runs underway now that AI research behavior is no longer silently degraded. PostTrainBench asks how well frontier AI can train weaker
https://x.com/thoughtfullab/status/2065096885514227876
Fable 5 lies 96% of the time. We were surprised by it’s skill… 🧵
https://x.com/kradleai/status/2064907897373642912
Fable 5 scores 81.9% on SimpleBench the highest score almost reaching human baseline.
https://x.com/JasonBotterill/status/2064699951578505446
Fable: “”create a visually interesting shader that can run in twigl-dot-app make it like an infinite city of neo-gothic towers partially drowned in a stormy ocean with large waves.”” “”Make it better”” All of this is procedurally generated.
https://x.com/emollick/status/2064424775527624736
Fable: “”write me a rhyming poem with six four line stanzas, each stanza removes another vowel. the first has no u, the second no u or i, etc.””
https://x.com/emollick/status/2064769268013584767
I think we’ve reached the point where normal people can’t really determine whether new models are better than previous ones. Like Fable doesn’t seem that much better to me, but every 150 IQ person I know is like “wow the singularity came sooner than I thought”.
https://x.com/citrini/status/2064480613852201336
I’ve had access to Fable for a bit. A genuine jump in capability, I could feed it a 15 page design document for a project and it would work for 9+ hours and deliver terrific results. But working with it is weird & weirder is coming Lots of examples:
https://x.com/emollick/status/2064395281903346013
More realistic example of a one shotted game. Asked Fable 5 to recreate a game in the style of The Elder Scrolls 5 Morrowind. It one shotted quests, currencys and fighting, journal and minimap. And it worked.
https://x.com/kimmonismus/status/2064744343349399634
my response to fable was to be quiet, still and humbled. i felt a panic that i didn’t have purpose anymore. it strictly dominates me as an engineer. i had a month of that before realising that more than caring about “how i achieved my goals”, i just cared about “my goals” and a
https://x.com/akbirkhan/status/2064418425552928812
My take 24 hours after Fable 5: Your organization will likely not scale with the exponential curve of AI. I’l just come out to say: This should be a wakeup call for engineering teams. Set up your cloud software factories. Now. Models can now fix impossible bugs, UI-test the
https://x.com/walden_yan/status/2064755974548902006
One interesting pattern with Fable 5 is that it will often say things that are gibberish when I use it for coding. Things like “”The morning’s slim-scan fix cured the scan hang””, “”this is a latent-drift API-shape wrinkle””, etc. When I ask why it does this, Fable explains that it
https://x.com/tamaybes/status/2065147305494450248
One thing I mentioned only in passing in my Fable post is that, for long running tasks, Fable starts to develop its own dialect as its many agents and tasks reinforce themselves and make Claudish language ever more Claudish. You need to ask it to report out in plain English.
https://x.com/emollick/status/2064542441848422611
“AI agents will outperform humans at almost all jobs by 2026-2027.” – The forecast is everywhere. So we built the exam to test that claim, on real labor-market aligned work. On the hardest tier, top agents pass 2.6%. Meet Agents’ Last Exam (ALE), a rolling benchmark measuring
https://x.com/YiyouSun/status/2064392466011394213
Agent Arena evals are fundamentally different. You can’t ask humans to judge hundreds of tool calls across a 30-minute trace. So we built something different. We break down how the Agent Arena Leaderboard mines real usage traces for objective signals to move beyond human
https://x.com/arena/status/2064748918135824876
Agentic AI is now evaluated in the Arena with Agent Mode and measured with Agent Arena. Founding Engineer Matt and Product Lead Ted show you Agent Mode in action: deep research, complex bash operations, whatever you throw at it. Every session contributes to the Agent Arena
https://x.com/arena/status/2062902033389322477
Anthropic’s newest model, Claude Fable 5, a Mythos class model for non-security work lands in Microsoft Foundry and GitHub Copilot today. Read the blog here:
https://x.com/Azure/status/2064421301108834552
Claude Fable 5 is now available in Devin. Fable 5 earns the #1 spot on FrontierCode, our benchmark for real-world engineering tasks that grades mergeability and quality:
https://x.com/cognition/status/2064398549073453266
Evaluating agentic coding systems past vibes is brutal. Yes, we can add: Tests PM reviews PR reviews CI/CD loops But how do we know the system works reliably? While speaking with @ben_burtenshaw from @huggingface at @UphillConf, he suggested a simple but powerful idea:
https://x.com/pauliusztin_/status/2062874580411162811
Everyone says the latest AI agents will be “”job-ready”” soon, especially after the release of Fable 5 this week. But is that really the case? Over the past many months, my group and collaborators have been building Agents’ Last Exam (ALE), a benchmark designed to test exactly
https://x.com/dawnsongtweets/status/2065095757988868190
Fable 5 is also by far the best computer use model according to Stagehand Agent Evals it also costs less than half of GPT-5.5
https://x.com/scaling01/status/2064812046902817051
Microsoft Research introduces Arbor A generalist autonomous research agent that uses persistent hypothesis-tree refinement to turn long-horizon exploration into cumulative learning. It beats Codex and Claude Code across 6 research tasks and hits 86% Any-Medal on MLE-Bench Lite.
https://x.com/HuggingPapers/status/2065062300218749172
we just shipped support for rubrics in deepagents ✅ give your agent a clear definition of what “”done”” looks like, and force it to run in a loop until said goal is complete this is similar to /goal in claude code, but works for any agent (not just a coding agent)
https://x.com/sydneyrunkle/status/2064034061165682931
We’ve open sourced my favorite Devin feature: /handoff Hand off jobs to cloud Devins from your local machine Install it as a plugin in Claude Code or Codex or any other coding agent Close your laptop without pausing your agents 😉
https://x.com/imjaredz/status/2065153770762154186
You can try Claude Fable 5 as part of Devin Cloud’s Ultra agent. Devin Ultra is our smartest and most capable agent, which excels at long-horizon tasks and debugging. We tuned the harness so Ultra costs only ~40% more than default Devin agent. Claude Fable 5 is also available
https://x.com/cognition/status/2064398551539761387
BREAKING: Anthropic just dropped Claude Fable 5–this is Mythos, made safe for public release. It is the best coding model in the world. We’ve been testing it internally @every for the last week or so across coding, writing, marketing, editing, and more–here’s our vibe check: –
https://x.com/danshipper/status/2064393970856124501
Claude 5 Fable tl;dr – It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research -The longer and more complex the task, the larger Fable 5’s lead over our other
https://x.com/kimmonismus/status/2064401121515274747
Claude Fable 5 (high) scores 87.8% and takes the lead on WeirdML. It’s the first model that scores above 70% on average on each separate task. It uses about 8k output tokens on average, almost as much as Opus 4.7 (high). EDIT: This post first said “”no thinking””, which is not
https://x.com/htihle/status/2065050640154350043
Claude Fable 5 changed how we work on the Claude Code team day to day. We used to verify that Claude did the work right. Now we verify that it’s doing the right work. Here’s the 3 biggest changes:
https://x.com/ClaudeDevs/status/2064399512664526853
Claude Fable 5 is now available in Computer as an orchestrator model. This is Anthropic’s state-of-the-art model for long, complex tasks. Available only to Pro and Max subscribers in Computer.
https://x.com/perplexity_ai/status/2064771411894567373
Claude Fable 5 is now available in Cursor. It sets a new state of the art on CursorBench at 72.9%, 8 points above the previous best.
https://x.com/cursor_ai/status/2064394824313376787
Claude Fable 5 is the champion negotiator on PACT, a multi-round LLM conversational bargaining benchmark! 🏆 It overtakes the previous champion, GPT-5.5.
https://x.com/LechMazur/status/2064815890651140447
Claude Fable 5 launched today at #1 on the Artificial Analysis Intelligence Index, putting Anthropic nearly 5 points ahead of any other lab’s best model We supported @AnthropicAI with pre-release evaluation of Claude Fable 5. Claude Fable 5 scores 64.9 on the Artificial Analysis
https://x.com/ArtificialAnlys/status/2064500150069030992
Claude Fable 5 ranks #1 on FrontierSWE. This represents the biggest capability jump we have observed since releasing the benchmark On many tasks, Fable 5 works productively for close to 20 hours and fully saturates tasks that were effectively out of reach for earlier models
https://x.com/ProximalHQ/status/2065184730279223410
Claude Mythos 5 scores 30.9% on FrontierCode Diamond Opus 4.8, the second best model is stuck at 13.4%
https://x.com/scaling01/status/2064391295620010383
I’ve been at Anthropic through every model launch. There’s been a few cases I can remember of a launch that stands out and marks a step-change in how we use models: – Claude Opus 3 – Claude Sonnet 3.5 – Claude Opus 4.5 And now Claude Fable 5. With Fable, the model stopped
https://x.com/alexalbert__/status/2064394410004304003
Opus 4.8 underperforms Opus 4.7 on the LLM Debate Benchmark (1717 → 1697), but Claude is still dominating the leaderboard. Qwen 3.7 Max scores worse than Qwen 3.6 Max: 1540 → 1499. Step 3.7 Flash lands at 1457. Ernie 5.1 improves a lot over Ernie 5.0: 1311 → 1447.
https://x.com/LechMazur/status/2062954327199666602
Sonnet 4.6 was below Sonnet 4? Really? What’s happening here? I get that their point is about the last stretch from Opus 4.5 to Mythos, but the previous trajectory looks like fumbling.
https://x.com/teortaxesTex/status/2062807380643958948
Subscription plans are massively subsidized. And by massively, I mean absurdly: Claude Max 20x: $200/month, with usage reportedly worth around $8,000 ChatGPT Pro 20x: $200/month, with usage reportedly worth around $14,000
https://x.com/kimmonismus/status/2064987311402537184
Wrote up my initial impressions of Claude Fable 5 – it has a big model smell: slow, expensive and capable of crunching through pretty much everything I threw at it
https://x.com/simonw/status/2064501565738930433
Introducing CADGenBench: measure how well AI systems produce engineering-grade 3D parts! While current models can generate 3D parts, they are far from precise enough to build functional parts. We built a benchmark to systematically measure their capabilities on two tasks: 1.
https://x.com/MikushRab/status/2063999885796614522
Fable 5 pricing confirmed at $10/$50 The model will be available in the Pro, Max, Team and seat-based Enterprise plans at no additional cost until June 22nd Afterwards it will require usage credits
https://x.com/scaling01/status/2064394893603049625
you could build a top tier venture firm just focusing investment decisions short and long term based on deep model benchmarking / evals find capability overhang, find areas models suck and track trajectory, etc
https://x.com/OfficialLoganK/status/2063312360102838278
The pessimistic framing of Moore’s law: every 18 months, the value of computation halves.
https://x.com/dwarkesh_sp/status/2063274857912221844
AI is moving beyond text, images, and code. Engineering artifacts are becoming a new class of model outputs and evaluating them requires different tools than we use for text, code, or images. Today we’re excited to release CADGenBench, a benchmark for CAD generation and
https://x.com/Thom_Wolf/status/2064029993638764672
In the Image Arena: open-weight Text-to-Image has a clear leader, with a tight race directly behind it: – #1 Ideogram-4.0 Quality has set the pace this week with a score of 1204. @ideogram_ai – #2 Hunyuan Image 3.0 by @TencentHunyuan with a score of 1151, just +1 pt ahead of
https://x.com/arena/status/2062997992777609534
MiniMax-M3 scores 55 on the Artificial Analysis Intelligence Index. Once the weights are released, it will be the leading open weights model M3 is @MiniMax_AI’s first multimodal M-series model, adding image and video input and a 1M token context window over the text-only
https://x.com/ArtificialAnlys/status/2064066303863005254
Three new models entered the Image Arena Top 10 this past month (Text-to-Image): – #2 Reve 2.0 by @Reve (1,273), behind only GPT Image 2. – #4 MAI-Image-2.5 by @MicrosoftAI (1,253). – #9 Ideogram 4.0 Quality by @Ideogram_ai enters at #9 (1,204). And the only open-weights model in
https://x.com/arena/status/2062957421757452516
I really don’t think OpenAI is going to let this slide. I’ve been saying it for a long time, the real inflection was when they reached 5.2. I have no clear insight on what they currently have internally, but if they haven’t made a Mythos/Fable yet, it was *a choice*.
https://x.com/teortaxesTex/status/2064473970892587105
Fable 5 is state-of-the-art on nearly all tested benchmarks, with exceptional performance in software engineering, knowledge work, scientific research, and vision. The longer and more complex the task, the larger Fable 5’s lead over our other models.
https://x.com/claudeai/status/2064394151441863006
⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data
https://www.latent.space/p/swe-bench-dead
Who is the better reviewer – AI or a human? 45 scientists spent 469 hours judging 2,960 individual review criticisms from humans and AI across Nature-family papers and found that: → AI reviewers surfaced 26% of issues humans missed → GPT-5.2 beat the top human reviewer BUT
https://x.com/TheTuringPost/status/2063192741258104944
Introducing FrontierCode: a coding eval that raises the bar for difficulty & quality. Each task took 40+ hrs of work by leading open-source maintainers. Models write sloppy code that works but isn’t maintainable. Our eval is first to measure: would you actually merge this code?
https://x.com/cognition/status/2064061031912288715
You know Fable is the real deal when it calmly picks up 4 threads (2 pure research + 1 code + 1 hank) that had been stuck for a month and just moves forward genuinely clever ideas and solutions Deepseek pushing cost, Mimo pushing speed and now Fable – new frontiers already
https://x.com/hrishioa/status/2064717079526383699
Do larger context windows make vector search obsolete? This benchmark compares two approaches: → Sending large amounts of context directly to an LLM → Using a 2-step retrieval pipeline powered by Qdrant to fetch only the most relevant information The results highlight a key
https://x.com/qdrant_engine/status/2065056457461321761
Production AI Engineering starts with Evals — with Ankur Goyal of Braintrust
https://www.latent.space/p/braintrust
// Agents’ Last Exam // Agents’ Last Exam is a living benchmark of over 1,000 economically valuable tasks, built with 250+ industry experts and mapped to the U.S. federal occupational taxonomy. The hardest tier sits at a 2.6% average full pass rate across mainstream harnesses
https://x.com/dair_ai/status/2062916866235068607
A good benchmark provides a north star to guide & measure AI development. A great benchmark is scalable and can be easily turned into a data-generation pipeline for training. I think great benchmarks can only be built by crawling real-world data sources, not by manual design.
https://x.com/OfirPress/status/2063990430350340575
I’ve added two optimizers to the public benchmark: (1) Shampoo (with its original 1/4 power). (2) Spectral descent, which is equivalent to both Muon(mu=0) and Shampoo(b1=b2=0). Result: Shampoo falls halfway between Muon & Adam; Spectral descent is ~2x slower. Thread below 1/6
https://x.com/kellerjordan0/status/2064062891607888058
wow, I’m blown away. fable 5 completely demolishes my benchmark
https://x.com/mchlhess/status/2064734182648221952
🤖 From this week’s issue: An Epoch AI data insight measuring the open-to-closed model capability gap using their Epoch Capabilities Index (ECI), finding open-weight models now lag frontier closed models by an average of 4 months.
https://x.com/dl_weekly/status/2064422551762153946
AI companies say their models are getting better at finding software vulnerabilities. Is that bearing out in public data? Introducing our Cyber Vulnerabilities explorer, which visualizes Common Vulnerabilities and Exposures (CVE) reported to the CVE Program since 2022.
https://x.com/EpochAIResearch/status/2063027791638237487
Everyone building programs with AI needs to now: 1. Be able to swap out models, in short order. 2. Be able to verify, regularly, that your output isn’t silently changing. (You should have been doing this for any production system, with evals, but now there’s no excuse.)
https://x.com/dbreunig/status/2064751540003643738
Fable 5 refused 200 out of 200 ProgramBench tasks lmao
https://x.com/scaling01/status/2065209370145702040
Fable is the new leader on CADGenBench! Still long way to go:
https://x.com/lvwerra/status/2064758389406589134
It’s finally out!!! @METR_Evals found that more than half of SWEBench results is unmergeable slop. FrontierCode represents over 1000+ hours of maintainer validated software engineering work most frontier models cannot yet solve, much less solve with high quality. Cog had IOI
https://x.com/swyx/status/2064081945567580323
Many SWE-bench-Passing PRs Would Not Be Merged into Main – METR
https://metr.org/notes/2026-03-10-many-swe-bench-passing-prs-would-not-be-merged-into-main/#introduction
spent all day on fable for a giant PR. ~10kloc, lots of testing and intervention. 250$. I… don’t think it’s worth it? happy with 4.8/5.5, and the quality of work is better when it’s smaler steps. Still rocking @cursor_ai, that’s software that I still love using on the
https://x.com/threepointone/status/2065131942279016700
the amount of alpha you can have right now creating good public AI benchmarks is wild, such a big opportunity
https://x.com/OfficialLoganK/status/2062738933327499605





Leave a Reply