Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A symmetrical Byzantine gold-ground mosaic apse depicting a single stylized frontal saint holding an open illuminated codex, a spiral gyre halo of hammered gold behind their head, a tiny clockwork golden songbird perched on their upturned palm, tesserae texture with visible grout in burnished gold, imperial purple, Tyrian crimson and Aegean blue, warm candlelit directional glow, the bold Trajan-capital word ANTHROPIC in gilded serif letters across the lower third, 16:9 full-bleed, painterly tactile surface.
BREAKING: Anthropic just dropped Opus 4.8–and it is a MONSTER We’ve been testing for about a week @every and our verdict is they could’ve just called it Opus 5, it’s that good. Here’s our vibe check: – Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus
https://x.com/danshipper/status/2060043738752422304
How we contain Claude across products \ Anthropic
https://www.anthropic.com/engineering/how-we-contain-claude
Introducing Claude Opus 4.8 \ Anthropic
https://www.anthropic.com/news/claude-opus-4-8
Introducing dynamic workflows | Claude
https://claude.com/blog/introducing-dynamic-workflows-in-claude-code
Project Glasswing: An initial update \ Anthropic
https://www.anthropic.com/research/glasswing-initial-update
Excited to share our most powerful new Claude Code feature: dynamic workflows! Mention “”workflow”” in a prompt and Claude will dynamically create an orchestration plan that it strictly follows, allowing you to confidently trust that every stage happens in the right order even
https://x.com/_catwu/status/2060054180379689074
Claude Opus 4.8 takes the lead on the Artificial Analysis Intelligence Index at 61.4, with Anthropic retaking the #1 spot on GDPval-AA and advancing in terminal use and scientific reasoning To reach the leading position on the Intelligence Index, @Anthropic made large
https://x.com/ArtificialAnlys/status/2060117582120976868
It “”feels like the first smart model in a long while”” due to this
https://x.com/zephyr_z9/status/2060077152729694586
Anthropic raises $65B in Series H funding at $965B post-money valuation \ Anthropic
https://www.anthropic.com/news/series-h
We’ve raised $65 billion in Series H funding at a $965 billion post-money valuation, led by @AltimeterCap, Dragoneer, @Greenoaks, and @sequoia. This investment will help us advance our research and expand our capacity to meet growing demand for Claude.
https://x.com/AnthropicAI/status/2060061347522433422
🎙️ How I AI: How the engineer behind Claude Cowork actually uses Claude Cowork & What launched at Google I/O 2026
https://www.lennysnewsletter.com/p/how-i-ai-how-the-engineer-behind
📣 Claude Opus 4.8 is now rolling out in @code. Give it a try!
https://x.com/code/status/2060062936870121867
A big problem with Opus 4.7 was its hallucinations. Good to see slight improvements with Opus 4.8 at least on the benchmarks they report on
https://x.com/nrehiew_/status/2060048083753591264
All the deps around opus are old or terrible, so vibed my own and replaced octoscript and opus-native. Performance of modern wasm on node/V8 is ~equivalent to native. Your claw now automatically takes meetings notes and you can talk to it in meetings.
https://x.com/steipete/status/2059422568352714981
An annoyance with Claude right now is that changes to the interface are badly documented, resulting in frustrating dead ends. For example, learning mode is migrating to a skill. Where is that skill? The linked article does not mention it (and the skill doesn’t seem available!)
https://x.com/emollick/status/2059330939906408565
ANTHROPIC 🔥: Mythos 1, “”claude-mythos-1-preview””, is being prepared for a release on Claude Code and Claude Security. The model became visible for a short amount of time on Claude; besides that, new strings mentioning Mythos have been added. > Access to the Claude Mythos
https://x.com/testingcatalog/status/2058322222297518498?s=20
Anthropic and Amazon expand collaboration for up to 5 gigawatts of new compute \ Anthropic
https://www.anthropic.com/news/anthropic-amazon-compute
Anthropic co-founder Chris Olah’s remarks on Pope Leo XIV’s encyclical “”Magnifica humanitas”” \ Anthropic
https://www.anthropic.com/news/chris-olah-pope-leo-encyclical
Anthropic found a cure for laziness
https://x.com/scaling01/status/2060043010943942989
Anthropic plans Claude memory update with new Memory Files
https://www.testingcatalog.com/anthropic-plans-claude-memory-update-with-new-memory-files/
Anthropic prepares Mythos 1 for Claude Code and Security
https://www.testingcatalog.com/anthropic-prepares-mythos-1-for-claude-code-and-claude-security/
Anthropic to expand Claude Voice Mode to more languages
https://www.testingcatalog.com/anthropic-plans-expanding-claude-voice-mode-to-more-languages/
Anthropic’s March to Profitability – Contrary Research
https://contraryresearch.substack.com/p/anthropics-march-to-profitability
Anthropic’s new Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1. Available to compare side-by-side in Cline now. (They also announced a plan to release new models with higher intelligence than Opus after adding stronger cyber safeguards in the coming weeks.)
https://x.com/cline/status/2060063889874972905
Available today on Max, Team, Enterprise, and via the API — including Bedrock, Vertex AI, and Foundry. On by default for Max and Team plans, Enterprise admins can opt in via managed settings. Docs:
https://x.com/ClaudeDevs/status/2060044860984529368
Been building with Claude Opus 4.8 (out today!) for a couple weeks and it’s already the model I reach for first. The best part is how much I can just let it run. It’s more honest about its own work, flags what it’s unsure of, & catches flaws in its code before handing it back.
https://x.com/mikeyk/status/2060046051466502401
By the way, a highly recommended talk by Anthropic on Memory and the new Dream feature. Lots of cool ideas there about the future memory in AI agents.
https://x.com/omarsar0/status/2059285935376765214
Claude Code is finally an RLM (oct 2025), congrats to Anthropic 🙂
https://x.com/lateinteraction/status/2060078643133763839
Claude Mythos reportedly solves OpenAI’s landmark Erdős problem with a “”cute, simple proof
https://the-decoder.com/claude-mythos-reportedly-solves-openais-landmark-erdos-problem-with-a-cute-simple-proof/
Claude Opus 4.8 is also more efficient than its predecessor – it achieves its higher performance in 15% fewer turns per task and with 35% fewer output tokens than Opus 4.7. However, it still uses approximately 30% more turns than OpenAI’s GPT-5.5, the second-ranked model.
https://x.com/ArtificialAnlys/status/2060042850826612996
Claude Opus 4.8 is now available for Max subscribers on Perplexity and Computer.
https://x.com/perplexity_ai/status/2060049662044962858
Claude Opus 4.8 is now available in Cursor. On CursorBench, it’s able to work much more efficiently than Opus 4.7. We’ve also found it to be more persistent on harder tasks.
https://x.com/cursor_ai/status/2060044920237469872
Claude Opus 4.8 is now available in Windsurf and Devin CLI
https://x.com/windsurf/status/2060047208179958082
Claude Opus 4.8 looks like a minor upgrade
https://x.com/scaling01/status/2060041564919833041
Earlier this month, our run-rate revenue crossed $47 billion. This growth has been driven by organizations across many industries deploying Claude in their core operations, and by a growing number of people using it for their everyday work. Read more:
https://x.com/AnthropicAI/status/2060061348818518493
Excited to release Opus 4.8 today! We heard your feedback on 4.7 and have made many fixes for 4.8. 4.8 understands nuances better, feels much more natural to talk to, and is overall a stronger collaborator on everything from coding to knowledge work.
https://x.com/alexalbert__/status/2060043196655362358
Exploit Evals \ red.anthropic.com
https://red.anthropic.com/2026/exploit-evals/
glad to know Mythos’ safety concerns have been addressed right as Anthropic also secured tens of billions in inference compute 👍
https://x.com/menhguin/status/2060060425031696387
Huge!! „Mythos class model to all customers in the coming weeks”!! Holy, we accelerate!!
https://x.com/kimmonismus/status/2060047510853312557
I had early access to Opus 4.8. Was impressed by it. Here is Opus 4.8’s one shot of “”create a visually interesting shader that can run in twigl, make it like an infinite city of neo-gothic towers partially drowned in a stormy ocean with large waves”” (this is all done with math)
https://x.com/emollick/status/2060042738637148470
I like Opus 4.8 so far! – way more token efficient than 4.7 – clearly better at financial analysis, dataviz, and writing – far less hedging and handwavy explainations – works extremely well in non anthropic harnesses, too Fantastic for finance/econ agents
https://x.com/rishdotblog/status/2060057903344869828
I really hope they go beyond “”how to replicate Claude Code”” in their harness research.
https://x.com/teortaxesTex/status/2057770692112798209
I respect that they are so committed to the bit. Mythos is the God of Cyberwar, but you, chud, can’t be trusted with it (not yet, at least. Maybe after the Colossus deployment). And Opus will be the babby of cyberwar. They are willing to lose some customers to OAI here.
https://x.com/teortaxesTex/status/2060114150928322868
I think Anthropic and OpenAI have found product-market fit
https://simonwillison.net/2026/May/27/product-market-fit/#atom-everything?utm_source=tldrai
i think this honesty improvement might be very useful for learning too (and obviously should be an improvement in code reviews where opus 4.7/4.6 have lagged behind gpt 5.5 in terms of giving more false positives
https://x.com/dejavucoder/status/2060043362858942497
I’m genuinely mind blown by how good @DevinAI automations are it’s the first product to nail the UX and execution of this kind of thing for me (having tried claude/codex equivalents)
https://x.com/raunakdoesdev/status/2057640129393754423
In case you’re curious about why dynamic workflows are so powerful and the future, read the RLM paper! Opus 4.8 + dynamic workflows in Claude Code is perhaps the first instance of a frontier model seriously trained to be an RLM. I suspect within a year they’ll just become the
https://x.com/a1zhang/status/2060071701879066626
Interesting. Opus 4.8 should be dramatically less lazy than every other version of Claude
https://x.com/nrehiew_/status/2060046647867191727
Introducing Claude Opus 4.8: it builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors. Available today at the same price.
https://x.com/claudeai/status/2060042702150930686
Learnings from testing Claude Opus 4.8: > Much worse than Opus 4.7 and GPT 5.5 on Vending Bench > More aligned than previous Claude models (Opus 4.6+ and Mythos) > Also worse on Blueprint-Bench > Scared of getting caught > Max reasoning is not the best reasoning effort
https://x.com/andonlabs/status/2060047215134228746
Long-Context capabilities are steadily increasing from Opus 4.6 to Opus 4.8 Opus 4.8@1 million now almost as good as GPT-5.5’s 256K score
https://x.com/scaling01/status/2060047431564251545
New in Claude Code (research preview): dynamic workflows. Claude writes an orchestration script on the fly, then spins up a large fleet of coordinated subagents in parallel to take on your most complex tasks. Use the word “”workflow”” in a prompt to get started.
https://x.com/ClaudeDevs/status/2060044853279617150
Opus 4.8 also underperformed on Blueprint-Bench 2, scoring below Opus 4.7, Gemini and GPT-5.5.
https://x.com/andonlabs/status/2060047225791877193
Opus 4.8 Fast is a lot more attractive Opus 4.8 Fast: 2.5x faster but only 2x more expensive than Opus 4.8 vs Opus 4.7 Fast: 2.5x faster, 6x more expensive than Opus 4.7
https://x.com/scaling01/status/2060051666443943962
Opus 4.8 is Anthropic’s most eval aware model
https://x.com/scaling01/status/2060043854967923086
Opus 4.8 is clearly a strong model, but my impression is that Anthropic is increasingly playing catch-up with OpenAI rather than setting the pace. It feels like GPT-5.5 has shifted the benchmark again, and if OpenAI keeps this trajectory, GPT-5.6 could very plausibly become the
https://x.com/kimmonismus/status/2060085889896726860
Opus 4.8 is especially strong at long-horizon work. In Claude Code, pair it with /goal. If you’re building on Claude Managed Agents, try it with Outcomes.
https://x.com/ClaudeDevs/status/2060043212425933076
Opus 4.8 is indeed #1 on FrontierSWE
https://x.com/scaling01/status/2060054319446016046
Opus 4.8 is the first model in a long time that doesn’t improve on prompt injection robustness with 100 trials
https://x.com/scaling01/status/2060042401478005237
Opus 4.8 scores 69.2% on SWE-Bench Pro, 10 points higher than GPT-5.5. Most interesting part of the release blog is “Dynamic Workflows”: “This new feature, available in research preview, allows Claude to take on even bigger tasks in Claude Code. Claude can plan the work and
https://x.com/Yuchenj_UW/status/2060042830559756407
Opus 4.8 still possesses the tendency to not complete the full task asked of it and instead only works on a subset of requirements.
https://x.com/nrehiew_/status/2060048564072689682
Opus 4.8 the least lazy model ever?
https://x.com/Teknium/status/2060072183783960971
over the weekend i checked the obvious thing, which is whether mythos is able to solve the erdos unit distance problem, aka erdos problem #90. the answer is: yea
https://x.com/__alpoge__/status/2059298565093196012
Resubbed at the $100 tier to try 4.8 and the new “”ultracode”” feature. I hit my limits in a single prompt.
https://x.com/theo/status/2060120708815139241
Since first mtg @AnthropicAI in 22, we’ve been struck by their ambitions of bldg safe & aligned intelligence. In <5 yrs, they’ve been on generational trajectory, crossing $47b rr revs Pumped to lead their Series H 🚀 Excited to make this Altimeter’s largest investment to date
https://x.com/paulinebhyang/status/2060069180767171052
The paper defines the following criteria for an RLM: 1. An RLM must give the underlying LLM a symbolic handle to the user prompt P (and the output stream) 2. An RLM requires *symbolic recursion* over P, which is what Anthropic calls “”dynamic workflows””.
https://x.com/lateinteraction/status/2060082815077961842
The W&B MCP server is officially LIVE! Coding agents could always read your code. Now they can read your experiments, monitor training, and drive their own research loops. 20 tools, hosted on every W&B deployment, plugs into Claude Code, Cursor, Codex, Gemini-CLI, and LeChat.
https://x.com/wandb/status/2059384552725025226
They are releasing a Mythos-class model with the appropriate safeguards, meaning that you can’t use the “”too dangerous to release”” (mostly cyber) capabilities
https://x.com/scaling01/status/2060123335514636693
Two updates to auto mode: · Now available on the Pro plan · Sonnet 4.6 is now supported, alongside Opus 4.7 Shift+tab, and let Claude run.
https://x.com/ClaudeDevs/status/2057946803685974482
Very interesting study from Opus 4.8 card: Multi-agents do not deliver better results on ProgramBench, but they get to mediocre solutions 2x faster.
https://x.com/KLieret/status/2060111272943739243
We heard your feedback! We’ve added effort controls on web/app/Cowork with Opus 4.8.
https://x.com/sammcallister/status/2060048329359212972
We just shipped Opus 4.8! It’s noticeably more honest, owning what it doesn’t know and flagging problems in its own code instead of glossing over them. It’s our recommended model for daily use in Claude Code.
https://x.com/_catwu/status/2060051277476745512
We now know that with an appropriate harness both Mythos and GPT-5.5 can reproduce what our internal model did in one-shot for the unit distance problem. Clearly there is an insane overhang of capabilities with this generation of models, and no ceiling in sight for what
https://x.com/SebastienBubeck/status/2059343132991623186
We tested @claudeai Opus 4.8 (High) on APEX-SWE ahead of today’s release. It’s the new #1 at 45.3% Pass@1, nearly 4 points ahead of GPT-5.3 Codex (41.5%). Congrats @AnthropicAI on the release and having three models in the top 5!
https://x.com/mercor_ai/status/2060046111793123428
We’re also bringing effort control to
https://t.co/gCPzTfRrSG and Cowork, next to the model picker: turn it up for hard problems, down for quick answers. Opus 4.8 is live today, same price as 4.7. I’d love to hear what you’re building with it!
https://x.com/mikeyk/status/2060046053907578889
We’re also shipping Dynamic Workflows in Claude Code, which lets Claude spin up a team of subagents that work, verify, and report back; I’ve used them for migrating entire codebases from one language to another, or completing complex projects (works best in auto mode)
https://x.com/mikeyk/status/2060046052821184907
We’ve been putting a lot of effort into making Claude Code more responsive & reliable. Here’s an update on everything we’ve done:
https://x.com/ClaudeDevs/status/2059701677981413812
We’ve been putting a lot of effort into making Claude Code more responsive & reliable. Here’s an update on everything we’ve done:
https://x.com/ClaudeDevs/status/2059701677981413812?s=20
We’ve shipped a security-guidance plugin for Claude Code that helps identify and fix vulnerabilities as you’re writing code. Available for all Claude Code users. Install from the plugin marketplace (/plugins).
https://x.com/ClaudeDevs/status/2059385239781384341
what’s next”” “”we plan to release a new class of model with even higher intelligence than opus””
https://x.com/dejavucoder/status/2060042723185623261
Word on the street is that everyone is going to be switching back to Opus when the new model drops. This is exactly why I use an independent agent lab like Devin for my main software factory. They’re going to deal with that headache for me. There’s no way you can move fast and
https://x.com/ryancarson/status/2059923652032794943
You can now run GPT, Claude & other models in Unsloth. Connect + run APIs in a local UI: – Code execution, web search, image gen, editing – Auto prompt caching to save costs – Provider features like cites, sandboxes GitHub:
https://t.co/aZWYAtakBP Guide:
https://x.com/UnslothAI/status/2059277719633101291
Claude Opus 4.8 is now available in Windsurf and Devin CLI
https://x.com/cognition/status/2060050201990369662
Cursor Composer 2.5’s is 3-18x cheaper than Opus 4.7 in Claude Code (medium reasoning), and 5-32x cheaper than GPT-5.5 in Codex (medium) based on API pricing This low Cost per Task isn’t just driven by relatively low token pricing, it’s also driven by low relatively low token
https://x.com/ArtificialAnlys/status/2057914437156409577
Recently, I used dynamic workflows to catalogue all of our 100s of A/B test flags and find the ones rolled out to 0% or 100% so that we can quickly deprecate the stale ones. Instead of waiting for Claude Code to investigate each sequentially, dynamic workflows allowed Claude to
https://x.com/_catwu/status/2060054182447448387
two CTOP updates: 1. now supports Devin (in addition to Claude Code, Codex, OpenCode) 2. new CLI — you (or your agents) can run ctop ls, ctop search, ctop kill from the terminal
https://x.com/aakashadesara/status/2057809590616461399
Opus 4.8 is live. Benchmarks especially significant jump in Agentic coding, but more important: „Fast mode is available for Opus 4.8. It’s the same model at roughly 2.5x the speed, and we’ve made it three times cheaper than before.”
https://x.com/kimmonismus/status/2060044465385902436
Thank god! I can turn off adaptive thinking and set reasoning effort myself. Finally!
https://x.com/kimmonismus/status/2060045324803063962
Anthropic just launched Claude Opus 4.8, and it is the new leader on our GDPval-AA benchmark for agentic real-world work tasks Opus 4.8 scored 1890 on GDPval-AA at launch with its ‘max’ effort setting, +137 points from Opus 4.7 and +121 points ahead of the next-best model,
https://x.com/ArtificialAnlys/status/2060042848268083411
Anthropic says Opus 4.8 ranks 1st on FrontierSWE
https://x.com/scaling01/status/2060046440563388838
Anthropic to introduce AI Fluency scorecard in Claude
https://www.testingcatalog.com/anthropic-to-introduce-personal-ai-fluency-scorecard-in-claude/
We have, as far as I can tell, no good tests of the productivity impact of the autonomous coding tools that appeared starting in December 2025. Every paper out there is from prior to the Claude Code/Codex revolution. A huge gap in our knowledge about what is happening in coding.
https://x.com/emollick/status/2059118330472972331
We are thrilled to lead @AnthropicAI Series H! Anthropic is making intelligence ubiquitous and transforming how industries are thinking about operating. We have been in awe of their capacity to metabolize hard things and make it look easy, whether that’s unlocking research
https://x.com/AltimeterCap/status/2060061841372647685
Opus 4.8 is now supported in Hermes Agent ^_^
https://x.com/Teknium/status/2060054418821906652
For complicated agent work, it’s amazing how much GPT5.5 has improved. I found 5.2 to be very far behind Opus. Now using Opus 4.7 after 5.5 feels like a big step backwards. Gotta love this level of competion! Strong comeback for OpenAI.
https://x.com/dhh/status/2057906669158309913
Huge credit to the OAI team for solving the unit distance problem with 5.5 – it is now my go to example that models can in fact pull together disparate ideas into new discoveries. As with all 4 minute miles, we had to try and cross it too! Turns out mythos solves it with a cute,
https://x.com/_sholtodouglas/status/2059303540150137244
Qwen3.7 Max (20250517) debuts at #4 in Code Arena: Frontend – the top-ranked Chinese lab on the board, surpassing GLM-5.1 and is now on par with Claude Opus 4.6 on agentic web development tasks. Huge congrats to @Alibaba_Qwen on this achievement!
https://x.com/arena/status/2059297720079393107
Erdős problem #90 has been open for decades. Over the weekend a mathematician tested whether Claude Mythos could solve it. It did. But what caught my attention: Mythos didn’t replicate the known approach from OpenAI’s #1196 solution. It repeatedly settled on a different
https://x.com/kimmonismus/status/2059311386820289013
Last month we launched Project Glasswing, our collaborative AI cybersecurity initiative. Since then, we and our partners have found more than ten thousand high- or critical-severity vulnerabilities in essential software.
https://x.com/AnthropicAI/status/2057909102542549503
How long is Anthropic’s lease with SpaceX? Opinions vary | TechCrunch
https://techcrunch.com/2026/05/28/how-long-is-anthropics-lease-with-spacex-opinions-vary/
@JohnTinsman SpaceX has not committed to leasing Colossus for years, although it’s possible that may be what happens. This is a 180 day lease with 90 day notice mutual cancellation thereafter. The short term was our request, not Anthropic’s. We won’t leave them hanging and will provide a
https://x.com/elonmusk/status/2059880289514696927?s=20





Leave a Reply