Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A single Julio Le Parc-style ribbon of concentric ROYGBIV bands — violet outside stepping through blue, green, yellow, orange to red — spirals inward from the left and resolves into the word ‘ANTHROPIC’ lettered in the same nested rainbow line-work, centered on a clean off-white background with generous negative space, flat matte finish, crisp printed edges, no shadows or gradients.
Here Opus 4.8 built and play-tested a new RPG in Claude Code, including 3 PDF manuals and adventures, playtest notes, a website, and a playable solo adventure – then put it all on Netlify. No feedback from me at all.
https://x.com/emollick/status/2060045063275573723
In early May, the best superforecasters predicted that, by the end of the year, the longest METR 80% task horizons would reach 3-4 hours. In late May, Claude Mythos achieved that number.
https://x.com/emollick/status/2062235461364445204
I had Opus 4.8 in Claude Code write a sophisticated, if minor, academic paper from a archive of hundreds of de-identified research files from years ago I had to use GPT-5.5 Pro as a reviewer, it spotted one major error & some minor points. Opus corrected
https://x.com/emollick/status/2060098885561778341
AI research is a series of next-step decisions. We looked at sessions where a human researcher took a wrong turn, showed Claude the session up to that point, and asked it what to do next. Mythos Preview improved on humans 64% of the time–up from 22% in 2024.
https://x.com/AnthropicAI/status/2062568870872003021
Expanding Project Glasswing \ Anthropic
https://www.anthropic.com/news/expanding-project-glasswing
Our internal data shows Claude is accelerating AI development–a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention.
https://x.com/AnthropicAI/status/2062568862479208923
How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever lived. In addition to being neat, it is an interesting test of combining research, code, design and stats for an AI.
https://x.com/emollick/status/2060165879908749490
Introducing Claude Opus 4.8 \ Anthropic
https://www.anthropic.com/news/claude-opus-4-8
Anthropic confidentially submits draft S-1 to the SEC \ Anthropic
https://www.anthropic.com/news/confidential-draft-s1-sec
Anthropic has confidentially submitted a draft S-1 registration statement to the Securities and Exchange Commission. Pending completion of SEC review, this gives us the option to pursue an initial public offering. Read more:
https://x.com/AnthropicAI/status/2061478052257841495
Today, Anthropic engineers on average ship 8x as much code per quarter as they did compared to 2021-2025.
https://x.com/AnthropicAI/status/2062568864240836995
When AI builds itself \ Anthropic
https://www.anthropic.com/institute/recursive-self-improvement
OpenAI, DeepMind, Anthropic CEOs back mandatory DNA synthesis screening A coalition of AI leaders, synthesis-industry executives, biosecurity researchers, and former national-security officials published an open letter in June 2026 urging Congress to make screening and
https://x.com/kimmonismus/status/2062485389949145457
🟢 MiniMax M3 is now in OpenClaude OpenClaude (coding CLI that works with any LLM) just shipped first-class support for MiniMax M3, MiniMax’s next-gen coding/agentic model with a 1,048,576-token (1M) context window. @MiniMax_AI
https://x.com/gitlawb/status/2061581678871806083
Anthropic: “”We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development
https://x.com/scaling01/status/2062572962117562507
Each time we release a model, we run the same test: give it code that trains a small AI model, ask the new model to speed it up. It takes a skilled human 4-8 hours to reach 4x faster. In May 2024, Claude Opus 4 averaged a ~3x speedup. This April, Mythos Preview achieved ~52x.
https://x.com/AnthropicAI/status/2062568869240476050
Anthropic says 80% of its new production code is now authored by Claude — how your enterprise can keep up | VentureBeat
https://venturebeat.com/technology/anthropic-says-80-of-its-new-production-code-is-now-authored-by-claude-how-your-enterprise-can-keep-up
Big development – Anthropic is now advocating to build verification mechanisms to enable the option to pause AI development.
https://x.com/a_karvonen/status/2062572851916574730
Correction: Claude Opus 4’s ~3x average speedup dates to May 2025, not May 2024. This evaluation has only existed since September 2024, but we backtested it on earlier models: those from May 2024 showed no speedup whatsoever.
https://x.com/AnthropicAI/status/2062634151556292775
Had Claude Code build a snake game where the snake becomes aware it is in the game and then… stuff happens. Some impressive creative decisions by the AI (& also some very AI ones), I just gave a first prompt and some feedback on the game as it went.
https://x.com/emollick/status/2062039734453416361
Re Opus/sonnet: from what I understood they compare it to sonnet 4.6. only on SWE pro comparable to opus. If I’m correct the quote was „side by side with sonnet 4.6″
https://x.com/kimmonismus/status/2061918020843557110
We’ve updated /fork in Claude Code /fork now runs a background agent with your exact context (system prompt, tools, history, model) and prompt cache. The result gets returned to your session. /branch (the old /fork) still copies the transcript to a new session you drive.
https://x.com/ClaudeDevs/status/2061947411141169494
Claude really can roleplay an economist. I love this little comment Claude made after some robustness checks on the paper it wrote: “”On a 1-10 identification scale, I’d now put the paper at about 4.5 — better than the 3.5 I’d have given before these tests, but well short of
https://x.com/emollick/status/2060168513176658003
Running an AI-native engineering org | Claude
https://claude.com/blog/running-an-ai-native-engineering-org
In WeirdML we see opus models increasingly use submissions to just explore the data, without actually trying to solve the problem (no predictions for the test set). It seems like, with no or low thinking, at least for some tasks, the prior for Opus to just explore the data
https://x.com/htihle/status/2061412097720774679
None of this guarantees recursive self-improvement is on the horizon. It’s not yet clear that Claude is capable of research judgment–of choosing the right problems to work on. But if these trends continue, AI systems designing and building their own successors is plausible. This
https://x.com/AnthropicAI/status/2062568873321513443
OPUS PSYCHOSIS–Claudes Opus 4.6 and 4.7 make stuff up all the time, constantly. Using Opus too much gives you AI psychosis, it makes you believe in fringe scientific and medical theories. I think it’s a very serious credibility and reliability problem for non-coding Claude usage
https://x.com/distributionat/status/2061362406971060244
Anthropic Opus 4.8 is new SOTA on ARC-AGI-3 Score: 1.5%, ~$10K ARC-AGI-3 analysis notes: * Opus 4.8 read the environment an abstraction *above* Opus 4.7, as objects & systems, not pictures * Opus 4.8 succeeded on early levels, but still committed to a wrong sub-goal
https://x.com/arcprize/status/2061512025638121516
Claude Opus 4.8: The System Card | Don’t Worry About the Vase
https://thezvi.wordpress.com/2026/05/29/claude-opus-4-8-the-system-card/
I had early access to Opus 4.8. Was impressed by it. Here is Opus 4.8’s one shot of “”create a visually interesting shader that can run in twigl, make it like an infinite city of neo-gothic towers partially drowned in a stormy ocean with large waves”” (this is all done with math)
https://x.com/emollick/status/2060042738637148470
Opus 4.8 Part 2: Model Welfare | Don’t Worry About the Vase
https://thezvi.wordpress.com/2026/06/01/opus-4-8-part-2-model-welfare/
Opus 4.8 vs MiniMax M3 tested both on default settings with the same prompt > Opus one shotted everything in 7 minutes > M3 needed an extra prompt to fix the “”break block”” feature and took 20+ minutes both got super close, judge both and lemme know which one looks better?
https://x.com/notjazii/status/2061407087293313210
The issue affected how Opus 4.8 requests were handled, causing the model to trigger more parallel tool calls than intended. It was unrelated to dynamic workflows.
https://x.com/ClaudeDevs/status/2061501790131265803
Anthropic expands Mythos to 150 additional organizations
https://www.cnbc.com/2026/06/02/anthropic-mythos-ai-project-glasswing.html
This story was so implausible that the only way it even (kind of) made sense if it is some sort of internal accounting placeholder at a cloud provider using their own compute. And even then it seems unbelievable for a wide number of reasons.
https://x.com/emollick/status/2062257554609033528
My timeline seems to have people surprised that U Chicago is getting Claude, but tons of schools (including U Penn where I teach) have school-wide AI There are lots of things that need to be figured out about AI & scholarship but safe & equitable access is a necessary foundation
https://x.com/emollick/status/2061977762861092994
.@GoogleDeepMind’s Gemma 4 – 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes –model gemma4:12b-mlx Claude Code: ollama launch claude –model gemma4:12b-mlx and more 👇👇👇 (Note, this currently works via MLX)
https://x.com/ollama/status/2062250522598572345
Part of the work was rebuilding leaner and faster dependencies: –
https://t.co/QIuyWGhe3V – proxy layer –
https://t.co/c177W6KYqH – filesystem safety –
https://t.co/YW6iCQZX4V – Image engine in WASM –
https://t.co/45NHImVgb7 – Opus in WASM –
https://t.co/B3ELm7gF7x – PDF in WASM
https://x.com/steipete/status/2060133435423789092
Anthropic modified its RSP (v3.3) to raise the bio/chemical threshold: it’s no longer enough for a model to significantly help malicious actors; now it must functionally replace the rare expertise of top-tier world specialists. Anthropic calls this a mere revision; Zvi and Opus
https://x.com/CRSegerie/status/2062474945377218819
We just published internal data on how much of Claude’s development is already being done by Claude: – Over 80% of all code merged into our codebase is now written by Claude – It’s been months since many researchers at Anthropic hand-wrote code – The typical Anthropic engineer
https://x.com/alexalbert__/status/2062580571214389510
We’ve added a CLI for Claude Platform to make every API endpoint runnable from your terminal. Call the Messages API, stand up Claude Managed Agents, pipe results straight into your shell. The ant CLI is well understood by coding agents (Claude Code) using the claude-api skill.
https://x.com/ClaudeDevs/status/2061877343078244459
We’ve reset 5-hour and weekly rate limits for all users on Pro and Max plans. We fixed an issue that caused some Claude Code sessions to spawn excessive parallel subagents, burning through usage faster than expected.
https://x.com/ClaudeDevs/status/2061501787769893055
Microsoft leaked the training FLOPS for Claude Mythos based on their slide Claude Mythos used: 6.1*10^27 FLOPs (with 95% CI at 5.3*10^27 and 7.1*10^27, assuming 1 px measurement error)
https://x.com/scaling01/status/2061897540161728791
The 6.1e27 FLOP figure for Claude Mythos is not realistic. The Microsoft intern just threw some darts. Instead, let me throw some darts. This is my realistic estimate for Claude Mythos compute, total params, active params and training tokens! Mythos was likely a training run
https://x.com/scaling01/status/2061989029025853757
It does seem like meaningfully better AI releases are accelerating, especially from OpenAI & Anthropic. To illustrate, I caused this timeline to be created. It only lists new models that scored 3 points or higher over previous models in the Artificial Analysis index.
https://x.com/emollick/status/2060867599869649097
John (@jyangballin) talking about the wide behavioral differences between GPT and Claude on ProgramBench.
https://x.com/OfirPress/status/2061458258821251081
Lots of companies are in the “”encourage AI adoption”” phase, whether teaching them ChatGPT/Claude or (sigh) tokenmaxxing. That dodges the harder problems of firm leadership: What do you want people to use AI for? What work should be reserved for people? What else needs to change?
https://x.com/emollick/status/2061483942310477999





Leave a Reply