Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: The word ‘BENCHMARKS’ centered on a clean off-white background, each letter built from concentric ROYGBIV rainbow bands in Julio Le Parc’s kinetic op-art style, letters set at varying heights like bars in a chart above a single thin horizontal baseline, flat matte screen-print finish with crisp edges and generous negative space, no shadows or gradients.

In early May, the best superforecasters predicted that, by the end of the year, the longest METR 80% task horizons would reach 3-4 hours. In late May, Claude Mythos achieved that number.
https://x.com/emollick/status/2062235461364445204

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever lived. In addition to being neat, it is an interesting test of combining research, code, design and stats for an AI.
https://x.com/emollick/status/2060165879908749490

Introducing Claude Opus 4.8 \ Anthropic
https://www.anthropic.com/news/claude-opus-4-8

Today, Anthropic engineers on average ship 8x as much code per quarter as they did compared to 2021-2025.
https://x.com/AnthropicAI/status/2062568864240836995

When AI builds itself \ Anthropic
https://www.anthropic.com/institute/recursive-self-improvement

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks. – It’s a
https://x.com/mustafasuleyman/status/2061880164498428188

Building a hill-climbing machine: Launching seven new MAI models | Microsoft AI
https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/

We took another look at the capability gap between open-weight and proprietary models. Since the start of the year, open-weight models have lagged the state of the art by four months.
https://x.com/EpochAIResearch/status/2060451576779886942

Introducing Agent Arena: real-world agentic evals at scale. How do you evaluate agents doing actual work? We measure millions of live sessions where real users accomplish real tasks. On Arena, models now get web search, filesystem, and terminal tools to complete complex
https://x.com/arena/status/2062566749418233981

Anthropic Opus 4.8 is new SOTA on ARC-AGI-3 Score: 1.5%, ~$10K ARC-AGI-3 analysis notes: * Opus 4.8 read the environment an abstraction *above* Opus 4.7, as objects & systems, not pictures * Opus 4.8 succeeded on early levels, but still committed to a wrong sub-goal
https://x.com/arcprize/status/2061512025638121516

Claude Opus 4.8: The System Card | Don’t Worry About the Vase
https://thezvi.wordpress.com/2026/05/29/claude-opus-4-8-the-system-card/

I had early access to Opus 4.8. Was impressed by it. Here is Opus 4.8’s one shot of “”create a visually interesting shader that can run in twigl, make it like an infinite city of neo-gothic towers partially drowned in a stormy ocean with large waves”” (this is all done with math)
https://x.com/emollick/status/2060042738637148470

Opus 4.8 Part 2: Model Welfare | Don’t Worry About the Vase
https://thezvi.wordpress.com/2026/06/01/opus-4-8-part-2-model-welfare/

Opus 4.8 vs MiniMax M3 tested both on default settings with the same prompt > Opus one shotted everything in 7 minutes > M3 needed an extra prompt to fix the “”break block”” feature and took 20+ minutes both got super close, judge both and lemme know which one looks better?
https://x.com/notjazii/status/2061407087293313210

The issue affected how Opus 4.8 requests were handled, causing the model to trigger more parallel tool calls than intended. It was unrelated to dynamic workflows.
https://x.com/ClaudeDevs/status/2061501790131265803

Help us produce the most useful work on AI by taking our 5-minute survey:
https://t.co/W2tLu3e4WW (You can sign up at the end to join our compensated user research panel.)
https://x.com/EpochAIResearch/status/2059336781208924566

I think Epoch does a great job benchmarking, but I continue to believe that open weights models are much more fragile, especially out-of-distribution, than their benchmarks indicate. Vibe-wise, I don’t think they were only 3 months behind last year or only 4 months behind today.
https://x.com/emollick/status/2060736941453189622

Intelligence Per Dollar | Tomasz Tunguz
https://tomtunguz.com/tokens-per-result/

State of AI Engineering | Datadog
https://www.datadoghq.com/resources/state-of-ai-engineering/

State tracking is a core pillar of video understanding: it requires identifying entities and events, and mapping how their states evolve over time. Frontier multimodal models are surprisingly bad at it, so we built a benchmark to measure it. Meet VSTAT!
https://x.com/PinzhiHuang/status/2062004108249145442

We’ve added narrations to our long-form content on the Epoch AI website, including reports, Gradient Updates, and topic overviews. Look for the play button.
https://x.com/EpochAIResearch/status/2060096279808745759

We just published internal data on how much of Claude’s development is already being done by Claude: – Over 80% of all code merged into our codebase is now written by Claude – It’s been months since many researchers at Anthropic hand-wrote code – The typical Anthropic engineer
https://x.com/alexalbert__/status/2062580571214389510

We’ve added a CLI for Claude Platform to make every API endpoint runnable from your terminal. Call the Messages API, stand up Claude Managed Agents, pipe results straight into your shell. The ant CLI is well understood by coding agents (Claude Code) using the claude-api skill.
https://x.com/ClaudeDevs/status/2061877343078244459

We’ve reset 5-hour and weekly rate limits for all users on Pro and Max plans. We fixed an issue that caused some Claude Code sessions to spawn excessive parallel subagents, burning through usage faster than expected.
https://x.com/ClaudeDevs/status/2061501787769893055

As token budgets take on a larger part of operating expenses over time, model routing is the inevitable conclusion. This is also one of the biggest areas of differentiation for the applied AI layer over time. By understanding the different work patterns in your domain, and
https://x.com/levie/status/2061974298760495132

Excited to see the use of GEPA-optimized LLM judges for data filtering in MAI-Thinking-1 model’s pre-training pipeline!
https://x.com/LakshyAAAgrawal/status/2062013650639241403

Give MAI-Code-1-Flash a try and let us know what you think!
https://x.com/pierceboggan/status/2062220583786709163

MAI-Thinking-1 is out! Excited to share what we are building and how climbing from scratch (no distillation) actually works: simple recipes, rigorous science, self-distillation, patience, and great infra. Check out our tech report has the full story of our RL climbs.
https://x.com/HannaHajishirzi/status/2061901432627044430

Super detailed tech report for MAI-Thinking-1, with a ton of info on all stages of the pipeline. I’m surprised so much of this info is released 🙂 Super long thread on my notes:
https://x.com/nrehiew_/status/2062013300196700395

this was an insanely good read, i think this is the most detailed report i’ve read at this scale in some aspects. i really hope MAI continues releasing those tech reports, thanks a lot to the team for this gift 🥹
https://x.com/eliebakouch/status/2062004670017486912

Today, Baseten and @MicrosoftAI are excited to announce that MAI-Thinking-1 is coming to Baseten. MAI-Thinking-1 is a model you can fine-tune without giving your data to the lab. Key characteristics include: → Clean data lineage, with zero distillation from third-party models
https://x.com/baseten/status/2061878701823066431

We’re excited to work with @Baseten to make MAI-Thinking-1 available to developers and enterprises.
https://x.com/MicrosoftAI/status/2061923309344756043

WOW microsoft new “”MAI Thinking 1″” model comes with a 109 page tech report that looks REALLY detailed, this is amazing
https://x.com/eliebakouch/status/2061877335960281459

Microsoft has released MAI-Transcribe-1.5: an exceptionally fast speech transcription model at a speed factor of ~276x, while still achieving 2.4% on AA-WER (#3), leading the accuracy-speed Pareto frontier MAI-Transcribe-1.5 is Microsoft AI (MAI)’s latest speech transcription
https://x.com/ArtificialAnlys/status/2061878491860324402

Microsoft introduces MAI-Thinking-1 It’s a 1T@35B parameter model pre-trained on 30T tokens with a maximum context length of 256k tokens using 8192 GB200 GPUs. Based on benchmarks it seems to be around GLM-5 level. Microsoft also released a comprehensive 109 pages tech-report:
https://x.com/scaling01/status/2061889624847343825

microsoft used gepa / dspy to tune the LLM judge prompt for quality scoring. @lateinteraction stays winning. from the mai-thinking-1 report
https://x.com/bj2rn/status/2061941109828301241

Today we’re announcing MAI-Thinking-1 with Microsoft and it will be available on Baseten soon. Microsoft built something genuinely different here: a commercial-grade thinking model trained on clean data with no distillation from third-party models and designed to be fine-tuned
https://x.com/tuhinone/status/2061879239817969756

big congrats to the microsoft AI team on MAI-Thinking-1! this is the kind of thoughtful post-training the field needs more of – focused on what actually matters to users excited to see a new frontier model in the race 😎
https://x.com/echen/status/2061907282607100075

Introducing MAI-Code-1-Flash A new coding model from Microsoft for fast, efficient assistance in everyday workflows Rolling out to @code developers in model picker and Auto now!
https://x.com/pierceboggan/status/2061877165810131297

It is difficult to know how good MAI-Thinking-1 is from the scores alone (like weirdly low GPQA & Terminal Bench 2.0) But Microsoft makes it really hard to try its models upon release (a general issue with many Microsoft AI products), so I dunno. Stats below Meta Spark, though.
https://x.com/emollick/status/2061907785768489127

MAI-Image-2.5 has officially released from @MicrosoftAI landing at #2 in the Image Edit Arena (Single-Image-Edit) with a score of 1401 and advances the Pareto frontier! This puts the model +10 pts over Nano Banana 2, Grok Imagine Image Quality and ChatGPT-Image-Latest-High
https://x.com/arena/status/2061887242579382660

MAI-Image-2.5 ranks #2 in the Image Edit Arena and advances the Pareto frontier. That means: at its price tier, no model scores higher on Arena. Congrats again to @MicrosoftAI on this release!
https://x.com/arena/status/2061894541888962712

Microsoft AI has announced their very own reasoning model, MAI-Thinking-1, they have a detailed tech report too! I really appreciate they’ve reported health evals: HealthBench Professional and MedXpertQA. These are both very solid benchmark tasks that I recommend people use.
https://x.com/iScienceLuvr/status/2061926066453962952

MiniMax just dropped M3! It hits 59% on SWE-Bench Pro, edging out GPT-5.5 (58.6%) and beating Gemini 3.1 Pro (54.2%). Trails Opus 4.7 on coding, but leads it on autonomous browsing at 83.5% on BrowseComp. First open model to pack frontier coding, a 1M-token context, and native
https://x.com/kimmonismus/status/2061473350766170420

MiniMax M3 is now the leading open model on the Next.js agent evaluations (https://t.co/SnZ54XoRWV). Right behind Opus & GPT5, but 10× cheaper (And 20× cheaper right now on ▲ AI Gateway!)
https://x.com/rauchg/status/2061593874498531707

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading