Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: Using the provided reference images, keep the authentic Sonoran Desert trail setting, midday Arizona lighting, and the brown wooden post with weathered ranger-style sign, but replace the header with bold all-caps ‘BENCHMARKS’ and list trail entries like ‘MMLU Ridge → 4.2 mi’, ‘HumanEval Loop → 1.8 mi’, ‘ARC Saddle → 6.5 mi’ with arrows and mileages, and add a worn stopwatch resting on a volcanic rock beside the post and a surveyor’s measuring wheel leaning against the wood, with saguaro and the valley vista behind. Maintain photorealistic high desert documentary style, natural shadows, and the worn sun-baked sign texture of the reference.

Today we’re releasing Refactoring, the final leaderboard of our SWE Atlas suite. This new leaderboard is the ultimate test of an agent’s ability to restructure code without breaking the system. Claude Opus 4.7 with Claude Code takes the top spot🥇
https://x.com/ScaleAILabs/status/2052434456510878021

Gemma-4 lands in Code Arena: Frontend Webdev and shifts the Pareto Frontier! Among open models, Gemma-4-31b ranks #13 and Gemma-4-26b-a4b ranks #17. Congrats to @GoogleDeepMind on shifting the frontier!
https://x.com/arena/status/2052061349312921686

Congrats to @OpenAI for taking the top spot on our Audio MultiChallenge S2S leaderboard with the release of GPT‑Realtime‑2 🥇 GPT-Realtime-2 more than doubles GPT-Realtime-1.5 on instruction retention, rising from 36.7% to 70.8% APR, and also stands out on voice editing,
https://x.com/ScaleAILabs/status/2052451341071683732

All benchmarks are flawed, but GPQA has been fairly consistent & highly correlated with other measured benchmars. I think it’s a good way to see how far we’ve come that the free model from OpenAI, GPT 5.5 Instant, is at a level that even paid models did not reach until late 2025
https://x.com/emollick/status/2051801703209742734

Medicine | The 2026 AI Index Report | Stanford HAI
https://hai.stanford.edu/ai-index/2026-ai-index-report/medicine

New paper (on an old AI) tests o1 against doctors on medical benchmarks & real ER cases: “across a variety of scenarios and applications, the large language model outperformed both human physicians and older models” The potential suggests an “urgent need for prospective trials.”
https://x.com/emollick/status/2050197369250033813

Very excited to release Terminal-Bench 2.1! Coding agents are among the most economically consequential deployments of LLMs to date. As agents improve, benchmark reliability matters more. We audited TB2.0 and found and corrected issues in 28/89 tasks. 30% of the benchmark!
https://x.com/ekellbuch/status/2052165464655298866

We recently built HiL-Bench, the first benchmark to test a critical question: do AI agents know what they’re missing and when to ask? Frontier models perform well with perfect specs. But remove a few key details, and they confidently guess and ship plausible wrong answers. We
https://x.com/ScaleAILabs/status/2051333688798097567

PostTrainBench results for GPT-5.5 are in it doesn’t beat Opus 4.7 in the Claude Code harness even with almost 2 more hours of working time via reprompting
https://x.com/scaling01/status/2050289320699818417

A challenge with AI regulation and vetting is how bad our benchmarks of AI model performance and risks are. There is no benchmark for risks and red-teaming requires experiments from dedicated specialist organizations & is not easy to put metrics around. No clear objective numbers
https://x.com/emollick/status/2051431009766289734

A new very hard search benchmark that exposes bottlenecks in modern neural retrievers!
https://x.com/nlp_mit/status/2052069072607547892

Are AI benchmarks doomed? @GregHBurnham and @tmkadamcz join @ansonwhho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like. (0:00:00) – Preview (0:00:36) – Intro: Are AI benchmarks doomed? (0:03:13) – The costs and benefits
https://x.com/EpochAIResearch/status/2051330509989368211

Benchmarks aren’t just about showing where we are now, benchmarks are a treasure map that shows us how to get to the future the benchmark specifies. SWE-bench was about automating bugfixing and small feature reqs, this is about automating entire repo development.
https://x.com/OfirPress/status/2052106927908200957

How much of SQLite, FFmpeg, PHP compiler can LMs code from scratch? Given just an executable and no starter code or internet access. Introducing ProgramBench: 200 rigorous, whole-repo generation tasks where models design, build, and ship a working program end to end. 🧵
https://x.com/jyangballin/status/2051677497562210552

I’ve never been this excited about search. 6-7 years ago, IR got an influx of the paradigms we still use, all enabled by the big headroom MS MARCO and then BEIR created. Then progress slowed. Today, Diane releases perhaps the most ambitious IR benchmark to date: OBLIQ-Bench.
https://x.com/lateinteraction/status/2052055143038713875

Looking at average pass rate is *very* misleading- every task has a big chunk of tests that are very easy to pass and sometimes a minority of tests that are much harder to pass- so you can implement 10% of the program and get a 60% pass rate.
https://x.com/OfirPress/status/2051757679283143089

New research from @AISecurityInst and Goodfire: Models sometimes recognize they’re being evaluated, occasionally even identifying the benchmark. We show this verbalized eval awareness inflates safety scores, meaning safety benchmarks may not reflect real-world behavior. (1/7)
https://x.com/GoodfireAI/status/2051382876483231968

ProgramBench uses a not so useful / weird metric like ARC-AGI > headline score of all models -> 0% > looks inside > Opus 4.6 and 4.7 pass on average >50% of tests per task > why? > they only count a task as passed if 100% of tests are successful and as we all know software
https://x.com/scaling01/status/2051733949877985349

very impressive release with lots of care at every stage of training: custom arch with bigger experts, more expressive router, compressed attention, residual scaling, and much more on the post training side including test time compute etc.. benchmark scores are very competitive
https://x.com/eliebakouch/status/2052126118891729148

NEW paper from Sakana AI (ICLR 2026). A 7B Conductor model just hit SOTA on GPQA-Diamond and LiveCodeBench by orchestrating other LLMs instead of solving problems itself. (great paper! bookmark it!) The Conductor is trained with RL to do two things at once: design
https://x.com/omarsar0/status/2051306659021242635

Introducing a new paper! Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs Static benchmarks are no longer enough. Models improve too quickly and numbers become stale quickly. Instead, we argue for continuously maintained evaluation platforms.
https://x.com/j_dekoninck/status/2051268263150276872

ProgramBench
https://programbench.com/

ProgramBench: Can Language Models Rebuild Programs From Scratch? John Yang, Kilian Lieret, Jeffrey Ma, Parth Thakkar, Dmitrii Pedchenko, Sten Sootla, Emily McMilin, Pengcheng Yin, Rui Hou, Gabriel Synnaeve, Diyi Yang, Ofir Press
https://t.co/MRk3XwhsYA [𝚌𝚜.𝚂𝙴 𝚌𝚜.𝙰𝙸]
https://x.com/ComputerPapers/status/2051895799043215415

The artificial analysis index is a normalized score of several benchmarks (and has changed over time) it is fine for roughly comparing models, it is not useful for trend analysis and it is unclear what individual point differences in the scores mean.
https://x.com/emollick/status/2051061792667754507

The recipe for “classic” reasoning benchmarks is simple: text-only, several-hour time horizons, easy to grade, with expert human baselines. What next? In this week’s Gradient Update, @GregHBurnham argues it’s as easy as dropping one of these four ingredients.
https://x.com/EpochAIResearch/status/2051760424891392204

We are launching domain-specific capability scores, tracking the capabilities of models across SWE and Math benchmarks, using the same scale as the general ECI. We also support customization for users who want to create their own variants of the ECI. Link below!
https://x.com/EpochAIResearch/status/2052069897530933438

We set out to build a better retriever, so we looked for the hardest IR benchmarks. For each, we asked how much headroom remained by running oracle reranking with a frontier LLM. Most had little room left! So we built OBLIQ-Bench to study much harder search queries than before.
https://x.com/dianetc_/status/2052053806121140254

We’re releasing Terminal-Bench 2.1 to patch 28 of the 89 tasks in Terminal-Bench 2.0 TB2.1 includes • recalibrated limits • fixed solutions • realigned verifiers Per-task breakdowns in 🧵 We’ll continue to support TB2 and TB2.1 leaderboards (new submission process 🔜)
https://x.com/terminalbench/status/2052119174500220964

GPT-5.5 & Opus 4.7 on ARC-AGI-3 – GPT-5.5: 0.43% – Opus 4.7: 0.18% We found 3 failure modes: – True local effect, false world model – Wrong level of abstraction from training data – Solved the level, didn’t reinforce the reward See our full analysis 🧵
https://x.com/arcprize/status/2050261221165989969

💫Very happy to release NeuralBench, to benchmark Neuro AI models and datasets in the open! 🧵Thread, 💻Code, 📝White Paper below:
https://x.com/JeanRemiKing/status/2052034314120896582

Do people have some good evals that require context compaction to be achieved?
https://x.com/_philschmid/status/2051002064826724724

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading