Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A symmetrical Byzantine gold-mosaic icon of ornate hammered-gold balance scales centered against a burnished gold-ground apse, one pan holding a tiny clockwork nightingale of gold and enamel, the other pan stacked with tesserae tablets etched with tally marks and spiral gyres, deep imperial purple and Tyrian crimson accents, warm candlelit glow catching silver highlights on the mechanism, tactile grout-textured tesserae surface, with the bold Trajan-capital title ‘BENCHMARKS’ in gold across the lower third.
Claude Opus 4.8 takes the lead on the Artificial Analysis Intelligence Index at 61.4, with Anthropic retaking the #1 spot on GDPval-AA and advancing in terminal use and scientific reasoning To reach the leading position on the Intelligence Index, @Anthropic made large
https://x.com/ArtificialAnlys/status/2060117582120976868
It “”feels like the first smart model in a long while”” due to this
https://x.com/zephyr_z9/status/2060077152729694586
The cost per accepted line of code varies by roughly 7x across model families.
https://x.com/cursor_ai/status/2060025070425395562
Announcing AA-WER Streaming, our new benchmark measuring streaming Speech to Text models on accuracy and latency for voice agent use cases. Pareto optimal models on this new benchmark include those from Cartesia, ElevenLabs, and Deepgram Streaming Speech to Text (STT) powers
https://x.com/ArtificialAnlys/status/2060021901234458958
Agent Judge: Solving Long-Context Evals for Production Agents — Judgment Labs
https://www.judgmentlabs.ai/blogs/agent-judge-solving-long-context-evaluations
Artificial Analysis and IBM Research are launching ITBench-AA, the first in a new series of benchmarks evaluating models on agentic enterprise IT tasks, starting with Site Reliability Engineering tasks where frontier models score below 50% ITBench-AA’s SRE tasks benchmark model
https://x.com/ArtificialAnlys/status/2059698327235805258
Interesting new SWE/agentic benchmark (DeepSWE) was released yesterday. 113 tasks across 91 repos in 5 languages. Here are interesting things I noticed: – The evaluation harness (mini-swe-agent) gives every model a single bash tool and the same SI. No vendor editing primitives.
https://x.com/_philschmid/status/2059564676569076021
Opus 4.8 is live. Benchmarks especially significant jump in Agentic coding, but more important: „Fast mode is available for Opus 4.8. It’s the same model at roughly 2.5x the speed, and we’ve made it three times cheaper than before.”
https://x.com/kimmonismus/status/2060044465385902436
Anthropic just launched Claude Opus 4.8, and it is the new leader on our GDPval-AA benchmark for agentic real-world work tasks Opus 4.8 scored 1890 on GDPval-AA at launch with its ‘max’ effort setting, +137 points from Opus 4.7 and +121 points ahead of the next-best model,
https://x.com/ArtificialAnlys/status/2060042848268083411
Anthropic says Opus 4.8 ranks 1st on FrontierSWE
https://x.com/scaling01/status/2060046440563388838
Anthropic to introduce AI Fluency scorecard in Claude
https://www.testingcatalog.com/anthropic-to-introduce-personal-ai-fluency-scorecard-in-claude/
We have, as far as I can tell, no good tests of the productivity impact of the autonomous coding tools that appeared starting in December 2025. Every paper out there is from prior to the Claude Code/Codex revolution. A huge gap in our knowledge about what is happening in coding.
https://x.com/emollick/status/2059118330472972331
Qwen3.7 Max (20250517) debuts at #4 in Code Arena: Frontend – the top-ranked Chinese lab on the board, surpassing GLM-5.1 and is now on par with Claude Opus 4.6 on agentic web development tasks. Huge congrats to @Alibaba_Qwen on this achievement!
https://x.com/arena/status/2059297720079393107
Announcing Surya OCR 2: – 650M params – 83.3% olmocr bench score (top under 3B) – 87% on internal 91-lang benchmark – 5 pages/s on RTX 5090 – Runs on CPU, GPU, MPS
https://x.com/VikParuchuri/status/2059675773712167423
Help us produce the most useful work on AI by taking our 5-minute survey:
https://t.co/W2tLu3e4WW (You can sign up at the end to join our compensated user research panel.)
https://x.com/EpochAIResearch/status/2059336781208924566
Introducing BenchBench – by Rohit Krishnan
https://www.strangeloopcanon.com/p/introducing-benchbench
MaxSim v2 is out! Now finally with backprop so you can do training with this one. Bench compared to naive torch: – 10.33× faster on H200 – 11.94× faster on A100 And a bunch of other improvements like better bounds check, testing and better docs + examples of the APIs.
https://x.com/ErikKaum/status/2059659837219156453
MiniMax just teased their Sparse Attention architecture for M3. The benchmarks show 9.7x prefilling speedup and 15.6x decoding speedup at 1M tokens vs M2. MiniMax deliberately went back to full attention for M2 because efficient attention wasn’t production-ready. Their pretrain
https://x.com/kimmonismus/status/2059302121489486335
that’s an interesting eval code summary dishonesty at an all time low
https://x.com/scaling01/status/2060042892903678414
The experiments conducted in this post illustrate how early we are as an industry on eval tooling. Some takeaways and related thoughts: 1. Naively applying automation (which many current frameworks do) is likely to fail. 2. It’s easy to get fooled that automation (esp
https://x.com/HamelHusain/status/2057875320011882923
This is the first code bench that actually aligns with how it feels to use these models coding.
https://x.com/theo/status/2059352130289651925
I don’t think anyone has a good intuitive sense about what this means, and that failure of imagination is a generally bad thing for planning, investment, and policy. I also don’t have an easy solution (funnily enough, AIs are cliche at imagining the AI future, so no help there)
https://x.com/emollick/status/2057517583914389847
Most models are only evaluated on a fraction of the benchmarks out there. ArtifactLinker, our new system, predicts which ones would set a new state-of-the-art on benchmarks hosted on @HuggingFace, then runs the evaluation to verify. 🧵
https://x.com/allen_ai/status/2057838486204326078
Sonic 3.5 is now the #1 text to speech model on the @ArtificialAnlys leaderboard! You no longer have to trade off quality and latency – Sonic 3.5 also has the fastest time to first audio at 82ms end to end. See full benchmark results 👇
https://x.com/cartesia/status/2057880195403800633
Announcing ESMFold2, our new state-of-the-art structure prediction model capable of predicting structure from single sequences or MSAs. ESMFold2 improves on benchmarks of protein-protein interaction and is particularly strong on predictions of antibody-antigen complexes.
https://x.com/proteinrosh/status/2059633089702240598
MiniMax teases upcoming M3 model with new sparse attention mechanism and 15.6X long-context response speed boost | VentureBeat
https://venturebeat.com/technology/minimax-teases-upcoming-m3-model-with-new-sparse-attention-mechanism-and-15-6x-response-speed-boost
Gemini 3.5 Flash has made huge progress from 3.1 Pro on GDPval, Flash is competing at the frontier, post training going strong 🙂
https://x.com/OfficialLoganK/status/2057682092583227881
Gemini 3.5 Flash is on the Pareto frontier of cost per intelligence on Vending Bench (a measure of a models ability to run a simulated store)!
https://x.com/OfficialLoganK/status/2058210519291777087
Gemini 3.5 Flash Looks Good For How Fast It Is | Don’t Worry About the Vase
https://thezvi.wordpress.com/2026/05/22/gemini-3-5-flash-looks-good-for-how-fast-it-is/
Gemini 3.5 Flash outperforms 3.1 Pro on many vision use cases (like the below Roboflow eval) while being ~6x faster on average 🤯 Gemini multimodal understanding for the win.
https://x.com/OfficialLoganK/status/2057888362011463988
Gemini 3.5 Flash ranks #1 on the APEX-Agents-AA benchmark, outperforming much larger models a whole size above it.
https://x.com/OfficialLoganK/status/2057460544643404125
some more recent ALE-Bench results: Grok-4.3 is pretty terrible, basically worse than all the frontier chinese models like Kimi-K2.6, DeepSeek-V4, GLM-5.1 and even Grok-4.2 lol Gemini 3.5 Flash only gets good with multiple iterations but gets mogged by Kimi-K2.6 and is also
https://x.com/scaling01/status/2057937081070944709
Gemini 3.5 flash release is underwhelming for browser agents. Slight improvement in performance over Gemini 3.1 pro, at a small increase in total cost
https://x.com/Alezander907/status/2057686331380359566
Its very limiting that a big set of very hard problems that we have just lying around are Erdos problems. Don’t get me wrong, they are quite cool, but we really need hard problems repositories for many fields, including areas that have less specified answers & require judges.
https://x.com/emollick/status/2059012803009151444
Modern LLMs can do multiplication of 100-digit numbers without tools. So much for “”embers of autoregression””. Just scale the COT bro
https://x.com/teortaxesTex/status/2057826903721951273





Leave a Reply