The new version completely smashes GPT-5.5 and the previous Mythos version. Before Mythos Preview completed the cyber range 3 out of 10 times. The new version completed it 6 out of 10 times and is much more efficient!
https://x.com/scaling01/status/2054594892903436553
Google DeepMind is pushing medical AI into “”co-clinician”” research They shared an AI co-clinician research initiative that tests evidence-grounded clinical reasoning and real-time multimodal telemedicine simulations. The careful wording matters: supportive tool under physician
https://x.com/TheTuringPost/status/2052188488553079156
Google’s AI Drug Startup Isomorphic Labs Nears $2 Billion Capital Raise – Bloomberg
https://www.bloomberg.com/news/articles/2026-05-08/google-s-isomorphic-labs-to-raise-over-2-billion-in-new-funding
Meet physics-intern🧑🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA on one of the hardest benchmarks for LLMs. Theoretical physics is hard for humans and LLMs alike. But physics-intern decomposes problems and
https://x.com/dlouapre/status/2054217281895309480
NEW paper from Google DeepMind. (bookmark it) AI Co-Mathematician is an agentic research workbench for mathematicians, and it just hit 48% on FrontierMath Tier 4, a new high score among AI systems evaluated. The system is an asynchronous, stateful environment that supports
https://x.com/dair_ai/status/2054224343551639958
Kimi K2.6 is now open-weight #1 on Finance Agent Benchmark V2.
https://x.com/Kimi_Moonshot/status/2054803169994272819
I promise this will be the best 20 min you spend today! Robotics: Endgame, the sequel to my last year’s Sequoia AI Ascent talk, “”Physical Turing Test””. I laid out the roadmap for solving Physical AGI as a simple parallel to the LLM success story. Be a good scientist, copy
https://x.com/DrJimFan/status/2052758642781487237
Agent observability is a means to an end: making your agent better. But observability and evals tools have traditionally failed to connect traces to meaningful actions. Agent engineering teams are left combing through traces, guessing at root causes, and writing evals manually.
https://x.com/bentannyhill/status/2054949581679653326
Your customer support needs a voice agent built for the real world. Grok Voice Think Fast 1.0 handles complex workflows with speed and accuracy, even in hard-to-hear environments. From multi-step troubleshooting to high-volume tool calls, it keeps up.
https://x.com/xai/status/2052529102280880234
Working with agents for the past months has me convinced that outcome-only evaluation is a flawed approach to benchmarking. You need to look at the logs to understand if the agent really did its job! In our paper Log analysis is necessary for credible evaluation of AI agents, we
https://x.com/steverab/status/2054564579573698921
✅ Harness profiles: Per-model tuning + support for open models (@Kimi_Moonshot, @Alibaba_Qwen + @deepseek_ai) ✅ Code interpreter: A programmable runtime inside the agent loop ✅ Streaming-typed projections for messages, tool calls, + subagent events ✅ DeltaChannel:
https://x.com/LangChain_OSS/status/2054641656222388700
🚀 Excited to share our new preprint: Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs. To study research-level mathematical reasoning, we introduce Soohak, a benchmark of 439 research-level math problems created from scratch by
https://x.com/gson_AI/status/2054036114483392997
AI Gateway production index – Vercel
https://vercel.com/blog/ai-gateway-production-index
Announcing the Artificial Analysis Coding Agent Index! Our new coding agent benchmarks measure how combinations of agent harnesses and models perform on 3 leading benchmarks, token usage, cost and more When developers use AI to code they’re choosing a model, but also pairing it
https://x.com/ArtificialAnlys/status/2053865095076438427
Great read from the @RedHat_AI team — a comprehensive investigation into TurboQuant in vLLM, with FP8 and BF16 as reference baselines: 4 models (30B to 200B+, decoder-only and MoE) and 5 benchmarks covering long-context retrieval and reasoning, all on the stable vLLM 0.20.2
https://x.com/vllm_project/status/2053852636093239555
I love seeing a new eval with such low scores. When we announced GPT-5.5, almost every benchmark had a score above 50%. It’s time to retire evals like GQPA and bring in a new set.
https://x.com/polynoamial/status/2054255862441812099
Log analysis is not a “one and done” technique, it requires constant effort in validating benchmark results. One reason it’s hard to uncover evaluation bugs is that they become apparent only after models get good enough to solve tasks (or circumvent constraints in evaluations,
https://x.com/sayashk/status/2054569643080077576
The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls from 730.1µs to 438.5µs. For decode, GB200 sustains much higher throughput at high token speeds.
https://x.com/perplexity_ai/status/2054204425833726353
The first ProgramBench task was just solved by GPT 5.5 high/xhigh. Interestingly, high/xhigh picked two different languages for the task (C vs Python). GPT 5.5 xhigh was significantly better than Opus 4.7 xhigh in all metrics. 🧵
https://x.com/KLieret/status/2054215545663144217
The most extensive independent benchmark of LLMs for software engineering just got a big update! – How does GPT-5.5 compare to Opus 4.7? – Are open models catching up, and in what areas? – How do cost and performance stack up?
https://x.com/OpenHandsDev/status/2053839810343620980
We’re excited to release Medmarks v1.0 + a technical report! This is an update to our Medmarks benchmark suite, the largest open-source automated suite for evaluating the medical capabilities of LLMs. We added 10 benchmarks (20→30) and 15 models (46→61) to the leaderboard!
https://x.com/SophontAI/status/2054270239387627927
We Tested DeepSeek V4 Pro and Flash Against Claude Opus 4.7 and Kimi K2.6
https://blog.kilo.ai/p/we-tested-deepseek-v4-pro-and-flash
As much as the state of benchmarks in AI is flawed, it is so much easier to track AI progress than robotics. Not sure what you can make of all the videos of robots running races or doing laundry – are there any equivalents to independent AI benchmarks for robots? ARC-AGI-BOT?
https://x.com/emollick/status/2053104629282378061





Leave a Reply