Image created with Ideogram 3.0. Image prompt: Lower-East-Side street-corner photograph reminiscent of a late-80s album cover: weathered red-brick tenement with exterior fire-escapes, canvas awning shading racks of vintage clothes; above the awning, a hand-painted board reads ‘Benchmarks SPORTSWEAR’; a hanging blade sign in cursive script reads ‘Benchmarks Boutique’; a bulletin board displays scorecards titled ‘Benchmarks 1989’ fluttering in the breeze; warm golden-hour light, subtle 35mm film grain, muted yet punchy color palette, gritty NYC vibe.
2.5 Pro is now the best model for coding and learning. With a strong ELO score of 1415, it’s topping the WebDev Arena leaderboard – and it incorporates LearnLM, our family of models fine-tuned for learning built with educational experts. #GoogleIO https://x.com/GoogleDeepMind/status/1924878252172353851
Google announces AI Ultra subscription plan https://blog.google/products/google-one/google-ai-ultra/
StackOverflow questions over time, source SEDE; sadface, lunch has been eaten https://x.com/marcgravell/status/1922922817143660783
Report: Spring 2025 AI Model Usage Trends – Poe https://poe.com/blog/spring-2025-ai-model-usage-trends
Here’s my “”Dark Leisure”” theory of any potential productivity paradox in AI: – most AI use rn is bottom up and hidden (employee first, not company first): employees vibe code, vibe market, vibe write and get stuff done faster – in many orgs, there is too little incentive to”” / X https://x.com/fabianstelzer/status/1926000937702764635
I wish these skeptical AI articles would actually grapple with the growing body of research that AI can really do original research & perform key unstructured tasks across the spectrum of high-end white collar employment. AI criticism is important, but it should be clear-eyed.”” / X https://x.com/emollick/status/1923417536072241529
Anthropic closes $2.5 billion credit facility https://www.cnbc.com/amp/2025/05/16/anthropic-ai-credit-facility.html
📢We’re excited to share that we’ve raised $100M in seed funding to support LMArena and continue our research on reliable AI. Led by @a16z and UC Investments (@UofCalifornia), we’re proud to have the support of those that believe in both the science and the mission. We’re https://x.com/lmarena_ai/status/1925241333310189804
Google dropped AlphaEvolve, an AI that discovers algorithms for scientific and computational challenges —Uses Gemini models with auto-evaluation & iteration —Found the first improvement on 1969’s Strassen’s algorithm —Also boosting efficiency for Google https://x.com/adcock_brett/status/1924133683444793819
Very strong agentic performance by the new Claude 4 Opus and Sonnet, placing 1st and 3rd on the GAIA benchmark. Notice, this leaderboard unfortunately doesn’t include Google’s latest Gemini models nor OpenAI’s o4-mini or o3. https://x.com/scaling01/status/1926017165108375607
Agents taking actions in your ad account🐳 Coming next week to Moby Agents… https://x.com/AY_Orbach/status/1923425142039822517
II-Agent – Intelligent Internet https://ii.inc/web/blog/post/ii-agent
We released Devstral. It is a 24B model released under the Apache 2.0 license. It the best open model on SWE-Bench verified today. You can check our blog post or test it with OpenHands (from @allhands_ai ) following the instructions here: https://x.com/b_roziere/status/1925194095359676768
Deep Think’ boosts the performance of Google’s flagship Google Gemini AI model | TechCrunch https://techcrunch.com/2025/05/20/deep-think-boosts-the-performance-of-googles-flagship-google-gemini-ai-model/
6. Gemini AI updates: —Gemini 2.5 Pro Deep Think, a new mode that uses parallel thinking to solve math and coding problems —Gemini 2.5 Flash, upgraded with improved performance across benchmarks Both also include native audio outputs across languages https://x.com/rowancheung/status/1925084387894300762
ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems https://arxiv.org/pdf/2505.11831
Generative AI Adoption Index – US Press Center https://press.aboutamazon.com/aws/2025/5/generative-ai-adoption-index
Anthropic’s Claude 4 models are here. Opus 4 and Sonnet 4 both show strong (& improved) coding abilities, with Sonnet 4 at 72.7% on SWE-bench and Opus 4 at 72.5% What does this mean for developers using Cline? 🧵 https://x.com/cline/status/1925680741108613503
opus 4 review: its a good model i was an early tester and found that it combines much of what people loved about sonnet 3.6 and 3.7 (and some opus!) into something which is much greater than the parts amazing at long-term tasks, intelligent tool usage, and helping you write!”” / X https://x.com/nearcyan/status/1925698351661502810
Anthropic was indeed underpriced at 2%. It peaked at 31.6% 3 minutes after the release of Claude 4 Opus However, the dust has settled and betting markets still think that Google has the best model. Do you agree? https://x.com/scaling01/status/1925623699186589698
Claude-4 Opus is definitely not a frontier model for mathematics (screenshot from MathArena) Results for Claude-4 Sonnet haven’t been published. https://x.com/scaling01/status/1926018522372514037
LM Arena, the organization behind popular AI leaderboards, lands $100M | TechCrunch https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/
The three gulfs of LLM pipeline development is a powerful mental model to keep in mind while creating AI evals. @eugeneyan , @sh_reya and I discuss here (links to more resources in the replies) https://x.com/HamelHusain/status/1923382256867180606
The Top 5 domestic large models contend for supremacy, a decisive battle in AGI – Google Docs https://docs.google.com/document/d/1qzm3wVxr3m-vFXqSQQjI8MaP9wADddbOpEQkH2kZLfo/edit?tab=t.0#heading=h.8xti639xclnh
US lawmakers have concerns about Apple-Alibaba deal | TechCrunch https://techcrunch.com/2025/05/18/u-s-lawmakers-have-concerns-about-apple-alibaba-deal/
Chatbot Arena Group Goes From Academic Project to $600 Million Startup – Bloomberg https://www.bloomberg.com/news/articles/2025-05-21/lmarena-goes-from-academic-project-to-600-million-startup?embedded-checkout=true
Really cool how DeepSeek is now the benchmark for Nvidia”” / X https://x.com/teortaxesTex/status/1924588309688267139
This is the second time that this has happened. I really wish xAI would fully embrace the transparency they mention as a core value. That would include also posting system cards for models and explaining the processes they use to stop “”unauthorized modifications”” going forward.”” / X https://x.com/emollick/status/1923192977800802626
You asked us for a more efficient 2.5 Flash. You got it. It’s improved across key benchmarks for reasoning, multimodality, code and long context – while getting even more efficient, using 20-30% less tokens in our evaluations. https://x.com/GoogleDeepMind/status/1924879653048877192
Google’s new Imagen 4 Ultra comes 3rd in the Artificial Analysis Image Arena, falling just short of OpenAI’s GPT-4o and ByteDance’s Seedream 3.0 After a few days in the Artificial Analysis Image Arena, Imagen 4 Ultra appears to be a noticeable but small upgrade over Imagen 3, https://x.com/ArtificialAnlys/status/1925803621792227637
📰 News in Arena: Mistral Medium 3 makes a strong debut with the community! Highlights: 💠 #11 overall in chat: a +90 point leap from Mistral Large 💠Top-tier in technical domains (#5 in Math, #7 in Hard Prompts & Coding) 💠#9 in WebDev Arena Congrats to @MistralAI on the https://x.com/lmarena_ai/status/1924482515244622120
Salesforce introduces: BLIP3-o: A Family of Fully Open Unified Multimodal Models—Architecture, Training and Dataset “”we introduce a novel approach that employs a diffusion transformer to generate semantically rich CLIP image features, in contrast to conventional VAE-based https://x.com/iScienceLuvr/status/1922843713514193076
[2503.04474] Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges https://arxiv.org/abs/2503.04474
Gemini 2.5 Pro Deep Think is here! > Leads on LiveCodeBench > Scores 84.0% on MMMU Uses parallel thinking techniques that lead to significant performance improvements. https://x.com/omarsar0/status/1924885378692972554
It’s so over. Gemini 2.5 flash beats everything else (except 2.5 pro). 😎 I’ve only been back at Google for 5 months but together with @quocleix & @vqctran (and a few others) we landed some outrageous research idea (which I can’t talk about 😶) into this model. 🔥”” / X https://x.com/YiTayML/status/1924900546215084318
We’ve shared a lot of updates today, but there’s one we know you’ve been waiting for. Here’s our leaderboard that measures the number of times we’ve said “AI” in the keynote. Looks like we’ve got a new front-runner. 😉 https://x.com/Google/status/1924901352070676540
AlphaEvolve is deeply disturbing for RL diehards like yours truly Maybe midtrain + good search is all you need for AI for scientific innovation And what an alpha move to keep it secret for a year Congrats big G”” / X https://x.com/_jasonwei/status/1923091260354531612
With Google’s AlphaEvolve, we have evidence that LLMs can discover novel & useful ideas, when put together with the right tooling. These results are impressive: given 50 open math problems, the AI rediscovered the leading approach 75% of the time & improved on it 20% of the time https://x.com/emollick/status/1922860271208456588
🚨Breaking from Arena: @GoogleDeepMind’s new Gemini-2.5-Flash climbs to #2 overall in chat, a major jump from its April release (#5 → #2)! Highlights: – Top-2 across major categories (Hard, Coding, Math) – #3 in WebDev Arena, #2 in Vision Arena – New model at the https://x.com/lmarena_ai/status/1924894101918646510
Gemini Diffusion: diffusion-based LLM, much faster than autoregressive LLMs Gemini 2.5 Pro Deep Think: doubles o3’s score on 2025 USAMO math competition Imagen 4: can spell Veo 3: native audio generation, characters can speak Google is back. Artificial Pichai Intelligence.”” / X https://x.com/Yuchenj_UW/status/1924896740068753825
Lumina-Next on Qwen base, from Salesforce. Slightly surpasses Janus-Pro. I hope we start seeing actually multimodally pretrained unified models soon. https://x.com/teortaxesTex/status/1922961229233946869
🚀 Big news: Together AI has acquired @RefuelAI! Refuel specializes in models and tools that turn messy, unstructured data into clean, structured input—exactly what teams need to build high-quality, production-grade AI applications. Details below 👇 https://x.com/togethercompute/status/1923059615534616846
Who said Transformers couldn’t be good at forecasting? Datadog’s new open model tops forecasting benchmarks! 💥 And boy did they cook. They followed the playbook to build the best model: 1. The best benchmark They release a new benchmark named BOOM, based on observability https://x.com/AymericRoucher/status/1925844148478767243
codex-1 shows strong performance even without AGENTS .md or custom scaffolding. It consistently outperforms o3-high on SWE tasks. https://x.com/omarsar0/status/1923401100419387759




