Image created with Ideogram 3.0. Image prompt: Lower-East-Side street-corner photograph reminiscent of a late-80s album cover: weathered red-brick tenement with exterior fire-escapes, canvas awning shading racks of vintage clothes; above the awning, a hand-painted board reads ‘Benchmarks SPORTSWEAR’; a hanging blade sign in cursive script reads ‘Benchmarks Boutique’; a bulletin board displays scorecards titled ‘Benchmarks 1989’ fluttering in the breeze; warm golden-hour light, subtle 35mm film grain, muted yet punchy color palette, gritty NYC vibe.

2.5 Pro is now the best model for coding and learning. With a strong ELO score of 1415, it’s topping the WebDev Arena leaderboard – and it incorporates LearnLM, our family of models fine-tuned for learning built with educational experts. #GoogleIO https://x.com/GoogleDeepMind/status/1924878252172353851

Google announces AI Ultra subscription plan https://blog.google/products/google-one/google-ai-ultra/

StackOverflow questions over time, source SEDE; sadface, lunch has been eaten https://x.com/marcgravell/status/1922922817143660783

Report: Spring 2025 AI Model Usage Trends – Poe https://poe.com/blog/spring-2025-ai-model-usage-trends

Here’s my “”Dark Leisure”” theory of any potential productivity paradox in AI: – most AI use rn is bottom up and hidden (employee first, not company first): employees vibe code, vibe market, vibe write and get stuff done faster – in many orgs, there is too little incentive to”” / X https://x.com/fabianstelzer/status/1926000937702764635

I wish these skeptical AI articles would actually grapple with the growing body of research that AI can really do original research & perform key unstructured tasks across the spectrum of high-end white collar employment. AI criticism is important, but it should be clear-eyed.”” / X https://x.com/emollick/status/1923417536072241529

Anthropic closes $2.5 billion credit facility https://www.cnbc.com/amp/2025/05/16/anthropic-ai-credit-facility.html

📢We’re excited to share that we’ve raised $100M in seed funding to support LMArena and continue our research on reliable AI. Led by @a16z and UC Investments (@UofCalifornia), we’re proud to have the support of those that believe in both the science and the mission. We’re https://x.com/lmarena_ai/status/1925241333310189804

Google dropped AlphaEvolve, an AI that discovers algorithms for scientific and computational challenges —Uses Gemini models with auto-evaluation & iteration —Found the first improvement on 1969’s Strassen’s algorithm —Also boosting efficiency for Google https://x.com/adcock_brett/status/1924133683444793819

Very strong agentic performance by the new Claude 4 Opus and Sonnet, placing 1st and 3rd on the GAIA benchmark. Notice, this leaderboard unfortunately doesn’t include Google’s latest Gemini models nor OpenAI’s o4-mini or o3. https://x.com/scaling01/status/1926017165108375607

Agents taking actions in your ad account🐳 Coming next week to Moby Agents… https://x.com/AY_Orbach/status/1923425142039822517

II-Agent – Intelligent Internet https://ii.inc/web/blog/post/ii-agent

We released Devstral. It is a 24B model released under the Apache 2.0 license. It the best open model on SWE-Bench verified today. You can check our blog post or test it with OpenHands (from @allhands_ai ) following the instructions here: https://x.com/b_roziere/status/1925194095359676768

Deep Think’ boosts the performance of Google’s flagship Google Gemini AI model | TechCrunch https://techcrunch.com/2025/05/20/deep-think-boosts-the-performance-of-googles-flagship-google-gemini-ai-model/

6. Gemini AI updates: —Gemini 2.5 Pro Deep Think, a new mode that uses parallel thinking to solve math and coding problems —Gemini 2.5 Flash, upgraded with improved performance across benchmarks Both also include native audio outputs across languages https://x.com/rowancheung/status/1925084387894300762

ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems https://arxiv.org/pdf/2505.11831

Generative AI Adoption Index – US Press Center https://press.aboutamazon.com/aws/2025/5/generative-ai-adoption-index

Anthropic’s Claude 4 models are here. Opus 4 and Sonnet 4 both show strong (& improved) coding abilities, with Sonnet 4 at 72.7% on SWE-bench and Opus 4 at 72.5% What does this mean for developers using Cline? 🧵 https://x.com/cline/status/1925680741108613503

opus 4 review: its a good model i was an early tester and found that it combines much of what people loved about sonnet 3.6 and 3.7 (and some opus!) into something which is much greater than the parts amazing at long-term tasks, intelligent tool usage, and helping you write!”” / X https://x.com/nearcyan/status/1925698351661502810

Anthropic was indeed underpriced at 2%. It peaked at 31.6% 3 minutes after the release of Claude 4 Opus However, the dust has settled and betting markets still think that Google has the best model. Do you agree? https://x.com/scaling01/status/1925623699186589698

Claude-4 Opus is definitely not a frontier model for mathematics (screenshot from MathArena) Results for Claude-4 Sonnet haven’t been published. https://x.com/scaling01/status/1926018522372514037

LM Arena, the organization behind popular AI leaderboards, lands $100M | TechCrunch https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/

The three gulfs of LLM pipeline development is a powerful mental model to keep in mind while creating AI evals. @eugeneyan , @sh_reya and I discuss here (links to more resources in the replies) https://x.com/HamelHusain/status/1923382256867180606

The Top 5 domestic large models contend for supremacy, a decisive battle in AGI – Google Docs https://docs.google.com/document/d/1qzm3wVxr3m-vFXqSQQjI8MaP9wADddbOpEQkH2kZLfo/edit?tab=t.0#heading=h.8xti639xclnh

US lawmakers have concerns about Apple-Alibaba deal | TechCrunch https://techcrunch.com/2025/05/18/u-s-lawmakers-have-concerns-about-apple-alibaba-deal/

Chatbot Arena Group Goes From Academic Project to $600 Million Startup – Bloomberg https://www.bloomberg.com/news/articles/2025-05-21/lmarena-goes-from-academic-project-to-600-million-startup?embedded-checkout=true

Really cool how DeepSeek is now the benchmark for Nvidia”” / X https://x.com/teortaxesTex/status/1924588309688267139

This is the second time that this has happened. I really wish xAI would fully embrace the transparency they mention as a core value. That would include also posting system cards for models and explaining the processes they use to stop “”unauthorized modifications”” going forward.”” / X https://x.com/emollick/status/1923192977800802626

You asked us for a more efficient 2.5 Flash. You got it. It’s improved across key benchmarks for reasoning, multimodality, code and long context – while getting even more efficient, using 20-30% less tokens in our evaluations. https://x.com/GoogleDeepMind/status/1924879653048877192

Google’s new Imagen 4 Ultra comes 3rd in the Artificial Analysis Image Arena, falling just short of OpenAI’s GPT-4o and ByteDance’s Seedream 3.0 After a few days in the Artificial Analysis Image Arena, Imagen 4 Ultra appears to be a noticeable but small upgrade over Imagen 3, https://x.com/ArtificialAnlys/status/1925803621792227637

📰 News in Arena: Mistral Medium 3 makes a strong debut with the community! Highlights: 💠 #11 overall in chat: a +90 point leap from Mistral Large 💠Top-tier in technical domains (#5 in Math, #7 in Hard Prompts & Coding) 💠#9 in WebDev Arena Congrats to @MistralAI on the https://x.com/lmarena_ai/status/1924482515244622120

Salesforce introduces: BLIP3-o: A Family of Fully Open Unified Multimodal Models—Architecture, Training and Dataset “”we introduce a novel approach that employs a diffusion transformer to generate semantically rich CLIP image features, in contrast to conventional VAE-based https://x.com/iScienceLuvr/status/1922843713514193076

[2503.04474] Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges https://arxiv.org/abs/2503.04474

Gemini 2.5 Pro Deep Think is here! > Leads on LiveCodeBench > Scores 84.0% on MMMU Uses parallel thinking techniques that lead to significant performance improvements. https://x.com/omarsar0/status/1924885378692972554

It’s so over. Gemini 2.5 flash beats everything else (except 2.5 pro). 😎 I’ve only been back at Google for 5 months but together with @quocleix & @vqctran (and a few others) we landed some outrageous research idea (which I can’t talk about 😶) into this model. 🔥”” / X https://x.com/YiTayML/status/1924900546215084318

We’ve shared a lot of updates today, but there’s one we know you’ve been waiting for. Here’s our leaderboard that measures the number of times we’ve said “AI” in the keynote. Looks like we’ve got a new front-runner. 😉 https://x.com/Google/status/1924901352070676540

AlphaEvolve is deeply disturbing for RL diehards like yours truly Maybe midtrain + good search is all you need for AI for scientific innovation And what an alpha move to keep it secret for a year Congrats big G”” / X https://x.com/_jasonwei/status/1923091260354531612

With Google’s AlphaEvolve, we have evidence that LLMs can discover novel & useful ideas, when put together with the right tooling. These results are impressive: given 50 open math problems, the AI rediscovered the leading approach 75% of the time & improved on it 20% of the time https://x.com/emollick/status/1922860271208456588

🚨Breaking from Arena: @GoogleDeepMind’s new Gemini-2.5-Flash climbs to #2 overall in chat, a major jump from its April release (#5 → #2)! Highlights: – Top-2 across major categories (Hard, Coding, Math) – #3 in WebDev Arena, #2 in Vision Arena – New model at the https://x.com/lmarena_ai/status/1924894101918646510

Gemini Diffusion: diffusion-based LLM, much faster than autoregressive LLMs Gemini 2.5 Pro Deep Think: doubles o3’s score on 2025 USAMO math competition Imagen 4: can spell Veo 3: native audio generation, characters can speak Google is back. Artificial Pichai Intelligence.”” / X https://x.com/Yuchenj_UW/status/1924896740068753825

Lumina-Next on Qwen base, from Salesforce. Slightly surpasses Janus-Pro. I hope we start seeing actually multimodally pretrained unified models soon. https://x.com/teortaxesTex/status/1922961229233946869

🚀 Big news: Together AI has acquired @RefuelAI! Refuel specializes in models and tools that turn messy, unstructured data into clean, structured input—exactly what teams need to build high-quality, production-grade AI applications. Details below 👇 https://x.com/togethercompute/status/1923059615534616846

Who said Transformers couldn’t be good at forecasting? Datadog’s new open model tops forecasting benchmarks! 💥 And boy did they cook. They followed the playbook to build the best model: 1. The best benchmark They release a new benchmark named BOOM, based on observability https://x.com/AymericRoucher/status/1925844148478767243

codex-1 shows strong performance even without AGENTS .md or custom scaffolding. It consistently outperforms o3-high on SWE tasks. https://x.com/omarsar0/status/1923401100419387759

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading