Image created with Ideogram 3.0. Image prompt: Lower-East-Side street-corner photograph reminiscent of a late-80s album cover: weathered red-brick tenement with exterior fire-escapes, canvas awning shading racks of vintage clothes; above the awning, a hand-painted board reads ‘Tech SPORTSWEAR’; a hanging blade sign in cursive script reads ‘Tech Boutique’; a cardboard box stenciled ‘Tech Surplus’ sits beside stereo equipment; warm golden-hour light, subtle 35mm film grain, muted yet punchy color palette, gritty NYC vibe.
2.5 Pro is now the best model for coding and learning. With a strong ELO score of 1415, it’s topping the WebDev Arena leaderboard – and it incorporates LearnLM, our family of models fine-tuned for learning built with educational experts. #GoogleIO https://x.com/GoogleDeepMind/status/1924878252172353851
Google announces AI Ultra subscription plan https://blog.google/products/google-one/google-ai-ultra/
Tool-using LLMs can learn to reason—without reasoning traces. 🔥 We present Nemotron-Research-Tool-N1, a family of tool-using reasoning LLMs trained entirely via rule-based reinforcement learning—no reasoning supervision, no distillation. 📄 Paper: https://x.com/ShaokunZhang1/status/1922105694167433501
We’ve developed Gemini Diffusion: our state-of-the-art text diffusion model. Instead of predicting text directly, it learns to generate outputs by refining noise, step-by-step. This helps it excel at coding and math, where it can iterate over solutions quickly. #GoogleIO https://x.com/GoogleDeepMind/status/1924888095448825893
StackOverflow questions over time, source SEDE; sadface, lunch has been eaten https://x.com/marcgravell/status/1922922817143660783
Report: Spring 2025 AI Model Usage Trends – Poe https://poe.com/blog/spring-2025-ai-model-usage-trends
Here’s my “”Dark Leisure”” theory of any potential productivity paradox in AI: – most AI use rn is bottom up and hidden (employee first, not company first): employees vibe code, vibe market, vibe write and get stuff done faster – in many orgs, there is too little incentive to”” / X https://x.com/fabianstelzer/status/1926000937702764635
As someone involved in academic research on AI, it is notable to me that most of the key experiments showing the impressive abilities of AI on work, medicine, psychology, and so many other fields were done on GPT-4… a model that is now so obsolete that it is gone from ChatGPT. https://x.com/emollick/status/1923134492115365905
QoL Update: Starting today, you will see an AI generated summary for all papers of Hugging Face Papers! 🔥 GG @mishig25 🐐 https://x.com/reach_vb/status/1925517801197879737
Reasoning Generalization Reasoning fails to generalize across environments. Agents struggle with spatial coordination (Messenger), legal move inference (Hanoi), and adapting to opponent patterns (RPS). Even with reward shaping and hints, models often underperform random https://x.com/omarsar0/status/1924182841677709540
I wish these skeptical AI articles would actually grapple with the growing body of research that AI can really do original research & perform key unstructured tasks across the spectrum of high-end white collar employment. AI criticism is important, but it should be clear-eyed.”” / X https://x.com/emollick/status/1923417536072241529
AI learns how vision and sound are connected, without human intervention | MIT News | Massachusetts Institute of Technology https://news.mit.edu/2025/ai-learns-how-vision-and-sound-are-connected-without-human-intervention-0522
Anthropic closes $2.5 billion credit facility https://www.cnbc.com/amp/2025/05/16/anthropic-ai-credit-facility.html
📢We’re excited to share that we’ve raised $100M in seed funding to support LMArena and continue our research on reliable AI. Led by @a16z and UC Investments (@UofCalifornia), we’re proud to have the support of those that believe in both the science and the mission. We’re https://x.com/lmarena_ai/status/1925241333310189804
Google dropped AlphaEvolve, an AI that discovers algorithms for scientific and computational challenges —Uses Gemini models with auto-evaluation & iteration —Found the first improvement on 1969’s Strassen’s algorithm —Also boosting efficiency for Google https://x.com/adcock_brett/status/1924133683444793819
Gemini Diffusion https://simonwillison.net/2025/May/21/gemini-diffusion/
NVIDIA has published a paper on DREAMGEN – a powerful 4-step pipeline for generating synthetic data for humanoids that enables task and environment generalization. – Step 1: Fine-tune a video generation model using a small number of human teleoperation videos – Step 2: Prompt https://x.com/TheHumanoidHub/status/1925255036965408887
[2505.09662] Large Language Models Are More Persuasive Than Incentivized Human Persuaders https://arxiv.org/abs/2505.09662
Very strong agentic performance by the new Claude 4 Opus and Sonnet, placing 1st and 3rd on the GAIA benchmark. Notice, this leaderboard unfortunately doesn’t include Google’s latest Gemini models nor OpenAI’s o4-mini or o3. https://x.com/scaling01/status/1926017165108375607
Agents taking actions in your ad account🐳 Coming next week to Moby Agents… https://x.com/AY_Orbach/status/1923425142039822517
II-Agent – Intelligent Internet https://ii.inc/web/blog/post/ii-agent
We released Devstral. It is a 24B model released under the Apache 2.0 license. It the best open model on SWE-Bench verified today. You can check our blog post or test it with OpenHands (from @allhands_ai ) following the instructions here: https://x.com/b_roziere/status/1925194095359676768
Deep Think’ boosts the performance of Google’s flagship Google Gemini AI model | TechCrunch https://techcrunch.com/2025/05/20/deep-think-boosts-the-performance-of-googles-flagship-google-gemini-ai-model/
Semantic Layer Summit https://www.semanticlayersummit.com/
6. Gemini AI updates: —Gemini 2.5 Pro Deep Think, a new mode that uses parallel thinking to solve math and coding problems —Gemini 2.5 Flash, upgraded with improved performance across benchmarks Both also include native audio outputs across languages https://x.com/rowancheung/status/1925084387894300762
ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems https://arxiv.org/pdf/2505.11831
How far can reasoning models scale? | Epoch AI https://epoch.ai/gradient-updates/how-far-can-reasoning-models-scale
Generative AI Adoption Index – US Press Center https://press.aboutamazon.com/aws/2025/5/generative-ai-adoption-index
Anthropic’s Claude 4 models are here. Opus 4 and Sonnet 4 both show strong (& improved) coding abilities, with Sonnet 4 at 72.7% on SWE-bench and Opus 4 at 72.5% What does this mean for developers using Cline? 🧵 https://x.com/cline/status/1925680741108613503
opus 4 review: its a good model i was an early tester and found that it combines much of what people loved about sonnet 3.6 and 3.7 (and some opus!) into something which is much greater than the parts amazing at long-term tasks, intelligent tool usage, and helping you write!”” / X https://x.com/nearcyan/status/1925698351661502810
Anthropic was indeed underpriced at 2%. It peaked at 31.6% 3 minutes after the release of Claude 4 Opus However, the dust has settled and betting markets still think that Google has the best model. Do you agree? https://x.com/scaling01/status/1925623699186589698
Claude-4 Opus is definitely not a frontier model for mathematics (screenshot from MathArena) Results for Claude-4 Sonnet haven’t been published. https://x.com/scaling01/status/1926018522372514037
Claude 4 System Card https://www-cdn.anthropic.com/4263b940cabb546aa0e3283f35b686f4f3b2ff47.pdf
LM Arena, the organization behind popular AI leaderboards, lands $100M | TechCrunch https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/
The three gulfs of LLM pipeline development is a powerful mental model to keep in mind while creating AI evals. @eugeneyan , @sh_reya and I discuss here (links to more resources in the replies) https://x.com/HamelHusain/status/1923382256867180606
The Top 5 domestic large models contend for supremacy, a decisive battle in AGI – Google Docs https://docs.google.com/document/d/1qzm3wVxr3m-vFXqSQQjI8MaP9wADddbOpEQkH2kZLfo/edit?tab=t.0#heading=h.8xti639xclnh
US lawmakers have concerns about Apple-Alibaba deal | TechCrunch https://techcrunch.com/2025/05/18/u-s-lawmakers-have-concerns-about-apple-alibaba-deal/
Chatbot Arena Group Goes From Academic Project to $600 Million Startup – Bloomberg https://www.bloomberg.com/news/articles/2025-05-21/lmarena-goes-from-academic-project-to-600-million-startup?embedded-checkout=true
Really cool how DeepSeek is now the benchmark for Nvidia”” / X https://x.com/teortaxesTex/status/1924588309688267139
This is the second time that this has happened. I really wish xAI would fully embrace the transparency they mention as a core value. That would include also posting system cards for models and explaining the processes they use to stop “”unauthorized modifications”” going forward.”” / X https://x.com/emollick/status/1923192977800802626
You asked us for a more efficient 2.5 Flash. You got it. It’s improved across key benchmarks for reasoning, multimodality, code and long context – while getting even more efficient, using 20-30% less tokens in our evaluations. https://x.com/GoogleDeepMind/status/1924879653048877192
Google’s new Imagen 4 Ultra comes 3rd in the Artificial Analysis Image Arena, falling just short of OpenAI’s GPT-4o and ByteDance’s Seedream 3.0 After a few days in the Artificial Analysis Image Arena, Imagen 4 Ultra appears to be a noticeable but small upgrade over Imagen 3, https://x.com/ArtificialAnlys/status/1925803621792227637
📰 News in Arena: Mistral Medium 3 makes a strong debut with the community! Highlights: 💠 #11 overall in chat: a +90 point leap from Mistral Large 💠Top-tier in technical domains (#5 in Math, #7 in Hard Prompts & Coding) 💠#9 in WebDev Arena Congrats to @MistralAI on the https://x.com/lmarena_ai/status/1924482515244622120
Salesforce introduces: BLIP3-o: A Family of Fully Open Unified Multimodal Models—Architecture, Training and Dataset “”we introduce a novel approach that employs a diffusion transformer to generate semantically rich CLIP image features, in contrast to conventional VAE-based https://x.com/iScienceLuvr/status/1922843713514193076
[2503.04474] Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges https://arxiv.org/abs/2503.04474
Insights into DeepSeek-V3 Scaling Challenges and Reflections on Hardware for AI Architectures https://x.com/_akhaliq/status/1923001697498006016
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures Overview: DeepSeek-V3 which is an LLM trained on 2,048 H800 GPUs, utilizes hardware-aware co-design incorporating Multi-head Latent Attention, MoE, FP8 training, and a Multi-Plane https://x.com/TheAITimeline/status/1924232113101890003
Do LLMs Really Understand Cell Biology? Interesting paper evaluating LLMs potential in understanding cell biology. Finding: It finds that specialist models don’t work so great. Generalist models, such as Qwen and DeepSeek, exhibit preliminary understanding capabilities within https://x.com/omarsar0/status/1922662317986099522
Everything you need to know to understand GRPO: GRPO (Group Relative Policy Optimization) is a reinforcement learning algorithm created by DeepSeek specifically for LLMs. It drops the need to use critic network like in PPO and so it doesn’t use absolute value estimate to https://x.com/TheTuringPost/status/1925146257372381485
The Strange Behavior of LLMs in Hiring Decisions: Systemic Gender and Positional Biases in Candidate Selection https://davidrozado.substack.com/p/the-strange-behavior-of-llms-in-hiring
Gemini 2.5 Pro Deep Think is here! > Leads on LiveCodeBench > Scores 84.0% on MMMU Uses parallel thinking techniques that lead to significant performance improvements. https://x.com/omarsar0/status/1924885378692972554
It’s so over. Gemini 2.5 flash beats everything else (except 2.5 pro). 😎 I’ve only been back at Google for 5 months but together with @quocleix & @vqctran (and a few others) we landed some outrageous research idea (which I can’t talk about 😶) into this model. 🔥”” / X https://x.com/YiTayML/status/1924900546215084318
We’ve shared a lot of updates today, but there’s one we know you’ve been waiting for. Here’s our leaderboard that measures the number of times we’ve said “AI” in the keynote. Looks like we’ve got a new front-runner. 😉 https://x.com/Google/status/1924901352070676540
AlphaEvolve is deeply disturbing for RL diehards like yours truly Maybe midtrain + good search is all you need for AI for scientific innovation And what an alpha move to keep it secret for a year Congrats big G”” / X https://x.com/_jasonwei/status/1923091260354531612
With Google’s AlphaEvolve, we have evidence that LLMs can discover novel & useful ideas, when put together with the right tooling. These results are impressive: given 50 open math problems, the AI rediscovered the leading approach 75% of the time & improved on it 20% of the time https://x.com/emollick/status/1922860271208456588
🚨Breaking from Arena: @GoogleDeepMind’s new Gemini-2.5-Flash climbs to #2 overall in chat, a major jump from its April release (#5 → #2)! Highlights: – Top-2 across major categories (Hard, Coding, Math) – #3 in WebDev Arena, #2 in Vision Arena – New model at the https://x.com/lmarena_ai/status/1924894101918646510
Gemini Diffusion: diffusion-based LLM, much faster than autoregressive LLMs Gemini 2.5 Pro Deep Think: doubles o3’s score on 2025 USAMO math competition Imagen 4: can spell Veo 3: native audio generation, characters can speak Google is back. Artificial Pichai Intelligence.”” / X https://x.com/Yuchenj_UW/status/1924896740068753825
Analog Foundation Models “”In this work, we introduce a general and scalable method to robustly adapt LLMs for execution on noisy, low-precision analog hardware. Our approach enables state-of-the-art models – including Phi-3-mini-4k-instruct and Llama-3.2-1B-Instruct to – https://x.com/iScienceLuvr/status/1923269433751158884
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset Author’s Explanation: https://x.com/TheAITimeline/status/1924232118755824119
Emerging Properties in Unified Multimodal Pretraining https://arxiv.org/pdf/2505.14683
Wolfram’s guess about why ChatGPT was suddenly so much more capable than expected is that, as the scale of the model increased, it found the hidden deep structural patterns of writing which is similar to the structure of human thought (“writing is thinking” is a common assertion) https://x.com/emollick/status/1923533622071480727
Lumina-Next on Qwen base, from Salesforce. Slightly surpasses Janus-Pro. I hope we start seeing actually multimodally pretrained unified models soon. https://x.com/teortaxesTex/status/1922961229233946869
🚀 Big news: Together AI has acquired @RefuelAI! Refuel specializes in models and tools that turn messy, unstructured data into clean, structured input—exactly what teams need to build high-quality, production-grade AI applications. Details below 👇 https://x.com/togethercompute/status/1923059615534616846
[2504.05652] Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking https://arxiv.org/abs/2504.05652
[2505.09558v1] WavReward: Spoken Dialogue Models With Generalist Reward Evaluators https://arxiv.org/abs/2505.09558v1
[2505.14489v1] Reasoning Models Better Express Their Confidence https://arxiv.org/abs/2505.14489v1
🚨 New paper 🚨 J1: Incentivizing Thinking in LLM-as-a-Judge via RL – Converts judgement task into a verifiable one for both verifiable and non-verifiable prompts. Uses only synthetic pairwise data – Optimizes thoughts, scores, and judgments using GRPO – Outperforms all https://x.com/jaseweston/status/1923186392420450545
a lot of past research (relative representations, The Platonic Representation Hypothesis, comparison metrics like CCA, SVCCA, …) has asserted that once they reach a certain scale, different models learn the same thing this has been shown using various metrics of comparison”” / X https://x.com/jxmnop/status/1925224614289908107
and practically, this is bad for vector databases. this means that even if you fine-tune your own model, and keep the model secret, someone with access to embeddings alone can decode their text embedding inversion without model access 😬 https://x.com/jxmnop/status/1925224622401692134
Cost-Effective, Low Latency Vector Search It looks like a pretty good read for developers looking to scale vector searches. Proposal: Integrates DiskANN (a vector indexing library) inside of Azure Cosmos DB NoSQL (an operational dataset) that uses a single vector index per https://x.com/omarsar0/status/1921938925142384736
Deep Learning is Applied Topology – by theahura https://theahura.substack.com/p/deep-learning-is-applied-topology
did you know people have been training neural networks on text since 2003? everyone talks about Attention Is All You Need. but this is the real paper that got our field started. it was in 2003, in montreal. i read it, and it was even more forward-thinking than i expected: https://x.com/jxmnop/status/1924207755956478400
Do LLMs struggle in long, multi-turn conversations? Yes, they do, performance degrades in multi-turn conversations, due to a increase in unreliability. A new paper found an drop of 39% in multi-turn settings, showing LLMs make premature assumptions and struggle to recover from https://x.com/_philschmid/status/1923282796081987848
Emergent social conventions and collective bias in LLM populations | Science Advances https://www.science.org/doi/10.1126/sciadv.adu9368
excited to finally share on arxiv what we’ve known for a while now: All Embedding Models Learn The Same Thing embeddings from different models are SO similar that we can map between them based on structure alone. without *any* paired data feels like magic, but it’s real:🧵”” / X https://x.com/jxmnop/status/1925224612872233081
Exploring Quantization Backends in Diffusers https://huggingface.co/blog/diffusers-quantization
Helping Characters Remember What Matters Most https://blog.character.ai/helping-characters-remember-what-matters-most/
How Fast Can Algorithms Advance Capabilities? – Epoch AI https://epochai.substack.com/p/how-fast-can-algorithms-advance-capabilities
HuB is a unified framework that enables humanoids to execute complex balancing tasks. It integrates reference motion refinement, balance-aware policy learning, and sim-to-real robustness training. Project page: https://x.com/TheHumanoidHub/status/1922153643555406058
Improving Assembly Code Performance with Large Language Models via Reinforcement Learning https://x.com/_akhaliq/status/1924505603403047208
Introducing KumoRFM: A Foundation Model for In-Context Learning on Relational Data – Kumo https://kumo.ai/company/news/kumo-relational-foundation-model/
LLMs Get Lost in Multi-turn Conversation The cat is out of the bag. Pay attention, devs. This is one of the most common issues when building with LLMs today. Glad there is now paper to share insights. Here are my notes: https://x.com/omarsar0/status/1922755721428598988
Mitigation strategies: Four strategies are tested to counteract the reasoning-induced failures: 1. few-shot in-context learning with carefully chosen examples 2. self-reflection (models critique and revise their own answers) 3. self-selective reasoning (models decide when”” / X https://x.com/omarsar0/status/1924458176096751806
Model Merging in Pre-training of Large Language Models “”We present the Pre-trained Model Averaging (PMA) strategy, a novel framework for model merging during LLM pre-training. Through extensive experiments across model scales (from millions to over 100B parameters), we https://x.com/iScienceLuvr/status/1924804322568896800
Reasoning Reliability Prompting strategies yield high variance. The same prompt can both outperform or underperform baselines depending on execution. This brittleness undermines the reliability of advanced reasoning techniques for consistent improvement.”” / X https://x.com/omarsar0/status/1924182837307089216
Researchers significantly improved a large language model’s reasoning by fine-tuning it on just 1,000 examples. Their model, s1, achieved strong performance on benchmarks like AIME and MATH 500 by appending the word “Wait” during inference to extend the model’s reasoning https://x.com/DeepLearningAI/status/1923366576549433520
RL for Search-Efficient LLMs Presents a new post-training RL framework that explicitly trains LLMs to optimize search usage. Recipe: structured reasoning template and reward policy + GRPO Leads to smarter and more efficient reasoning and retrieval of external knowledge. https://x.com/omarsar0/status/1922665313117552664
Scaling Reasoning can Improve Factuality in Large Language Models https://x.com/_akhaliq/status/1924477447120068895
so yes, we’re using a GAN. adversarial loss (to align representations) and cycle consistency loss (to make sure we align the *right* representations) and it works. here’s embeddings from GTR (a T5-based model) and GTE (a BERT-based model), after training our GAN for 50 epochs: https://x.com/jxmnop/status/1925224618060587523
The Pitfalls of Reasoning for Instruction-Following in LLMs If you’re dev using reasoning models, read this one. Lots of great insights and mitigation tactics. Here are my notes: https://x.com/omarsar0/status/1924458157444579700
The study finds a strong correlation between question similarity and strategy similarity, meaning models tend to use similar reasoning patterns for similar problems. This enables accurate prediction of optimal strategies for new, unseen questions based on their resemblance to”” / X https://x.com/omarsar0/status/1923383967652127226
The whole trick to building AI systems is scaling test time compute. You do this by building search. Little search. Greedy as hell search. Narrow and deep search. Shallow and broad search. Approximate search. Exact search. Hybrid search. Searching by offloaded computation.”” / X https://x.com/mbusigin/status/1922809557459349849
theoretically, the implications of this seem big. we call it The Strong Platonic Representation Hypothesis: models of a certain scale learn representations that are so similar that we can learn to translate between them, using *no* paired data (just our version of CycleGAN) https://x.com/jxmnop/status/1925224620166128039
We love to focus on models and algorithms, but data quality makes the real difference when training LLMs! Here’s a practical guide for debugging your LLM’s training dataset… Developing an LLM. When training an LLM, we follow an iterative, two-step process: 1. Train our model https://x.com/cwolferesearch/status/1924482811899064451
we take things a step further. if models E1 and E2 are learning ‘similar’ representations, what if we were able to actually align them? and can we do this with just random samples from E1 and E2, by matching their structure? we take inspiration from 2017 GAN papers that https://x.com/jxmnop/status/1925224615866966459
We’ve had support for Dicts, a serverless KV store, for a while now on @modal_labs. @dshaar_ just overhauled them to make them 10x better: 🚀 No scale limits 📆 LRU-cache semantics 🔐 Distributed locking 💎 Durability With this, they feels much more like a primitive that”” / X https://x.com/akshat_b/status/1924552967673545055
CGS-GAN 3D Consistent Gaussian Splatting GANs for High Resolution Human Head Synthesishttps://fraunhoferhhi.github.io/cgs-gan/
DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
https://ziqiaopeng.github.io/dualtalk/
Who said Transformers couldn’t be good at forecasting? Datadog’s new open model tops forecasting benchmarks! 💥 And boy did they cook. They followed the playbook to build the best model: 1. The best benchmark They release a new benchmark named BOOM, based on observability https://x.com/AymericRoucher/status/1925844148478767243
With the retraction of the MIT paper, we now have no clear experimental evidence that LLMs act as a multiplier for high performers (even though that would still seem to make sense). Lots of evidence that low performers get a big performance boost, though, across many studies.”” / X https://x.com/emollick/status/1923571343590633839
This paper was apparently fraudulent. From MIT. https://x.com/emollick/status/1923411824893968601
Larger models benefit less from strategic prompting. While strategies improve smaller models on long-text understanding and planning. Llama3.3-70B shows only marginal gains and, in some cases, experiences performance drops due to overcautious or inefficient reasoning paths.”” / X https://x.com/omarsar0/status/1924182839081218092
Avoid Excessive Reasoning Excessive reasoning hurts smaller models on simple tasks. A too-long prompt can negatively impact smaller models on basic reactive tasks, while larger models show more robust behaviour. Adding reflections or plans dilutes key information and often”” / X https://x.com/omarsar0/status/1924182835289620950
codex-1 shows strong performance even without AGENTS .md or custom scaffolding. It consistently outperforms o3-high on SWE tasks. https://x.com/omarsar0/status/1923401100419387759
Qwen3 Technical Report Author’s Explanation: https://x.com/TheAITimeline/status/1924232110383960163
Tsinghua University researchers detailed HuB, a unified framework to help humanoids handle extreme balancing tasks It integrates reference motion refinement, balance-aware policy learning, and robustness training to improve sim-to-real consistency https://x.com/adcock_brett/status/1924133916971020739
This is a interesting but appropriately hard to generalize It shows that AIs tend to generalize scientific abstracts in ways that may be less precise (but not actually erroneous) than the work of experts. Hard to know how much it matters, but good to see study of subtle effects https://x.com/emollick/status/1924513124096377214
We’re starting to see more and more AI for chemistry and biology which I’m super excited about given the potential for good! @AIatMeta just released OMol25 on @huggingface, a dataset of 𝟭𝟬𝟬𝗠+ 𝗺𝗼𝗹𝗲𝗰𝘂𝗹𝗮𝗿 𝗰𝗼𝗻𝗳𝗼𝗿𝗺𝗲𝗿𝘀 spanning 83 elements and diverse chemical https://x.com/ClementDelangue/status/1924836697373565191




