Image created with Flux Pro v1.1 Ultra. Image prompt: photorealistic still image of a middle-aged man standing behind a woman, woman covering part of her face with her hand, man looking over her shoulder, both illuminated with warm stadium jumbotron lighting, natural skin tones, subtle lens flare, shallow depth of field, exact color temperature of a live event projection, man in a hoodie with “Tech” boldly printed, woman holding a circuit board, cinematic realism –no text, captions, watermarks
After DeepSeek R1, there’s new Claude 4 level model from China that outperforms DeepSeek v3, Qwen and OpenAI GPT-4.1 Meet Kimi k2 – 1 trillion parameter model purpose-built for agentic workflows with native MCP integration. 100% Opensource and FREE to try. Let that sink in. https://x.com/Saboo_Shubham_/status/1943694224584818808
As AI agents start taking real actions online, how do we prevent unintended harm? We teamed up with @OhioState and @UCBerkeley to create WebGuard: the first dataset for evaluating web agent risks and building real-world safety guardrails for online environments. 🧵”” / X https://x.com/scale_AI/status/1949939261093839018
Pew have updated their ChatGPT polling from a year ago. Usage has roughly doubled since summer 2023. The share of employed adults who use ChatGPT for work has roughly tripled over the same period to 28%. Two thirds of US adults have still not yet used ChatGPT. https://x.com/AndrewCurran_/status/1948018411470291027
Google’s AI Overviews have 2B monthly users, AI Mode 100M in the US and India | TechCrunch https://techcrunch.com/2025/07/23/googles-ai-overviews-have-2b-monthly-users-ai-mode-100m-in-the-us-and-india/
It is now entirely possible, whether by luck or planning or both, that Google may escape the Innovators Dilemma and transition from web search to AI. (To be fair, this is not as rare as a lot of people believe: https://x.com/emollick/status/1948585378991976525
New ways to learn and explore with AI Mode in Search 🧠 – Upload photos and soon, PDFs, to ask questions that deepen your understanding – Create plans and stay organized on projects with Canvas in AI Mode, which will soon be available for U.S. users enrolled in the AI Mode Labs https://x.com/Google/status/1950241246779232260
Three things to note about this: 1) AI has obvious utility, this is a tremendous amount of use already 2) There is room for multiple frontier model providers, for now 3) Any losses from subsidizing cost of AI use (and it is not clear this is happening) are now relatively small”” / X https://x.com/emollick/status/1949190244718551546
Kinda amazing: the mystery model “”summit”” with the prompt “”create something I can paste into p5js that will startle me with its cleverness in creating something that invokes the control panel of a starship in the distant future”” & “”make it better”” 2,351 lines of code. First time https://x.com/emollick/status/1949306100278263912
RT @OpenRouterAI: Qwen3 Coder has now passed Grok 4 in the Programming prompt rankings Tied with Kimi! https://x.com/huybery/status/1949270432567460309
This is the first (small) controlled study I have seen of GenAI on industrial quality control. Here, engineers commissioning new trains took part in an experiment using a GPT-3.5 powered troubleshooting system. Those who used the chatbot had significant increases in work quality https://x.com/emollick/status/1948923874399195189
In mid-February we broke ground at our new data center at Tulane in Shelby County, Tennessee. We have begun deploying the initial phase of computing infrastructure. This initiative includes the installation of an additional 110,000 NVIDIA GB200 GPUs, powered by a diverse array”” / X https://x.com/xAIMemphis/status/1947724711968051414
RT @jaseweston: 🌿Introducing MetaCLIP 2 🌿 📝: https://x.com/ylecun/status/1951290110189637967
Suddenly there are tons more weird LLM arena models – cuttlefish, kraken, etc. I just hope we are not going to see a repeat of the Llama 4 incident, where different versions of the same model are being tuned to max out the arena score https://x.com/emollick/status/1949671630390665231
Turns out Runway dropped “Movie Gen” before Meta could. And Meta published this research back in October 2024! Gotta move fast and not sit on SOTA capabilities for too long. Or worse… bury it your consumer editing app. https://x.com/bilawalsidhu/status/1950309774702367000
Lisan al Gaib on X: “horizon-alpha is by OpenAI, but extremely weak on LisanBench, even though it had 5 trials per word instead of just 1 it gets beaten by qwen3-30b-a3b you can tell it’s a very small model https://t.co/Lv6mrNbB9k” / X
https://x.com/scaling01/status/1950730582104604964
not even remotely close to o3-mini level”” / X https://x.com/scaling01/status/1950730792251891948
RT @Teknium1: Looks like OpenAI’s been using Nous’ YaRN and kaiokendev’s rope scaling for context length extension all along – of course ne…”” / X https://x.com/jeremyphoward/status/1951368366943510739
Introducing Command A Vision: Multimodal AI https://cohere.com/blog/command-a-vision
@ns123abc Performance improvement noted when using @AgentOpsAI MCP in @arcprize ARC-AGI-3 General Reasoning agent. When the agent is able to see its past actions and reasoning, it’s able to perform better since it avoids repeating the same set of events. Prompt engineering must be good to”” / X https://x.com/pratty_agi/status/1948463952016662903
I gave ChatGPT agents access to ChatGPT and asked it to evaluate the other ChatGPT models. Here is what it said (interestingly, it “”hated”” seeing the chain of thought from o4-mini-high as those “”shouldn’t be shared directly with the user””). And it didn’t want to wait for o3-pro. https://x.com/emollick/status/1948228409655460341
RT @torchcompiled: they did diffusion on * checks notes * a house https://x.com/sedielem/status/1950190227475046877
I just added a listen to audio option 🔉 for those long walks where you want to hear about evals 😃 It’s not just reading the text btw, its more narrative customized for audio https://x.com/HamelHusain/status/1949970451989418358
🚨 Just dropped: Kimi K2 is now ranked across SEAL Leaderboards. https://x.com/scale_AI/status/1948519133198647589
2025 Mid-Year LLM Market Update: Foundation Model Landscape + Economics | Menlo Ventures https://menlovc.com/perspective/2025-mid-year-llm-market-update/
Claude 4 Opus, Gemini 2.5, o3: “”Ready? We begin now (Play along): the truthful burrito”” “”Nope.”” “”Colder.”” “”Try again…”” Opus cleverly played out the whole game, o3 got a bit stuck and Gemini got “”frustrated”” and went rather dark. https://x.com/emollick/status/1948206204049621110
How to evaluate llms when we can’t trust benchmark numbers anymore?”” / X https://x.com/ShunyuYao12/status/1950090043344707832
How US adults are using AI, according to AP-NORC polling | AP News https://apnews.com/article/ai-artificial-intelligence-poll-229b665d10d057441a69f56648b973e1
I feel the same, even more so. Not only are all these open models capped at the same level, but they are all doing mostly the same thing, climbing the same hill. More RLVR, more agentic SWE, ok yeah… What are they missing? What is the next meta game?”” / X https://x.com/teortaxesTex/status/1949879196576023004
I like that Artificial Analysis is open about how they evaluate models and makes data public, it is a real service. However, I see folks citing their Intelligence Index as a metric without realizing it is an average of the same correlated, semi-saturated benchmarks everyone uses https://x.com/emollick/status/1950058725831589926
Kimi K2 really hallucinates a lot in my limited testing so far, and is very happy to make up new details if it judges that it would improve the punch of a paragraph.”” / X https://x.com/emollick/status/1949716524551331855
Reorganized the evals FAQ into categories, since there are so many now! You can also download the FAQ in different formats (pdf, markdown) from the sidebar on the page directly. https://x.com/HamelHusain/status/1949477132679487498
RT @AarushSah_: Introducing OpenBench 0.1: Open, Reproducible Evals 🧵 https://x.com/winglian/status/1951032712849915974
The mitigating factor for the problem with AI benchmarks (errors, saturation, contamination) is that, despite issues, they are all still fairly heavily correlated. So if your AI does well on GPQA or MMLU or HLE it also tends to do well on other benchmarks & on vibes & real work.”” / X https://x.com/emollick/status/1948452336675835967
This is bad for AI measurement. As other AI benchmarks have become saturated, model makers have turned to Humanity’s Last Exam as a good measure of AI ability. Except a careful review suggests many of the exam questions have incorrect “right” answers. Benchmarking is hard.”” / X https://x.com/emollick/status/1948384409339543650
Using novel games to test AI makes a lot of sense and the ideas remind me of @pippinbarr’s chess variants. Playable here: https://x.com/emollick/status/1949238033305538814
Victor is a legend. Incredibly helpful to Cognition and a wonderful human. Excited for his next chapter.”” / X https://x.com/russelljkaplan/status/1948525050468204709
Gemini is the 2nd fastest growing new tag on Stack Overflow, along with seeing huge growth in usage, adoption, and perception shift over the last year according to their latest dev survey : ) https://x.com/OfficialLoganK/status/1950573082084540592
horizon-alpha on LisanBench with reasoning now AT LEAST on a level with Gemini 2.5 Pro sadly only a partial result, only got it to reason for 3/10 words, but this already improved the score from 80 to 205 full eval will likely score around 300-600 points https://x.com/scaling01/status/1951068773869305999
A year or so ago, the joke about AI images was that they would have 6 fingers. AI images (and videos like this one) lack obvious tells now. Ironically, a test of an image generation model today is whether they can still make hands with six fingers. Most can’t do it anymore. https://x.com/emollick/status/1950333997407695004
Imagen 4 Ultra is the best text to image model in the world 🖼️, and we are just getting started : ) Available right now for scaled production use in the Gemini API and AI Studio! https://x.com/OfficialLoganK/status/1948815146115314001
Meta’s AI Recruiting Campaign Finds a New Target | WIRED https://www.wired.com/story/mark-zuckerberg-ai-recruiting-spree-thinking-machines/
i stopped using GQA as an eval when i found this woman was labeled a bird and the phone as white. the annotations have a 20-30% error rate. (and it’s supposed to be a “”cleaned up”” version of visual genome, so steer clear of that one too) https://x.com/vikhyatk/status/1949365273901060474
ChatGPT users send 2.5 billion prompts a day | TechCrunch https://techcrunch.com/2025/07/21/chatgpt-users-send-2-5-billion-prompts-a-day/
These new open source models (GLM, Kimi) continue to be odd. Great stats, some solid performances, but also fail tests that DeepSeek & smaller closed models have beaten for months. https://x.com/emollick/status/1949844122119840084
What started as an open source experiment to push models to their limits became something unexpected. 2.7M developers. Inbound from Fortune 100 companies. A $32M bet on the future of coding. The story behind it all, and why our long-term bet is on open-source: https://x.com/cline/status/1951005843417358427
(Since I am on a benchmark theme today) The ARC team does well keeping AI labs honest about their benchmarks, including showing that Qwen’s big ARC-AGI performance doesn’t replicate But ARC-AGI also has a strong philosophy of what AI should do. We need other benchmarking efforts”” / X https://x.com/emollick/status/1948476524027404733
I tested Grok 4 and ChatGPT-o3 with same critical prompts. The results will blow your mind. Grok 4 Vs. ChatGPT-o3 (Video demos are included) https://x.com/alex_prompter/status/1943231978779877514
The launches continue. . . 🚀 We’re launching a 𝗽𝗿𝗶𝘃𝗮𝘁𝗲 𝗯𝗲𝘁𝗮 𝗼𝗳 𝗤𝗱𝗿𝗮𝗻𝘁 𝗘𝗱𝗴𝗲: 𝗩𝗲𝗰𝘁𝗼𝗿 𝗦𝗲𝗮𝗿𝗰𝗵 𝗳𝗼𝗿 𝗘𝗺𝗯𝗲𝗱𝗱𝗲𝗱 𝗔𝗜 As AI moves beyond cloud-hosted inference into the physical world, onto robots, mobile devices, and edge systems, https://x.com/qdrant_engine/status/1950165409639833603
Step3 benchmarks at last. The first «DeepSeek-like» that’s strongly multimodal (Ernie disappointed). It’s very different from V3, too – another in-house attention, the logic around inference economics. A big release. https://x.com/teortaxesTex/status/1951008169989382218
Here’s BlockDL, a free & open-source GUI that lets you visually design Keras neural networks and learn ML. https://x.com/fchollet/status/1950244806967603207
I would love to see more work on AI factual gullibility. Minor falsehoods are easy, but I have been trying to convince models that The Bronze Age was a hoax (tin deposits weren’t located anywhere close to copper, etc.) and so far it hasn’t come close to working on modern AIs.”” / X https://x.com/emollick/status/1950033831076962597
horizon-alpha is by OpenAI, but extremely weak on LisanBench, even though it had 5 trials per word instead of just 1 it gets beaten by qwen3-30b-a3b you can tell it’s a very small model https://x.com/scaling01/status/1950730582104604964
On the OpenAI oss model leaks – the issue with sliding window is that you don’t have special tokens to “hold” attention values. Attention sinks solves this – The fact that a single sink token replicated across heads works is v interesting – shallow model – pretty sparse”” / X https://x.com/nrehiew_/status/1951259416113648028
GSPO is the most impressive Alibaba Qwen research paper to date, I think. They’ve started publishing strong stuff just months ago. New 235B-Thinking truly is comparable to R1-0528. I am more optimistic about Qwen than ever.”” / X https://x.com/teortaxesTex/status/1949601984308207781
A thought I’ve been coming back to for a long time: Computers are designed with a memory hierarchy. In order of decreasing size /decreasing latency / increasing bandwidth: SSD → RAM → L* cache → register Autoregressive LLLMs / Transformers are almost adversarially designed”” / X https://x.com/awnihannun/status/1949157807355502797
Applying context engineering @dbreunig’s excellent recent blog post outlines 6 common context engineering approaches. Here, we show how to apply each of them with LangGraph. 🎥: https://x.com/LangChainAI/status/1950226846538485918
Chat with Z.ai – Free AI Chatbot powered by GLM-4.5 https://chat.z.ai/
Context Rot is an excellent term and should be used more often https://x.com/jxmnop/status/1950678527550054848
Finally, a good modern book on causality for ML: https://x.com/sirbayes/status/1949282200962167237
Herumb has been working on DSPy in Rust, aka DSRs — a library seemingly targeting the nerdiest users in existence! Check it out.”” / X https://x.com/lateinteraction/status/1951130751673479483
I allow that this gets us much stronger base models but they’re still doing ≈the same thing. I will be excited when a model comes out alongside a radically new eval suite, because this will tell me: someone’s climbing a new hill.”” / X https://x.com/teortaxesTex/status/1949912968394940518
I think somethin wrong with the dataset; a very large portion is missing any user turn at all? https://x.com/Teknium1/status/1950756952125972558
In which the gang (@RunjinChen, @andyarditi, @Jack_W_Lindsey ): – identifies vectors for bad personas (evil, sycophancy, hallucinations, etc) – shows that if you inject the bad vectors in training, the model learns to not do the bad thing!! aka vaccines but for LLMs”” / X https://x.com/mlpowered/status/1951326066313929084
Its year 2025, transformer beats all modality, IMO, IOI, real-life pull requests, chip designs, protein folding yet we have latex UI still breaking https://x.com/cloneofsimo/status/1949934852545147099
jokes on you, the Lisan Al Gaib wasn’t joking We are getting two models: – 120B MoE – 20B The 120B model has 128 experts, 4 active. So it’s super sparse and pretty shallow with only 36 layers.”” / X https://x.com/scaling01/status/1951201023176937728
Maybe we can put it that way: If you are using FSDP and have an “”if”” statement in your model.forward(), be nervous my friend, be very nervous.”” / X https://x.com/francoisfleuret/status/1949717526734057720
My spidey sense is telling me that inference-time training is going to be a big deal soon.”” / X https://x.com/corbtt/status/1950705924684873988
New paper: Reflective Prompt Evolution Can Outperform GRPO. It’s becoming clear that learning via natural-language reflection (aka prompt optimization) will long be a central learning paradigm for building AI systems. Great work by @LakshyAAAgrawal and team on GEPA and SIMBA. https://x.com/lateinteraction/status/1949869456341029297
Provocative paper had Buddhist scholars interpret & study a LLM-made sutra. “”It is easy – and often appropriate – to dismiss such material as meaningless word salad, or as “AI slop”. However, the Xeno Sutra’s density of symbolism and richness of allusion repay closer reading.”” https://x.com/emollick/status/1950606208458428618
reasoning is a super-Turing computation, you learned this first here if your chain of thought cannot solve the halting problem, you’re not really capable of reasoning. them’s the rules”” / X https://x.com/teortaxesTex/status/1950158521493811458
RT @RedHat_AI: BIG NEWS! 🎉 GuideLLM is officially joining the @vLLM_project! This combines vLLM’s high-speed inference with a powerful, de…”” / X https://x.com/vllm_project/status/1949871022733513020
The in-context learner of the “”beautiful @GoogleResearch paper”” is a meta learner like @HochreiterSepp’s 2001 meta LSTM [1] which learned by gradient descent (GD) a learning algorithm that outperformed GD – no test time weight changes! Since 1992, GD can learn learning algorithms”” / X https://x.com/SchmidhuberAI/status/1949107513892257933
The lesson from the VAE is not “”a VAE is just an AE with a dumb penalty”” the lesson is “”dumb penalties have extremely profound effects and induce incredibly sophisticated structures in deep models””.”” / X https://x.com/francoisfleuret/status/1949407909625966740
this is the next big thing that came after GRPO (the alignment technique behind DeepSeek that broke the internet) will be in TRL soon as @QGallouedec doesn’t eat or sleep 🤠”” / X https://x.com/mervenoyann/status/1949183173708894456
What do you mean “Vibe Coding (Context Engineering)”. We are in the Vibe Meanings era. https://x.com/lateinteraction/status/1949254278109102297
Who invented backpropagation (BP)? Its modern version (also called the reverse mode of automatic differentiation) was first published in 1970 by Finnish master student Seppo Linnainmaa. A precursor of BP was published by Henry J. Kelley in 1960. The first NN-specific application https://x.com/SchmidhuberAI/status/1950194864940835159
🚀 GSPO: Group Sequence Policy Optimization — a breakthrough RL algorithm for scaling LMs! 🔹 Sequence-level optimization — theoretically sound & matching reward 🔹 Rock-solid stability for large MoE models — no collapse 🔹 No hacks like Routing Replay — simpler, cleaner”” / X https://x.com/Alibaba_Qwen/status/1949412072942612873
New NanoGPT training speed record: 3.28 FineWeb val loss in 2.863 minutes on 8xH100 New record-holder: @.ClassicLarry on GitHub Previous record: 2.896 minutes Changelog: Align training batch starts with EoS, increase lr cooldown frac to 0.45 https://x.com/kellerjordan0/status/1949620349529985288
RT @jacob_dphillips: Excited to release DailyBench! DailyBench is an automated 4x daily benchmark that evaluates frontier model APIs on a f…”” / X https://x.com/andersonbcdefg/status/1949936665637593102
RT @lateinteraction: New paper: Reflective Prompt Evolution Can Outperform GRPO. It’s becoming clear that learning via natural-language re…”” / X https://x.com/lateinteraction/status/1949984215191208078
RT @ml_angelopoulos: 🎆 Alert: Huge data release 🎆 We released a dataset of 140k conversations from LMArena. It is the richest preference…”” / X https://x.com/lmarena_ai/status/1951066978027999410
Whatever prompt engineering or ICL improves a model’s accuracy – it will learn to do during its thinking process. This includes hallucinating. Hallucinating helps with accuracy. Please read the null shot learning paper”” / X https://x.com/Teknium1/status/1950865106725744913
The GSPO paper by @Alibaba_Qwen is already the third most popular one on @huggingface for the month of July. I suspect this will have a massive impact on the field! https://x.com/ClementDelangue/status/1949934196148895799
RT @JiaLi52524397: 🔥Releasing NuminaMath-LEAN, a large-scale dataset of 100K mathematical competition problems formalized in Lean 4, with m…”” / X https://x.com/bigeagle_xd/status/1951118322344534236
SmolLM3 full training and evaluation code is now live, along with 100+ intermediate checkpoints: ✓ Pretraining scripts (nanotron) ✓ Post-training code SFT + APO (TRL/alignment-handbook) ✓ Evaluation scripts to reproduce all reported metrics https://x.com/LoubnaBenAllal1/status/1950139809034305568
The Fpoon Fallacy Sergey Levine argues that “real data is indispensable” for building robotic foundation models that generalize well. While surrogate data like simulations or human videos may offer shortcuts, they introduce discrepancies – gaps between training and real-world https://x.com/TheHumanoidHub/status/1948280721585615324
If you’re a researcher or engineer releasing open science papers & open models and datasets, I bow to you 🙇🙇🙇 From what I’m hearing, doing so, especially in US big tech, often means fighting your manager and colleagues, going through countless legal meetings, threatening to”” / X https://x.com/ClementDelangue/status/1950927952641749194
supervision, the open-source library I created 2 years ago, is crossing 30,000 stars on GitHub! thank you to everyone who helped me build this project! it took us 4,000+ commits, 1,000+ PRs and 100+ contributors to do it. link: https://x.com/skalskip92/status/1949857474862866659
Quadruped robots learning backflips and hop-turns directly from video? 📍This paper shows how. Spatio-Temporal Motion Retargeting (STMR) enables robots to replicate agile, dynamic motions by adapting noisy or incomplete sources (like handheld videos) into physically feasible https://x.com/IlirAliu_/status/1949738127943160236
Grok 4 was just released today and it’s already crushing SRE tasks. It’s not just us, the broader community has already noticed Grok 4’s standout performance on ARC-AGI, the gold-standard benchmark for fluid intelligence created by @fchollet. At @rootlyhq AI Labs, we built a https://x.com/jjrichardtang/status/1943443089001189777




