Image created with OpenAI GPT-Image-1. Image prompt: vintage Sly & the Family Stone album-cover style, glowing neon marquee sign reading I’M BACK featuring towering motherboard cityscape; grainy retro print texture, vibrant 60s funk color palette, high-resolution

Veo 3 is now the first model to top both the Image to Video and Text to Video leaderboards, outperforming Kling 2.0 and Runway Gen 4 to secure the #1 spot across both modalities! Veo 3 represents a significant leap in Image to Video generation, where Google’s previous Veo 2 had https://x.com/ArtificialAnlys/status/1928318831761707224

Exciting news: @OpenAI’s GPT-Image-1 takes the #1 spot in the Text-to-Image Arena! 🖼️🏆 ➤ Outperforms Google’s Imagen-3.0 by 50+ points ➤ Major leap over DALL·E 3 Huge congrats to @OpenAI! 👏 https://x.com/lmarena_ai/status/1930296340648735147

DeepSeek’s R1 leaps over xAI, Meta and Anthropic to be tied as the world’s #2 AI Lab and the undisputed open-weights leader DeepSeek R1 0528 has jumped from 60 to 68 in the Artificial Analysis Intelligence Index, our index of 7 leading evaluations that we run independently https://x.com/ArtificialAnlys/status/1928071179115581671

I wrote a history of AI in 32 images of otters using wifi on airplanes, from images to video to code. It shows two big trends: rapid improvements in AI models of all types and the growth of open weights AI models. Link in the comments. https://x.com/emollick/status/1929306757903319089

I’ve been using prompts of otters as a test of AI ability. It has taken less than three years to go from a text prompt producing images of abstract masses of fur to producing realistic videos with sound (including “”like the musical Cats but for otters”). https://x.com/emollick/status/1929612980041253132

New Paper! Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents A longstanding goal of AI research has been the creation of AI that can learn indefinitely. One path toward that goal is an AI that improves itself by rewriting its own code, including any code https://x.com/hardmaru/status/1928284568756629756

Anthropic’s progress on SWE-bench Verified really stands out. Naive extrapolation suggests they’ll have it mostly solved in another year. https://x.com/i/web/status/1929568948086800798

An AI agent upgraded its own tools and doubled its bug-fix score. Darwin-style search plus Gödel-style self-reference cracked coding tasks. Pass rate jumps from 20 % to 50 % on SWE-bench-Verified Darwin Gödel Machine (DGM) is a coding agent that rewrites its own code, tests https://x.com/rohanpaul_ai/status/1929461153182122442

in 2026 we will no longer have to deal with call centers and phone menus. AI agents will own the space”” / X https://x.com/braelyn_ai/status/1927888909121507367

HOT: MiMo-VL new 7B vision LMs by Xiaomi surpassing gpt-4o (Mar), competitive in GUI agentic + reasoning tasks ❤️‍🔥 not only that, but also MIT license & usable with transformers 🔥 available on @huggingface 🤗 https://x.com/mervenoyann/status/1928475979753619663

To the extent that this and other recent RLVR findings are true, I think it’s been rather clear but also conceptually sad that: In essence, the reason RLVR post-training for math and coding is so trivially powerful in 2025 but not 2023 is “recent models all now do supervised”” / X https://x.com/i/web/status/1930045203681030248

Our interpretability team recently released research that traced the thoughts of a large language model. Now we’re open-sourcing the method. Researchers can generate “attribution graphs” like those in our study, and explore them interactively.”” / X https://x.com/AnthropicAI/status/1928119229384970244

This is a strawman. We don’t use the phrase “”AGI”” in the MAIM paper (Superintelligence Strategy). In fact, we discuss how the concept of AGI is too vague to be useful in the appendix. We make it clear that the first thing we want to deter is an intelligence recursion—thousands”” / X https://x.com/i/web/status/1929713070265516459

🎉 After 2 years in production serving millions of requests, we’re open sourcing Chatterbox – our state-of-the-art TTS model that just beat ElevenLabs in blind evaluations. In recent testing, 63.75% of listeners preferred Chatterbox over ElevenLabs. Not only is it free and open https://x.com/resembleai/status/1927755087620796668

The team at @podonos did a subjective evaluation where they found that Chatterbox outperforms other proprietary models like ElevenLabs. https://x.com/resembleai/status/1927755092507144348

🚨 Wow. MOSTLY AI launches $100K synthetic data competition, two challenges, $50K each, pushes privacy-safe data sharing. @mostly_ai 🧵1/n The competition awards $50 K per track to teams whose synthetic data best balance realism and privacy. They’re spending large sums to https://x.com/rohanpaul_ai/status/1929527535445889063

I do like these sorts of tests but wish they made it clearer that they represent the minimum bounds of AI. They are representative of what a naive user would experience (which is important!) but not what you could do with Gemini 2.5 Pro or o3. Backward, not forward looking. https://x.com/emollick/status/1930346576125288868

IBM Unveils watsonx AI Labs: The Ultimate Accelerator for AI Builders, Startups and Enterprises in New York City https://newsroom.ibm.com/2025-06-02-ibm-unveils-watsonx-ai-labs-the-ultimate-accelerator-for-ai-builders,-startups-and-enterprises-in-new-york-city

After over a year of saying i need to do an evals conference, we finally have the speakers (and practitioners who lead these evals at work instead of trying to sell you on their evals) to do a dedicated evals track for the first time ever! every AI engineer serious enough about https://x.com/i/web/status/1929609793104499152

Large language models are proficient in solving and creating emotional intelligence tests | Communications Psychology https://www.nature.com/articles/s44271-025-00258-x

🚀 DeepSeek-R1-0528 is here! 🔹 Improved benchmark performance 🔹 Enhanced front-end capabilities 🔹 Reduced hallucinations 🔹 Supports JSON output & function calling ✅ Try it now: https://x.com/deepseek_ai/status/1928061589107900779

DeepSeek has released DeepSeek-R1-0528, an updated version of DeepSeek-R1. How does the new model stack up in benchmarks? We ran our own evaluations on a suite of math, science, and coding benchmarks. Full results in thread! https://x.com/EpochAIResearch/status/1928489524616630483

New DeepSeek just dropped. Proud to serve the fastest DeepSeek R1 0528 inference on OpenRouter (#1 on TTFT and TPS) with our Model APIs. https://x.com/basetenco/status/1928195639822700898

The DeepSeek-R1-0528 model card just dropped. Up 17.5 points on the AIME 2025 test. https://x.com/fdaudens/status/1928055679182352461

Today’s open weights frontier is led by DeepSeek (both reasoning and non-reasoning models) https://x.com/ArtificialAnlys/status/1928477951365939328

We made dynamic 1bit quants for DeepSeek-R1-0528 – 74% smaller 713GB to 185GB. Use the magic incantation -ot “”.ffn_.*_exps.=CPU”” to offload MoE layers to RAM, allowing non MoEs to fit < 24GB VRAM on 16K context! The rest sits in RAM & disk. Quants here: https://x.com/danielhanchen/status/1928278088951157116

On GPQA Diamond, a set of PhD-level multiple-choice science questions, DeepSeek-R1-0528 scores 76% (±2%), outperforming the previous R1’s 72% (±3%). This is generally competitive with other frontier models, but below Gemini 2.5 Pro’s 84% (±3%). https://x.com/EpochAIResearch/status/1928489527204589680

DeepSeek R1 05-28 LiveBench results: – 8th in the Overall ahead of o4-mini, Gemini 2.5 Flash Preview and Qwen3-235B-A22B (biggest competitors) – 1st on Data Analysis !!! – 3rd on Reasoning !! – 4th on Mathematics ! – 11th on Language – 20th on Instruction Following – 23rd on https://x.com/scaling01/status/1928173385399308639

sharing negative results >>> 20k google scholar citations”” / X https://x.com/i/web/status/1929449598524821509

Releasing our Q2 2025 State of AI – China Report 🇨🇳: Chinese AI labs have achieved close to parity with US labs, led by DeepSeek’s leap to world #2 in intelligence and backed by a deep ecosystem of 10+ players Key findings from our analysis: 🇨🇳 The Chinese AI Ecosystem has depth https://x.com/ArtificialAnlys/status/1928477941715079175

Predicting and explaining AI model performance: A new approach to evaluation – Microsoft Research https://www.microsoft.com/en-us/research/blog/predicting-and-explaining-ai-model-performance-a-new-approach-to-evaluation/

The latest mlx-lm has a new dynamic quantization method (made with @angeloskath). It consistently results in better model quality with no increase in size. Some perplexity results (lower is better) for a few Qwen3 base models: https://x.com/i/web/status/1929633379504493048

Nvidia presents ProRL Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models https://x.com/i/web/status/1929540706374201756

new paper from our work at Meta! **GPT-style language models memorize 3.6 bits per param** we compute capacity by measuring total bits memorized, using some theory from Shannon (1953) shockingly, the memorization-datasize curves look like this: ___________ / / (🧵) https://x.com/i/web/status/1929903028372459909

To sum up: 1. Transformers can learn variable binding via emergent mechanisms, w/o explicit symbolic machinery 2. Learning is cumulative, with a general mechanism learned on top of heuristics. This challenges traditional narratives about grokking https://x.com/i/web/status/1929887222553002193

[2506.01833] SPACE: Your Genomic Profile Predictor is a Powerful DNA Foundation Model https://arxiv.org/abs/2506.01833

Last week, @Google dropped a paper on ATLAS, a new architecture that reimagines how models learn and use memory. Unfortunately, it flew under everyone’s radar – but it shouldn’t have! So what’s Atlas bringing to the table? ▪️ Active memory via Google’s so-called Omega rule. It https://x.com/i/web/status/1929992259019432115

Shisa V2 405B: Japan’s Highest Performing LLM https://simonwillison.net/2025/Jun/3/shisa-v2/

Verified Auto Labeling: Smarter Annotation at Scale – June 24, 2025 https://voxel51.com/events/verified-auto-labeling-smarter-annotation-at-scale-june-24-2025

Selecting a Model Based on Stripe Conversion – A Practical Eval for Startups https://cookbook.openai.com/examples/stripe_model_eval/selecting_a_model_based_on_stripe_conversion

The Stripe eval: How @HyperwriteAI A/B tested models and chose GPT-4.1—the one that drove the most customer purchases for them: https://x.com/i/web/status/1929632332837015833

[2506.01533] A Diffusion-Based Method for Learning the Multi-Outcome Distribution of Medical Treatments https://arxiv.org/abs/2506.01533

BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model https://bowang-lab.github.io/BioReason/

There are traditionally two types of research: problem-driven research and method-driven research. As we’ve seen with large language models and now AlphaEvolve, it should be very clear now that total method-driven research is a huge opportunity. Problem-driven research is nice”” / X https://x.com/i/web/status/1929621539881996607

An Operating System for Memory-Augmented Generation in LLMs Lots of great ideas on how to think about memory and better manage it in LLM-based agents. Must read! Here are my notes: https://x.com/omarsar0/status/1928116365640225222

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning “”By examining token entropy patterns in Chain-of-Thought (CoT) reasoning, we observe that only a small fraction of tokens exhibit high entropy, and these tokens act as https://x.com/i/web/status/1929750117927797143

The problem with giving AIs a default personal system prompt is that you have no idea how that prompt interacts with various AI tasks. Our research shows that even small prompt changes (like saying “”please””) can backfire on some problems and lower accuracy in unexpected ways. https://x.com/emollick/status/1928523926490972586

Just as scaling pretraining unlocked powerful meta-learning in LLMs, scaling the volume and diversity of RL environments may be key to giving AIs the meta-skills to continually learn, adapt to new tools, and operate productively in open-ended settings. https://x.com/i/web/status/1929683184163447141

Auto-Labeling Tool for Computer Vision – Voxel51 Annotation https://voxel51.com/annotation

[2505.23836] Large Language Models Often Know When They Are Being Evaluated https://arxiv.org/abs/2505.23836

[2505.24832] How much do language models memorize? https://arxiv.org/abs/2505.24832

LLMs use sparse attention for speed but lose accuracy due to a distributional shift in outputs. This paper proposes Delta Attention, a simple correction method that realigns sparse outputs with full attention, recovering significant performance. Methods 🔧: → Delta Attention https://x.com/rohanpaul_ai/status/1929491684326244470

therapy practice template built with @v0 @inkko44 prompts have become my study guides https://x.com/heisenbit/status/1926487424571490473

“Beyond the 80/20 Rule High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning” https://x.com/i/web/status/1929900050638852479

RL 2018-20 was such a mess. Is RL-tuning 2024-26 headed the same way?”” / X https://x.com/giffmana/status/1928314882761334871

at this point i think it might be best to just not read any LLM RL papers”” / X https://x.com/vikhyatk/status/1928268671979565330

CVPR 2025 starts in less than few weeks; I’m working on a list of must-see CVPR papers / projects any important papers I should add? link: https://x.com/i/web/status/1929572784599859220

Effective AI use requires conscious control of the context an AI has access to in each conversation Universal memory, custom system prompts, web access, projects, access to your emails, connections to document storage – all have their place, but also can result in worse outcomes”” / X https://x.com/emollick/status/1928865675423928637

Good post from @balajis on the “”verification gap””. You could see it as there being two modes in creation. Borrowing GAN terminology: 1) generation and 2) discrimination. e.g. painting – you make a brush stroke (1) and then you look for a while to see if you improved the”” / X https://x.com/karpathy/status/1930305209747812559

In inference one usually gets either high throughput or low latency, but not both – enter shift parallelism which automatically adapts for the best performance!”” / X https://x.com/StasBekman/status/1928571964647682400

Nice empirical approach to critical batch size by AI2. The recipe in practice is quite simple: 1) Take your latest checkpoint. 2) Double the batch size, and (with learning rate scaling down) see if the loss recovers after 2B tokens. 3) If yes, you can double your batch size and https://x.com/i/web/status/1930081408657051745

Tokasaurus: An LLM Inference Engine for High-Throughput Workloads | Scaling Intelligence Lab at Stanford University https://scalingintelligence.stanford.edu/blogs/tokasaurus/

We’ve been thinking about what the “”ideal”” architecture should look like in the era where inference is driving AI progress. GTA & GLA are steps in this direction: attention variants tailored for inference: high arithmetic intensity (make GPUs go brr even during decoding), easy to”” / X https://x.com/tri_dao/status/1928170648863473892

what people miss about memory creation for AI entities is that its not just remembering facts, its performing analysis of the facts learned what does this information *mean* for its future behavior. e.g. how should the AI change what it does based on what it has learned. this”” / X https://x.com/i/web/status/1930053753807483216

LLMs struggle with real-world software engineering tasks like fixing GitHub issues, especially smaller models, because correct solutions are rare in their initial outputs. Evolutionary Test-Time Scaling refines patch generations iteratively. This Paper finds correct solutions https://x.com/rohanpaul_ai/status/1929812296848826726

Software development is more than just coding. Introducing Droids — the world’s first software development agents. 🤖 Starting today, Droids are available for general access. Factory integrates with your entire engineering system (GitHub, Slack, Linear, Notion, Sentry) and https://x.com/FactoryAI/status/1927754706014630357

Impromptu VLA https://impromptu-vla.c7w.tech/

Introducing LisanBench LisanBench is a simple, scalable, and precise benchmark designed to evaluate large language models on knowledge, forward-planning, constraint adherence, memory and attention, and long context reasoning and “”stamina””. “”I see possible futures, all at once. https://x.com/scaling01/status/1928510435164037342

Models can already tell when you are grading them. 😯 Your evaluation prompt has a scent; top LLMs smell it fast. Frontier language models can already sense when they are being tested. A new 1 000-item benchmark shows top systems spot evaluation prompts almost as well as https://x.com/rohanpaul_ai/status/1930579137581723905

This paper introduces EfficientLLM, the first large-scale benchmark evaluating efficiency techniques across the LLM lifecycle with fine-grained metrics. Methods 🔧: → The benchmark evaluates efficient attention variants, sparse Mixture-of-Experts, and attention-free https://x.com/rohanpaul_ai/status/1929522638403297582

Our experiments demonstrate that the Darwin Gödel Machine can continuously self-improve by modifying its own codebase. On SWE-bench, DGM automatically improved its performance from 20% to 50%. The figure here shows the performance progress over iterations, and also a summary of https://x.com/SakanaAILabs/status/1928447873362153669

[2506.03857] Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation https://arxiv.org/abs/2506.03857

#NLProc and LLMs: Ready for some summer learning? The 2024 version of cs224n is out, with new content on pre-training, post-training, benchmarking, reasoning, agents, and more https://x.com/i/web/status/1929557373213118869

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training “”It has grown to become the largest open dataset for pre-training Large Language Models at about 2 trillion tokens and the only one in its size range available in multiple languages”” https://x.com/i/web/status/1929751110723805525

Seems like no one saw this either, scraping arxiv manually seems to be the way. Pretty cool paper on rl for creative writing on Qwen3 32B base, and most interestingly it’s one author from the Star Writing Team (haven’t heard of them). They seem to have access to the 32B base tho https://x.com/i/web/status/1929996614883783170

Large language models struggle with theorem proving using formal proof systems that poorly align with their natural language strengths. The DeepTheorem framework uses natural language and reinforcement learning to improve this capability. Methods 🔧: → DeepTheorem dataset https://x.com/rohanpaul_ai/status/1929476333090029879

X changes its terms to bar training of AI models using its content | TechCrunch https://techcrunch.com/2025/06/05/x-changes-its-terms-to-bar-training-of-ai-models-using-its-content/

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading