Image created with Ideogram V2. Image prompt: A vibrant spring meadow with exaggerated blooming flowers in bright colors. Hidden comically in the middle is a tech startup office complete with standing desks, servers, and programmers, all attempting to blend in with plants decorating their equipment. Circuit patterns run through the grass. Gadgets of all kinds hang from tree branches. Woodland animals wear VR headsets and use tiny laptops. A whiteboard with tech jargon stands among tall flowers. The whole scene is bathed in golden sunshine with lens flares. Vibrant colors and high detail. The word “TECH” integrated into the scene.

“Here are the top AI Papers of the Week (April 7 – 13): – NoProp – The AI Scientist V2 – Concise Reasoning via RL – Rethinking Reflection in Pre-Training – Efficient KG Reasoning for Small LLMs – Agentic Knowledgeable Self-awareness (bookmark for later) Read on for more:” / X https://x.com/dair_ai/status/1911444942523621550

IQ Test | Tracking AI https://trackingai.org/home

amazon-agi/SIFT-50M · Datasets at Hugging Face https://huggingface.co/datasets/amazon-agi/SIFT-50M

“🚨This week’s top AI/ML research papers: – Scaling Laws for Native Multimodal Models – Quantization Hurts Reasoning? – The AI Scientist -v2 – Parallel LLM Generation via Concurrent Attention – Gaussian Mixture Flow Matching Models – VAPO – Are Reasoning Models losing Critical https://x.com/TheAITimeline/status/1911633257952575492

“LLMs often lack reasoning accuracy and factual reliability for complex, knowledge-intensive questions like those in medicine. This paper introduces Retrieval-Augmented Reasoning Enhancement (RARE) to address this by adding dynamic information retrieval into the reasoning process https://x.com/rohanpaul_ai/status/1911348131721081311

“THIS IS BAD NEWS o3 is worse at replicating research papers than o1 https://x.com/scaling01/status/1912554822454116736

“How new data permeates LLM knowledge and how to dilute it – Learning a new fact can cause the model to inappropriately apply that knowledge in unrelated contexts – Alleviates effects by 50-95% while preserving the model’s ability to learn new information https://x.com/arankomatsuzaki/status/1911992299191669184

“One of the most interesting challenges with demo’ing AI live in front of a crowd is that sometimes the models, capabilities, or UX changes in the middle of a talk without any warning. I had that experience once while demoing “it can’t count the words in a sentence” & many others” / X https://x.com/emollick/status/1910371257322811556

“(4/) Fascinating early insights from our SAEs: – Effective steering must wait until after phrases like “Okay, so the user has asked a question about…”—not explicit tags like “<think>”—highlighting unintuitive internal markers of reasoning – Oversteering can paradoxically https://x.com/GoodfireAI/status/1912217319537099195

“Concise Reasoning via Reinforcement Learning Author’s Explanation: https://x.com/TheAITimeline/status/1911633319910928756

[2504.07807v1] Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models https://arxiv.org/abs/2504.07807v1

“Turns out you can use reasoning models to enhance non-reasoning models. Researchers explore how to distill reasoning-intensive outputs (answers and explanations) from top-tier LLMs into more lightweight models that don’t explicitly reason step by step. By fine-tuning smaller https://x.com/omarsar0/status/1912149669897187579

“all this drama over a paper claiming synthetic data causes model collapse meanwhile in the real world synth data pipelines are going brrr” / X https://x.com/vikhyatk/status/1911684113628889350

“Reinforcement Learning (RL) is quickly becoming the most important skill for AI researchers. Here are the best resources for learning RL for LLMs… TL;DR: RL is more important now than it has ever been, but (probably due to its complexity) there aren’t a ton of great resources https://x.com/cwolferesearch/status/1912566886509817965

“Scaling is incredibly hard and demanding and leaves very little room for error in every little part of the training stack But once it works, it’s beautiful to see it https://x.com/MillionInt/status/1912568397419954642

“Integrating LLMs with recommendation systems often involves disconnected optimization, preventing LLMs from directly improving recommendation quality. This paper introduces REC-R1, a framework using reinforcement learning (RL) to directly optimize LLMs. It uses feedback signals https://x.com/rohanpaul_ai/status/1911363482894975202

“Selecting the best SQL query from multiple LLM generations is hard, as correct queries can differ structurally but yield identical results. This paper proposes executing candidate SQL queries. It then compares their output data tables to pick the most semantically consistent https://x.com/rohanpaul_ai/status/1911302078548296103

[2504.00050] JudgeLRM: Large Reasoning Models as a Judge https://arxiv.org/abs/2504.00050

INTELLECT-2: The First Globally Distributed Reinforcement Learning Training of a 32B Parameter Model https://www.primeintellect.ai/blog/intellect-2

“Large language model (LLM) training suffers from gradient instability and loss spikes, and fixed gradient clipping methods fail to adapt dynamically. This paper proposes ZClip, an adaptive algorithm that adjusts the clipping threshold based on the gradient norm’s recent https://x.com/rohanpaul_ai/status/1911605578323091682

[2504.07934v1] SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement https://arxiv.org/abs/2504.07934v1

“An untested strategy for AI labs: set up a clear release plan for the next 12 months, explaining goals for future models in ways that don’t involve Delphic riddles, give expected release dates (including caveats that training involves some luck and dates may change), follow that.” / X https://x.com/emollick/status/1911962140262416753

SHeaP: Self-Supervised Head Geometry Predictor Learned via 2D Gaussians https://nlml.github.io/sheap/

[2504.07439v1] LLM4Ranking: An Easy-to-use Framework of Utilizing Large Language Models for Document Reranking https://arxiv.org/abs/2504.07439v1

“my base case for LLMs is that over the next few years they’ll evolve into hyper-specialized autistic superintelligences that excel in domains where verification is straightforward that is simply because we are compute bottlenecked we can’t create synthetic data for everything” / X https://x.com/scaling01/status/1911187189548933143

DataDecide: How to predict best pretraining data with small experiments | Ai2 https://allenai.org/blog/datadecide

[2504.07963] PixelFlow: Pixel-Space Generative Models with Flow https://arxiv.org/abs/2504.07963

“Rethinking Reflection in Pre-Training This is one of the more interesting papers I read this week. It argues that reflection emerges during pre-training. It introduces adversarial reasoning tasks to show that self-reflection and correction capabilities steadily improve as https://x.com/omarsar0/status/1911442761238340095

“The weird thing about the fact that AI models are more grown & gardened than built, and that training runs are expensive and have high variance, is that there is a lot of odd variations in how AI models turn out, even within labs. Some just seem to be unexpectedly good or bad.” / X https://x.com/emollick/status/1911497988910063620

DNF-Avatar: Distilling Neural Fields for Real-time Animatable Avatar Relighting https://jzr99.github.io/DNF-Avatar/

Introducing HELMET: Holistically Evaluating Long-context Language Models https://huggingface.co/blog/helmet

Introducing the Search Arena: Evaluating Search-Enabled AI | LM Arena https://blog.lmarena.ai/blog/2025/search-arena/

[2504.08169] On the Practice of Deep Hierarchical Ensemble Network for Ad Conversion Rate Prediction https://arxiv.org/abs/2504.08169

Popular AI-Ranking Website Chatbot Arena Is Becoming a Real Company – Bloomberg https://www.bloomberg.com/news/articles/2025-04-17/popular-ai-ranking-website-chatbot-arena-is-becoming-a-real-company?embedded-checkout=true

[2410.16144] 1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs https://arxiv.org/abs/2410.16144

“Another cool paper from @GoogleDeepMind Using LLMs for ranking, evaluation (judging), and content creation in search systems creates risks of biased evaluation outcomes. This paper empirically studies these risks. It finds LLM judges favor LLM rankers and struggle to https://x.com/rohanpaul_ai/status/1911332780522483858

“Super cool paper from @GoogleDeepMind Real-world queries for LLMs often lack necessary information for reasoning tasks. This paper tackles this by framing underspecification as a Constraint Satisfaction Problem where one variable is missing. It introduces QuestBench, a https://x.com/rohanpaul_ai/status/1911317429470318909

“vLLM🤝🤗! You can now deploy any @huggingface language model with vLLM’s speed. This integration makes it possible for one consistent implementation of the model in HF for both training and inference. 🧵 https://x.com/vllm_project/status/1912958639633277218

Cobra https://zhuang2002.github.io/Cobra/

InstantCharacter : Personalize Any Characters with a Scalable Diffusion Transformer Framework https://instantcharacter.github.io/

FireEdit: Fine-grained Instruction-based Image Editing via Region-aware
Vision Language Model https://arxiv.org/pdf/2503.19839

“We challenged Reachy 2 to sort healthy & unhealthy foods—no AI training, just real-time object detection! Built using pollen_vision & our self-development kit. simple version : https://x.com/pollenrobotics/status/1897286630249259310

“Good benchmark to test the scientific equation discovery capabilities of LLMs. “Extensive experiments on state-of-the-art methods reveal performance peaks at 31%.” https://x.com/omarsar0/status/1912144486970630512

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading