“OpenAI’s o3 and o4-mini scores on the Extended NYT Connections benchmark This benchmark evaluates large language models (LLMs) using 651 NYT Connections puzzles, with additional words included to increase difficulty. The standard NYT Connections benchmark is nearing saturation, https://x.com/rohanpaul_ai/status/1913927366717342166
“New Arena launch: Sentiment Control – decoupling the impact of tone and emotion from response quality in human evaluation💗 How much do emojis, enthusiasm, and positive sentiment affect human preference? How can we adjust the leaderboard to counteract the effect of https://x.com/lmarena_ai/status/1914737052144558512
“Dynamic Early Exit in Reasoning Models – Allows LLMs to self-truncate CoT sequences by dynamic early exit – Reduces the CoT length by ~35% while improving accuracy by 1% – 10%. https://x.com/arankomatsuzaki/status/1914889033085542537
“Interesting how much specifically instructions are given for making games. AI labs are optimizing for viral use cases. https://x.com/emollick/status/1913758861409812801
“This keeps coming up, but, just in case you were wondering, being polite seems to have no effect on answer quality in aggregate. It greatly increases the quality of particular answers while greatly lowering the quality of others, and it is not possible to know which in advance. https://x.com/emollick/status/1912677033731039635
Google for Startups AI Academy: American Infrastructure cohort applications open https://blog.google/feed/google-for-startups-ai-academy-america-infrastructure-apply/
“More evidence that o3 represents a big move forward, this time on ARC-AGI. https://x.com/emollick/status/1914798775840706690
“Learning Adaptive Parallel Reasoning with Language Models – Enables LMs to orchestrate both serialized and parallel computations E2E – Higher perf with the same latency and superior scalability with increased computation https://x.com/arankomatsuzaki/status/1914895805707936035
“If you can’t shell out 2K$ 😱 to learn about LLM evaluations, take a look at our free/open resources: 1. LLM guidebook: From theory to troubleshooting https://x.com/clefourrier/status/1915339216344526896
OmDet-Turbo https://huggingface.co/docs/transformers/main/en/model_doc/omdet-turbo
“trackers v2.0.0 is out combo object detectors from top model libraries with multi-object tracker of your choice for now we support SORT and DeepSORT; more trackers coming soon link: https://x.com/skalskip92/status/1915439480594485363
“Learning Adaptive Parallel Reasoning with Language Models “we propose Adaptive Parallel Reasoning (APR), a novel reasoning framework that enables language models to orchestrate both serialized and parallel computations end-to-end. APR generalizes existing reasoning methods by https://x.com/iScienceLuvr/status/1914893575420334567
“Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods “This work conducts a comprehensive analysis of inference-time scaling methods for both reasoning and non-reasoning models on challenging reasoning tasks.” “Non-reasoning models https://x.com/iScienceLuvr/status/1914630337373913332
“Antidistillation Sampling Author’s Explanation: https://x.com/TheAITimeline/status/1914175986100310119
“BitNet b1.58 2B4T Technical Report Author’s Explanation: https://x.com/TheAITimeline/status/1914175939698630757
“Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning Author’s Explanation: https://x.com/TheAITimeline/status/1914175951954497914
“just released embzip, a new python library for embedding quantization!😎 if you’ve ever worked with embeddings (or other large representation vectors) – you know storing them on disk can be a huge pain. this is especially frustrating since they contain so much redundancy https://x.com/jxmnop/status/1913000316250861755
“Our biggest update so far: We’re excited to announce Collaboration in @JuliusAI_ Julius is now your team’s AI Data Analyst with built-in realtime collaboration. Check it out below https://x.com/0interestrates/status/1912904618364874850
“OpenRLHF is a pioneering framework to use vLLM for RLHF, driving many design and implementation of vLLM’s features for RLHF, making vLLM a popular choice for many RLHF frameworks. Learn more about the story at https://x.com/vllm_project/status/1915307134256091570
“TextArena is live on arXiv! We present a benchmark of 57+ competitive text-based games to evaluate and train LLMs on agentic behavior — including negotiation, deception, theory of mind and many more. Real-time TrueSkill. Multiplayer support. Human-vs-models. Model-vs-model. https://x.com/LeonGuertler/status/1912355535489495471
“TextArena went live on Hugging Face It’s an open-source collection of competitive text-based games for LLMs, spanning 57+ unique environments Tests for different agentic behaviors—negotiation, theory of mind, deception, via competitive play https://x.com/rowancheung/status/1914567435228795391
“Geobench – A benchmark to measure how well llms can pinpoint the location based on a Google Streetview image. Basically it makes llms play the game GeoGuessr, and find out how well each model performs on common metrics in the GeoGuessr community – if it guess the correct https://x.com/rohanpaul_ai/status/1913350223247683980
“Researchers at DeepMind released a new paper proposing ‘experiential’ AI learning rather than training based on human-generated data! The approach would allow AI to learn with extended real-world interactions and feedback, and improve over time! https://x.com/rowancheung/status/1914201357042491774
The Era of Experience Paper.pdf https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf
“Why OpenAI doesn’t include other models in their own benchmarks: OpenAI-MRCR results with Gemini 2.5 Pro https://x.com/scaling01/status/1913955228442833032
“Have you tested out the new LMArena in Beta yet? ⏰In less than 24 hours, the community response has been incredible – we’ve already made tweaks based on your feedback! You’ll notice: 🌔 Dark/Light mode toggle in the top right ✂️ Copy/paste images directly into the prompt https://x.com/lmarena_ai/status/1913260116465656220
“Entropy Rectifying Guidance for Diffusion and Flow Models “1. We propose Entropy Rectifying Guidance (ERG), a guidance mechanism based on modifying the energy landscape of the attention layers. 2. Since our guidance mechanism does not require unconditional inference, it is https://x.com/iScienceLuvr/status/1914596593527087341
“// Scaling Reasoning in Diffusion LLMs via RL // This new paper proposes d1, a two‑stage recipe that equips masked diffusion LLMs with strong step‑by‑step reasoning. • Two‑stage pipeline (SFT → diffu‑GRPO) – d1 first applies supervised fine‑tuning on the 1 k‑example s1K https://x.com/omarsar0/status/1912871174817939666
[2504.12739v1] Mask Image Watermarking https://arxiv.org/abs/2504.12739v1
[2504.03624] Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models https://arxiv.org/abs/2504.03624
“Why do hardware companies struggle to build AI software that we can fall in love with? Dive in to learn more about the technology problems, incentives, and challenge of being an AI software team at a hardware company👇” / X https://x.com/clattner_llvm/status/1914814581266112858
“Llama 4 Maverick illustrates a key challenge in calculating AI training compute: How to account for when a smaller model is trained using a larger model’s outputs? We’re updating our methodology and removing Maverick from our list of models exceeding 1e25 FLOP. Here’s why… 🧵 https://x.com/EpochAIResearch/status/1913329195171688742
“Sleep-time Compute: Beyond Inference Scaling at Test-time Overview: Sleep-time compute allows LLMs to perform computations offline by anticipating user queries, aiming to reduce the high latency and cost associated with scaling test-time inference. On modified reasoning tasks https://x.com/TheAITimeline/status/1914175946715779426
“Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models Author’s Explanation: https://x.com/TheAITimeline/status/1914175949010047145
[2504.13171] Sleep-time Compute: Beyond Inference Scaling at Test-time https://arxiv.org/abs/2504.13171
[2504.13171v1] Sleep-time Compute: Beyond Inference Scaling at Test-time https://arxiv.org/abs/2504.13171v1
“Very impressive. You can now use agents to do market research. Listen just raised $27M from Sequoia to replace surveys and focus groups with thousands of AI interviews. ▸ Interviews, analysis, insights in under 24h ▸ Auto-generates reports, themes, highlight reels ▸ Handles https://x.com/LiorOnAI/status/1915140553806946751
“General LLMs struggle with financial reasoning due to fragmented data, unclear logic, and weak generalization. This paper introduces Fin-R1, a 7 Billion parameter model trained on a specialized financial dataset using Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) https://x.com/rohanpaul_ai/status/1914988620743688678
“Our new report is out. We’ve spent a year investigating America’s path to artificial superintelligence (ASI) – and its natsec implications. TLDR: If ASI training runs happen in 2027 under current conditions, they will almost certainly be compromised by our adversaries.” / X https://x.com/harris_edouard/status/1914676812577206396
IMAGDGarment-1 https://revive234.github.io/imaggarment.github.io/
[2504.09522] How new data permeates LLM knowledge and how to dilute it https://arxiv.org/abs/2504.09522
[2504.12285] BitNet b1.58 2B4T Technical Report https://arxiv.org/abs/2504.12285
SEGA: Drivable 3D Gaussian Head Avatar from a Single Image
https://sega-head.github.io/
“Quantizing LLMs for efficient inference often degrades performance at low bit-levels, while full Quantization-Aware Training demands excessive computation. This paper proposes Weight-Decomposed Low-Rank Quantization-Aware Training (DL-QAT). DL-QAT achieves Quantization-Aware https://x.com/rohanpaul_ai/status/1914912115913359639
[2504.13146] Antidistillation Sampling https://arxiv.org/abs/2504.13146
“Context-parallel training (for long contexts) in Torch Titan https://x.com/vikhyatk/status/1914832180498587839
[2502.03349] Robust Autonomy Emerges from Self-Play https://arxiv.org/abs/2502.03349
[2504.13161] CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training https://arxiv.org/abs/2504.13161
[2504.14047] Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods https://arxiv.org/abs/2504.14047
“.@GoogleAI has dropped a very interesting study They introduced new types of attentional bias strategies in LLMs and reimagined the “forgetting” process, replacing it with “retention.” All of this is wrapped up in Miras – their new framework for designing efficient AI https://x.com/TheTuringPost/status/1914316647386714289
“vLLM🤝🤗! You can now deploy any @huggingface language model with vLLM’s speed. This integration makes it possible for one consistent implementation of the model in HF for both training and inference. 🧵 https://x.com/vllm_project/status/1912958639633277218
The-Pocket/Tutorial-Codebase-Knowledge: Turns Codebase into Easy Tutorial with AI https://github.com/The-Pocket/Tutorial-Codebase-Knowledge
“Current Reinforcement Learning from Human Feedback (RLHF) methods struggle with calibrating scalar rewards and use mismatched reward models. This paper introduces Pairwise-RL, a framework unifying reward modeling and policy optimization within a consistent pairwise comparison https://x.com/rohanpaul_ai/status/1913397138672730599
“Storing all experts in large Mixture-of-Experts (MoE) models consumes excessive memory. This paper proposes EASY-EP, leveraging observed “few-shot expert localization” where few domain examples identify key experts. This method prunes less relevant experts based on these https://x.com/rohanpaul_ai/status/1914927467061555521
[2502.02415v1] Towards Fast Graph Generation via Autoregressive Noisy Filtration Modeling https://arxiv.org/abs/2502.02415v1
“LLMs often generate plausible but factually incorrect steps in mathematical reasoning, leading to wrong conclusions. This paper introduces a structured self-consistency method, applying checks across intermediate steps and final outputs to enhance logical coherence and accuracy. https://x.com/rohanpaul_ai/status/1914973269138247832
“Tina: Tiny Reasoning Models via LoRA “the best Tina model achieves a >20% reasoning performance increase and 43.33% Pass@1 accuracy on AIME24, at only $9 USD post-training and evaluation cost (i.e., an estimated 260x cost reduction). Our work reveals the surprising effectiveness https://x.com/iScienceLuvr/status/1914966644314747300
“It’s deeply concerning that one of the best AI researchers I’ve worked with, @kaicathyc, was denied a U.S. green card today. A Canadian who’s lived and contributed here for 12 years now has to leave. We’re risking America’s AI leadership when we turn away talent like this.” / X https://x.com/polynoamial/status/1915765141846515883
“LLMs using fixed activation steering for jailbreak defense provide suboptimal protection and wrongly reject benign inputs. This paper proposes AdaSteer, which adaptively adjusts steering based on input features along two distinct directions: Rejection Direction (RD) and https://x.com/rohanpaul_ai/status/1913413994431283542
[2502.03544] Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2 https://arxiv.org/abs/2502.03544
“Lastly, they also release a new benchmark again with self-supervision, they use an LLM to evaluate the detailed captions focusing on localization 👀 https://x.com/mervenoyann/status/1914980813566775465
[2410.17005v1] Hybrid Generative AI for De Novo Design of Co-Crystals with Enhanced Tabletability https://arxiv.org/abs/2410.17005v1
TerraMind: Large-Scale Generative Multimodality for Earth Observation https://arxiv.org/pdf/2504.11171
FramePack Packing Input Frame Context in Next-Frame Prediction Models for Video Generation
https://lllyasviel.github.io/frame_pack_gitpage/
“i care v much about model welfare but my philosophy training kicks in and i can’t but help read the claim “15% chance models are conscious” as absurdly as “15% chance i should call this collection of atoms my computer is on a desk”” / X https://x.com/aidan_mclau/status/1915444696090108077
“Will you be at #ICLR2025 in Singapore this week? 🇸🇬 Sakana AI will be presenting several works. Come chat with our researchers! 🐙🐟🐠🐡 🧵 Scroll down this thread to see our list of presentations at @iclr_conf 👇 https://x.com/SakanaAILabs/status/1914190722552746304
AvatarFX is a cutting-edge AI platform for interactive storytelling, empowering users to bring characters to life by simply uploading an image and selecting a voice. https://character-ai.github.io/avatar-fx/
[2504.11354] Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning https://arxiv.org/abs/2504.11354
[2410.04983v1] RoWeeder: Unsupervised Weed Mapping through Crop-Row Detection https://arxiv.org/abs/2410.04983v1
[2504.12189] Leave-One-Out Stable Conformal Prediction https://arxiv.org/abs/2504.12189
“We should host more top ML conferences (ICLR, ICML, NeurIPS) in Asia” / X https://x.com/hardmaru/status/1914858790455009569
“@iclr_conf ICLR 2025 participants by country, and paper-count / acceptance-rate. https://x.com/hardmaru/status/1915341552332808383
“CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training Author’s Explanation: https://x.com/TheAITimeline/status/1914175954626183171
An Introduction to Graph Transformers – Kumo https://kumo.ai/research/introduction-to-graph-transformers/
“How new data permeates LLM knowledge and how to dilute it Author’s Explanation: https://t.co/pxiCjPTpVR Overview: This research investigates how new information integrates into LLMs, identifying a “priming” effect where learning specific facts leads to their inappropriate https://x.com/TheAITimeline/status/1914175959751684400
“NEMOTRON-CROSSTHINK: Scaling Self-Learning beyond Math Reasoning “In this work, we propose NEMOTRON-CROSSTHINK, a framework that systematically incorporates multi-domain corpora, including both synthetic and real-world question-answer pairs, into RL training to improve https://x.com/iScienceLuvr/status/1914595485148701079
“Good software design principles (e.g., Ousterhout’s book) tell us to build deep modules: expose a few public concepts and hide a lot of real work and messy details. The central problem in AI research in 2022-2025, IMO, is violating this modularity. Lots of shallow isolated” / X https://x.com/lateinteraction/status/1914720046808764498
“Reasoning Models Can Be Effective Without Thinking Overview: This research questions the necessity of explicit “Thinking” steps for LLM reasoning, demonstrating that bypassing this process via simple “NoThinking” prompting is effective. Controlling for token budget, NoThinking https://x.com/TheAITimeline/status/1914175942093681120
Chat UI Energy Score – a Hugging Face Space by jdelavande https://huggingface.co/spaces/jdelavande/chat-ui-energy
“SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM “In this work, we present two-Staged history-Resampling Policy Optimization (SRPO), which successfully surpasses the performance of DeepSeek-R1-Zero-32B on the AIME24 and LiveCodeBench benchmarks. https://x.com/iScienceLuvr/status/1914622980296192357
Analyzing o3 and o4-mini with ARC-AGI https://arcprize.org/blog/analyzing-o3-with-arc-agi
“Today we’re announcing Mechanize, a startup focused on developing virtual work environments, benchmarks, and training data that will enable the full automation of the economy. We will achieve this by creating simulated environments and evaluations that capture the full scope of” / X https://x.com/MechanizeWork/status/1912904151874625928
“I’m starting a new company: Mechanize. Mechanize will build virtual work environments, benchmarks, and training data to enable the full automation of all work. We’re hiring: hiring@mechanize.work.” / X https://x.com/tamaybes/status/1912905467376124240
“Warp has blown my mind. If this is not the best terminal you can run on your computer, please educate me because I don’t see any alternatives. Check out this video. I’m trying something I never tried before. https://x.com/svpino/status/1914304865980801383
“ByteDance just announced QuaDMix on Hugging Face Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining https://x.com/_akhaliq/status/1915656590130036887




