“Dedicated Feedback and Edit Models Empower Inference-Time Scaling for Open-Ended General-Domain Tasks https://x.com/_akhaliq/status/1897847351366029704
“@DaveShapi wait are you saying that 4.5 is overfit to benchmarks? or that every other model is?” / X https://x.com/aidan_mclau/status/1896610660840259866
“Congratulations to @togethercompute , @FireworksAI_HQ , @hyperbolic_labs and @GroqInc for being the first to launch serverless endpoints! Live performance benchmarks available now on Artificial Analysis. https://x.com/ArtificialAnlys/status/1897701018231881772
“More documentation here: https://x.com/awnihannun/status/1896371097240834188
“🚨Our Generative AI Lab at Wharton is releasing its first Prompt Engineering Report, empirically testing prompting approaches. This time we find: 1) Prompting “tricks” like saying “please” do not help consistently or predictably 2) How you measure against benchmarks matters a lot https://x.com/emollick/status/1897046511101599797
“SynaLinks: https://x.com/fchollet/status/1896663405462978611
“Why the hell is there no Oauth gateway that allows users to use their own LLM API tokens by just signing in? Like avoid friction of copy pasting keys” / X https://x.com/HamelHusain/status/1897486751696085093
“One of the biggest workflows we see people doing in langsmith is using it “to transform user feedback (regressions) into evals” Why having good observability IS evals tooling” / X https://x.com/hwchase17/status/1896425160414314695
“Why do frontier labs celebrate stuff like this? What fun is there in playing games in which you claim your thing is clearly best because of a +0.6% margin on whatever kind of scale? https://x.com/lateinteraction/status/1896682075585220737
“Yes there’s an evals crisis, but evaluating *models* is not even the right question most of the time LangProBe from @ShangyinT @LakshyAAAgrawal @arnav_thebigman @LaiLiheng112 @michaelryan207 et al begins to ask what complete *AI systems* we should build and under what settings” / X https://x.com/lateinteraction/status/1896632714054467763
“Many works discussed in the interview: (1) Large Language Models As Evolution Strategies https://x.com/SakanaAILabs/status/1896483381183205420
“@karpathy Ok, according to the system card (https://t.co/NtugtBMkTD), it’s trained “using new supervision techniques” https://x.com/rasbt/status/1895511885950357888
“Improvements in DiLoCo / era of RL means our focus will not shift from VRAM /Flops per gpu per dollar, but ultimately flops / watt. Then chipmakers will focus on making GPUs more efficient per machine, and we might have a different winner. With 1000 caveats, but something to” / X https://x.com/cloneofsimo/status/1897557416117686772
“by a staggering amount. This will change stuff.. https://x.com/cloneofsimo/status/1897757653436432494
“abs: https://x.com/iScienceLuvr/status/1895331436171010152
“abs: https://x.com/iScienceLuvr/status/1895334574160519522
““residual stream saturation” is an underdiscussed problem in all long-context talks” / X https://x.com/teortaxesTex/status/1898028373542150325
“Chain of Draft (CoD) vs. Chain of Thought (CoT) CoD encourages models to generate short but meaningful reasoning steps instead of the long explanations used in CoT. By using only 7.6% of the words, CoD helps to: • Reduce costs • Make the model faster, maintaining (or even https://x.com/TheTuringPost/status/1895398797808959963
“Today we announced that development of Orion Browser for Linux, a WebKit based browser, has officially started. This will expand Kagi’s eco-system of user-centric products to a platform we all love! https://x.com/vladquant/status/1897797849091653778
“No I don’t care if it draws a triangle using SVG really well. Seriously that’s what you show for the *largest* language model as capabilities?” / X https://x.com/abacaj/status/1895520453520970054
“not to be a relentless RL shill (*cough* guilty) but this is why we need RL- supervised learning is fundamentally the wrong paradigm, although it will take us to the border of the promised land” / X https://x.com/finbarrtimbers/status/1897779253665841370
“if you’ve only ever worked in software you have no idea how bad the state of inventory tracking in much of the economy is. generalize this a bit and you’ll quickly understand where the demand for the next trillion tokens per day is going to come from.” / X https://x.com/gallabytes/status/1896434971386236951
uv + Ray: Pain-Free Python Dependencies in Clusters | Anyscale https://www.anyscale.com/blog/uv-ray-pain-free-python-dependencies-in-clusters
“there’s an easy fix to all of these problems: more data, more parameters, more flops scale is all you need” / X https://x.com/vikhyatk/status/1897858708509802521
AK on X: “LongRoPE2 Near-Lossless LLM Context Window Scaling https://t.co/02lgp26FFT” / X
https://x.com/_akhaliq/status/1895338193303806287
AK on X: “R1-T1 Fully Incentivizing Translation Capability in LLMs via Reasoning Learning https://t.co/7APHYo6W3A” / X
https://x.com/_akhaliq/status/1895339378257600744
(1) AK on X: “FlexiDiT Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute https://t.co/ZDMx9iONQA” / X
https://x.com/_akhaliq/status/1895340908134178897
(1) Aran Komatsuzaki on X: “Nvidia presents: Token-Efficient Long Video Understanding for Multimodal LLMs SotA results across various long video understanding benchmarks while reducing the computation costs by up to 8x and the decoding latency by 2.4-2.9x for the fixed numbers of input frames https://t.co/0M6MQliA11” / X
https://x.com/arankomatsuzaki/status/1897854511814770733
(1) Tanishq Mathew Abraham, Ph.D. on X: “Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners Distilling Llama-1B and -3B models with only 8 billion tokens into subquadratic models like Mamba to achieve better and faster scaling of inference-time compute with minimal performance loss. https://t.co/E3QXKUXkvj” / X
https://x.com/iScienceLuvr/status/1895331433897697689
(1) Tanishq Mathew Abraham, Ph.D. on X: “Self-Training Elicits Concise Reasoning in Large Language Models “Upon examination of the output distribution of current LLMs, we find evidence on their latent ability to reason more concisely, relative to their default behavior. To elicit this capability, we propose simple https://t.co/cNEtpWHIst” / X
https://x.com/iScienceLuvr/status/1895333239608508473
(1) elvis on X: “Test-Time Scaling on Chart Generation Applies test-time scaling and multi-agents on chart generation. The results are impressive! Read on for more: https://t.co/CxpJmAtelR” / X
https://x.com/omarsar0/status/1895528398820425741
(1) elvis on X: “Modality-tailored critiques enhance self-correction Separate visual and code critique mechanisms substantially boost the self-correction capability of VLMs. Shows a 5.16% improvement in accuracy when modality-specific feedback was employed. https://t.co/SrcNE73jCF” / X
https://x.com/omarsar0/status/1895528446882955439
(1) elvis on X: “METAL achieved superior performance improvements over SoTA methods. Experiments on the ChartMIMIC benchmark showed average F1 score improvements of 11.33% with open-source models (LLAMA 3.2-11B) and 5.2% with closed-source models (GPT-4o). https://t.co/OnlRt20fMC” / X
https://x.com/omarsar0/status/1895528463504982115
amphion/Emilia-Dataset · Datasets at Hugging Face https://huggingface.co/datasets/amphion/Emilia-Dataset
“Human: “Put the Ketchup away, where you think it belongs” (Ketchup was not present in the training set) https://x.com/adcock_brett/status/1897399420595134526
“What’s the vibe on new qwq – benchmark maxxing or good model?” / X https://x.com/abacaj/status/1897645343233241497
2503.01610 https://arxiv.org/pdf/2503.01610
BodyGen https://genesisorigin.github.io/
“Another great paper on reasoning LLM efficiency. This one focuses on the relationship between reasoning length and model performance using diverse compression instructions (e.g., ‘use 10 words or less’). These papers provide good tips for how to leverage reasoning LLMs. https://x.com/omarsar0/status/1896939453069074907
ACM A.M. Turing Award Honors Two Researchers Who Led the Development of Cornerstone AI Technology https://www.acm.org/media-center/2025/march/turing-award-2024
“🚨 New dataset release: 1.9M scanned pages transcribed using Pixtral. https://x.com/vikhyatk/status/1897951196079353970
“Many noise from recent 4.5 release, but two important take from vid imo: 1. “GPT4.5 was trained on multiple datacenters” Translate that to “diloco goes brr for largest LLM on the market”, bullish on async, low bandwidth training in 2025. 2. “We aggressively used low precision” / X https://x.com/cloneofsimo/status/1895319178116243763
Ai2 Case Study – Allen Institute for Artificial Intelligence https://www.cirrascale.com/ai2-case-study
VARGPT https://vargpt-1.github.io/
[2502.19962v1] ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence Learning https://arxiv.org/abs/2502.19962v1
DEMO³ https://adrialopezescoriza.github.io/demo3/
“We found that when under pressure, some AI systems lie more readily than others. We’re releasing MASK, a benchmark of 1,000+ scenarios to systematically measure AI honesty. @ai_risks @scale_AI https://x.com/DanHendrycks/status/1896972178387841140
“@karpathy apparently they didn’t use character-training” / X https://x.com/rasbt/status/1895502063154733239
NotaGen https://electricalexis.github.io/notagen-demo/
“A Deep Dive into Reasoning LLMs This is a really nice summary of the progress made in post-training and reasoning LLMs. Highly recommend this one! https://x.com/omarsar0/status/1896572276461703193
[2502.20321] UniTok: A Unified Tokenizer for Visual Generation and Understanding https://arxiv.org/abs/2502.20321
“What I see here is that non-thinking LLMs, pretrained mostly on natural data, have hit their practical limit. I genuinely do not believe that a $1T training run would much improve this. This ceiling is imposed by the nature of data, architecture and objective function.” / X https://x.com/teortaxesTex/status/1895496203401629933
[2502.18770v1] Reward Shaping to Mitigate Reward Hacking in RLHF https://arxiv.org/abs/2502.18770v1
[2503.01328] PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization https://arxiv.org/abs/2503.01328
“Checkout the project over at Github: https://x.com/ggerganov/status/1896592517686313170
Learning Pokémon With Reinforcement Learning | Pokémon RL https://drubinstein.github.io/pokerl/
“The training set we collected for the initial Helix unveil was ~20TB This is ~2x the size of a LLM training set assuming 1×10^13 tokens There is so much data in the real world, we haven’t even scratched the surface” / X https://x.com/adcock_brett/status/1896035080327610706
“An interesting paper on explaining generalization behavior & other observed phenomena in deep learning like benign overfitting, double descent, overparametrization with “soft inductive biases” The polynomial regression with order-dependent regularization experiments are” / X https://x.com/iScienceLuvr/status/1897619075364487338
[2502.20766] FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference https://arxiv.org/abs/2502.20766
[2502.21186] Scalable Decision-Making in Stochastic Environments through Learned Temporal Abstraction https://arxiv.org/abs/2502.21186
World’s first “Synthetic Biological Intelligence” runs on living human cells https://newatlas.com/brain/cortical-bioengineered-intelligence/
[2502.20957] Reward Dimension Reduction for Scalable Multi-Objective Reinforcement Learning https://arxiv.org/abs/2502.20957
“We are excited to introduce Mercury, the first commercial-grade diffusion large language model (dLLM)! dLLMs push the frontier of intelligence and speed with parallel, coarse-to-fine text generation. https://x.com/InceptionAILabs/status/1894847919624462794
“This is interesting as a first large diffusion-based LLM. Most of the LLMs you’ve been seeing are ~clones as far as the core modeling approach goes. They’re all trained “autoregressively”, i.e. predicting tokens from left to right. Diffusion is different – it doesn’t go left to” / X https://x.com/karpathy/status/1894923254864978091
“Shipping tools for quality data curation is just as important as shipping quality models 🚚 From the diffusers bandwagon, we’re shipping a smol shot categorizer, capable of producing accurate shot info … that too fast (<1s on CPU) 🔥 Use it for video data curation amongst https://x.com/RisingSayak/status/1897590118736957442
“Was having trouble falling asleep. Checked my phone. CogView4 dropped. Tried to go to sleep. So anyway, now I am adding that to AI Toolkit at 2am.” / X https://x.com/ostrisai/status/1896845726539513930
“LADDER is a framework enabling LLMs to recursively generate and solve progressively simpler variants of complex problems—boosting math integration accuracy. Key insights include: • Autonomous difficulty-driven learning – LADDER lets models create easier problem variants of an https://x.com/dair_ai/status/1898037429434826795




