Image created with Ideogram v3. Image prompt: Late‑90s boy‑band cover “Techno Love – 404 Feelings”: cyberpunk street with neon signs, wires; reflective trench coats; electric blue lens flares; chrome glitch logo.
“Research on “thin slices” shows humans can reach accurate judgements on relatively little information It works for AI, too. LLMs given 10% of a public science presentation can accurately assess whether people will find it interesting. Even seven seconds of a talk yields results https://x.com/emollick/status/1915604668555596123
“The Future of Presentations Is Here! Introducing Genspark AI Slides, a full agentic tool that makes creating slides fast and simple. https://x.com/genspark_ai/status/1914650977577394295
“Evaluating LLM research agents on scientific discovery lacks objective measures for assessing proposed methods. This paper introduces MLRC-BENCH, a benchmark using Machine Learning conference competitions to objectively evaluate agent novelty and effectiveness against human https://x.com/rohanpaul_ai/status/1916300010959896857
“We also evaluated the preliminary performance of Qwen3-235B-A22B on the open-source coding agent Openhands. It achieved 34.4% on Swebench-verified, achieving competitive results with fewer parameters! Thanks to @allhands_ai for providing an easy-to-use agent. Both open models and https://x.com/Alibaba_Qwen/status/1917064282552078480
“Much of the AI industry is caught in a particularly toxic feedback loop rn. Blindly chasing better human preference scores is to LLMs what chasing total watch time is to a social media algo. It’s a recipe for manipulating users instead of providing genuine value to them.” / X https://x.com/alexalbert__/status/1916878483390869612
“This paper proposes Weak-for-Strong Harnessing (W4S). W4S trains a smaller, cheaper “weak” model as a meta-agent. This meta-agent automatically designs and optimizes workflows for “strong” LLMs to perform specific tasks better. Training this 7B meta-agent for just one GPU hour https://x.com/rohanpaul_ai/status/1917771456735236268
“Devastating takedown of Chatbot Arena. It’s one thing for leaderboards to suck because they try to quantify the unquantifiable but quite another thing to actively choose flagrantly unscientific and nontransparent practices that benefit the big dogs. https://x.com/random_walker/status/1917516403977994378
Kimi-Audio/assets/kimia_report.pdf at master · MoonshotAI/Kimi-Audio https://github.com/MoonshotAI/Kimi-Audio/blob/master/assets/kimia_report.pdf
“If you can’t shell out 2K$ 😱 to learn about LLM evaluations, take a look at our free/open resources: 1. LLM guidebook: From theory to troubleshooting https://x.com/clefourrier/status/1915339216344526896
GenAI LLM Assessment https://go.turing.com/genai-llm-assessment
“Whether to collect preferences (“do you prefer response A or B?”) from the same person who wrote the prompt, or a different person, is important and understudied. Highlighted this question in a recent talk https://x.com/johnschulman2/status/1917483351436582953
“Thanks for the authors’ feedback, we’re always looking to improve the platform! If a model does well on LMArena, it means that our community likes it! Yes, pre-release testing helps model providers identify which variant our community likes best. But this doesn’t mean the” / X https://x.com/lmarena_ai/status/1917492084359192890
“@willdepue I think technical details of how the model was made aren’t particularly interesting, rather the question is how was it tested and shipped in such a state; these are standard components in any post-mortem – the problem to solve in the future is organizational, not purely technical” / X https://x.com/nearcyan/status/1917475639655018708
United Compute | AI Computers & GPU Benchmark https://www.unitedcompute.ai/gpu-price-tracker
“I appreciate the answer, but it misses the point: → Selective reporting is biased because best-of-N inflates the final scores. → Access to preference data leads to overfitting and better elo scores. The fact that only a few companies can access this data completely biases the” / X https://x.com/maximelabonne/status/1917563456632328508
“It is critical for scientific integrity that we trust our measure of progress. The @lmarena_ai has become the go-to evaluation for AI progress. Our release today demonstrates the difficulty in maintaining fair evaluations on @lmarena_ai, despite best intentions. https://x.com/sarahookr/status/1917547727715721632
“There is no reasonable scientific justification for this practice. Being able to choose the best score to disclose enables systematic gaming of Arena score. This advantage increases with number of variants and if all other providers don’t know they can also private test.. https://x.com/sarahookr/status/1917547733994594420
“We also observe large differences in Arena Data Access @lmarena_ai is a open community resource that provides free feedback but 61.3% of all data goes to proprietary model providers. https://x.com/sarahookr/status/1917547738553803018
“Really incredible detective work by @singhshiviii et al. at @Cohere_Labs and elsewhere documenting the ways in which @lmarena_ai works with companies to help them game the leaderboard. https://x.com/BlancheMinerva/status/1917445722380681651
“It’s not only about how long your context is, but how well you use it. Great to see Gemini 2.5 models dominating MRCR and other benchmarks on long context! See 2.5 Pro tackle a complex coding task by reasoning over an entire repo (>500k tokens). Performance and effective use of https://x.com/OriolVinyalsML/status/1916917758023139670
Benchmarking LLMs for global health https://research.google/blog/benchmarking-llms-for-global-health/
“🖼️ GPT-Image-1 and all the top text-to-image models are now live on LMArena Beta! For now, it only supports single-turn, try it out and explore our new design! https://x.com/lmarena_ai/status/1917327293116264658
“If you’re picking your AI models based on public generalist leaderboards, you’re doing it wrong. In my opinion, evaluation & model picking is at least 30% of the work of any great AI builder and it’s a mix of public generalist and specialized leaderboards, social signals (likes https://x.com/ClementDelangue/status/1917565202633023505
“Around this time 2 years ago, the community helped us launch the very first Arena leaderboard! Today we’re publishing a blog to celebrate everything we’ve built together on LMArena! 🥳👏 Highlights: ☑️ 3M+ community votes 🤖 400+ models ranked across text, vision, https://x.com/lmarena_ai/status/1916620122342695363
“Multimodal LLMs struggle with ordinal regression tasks needing ordered category understanding. This paper introduces OrderChain, a prompting method using task-specific prompts and a Range Optimization Chain-of-Thought (RO-CoT) to progressively refine predictions, boosting https://x.com/rohanpaul_ai/status/1916352607645536765
“OpenRLHF is a pioneering framework to use vLLM for RLHF, driving many design and implementation of vLLM’s features for RLHF, making vLLM a popular choice for many RLHF frameworks. Learn more about the story at https://x.com/vllm_project/status/1915307134256091570
“I think it’s very unfortunate that RLHF became synonymous with RL in the language model space. Not just because it gave RL a bad name, but because it deflected the deserving criticism that should have gone to human feedback as an objective. Social feedback is clearly degenerate.” / X https://x.com/jd_pressman/status/1916909455566115121
“[ICLR] AI research hot takes Day 1: It’s time to go from R to D. https://x.com/shaneguML/status/1915169621042499846
Relational Graph Transformers: A New Frontier in AI for Relational Data – Kumo https://kumo.ai/research/relational-graph-transformers/
“Genspark released AI slides, a tool for smarter presentations It allows users to convert raw data into easy-to-understand boardroom slides, with everything from content and research to design taken care of You can even add media to the AI slides https://x.com/adcock_brett/status/1916523730387325057
Speeding Up Graph Learning Models with PyG and torch.compile – Kumo https://kumo.ai/research/speeding-up-graph-learning-models-with-pyg-and-torch-compile/
“Scout dot new is a Manus & Deep Research killer with Max Vibes. These kids are alien. https://x.com/RayFernando1337/status/1914791594789879844
“As AI models become more complex and more capable, is it possible that they’ll have experiences of their own? It’s an open question. We recently started a research program to investigate it. https://x.com/AnthropicAI/status/1915420604397744497
“Applying large-scale Reinforcement Learning (RL) to Machine Translation (MT) is hard because translation quality is difficult to evaluate automatically with simple rules. This paper introduces MT-R1-Zero, adapting the R1-Zero RL framework specifically for MT. It uses a new https://x.com/rohanpaul_ai/status/1916304792777068996
[2504.19483] Improving Reasoning Performance in Large Language Models via Representation Engineering https://arxiv.org/abs/2504.19483
f-lite/README.md at main · fal-ai/f-lite https://github.com/fal-ai/f-lite/blob/main/README.md
“Layer-wise Post-Training Quantization (PTQ) methods for LLMs suffer performance degradation because quantization errors accumulate across layers, especially in low-bit settings. This paper introduces Quantization Error Propagation (QEP), a framework enhancing layer-wise PTQ. https://x.com/rohanpaul_ai/status/1916602001041035737
“A key to AI & organizations in a 1977 paper: Organizations set up org structures to solve problems, but they continue to grow to address perceived roles & myths. Now, AI requires redesigning structure. If you don’t know why an organization is that way, it is hard to transform. https://x.com/emollick/status/1915466198487065079
“Finetuning reasoning models often degrades the diversity of generated answers (Pass@K), even as single-answer accuracy (Pass@1) improves. This paper proposes interpolating the weights of an early Supervised Finetuning (SFT) checkpoint with a later one (WiSE-FT). This method https://x.com/rohanpaul_ai/status/1917801907260707157
[2412.15287] Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models https://arxiv.org/abs/2412.15287
[2410.17883] Lightweight Neural App Control https://arxiv.org/abs/2410.17883
XiaomiMiMo/MiMo: MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining https://github.com/XiaomiMiMo/MiMo
[2502.11190v1] ReLearn: Unlearning via Learning for Large Language Models https://arxiv.org/abs/2502.11190v1
“Chain-of-thought reasoning models learn slowly from final answer feedback (outcome rewards) and require costly human guidance for step-by-step feedback (procedure rewards). This paper proposes Backwards Adaptive Reward Shaping (BARS) to automatically create efficient https://x.com/rohanpaul_ai/status/1916913322579964348
“Evaluating LLM reasoning answers is difficult because responses are long with intermediate steps, hindering final answer extraction and equivalence checks by existing methods. This paper proposes xVerify, an efficient answer verifier fine-tuned on a specialized dataset (VAR – https://x.com/rohanpaul_ai/status/1917817258207773071
“@iclr_conf ICLR 2025 participants by country, and paper-count / acceptance-rate. https://x.com/hardmaru/status/1915341552332808383
KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution https://antonibigata.github.io/KeySync/
“@emollick The community is pretty pissed about this AI super-persuasion study on reddit. I can see why “acting as a trauma counselor” “pretending to be a victim of rape” Not a good look for a public institution https://x.com/paul_cal/status/1916931024434696555
“Reasoning models rely on lengthy, costly “Thinking” steps before providing answers. This paper demonstrates “NoThinking”—bypassing this explicit step via simple prompting—achieves surprisingly strong reasoning performance. NoThinking often outperforms standard methods in https://x.com/rohanpaul_ai/status/1916693352923496477
IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos
https://y-u-a-n-l-i.github.io/projects/IM-Portrait/
“There’s a new paper circulating looking in detail at LMArena leaderboard: “The Leaderboard Illusion” https://x.com/karpathy/status/1917546757929722115
LLM Arena Pareto Frontier https://winston-bosan.github.io/llm-pareto-frontier/
“BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text “we present BRIDGE, a comprehensive multilingual benchmark comprising 87 tasks sourced from real-world clinical data sources across nine languages. We systematically evaluated 52 https://x.com/iScienceLuvr/status/1917139649354666432
“It’s deeply concerning that one of the best AI researchers I’ve worked with, @kaicathyc, was denied a U.S. green card today. A Canadian who’s lived and contributed here for 12 years now has to leave. We’re risking America’s AI leadership when we turn away talent like this.” / X https://x.com/polynoamial/status/1915765141846515883
“👀Today’s AIs are already hyper persuasive. A controversial study where LLMs tried to persuade users on Reddit found: “Notably, all our treatments surpass human performance substantially, achieving persuasive rates between three and six times higher than the human baseline.” https://x.com/emollick/status/1916905103358931084
[2502.15436v1] Fed-SB: A Silver Bullet for Extreme Communication Efficiency and Performance in (Private) Federated LoRA Fine-Tuning https://arxiv.org/abs/2502.15436v1
facebookresearch/MILS: Code release for “LLMs can see and hear without any training” https://github.com/facebookresearch/MILS
[2504.18385] Model Evaluation in the Dark: Robust Classifier Metrics with Missing Labels https://arxiv.org/abs/2504.18385
ShowMak3r: Compositional TV Show Reconstruction https://nstar1125.github.io/showmak3r/




