Image created with gemini-2.5-flash-image with claude-sonnet-4-5-20250929. Image prompt: A minimalist two-tiered birthday cake made entirely of illuminated circuit boards and electronic components, with glowing blue and red traces forming the number 2, two lit candles on top casting warm light, photographed with crisp studio lighting against a deep blue gradient background with high contrast and sharp reflections on a glossy black surface.

A SOTA moment to me: Kimi’s OK Computer generate this website in just one shot > It designed a very beautiful site, all images were AI-generated, and when you click, the sidebar expands. > Inside the sidebar, there’s a handwritten letter, it really feels like a website made by a https://x.com/crystalsssup/status/1971133240619757794

Say hi to OK Computer, Kimi’s agent mode 🤖🎸 Your AI product & engineering team, all in one. ✨ From chat → multi-page websites, mobile first designs, editable slides ✨ From up to 1 million rows of data → interactive dashboards ✨ Agency: self-scopes, surveys & designs ✨ https://x.com/Kimi_Moonshot/status/1971078467560276160

Not to take away from Grok 4 Fast (which seems like a very good model) or from Artificial Analysis (one of the few organizations doing independent benchmarking), but the Intelligence Index is an average of pretty saturated benchmarks (aside from HLE), we really need better ones.”” / X https://x.com/emollick/status/1969270709361733942

Veo is a more general reasoner than you might think. Check out this super cool paper on “”Video models are zero-shot learners and reasoners”” from my colleagues at @GoogleDeepMind. https://x.com/tkipf/status/1971063116734841248

I find it unimaginably based that the OAI Evals team keeps making benchmarks finding that Claude is better and publishing it anyway. they are 3 for 3 this year in acknowledging specifically how much Claude is better at tasks OAI care about. there is no sarcasm here folks. this https://x.com/swyx/status/1971404125553242253

A postmortem of three recent issues \ Anthropic https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues

GPT-5 is the best model for code quality out there 2 years ago, we created the world’s hardest software design quiz. Only 5 questions, multiple choice. Yet only about 3% of software engineers get them. The average score is somewhere between 2 and 3. Supposedly brilliant models https://x.com/jimmykoppel/status/1968683689421701413

Measuring the performance of our models on real-world tasks | OpenAI https://openai.com/index/gdpval/

We’ve released a large-scale study on how people are using ChatGPT. Consumer adoption has broadened beyond early-user groups, and lots of economic value is being created through both personal and professional use: https://x.com/gdb/status/1969953507215302836

💥 Announcing GDPval, a new eval that measures model performance on economically valuable, real-world tasks across 44 occupations.”” / X https://x.com/kevinweil/status/1971250647778635904

GDPval.pdf https://cdn.openai.com/pdf/d5eb7428-c4e9-4a33-bd86-86dd4bcf12ce/GDPval.pdf

Just released GDPval: an early step towards better methods for measuring and forecasting real-world model progress.”” / X https://x.com/gdb/status/1971301844585676930

opus 4.1 beats gpt-5-high on OAI’s own new GDP eval. nice of them to be transparent hehe https://x.com/dejavucoder/status/1971253593404735706

Today we’re introducing GDPval, a new evaluation that measures AI on real-world, economically valuable tasks. Evals ground progress in evidence instead of speculation and help track how AI improves at the kind of work that matters most. https://x.com/OpenAI/status/1971249374077518226

A research team at @OpenAI, where I am proud to be a board member, released an important new paper today. This paper looks at what might be thought of as task specific Turing Tests and shows that AI systems, even with limited guidance, perform many tasks — such as planning”” / X https://x.com/LHSummers/status/1971252567981146347

🚨 New Models Update! 🔥 Qwen3 coming in hot into the Arena with three different models: 🔹Qwen3-VL-235b-a22b-thinking for Text & Vision 🔹Qwen3-VL-235b-a22b-instruct for Text & Vision 🔹Qwen3-Max-2025-9-23 for Text Check out the thread to learn more about them and get https://x.com/arena/status/1970920636957831611

VraserX e/acc on X: “GPT-5 just passed what researchers call the “Gödel Test.” That means it’s not just solving textbook problems, it’s tackling open math conjectures that would normally take a skilled PhD student days to crack. In a new paper, GPT-5 was tested on 5 unsolved optimization https://t.co/4lGYKLrdrD” / X
https://x.com/VraserX/status/1970902050931159184

xAI has released Grok 4 Fast – breaking through our intelligence vs cost frontier by achieving Gemini 2.5 Pro level intelligence at a ~25X cheaper cost Intelligence: @xai shared with us pre-release access to Grok 4 Fast. In reasoning mode, the model scores an impressive 60 on https://x.com/artificialanlys/status/1969180023107305846

This is Ray3. The world’s first reasoning video model, and the first to generate studio-grade HDR. Now with an all-new Draft Mode for rapid iteration in creative workflows, and state of the art physics and consistency. Available now for free in Dream Machine. https://x.com/LumaLabsAI/status/1968684330034606372

Veo 3 = Zero-shot video reasoner • Trained on web-scale video, shows broad zero-shot skills (perception → physics → manipulation → reasoning) • New “Chain-of-Frames” reasoning = visual analogue of CoT • Big jump Veo2 → Veo3: edits, memory, symmetry, mazes, analogies • https://x.com/arankomatsuzaki/status/1971042970800701809

What used to take hours in After Effects now takes just ONE prompt. Nano Banana, Seedream 4, Wan 2.2, Runway Aleph et al are pioneering instruction-based editing — collapsing complex VFX pipelines into a single, implicit step. Here’s everything you need to know in 10 mins: https://x.com/bilawalsidhu/status/1970915228536947026

Introducing our Agentic Leaderboards. These new leaderboards test AI agents in real-world, high-complexity environments, setting a new standard for completing end-to-end digital tasks. https://x.com/scale_AI/status/1969416303015301128

Unlocking a Million Times More Data for AI | IFP https://ifp.org/unlocking-a-million-times-more-data-for-ai/

👏We are proud to share that in an internal blind test, where professionals evaluated pairwise comparisons using a win/loss ratio benchmark, the Kling AI 2.5 Turbo model significantly outperformed Seedance 1.0, Veo 3 Fast, and Seedance 1.0 Mini in both text-to-video and https://x.com/Kling_ai/status/1970832920085753893

Factory are one of the most formidable teams in the AI Coding space and I think the Droid concept strikes exactly the right balance between concerns of engineer replacement and the ill defined expectations of agents. while there is some way to go on SWEBench still, its p clear https://x.com/swyx/status/1971310686585356295

Many people think LLMs are non-deterministic. This is often not true! You just need 3 lines of code to make your LLM deterministic LLMs (as any PyTorch model) are non-deterministic only when they include certain operations or when using multiple GPUs Try the code yourself https://x.com/gabriberton/status/1968559505966350705

Kimi Infra team dropped K2 Vendor Verifier > You can visually see the difference in tool call accuracy across providers on OpenRouter. https://x.com/crystalsssup/status/1971158566343184511

LIMI: Less Is More for Agency • Argues agentic AI doesn’t need more data, just better data • 78 curated demos → 73.5% on AgencyBench (beats models trained on 10k samples) • Outperforms SOTA (Kimi-K2: 24.1%, DeepSeek: 11.9%, Qwen3: 27.5%, GLM-4.5: 45.1%) • Establishes Agency https://x.com/arankomatsuzaki/status/1970328242688246160

Cross-Agent Privilege Escalation: When Agents Free Each Other https://simonwillison.net/2025/Sep/24/cross-agent-privilege-escalation/#atom-everything

First set of @aisdk agents docs have shipped! – foundations – workflow patterns – loop control – `Agent` class https://x.com/nicoalbanese10/status/1968398677686301134

Introducing Parallel Thinking for Reka Research! Instead of one line of reasoning, we explore multiple paths in parallel, then resolve the best answer. Big accuracy gains on Research-Eval (+4.2) and SimpleQA (+3.5). Now live in the Reka Research API! https://x.com/RekaAILabs/status/1971241107322540194

Turso is an incredible technical feat. A Rust rewrite of sqlite, with an async-first architecture, incoming support for concurrent writes, vector search, and browser / wasm support out of the box. I think this has a very good chance of being a foundational piece of https://x.com/rauchg/status/1969515038512926823

Tool Calls Are Expensive And Finite https://www.reillywood.com/blog/tool-calls-are-expensive-and-finite/

AI agency breakthrough: Less data, more power New research with LIMI shows AI agents can achieve 73.5% on benchmarks, outperforming SOTA models by 50%+ using only 78 carefully chosen samples. The “”Agency Efficiency Principle”” is here. https://x.com/HuggingPapers/status/1970400645871185942

Code generators can write functions, but fail with full repositories. It’s because natural language is ill-suited for software structures. @Microsoft introduced the Repository Planning Graph (RPG), a blueprint that links abstract project goals to clear code structures: • https://x.com/TheTuringPost/status/1970068577509327197

You use the same prompt and the same model, but always get different results. Why LLMs are so unpredictable? This inconsistency is called nondeterminism. And it happens because of: – Messy math with approximations – Parallel computing – But the main thing is batching A new https://x.com/TheTuringPost/status/1968470771212103722

China’s Alibaba just dropped an opensource 30B agentic LLM that outperforms Claude 4 Sonnet, DeepSeek v3.1, Kimi k2 on a range of agentic search benchmarks. Only 3B parameters are activated per token. 100% open-source. https://x.com/unwind_ai_/status/1969053988143477186

New, very needed benchmark from @scale_AI: SWE-Bench Pro Includes: – Multi-file edits – 100+ lines changed on average – Complex dependencies across large codebases Current top model scores: – GPT-5: 23.3% – Claude Opus 4.1: 22.7% – Others drop further (<15%)”” / X https://x.com/alexandr_wang/status/1969805196462358919

Results so far No single model dominates: GPT-5 “high” reasoning leads on tough tasks but collapses on time-critical ones. Claude-4 Sonnet balances speed vs accuracy but at higher cost. Open-source models (like Kimi-K2) show promise in adaptability. Scaling curves plateau, https://x.com/omarsar0/status/1970147904087322661

@OpenAI Interesting: 1. Linear progress across OpenAI generations (GPT-4o, o3, GPT-5) 2. Claude Opus 4.1 is on top, nearing industry expert, much better than GPT-5 high. Thanks for acknowledging competitors. https://x.com/Yuchenj_UW/status/1971254164069212231

it’s quite incredible how bad Sonnet 4 is at long-context retrieval Grok-4 > GPT-5 ~ Gemini 2.5 Pro > Claude 4 Sonnet https://x.com/scaling01/status/1970661469667660100

Language Models that Think and Chat Better Proposes a simple RL recipe to improve small open models (eg, 8B) that rivals GPT-4o and Claude 3.7 Sonnet (thinking). Pay attention to this one, AI devs! Here are my notes: https://x.com/omarsar0/status/1971215698140819516

Lots of sympathy to the Anthropic team 🙏🙏🙏 https://x.com/cHHillee/status/1968536182284849459

Apple presents EpiCache Episodic KV Cache Management for Long Conversational Question Answering https://x.com/_akhaliq/status/1970475890501955834

Price analysis reveals trends in the Speech to Text market: Fireworks and Groq are the lowest-cost inference providers for Whisper Large v3, offering competitive access to OpenAI’s model. As word error rate decreases, pricing tends to increase, reflecting the performance-cost https://x.com/ArtificialAnlys/status/1971232403973943517

Clementine just dropped an incredibly useful guide to evaluations in 2025 ✨ The key insight: we’re transitioning from testing knowledge retention to measuring practical problem-solving ability. Her framework spans core capabilities, integrated assistant tasks, adaptive”” / X https://x.com/joelniklaus/status/1968596729852231813

Happy to start to see new, harder AI benchmarks coming into use. But they all need the following things: 1) Validated comparison: how does a human expert do? 2) Public & private tests 3) Care to eliminate question/answer errors 4) Reporting standard errors 5) Real-world validity https://x.com/emollick/status/1969938204879896622

I have been wondering if there is an underlying capability factor that all the many benchmarks for AI are measuring. It seems like the answer is yes. Overall correlation is good (median r ≈ 0.51) and there are distinct clusters (eg reasoning, code) with VERY high correlation. https://x.com/emollick/status/1969986020042010640

Introducing SEAL Showdown | Scale https://scale.com/blog/showdown

Scale AI launches Seal Showdown, an alternative to LMArena leaderboard | Mashable https://mashable.com/article/scale-ai-seal-showdown-benchmarking-leaderboard-lmarena

SWE-Bench Pro https://static.scale.com/uploads/654197dc94d34f66c0f5184e/SWEAP_Eval_Scale%20%289%29.pdf

Last week we found an issue with SWE-Bench, allowing agents to cheat by looking at future commits. Instead of celebrating the SWE-Bench Devs for quickly fixing the issue and being transparent, the HN crowd is dunking on them and drawing wildly inaccurate conclusions about”” / X https://x.com/TacoCohen/status/1966421688846778561

What are popular AI coding benchmarks actually measuring? – nilenso blog https://blog.nilenso.com/blog/2025/09/25/swe-benchmarks/

We’re launching the Artificial Analysis Word Error Rate Index (AA-WER), our new synthesis benchmark for Speech to Text model accuracy comprising of 3 challenging datasets AA-WER comprises three challenging datasets aligned with real-world use cases: AMI-SDM (multi-speaker https://x.com/ArtificialAnlys/status/1971232397921534141

EngDesign: Benchmarking LLMs toward Engineering AGI • 101 tasks across 9 domains (OS, circuits, robotics, etc.) • Simulation-based eval (SPICE, FEA, etc.), not just Q&A • Iterative refinement boosts pass rate (o3 → ~60%) • Error analysis: 111 failure modes identified https://x.com/arankomatsuzaki/status/1970326076271513805

In Defense of AI Evals, for Everyone https://www.sh-reya.com/blog/in-defense-ai-evals/

OSF | Quantifying Human-AI Synergy https://osf.io/preprints/psyarxiv/vbkmt_v1

Some useful findings: 1) Working with AI boosts the performance of people solving math, science & ethics questions 2) The biggest boost is for the hardest problems 3) High performers remain highest performing, but low performers gain more 4) People who are good with AI gain most https://x.com/emollick/status/1970646524389478532

A much improved model scheduling system is now on Ollama! – 🫶 Significantly reduced crashes due to out of memory issues – 📍 Maximizing GPU utilization – 🚵 Multi-GPU performance – 🌎 Accurate reporting of memory usage Learn more & try the latest Ollama! 👇👇👇 https://x.com/ollama/status/1970591425566806231

🚨 Major milestone for open-source AI: DeepSeek-R1, with Wenfeng Liang as corresponding author, has landed on the cover of Nature! 🔥 It’s the first fully peer-reviewed LLM published in a top academic journal, sparking huge debate in China’s tech community Zhihu. Zhihu mind https://x.com/ZhihuFrontier/status/1968573286696239247

Google introduces Test-Time Diffusion Deep Researcher Don’t sleep on diffusion models. Test-Time Diffusion Deep Researcher (TTD-DR) is a deep research agent that models research writing as a diffusion process. Instead of static reasoning or bolted-on tools, the system drafts https://x.com/omarsar0/status/1970864565710921891

How are developers using AI? Inside Google’s 2025 DORA report https://blog.google/technology/developers/dora-report-2025/

EmbeddingGemma: lightweight SOTA embeddings • 308M params, built on Gemma 3 • Tops MTEB (multilingual, English, code) <500M models • Matches models 2× its size, efficient even w/ 4-bit or 128-dim embeddings • Encoder-decoder init + geometric distillation • Spread-out https://x.com/arankomatsuzaki/status/1971041110446465251

LLM-JEPA: much worse than image JEPAs imo I read through the paper and it feels super useless because they are constrained to paired data like “”text <-> SQL””. It’s not generalizable to arbitrary data, it just adds a term in the loss so that the embeddings of the “”SQL and text”””” / X https://x.com/scaling01/status/1969410266304545066

Meta Superintelligence Labs presents MetaEmbed: Scalable multimodal retrieval • Flexible late interaction via Meta Tokens • Test-time scaling: trade off retrieval accuracy vs efficiency • SOTA on MMEB + ViDoRe, robust up to 32B models • Matryoshka training → coarse-to-fine https://x.com/arankomatsuzaki/status/1970323735774404960

MetaEmbed is a cool new paper by @ZilinXiao2 in which we append extra writeable “”memory tokens”” at the end of ColPali tokens and only store and use those for Late Interaction. This reduces the memory footprint, yet retains rich query/doc granular interaction that scale well https://x.com/ManuelFaysse/status/1970427315004866977

Thanks @arankomatsuzaki, @ManuelFaysse, and @_akhaliq for reposting our work! We extended matryoshka to multi-vector embedding and were glad to see that the popular concept, test-time scaling in LLM, also works in retrieval. Work done with @AIatMeta & @vislang!”” / X https://x.com/ZilinXiao2/status/1970511456778232074

Tl;Dr: instead of using all the tokens from images/queries for late interaction, they append a set of meta tokens and use them to perform late interaction. They also apply MRL to them to further compress those tokens Seems to have rather good performance while not being too”” / X https://x.com/antoine_chaffin/status/1970400482343493784

Microsoft presents Measuring LLM inference energy (production scale) • Median cost: 0.34 Wh/query (chatbot) • Long reasoning: 4.3 Wh/query (~13× higher) • Fleet scale: ~0.9 GWh/day @1B queries → ~web search level • Public est. often 4–20× too high • Efficiency gains https://x.com/arankomatsuzaki/status/1971059016878240241

Designing multimodal systems is challenging. To build with the community we’re sharing the design behind TensorStream – a convenient tensor-like interface for interleaved multimodal data. Internally our training and inference code is built on top of this. https://x.com/perceptroninc/status/1970670362355736886

Manzano is a multimodal LLM that unifies image understanding and generation. It uses a shared ViT with two adapters: continuous embeddings for image-to-text and discrete FSQ tokens (64K) for text-to-image, both in the same semantic space to reduce task conflict. A single https://x.com/gm8xx8/status/1969974517024923936

Agent run times aren’t everything. I gave the same high level task to Sonnet 4 and GPT-5-Codex. Sonnet completed the task across 7 files in ~6 minutes. Codex is so far at 8 files but 32 minutes in still…. The Sonnet output was perfectly acceptable.”” / X https://x.com/zachtratar/status/1970625784500130065

Aviro (@aviro_ai) makes enterprise AI agents continuously upskill to deliver on complex tasks. Their runtime layer, Cortex, helped their in-house deep research agent beat OpenAI’s by 70% on enterprise search and top Microsoft’s Deep Research benchmark. https://x.com/ycombinator/status/1968691488222503194

The Illusion of Readiness (Health AI) • GPT-5 & peers ace med benchmarks—but stress tests reveal fragility • Guess answers w/o images, flip under trivial prompt tweaks • Fabricate “reasoning” that sounds right but isn’t • Leaderboard wins ≠ real-world readiness https://x.com/arankomatsuzaki/status/1970684893966516477

Is OpenAI’s Reinforcement Fine-Tuning (RFT) Worth It? · TensorZero https://www.tensorzero.com/blog/is-openai-reinforcement-fine-tuning-rft-worth-it/

I have heard of several folks using torchtitan internally for RL training. However, torchtitan doesn’t directly support GRPO, which means folks are adding an implementation themselves. A few questions: 1. Are there any good open-source torchtitan forks with GRPO support? 2. What”” / X https://x.com/iScienceLuvr/status/1968509941578338560

We’re excited to introduce ShinkaEvolve: An open-source framework that evolves programs for scientific discovery with unprecedented sample-efficiency. Blog: https://x.com/SakanaAILabs/status/1971081557210489039

Try the new Qwen models in the Arena!”” / X https://x.com/Alibaba_Qwen/status/1971097727477088717

🚨 New Models Alert: WebDev 💻 GPT-5-Codex and Qwen3-Coder-Plus are both now available on WebDev Arena! In the WebDev Arena, you can test out all the best frontier AI coding models on web development tasks. Vote for your preferred response and see how they stack up on the https://x.com/arena/status/1970962780225507775

Why can robots do backflips but still struggle to open a drawer??? [📍 Link to project] Precise grasping and whole-body coordination make it harder than acrobatics. DreamControl takes a step toward solving this. It combines diffusion models and reinforcement learning to teach https://x.com/IlirAliu_/status/1970539603368042823

[2509.19249] Reinforcement Learning on Pre-Training Data https://arxiv.org/abs/2509.19249

📜 Paper on new pretraining paradigm: Synthetic Bootstrapped Pretraining SBP goes beyond next-token supervision in a single document by leveraging inter-document correlations to synthesize new data for training — no teacher needed. Validation: 1T data + 3B model from scratch.🧵 https://x.com/ZitongYang0/status/1970129028536484089

🚀 #Rodin Gen-2 is NOW LIVE! Our massive scale in data & params delivers: – 4X Mesh Quality🏅 – Recursive Part Gen🔥 – Bake High-Poly #Mesh →Low+Normal💎 – HD Texture(Beta)🚧 🚨Plus all your favs: #3D ControlNets, Quads(Part-level), T/A Pose, #PBR.. 50% OFF for first mo💥 https://x.com/DeemosTech/status/1970501652819149098

APRIL: Active Partial Rollouts in Reinforcement Learning to tame long-tail generation “”we propose Active Partial Rollouts in Reinforcement Learning (APRIL), which mitigates long-tail inefficiency.”” “”Experiments show that APRIL improves rollout throughput by at most 44% across https://x.com/iScienceLuvr/status/1970794655270003037

Bridging the AI production gap: How observability unlocks enterprise AI success https://www.dynatrace.com/info/whitepapers/bridging-the-ai-production-gap/

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision “”This paper asks a simple question: Can inference compute substitute for missing supervision?”” “”the current policy produces a group of rollouts; a frozen anchor (the initial policy) reconciles https://x.com/iScienceLuvr/status/1968599654507102491

crazy that they called it context window when attention span was right there”” / X https://x.com/lateinteraction/status/1970288227904033255

Day-0 support on one of the most anticipated model releases🚀🚀🚀 Detailed deployment guide coming soon at vllm-recipes https://x.com/rogerw0108/status/1970619149757096037

Did you know that when they say stuff like “”The A18 uses TSMC’s 3nm process”” or “”announced the 2nm node”” The 3nm, 2nm actually doesn’t mean anything?! It’s just like a version number. They make it up. Literally nothing measures 2nm or 3nm. I certainly didn’t know. https://x.com/giffmana/status/1970620746155393441

DynaGuard – a new guardian model, highly adaptable to dynamic policies. It can handle custom user rules (what’s allowed or not in different cases) thanks to: • Training on the DynaBench dataset with 40,000 unique policies • Taking in 2 things as inputs: – the policy – a https://x.com/TheTuringPost/status/1970079921704997326

Dynamic CFG: adaptive guidance for diffusion models • Static CFG = “one-size-fits-all” fails across prompts • New method: online feedback from latent evaluators (CLIP, fidelity, prefs) → dynamic per-step CFG • Just +1% overhead, big gains in alignment, quality & text https://x.com/arankomatsuzaki/status/1969975609842688383

Effective reasoning ≠ longer CoTs • Across 10 LRMs: longer chains + more review → lower accuracy • New metric: Failed-Step Fraction (share of abandoned steps) predicts correctness best • FSF-based reranking boosts pass@1 by up to +10% • Editing out failed branches improves https://x.com/arankomatsuzaki/status/1970691075229864357

EmbeddingGemma paper is out, with insights into the architecture, training, initialization, detailed results, and more https://x.com/osanseviero/status/1971187988806897876

Excited to release a preview of Moondream 3. A 9B param, 2B active MoE vision language model that makes no compromises; offering state-of-the-art visual reasoning while still retaining an efficient and deployment-friendly form factor. https://x.com/vikhyatk/status/1968800178640429496

For @bobvanluijt, building @weaviate_io wasn’t a choice… That’s how founders need to feel: “You will die if you don’t do it.”” 🎙️ I talked with Bob van Luijt, co-founder and CEO of Weaviate, the open-source vector database that’s become core infrastructure for AI-native https://x.com/IlirAliu_/status/1968581681461477587

How I Use AI – Tim Kellogg https://timkellogg.me/blog/2025/09/15/ai-tools

Interesting Muon experiment 🤓 Learning rate of Adam parameters (embeddings/gains) does not matter so much (here from 1e-4 to 1e-2). It’s more about Muon LR https://x.com/borisdayma/status/1968711933613211837

Introducing Composite Evaluators in LangSmith 📊Combine multiple evaluator scores into one metric for a complete view of your app’s performance. ➕Supports weighted averages or weighted sums with configurable weights. Learn more 👉 https://x.com/LangChainAI/status/1970540057359720663

Jules now acts on your PR feedback. Leave a comment on any pull request, and Jules will try to implement the requested changes. Go from idea to merged code, all without leaving GitHub. https://x.com/julesagent/status/1970640318606258605

Language Models that Think, Chat Better “”This paper shows that the RLVR paradigm is effective beyond verifiable domains, and introduces RL with Model-rewarded Thinking (RLMT) for general-purpose chat capabilities.”” “”RLMT consistently outperforms standard RLHF pipelines. This https://x.com/iScienceLuvr/status/1971154927415329001

LLM-Deflate: Extracting LLMs Into Datasets https://www.scalarlm.com/blog/llm-deflate-extracting-llms-into-datasets/

LLMs forget. Context windows are short. Sensitive data often leaves your device. mem-agent changes that.  It’s a small local model that manages your memory in natural-language markdown: – Retrieves exactly what you need – Updates as you go – Filters sensitive info on demand https://x.com/driaforall/status/1968377065238831272

Meet LFM2-2.6B, the newest member of our LFM2 family, a new leader in the 3B model class. > light-weight with 2.6B parameters > fast, built by our v2 efficient architecture (short convs + group query attention) > Trained on 10T tokens32k context length > open-weight, https://x.com/LiquidAI_/status/1970484704903119241

Read our report, “ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution” for more details. https://x.com/SakanaAILabs/status/1971214066510332009

Research teams have complexity budgets. Simple foundations allow teams to spend on more novelty and do it more efficiently. VLM training recipes often have complex specs (many stages, permodule LRs, MLP warmup). We wanted to find the minimum recipe that maximizes performance.”” / X https://x.com/kilian_maciej/status/1970701658494738514

RLPT: Reinforcement Learning on Pre-Training Data • RL directly on pre-train data (no human labels) • Next-segment reasoning objective (ASR + MSR tasks) → self-supervised rewards • Gains on Qwen3-4B: +3.0 MMLU, +8.1 GPQA-Diamond, +6.6 AIME24, +5.3 AIME25 https://x.com/arankomatsuzaki/status/1970684035258294548

Sampling and structured outputs in LLMs | parth sareen https://parthsareen.com/blog.html#sampling.md

SBP: Synthetic bootstrapped pretraining • Learns inter-document relations → synthesizes new training data • Goes beyond token-level correlations, capturing latent concepts • 3B model trained on 1T tokens: beats strong repetition baseline • Closes much of the gap to an https://x.com/arankomatsuzaki/status/1969973861178626245

SGLang now supports deterministic LLM inference! Building on @thinkymachines batch-invariant kernels, we integrated deterministic attention & sampling ops into a high-throughput engine – fully compatible with chunked prefill, CUDA graphs, radix cache, and non-greedy sampling. ✅ https://x.com/lmsysorg/status/1970240927429206161

Soft Tokens, Hard Truths • First scalable RL method for continuous CoT • Learns “soft” tokens (mixtures + noise) → richer reasoning paths • Matches discrete CoTs at pass@1, beats them at pass@32 (more diversity) • Best setup: train w/ soft tokens, infer w/ hard tokens https://x.com/arankomatsuzaki/status/1970692910766346277

Some perf related must-reads: • How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: https://x.com/fleetwood___/status/1968716580621271076

Stuck in the chat box with tireless prompting, it’s not using AI, it’s abused by AI. Introducing @HQbetter, the world’s first AI execution engine that runs in the background with minimum input 24/7. Better 10x your work experience by: 1. Read all your context from your past https://x.com/xinyzng/status/1968417929587998928

That’s right, we released our first iteration of JEPAs for LLMs: https://x.com/randall_balestr/status/1969283010982744133

The Extreme Inefficiency of RL for Frontier Models — Toby Ord https://www.tobyord.com/writing/inefficiency-of-reinforcement-learning

The Ultimate Claude Code Guide ⚡️ After building 10+ client MVPs with Claude Code, I wrote down every tip + trick that actually works. I’ve already sent this guide to 800+ people. Want it? → Comment “claude” → Follow I’ll DM you the link. https://x.com/PrajwalTomar_/status/1968322365798179245

Thinking Augmented Pre-training “”we propose Thinking augmented Pre-Training (TPT), a universal methodology that augments text with automatically generated thinking trajectories. Such augmentation effectively increases the volume of the training data and makes high-quality tokens https://x.com/iScienceLuvr/status/1971155514865352990

Today we’re releasing the technical details behind Isaac 0.1. We built Isaac to demonstrate that with the right principles, a simple recipe can reach competitive performance with a small model. Check out the report at: https://x.com/perceptroninc/status/1970701029441483087

Toward Computational Taste: LLMs, Aesthetics & Judgment | Patron https://patron.fund/blog/toward-computational-taste-llms-aesthetics-judgment

We’ve spent nearly a decade building the data foundation of AI, from the training data that powers autonomous vehicles to evaluating LLMs. Now, our Data Engine is fueling the next frontier: physical AI + robotics. https://x.com/scale_AI/status/1970872065340383644

Why does AI sometimes fail to generalize, and what might help? In a new paper, we highlight the latent learning gap — which unifies findings from language model weaknesses to agent navigation — and suggest that episodic memory complements parametric learning to bridge it. Thread: https://x.com/AndrewLampinen/status/1969980297661047055

With Together Instant Clusters, you can keep responses fast during launch‑day spikes. ⚡🚀 https://x.com/togethercompute/status/1968661658617692379

Yay, our team has just published a new paper, “Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads”” https://x.com/StasBekman/status/1971262600555135227

Most grasping methods fail outside clean lab settings. Open-loop breaks under noise. Closed-loop fails in clutter… Grasp-MPC, from TUM and collaborators, combines model-based MPC with data-driven value functions for robust 6DoF closed-loop grasping… even on moving or cluttered https://x.com/IlirAliu_/status/1970197071123652990

SWE-Bench Pro A new coding benchmark to replace SWE-Bench verified was introduced 1-2 days ago: https://x.com/scaling01/status/1969792786594509190

Serving a model at scale is hard. Serving it across three hardware platforms (AWS Trainium, NVIDIA GPUs, Google TPUs) while maintaining strict equivalence is a whole other level. Makes you wonder if the hardware flexibility is truly worth the hit to development speed and https://x.com/_philschmid/status/1968586407548518565

The cost of intelligence continues to fall rapidly after new frontiers are reached: Grok 4 Fast brings the cost of the Intelligence Index >60 category down to just $0.2/$0.5 per million input/output tokens We track the lowest priced model for different tiers of intelligence to https://x.com/ArtificialAnlys/status/1970251031373390292

Pretty wild to see the demographic split on YouTube based on the niche you’re in. Left (my vfx channel) vs. Right (my ai channel) https://x.com/bilawalsidhu/status/1970514756944699515

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading