Image created with Flux Pro v1.1 Ultra. Image prompt: Tech, resort café table, printed research papers with neat equations weighed by a stone, quiet morning, photorealistic, editorial, minimal, landscape, vacation, no text overlays

📊 @Kimi_Moonshot’s K2-0905 on @GroqInc scored 7th overall at 94% on Roo Code evals, the 1st open-source model to break the 90+ barrier. It’s also the fastest and cheapest in the top 10, while holding its own on accuracy. View the full leaderboard: https://x.com/roo_code/status/1965098976677658630

It feels the coding agent frontier is now open-weights: GLM 4.5 costs only $3/month and is on par with Sonnet Kimi K2.1 Turbo is 3x speed, 7x cheaper vs Opus 4.1, but as good Kimi K2.1 feels clean. The best model for me. GPT-5 is only good for complicated specs — too slow.”” / X https://x.com/Tim_Dettmers/status/1965021602267217972

Kimi K2 0905 upgrade: Substantial improvement in agentic capabilities, modest change in overall intelligence Key takeaways: ➤ Intelligence increased +2 pts in our Artificial Analysis Intelligence Index ➤ Agentic capabilities substantially improved as shown by our two new https://x.com/ArtificialAnlys/status/1965010554499788841

🚨 Leaderboard Disrupted! Two new models have entered the Top 10 Text leaderboard: 🔸#6 Qwen3-max-preview (Proprietary) by @Alibaba_Qwen 🔸#8 Kimi-K2-0905-preview (Modified MIT) by @Kimi_Moonshot tied with 7 others. Note that this puts Kimi-K2-0905-preview in a tight race for https://x.com/arena/status/1965115050273976703

Seedream 4.0 is the new leading image model across both the Artificial Analysis Text to Image and Image Editing Arena, surpassing Google’s Gemini 2.5 Flash (Nano-Banana), across both! Seedream 4.0 is the latest release from Bytedance Seed, and is a substantial improvement on https://x.com/ArtificialAnlys/status/1966167814512980210

With reinforcement learning, it learns general principles of coordination – allowing it to generate efficient plans for new workflows in seconds. 💡 This research is a key step towards more adaptable manufacturing lines of the future. Find out more ↓ https://x.com/GoogleDeepMind/status/1965040648400351337

LLMs do many things, to different levels of quality, the “jagged frontier” of ability that my coauthors and I discussed in 2023. One weak part of multimodal LLMs has been seeing fine visual details. So this is an interesting benchmark to watch to follow progress in this area.”” / X https://x.com/emollick/status/1964758268930379794

Warp Code launched yesterday — here’s what’s new: – Top coding agent: #3 SWE-bench, 52% Terminal-Bench – Built-in code review – Native editor – Slash Commands, Project Rules, and more We’re already seeing millions more lines of code shipped through Warp. https://x.com/warpdotdev/status/1963683282538688694

MBZUAI and G42 Launch K2 Think: A Leading Open-Source System for Advanced AI Reasoning https://www.prnewswire.com/news-releases/mbzuai-and-g42-launch-k2-think-a-leading-open-source-system-for-advanced-ai-reasoning-302551074.html

⚡️ Efficient weight updates for RL at trillion-parameter scale 💡 Best practice from Kimi @Kimi_Moonshot vLLM is proud to collaborate with checkpoint-engine: • Broadcast weight sync for 1T params in ~20s across 1000s of GPUs • Dynamic P2P updates for elastic clusters •”” / X https://x.com/vllm_project/status/1965824120920342916

Introducing checkpoint-engine: our open-source, lightweight middleware for efficient, in-place weight updates in LLM inference engines, especially effective for RL. ✅ Update a 1T model on thousands of GPUs in ~20s ✅ Supports both broadcast (sync) & P2P (dynamic) updates ✅ https://x.com/Kimi_Moonshot/status/1965785427530629243

Updated & turned my Big LLM Architecture Comparison article into a narrated video lecture. The 11 LLM architectures covered in this video: 1. DeepSeek V3/R1 2. OLMo 2 3. Gemma 3 4. Mistral Small 3.1 5. Llama 4 6. Qwen3 7. SmolLM3 8. Kimi 2 9. GPT-OSS 10. Grok 2.5 11. GLM-4.5 https://x.com/rasbt/status/1965798055141429523

💁‍♂️ Introducing human-in-the-loop Middleware! 💵 Tool calls can be risky and expensive! For certain tool calls you may want to get user feedback to approve, deny, or modify them before execution. Our new middleware provides an easy off-the-shelf way to build this into your https://x.com/sydneyrunkle/status/1966184060360757340

Today we want to share a hot & impressive debate: Why do today’s AI Agents often feel all hype, no results? 🧠 Zhihu mind explorer, Prof. 俞扬 from Nanjing University explains: LLM-based Agents ≠ LLMs themselves: LLMs focus on generation/prediction, while agents focus on https://x.com/ZhihuFrontier/status/1964928650081698167

Wow, thanks to @charles_irl , you can understand internals of vLLM with a live notebook from @modal 🥰”” / X https://x.com/vllm_project/status/1965708611222684017

People are obsessed with this “”machine break out of the simulation”” story but it’s just not real. This affected a few trajectories in 4 submissions. The bug has now been fixed by @_carlosejimenez. The overall picture and the trends on SWE-bench are not affected at all. https://x.com/OfirPress/status/1966227423252595056

🗣️ Evals now support native audio inputs and audio graders. Evaluate model audio responses, with no text transcription needed. Get started in the Cookbook guide: https://x.com/OpenAIDevs/status/1965923707085533368

Congrats to @Zai_org GLM-4.5 on getting the 7th spot on our SWE-bench Verified [Bash Only] leaderboard! w/ @KLieret @_carlosejimenez @jyangballin https://x.com/OfirPress/status/1965889864395899262

Results – Outperformed GPT-4o on web navigation (26% vs. 16%). – Did very well on deep search (38% vs. GPT-4o’s 26%), even topping some subtasks. – Reached 91% in the TextCraft game and was one of the few models to handle the hardest level. – Hit 96.7% in BabyAI (a simulated https://x.com/omarsar0/status/1966167191805734978

.@_carlosejimenez just merged a PR that fixes the SWE-bench bug that allowed agents to ‘look into the future’. Our analysis showed that this bug was only exploited by a few agents a handful of times. Version update coming soon with a bunch of other extra things! https://x.com/OfirPress/status/1965978758336163907

With the economic value at stake, it would be worthwhile to assemble a few standing bodies of experts to do very fast evaluations for new AI models so that the world doesn’t need to rely on benchmarks that consist of math problems, trivia questions, & the vibes of people like me.”” / X https://x.com/emollick/status/1964044908030816378

✨ NEW: Feature update: 🖼️ Image edit models now support multi-turn editing! Instead of trying to fit every edit into one mega-prompt, you can now refine your image step by step. Like a natural back-and-forth conversation. Do it in Battle mode, Side by Side or Direct. Just https://x.com/arena/status/1965150440401809436

Don’t forget: Multi-turn is now available in Image Edit! You don’t have to prompt for a one-shot. Iterate and edit via a natural conversation. https://x.com/arena/status/1965929101799399757

One challenge that no AI model has solved yet is “”Create a compelling puzzle that is solvable by players for a D&D game that isn’t boring or trite and where choices matter.”” It just involves too much planning & detail. Here, GPT-5 Pro comes very close, but there are still flaws. https://x.com/emollick/status/1964882159157784961

We just released AlgoPerf v0.6! 🎉 ✅ Rolling leaderboard ✅ Lower compute costs ✅ JAX jit migration ✅ Bug fixes & flexible API Coming soon: More contemporary baselines + an LM workload… https://x.com/algoperf/status/1965044626626342993

📢 New Model Drop: Seedream 4.0 is live on Yupp! This image model from ByteDance offers text-to-image generation as well as image editing. We dove in with some prompts: https://x.com/yupp_ai/status/1965827081826422990

🚨 ByteDance just released Seedream 4.0 — how does its AI image generation perform? Zhihu contributors share their feedbacks. Let’s have a quick view👇 🎨 Trisimo 崔思莫: ➤ Seed 4.0 vs Nano Banana: Different tech paths. • Nano Banana: multimodal, stronger understanding & https://x.com/ZhihuFrontier/status/1965681077231727069

🚨 New Model Alert! ByteDance’s latest Seedream 4 is ready in the Arena! 🖼️Seedream 4 merges the capabilities of Seedream 3 (Text-to-Image) with SeedEdit 3 (Image Edit). Come and test out your hardest Text-to-Image and Image Edit prompts! https://x.com/arena/status/1965929099370889432

Heads up: if you’re looking to try CUDA-13 (most useful for Blackwell gpus), PyTorch nightly already has cu130 builds: https://x.com/StasBekman/status/1965826539540590791

This is a handy report on the state of cloud GPUs in 2025: costs, performance, playbooks by @dstackai https://x.com/StasBekman/status/1965817531043811339

DeepSeek V3.1 dynamic @UnslothAI quants on Aider Polyglot benchmarks are here! 1. 3-bit thinking gets 75.6% vs 76.1% un-quantized 2. Leaving attn_k_b in 8-bit gets +2% accuracy vs 4-bit 3. Dynamic quants beat other similar imatrix quants 4. AMA r/LocalLlama today 10AM PST! https://x.com/danielhanchen/status/1965800675105017980

Evals are a scam. And we’re being gaslit into believing they aren’t. New post just dropped (🧵). https://x.com/AlexReibman/status/1964116243847565526

Evals are a scam. This is what Twitter was fighting about all weekend, and honestly? Both sides are missing the point. The real problem isn’t whether evals work or don’t work. It’s that everyone uses “”evals”” to mean 6 different things and then acts shocked when the conversation https://x.com/bnicholehopkins/status/1965130607790264452

We challenged ourselves to build the cleanest, highest-signal factuality benchmark out there. Today, we’re releasing the result: SimpleQA Verified ✅🥇 On this more reliable, 1,000-prompt eval, Gemini 2.5 Pro establishes a new SOTA, outperforming other frontier models. We’re https://x.com/lkshaas/status/1965799946621202719

SimpleQA Verified Benchmark! A New reliable factuality benchmark for measuring knowledge in LLMs from our Research Team @GoogleDeepMind! – Includes exactly 1,000 prompts for evaluating short-form factuality. – Reduced from original 4,326 questions to address incorrect labels. – https://x.com/_philschmid/status/1965806183970652368

Gemini, just like everybody else. From a fascinating blog post about AI agents assigned to play web games, and failing, in large part because vision and computer use tools aren’t good enough: https://x.com/emollick/status/1963968533051617322

Apologies that I haven’t written anything since joining Thinking Machines but I hope this blog post on a topic very near and dear to my heart (reproducible floating point numerics in LLM inference) will make up for it!”” / X https://x.com/cHHillee/status/1965828670167331010

4B OCR with Apache-2.0 license outperforming Mistral OCR 🔥 Tencent released Points-Reader, it’s a new model firstly trained on Qwen2.5VL annotations and then self-trained on real data in many benchmarks, it performs better than Qwen2.5VL and MistralOCR! https://x.com/mervenoyann/status/1966176133894098944

i’m hiring for a new team @openai: Applied Evals our goal is to build the world’s best evals for the economically valuable tasks our customers care about most. we’ll execute as a group of high‑taste engineers, combining hands-on, unscalable efforts with systems that others can”” / X https://x.com/shyamalanadkat/status/1965807750916812803

OpenAI referenced Artificial Analysis’ Big Bench Audio benchmark in their recent GPT-Realtime release, where they secured the #1 position with a score of 83% Benchmark context: Big Bench Audio is the first dedicated dataset for evaluating reasoning performance of speech models. https://x.com/ArtificialAnlys/status/1966116575851028970

[1701.06538] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer https://arxiv.org/abs/1701.06538

@suchenzang Always depends on the context. This has always been a big discussion, even in the convolutional neural network days (non-deterministic by default in CUDA)”” / X https://x.com/rasbt/status/1965918363928211459

@winglian So a PyTorch backend is effectively a list of many PyTorch operators you need to implement, this eval is closest in spirit to a PyTorch focused kernelbench level 1. Fusions were out of scope for our first release but something we’d like to visit soon”” / X https://x.com/marksaroufim/status/1965803842697875746

🔥AutoRound is now part of SGLang. Thanks to the team work, esp. Weiwei, Wenhua, and Yi from @intel , and Yineng, Jiexin, and FlamingoPg from @sgl_project. 👏PR https://x.com/HaihaoShen/status/1964926924880523701

🤔 Struggling with slow weight sync across nodes in RL setups? This deep dive by Zhihu contributor 陈乐群 @abcdabcd987 shows how he achieved ~2s weight sync for Qwen3‑235B (BF16 → FP8) from 128 training GPUs → 32 inference GPUs, using raw RDMA WRITEs. – No disk I/O – No host https://x.com/ZhihuFrontier/status/1965620593463861638

🚀 Just shipped TRL v0.23 – train with *any* context length This release brings Context Parallelism which allow to train with arbitrary context length along with major improvements for post-training Here’s what’s new 🧵👇”” / X https://x.com/QGallouedec/status/1965798094660084017

a lightweight RFT solution integrated with prime-rl, verifiers and the Environments Hub is now live on our platform too more coming soon as we scale this to the full-stack SOTA RL infra currently locked behind the walls of closed labs https://x.com/johannes_hage/status/1965868932545593845

Check out Set Block Decoding, combining next-token-prediction and masked (or discrete diffusion) models, allowing parallel decoding (x3-5 speedup) without any architectural changes and with exact KV cache. Matches NTP performance!”” / X https://x.com/itai_gat/status/1965112129499046230

Deep dive into optimizing weight transfer step by step and improving it 60x!”” / X https://x.com/vllm_project/status/1965879465688576110

Defeating Nondeterminism in LLM Inference – Thinking Machines Lab https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/

General VLMs struggle with precise text localization & often hallucinate on dense docs. 🔥PP-OCRv5 solves this with a modular, two-stage pipeline: ✅Smaller, faster, lighter ✅Accurate bounding boxes ✅Edge-device friendly Read more 👇 https://x.com/PaddlePaddle/status/1965957482716832193

HierMoE is a topology-aware MoE training system that cuts All-to-All overhead by deduplicating tokens across hierarchy levels and swapping experts with a comm-time model on duplicate-free tokens. All-to-All dominates MoE training because tokens must be shuffled across GPUs; https://x.com/gm8xx8/status/1965926377279902022

Is there truly a gap in performance between online and offline RL training for LLMs? Here’s what the research says… TL;DR: There is a clear performance gap between online and offline RL algorithms, especially in large-scale LLM training. However, this gap can be minimized by https://x.com/cwolferesearch/status/1965088925510520853

KV cache compression techniques ▪️KV caching (basic) – stores previously computed Keys and Values in memory and calculates attention only for new tokens. ▪️ Quantization – represents KV cache with fewer bits. ▪️ Low-rank decomposition – compresses the KV cache into smaller https://x.com/TheTuringPost/status/1964971207188791464

Kyutai presents DSM: Streaming seq2seq with delayed streams • Handles ASR ↔ TTS with SOTA latency/quality (few 100ms) • Competitive with offline baselines • Decoder-only LM + pre-aligned streams → simple & flexible • Supports infinite sequences, batching https://x.com/arankomatsuzaki/status/1965984604558700751

LLM inference can be made deterministic with a special care on kernels. Check out the vllm example on how we achieve this!”” / X https://x.com/woosuk_k/status/1965843109876797935

Need to read this but fwiw when things are just text to text, you can get perfect repeatability with a cache. Sometimes things are simpler than they look.”” / X https://x.com/lateinteraction/status/1965919773193380290

Qwen3-Next: > Gated DeltaNet with standard attention, 3:1 ratio > 512 experts, 10 routed + 1 shared > Zero-Centered RMSNorm + weight decay to norm weights, no need for attention sink tricks like with SWA in OSS > optimized MTP > “”only”” 15T tokens > perf closer to 235B-22AB@36T https://x.com/teortaxesTex/status/1966201258404204568

Really great deep dive on sources of nondeterminism in LLM inference. Before reading, I also believed atomicAdd was to blame for all of it, but it seems like that’s mostly a red herring nowadays!”” / X https://x.com/sedielem/status/1966103855508169006

Recently finished writing a new blogpost about @PyTorch compilation in ZeroGPU Spaces. Worth reading if you’re interested in learning about : – PyTorch ahead-of-time compilation – ZeroGPU internals https://x.com/charlesbben/status/1965046090945954104

SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge: https://x.com/_philschmid/status/1965806186827330035

Someone has finally published nuances of how NCCL algorithms and protocols work. Thank you so much to the authors since documentation is so scarce! https://x.com/StasBekman/status/1966194963194257759

Today Thinking Machines Lab is launching our research blog, Connectionism. Our first blog post is “Defeating Nondeterminism in LLM Inference” We believe that science is better when shared. Connectionism will cover topics as varied as our research is: from kernel numerics to https://x.com/thinkymachines/status/1965826369721623001

AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
https://agentgym-rl.github.io/

building pretraining infrastructure is an exercise in complexity management, abstraction design, operability/observability, and deep systems and ML understanding. reflects some of the trickiest and most rewarding problems in software engineering. which makes it really fun!”” / X https://x.com/gdb/status/1964538494895993162

is “”numerical determinism”” worth a 60%* latency hit? 🤔 (* = unoptimized upper-bound by some of the best talent in the industry) https://x.com/suchenzang/status/1965914700786622533

Writing fast GPU kernels is important, though not nearly as important as writing correct ones. That’s why the folks from Meta have released BackendBench.https://x.com/johannes_hage/status/1965945249274151107

Fine-tune Any LLM from the Hugging Face Hub with Together AI https://huggingface.co/blog/togethercomputer/together-ft

Thank you @cHHillee for the great explanation and demo of how to implement deterministic inference on vLLM!!! https://x.com/vllm_project/status/1965832975297503662

Meta researchers just unveiled Set Block Decoding on Hugging Face. It’s a game-changer for language model inference, delivering 3-5x speedup in token generation with existing models. No architectural changes needed, matches previous performance. https://x.com/HuggingPapers/status/1965084731839513059

Announcing Genkit Go 1.0 ✨ The SDK is now stable and production-ready. This release introduces the genkit init:ai-tools command for seamless integration with AI coding tools, plus built-in support for tool calling, RAG, and more. Read the blog → https://x.com/googledevs/status/1965778301949022441

The pi-05 model is now in openpi: https://x.com/svlevine/status/1965161524722630734

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading