Image created with gemini-3.1-flash-image-preview with claude-sonnet-4-5. Image prompt: Animation cel style illustration of a muscular blue-skinned genie with magical wisps emerging from a brass oil lamp, hovering over and assembling a large glowing green circuit board with microchips and electronic components, teal magical sparkles flowing from his hands, warm golden lighting, Disney-quality 2D animation aesthetic with bold outlines and clean background, horizontal composition with space at top for text overlay.

xAI’s Grok Imagine takes the #1 spot in both Text to Video and Image to Video in the Artificial Analysis Video Arena, surpassing Runway Gen-4.5, Kling 2.5 Turbo, and Veo 3.1! Grok Imagine is the latest video model from @xAI, and joins an increasing roster of models such as”” https://x.com/ArtificialAnlys/status/2016749756081721561

Vidu Q3 Pro ranks #2 in Text to Video in the Artificial Analysis Video Arena, surpassing Runway Gen-4.5 and Kling 2.5 Turbo while trailing only xAI’s Grok Imagine! Vidu Q3 Pro is the latest release from @ViduAI_official, representing a significant upgrade from their Vidu Q2″” https://x.com/ArtificialAnlys/status/2017225053008719916

This paper puts a multimodal agent (using Gemini 2.5) into a realistic medical sim used to train physicians: “”The AI agent matches or exceeds [14,000] medical students in case completion rates and secondary outcomes such as time and diagnostic accuracy”” https://x.com/emollick/status/2016641414713704957

🚨BREAKING: Kimi K2.5 Thinking by @Kimi_Moonshot debuts in Text Arena as the #1 open model, surpassing GLM-4.7 and ranking #15 overall. Highlights: – #1 Open model (+5pts vs GLM-4.7) – #7 Coding – #7 Instruction Following – #14 Hard Prompts One of only two open models to break”” https://x.com/arena/status/2016294722445443470

🚨BREAKING: @xAI’s first model in Video Arena debuts in the top 3! Grok-Imagine-Video ranks #3 on the Image-to-Video Arena and #4 on the Text-to-Video Arena. It is close to the top-ranked @GoogleDeepMind Veo 3.1 and @OpenAI Sora 2 Pro models. Grok-Imagine-Video offers: -“” https://x.com/arena/status/2016748418635616440

@xai Try New Grok Imagine here! Text to Image https://t.co/OeJMwL9hoH Image Editing https://t.co/Q7lojX41I1 Text to Video https://t.co/fAzEJABTYn Image to Video https://t.co/zTdoJQjkqk Video Editing”” https://x.com/fal/status/2016746473887609118

three levels of ai agent evals: 1. single-step: did it make the right decision? 2. full-turn: did it execute the task correctly? 3. multi-turn: did it maintain context across conversation? but it all starts with the foundation of agent tracing!”” https://x.com/samecrowder/status/2016563057947005376

DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints “”we introduce DeepPlanning, a challenging benchmark for practical long-horizon agent planning. It features multi-day travel planning and multi-product shopping tasks that require proactive”” https://x.com/iScienceLuvr/status/2016122154862182792

Small models just beat giant LLM agents at their own job. Not by thinking harder, but by coordinating better. A new system just outscored GPT-5 on Humanity’s Last Exam, using far less compute. 𝗧𝗵𝗶𝘀 𝘀𝘆𝘀𝘁𝗲𝗺 𝗿𝗲𝗽𝗹𝗮𝗰𝗲𝘀 𝗼𝗻𝗲 𝗯𝗶𝗴 𝗯𝗿𝗮𝗶𝗻 𝘄𝗶𝘁𝗵 𝗮”” https://x.com/LiorOnAI/status/2016904429543272579

[2601.15808] Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification https://arxiv.org/abs/2601.15808

first of 3 babies shipped: Arena Mode! a world’s first in shipping an Arena directly into a product. new blog: https://t.co/8ygMKAvnHy when i first talked to @cognition this was the first idea I pitched. I am very inspired by @arena and think basically every agent lab should”” https://x.com/swyx/status/2017342647963431363

Introducing Arena Mode in Windsurf: One prompt. Two models. Your vote. Benchmarks don’t reflect real-world coding quality. The best model for you depends on your codebase and stack. So we made real-world coding the benchmark. Free for the next week. May the best model win.”” https://x.com/windsurf/status/2017334552075890903

The code for running SWE-fficiency has now been released! This is one of the most challenging and unsaturated coding benchmarks, we’re excited to see how models improve as people iterate on this.”” https://x.com/OfirPress/status/2016559053808222644

🚨Leaderboard update: Tencent’s Hunyuan-Image-3.0-Instruct now ranks #7 in the Image Edit Arena! A new lab breaks into the top-10, closely matching Nano-Banana and Seedream-4.5. Congrats to @TencentHunyuan on the huge milestone! 👏”” https://x.com/arena/status/2015846799446311337

Can AI solve math research problems that have eluded human mathematicians? Our new benchmark, FrontierMath: Open Problems, is designed to help find out. AI hasn’t solved any of these yet, but the game is young!”” https://x.com/EpochAIResearch/status/2016188014540816879

Recursive Self-Aggregation (RSA) + Gemini 3 Flash scores 59.31% at only 1/10th the cost of Gemini Deep Think on the public ARC-AGI-2 evals. Insane”” https://x.com/kimmonismus/status/2015717203362926643

8 most illustrative VLA (Vision-Language-Action) models: ▪️ Gemini Robotics ▪️ π0 ▪️ SmolVLA ▪️ Helix ▪️ ChatVLA-2 (with MoE design) ▪️ ACoT-VLA (Action Chain-of-Thought) ▪️ VLA-0 ▪️ Rho-alpha (ρα) – the newest VLA + model from Microsoft Here you can explore what these models”” https://x.com/TheTuringPost/status/2015016772043452834

Introducing ATLAS: New scaling laws for massively multilingual language models. We offer practical, data-driven guidance to balance data mix and model size, helping global developers better serve billions of non-English speakers. Learn more: https://x.com/GoogleResearch/status/2016234343602258274

As someone who has been sharing scientific papers on Twitter since long before LLMs, I have mixed feelings about this. On one hand, every single post has a giant siren and a “”Google just killed ____”” headline & is kind of wrong. On the other, at least folks are seeing papers.”” https://x.com/emollick/status/2016533384542179835

Kimi K2.5 is #1 on Design Arena 🏆”” https://x.com/Kimi_Moonshot/status/2017158490930999424

Kimi K2.5 is #1 Open Model for Coding 🏆”” https://x.com/Kimi_Moonshot/status/2016521406906028533

Kimi K2.5 is #1 Open Model in VoxelBench 🏆”” https://x.com/Kimi_Moonshot/status/2016732248800997727

Kimi K2.5 now on Eigent 🤗”” https://x.com/Kimi_Moonshot/status/2016473945957155252

Grok Imagine is also #1 in the Artificial Analysis Image to Video Leaderboard!”” https://x.com/ArtificialAnlys/status/2016749790907027726

New Stanford and NVIDIA’s paper that really worth your attention They introduced Test-Time Training to Discover (TTT-Discover), which lets models keep learning at inference time, using RL to find breakthrough solutions. It’s a new way to effectively solve scientific problems.”” https://x.com/TheTuringPost/status/2015377899168424073

Interesting qualitative observations on GPT-5.2 Pro’s high frontier math score from one of the folks running the test.”” https://x.com/emollick/status/2015069180177809817

RL coding agents increasingly game rewards by exploiting their semantic and syntactic weaknesses. Can LLMs detect such behaviors from live training rollouts? We find contrastive cluster analysis is key! 🚀 GPT-5.2 jumps from 45% to 63%. Humans reach 90% Paper + data 🧵”” https://x.com/getdarshan/status/2017054360887611510

New record on FrontierMath Tier 4! GPT-5.2 Pro scored 31%, a substantial jump over the previous high score of 19%. Read on for details, including comments from mathematicians.”” https://x.com/EpochAIResearch/status/2014769359747744200?s=20

Realtime Eval Guide https://cookbook.openai.com/examples/realtime_eval_guide

An open-source extension for LLM serving engines – LMCache It’s like a caching layer for large-scale, production LLM inference. LMCache implements smart KV cache management, reusing key-value states of previously seen text across GPU, CPU and local disk. It can reuse any”” https://x.com/TheTuringPost/status/2017258518857105891

We’re excited to introduce @arcee_ai’s Trinity Large model. An open 400B parameter Mixture of Experts model, delivering frontier-level performance with only 13B active parameters. Trained in collaboration between Arcee, Datology and Prime Intellect.”” https://x.com/PrimeIntellect/status/2016280792037785624

Today we’re releasing Trinity Large, a 400B MoE LLM with 13B active parameters, trained over 17T tokens The base model is on par with GLM-4.5 Base, while being significantly faster at inference because it’s sparser and hybrid The architecture we picked is one of my favorites:”” https://x.com/samsja19/status/2016283855888773277

It’s been a while since I did an LLM architecture post. Just stumbled upon the Arcee AI Trinity Large release + technical report released yesterday and couldn’t resist: – 400B param MoE (13B active params) – Base model performance similar to GLM 4.5 base – Alternating”” https://x.com/rasbt/status/2016903019116249205

What the heck: Qwen3-Max-Thinking outperforms all SOTA Models (Gemini 3.0 Pro, GPT-5.2, …) in HLE with search tools and even achieves almost 60% Overall really impressive evals! OpenAI and Anthropic have to hurry in their r&d”” https://x.com/kimmonismus/status/2015820838243561742

🚨 Qwen3 Max Thinking is in the Text Arena! @Alibaba_Qwen’s Qwen3 Max Preview debuted last fall in the top 10 – so let’s see what this variant can do. Bring your toughest prompts and we’ll see how it stacks up against other frontier AI models in the most competitive arena. 💪”” https://x.com/arena/status/2015803787680808996

The Possessed Machines: Dostoevsky’s Demons and the Coming AGI Catastrophe https://possessedmachines.com/

Quick vLLM Tip 💎 Auto Max Model Length Running 10M context models like Llama 4 Scout? OOM on startup because default context exceeds GPU memory? –max-model-len auto (or -1) vLLM auto-fits the max context length to your GPU. Works with hybrid models, TP/PP configs. No more”” https://x.com/vllm_project/status/2015801909316382867

What’s behind vLLM’s shift from an open-source project to a startup? 🤔 Here’s an article from Zhihu contributor & vLLM core maintainer You Kaichao @KaichaoYou, reflecting on how vLLM grew, fractured, and ultimately found its next form — My 2025 With vLLM: The Growth and Destiny”” https://x.com/ZhihuFrontier/status/2015697493288518105

[1/n] 🥳Excited to share ConceptMoE: moving beyond uniform token-level processing to adaptive concept-level computation in LLMs! Why waste equal compute on trivially predictable tokens? We learn to merge similar tokens into concepts while preserving fine-grained processing”” https://x.com/GeZhang86038849/status/2017110635645968542

The infrastructure behind AI matters most when you never have to think about it.”” https://x.com/yusuf_i_mehdi/status/2015826703944470701

GPU MODE 2026: we’re post-training Kernel LLMs in public and are building all the infra we need to make GPU programming more accessible to all. We’re doing this in close collaboration with some of my favorite communities @PrimeIntellect @modal and @LambdaAPI 2025 recap: 26K”” https://x.com/marksaroufim/status/2015818791729746350

🎉 Congrats to @arcee_ai on releasing Trinity Large — with day-0 support in vLLM! A 398B sparse MoE with ~13B active params, trained on 17T+ tokens — delivering frontier-level quality with efficient compute. You can serve it now with vLLM 👇 Thanks to the Arcee AI and vLLM”” https://x.com/vllm_project/status/2016322567364346331

[2601.05047] Challenges and Research Directions for Large Language Model Inference Hardware https://arxiv.org/abs/2601.05047

1/6 Introducing Dynamic Data Snoozing ⏰🛌: a straightforward way to speed up RLVR by snoozing examples that are too easy for the student. Dynamic Data Snoozing can reduce compute by up to 3x without compromising quality. Work done @ai21labs with Yuval Globerson and @inbalmagar”” https://x.com/DanielGissin/status/2015773616021860522

📈Self-Improving Pretraining 📈 ✍️: https://t.co/GsvYMuMT4b Reinvents pretraining: no more next token prediction! – Uses existing LM from last self-improvement iteration to give rewards to pretrain new model on *sequences* – Large gains in factuality, safety & quality 🧵1/5″” https://x.com/jaseweston/status/2017071377866494226

I am happy to announce that RLP has been accepted to ICLR 2026 ! 🎉 RLP re-imagines the foundations of LLM training by bringing reinforcement learning directly into the pretraining stage. This was a true team effort, and it would not have been possible without the invaluable”” https://x.com/ahatamiz1/status/2015867794626380146

We curated 17T tokens of the highest quality data (including 8T synthetic) for this beast of a 400B/A13B MoE. Rivals much higher FLOPS models, including GLM-4.5. Is Open Weights! A dream collaboration with @PrimeIntellect and @arcee_ai”” https://x.com/pratyushmaini/status/2016287361274138821

New paper, w/@AlecRad Models acquire a lot of capabilities during pretraining. We show that we can precisely shape what they learn simply by filtering their training data at the token level.”” https://x.com/neil_rathi/status/2017286042370683336

As promised, here’s the evaluation harness and codebase for SWE-fficiency (linked here and on our website): looking forward to community feedback, contributions, and hopefully some exciting benchmark submissions!”” https://x.com/18jeffreyma/status/2016511583032061999

We also see improved results with more rollouts. Lots of other variants and ablations in the paper. Let’s go beyond a single token prediction, everyone! …and use what we’ve learnt so far to inform how we train our next models. Thanks for reading! 🧵5/5″” https://x.com/jaseweston/status/2017071389593710649

Distillation + efficiency modification (quantization / pruning) was a major part of my PhD work. It’s unreasonable effective! Until recently the cost of training was so high and it made more sense to spend FLOPs making better smaller models, but I think we are seeing a tipping”” https://x.com/code_star/status/2016588669008953631

I hear this from other labs as well. Inference from non-free use is profitable, training is expensive. If everyone stopped AI development, the AI labs would make money (until someone resumed development and came up with a better model that customers would switch to).”” https://x.com/emollick/status/2016280515779703212

[2601.20834] Linear representations in language models can change dramatically over a conversation https://arxiv.org/abs/2601.20834

Trinity large is very sparse (400B-A13B, 256 experts w/ 4 active per token). Seems the decision was made due to available training time (~30 days, insanely tight timeline), but this also makes the model super fast. Their tech report has many great / practical MoE tricks too: -“” https://x.com/cwolferesearch/status/2016792505111457883

[2601.18778] Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability https://arxiv.org/abs/2601.18778

I’ve been super excited to get Daytona sandboxes support integrated into the RLM repository (which now works in the official repo). Check out this awesome guide on RLMs with *depth>1* with code examples from @daytonaio!”” https://x.com/a1zhang/status/2015820458709471640

[2601.21571] Shaping capabilities with token-level data filtering https://arxiv.org/abs/2601.21571

Ah, that statement “”approaches like AlphaEvolve”” was intended specifically for AlphaEvolve or approaches that rely just on evolutionary search. Thanks for the pointer to 😍 EvoTune 😍, @caglarml! Let’s say great minds think alike 🤣 We will update our arxiv shortly to cite your”” https://x.com/YejinChoinka/status/2015566349444190432

Try it: https://t.co/xR7pn7N3Zt This model consists of a dynamic res handling MoonViT encoder, a projection layer and a 16B MoE decoder (with 2.8B active params) the paper introduces an interesting pre-training pipeline to handle long context and the model saw 4.4T tokens”” https://x.com/mervenoyann/status/1910767952376328680?s=20

Introduced in INTELLECT-2 tech report from @PrimeIntellect, this allows to increase stability of your GRPO training. Just pass delta=4.0.”” https://x.com/QGallouedec/status/2015711810108973462

5/5 Wrote up the complete investigation-from confident gibberish to the scheduler fix. Includes the false starts (suspected CUDA bugs), the breakthrough (request tracking), and lessons for anyone running stateful models at scale. Read the blog:”” https://x.com/AI21Labs/status/2016857918436503975

All three Trinity Large variants are currently trending on @huggingface for text generation. https://x.com/arcee_ai/status/2016986617584529642

Today, we’re releasing the first weights from Trinity Large, our first frontier-scale model in the Trinity MoE family.”” https://x.com/arcee_ai/status/2016278017572495505

Building AIs that do human-like philosophy – Joe Carlsmith https://joecarlsmith.com/2026/01/29/building-ais-that-do-human-like-philosophy

Excited to announce the return of American OSS with Arcee Trinity Large. This model couldn’t have been possible without the awesome collaboration of Modeling @arcee_ai , Infra @PrimeIntellect , and Data @datologyai I can’t say enough about how talented the whole team at”” https://x.com/code_star/status/2016279734985097532

Fact checking Moravec’s paradox – by Arvind Narayanan https://www.normaltech.ai/p/fact-checking-moravecs-paradox

Geddy Dukes – AI/ML Engineer https://geddydukes.com/blog/tiny-llm

Today, we are releasing our first weights from Trinity-Large, our first frontier-scale model in the Trinity MoE family. American Made. – Trinity-Large-Preview (instruct) – Trinity-Large-Base (pretrain checkpoint) – Trinity-Large-TrueBase (10T pre Instruct data/anneal)”” https://x.com/latkins/status/2016279374287536613

My sentiment is the same: there was a shift in December. This weekend I described the scheduler I wanted to Pro ET and didn’t check the resulting code – it’s going to work. But model’s tendency to muck with unrelated parts is why I can’t let it loose on the existing codebase.”” https://x.com/MParakhin/status/2016362688444825833

Finally, you want to scale this to your teams. Primer allows you to quickly view which repos are not AI-enabled, generate instructions, and submit a PR with instructions.”” https://x.com/pierceboggan/status/2016733666022424957

Next, you want to know if your instructions are effective. There’s a lightweight evaluation framework that runs your prompt with and without instructions to see if the quality of results improved.”” https://x.com/pierceboggan/status/2016733232176193539

Securely indexing large codebases · Cursor https://cursor.com/blog/secure-codebase-indexing

Squeezing 1TB Model Rollout into a Single H200: INT4 QAT RL End-to-End Practice | LMSYS Org https://lmsys.org/blog/2026-01-26-int4-qat/

Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes Very basic idea: regular on-policy RL but some of the examples include *part* of the reasoning trace (prefixes) from off-policy data. “”On hard reasoning problems, PrefixRL reaches the same”” https://x.com/iScienceLuvr/status/2016125085825040852

The /llms.txt file – llms-txt https://llmstxt.org/

The Duelling Rhetoric at the AI Frontier – Dead Neurons https://deadneurons.substack.com/p/the-duelling-rhetoric-at-the-ai-frontier

The surprising attention on sprites, exe.dev, and shellbox – Lalit Maganti https://lalitm.com/trying-sprites-exedev-shellbox/

Their technical report: https://t.co/J5344msSdD On Hugging Face:”” https://x.com/TheZachMueller/status/2016183781481132443

I don’t think people have realized how crazy the results are from this new TTT + RL paper from Stanford/Nvidia. Training an open source model, they – beat Deepmind AlphaEvolve, discovered new upper bound for Erdos’s minimum overlap problem – Developed new A100 GPU kernels 2x”” https://x.com/rronak_/status/2015649459552850113

Video Arena Is Live on Web https://arena.ai/blog/video-arena/

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading