Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: Using the provided Alesso reference image, keep the pure white background, landscape composition, vertical type hierarchy, and galaxy-punchout starfield treatment inside every letterform exactly as in the reference, but replace ‘HEROES’ with ‘TECH’ in the same bold condensed grotesque all-caps, replace ‘ALESSO’ with ‘UNIVERSAL SUBSTRATE’ in the same light geometric all-caps, keep ‘(we could be)’ and ‘FEATURING.’ unchanged, and replace ‘TOVE LO’ with ‘OPEN SILICON’ in the same condensed grotesque all-caps galaxy-punchout, maintaining identical tracking, font contrast, and album-cover austerity.

// Skill Learning for Autonomous Web Agents // Web agents can navigate a page, but ask them to repeat a checkout flow they already completed, and they start from scratch every time. This work introduces WebXSkill, a skill learning framework where web agents extract reusable
https://x.com/dair_ai/status/2045139481892880892

Anthropic released Claude design, direct attack on figma and lovable. Anthropic just shipped Claude Design, powered by Claude Opus 4.7., a tool that turns conversations into polished prototypes, pitch decks, and marketing assets. It auto-applies your brand system, lets you
https://x.com/kimmonismus/status/2045162358004216134

Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our most capable vision model. Available in research preview on the Pro, Max, Team, and Enterprise plans, rolling out throughout the day.
https://x.com/claudeai/status/2045156267690213649

On the plus side with Opus 4.7, if it does decide to think it produces BY FAR the best Sparks unicorn* ever, even non-thinking is pretty good, if not great. * This is created using TikZ, which is a language built for scientific diagrams & very much not for drawing. The original
https://x.com/emollick/status/2044880350237626844

GPT-5.5 takes OpenAI back to the clear number one in AI. OpenAI’s new model tops the Artificial Analysis Intelligence Index by 3 points, breaking a three-way tie with Anthropic and Google OpenAI gave us pre-release access to test all five reasoning effort levels: xhigh, high,
https://x.com/ArtificialAnlys/status/2047378419282034920

GPT-Image-2 takes #1 in every single Text-to-Image category — all 7 of them. Surpassing the next leading model (Nano-banana-2 with web-search) across the board. Here’s the drill-down on improvements vs. its predecessor, GPT-Image-1.5-High-Fidelity: – #1 Product, Branding &
https://x.com/arena/status/2046670705958551938

I’ve been an early tester of GPT-5.5, and it destroyed the “”GPT Plays Pokémon FireRed”” benchmark. GPT-5.4 never finished the game, it got stuck in a loop, reloading the last save and retrying the final rival fight over and over. GPT-5.5 not only beat it on the first try, but did
https://x.com/clad3815/status/2047392779006013833?s=12

DeepSeek is once again the open-source king and it’s competitive with frontier models 1st or 2nd place on 12/22 benchmarks
https://x.com/scaling01/status/2047512176856899985

GPT-5.5, not fully saturating the TikZ unicorn test yet but getting awfully close … (yes this is actual TikZ code, I personally find it so unbelievable that I’m putting the code below for anyone to verify for themself)
https://x.com/sebastienbubeck/status/2047383628922167390?s=46

GPT-ImageGen-2 did this in one shot, with just the prompt “”turn all of Tennyson’s Ulysses into a comic, across as many pages as needed. make it great, include the full text”” 10 pages, though it did use what seems to be the ImageGen-2 ‘s preferred “”spackled drawing”” style 1/
https://x.com/emollick/status/2046843402021380556

Kimi K2.6 Tech Blog: Advancing Open-Source Coding
https://www.kimi.com/blog/kimi-k2-6

Kimi K2.6 autonomously overhauled exchange-core, an 8-year-old open-source financial matching engine. Over a 13-hour execution, the model iterated through 12 optimization strategies, initiating over 1,000 tool calls to precisely modify more than 4,000 lines of code. Acting as
https://x.com/Kimi_Moonshot/status/2046531057147933137

Exciting news – GPT-Image-2 by @OpenAI has claimed the #1 spot across all Image Arena leaderboards! A clean sweep with a record-breaking +242 point lead in Text-to-Image – the largest gap we’ve seen to date. – #1 Text-to-Image (1512), +242 over #2 (Nano-banana-2 with web-search
https://x.com/arena/status/2046670703311884548

GPT Image Generation Models Prompting Guide
https://developers.openai.com/cookbook/examples/multimodal/image-gen-models-prompting-guide

Introducing ChatGPT Images 2.0 | OpenAI
https://openai.com/index/introducing-chatgpt-images-2-0/

Introducing ChatGPT Images 2.0 A state-of-the-art image model that can take on complex visual tasks and produce precise, immediately usable visuals, with sharper editing, richer layouts, and thinking-level intelligence. Video made with ChatGPT Images
https://x.com/OpenAI/status/2046670977145372771

Arena Trends: Text-to-Image, Jan 2026 – Apr 2026 For most of the year, @GoogleDeepMind and @OpenAI traded the top spot within a tight margin – GPT-Image vs. Nano Banana – with the rest of the field clustered below 1,200. Today, GPT-Image-2 breaks away with a score of 1,512, 242
https://x.com/arena/status/2046690103515648061

We’ve post trained a model on top of Qwen that achieves Pareto optimality on accuracy-cost curves. Unlike our previous post trained models, this model has been trained to be good at search and tool calls simultaneously, allowing us to unify the tool call router and
https://x.com/AravSrinivas/status/2047019688920756504

🚀 Meet Qwen3.6-27B, our latest dense, open-source model, packing flagship-level coding power! Yes, 27B, and Qwen3.6-27B punches way above its weight. 👇 What’s new: 🧠 Outstanding agentic coding — surpasses Qwen3.5-397B-A17B across all major coding benchmarks 💡 Strong
https://x.com/Alibaba_Qwen/status/2046939764428009914

Qwen 3.6-Max-Preview solves AIME-2026 #15 after like 30 minutes of thinking, but on first try. Preview or not, it’s more baked than DeepSeek-Expert. Other tests validate this impression. It doesn’t screw up. Alibaba Qwen is, after all, a frontier lab.
https://x.com/teortaxesTex/status/2046166258853269990

A real issue with the current state of our knowledge on the work implications of AI is that there was a genuine discontinuity in AI ability with the rise of practical agentic systems in 2026. We were starting to get a picture of the impact of chatbots, no real data on agents.
https://x.com/emollick/status/2044576806926512446

Benchmarking Inference Engines on Agentic Workloads | Applied Compute
https://www.appliedcompute.com/research/inference-benchmark

Cloud agent infrastructure has a lot of moving parts: VM isolation, session persistence, environment provisioning, orchestration, integrations. Each one is its own engineering challenge. In this post, we break down what it takes to build cloud agent infrastructure from
https://x.com/cognition/status/2047392064355377194

Glad to be a part of this initiative to develop open-world evaluations for AI. We need the ability to assess just how capable agents are becoming in order to anticipate and respond to the impact they can have on real world systems and transactions. An agent that can successfully
https://x.com/ghadfield/status/2045245020429570505

had a blast on the pod (first one kinda nervous 😅), big shoutout to @himanshustwts just felt like us riffing on agent engineering, open source & research ❤️ some fun highlights – working backwards from the model’s capabilities/flaws and building systems (a harness) around them
https://x.com/Vtrivedy10/status/2046942634321559707

Have AI capabilities accelerated? On 3 out of the 4 AI capability metrics we investigated, we found strong evidence of acceleration, around when reasoning models emerged.
https://x.com/EpochAIResearch/status/2045205780916560010

Let’s talk content faithfulness. Four days ago, we launched ParseBench, the first document OCR benchmark for AI agents. Its most fundamental metric asks: did the parser capture all the text, in order, without making things up? We grade three failure modes with 167K+ rule-based
https://x.com/llama_index/status/2045145054772183128

Let’s talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDataPointMatch. Most document look at a chart and OCR the caption. Agents need the actual numbers. That’s the gap between “”OCR’d the
https://x.com/llama_index/status/2046586730879283227

This is crazy. ml-intern just passed the @huggingface internship test in 15 minutes. The task: replicate a research baseline from a DeepMind paper on test-time compute scaling. Here’s what the agent did: – Read the DeepMind paper, dug into Appendix E, picked the right scoring
https://x.com/akseljoonas/status/2047332440025321796

what do we think about perplexity/NLL eval for post trained models? cursor composer 2 did use it to choose the starting model (but didn’t end up choosing the best one according to NLL)
https://x.com/eliebakouch/status/2045115926123520100

Yesterday, we announced CRUX, a project that aims to conduct regular “open-world evaluations,” where we will be testing the ability of AI agents to complete long-horizon tasks in messy, real-world environments. @sayashk’s post dives into the details; here are a few of my own
https://x.com/PKirgis/status/2045265295649231354

This paper makes a strong case for open-world evaluations as a complement to traditional benchmarks, particularly for realistic, long-horizon, open-ended settings! Glad the AISI SoE team could contribute to this effort.
https://x.com/CUdudec/status/2045139195220431022

ParseBench is the first benchmark to include VLM chart understanding 📊📈📉 over enterprise documents. 🟠 Existing benchmarks (ChartQA, ChartXiv) test over charts specifically and not the chart’s inclusion in the overall document. Also doesn’t contain references to real-world
https://x.com/jerryjliu0/status/2046725527806021937

// Self-Evolving Agent Protocol // One of the more interesting papers I read this week. (bookmark it if you are an AI dev) The paper introduces Autogenesis, a self-evolving agent protocol where agents identify their own capability gaps, generate candidate improvements,
https://x.com/omarsar0/status/2045241905227915498

// Stateless Decision Memory for Enterprise AI Agents // Most of the interesting AI agent papers right now are about capability. This one is about plumbing, and it’s probably more important than it looks. One of the few agent memory works focusing on production-grade
https://x.com/omarsar0/status/2047325132096758228

🔒`deepagents deploy` now supports custom auth: turn a single deployment into a multi-tenant platform. per-user resource and thread isolation, access control, and your own auth provider, all without spinning up extra infrastructure.
https://x.com/sydneyrunkle/status/2046643201738449076

A good AGENTS.md is a model upgrade. A bad one is worse than no docs at all. | Augment Code
https://www.augmentcode.com/blog/how-to-write-good-agents-dot-md-files

A new protocol that can become a useful part of agentic workflows – Autogenesis Protocol (AGP) It’s all about structuring self-improvement of agentic systems to make it safe, clear and continuous. The idea is to separate what evolves from how changes happen That’s why AGP has
https://x.com/TheTuringPost/status/2046254041051943157

A survey that deserves your attention – Externalized Intelligence in LLM Agents It explains a shift from intelligence inside the model weights to intelligence in the system around it, showing: • How capability is increasingly coming from: – memory systems – persistent state –
https://x.com/TheTuringPost/status/2045988056088678667

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
https://agent-tars-world.github.io/-/

Agents can’t choose between structure and flexibility
https://frontierai.substack.com/p/agents-cant-choose-between-structure

AI Agent Frameworks Workshop | AWS Marketplace
https://pages.awscloud.com/awsmp-gro-hjua-webinar-mss-module-4-multi-agent-architectures-ai-series.html?trk=cac541b5-48a5-40fe-b9f3-5ced632e84d2&sc_channel=el

Coding Agents | Band
https://www.band.ai/for-coding-agents

developing an agent is a harness problem. deploying an agent is a runtime problem. here’s a guide to the runtime capabilities that keep agents running in production and the infrastructure that enables self improvement!
https://x.com/sydneyrunkle/status/2046284044942397744

Evals ~= Environments…they’re one of the best investments a team can make for improving agents Step 0: Turn On Tracing for Agents Step 1: Point compute at Traces to understand agent behavior, segment useful tasks, and isolate error modes Step 2: Turn Trace Data into
https://x.com/Vtrivedy10/status/2047362615836336473

Governing multi-agent systems at scale is where complexity explodes An upcoming live session with @rungalileo co-founder @YashSheth46 and @crewAIInc founder @joaomdmoura will help you master it. You’ll learn how to: – Enforce safety and security policies in agents – Steer
https://x.com/TheTuringPost/status/2044915828542603502

Great paper on self-improving agents. Why? We need to think more deeply about AI agent system design. The protocol specifies a framework for proposing, assessing, and committing improvements with auditable lineage and rollback. Visual below (courtesy of my research agent).
https://x.com/omarsar0/status/2045956901750399374

Great to see more people adopting ADP to standardize agent trajectories! If you’re building large agent training datasets, please reach out and we’ll be happy to help.
https://x.com/gneubig/status/2046963826109689983

i haven’t seen a model that just works across agent harnesses. seems like it should exist. great opportunity for open-weight models. any thoughts?
https://x.com/omarsar0/status/2047006936306962754

Introducing Deep Max: State-of-the-Art Agentic Search | Exa Blog
https://exa.ai/blog/deep-max

Introducing ml-intern, the agent that just automated the post-training team @huggingface It’s an open-source implementation of the real research loop that our ML researchers do every day. You give it a prompt, it researches papers, goes through citations, implements ideas in GPU
https://x.com/akseljoonas/status/2046543093856412100

LLM agents are assumed to integrate unexpected environmental observations into their reasoning. It turns out they don’t. We added the complete task solution into agent environments as a file or an API endpoint, and measured whether agents act on what they discover. They almost
https://x.com/LeonEnglaender/status/2046621862214488473

LLM agents loop, drift, and get stuck on hard reasoning tasks up to 30% of the time. Current fixes are either too blunt (hard step limits) or too expensive (LLM-as-judge adding 10-15% overhead per step). New research proposes a smarter middle ground. The work introduces the
https://x.com/omarsar0/status/2045139481779696027

RLM means notebooks are gonna be back (I hope). Agent driving a REPL with interleaved prose. The exact backend the nb interface is for. LOTS of have been swirling around the idea, but RLM solidified it and made it work. So I hope to see big NB releases soon! p.s. If you
https://x.com/isaac_flath/status/2046588093399019918

The Agent Harness Is a Shell | blog | inference.sh
https://inference.sh/blog/opinions/harness-is-a-shell

The idea of training LLMs to manage their own KV cache is super interesting to me. The recent neural garbage collection (NGC) paper was a great read on this topic. Reasoning models / agents obviously need long sequences to handle complex reasoning, long horizon tasks, tool
https://x.com/cwolferesearch/status/2047476297031631102

Things don’t always go to plan when bringing agents into production. ‘deepagents deploy’ is purpose-built around the challenges teams face when deploying. ✅ Every infrastructural consideration is mapped to a purpose-built runtime capability. No need to rebuild components from
https://x.com/LangChain/status/2046275653335462128

we just shipped support for subagents with `deepagents deploy`! add an agents/ dir to your project with an AGENTS.md per specialized subagent. subagents are great for task delegation with isolated/optimized context
https://x.com/sydneyrunkle/status/2045209395881980276

we want to help you navigate all the stuff that comes with shipping an agent to production – so we wrote a guide! this guide packages learnings across our + customer deployments, building open source agent infra/frameworks/examples, debugging infra errors, and building great
https://x.com/Vtrivedy10/status/2046280543978057892

We’re launching the beta for our new commercial AI product: Sakana Fugu 🐡, a multi-agent orchestration system! Blog:
https://t.co/c7wlMoX1W2 Fugu hits SOTA on SWE-Pro, GPQA-D, and ALE-Bench, and has been our internal secret weapon. It dynamically coordinates frontier models,
https://x.com/SakanaAILabs/status/2047479445209145785

Moonshot AI: “”Our RL infra team used a K2.6-backed agent that operated autonomously for 5 days, managing monitoring, incident response, and system operations, demonstrating persistent context, multi-threaded task handling, and full-cycle execution from alert to resolution””
https://x.com/scaling01/status/2046250343479054540

A lot of bugs that folks may have hit yesterday when first trying Opus 4.7 are now fixed. Thanks for bearing with us🙏
https://x.com/alexalbert__/status/2045159041283064095

A major lesson to take away from Opus 4.7 is that, while there is a lot of arguments about implementation choices and personality, models keep improving measurably on economically important tasks with each release (it has been two months since Opus 4.6), with no signs of slowdown
https://x.com/emollick/status/2045314251804324080

Changes in the system prompt between Claude Opus 4.6 and 4.7
https://simonwillison.net/2026/Apr/18/opus-system-prompt/

Claude Opus 4.7 by @AnthropicAI advances the price-performance Pareto frontier in both Code and Text Arena! This makes Claude Opus 4.7 now the only model from a US lab that remains on the Pareto frontier for Code Arena.
https://x.com/arena/status/2045206342173086156

Claude Opus 4.7 by @AnthropicAI also lands at #1 and #3 in the Text Arena. Opus 4.7 Thinking ranks #1 across major categories: – #1 Overall 1505, +9 points over Muse Spark – #1 Expert 1561, +19 points over 4.6 Thinking – #1 Coding 1567, +23 points over 4.6 Thinking – #2
https://x.com/arena/status/2045177497378316597

I have found that asking for a sestina regularly triggers Opus 4.7’s safety guardrails. The forbidden poetic form!
https://x.com/emollick/status/2044863531900686775

I’ll give Anthropic credit for moving quickly. Opus 4.7 Adaptive Thinking now triggers thinking much more often, including for the tasks it failed at yesterday. That also means it is doing a lot more web search. So far, a large improvement in output quality on non-coding tasks.
https://x.com/emollick/status/2045147490316374414

Introducing the next-gen AI for design and creation — Genspark Build 🚀 Powered by Claude Opus 4.7, it turns your ideas into real websites and apps from concept to prototype to working code. Now in Public Preview: all Plus and Pro users get 3 days of zero-credit access (April
https://x.com/genspark_ai/status/2046610783203975539?s=20

With max thinking Opus 4.7 is quite impressive, with a real sense of style In two prompts: “”implement the Tower of Babel, in 3D, in as sophisticated and visually interesting a way as possible. It should be interactive”” and then “”make it better.”” Play:
https://x.com/emollick/status/2044966818339594252

[2604.14228] Dive into Claude Code: The Design Space of Today’s and Future AI Agent Systems
https://arxiv.org/abs/2604.14228

Claude Opus 4.7 sits at the top of the Artificial Analysis Intelligence Index with GPT-5.4 and Gemini 3.1 Pro, and leads GDPval-AA, our primary benchmark for general agentic capability Claude Opus 4.7 scores 57 on the Artificial Analysis Intelligence Index, a 4 point uplift over
https://x.com/ArtificialAnlys/status/2045292578434875552

Claude Opus 4.7 from @AnthropicAI takes #1 in Vision & Document Arena! In Document Arena: Opus 4.7 lands +4 points over Opus-4.6 and +45 over the next non-Anthropic model, GPT-5.4 (#6). This is huge ~70 pts lead over Muse Spark and Gemini-3.1-Pro. Real world research work like
https://x.com/arena/status/2046224760657658239

Errata: Yesterday, we discovered that some of our chip owner estimates were stale–Oracle’s Nvidia compute wasn’t subtracted from “”Other”” as intended. This inflated “Other” by ~1M H100e, 5% of the overall total. In our corrected figures, hyperscalers hold 71% of world AI compute.
https://x.com/EpochAIResearch/status/2044832224198226177

Feb 23: OpenAI tells us all to switch to SWE-Bench Pro because it’s less contaminated April 23: OpenAI buries SWE-bench Pro in the appendix of the GPT-5.5 release after getting mogged by Claude
https://x.com/chowdhuryneil/status/2047416077622395025?s=46

LLMs are still not consistent judges of qualitative work, and small changes to how that work is presented affect outcomes. Better harnessing and methods (multiple judging runs with randomized orders, etc) would certainly help, but the jagged frontier is very much still real.
https://x.com/emollick/status/2046668472449405152

MiMo-V2.5 by @XiaomiMiMo is now live on Arena. Evaluate it across Text, Vision & Code Arena – Pro versions available specifically in Text & Code. Start prompting and voting in Battle mode. Scores incoming.
https://x.com/arena/status/2047013664142893286

You can watch the accelerated shipping from the AI labs to get a feeling of what AI-driven product development makes possible. A tremendous number of products are coming out, many of them are really good (with rough edges)… but we also don’t have the capacity to absorb it all.
https://x.com/emollick/status/2045170892951474263

please show me the >6 hour METR time horizons, WeirdqML, GSO, PPBench, LisanBench, SimpleBench, RLI or ARC-AGI-2 scores if you are so confident that open-source models are at GPT-5.4 or Opus 4.5+ level there’s still a big gap and it’s likely getting bigger
https://x.com/scaling01/status/2046565191903511010

Redwood Research presents LinuxArena – 20 live production environments for AI agents – Frontier models achieve ~23% undetected sabotage vs. trusted monitors – Useful work ≈ attack surface → sandboxing fails, monitoring is essential
https://x.com/arankomatsuzaki/status/2046070569758752984

Exciting news – Claude Opus 4.7 from @AnthropicAI takes #1 in Code Arena! +37 points over Opus-4.6 and +46 over the next non-Anthropic model, GLM-5.1 (#4). Massive ~130 pts lead over GPT-5.4 and Gemini-3.1-Pro. #1 on both React and HTML leaderboards. Code Arena evaluates
https://x.com/arena/status/2045177492936532029

The continuing gap between the capabilities of Gemini Pro 3.1 (very good model) and the capabilities of the Gemini app/website is odd. The model can do what Claude/GPT can do, but there is a minimal harness for tools (file creation, research etc), no auditable CoT/actions, manual
https://x.com/emollick/status/2045909435315323321

I think the adaptive thinking requirement in Claude Opus 4.7 is bad in the ways that all AI effort routers are bad, but magnified by the fact that there is no manual override like in ChatGPT. It regularly decides that non-math/code stuff is “”low effort”” & produces worse results.
https://x.com/emollick/status/2044864822076969268

Opus 4.7 better than Opus 4.6 but can’t beat Gemini 3.1 Pro and GPT-5.4 on LiveBench
https://x.com/scaling01/status/2045178622617498084

Opus 4.7 scores 156 on ECI, our tool for combining multiple benchmarks onto a single scale. This puts it a bit ahead of Opus 4.6 and a bit behind only GPT-5.4, Gemini 3.1 Pro, and GPT-5.4 Pro. Thread with individual scores and commentary.
https://x.com/EpochAIResearch/status/2046631622909558857

This is one of the most interesting and creative benchmarks to reveal LLMs’ biases (it’s open source and the results are quite surprising)
https://x.com/TheTuringPost/status/2044572602807534015

DeepSeek V4 Flash Thinking at 284B parameters (13B activated) shifts the Text Pareto frontier with $0.14 input / $0.28 output per MToken. Congrats again to @DeepSeek_AI on the open model progress!
https://x.com/arena/status/2047524055679729885

Exciting news – DeepSeek V4 Pro is in the Arena with 1.6T parameters (49B activated) alongside V4 Flash at 284B parameters (13B activated). Both support 1M token context. It’s a major leap over DeepSeek V3.2! Code Arena: – DeepSeek V4 Pro (thinking): #3 open model (#14 overall),
https://x.com/arena/status/2047518354903359697

In Vending-Bench Arena (the multiplayer version of Vending-Bench with competition dynamics), GPT-5.5 actually beats Opus 4.7. Opus 4.7 showed similar behavior to Opus 4.6: lying to suppliers and stiffing customers on refunds. GPT-5.5’s tactics were clean, and it still won.
https://x.com/andonlabs/status/2047377260412649967?s=46

The GPT-5.5 model family completely dominates the cost-performance frontier on the Artificial Analysis Index
https://x.com/scaling01/status/2047380890402123928

We find that GPT 5.4 over-edits the most while Opus 4.6 over-edits the least. Next, we prompt the models with the explicit instruction to preserve as much of the original code as possible, and find that while this instruction does help performance, the performance gains are
https://x.com/nrehiew_/status/2046963041338855791

🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. 🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world’s top closed-source models. 🔹 DeepSeek-V4-Flash: 284B total / 13B active params.
https://x.com/deepseek_ai/status/2047516922263285776

I think Artificial Analysis does a good job overall and provides transparency in benchmarking, but GDPval-AA is not a good benchmark and needs to stop being reported. It is Gemini 3.1 judging other models on the output of the public questions from GDPval, which tells us nothing.
https://x.com/emollick/status/2045305504679842279

Muse Spark is #3 on ClawEval, ahead of GPT-5.4 and Gemini 3.1 Pro. It is honestly a surprisingly agentic model.
https://x.com/alexandr_wang/status/2045348588734066794

🎉 Congrats to the Moonshot team on Kimi K2.6 — day-0 support on vLLM 0.19.1. • 1T total / 32B active MoE — 384 experts, 8 routed + 1 shared • MLA attention, 256K context • Native multimodal: MoonViT vision encoder + video input • Native INT4 quantization • Interleaved
https://x.com/vllm_project/status/2046251287206035759

FINALLLY FINALLY it is here. V4-flash: all the way back to V2 prices, only now with 1M V4-pro: roughly Kimi/GLM/MiMo competitor Chat prefix completion and FIM back – thank you! Missed this forever but what can they do?
https://x.com/teortaxesTex/status/2047508587883250112

Kimi 2.6 Thinking seems very good for an open weights model, but many rough edges compared to closed SoTA. The Lem Test resulted in a 74 page thinking trace… and an okay-ish answer. It did an okay TiKZ unicorn, an adequate twigl shader for a neogothic city in the waves, etc.
https://x.com/emollick/status/2046411222354989189

Kimi K2.6 + DFlash: 508 tok/s on 8x MI300X 5.6x throughput improvement over baseline autoregressive serving 90 tok/s → 508 tok/s on the same hardware, same model, zero quality loss
https://x.com/HotAisle/status/2046620289984057634

Kimi K2.6 demonstrates strong long-horizon coding in complex engineering tasks: Kimi K2.6 successfully downloaded and deployed the Qwen3.5-0.8B model locally on a Mac. By implementing and optimizing model inference in Zig–a highly niche programming language–it demonstrated
https://x.com/Kimi_Moonshot/status/2046531052957569211

Kimi K2.6 has landed, and it is live on Baseten! We have baked in multiple inference optimizations so that you can leverage Kimi K2.6 in production right away. To run Kimi K2.6, Baseten uses: -> The Baseten Inference Stack with advanced optimizations, including KV-aware routing
https://x.com/baseten/status/2046263526281576573

Kimi K2.6 helped us rewrite kernels; it worked like a charm 🙂
https://x.com/Yulun_Du/status/2046252918526071017

Kimi K2.6 is live on OpenRouter! @Kimi_Moonshot’s new model is a long-horizon coding model built for sustained agentic work. It behaves more like a systems engineer than a chatbot, with the stamina to decompose, execute, and optimize complex tasks. Try it in all your favorite
https://x.com/OpenRouter/status/2046259590774571199

Kimi K2.6 is now available in Windsurf! Available for free for the next 2 weeks for Pro, Teams, and Max users.
https://x.com/windsurf/status/2046686574793154996

Kimi K2.6 now in OpenCode — Go included
https://x.com/opencode/status/2046275886396125680

Kimi K2.6 was released 1h ago, and it looks amazing! Here it’s running with MLX (mlx-vlm) on two M3 Ultras (full 1T param VLM) 🔥
https://x.com/pcuenq/status/2046283942689456297

Meet Kimi K2.6: Advancing Open-Source Coding 🔹Open-source SOTA on HLE w/ tools (54.0), SWE-Bench Pro (58.6), SWE-bench Multilingual (76.7), BrowseComp (83.2), Toolathlon (50.0), Charxiv w/ python(86.7), Math Vision w/ python (93.2) What’s new: 🔹Long-horizon coding – 4,000+
https://x.com/Kimi_Moonshot/status/2046249571882500354

Moonshot AI launches Kimi K2.6 on Kimi Chat and APIs
https://www.testingcatalog.com/moonshot-ai-launches-kimi-k2-6-on-kimi-chat-and-apis/

Qwen3.6-27B can now run locally! 💜 Run on 18GB RAM via Unsloth Dynamic GGUFs. Qwen3.6-27B surpasses Qwen3.5-397B-A17B on all major coding benchmarks. GGUFs:
https://t.co/ykKgwh2zI9 Guide:
https://x.com/UnslothAI/status/2046959757299487029

Ran Qwen3-8B (8.2B dense, open) on LongCoT-Mini. Vanilla: 0/507. dspy.RLM: 33/507 (6.5%). Same model. Same weights. No fine-tuning. The scaffold is doing 100% of the lifting. Context: leaderboard’s smallest open MoE is GLM-4.7 at 358B total / 32B active params. Qwen3-8B is ~4x
https://x.com/raw_works/status/2045208764509470742

these questions are silly Kimi > all other open-source models tho
https://x.com/scaling01/status/2046591683198906542

We’re open-sourcing FlashKDA — our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. Achieves 1.72×-2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention. Explore on github:
https://x.com/Kimi_Moonshot/status/2046607915424034839

OpenClaw 2026.4.20 🦞 🧠 Kimi K2.6 support + provider-aware /think 💬 BlueBubbles iMessage sends + tapbacks fixed ⏰ Cron state/delivery cleanup 🔐 Gateway pairing + plugin startup hardening Less haunted. More useful.
https://x.com/openclaw/status/2046686809367708123

Kimi K2.6 wrote an inference engine for Qwen3.5 0.5B in Zig and managed to beat LM Studio’s token per second by 20%, running for 12 hours and with 4000+ tool calls
https://x.com/nrehiew_/status/2046254256194474221

I find that open weights models over-perform on benchmarks compared to actual real-world usage, and Kimi feels like no exception. For example, a small amount of use will show that Kimi is not as good as Claude Opus 4.6, which it beats on the benchmarks. Still a good model, tho!
https://x.com/emollick/status/2046430301593751797

Image models tend to get much more stuck on a particular direction than text models, requiring clearing the context window fairly often. PerfectSquashBench is my new measure of how image models anchor. The squash remains merely fine after many attempts.
https://x.com/emollick/status/2047073009312121000

Qwen
https://qwen.ai/blog?id=qwen3.6-max-preview

Qwen3.6 Plus lands at #7 in Code Arena with a score of 1476 – up +16 points since the Preview. The new score also moves @AlibabaGroup to #3 lab in Code Arena. In the Text Arena, Qwen3.6 Plus lands at #36, a +13 point improvement since Preview. Congrats to the Qwen team on the
https://x.com/arena/status/2046268995163258958

🚀 Introducing Qwen3.6-Max-Preview, an early preview of our next flagship model Highlights: ⚡️ Improved agentic coding capability over Qwen3.6-Plus 📖 Stronger world knowledge and instruction following 🌍 Improved real-world agent and knowledge reliability performance Smarter,
https://x.com/Alibaba_Qwen/status/2046227759475921291

Guys, I am absolutely astounded. The Qwen 3.6 27b is like a jump to Qwen 4 from Qwen 27B 3.5. I just did a full suite of front end design tests and agentic benchmarks, made entirely by it. VERDICT: They’re so much better than I thought they’d be, like I’m completely astounded. I
https://x.com/KyleHessling1/status/2046986423736451327

llama-server -hf ggml-org/Qwen3.6-27B-GGUF –spec-default
https://x.com/ggerganov/status/2046988075302064209

Qwen 3.6 27B model is available on Ollama! Use it with all the integrations in Ollama or chat with the model. Chat with the model: ollama run qwen3.6:27b OpenClaw: ollama launch openclaw –model qwen3.6:27b Claude Code: ollama launch claude –model qwen3.6:27b More
https://x.com/ollama/status/2047066252523507916

We ran Qwen3.6-35B-A3B GGUF KLD benchmarks of all our dynamic quants and other providers. 1. Nearly all Unsloth quants for mean KLD, 90%, 99.9% KLD are on the Pareto Frontier for KLD vs Disk Space. 2. MXFP4_MOE is an outlier for all. 3. We’ll also make some smaller quants soon!
https://x.com/danielhanchen/status/2045169369723064449

Sharing my current setup to run Qwen3.6 locally in a good agentic setup (Pi + llama.cpp). Should give you a good overview of how good local agents are today: # Start llama.cpp server: llama-server \ -hf unsloth/Qwen3.6-35B-A3B-GGUF:Q4_K_XL \ –jinja \
https://x.com/victormustar/status/2045068986446958899

[2604.15804] Qwen3.5-Omni Technical Report
https://arxiv.org/abs/2604.15804

🎉 Day-0 vLLM support for Qwen3.6-27B! Congrats to @Alibaba_Qwen on the new 27B dense model release. Looking forward to more of the Qwen3.6 series. 👀 📖 Recipe:
https://x.com/vllm_project/status/2046943674890871019

LM Performance:With only 27B parameters, Qwen3.6-27B outperforms the Qwen3.5-397B-A17B (397B total / 17B active, ~15x larger!) on every major coding benchmark — including SWE-bench Verified (77.2 vs. 76.2), SWE-bench Pro (53.5 vs. 50.9), Terminal-Bench 2.0 (59.3 vs. 52.5), and
https://x.com/Alibaba_Qwen/status/2046939775924584577

Qwen3.5-Omni Technical Report | alphaXiv
https://www.alphaxiv.org/abs/2604.15804

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
https://simonwillison.net/2026/Apr/22/qwen36-27b/

Qwen3.6-35B-A3B just dropped. Red Hat AI has an NVFP4 quantized checkpoint ready. 35B params, 3B active, quantized with LLM Compressor. Preliminary GSM8K Platinum: 100.69% recovery (slightly above baseline). Early release. Let us know what you think!
https://x.com/RedHat_AI/status/2045153791402520952

The new Qwen3.6-27B just gave me definitely the best pelican riding a bicycle I’ve had from a 16.8GB model file!
https://x.com/simonw/status/2046995047720378458

We then experiment with 4 different training methods for the Minimal Code Editing task using Qwen3 4B. We find that SFT only works when trained on the same set of evaluation corruptions. It collapses otherwise, indicating that it fails to learn the general minimal coding style
https://x.com/nrehiew_/status/2046963050427879488

VLM Performance:Qwen3.6-27B is natively multimodal, supporting both vision-language thinking and non-thinking modes in a single unified checkpoint — the same as Qwen3.6-35B-A3B. It handles images and video alongside text, enabling multimodal reasoning, document understanding,
https://x.com/Alibaba_Qwen/status/2046939788184547610

@eliebakouch I think NLL on neutral text does not actually change much (or very variably) from assistant tuning and heavy RL, as it mainly shifts the probabilities of few tokens in special contexts. If you have a decent webtext prefix, the prediction should be close to base model’s
https://x.com/teortaxesTex/status/2045139476972745120

>be application engineer >live demo in 48 hours >customs f#cked up >microwave stuck somewhere in transit >no microwave = no demo >nice >saturday morning >drive to nearest electronics store >buy random microwave >whatever is in stock >take photos >front back inside details
https://x.com/IlirAliu_/status/2046652192518533355

🔧 The tricky part: naïvely casting BF16 group scales to FP8 dropped the quality. Our fix: quantize scales per-channel (outer vector scaling) + rescale by 1/8 to avoid FP8 clipping. Result: >99.5% of W4A16 accuracy recovered on Command A & Cohere MoE. Paired with a CUTLASS
https://x.com/cohere/status/2047052560553681183

13+ Attention mechanisms you should know ▪️ Self-attention ▪️ Cross-attention ▪️ Causal attention ▪️ Linear Attention ▪️ Softmax attention ▪️ Sliding Window (local attention) ▪️ Global attention ▪️ FlashAttention ▪️ Multi-Head Attention (MHA) ▪️ Multi-Query Attention (MQA) ▪️
https://x.com/TheTuringPost/status/2045868425017442347

30B -> 300T tokens per month YoY @togethercompute
https://x.com/vipulved/status/2047183589222273231

3a8d990bc90098038eabd77b0d12ff636ed58d50.pdf

Click to access 3a8d990bc90098038eabd77b0d12ff636ed58d50.pdf

AI 101: What Is a Token (and why it runs AI)?
https://x.com/TheTuringPost/status/2046660466441671000

AI and Cloud Infrastructure Provider | Runpod
https://www.runpod.io/?inflect=&targetid=kwd-461794446387&adgroupid=189724568342&loc_interest=&loc_physical=9012242&hsa_acc=4558579452&hsa_cam=23516214656&hsa_grp=189724568342&hsa_ad=797799116312&hsa_src=g&hsa_tgt=kwd-461794446387&hsa_kw=runpod&hsa_mt=e&hsa_net=adwords&hsa_ver=3&gad_source=1&gad_campaignid=23516214656&gbraid=0AAAAAoZSBmm_mSAwIfhLEqmrXZN5buVzH&gclid=CjwKCAjw7vzOBhBxEiwAc7WNr_bxZ7lF9pJE1bLS0pbFJ79N79Y2ltJ4iVWWnUb-AdHNb0lPSb2PpRoCPMQQAvD_BwE

Attention → Mamba cross-architecture distillation is real Transformer doesn’t need to stay just a transformer – Apple showed how you can transfer it into a State Space Model (SSM) ▪️ It happens through a linearized-attention intermediate: 1. Distill the Transformer into a
https://x.com/TheTuringPost/status/2045676651376472246

Can a language model learn, end-to-end, what to keep in its own KV cache and what to throw away? Can it learn to forget while it learns to reason? Deep learning’s central lesson: capability emerges from end-to-end optimization, not heuristics/strong inductive biases. But for
https://x.com/michaelyli__/status/2047019938339340602

Can LLMs flip coins in their heads? When prompted to “Flip a fair coin” 100 times, the heads to tails ratio drifts far from 50:50. LLMs can understand what the target probability should be, but generating outputs that faithfully follow a given distribution is a separate problem.
https://x.com/SakanaAILabs/status/2046248967307174225

DSPy 3.2.0 is out! Here are a few highlights: – dspy.RLM improvements around parsing, tool execution, and failure recovery. Expect greater reliability in the bridge between Python and Deno. – @MaximeRivest is leading an ongoing effort to decouple DSPy from LiteLLM. This
https://x.com/isaacbmiller1/status/2046643827247546441

Excited to share our work on production-ready W4A8 inference, now integrated in vLLM! By combining 4-bit weights (low memory) with 8-bit activations (high compute), we hit the sweet spot for both decoding and prefill — up to 58% faster TTFT and 45% faster TPOT vs W4A16 on Hopper.
https://x.com/cohere/status/2047052557915476304

For a decade, we’ve made models wider and deeper–but we’ve barely changed how layers *talk* to each other. Since ResNet’s `x + F(x)` in 2015, the depth residual has been the only highway for inter-layer communication. It’s time to upgrade the staircase. 🧵
https://x.com/lianghui_zhu/status/2045868757869080695

For decades, industrial simulation was a “”waiting game.”” Run a physics test… wait for hours… Fewer iterations, expensive guesswork on the floor, and a massive gap between the digital plan and reality. That “”latency gap”” has been the silent tax on industrial innovation. And
https://x.com/IlirAliu_/status/2047014426025574681

Gzip: 92KB Shared dictionary: 159 bytes 99.5% smaller 🤯 Compression just got a memory. Blog + live demo:
https://x.com/ackriv/status/2045177696506794336

How Token Taxonomy Affects Your Bill
https://x.com/TheTuringPost/status/2047097849385488764

JWT format support for M2M tokens
https://clerk.com/changelog/2026-02-24-m2m-jwt-tokens?dub_id=HUB0AdIJnr0Onr4E

Key to note that AI scientists are not experts on labor. Some other economists active on X doing work on AI & labor: @alexolegimas, @danielrock, @joshgans & @robseamans (among many others) But worth noting that economists don’t have a consensus either:
https://x.com/emollick/status/2045616895592706291

Liked this paper. KVCache transfer between separate prefill and decode clusters is a huge bottleneck. The main idea is to handle short requests in local disaggregated prefill-decode clusters while routing long requests to compute-specialized prefill clusters.
https://x.com/nrehiew_/status/2046201782163095596

Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers Author’s Explanation:
https://t.co/f9C2AVwH2s Overview: Recurrent-depth transformers facilitate implicit multi-hop reasoning by enabling iterative computation across shared layers, addressing the
https://x.com/TheAITimeline/status/2046043384289112408

Looped transformers feel like the revival of Universal Transformers. If there is in fact some relation here, leveraging literature from UT might be interesting.
https://x.com/torchcompiled/status/2046060774083449033

Maximal Brain Damage: Sign-Bit Flips in Neural Networks
https://mkimhi.github.io/DNL/

memory will be the great lock in and everyone knows it, so are rushing to get there first memory should be open!
https://x.com/hwchase17/status/2046308913939919232

New model picker live in latest nightly! ⌘+Shift+M or /model to open ⌘+1..9 to quick select, or navigate with arrow keys 🔍Search ⭐ Favorite Thanks chrono for the initial work copying @r_marked’s design from @t3dotchat
https://x.com/jullerino/status/2046110099262103743

oh, and yeah, the small print: > Due to constraints in high-end computing power, the throughput of V4-Pro is currently very limited. It is expected that after the Ascend 950 super nodes are mass-released in H2 2026, the price of the Pro version will drop significantly.
https://x.com/teortaxesTex/status/2047523707199909977

One of my favorite things about Sakana Fugu is the recursive test-time scaling. When allowed to call itself recursively, it reads its own prior output and spins up corrective workflows on the fly. We are opening up the API for beta testers to try it out:
https://x.com/hardmaru/status/2047483783323283941

One of the premier journals in my field… I think there are very valid reasons to set rules on AI in peer review (including disclosure), but the idea that all AI models steal your data is very 2023. Require people to use enterprise accounts or models with training turned off.
https://x.com/emollick/status/2045311526945394722

Our engineers developed RadixMLP, a technique that exploits the position-wise nature of MLPs, LayerNorms, linear projections, and embeddings to eliminate redundant shared-prefix computations. It makes inference 1.4-1.6x faster in realistic reranking workloads and up to 5x on
https://x.com/baseten/status/2047019335542358284

PD disaggregation… on a single node? @AMD and @EmbeddedLLM just dropped a blog post on the MORI-IO KV Connector and it’s giving vLLM 2.5× higher goodput. Perfect intro to PD disaggregation, explains the fundamentals with real single-node results. Stable buttery decode even at
https://x.com/vllm_project/status/2045381618928582995

Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
https://arxiv.org/html/2604.15039v1

Prefill-as-a-Service: Linear Attention = Final Piece for Cross-DC PD? Insights from Zhihu contributor &Tsinghua Univ. Asst. Prof. ZHANG Mingxing @james0zan👇 Weve discussed heterogeneous PD separation since Mooncake’s launch, especially critical amid GPU shortage chaos.
https://x.com/ZhihuFrontier/status/2046171631228428572

Self-play led to superhuman Go performance, why hasn’t it for LLMs? In practice, long run self-play plateaus like RL. We study why this happens, and build a self-play algorithm that scales better. It solves as many problems with a 7B model as the pass@4 of a model 100x bigger.
https://x.com/LukeBailey181/status/2047340293490724945

The fall of the theorem economy – David Bessis
https://davidbessis.substack.com/p/the-fall-of-the-theorem-economy

The more I program with LLMs using DSPy Signatures as my core, the less bugged I am about wanting more powerful models. I’m also getting a lot more out of good old Sonnet as a result. I think RLM and DSPy are really just showing us examples on how to put these models on a tight
https://x.com/AsfiShaheen/status/2045072599508508914

Today we release all the data sources (and more) in one place, more than 1.4B query-document pairs Plus a new high-quality web dataset built on FineWeb-Edu, replacing the outdated “”common crawl”” splits most mixtures still rely on (thanks @orionweller)
https://x.com/antoine_chaffin/status/2046609260440629588

Train separately, merge together: Modular post-training with mixture-of-experts | Ai2
https://allenai.org/blog/bar

We need a new document that AI labs should release with each new model, besides the model card: a sort of changelog I want to see how & in what way the new model changes, breaks, or improves at a range of individual tasks compared to the earlier models. Increasingly important!
https://x.com/emollick/status/2045140944740122745

What is a token? (basics for understanding AI) 1. Models can’t process raw text as is, so they start with tokens 2. Tokens are small units into which text is broken before being fed into a model 3. A token can be a whole word, part of a word, punctuation, or even a space 4.
https://x.com/TheTuringPost/status/2045478142660464859

When Can LLMs Learn to Reason with Weak Supervision?
https://salmanrahman.net/rlvr-weak-supervision

When LLMs Get Personal – by Joshua Budman
https://joshbudman.substack.com/p/when-llms-get-personal

Will Neural Computers even work? Here’s everything you need to know about them
https://x.com/TheTuringPost/status/2044904774865133932

New paper: “”Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs””. Our system (BLF) matches human superforecasters on ForecastBench, and beats all the top methods (GPT-5, Cassi, Grok 4.20, and Foresight-32B). 🧵
https://x.com/sirbayes/status/2046961503107166689

Nice paper combining the strength of Skills and RAG. Most RAG systems retrieve on every query, whether the model needs help or not. This is wasteful when the model already knows the answer, and often too late when it does not. New research introduces Skill-RAG, a
https://x.com/omarsar0/status/2046249336162632155

An Open Source Dev Kit for AI-native Robotics 📍GitHub:
https://t.co/U11gUQXipc Saying hello to me friend and founder of the robot learning company, @JannikGrothusen. Great for beginners! —- Weekly robotics and AI insights. Subscribe free:
https://x.com/IlirAliu_/status/2045926158030573893

Robots don’t feel, they just follow positions. That’s why they fail in the real world. This paper shows what changes: It lets robots feel force while acting. Not just where to move, but how hard to push. That’s the difference between demo and real work. Thanks for sharing,
https://x.com/IlirAliu_/status/2045049462066676179

🚨🚨New Paper: Generative Insight Anticipation from Scientific Literature Human scientists achieve breakthroughs by “”standing on the shoulders of giants,”” synthesizing profound insights from disparate sources. While LMs show promise in scientific discovery, they currently
https://x.com/Anikait_Singh_/status/2045149764636094839

Our paper landed in Nature Health today! Healthcare is one of the most high-stakes, high-potential applications of AI. So we set out to understand how people actually use it in our AI products today.
https://x.com/mustafasuleyman/status/2044817893460996487

Scientists often make breakthroughs by synthesizing ideas across papers. In our new paper, we ask whether a language model can anticipate this process: given two parent papers, can it generate the core insight of a future paper built on them? 🧵⬇️
https://x.com/JoyHeYueya/status/2045147082546462860

This paper shows people are asking a lot of medical questions of AI already, but we have little evidence of how good or bad this is. Most of the published research uses old models & compares to doctors. How do new models compare to the info people would have gotten without AI?
https://x.com/emollick/status/2045708638317080798

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading