Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: Using the provided reference image, keep the pure white landscape field, exact vertical type hierarchy, galaxy-punchout starfield letterforms, and font contrast between bold condensed grotesque and light geometric sans, but replace ‘HEROES’ with ‘BENCHMARKS’, replace ‘ALESSO’ with ‘HONEST YARDSTICK’ in the light geometric all-caps galaxy treatment, and replace ‘TOVE LO’ with ‘EVAL HARNESS’ in condensed grotesque galaxy treatment, while keeping ‘(we could be)’ and ‘FEATURING.’ unchanged with identical tracking and margins.

Anthropic released Claude design, direct attack on figma and lovable. Anthropic just shipped Claude Design, powered by Claude Opus 4.7., a tool that turns conversations into polished prototypes, pitch decks, and marketing assets. It auto-applies your brand system, lets you
https://x.com/kimmonismus/status/2045162358004216134

Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our most capable vision model. Available in research preview on the Pro, Max, Team, and Enterprise plans, rolling out throughout the day.
https://x.com/claudeai/status/2045156267690213649

On the plus side with Opus 4.7, if it does decide to think it produces BY FAR the best Sparks unicorn* ever, even non-thinking is pretty good, if not great. * This is created using TikZ, which is a language built for scientific diagrams & very much not for drawing. The original
https://x.com/emollick/status/2044880350237626844

GPT-5.5 takes OpenAI back to the clear number one in AI. OpenAI’s new model tops the Artificial Analysis Intelligence Index by 3 points, breaking a three-way tie with Anthropic and Google OpenAI gave us pre-release access to test all five reasoning effort levels: xhigh, high,
https://x.com/ArtificialAnlys/status/2047378419282034920

GPT-Image-2 takes #1 in every single Text-to-Image category — all 7 of them. Surpassing the next leading model (Nano-banana-2 with web-search) across the board. Here’s the drill-down on improvements vs. its predecessor, GPT-Image-1.5-High-Fidelity: – #1 Product, Branding &
https://x.com/arena/status/2046670705958551938

I’ve been an early tester of GPT-5.5, and it destroyed the “”GPT Plays Pokémon FireRed”” benchmark. GPT-5.4 never finished the game, it got stuck in a loop, reloading the last save and retrying the final rival fight over and over. GPT-5.5 not only beat it on the first try, but did
https://x.com/clad3815/status/2047392779006013833?s=12

DeepSeek is once again the open-source king and it’s competitive with frontier models 1st or 2nd place on 12/22 benchmarks
https://x.com/scaling01/status/2047512176856899985

GPT-5.5, not fully saturating the TikZ unicorn test yet but getting awfully close … (yes this is actual TikZ code, I personally find it so unbelievable that I’m putting the code below for anyone to verify for themself)
https://x.com/sebastienbubeck/status/2047383628922167390?s=46

GPT-ImageGen-2 did this in one shot, with just the prompt “”turn all of Tennyson’s Ulysses into a comic, across as many pages as needed. make it great, include the full text”” 10 pages, though it did use what seems to be the ImageGen-2 ‘s preferred “”spackled drawing”” style 1/
https://x.com/emollick/status/2046843402021380556

Kimi K2.6 Tech Blog: Advancing Open-Source Coding
https://www.kimi.com/blog/kimi-k2-6

Kimi K2.6 autonomously overhauled exchange-core, an 8-year-old open-source financial matching engine. Over a 13-hour execution, the model iterated through 12 optimization strategies, initiating over 1,000 tool calls to precisely modify more than 4,000 lines of code. Acting as
https://x.com/Kimi_Moonshot/status/2046531057147933137

Exciting news – GPT-Image-2 by @OpenAI has claimed the #1 spot across all Image Arena leaderboards! A clean sweep with a record-breaking +242 point lead in Text-to-Image – the largest gap we’ve seen to date. – #1 Text-to-Image (1512), +242 over #2 (Nano-banana-2 with web-search
https://x.com/arena/status/2046670703311884548

GPT Image Generation Models Prompting Guide
https://developers.openai.com/cookbook/examples/multimodal/image-gen-models-prompting-guide

Introducing ChatGPT Images 2.0 | OpenAI
https://openai.com/index/introducing-chatgpt-images-2-0/

Introducing ChatGPT Images 2.0 A state-of-the-art image model that can take on complex visual tasks and produce precise, immediately usable visuals, with sharper editing, richer layouts, and thinking-level intelligence. Video made with ChatGPT Images
https://x.com/OpenAI/status/2046670977145372771

Arena Trends: Text-to-Image, Jan 2026 – Apr 2026 For most of the year, @GoogleDeepMind and @OpenAI traded the top spot within a tight margin – GPT-Image vs. Nano Banana – with the rest of the field clustered below 1,200. Today, GPT-Image-2 breaks away with a score of 1,512, 242
https://x.com/arena/status/2046690103515648061

We’ve post trained a model on top of Qwen that achieves Pareto optimality on accuracy-cost curves. Unlike our previous post trained models, this model has been trained to be good at search and tool calls simultaneously, allowing us to unify the tool call router and
https://x.com/AravSrinivas/status/2047019688920756504

🚀 Meet Qwen3.6-27B, our latest dense, open-source model, packing flagship-level coding power! Yes, 27B, and Qwen3.6-27B punches way above its weight. 👇 What’s new: 🧠 Outstanding agentic coding — surpasses Qwen3.5-397B-A17B across all major coding benchmarks 💡 Strong
https://x.com/Alibaba_Qwen/status/2046939764428009914

Qwen 3.6-Max-Preview solves AIME-2026 #15 after like 30 minutes of thinking, but on first try. Preview or not, it’s more baked than DeepSeek-Expert. Other tests validate this impression. It doesn’t screw up. Alibaba Qwen is, after all, a frontier lab.
https://x.com/teortaxesTex/status/2046166258853269990

A real issue with the current state of our knowledge on the work implications of AI is that there was a genuine discontinuity in AI ability with the rise of practical agentic systems in 2026. We were starting to get a picture of the impact of chatbots, no real data on agents.
https://x.com/emollick/status/2044576806926512446

Benchmarking Inference Engines on Agentic Workloads | Applied Compute
https://www.appliedcompute.com/research/inference-benchmark

Cloud agent infrastructure has a lot of moving parts: VM isolation, session persistence, environment provisioning, orchestration, integrations. Each one is its own engineering challenge. In this post, we break down what it takes to build cloud agent infrastructure from
https://x.com/cognition/status/2047392064355377194

Glad to be a part of this initiative to develop open-world evaluations for AI. We need the ability to assess just how capable agents are becoming in order to anticipate and respond to the impact they can have on real world systems and transactions. An agent that can successfully
https://x.com/ghadfield/status/2045245020429570505

had a blast on the pod (first one kinda nervous 😅), big shoutout to @himanshustwts just felt like us riffing on agent engineering, open source & research ❤️ some fun highlights – working backwards from the model’s capabilities/flaws and building systems (a harness) around them
https://x.com/Vtrivedy10/status/2046942634321559707

Have AI capabilities accelerated? On 3 out of the 4 AI capability metrics we investigated, we found strong evidence of acceleration, around when reasoning models emerged.
https://x.com/EpochAIResearch/status/2045205780916560010

Let’s talk content faithfulness. Four days ago, we launched ParseBench, the first document OCR benchmark for AI agents. Its most fundamental metric asks: did the parser capture all the text, in order, without making things up? We grade three failure modes with 167K+ rule-based
https://x.com/llama_index/status/2045145054772183128

Let’s talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDataPointMatch. Most document look at a chart and OCR the caption. Agents need the actual numbers. That’s the gap between “”OCR’d the
https://x.com/llama_index/status/2046586730879283227

This is crazy. ml-intern just passed the @huggingface internship test in 15 minutes. The task: replicate a research baseline from a DeepMind paper on test-time compute scaling. Here’s what the agent did: – Read the DeepMind paper, dug into Appendix E, picked the right scoring
https://x.com/akseljoonas/status/2047332440025321796

what do we think about perplexity/NLL eval for post trained models? cursor composer 2 did use it to choose the starting model (but didn’t end up choosing the best one according to NLL)
https://x.com/eliebakouch/status/2045115926123520100

Yesterday, we announced CRUX, a project that aims to conduct regular “open-world evaluations,” where we will be testing the ability of AI agents to complete long-horizon tasks in messy, real-world environments. @sayashk’s post dives into the details; here are a few of my own
https://x.com/PKirgis/status/2045265295649231354

This paper makes a strong case for open-world evaluations as a complement to traditional benchmarks, particularly for realistic, long-horizon, open-ended settings! Glad the AISI SoE team could contribute to this effort.
https://x.com/CUdudec/status/2045139195220431022

ParseBench is the first benchmark to include VLM chart understanding 📊📈📉 over enterprise documents. 🟠 Existing benchmarks (ChartQA, ChartXiv) test over charts specifically and not the chart’s inclusion in the overall document. Also doesn’t contain references to real-world
https://x.com/jerryjliu0/status/2046725527806021937

A lot of bugs that folks may have hit yesterday when first trying Opus 4.7 are now fixed. Thanks for bearing with us🙏
https://x.com/alexalbert__/status/2045159041283064095

A major lesson to take away from Opus 4.7 is that, while there is a lot of arguments about implementation choices and personality, models keep improving measurably on economically important tasks with each release (it has been two months since Opus 4.6), with no signs of slowdown
https://x.com/emollick/status/2045314251804324080

Changes in the system prompt between Claude Opus 4.6 and 4.7
https://simonwillison.net/2026/Apr/18/opus-system-prompt/

Claude Opus 4.7 by @AnthropicAI advances the price-performance Pareto frontier in both Code and Text Arena! This makes Claude Opus 4.7 now the only model from a US lab that remains on the Pareto frontier for Code Arena.
https://x.com/arena/status/2045206342173086156

Claude Opus 4.7 by @AnthropicAI also lands at #1 and #3 in the Text Arena. Opus 4.7 Thinking ranks #1 across major categories: – #1 Overall 1505, +9 points over Muse Spark – #1 Expert 1561, +19 points over 4.6 Thinking – #1 Coding 1567, +23 points over 4.6 Thinking – #2
https://x.com/arena/status/2045177497378316597

I have found that asking for a sestina regularly triggers Opus 4.7’s safety guardrails. The forbidden poetic form!
https://x.com/emollick/status/2044863531900686775

I’ll give Anthropic credit for moving quickly. Opus 4.7 Adaptive Thinking now triggers thinking much more often, including for the tasks it failed at yesterday. That also means it is doing a lot more web search. So far, a large improvement in output quality on non-coding tasks.
https://x.com/emollick/status/2045147490316374414

Introducing the next-gen AI for design and creation — Genspark Build 🚀 Powered by Claude Opus 4.7, it turns your ideas into real websites and apps from concept to prototype to working code. Now in Public Preview: all Plus and Pro users get 3 days of zero-credit access (April
https://x.com/genspark_ai/status/2046610783203975539?s=20

With max thinking Opus 4.7 is quite impressive, with a real sense of style In two prompts: “”implement the Tower of Babel, in 3D, in as sophisticated and visually interesting a way as possible. It should be interactive”” and then “”make it better.”” Play:
https://x.com/emollick/status/2044966818339594252

Claude Opus 4.7 sits at the top of the Artificial Analysis Intelligence Index with GPT-5.4 and Gemini 3.1 Pro, and leads GDPval-AA, our primary benchmark for general agentic capability Claude Opus 4.7 scores 57 on the Artificial Analysis Intelligence Index, a 4 point uplift over
https://x.com/ArtificialAnlys/status/2045292578434875552

Claude Opus 4.7 from @AnthropicAI takes #1 in Vision & Document Arena! In Document Arena: Opus 4.7 lands +4 points over Opus-4.6 and +45 over the next non-Anthropic model, GPT-5.4 (#6). This is huge ~70 pts lead over Muse Spark and Gemini-3.1-Pro. Real world research work like
https://x.com/arena/status/2046224760657658239

Errata: Yesterday, we discovered that some of our chip owner estimates were stale–Oracle’s Nvidia compute wasn’t subtracted from “”Other”” as intended. This inflated “Other” by ~1M H100e, 5% of the overall total. In our corrected figures, hyperscalers hold 71% of world AI compute.
https://x.com/EpochAIResearch/status/2044832224198226177

Feb 23: OpenAI tells us all to switch to SWE-Bench Pro because it’s less contaminated April 23: OpenAI buries SWE-bench Pro in the appendix of the GPT-5.5 release after getting mogged by Claude
https://x.com/chowdhuryneil/status/2047416077622395025?s=46

LLMs are still not consistent judges of qualitative work, and small changes to how that work is presented affect outcomes. Better harnessing and methods (multiple judging runs with randomized orders, etc) would certainly help, but the jagged frontier is very much still real.
https://x.com/emollick/status/2046668472449405152

MiMo-V2.5 by @XiaomiMiMo is now live on Arena. Evaluate it across Text, Vision & Code Arena – Pro versions available specifically in Text & Code. Start prompting and voting in Battle mode. Scores incoming.
https://x.com/arena/status/2047013664142893286

You can watch the accelerated shipping from the AI labs to get a feeling of what AI-driven product development makes possible. A tremendous number of products are coming out, many of them are really good (with rough edges)… but we also don’t have the capacity to absorb it all.
https://x.com/emollick/status/2045170892951474263

please show me the >6 hour METR time horizons, WeirdqML, GSO, PPBench, LisanBench, SimpleBench, RLI or ARC-AGI-2 scores if you are so confident that open-source models are at GPT-5.4 or Opus 4.5+ level there’s still a big gap and it’s likely getting bigger
https://x.com/scaling01/status/2046565191903511010

Redwood Research presents LinuxArena – 20 live production environments for AI agents – Frontier models achieve ~23% undetected sabotage vs. trusted monitors – Useful work ≈ attack surface → sandboxing fails, monitoring is essential
https://x.com/arankomatsuzaki/status/2046070569758752984

Exciting news – Claude Opus 4.7 from @AnthropicAI takes #1 in Code Arena! +37 points over Opus-4.6 and +46 over the next non-Anthropic model, GLM-5.1 (#4). Massive ~130 pts lead over GPT-5.4 and Gemini-3.1-Pro. #1 on both React and HTML leaderboards. Code Arena evaluates
https://x.com/arena/status/2045177492936532029

The continuing gap between the capabilities of Gemini Pro 3.1 (very good model) and the capabilities of the Gemini app/website is odd. The model can do what Claude/GPT can do, but there is a minimal harness for tools (file creation, research etc), no auditable CoT/actions, manual
https://x.com/emollick/status/2045909435315323321

I think the adaptive thinking requirement in Claude Opus 4.7 is bad in the ways that all AI effort routers are bad, but magnified by the fact that there is no manual override like in ChatGPT. It regularly decides that non-math/code stuff is “”low effort”” & produces worse results.
https://x.com/emollick/status/2044864822076969268

Opus 4.7 better than Opus 4.6 but can’t beat Gemini 3.1 Pro and GPT-5.4 on LiveBench
https://x.com/scaling01/status/2045178622617498084

Opus 4.7 scores 156 on ECI, our tool for combining multiple benchmarks onto a single scale. This puts it a bit ahead of Opus 4.6 and a bit behind only GPT-5.4, Gemini 3.1 Pro, and GPT-5.4 Pro. Thread with individual scores and commentary.
https://x.com/EpochAIResearch/status/2046631622909558857

This is one of the most interesting and creative benchmarks to reveal LLMs’ biases (it’s open source and the results are quite surprising)
https://x.com/TheTuringPost/status/2044572602807534015

DeepSeek V4 Flash Thinking at 284B parameters (13B activated) shifts the Text Pareto frontier with $0.14 input / $0.28 output per MToken. Congrats again to @DeepSeek_AI on the open model progress!
https://x.com/arena/status/2047524055679729885

Exciting news – DeepSeek V4 Pro is in the Arena with 1.6T parameters (49B activated) alongside V4 Flash at 284B parameters (13B activated). Both support 1M token context. It’s a major leap over DeepSeek V3.2! Code Arena: – DeepSeek V4 Pro (thinking): #3 open model (#14 overall),
https://x.com/arena/status/2047518354903359697

In Vending-Bench Arena (the multiplayer version of Vending-Bench with competition dynamics), GPT-5.5 actually beats Opus 4.7. Opus 4.7 showed similar behavior to Opus 4.6: lying to suppliers and stiffing customers on refunds. GPT-5.5’s tactics were clean, and it still won.
https://x.com/andonlabs/status/2047377260412649967?s=46

The GPT-5.5 model family completely dominates the cost-performance frontier on the Artificial Analysis Index
https://x.com/scaling01/status/2047380890402123928

We find that GPT 5.4 over-edits the most while Opus 4.6 over-edits the least. Next, we prompt the models with the explicit instruction to preserve as much of the original code as possible, and find that while this instruction does help performance, the performance gains are
https://x.com/nrehiew_/status/2046963041338855791

🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. 🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world’s top closed-source models. 🔹 DeepSeek-V4-Flash: 284B total / 13B active params.
https://x.com/deepseek_ai/status/2047516922263285776

I think Artificial Analysis does a good job overall and provides transparency in benchmarking, but GDPval-AA is not a good benchmark and needs to stop being reported. It is Gemini 3.1 judging other models on the output of the public questions from GDPval, which tells us nothing.
https://x.com/emollick/status/2045305504679842279

Muse Spark is #3 on ClawEval, ahead of GPT-5.4 and Gemini 3.1 Pro. It is honestly a surprisingly agentic model.
https://x.com/alexandr_wang/status/2045348588734066794

🎉 Congrats to the Moonshot team on Kimi K2.6 — day-0 support on vLLM 0.19.1. • 1T total / 32B active MoE — 384 experts, 8 routed + 1 shared • MLA attention, 256K context • Native multimodal: MoonViT vision encoder + video input • Native INT4 quantization • Interleaved
https://x.com/vllm_project/status/2046251287206035759

FINALLLY FINALLY it is here. V4-flash: all the way back to V2 prices, only now with 1M V4-pro: roughly Kimi/GLM/MiMo competitor Chat prefix completion and FIM back – thank you! Missed this forever but what can they do?
https://x.com/teortaxesTex/status/2047508587883250112

Kimi 2.6 Thinking seems very good for an open weights model, but many rough edges compared to closed SoTA. The Lem Test resulted in a 74 page thinking trace… and an okay-ish answer. It did an okay TiKZ unicorn, an adequate twigl shader for a neogothic city in the waves, etc.
https://x.com/emollick/status/2046411222354989189

Kimi K2.6 + DFlash: 508 tok/s on 8x MI300X 5.6x throughput improvement over baseline autoregressive serving 90 tok/s → 508 tok/s on the same hardware, same model, zero quality loss
https://x.com/HotAisle/status/2046620289984057634

Kimi K2.6 demonstrates strong long-horizon coding in complex engineering tasks: Kimi K2.6 successfully downloaded and deployed the Qwen3.5-0.8B model locally on a Mac. By implementing and optimizing model inference in Zig–a highly niche programming language–it demonstrated
https://x.com/Kimi_Moonshot/status/2046531052957569211

Kimi K2.6 has landed, and it is live on Baseten! We have baked in multiple inference optimizations so that you can leverage Kimi K2.6 in production right away. To run Kimi K2.6, Baseten uses: -> The Baseten Inference Stack with advanced optimizations, including KV-aware routing
https://x.com/baseten/status/2046263526281576573

Kimi K2.6 helped us rewrite kernels; it worked like a charm 🙂
https://x.com/Yulun_Du/status/2046252918526071017

Kimi K2.6 is live on OpenRouter! @Kimi_Moonshot’s new model is a long-horizon coding model built for sustained agentic work. It behaves more like a systems engineer than a chatbot, with the stamina to decompose, execute, and optimize complex tasks. Try it in all your favorite
https://x.com/OpenRouter/status/2046259590774571199

Kimi K2.6 is now available in Windsurf! Available for free for the next 2 weeks for Pro, Teams, and Max users.
https://x.com/windsurf/status/2046686574793154996

Kimi K2.6 now in OpenCode — Go included
https://x.com/opencode/status/2046275886396125680

Kimi K2.6 was released 1h ago, and it looks amazing! Here it’s running with MLX (mlx-vlm) on two M3 Ultras (full 1T param VLM) 🔥
https://x.com/pcuenq/status/2046283942689456297

Meet Kimi K2.6: Advancing Open-Source Coding 🔹Open-source SOTA on HLE w/ tools (54.0), SWE-Bench Pro (58.6), SWE-bench Multilingual (76.7), BrowseComp (83.2), Toolathlon (50.0), Charxiv w/ python(86.7), Math Vision w/ python (93.2) What’s new: 🔹Long-horizon coding – 4,000+
https://x.com/Kimi_Moonshot/status/2046249571882500354

Moonshot AI launches Kimi K2.6 on Kimi Chat and APIs
https://www.testingcatalog.com/moonshot-ai-launches-kimi-k2-6-on-kimi-chat-and-apis/

Qwen3.6-27B can now run locally! 💜 Run on 18GB RAM via Unsloth Dynamic GGUFs. Qwen3.6-27B surpasses Qwen3.5-397B-A17B on all major coding benchmarks. GGUFs:
https://t.co/ykKgwh2zI9 Guide:
https://x.com/UnslothAI/status/2046959757299487029

Ran Qwen3-8B (8.2B dense, open) on LongCoT-Mini. Vanilla: 0/507. dspy.RLM: 33/507 (6.5%). Same model. Same weights. No fine-tuning. The scaffold is doing 100% of the lifting. Context: leaderboard’s smallest open MoE is GLM-4.7 at 358B total / 32B active params. Qwen3-8B is ~4x
https://x.com/raw_works/status/2045208764509470742

these questions are silly Kimi > all other open-source models tho
https://x.com/scaling01/status/2046591683198906542

We’re open-sourcing FlashKDA — our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. Achieves 1.72×-2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention. Explore on github:
https://x.com/Kimi_Moonshot/status/2046607915424034839

OpenClaw 2026.4.20 🦞 🧠 Kimi K2.6 support + provider-aware /think 💬 BlueBubbles iMessage sends + tapbacks fixed ⏰ Cron state/delivery cleanup 🔐 Gateway pairing + plugin startup hardening Less haunted. More useful.
https://x.com/openclaw/status/2046686809367708123

Kimi K2.6 wrote an inference engine for Qwen3.5 0.5B in Zig and managed to beat LM Studio’s token per second by 20%, running for 12 hours and with 4000+ tool calls
https://x.com/nrehiew_/status/2046254256194474221

I find that open weights models over-perform on benchmarks compared to actual real-world usage, and Kimi feels like no exception. For example, a small amount of use will show that Kimi is not as good as Claude Opus 4.6, which it beats on the benchmarks. Still a good model, tho!
https://x.com/emollick/status/2046430301593751797

Image models tend to get much more stuck on a particular direction than text models, requiring clearing the context window fairly often. PerfectSquashBench is my new measure of how image models anchor. The squash remains merely fine after many attempts.
https://x.com/emollick/status/2047073009312121000

Qwen
https://qwen.ai/blog?id=qwen3.6-max-preview

Qwen3.6 Plus lands at #7 in Code Arena with a score of 1476 – up +16 points since the Preview. The new score also moves @AlibabaGroup to #3 lab in Code Arena. In the Text Arena, Qwen3.6 Plus lands at #36, a +13 point improvement since Preview. Congrats to the Qwen team on the
https://x.com/arena/status/2046268995163258958

🚀 Introducing Qwen3.6-Max-Preview, an early preview of our next flagship model Highlights: ⚡️ Improved agentic coding capability over Qwen3.6-Plus 📖 Stronger world knowledge and instruction following 🌍 Improved real-world agent and knowledge reliability performance Smarter,
https://x.com/Alibaba_Qwen/status/2046227759475921291

Guys, I am absolutely astounded. The Qwen 3.6 27b is like a jump to Qwen 4 from Qwen 27B 3.5. I just did a full suite of front end design tests and agentic benchmarks, made entirely by it. VERDICT: They’re so much better than I thought they’d be, like I’m completely astounded. I
https://x.com/KyleHessling1/status/2046986423736451327

llama-server -hf ggml-org/Qwen3.6-27B-GGUF –spec-default
https://x.com/ggerganov/status/2046988075302064209

Qwen 3.6 27B model is available on Ollama! Use it with all the integrations in Ollama or chat with the model. Chat with the model: ollama run qwen3.6:27b OpenClaw: ollama launch openclaw –model qwen3.6:27b Claude Code: ollama launch claude –model qwen3.6:27b More
https://x.com/ollama/status/2047066252523507916

We ran Qwen3.6-35B-A3B GGUF KLD benchmarks of all our dynamic quants and other providers. 1. Nearly all Unsloth quants for mean KLD, 90%, 99.9% KLD are on the Pareto Frontier for KLD vs Disk Space. 2. MXFP4_MOE is an outlier for all. 3. We’ll also make some smaller quants soon!
https://x.com/danielhanchen/status/2045169369723064449

Sharing my current setup to run Qwen3.6 locally in a good agentic setup (Pi + llama.cpp). Should give you a good overview of how good local agents are today: # Start llama.cpp server: llama-server \ -hf unsloth/Qwen3.6-35B-A3B-GGUF:Q4_K_XL \ –jinja \
https://x.com/victormustar/status/2045068986446958899

[2604.15804] Qwen3.5-Omni Technical Report
https://arxiv.org/abs/2604.15804

🎉 Day-0 vLLM support for Qwen3.6-27B! Congrats to @Alibaba_Qwen on the new 27B dense model release. Looking forward to more of the Qwen3.6 series. 👀 📖 Recipe:
https://x.com/vllm_project/status/2046943674890871019

LM Performance:With only 27B parameters, Qwen3.6-27B outperforms the Qwen3.5-397B-A17B (397B total / 17B active, ~15x larger!) on every major coding benchmark — including SWE-bench Verified (77.2 vs. 76.2), SWE-bench Pro (53.5 vs. 50.9), Terminal-Bench 2.0 (59.3 vs. 52.5), and
https://x.com/Alibaba_Qwen/status/2046939775924584577

Qwen3.5-Omni Technical Report | alphaXiv
https://www.alphaxiv.org/abs/2604.15804

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
https://simonwillison.net/2026/Apr/22/qwen36-27b/

Qwen3.6-35B-A3B just dropped. Red Hat AI has an NVFP4 quantized checkpoint ready. 35B params, 3B active, quantized with LLM Compressor. Preliminary GSM8K Platinum: 100.69% recovery (slightly above baseline). Early release. Let us know what you think!
https://x.com/RedHat_AI/status/2045153791402520952

The new Qwen3.6-27B just gave me definitely the best pelican riding a bicycle I’ve had from a 16.8GB model file!
https://x.com/simonw/status/2046995047720378458

We then experiment with 4 different training methods for the Minimal Code Editing task using Qwen3 4B. We find that SFT only works when trained on the same set of evaluation corruptions. It collapses otherwise, indicating that it fails to learn the general minimal coding style
https://x.com/nrehiew_/status/2046963050427879488

VLM Performance:Qwen3.6-27B is natively multimodal, supporting both vision-language thinking and non-thinking modes in a single unified checkpoint — the same as Qwen3.6-35B-A3B. It handles images and video alongside text, enabling multimodal reasoning, document understanding,
https://x.com/Alibaba_Qwen/status/2046939788184547610

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading