Image created with gemini-2.5-flash-image with claude-sonnet-4-5-20250929. Image prompt: A modern birthday cake with two candles whose flames merge into one brilliant light, surrounded by diverse hands reaching in from all sides adding colorful sprinkles and tiny code symbols as decorations, photographed from above with dramatic lighting against a deep blue background, celebrating collective creation and shared joy.
A SOTA moment to me: Kimi’s OK Computer generate this website in just one shot > It designed a very beautiful site, all images were AI-generated, and when you click, the sidebar expands. > Inside the sidebar, there’s a handwritten letter, it really feels like a website made by a https://x.com/crystalsssup/status/1971133240619757794
Say hi to OK Computer, Kimi’s agent mode 🤖🎸 Your AI product & engineering team, all in one. ✨ From chat → multi-page websites, mobile first designs, editable slides ✨ From up to 1 million rows of data → interactive dashboards ✨ Agency: self-scopes, surveys & designs ✨ https://x.com/Kimi_Moonshot/status/1971078467560276160
Nvidia just released Lyra on Hugging Face Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation TL;DR: Feed-forward 3D and 4D scene generation from a single image/video trained with synthetic data generated by a camera-controlled video diffusion model https://x.com/_akhaliq/status/1970949464606245139
It’s becoming increasingly clear that gpt5 can solve MINOR open math problems, those that would require a day/few days of a good PhD student. Ofc it’s not a 100% guarantee, eg below gpt5 solves 3/5 optimization conjectures. Imo full impact of this has yet to be internalized… https://x.com/SebastienBubeck/status/1970875019803910478
Two years ago. No model had surpassed GPT-4 & it wasn’t clear that was possible. Now you can get better than GPT-4 level performance on open weights models running on consumer hardware, and the state of the art in LLMs is cheaper & faster and very much more capable than GPT-4.”” / X https://x.com/emollick/status/1970790843868213361
Flooding the AI Frontier – Chinese models are DOMINATING the open-weight LLM space.
– by Ben https://bturtel.substack.com/p/flooding-the-ai-frontier
Qwen3-VL is finally released and open-sourced, available in both Thinking and Instruct versions! This time, we’ve placed special emphasis on strengthening Visual Agent and Visual Coding, which are crucial steps toward building a true Digital Agent 🚀”” / X https://x.com/huybery/status/1970650821747712209
🎙️ Meet Qwen3-TTS-Flash — the new text-to-speech model that’s redefining voice AI! Demo: https://x.com/Alibaba_Qwen/status/1970163551676592430
🔥 Qwen-Image-Edit-2509 IS LIVE — and it’s a GAME CHANGER. 🔥 We didn’t just upgrade it. We rebuilt it for creators, designers, and AI tinkerers who demand pixel-perfect control. ✅ Multi-Image Editing? YES. Drag in “person + product” or “person + scene” — it blends them like https://x.com/Alibaba_Qwen/status/1970189775467647266
🚀 Introducing Qwen3-LiveTranslate-Flash — Real‑Time Multimodal Interpretation — See It, Hear It, Speak It! 🌐 Wide language coverage — Understands 18 languages & 6 dialects, speaks 10 languages. 👁️ Vision‑Enhanced Comprehension — Reads lips, gestures, on‑screen text and https://x.com/Alibaba_Qwen/status/1970565641594867973
🚀 Introducing Qwen3-Omni — the first natively end-to-end omni-modal AI unifying text, image, audio & video in one model — no modality trade-offs! 🏆 SOTA on 22/36 audio & AV benchmarks 🌍 119L text / 19L speech in / 10L speech out ⚡ 211ms latency | 🎧 30-min audio https://x.com/Alibaba_Qwen/status/1970181599133344172
🚀 We’re thrilled to unveil Qwen3-VL — the most powerful vision-language model in the Qwen series yet! 🔥 The flagship model Qwen3-VL-235B-A22B is now open-sourced and available in both Instruct and Thinking versions: ✅ Instruct outperforms Gemini 2.5 Pro on key vision https://x.com/Alibaba_Qwen/status/1970594923503391182
🛡️ Meet Qwen3Guard — the Qwen3-based safety moderation model series built for global, real-time AI safety! 🌍 Supports 119 languages and dialects ✅ 3 sizes available: 0.6B, 4B, 8B ⚡ Low-latency, Real-time streaming detection with Qwen3Guard-Stream 📝 Robust Full-context safety https://x.com/Alibaba_Qwen/status/1970510193537753397
Alibaba Qwen officially achieves frontier lab status LFG https://x.com/zephyr_z9/status/1970587657421156622
Alibaba released Qwen3-Next-80B-A3B in Base, Instruct, and Thinking variants under an open-weights Apache 2.0 license, targeting faster long-context inference. The 80-billion-parameter mixture-of-experts design swaps most vanilla attention layers for Gated DeltaNet ones and the https://x.com/DeepLearningAI/status/1970254860416131146
Announcing the open-source release of Qwen3-VL! A powerful vision-language model that can operate GUIs, code https://t.co/ww8tsXcd1u charts from mockups, and recognize “”everything”” from daily life to specialized fields. Highlights: 🔹 Precise event location in videos up to 2 https://x.com/Ali_TongyiLab/status/1970665194390220864
NEW: Qwen 235B A22B Vision Language Model is OUTT! Apache 2.0 licensed and upto 1 Million context length 🤯 https://x.com/reach_vb/status/1970589927134937309
Qwen https://qwen.ai/blog?id=1675c295dc29dd31073e5b3f72876e9d684e41c6&from=research.research-list
Qwen https://qwen.ai/blog?id=f0bbad0677edf58ba93d80a1e12ce458f7a80548&from=research.research-list
Qwen https://qwen.ai/blog?id=f50261eff44dfc0dcbade2baf1b527692bdca4cd&from=research.research-list
Qwen https://qwen.ai/blog?id=fdfbaf2907a36b7659a470c77fb135e381302028&from=research.research-list
Qwen3 VL might be the best multimodal (vision) model on the planet”” / X https://x.com/scaling01/status/1970591728433283354
Qwen3-Omni is new sota any-to-any model🔥 everything you have to know ⤵️ > a 30B MoE model with 3B active params, comes in three variants: instruct, thinking and captioner 🤩 thinking is for reasoning and captioner is for robust speech generation 🗣️ > it understands everything https://x.com/mervenoyann/status/1970444546216444022
Qwen3-Omni Technical Report A unified multimodal model that matches same-size Qwen text-only and vision-only baselines while pushing audio and audio-visual SOTA. Key technical details below: https://x.com/omarsar0/status/1970502225379381662
We’re excited to announce the upgrade of Qwen3-Coder, and the upgraded API `qwen3-coder-plus` is now available on Alibaba Cloud Model Studio with major improvements: 💻 Enhanced terminal task capabilities and better performance on Terminal Bench (w/ Qwen Code / Claude Code) 🏆 https://x.com/Alibaba_Qwen/status/1970582211993927774
🚨 New Models Update! 🔥 Qwen3 coming in hot into the Arena with three different models: 🔹Qwen3-VL-235b-a22b-thinking for Text & Vision 🔹Qwen3-VL-235b-a22b-instruct for Text & Vision 🔹Qwen3-Max-2025-9-23 for Text Check out the thread to learn more about them and get https://x.com/arena/status/1970920636957831611
Qwen just released Qwen3Guard-Gen-8B on Hugging Face This new safety moderation model offers three-tiered severity classification and multilingual support for AI content. https://x.com/HuggingPapers/status/1970504452466413639
Wow. Qwen Image Edit now has native support for ControlNet (depth maps, edge maps, keypoint maps etc)”” / X https://x.com/bilawalsidhu/status/1970193454505541755
Metastone unveils MCP-AgentBench A new benchmark evaluating real-world language agent performance with MCP-mediated tools. It features 33 live servers & 188 tools to rigorously test agent capabilities beyond traditional metrics. https://x.com/HuggingPapers/status/1969864853238985001
TorchAO Quantized Models and Quantization Recipes Now Available on HuggingFace Hub – PyTorch https://pytorch.org/blog/torchao-quantized-models-and-quantization-recipes-now-available-on-huggingface-hub/
🚀 ARE: scaling up agent environments and evaluations Everyone talks about RL envs so we built one we actually use. In the second half of AI, evals & envs are the bottleneck. Today we OSS it all: Meta Agent Research Environment + GAIA-2 (code, demo, evals). 🔗Links👇 https://x.com/ThomasScialom/status/1970122143993037170
Kimi Infra team dropped K2 Vendor Verifier > You can visually see the difference in tool call accuracy across providers on OpenRouter. https://x.com/crystalsssup/status/1971158566343184511
LIMI: Less Is More for Agency • Argues agentic AI doesn’t need more data, just better data • 78 curated demos → 73.5% on AgencyBench (beats models trained on 10k samples) • Outperforms SOTA (Kimi-K2: 24.1%, DeepSeek: 11.9%, Qwen3: 27.5%, GLM-4.5: 45.1%) • Establishes Agency https://x.com/arankomatsuzaki/status/1970328242688246160
🌟Our latest LangChain Academy course – Deep Agents with LangGraph – is now live!🌟 Many agents today follow the same simple pattern: run in a loop, call tools. That architecture works well enough, but it breaks down as tasks get more complex. Today, companies of all sizes – https://x.com/LangChainAI/status/1968708505201951029
After gpt-oss, the latest gpt-5-codex is the second model to be Responses API only! Makes sense, since Responses is objectively a better choice than Chat Completions for Agentic use-cases”” / X https://x.com/reach_vb/status/1970585119900528964
AI agency breakthrough: Less data, more power New research with LIMI shows AI agents can achieve 73.5% on benchmarks, outperforming SOTA models by 50%+ using only 78 carefully chosen samples. The “”Agency Efficiency Principle”” is here. https://x.com/HuggingPapers/status/1970400645871185942
China’s Alibaba just dropped a Python framework for building multi-agent apps. AgentScope lets you build AI agents visually with MCP tools, memory, rag, and reasoning capabilities. Works with any LLM and supports real-time steering. 100% Opensource. https://x.com/Saboo_Shubham_/status/1967274908742025356
China’s Alibaba just dropped an opensource 30B agentic LLM that outperforms Claude 4 Sonnet, DeepSeek v3.1, Kimi k2 on a range of agentic search benchmarks. Only 3B parameters are activated per token. 100% open-source. https://x.com/unwind_ai_/status/1969053988143477186
Results so far No single model dominates: GPT-5 “high” reasoning leads on tough tasks but collapses on time-critical ones. Claude-4 Sonnet balances speed vs accuracy but at higher cost. Open-source models (like Kimi-K2) show promise in adaptability. Scaling curves plateau, https://x.com/omarsar0/status/1970147904087322661
🚀 DeepSeek-V3.1 → DeepSeek-V3.1-Terminus The latest update builds on V3.1’s strengths while addressing key user feedback. ✨ What’s improved? 🌐 Language consistency: fewer CN/EN mix-ups & no more random chars. 🤖 Agent upgrades: stronger Code Agent & Search Agent performance.”” / X https://x.com/deepseek_ai/status/1970117808035074215
Congrats to @deepseek_ai ! DeepSeek-R1 was published in Nature yesterday as the cover article, and vLLM is proud to have supported its RL training and inference🥰 https://x.com/vllm_project/status/1968506474709270844
Finetune DeepSeek 🐳 with two Mac Studios + MLX 🚀 We use pipeline parallelism to split the full 671GB model across two devices connected by a single TB5 cable. LoRA reduces the number of parameters to train from 671 billion down to 37 million, reducing the memory overhead from https://x.com/MattBeton/status/1968739407260742069
Nature Portfolio also addressed this on Zhihu: Publishing this paper is itself a significant milestone👏 🤔 DeepSeek-R1 learns step-by-step reasoning with minimal human help: • Reinforcement learning: correct answers get rewards, mistakes penalized • Learns to self-verify & https://x.com/ZhihuFrontier/status/1968603082167828494
Pro Tip💡Fast and simple way to deploy DeepSeek-V3.1-Terminus with vLLM ⚡️ Run it with: vllm serve deepseek-ai/DeepSeek-V3.1-Terminus -tp 8 -dcp 8 (as simple as appending -dcp 8 after -tp 8) Thanks to the @Kimi_Moonshot team, vLLM 0.10.2 adds Decode Context Parallel (DCP) https://x.com/vllm_project/status/1970814441718755685
PSA: you can run the new DeepSeek v3.1 Terminus on a single M3 Ultra with mlx-lm at very usable speed. 4-bit quant one-shotted space-invaders in HTML/CSS: https://x.com/awnihannun/status/1970151204102750573
🚨 Major milestone for open-source AI: DeepSeek-R1, with Wenfeng Liang as corresponding author, has landed on the cover of Nature! 🔥 It’s the first fully peer-reviewed LLM published in a top academic journal, sparking huge debate in China’s tech community Zhihu. Zhihu mind https://x.com/ZhihuFrontier/status/1968573286696239247
Scaleway on Hugging Face Inference Providers 🔥 https://huggingface.co/blog/inference-providers-scaleway
Xet by Hugging Face is the most important AI technology that nobody is talking about! Under the hood, it now powers 5M Xet-enabled AI models & datasets on HF which see hundreds of terabytes of uploads and downloads every single day. What makes it super powerful is that it https://x.com/ClementDelangue/status/1970512794303807724
GenExam: The first multidisciplinary text-to-image exam is now on Hugging Face This new benchmark challenges T2I models with 1,000 rigorous, exam-style prompts across 10 subjects. It comes with ground-truth images and detailed scoring for semantic correctness and visual https://x.com/HuggingPapers/status/1968527551703433595
Ollama now has a web search API and MCP server! ⚡️ Augment local and cloud models with the latest content to improve accuracy 🔧 Build your own search agent 🔍 Directly plugs into existing MCP clients like @OpenAI Codex, @cline, Goose (@jack) and more! Let’s go!!!! 🧵👇 https://x.com/ollama/status/1971085470785319349
Meta’s AI system Llama approved for use by US government agencies | Reuters https://www.reuters.com/world/us/metas-ai-system-llama-approved-use-by-us-government-agencies-2025-09-22/
CWM: An Open-Weights LLM for Research on Code Generation with World Models | Research – AI at Meta https://ai.meta.com/research/publications/cwm-an-open-weights-llm-for-research-on-code-generation-with-world-models/
New from Meta FAIR: Code World Model (CWM), a 32B-parameter research model designed to explore how world models can transform code generation and reasoning about code. We believe in advancing research in world modeling and are sharing CWM under a research license to help empower https://x.com/AIatMeta/status/1970963571753222319
new research from Meta FAIR: Code World Model (CWM), a 32B research model we encourage the research community to research this open-weight model! pass@1 evals, for the curious: 65.8 % on SWE-bench Verified 68.6 % on LiveCodeBench 96.6 % on Math-500 76.0 % on AIME 2024 🧵 https://x.com/alexandr_wang/status/1970973317227225433
We’re excited to share our preparedness report on Code World Model (CWM), FAIR’s latest open-weight model for code generation and reasoning. This report was developed by the SEAL team and the AI Security team, marking our first external publication since part of SEAL joined Meta”” / X https://x.com/summeryue0/status/1970971944557346851
Introducing Magistral Small 1.2 & Magistral Medium 1.2, minor updates to our Magistral 1.1 models! – Multimodality: Now equipped with a vision encoder, these models handle both text and images seamlessly. – Performance Boost: 15% improvements on math and coding benchmarks such https://x.com/MistralAI/status/1968670593412190381
Magistral Medium 1.2 is out Multimodality: Now equipped with a vision encoder, these models handle both text and images seamlessly Performance Boost: 15% improvements on math and coding benchmarks such as AIME 24/25 and LiveCodeBench v5/v6 default in anycoder under Magistral https://x.com/_akhaliq/status/1968708201236381858
@NVIDIA contributes extensively to open-source models on Hugging Face – with over 57 collections published – most within the past year: https://x.com/PavloMolchanov/status/1970553850173255895
DeepSeek’s updated V3.1 Terminus ties with gpt-oss-120b (high) as the most intelligent open weights model and offers increased instruction following and long context reasoning capabilities 🧠 Our benchmarking results indicate DeepSeek V3.1 Terminus shows a greater intelligence https://x.com/ArtificialAnlys/status/1971114096008495501
Shipped. For instance, you can see that gpt-oss-120b is 196 GB right from the “Files” tab https://x.com/mishig25/status/1968598133543256151
Excited to announce @firecrawl_dev Open Source Bounties 💰 We’re offering over $10k+ in rewards through this program over the next few weeks. Check out the first 5 bounties at the link below 🚀 https://x.com/nickscamara_/status/1968340679953989699
🚀 LongCat-Flash-Thinking: Smarter reasoning, leaner costs! 🏆 Performance: SOTA open-source models on Logic/Math/Coding/Agent tasks 📊 Efficiency: 64.5% fewer tokens to hit top-tier accuracy on AIME25 with native tool use, agent-friendly ⚙️ Infrastructure: Async RL achieves a https://x.com/Meituan_LongCat/status/1969823529760874935
An open-source extension for LLM serving engines – LMCache It’s like a caching layer for large-scale, production LLM inference. LMCache implements smart KV cache management, reusing key–value states of previously seen text across GPU, CPU and local disk. It can reuse any https://x.com/TheTuringPost/status/1971318599253098559
What GPT-oss Leaks About
OpenAI’s Training Data https://fi-le.net/oss/
I have heard of several folks using torchtitan internally for RL training. However, torchtitan doesn’t directly support GRPO, which means folks are adding an implementation themselves. A few questions: 1. Are there any good open-source torchtitan forks with GRPO support? 2. What”” / X https://x.com/iScienceLuvr/status/1968509941578338560
We’re excited to introduce ShinkaEvolve: An open-source framework that evolves programs for scientific discovery with unprecedented sample-efficiency. Blog: https://x.com/SakanaAILabs/status/1971081557210489039
Qwen team released two demos along with three models, find all of them in this collection https://x.com/mervenoyann/status/1970445595887161817
We have released Qwen3-Max, the most powerful Qwen model to date! By continuously scaling up model size, data, and RL tasks, great things have happened. This time, coding and agent capabilities have also been significantly enhanced—enjoy!”” / X https://x.com/huybery/status/1970649341582024953
We’ve updated the Qwen3-Coder-Plus API, fixing known issues and making further improvements, especially with adaptations for different scaffolds. Looking ahead, we’ll continue exploring how agent systems can leverage RL to push vibe coding even further!”” / X https://x.com/huybery/status/1970652792848293926
.@Alibaba_Qwen shipping velocity is unmatched Avg 3.5 releases per month, or almost 1 release per week And the majority are open-weights models. Image credt @Smol_AI https://x.com/awnihannun/status/1970839682503348623
[23 Sept 2025] Alibaba Yunqi: 7 models released in 4 days (Qwen3-Max, Qwen3-Omni, Qwen3-VL) and $52B roadmap congrats @Alibaba_Qwen ! https://x.com/Smol_AI/status/1970842828512088486
📢 New Model(s) Drop: Qwen3 VL 235B A22B Instruct & Thinking are now on Yupp! These latest models are @Alibaba_Qwen’s most powerful vision-language models yet. https://x.com/yupp_ai/status/1970640795259851079
🚀 Qwen3-Max is here—no preview, just power! Qwen Chat: https://x.com/Alibaba_Qwen/status/1970599097297183035
🚀 Your Personal AI Travel Designer Is Here! 🍁 Stop wasting hours planning trips. Qwen Chat Travel Planner crafts complete, day-by-day itineraries tailored JUST for you — powered by Amap, Fliggy APIs and Search. ✅ Recommends perfect hotels & transport routes ✅ Builds https://x.com/Alibaba_Qwen/status/1970554287202935159
🚨 Top 10 Open Model Leaderboard Update New open models have entered the Text Arena, and the top 10 rankings by provider have shifted for September! 🔹Qwen-3-235b-a22b-instruct from @Alibaba_Qwen holds the crown at #1 🏆 🔹Longcat-flash-chat from @Meituan_LongCat makes a strong https://x.com/arena/status/1968705194868535749
ChatGPT – Qwen Aug–sep 2025 Timeline (interactive) https://chatgpt.com/canvas/shared/68d3972d363881918f24524394a87d87
Four new releases from Qwen https://simonwillison.net/2025/Sep/22/qwen/#atom-everything
Just enabled full cudagraphs by default on @vllm_project! This change should offer a huge improvement for low latency workloads on small models and efficient MoEs For Qwen3-30B-A3B-FP8 on H100 at bs=10 1024/128, I was able to see a speedup of 47% 🔥 https://x.com/mgoin_/status/1970601094142439761
qwen3-coder-plus is now available on Anycoder Enhanced terminal task capabilities and better performance on Terminal Bench (w/ Qwen Code / Claude Code) SWE-Bench performance up to 69.6 Safer code generation available as Qwen3-Coder-Plus-2025-09-23 https://x.com/_akhaliq/status/1970595669896503462
Qwen3-Max The Tau bench score is insane”” / X https://x.com/scaling01/status/1970599394337587671
Qwen3-Omni-30B-A3B: – instruct – thinking and -captioner https://x.com/scaling01/status/1970182151019659493
So far, Qwen3-Max seems impressive for a non-reasoning model, doing a good job at a lot of my weird tests that even some reasoners struggle with. https://x.com/emollick/status/1970847381966180685
the new Qwen3-Max is now available in anycoder as default as Qwen3-Max-2025-09-23 https://x.com/_akhaliq/status/1970618469344235677
Try the new Qwen models in the Arena!”” / X https://x.com/Alibaba_Qwen/status/1971097727477088717
Qwen (Qwen) https://huggingface.co/Qwen
Inference Providers @huggingface powered by @novita_labs supports Qwen3-VL, the bleeding-edge vision LM 🔥 the model is quite large (22B active 235B total params) so this makes it super easy to try 💚 https://x.com/mervenoyann/status/1971168938848551021
🚨 New Models Alert: WebDev 💻 GPT-5-Codex and Qwen3-Coder-Plus are both now available on WebDev Arena! In the WebDev Arena, you can test out all the best frontier AI coding models on web development tasks. Vote for your preferred response and see how they stack up on the https://x.com/arena/status/1970962780225507775
Proud to release ShinkaEvolve, our open-source framework that evolves programs for scientific discovery with very good sample-efficiency! 🐙 Paper: https://x.com/hardmaru/status/1971081987818745930
We are building “Open Source Nano Banana for Video” – here is open source demo v0.1 We are open sourcing Lucy Edit, the first foundation model for text-guided video editing! Get the model on @huggingface 🤗, API on @FAL, and nodes on @ComfyUI 🧵 https://x.com/DecartAI/status/1968769793567207528




