Image created with gemini-2.5-flash-image with claude-sonnet-4-5-20250929. Image prompt: A grand stone courtyard in the style of Lincoln’s Inn with classical arches and columns, centered on a glowing communal forge where multiple diverse hands collaborate on a single radiant artifact, while stone tablets inscribed with code line the walls like ancient legal precedents under an open sky, cinematic photograph with warm golden light and cool stone textures.
🏖️ Summarization Middleware As agent loops get long (either because lots of messages or lots of tool calls) you want to summarize what has occurred so you don’t overflow context (and break your workflow). LangChain’s new middleware automatically summarizes history to keep you https://x.com/sydneyrunkle/status/1967991069368275282
Avoid overflowing context windows with LangChain’s SummarizationMiddleware. This is especially important for long running conversations that have lots of messages and agent loops with lots of tool calls.”” / X https://x.com/LangChainAI/status/1967993889958031560
Why do AI models keep “hallucinating”? @OpenAI’s paper argues: Models aren’t broken. Training and benchmarks reward confident guesses over honesty. Proposed solutions: – Change benchmark scoring: not penalize models for “”I don’t know”” – Realign current leaderboards instead of https://x.com/TheTuringPost/status/1966638472854483129
💬 From system architecture to AI infra, Zhihu contributor 想养一只猫 breaks down why @Alibaba_Qwen Qwen3-Next-80B-A3B matters. 🌟Might be the first real shot at high-complexity hybrid architecture in open-source models. Big players like NVIDIA (Nemotron-H+) & MiniMax+ have https://x.com/ZhihuFrontier/status/1966419946885493098
🐻Qwen3-Next just dropped on Together AI 80B parameters, 3B activated. Two models: ⚡Thinking: Outperforms Gemini-2.5-Flash-Thinking on reasoning benchmarks 🧠Instruct: Matches 235B model performance on key tasks Available now via our API 🚀 https://x.com/togethercompute/status/1966932629078634543
Qwen3 Next 80B A3B Thinking outperforms higher-cost and closed models like Gemini 2.5 Flash Thinking on benchmarks, nearing Qwen’s flagship model quality at a fraction the size. We have it ready to deploy in our model library, running on @nvidia and the Baseten Inference Stack. https://x.com/basetenco/status/1967688601640288288
📢 @Alibaba_Qwen new open-source model Qwen3-Next-80B-A3B is making waves. With a hybrid architecture & strong long-context reasoning, it’s sparking intense debate in the Zhihu community🔥 🔧 Zhihu contributor toyama nao with evalution: TLDR: A new “”gatekeeper”” for open-source https://x.com/ZhihuFrontier/status/1966415278922989813
🚨 Top 10 Open Model Leaderboard Update New open models have entered the Text Arena, and the top 10 rankings by provider have shifted for September! 🔹Qwen-3-235b-a22b-instruct from @Alibaba_Qwen holds the crown at #1 🏆 🔹Longcat-flash-chat from @Meituan_LongCat makes a strong https://x.com/arena/status/1968705194868535749
Alibaba has released Qwen3 Next 80B: an open weights hybrid reasoning model that achieves DeepSeek V3.1-level intelligence with only 3B active parameters Key takeaways: 💡 Novel architecture: First model to introduce @Alibaba_Qwen’s ‘Qwen3-Next’ foundation models, with several https://x.com/ArtificialAnlys/status/1966523300781428788
The new open-source Qwen3-Next Instruct and Thinking models put state-of-the-art long-context reasoning into the hands of everyone. We collaborated with #opensource frameworks from SGLang (@lmsysorg) and @vllm_project to enable communities to deploy Qwen3-Next across the https://x.com/NVIDIAAIDev/status/1967575419638468667
Introducing SpatialVID: A massive new video dataset for 3D spatial intelligence Crucial for training next-gen models, it features over 7,000 hours of diverse, in-the-wild video with dense annotations like camera poses, depth maps, and dynamic masks. https://x.com/HuggingPapers/status/1967260292569845885
ever wondered how the text search on your phone image gallery works? @AIatMeta released MetaCLIP2, and we’ve added it to @huggingface transformers 🔥 it’s a multilingual model that can understand image + text! find the notebook for text-to-image search on the next ⤵️ https://x.com/mervenoyann/status/1966544046744011242
🚀 Kimi K2 Official Turbo API — 50% OFF for 30 days Code faster, ship sooner. Try it now: https://x.com/Kimi_Moonshot/status/1967829577037910427
Our engineer wrote about the thinking and technical story behind Checkpoint Engine. 👉 https://x.com/Kimi_Moonshot/status/1967923416008462785
Excited to release a preview of Moondream 3. A 9B param, 2B active MoE vision language model that makes no compromises; offering state-of-the-art visual reasoning while still retaining an efficient and deployment-friendly form factor. https://x.com/vikhyatk/status/1968800178640429496
when chatgpt said moondream wasn’t a frontier model, i took it personally”” / X https://x.com/vikhyatk/status/1968811248381784167
Model-agnostic plug-n-play LangChain/LangGraph agents powered entirely by MCP tools over HTTP/SSE. https://x.com/_avichawla/status/1967476110285021213
AWS released an open-source framework that lets you orchestrate multiple AI agents and handle complex conversations.
https://x.com/LiorOnAI/status/1964403536067743764
Qwen3 Next 80B used ~100M tokens with reasoning and ~25M without reasoning to run the Artificial Analysis Intelligence Index, slightly less verbose than Qwen3 235B 2507 with reasoning, and similar to it without reasoning https://x.com/ArtificialAnlys/status/1966523306338893979
🚨 Major milestone for open-source AI: DeepSeek-R1, with Wenfeng Liang as corresponding author, has landed on the cover of Nature! 🔥 It’s the first fully peer-reviewed LLM published in a top academic journal, sparking huge debate in China’s tech community Zhihu. Zhihu mind https://x.com/ZhihuFrontier/status/1968573286696239247
Finetune DeepSeek 🐳 with two Mac Studios + MLX 🚀 We use pipeline parallelism to split the full 671GB model across two devices connected by a single TB5 cable. LoRA reduces the number of parameters to train from 671 billion down to 37 million, reducing the memory overhead from https://x.com/MattBeton/status/1968739407260742069
Congrats to @deepseek_ai ! DeepSeek-R1 was published in Nature yesterday as the cover article, and vLLM is proud to have supported its RL training and inference🥰 https://x.com/vllm_project/status/1968506474709270844
DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning | Nature https://www.nature.com/articles/s41586-025-09422-z
Nature Portfolio also addressed this on Zhihu: Publishing this paper is itself a significant milestone👏 🤔 DeepSeek-R1 learns step-by-step reasoning with minimal human help: • Reinforcement learning: correct answers get rewards, mistakes penalized • Learns to self-verify & https://x.com/ZhihuFrontier/status/1968603082167828494
Introducing VaultGemma 🧠Gemma pre-trained with differential privacy (largest open model trained from scratch like this) 🔒Strong, mathematically-backed privacy guarantees 🤏Just 1B parameters 📈Novel research on scaling laws”” / X https://x.com/osanseviero/status/1966534013511672148
VaultGemma: The world’s most capable differentially private LLM https://research.google/blog/vaultgemma-the-worlds-most-capable-differentially-private-llm/
@huggingface Cerebras Inference powers the world’s top coding models, making code generation pretty much instant. And anyone can get a free API key from https://x.com/code/status/1966638514100924846
transformers @huggingface comes with Kosmos2.5 model by @MicrosoftAI 🤩 to celebrate this, we built a notebook for fine-tuning with OCR with detection 🔥 you can also try the model out of the box with the demo 🫡 https://x.com/mervenoyann/status/1966487632659005667
Tencent released SRPO on Hugging Face Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference By fine-tuning the FLUX1dev model with optimized denoising and online reward adjustment, improve its human-evaluated realism and aesthetic quality by over 3x https://x.com/_akhaliq/status/1966911634657390890
Our lightweight open-source eval library “”lighteval”” now ships with 7,000+ (!!) benchmarks baked in. Running it locally is literally a one-liner: >> lighteval vllm “”model_name=gpt2″” “”leaderboard|truthfulqa:mc|0″” (there is also a Python API for in/post-training evals ofc)”” / X https://x.com/Thom_Wolf/status/1967926861889163304
Meta announced LlamaFirewall, a toolkit to protect LLM agents from jailbreaking, goal hijacking, and exploiting vulnerabilities in generated code. The toolkit is now free to use for projects with up to 700 million monthly active users. Read our summary of the paper in The https://x.com/DeepLearningAI/status/1967986588312539272
Introducing Magistral Small 1.2 & Magistral Medium 1.2, minor updates to our Magistral 1.1 models! – Multimodality: Now equipped with a vision encoder, these models handle both text and images seamlessly. – Performance Boost: 15% improvements on math and coding benchmarks such https://x.com/MistralAI/status/1968670593412190381
Magistral Medium 1.2 is out Multimodality: Now equipped with a vision encoder, these models handle both text and images seamlessly Performance Boost: 15% improvements on math and coding benchmarks such as AIME 24/25 and LiveCodeBench v5/v6 default in anycoder under Magistral https://x.com/_akhaliq/status/1968708201236381858
here’s a notebook tutorial to walk you through how to do this step-by-step 💗 GH → https://x.com/mervenoyann/status/1966544570436424074
1/ Introducing Isaac 0.1 — our first perceptive-language model. 2B params, open weights. Matches or beats models significantly larger on core perception. We are pushing the efficient frontier for physical AI. https://x.com/perceptroninc/status/1968365052270150077
Excited to introduce Isaac – our first open-weights model which excels at localization and visual understanding.”” / X https://x.com/kilian_maciej/status/1968396992104874452
Shipped. For instance, you can see that gpt-oss-120b is 196 GB right from the “Files” tab https://x.com/mishig25/status/1968598133543256151
⚡️Ling-flash-2.0⚡️ is now open source. 100B MoE LLM • only 6.1B active params –> 3x faster than 36B dense (200+ tok/s on H20) –> Beats ~40B dense LLM on complex reasoning –> Powerful coding and frontend development Small activation. Big performance. https://x.com/AntLingAGI/status/1968323481730433439
🤖📰 Intelligent News Agent A powerful news curation system that transforms information overload into personalized insights using LangGraph’s reactive agents. Features smart deduplication and multi-source synthesis for automated news processing. Explore the system here 🔍 https://x.com/LangChainAI/status/1966909743383146735
We have updated ChatGPT’s personalization page: personality configuration, custom instructions, and memories are now all in one place. Going live over the next couple of days. https://x.com/sama/status/1967789125702140021
Struggling with the 3-minute limit on Qwen3-ASR-Flash? No more! Introducing the Qwen3-ASR-Toolkit 🚀 A free, open-source CLI to transcribe HOURS-long audio/video files at high speed. Unleash the full power of the Qwen3-ASR-Flash API! 💥 🧠 Smart VAD splitting (no awkward”” / X https://x.com/Alibaba_Qwen/status/1968230660973396024
Here’s the list: Qwen3 Next: Hybrid SSM, MoE Ling Mini: Regular Attention, MoE Granite 4: Hybrid SSM, MoE Kwai Klear: Reg attn, MoE Nemotron-H: Hybrid SSM, MoE Long Cat Flash: Reg attn, MoE Apertus: Regular attn, Dense”” / X https://x.com/awnihannun/status/1966937464834314614
LM Studio now supports Qwen3-Next with MLX on Mac — so cool! 🎉 And that Qwen capybara and LM Studio purple little guy are just too cute 😍”” / X https://x.com/Alibaba_Qwen/status/1968131326034448442
“Welcome Back” project summary on reopen! https://x.com/Alibaba_Qwen/status/1966451500340703418
📢 New Model Drop: Qwen3 Coder Flash & Qwen3 Coder Plus are now on Yupp! These models combine coding proficiency with versatile general-purpose abilities. We tossed some prompts its way: https://x.com/yupp_ai/status/1968387335651000324
Excited to see Qwen3-Next-80B on Poe! 🚀”” / X https://x.com/Alibaba_Qwen/status/1967835503308443687
Thanks for the evaluation! Qwen3-Next 80B achieves strong performance with only 3B active parameters!”” / X https://x.com/Alibaba_Qwen/status/1966831435756794071
The new batch generation in MLX LM is pretty fast. Here’s 4 simultaneous generations with Qwen3 4B on my M4 max: https://x.com/awnihannun/status/1967966714173534494
🆙Qwen Code v0.0.10 & v0.0.11 bring new features and dev-friendly improvements: ✨New UX & Productivity · Subagents for smarter task decomposition · Todo Write tool for task tracking · “Welcome Back” project summary on reopen! · Customizable cache Strategy ⚡Performance & Dev https://x.com/Alibaba_Qwen/status/1966451235328008563
tldr: you can RL qwen3 8b to fool gpt-4o that it’s not doing a hidden side task (when it is) this is somewhat surprising given the disparity in model capabilities between an 8b agent and gpt-4o as a relatively strong monitor https://x.com/neev_parikh/status/1967767438243876924
HunyuanImage 2.1 is the new leading open weights text to image model from @TencentHunyuan , surpassing HiDream-I1-Dev and Qwen-Image in the Artificial Analysis Image Arena! HunyuanImage 2.1 is the latest release from Tencent – a 17B DiT text-to-image model natively supporting https://x.com/ArtificialAnlys/status/1967800071115903358
First test of MLX batch generation PR on Mac Studio M3 Ultra 512GB with Qwen3-1.7B (4K ctx, 64 tokens) 🔥 Batch generation = WOW bf16 vs 4bit (avg of 3 runs) Batch of 1 → 127 vs 237 t/s 5 → 365 vs 515 t/s 10 → 556 vs 625 t/s 15 → 672 vs 617 t/s MLX vllm not a dream anymore! https://x.com/ivanfioravanti/status/1966903782400545196
LM Studio now supports Qwen3-Next with MLX on Mac! 🧵 https://x.com/lmstudio/status/1967985102845366280
Woah, 66 tok/s on a Macbook M4 Max 64GB with qwen3-next-80b-a3b-instruct-mlx@4bit, which uses about 41GB. Amazing job to the folks working on MLX, aware of at least these guys: @ivanfioravanti @ActuallyIsaak @awnihannun https://x.com/rwojo/status/1967767157250592899
Check out the actual speed (not yet the final version) of Qwen3-Next-80B-A3B-Instruct on Apple MLX! 🔥 4-bit: 67 TPS 8-bit: 58 TPS bf16: 48 TPS Movie normal speed, only waiting times removed. @awnihannun and @ActuallyIsaak did it and I bet there is still room for improvement 💪 https://x.com/ivanfioravanti/status/1966866942461177925
Big mlx-lm release: pip install -U mlx-lm – A bunch of new models: Qwen3 Next, Ling Mini, Meta’s MobileLLM, and more – Batch generation – Nice speedups for SSM and hybrid SSM models – Faster prompt processing for GPT-OSS https://x.com/awnihannun/status/1968426979838869789
@Alibaba_Qwen Massive efficiency gains for long contexts. 262K context native, extensible to 1M+ tokens. Perfect for: ⚡ Repository-scale code analysis 🧠 Complex reasoning tasks 📄 Long document processing Both models available now → Instruct: https://x.com/togethercompute/status/1966933240683319556
Build an LLM from scratch https://x.com/rasbt/status/1966876565788135837
1. Model agnostic. 2. Inference agnostic. & now, 3. Platform agnostic. Cline for JetBrains is here. (install it below) https://x.com/cline/status/1968360125686759505
We are building “Open Source Nano Banana for Video” – here is open source demo v0.1 We are open sourcing Lucy Edit, the first foundation model for text-guided video editing! Get the model on @huggingface 🤗, API on @FAL, and nodes on @ComfyUI 🧵 https://x.com/DecartAI/status/1968769793567207528




