Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Minimalist luxury home office with single MacBook on white marble desk displaying GitHub repository with many stars, dramatic moonlight through tall windows, cold blue-grey lighting, pristine empty room, architectural photography style, bold white text ‘OPEN SOURCE’ overlaid across image
🚨BREAKING: Text Leaderboard Update: A new open source model has landed on the leaderboard! Mistral-Large-3 lands at #6 among open models and #28 overall on the Text leaderboard. Mistral 3 is the next generation of Mistral AI models and their most capable model family to date. https://x.com/arena/status/1995877395510051253
Introducing Mistral 3 | Mistral AI https://mistral.ai/news/mistral-3
Introducing Mistral Code | Mistral AI https://mistral.ai/news/mistral-code
Introducing the Mistral 3 family of models: Frontier intelligence at all sizes. Apache 2.0. Details in 🧵 https://x.com/MistralAI/status/1995872766177018340
Magistral | Mistral AI https://mistral.ai/news/magistral
Mistral Small 3 | Mistral AI https://mistral.ai/news/mistral-small-3
Mistral Small 3.1 | Mistral AI https://mistral.ai/news/mistral-small-3-1
Voxtral | Mistral AI https://mistral.ai/news/voxtral
Aerial intelligence with prompts. Moondream segments pools, tennis courts, and even solar panels, with pixel-perfect accuracy. https://x.com/moondreamai/status/1997058204589871395
Moondream’s new segmentation just dropped. Prompt: “dirty laundry items on the bed.” Moondream: pixel-perfect + actually understands the scene. SAM 3: grabs the floor. https://x.com/moondreamai/status/1996001944838832501
Open-Vocabulary Image Segmentation | Moondream
https://moondream.ai/skills/segment
We used Claude Code to train open LLMs. Check out the tutorial. basically, we plugged HF skills into claude code and it was able to train LLMs end-to-end. Best thing, this works on all major coding agents: Codex, Cursor, and Gemini CLI. – You tell the agent to fine-tune a model https://x.com/ben_burtenshaw/status/1996602896436375822
🚨New Models in the Arena! 🐳DeepSeek V3.2: a new family of reasoning-first, agent-oriented models from @deepseek_ai are now live in the Arena. Standard, Thinking, and Speciale are all in the Text Arena, waiting for your toughest prompts! Get your votes in: we’ll see how they https://x.com/arena/status/1995564824718442620
NEW RELEASE – huggingface/skills is a universal implementation of agent context for AI tasks like training models, building datasets, and generating datasets. – compatible with all major coding agent tools: Codex, Cursor, Claude Code, Gemini CLI. – has integrated local script https://x.com/ben_burtenshaw/status/1995877869562855687
At this point, papers testing whether AI can or cannot do something should try to test the strongest case, as well as a default. It is fine to say Llama 2 failed, but did a serious attempt to use GPT-5.1 Thinking in an agentic harness work? It would help better map the frontier.”” / X https://x.com/emollick/status/1994913383871586563
WOW! @AnthropicAI released interviews with 1,250 professionals about how they use AI for work. You can find it on @huggingface as an open dataset! https://x.com/calebfahlgren/status/1996646452509266266
Apple just released CLaRa-7B-Instruct https://x.com/_akhaliq/status/1995899476624458002
apple/CLaRa-7B-Instruct · Hugging Face https://huggingface.co/apple/CLaRa-7B-Instruct
apple/starflow · Hugging Face https://huggingface.co/apple/starflow
More inference workloads now mix autoregressive and diffusion models in a single pipeline to process and generate multiple modalities – text, image, audio, and video. Today we’re releasing vLLM-Omni: an open-source framework that extends vLLM’s easy, fast, and cost-efficient”” / X https://x.com/vllm_project/status/1995566791234629989
🚨Top 10 Open Models by Provider for November The open model race continues with new models entering the Text Arena. Confidence intervals are getting tighter and the competition is heating up! Here are the November Top 3: 🥇 #1 Kimi-K2-Thinking-Turbo by @Kimi_Moonshot (Modified https://x.com/arena/status/1995534475070243043
@awnihannun added batched generation to MLX-LM >2 months ago. Everybody, since, has been asking for batching in the MLX-LM server. Well, enjoy the first version in the latest MLX-LM release. The following video is serving 4 consecutive requests for Qwen3 30B on an M2 Ultra. https://x.com/angeloskath/status/1996364526749639032
🚀 @deepseek_ai just dropped two official models — V3.2 & V3.2-Speciale, and Chinese tech circles are buzzing. What do they really achieve? Zhihu contributor toyama nao breaks it down, closely aligning with DeepSeek’s own published scores👇 DeepSeek has already shaken China’s AI https://x.com/ZhihuFrontier/status/1995689116999311455
🚀 Day 0 Deepseek v3.2 launch on @FireworksAI_HQ ! Congrat @deepseek_ai team on releasing another SOTA model! Continuing our promise, you can access DSV3.2 now on our platform. We heavily focus on quality first. A ton of perf optimization will come shortly. Below are the”” / X https://x.com/lqiao/status/1995915147714723974
🚀 Launching DeepSeek-V3.2 & DeepSeek-V3.2-Speciale — Reasoning-first models built for agents! 🔹 DeepSeek-V3.2: Official successor to V3.2-Exp. Now live on App, Web & API. 🔹 DeepSeek-V3.2-Speciale: Pushing the boundaries of reasoning capabilities. API-only for now. 📄 Tech https://x.com/deepseek_ai/status/1995452641430651132
🚀 vLLM now offers an optimized inference recipe for DeepSeek-V3.2. ⚙️ Startup details Run vLLM with DeepSeek-specific components: –tokenizer-mode deepseek_v32 \ –tool-call-parser deepseek_v32 🧰 Usage tips Enable thinking mode in vLLM: – https://x.com/vllm_project/status/1996760535908642986
DeepSeek V3.2 is the #2 most intelligent open weights model and also ranks ahead of Grok 4 and Claude Sonnet 4.5 (Thinking) – it takes DeepSeek Sparse Attention out of ‘experimental’ status and couples it with a material boost to intelligence @deepseek_ai V3.2 scores 66 on the https://x.com/ArtificialAnlys/status/1996110256628539409
deepseek-ai/DeepSeek-V3.2 · Hugging Face https://huggingface.co/deepseek-ai/DeepSeek-V3.2
Game over https://x.com/Yuchenj_UW/status/1995523554679673180
If you need an adrenaline rush to wake up from your post-Thanksgiving stupor… we got you. @deepseek_ai V3.2 dropped this week and is now available on Baseten. It’s so smart your mother will ask why you can’t be more like DeepSeek. V3.2 is currently on par with GPT-5 all whilst https://x.com/basetenco/status/1996623218040254793
Incredible writeup! Some notable 💎s: Deepseek reduced attention complexity from quadratic to ~linear through warm-starting (w/ separate init + opt dynamics) and adapting the change over ~1T tokens. They also use separate attention modes for disaggregated prefill vs decode (is https://x.com/suchenzang/status/1995535496421015741
Introducing DeepSeek-V3.2-Exp | DeepSeek API Docs https://api-docs.deepseek.com/news/news250929
Link to DeepSeek’s technical paper: https://x.com/ArtificialAnlys/status/1996110267353325748
LisanBench results for DeepSeek-V3.2 DeepSeek-V3.2 and V3.2 Speciale are affordable frontier models* *the caveat is that they are pretty slow at ~30-40tks/s and produce by far the longest reasoning chains at 20k and 47k average output tokens (incl. reasoning) – which results in https://x.com/scaling01/status/1995895894219100462
New Model(s) Drop: DeepSeek V3.2 is now live on Yupp! From @deepseek_ai, these are open-source models with enhanced math, coding and logic capabilities – offered in Chat, Thinking and Speciale versions. Let’s see how they perform: https://x.com/yupp_ai/status/1995538168146526274
Speciale is the first DeepSeek model of all time that gets my dumb bilingual joke about God of War. V3.2 flails and invents cringe fake etymology, just like R1. 4o could do this already. The knowledge gap is wide and deep indeed. Still. Frontier at last. https://x.com/teortaxesTex/status/1995527632578834829
While reviewing the results again, I noticed a misjudged part in the score for the deepseek v3.2 Speciale model and corrected it. The revised result is a very impressive 8.81, which is top tier and achieved a perfect 10 across all quantitative metrics. However, as I mentioned https://x.com/Hangsiin/status/1995899545339990042
🚨BREAKING: Text Leaderboard Update 🐳 Deepseek-v3.2 enters the leaderboard at #38, and Deepseek-v3.2-thinking lands at #41. For comparison, previous versions ranked higher: 🔹 v3.2 ranks #38 (-5 pts v3.1 and -14 pts v3.2-exp) 🔹 v3.2-thinking ranks #41 (-7 pts vs v3.1-thinking https://x.com/arena/status/1996707563208167881
Compare how DeepSeek V3.2 performs relative to models you are using or considering at: https://x.com/ArtificialAnlys/status/1996110266065715249
DeepSeek’s new DeepSeekMath-V2 hits gold-medal performance on IMO and Putnam. It’s the first open model that can check its own proofs, fix mistakes, and improve itself. DeepSeekMath-V2 uses two “minds” in one model: ▪️ A verifier – Reads a proof and points out issues. – https://x.com/TheTuringPost/status/1994926897248288813
very smart choices by @stochasticchasm and the arcee team. in terms of arch, this is pretty much the perfect setup if you’re a bit constrained by compute/time and can’t do 100s of ablations hybrid nope, gated attention, norms to stabilize everything, muon, deepseek routing this”” / X https://x.com/eliebakouch/status/1995600008603697346
We managed to get Claude code, Codex and Gemini CLI to train good AI models thanks to @huggingface skills and you can too even (especially?) if you’ve never trained a model before 🤯🤯🤯 After changing the way we build software, AI might start to change the way we build AI https://x.com/ClementDelangue/status/1996718490435174435
LongCat-Image is out on Hugging Face https://x.com/_akhaliq/status/1996946556834959663
Out now: LongCat-Image-Edit – seems very very good at image editing (+ Apache 2.0 license) 🔥 ⬇️ Demo available on Hugging Face https://x.com/victormustar/status/1997012462252732882
American open-source is making a comeback in 2026 Arcee just started cooking Trinity Large which will be released in early 2026 It will have 420@13B params and is trained on 2048 B300 with 20T tokens https://x.com/scaling01/status/1995616210109825447
Introducing Trinity Mini from @arcee_ai, an open-weight 26B sparse MoE model that activates just 3B parameters per token while delivering frontier-class reasoning. AI natives can now use Trinity Mini on Together AI — and benefit from reliable inference for production-scale https://x.com/togethercompute/status/1995594629505573338
Introducing Trinity, the start of a new open-weight MoE family. Rolling out today today: Trinity-Mini (26B-A3B) Trinity-Nano-Preview (6B-A1B) Download on HuggingFace. Free for limited time on OpenRouter. https://x.com/arcee_ai/status/1995600354374025395
Today, we are introducing Trinity, the start of an open-weight MoE family that businesses and developers can own. Trinity-Mini (26B-A3B) Trinity-Nano-Preview (6B-A1B) Available Today on Huggingface. https://x.com/latkins/status/1995592664637665702
@swyx yeah all public info – the 3K cluster is in the main mistral 3 blog post today, and the 18K cluster was announced by nvidia earlier this year https://x.com/AnjneyMidha/status/1996000762904936755
And Mistral Large 3, a frontier class open source MoE. https://x.com/MistralAI/status/1995872771516354828
🎉 Congratulations to the Mistral team on launching the Mistral 3 family! We’re proud to share that @MistralAI, @NVIDIAAIDev, @RedHat_AI, and vLLM worked closely together to deliver full Day-0 support for the entire Mistral 3 lineup. This collaboration enabled: • NVFP4 https://x.com/vllm_project/status/1995890057224618154
Europe still has one frontier model maker that can generally keep pace with Chinese open weights models, though no reasoner for Mistral 3 yet means they are behind the curve of actual performance – DeepSeek r1 got 71.5% on GPQA Diamond (& 1-shot, not 5-shot) back in January. https://x.com/emollick/status/1996068920596594932
I want to especially thank @MistralAI for releasing the base models for Mistral 3. Fewer companies are sharing base models and this opens many use cases from custom instruct to non-instruct cases”” / X https://x.com/QuixiAI/status/1996272948378804326
Meet the Ministral 3 models from @MistralAI! – 3B, 8B, and 14B models – Instruct, reasoning, and base variants – Supports tool use and vision input – Open-weights, Apache 2.0 licensed https://x.com/lmstudio/status/1995908228526604451
Mistral 3 is now available on Ollama v0.13.1 (currently in pre-release on GitHub). 14B: ollama run ministral-3:14b 8B: ollama run ministral-3:8b 3B: ollama run ministral-3:3b Please update to the latest Ollama. https://x.com/ollama/status/1995885696360566885
Mistral releases Ministral 3, their new reasoning and instruct models! 🔥 Ministral 3 comes in 3B, 8B, and 14B with vision support and best-in-class performance. Run the 14B models locally with 24GB RAM. Guide + Notebook: https://x.com/UnslothAI/status/1995874975631503479
NEW: @MistralAI released a fantastic family of multimodal models, Ministral 3. You can fine-tune them for free on Colab using TRL ⚡️, supporting both SFT and GRPO https://x.com/SergioPaniego/status/1996257877871509896
NEW: @MistralAI releases Mistral 3, a family of multimodal models, including three start-of-the-art dense models (3B, 8B, and 14B) and Mistral Large 3 (675B, 41B active). All Apache 2.0! 🤗 Surprisingly, the 3B is small enough to run 100% locally in your browser on WebGPU! 🤯 https://x.com/xenovacom/status/1995879338583945635
Run Mistral Large 3 on Ollama’s cloud: ollama run mistral-large-3:675b-cloud”” / X https://x.com/ollama/status/1996682858933768691
Super nice to see Mistral Large 3 as the #1 OSS model for coding on lmarena 🥳😎🙌 And the spoiler alert! 👀👀”” / X https://x.com/sophiamyang/status/1996587296666128398
Support for running Mistral Large 3 locally will be available in Ollama soon.”” / X https://x.com/ollama/status/1996683156817416667
The Bert-Nebulon Alpha Stealth model is live now as @MistralAI’s new Mistral Large 3! Try the full release now on OpenRouter: https://x.com/OpenRouterAI/status/1995904288560988617
The world’s best small models–Ministral 3 (14B, 8B, 3B), each released with base, instruct and reasoning versions. https://x.com/MistralAI/status/1995872768601325836
Mistral Large 3 debuts as the #1 open source coding model on the @arena leaderboard. We’d love for you to try it! More on coding in a few days… 👀 https://x.com/MistralAI/status/1996580307336638951
Mistral AI raises 1.7B€ to accelerate technological progress with AI | Mistral AI https://mistral.ai/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai
🧊 Off-policy RL for LLMs is hard. Dr. GRPO collapses at 10 steps off-policy. TBA doesn’t. @Kimi_Moonshot K2’s approach is robust too – both independently landed on the same key ingredients 🤝 We ablate RL recipe ingredients + show the 2 small changes giving off-policy https://x.com/bartoldson/status/1996769053420265959
You can now integrate Kimi CLI into JetBrains via the ACP. For details, check out the Kimi CLI GitHub repo: https://x.com/Kimi_Moonshot/status/1996953835080966390
Comparing Openness with Intelligence, we see a negative correlation. This relationship is substantially driven by frontier lab releases and gaps in transparency for some leading open weights models. https://x.com/ArtificialAnlys/status/1995523186394575204
Just shipped Tangle – the first open source experimentation platform with content-based caching and visual editor that’s actually pleasant to use. The CPU time savings alone are ridiculous (seeing 1+ year saved at Shopify). We’re at NeurIPS, booth #1713, if you want to see it. https://x.com/MParakhin/status/1995910229641629849
The moment open-source models were close to 30% of OpenRouter traffic and almost all of them came from China with the notable models being: DeepSeek V3/R1, Qwen3 family, Kimi-K2 and GLM-4.5 + Air Minimax M2 is now also a major player, but open-weights models token-usage https://x.com/scaling01/status/1996975947082289418
Today we’re releasing BrowseSafe and BrowseSafe-Bench: an open-source detection model and benchmark to catch and prevent malicious prompt-injection instructions in real-time. https://x.com/perplexity_ai/status/1995965227494699339
We’re taking CUDA debugging to the next level. 🚀 Building on our previous work with CUDA Core Dumps, we are releasing a new guide on tracing hanging and complicated kernels down to the source code. As kernels get more complex (deep inlining, async memory access), standard https://x.com/vllm_project/status/1996256049368793218
> be arcee > look around > realize open-weight frontier MoE is basically a Qwen/DeepSeek monopoly > decide “nah, we’re building our own” > actual end-to-end pretraining > on US soil > introducing Trinity > Nano (6B MoE) and Mini (26B MoE) > open weights, Apache 2.0 > free on https://x.com/TheAhmadOsman/status/1995613231629381935
Our new Qwen3-TTS (version 2025-11-27) is here! 🚀 We’ve leveled up on what matters most: ✨ More Personalities: Over 49 high-quality voices, from cute and playful to wise and stern. Find your perfect match! 🌍 Global Reach: Now speaks 10 languages (zh, en, de, it, pt, es, ja, https://x.com/Alibaba_Qwen/status/1996947806138126547
The latest mlx-lm is out and it has continuous batching with mlx_lm.server! Added by @angeloskath Check-out the video of 4 simultaneous requests running with Qwen3 30B on the same M2 Ultra:”” / X https://x.com/awnihannun/status/1996365940343402596
TIL you can compile quantized models thanks to quanto although memory blows up a bit on Qwen3-VL https://x.com/mervenoyann/status/1996998362118201850





Leave a Reply