Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Photorealistic 35mm cinema shot of child aged 6-8 in cozy bedroom with panoramic arc of TV screens, shallow depth of field, one screen showing colorful syntax-highlighted code with annotations, scattered programming books with bookmarks on plush rug, small cardboard box labeled CONTRIBUTIONS filled with USB drives, warm peach and cream lighting contrasted with cool blue-white screen glow, Raspberry Pi visible on shelf, gentle side angle, soft focus, large bold text OPEN SOURCE at top of frame.

IBM dropped CUGA, open-source enterprise agent to automate boring tasks 🔥 > given workspace files, it writes and executes code to accomplish any task 🤯 > comes with a ton of tools built for enterprise tasks, supports MCPs > plug in your favorite LLM 👏 here’s a small demo https://x.com/mervenoyann/status/2000599316121924052

Tinker is now generally available. We also added support for advanced vision input models, Kimi K2 Thinking, and a simpler way to sample from models. https://x.com/thinkymachines/status/1999543421631946888

Tinker: General Availability and Vision Input – Thinking Machines Lab
https://thinkingmachines.ai/blog/tinker-general-availability/

Tinker: General Availability and Vision Input – Thinking Machines Lab https://thinkingmachines.ai/blog/tinker-general-availability/

Today we are releasing Tinker to everyone, and now with vision input! You can now finetune a frontier Qwen3-VL-235B on your own image+text data, bringing your own algorithm (sft, RL, something else?). We’ll take care of the GPU infra. Full update: https://x.com/rown/status/1999544121984245872

NVIDIA Debuts Nemotron 3 Family of Open Models | NVIDIA Newsroom https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models

🎨 Qwen-Image-Layered is LIVE — native image decomposition, fully open-sourced! ✨ Why it stands out ✅ Photoshop-grade layering Physically isolated RGBA layers with true native editability ✅ Prompt-controlled structure Explicitly specify 3-10 layers — from coarse layouts to https://x.com/Alibaba_Qwen/status/2002034611229229388

🚨 Qwen Image Layered is live on fal! ✨ Photoshop-grade layering – Native Decomposition 👑 Physically isolated RGBA layers with true native editability 🎨 Explicitly specify layers, from coarse layouts to fine-grained details https://x.com/fal/status/2002055913390195137

DeepCode: Open Agentic Coding AI coding agents still can’t reliably turn research papers into working code. The best LLM agents achieve only 42% replication scores on scientific papers, while human PhD experts hit 72%. But the problem isn’t model capability. This new paper https://x.com/omarsar0/status/2000385348413850055

Devstral 2 is now available in MLX on Apple Silicon. Run a local SOTA coding model on your MacBook. Please update to LM Studio 0.3.35 first! Have a nice weekend! 👾 https://x.com/lmstudio/status/1999648656958296119

@OpenAI Super cool to see the eval on the Hugging Face hub too – OPEN SOURCE EVALS FTW! 🔥 https://x.com/reach_vb/status/2000982838171328882

Important new eval!”” / X https://x.com/sama/status/2000980694588383434

TIL Xiaomi MiMo-V2-Flash lead was previously one of the key researcher on deepseek V2! https://x.com/eliebakouch/status/2001006476245262395

💡 LMArena Deep Dive: DeepSeek v3.2 (Text Arena) Leaderboard rank doesn’t always tell the full story. As previously reported, DeepSeek released v3.2 two weeks ago. Its results varied across categories and, overall, ranked lower than earlier v3.1 and v3.2 Experimental versions. https://x.com/arena/status/2000637978662821942

NEW: Google releases FunctionGemma, a lightweight (270M), open foundation model built for creating specialized function calling models! 🤯 To test it out, I built a small game: use natural language to solve fun physics simulation puzzles, running 100% locally in your browser! 🕹️ https://x.com/xenovacom/status/2001703932968452365

Introducing Gemma Scope 2 🤗Largest open release of interpretability tools (over 1 trillion parameters trained!) 🔬Works as a microscope to analyze all Gemma 3 models’ internal activations 🗣️Advanced tools for analyzing chat behaviors https://x.com/osanseviero/status/2001989567998836818

FunctionGemma has day-0 support on MLX 🔥🚀 A tiny but mighty single-turn function calling model. Great for on-device tool use, MCP, RAG, routing and more. Get started today: > pip install -U mlx-lm Or run it on your iPhone using MLX-Swift. Notebook example: https://x.com/Prince_Canuma/status/2001713991115026738

🎉 llama.cpp now has Ollama-style model management. • Auto-discover GGUFs from cache • Load on first request • Each model runs in its own process • Route by `model` (OpenAI-compatible API) • LRU unload at `–models-max` https://x.com/victormustar/status/1999484435910263256

.@MistralAI’s Devstral 2 family of models are now available in Ollama. 24B: ollama run devstral-small-2 123B: ollama run devstral-2 Ollama’s cloud: ollama run devstral-2:123b-cloud https://x.com/ollama/status/1999590723373662612

(13) Mistral AI Podcast: Arthur Mensch, Co-founder and CEO (Episode 1) – YouTube https://www.youtube.com/watch?v=xgaLsQTFUEw

Mistral OCR 3 sets new benchmarks in both accuracy and efficiency, outperforming enterprise document processing solutions as well as AI-native OCR. https://x.com/MistralAI/status/2001669583296712970

Very happy to announce the release of our latest Mistral OCR, which significantly outperforms existing solutions! A lot of effort was done to improve handwritten content, low quality scans, and complex tables & forms commonly found in enterprise documents. https://x.com/GuillaumeLample/status/2001719413649617404

Introducing Mistral OCR 3 | Mistral AI https://mistral.ai/news/mistral-ocr-3

Introducing Mistral OCR 3, a new frontier in document intelligence! 🧵👇 https://x.com/MistralAI/status/2001669581275033741

🚀🚀🚀 We’re excited to support @NVIDIA and their new open family of models: NVIDIA Nemotron 3! Open in weights, data, tools, and training, Nemotron 3 is built for multi-agent apps and features: ⚡️An efficient hybrid Mamba‑Transformer MoE architecture 🧾1M token context for”” / X https://x.com/vllm_project/status/2000623058076492276

Agent demos often fail for reasons that are hard to see: unclear tool traces, silent failures, and changes that improve one behavior but break another. Our new course with @Nvidia shows how to use their NeMo Agent Toolkit to surface these issues with OpenTelemetry tracing, run https://x.com/DeepLearningAI/status/2001329113622073611

Baseten supports @nvidia Nemotron 3 Nano on day zero Up to 4× faster token generation, high accuracy, and predictable inference built for agentic AI. Available to deploy today on Baseten for high-performance inference. Read more here: https://x.com/basetenco/status/2000582868532121688

Introducing NVIDIA Nemotron 3 Nano, a fully open 30B with 3B active parameter hybrid MoE model engineered for maximum efficiency and benchmark-leading accuracy. AI natives can now use Nemotron 3 Nano on Together AI — with fast, reliable inference for specialized agentic systems https://x.com/togethercompute/status/2000572943718314392

NEWS: NVIDIA announces the NVIDIA Nemotron 3 family of open models, data, and libraries, offering a transparent and efficient foundation for building specialized agentic AI across industries. Nemotron 3 features a hybrid mixture-of-experts (MoE) architecture and new open https://x.com/nvidianewsroom/status/2000588337896198481

.@nvidia Nemotron 3 Nano is now available on Ollama! Local ollama run nemotron-3-nano Cloud ollama run nemotron-3-nano:30b-cloud https://x.com/ollama/status/2000820163231232167

🚀 Day-0 support for @NVIDIA Nemotron 3 Nano in SGLang SGLang now supports Nemotron 3 Nano on Day 0 🎉 A highly efficient, fully open Hybrid MoE model with 1M context, thinking budget, and industry-leading accuracy per compute. ✅ Open weights, data, and recipes ⚡ Fast, https://x.com/lmsysorg/status/2000567938949243111

As AI Grows More Complex, Model Builders Rely on NVIDIA | NVIDIA Blog https://blogs.nvidia.com/blog/leading-models-nvidia/

BREAKING CUDA MOAT EXPANDS: Today, NVIDIA has acquired SchedMD, makers of SLURM, a widely used “”open source”” workload scheduler. Many AI companies such as Mistral, Thinking Machines, parts of Meta’s FAIR division, university academic labs use SLURM. NVIDIA’s acquisition expands https://x.com/SemiAnalysis_/status/2000620209262985641

BREAKING: NVIDIA just dropped an open 30B model that beats GPT-OSS and Qwen3-30B — and runs 2.2-3.3× faster Nemotron 3 Nano: • Up to 1M-token context • MoE: 31.6B total params, 3.6B active • Best-in-class performance for SWE-Bench • Open weights + training recipe + https://x.com/AskPerplexity/status/2000589984818954719

First time I see a major org release @huggingface collections inside collections 🤯 Kudos @nvidia for this brilliant release https://x.com/NielsRogge/status/2000639749514760465

In collaboration with NVIDIA, the new Nemotron 3 Nano model is fully supported in llama.cpp Nemotron 3 Nano features an efficient hybrid, Mamba, MoE architecture. It’s a promising model, suitable for local AI applications on mid-range hardware. The large context window makes it”” / X https://x.com/ggerganov/status/2000574990425415765

Inside NVIDIA Nemotron 3: Techniques, Tools, and Data That Make It Efficient and Accurate | NVIDIA Technical Blog https://developer.nvidia.com/blog/inside-nvidia-nemotron-3-techniques-tools-and-data-that-make-it-efficient-and-accurate/

New mlx-lm release: pip install -U mlx-lm Includes support for a few new models: – Nemotron 3 Nano (Nvidia) – Devstral (Mistral) – rnj-1 (Essential AI) https://x.com/awnihannun/status/2000974327660077298

Nvidia continues to put out some of the strongest and fastest open models. Pretraining and post training data are released as well, something very few orgs have done”” / X https://x.com/tri_dao/status/2000707760288092655

NVIDIA Debuts Nemotron 3 Family of Open Models | NVIDIA Newsroom https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models/?ncid=so-twit-561360

NVIDIA has just released Nemotron 3 Nano, a ~30B MoE model that scores 52 on the Artificial Analysis Intelligence Index with just ~3B active parameters Hybrid Mamba-Transformer architecture: Nemotron 3 Nano combines the hybrid Mamba-Transformer approach @NVIDIAAI has used on https://x.com/ArtificialAnlys/status/2000602570092675402

NVIDIA just released Nemotron-Agentic-v1 on Hugging Face This dataset empowers LLMs as interactive, tool-using agents for multi-turn conversations and reliable task completion. Ready for commercial use. https://x.com/HuggingPapers/status/2000628009049760072

NVIDIA just released Nemotron-Cascade-8B on Hugging Face A powerful 8B general-purpose reasoning model that achieves best-in-class performance across diverse benchmarks, from math to coding, by using novel Cascade RL. https://x.com/HuggingPapers/status/2001065870676603333

NVIDIA releases Nemotron 3 Nano, a new 30B hybrid reasoning model! 🔥 Nemotron 3 has a 1M context window and the best in class performance for SWE-Bench, reasoning and chat. Run the MoE model locally with 24GB RAM. Guide: https://x.com/UnslothAI/status/2000568378407452746

Really impressive release from NVIDIA, who not only went head-to-head with Qwen3, but: – innovated on the architecture (risky for most open labs) – did legit multi-env RL, complete with agentic evals (first time I see this from an open lab) – plan to open source the pretraining”” / X https://x.com/_lewtun/status/2000599470099099990

SemiAnalysis InferenceMAX showing GPT OSS on Blackwell is 33% more tokens per $ in just 1 month thanks to the awesome work of @vllm_project and @nvidia”” / X https://x.com/dylan522p/status/2002135815233970295

This is not just another strong open model. Nemotron actually releases training data (!), RL environments, and training code. This is a big difference: almost all model developers just want people to use their models; NVIDIA is enabling people to make their own models. We are”” / X https://x.com/percyliang/status/2000608134205985169

Today, @NVIDIA is launching the open Nemotron 3 model family, starting with Nano (30B-3A), which pushes the frontier of accuracy and inference efficiency with a novel hybrid SSM Mixture of Experts architecture. Super and Ultra are coming in the next few months. https://x.com/ctnzr/status/2000567572065091791

vLLM delivers even more inference performance with the same GPU platform. In just 1 month, we’ve worked with NVIDIA to increase @nvidia Blackwell maximum throughput per GPU by up to 33% — significantly reducing cost per token — while also enabling even higher peak speed for https://x.com/vllm_project/status/2001449658984632699

When @NVIDIA announced Nemotron 3 – it marked a symbolic turning point in a year that fundamentally reshaped open-source AI leadership. Is NVIDIA the new open-source king? What’s behind this strategy? Let’s see. ▪️ It releases 3 trillion tokens of new pretraining, 18 million https://x.com/TheTuringPost/status/2001087448299065372

Olmo 3 and the Open LLM Renaissance https://cameronrwolfe.substack.com/p/olmo-3

today we’re open-sourcing nmoe: https://x.com/_xjdr/status/2001434891087671779

🚀 Qwen Code v0.5.0 is here! ✨ What’s new: • VSCode Integration: Bundled CLI into VSCode release package with improved cross-platform compatibility • Native TypeScript SDK: Seamlessly integrate with Node/TS • Smart Session Management: Auto-saves and continue conversations •”” / X https://x.com/Alibaba_Qwen/status/2000556828690624685

Open models year in review What a year! We’re back with an updated open model builder tier list, our top models of the year, and our predictions for 2026. First, the winning models: 1. DeepSeek R1 (@deepseek_ai): Transformed the AI world 2. Qwen 3 Family (@AlibabaGroup): The new https://x.com/natolambert/status/2000299636863734026

ITS LIVE photoshop-grade layering physically isolated RGBA layers with native editability 🤯 https://x.com/linoy_tsaban/status/2002038877511377393

You can now fine-tune LLMs and deploy them directly on your phone! 🚀 We collabed with PyTorch so you can export and run your trained model 100% locally on your iOS or Android device. Deploy Qwen3 on Pixel 8 and iPhone 15 Pro at ~40 tokens/sec. Guide: https://x.com/UnslothAI/status/2001305185206091917

Low latency communication is crucial for tensor parallel inference which is now available on the latest mlx-lm (not on pypi yet). In the following video Devstral is generating a quicksort in C++ 1.7x faster on 2 M3 Ultras (right) vs on 1 (left). https://x.com/angeloskath/status/2001739468425040002

I’m open-sourcing jax-js — a machine learning library for the web, in pure JavaScript jax-js is the first ML compiler that runs in the browser, generating fast WebGPU kernels. Built from scratch over the past year as a personal side project Details: https://x.com/ekzhang1/status/2001680771363254646

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading