Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A flat matte 16:9 op-art poster in Julio Le Parc’s style, with the word ‘HuggingFace’ centered and lettered from concentric ROYGBIV rainbow bands violet-outward-to-red-inward, flanked by two thick spectrum ribbons curving in from left and right like embracing arms that meet in a clean overlapping clasp below the word, set on a single pale off-white background with abundant negative space, crisp printed edges, no shadows or gradients.

NVIDIA Nemotron 3 Ultra Powers Faster, More Efficient Reasoning for Long-Running Agents | NVIDIA Technical Blog
https://developer.nvidia.com/blog/nvidia-nemotron-3-ultra-powers-faster-more-efficient-reasoning-for-long-running-agents/

Shared my first trace from @NanoClaw_AI to @huggingface yesterday. Very cool! By default, all agents should store their traces on HF (in private) so that you can keep a history of them, analyze them,… & share them and post-train better models, harnesses and more. Excited
https://x.com/ClementDelangue/status/2062542713463980303

I just shipped oMLX v0.4.0, the first official release with the new native Swift macOS app.
https://t.co/kJeAGjGiSp oMLX now ships with a redesigned onboarding flow, settings UI, Hugging Face cache discovery, and a much more native-feeling way to manage and run local models on
https://x.com/jundotkim/status/2061863850874634242

My new favorite @huggingface feature: hardware compatibility check for every model 🤗 They really are betting on local AI and I love it. Also check out @dphnAI X1 Trinity Nano. 6B params with 3-8-bit versions!
https://x.com/m_newhaus/status/2061824017510584630

Routing and post-training open-source models won’t only give you more accurate systems but also meaningfully faster and cheaper systems as most companies are currently learning (in addition to giving you more control and privacy). The idea that a “”frontier”” model (by frontier we
https://x.com/ClementDelangue/status/2062248714945630632

So much great work lately from Nvidia, the “”King of American Open-source AI””! – Crossed 1,000 total public repositories on @huggingface (820 models, 249 datasets & 57 spaces) & almost 60,000 followers – Current #1 trending model on HF with LocateAnything and #5 trending with PiD
https://x.com/ClementDelangue/status/2061487081315094906

We want to work with kernel developers to help them publish their cool kernels on the @huggingface Hub via🤗 Kernels. This has several advantages: * A consistent build structure * Extreme ease of use * Standardized distribution * Reproducibility Reach out if interested 🤗
https://x.com/RisingSayak/status/2062471134260687264

Automatic behind the scene routing in user interfaces (instead of model picker) will redistribute value capture and usage towards many more models than just frontier ones (especially towards open-source/smaller/cheaper ones). Because it removes the cognitive load for the final
https://x.com/ClementDelangue/status/2061871024627482964

@openclaw @NousResearch @LangChain As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, and post-training recipes. Available now on @huggingface →
https://x.com/NVIDIAAI/status/2062521383582646537

🚀 Day-0 support for NVIDIA Nemotron 3 Ultra on vLLM! Ready to be served with the latest vLLM stable release, the new open frontier reasoning model is built for long-running autonomous agents: 🧠 550B total / 55B active — Hybrid Transformer-Mamba MoE 📚 Up to 1M token context
https://x.com/vllm_project/status/2062574262163280172

420.2 tok/s on a 550B model. ⚡️ Nemotron-3-Ultra-550B-A55B reaches 420.2 tok/s powered by BLACKBOX AI Inference Engine. Blackbox now delivers the fastest inference in the industry, outperforming every other provider, including on smaller-parameter models. Check our blog in
https://x.com/blackboxai/status/2062546216949588001

Are you tired of waiting 17 minutes for an AI agent to finish a code change? As an agent’s context grows, standard transformer attention can turn long runs into a bottleneck. @NVIDIAAI Nemotron 3 Ultra addresses this with a hybrid architecture that replaces several
https://x.com/baseten/status/2062609272815685759

Big day for American open models… Nemotron 3 Ultra is now the strongest US open-weight model tested, while apparently serving 300+ tok/s 🤯 Comparable large DeepSeek/Kimi models are usually 50-100 tok/s btw
https://x.com/caspar_br/status/2061505720907182280

Introducing NVIDIA Nemotron 3 Ultra. A frontier smart open model built for long-running agents that need to plan, reason, use tools and keep working across complex coding, research and enterprise workflows. Up to 5x faster inference and up to 30% lower cost for agentic tasks.
https://x.com/nvidia/status/2062522316672667770

Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-latency multilingual speech recognition. AI natives can now build coding agents, deep research agents, and real-time voice systems on the AI
https://x.com/togethercompute/status/2062520009893576974

nemotron 3 is significantly less sparse than other models (~10% active vs ~3% for kimi K2/deepseek v4)
https://x.com/eliebakouch/status/2061607195268038777

Nemotron 3 Ultra (550B-A55B) is here – our strongest open-weight model and full training recipe to date. Heavy emphasis on real-world inference efficiency for long-context agentic workloads. Everything is open 🤗: base, post-trained, reward checkpoints, NVFP4 quantized
https://x.com/PavloMolchanov/status/2062538679470657727

Nemotron 3 Ultra is now the best open weight model on
https://t.co/EJXiSfWv2O 💚
https://x.com/ctnzr/status/2061483152741175757

Nemotron 3 Ultra was launched today, including a focus on low latency agentic performance. We tested it against peers under restricted turn-usage limits on Terminal-Bench v2.1 – @NVIDIA Nemotron 3 Ultra completes tasks at a much faster pace than peers due to its high inference
https://x.com/ArtificialAnlys/status/2062598349757567359

NVIDIA has just released Nemotron 3 Ultra, the new most intelligent US open weights model, with leading speed for its intelligence Nemotron 3 Ultra scores 47.7 on the Artificial Analysis Intelligence Index, well ahead of the next strongest US open weights models, Gemma 4 31B
https://x.com/ArtificialAnlys/status/2062527871529439438

NVIDIA just announced the release of Nemotron 3 Ultra in Jensen Huang’s Computex keynote: at 550B parameters (55B active), this is the largest Nemotron 3 model to date, and it is the most intelligent US open weights model We partnered with @nvidia to evaluate this model for
https://x.com/ArtificialAnlys/status/2061304911565144230?s=20

NVIDIA just released Nemotron 3 Ultra, a 550b-parameter agentic coding model with a 1m context window. It was built for token efficiency, and is up to 5x faster and 30% cheaper than other similar models. It’s the largest US open-weights model release ever. Free in Cline now!
https://x.com/cline/status/2062620668085297214

NVIDIA Nemotron 3 Ultra is here We have Day‑0 support for Nemotron 3 Ultra in prime-rl and Lab. Specialize Nemotron 3 Ultra for your use case.
https://x.com/PrimeIntellect/status/2062622550300275088

NVIDIA Nemotron 3 Ultra is on Fireworks, day zero. Nemotron Ultra is an open model for frontier reasoning and orchestration in long-running autonomous agents. Think use cases like coding agents, deep research, and complex enterprise workflows. Read on:
https://x.com/FireworksAI_HQ/status/2062568688201646321

NVIDIA’s Nemotron 3 Ultra is available on Ollama’s cloud! Try it 👇 Claude Code: ollama launch claude –model nemotron-3-ultra:cloud Hermes Agent: ollama launch hermes –model nemotron-3-ultra:cloud OpenClaw: ollama launch openclaw –model nemotron-3-ultra:cloud
https://x.com/ollama/status/2062591290743853291

Oh wow, they pre-trained Nemotron 3 Ultra in NVFP4 big update for estimating future model sizes and flops, especially for OpenAI models
https://x.com/scaling01/status/2062540298933219832

The @nvidia Nemotron 3 Ultra is live on CW Serverless Inference 🚀 Open, frontier-reasoning, built for long-running agents — 550B params (55B active), hybrid Transformer-Mamba MoE, up to 1M context. For orchestration, coding agents & deep research. No infra to manage.
https://x.com/wandb/status/2062577626242580896

We are excited to join Nvidia’s Nemotron Coalition of leading AI labs working together to advance open frontier foundation models. To celebrate we have partnered with @nvidia and @nebiustf to provide 2 free weeks of the new Nemotron 3 Ultra model on the Nous Portal!
https://x.com/NousResearch/status/2062554136625766409

With today’s launch of Nemotron 3 Ultra, @nvidia continues to expand its investment in open-source AI. Their flagship frontier-reasoning model, built for long-running autonomous agents, is available Day 0 on Modal. – 550B with 55B active parameters – Hybrid Transformer-Mamba MoE
https://x.com/modal/status/2062528720104227149

Nemotron 3.5 ASR is built for streaming multilingual speech recognition and voice agents. One 0.6B checkpoint. 40 language-locales. Sub-100ms latency. Cache-aware FastConformer carries context forward between chunks instead of repeatedly reprocessing overlapping audio. Try it:
https://x.com/togethercompute/status/2062520605102993436

Open speech models for real-time voice agents. @NVIDIAAI Nemotron 3.5 ASR is now available on fal. An open streaming speech recognition model supporting 40 language-locale combinations with ultra-low latency, native punctuation, and capitalization. Built for voice agents,
https://x.com/fal/status/2062521027020611933

Second big release from us today: Nemotron-3.5-ASR-Streaming! 🌎40 languages ⚡️80ms – 1s controllable latency 🔥240 – 2400 concurrent streams on 1xH100 🧱FastConformer Cache-Aware RNN-T architecture
https://x.com/PiotrZelasko/status/2062538923776290909

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading