Image created with Flux Pro v1.1 Ultra. Image prompt: Maker space with wafer under a loupe and EUV masks on a light table; the word “Chips” stamped on a metal tool drawer in condensed sans; semiconductor clippings pinned beside microscope shots; precise macro, violet highlights, lab calm

Nvidia announced Cosmos Reason 7B, an open-source VLM to enable robots to see, reason, and act in the physical world, solving multistep tasks The company also made Isaac Sim 5.0 and Isaac Lab 2.2 generally available https://x.com/adcock_brett/status/1957111085481242892

when you engage “”hovercraft mode”” on your new whip (made w/ nvidia cosmos) https://x.com/bilawalsidhu/status/1956160140404777142

NVIDIA ON A ROLL! Canary 1B and Parakeet TDT (0.6B) SoTA ASR models – Multilingual, Open Source 🔥 – 1B and 600M parameters – 25 languages – automatic language detection and translation – word and sentence timestamps – transcribe up to 3 hours of audio in one go – trained on 1 https://x.com/reach_vb/status/1957148807562723809

Colossus 2, built by @xAI, will be the world’s first Gigawatt+ AI training supercomputer”” / X https://x.com/elonmusk/status/1958846872157921546

The Economic Daily (Taiwan): Foxconn will unveil a humanoid robot with an “LLM-powered brain” at Foxconn Tech Day in November. NVIDIA will support the development of the robotic brain and provide AI expertise. The robots will be deployed in Q1 2026 in Foxconn’s new Houston https://x.com/TheHumanoidHub/status/1957890628152693081

NEW: Run @huggingface models locally on your @AMD Ryzen AI and Radeon PCs, with Lemonade! 🍋 + 🤗 = 💻🧠 GG @roaner @reach_vb @AIatAMD https://x.com/jeffboudier/status/1957972077002035405

We’re excited to share that @TectonAI will soon join Databricks, providing enterprises with fast, reliable, real-time data for deploying AI agents. Tecton’s technology helps enterprises leverage their mission-critical data to power AI agents for critical use cases. Bringing https://x.com/databricks/status/1959041076087726523

Agent Bricks | Databricks on AWS https://docs.databricks.com/aws/en/generative-ai/agent-bricks/

Introducing Parallel | Web Search Infrastructure for AIs | Parallel Web Systems | Enterprise Deep Research API https://parallel.ai/blog/introducing-parallel

This is actually a pretty surprising and something that should lead companies to change how they are thinking about hosting. Model performance for the open weights GPT model vary by meaningful amounts depending on who is hosting it, with Azure & AWS being low. Worth watching.”” / X https://x.com/emollick/status/1955365624613630349

We dove into the H100’s performance improvement over time from software over 2 years. Covered power usage + $ cost for training in a very detailed way for training runs on thousands of GPUs Equating this to US household power consumption @JeffDean + GB200 reliability challenges”” / X https://x.com/dylan522p/status/1958034446789095613

🚀 Spin up Qdrant in under 10 minutes. With Docker or Python you can go from zero to a production-ready vector database: • High-throughput similarity search • Structured payload filters • A working city-similarity finder Sustaining ~24 ms search latency across millions of https://x.com/qdrant_engine/status/1957027122133835847

Databricks is raising a Series K Investment at >$100 billion valuation – Databricks https://www.databricks.com/company/newsroom/press-releases/databricks-raising-series-k-investment-100-billion-valuation

MoE layers can be really slow. When training our coding models @cursor_ai, they ate up 27–53% of training time. So we completely rebuilt it at the kernel level and transitioned to MXFP8. The result: 3.5x faster MoE layer and 1.5x end-to-end training speedup. We believe our https://x.com/stuart_sul/status/1957927497351467372

The world’s fastest MXFP8 MoE kernels!”” / X https://x.com/amanrsanger/status/1957932614746304898

Today we’re putting out an update to the JAX TPU book, this time on GPUs. How do GPUs work, especially compared to TPUs? How are they networked? And how does this affect LLM training? 1/n https://x.com/jacobaustin132/status/1957447351011840336

When it comes to hardware that’s meant for training or inference, most think about in hardware specs like memory bandwidth even though dev velocity is often a more important factor. One implication is that RL training and prod. inference are meaningfully different workloads.”” / X https://x.com/cHHillee/status/1956911060646072677

Databricks just signed a Series K term sheet at >$100B valuation to scale two flagship products: 🔥 Lakebase — serverless Postgres with true compute/storage separation 🧠 Agent Bricks — agentic framework with built-in reasoning guardrails for enterprise data”” / X https://x.com/alighodsi/status/1957795160416309717

@TheZachMueller Thunderbolt 4 has a maximum bandwidth of 40 Gb/s. However, this is a bit misleading because not all of that bandwidth can be used for data transfer. Approximately 8 Gbp/s can only be used for video, leaving 32 Gb/s for non-video data (PCIe 3.0: 4 lanes x 8 Gb/s). you good”” / X https://x.com/alphatozeta8148/status/1958930594370658369

Frontier AI performance typically reaches consumer hardware in just 9 months. With a single gaming GPU, you can run open-weight models matching the benchmark performance of the absolute frontier from less than a year ago. 🧵 https://x.com/EpochAIResearch/status/1956468453399044375

We’re starting a series on Multi GPU talks on Aug 16 at noon PST On Aug 16 we’ll have Jeff Hammond one of the maintainers of NCCL giving us a talk on Multi GPU programming, On Aug 22 we’ll be hearing from Didem Unat who will give us a broad overview of all GPU centric”” / X https://x.com/GPU_MODE/status/1956590989575119048

FlashAttention v4 is coming to Blackwell GPUs”” / X https://x.com/scaling01/status/1957397971479200083

We started the distributed summer at @GPU_MODE pretty strong, with Jeff Hammond from @nvidia talking about nccl/nvshmem. One of the best talks I had the pleasure to see til now. https://x.com/m_sirovatka/status/1956824361819652175

🚀 Exciting news: DeepSeek-V3.1 from @deepseek_ai now runs on vLLM! 🧠 Seamlessly toggle Think / Non-Think mode per request ⚡ Powered by vLLM’s efficient serving — scale to multi-GPU with ease 🛠️ Perfect for agents, tools, and fast reasoning workloads 👉 Guide & examples: https://x.com/vllm_project/status/1958580047658491947

We’re setting up a LOAD of TPUs today and warming them up. Ever wanted to try Veo 3 in @GeminiApp?”” / X https://x.com/joshwoodward/status/1958555951344345461

No names necessary. You know the man. You know the jacket. You know the company. But how did a napkin sketch turn into one of the most valuable companies on earth… Now powering the AI & Robotics infrastructure? 🧵👇 https://x.com/IlirAliu_/status/1957081211970396482

HUGE RELEASE! Nvidia just droppped: > Granary: the largest open-source speech dataset for European languages 🗣️🇪🇺 > Canary-1b-v2: 25 languages, ASR + En↔X translation > Parakeet-tdt-0.6b-v3: SOTA multilingual ASR You can now train your ASR model to understand European https://x.com/Tu7uruu/status/1956350036343701583

Nvidia Parakeet v3 is out! Enjoy Day 0 support with Argmax SDK – What changed from v2? – How do I use it? – Should I upgrade to this model right away? Answers in comments https://x.com/argmaxinc/status/1956385793892917288

Jensen visited the Figure HQ. https://x.com/TheHumanoidHub/status/1957690025057087933

NVIDIA Nemotron Nano v2 – a 9B hybrid SSM that is 6X faster than similarly sized models, while also being more accurate. 💚💚💚 9B: https://x.com/ClementDelangue/status/1957519608992407848

nvidia parakeet-tdt-0.6b-v3 600M model here: https://x.com/reach_vb/status/1957149090913128598

🚨 New: We built @a16z’s personal GPU AI Workstation Founders Edition – 4x NVIDIA RTX 6000 PRO Blackwell Max-Q (384GB total VRAM) – 8TB of NVMe PCIe 5.0 storage – AMD Threadripper PRO 7975WX (32 cores, 64 threads) – 256GB ECC DDR5 RAM – 1650Watts at peak (runs on a standard https://x.com/Mascobot/status/1958925710988582998

NVIDIA Nemotron Nano 2 An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model https://x.com/_akhaliq/status/1958545622618788174

Today we’re releasing NVIDIA Nemotron Nano v2 – a 9B hybrid SSM that is 6X faster than similarly sized models, while also being more accurate. Along with this model, we are also releasing most of the data we used to create it, including the pretraining corpus. Links to the https://x.com/ctnzr/status/1957504768156561413

Nvidia dropping model that rivals qwen 3 8b, with data, with base model, not that bad of a license (could be better to be clear) a big win, love to see it. Hopefully is well integrated into open tools and “”easy to finetune”” etc, which is hard to measure”” / X https://x.com/natolambert/status/1957517030929887284

[episode 120 of frontier lab gossip: OH in the mission] > **** folks recently bragged about a 100k H100 training run > wrote a post, got all the likes and shares internally > then some of them ran the exact same job on 20k H100s instead of 100k, and ended up with the exact same”” / X https://x.com/suchenzang/status/1956851798221996178

China reportedly discouraged purchase of NVIDIA AI chips due to ‘insulting’ Lutnick statements https://www.engadget.com/ai/china-reportedly-discouraged-purchase-of-nvidia-ai-chips-due-to-insulting-lutnick-statements-123055120.html

🚀 New example: multi-node serving for 1T+ param models! • Serve large models like @Kimi_Moonshot K2 with full context length • Combine tensor and pipeline parallelism with @vllm_project • SkyPilot simplifies multi-node setup and scales up replicas https://x.com/skypilot_org/status/1957831495462379743

One of the most underrated players in AI models, @IBM, released 2 new extremely efficient embedding models: granite-embedding-english-r2 & granite-embedding-small-english-r2, commercially viable. Details in 🧵: https://x.com/tomaarsen/status/1957389356412330282

NVIDIA-Nemotron-Nano-2-Technical-Report.pdf https://research.nvidia.com/labs/adlr/files/NVIDIA-Nemotron-Nano-2-Technical-Report.pdf

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading