Image created with Flux Pro v1.1 Ultra. Image prompt: Maker space with wafer under a loupe and EUV masks on a light table; the word “Chips” stamped on a metal tool drawer in condensed sans; semiconductor clippings pinned beside microscope shots; precise macro, violet highlights, lab calm
Nvidia announced Cosmos Reason 7B, an open-source VLM to enable robots to see, reason, and act in the physical world, solving multistep tasks The company also made Isaac Sim 5.0 and Isaac Lab 2.2 generally available https://x.com/adcock_brett/status/1957111085481242892
when you engage “”hovercraft mode”” on your new whip (made w/ nvidia cosmos) https://x.com/bilawalsidhu/status/1956160140404777142
NVIDIA ON A ROLL! Canary 1B and Parakeet TDT (0.6B) SoTA ASR models – Multilingual, Open Source 🔥 – 1B and 600M parameters – 25 languages – automatic language detection and translation – word and sentence timestamps – transcribe up to 3 hours of audio in one go – trained on 1 https://x.com/reach_vb/status/1957148807562723809
Colossus 2, built by @xAI, will be the world’s first Gigawatt+ AI training supercomputer”” / X https://x.com/elonmusk/status/1958846872157921546
The Economic Daily (Taiwan): Foxconn will unveil a humanoid robot with an “LLM-powered brain” at Foxconn Tech Day in November. NVIDIA will support the development of the robotic brain and provide AI expertise. The robots will be deployed in Q1 2026 in Foxconn’s new Houston https://x.com/TheHumanoidHub/status/1957890628152693081
NEW: Run @huggingface models locally on your @AMD Ryzen AI and Radeon PCs, with Lemonade! 🍋 + 🤗 = 💻🧠 GG @roaner @reach_vb @AIatAMD https://x.com/jeffboudier/status/1957972077002035405
We’re excited to share that @TectonAI will soon join Databricks, providing enterprises with fast, reliable, real-time data for deploying AI agents. Tecton’s technology helps enterprises leverage their mission-critical data to power AI agents for critical use cases. Bringing https://x.com/databricks/status/1959041076087726523
Agent Bricks | Databricks on AWS https://docs.databricks.com/aws/en/generative-ai/agent-bricks/
Introducing Parallel | Web Search Infrastructure for AIs | Parallel Web Systems | Enterprise Deep Research API https://parallel.ai/blog/introducing-parallel
This is actually a pretty surprising and something that should lead companies to change how they are thinking about hosting. Model performance for the open weights GPT model vary by meaningful amounts depending on who is hosting it, with Azure & AWS being low. Worth watching.”” / X https://x.com/emollick/status/1955365624613630349
We dove into the H100’s performance improvement over time from software over 2 years. Covered power usage + $ cost for training in a very detailed way for training runs on thousands of GPUs Equating this to US household power consumption @JeffDean + GB200 reliability challenges”” / X https://x.com/dylan522p/status/1958034446789095613
🚀 Spin up Qdrant in under 10 minutes. With Docker or Python you can go from zero to a production-ready vector database: • High-throughput similarity search • Structured payload filters • A working city-similarity finder Sustaining ~24 ms search latency across millions of https://x.com/qdrant_engine/status/1957027122133835847
Databricks is raising a Series K Investment at >$100 billion valuation – Databricks https://www.databricks.com/company/newsroom/press-releases/databricks-raising-series-k-investment-100-billion-valuation
MoE layers can be really slow. When training our coding models @cursor_ai, they ate up 27–53% of training time. So we completely rebuilt it at the kernel level and transitioned to MXFP8. The result: 3.5x faster MoE layer and 1.5x end-to-end training speedup. We believe our https://x.com/stuart_sul/status/1957927497351467372
The world’s fastest MXFP8 MoE kernels!”” / X https://x.com/amanrsanger/status/1957932614746304898
Today we’re putting out an update to the JAX TPU book, this time on GPUs. How do GPUs work, especially compared to TPUs? How are they networked? And how does this affect LLM training? 1/n https://x.com/jacobaustin132/status/1957447351011840336
When it comes to hardware that’s meant for training or inference, most think about in hardware specs like memory bandwidth even though dev velocity is often a more important factor. One implication is that RL training and prod. inference are meaningfully different workloads.”” / X https://x.com/cHHillee/status/1956911060646072677
Databricks just signed a Series K term sheet at >$100B valuation to scale two flagship products: 🔥 Lakebase — serverless Postgres with true compute/storage separation 🧠 Agent Bricks — agentic framework with built-in reasoning guardrails for enterprise data”” / X https://x.com/alighodsi/status/1957795160416309717
@TheZachMueller Thunderbolt 4 has a maximum bandwidth of 40 Gb/s. However, this is a bit misleading because not all of that bandwidth can be used for data transfer. Approximately 8 Gbp/s can only be used for video, leaving 32 Gb/s for non-video data (PCIe 3.0: 4 lanes x 8 Gb/s). you good”” / X https://x.com/alphatozeta8148/status/1958930594370658369
Frontier AI performance typically reaches consumer hardware in just 9 months. With a single gaming GPU, you can run open-weight models matching the benchmark performance of the absolute frontier from less than a year ago. 🧵 https://x.com/EpochAIResearch/status/1956468453399044375
We’re starting a series on Multi GPU talks on Aug 16 at noon PST On Aug 16 we’ll have Jeff Hammond one of the maintainers of NCCL giving us a talk on Multi GPU programming, On Aug 22 we’ll be hearing from Didem Unat who will give us a broad overview of all GPU centric”” / X https://x.com/GPU_MODE/status/1956590989575119048
FlashAttention v4 is coming to Blackwell GPUs”” / X https://x.com/scaling01/status/1957397971479200083
We started the distributed summer at @GPU_MODE pretty strong, with Jeff Hammond from @nvidia talking about nccl/nvshmem. One of the best talks I had the pleasure to see til now. https://x.com/m_sirovatka/status/1956824361819652175
🚀 Exciting news: DeepSeek-V3.1 from @deepseek_ai now runs on vLLM! 🧠 Seamlessly toggle Think / Non-Think mode per request ⚡ Powered by vLLM’s efficient serving — scale to multi-GPU with ease 🛠️ Perfect for agents, tools, and fast reasoning workloads 👉 Guide & examples: https://x.com/vllm_project/status/1958580047658491947
We’re setting up a LOAD of TPUs today and warming them up. Ever wanted to try Veo 3 in @GeminiApp?”” / X https://x.com/joshwoodward/status/1958555951344345461
No names necessary. You know the man. You know the jacket. You know the company. But how did a napkin sketch turn into one of the most valuable companies on earth… Now powering the AI & Robotics infrastructure? 🧵👇 https://x.com/IlirAliu_/status/1957081211970396482
HUGE RELEASE! Nvidia just droppped: > Granary: the largest open-source speech dataset for European languages 🗣️🇪🇺 > Canary-1b-v2: 25 languages, ASR + En↔X translation > Parakeet-tdt-0.6b-v3: SOTA multilingual ASR You can now train your ASR model to understand European https://x.com/Tu7uruu/status/1956350036343701583
Nvidia Parakeet v3 is out! Enjoy Day 0 support with Argmax SDK – What changed from v2? – How do I use it? – Should I upgrade to this model right away? Answers in comments https://x.com/argmaxinc/status/1956385793892917288
Jensen visited the Figure HQ. https://x.com/TheHumanoidHub/status/1957690025057087933
NVIDIA Nemotron Nano v2 – a 9B hybrid SSM that is 6X faster than similarly sized models, while also being more accurate. 💚💚💚 9B: https://x.com/ClementDelangue/status/1957519608992407848
nvidia parakeet-tdt-0.6b-v3 600M model here: https://x.com/reach_vb/status/1957149090913128598
🚨 New: We built @a16z’s personal GPU AI Workstation Founders Edition – 4x NVIDIA RTX 6000 PRO Blackwell Max-Q (384GB total VRAM) – 8TB of NVMe PCIe 5.0 storage – AMD Threadripper PRO 7975WX (32 cores, 64 threads) – 256GB ECC DDR5 RAM – 1650Watts at peak (runs on a standard https://x.com/Mascobot/status/1958925710988582998
NVIDIA Nemotron Nano 2 An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model https://x.com/_akhaliq/status/1958545622618788174
Today we’re releasing NVIDIA Nemotron Nano v2 – a 9B hybrid SSM that is 6X faster than similarly sized models, while also being more accurate. Along with this model, we are also releasing most of the data we used to create it, including the pretraining corpus. Links to the https://x.com/ctnzr/status/1957504768156561413
Nvidia dropping model that rivals qwen 3 8b, with data, with base model, not that bad of a license (could be better to be clear) a big win, love to see it. Hopefully is well integrated into open tools and “”easy to finetune”” etc, which is hard to measure”” / X https://x.com/natolambert/status/1957517030929887284
[episode 120 of frontier lab gossip: OH in the mission] > **** folks recently bragged about a 100k H100 training run > wrote a post, got all the likes and shares internally > then some of them ran the exact same job on 20k H100s instead of 100k, and ended up with the exact same”” / X https://x.com/suchenzang/status/1956851798221996178
China reportedly discouraged purchase of NVIDIA AI chips due to ‘insulting’ Lutnick statements https://www.engadget.com/ai/china-reportedly-discouraged-purchase-of-nvidia-ai-chips-due-to-insulting-lutnick-statements-123055120.html
🚀 New example: multi-node serving for 1T+ param models! • Serve large models like @Kimi_Moonshot K2 with full context length • Combine tensor and pipeline parallelism with @vllm_project • SkyPilot simplifies multi-node setup and scales up replicas https://x.com/skypilot_org/status/1957831495462379743
One of the most underrated players in AI models, @IBM, released 2 new extremely efficient embedding models: granite-embedding-english-r2 & granite-embedding-small-english-r2, commercially viable. Details in 🧵: https://x.com/tomaarsen/status/1957389356412330282
NVIDIA-Nemotron-Nano-2-Technical-Report.pdf https://research.nvidia.com/labs/adlr/files/NVIDIA-Nemotron-Nano-2-Technical-Report.pdf




