Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Create a 16:9 cinematic split-screen poster. LEFT SIDE (40% width): – A close view of a small antistatic tray holding several computer chips and circuit boards on a clean workbench, with a magnifying glass and a tiny screwdriver nearby. – The background is a turquoise / teal abstract field made of stylized blue rods or data fibers, subtly reinforcing the idea of data pathways. – Use natural or soft bench lighting. No glowing traces, no neon circuits. RIGHT SIDE (60% width): – A green-toned abstract aerial forest canopy texture, suggesting energy and life. – Two clean rounded rectangles stacked vertically near the center-right. – The TOP rectangle contains the text: “Chips”. – The BOTTOM rectangle contains the text: “2025/10/10”. – Clean sans-serif font with dark green or charcoal text. OVERALL STYLE: – Very tactile and physical, focusing on real hardware. – No logos or extra copy. – Keep the turquoise/forest split-screen layout.

Congrats to Michel Devoret, John Martinis, and John Clarke on the Nobel Prize in Physics. 🔬🥼 Michel is chief scientist of hardware at our Quantum AI lab and John Martinis led the hardware team for many years. Their pioneering work in quantum mechanics in the 1980s made recent”” / X https://x.com/sundarpichai/status/1975590130690781463

Congratulations to Michel Devoret, Google Quantum AI’s Chief Scientist of Quantum Hardware, who was awarded the 2025 Nobel Prize in Physics today. Google now has five Nobel laureates among our ranks, including three prizes in the past two years. https://x.com/Google/status/1975623817943752714

BREAKING NEWS The Royal Swedish Academy of Sciences has decided to award the 2025 #NobelPrize in Chemistry to Susumu Kitagawa, Richard Robson and Omar M. Yaghi “for the development of metal–organic frameworks.” https://x.com/NobelPrize/status/1975860703857680729

BREAKING NEWS The Royal Swedish Academy of Sciences has decided to award the 2025 #NobelPrize in Physics to John Clarke, Michel H. Devoret and John M. Martinis “for the discovery of macroscopic quantum mechanical tunnelling and energy quantisation in an electric circuit.” https://x.com/NobelPrize/status/1975498493218394168

5 things: Nvidia’s Huang on the state of the AI race with China https://www.cnbc.com/2025/10/08/nvidia-huang-ai-race-china-us-trump.html

🚨Breaking: OpenAI has partnered with AMD in a massive five-year GPU supply agreement to deploy 6 Gigawatts of AMD GPUs over multiple years. Deal Highlights: – The deal gives AMD a seat at the table as one of OpenAI’s core compute suppliers – AMD issued OpenAI a warrant for up https://x.com/TheRundownAI/status/1975205711715181022

AMD and OpenAI announce strategic partnership to deploy 6 gigawatts of AMD GPUs | OpenAI https://openai.com/index/openai-amd-strategic-partnership/

Exciting day today! Thrilled to partner with @OpenAI to deploy 6GWs of AMD Instinct GPUs. The world needs more AI compute. Together, we’re bringing the best of both companies to accelerate the global AI infrastructure buildout. Thanks @sama @gdb for the trust and partnership.”” / X https://x.com/LisaSu/status/1975210493796385233

Wall Street analysts explain how AMD’s own stock will pay for OpenAI’s billions in chip purchases  | TechCrunch https://techcrunch.com/2025/10/07/wall-street-analysts-explain-how-amds-own-stock-will-pay-for-openais-billions-in-chip-purchases/

Even after Stargate, Oracle, Nvidia, and AMD, OpenAI has more big deals coming soon, Sam Altman says | TechCrunch https://techcrunch.com/2025/10/08/even-after-stargate-oracle-nvidia-and-amd-openai-has-more-big-deals-coming-soon-sam-altman-says/

Musk’s xAI nears $20 billion capital raise tied to Nvidia chips, Bloomberg News reports https://finance.yahoo.com/news/musks-xai-nears-20-billion-232913241.html

RL X-mas came early. 🎄 For too long, building powerful AI agents with Reinforcement Learning has been blocked by GPU scarcity and complex infrastructure. That ends today. Introducing Serverless RL from wandb, powered by @CoreWeave! We’re making RL accessible to all. https://x.com/weights_biases/status/1975996733269344571

Sept 8: @CoreWeave acquires @OpenPipeAI Oct 8: @OpenPipeAI ships Serverless RL”” / X https://x.com/shawnup/status/1975993379826827697

I wonder how much of the spend on AI infrastructure is because it is otherwise very hard to get market exposure to the possibility of transformative AI. There are only a few companies in that AI race, so if you want “AGI” hedges in your portfolio, it is data centers or nothing?”” / X https://x.com/emollick/status/1975208098362200223

Groq has come out with OpenBench – a cool way to run a bunch of benchmarks They added support for ARC-AGI”” / X https://x.com/GregKamradt/status/1976718318544601573

This is the power of a unified AI stack. Go from idea to training in minutes, not days. We handle the complexity so you can focus on building. Dive into the full announcement and start training your first agent today: https://x.com/CoreWeave/status/1976008192816513455

today i learned that H100 instances will cost you like $9+ per hour on Azure, AWS and so on i was living in a $2/hour world”” / X https://x.com/scaling01/status/1975598023834280111

Today we are launching InferenceMAX! We have support from Nvidia, AMD, OpenAI, Microsoft, Pytorch, SGLang, vLLM, Oracle, CoreWeave, TogetherAI, Nebius, Crusoe, HPE, SuperMicro, Dell It runs every day on the latest software (vLLM, SGLang, etc) across hundreds of GPUs, $10Ms of”” / X https://x.com/dylan522p/status/1976422855928680454

Here’s some MLX code that can help understand what a matmul with CUDA tensor cores is doing on the GPU. If you understand this you understand the core algorithm. Most of the rest of the complexity comes from efficiently loading the data and feeding the GPU in a pipeline and https://x.com/awnihannun/status/1976347648014811634

On the surface you’d think that the convergence of model architecture to the Transformer would open the door for specialized hardware. But somehow it feels like general purpose hardware (GP in GPGPU) is more useful now than ever. Like back in the RNN and conv days it was”” / X https://x.com/awnihannun/status/1976715815019037101

Online training methods (e.g., GRPO) require real-time completion generation, a compute- and memory-heavy bottleneck. TRL has built-in vLLM support and in this new recipe, we show how to leverage it for efficient online training. Run on Colab ⚡, scale to multi-GPU/multi-node! https://x.com/SergioPaniego/status/1975498366084923899

Shout out to ThunderKittens for writing simple yet very performant GPU code. We’re working on “”tinykittens”” which uses the same insight but in tinygrad’s language. The insight is that GPU “”registers”” are the wrong primitive and TK’s “”register tile”” is a lot more sensible. https://x.com/__tinygrad__/status/1976084605141909845

Super happy to announce that University of Zurich (@UZH_en) just joined @huggingface Academia Hub 🇨🇭🎉 Their students and educators get a better access to collaboration and compute features (including ZeroGPU power) on the Hub. 🔥 https://x.com/julien_c/status/1975515541700841935

@finbarrtimbers Not sure if this fits the bill, but it’s what we use to stress test new nodes on our cluster https://x.com/_lewtun/status/1975403104586563625

4. In terms of efficiency, Delethink is faster and cheaper than LongCoT-RL: – 1 RL step takes 215s vs. 249s for LongCoT-RL – It generates 8,500 vs. 6,000 tokens/sec on an H100 GPU – At test time, Delethink keeps improving even when reasoning far beyond its training length, https://x.com/TheTuringPost/status/1976798717094379588

Big shoutout to the @vllm_project team for an exceptional showing in the SemiAnalysis InferenceMAX benchmark on NVIDIA Blackwell GPUs 👏 Built through close collaboration with our engineers, vLLM delivered consistently strong Blackwell performance gains across the Pareto”” / X https://x.com/NVIDIAAIDev/status/1976686560398426456

Every Together Instant Cluster is stress-tested for reliability: burn‑in → NVIDIA NVLink checks → NCCL all‑reduce validation. ✅ https://x.com/togethercompute/status/1975965240144888301

Happy that InferenceMAX is here because it signals a milestone for vLLM’s SOTA performance on NVIDIA Blackwell! 🥳 It has been a pleasure to deeply collaborate with @nvidia in @vllm_project, and we have much more to do Read about the work we did here: https://x.com/mgoin_/status/1976452383258648972

Nvidia B200s are now available in @huggingface Inference Endpoints! The world needs more compute 😅😅😅 https://x.com/ClementDelangue/status/1975266333949604237

Excited to partner with AMD to use their chips to serve our users! This is all incremental to our work with NVIDIA (and we plan to increase our NVIDIA purchasing over time). The world needs much more compute…”” / X https://x.com/sama/status/1975185516225278428

New data insight: How does OpenAI allocate its compute? OpenAI spent ~$7 billion on compute last year. Most of this went to R&D, meaning all research, experiments, and training. Only a minority of this R&D compute went to the final training runs of released models. https://x.com/EpochAIResearch/status/1976714284349767990

Overall compute spend is from media accounts of OpenAI’s investor reporting. We estimated training costs for released models using evidence on GPT-4.5’s training cluster and compute estimates for other models. These likely add up to <$1 billion, compared to a ~$5B R&D total.”” / X https://x.com/EpochAIResearch/status/1976714297255588053

Two visionaries in robotics are joining forces this Tuesday at 10 AM PT. ⭐ @DrJimFan x @drfeifei They’ll dive into BEHAVIOR, a groundbreaking new benchmark reshaping the future of embodied AI. Come ready to learn, get inspired, and ask your biggest questions. 💡 🗓️ Add to https://x.com/NVIDIARobotics/status/1975367246265414071

ATLAS delivers 400% faster LLM inference by learning from your workloads in real-time ⚡ From today’s coverage on @VentureBeat: “”The shift from static to adaptive optimization represents a fundamental rethinking of how inference platforms should work.”” We couldn’t agree more! https://x.com/togethercompute/status/1976743626685530540

Operations: isend/irecv async collectives allow work to continue around the actual movement of data, to help reduce your wait time as GPUs gather data from other processes. One such example of this is the isend/irecv paradigm (as opposed to send/recv from the other day) With https://x.com/TheZachMueller/status/1975558921193484423

Remembering how each torch.distributed operation behaves can be difficult. Here’s an easy PDF with visuals of each kind showcasing how tensors get modified throughout all of the common MPI operations. Download the PDF for free in the comments! https://x.com/TheZachMueller/status/1975624506262851676

Unlike static speculators, ATLAS (AdapTive-LeArning Speculator System) learns from your live traffic in real time, automatically adapting as your workload evolves.”” / X https://x.com/togethercompute/status/1976655647925215339

🚀 Big launch from @OpenPipeAI: We just launched Serverless RL — train agents faster and cheaper with zero infra headaches. Compared to running your own GPUs, Serverless RL is: – 40% cheaper – 28% faster wall‑clock – instantly deployed to prod via @weights_biases Inference https://x.com/corbtt/status/1975990784404115692

🚀 TensorRT-LLM hit its v1.0 milestone — a culmination of 4 years of architecture pivots, cross-continent teamwork, and relentless optimization at NVIDIA. What started as a small team optimizing ONNX runtime has grown into a full-scale, PyTorch-native inference system powering https://x.com/ZhihuFrontier/status/1974559265273639349

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading