Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Create a 16:9 cinematic split-screen poster. LEFT SIDE (40% width): – A close view of a small antistatic tray holding several computer chips and circuit boards on a clean workbench, with a magnifying glass and a tiny screwdriver nearby. – The background is a turquoise / teal abstract field made of stylized blue rods or data fibers, subtly reinforcing the idea of data pathways. – Use natural or soft bench lighting. No glowing traces, no neon circuits. RIGHT SIDE (60% width): – A green-toned abstract aerial forest canopy texture, suggesting energy and life. – Two clean rounded rectangles stacked vertically near the center-right. – The TOP rectangle contains the text: “Chips”. – The BOTTOM rectangle contains the text: “2025/10/10”. – Clean sans-serif font with dark green or charcoal text. OVERALL STYLE: – Very tactile and physical, focusing on real hardware. – No logos or extra copy. – Keep the turquoise/forest split-screen layout.
Congrats to Michel Devoret, John Martinis, and John Clarke on the Nobel Prize in Physics. 🔬🥼 Michel is chief scientist of hardware at our Quantum AI lab and John Martinis led the hardware team for many years. Their pioneering work in quantum mechanics in the 1980s made recent”” / X https://x.com/sundarpichai/status/1975590130690781463
Congratulations to Michel Devoret, Google Quantum AI’s Chief Scientist of Quantum Hardware, who was awarded the 2025 Nobel Prize in Physics today. Google now has five Nobel laureates among our ranks, including three prizes in the past two years. https://x.com/Google/status/1975623817943752714
BREAKING NEWS The Royal Swedish Academy of Sciences has decided to award the 2025 #NobelPrize in Chemistry to Susumu Kitagawa, Richard Robson and Omar M. Yaghi “for the development of metal–organic frameworks.” https://x.com/NobelPrize/status/1975860703857680729
BREAKING NEWS The Royal Swedish Academy of Sciences has decided to award the 2025 #NobelPrize in Physics to John Clarke, Michel H. Devoret and John M. Martinis “for the discovery of macroscopic quantum mechanical tunnelling and energy quantisation in an electric circuit.” https://x.com/NobelPrize/status/1975498493218394168
5 things: Nvidia’s Huang on the state of the AI race with China https://www.cnbc.com/2025/10/08/nvidia-huang-ai-race-china-us-trump.html
🚨Breaking: OpenAI has partnered with AMD in a massive five-year GPU supply agreement to deploy 6 Gigawatts of AMD GPUs over multiple years. Deal Highlights: – The deal gives AMD a seat at the table as one of OpenAI’s core compute suppliers – AMD issued OpenAI a warrant for up https://x.com/TheRundownAI/status/1975205711715181022
AMD and OpenAI announce strategic partnership to deploy 6 gigawatts of AMD GPUs | OpenAI https://openai.com/index/openai-amd-strategic-partnership/
Exciting day today! Thrilled to partner with @OpenAI to deploy 6GWs of AMD Instinct GPUs. The world needs more AI compute. Together, we’re bringing the best of both companies to accelerate the global AI infrastructure buildout. Thanks @sama @gdb for the trust and partnership.”” / X https://x.com/LisaSu/status/1975210493796385233
Wall Street analysts explain how AMD’s own stock will pay for OpenAI’s billions in chip purchases | TechCrunch https://techcrunch.com/2025/10/07/wall-street-analysts-explain-how-amds-own-stock-will-pay-for-openais-billions-in-chip-purchases/
Even after Stargate, Oracle, Nvidia, and AMD, OpenAI has more big deals coming soon, Sam Altman says | TechCrunch https://techcrunch.com/2025/10/08/even-after-stargate-oracle-nvidia-and-amd-openai-has-more-big-deals-coming-soon-sam-altman-says/
Musk’s xAI nears $20 billion capital raise tied to Nvidia chips, Bloomberg News reports https://finance.yahoo.com/news/musks-xai-nears-20-billion-232913241.html
RL X-mas came early. 🎄 For too long, building powerful AI agents with Reinforcement Learning has been blocked by GPU scarcity and complex infrastructure. That ends today. Introducing Serverless RL from wandb, powered by @CoreWeave! We’re making RL accessible to all. https://x.com/weights_biases/status/1975996733269344571
Sept 8: @CoreWeave acquires @OpenPipeAI Oct 8: @OpenPipeAI ships Serverless RL”” / X https://x.com/shawnup/status/1975993379826827697
I wonder how much of the spend on AI infrastructure is because it is otherwise very hard to get market exposure to the possibility of transformative AI. There are only a few companies in that AI race, so if you want “AGI” hedges in your portfolio, it is data centers or nothing?”” / X https://x.com/emollick/status/1975208098362200223
Groq has come out with OpenBench – a cool way to run a bunch of benchmarks They added support for ARC-AGI”” / X https://x.com/GregKamradt/status/1976718318544601573
This is the power of a unified AI stack. Go from idea to training in minutes, not days. We handle the complexity so you can focus on building. Dive into the full announcement and start training your first agent today: https://x.com/CoreWeave/status/1976008192816513455
today i learned that H100 instances will cost you like $9+ per hour on Azure, AWS and so on i was living in a $2/hour world”” / X https://x.com/scaling01/status/1975598023834280111
Today we are launching InferenceMAX! We have support from Nvidia, AMD, OpenAI, Microsoft, Pytorch, SGLang, vLLM, Oracle, CoreWeave, TogetherAI, Nebius, Crusoe, HPE, SuperMicro, Dell It runs every day on the latest software (vLLM, SGLang, etc) across hundreds of GPUs, $10Ms of”” / X https://x.com/dylan522p/status/1976422855928680454
Here’s some MLX code that can help understand what a matmul with CUDA tensor cores is doing on the GPU. If you understand this you understand the core algorithm. Most of the rest of the complexity comes from efficiently loading the data and feeding the GPU in a pipeline and https://x.com/awnihannun/status/1976347648014811634
On the surface you’d think that the convergence of model architecture to the Transformer would open the door for specialized hardware. But somehow it feels like general purpose hardware (GP in GPGPU) is more useful now than ever. Like back in the RNN and conv days it was”” / X https://x.com/awnihannun/status/1976715815019037101
Online training methods (e.g., GRPO) require real-time completion generation, a compute- and memory-heavy bottleneck. TRL has built-in vLLM support and in this new recipe, we show how to leverage it for efficient online training. Run on Colab ⚡, scale to multi-GPU/multi-node! https://x.com/SergioPaniego/status/1975498366084923899
Shout out to ThunderKittens for writing simple yet very performant GPU code. We’re working on “”tinykittens”” which uses the same insight but in tinygrad’s language. The insight is that GPU “”registers”” are the wrong primitive and TK’s “”register tile”” is a lot more sensible. https://x.com/__tinygrad__/status/1976084605141909845
Super happy to announce that University of Zurich (@UZH_en) just joined @huggingface Academia Hub 🇨🇭🎉 Their students and educators get a better access to collaboration and compute features (including ZeroGPU power) on the Hub. 🔥 https://x.com/julien_c/status/1975515541700841935
@finbarrtimbers Not sure if this fits the bill, but it’s what we use to stress test new nodes on our cluster https://x.com/_lewtun/status/1975403104586563625
4. In terms of efficiency, Delethink is faster and cheaper than LongCoT-RL: – 1 RL step takes 215s vs. 249s for LongCoT-RL – It generates 8,500 vs. 6,000 tokens/sec on an H100 GPU – At test time, Delethink keeps improving even when reasoning far beyond its training length, https://x.com/TheTuringPost/status/1976798717094379588
Big shoutout to the @vllm_project team for an exceptional showing in the SemiAnalysis InferenceMAX benchmark on NVIDIA Blackwell GPUs 👏 Built through close collaboration with our engineers, vLLM delivered consistently strong Blackwell performance gains across the Pareto”” / X https://x.com/NVIDIAAIDev/status/1976686560398426456
Every Together Instant Cluster is stress-tested for reliability: burn‑in → NVIDIA NVLink checks → NCCL all‑reduce validation. ✅ https://x.com/togethercompute/status/1975965240144888301
Happy that InferenceMAX is here because it signals a milestone for vLLM’s SOTA performance on NVIDIA Blackwell! 🥳 It has been a pleasure to deeply collaborate with @nvidia in @vllm_project, and we have much more to do Read about the work we did here: https://x.com/mgoin_/status/1976452383258648972
Nvidia B200s are now available in @huggingface Inference Endpoints! The world needs more compute 😅😅😅 https://x.com/ClementDelangue/status/1975266333949604237
Excited to partner with AMD to use their chips to serve our users! This is all incremental to our work with NVIDIA (and we plan to increase our NVIDIA purchasing over time). The world needs much more compute…”” / X https://x.com/sama/status/1975185516225278428
New data insight: How does OpenAI allocate its compute? OpenAI spent ~$7 billion on compute last year. Most of this went to R&D, meaning all research, experiments, and training. Only a minority of this R&D compute went to the final training runs of released models. https://x.com/EpochAIResearch/status/1976714284349767990
Overall compute spend is from media accounts of OpenAI’s investor reporting. We estimated training costs for released models using evidence on GPT-4.5’s training cluster and compute estimates for other models. These likely add up to <$1 billion, compared to a ~$5B R&D total.”” / X https://x.com/EpochAIResearch/status/1976714297255588053
Two visionaries in robotics are joining forces this Tuesday at 10 AM PT. ⭐ @DrJimFan x @drfeifei They’ll dive into BEHAVIOR, a groundbreaking new benchmark reshaping the future of embodied AI. Come ready to learn, get inspired, and ask your biggest questions. 💡 🗓️ Add to https://x.com/NVIDIARobotics/status/1975367246265414071
ATLAS delivers 400% faster LLM inference by learning from your workloads in real-time ⚡ From today’s coverage on @VentureBeat: “”The shift from static to adaptive optimization represents a fundamental rethinking of how inference platforms should work.”” We couldn’t agree more! https://x.com/togethercompute/status/1976743626685530540
Operations: isend/irecv async collectives allow work to continue around the actual movement of data, to help reduce your wait time as GPUs gather data from other processes. One such example of this is the isend/irecv paradigm (as opposed to send/recv from the other day) With https://x.com/TheZachMueller/status/1975558921193484423
Remembering how each torch.distributed operation behaves can be difficult. Here’s an easy PDF with visuals of each kind showcasing how tensors get modified throughout all of the common MPI operations. Download the PDF for free in the comments! https://x.com/TheZachMueller/status/1975624506262851676
Unlike static speculators, ATLAS (AdapTive-LeArning Speculator System) learns from your live traffic in real time, automatically adapting as your workload evolves.”” / X https://x.com/togethercompute/status/1976655647925215339
🚀 Big launch from @OpenPipeAI: We just launched Serverless RL — train agents faster and cheaper with zero infra headaches. Compared to running your own GPUs, Serverless RL is: – 40% cheaper – 28% faster wall‑clock – instantly deployed to prod via @weights_biases Inference https://x.com/corbtt/status/1975990784404115692
🚀 TensorRT-LLM hit its v1.0 milestone — a culmination of 4 years of architecture pivots, cross-continent teamwork, and relentless optimization at NVIDIA. What started as a small team optimizing ONNX runtime has grown into a full-scale, PyTorch-native inference system powering https://x.com/ZhihuFrontier/status/1974559265273639349




