Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Photorealistic Times Square at golden hour dusk transformed into a semiconductor showcase, every billboard displaying massive illuminated silicon wafer cross-sections and GPU die photography in electric blue and amber, the streets reflecting neon circuit traces, pedestrians dwarfed by towering 3D microchip architecture projections between buildings, cinematic lighting with warm sunset mixing with cool LED glow, ultra-detailed 8k photography style.
Microsoft inks $33 billion in deals with ‘neoclouds’ like Nebius, CoreWeave — Nebius deal alone secures 100,000 Nvidia GB300 chips for internal use | Tom’s Hardware https://www.tomshardware.com/tech-industry/artificial-intelligence/microsoft-inks-usd33-billion-in-deals-with-neoclouds-like-nebius-coreweave-nebius-deal-alone-secures-100-000-nvidia-gb300-chips-for-internal-use
Cerebras Raises $1.1 Billion at $8.1 Billion Valuation https://www.cerebras.ai/press-release/series-g
OpenAI’s Stargate project to consume up to 40% of global DRAM output — inks deal with Samsung and SK hynix to the tune of up to 900,000 wafers per month | Tom’s Hardware https://www.tomshardware.com/pc-components/dram/openais-stargate-project-to-consume-up-to-40-percent-of-global-dram-output-inks-deal-with-samsung-and-sk-hynix-to-the-tune-of-up-to-900-000-wafers-per-month
CoreWeave stock climbs 12% after $14 billion deal with Meta https://www.cnbc.com/2025/09/30/coreweave-meta-deal-ai.html
As Jensen mentioned with @altcap @BG2Pod @bgurley, something that few people know is that @nvidia is becoming the American open-source leader in AI, with over 300 contributions of models, datasets and apps on @huggingface in the past year. And I have a feeling they’re just https://x.com/ClementDelangue/status/1971698860146999502
Bring me the healthiest snack.”” The robot goes to the kitchen and gets the snack, fully autonomously. This is the first public demo of NVIDIA’s Isaac GR00T N1.6 foundation model presented by Yuke Zhu at CoRL 2025. The previous versions focused only on bimanual stationary https://x.com/TheHumanoidHub/status/1972698708975440349
Go check out @yukez’s talk at CoRL! Project GR00T is cooking 🍳”” / X https://x.com/DrJimFan/status/1971370444474417658
Why did OpenAI train GPT-5 with less compute than GPT-4.5? Due to the higher returns to post-training, they scaled post-training as much as possible on a smaller model And since post-training started from a much lower base, this meant a decrease in total training FLOP 🧵 https://x.com/EpochAIResearch/status/1971675079219282422
2025 State of AI Infrastructure Report | Google Cloud https://cloud.google.com/resources/content/state-of-ai-infrastructure?e=48754805
It’s true – @modal has raised a $87M Series B at a $1.1B valuation to advance the future of AI infrastructure. Thank you to @Lux_Capital, @Redpoint, @AmplifyPartners, and others. Now more than ever, AI demands a complete reinvention of traditional compute infrastructure https://x.com/bernhardsson/status/1972649681701486821
Our resident GPU enjoyers have pulled off an incredible feat here. First look into how FA4 gets a 20% speed up: marshaling bits across specialized warps, using cubics to approximate exponentials, and more.”” / X https://x.com/akshat_b/status/1971617146930450758
there is no infrastructure product I can recommend as strongly as @modal one of the smartest teams we work with, and there are basically no tradeoffs – we own the code and models, they manage the scaling. it’s literally a superpower for our ml/product team.”” / X https://x.com/raunakdoesdev/status/1972755957047587178
GPU cloud costs shouldn’t surprise you. With Together Instant Clusters, you get pricing you can easily understand and predict. Hourly rates for on‑demand; discounts for committed terms. https://x.com/togethercompute/status/1974167802337730854
I quite enjoyed this and it covers a bunch of topics without good introductory resources! 1. A bunch of GPU hardware details in one place (warp schedulers, shared memory, etc.) 2. A breakdown/walkthrough of reading PTX and SASS. 3. Some details/walkthroughs of a number of other”” / X https://x.com/cHHillee/status/1972782935859401155
Good morning! 🌅 We’re excited to announce the immediate availability of no-contract, on-demand, 2x and 4x @AMD MI300x VMs, making us the first AMD-exclusive cloud to offer this. Now, developers and businesses can build and test their applications on best-in-class AMD compute, https://x.com/HotAisle/status/1973768786965639643
GPUs are expensive and setting up the infrastructure to make GPUs work for you properly is complex, making experimentation on cutting-edge models challenging for researchers and ML practitioners. Providing high quality research tooling is one of the most effective ways to https://x.com/lilianweng/status/1973455232341516731
Meta Is Said to Acquire Chips Startup Rivos to Push AI Effort – Bloomberg https://www.bloomberg.com/news/articles/2025-09-30/meta-is-said-to-acquire-chips-startup-rivos-to-push-ai-effort
IBM is back! They just joined Hugging Face Enterprise & released Granite 4.0 in open-source with a new hybrid Mamba/transformer architecture that reduces memory requirements without reducing accuracy much. This set of models is great for agentic workflows like tool calling, https://x.com/ClementDelangue/status/1973798540389355903
AMD is using Cline as their coding agent for local models. After testing 20+ models, they found what actually works: > 32GB RAM: Qwen3-Coder 30B (4-bit) > 64GB RAM: Qwen3-Coder 30B (8-bit) > 128GB+ RAM: GLM-4.5-Air 10-minute setup with @lmstudio + Cline, linked below”” / X https://x.com/cline/status/1973035211379310708
new latent cot architecture just dropped 👀 CoT is fundamentally about freeing the model to use extra compute when it wants to this method lets the model learn to use extra compute by inserting latent tokens, without needing to bootstrap CoTs from SFT”” / X https://x.com/khoomeik/status/1973785079932727760
Thank you @NVIDIADC for supporting vLLM for @deepseek_ai model launch. Blackwell is now the go to release platform for new MoEs, and we cannot do it without the amazing team from NVIDIA.”” / X https://x.com/vllm_project/status/1973238773090754735
Microsoft aims to swap AMD, Nvidia GPUs for its own AI chips • The Register https://www.theregister.com/2025/10/02/microsoft_maia_dc/
New in-depth blog post time: “”Inside NVIDIA GPUs: Anatomy of high performance matmul kernels””. If you want to deeply understand how one writes state of the art matmul kernels in CUDA read along. (Remember matmul is the single most important operation that transformers execute https://x.com/gordic_aleksa/status/1972696444101632151
Alibaba bets big on AI with Nvidia tie-up, new data center plans https://interestingengineering.com/culture/alibaba-nvidia-ai-partnership-expansion
the points show RF-DETR evaluated at input resolutions of 312, 384, and 432 using the same trained weights. latency is end-to-end, measured with TensorRT 10.4 on a T4 GPU with CUDA 12.4. blog: https://x.com/skalskip92/status/1974160481192747039
The latest from our GPU engineers is a 4,000 word deep dive on Flash Attention 4. No performance-enhancing drugs were involved, just some insanely deep CUDA spelunking.”” / X https://x.com/bernhardsson/status/1971603562355716160
Pre-training under infinite compute https://arxiv.org/pdf/2509.14786
Modal is a really impressive product – for example its insanely fast “remote but feels local” On top of that it’s a cloud purpose built for ML from the ground up. Highly recommend watching this talk from Erik on the infra they built https://x.com/HamelHusain/status/1972660215658189049
This is one of my favorite results from our work on scaling laws for QAT. It helps answer the question: should you train an 8-bit model or a 4 bit model that has twice as many parameters? Both are the same size in RAM which is highly correlated with generation latency. Or even”” / X https://x.com/awnihannun/status/1974245339512385784
.@nvidia presented Reinforcement Learning with Binary Flexible Feedback (RLBFF). It combines adaptable human feedback (RLHF) with rule-based checks (RLVR), turning human feedback into binary principles – yes/no. ! Users can also plug in their own principles. This method https://x.com/TheTuringPost/status/1972631189136806257
Beautiful @nvidia paper. 👏 💾 NVFP4 shows 4-bit pretraining of a 12B Mamba Transformer on 10T tokens can match FP8 accuracy while cutting compute and memory. 🔥 NVFP4 is a way to store numbers for training large models using just 4 bits instead of 8 or 16. This makes training https://x.com/rohanpaul_ai/status/1973017414011932791
Nvidia presents RLP Reinforcement as a Pretraining Objective https://x.com/_akhaliq/status/1974190336256962812
Inside NVIDIA GPUs: Anatomy of high performance matmul kernels – Aleksa Gordić https://www.aleksagordic.com/blog/matmul




