“Congratulations to @togethercompute , @FireworksAI_HQ , @hyperbolic_labs and @GroqInc for being the first to launch serverless endpoints! Live performance benchmarks available now on Artificial Analysis. https://x.com/ArtificialAnlys/status/1897701018231881772
“Improvements in DiLoCo / era of RL means our focus will not shift from VRAM /Flops per gpu per dollar, but ultimately flops / watt. Then chipmakers will focus on making GPUs more efficient per machine, and we might have a different winner. With 1000 caveats, but something to” / X https://x.com/cloneofsimo/status/1897557416117686772
“356 mm² for 54B transistors on N4. 826 mm² for 54B on N7 (in GA100). N4 basically allows 2.32x denser chips, and this is not that far off from 49/16 implied by straightforward interpretation of “X nm process”. White pill for Moore enthusiasts, a surprise for me.” / X https://x.com/teortaxesTex/status/1895494637231763464
“At Together AI, we are thrilled to be building a world-class kernels team. If you’d like to come build with us, please reach out!” / X https://x.com/togethercompute/status/1897703705790542137
Trump to make investment announcement as he meets with TSMC CEO | AP News https://apnews.com/article/trump-tsmc-chip-manufacturing-tariffs-42980704ffca62e823182422ee4b7b83
“🔥 The Weaviate team is on fire! 🤯 The month is almost over, but look at this laundry list of things shipped! 🚢 𝗪𝗲𝗮𝘃𝗶𝗮𝘁𝗲 𝗘𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴𝘀 𝗶𝘀 𝗻𝗼𝘄 𝗚𝗔 𝗶𝗻 𝗪𝗲𝗮𝘃𝗶𝗮𝘁𝗲 𝗖𝗹𝗼𝘂𝗱, featuring @SnowflakeDB’s Arctic Embed 2.0 🚢 Dedicated Enterprise Deployment https://x.com/bobvanluijt/status/1895463589915353467
“🚨 @POTUS announces a $100 BILLION new investment by TSMC in U.S. chips manufacturing! https://x.com/WhiteHouse/status/1896665779095093525
“actually very surprised with this topping every single category given 4.5 doesn’t have use any test-time compute. pretraining is still alive my friends!” / X https://x.com/willdepue/status/1896603985861378440
“FlexiDiT Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute https://x.com/_akhaliq/status/1895340908134178897
“@jeremyphoward what I am curious about: is GPT4.5 more expensive and slower than o1 (assuming o1 is GPT4-sized + inference-compute scaling)? Would be interesting in the context of what GPT4.5 with o1-style inference-compute scaling will look like.” / X https://x.com/rasbt/status/1895496476056559811
“The CUDA MOAT is so great that even AMD engineers have to write their own tensor engines primarily in CUDA https://x.com/casper_hansen_/status/1895393985847517313
“TL;DR: Chains-of-thought make the amount of compute during inference the key bottleneck. We show that (1) you can distill a large model into a faster SSM or hybrid one for that task, and (2) the resulting setup reach a better trade-off than using a full-fledge model.” / X https://x.com/francoisfleuret/status/1896584551117594815
“I first read this book about Seymour Cray and the supercomputer space a couple decades ago. Rereading it now, with some big-company experience and after the extended-and-disappointing Rage development at Id makes me feel even more kinship with Seymour. The ability to start each https://x.com/ID_AA_Carmack/status/1897671486229414017
“Jensen on the next wave of AI in yesterday’s NVIDIA earnings call: https://x.com/TheHumanoidHub/status/1895199776079229347
“Senior Researcher at NVIDIA, Yuke Zhu, discusses the data flywheel for general-purpose humanoids and how the data pyramid will flip upside down as the flywheel spins. https://x.com/TheHumanoidHub/status/1897359739232903386
“The countdown to @nvidia GTC is on, and we’re excited to connect with AI builders, researchers, and enterprises shaping the future of AI infrastructure. Join us for the AI Pioneers Happy Hour, hosted by Together AI, @SemiAnalysis_ , and @HypertecGroup. 🧵👇 https://x.com/togethercompute/status/1898094750047031651
“FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute “Our simple and sample-efficient framework enables pre-trained DiT models to be converted into flexible ones — dubbed FlexiDiT— allowing them to process inputs at varying compute https://x.com/iScienceLuvr/status/1895334570704470281
“NEW: AMD introduces Instella, a series of fully open-source, state-of-the-art 3B parameter language models. It’s nice to see AMD pushing out its own small language models. It looks like a very competitive model for its size. Here’s everything you need to know: • https://x.com/omarsar0/status/1897642582966165523
“A beautiful idea I try to share often is that we do have a way to make many problem-specific systems “scale with data and compute”. It’s called compilers: (1) Write high-level specs or code at the level of abstraction of the problem domain, not the machine or a narrow paradigm.” / X https://x.com/lateinteraction/status/1897699917801701512
“GPT-4.5 release shows LLM pre-training scaling has plateaued: A much bigger model, trained with 10X more compute, can only give us so much (or so little?) improvement. That’s why xAI can catch up within 2 years. DeepSeek drastically reduced GPU requirements by optimizing” / X https://x.com/Yuchenj_UW/status/1895531031475920978
“Ready for an inside look at how the @nvidia GB200 NVL72 performs under real-world conditions? CoreWeave is proud to be one of the first cloud providers to present a demonstration of this state-of-the-art system. https://x.com/CoreWeave/status/1861607131046162868
“Gotta love MoEs on Apple silicon with MLX. Kimi’s new 16B (3B active) Moonshot model runs very nicely on an M4 Max. As good or better than some of the best dense 7Bs and 1.5x faster inference (154 toks/sec!): https://x.com/awnihannun/status/1895546249698558363
CoreWeave to Acquire Weights & Biases – Industry Leading AI Developer Platform for Building and Deploying AI Applications https://www.prnewswire.com/news-releases/coreweave-to-acquire-weights–biases—industry-leading-ai-developer-platform-for-building-and-deploying-ai-applications-302392342.html




