“Just announced new versions of Gemma 3 – the most capable model to run just one H100 GPU – can now run on just one *desktop* GPU! Our Quantization-Aware Training (QAT) method drastically brings down memory use while maintaining high quality. Excited to make Gemma 3 even more https://x.com/sundarpichai/status/1913260423622656432

“Performance has grown drastically – FLOP/s of leading AI supercomputers have doubled every 9 months. Driven by: – Deploying more chips (1.6x/year) – Higher performance per chip (1.6x/year) Systems with 10,000 AI chips were rare in 2019. Now, leading companies have clusters 10x https://x.com/EpochAIResearch/status/1915098225465245931

“Nvidia just released Describe Anything 3B – Multimodal LLM for Detailed Localized Image and Video Captioning ⚡ > integrates full-image/ video context with fine-grained local details using a focal prompt and a localised vision backbone with gated cross-attention DAM-3B > https://x.com/reach_vb/status/1914962078571356656

“Nvidia presents Eagle 2.5! – A family of frontier VLMs for long-context multimodal learning – Eagle 2.5-8B matches the results of GPT-4o and Qwen2.5-VL-72B on long-video understanding https://x.com/arankomatsuzaki/status/1914517474370052425

“Keeps getting better: Nvidia also dropped OpenMath Nemotron 32B & 14B – secured FIRST prize in AIMO-2 competition 🤯 > beats DeepSeek R1, QwQ and more on AIME, HLE-Math and more So cool to see Nvidia not just releasing model checkpoints, but also the code and the datasets too https://x.com/reach_vb/status/1915153226145427574

“Google also released a new version of Gemma 3, optimized with ‘Quantization-Aware Training’ This reduces the 27B model’s memory requirements, enabling it to run on consumer GPUs, like the NVIDIA RTX 3090, with maintained performance https://x.com/rowancheung/status/1914201267053777278

“Wait, Nvidia dropped a 4 MILLION context length Llama 3.1 Nemotron 🤯 could literally drop entire codebases in it! https://x.com/reach_vb/status/1912743420851875986

“New foundation model on image and video captioning just dropped by @NVIDIAAI 🔥 Describe Anything Model (DAM) is a 3B vision language model to generate detailed captions with localized references 😮 The team released the models, the dataset, a new benchmark and a demo 🤩 https://x.com/mervenoyann/status/1914980803055862176

“How quickly are AI supercomputers scaling, where are they, and who owns them? Our new dataset covers 500+ of the largest AI supercomputers (aka GPU clusters or AI data centers) over the last six years. Here is what we found🧵 https://x.com/EpochAIResearch/status/1915098223082873015

Gemma 3 QAT Models: Bringing state-of-the-Art AI to consumer GPUs – Google Developers Blog https://developers.googleblog.com/en/gemma-3-quantized-aware-trained-state-of-the-art-ai-to-consumer-gpus/

“I view SB 813 as a step forward—particularly in fostering innovation in AI safety through the creation of independent multi-stakeholder regulatory organizations (MROs). It would also set in motion crucial work to establish the legal infrastructure and adaptive standards needed to” / X https://x.com/Yoshua_Bengio/status/1914828175772700716

“A dual-FPGA LoopLynx achieves 2.52x speedup over an Nvidia A100 with 48.1% energy. Field-Programmable Gate Array (FPGA) architectures for LLM inference are often inefficient during sequential token decoding. LoopLynx presents a hybrid spatial-temporal architecture using large https://x.com/rohanpaul_ai/status/1915727488946209224

[2504.03624] Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models https://arxiv.org/abs/2504.03624

China’s Baidu says its Kunlun chip cluster can train DeepSeek-like models | Reuters https://www.reuters.com/world/china/chinas-baidu-says-its-kunlun-chip-cluster-can-train-deepseek-like-models-2025-04-25/

“Llama 4 Maverick illustrates a key challenge in calculating AI training compute: How to account for when a smaller model is trained using a larger model’s outputs? We’re updating our methodology and removing Maverick from our list of models exceeding 1e25 FLOP. Here’s why… 🧵 https://x.com/EpochAIResearch/status/1913329195171688742

“Sleep-time Compute: Beyond Inference Scaling at Test-time Overview: Sleep-time compute allows LLMs to perform computations offline by anticipating user queries, aiming to reduce the high latency and cost associated with scaling test-time inference. On modified reasoning tasks https://x.com/TheAITimeline/status/1914175946715779426

“Nvidia just dropped Describe Anything on Hugging Face Detailed Localized Image and Video Captioning https://x.com/_akhaliq/status/1914917564137828622

“Ownership has rapidly shifted from the public to the private sector. In 2019, the two sectors had roughly equal compute shares. We estimate that companies now control over 80% of global AI computing capacity, while the share of governments and academia has fallen below 20%. https://x.com/EpochAIResearch/status/1915098234537603329

“TSMC responds to the allegation that it knowingly supplied Huawei with 910B dies: if it ever happened, it happened legally, in… 2020. Lol. Ok? (p.s. it’s funny how everything is about DeepSeek. I keep telling you bros, more where that came from) https://x.com/teortaxesTex/status/1915295046930223312

Exclusive: Huawei readies new AI chip for mass shipment as China seeks Nvidia alternatives, sources say | Reuters https://www.reuters.com/world/china/huawei-readies-new-ai-chip-mass-shipment-china-seeks-nvidia-alternatives-sources-2025-04-21/

“seems kinda clear we are definitely getting the ‘software singularity’ far, far before the ‘hardware singularity’, which seems.. more delayed than ever and paul christiano-esque takes seem to have panned out very well so, kinda full cyberpunk timeline? almost too predictable?” / X https://x.com/nearcyan/status/1912692168764109098

[2504.13171] Sleep-time Compute: Beyond Inference Scaling at Test-time https://arxiv.org/abs/2504.13171

“Geographically, the US dominates with 75% of global AI supercomputer performance in our dataset. China is in second place with 15%, while Europe plays only a small role. (Our data only covers ~15% of global AI compute, so there’s some uncertainty.) https://x.com/EpochAIResearch/status/1915098237083541920

“What could scaling unlock for biology? Introducing ProGen3- our next AI foundation models for protein generation. We develop compute-optimal scaling laws up to 46B parameters on 1.5T tokens with real evidence in the wet lab. +we solve a new set of challenges for drug discovery https://x.com/thisismadani/status/1912502918328315937

Enterprises Onboard AI Teammates Faster With NVIDIA NeMo Tools to Scale Employee Productivity | NVIDIA Blog https://blogs.nvidia.com/blog/nemo-enterprises-ai-teammates-employee-productivity/

“Vector search with large databases slows down RAG pipelines. Moving the entire search to the GPU hurts LLM performance by using up critical memory. VectorLiteRAG solves this by adaptively partitioning the vector index, placing only frequently accessed “hot” clusters onto the https://x.com/rohanpaul_ai/status/1913951536926834740

[2504.13171v1] Sleep-time Compute: Beyond Inference Scaling at Test-time https://arxiv.org/abs/2504.13171v1

South Korea’s AI Chip Champion Is Poised To Carve Out Global Niche https://www.forbes.com/sites/johnkang/2025/04/14/south-koreas-ai-chip-champion-is-poised-to-carve-out-global-niche/

“Why did AI supercomputers grow so quickly? They transitioned from research tools to industrial machines delivering economic value. Bigger AI Supercomputers -> more capable models -> more investment -> bigger AI supercomputers” / X https://x.com/EpochAIResearch/status/1915098240321470943

After years of failed AI deals, Intel plans homegrown challenge to Nvidia | Reuters https://www.reuters.com/business/after-years-failed-ai-deals-intel-plans-homegrown-challenge-nvidia-2025-04-25/

Databricks to boost hiring, invest $250 million in India for AI expansion | Reuters https://www.reuters.com/world/india/databricks-boost-hiring-invest-250-million-india-ai-expansion-2025-04-24/

“Some day in the next decade, we will have robots in every home, every hospital and factory, doing every dull and dangerous jobs with superhuman dexterity. That day will be known as “Thursday”. Not even Turing would dare to dream up our lifetime in his wildest dreams.” / X https://x.com/DrJimFan/status/1914681494385360959

“First Luma AI, now Odyssey ML. All of the 3D-first AI startups are pivoting to video diffusion as the path to interactive world building. Maybe Jensen was right. Every pixel will be generated, not rendered. Makes me wonder if World Labs stays the course.” / X https://x.com/bilawalsidhu/status/1913046871112573146

“Today we’re announcing Mechanize, a startup focused on developing virtual work environments, benchmarks, and training data that will enable the full automation of the economy. We will achieve this by creating simulated environments and evaluations that capture the full scope of” / X https://x.com/MechanizeWork/status/1912904151874625928

“I’m starting a new company: Mechanize. Mechanize will build virtual work environments, benchmarks, and training data to enable the full automation of all work. We’re hiring: hiring@mechanize.work.” / X https://x.com/tamaybes/status/1912905467376124240

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading