Image created with Flux Pro v1.1 Ultra. Image prompt: Education, graduation cap with tassel made of tiny banana strands and brim stitched with small-banana motif, photorealistic, editorial, minimal, high detail, 3:2 landscape

We can now say pretty definitively that AI progress is well ahead of expectations from a few years ago. In 2022, the Forecasting Research Institute had super forecasters & experts to predict AI progress. They gave a 2.3% & 8.6% probability of an AI Math Olympiad gold by 2025… https://x.com/emollick/status/1962859757674344823

Gemini 2.5 Flash Image (Nano Banana) best practices 🍌🍌🍌
https://x.com/_philschmid/status/1961809165191397863

Notebook LM Rolling out NEW audio overview formats:
(Default) Deep Dive: a thorough examination of your sources
Brief: 1-2 minute, bite-sized overviews
Critique: an expert review, offering constructive feedback on your material
Debate: a thoughtful debate between two hosts https://x.com/NotebookLM/status/1962949985546187120

Really excited about this new AI research that’s pushing the boundaries of what’s currently possible in astrophysics. 🌌”” / X https://x.com/sundarpichai/status/1963668228481159371

Using AI to advance our understanding of fundamental physics is the dream. Excited to see our latest AI model ‘Deep Loop Shaping’ help @LIGO and @Caltech detect the gravitational waves of intermediate-mass black holes better! Published in @ScienceMagazine”” / X https://x.com/demishassabis/status/1963795824854335528

We’re helping to unlock the mysteries of the universe with AI. 🌌 Our novel Deep Loop Shaping method published in @ScienceMagazine could help astronomers observe more events like collisions and mergers of black holes in greater detail, and gather more data about rare space https://x.com/GoogleDeepMind/status/1963664018515849285

Get a free visual guidebook to learn MCPs from scratch (with 11 projects):
https://x.com/_avichawla/status/1961677843903185078

A 14B model just beat a 671B model on math reasoning. Here’s how Microsoft’s rStar2-Agent achieves frontier math performance in 1 week of RL training
https://x.com/FrankYouChill/status/1962180218053144655

There is significant unmet demand for developers who understand AI. At the same time, because most universities have not yet adapted their curricula to the new reality of programming jobs being much more productive with AI tools, there is also an uptick in unemployment of recent https://x.com/AndrewYNg/status/1963631698987684272

Cool research from Microsoft! They release rStar2-Agent, a 14B math reasoning models trained with agentic RL. It reaches frontier-level math reasoning in just 510 RL training steps. Here are my notes: https://x.com/omarsar0/status/1964045125115662847

rStar2-Agent: Agentic Reasoning Technical Report “”We introduce rStar2-Agent, a 14B math reasoning model trained with agentic reinforcement learning to achieve frontier-level performance.”” “”three key innovations that makes agentic RL effective at scale: (i) an efficient RL https://x.com/iScienceLuvr/status/1962798181059817480

New White House commitments empower teachers, students, and job seekers through AI skilling and learning  – Microsoft On the Issues https://blogs.microsoft.com/on-the-issues/2025/09/04/new-white-house-commitments/

We are rolling out Comet to all students worldwide. Ask Comet to manage your schedule, order textbooks, or prepare for exams with Study Mode. https://x.com/perplexity_ai/status/1963285255198314951

Interested in building and benchmarking deep research systems? Excited to introduce DeepScholar-Bench, a live benchmark for generative research synthesis, from our team at Stanford and Berkeley! 🏆Live Leaderboard https://x.com/lianapatel_/status/1961487232331911651

A really useful prompt for writing: “”review this for accuracy, look up any facts you may want to challenge or explore.”” Even if not perfect, it is a good sanity check. Works well with Claude 4.1, GPT-5 Thinking, and Grok 4. Weirdly, Gemini 2.5 Pro often won’t do web searches. https://x.com/emollick/status/1961257429846691881

✍️ When it comes to creative writing optimization, you can’t ignore Zhi-Create-Qwen3-32B, a fine-tuned variant of Qwen3-32B. On WritingBench, it scores 82.08, outperforming the base model (78.97), showing notable gains across 6 domains (Fig.1) What powers its performance boost? https://x.com/ZhihuFrontier/status/1963441300692402659

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading