Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Photorealistic wide shot of six Ionic limestone columns on a university quad topped with a classical stone entablature, the word SCIENCE carved in Roman serif letters centered on the architrave, and a detailed bas-relief frieze below showing a microscope, telescope, DNA helix, and molecular structures carved into the limestone, late afternoon golden light casting dramatic shadows on the carvings, red brick buildings and green lawn in background, architectural photography style.

After two years of work, we’ve made an AI Scientist that runs for days and makes genuine discoveries. Working with external collaborators, we report seven externally validated discoveries across multiple fields. It is available right now for anyone to use. 1/5 https://x.com/andrewwhite01/status/1986094948048093389

Edison Scientific (a brand-new company spun out of FutureHouse) releases Kosmos: An AI Scientist for Autonomous Discovery “”Our beta users estimate that Kosmos can do in one day what would take them 6 months, and we find that 79.4% of its conclusions are accurate.”” The paper https://x.com/iScienceLuvr/status/1986023952037417109

Kosmos: An AI Scientist for Autonomous Discovery https://edisonscientific.com/articles/announcing-kosmos

“ChatGPT-o1 & DeepSeek-R1, achieved diagnostic accuracy up to 93.75%. For context, this figure approaches the 96% accuracy benchmark reported for primary care physicians on the same vignette set” Except they told folks to get urgent care too often. Not unexpected given alignment”” / X https://x.com/emollick/status/1985164511947682070

A year ago, I would not have expected the first academic field to seem to reach a consensus that AIs will accelerate research (which is not the same thing as autonomous research) would be math But that appears to be happening based on math professors in my feed and elsewhere.”” / X https://x.com/emollick/status/1984388281061282081

Here is the story of a remarkable, independent treatment suggestion by GPT-5 Pro: repurposing a known drug for a patient with food protein-induced enterocolitis syndrome (FPIES). First, how we came to test this. My close friend, physician-scientist Dr. Oral Alpan, treated the https://x.com/DeryaTR_/status/1984083644437192737

Over the past few months, OpenAI models crossed a threshold: we’re seeing early/small-scale but repeated examples of GPT-5 meaningfully contributing to novel research. AI is the next great scientific instrument, and it benefits every field. Progress accelerates when researchers”” / X https://x.com/kevinweil/status/1986115564868186288

The Royal Surrey Hospital performs 10,000th robotic surgery https://www.bbc.com/news/articles/c7v8176z7dlo

I firmly believe we are at a watershed moment in the history of mathematics. In the coming years, using LLMs for math research will become mainstream, and so will Lean formalization, made easier by LLMs. (1/4)”” / X https://x.com/ErnestRyu/status/1984033423586160889

DS-STAR is a state-of-the-art data science agent designed to autonomously solve complex data science problems. It automates tasks from analysis to data wrangling across diverse data types to achieve top performance on challenging benchmarks. Learn more: https://x.com/GoogleResearch/status/1986491681571807584

WindBorne Systems has developed autonomous weather balloons that stay aloft for over 50 days, collecting atmospheric data from regions too dangerous or remote for traditional methods. https://www.instagram.com/reel/DQcoZplDcB1/

Google DeepMind release: Towards Robust Mathematical Reasoning Introduces IMO-Bench, a suite of advanced reasoning benchmarks that played a crucial role in GDM’s IMO-gold journey. Vetted by a panel of IMO medalists and mathematicians. IMO-AnswerBench – a large-scale test on https://x.com/iScienceLuvr/status/1985685404276965481

While human expert evaluation remains the gold standard for mathematical proofs, its cost and time intensity limit scalable research. To address this, we built #ProofAutoGrader, an automatic grader for IMO-ProofBench. The autograder leverages Gemini 2.5 Pro, providing it with a https://x.com/lmthang/status/1985772094085595570

Continuing our IMO-gold journey, I’m delighted to share our #EMNLP2025 paper “Towards Robust Mathematical Reasoning”, which tells some of the key stories behind the success of our advanced Gemini #DeepThink at this year IMO. Finding the right north-star metrics was highly https://x.com/lmthang/status/1985760224612057092

Google and Mombak collaborate on CO2 removal https://blog.google/outreach-initiatives/sustainability/mombak-co2-removal/

AquaWomb: Artificial womb in glass tank could save premature babies https://interestingengineering.com/science/aquawomb-artificial-womb-premature-babies

@_ivyzhang @risi1979 @LearningLukeD What happens when we make multiple different Neural Cellular Automata compete for space? Our GitHub implementation of Petri Dish NCA: https://x.com/SakanaAILabs/status/1986041771458261477

AI coaching is going to revolutionize preventative health Whoop now tracks bloodwork and biomarkers, combining with sleep, strain, and recovery data for personalized AI plans It’s just the start, but in 10 years everyone will probably have a personal AI coach/doctor that’s https://x.com/rowancheung/status/1985377261004992879

SKATE Enhances Volcano Monitoring Safety – IEEE Spectrum https://spectrum.ieee.org/volcano-monitoring-stromboli-skate

The “AI will replace radiologists” prediction remains a rich example. A lot of folks have pointed out the problems with confusing a task (“reading a scan”) with a job (“radiologist”) with many tasks. That is true. But there was a human problem. Radiologists rejected (pre-LLM) AI”” / X https://x.com/emollick/status/1984696156140470530

Do more with less strain: UTA’s robotic arm – News Center – The University of Texas at Arlington https://www.uta.edu/news/news-releases/2025/10/24/do-more-with-less-utas-robotic-arm

New pathology foundation model from PathAI, PLUTO-4 Interesting FlexiViT arch for their smaller model, and good discussion of some of the training infra/scaling. Trained on 32 H200s with DINOv2 framework, on 551k whole slide images. Great to see EVA (from https://x.com/iScienceLuvr/status/1986031522231865571

Cross-continent teleop is quite insane when you think about it. It’s remote work for atoms, not bits. For the first time, we can apply the cost advantages of outsourced labor to the physical economy. Starlink is only going to accelerate this.”” / X https://x.com/aryxnsharma/status/1985427799541457043

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading