Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Create a 16:9 cinematic split-screen poster. LEFT SIDE (40% width): – A teacher’s desk with stacked textbooks, an open notebook, a tablet showing a calm lesson plan, and a cup of pens, suggesting AI quietly assisting in curriculum planning. – The background is a turquoise / teal abstract field made of stylized blue rods or data fibers, hinting at supportive learning tools behind the scenes. – Use warm classroom lighting. No glowing screens or neon. RIGHT SIDE (60% width): – A green-toned abstract aerial forest canopy texture, representing growth and curiosity. – Two clean rounded rectangles stacked vertically near the center-right. – The TOP rectangle contains the text: “Education”. – The BOTTOM rectangle contains the text: “2025/10/10”. – Clean sans-serif font, dark green or charcoal text. OVERALL STYLE: – Gentle, inclusive, and optimistic. – No logos or extra slogans. – Preserve the turquoise/forest split-screen.

My co-author Lennart Meincke had GPT-5 Pro look over a paper before we submitted it to a journal. It caught a tiny error in the citations that we missed (apparently it estimated the volume) A big difference from constant hallucinations, especially GPT5 Pro; though not error-free https://x.com/emollick/status/1973910542072102962

GPT-5 Pro for catching subtle errors in academic work:”” / X https://x.com/gdb/status/1974018754657837083

POV: Your LLM agent is dividing a by b https://x.com/karpathy/status/1976082963382272334

Evals in the wild”” / X https://x.com/lateinteraction/status/1976439833158615345

Super happy to announce that University of Zurich (@UZH_en) just joined @huggingface Academia Hub 🇨🇭🎉 Their students and educators get a better access to collaboration and compute features (including ZeroGPU power) on the Hub. 🔥 https://x.com/julien_c/status/1975515541700841935

A question relevant to a lot of academics, if you think AI will be able to meaningfully contribute to scientific research in the near future.”” / X https://x.com/emollick/status/1973912881096995276

For more on FrontierMath, and more analysis of AI math capabilities, check out our website! https://x.com/EpochAIResearch/status/1976685780862144978

Mathematical discovery in the age of artificial intelligence https://uva.theopenscholar.com/files/ken-ono/files/documents/naturephysics.pdf

Gemini 2.5 Deep Think is SoTA on FrontierMath! 🔥 Thank you for testing it @EpochAIResearch. https://x.com/_philschmid/status/1976626257090535432

Gemini 2.5 Deep Think is SoTA on FrontierMath! 🔥🫡”” / X https://x.com/YiTayML/status/1976470535308734575

Just getting answers is easy. Understanding them? That’s harder. That’s why I’m excited Study and learn mode + Quizzes are live now in Copilot, giving every student a tutor in their pocket 🧵 https://x.com/mustafasuleyman/status/1973791482369937533

OpenAI, which funded FrontierMath, has access to 28/48 problems and solutions. Epoch holds out the remaining 20 problems and solutions. Of the eight problems solved at least once by GPT-5 Pro, five are in the held-out set.”” / X https://x.com/EpochAIResearch/status/1976685757369851990

Getting started with datasets – OpenAI API https://platform.openai.com/docs/guides/evaluation-getting-started?api-mode=responses

This is an interesting debate about AI text between an OpenAI researcher who thinks about AI writing and one of the great short story masters Now that we have machines that can write stories, occasionally very good or moving stories, we need to think more about what that means”” / X https://x.com/emollick/status/1975371420252442715

Previously, eight Tier 4 problems had been solved at least once. These eight were all solved across these high-compute runs as well. Adding the one new problem solved by GPT-5 Pro brings the total number ever solved to nine, or 19% of the benchmark.”” / X https://x.com/EpochAIResearch/status/1976685769130705300

I am hearing similar things in economics & the social sciences. Not autonomous work, but expert-directed AI is absolutely helping academics do novel research in significant ways. Especially Pro/High Thinking models.”” / X https://x.com/emollick/status/1975620277490032704

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading