Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Create a 16:9 cinematic split-screen poster. LEFT SIDE (40% width): – A teacher’s desk with stacked textbooks, an open notebook, a tablet showing a calm lesson plan, and a cup of pens, suggesting AI quietly assisting in curriculum planning. – The background is a turquoise / teal abstract field made of stylized blue rods or data fibers, hinting at supportive learning tools behind the scenes. – Use warm classroom lighting. No glowing screens or neon. RIGHT SIDE (60% width): – A green-toned abstract aerial forest canopy texture, representing growth and curiosity. – Two clean rounded rectangles stacked vertically near the center-right. – The TOP rectangle contains the text: “Education”. – The BOTTOM rectangle contains the text: “2025/10/10”. – Clean sans-serif font, dark green or charcoal text. OVERALL STYLE: – Gentle, inclusive, and optimistic. – No logos or extra slogans. – Preserve the turquoise/forest split-screen.
My co-author Lennart Meincke had GPT-5 Pro look over a paper before we submitted it to a journal. It caught a tiny error in the citations that we missed (apparently it estimated the volume) A big difference from constant hallucinations, especially GPT5 Pro; though not error-free https://x.com/emollick/status/1973910542072102962
GPT-5 Pro for catching subtle errors in academic work:”” / X https://x.com/gdb/status/1974018754657837083
POV: Your LLM agent is dividing a by b https://x.com/karpathy/status/1976082963382272334
Evals in the wild”” / X https://x.com/lateinteraction/status/1976439833158615345
Super happy to announce that University of Zurich (@UZH_en) just joined @huggingface Academia Hub 🇨🇭🎉 Their students and educators get a better access to collaboration and compute features (including ZeroGPU power) on the Hub. 🔥 https://x.com/julien_c/status/1975515541700841935
A question relevant to a lot of academics, if you think AI will be able to meaningfully contribute to scientific research in the near future.”” / X https://x.com/emollick/status/1973912881096995276
For more on FrontierMath, and more analysis of AI math capabilities, check out our website! https://x.com/EpochAIResearch/status/1976685780862144978
Mathematical discovery in the age of artificial intelligence https://uva.theopenscholar.com/files/ken-ono/files/documents/naturephysics.pdf
Gemini 2.5 Deep Think is SoTA on FrontierMath! 🔥 Thank you for testing it @EpochAIResearch. https://x.com/_philschmid/status/1976626257090535432
Gemini 2.5 Deep Think is SoTA on FrontierMath! 🔥🫡”” / X https://x.com/YiTayML/status/1976470535308734575
Just getting answers is easy. Understanding them? That’s harder. That’s why I’m excited Study and learn mode + Quizzes are live now in Copilot, giving every student a tutor in their pocket 🧵 https://x.com/mustafasuleyman/status/1973791482369937533
OpenAI, which funded FrontierMath, has access to 28/48 problems and solutions. Epoch holds out the remaining 20 problems and solutions. Of the eight problems solved at least once by GPT-5 Pro, five are in the held-out set.”” / X https://x.com/EpochAIResearch/status/1976685757369851990
Getting started with datasets – OpenAI API https://platform.openai.com/docs/guides/evaluation-getting-started?api-mode=responses
This is an interesting debate about AI text between an OpenAI researcher who thinks about AI writing and one of the great short story masters Now that we have machines that can write stories, occasionally very good or moving stories, we need to think more about what that means”” / X https://x.com/emollick/status/1975371420252442715
Previously, eight Tier 4 problems had been solved at least once. These eight were all solved across these high-compute runs as well. Adding the one new problem solved by GPT-5 Pro brings the total number ever solved to nine, or 19% of the benchmark.”” / X https://x.com/EpochAIResearch/status/1976685769130705300
I am hearing similar things in economics & the social sciences. Not autonomous work, but expert-directed AI is absolutely helping academics do novel research in significant ways. Especially Pro/High Thinking models.”” / X https://x.com/emollick/status/1975620277490032704




