Image created with GPT Image 1. Image prompt: split-screen collage of burnt-orange blaze and cobalt surf, Republic split-color palette, minimalist graphic design inspired by New Order’s ‘Republic’, metaphor for general intelligence core orb, flat color, subtle texture, 1980s Saville typography style
Introduction – SITUATIONAL AWARENESS: The Decade Ahead https://situational-awareness.ai/
New HealthBench eval! Very excited we (@OpenAI) are investing in AI for health, a defining use case for AGI. Favorite plot is how the performance-cost frontier has improved over time. Congrats @rahularoradfs @thekaransinghal & team! Follow them for more exciting work to come https://x.com/_jasonwei/status/1922002699240775994
Rob Fergus is the new head of Meta-FAIR! FAIR is refocusing on Advanced Machine Intelligence: what others would call human-level AI or AGI. https://x.com/ylecun/status/1920556537233207483
Anthropic’s Upcoming Models Will Think… And Think Some More — The Information https://www.theinformation.com/articles/anthropics-upcoming-models-will-think-think
Welcome @fidjissimo! Fidji has been an amazing friend and colleague, with unique insights and advice on OpenAI. I’m super excited to work with her to deliver AGI that benefits all of humanity.”” / X https://x.com/gdb/status/1920344903466529193
is this.. AGI? 😮 meet any-to-any models on @huggingface, models that take in and output multiple modalities (e.g. a model that takes image + text input and responds with speech!) we’ve shipped a beginner friendly doc on everything you need to know, on the next one ⤵️ https://x.com/mervenoyann/status/1923053505704493311
the AI labs spent a few years quietly scaling up supervised learning, where the best-case outcome was obvious: an excellent simulator of human text now they are scaling up reinforcement learning, which is something fundamentally different. and no one knows what happens next”” / X https://x.com/jxmnop/status/1922078186864566491
How far can reasoning models scale? | Epoch AI https://epoch.ai/gradient-updates/how-far-can-reasoning-models-scale
AI models are dramatically improving at IQ tests (70 IQ → 120), yet they don’t feel vastly smarter than two years ago. At their current level of intelligence, rehashing existing human writings will work better than leaning on their own intelligence to produce novel analysis. https://x.com/DanHendrycks/status/1921429850432405827
How to simulate and evaluate multi-turn conversations 💬 Most LLM applications today are chat-based. How would you evaluate the conversations? 🔧 We’re excited to launch OpenEvals — a set of utilities to simulate full conversations and evaluate your LLM application’s https://x.com/LangChainAI/status/1922747560483226041
Introducing Continuous Thought Machines https://sakana.ai/ctm/




