Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Seamless repeating wrapping paper pattern of ornate Victorian keyholes in cascading diagonal arrangement, elaborate filigree and Art Nouveau decorative flourishes, OpenAI integrated as damask monogram between elements, deep navy blue background with antique gold and cream details, subtle embossed texture, elegant gift wrap quality, hand-drawn refinement meets textile design precision.

Bloody hell, I’ll say this GPT 5.2 Codex Extra High is a methodical beast It’s updating the OpenCode OpenAI Codex OAuth plugin Literally not leaving any stone unturned This is the first model that feels like it’s building for itself ie leaving the door open for future work”” / X https://x.com/nummanali/status/2002116277666803917

A lot of people underestimate AI due to the confluence of 4 OpenAI choices: 1) GPT-5.x instant is not a very smart model 2) Most users are free users & the ChatGPT router sends them to instant often 3) The router calls everything GPT-5.2 4) Most people don’t know Reasoners exist https://x.com/emollick/status/2001840267155153362

Today @OpenAI updated the Model Spec, laying out how models are ‘intended to behave.’ Not marketing. Just explicit rules, priorities, and tradeoffs. Great reading if you’re wondering why models respond the way they do. Changelog + teen protections in 🧵👇 https://x.com/shaunralston/status/2001744269128954350

We finally had a moment to run our system with GPT-5.2 X-High on ARC-AGI-2! Using the same Poetiq harness as before, we saw results as high as 75% at under $8 / problem using GPT-5.2 X-High on the full PUBLIC-EVAL dataset. This beats the previous SOTA by ~15 percentage points. https://x.com/poetiq_ai/status/2003546910427361402

I would judge this a win by Gemini and a close second from Claude. ChatGPT-5.2 missed the reference (though, to be fair, it did write a surprising amount of successful code to actually enhance the image) and Grok wasn’t in the ballpark. https://x.com/emollick/status/2002961280534303206

So, Claude 4.5 came in far above trend in the much-watched METR measure of the task duration that AI can accomplish autonomously at 4 hours 49 minutes. Interestingly, at the harder 80% success threshold, it is GPT-5.1 Codex Max that breaks the trend. In 2023, GPT-4 was a minute. https://x.com/emollick/status/2002208335991337467

You can now adjust specific characteristics in ChatGPT, like warmth, enthusiasm, and emoji use. Now available in your “”Personalization”” settings. https://x.com/OpenAI/status/2002099459883479311

SoftBank races to fulfill $22.5 billion funding commitment to OpenAI by year-end https://finance.yahoo.com/news/exclusive-softbank-races-fulfill-22-233202534.html

Softbank fulfills $40 billion OpenAI backing, sources tell CNBC https://www.cnbc.com/2025/12/30/softbank-openai-investment.html?taid=6953e10b1534bb0001d46587

Documents: OpenAI has sold 700K+ ChatGPT licenses to ~35 US public universities for students and faculty, who used it 14M+ times in September, beating Copilot (Bloomberg) https://x.com/Techmeme/status/2001633781388648559

We are hiring a Head of Preparedness. This is a critical role at an important time; models are improving quickly and are now capable of many great things, but they are also starting to present some real challenges. The potential impact of models on mental health was something we”” / X https://x.com/sama/status/2004939524216910323?s=20

OpenAI and the U.S. Department of Energy are expanding their collaboration on AI and advanced computing in support of national scientific priorities. The agreement builds on our work with DOE’s national labs and advances the Genesis Mission to accelerate scientific discovery.”” / X https://x.com/OpenAINewsroom/status/2001731892253724689

In fact the average ChatGPT query takes almost exactly as much energy as a Google search in 2008 (that is the last time Google clearly indicated the electrical consumption of a search). https://x.com/emollick/status/2003749085468311853

Evaluating chain-of-thought monitorability | OpenAI https://openai.com/index/evaluating-chain-of-thought-monitorability/

OpenAI built the Sora Android app (which hit #1 app in the world) in just 18 days with the help of Codex https://x.com/lennysan/status/2001074732293300301

A Redditor fed his MRI into ChatGPT and it appears to have correctly identified the cause of his sciatic leg pain. This could be a watershed moment for AI. https://x.com/reddit_lies/status/2003512194672025826

GPT 5.2 has felt like a more dramatic step-change to me than even going from 3.5 to 4 did.”” / X https://x.com/Javi/status/2001837508951445794

🆕 Writing blocks make it easier to craft the perfect email in ChatGPT. ∙Update & format text right in chat ∙Highlight to ask for changes, and accept or reject suggestions ∙Open in your email client once you’re ready to send Try it & please let us know what you think! https://x.com/jamesfzhang/status/2002104182397153358

To preserve chain-of-thought (CoT) monitorability, we must be able to measure it. We built a framework + evaluation suite to measure CoT monitorability — 13 evaluations across 24 environments — so that we can actually tell when models verbalize targeted aspects of their”” / X https://x.com/OpenAI/status/2001791131353542788

Your Year with ChatGPT! Now rolling out to everyone in the US, UK, Canada, New Zealand, and Australia who have reference saved memory and reference chat history turned on. Just make sure your app is updated. https://x.com/OpenAI/status/2003190103729144224

GPT-5.2-Codex launches today. It is trained specifically for agentic coding and terminal use, and people at OpenAI have been having great success with it.”” / X https://x.com/sama/status/2001724019188408352

Just launched GPT-5.2-Codex! The best model for long-horizon agentic coding, including strong performance on refactors and migrations. Codex becoming very magical. https://x.com/gdb/status/2001758275998785743

“”Having a *feel the AGI* moment with @OpenAI ‘s GPT 5.2 on extra high reasoning in Codex… It actually feels like a great junior engineer. Honestly, mind blown. 🤯”” / X https://x.com/AjaySohmshetty/status/2003223257655443840

🚨BREAKING: You can now build real apps inside ChatGPT. No setup. No switching tabs. Just describe what you want — and watch it come to life. Meet Replit in ChatGPT / @Replit💈 https://x.com/details_with_ai/status/2003393465208754334

Codex has (finally) a new /experimental setting that enables background terminals. (useful for long running processes) Especially when you run a dev server or logs, you won’t be blocked and you can resume working in Codex. https://x.com/kevinkern/status/2003118604808786086

gpt-5.2 codex has been rock-solid, even on big, messy codebases it can run forever without going off track, and i rarely have to throw away what it produces the only downside is that it takes long enough that i start doing other stuff while it runs and burn through my quota”” / X https://x.com/slow_developer/status/2002250108348379605

🆕 Codex now officially supports skills Skills are reusable bundles of instructions, scripts, and resources that help Codex complete specific tasks. You can call a skill directly with $.skill-name, or let Codex choose the right one based on your prompt. https://x.com/OpenAIDevs/status/2002099762536010235

Prompting GPT 5.2 Codex for Continuity It excels at long running tasks but without explicit guidance can lose track of outcomes Put this at the top of your AGENTS .md file, it will let Codex work on even larger scale tasks It’s how I let it run for 3 hours coherently https://x.com/nummanali/status/2002724188436738459

.@OpenAI introduced a rigorous framework for evaluating “chain-of-thought monitorability” It’s a fancy way of asking: Can we understand what our AIs are thinking before they act? The answer: yes, but not without nuance. – Longer reasoning helps – Bigger models muddle things – https://x.com/TheTuringPost/status/2003636642767384639

I think there is likely too much emphasis on the METR long-task measurement as a sign of AI progress… … but it doesn’t matter. With a little help from GPT-5.2 Pro, I calculated the correlations between log(METR) & other key benchmarks, and they basically all correlate highly https://x.com/emollick/status/2002861706658398211

I’ll work to make ChatGPT a better tool for accelerating scientific and mathematical discoveries. If you come across failure cases to improve upon (or exciting success stories) please send them my way!”” / X https://x.com/ErnestRyu/status/2003542931568025676

Capability overhang means too many gaps today between what the models can do and what most people actually do with them. 2026 Prediction: Progress towards AGI will depend as much on helping people use AI well, in ways that directly benefit them as on progress in frontier models https://x.com/OpenAI/status/2003594025098785145

I had some questions about whether these modifications impacted accuracy of outputs, but was told by the OpenAI team that worked on this that tone does not impact that. (Also I like that we are moving away from discussing prompts to modify AI personality to prompts changing tone)”” / X https://x.com/emollick/status/2002452909657895115

The ChatGPT apps are just hard to figure out in a way that feels like the previous GPT Store, some work exactly as you might hope (the Canva integration) and some feels remarkably non-magical (the Apple Music integration can’t access my playlists despite linking my Apple account) https://x.com/emollick/status/2002580968071213455

Context Arena Update: Added @ByteDance’s Seed 1.6 and Seed 1.6 Flash to the MRCR leaderboards. Seed 1.6 closely mimics the retrieval curve of @OpenAI ‘s reasoning models (o3 / o4-mini). It offers high fidelity at start, but follows a similar degradation slope as complexity https://x.com/DillonUzar/status/2005671520488640587

Message from Welcome to Sonar Chat! https://www.sonarsource.com/blog/new-data-on-code-quality-gpt-5-2-high-opus-4-5-gemini-3-and-more/

The browser as a body for AGI? “”AGI is something that can take action for you. The browser itself is an environment where that can happen.”” @bengoodger, head of engineering for ChatGPT Atlas at @OpenAI (and former Firefox and Chrome builder), in our interview https://x.com/TheTuringPost/status/2002891352103907465

Last week, a security researcher using our previous model found and disclosed a vulnerability in React that could lead to source code exposure. I believe these models will be a net win for cybersecurity, but we are in the ‘real impact phase’ as they improve. https://x.com/sama/status/2001724828567400700

The Sparks paper was an innovative attempt at trying to figure out ways of pointing at GPT-4 & saying “there is something unexpected here that is hard to measure right now” I think Early Science Acceleration feels similar. A blurry picture that will become clearer coming years https://x.com/emollick/status/2001456094418256077

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading