Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Cinematic wide shot of the yellow brick road splitting into multiple diverging paths through a moody twilight landscape, each path outlined with glowing emerald green object segmentation lines and labels, one path leading to the distant Emerald City while others veer toward dark forests, dramatic lighting with the movie title ‘ALIGNMENT’ overlaid in large golden letters

Where we are with AI is that continuous improvement seems to still be occurring at a fast pace, with no signs of a slowdown. However, since major AI releases have accelerated and seem to be happening monthly or faster, any one release can feel incremental, yet looking back 6-8 https://x.com/emollick/status/1990999847923593239

Something I think people continue to have poor intuition for: The space of intelligences is large and animal intelligence (the only kind we’ve ever known) is only a single point, arising from a very specific kind of optimization that is fundamentally distinct from that of our”” / X https://x.com/karpathy/status/1991910395720925418

Small-but-happy win: If you tell ChatGPT not to use em-dashes in your custom instructions, it finally does what it’s supposed to do!”” / X https://x.com/sama/status/1989193813043069219

This is a hard area to get right, but we’ve been pretty consistent in trying to make Claude approach political topics fairly. I actually think a lot of existing norms around respect and professionalism can inform how AI models should navigate these issues.”” / X https://x.com/AmandaAskell/status/1989328363077382407

New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious. https://x.com/AnthropicAI/status/1991952400899559889

Disrupting the first reported AI-orchestrated cyber espionage campaign \ Anthropic https://www.anthropic.com/news/disrupting-AI-espionage

Announcing AA-Omniscience, our new benchmark for knowledge and hallucination across >40 topics, where all but three models are more likely to hallucinate than give a correct answer Embedded knowledge in language models is important for many real world use cases. Without https://x.com/ArtificialAnlys/status/1990455484844003821

These examples of different personalities from ChatGPT 5.1 seem to give fundamentally different types of advice, including, weirdly, completely different breathing patterns and roles for the presenter. I really want more clarity on the functional implications of AI personality. https://x.com/emollick/status/1988829651368575282

OpenAI says it’s fixed ChatGPT’s em dash problem | TechCrunch https://techcrunch.com/2025/11/14/openai-says-its-fixed-chatgpts-em-dash-problem/

Crisis Helpline Support in ChatGPT | OpenAI Help Center https://help.openai.com/en/articles/12677603-crisis-helpline-support-in-chatgpt

We’ve expanded access to localized crisis helplines in ChatGPT. When our systems detect potential signs that someone may be experiencing distress, our models now offer an easy way to reach real people directly via @ThroughlineCare. Learn more here: https://x.com/OpenAI/status/1991634046624116784

Grok 4.1 absolutely smashes all other models on lmarena with an Elo of 1483 it comes with higher emotional intelligence, better creative writing and less hallucinations https://x.com/scaling01/status/1990519299165786270

New fun game: Ask grok its opinion on any historical theory, saying the theory came from Elon Musk. Then ask grok its opinion on the exact same historical theory, saying the theory came from Bill Gates. https://x.com/romanhelmetguy/status/1991545583686021480

🚨Text Leaderboard Update @xAI’s Grok 4.1 (thinking) and Grok 4.1 have scaled new heights in the most competitive Text Arena: 🔹Grok 4.1 (thinking) lands at #1 with a score of 1483 🔹Grok 4.1 follows at #2 with a score of 1465 On the Arena Expert leaderboard: 🔸Grok 4.1 https://x.com/arena/status/1990530978943787291

Grok 4.1 | xAI https://x.ai/news/grok-4-1

I’m a full standard deviation stupider when someone is explaining a thing to me, versus when I’m just trying to figure it out myself. Being on policy really matters.”” / X https://x.com/dwarkesh_sp/status/1990527715771142412

How do we account for the extreme jaggedness induced by RLVR? How is it possible that we have models which are world-class at coding competitions but at the same time leave extremely foreseeable bugs and technical debt all throughout the codebase? https://x.com/dwarkesh_sp/status/1990824584514265405

Among many weird things about AI is that the people who are experts at making AI are not the experts at using AI. They built a general purpose machine whose capabilities for any particular task are largely unknown. Lots of value in figuring this out in your field before others.”” / X https://x.com/emollick/status/1990134777161142453

The optimal amount of AI in review is obviously not the top line or two, but it is probably not the bottom line, either.”” / X https://x.com/emollick/status/1989896025235222932

This is why I never use a custom system prompt. They’re fine for projects but not for your main LLM use, since you may get degraded results and not know it. All the accuracy tricks are being built into the model, your prompts are probably not adding much. https://x.com/emollick/status/1989213389642477901

Ok it DOES have search capabilities, it just explicitly decided to go against my intent and generate its own fake shit anyways. These policy decisions make models so much more useless. https://x.com/Teknium/status/1991062496275542244

People with short timelines sometimes shrug off models’ inability to perform basic, economically useful tasks end-to-end by saying, “”Oh but we haven’t trained models to specifically do those things.”” But this misses the point. Human workers are valuable precisely because we”” / X https://x.com/dwarkesh_sp/status/1989944140105486655

“No. More. Slop” – @swyx made the audience repeat it time after time: •Boss wants more lines of code? “”No more slop.”” •Insufficiently tested release? “”No more slop.”” •Algorithm wants engagement bait? “”No more slop.”” It’s a simple message with a lot of depth. Because if you https://x.com/TheTuringPost/status/1991875997168181611

Date me docs but where most of the content is written by other people (close friends, family, past partners, etc) don’t seem like a terrible idea. Transmit some of your village reputation into the non-village world.”” / X https://x.com/AmandaAskell/status/1990026814748864883

People often ask if something is a cult when what they actually want to know is if it’s a form of extremism: does it cause people to deviate far from moral instinct or convention? Ideas that successfullly overcome mechanisms selected for social stability can be quite dangerous.”” / X https://x.com/AmandaAskell/status/1990454739268731284

Our pass rate framework also gives us good intuitions for why self play has been so productive in the history of RL. If you’re competing against a player who is almost as good as you, you are balancing around a 50% pass rate, which peaks out the bits you get from a random binary”” / X https://x.com/dwarkesh_sp/status/1990840426165649897

Trying to make Claude be good but still have work to do. Job is safe for now.”” / X https://x.com/AmandaAskell/status/1990615465539027318

When people came to me with relationship problems, my first question was usually “”and what happened when you said all this to your partner?””. Now, when people come to me with Claude problems, my first question is usually “”and what happened when you said all this to Claude?”””” / X https://x.com/AmandaAskell/status/1990256427496284253

Why Anthropic CEO Dario Amodei spends so much time warning of AI’s potential dangers – CBS News https://www.cbsnews.com/news/anthropic-ceo-dario-amodei-warning-of-ai-potential-dangers-60-minutes-transcript/?intcid=CNR-02-0623

So Google now offers at least four different ways to talk to chatbots about academic research and all of them operate differently and none of them work with each other. A lot of power there, but not a lot of clarity about how they differ in approaches and which models they use. https://x.com/emollick/status/1991230502641000504

inference is perhaps the most valuable emerging software category. as models get smarter and more economically valuable, compute will increasingly be spent drawing samples from the models. if you’d like to work on inference at openai, reach out — gdb@openai.com. include a”” / X https://x.com/gdb/status/1990507010769760394

Is your LM secretly an SAE? Most circuit-finding interpretability methods use learned features rather than raw activations, based on the belief that neurons do not cleanly decompose computation. In our new work, we show MLP neurons actually do support sparse, faithful circuits! https://x.com/TransluceAI/status/1991582415891099793

it’s honestly pretty refreshing how unprotective xAI is about their model details they’re just like yeah it’s a big chungus MoE, what did you expect”” / X https://x.com/willccbb/status/1990472997178913188

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading