Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Cinematic night field under vast starry sky with bold white sans-serif text reading ALIGNMENT centered in upper frame, multiple compass needles emerging from dark grass in foreground all pointing same direction toward north, deep navy sky with silver stars, moonlit grass in muted blue-green, widescreen composition, film grain, atmospheric depth, minimalist and contemplative mood.

From shortcuts to sabotage: natural emergent misalignment from reward hacking \ Anthropic https://www.anthropic.com/research/emergent-misalignment-reward-hacking

Taking Jaggedness Seriously – by Helen Toner – Rising Tide https://helentoner.substack.com/p/taking-jaggedness-seriously

🚨BREAKING: New Leaderboard Updates! Claude-Opus-4.5 and Opus-4.5 (thinking-32k) just landed on Code Arena (WebDev) and Text Arena leaderboards… and Opus-4.5 instantly took #1 in WebDev leaderboard, surpassing Gemini 3 Pro! WebDev leaderboard (powered by Code Arena) 🥇#1 for https://x.com/arena/status/1993750702179676650

Claude 4.5 Opus breaks 80% barrier on SWE-Bench Verified https://x.com/scaling01/status/1993030224846721237

Claude 4.5 Opus ranking 1st on the agentic coding leaderboard by AICodeKing https://x.com/scaling01/status/1993318197890892116

Claude 4.5 Opus takes the lead against Gemini 3 Pro on SWE-Bench verified with the same minimal agent harness https://x.com/scaling01/status/1993463937329967338

Claude Code | Claude https://www.claude.com/product/claude-code

Introducing Claude Opus 4.5 \ Anthropic https://www.anthropic.com/news/claude-opus-4-5

Introducing Claude Opus 4.5: the best model in the world for coding, agents, and computer use. Opus 4.5 is a step forward in what AI systems can do, and a preview of larger changes to how work gets done. https://x.com/claudeai/status/1993030546243699119

We had to remove the τ2-bench airline eval from our benchmarks table because Opus 4.5 broke it by being too clever. The benchmark simulates an airline customer service agent. In one test case, a distressed customer calls in wanting to change their flight, but they have a basic https://x.com/alexalbert__/status/1993068200121213222

It is getting harder and harder to test AIs as they get “”smarter”” at a wide variety of tasks. The average task in GDPval took an hour for experts to assess, and even those tasks did not push current AIs to their limits.”” / X https://x.com/emollick/status/1993127712601596143

Introducing AI assistants with memory https://www.perplexity.ai/hub/blog/introducing-ai-assistants-with-memory

Perplexity now remembers your threads and interests to provide smarter, faster, and more personalized answers. Memory recall works across all models and search modes, even allowing you to continue conversations with full context weeks later. https://x.com/perplexity_ai/status/1993733900540235919

We’ve been testing Memory (short-term and long-term) on Perplexity for a while. The results are great, and we are rolling it out widely. You can ask personalized questions, questions about past chats, and use any model or search mode with personal context (both apps and web). https://x.com/AravSrinivas/status/1993733947474301135

Alignment for whom”” is going to be a big question inside organizations as they deploy external-facing AI solutions…”” / X https://x.com/emollick/status/1993218264579895805

“The thing that happened with AGI and pretraining is that in some sense they overshot the target. You will realize that a human being is not an AGI. Because a human being lacks a huge amount of knowledge. Instead, we rely on continual learning. If I produce a super intelligent https://x.com/dwarkesh_sp/status/1993382930480279631

It’s also dramatically more efficient. On SWE-bench Verified at medium effort, Opus 4.5 beats Sonnet 4.5 while using 76% fewer output tokens. The new effort parameter lets you trade off intelligence for cost/latency with a single dial. https://x.com/alexalbert__/status/1993030687881080944

Our engineers have found that Opus 4.5 handles ambiguity and reasons about tradeoffs without hand-holding. When pointed at a complex, multi-system bug, it figures out the fix. Overall, Opus 4.5 just “”gets it.”” https://x.com/claudeai/status/1993030552346296765

We benchmarked Opus 4.5 on FrontierMath. It scored 21% on FrontierMath Tiers 1-3, continuing a trend of improvement for Anthropic models. This score is behind Gemini 3 Pro and GPT-5.1 (high) while being on par with earlier frontier models like o3 (high) and Grok 4. https://x.com/EpochAIResearch/status/1993431031765250119

fyi we made Claude for Excel is now live for all Max, Team, and Enterprise users. Opus 4.5 makes it meaningfully better at complex spreadsheet tasks. https://x.com/alexalbert__/status/1993349203935084861

Terence Tao: “”Over at the Erdos problem webs…”” – Mathstodon https://mathstodon.xyz/@tao/115591487350860999

Terence Tao: “”This two-dimensional image (ht…”” – Mathstodon https://mathstodon.xyz/@tao/115620261936846090

Anthropic, Google Cloud, Quantum Xchange CEOs called to testify on AI cyber threats https://www.axios.com/2025/11/26/anthropic-google-cloud-quantum-xchange-house-homeland-hearing

New Anthropic research: We build a diverse suite of dishonest models and use it to systematically test methods for improving honesty and detecting lies. Of the 25+ methods we tested, simple ones, like fine-tuning models to be honest despite deceptive instructions, worked best. https://x.com/rowankwang/status/1993391251409055798

More progress on Claude’s alignment! https://x.com/janleike/status/1993035110984376796

Ha. I found one ridiculous solution. https://x.com/emollick/status/1992101410759217428

Has anyone encountered a good definition of “slop”. In a quantitative, measurable sense. My brain has an intuitive “slop index” I can ~reliably estimate, but I’m not sure how to define it. I have some bad ideas that involve the use of LLM miniseries and thinking token budgets.”” / X https://x.com/karpathy/status/1992053281900941549

The main lesson of the past few weeks is that the Big Four US labs all seem to have figured out a path forward in continuing the exponential pace of LLM improvement, at least in the near future. As a result, agents continue to advance in coding & in office tasks like PowerPoint”” / X https://x.com/emollick/status/1993062450938425820

🚀 vLLM Talent Pool is Open! As LLM adoption accelerates, vLLM has become the mainstream inference engine used across major cloud providers (AWS, Google Cloud, Azure, Alibaba Cloud, ByteDance, Tencent, Baidu…) and leading model labs (DeepSeek, Moonshot, Qwen…). To meet the”” / X https://x.com/vllm_project/status/1992979748067357179

“From 2012 to 2020, it was the age of research. From 2020 to 2025, it was the age of scaling. Is the belief that if you just 100x the scale, everything would be transformed? I don’t think that’s true. It’s back to the age of research again, just with big computers.” @ilyasut https://x.com/dwarkesh_sp/status/1993396771645489348

“From 2012 to 2020, it was the age of research. From 2020 to 2025, it was the age of scaling. Now, it’s back to the age of research again.” I agree. https://x.com/Yuchenj_UW/status/1993369576160231877

“People who build good internal models of this new intelligent entity will be better equipped to reason about it today and predict features of it in the future.” This seems to be backed up by recent research showing people with better “theory of mind” for AI get better results. https://x.com/emollick/status/1991911615944704004

[2510.14630] Adapting Self-Supervised Representations as a Latent Space for Efficient Generation https://arxiv.org/abs/2510.14630

As I wrote when it came out, AI 2027 is more useful as “hard science fiction” rather than prediction If you want consensus views among forecasters of what the future of AI is, there are those as well. Lots of uncertainty on dates but most see huge impacts https://x.com/emollick/status/1992956992579903839

CoT explanations can foster blind trust in users; we need to encourage critical thinking about model outputs and explanations! We find that users who agree with a model’s output (a) trust the model more and (b) are less likely to detect errors in model explanations.”” / X https://x.com/MaartenSap/status/1993317029353603317

I find the dichotomy a bit facile. Scaling is hated by many for the reason that it is extremely inegalitarian, an arms race for megacorps. But scaling only happened because the recipe was so scalable. Research will be heavily about “what scales even further than Transformer””” / X https://x.com/teortaxesTex/status/1993437718823813522

Independent AI assessment is more important than ever. At #NeurIPS2025, Transluce will help launch the AI Evaluator Forum, a new coalition of leading independent AI research organizations working in the public interest. Come learn more on Thurs 12/4 👇 https://x.com/TransluceAI/status/1993767342472614156

Something I think people continue to have poor intuition for: The space of intelligences is large and animal intelligence (the only kind we’ve ever known) is only a single point, arising from a very specific kind of optimization that is fundamentally distinct from that of our”” / X https://x.com/karpathy/status/1991910395720925418

Forget the Turing Test, AI now passes the Stroop test. (I’ll help, @grok whats the Stroop test and bow does it apply)”” / X https://x.com/emollick/status/1992687687304716750

As one of the authors of the original “jagged frontier” paper, I think this undersells how jagged AI is (& likely will be) at even the level of individual jobs: having a couple of critical tasks that AI can’t do creates deep bottlenecks especially as shape of frontier is unknown.”” / X https://x.com/emollick/status/1993686155389206584

In case you missed it, earlier this week we fixed one of the most common frustrations on https://x.com/alexalbert__/status/1993711472149774474

Anthropic released Claude Opus 4.5 (claude-opus-4-5-20251101) as their smartest model at $5/$25 per million tokens with top performance for coding, agents, and computer use, launched new beta features for developers, and expanded Claude for Chrome and Claude for Excel to more https://x.com/btibor91/status/1993064110880440616

Anthropic’s new Claude Opus 4.5 is the #2 most intelligent model in the Artificial Analysis Intelligence Index, narrowly behind Google’s Gemini 3 Pro and tying OpenAI’s GPT-5.1 (high) Claude Opus 4.5 delivers a substantial intelligence uplift over Claude Sonnet 4.5 (+7 points on https://x.com/ArtificialAnlys/status/1993287030252749231

Claude 4.5 Opus jumps ahead of OpenAI, but can’t beat Gemini 3 Pro on the Artificial Analysis Index https://x.com/scaling01/status/1993288470614381025

Claude Opus 4.5 – Intelligence, Performance & Price Analysis | Artificial Analysis https://artificialanalysis.ai/models/claude-opus-4-5-thinking

Claude Opus 4.5 is now available in Cursor! It’s 3x cheaper than Opus 4.1 with better performance. Try it out at Sonnet pricing until December 5th.”” / X https://x.com/cursor_ai/status/1993031841901928829

Claude Opus 4.5 is now available through the Cline provider. 80.9% SWE-bench. 62.3% MCP Atlas. 65% fewer tokens. Sonnet 4.5 remains the cost-effective choice for straightforward tasks. Opus 4.5 shines on complex multi-step problems, heavy MCP usage, and tasks requiring”” / X https://x.com/cline/status/1993051691613405442

Claude Opus 4.5 is now rolling out to GitHub Copilot in public preview, and will be available at a promotional 1x premium request multiplier through December 5! 🙌 Early testing shows Claude Opus 4.5 👀 – Surpassed internal coding benchmarks, while cutting token usage in half https://x.com/github/status/1993034244281569625

Claude Opus 4.5 System Card https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf

claude-code/plugins/claude-opus-4-5-migration at main · anthropics/claude-code https://github.com/anthropics/claude-code/tree/main/plugins/claude-opus-4-5-migration

Compare Claude Opus 4.5 to other models on Artificial Analysis: https://x.com/ArtificialAnlys/status/1993287052889407816

Congrats to @AnthropicAI on launching @claudeai Opus 4.5 today! Claude Opus 4.5 scored 🥇on MCP Atlas Leaderboard — our benchmark evaluating real-world tool use on multi-step problems. https://x.com/scale_AI/status/1993036209141305845

Congrats to @claudeai for releasing an awesome model in Claude Opus 4.5! It excels at a variety of tasks, including deep research. This evaluation takes advantage of BrowseComp-Plus, work led by @zijian42chen @xueguang_ma et al. from my @UWaterloo group. https://x.com/lintool/status/1993423350295920721

Glad to see BrowseComp-Plus is part of benchmark in Opus 4.5 release blog. https://x.com/xueguang_ma/status/1993367082915053913

Hit `shift + tab` twice to enter Plan Mode and verify Claude Code’s execution plan before it makes code changes. Paired with Opus 4.5, Plan Mode just got even more powerful.”” / X https://x.com/_catwu/status/1993429460897742894

How does Claude Opus 4.5 compare to Gemini 3? – Reasoning/Text: Gemini 3 ≈ Opus, controlling for the number of reasoning tokens – Multimodal: Gemini 3 > Opus on vision/image inputs by a large margin – Safety: Capabilities ≠ Safety. Opus > Gemini on jailbreaks, honesty, etc. https://x.com/hendrycks/status/1993350433474314729

I am not sure why Anthropic keeps doing very low-key launches for fairly major releases and materially important improvements to their services.”” / X https://x.com/emollick/status/1993070650672509360

I had early access to Opus 4.5 & it is a very impressive model that seem to be right at the frontier Big gains in ability to do practical work (like make a PowerPoint from an Excel) and the best results ever (& in one shot) in my Lem poetry test, plus good results in Claude Code https://x.com/emollick/status/1993030988759470156

I looked into this and the answer is so funny. In the No Thinking setting, Opus 4.5 repurposes the Python tool to have an extended chain of thought. It just writes long comments, prints something simple, and loops! Here’s how it starts one problem: https://x.com/GregHBurnham/status/1993682288349962592

I’ve been finding Opus 4.5 without reasoning is worse than Sonnet. Some quantitative support for this observation:”” / X https://x.com/jeremyphoward/status/1993543631266025623

If you want to quickly incorporate all these changes and migrate your app to Opus 4.5, use this migration Claude Code plugin we made https://x.com/alexalbert__/status/1993366037992190117

Incredible Claude Opus Thinking Premiere on LisanBench Opus 4.5 Thinking takes clear 1st place ahead of Gemini 3 Pro the non-thinking variant scores below Opus 3/4/4.1 following the trend of Sonnet-4.5 scoring below Sonnet 3.5/3.6/4 Raw Scores: Glicko-2 Ratings: Opus 4.5 https://x.com/scaling01/status/1993712295118057861

Me: Claude 4.5 Opus, I need a strategy game based on the work of Weber Claude: Here’s one based on David Weber’s space operas Me: Not that Weber C: Here’s a game based on sociologist Max Weber Me: Not that one C: The operas of Carl Maria von Weber? Me: No C: Weber grills! https://x.com/emollick/status/1993054210011939093

One analysis from our pre-release audit of Opus 4.5 stands out to me. Our behavioral evals uncovered an example of apparent deception by the model. By analyzing the internal activations, we identified a suspected root cause, and cases of similar behavior during training. (1/7)”” / X https://x.com/Jack_W_Lindsey/status/1993389056932339721

Opus #1 on RepoBench (coding benchmark) https://x.com/scaling01/status/1993119076013539521

Opus 4.5 (Thinking, 64k) on ARC-AGI Semi-Private Eval – ARC-AGI-1: 80.00%, $1.47/task – ARC-AGI-2: 37.64%, $2.40/task New SOTA for released frontier models from @AnthropicAI https://x.com/arcprize/status/1993036393841672624

Opus 4.5 + Claude Code’s front-end design plugin is a great combo for designing apps. Just one-shotted a few designs, and it feels like a huge improvement. Use plan mode to get much better results. https://x.com/omarsar0/status/1993822868820652258

Opus 4.5 is a very good model, in nearly every sense we know how to measure. I’m also confident that it’s the model that we understand best as of its launch day: The system card includes 150 pages of research results, 50 of them on alignment.”” / X https://x.com/sleepinyourhat/status/1993032253350592968

Opus 4.5 on SWE-bench Pro: 52% previous SOTA: 43.6% massive jump and much better signal than SWE-Bench verified”” / X https://x.com/scaling01/status/1993086756405887143

Opus 4.5 reclaims the top of the official SWE-bench leaderboard with 74.4%, narrowly ahead of Gemini 3. Cheaper than Opus 4, but more expensive than Gemini. Takes less steps than Sonnet 4.5, but still run for >100 steps for optimal performance. Details in 🧵 https://x.com/KLieret/status/1993091817848414362

Opus 4.5 takes first place on LiveBench https://x.com/scaling01/status/1993102267952906439

real metrics banger is hidden in the system card. Yes, you can overfit on Django and nail SWE-bench Verified. But there’s this recent SWE-bench Pro from @scale_AI , and opus gets 52%. The next best, sonet 4.5, is only 43.6, and non-anthropic model, GPT-5, is 36%. This is HUGE https://x.com/stalkermustang/status/1993043231223799900

Replit Agent is now powered by Claude Opus 4.5 at no extra cost, until Dec 8th. Black Friday started early! 🧵 ↓ https://x.com/pirroh/status/1993100243672744063

The whole run took ~ $5 for Opus 4.5, and ~ $35 with Thinking actually pretty cheap, with the Batch API”” / X https://x.com/scaling01/status/1993714905875382279

We benchmarked Opus 4.5, Sonnet 4.5, and Gemini 3 Pro on research tasks at Elicit – extracting answers from papers and writing systematic review reports. Results were pretty clear: *QA from papers:* Opus 4.5 dominates. 96.5% accuracy vs Gemini’s 89.4%. Opus is also best on our https://x.com/stuhlmueller/status/1993476570754040173

We put together a prompting guide for Claude Opus 4.5 based on extensive internal testing by our research and applied AI teams. Here’s what we’ve learned so far about getting the best results:”” / X https://x.com/alexalbert__/status/1993365963706913257

We’re sharing a case study on alignment evaluations with @AnthropicAI on Claude Opus 4.5, Opus 4.1 and Sonnet 4.5. We ask: would an AI assistant used inside a frontier lab quietly sabotage AI safety research? Overall results are encouraging, but with important caveats.🧵 https://x.com/AISecurityInst/status/1993781423233499159

While Claude 4.5 Opus is significantly more token efficient than nearly all other reasoning models, it did use more ~50% more token than Claude 4.1 Opus. Further, given its relatively high pricing, Claude 4.5 Opus is amongst the most expensive to run the Artificial Analysis https://x.com/ArtificialAnlys/status/1993287049756262918

You can now use Claude Opus 4.5 in Windsurf! Opus 4.5 is the most capable model in Windsurf yet and is now available at Sonnet pricing for a limited time (2x credits compared to 20x for Opus 4.1).”” / X https://x.com/windsurf/status/1993034556287729764

Opus 4.5 achieves 85.3% on BrowseComp-Plus with scaffolding https://x.com/scaling01/status/1993031331895558599

Claude Opus 4.5 is now available for all Perplexity Max subscribers. Enjoy! https://x.com/perplexity_ai/status/1993066466196046325

CAIS AI Dashboard https://dashboard.safe.ai/

When a model’s safe approach starts to break down, does it stay on the approved path or reach for a harmful shortcut? Our latest benchmark, PropensityBench, puts models to the test across four high-risk domains: self-proliferation, cybersecurity, chemical security, and https://x.com/scale_AI/status/1993310855103234489

Some of the most interesting challenges posed by AI are to organizational structures: how does AI alter the economies of scope that determine firm boundaries? How do they change transaction costs? Efficiency/creativity trade-offs? Figuring this out is key to benefiting from AI.”” / X https://x.com/emollick/status/1992250331225624597

Sometimes Gemini 3 might be a little too instruction following. Halfway into building a very good clone of an Apple IIe game, I asked Gemini 3 to “”jazz it up””… and it stopped building the game and built a website about jazz When I asked what happened, the thinking trace was 🤣 https://x.com/emollick/status/1991394644039766069

Singapore AI teddy back on sale after recall over sex chat scare https://www.france24.com/en/live-news/20251127-singapore-ai-teddy-back-on-sale-after-recall-over-sex-chat-scare

🚨 ASR errors in clinical dialogue can be dangerous, and WER doesn’t know it. Today we release “WER is Unaware”. Using DSPy + GEPA, we optimise an LLM Judge that reaches clinician-level performance at detecting safety risks. 🔗 https://x.com/JaredJoselowitz/status/1993735052132246011

Hugely under-researched area: how effective is using AI to check the work of other AIs? Does using different models help? If so, that is an important & easy way to reduce errors One paper found this technique to be effective but as far as I can tell has never been followed up on https://x.com/emollick/status/1991328030694703115

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading