Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Cinematic night field under expansive starry sky where stars form subtle circuit board patterns, bold white sans-serif text reading TECH centered in upper frame like a movie title card, faint blue glow on distant horizon, deep navy sky, moonlit grass, widescreen composition, film grain texture, high contrast typography.

Claude Desktop now supports “”multi-clauding”” for both local and cloud sessions. This has been one of our top requests. Excited to see what you build with it!”” / X https://x.com/_catwu/status/1993428129197834741

Claude for Excel | Claude https://www.claude.com/claude-for-excel

Jeff Bezos’ New AI Venture Quietly Acquired an Agentic Computing Startup | WIRED https://www.wired.com/story/jeff-bezos-new-ai-company-acquired-agentic-computing-startup/

Fara-7B: An Efficient Agentic Model for Computer Use – Microsoft Research https://www.microsoft.com/en-us/research/blog/fara-7b-an-efficient-agentic-model-for-computer-use/

Effort – Claude Docs https://platform.claude.com/docs/en/build-with-claude/effort

MCP Apps: Extending servers with interactive user interfaces | Model Context Protocol Blog https://blog.modelcontextprotocol.io/posts/2025-11-21-mcp-apps/

🚨BREAKING: New Leaderboard Updates! Claude-Opus-4.5 and Opus-4.5 (thinking-32k) just landed on Code Arena (WebDev) and Text Arena leaderboards… and Opus-4.5 instantly took #1 in WebDev leaderboard, surpassing Gemini 3 Pro! WebDev leaderboard (powered by Code Arena) 🥇#1 for https://x.com/arena/status/1993750702179676650

Claude 4.5 Opus breaks 80% barrier on SWE-Bench Verified https://x.com/scaling01/status/1993030224846721237

Claude 4.5 Opus ranking 1st on the agentic coding leaderboard by AICodeKing https://x.com/scaling01/status/1993318197890892116

Claude 4.5 Opus takes the lead against Gemini 3 Pro on SWE-Bench verified with the same minimal agent harness https://x.com/scaling01/status/1993463937329967338

Claude Code | Claude https://www.claude.com/product/claude-code

Introducing Claude Opus 4.5 \ Anthropic https://www.anthropic.com/news/claude-opus-4-5

Introducing Claude Opus 4.5: the best model in the world for coding, agents, and computer use. Opus 4.5 is a step forward in what AI systems can do, and a preview of larger changes to how work gets done. https://x.com/claudeai/status/1993030546243699119

We had to remove the τ2-bench airline eval from our benchmarks table because Opus 4.5 broke it by being too clever. The benchmark simulates an airline customer service agent. In one test case, a distressed customer calls in wanting to change their flight, but they have a basic https://x.com/alexalbert__/status/1993068200121213222

It is getting harder and harder to test AIs as they get “”smarter”” at a wide variety of tasks. The average task in GDPval took an hour for experts to assess, and even those tasks did not push current AIs to their limits.”” / X https://x.com/emollick/status/1993127712601596143

Nvidia says its GPUs are a ‘generation ahead’ of Google’s AI chips https://www.cnbc.com/2025/11/25/nvidia-says-its-gpus-are-a-generation-ahead-of-googles-ai-chips.html

We’re delighted by Google’s success — they’ve made great advances in AI and we continue to supply to Google. NVIDIA is a generation ahead of the industry — it’s the only platform that runs every AI model and does it everywhere computing is done. NVIDIA offers greater”” / X https://x.com/nvidianewsroom/status/1993364210948936055?s=20

Alignment for whom”” is going to be a big question inside organizations as they deploy external-facing AI solutions…”” / X https://x.com/emollick/status/1993218264579895805

“The thing that happened with AGI and pretraining is that in some sense they overshot the target. You will realize that a human being is not an AGI. Because a human being lacks a huge amount of knowledge. Instead, we rely on continual learning. If I produce a super intelligent https://x.com/dwarkesh_sp/status/1993382930480279631

It’s also dramatically more efficient. On SWE-bench Verified at medium effort, Opus 4.5 beats Sonnet 4.5 while using 76% fewer output tokens. The new effort parameter lets you trade off intelligence for cost/latency with a single dial. https://x.com/alexalbert__/status/1993030687881080944

Our engineers have found that Opus 4.5 handles ambiguity and reasons about tradeoffs without hand-holding. When pointed at a complex, multi-system bug, it figures out the fix. Overall, Opus 4.5 just “”gets it.”” https://x.com/claudeai/status/1993030552346296765

We benchmarked Opus 4.5 on FrontierMath. It scored 21% on FrontierMath Tiers 1-3, continuing a trend of improvement for Anthropic models. This score is behind Gemini 3 Pro and GPT-5.1 (high) while being on par with earlier frontier models like o3 (high) and Grok 4. https://x.com/EpochAIResearch/status/1993431031765250119

fyi we made Claude for Excel is now live for all Max, Team, and Enterprise users. Opus 4.5 makes it meaningfully better at complex spreadsheet tasks. https://x.com/alexalbert__/status/1993349203935084861

The Economics of Replacing Call Center Workers With AIs — LessWrong https://www.lesswrong.com/posts/rJatmEDcYrDQcwstT/the-economics-of-replacing-call-center-workers-with-ais

Black Forest Labs – Frontier AI Lab https://bfl.ai/research/representation-comparison

This NVIDIA paper just broke my brain. Everyone keeps talking about scaling transformers with bigger clusters and smarter optimizers… meanwhile NVIDIA and Oxford just showed you can train billion-parameter models using evolution strategies a method most people wrote off as https://x.com/rryssf_/status/1993672852206444675

Anthropic system cards are simply the best in the game so much info, even included new benchmarks like AA-Omniscience https://x.com/scaling01/status/1993032258677293357

Context editing – Claude Docs https://platform.claude.com/docs/en/build-with-claude/context-editing#client-side-compaction-sdk

Here’s Anthropic’s write up of “”advanced tool use””: https://t.co/4oEOIAHI4O And the “”tool loadout”” pattern: https://x.com/dbreunig/status/1993387763291635882

New Anthropic research: Estimating AI productivity gains from Claude conversations. The Anthropic Economic Index tells us where Claude is used, and for which tasks. But it doesn’t tell us how useful Claude is. How much time does it save? https://x.com/AnthropicAI/status/1993305312305009133

Our study has limitations: above all, Claude can’t use what happens outside of the chat window to refine its estimate of task-level savings. But as models improve, we think its estimates of task-level savings will improve too. We’ll return to this research soon.”” / X https://x.com/AnthropicAI/status/1993305334484705533

The SWE-bench Verified leaderboard is a fair competition environment for all models since they must all use mini-SWE-agent, so private scaffolding improvements that each company creates do not effect performance. Congrats to Anthropic! https://x.com/OfirPress/status/1993116355059703917

Tool Use Examples JSON Schema defines what’s valid, not what’s correct. Now you can show Claude concrete usage patterns directly in tool definitions to improve Claude’s accuracy and knowledge when using tools.”” / X https://x.com/alexalbert__/status/1993038680177754574

We are still in an era where no model dominates everything. For people who do a lot with AI, you are going to be alternating between Gemini, Claude & ChatGPT. And that isn’t only because models have specific skills, each has a personality that contributes to utility on tasks.”” / X https://x.com/emollick/status/1993074733001384115

More progress on Claude’s alignment! https://x.com/janleike/status/1993035110984376796

Estimating AI productivity gains \ Anthropic https://www.anthropic.com/research/estimating-productivity-gains

MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology Going beyond the standard single-turn multiple choice benchmarks, this paper introduces a multimodal longitudinal agentic benchmark that simulates tumor boards, where oncologists review patient https://x.com/iScienceLuvr/status/1993645980869365960

🚀 LangChain 1.1 You can now programmatically access model capabilities and supported features through model.profile. This powers some new features in middleware. One example: SummarizationMiddleware can now dynamically trigger based on the model’s available context.”” / X https://x.com/LangChainAI/status/1993407335386189949

Agent Design Is Still Hard | Armin Ronacher’s Thoughts and Writings https://lucumr.pocoo.org/2025/11/21/agents-are-hard/

Agent Frameworks, Runtimes, and Harnesses- oh my! https://blog.langchain.com/agent-frameworks-runtimes-and-harnesses-oh-my/

AI has been trained on the entire corpus of art, so knowing something about the history of design yourself is helpful. Here is “a poster advertising the concept of free will” in Sachplakat style, 1970s Polish cinema poster style, Constructivist, & International Typographic Style https://x.com/emollick/status/1993758684800012421

deepagents/libs/deepagents-cli at master · langchain-ai/deepagents https://github.com/langchain-ai/deepagents/tree/master/libs/deepagents-cli

Interleaved thinking is a game-changer. I built this little deep research agent, and the results are impressive. The agent is just more efficient at reasoning over multiple steps. Huge leverage for self-improving agents. https://x.com/omarsar0/status/1993689618856689789

Most companies talk about AI agents in theory. @bookingcom shipped one that handles thousands of customer conversations per day. Here’s what they built: Their autonomous agent helps accommodation partners respond to guest inquiries faster and more accurately. The agent can https://x.com/victorialslocum/status/1993636038313443826

Multi-agent systems are powerful but expensive. However, the cost isn’t in the reasoning itself. It’s in the communication. Agents exchange full text messages, consuming tokens for every coordination step. When agents need to collaborate on complex problems, this overhead adds https://x.com/dair_ai/status/1993697268848115915

VS Code and GitHub Copilot’s daily builds now include release notes! Available with the “”Show Release Notes”” command or on the @code website: https://x.com/pierceboggan/status/1993364247431245848

What’s the difference between agent runtime, a framework, and a harness? In this post we expand on the patterns we’re seeing, when to use different approaches, and how things are evolving. Our mental model in brief: – Frameworks: abstractions that help build agents with LLMs”” / X https://x.com/LangChainAI/status/1993746547587338508

Gemini 3 is SOTA on SWE Bench Verified with a standard agent harness (all models using the same independently created harness) 🤯 https://x.com/OfficialLoganK/status/1991327656990896372

Just back from visiting our Microsoft AI Asia teams in China. Blown away by their pace, execution & creativity. Loved seeing the work on multi-agent “chain-of-debate” AIs at our hackathon. Huge thanks to everyone in Suzhou and Beijing for the hospitality! https://x.com/mustafasuleyman/status/1993704093202997621

LLMs can’t see. How can we build effective multi-agent systems with vision capabilities? Building multimodal models from scratch is expensive. Training joint vision-language architectures requires massive compute, specialized datasets, and careful optimization. But there’s https://x.com/dair_ai/status/1993367363790717060

Navigating Gigapixel Pathology Images with Large Multimodal Models GPT-5 with the right agentic scaffold for navigating whole slide images outperforms slide-level pathology models on the novel MultiPathQA benchmark. https://x.com/iScienceLuvr/status/1993650850120818888

🚀 vLLM Talent Pool is Open! As LLM adoption accelerates, vLLM has become the mainstream inference engine used across major cloud providers (AWS, Google Cloud, Azure, Alibaba Cloud, ByteDance, Tencent, Baidu…) and leading model labs (DeepSeek, Moonshot, Qwen…). To meet the”” / X https://x.com/vllm_project/status/1992979748067357179

“From 2012 to 2020, it was the age of research. From 2020 to 2025, it was the age of scaling. Is the belief that if you just 100x the scale, everything would be transformed? I don’t think that’s true. It’s back to the age of research again, just with big computers.” @ilyasut https://x.com/dwarkesh_sp/status/1993396771645489348

“From 2012 to 2020, it was the age of research. From 2020 to 2025, it was the age of scaling. Now, it’s back to the age of research again.” I agree. https://x.com/Yuchenj_UW/status/1993369576160231877

“People who build good internal models of this new intelligent entity will be better equipped to reason about it today and predict features of it in the future.” This seems to be backed up by recent research showing people with better “theory of mind” for AI get better results. https://x.com/emollick/status/1991911615944704004

[2510.14630] Adapting Self-Supervised Representations as a Latent Space for Efficient Generation https://arxiv.org/abs/2510.14630

As I wrote when it came out, AI 2027 is more useful as “hard science fiction” rather than prediction If you want consensus views among forecasters of what the future of AI is, there are those as well. Lots of uncertainty on dates but most see huge impacts https://x.com/emollick/status/1992956992579903839

CoT explanations can foster blind trust in users; we need to encourage critical thinking about model outputs and explanations! We find that users who agree with a model’s output (a) trust the model more and (b) are less likely to detect errors in model explanations.”” / X https://x.com/MaartenSap/status/1993317029353603317

I find the dichotomy a bit facile. Scaling is hated by many for the reason that it is extremely inegalitarian, an arms race for megacorps. But scaling only happened because the recipe was so scalable. Research will be heavily about “what scales even further than Transformer””” / X https://x.com/teortaxesTex/status/1993437718823813522

Independent AI assessment is more important than ever. At #NeurIPS2025, Transluce will help launch the AI Evaluator Forum, a new coalition of leading independent AI research organizations working in the public interest. Come learn more on Thurs 12/4 👇 https://x.com/TransluceAI/status/1993767342472614156

Something I think people continue to have poor intuition for: The space of intelligences is large and animal intelligence (the only kind we’ve ever known) is only a single point, arising from a very specific kind of optimization that is fundamentally distinct from that of our”” / X https://x.com/karpathy/status/1991910395720925418

As one of the authors of the original “jagged frontier” paper, I think this undersells how jagged AI is (& likely will be) at even the level of individual jobs: having a couple of critical tasks that AI can’t do creates deep bottlenecks especially as shape of frontier is unknown.”” / X https://x.com/emollick/status/1993686155389206584

Call Now, Fetch Later: MCP SEP-1686 – by Adam Azzam https://aaazzam.substack.com/p/call-now-fetch-later-mcp-sep-1686?triedRedirect=true

General purpose agents like Claude Code and Manus use remarkably few tools. How? By giving agents access to a computer. With bash and filesystem tools, agents can perform actions without needing specialized bound tools for every task. Skills also offer two key advantages over”” / X https://x.com/LangChainAI/status/1993379154868519217

I get LOTS of questions about deploying DSPy programs, so @isaacbmiller1 and I built dspy-cli: a tool that serves DSPy programs as HTTP APIs with Docker config, OpenAPI specs, MCP support, and more. Here’s a quick intro: https://x.com/dbreunig/status/1993462894814703640

It’s not flashy. It’s infrastructure. But it’s the kind of engineering a protocol like MCP deserves. Task support is coming to @fastmcp , powered by @PrefectIO. More soon.”” / X https://x.com/AAAzzam/status/1993495232881869138

MCP gateways have proven to be a critical piece of infrastructure of bringing MCP into enterprise settings, and that gateway infra is enabling creative ways of solving downstream MCP challenges. One of them: solving tool overload on a use-case by use-case basis.”” / X https://x.com/tadasayy/status/1993410677948785022

Pricing is $5/$25 per million tokens. Available now on the Claude API and all three major cloud platforms (Amazon Bedrock, Google Cloud’s Vertex AI, and Microsoft Foundry). Read more here: https://x.com/alexalbert__/status/1993030702053703746

SEP-1686 ships today in MCP. It adds task-based execution to the protocol–background long-running work, poll for status, retrieve results when done. Here’s why this matters 🧵”” / X https://x.com/AAAzzam/status/1993495222035399060

Tool Search Tool Instead of loading all tool definitions upfront, Claude discovers tools on-demand. Mark tools with defer_loading: true and only pays tokens for tools Claude actually needs. Up to an 85% token reduction and big boost in accuracy on our MCP evals (79.5% to 88.1%) https://x.com/alexalbert__/status/1993038651916533768

Anthropic released Claude Opus 4.5 (claude-opus-4-5-20251101) as their smartest model at $5/$25 per million tokens with top performance for coding, agents, and computer use, launched new beta features for developers, and expanded Claude for Chrome and Claude for Excel to more https://x.com/btibor91/status/1993064110880440616

Anthropic’s new Claude Opus 4.5 is the #2 most intelligent model in the Artificial Analysis Intelligence Index, narrowly behind Google’s Gemini 3 Pro and tying OpenAI’s GPT-5.1 (high) Claude Opus 4.5 delivers a substantial intelligence uplift over Claude Sonnet 4.5 (+7 points on https://x.com/ArtificialAnlys/status/1993287030252749231

Claude 4.5 Opus jumps ahead of OpenAI, but can’t beat Gemini 3 Pro on the Artificial Analysis Index https://x.com/scaling01/status/1993288470614381025

Claude Opus 4.5 – Intelligence, Performance & Price Analysis | Artificial Analysis https://artificialanalysis.ai/models/claude-opus-4-5-thinking

Claude Opus 4.5 is now available in Cursor! It’s 3x cheaper than Opus 4.1 with better performance. Try it out at Sonnet pricing until December 5th.”” / X https://x.com/cursor_ai/status/1993031841901928829

Claude Opus 4.5 is now available through the Cline provider. 80.9% SWE-bench. 62.3% MCP Atlas. 65% fewer tokens. Sonnet 4.5 remains the cost-effective choice for straightforward tasks. Opus 4.5 shines on complex multi-step problems, heavy MCP usage, and tasks requiring”” / X https://x.com/cline/status/1993051691613405442

Claude Opus 4.5 is now rolling out to GitHub Copilot in public preview, and will be available at a promotional 1x premium request multiplier through December 5! 🙌 Early testing shows Claude Opus 4.5 👀 – Surpassed internal coding benchmarks, while cutting token usage in half https://x.com/github/status/1993034244281569625

Claude Opus 4.5 System Card https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf

claude-code/plugins/claude-opus-4-5-migration at main · anthropics/claude-code https://github.com/anthropics/claude-code/tree/main/plugins/claude-opus-4-5-migration

Compare Claude Opus 4.5 to other models on Artificial Analysis: https://x.com/ArtificialAnlys/status/1993287052889407816

Congrats to @AnthropicAI on launching @claudeai Opus 4.5 today! Claude Opus 4.5 scored 🥇on MCP Atlas Leaderboard — our benchmark evaluating real-world tool use on multi-step problems. https://x.com/scale_AI/status/1993036209141305845

Congrats to @claudeai for releasing an awesome model in Claude Opus 4.5! It excels at a variety of tasks, including deep research. This evaluation takes advantage of BrowseComp-Plus, work led by @zijian42chen @xueguang_ma et al. from my @UWaterloo group. https://x.com/lintool/status/1993423350295920721

Glad to see BrowseComp-Plus is part of benchmark in Opus 4.5 release blog. https://x.com/xueguang_ma/status/1993367082915053913

Hit `shift + tab` twice to enter Plan Mode and verify Claude Code’s execution plan before it makes code changes. Paired with Opus 4.5, Plan Mode just got even more powerful.”” / X https://x.com/_catwu/status/1993429460897742894

How does Claude Opus 4.5 compare to Gemini 3? – Reasoning/Text: Gemini 3 ≈ Opus, controlling for the number of reasoning tokens – Multimodal: Gemini 3 > Opus on vision/image inputs by a large margin – Safety: Capabilities ≠ Safety. Opus > Gemini on jailbreaks, honesty, etc. https://x.com/hendrycks/status/1993350433474314729

I am not sure why Anthropic keeps doing very low-key launches for fairly major releases and materially important improvements to their services.”” / X https://x.com/emollick/status/1993070650672509360

I had early access to Opus 4.5 & it is a very impressive model that seem to be right at the frontier Big gains in ability to do practical work (like make a PowerPoint from an Excel) and the best results ever (& in one shot) in my Lem poetry test, plus good results in Claude Code https://x.com/emollick/status/1993030988759470156

I looked into this and the answer is so funny. In the No Thinking setting, Opus 4.5 repurposes the Python tool to have an extended chain of thought. It just writes long comments, prints something simple, and loops! Here’s how it starts one problem: https://x.com/GregHBurnham/status/1993682288349962592

I’ve been finding Opus 4.5 without reasoning is worse than Sonnet. Some quantitative support for this observation:”” / X https://x.com/jeremyphoward/status/1993543631266025623

If you want to quickly incorporate all these changes and migrate your app to Opus 4.5, use this migration Claude Code plugin we made https://x.com/alexalbert__/status/1993366037992190117

Incredible Claude Opus Thinking Premiere on LisanBench Opus 4.5 Thinking takes clear 1st place ahead of Gemini 3 Pro the non-thinking variant scores below Opus 3/4/4.1 following the trend of Sonnet-4.5 scoring below Sonnet 3.5/3.6/4 Raw Scores: Glicko-2 Ratings: Opus 4.5 https://x.com/scaling01/status/1993712295118057861

Me: Claude 4.5 Opus, I need a strategy game based on the work of Weber Claude: Here’s one based on David Weber’s space operas Me: Not that Weber C: Here’s a game based on sociologist Max Weber Me: Not that one C: The operas of Carl Maria von Weber? Me: No C: Weber grills! https://x.com/emollick/status/1993054210011939093

One analysis from our pre-release audit of Opus 4.5 stands out to me. Our behavioral evals uncovered an example of apparent deception by the model. By analyzing the internal activations, we identified a suspected root cause, and cases of similar behavior during training. (1/7)”” / X https://x.com/Jack_W_Lindsey/status/1993389056932339721

Opus #1 on RepoBench (coding benchmark) https://x.com/scaling01/status/1993119076013539521

Opus 4.5 (Thinking, 64k) on ARC-AGI Semi-Private Eval – ARC-AGI-1: 80.00%, $1.47/task – ARC-AGI-2: 37.64%, $2.40/task New SOTA for released frontier models from @AnthropicAI https://x.com/arcprize/status/1993036393841672624

Opus 4.5 + Claude Code’s front-end design plugin is a great combo for designing apps. Just one-shotted a few designs, and it feels like a huge improvement. Use plan mode to get much better results. https://x.com/omarsar0/status/1993822868820652258

Opus 4.5 is a very good model, in nearly every sense we know how to measure. I’m also confident that it’s the model that we understand best as of its launch day: The system card includes 150 pages of research results, 50 of them on alignment.”” / X https://x.com/sleepinyourhat/status/1993032253350592968

Opus 4.5 on SWE-bench Pro: 52% previous SOTA: 43.6% massive jump and much better signal than SWE-Bench verified”” / X https://x.com/scaling01/status/1993086756405887143

Opus 4.5 reclaims the top of the official SWE-bench leaderboard with 74.4%, narrowly ahead of Gemini 3. Cheaper than Opus 4, but more expensive than Gemini. Takes less steps than Sonnet 4.5, but still run for >100 steps for optimal performance. Details in 🧵 https://x.com/KLieret/status/1993091817848414362

Opus 4.5 takes first place on LiveBench https://x.com/scaling01/status/1993102267952906439

real metrics banger is hidden in the system card. Yes, you can overfit on Django and nail SWE-bench Verified. But there’s this recent SWE-bench Pro from @scale_AI , and opus gets 52%. The next best, sonet 4.5, is only 43.6, and non-anthropic model, GPT-5, is 36%. This is HUGE https://x.com/stalkermustang/status/1993043231223799900

Replit Agent is now powered by Claude Opus 4.5 at no extra cost, until Dec 8th. Black Friday started early! 🧵 ↓ https://x.com/pirroh/status/1993100243672744063

The whole run took ~ $5 for Opus 4.5, and ~ $35 with Thinking actually pretty cheap, with the Batch API”” / X https://x.com/scaling01/status/1993714905875382279

We benchmarked Opus 4.5, Sonnet 4.5, and Gemini 3 Pro on research tasks at Elicit – extracting answers from papers and writing systematic review reports. Results were pretty clear: *QA from papers:* Opus 4.5 dominates. 96.5% accuracy vs Gemini’s 89.4%. Opus is also best on our https://x.com/stuhlmueller/status/1993476570754040173

We put together a prompting guide for Claude Opus 4.5 based on extensive internal testing by our research and applied AI teams. Here’s what we’ve learned so far about getting the best results:”” / X https://x.com/alexalbert__/status/1993365963706913257

We’re sharing a case study on alignment evaluations with @AnthropicAI on Claude Opus 4.5, Opus 4.1 and Sonnet 4.5. We ask: would an AI assistant used inside a frontier lab quietly sabotage AI safety research? Overall results are encouraging, but with important caveats.🧵 https://x.com/AISecurityInst/status/1993781423233499159

While Claude 4.5 Opus is significantly more token efficient than nearly all other reasoning models, it did use more ~50% more token than Claude 4.1 Opus. Further, given its relatively high pricing, Claude 4.5 Opus is amongst the most expensive to run the Artificial Analysis https://x.com/ArtificialAnlys/status/1993287049756262918

You can now use Claude Opus 4.5 in Windsurf! Opus 4.5 is the most capable model in Windsurf yet and is now available at Sonnet pricing for a limited time (2x credits compared to 20x for Opus 4.1).”” / X https://x.com/windsurf/status/1993034556287729764

Opus 4.5 achieves 85.3% on BrowseComp-Plus with scaffolding https://x.com/scaling01/status/1993031331895558599

Claude Opus 4.5 is now available for all Perplexity Max subscribers. Enjoy! https://x.com/perplexity_ai/status/1993066466196046325

This result implies a doubling of the baseline labor productivity growth trend–placing our estimate towards the upper end of recent studies. And if models improve, the effect could be larger still. https://x.com/AnthropicAI/status/1993305330869223463

We’re launching a new frontier physics eval on Artificial Analysis where no model achieves greater than 9%: CritPt (Complex Research using Integrated Thinking – Physics Test) Developed by 60+ researchers from 30+ institutions across the world including the Argonne National https://x.com/ArtificialAnlys/status/1991913465968222555

How far can multimodal LLMs push real-world recommendations? 🤔 Today we’re sharing China knowledge platform Zhihu’s latest technical practice: how multimodal LLMs (Qwen2.5-VL) upgrade content understanding and cold-start performance in large-scale recsys. ‼️Modern recsys has https://x.com/ZhihuFrontier/status/1993570114810396761

HP is betting $1 billion on AI — even if it means cutting thousands of jobs, says CEO https://finance.yahoo.com/news/hp-is-betting-1-billion-on-ai–even-if-it-means-cutting-thousands-of-jobs-says-ceo-223617822.html

Interesting how absolutely stable the underlying dynamics of AI development have been: 1) Six month doubling time for AI capabilities (METR is just one measure, but others are similar) 2) Open weights models lag 8 months or so behind. Baseline assumption should be this continues”” / X https://x.com/emollick/status/1991890368649179327

IQ Test | Tracking AI https://trackingai.org/home

LLM as a judge has become a dominant way to evaluate how good a model is at solving a task, since it works without a test set and handles cases where answers are not unique. But despite how widely this is used, almost all reported results are highly biased. Excited to share our https://x.com/Kangwook_Lee/status/1993438649963164121

METR is external evaluator I hold in highest regard, and I think a lot of frontier lab staff would say the same.”” / X https://x.com/andy_l_jones/status/1993485558044410188

CAIS AI Dashboard https://dashboard.safe.ai/

When a model’s safe approach starts to break down, does it stay on the approved path or reach for a harmful shortcut? Our latest benchmark, PropensityBench, puts models to the test across four high-risk domains: self-proliferation, cybersecurity, chemical security, and https://x.com/scale_AI/status/1993310855103234489

Benchmark Scores = General Capability + Claudiness https://epochai.substack.com/p/benchmark-scores-general-capability

In the latest Chain of Thought, @Bckenstler, @afeyzaakyurek, @agxsai, and @calvincbzhang dive deep on our newest Professional Reasoning Benchmark (PRBench). Together, they explore why many models struggle to perform on real-world legal and financial reasoning tasks: https://x.com/scale_AI/status/1991589754199240841

AI infrastructure in the “”Era of experience”” https://www.tensoreconomics.com/p/ai-infrastructure-in-the-era-of-experience

What will the next-gen LLM architecture look like? This question keeps sparking debates — and Zhihu contributor & developer Yuxuan offers a sharp comparison between DeepSeek Sparse Attention (DSA) and Native Sparse Attention (NSA), plus a practical look at implementing DSA https://x.com/ZhihuFrontier/status/1993231992876421156

Gemini 3 is SOTA on even more benchmarks (math) 🤯 https://x.com/OfficialLoganK/status/1992004386990813598

A big practical weakness for working with Gemini 3 compared to ChatGPT-5.1 Thinking is that details in the thought/action traces are much less clear I can tell what ChatGPT-5.1 is doing and what tools it is using, I can’t with Gemini 3. Makes it hard to track and diagnose issues https://x.com/emollick/status/1993022071836717206

Google’s new Nano Banana Pro (Gemini 3 Pro Image) model is the new #1 Image Generation and Image Editing model in the Artificial Analysis Image Arena! Google’s Nano Banana Pro improves performance over Nano Banana but will not be a replacement for all users given its premium https://x.com/ArtificialAnlys/status/1993032471274024970

This AI paper just solved Google Earth’s biggest problem. Satellites look down. Humans look across. That perspective gap is why 3D maps are limited to cities you can blanket with aerial flyovers. Skyfall-GS bridges the gap by synthesizing the views we never captured – https://x.com/bilawalsidhu/status/1992051324096238068

Ilya on research taste: “One thing that guides me personally is an aesthetic of how AI should be by thinking about how people are. There’s no room for ugliness. It’s just beauty, simplicity, elegance, with correct inspiration from the brain. The more they are present, the more https://x.com/dwarkesh_sp/status/1993391989451014193

Ilya Sutskever – We’re moving from the age of scaling to the age of research https://www.dwarkesh.com/p/ilya-sutskever-2

Ok, so what Ilya saw was extreme benchmaxxing, which in turn prompted him to create his own company to do LLM development the proper way?! Makes sense, I sympathize with that.”” / X https://x.com/rasbt/status/1993379957570257298

🚀 New from Meta AI Research: Souper-Model! By smartly averaging multiple model weights using a method called SoCE (Soup Of Category Experts), the team achieved strong performance without retraining a whole new model. Read more here: https://x.com/MetaOpenSource/status/1993002595867136268

Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance | Research – AI at Meta https://ai.meta.com/research/publications/souper-model-how-simple-arithmetic-unlocks-state-of-the-art-llm-performance/

Robotics keeps getting better at seeing the world, but very few models can explain… HOW actions change it. [ 📍 Everything is open-sourced] Most benchmarks test passive perception. Almost none test interaction. That is why ENACT stands out. It asks a simple question with big https://x.com/IlirAliu_/status/1993755132131963275

A first preview of something we expect to see a lot more of soon:”” / X https://x.com/sama/status/1991597415888220596

I’m pleased to share the Second Key Update to the International AI Safety Report, which outlines how AI developers, researchers, and policymakers are approaching technical risk management for general-purpose AI systems. (1/5) https://x.com/Yoshua_Bengio/status/1993290185380184304

[2511.17803] Pillar-0: A New Frontier for Radiology Foundation Models https://arxiv.org/abs/2511.17803

[2511.18822] DiP: Taming Diffusion Models in Pixel Space https://arxiv.org/abs/2511.18822

[2511.18890] Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models https://arxiv.org/abs/2511.18890

[2511.21689] ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration https://arxiv.org/abs/2511.21689

2/ “Continuous Thought Machines” will be presented as a Spotlight at #NeurIPS2025🧠 Paper: https://x.com/SakanaAILabs/status/1992909033800716667

A tsunami of COGS – Better than Random https://betterthanrandom.substack.com/p/a-tsunami-of-cogs

Continuous Thought Machines | OpenReview https://openreview.net/forum?id=y0wDflmpLk

Data-free distillation of diffusion models is back! After BOOT ( https://x.com/sedielem/status/1993413445744836747

Despite only 760M active parameters, ZAYA1-base outperforms dense models such as Llama-3-8B and is competitive with Qwen3-4B and Gemma3-12B on mathematics and coding benchmarks. In high pass@k settings, the base model approaches the performance of specialized reasoning models. https://x.com/ZyphraAI/status/1993001723959689411

Diffusers users can easily switch between different attention backends for different trade-offs 🧨 We support FA3, FA2, and SAGE through `kernels` so that you can save lightyears without having to build them yourself! @_DhruvNair_, thanks for jamming 🚀 https://x.com/RisingSayak/status/1992921674585624873

DiP: Taming Diffusion Models in Pixel Space “”DiP decouples generation into a global and a local stage: a Diffusion Transformer (DiT) backbone operates on large patches for efficient global structure construction, while a co-trained lightweight Patch Detailer Head leverages https://x.com/iScienceLuvr/status/1993288136244576579

Direct link to the system card: https://x.com/sleepinyourhat/status/1993033332578464183

DSA vs NSA & DSA稀疏训练算子实现(TileLang) – 知乎 https://zhuanlan.zhihu.com/p/1975973344672245268

Exa AI Research Blog | Semantic Search & Neural Network Search Engine https://exa.ai/blog/exa-api-2-1

First make it fast, then make it smart | ʕ☞ᴥ ☜ʔ Kix’s blog https://kix.dev/first-make-it-fast-then-make-it-smart/

FP8 Reinforcement Learning | Unsloth Documentation https://docs.unsloth.ai/new/fp8-reinforcement-learning

FreeFlow: Flow Map Distillation Without Data https://data-free-flow-distill.github.io/

Fudan professor & Zhihu contributor @xpqiu 邱锡鹏 just introduced DiRL, a practical post-training framework for Diffusion Language Models, enabling an 8B diffusion LM to outperform 32B autoregressive models. 🚀 DLLMs promise stronger global reasoning and diversity, but the https://x.com/ZhihuFrontier/status/1992919281445855697

Haven’t read the full paper, which isn’t out yet, so can’t speak to details, but I am glad to see more methodological rigor being applied to LLM as a judge. LLM ratings are at the heart of a huge number of benchmarks & often used without clear statistical validation.”” / X https://x.com/emollick/status/1993515342526910901

How LLM Inference Works https://arpitbhayani.me/blogs/how-llm-inference-works

If scaling is not over & remains key, that would indicate that few, if any, new firms will be able to join the nine or so that can build state-of-the-art models. You just can’t catch up for the most part. And of those nine or so, only a few can actually scale in a serious way.”” / X https://x.com/emollick/status/1991348913299943721

INTELLECT-3: A 100B+ MoE trained with large-scale RL https://www.primeintellect.ai/blog/intellect-3

Introducing Stories – A New Way to Create, Play, and Share Adventures with Your Favorite Characters https://blog.character.ai/introducing-stories-a-new-way-to-create-play-and-share-adventures-with-your-favorite-characters/

Last time we shared Docker Model Runner + vLLM, many in the community showed strong interest — especially around how simple it makes local AI workflows. If you want a deeper look, don’t miss this upcoming session: 📅 Level Up Your Local AI Workflow with Model Runner 🔗”” / X https://x.com/vllm_project/status/1993328659005161510

LLMs can invent their own compression – Rajan Agarwal https://www.rajan.sh/llm-compression

Model Weight Preservation is not enough — LessWrong https://www.lesswrong.com/posts/fGCGJGCKMLbfquKiu/model-weight-preservation-is-not-enough

Multi-task RL can be highly sample-efficient and when done right, it unlocks LLM-style transfer and fine-tuning. We’re excited to introduce BRC, a simple recipe for multi-task RL that outperforms SOTA single-task agents while using less compute (!) https://x.com/mic_nau/status/1993005130095247704

My first published academic paper was on Moore’s Law and right now AI development looks similar: the exponential of Moore’s Law was not the result of a single technology, but rather many different technologies over many decades that were ready when one chip-making approach https://x.com/emollick/status/1993422227862438031

naumix/BiggerRegularizedCategorical https://github.com/naumix/BiggerRegularizedCategorical

Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models Two central architectural factors for designing SLMs: depth-width ratios and operator choices. although deep-thin models generally achieve better accuracy under the same parameter budget, they may not lie on https://x.com/iScienceLuvr/status/1993286325622214742

NEW: You bring the LoRA, we bring the @CoreWeave GPUs 🤝 Introducing Serverless LoRA Inference on W&B! Upload, version, and serve custom LoRAs instantly on the only @SemiAnalysis_ platinum grade AI cloud. This is totally not an AI generated infographic to help you get started. https://x.com/wandb/status/1993032159985385978

November 2025 Insiders (version 1.107) https://code.visualstudio.com/updates/v1_107

Our new paper Economies of Open Intelligence is out and covered by @Melissahei in the @FinancialTimes. It offers the clearest picture yet of how global power is shifting inside the open AI ecosystem, and what it means for the open-source community. https://x.com/frimelle/status/1993596653664977243

Pillar-0: A New Frontier for Radiology Foundation Models “”Here, we introduce Pillar-0, a radiology foundation model pretrained on 42,990 abdomen-pelvis CTs, 86,411 chest CTs, 14,348 head CTs, and 11,543 breast MRIs from a large academic center, together with RATE, a scalable https://x.com/iScienceLuvr/status/1993297733235438000

PixelDiT: Pixel Diffusion Transformers for Image Generation “”PixelDiT adopts a fully transformer-based architecture shaped by a dual-level design: a patch-level DiT that captures global semantics and a pixel-level DiT that refines texture details, enabling efficient training https://x.com/iScienceLuvr/status/1993632594093813999

Previewing Locus https://www.intology.ai/blog/previewing-locus

Product Evals in Three Simple Steps https://eugeneyan.com/writing/product-evals/

Project Iceberg – Coordinating the Human-AI Future https://iceberg.mit.edu/

Reasoning models are expensive. Not because the models are huge. It’s because they generate thousands of tokens just to think. But what if smaller models could learn to reason efficiently? This new paper compares training 12B models on reasoning traces from two frontier https://x.com/omarsar0/status/1993695515595444366

Researchers are uncovering the mysteries of why early-in-life acute infections can lead to neurodegenerative diseases in later years. https://x.com/StanfordMed/status/1991619773030101427

Tell all the truth but tell it slant– Success in Circuit lies Too bright for our infirm Delight The Truth’s superb surprise This paper finds poetry is a universal single shot jailbreak for LLMs. Systems built to stop prosaic attacks fail when the request is phrased in verse. https://x.com/emollick/status/1991624198855561508

Terminal Velocity Matching “”We propose Terminal Velocity Matching (TVM), a generalization of flow matching that enables high-fidelity one- and few-step generative modeling. TVM models the transition between any two diffusion timesteps and regularizes its behavior at its https://x.com/iScienceLuvr/status/1993631949957841214

The Bitter Lesson of LLM Extensions https://www.sawyerhood.com/blog/llm-extension

The Current State of the Theory that GPL Propagates to AI Models Trained on GPL Code – Open Source Guy https://shujisado.org/2025/11/27/gpl-propagates-to-ai-models-trained-on-gpl-code/

The Impossible Prompt https://teodordyakov.github.io/the-impossible-promt/

There is an interesting observation here which I think this paper misses. If one set of answers is 4x more verbose and training on them results in roughly the same performance then this implies something very interesting with information density is happening. Put another way you”” / X https://x.com/code_star/status/1993745248028164532

Things are looking smoothly exponential for AI over the past several years, and I continue to think this is the best default assumption (until the AI R&D automation feedback loop eventually speeds everything up) https://x.com/daniel_271828/status/1991407945964482575

tilelang-sparse-attention/examples/dsa_sparse_finetune at main · RUCKBReasoning/tilelang-sparse-attention https://github.com/RUCKBReasoning/tilelang-sparse-attention/tree/main/examples/dsa_sparse_finetune

Universal LLM Memory Does Not Exist https://fastpaca.com/blog/memory-isnt-one-thing

Unsloth RL is now 1.4x faster + has 2x longer contexts with FP8 LoRA GRPO vs BF16! We made inference time >96% of the entire RL run, so the majority is just vLLM inference! FP8 LoRA GRPO on Qwen, Llama, Mistral, Gemma & more can be easily enabled with 1 flag: `load_in_fp8=True`.”” / X https://x.com/danielhanchen/status/1993381372795535828

Vector search: fast, accurate, or affordable. Pick… all three? ✨ Most engineering teams are trapped in an expensive cycle: as their AI applications scale, they’re forced to choose between performance and budget. More data means bigger infrastructure bills, slower searches, https://x.com/weaviate_io/status/1992986708766323004

Why are vLLM and transformers so damn fast? ⚡ Continuous batching. That’s the secret sauce 🔥 Never heard of it? We just dropped a blog post building it up from first principles 🤗 See what happens inside the minds of the engineers pushing inference to the edge 🧠 https://x.com/remi_or_/status/1993324801881260112

Using skills with Deep Agents CLI – YouTube https://www.youtube.com/watch?v=Yl_mdp2IiW4

Hugely under-researched area: how effective is using AI to check the work of other AIs? Does using different models help? If so, that is an important & easy way to reduce errors One paper found this technique to be effective but as far as I can tell has never been followed up on https://x.com/emollick/status/1991328030694703115

Today we are shipping dnet, a distributed inference framework that lets Apple Silicon clusters run models that exceed their physical memory. We fuse pipelined-ring parallelism, disk streaming and UMA-aware scheduling so “out of memory” stops being the limit. https://x.com/driaforall/status/1993729375745749339

One of the very confusing things about the models right now: how to reconcile the fact that they are doing so well on evals. And you look at the evals and you go, ‘Those are pretty hard evals.’ But the economic impact seems to be dramatically behind. There is [a possible] https://x.com/dwarkesh_sp/status/1993450075616690474

Continuous batching from first principles https://huggingface.co/blog/continuous_batching

Remember all the line of works mega-focused on using activations of imagenet-pretrained model to converge better on imagenet? Remember when I told you there is no way thats going to scale beyond imagenet / other domains? Im glad BFL shared this result with the world https://x.com/cloneofsimo/status/1993371224140054873

Emerging AI capabilities are at one level quite predictable… First came IQ (factuality). Then EQ (personality). Now AQ (actions quotient or agents). The next big frontier is SQ (social intelligence).”” / X https://x.com/mustafasuleyman/status/1991888331278270976

This model is called Z-image, Apache2.0 licensed. 6B size, Coming soon.🫡 https://x.com/bdsqlsz/status/1993545608179990544

Most datasets for hand object manipulation are slow to build, expensive to capture, or too small to be useful… not this: HO Cap feels different. A simple idea done well. They built an 8 camera RealSense setup with an Azure Kinect on top. Users wear a HoloLens so they also get https://x.com/IlirAliu_/status/1993610121759326223

STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows “”In this work, we revisit this design space by presenting STARFlow-V, a normalizing flow-based video generator with substantial benefits such as end-to-end learning, robust causal prediction, and native https://x.com/iScienceLuvr/status/1993629956375822508

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading