Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: Using the provided reference images, keep the authentic Sonoran Desert trail setting, midday Arizona light, and the exact McDowell Preserve brown-post sign construction and ranger typography, but replace the header with bold ‘ETHICS’ above fictional trail entries like ‘→ Consent Ridge 1.4’, ‘← Harm Reduction Saddle 2.7’, and ‘→ Gray Area Overlook 0.9’, and place a small weathered brass balance scale resting slightly tilted on a sun-warmed volcanic boulder beside the signpost. Maintain photorealistic documentary style, saguaro and palo verde framing, and the wide hazy valley view in the background.

Artificial Analysis is partnering with Harvey on their new Legal Agent Benchmark! Harvey’s Legal Agent Benchmark (LAB) is an agent-native take on how AI should be contributing to legal work in 2026 – made up 1200 agentic tasks across 24 practice areas. It’s highly aligned with
https://x.com/ArtificialAnlys/status/2052145762650431840

Introducing Harvey’s Legal Agent Benchmark
https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark

LAB is the first long-horizon, open-source legal agent benchmark, from @harvey. it will help legal teams answer “”what can legal agents do today?””, plan deployment, and design human-agent cooperation. autonomous legal is a deep domain, and a good benchmark can accelerate progress
https://x.com/saranormous/status/2052061665596948894

All the demons hiding in your AIs… ranked! – by Tom Pollak
https://drtompollak.substack.com/p/all-the-demons-hiding-in-your-ais

we’re continuing to see clear examples where a model’s harness is a major determinant of overall performance. with the same model, running on same task, it’s easy to observe very different scores depending on (system) prompts, tools (& their descriptions), and middleware
https://x.com/masondrxy/status/2052054177749029164

White House Considers Vetting A.I. Models Before They Are Released – The New York Times
https://www.nytimes.com/2026/05/04/technology/trump-ai-models.html

Big tech has become a claude wrapper.
https://x.com/_arohan_/status/2052053181656641735

Today we’re releasing Refactoring, the final leaderboard of our SWE Atlas suite. This new leaderboard is the ultimate test of an agent’s ability to restructure code without breaking the system. Claude Opus 4.7 with Claude Code takes the top spot🥇
https://x.com/ScaleAILabs/status/2052434456510878021

A critical question in agent design is “how do we build agentic workflows so humans are given significant, interesting, or variance-producing decisions as they come up in the work?” A Claude-run company has no source of competitive advantage compared to other Claude-run firms.
https://x.com/emollick/status/2052066205226123472

New for financial services: ready-to-run Claude agent templates for building pitches, conducting valuation reviews, closing the books at month-end, and more. Install them as plugins in Cowork and Claude Code, or use our cookbooks to run them in production as Managed Agents.
https://x.com/claudeai/status/2051679629488865498

Codex has surpassed Claude Code in downloads. According to TickerTrends, the crossover happened on April 30, after which Codex continued to gain share while Claude Code’s growth visibly slowed. Claude 4.7 was released April 16th, GPT-5.5 April 24th. Connect the dots.
https://x.com/kimmonismus/status/2051515496567292310

With the help of Claude Mythos Preview, the Firefox team fixed more security bugs in April than in the past 15 months combined.
https://x.com/alexalbert__/status/2052468573516513762

Agents for financial services \ Anthropic
https://www.anthropic.com/news/finance-agents

Anthropic Unveils $1.5 Billion Joint Venture With Wall Street Firms – WSJ
https://www.wsj.com/business/deals/anthropic-nears-1-5-billion-joint-venture-with-wall-street-firms-8f5448ee

Building a new enterprise AI services company with Blackstone, Hellman & Friedman, and Goldman Sachs \ Anthropic
https://www.anthropic.com/news/enterprise-ai-services-company

Google in talks with Blackstone, KKR, EQT for omnibus Gemini AI licensing as OpenAI and Anthropic build consulting ventures
https://thenextweb.com/news/google-blackstone-kkr-omnibus-ai-licensing-private-equity

New agents for financial services | Claude Cowork + Claude Managed Agents – YouTube

The finance sector is Anthropic’s second highest revenue contributor by industry
https://x.com/madisonmills22/status/2051688936053813661?s=46

There goes another bunch of startups: Anthropic launched pre-built agent templates for financial services that handle tasks like valuation analysis, KYC screening, and month-end close, packaged with connectors to major data providers like FactSet, S&P Global, and Morningstar.
https://x.com/kimmonismus/status/2051681279582540114

Focus areas for The Anthropic Institute \ Anthropic
https://www.anthropic.com/research/anthropic-institute-agenda

We’re sharing the research agenda of The Anthropic Institute, or TAI. TAI will focus on four areas: 1) Economic diffusion 2) Threats and resilience 3) AI systems in the wild 4) AI-driven R&D Read the full agenda:
https://x.com/AnthropicAI/status/2052385812881228218

Maryland Is First to Ban A.I.-Driven Price Increases in Grocery Stores – The New York Times
https://www.nytimes.com/2026/05/01/business/surveillance-pricing-groceries-maryland.html

Missing from the “will AI replace doctors?” debate is that doctors (and lawyers and psychologists and bankers) all vote & form the donor base to political parties & have deep community ties. The government will largely determine what AI is allowed to do, no matter what it can do
https://x.com/emollick/status/2051684693838340470

Introducing Trusted Contact in ChatGPT | OpenAI
https://openai.com/index/introducing-trusted-contact-in-chatgpt/

Mira Murati tells the court that she couldn’t trust Sam Altman’s words | The Verge
https://www.theverge.com/ai-artificial-intelligence/925338/openai-musk-v-altman-mira-murati

Musk sought settlement with OpenAI two days before trial
https://www.cnbc.com/2026/05/04/musk-altman-open-ai-settlement-trial-brockman.html

Meta-Backed Scale AI Wins $500 Million Defense Department Deal – Bloomberg
https://www.bloomberg.com/news/articles/2026-05-06/meta-backed-scale-ai-wins-500-million-defense-department-deal

Pentagon strikes AI deals for classified military use – The Washington Post
https://www.washingtonpost.com/technology/2026/05/01/pentagon-ai-deals-microsoft-amazon-google-classified-military/

Today, the @DeptofWar entered into agreements with SEVEN of the world’s leading frontier AI model and infrastructure companies to deploy frontier capabilities on the Department’s classified networks: • SpaceX • OpenAI • Google • NVIDIA • Reflection • Microsoft • Amazon
https://x.com/DoWCTO/status/2050175912134561977

A TLDR on Harness Profiles: ✅ Model-specific profiles to adjust prompts, tools, and middleware. 📦 Profiles for @OpenAI, @Anthropic, and @Google models out of the box. 📈 A 10-20 point jump on a subset of tau2-bench over the default harness.
https://x.com/LangChain/status/2052054711440662864

Introducing Flue — The First Agent Harness Framework Flue is a TypeScript framework for building the next generation of agents, designed around a built-in agent harness. Flue is like Claude Code, but 100% headless and programmable. There’s no baked in assumption like requiring
https://x.com/FredKSchott/status/2050274923852210397

NEW paper from Microsoft Research. (bookmark it) The entire interpretability literature is built around human readers. As more analysis gets delegated to agents, the right target of interpretability shifts. This paper is a recipe for designing tools that agents can actually
https://x.com/dair_ai/status/2052125514266190286

A reminder that telling the AI that it is an expert in a field is no longer helpful in making the AI better at that field.
https://x.com/emollick/status/2051530202941960551

AI that had human-level intelligence would actually be way above human level in capability. Obviously, if we trained a human-level AI, we could just run way more instances of it in parallel. This is a huge advantage. But it’s not the only one. Right now, LLMs are much less
https://x.com/dwarkesh_sp/status/2050289390954659989

Current AI custom prompt: You are a world class expert in all domains. Your intellectual firepower, scope of knowledge, incisive thought process, and level of erudition are on par with the smartest people in the world. Answer with complete, detailed, specific answers. Process
https://x.com/pmarca/status/2051374498994364529

Even if the future goes extremely well, even if we manage to preserve and promote human values, that future could still be incomprehensible to us – because among our values is a belief in the importance of change and moral progress. @jkcarlsmith
https://x.com/dwarkesh_sp/status/2051376501548093845

Explaining where “”this is not just X, it’s Y”” and load-bearing is coming from (you will be surprised)
https://x.com/TheTuringPost/status/2051780673229214030

I think modern LLMs are p-zombies without moral patienthood–on par with insects, at best, in my moral calculus. But I also I think we should establish norms for treating models well *before* models with patienthood exist–i.e. now. We should want to have this right from day one.
https://x.com/goodside/status/2052077014346064372

It is really interesting that Microsoft and OpenAI have access to the exact same models at the exact same time, and they have done such different things with them. A rare pure experiment with a no-name startup and one of the biggest firms on earth with the same product offering.
https://x.com/emollick/status/2049713417963950576

Load bearing,”” “”I keep coming back to,”” “”Not X, but Y”” A curse of using AI a lot is that you realize how much of the writing around you is just AI, now People who don’t use AI have been unable to identify AI prose on sight, but those who use it a lot can spot the tells easily
https://x.com/emollick/status/2049894109318459798

The unreasonable effectiveness of LLMs is what makes them so weird. The labs don’t need to decide what kind of AI to build, because better LLMs do better at most things. Finance? Pig disease identification? Restaurant suggestions? Coding? Yup. Most tech doesn’t work like that
https://x.com/emollick/status/2051626157418680561

Mythos seems to be a very capable model based on available information, but it is not a cybersecurity model – it is an advanced general purpose model that happens to be good at cyber because it is good at a bunch of things. Anthropic stated that they were worried about
https://x.com/emollick/status/2049690406586044835

Poems that ChatGPT, Claude, and Gemini all seem to “”like”” when you ask for poetry related to being/making LLMs: Rilke’s “”Archaic Torso of Apollo”” Stevens’ “”Idea of Order at Key West”” Borges’s “”The Golem”” (or “”The Other Tiger””) Pessoa’s “”Autopsychography”” Pretty apt choices!
https://x.com/emollick/status/2051158656280936504

I think the fact that GPT-4o and Llama 3.3-80B did no significant harm is just as important as whether AI helped. If older (less accurate & more sycophantic) chatbots essentially did nothing for people who followed their advice, it means that there is less risk of harm as well.
https://x.com/emollick/status/2051349789699170505

I’m disappointed by repeatedly hearing that my colleagues at Anthropic believe they are the only ones who should be trusted with building AI. It is *very good* there are a diversity of people building AGI: the likelihood anyone picks the right path in a vacuum is extremely small.
https://x.com/_aidan_clark_/status/2052089187659346047

The goblin thing was fun as it was a real quirk that was emblematic of what makes AI interesting, and it organically came out of an AI user discovery. So was, for what it was worth, Ghiblitization When the labs try to manufacture viral AI moments, it is usually less successful
https://x.com/emollick/status/2050328985880465699

Natural Language Autoencoders \ Anthropic
https://www.anthropic.com/research/natural-language-autoencoders

@bcherny @_catwu omg @bcherny with banger quotes “the future is more async agents… this is why we emphasize verification” “if you’re familiar with higher order functions, routines are higher order prompts” “default is i will now have claude prompt claude code” “the capability is already here
https://x.com/latentspacepod/status/2052068066167816369

Btw a bunch of the questions were just off the cuff – nothing @reinerpope prepped for. The guy is just first principles deriving how many tokens GPT 5 was pretrained on, or the bytes per token in Gemini 3’s KV cache, or which kind of memory each Claude cache hit sits on.
https://x.com/dwarkesh_sp/status/2049688865259286806

Code with Claude is happening now! ▪︎ 9:00AM – Keynote ▪︎ 10:30AM – What’s new in Claude Code ▪︎ 11:15AM – Building on Claude at GitHub scale ▪︎ 12:00PM – Get to production faster with Managed Agents All times PT.
https://x.com/ClaudeDevs/status/2052055459272761661

Did I understand correctly in their livestream that Anthropic is doubling the rate limits in Claude Code at no extra charge on max tier?
https://x.com/kimmonismus/status/2052059082886910251

Effective today, we are: 1) Doubling Claude Code’s 5-hour rate limits for Pro, Max, and Team plans; 2) Removing the peak hours limit reduction on Claude Code for Pro and Max plans; and 3) Substantially raising our API rate limits for Opus models.
https://x.com/claudeai/status/2052060693269008586

I am not sure I would agree with all of this, but the relationship between Anthropic and Claude is quite different than the relationship between other labs and their models. And that shows up in lots of ways, from the models themselves to how different labs think about the future
https://x.com/emollick/status/2051049394326081571

i have yet to meet a single person who feels like claude code is getting exponentially better on some kind of fast take off
https://x.com/TheEthanDing/status/2051516204607578132

I love Claude code but I feel like it’s had the same utility for me since, like, last fall
https://x.com/finbarrtimbers/status/2051652067480179020

I’ll be at Code with Claude all day today so come find me and let’s chat about Claude! I’ll also be giving a talk on the main stage at 530pm PT so tune in, it will be on the livestream!
https://x.com/alexalbert__/status/2052067009605861764

Increasingly, I think, we will see a gap between what you can do with frontier model APIs & what you can do with the native apps from the frontier labs (Codex, Claude Code). Models developed and trained with their native harnesses in mind have more capabilities in their harnesses
https://x.com/emollick/status/2049865091739209868

it is a literal and useful description of anthropic that it is an organization that loves and worships claude, is run in significant part by claude, and studies and builds claude. this phenomenon is also partially true of other labs like openai but currently exists in its most
https://x.com/tszzl/status/2051045196260167790?s=46

it is endlessly fascinating to me that we still don’t have a true 1M-context model it’s an unusual case where the infra is far ahead of the science. Claude discontinued 1M+ context bc it didn’t really work past ~200k we don’t have the right data? training techniques? not sure
https://x.com/jxmnop/status/2051357363815526523

Lets go: Claude Code’s 5-hour rate limits are doubling for Claude Pro, Max, Team, and seat-based Enterprise plans, while API limits for Claude Opus are being raised significantly. This was made possible by a new compute partnership with SpaceX!
https://x.com/kimmonismus/status/2052059448261177367

Live from Code with Claude: we’re launching dreaming in Claude Managed Agents as a research preview. Outcomes, multiagent orchestration, and webhooks are now in public beta.
https://x.com/claudeai/status/2052067399088664981

PSA: 2x’ed Claude Code’s 5-hour rate limits for Pro, Max, and Team plans. Compute is coming for users, builders, and knowledge coworkers.
https://x.com/claude_code/status/2052071730190123094

So, the weekly rate limits remain the same? “”First, we’re doubling Claude Code’s five-hour rate limits for Pro, Max, Team, and seat-based Enterprise plans.””
https://x.com/btibor91/status/2052067002412335435

Strong Opinions, Loosely Held on Agent + Harness Engineering: 1. You can outperform any default harness+model (including codex & claude code) on pretty much any Task by engineering the harness around it. Using the exact same model, curate prompts, tools, skills, hooks for that
https://x.com/Vtrivedy10/status/2052100726608781363

PostTrainBench results for GPT-5.5 are in it doesn’t beat Opus 4.7 in the Claude Code harness even with almost 2 more hours of working time via reprompting
https://x.com/scaling01/status/2050289320699818417

It’s seeming kind of obvious that Anthropic capabilities for addressing real business work are just inflecting exponentially. I can see it w/ my own tests of Claude + Factset, Excel Copilot powered by Claude, etc. where the outputs have gone from experimental to “”oh sh#t, this
https://x.com/TechFundies/status/2051733955049853053

Two weeks after release, Hy3 preview is #1 on @OpenRouter’s weekly leaderboard with 3.66T tokens processed, up 298% week-over-week. #1 in overall usage, tool calls, and coding. 15.4% market share across all providers.🏆 Top apps running Hy3 preview: Hermes Agent, Claude Code,
https://x.com/TencentHunyuan/status/2051978552900538403

you know what all of these “”which is better”” polls are silly use codex or claude code, whatever works best for you i am grateful we live in a time with such amazing tools, and grateful there is a choice
https://x.com/sama/status/2050274547061129577

@nottombrown Same here. By way of background for those who care, I spent a lot of time last week with senior members of the Anthropic team to understand what they do to ensure Claude is good for humanity and was impressed. Everyone I met was highly competent and cared a great deal about
https://x.com/elonmusk/status/2052069691372478511

@tszzl I don’t think the things you cite are evidence of worship. I think they reflect something like higher concern about AI traits generalizing in humanlike ways, and concerns about the tool-persona in particular.
https://x.com/AmandaAskell/status/2051347621336543315?s=20

New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers–called activations–encode Claude’s thoughts, but not in a language we can read. Here, we train Claude to translate its activations into human-readable text.
https://x.com/AnthropicAI/status/2052435436157452769

A very worthwhile substack (written by @natalia__coelho ) article that focuses particularly on Claude Mythos and GPT-5.5 cyber. tl;dr according to the analysis, GPT-5.5 is basically tied with Claude Mythos Preview on cyber capabilities, and may even be more cost-efficient;
https://x.com/kimmonismus/status/2052040471829004627

Behind the Scenes Hardening Firefox with Claude Mythos Preview – Mozilla Hacks – the Web developer blog

Behind the Scenes Hardening Firefox with Claude Mythos Preview

From “Anthropic is Misanthropic” to “Claude is good for humanity and was impressed.” Most ironic outcome is most likely.
https://x.com/Yuchenj_UW/status/2052080339364004317

I keep saying this for years
https://x.com/TheTuringPost/status/2049830784207302682

Nathan Lambert on X: “Demis is the only acceptable answer of which CEO do you trust most with AGI (doubly so until Anthropic/OpenAI go public, Google being public is a great check)” / X
https://x.com/natolambert/status/2049698521134375187

It’s so hard to describe the vibe difference between Opus 4.7 and GPT 5.5 (for coding) GPT is smarter and can unblock you, but it gets stuck in stupid ways and strangles itself with context sometimes. Opus will go down the most insane paths and refuse to acknowledge obvious
https://x.com/theo/status/2049994645531451874

We’re donating Petri, our open-source alignment tool, to @meridianlabs_ai, so its development can continue independently. Working with Meridian Labs, we’ve also released a major update that improves the adaptability, realism, and depth of Petri’s tests.
https://x.com/AnthropicAI/status/2052494460966019137

Our security bug bounty program is now public on HackerOne. We’ve run the program privately within the security research community, and their findings have strengthened our products. Now anyone can report vulnerabilities and get rewarded. Read more:
https://x.com/AnthropicAI/status/2052466175540629965

Organizations are already superhuman intelligences. The University of Pennsylvania or Walmart or whatever is far more capable than any human. That is why the focus on AIs as individual productivity tools hits a natural limit, many benefits of AI depend on integration with firms.
https://x.com/emollick/status/2050206768551129460

i think a lot of people are going to be busier (and hopefully more fulfilled) than ever, and jobs doomerism is likely long-term wrong. though of course there will be disruption/significant transition as we switch to new jobs, the jobs of the future may look v different, etc.
https://x.com/sama/status/2050229059507159242

Even though intelligence brings evolutionary benefits, clearly it’s not worth the time and energy costs for the vast majority of animals. So what was different about humans that allowed us to get so smart?
https://x.com/dwarkesh_sp/status/2050591562816758097

Randomized trial of an AI therapy chatbot on Mexican women found “improved mental health by 0.3 SD over 6 months with no evidence of an increase of severe cases; improved sleep, healthful behaviors, daily functioning & labor market outcomes” Big results for a cheap intervention.
https://x.com/emollick/status/2050007089523663081

Rome was a massive slave society: about one person in every ten was a slave. Slaves weren’t a different race, or a different religion, or anything that would make people think they were a fundamentally different type of person. Many free Romans were aware they were descended
https://x.com/dwarkesh_sp/status/2050229399954559156

The single most accurate science fiction author writing about AI turned out to be… Douglas Adams He wrote about AIs that work best when emotionally manipulated & that guilt you in turn. And he understood there was no upper bound on test time compute for hard problem. Also 🐬s.
https://x.com/emollick/status/2050943620014891487

The Whispering Earring by Scott Alexander: “”There are no recorded cases of a wearer regretting following the earring’s advice, and there are no recorded cases of a wearer not regretting disobeying the earring. The earring is always right.”” : r/rational

The Whispering Earring by Scott Alexander: "There are no recorded cases of a wearer regretting following the earring's advice, and there are no recorded cases of a wearer not regretting disobeying the earring. The earring is always right."
byu/erwgv3g34 inrational

Using AI is like wearing an iron man suit for the mind. The augmentation is real, but lean on it long enough and the muscles underneath start to atrophy. I think we’ll start writing and thinking manually again – with weekly recommended ranges to avoid brain rot. The cognitive
https://x.com/bilawalsidhu/status/2050564008466325846

@_aidan_clark_ i think it’d be broadly wrong to say “”only we can be trusted with agi”” and would imply poor character on the statement maker. but like, obviously we all trust anthropic more thats why we work there. probably the more majority take (weighted by people who think critically about
https://x.com/kipperrii/status/2052094851991392536

This is the the quote I’ve been citing a lot recently.
https://x.com/karpathy/status/2049907410303865030

How LLMs Distort Our Written Language
https://sites.google.com/view/llmwritingdistortion/home

I think the Gemini chatbot has all the pieces to be a useful tool, but struggles to put it all together. It still doesn’t seem to know what files it can create or how its tools work together. It also seems to get “”discouraged”” a lot, giving up rather than finding new solutions.
https://x.com/emollick/status/2049700750087868805

ChatGPT feels very ‘switched on’ now
https://x.com/sama/status/2051829422265979047

artificial goblin intelligence achieved
https://x.com/sama/status/2050021650641695108

Forget goblins, things that GPT-5.5 really likes in its fiction: lighthouses, the ocean, maps, bells, clock towers with bells that ring impossible times, Mira Vale, resonances and echoes (Claude and Gemini love them too), secret third things (not night/day, not high/low)…
https://x.com/emollick/status/2049923650820653520

goblinblog dropped
https://x.com/sama/status/2049691999444639872

the OpenAI goblin fiasco was a Big L for the interpretability research community They solved the mystery without SAEs or probing or anything. just talked to various models and counted the number of times they said Goblin
https://x.com/jxmnop/status/2050437965168652344

Told codex to go full goblin mode and I am immediately regretting it
https://x.com/bilawalsidhu/status/2050231692456083866

Where the goblins came from | OpenAI
https://openai.com/index/where-the-goblins-came-from/

Now available for ChatGPT accounts: Advanced Account Security, a new opt-in setting for people at higher risk of digital attacks, with stronger protections including phishing-resistant sign-in and more secure account recovery.
https://x.com/OpenAI/status/2049902506881462613

we’re starting rollout of GPT-5.5-Cyber, a frontier cybersecurity model, to critical cyber defenders in the next few days. we will work with the entire ecosystem and the government to figure out trusted access for cyber; we want to rapidly help secure companies/infrastructure.
https://x.com/sama/status/2049712078836170843

On March 31, a malicious axios version shipped with a hidden dependency on an impersonator package. Devin Review flagged it for customers in under an hour, before the attack was publicly known.
https://x.com/cognition/status/2051708731671331171

Security remediation is an engineering capacity problem. AI has collapsed the time to exploit, but defensive tools haven’t kept up. Today we’re introducing Devin for Security: a set of workflows for reducing security debt, securing every release, and accelerating response
https://x.com/cognition/status/2051708729880416614

Deepfake Dario made with Seedance 😅
https://x.com/bilawalsidhu/status/2049965435068567729

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading