a forest entrance with a technology inspired trail sign that reads “Anthropic” –ar 5:3 –style raw
This week’s category cover theme is a sign in a forest. Each category image prompt is a derivative of the formula “an [category themed object] in a forest with a trail sign that reads “[category name]”. Using a theme each week takes the cover creation time down to about 20 minutes, rather than several hours.
“Life update: After ~2 years at Anthropic, I joined OpenAI! This wasn’t the easiest decision and I’m very grateful to everyone who is supporting me through this transition, especially John, Barret, Boris, Mira, and Sam. I joined Anthropic as the first designer / front-end” / X
“If this work interests you, our interpretability team is hiring. We’re looking for: • Managers:
Reflections on our Responsible Scaling Policy \ Anthropic
“We published our Responsible Scaling Policy last year. As we continue to iterate on our empirically grounded framework, we’re gaining valuable insights. Today we share reflections on our progress:
Responsible Scaling Policy Evaluations Report – Claude 3 Opus
Anthropic publishes incredibly transparent research on the inner workings of Claude 3 Sonnet
Mapping the Mind of a Large Language Model \ Anthropic
“New Anthropic research paper: Scaling Monosemanticity. The first ever detailed look inside a leading large language model. Read the blog post here:
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
New Anthropic Research Sheds Light on AI’s ‘Black Box’
Anthropic’s “LLM Genome Project”: learning & clamping 34m features on Claude Sonnet • Buttondown
“Among these millions of features, we find several that are relevant to questions of model safety and reliability. These include features related to code vulnerabilities, deception, bias, sycophancy, power-seeking, and criminal activity.
“This week, we showed how altering internal “features” in our AI, Claude, could change its behavior. We found a feature that can make Claude focus intensely on the Golden Gate Bridge. Now, for a limited time, you can chat with Golden Gate Claude:
“The problem: most LLM neurons are uninterpretable, stopping us from mechanistically understanding the models. In October, we showed that dictionary learning could decompose a small model into “monosemantic” components we call “features”—making the model more interpretable.” / X
“These features are remarkably abstract, often representing the same concept across contexts and languages, even generalizing to image inputs. Importantly, they also causally influence the model’s outputs in intuitive ways.” / X
“This work is preliminary. Whereas we show that there are many features that seem *plausibly* relevant to safety applications, much more work is needed to establish that our approach is useful in practice.” / X
“Our goal is to let people see the impact our interpretability work can have. The fact that we can find and alter these features within Claude makes us more confident that we’re beginning to understand how large language models really work. Read more:
“Our new interpretability paper offers the first ever detailed look inside a frontier LLM and has amazing stories. I want to share two of them that have stuck with me ever since I read it. For background, the paper shows our latest work on interpreting the “features” of Claude 3
“Today, we announced that we’ve gotten dictionary learning working on Sonnet, extracting millions of features from one of the best models in the world. This is the first time this has been successfully done on a frontier model. I wanted to share some highlights 🧵
Anthropic’s Alex Albert answers questions about Claude
“”We put the model in the oven and then we wait to see what pops out” I asked @alexalbert__ of @AnthropicAI why Claude is the best writer of all current LLMs There’s a lot more to it, of course, but I appreciate the candor! Great new episode out today 🔗👇
“”We’re honest with Claude about what we know and don’t know. We recognize that these questions are very tricky and there’s no perfect answer – and we tell Claude that.” I asked @AnthropicAI’s @alexalbert__ why Claude-3 claims subjective experience, and I loved his answer 🔗👇
“Many people have asked me why we “chose” to let Claude to speculate on tricky philosophical questions. It isn’t so much that we purposefully made a decision one way or the other. We are just honest with Claude about what we know and what we don’t. https://twitter.com/alexalbert__/status/1793683229595341182

Heads up! You’ve scrolled to the end of this category. There may have been just one or two links (above), so go back up and double check to be sure you didn’t quickly scroll down past it.
Be Sure To Read This Week’s Main Post:
This week’s executive overview and top links are here:
AI News #34: Week Ending 05/24/2024 with Executive Summary and Top 47 Links
The post you just read is an deep dive extension of my weekly newsletter, This Week In AI, an executive summary of the top things to know in AI. Each week, I create an accessible overview for laypeople to feel confident they are conversant with the week’s AI developments. I include a curated list of must-click links of the week, to offer everyone a hands-on opportunity to explore the most intriguing updates in artificial intelligence across various categories, including robotics, imagery, video, AR/VR, science, ethics, and more. Beyond the overview, I post these topic-based deeper dives (below). If you haven’t read this week’s overview, I recommend starting there.
- Agents/Copilots
- Amazon
- Apple
- Artificial General Intelligence (AGI)
- Augmented and Virtual Reality (AR/VR)
- Autonomous Vehicles
- AI Audio
- Business and Enterprise AI
- Chips and Hardware
- Consumer Products
- Education
- Ethics/Legal Security
- Images/Photos
- International AI News
- Locally Run AI Models
- Mobile
- Meta
- Microsoft
- OpenAI
- Open Source
- Podcasts/YouTube
- Publishing and News
- Retrieval-Augmented Generation (RAG) News
- Robots and Embodiment
- Science and Medicine
- Video
- Vision/Multimodality
- X/Twitter/Grok
- Tech and Development
Credits/Sources

Most of these weekly links come from just a few prolific oversharing sources. Please follow them, as they work hard to find the news each week and they make it a lot easier for me to compile.
- Robert Scoble: https://x.com/Scobleizer
- Ethan Mollick: https://www.linkedin.com/in/emollick/
- Alan Thompson: https://lifearchitect.ai/
- Theoretically Media: https://www.youtube.com/@TheoreticallyMedia
- The Rundown: https://www.therundown.ai/
- Bilawal Sidhu: https://twitter.com/bilawalsidhu/
- TLDR: https://tldr.tech/ai
- Jeremiah Owyang: https://twitter.com/jowyang
- Nick St. Pierre: https://twitter.com/nickfloats
- Dr. Jim Fan: https://twitter.com/DrJimFan
- All About AI: https://www.youtube.com/@AllAboutAI
- Marshall Kirkpatrick: https://aitimetoimpact.com/
- AI News (Smol Talk): https://buttondown.email/ainews/archive/
For previous issues, please visit the archives!

Thanks for reading!





Leave a Reply