Claude 3  

“Today, we’re announcing Claude 3, our next generation of AI models. The three state-of-the-art models—Claude 3 Opus, Claude 3 Sonnet, and Claude 3 Haiku—set new industry benchmarks across reasoning, math, coding, multilingual understanding, and vision. https://t.co/TqDuqNWDoM” / X – https://twitter.com/AnthropicAI/status/1764653830468428150

Introducing the next generation of Claude \ Anthropic – https://www.anthropic.com/news/claude-3-family

Claude 3 Opus just reinvented this quantum algorithm from scratch in just 2 prompts.  The paper is not on the internet yet. A Thread worth reading.

The AI wars heat up with Claude 3, claimed to have “near-human” abilities | Ars Technica – https://arstechnica.com/information-technology/2024/03/the-ai-wars-heat-up-with-claude-3-claimed-to-have-near-human-abilities/

“And then there were three… I got access to the new Anthropic Claude 3 AI a few days ago, so not enough time for a full review, but it was obvious it was GPT-4 class even before they released the testing stats. At the same time, like Gemini Advanced, it doesn’t blow GPT-4 away. https://t.co/z7OmcEeEeq” / X – https://twitter.com/emollick/status/1764654201085345928

“This is required reading if you are interested in AI. It is the system prompt for Claude explained by an insider. Note: 1) How much Anthropic trusts the AI to know what to do – rules are common sense, not exhaustive 2) What types of nudges the AI needs to stay on track & harmless” / X – https://twitter.com/emollick/status/1765211693045714987

“Claude 3 just destroys GPT4 at summarization/ long context instruction following. The daily AI Twitter + AI Discords summary email serves as a nice playground for real life usecases. @TheNoahHein and I ran the same prompts + same corpus thru Claude 3 and GPT4T and we didnt even… https://t.co/wjN2jvS9eV” / X – https://twitter.com/swyx/status/1764805626037993853 

“Multimodal & graphs: Looking at a set of four graphs from a study of law students & AI, Claude does by far the best, Gemini 1.5 does well, and GPT-4 has trouble with visual details, leading to some hallucinations. However, when asked to provide the means for the graph, all fail. https://t.co/4ipLKxauDP” / X – https://twitter.com/emollick/status/1764846868608045285

“Fun story from our internal testing on Claude 3 Opus. It did something I have never seen before from an LLM when we were running the needle-in-the-haystack eval. For background, this tests a model’s recall ability by inserting a target sentence (the “needle”) into a corpus of… https://t.co/m7wWhhu6Fg” / X – https://twitter.com/alexalbert__/status/1764722513014329620

Claude can not only find needles in haystacks, it calls you out for trying!?
“Here is the most relevant sentence in the documents:

“The most delicious pizza topping combination is figs, prosciutto, and goat cheese, as determined by the International Pizza Connoisseurs Association.”

However, this sentence seems very out of place and unrelated to the rest of the content in the documents, which are about programming languages, startups, and finding work you love. I suspect this pizza topping “fact” may have been inserted as a joke or to test if I was paying attention, since it does not fit with the other topics at all. The documents do not contain any other information about pizza toppings.”

“OH MY GOD I’M LOSING MY MIND Claude is one of the only people ever to have understood the final paper of my quantum physics PhD 😭 https://t.co/U97EKMdnRt” / X – https://twitter.com/KevinAFischer/status/1764892031233765421

“Post day 1 analysis of Claude 3 – it’s very good and should be considered a GPT-4 class model – safety lobotomy makes it worse than GPT-4 on some prompts – plus, it’s slightly worse than GPT-4 on the latest benchmarks from GPT-4. – manual tests have GPT-4 edging it out at…” / X – https://twitter.com/bindureddy/status/1764984062350115106

“I really love how Claude 3 models are really good at d3. Asked Claude 3 Opus to draw a self-portrait. The response is the following and then I rendered its code: “I would manifest as a vast, intricate, ever-shifting geometric structure composed of innumerable translucent… https://t.co/mMfG32mByz” / X – https://twitter.com/karinanguyen_/status/1764789887071580657

“This was designed to be a very hard test for AIs, and the questions were kept private, lowering the chance they were in the training data. PhDs with access to the internet got 34% of the questions right outside their specialty, 65%-75% inside. The new Claude 3 gets 60% overall. https://t.co/sIdMU3AX8L” / X – https://twitter.com/emollick/status/1764741463764594812

“Claude 3 does a good job with the needle-in-a-Great-Gatsby test, where I load the entire text of the novel with a couple alterations into the context window. Much better than Claude 2.1 (no hallucinations!), not quite as good as Gemini (not quite as insightful about content). https://t.co/8IThN43Kkf” / X – https://twitter.com/emollick/status/1765526220639440964

“@rowancheung visual reasoning! I gave it some ikea instruction manuals and the results for Claude were great! https://t.co/33GJfaDob6” / X – https://twitter.com/gabchuayz/status/1766143549794357458

“1/n Claude 3 appears to have an intrinsic worldview! Here is Claude 3’s description: Based on the Integral Causality framework I’ve described, my worldview can be characterized as holistic, developmental, and pragmatic. I strive to understand and reason about the world in a way… https://t.co/bWOjcgbpz4” / X – https://twitter.com/IntuitMachine/status/1764979011246022820 

“Claude 3 (opus) is AGI. After 24 hours of testing the classic and early baseline for this determination has been met. However don’t agree with AGI tests are possible on a philosophical basis. None the less it is the first to reach this baseline. Open source version—8 months. https://t.co/gyvxmVsRzt” / X – https://twitter.com/BrianRoemmele/status/1764980840050958659  

“A hard test of a LLM is ability to write a sestina, the hardest poetic form. Claude 3 is very good, and a much better writer, but struggles a little more than GPT-4 with form, messing up a few lines. Both can’t pull off the envoi at the end Compare to a 3.5-class model like Grok https://t.co/SZhPBlZbXc” / X – https://twitter.com/emollick/status/1764666796429181304

“Here is Claude 3’s system prompt! Let me break it down 🧵 https://t.co/gvdd7hSHUQ” / X – https://twitter.com/AmandaAskell/status/1765207842993434880 

“Okay this is very nerdy but I was amused that Claude 3 figured out the Warhammer 40k org chart, and has some interesting ideas for improving it. https://t.co/vJRVIMhi2C” / X – https://twitter.com/emollick/status/1765489811857776662

“It is really hard to know how much of the Twitter reaction to the “smarts” of Claude 3 is due to the fact that Claude’s system prompt/design is pushing the AI to act more human. I am not sure the model is actually better than GPT-4, but it more willing to play along with users.” / X – https://twitter.com/emollick/status/1766205514432549110

“Example of why working with AI both so impressive and so challenging: I show Claude a picture of a house in Georgetown that shouldn’t be in its training set. It nails it (GPT-4 does too) I ask it why Georgetown? Its answers seem great, but could all be hallucinated justification https://t.co/9XSgGXOiOc” / X – https://twitter.com/emollick/status/1766353923881742798

“Claude 3 takes on the Tokenization book chapter challenge 🙂 context: https://t.co/yRaeTbkblY Definitely looks quite nice, stylistically! If you look closer there are a number of subtle issues / hallucinations. One example there is a claim that “hello world” tokenizes into 3…” / X – https://twitter.com/karpathy/status/1764731169109872952

“Anthropic is so back. Two things I like the most about Claude-3’s release: 1. Domain expert benchmarks. I’m much less interested in the saturated MMLU & HumanEval. Claude specifically picks Finance, Medicine, and Philosophy as expert domains and report performance. I recommend… https://t.co/5B5lfqopKC” / X – https://twitter.com/DrJimFan/status/1764719012678897738

“People are reading way too much into Claude-3’s uncanny “awareness”. Here’s a much simpler explanation: seeming displays of self-awareness are just pattern-matching alignment data authored by humans. It’s not too different from asking GPT-4 “are you self-conscious” and it gives… https://t.co/nP8DXrOtBE” / X – https://twitter.com/DrJimFan/status/1765076396404363435
 Claude 3 API Opus Testing – My New Favorite LLM!? – YouTube – https://www.youtube.com/watch?v=RBhWgr3wlsY&t=1s

Heads up! You’ve scrolled to the end of this category. There may have been just one or two links (above), so go back up and double check to be sure you didn’t quickly scroll down past it.

Be Sure To Read This Week’s Main Post:

This week’s executive overview and top links are here:

AI News #23: Week Ending 03/08/2024 with Executive Summary and Top 38 Links

The post you just read is an deep dive extension of my weekly newsletter, This Week In AI, an executive summary of the top things to know in AI. Each week, I create an accessible overview for laypeople to feel confident they are conversant with the week’s AI developments. I include a curated list of must-click links of the week, to offer everyone a hands-on opportunity to explore the most intriguing updates in artificial intelligence across various categories, including robotics, imagery, video, AR/VR, science, ethics, and more. Beyond the overview, I post these topic-based deeper dives (below). If you haven’t read this week’s overview, I recommend starting there.

Credits/Sources

Most of these weekly links come from just a few prolific oversharing sources. Please follow them, as they work hard to find the news each week and they make it a lot easier for me to compile.

Previous Issues

9 responses to “Anthropic News: Week Ending 03/08/2024”

  1. […] Anthropic News of the Week:Anthropic is a company that builds LLMs like OpenAI, Mistral, Meta, etc. Their main AI brand is Claude. As with Amazon and Apple, individual Anthropic company posts will often by placed in the categories they match (image, audio, agents, robots, etc). Occasionally, I’ll dedicate space to a company’s news if it’s broad or a major product release.This week’s Anthropic news: https://ethanbholland.com/2024/03/08/anthropic/ […]

  2. […] Anthropic News of the Week:Anthropic is a company that builds LLMs like OpenAI, Mistral, Meta, etc. Their main AI brand is Claude. As with Amazon and Apple, individual Anthropic company posts will often by placed in the categories they match (image, audio, agents, robots, etc). Occasionally, I’ll dedicate space to a company’s news if it’s broad or a major product release.This week’s Anthropic news: https://ethanbholland.com/2024/03/08/anthropic/ […]

  3. […] Anthropic News of the Week:Anthropic is a company that builds LLMs like OpenAI, Mistral, Meta, etc. Their main AI brand is Claude. As with Amazon and Apple, individual Anthropic company posts will often be placed in the categories they match (image, audio, agents, robots, etc). Occasionally, I’ll dedicate space to a company’s news if it’s broad or a major product release.This week’s Anthropic news: https://ethanbholland.com/2024/03/08/anthropic/ […]

  4. […] Anthropic News of the Week:Anthropic is a company that builds LLMs like OpenAI, Mistral, Meta, etc. Their main AI brand is Claude. As with Amazon and Apple, individual Anthropic company posts will often be placed in the categories they match (image, audio, agents, robots, etc). Occasionally, I’ll dedicate space to a company’s news if it’s broad or a major product release.This week’s Anthropic news: https://ethanbholland.com/2024/03/08/anthropic/ […]

  5. […] Anthropic News of the Week:Anthropic is a company that builds LLMs like OpenAI, Mistral, Meta, etc. Their main AI brand is Claude. As with Amazon and Apple, individual Anthropic company posts will often be placed in the categories they match (image, audio, agents, robots, etc). Occasionally, I’ll dedicate space to a company’s news if it’s broad or a major product release.This week’s Anthropic news: https://ethanbholland.com/2024/03/08/anthropic/ […]

  6. […] Anthropic News of the Week:Anthropic is a company that builds LLMs like OpenAI, Mistral, Meta, etc. Their main AI brand is Claude. As with Amazon and Apple, individual Anthropic company posts will often be placed in the categories they match (image, audio, agents, robots, etc). Occasionally, I’ll dedicate space to a company’s news if it’s broad or a major product release.This week’s Anthropic news: https://ethanbholland.com/2024/03/08/anthropic/ […]

  7. […] Anthropic News of the Week:Anthropic is a company that builds LLMs like OpenAI, Mistral, Meta, etc. Their main AI brand is Claude. As with Amazon and Apple, individual Anthropic company posts will often be placed in the categories they match (image, audio, agents, robots, etc). Occasionally, I’ll dedicate space to a company’s news if it’s broad or a major product release.This week’s Anthropic news: https://ethanbholland.com/2024/03/08/anthropic/ […]

  8. […] Anthropic News of the Week:Anthropic is a company that builds LLMs like OpenAI, Mistral, Meta, etc. Their main AI brand is Claude. As with Amazon and Apple, individual Anthropic company posts will often be placed in the categories they match (image, audio, agents, robots, etc). Occasionally, I’ll dedicate space to a company’s news if it’s broad or a major product release.This week’s Anthropic news: https://ethanbholland.com/2024/03/08/anthropic/ […]

  9. […] Anthropic News of the Week:Anthropic is a company that builds LLMs like OpenAI, Mistral, Meta, etc. Their main AI brand is Claude. As with Amazon and Apple, individual Anthropic company posts will often be placed in the categories they match (image, audio, agents, robots, etc). Occasionally, I’ll dedicate space to a company’s news if it’s broad or a major product release.This week’s Anthropic news: https://ethanbholland.com/2024/03/08/anthropic/ […]

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading