Statement from Dario Amodei on the Paris AI Action Summit \ Anthropic https://www.anthropic.com/news/paris-ai-summit

“The new ChatGPT-4o upgrades, whatever they were, does make it seem to pull off some funny writing for the first time. Claude still feels more light and charming, but the 4o update is very noticeable. “the most disturbing yet absurd corporate memo you can come up with” “MORE” https://x.com/emollick/status/1891586403126919413

“GPQA: 448 multiple choice questions in 16 subdomains SuperGPQA: 26,529 mutiple choice questions across 285 graduate disciplines 😲 DeepSeek-R1 outperforms o1, o2-mini, Claude 3.5 Sonnet, etc. on this benchmark 🤔 https://x.com/iScienceLuvr/status/1892879645223375319

“LLMs are still incredibly bad at long context, severe drop in response quality from the best of the best models (o1, Claude, grok, DeepSeek), doesn’t really matter what model – it will choke” / X https://x.com/abacaj/status/1893024046469493212

“Openrouter is now supported in ai-gradio in a few lines of code you can use deepseek-r1, claude, gemini and more with coder mode pip install –upgrade “ai-gradio[openrouter]” import gradio as gr import ai_gradio gr.load( name=’openrouter:anthropic/claude-3.5-sonnet”, https://x.com/_akhaliq/status/1890543241017405695

“I can’t believe X users are so stupid. Not voting for o3-mini is insane. You can literally already distill 4o, Claude 3.5, Deepseekv3, etc into sizes that will run on phones.” / X https://x.com/dylan522p/status/1891682135255154775

Introducing the Anthropic Economic Index \ Anthropic https://www.anthropic.com/news/the-anthropic-economic-index

“Proxy for Anthropic revenue growth. But as I said, Gemini 2.0 Flash will cut into Anthropic market share. If you look at Feb 13th token usage from Sonnet went down 7B and token usage from Gemini 2.0 Flash went up 7B. Anthropic needs to lower prices, release a better model for https://x.com/scaling01/status/1891849320321720399

“After the Grok-3 launch you have to consider xAI as a real competitor for SOTA models. Everything else is just cope. However, internally OpenAI, Anthropic and Google are likely ahead. But honestly not so sure about Google anymore, they need to drop a banger (Pro/Ultra with” / X https://x.com/scaling01/status/1891846484791820502

“i don’t recall ever caring about the algorithm powering google search i think we’re quickly arriving at the point where we also won’t care what model (gemini, claude, openai,…) powers our ai systems the better product or dev experience will win at the end of the day” / X https://x.com/omarsar0/status/1891570913327374496

“Anthropic introduced the idea of sandbagging a few years ago – showing that AIs may give you worse answers if you ask questions in an ignorant way. This new research shows sandbagging extends to alignment. Ask the AI to do something bad and it might pretend it can’t do that. https://x.com/emollick/status/1891905364376862966

“One concern about the present round of new models (starting with Claude 3.5 update) is that the labs have realized how much of a role personality plays in how smart the AI appears, making it harder to tell the difference between how much improvement is real and how much is vibes” / X https://x.com/emollick/status/1891541312018538755

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading