“Some notes on the DeepSeek-V3 Technical Report 🙂 The most insane thing to me: The whole training only cost $5.576 million or ~55 days on a 2048xH800 cluster. This is TINY compared to the Llama, GPT or Claude training runs. – 671B MoE with 37B activate params – DeepSeek MoE
https://x.com/scaling01/status/1872276861675286864

“InternVL team silently dropped a Vision Language Model that beats Sonnet 3.5, Gemini 1.5 Pro AND Qwen 2 VL, rivals O1 🔥

https://x.com/reach_vb/status/1871236295424655458
“The DeepSeek Technical Report is out!! 🔥 Trained on 14.8 Trillion Tokens, outperforms all open-source models, comparable to GPT-4o and Claude-Sonnet-3.5 Key contributions: > Load Balancing Strategy: Introduced an auxiliary-loss-free approach to minimize performance

https://x.com/reach_vb/status/1872245481427837169
“DeepSeek V3 is now live in the Arena🔥 Congrats @deepseek_ai on the impressive release, matching top proprietary models like Claude Sonnet/GPT-4o across standard benchmarks. Now it’s time for the human test – come challenge it with your toughest prompts at lmarena and stay

https://x.com/lmarena_ai/status/1872348623897514446
“CodeLLM – AI CODE EDITOR THAT COMBINES ALL THE BEST CODING LLMs We combined all the best coding LLMs, including o1, Sonnet 3.5, Gemini, and Qwen, and created CodeLLM. The new CodeLLM is available in a VS code-based client by the same name. We also get UNLIMITED INTRODUCTORY

https://x.com/bindureddy/status/1870218259334869327
“I’m still an Anthropic believer even after o3” / X

https://x.com/scaling01/status/1870980302128271531
“Wait the updated Claude Sonnet 3.5 can also do sestinas without test time compute? (I last tried this a couple months ago). What is the deal with Sonnet? In his workshop, he measures time with care, Each gear and spring a universe contained Within brass walls that tick” / X

https://x.com/emollick/status/1869949581171273929
“hmm wonder if I should make a little script to set up dialogues between models on a topic. @repligate curious which model you’d pair with o1 here for thinking about research ideas. I feel like there’s a good argument for opus and a good argument for sonnet.” / X

https://x.com/gallabytes/status/1871015610827800576
Anthropic’s 3.5 Haiku model comes to Claude users | TechCrunch

Anthropic’s 3.5 Haiku model comes to Claude users


“Our co-founders discuss the past, present, and future of Anthropic. Timestamps: 00:00 Why work on AI? 02:08 Scaling breakthroughs 10:57 Sentiment shifting 18:30 The Responsible Scaling Policy 30:42 Founding story 39:08 Racing to the top 43:43 Looking to the future

https://x.com/AnthropicAI/status/1870120288601456752
Building Python tools with a one-shot prompt using uv run and Claude Projects

https://simonwillison.net/2024/Dec/19/one-shot-python-tools/
Alignment faking in large language models \ Anthropic

https://www.anthropic.com/research/alignment-faking
How Claude uses AI to identify new threats

https://www.platformer.news/how-claude-uses-ai-to-identify-new-threats/
Pre-Deployment Evaluation of Anthropic’s Upgraded Claude 3.5 Sonnet | NIST

https://www.nist.gov/news-events/news/2024/11/pre-deployment-evaluation-anthropics-upgraded-claude-35-sonnet
“@TheZvi Let me explain why I don’t believe you. Primarily it’s just you’re smart, and the 1st highlight here is a retarded strawman of objections (Scott is in meh condition lately, but I think you aren’t). The issue isn’t that Claude «did good». It’s that goal-guarding behavior [for

https://x.com/teortaxesTex/status/1871576621691371918
Building effective agents \ Anthropic

https://www.anthropic.com/research/building-effective-agents
“Couple of updates to the analysis tool today: – Claude can now analyze large Excel files (up to 30MB) that would typically exceed its context window – The analysis tool is now available in the Claude mobile apps

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading