Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Sophisticated damask wrapping paper pattern featuring ornate scrollwork formed from constitutional text fragments and ethical guidelines, interconnected clauses creating Victorian wallpaper-style lattice, ‘Anthropic’ subtly woven as decorative monogram, deep navy and antique gold on cream background with embossed texture, elegant repeating tile design, Liberty of London meets legal manuscript aesthetic, museum-quality decorative arts style.
Bloom – an open-source agentic tool that auto-generates behavioral evaluations for AI models by @AnthropicAI It turns what was once painstaking alignment work into a matter of configuration. – Bloom crafts and judges hundreds of scenarios targeting specific traits like https://x.com/TheTuringPost/status/2003629256522498061
@YashGouravKar1 Correct. In the last thirty days, 100% of my contributions to Claude Code were written by Claude Code”” / X https://x.com/bcherny/status/2004897269674639461?s=20
Codex vs. Claude Code (Today) https://build.ms/2025/12/22/codex-vs-claude-code-today/
Introducing Bloom: an open source tool for automated behavioral evaluations \ Anthropic https://www.anthropic.com/research/bloom
We estimate that, on our tasks, Claude Opus 4.5 has a 50%-time horizon of around 4 hrs 49 mins (95% confidence interval of 1 hr 49 mins to 20 hrs 25 mins). While we’re still working through evaluations for other recent models, this is our highest published time horizon to date. https://x.com/METR_Evals/status/2002203627377574113?s=20
Supercharge Claude Code with better Excel understanding 📊 Coding agents are general enough to do any type of knowledge work, including reading/creating docs. There are some pre-built skills for Claude Code to read Excel sheets, but they kind of suck 🚫 – it requires the agent https://x.com/jerryjliu0/status/2005709989558775919
I spent all of Christmas reverse engineering Claude Chrome so it would work with remote browsers. Here’s how Anthropic taught Claude how to browse the web (1/7) https://x.com/pk_iv/status/2005694082627297735
I’m hearing from many folks across finance industry that Claude for Excel is blowing their minds. The agentic coding takeoff but for other fields is coming in 2026.”” / X https://x.com/alexalbert__/status/2005670179045523595
I would judge this a win by Gemini and a close second from Claude. ChatGPT-5.2 missed the reference (though, to be fair, it did write a surprising amount of successful code to actually enhance the image) and Grok wasn’t in the ballpark. https://x.com/emollick/status/2002961280534303206
So, Claude 4.5 came in far above trend in the much-watched METR measure of the task duration that AI can accomplish autonomously at 4 hours 49 minutes. Interestingly, at the harder 80% success threshold, it is GPT-5.1 Codex Max that breaks the trend. In 2023, GPT-4 was a minute. https://x.com/emollick/status/2002208335991337467
My talk from the @aiDotEngineer Code Summit is out! 🚨 “”How Claude Code Works”” and what we can learn about frontier agent architectures. Coding agents are suddenly really really good, and I’m trying to understand why. In short: better models, simple loop design, and bash tools https://x.com/imjaredz/status/2005731826699063657
Fun article on the failure of a Claude-run vending machine in the WSJ newsroom. Reporters are amazing red teamers, creating fake policies & convincing Claude to order (and give away) Playstations & live fish. And yet… there are some hints of very viable paths forward from here https://x.com/emollick/status/2001755082510012750
I guess this (from a thinking trace of Claude 4.5 Opus) suggests @tylercowen’s strategy of writing for AI is paying off. https://x.com/emollick/status/2002546946721112173
Tired of leaving your IDE to explore GitHub repos? With Zread MCP in GLM coding plan, you can now stay in your flow: dive right into repos, explore their structure, search docs, and read files. Code smarter, not harder. https://x.com/Zai_org/status/2003872419791229285
We Let Anthropic’s Claude AI Run Our Office Vending Machine. It Lost Hundreds of Dollars. – WSJ https://www.wsj.com/tech/ai/anthropic-claude-ai-vending-machine-agent-b7e84e34
this is a brilliant read for anyone building with code agents like codex/ claude code – quick notes: > default to building CLIs first (easier for agents to verify) and progressively add other surfaces (UI) > for macOS/iOS apps default to using Swift build tooling + codex”” / X https://x.com/reach_vb/status/2005554360307065023
Message from Welcome to Sonar Chat! https://www.sonarsource.com/blog/new-data-on-code-quality-gpt-5-2-high-opus-4-5-gemini-3-and-more/





Leave a Reply