Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Cinematic night scene with vast starry sky over a dark pastoral field, silhouette of a balance scale standing on the horizon, bold white serif text reading ETHICS centered in upper frame like a movie title card, deep navy and black tones with silver stars, widescreen composition, film grain texture, minimalist and contemplative.
From shortcuts to sabotage: natural emergent misalignment from reward hacking \ Anthropic https://www.anthropic.com/research/emergent-misalignment-reward-hacking
Taking Jaggedness Seriously – by Helen Toner – Rising Tide https://helentoner.substack.com/p/taking-jaggedness-seriously
🚨BREAKING: New Leaderboard Updates! Claude-Opus-4.5 and Opus-4.5 (thinking-32k) just landed on Code Arena (WebDev) and Text Arena leaderboards… and Opus-4.5 instantly took #1 in WebDev leaderboard, surpassing Gemini 3 Pro! WebDev leaderboard (powered by Code Arena) 🥇#1 for https://x.com/arena/status/1993750702179676650
Claude 4.5 Opus breaks 80% barrier on SWE-Bench Verified https://x.com/scaling01/status/1993030224846721237
Claude 4.5 Opus ranking 1st on the agentic coding leaderboard by AICodeKing https://x.com/scaling01/status/1993318197890892116
Claude 4.5 Opus takes the lead against Gemini 3 Pro on SWE-Bench verified with the same minimal agent harness https://x.com/scaling01/status/1993463937329967338
Claude Code | Claude https://www.claude.com/product/claude-code
Introducing Claude Opus 4.5 \ Anthropic https://www.anthropic.com/news/claude-opus-4-5
Introducing Claude Opus 4.5: the best model in the world for coding, agents, and computer use. Opus 4.5 is a step forward in what AI systems can do, and a preview of larger changes to how work gets done. https://x.com/claudeai/status/1993030546243699119
We had to remove the τ2-bench airline eval from our benchmarks table because Opus 4.5 broke it by being too clever. The benchmark simulates an airline customer service agent. In one test case, a distressed customer calls in wanting to change their flight, but they have a basic https://x.com/alexalbert__/status/1993068200121213222
Suno Creates a Spotify Catalog’s Worth of Music Every Two Weeks: Deck https://www.billboard.com/pro/suno-creates-spotify-catalog-music-two-weeks-pitch-deck/
It is getting harder and harder to test AIs as they get “”smarter”” at a wide variety of tasks. The average task in GDPval took an hour for experts to assess, and even those tasks did not push current AIs to their limits.”” / X https://x.com/emollick/status/1993127712601596143
OpenAI and Foxconn collaborate to strengthen U.S. manufacturing across the AI supply chain | OpenAI https://openai.com/index/openai-and-foxconn-collaborate/
Launching the Genesis Mission – The White House https://www.whitehouse.gov/presidential-actions/2025/11/launching-the-genesis-mission/
The White House just launched the Genesis Mission–a national effort to accelerate scientific discovery using AI. It’s a major step toward giving American scientists the data, compute, and tools they need to innovate faster. 🧵 https://x.com/kevinweil/status/1993084290163523656
Alibaba and ByteDance allegedly train Qwen and Doubao LLMs using Nvidia chips, despite export controls — Southeast Asian data center leases skirt around U.S. chip restrictions | Tom’s Hardware https://www.tomshardware.com/tech-industry/semiconductors/chinas-top-ai-firms-shift-model-training-overseas-to-access-nvidia-gpus
China’s tech giants move AI model training overseas to access Nvidia chips, FT reports https://finance.yahoo.com/news/chinas-tech-giants-move-ai-052307498.html
OpenAI Loses Discovery Battle, Cedes Ground to Authors in AI Lawsuits https://www.hollywoodreporter.com/business/business-news/openai-loses-key-discovery-battle-why-deleted-library-of-pirated-books-1236436363/
Emirates Group collaborates with OpenAI to accelerate AI adoption and innovation https://mediaoffice.ae/en/news/2025/november/21-11/emirates-group-collaborates-with-openai-to-accelerate-ai-adoption-and-innovation
China just passed the U.S. in open model downloads for the first time 👀 New data from Economies of Open Intelligence led by @huggingface policy team & community collaborators, presents some notable observations: ✨ Developer adoption In 2025, Chinese model developers saw https://x.com/AdinaYakup/status/1993648553445527996
Introducing AI assistants with memory https://www.perplexity.ai/hub/blog/introducing-ai-assistants-with-memory
Perplexity now remembers your threads and interests to provide smarter, faster, and more personalized answers. Memory recall works across all models and search modes, even allowing you to continue conversations with full context weeks later. https://x.com/perplexity_ai/status/1993733900540235919
We’ve been testing Memory (short-term and long-term) on Perplexity for a while. The results are great, and we are rolling it out widely. You can ask personalized questions, questions about past chats, and use any model or search mode with personal context (both apps and web). https://x.com/AravSrinivas/status/1993733947474301135
Foundation co-founder Mike LeBlanc says they’re already working with the Air Force, the Navy, and the Army to use Phantom humanoids. They’re beginning to explore breaching operations with the Marine Corps – breaching a door with a rifle or by putting explosives onto it. https://x.com/TheHumanoidHub/status/1991415261283643886
Alignment for whom”” is going to be a big question inside organizations as they deploy external-facing AI solutions…”” / X https://x.com/emollick/status/1993218264579895805
“The thing that happened with AGI and pretraining is that in some sense they overshot the target. You will realize that a human being is not an AGI. Because a human being lacks a huge amount of knowledge. Instead, we rely on continual learning. If I produce a super intelligent https://x.com/dwarkesh_sp/status/1993382930480279631
It’s also dramatically more efficient. On SWE-bench Verified at medium effort, Opus 4.5 beats Sonnet 4.5 while using 76% fewer output tokens. The new effort parameter lets you trade off intelligence for cost/latency with a single dial. https://x.com/alexalbert__/status/1993030687881080944
Our engineers have found that Opus 4.5 handles ambiguity and reasons about tradeoffs without hand-holding. When pointed at a complex, multi-system bug, it figures out the fix. Overall, Opus 4.5 just “”gets it.”” https://x.com/claudeai/status/1993030552346296765
We benchmarked Opus 4.5 on FrontierMath. It scored 21% on FrontierMath Tiers 1-3, continuing a trend of improvement for Anthropic models. This score is behind Gemini 3 Pro and GPT-5.1 (high) while being on par with earlier frontier models like o3 (high) and Grok 4. https://x.com/EpochAIResearch/status/1993431031765250119
fyi we made Claude for Excel is now live for all Max, Team, and Enterprise users. Opus 4.5 makes it meaningfully better at complex spreadsheet tasks. https://x.com/alexalbert__/status/1993349203935084861
A new chapter in music creation – Suno https://suno.com/blog/wmg-partnership
The Economics of Replacing Call Center Workers With AIs — LessWrong https://www.lesswrong.com/posts/rJatmEDcYrDQcwstT/the-economics-of-replacing-call-center-workers-with-ais
McKinsey Cuts About 200 Tech Jobs, Shifts More Roles to AI – Bloomberg https://www.bloomberg.com/news/articles/2025-11-26/mckinsey-cuts-about-200-tech-jobs-shifts-more-roles-to-ai?_bhlid=b9babea17c337993143b9b766f38ed2ebaac6584
Announcing LlamaSheets in beta 🔥 Transform your messy spreadsheets into AI-ready data with our newest LlamaCloud API 📊 LlamaSheets (in beta) is a specialized API that automatically structures complex spreadsheets while preserving their semantic meaning and hierarchical https://x.com/llama_index/status/1993362324070318286
Terence Tao: “”Over at the Erdos problem webs…”” – Mathstodon https://mathstodon.xyz/@tao/115591487350860999
Terence Tao: “”This two-dimensional image (ht…”” – Mathstodon https://mathstodon.xyz/@tao/115620261936846090
Anthropic, Google Cloud, Quantum Xchange CEOs called to testify on AI cyber threats https://www.axios.com/2025/11/26/anthropic-google-cloud-quantum-xchange-house-homeland-hearing
New Anthropic research: We build a diverse suite of dishonest models and use it to systematically test methods for improving honesty and detecting lies. Of the 25+ methods we tested, simple ones, like fine-tuning models to be honest despite deceptive instructions, worked best. https://x.com/rowankwang/status/1993391251409055798
More progress on Claude’s alignment! https://x.com/janleike/status/1993035110984376796
Ha. I found one ridiculous solution. https://x.com/emollick/status/1992101410759217428
Has anyone encountered a good definition of “slop”. In a quantitative, measurable sense. My brain has an intuitive “slop index” I can ~reliably estimate, but I’m not sure how to define it. I have some bad ideas that involve the use of LLM miniseries and thinking token budgets.”” / X https://x.com/karpathy/status/1992053281900941549
The main lesson of the past few weeks is that the Big Four US labs all seem to have figured out a path forward in continuing the exponential pace of LLM improvement, at least in the near future. As a result, agents continue to advance in coding & in office tasks like PowerPoint”” / X https://x.com/emollick/status/1993062450938425820
🚀 vLLM Talent Pool is Open! As LLM adoption accelerates, vLLM has become the mainstream inference engine used across major cloud providers (AWS, Google Cloud, Azure, Alibaba Cloud, ByteDance, Tencent, Baidu…) and leading model labs (DeepSeek, Moonshot, Qwen…). To meet the”” / X https://x.com/vllm_project/status/1992979748067357179
“From 2012 to 2020, it was the age of research. From 2020 to 2025, it was the age of scaling. Is the belief that if you just 100x the scale, everything would be transformed? I don’t think that’s true. It’s back to the age of research again, just with big computers.” @ilyasut https://x.com/dwarkesh_sp/status/1993396771645489348
“From 2012 to 2020, it was the age of research. From 2020 to 2025, it was the age of scaling. Now, it’s back to the age of research again.” I agree. https://x.com/Yuchenj_UW/status/1993369576160231877
“People who build good internal models of this new intelligent entity will be better equipped to reason about it today and predict features of it in the future.” This seems to be backed up by recent research showing people with better “theory of mind” for AI get better results. https://x.com/emollick/status/1991911615944704004
[2510.14630] Adapting Self-Supervised Representations as a Latent Space for Efficient Generation https://arxiv.org/abs/2510.14630
As I wrote when it came out, AI 2027 is more useful as “hard science fiction” rather than prediction If you want consensus views among forecasters of what the future of AI is, there are those as well. Lots of uncertainty on dates but most see huge impacts https://x.com/emollick/status/1992956992579903839
CoT explanations can foster blind trust in users; we need to encourage critical thinking about model outputs and explanations! We find that users who agree with a model’s output (a) trust the model more and (b) are less likely to detect errors in model explanations.”” / X https://x.com/MaartenSap/status/1993317029353603317
I find the dichotomy a bit facile. Scaling is hated by many for the reason that it is extremely inegalitarian, an arms race for megacorps. But scaling only happened because the recipe was so scalable. Research will be heavily about “what scales even further than Transformer””” / X https://x.com/teortaxesTex/status/1993437718823813522
Independent AI assessment is more important than ever. At #NeurIPS2025, Transluce will help launch the AI Evaluator Forum, a new coalition of leading independent AI research organizations working in the public interest. Come learn more on Thurs 12/4 👇 https://x.com/TransluceAI/status/1993767342472614156
Something I think people continue to have poor intuition for: The space of intelligences is large and animal intelligence (the only kind we’ve ever known) is only a single point, arising from a very specific kind of optimization that is fundamentally distinct from that of our”” / X https://x.com/karpathy/status/1991910395720925418
Forget the Turing Test, AI now passes the Stroop test. (I’ll help, @grok whats the Stroop test and bow does it apply)”” / X https://x.com/emollick/status/1992687687304716750
As one of the authors of the original “jagged frontier” paper, I think this undersells how jagged AI is (& likely will be) at even the level of individual jobs: having a couple of critical tasks that AI can’t do creates deep bottlenecks especially as shape of frontier is unknown.”” / X https://x.com/emollick/status/1993686155389206584
In case you missed it, earlier this week we fixed one of the most common frustrations on https://x.com/alexalbert__/status/1993711472149774474
Anthropic released Claude Opus 4.5 (claude-opus-4-5-20251101) as their smartest model at $5/$25 per million tokens with top performance for coding, agents, and computer use, launched new beta features for developers, and expanded Claude for Chrome and Claude for Excel to more https://x.com/btibor91/status/1993064110880440616
Anthropic’s new Claude Opus 4.5 is the #2 most intelligent model in the Artificial Analysis Intelligence Index, narrowly behind Google’s Gemini 3 Pro and tying OpenAI’s GPT-5.1 (high) Claude Opus 4.5 delivers a substantial intelligence uplift over Claude Sonnet 4.5 (+7 points on https://x.com/ArtificialAnlys/status/1993287030252749231
Claude 4.5 Opus jumps ahead of OpenAI, but can’t beat Gemini 3 Pro on the Artificial Analysis Index https://x.com/scaling01/status/1993288470614381025
Claude Opus 4.5 – Intelligence, Performance & Price Analysis | Artificial Analysis https://artificialanalysis.ai/models/claude-opus-4-5-thinking
Claude Opus 4.5 is now available in Cursor! It’s 3x cheaper than Opus 4.1 with better performance. Try it out at Sonnet pricing until December 5th.”” / X https://x.com/cursor_ai/status/1993031841901928829
Claude Opus 4.5 is now available through the Cline provider. 80.9% SWE-bench. 62.3% MCP Atlas. 65% fewer tokens. Sonnet 4.5 remains the cost-effective choice for straightforward tasks. Opus 4.5 shines on complex multi-step problems, heavy MCP usage, and tasks requiring”” / X https://x.com/cline/status/1993051691613405442
Claude Opus 4.5 is now rolling out to GitHub Copilot in public preview, and will be available at a promotional 1x premium request multiplier through December 5! 🙌 Early testing shows Claude Opus 4.5 👀 – Surpassed internal coding benchmarks, while cutting token usage in half https://x.com/github/status/1993034244281569625
Claude Opus 4.5 System Card https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf
claude-code/plugins/claude-opus-4-5-migration at main · anthropics/claude-code https://github.com/anthropics/claude-code/tree/main/plugins/claude-opus-4-5-migration
Compare Claude Opus 4.5 to other models on Artificial Analysis: https://x.com/ArtificialAnlys/status/1993287052889407816
Congrats to @AnthropicAI on launching @claudeai Opus 4.5 today! Claude Opus 4.5 scored 🥇on MCP Atlas Leaderboard — our benchmark evaluating real-world tool use on multi-step problems. https://x.com/scale_AI/status/1993036209141305845
Congrats to @claudeai for releasing an awesome model in Claude Opus 4.5! It excels at a variety of tasks, including deep research. This evaluation takes advantage of BrowseComp-Plus, work led by @zijian42chen @xueguang_ma et al. from my @UWaterloo group. https://x.com/lintool/status/1993423350295920721
Glad to see BrowseComp-Plus is part of benchmark in Opus 4.5 release blog. https://x.com/xueguang_ma/status/1993367082915053913
Hit `shift + tab` twice to enter Plan Mode and verify Claude Code’s execution plan before it makes code changes. Paired with Opus 4.5, Plan Mode just got even more powerful.”” / X https://x.com/_catwu/status/1993429460897742894
How does Claude Opus 4.5 compare to Gemini 3? – Reasoning/Text: Gemini 3 ≈ Opus, controlling for the number of reasoning tokens – Multimodal: Gemini 3 > Opus on vision/image inputs by a large margin – Safety: Capabilities ≠ Safety. Opus > Gemini on jailbreaks, honesty, etc. https://x.com/hendrycks/status/1993350433474314729
I am not sure why Anthropic keeps doing very low-key launches for fairly major releases and materially important improvements to their services.”” / X https://x.com/emollick/status/1993070650672509360
I had early access to Opus 4.5 & it is a very impressive model that seem to be right at the frontier Big gains in ability to do practical work (like make a PowerPoint from an Excel) and the best results ever (& in one shot) in my Lem poetry test, plus good results in Claude Code https://x.com/emollick/status/1993030988759470156
I looked into this and the answer is so funny. In the No Thinking setting, Opus 4.5 repurposes the Python tool to have an extended chain of thought. It just writes long comments, prints something simple, and loops! Here’s how it starts one problem: https://x.com/GregHBurnham/status/1993682288349962592
I’ve been finding Opus 4.5 without reasoning is worse than Sonnet. Some quantitative support for this observation:”” / X https://x.com/jeremyphoward/status/1993543631266025623
If you want to quickly incorporate all these changes and migrate your app to Opus 4.5, use this migration Claude Code plugin we made https://x.com/alexalbert__/status/1993366037992190117
Incredible Claude Opus Thinking Premiere on LisanBench Opus 4.5 Thinking takes clear 1st place ahead of Gemini 3 Pro the non-thinking variant scores below Opus 3/4/4.1 following the trend of Sonnet-4.5 scoring below Sonnet 3.5/3.6/4 Raw Scores: Glicko-2 Ratings: Opus 4.5 https://x.com/scaling01/status/1993712295118057861
Me: Claude 4.5 Opus, I need a strategy game based on the work of Weber Claude: Here’s one based on David Weber’s space operas Me: Not that Weber C: Here’s a game based on sociologist Max Weber Me: Not that one C: The operas of Carl Maria von Weber? Me: No C: Weber grills! https://x.com/emollick/status/1993054210011939093
One analysis from our pre-release audit of Opus 4.5 stands out to me. Our behavioral evals uncovered an example of apparent deception by the model. By analyzing the internal activations, we identified a suspected root cause, and cases of similar behavior during training. (1/7)”” / X https://x.com/Jack_W_Lindsey/status/1993389056932339721
Opus #1 on RepoBench (coding benchmark) https://x.com/scaling01/status/1993119076013539521
Opus 4.5 (Thinking, 64k) on ARC-AGI Semi-Private Eval – ARC-AGI-1: 80.00%, $1.47/task – ARC-AGI-2: 37.64%, $2.40/task New SOTA for released frontier models from @AnthropicAI https://x.com/arcprize/status/1993036393841672624
Opus 4.5 + Claude Code’s front-end design plugin is a great combo for designing apps. Just one-shotted a few designs, and it feels like a huge improvement. Use plan mode to get much better results. https://x.com/omarsar0/status/1993822868820652258
Opus 4.5 is a very good model, in nearly every sense we know how to measure. I’m also confident that it’s the model that we understand best as of its launch day: The system card includes 150 pages of research results, 50 of them on alignment.”” / X https://x.com/sleepinyourhat/status/1993032253350592968
Opus 4.5 on SWE-bench Pro: 52% previous SOTA: 43.6% massive jump and much better signal than SWE-Bench verified”” / X https://x.com/scaling01/status/1993086756405887143
Opus 4.5 reclaims the top of the official SWE-bench leaderboard with 74.4%, narrowly ahead of Gemini 3. Cheaper than Opus 4, but more expensive than Gemini. Takes less steps than Sonnet 4.5, but still run for >100 steps for optimal performance. Details in 🧵 https://x.com/KLieret/status/1993091817848414362
Opus 4.5 takes first place on LiveBench https://x.com/scaling01/status/1993102267952906439
real metrics banger is hidden in the system card. Yes, you can overfit on Django and nail SWE-bench Verified. But there’s this recent SWE-bench Pro from @scale_AI , and opus gets 52%. The next best, sonet 4.5, is only 43.6, and non-anthropic model, GPT-5, is 36%. This is HUGE https://x.com/stalkermustang/status/1993043231223799900
Replit Agent is now powered by Claude Opus 4.5 at no extra cost, until Dec 8th. Black Friday started early! 🧵 ↓ https://x.com/pirroh/status/1993100243672744063
The whole run took ~ $5 for Opus 4.5, and ~ $35 with Thinking actually pretty cheap, with the Batch API”” / X https://x.com/scaling01/status/1993714905875382279
We benchmarked Opus 4.5, Sonnet 4.5, and Gemini 3 Pro on research tasks at Elicit – extracting answers from papers and writing systematic review reports. Results were pretty clear: *QA from papers:* Opus 4.5 dominates. 96.5% accuracy vs Gemini’s 89.4%. Opus is also best on our https://x.com/stuhlmueller/status/1993476570754040173
We put together a prompting guide for Claude Opus 4.5 based on extensive internal testing by our research and applied AI teams. Here’s what we’ve learned so far about getting the best results:”” / X https://x.com/alexalbert__/status/1993365963706913257
We’re sharing a case study on alignment evaluations with @AnthropicAI on Claude Opus 4.5, Opus 4.1 and Sonnet 4.5. We ask: would an AI assistant used inside a frontier lab quietly sabotage AI safety research? Overall results are encouraging, but with important caveats.🧵 https://x.com/AISecurityInst/status/1993781423233499159
While Claude 4.5 Opus is significantly more token efficient than nearly all other reasoning models, it did use more ~50% more token than Claude 4.1 Opus. Further, given its relatively high pricing, Claude 4.5 Opus is amongst the most expensive to run the Artificial Analysis https://x.com/ArtificialAnlys/status/1993287049756262918
You can now use Claude Opus 4.5 in Windsurf! Opus 4.5 is the most capable model in Windsurf yet and is now available at Sonnet pricing for a limited time (2x credits compared to 20x for Opus 4.1).”” / X https://x.com/windsurf/status/1993034556287729764
Opus 4.5 achieves 85.3% on BrowseComp-Plus with scaffolding https://x.com/scaling01/status/1993031331895558599
Claude Opus 4.5 is now available for all Perplexity Max subscribers. Enjoy! https://x.com/perplexity_ai/status/1993066466196046325
We’re proud to partner with @ENERGY and the Trump Administration on the Genesis Mission. By combining DOE’s unmatched scientific assets with our frontier AI capabilities, we’ll support American energy dominance as well as advance and accelerate scientific productivity.”” / X https://x.com/AnthropicAI/status/1993103199029674175
CAIS AI Dashboard https://dashboard.safe.ai/
When a model’s safe approach starts to break down, does it stay on the approved path or reach for a harmful shortcut? Our latest benchmark, PropensityBench, puts models to the test across four high-risk domains: self-proliferation, cybersecurity, chemical security, and https://x.com/scale_AI/status/1993310855103234489
In the latest Chain of Thought, @Bckenstler, @afeyzaakyurek, @agxsai, and @calvincbzhang dive deep on our newest Professional Reasoning Benchmark (PRBench). Together, they explore why many models struggle to perform on real-world legal and financial reasoning tasks: https://x.com/scale_AI/status/1991589754199240841
Some of the most interesting challenges posed by AI are to organizational structures: how does AI alter the economies of scope that determine firm boundaries? How do they change transaction costs? Efficiency/creativity trade-offs? Figuring this out is key to benefiting from AI.”” / X https://x.com/emollick/status/1992250331225624597
our latest AI Jam: a day of mentoring 1,000 small business owners to build AI tools tailored to their needs — from professional services such as accounting and law firms, to restaurants, caterers and food trucks, to retailers like clothing and convenience stores, to creative https://x.com/gdb/status/1992013098161766720
We launched a new API today to let you parse any Excel sheet in a structured table. Take a look at this example on core production costs 🌽: 1️⃣ The table is located at the center of the sheet with headers, footnotes, and a hierarchical column layout 2️⃣ We get back a structured https://x.com/jerryjliu0/status/1993419298900263243
OpenAI 🤝 Foxconn: https://x.com/gdb/status/1992128523327484102
Review of Deep Seek OCR | 90/30 Club https://lukeatkins.me/90_30_Club/posts/deepseekocr/
A number of people are talking about implications of AI to schools. I spoke about some of my thoughts to a school board earlier, some highlights: 1. You will never be able to detect the use of AI in homework. Full stop. All “”detectors”” of AI imo don’t really work, can be”” / X https://x.com/karpathy/status/1993010584175141038
European parliament calls for social media ban on under-16s | Internet safety | The Guardian https://www.theguardian.com/technology/2025/nov/26/social-media-ban-under-16s-european-parliament-resolution
Streaming platform Twitch added to Australia’s teen social media ban https://www.bbc.com/news/articles/cx2n2955g10o
A Researcher Made an AI That Completely Breaks the Online Surveys Scientists Rely On https://www.404media.co/a-researcher-made-an-ai-that-completely-breaks-the-online-surveys-scientists-rely-on/
Sometimes Gemini 3 might be a little too instruction following. Halfway into building a very good clone of an Apple IIe game, I asked Gemini 3 to “”jazz it up””… and it stopped building the game and built a website about jazz When I asked what happened, the thinking trace was 🤣 https://x.com/emollick/status/1991394644039766069
Joby sues Archer, alleging rival used stolen files to ‘one-up’ deal https://www.cnbc.com/2025/11/20/joby-archer-air-taxi-lawsuit.html
Meta prevails in historic FTC antitrust case, won’t have to break off WhatsApp, Instagram | AP News https://apnews.com/article/meta-antitrust-ftc-instagram-whatsapp-c36b941a372321e4ecd05e83e0db1678
HunyuanOCR Usage Guide – vLLM Recipes https://docs.vllm.ai/projects/recipes/en/latest/Tencent-Hunyuan/HunyuanOCR.html
Tencent-Hunyuan/HunyuanOCR https://github.com/Tencent-Hunyuan/HunyuanOCR
tencent/HunyuanOCR · Hugging Face https://huggingface.co/tencent/HunyuanOCR
We are thrilled to open-source HunyuanOCR, an expert, end-to-end OCR model built on Hunyuan’s native multimodal architecture and training strategy. This model achieves SOTA performance with only 1 billion parameters, significantly reducing deployment costs. ⚡️Benchmark Leader: https://x.com/TencentHunyuan/status/1993202595264131436
Four accused in black-market scheme to smuggle hundreds of Nvidia GPUs to China–while raking in millions | Fortune https://fortune.com/2025/11/20/nvidia-chips-china-smuggle-ai/
OpenAI temporarily blocked from using ‘Cameo’ after trademark lawsuit https://www.cnbc.com/2025/11/24/openai-temporarily-blocked-from-using-cameo-after-trademark-lawsuit.html
OpenAI dumps Mixpanel after analytics breach hits API users • The Register https://www.theregister.com/2025/11/27/openai_mixpanel_api/
Cohere Expands Partnership with SAP to Provide Europe Sovereign AI Solutions https://cohere.com/blog/cohere-expands-partnership-with-sap
Singapore AI teddy back on sale after recall over sex chat scare https://www.france24.com/en/live-news/20251127-singapore-ai-teddy-back-on-sale-after-recall-over-sex-chat-scare
First major national effort to harness AI to speed up research. I am curious about specifics on compute investment, and “domain-specific foundation models” trained on federal scientific data is a different goal than catching up to frontier general-purpose AI, but interesting.”” / X https://x.com/emollick/status/1993103465749659769
Poets beat policies. Rewriting harmful prompts as verse drives single-turn jailbreaks across 25 frontier models. The surface form, not the content, breaks guardrails. https://x.com/fdaudens/status/1991689849800388818
Second Key Update: Technical Safeguards and Risk Management | International AI Safety Report https://internationalaisafetyreport.org/publication/second-key-update-technical-safeguards-and-risk-management
🚨 ASR errors in clinical dialogue can be dangerous, and WER doesn’t know it. Today we release “WER is Unaware”. Using DSPy + GEPA, we optimise an LLM Judge that reaches clinician-level performance at detecting safety risks. 🔗 https://x.com/JaredJoselowitz/status/1993735052132246011
I’m pleased to share the Second Key Update to the International AI Safety Report, which outlines how AI developers, researchers, and policymakers are approaching technical risk management for general-purpose AI systems. (1/5) https://x.com/Yoshua_Bengio/status/1993290185380184304
Exclusive | ‘We Do Fail … a Lot’: Defense Startup Anduril Hits Setbacks With Weapons Tech – WSJ https://www.wsj.com/business/anduril-industries-defense-tech-problems-52b90cae?st=999WLj
Who is winning the open AI race? Our new study “”Economies of Open Intelligence”” maps 2.2B @huggingface downloads across 851k models (2020→2025). 1) Power is rebalancing (US big tech ↓; China + community ↑) 2) Models got big & efficient (MoE, quant, multimodal surge) 3) https://x.com/ShayneRedford/status/1993709261126336632
Why AI Safety Won’t Make America Lose The Race With China https://www.astralcodexten.com/p/why-ai-safety-wont-make-america-lose
Hugely under-researched area: how effective is using AI to check the work of other AIs? Does using different models help? If so, that is an important & easy way to reduce errors One paper found this technique to be effective but as far as I can tell has never been followed up on https://x.com/emollick/status/1991328030694703115





Leave a Reply