Image created with gemini-2.5-flash-image with claude-sonnet-4-5-20250929. Image prompt: A modern digital vault door in slate grey and deep blue stands open, warm celebratory light with confetti and holographic ribbons spilling out from the illuminated interior in rich reds and whites, a subtle number ‘2’ etched into the metal surface, security scanner lights pulsing festively, high-contrast cinematic photograph with crisp edges and joyful energy.

DeepMind AI safety report explores the perils of “misaligned” AI – Ars Technica https://arstechnica.com/google/2025/09/deepmind-ai-safety-report-explores-the-perils-of-misaligned-ai/

Strengthening our Frontier Safety Framework – Google DeepMind https://deepmind.google/discover/blog/strengthening-our-frontier-safety-framework/

As AI capability increases, alignment work becomes much more important. In this work, we show that a model discovers that it shouldn’t be deployed, considers behavior to get deployed anyway, and then realizes it might be a test.”” / X https://x.com/sama/status/1968674357309223020

A postmortem of three recent issues \ Anthropic https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues

GSA and xAI Partner on $0.42 per Agency Agreement to Accelerate Federal AI Adoption | GSA https://www.gsa.gov/about-us/newsroom/news-releases/gsa-xai-partner-to-accelerate-federal-ai-adoption-09252025

10GW is about $340B of nvidia h100 at $30k/gpu (assuming 20% of power for non-gpus). if openai got a 30% volume discount, they’d pay nvidia $230b probably. so instead, maybe openai pays nvidia full price and nvidia invests the excess $100B into openai stock 😬 (just throwing”” / X https://x.com/soumithchintala/status/1970464906072801589

looking forward to what we’ll build together with NVIDIA!”” / X https://x.com/gdb/status/1970299081999426016

More compute in the making. Announcing 5 new Stargate sites with Oracle and SoftBank, putting us ahead of schedule on the 10-gigawatt commitment we announced in January. https://x.com/OpenAI/status/1970601342680084483

OpenAI & NVIDIA Announce Strategic Partnership to Deploy 10GW of NVIDIA Systems This enables OpenAI to build & deploy at least 10 gigawatts of AI datacenters with NVIDIA systems representing millions of GPUs for OpenAI’s next-gen AI infrastructure. https://x.com/OpenAINewsroom/status/1970157101633990895

OpenAI and NVIDIA Announce Strategic Partnership to Deploy 10 Gigawatts of NVIDIA Systems | NVIDIA Newsroom https://nvidianews.nvidia.com/news/openai-and-nvidia-announce-strategic-partnership-to-deploy-10gw-of-nvidia-systems

OpenAI and NVIDIA announce strategic partnership to deploy 10 gigawatts of NVIDIA systems | OpenAI https://openai.com/index/openai-nvidia-systems-partnership/

Together, NVIDIA and OpenAI are expanding the frontier of AI — transforming nearly every industry and unlocking use cases once unimaginable.   “There’s no partner but NVIDIA that can do this at this kind of scale, at this kind of speed,” said @OpenAI CEO Sam Altman. https://x.com/nvidianewsroom/status/1970223778937586043

OpenAI, SAP & Microsoft are launching OpenAI for Germany—a partnership to bring frontier AI to Germany’s public sector, through a sovereign, certified cloud environment. Built on SAP’s Delos Cloud and running on Microsoft Azure, this new initiative will help employees across”” / X https://x.com/OpenAINewsroom/status/1970844821624680801

SAP and OpenAI partner to launch sovereign ‘OpenAI for Germany’ | OpenAI https://openai.com/global-affairs/openai-for-germany/

Flooding the AI Frontier – Chinese models are DOMINATING the open-weight LLM space.
– by Ben https://bturtel.substack.com/p/flooding-the-ai-frontier

Bad news for AI safety: To fight against AI regulation, VC firm Andreessen Horowitz, AI billionaire Greg Brockman, and others recently started a >$100 million super PAC, one of the largest operating PACs in the US.”” / X https://x.com/janleike/status/1969115275837440206

Weaviate is ISO 27001 compliant! 🎉 But what does that actually mean for you? ISO 27001:2022 is the international standard for information security management systems. It requires organizations to systematically manage information security risks through continuous monitoring, https://x.com/weaviate_io/status/1970912361381843104

Infra bugs are evil. Kudos to the team at @AnthropicAI for finding the bugs, and then for transparently reporting them in their fairly detailed writeup.”” / X https://x.com/hyhieu226/status/1968708468820312435

Lots of sympathy to the Anthropic team 🙏🙏🙏 https://x.com/cHHillee/status/1968536182284849459

Meta’s AI system Llama approved for use by US government agencies | Reuters https://www.reuters.com/world/us/metas-ai-system-llama-approved-use-by-us-government-agencies-2025-09-22/

We’re excited to share our preparedness report on Code World Model (CWM), FAIR’s latest open-weight model for code generation and reasoning. This report was developed by the SEAL team and the AI Security team, marking our first external publication since part of SEAL joined Meta”” / X https://x.com/summeryue0/status/1970971944557346851

The Illusion of Readiness (Health AI) • GPT-5 & peers ace med benchmarks—but stress tests reveal fragility • Guess answers w/o images, flip under trivial prompt tweaks • Fabricate “reasoning” that sounds right but isn’t • Leaderboard wins ≠ real-world readiness https://x.com/arankomatsuzaki/status/1970684893966516477

We’ve made progress on the AI safety problem of detecting and reducing “”scheming””: – Created evaluation environments to detect scheming – Observed current models scheming in controlled settings – Found deliberative alignment ( https://x.com/gdb/status/1969437389027492333

Anti-Scheming https://www.antischeming.ai/

Endpoint Security for AI eBook https://www.delltechnologies.com/asset/en-us/solutions/business-solutions/briefs-summaries/endpoint-security-for-ai-ebook.pdf

What ensures safety in AI? Guardian models are the very safety layers that detect and filter harmful prompts and outputs, defending AI today. But they go beyond simple filtering. They can: – Serve as guardrails to block harmful content in real time – Act as evaluators to check https://x.com/TheTuringPost/status/1968635881004363969

A cautiously optimistic result on AI and disinformation. A week before 2024 UK elections 13% of all voters used AI for political topics. A randomized trial found this may be good: using AI led to similar gains in true knowledge as doing web search, regardless of model & prompts. https://x.com/emollick/status/1968770085973008621

This makes me think about @joshgans’ paper arguing that having authors sneak prompt injections into academic work IMPROVES science in a world of AI. Without the risk, reviewers would tend to rely heavily on AI reviews, with injections, they need to include some human review. https://x.com/emollick/status/1968526173442318618

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading