Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A 16:9 flat matte op-art poster in the style of Julio Le Parc, the word ETHICS centered in bold letters built from concentric ROYGBIV rainbow bands with crisp printed edges, a single continuous rainbow ribbon arcing above and looping down into two perfectly symmetrical hanging pans of a balance scale resting on the letters, violet on the outside stepping inward through blue, green, yellow, orange to red, on a clean off-white background with generous negative space, no shadows, no gradients, no glow.
Can we design legal agent verifiers that are up to 1,000x cheaper? Verifiers are LLM judges that check an agent’s work against rubric criteria: they’re used both in agent benchmarking and as reward signal in post-training. But verifiers can be a bottleneck at scale. For
https://x.com/harvey/status/2061866491033899371
Expanding Project Glasswing \ Anthropic
https://www.anthropic.com/news/expanding-project-glasswing
Our internal data shows Claude is accelerating AI development–a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention.
https://x.com/AnthropicAI/status/2062568862479208923
Promoting Advanced Artificial Intelligence Innovation and Security – The White House
https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/
Law professors wrote questions they were asked during office hours. Gemini 2.5 & humans answered them then other law professors blindly judged the results: -Gemini had a 75% win rate vs. professors -Gemini’s answers were rated LESS harmful than humans -Newer models do even better
https://x.com/emollick/status/2061876620638486584
Legendary filmmaker Martin Scorsese signed on last year as an adviser to Black Forest Labs, the German AI startup behind FLUX image models. On Tuesday he went public, testing the tool on a single scene during preproduction. His use is narrow: storyboarding only, complementing
https://x.com/TheRundownAI/status/2061834880917357011
Dreaming: Better memory for a more helpful ChatGPT | OpenAI
https://openai.com/index/chatgpt-memory-dreaming/
OpenAI, DeepMind, Anthropic CEOs back mandatory DNA synthesis screening A coalition of AI leaders, synthesis-industry executives, biosecurity researchers, and former national-security officials published an open letter in June 2026 urging Congress to make screening and
https://x.com/kimmonismus/status/2062485389949145457
Strengthening societal resilience with Rosalind Biodefense | OpenAI
https://openai.com/index/strengthening-societal-resilience-with-rosalind-biodefense/
Verifiers are important for scaling evals/RL But costs add up! So can we make them cheaper? Some great work by @Vtrivedy10 @jakebroekhuizen in conjunction with @nikogrupen @gabepereyra and the Harvey team on this
https://x.com/hwchase17/status/2061867746141356427
I see videos like this and get excited… it’s the old guard embracing new tech. Then I remember the polarizing reaction ahead – perhaps Scorsese is impervious to such pressures?
https://x.com/bilawalsidhu/status/2061811752786944074
Martin Scorsese × Black Forest Labs
https://bfl.ai/martin-scorsese-bfl-advisor
Today, we’re launching shift. We’re starting by cleaning your apartment in New York City, for free. Here’s how it works. Book a shift cleaning. A vetted shift operator comes to your home wearing one of our devices. They clean. They leave. You pay nothing. In exchange, we record
https://x.com/joinshiftX/status/2060044783519735987?s=20
When it comes to observability for agents, regular application tracing doesn’t cut it. We need tools that understand the specific semantics of agents (multi-turn sessions, tool calls, long context, etc.). This is what we built the new Weave for. Agent-first observability and
https://x.com/neutralino1/status/2061949197851742525
Role-specific plugins in Codex are built around the work teams actually do. Plugins for Data Analytics, Creative Production, and Product Design give Codex the tools and context to create reports, creative directions, and prototypes. Built and used by OpenAI teams.
https://x.com/OpenAIDevs/status/2061888366791246071
.@MukilLoganathan’s Interrupt keynote on Sandboxes.
https://t.co/oddQOs0Q6O In 20 minutes, you’ll learn how to run agent code safely. Isolated from your runtime, with network controls, persistent state, and snapshot/restore when things go wrong.
https://x.com/LangChain/status/2061448130806116827
Another thing about AI writing is that while a single instance of AI writing on a topic may be fine, any situation where lots of people use AI to respond to a particular prompt (comments sections, homework, admissions essays) the similarities among responses is tediously obvious.
https://x.com/emollick/status/2061799709275017711
Starting to be suspicious of any post that uses the words genuine or honest.
https://x.com/emollick/status/2062249984574005351
/goal and other fully automated AI agents are cool, but not a great model for the future of work with people. Instead you want your AI to know when to ask you GOOD questions, maybe because it is stuck, maybe because your taste matters, maybe because you would find it interesting.
https://x.com/emollick/status/2061192810422796321
Most people, including really accomplished people, don’t have an accurate mental model of how LLMs operate (and why would they?) You see this in wide beliefs that AI is just copying from known sources, or that it only produces average answers, or that it can’t generate new ideas
https://x.com/emollick/status/2062208940658618605
In WeirdML we see opus models increasingly use submissions to just explore the data, without actually trying to solve the problem (no predictions for the test set). It seems like, with no or low thinking, at least for some tasks, the prior for Opus to just explore the data
https://x.com/htihle/status/2061412097720774679
None of this guarantees recursive self-improvement is on the horizon. It’s not yet clear that Claude is capable of research judgment–of choosing the right problems to work on. But if these trends continue, AI systems designing and building their own successors is plausible. This
https://x.com/AnthropicAI/status/2062568873321513443
OPUS PSYCHOSIS–Claudes Opus 4.6 and 4.7 make stuff up all the time, constantly. Using Opus too much gives you AI psychosis, it makes you believe in fringe scientific and medical theories. I think it’s a very serious credibility and reliability problem for non-coding Claude usage
https://x.com/distributionat/status/2061362406971060244
Anthropic modified its RSP (v3.3) to raise the bio/chemical threshold: it’s no longer enough for a model to significantly help malicious actors; now it must functionally replace the rare expertise of top-tier world specialists. Anthropic calls this a mere revision; Zvi and Opus
https://x.com/CRSegerie/status/2062474945377218819
AI is advancing quickly. Society’s ability to manage its risks must advance just as fast. Today we’re sharing our vision for AI Resilience, with more than $130M in initial grants underway across bio-resilience, cyber-resilience, AI model safety, and AI’s impact on young people:
https://x.com/FoundationOAI/status/2061463726407155795
40 years… just imagine! How many ideas are sitting in someone’s garage right now… waiting for the technology to finally bring them to life? An old patent collected dust since 1985. William Freeman sketched a three-sided zipper. A fastener that shifts objects from flexible
https://x.com/IlirAliu_/status/2060059025363128784
Lem & Douglas Adams got AI right Presciently Golem XIV (from 1981) has an illustration of the jagged frontier as explained by an AI, Golem (GENERAL OPERATOR, LONG-RANGE, ETHICALLY STABILIZED, MULTIMODELING), discussing itself and a smarter AI (Honest Annie) compared to people
https://x.com/emollick/status/2059847527105462363
US moves to close the loophole letting Nvidia’s top chips reach Chinese firms abroad
https://thenextweb.com/news/us-nvidia-amd-china-subsidiary-chip-curbs
In the “”no prior”” live podcast, too, a significant amount of time is devoted to discussing the communities and data center expansion. There appears to be truly immense resistance. This topic occupies a substantial portion of the discussion. It is repeatedly stated that data
https://x.com/kimmonismus/status/2061903253890330639
Opinion | Bernie Sanders: A.I. Belongs to the People, Not to Billionaires – The New York Times
https://www.nytimes.com/2026/06/01/opinion/artificial-intelligence-bernie-sanders.html
Hackers Simply Asked Meta AI to Give Them Access to High-Profile Instagram Accounts. It Worked
https://www.404media.co/hackers-simply-asked-meta-ai-to-give-them-access-to-high-profile-instagram-accounts-it-worked/
We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-5.5’s agentic coding and tool use together with stronger intelligence for drug discovery, analysis, design, and experimental workflows.
https://x.com/OpenAI/status/2062281977122996256
We’re taking steps to accelerate defensive progress in biology: – Launching Rosalind Biodefense to help trusted builders develop new biodefense and pandemic preparedness capabilities. – Expanding trusted access to GPT-Rosalind for select U.S. government and allied partners
https://x.com/OpenAI/status/2060376598642405492
ChatGPT memory research has evolved over the past three years from saved memory, to introducing dreaming, to dreaming V3 rolling out today (first to Plus and Pro — but something here for Free soon!).
https://x.com/ChristinaHartW/status/2062585124450172956
we shipped a new version of gpt-5.5 instant today. the previous model was too bullet pilled. the new one improves on some other important dimensions: sycophancy, factuality, and multilingual performance. hope you’ll like it! always interested in feedback
https://x.com/michpokrass/status/2060219759682330970
We’ve been researching new ways for ChatGPT memory to carry context across conversations and keep it useful over time. Today, that work is rolling out as a more capable memory system in ChatGPT.
https://x.com/OpenAI/status/2062567556524003631
With the new memory system, you can review and steer what ChatGPT remembers through a memory summary, with more visibility and control over how context is used.
https://x.com/OpenAI/status/2062567559673856346
There’s real momentum right now for AI safety policy. Yesterday’s EO on cyber was an important step forward. We’re proposing a set of ideas for policymakers to consider next and to put the US out in front on frontier safety.
https://x.com/OpenAINewsroom/status/2062225755854274629
China rolled out a national digital ID system for humanoid robots in late May. The stated rationale is liability and traceability. – Mandatory 29-digit code assigned to each humanoid: 2-digit country, 4-digit manufacturer, 6-digit product model, 17-digit serial – Tracks the
https://x.com/TheHumanoidHub/status/2062276061862470071
Protecting against token theft – Vercel
https://vercel.com/blog/protecting-against-token-theft
// Reusable Context Engineering // Context bloat quietly kills long-horizon runs, but you can fix it from the outside without fine-tuning the underlying agent. (bookmark this) Context management is usually baked into an agent’s own prompt or weights, which does not transfer
https://x.com/dair_ai/status/2061455253325971789





Leave a Reply