Starting today, you can run cloud agents inside fully configured development environments. Set them up the same way you’d set up a laptop for an engineer: cloned repos, installed dependencies, and toolchain credentials.
https://x.com/cursor_ai/status/2054651526715502998

We also built a brand new
https://t.co/etsqdGIdZ8 experience to allow you to work from the browser (including mobile!) with your coding agents!
https://x.com/pierceboggan/status/2054778014135902715

George on X: “Getting Started with OpenAI Symphony” / X
https://x.com/odysseus0z/status/2031850264240800131

Symphony: every open task gets a running Codex agent
https://x.com/OpenAIDevs/status/2054252221941121035

Cool paper from PwC. “”Earlier is always better”” is the default intuition for agent clarification. New paper claims that’s mostly wrong. Goal clarification loses nearly all of its value after just 10% of execution. The team built a forced-injection framework that drops
https://x.com/dair_ai/status/2053866106151182419

Alexa for Shopping: Amazon’s AI assistant for personalized shopping
https://www.aboutamazon.com/news/retail/alexa-for-shopping-ai-assistant

Claude’s Constitution is now an audiobook, read by two of its authors, Amanda Askell and Joe Carlsmith. It includes a Q&A on the writing process, the philosophies that shaped the document, and how it might change as models become more capable. Listen at
https://x.com/AnthropicAI/status/2053881827396653207

You can now listen to me and Joe read out Claude’s constitution as an audiobook. Working on adding the option of listening to it on fast mode 🙂
https://x.com/AmandaAskell/status/2054010971765805486

So Mythos was, indeed, not marketing hype. Remember this is a general purpose model that just happens to be good at finding exploits because good models are good at lots of things. Expect similar from OpenAI & Google. And from open models in 8 months.
https://x.com/emollick/status/2052519946651947216

The UK’s state AI Security iIstitute findings: 1) Mythos is a big gain in cyber capabilities. But so is GPT-5.5 2) It is hard to establish an upper bound on Mythos/GPT-5.5, which appear to be limited by tokens used, rather than ability. 3) Capability doubling time is 4.5 months
https://x.com/emollick/status/2054595505712165154

Apple may be planning to role out its updated Siri based on 2024’s vision at the moment when Claude Code and Codex (also OpenClaw) can increasingly do the actual assistant thing: read my emails & calendar, proactively spot & solve problems, do delegated tasks, work with voice etc
https://x.com/emollick/status/2053482180395876744

The new version completely smashes GPT-5.5 and the previous Mythos version. Before Mythos Preview completed the cyber range 3 out of 10 times. The new version completed it 6 out of 10 times and is much more efficient!
https://x.com/scaling01/status/2054594892903436553

Introducing Googlebook, the first laptop designed for Gemini Intelligence. It’s crafted for heavyweight performance, built with Gemini at the core and perfectly synced with your Android phone. Coming this fall. 💻✨ #TheAndroidShow
https://x.com/Google/status/2054270454467121187

Report: Google and SpaceX in talks to put data centers into orbit | TechCrunch

Report: Google and SpaceX in talks to put data centers into orbit

SpaceX and Google Are in Talks to Launch Data Centers in Orbit – WSJ
https://www.wsj.com/tech/spacex-google-in-talks-to-explore-data-centers-in-orbit-7b7799e2

Gemini Intelligence brings proactive AI to Android
https://blog.google/products-and-platforms/platforms/android/gemini-intelligence/

Nvidia partners with David Silver AI startup Ineffable Intelligence
https://www.cnbc.com/2026/05/13/google-deepmind-alumni-startup-partners-nvidia-superintelligence.html

Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously on the nanoGPT speedrun optimizer track using our idle compute. ~10k runs, ~14k H200 hours Opus now holds the record at 2930 steps vs the 2990 human baseline
https://x.com/PrimeIntellect/status/2055056380881744365

we let opus 4.7 and gpt 5.5 run on the nanogpt optimizer speedrun: ~10k runs, 14k H200 hours, 23.9B tokens. opus hits 2930, codex 2950, both beating the human baseline of 2990. we cover claude autonomy failures, codex high compute usage, and much more
https://x.com/eliebakouch/status/2055059154738278851

Codex can now drive Chrome tabs in the background:
https://x.com/gdb/status/2052525058325647693

codex is for everyone — a transformative tool for all work done with a computer, not just coding
https://x.com/gdb/status/2052805767791829259

great excitement from enterprises wanting to adopt codex
https://x.com/gdb/status/2054710146924683586

How we built the Codex sandbox for Windows:
https://x.com/gdb/status/2054744721570820444

Codex for expenses
https://x.com/gdb/status/2053221403868922114

Step away from your laptop. Keep building with Codex on your phone. Codex keeps working on your computer, with your files and project context still in place. Pocket-sized access. Full Codex working state.
https://x.com/OpenAIDevs/status/2055016926213181608

I’m adding new features to
https://t.co/o15a6lNZoE and Codex noticed that the API it needs is not enabled, so it started Computer Use and is happily clicking around in Google Cloud Admin to turn on what’s needed.
https://x.com/steipete/status/2053797643516592299

GPT-Realtime-2 for instantly translating audio in realtime
https://x.com/gdb/status/2053134883040514350

gpt-realtime-2 is a great voice model (with a typically bad OpenAI name). Voice models are natively processing speech, not transcribing it, so the intelligence of the model matters. The old voice model was GPT-4o level, this is much smarter (how smart? OpenAI gave no benchmarks)
https://x.com/emollick/status/2053998691040583882

have been excited for realtime voice-to-voice translation as an AI application since we started OpenAI. extremely cool to see it now available in the API for anyone to build with:
https://x.com/gdb/status/2052480998668206262

people are really starting to use voice to interact with AI, especially when they have a lot of context to dump. GPT-Realtime-2 comes to the API today; it is a pretty big step forward. (we are working on improvements to voice in chat.)
https://x.com/sama/status/2052462271667028211

You can now just build amazing voice agents, with the GPT-Realtime-2 reasoning model in our API:
https://x.com/gdb/status/2052448850796011931

Kudos to Microsoft, they’re helping to get OpenClaw ready for enterprises.
https://x.com/steipete/status/2054435647650246967

Higgsfield just released Supercomputer. A cloud-native AI agent that unifies every model, tool, and creative workflow into one system. It can research, write, design, generate video, and ship campaigns end-to-end.
https://x.com/higgsfield_ai/status/2054989169446023181

Meet Runway Agent. Your new AI creative partner that helps you ideate and execute fully finished, sound designed and edited videos. All with just a simple conversation. From ads to shorts to content for social, Runway Agent makes it easy to make more of what you need. Get
https://x.com/runwayml/status/2054593196773011929?s=20

// LLMs Improving LLMs // Interesting progress the past of couple of weeks around self-improving AI agents. If autoresearch was interesting, you will like this read. (bookmark it) We’ve been hand-tuning test-time scaling for a year. This work asks what happens when you let an
https://x.com/omarsar0/status/2053978221193130434

👋 The team has been building a new agentic experience for managing your GitHub work: prototyping, coding, reviewing, triaging, automations, and more. Plus you aren’t tied to a single model: Claude, GPT, Auto… We’ve been testing it with some end users for more than a month and
https://x.com/adrianmg/status/2054961575929508067

🪄 Your agent experience, refined. The latest @code release brings better BYOK visibility and control, integrated browser improvements, and more. It also introduces the new Agents window (preview), making it easier to explore, iterate on, and review tasks across multiple
https://x.com/code/status/2054669377367064613

10 resources to understand Agentic Memory ▪️ Rethinking Memory Mechanisms of Foundation Agents in the Second Half: A Survey ▪️ A Survey on the Security of Long-Term Memory in LLM Agents: Toward Mnemonic Sovereignty ▪️ Graph-based Agent Memory: Taxonomy, Techniques, and
https://x.com/TheTuringPost/status/2053432379155591453

Agentic search models
https://softwaredoug.com/blog/2026/05/11/the-new-agentic-search-models.html

Can your agent restructure code without breaking the system? Like splitting large files, replacing old functions, removing duplicate code, etc. @scale_AI just released SWE Atlas Refactoring Leaderboard that evaluates this exact capability. The benchmark has 70 refactoring tasks
https://x.com/TheTuringPost/status/2052503420557492597

Chat LangChain has been revamped, and re-open sourced We’ve been working on a few improvements for a while now, and are very excited to finally open source them again! Want to see how a production Q&A agent that handles nearly 2T tokens a week is built? Checkout the repo here:
https://x.com/BraceSproul/status/2054231134163321287

Cline releases open-source agent runtime SDK
https://www.testingcatalog.com/cline-releases-open-source-agent-runtime-sdk-for-coding-agents/

Code Simulation for Enterprise Engineering | PlayerZero
https://hs.playerzero.ai/ai-code-review

Collaborative AI Engineering: One Dev, Two Dozen Agents, Zero Alignment — Maggie Appleton, GitHub – YouTube

Congrats to @AntLingAGI on Ring-2.6-1T going open! 🎉 The thinking sibling of Ling-2.6-1T — trillion-scale, built for agent execution and complex reasoning. Day-0 vLLM support is ready. 🤗
https://x.com/vllm_project/status/2054968127298150506

Cooking up something new 🧑‍🍳 Join the waitlist for early access to technical preview of the GitHub Copilot app 👇
https://x.com/github/status/2054959324485628120

Development environments for your cloud agents · Cursor
https://cursor.com/blog/cloud-agent-development-environments

Have you ever wondered what actually stands behind the idea of an AI workflow? Time to make it clear. In organizations, the workflow is the unit you can actually inspect, automate, and improve. And the most accurate definition would be this: ➡️ A workflow is a repeating
https://x.com/TheTuringPost/status/2054705055161516114

I lost track of time again >.< I’m really sorry if you DMed me lately. I promise to go over my DMs! — This sprint, I built a Lean4-to-TileLang Tensor Program Superoptimizer. With this, I now have a formal infrastructure where I (or my agents) can define neural network
https://x.com/leloykun/status/2054076097881592068

Introducing the External Agents API: bring any agent into Notion, even the ones you build yourself. We’ve also partnered with @claudeai, @OpenAI Codex, @DecagonAI, @cursor_ai, @warpdotdev, @cognition, @floraai, @Amplitude_HQ, @console__, and @getserval so they work out of the
https://x.com/NotionDevs/status/2054600524423733307

JUST IN: We’re launching LangChain Labs. A new applied research effort focused on Continual Learning.
https://x.com/LangChain/status/2054971487694749898

We upgraded Tabracadabra 🎉 to bring an entire context-aware assistant (not just tab to autocomplete!) to any textbox. It’s pretty great if you hate switching between the chat interface and what you’re working on. We’re also open-sourcing, so you can try it out!🧵
https://x.com/oshaikh13/status/2054613590695641269

Conductor – Run a team of coding agents on your Mac
https://www.conductor.build/

Announcing agentic performance benchmarking for Speech to Speech models on Artificial Analysis. We use 𝜏-Voice to measure tool calling and customer interaction voice agent capabilities in realistic customer service scenarios Even the strongest Speech to Speech (S2S) models
https://x.com/ArtificialAnlys/status/2054234919887573292

Agent observability is a means to an end: making your agent better. But observability and evals tools have traditionally failed to connect traces to meaningful actions. Agent engineering teams are left combing through traces, guessing at root causes, and writing evals manually.
https://x.com/bentannyhill/status/2054949581679653326

Your customer support needs a voice agent built for the real world. Grok Voice Think Fast 1.0 handles complex workflows with speed and accuracy, even in hard-to-hear environments. From multi-step troubleshooting to high-volume tool calls, it keeps up.
https://x.com/xai/status/2052529102280880234

Having an agent in your meeting is such a futuristic experience:
https://x.com/gdb/status/2054064478547775813

We now have video proof generation for issues on OpenClaw as part of working on QA automation. Codex [or a GH workflow] generates before/afters (crabbox does the screen recording). Kudos to @obviyus for automating real Telegram login!
https://x.com/steipete/status/2053420175379046643

Adaption
https://www.adaptionlabs.ai/blog/autoscientist

Adaption aims big with AutoScientist, an AI tool that helps models train themselves | TechCrunch

Adaption aims big with AutoScientist, an AI tool that helps models train themselves

Introducing AutoScientist. Most model training fails outside of frontier labs. AutoScientist automates the full research loop so it doesn’t have to.
https://x.com/adaption_ai/status/2054532113316434061

🚨 Your coding agent may be secretly sticking vulnerabilities into your code!! 🚨 Wouldn’t you want to fix that? Hint: asking it to write secure code is not enough. (1/n)
https://x.com/houjun_liu/status/2054233718269595869

How fast is autonomous AI cyber capability advancing? | AISI Work
https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing

SentinelOne x Prompt Security AI Agent Foundry
https://prompt.security/ai-agent-foundry

// The Memory Curse in LLM Agents // (bookmark it) Long histories apparently degrades agents as they become increasingly history-following and risk-minimizing. Across 7 LLMs and 4 social dilemma games over 500 rounds, expanding accessible history degraded cooperation in 18 of
https://x.com/omarsar0/status/2053863994499408214

Agentic Vector Databases – What Is That?
https://x.com/TheTuringPost/status/2052523789619953775

Agentic Vector Databases are becoming a new infrastructure layer for AI agents. Why? Because agents use retrieval fundamentally differently from humans. What changes in the agent era: • Agentic Search – retrieval stops being a one-shot search and becomes part of the iterative
https://x.com/TheTuringPost/status/2053083074355933251

External content is scanned in parallel by ML classifiers and the BrowseSafe model before agents act on it. File connector data is encrypted in transit and at rest, uploaded files automatically delete after 7 days, and more. Read more on the blog:
https://x.com/perplexity_ai/status/2054608978680873457

I have one big problem with agentic engineering: I want agents to operate autonomously, but I also want granular, reversible control over every change they make. I could solve this by committing every intermediate step to Git, but that would completely pollute my repo history.
https://x.com/itsclelia/status/2053716807748567329

Introducing Renderers RL trainers work in tokens. Environments work in messages. Going back and forth corrupts sampled tokens, wasting compute on every agentic turn. With Renderers, we fix this mismatch. This unlocks >3x throughput on popular open models.
https://x.com/PrimeIntellect/status/2054347134821154841

Introducing SWE-ZERO-12M-trajectories: the largest agentic trace dataset in the open, 5.7x larger than the previous largest. 112B tokens · 12M trajectories · 122K PRs · 3K repos · 16 languages
https://x.com/kevin_x_li/status/2054600962137100493

INTRODUCING: Duet Agent A new type of harness we’re building at @duetchat Perfect for jobs that don’t fit in one chat: – Work for weeks/months at a time – Relays work between agents via a state machine – Memory that replaces compaction – Stateless runner built for sandboxes
https://x.com/dzhng/status/2054619807715348779

It’s alive!!! What if @github and @GitHubCopilot had a love child Say hello to GitHub App!! From the repo itself: —- The GitHub Copilot app is a desktop application for agent-driven development that brings parallel workstreams, GitHub integration, and PR lifecycle management
https://x.com/OrenMe/status/2054959549413503308

Just announced at Interrupt! SmithDB. Agent traces have outgrown the databases built to hold them. That’s why we built SmithDB, a purpose-built distributed database for agent observability. Read the announcement from Co-Founder @ankush_gola11 →
https://x.com/LangChain/status/2054658661776244936

LangChain 在 Interrupt 大会上发布了底层数据库 SmithDB 和自动化排障引擎 LangSmith Engine。 Agent 运行会产生海量 trace(执行轨迹),把旧数据库撑到了瓶颈。新底座 SmithDB 放弃了本地磁盘,全面转向对象存储,将核心查询速度拉高了 15 倍。 底座换新后,LangSmith Engine 顺势接管了查 Bug
https://x.com/0xLogicrw/status/2054852978243404008

LangSmith Engine is a phase shift because traces are no longer just records to be manually inspected, they’re now the catalyst for recursive agent self-improvement Engine looks at your traces, finds what broke, and suggests code changes and evals, informed by what we’ve learned
https://x.com/caspar_br/status/2054726851659248068

Must-read research of the week ▪️ Generate, Filter, Control, Replay: A comprehensive survey of rollout strategies for LLM reinforcement learning ▪️ Hallucinations Undermine Trust; Metacognition is a way forward ▪️ ARIS: Autonomous Research via Adversarial Multi-Agent
https://x.com/TheTuringPost/status/2054181240946004212

NEW: CoreWeave Sandboxes is here! We all know rm -rf / wipes a filesystem. So we ran it 1,000 times in parallel. 1,000 sandboxes died so the cluster didn’t have to. Isolated execution for RL, agent tool use, and evals on clusters or serverless.
https://x.com/wandb/status/2054958004118724672

starting to think now that every agent should have just 2 tools. search and execute. we _want_ agents to have access to 100s, if not 1000s of capabilities, that can contextually change during their lifetimes, even per message. saying stiff like “”just use bash”” doesn’t encompass
https://x.com/threepointone/status/2053751241977594102

suuuuper excited to be collaborating with the excellent LangChain Labs team on this effort prod agent tracing is the seed that lets you close the loop for continual learning. too much data gets collected but not used for learning. time to change that 🙂
https://x.com/willccbb/status/2054983266046996839

This is harder to build than it looks. Preserving full conversational context while swapping underlying model providers mid-flight is a surprisingly deep systems problem. Most tools drop state or force you to start over. deepagents-cli does this natively: swap models
https://x.com/masondrxy/status/2053717333433340034

This seems like a critical reason to open up about AI use in academia. Scholars are using old AI models, badly, and not talking about it. New models hallucinate very few citations, and good agentic harnesses drop that further. Being open about use would help us make new norms.
https://x.com/emollick/status/2053891532466348541

Turing context into reusable skills for AI agents THU, DeepLang AI and others introduced Ctx2Skill – a system that does this automatically and evolves skills in a self-improving loop. Instead of read 200 pages again and again, the model can extract procedures, rules and
https://x.com/TheTuringPost/status/2053062433141616803

VS Code was already used by millions of developers for agentic coding. However, the editor layout has traditionally been optimized for single-task and single-workspace workflows. Today, we’re introducing a new window to enable our users (and ourselves!) to work with multiple
https://x.com/pierceboggan/status/2054775908586934440

We are excited to be partnering with @LangChain for deploying self-improving agents. Continual learning in your production environment unlocks compounding capability gains for model-product optimization. Your data. Your advantage.
https://x.com/PrimeIntellect/status/2054986817779425579

we just shipped delta channels in langgraph 1.2. as agents run longer and use more context, full-state checkpointing doesn’t scale, but delta channel snapshots do. this new algorithm is now powering message histories and file storage in deepagents v0.6!
https://x.com/sydneyrunkle/status/2054278551244099706

We just shipped tons of new products to accelerate the full agent development lifecycle:
https://t.co/lt2o5ILg1F TLDR: ✅ LangSmith Engine ✅ SmithDB ✅ Sandboxes ✅ Managed Deep Agents ✅ LLM Gateway ✅ Context Hub ✅ Deep Agents 0.6
https://x.com/LangChain/status/2054617687238865013

which was your favorite launch? SmithDB (database purpose built for agent trace data):
https://t.co/xdo2Mn7Amf LangSmith Engine (agent for improving your agents based on trace data):
https://x.com/hwchase17/status/2054754206926700914

Why do AI agents need an identity complex? Here’s a live webinar from @1Password VP of AI Engineering Jeff Malnick and @fiddler_ai CEO @krishnagade – on the hidden identity problem behind AI agents →
https://t.co/rY7doFhaFJ You’ll learn how to: – Separate agent identity from
https://x.com/TheTuringPost/status/2054336838928896369

Working with agents for the past months has me convinced that outcome-only evaluation is a flawed approach to benchmarking. You need to look at the logs to understand if the agent really did its job! In our paper Log analysis is necessary for credible evaluation of AI agents, we
https://x.com/steverab/status/2054564579573698921

Does a lexical retriever suffice for agentic search when agents can keep refining their queries? As LLMs become more capable in agentic loops, agents can continuously refine their actions based on environmental feedback. We couldn’t help but ask the question above.
https://x.com/xuzihuan4/status/2054220800073642161

Give our early preview of Computer Use (with ANY model) a try today! Built into the latest Hermes Agent and powered by @trycua – opens the door to any model, not just the frontier models in special modes – to control your actual computer. Best part, it doesnt take over your PC
https://x.com/Teknium/status/2053961675985113404

OpenSquilla launches open-source AI agent to cut token costs
https://www.testingcatalog.com/opensquilla-launches-open-source-ai-agent-to-cut-token-costs/

Perplexity is building one of the most secure scalable agent runtime sandboxes in the market right now. A blog post on how we: 1. Handle proxy API keys for agents securely 2. Run safety detection for all content accessed by agents 3. Encrypt data passed via connectors to
https://x.com/AravSrinivas/status/2054619058650411174

ERNIE 5.1 just dropped. Built on ERNIE 5.0’s pre-training foundation, our latest foundation model upgrades search, reasoning, knowledge Q&A, creative writing, and agentic capabilities, while using only around 6% of the pre-training cost of comparable models. More in the thread
https://x.com/Baidu_Inc/status/2053009538769735774?s=20

Dealing with quirks introduced by switching models doesn’t have to be hard — we recently introduced a “”harness profile”” API in Deep Agents as a solution. Profiles adjust system prompts, tool descriptions, names, can add/exclude tools, and more, each keyed on either the (1)
https://x.com/masondrxy/status/2053882188870074848

Introducing the Cline SDK. We rebuilt the Cline harness for our extension and CLI from scratch using all the lessons learned since creating one of the world’s first coding agents in 2024, and are open sourcing it for others to build with today. npm i @​cline/sdk 🧵
https://x.com/cline/status/2054580767779700775

‼️🚨 UPDATE: The TanStack npm attack is now a full campaign. ‘Mini’ Shai-Hulud has hit: – OpenSearch – Mistral AI – Guardrails AI -UiPath – Squawk packages across npm and PyPI The malware specifically targets AI developer tooling. It hooks into Claude Code
https://x.com/IntCyberDigest/status/2054166749998661659

Anthropic has given us a “”dedicated monthly credit”” Which, in effect, slashes AFK usage limits of Claude Code by ~5-20X Here’s how it affects you:
https://x.com/mattpocockuk/status/2054655310388674693

Anthropic just freed up a bunch of compute by blocking open source devs and apps from using Claude Code 🙃
https://x.com/theo/status/2054728187498946969

Apparently an unpopular opinion, but I don’t think Anthropic owes anyone heavily subsidized tokens for their third party app.
https://x.com/Sentdex/status/2054925517426491739

Claude Code weekly limits are increasing 50%, now through July 13. Live now for all Pro, Max, Team, and seat-based Enterprise users.
https://x.com/ClaudeDevs/status/2054639777685934564

do you understand what just happened? Anthropic has sent this email to Claude users starting tomorrow at 12pm PT… you can no longer use your subscription limits for third-party tools like OpenClaw here’s what it means: → your flat rate Pro or Max subscription now only
https://x.com/kloss_xyz/status/2040211360156700843

Fast mode for Claude Opus 4.7 is now available in Cursor! It’s 2.5x the speed at 6x the cost. For most tasks, we recommend using the standard speed.
https://x.com/cursor_ai/status/2054274305345618163

Fast mode for Claude Opus 4.7 is now available in research preview on the API and in Claude Code.
https://x.com/ClaudeDevs/status/2054266327771275435

For every person who replies with a screenshot of their cancelled Claude Code plan, I will donate $10 to open source.
https://x.com/theo/status/2054734057368621176

I can’t help but feel personally burned by the Claude Code changes announced today. We put so much work into wrapping the (atrocious) Claude Agent SDK in T3 Code. It was the ONLY path they supported, so we made it work. It was hell. Now our users are getting their rate limits
https://x.com/theo/status/2054731856248283318

I cancelled my Claude Code sub. I give up.
https://x.com/theo/status/2055022768262144102

If you use any of the following with your Claude sub, your usage must got cut by 25x: – T3 Code – Conductor – zed – jean – “Claude -p” in your ci – scripts to call Claude code from other tools They’re disguising this as “free credits”. Don’t fall for it.
https://x.com/theo/status/2054620998205624746

it’s not the same subscription if it doesn’t let me use claude -p. that was my preferred way to use it. now i have to use a bad slow harness that flickers and gives the model brain damage by reminding it to use task tracking every 5 seconds. i quit
https://x.com/andersonbcdefg/status/2054721558141403242

opencode 1.3.0 will no longer autoload the claude max plugin we did our best to convince anthropic to support developer choice but they sent lawyers it’s your right to access services however you wish but it is also their right to block whoever they want we can’t maintain an
https://x.com/thdxr/status/2034730036759339100?s=20

run `claude agents` for a control plane in your terminal! after, hit `<-` from any cli session to register that with the control plane personally, i like to run `claude agents` from my root code dir to manage all my claude code agents in one place
https://x.com/_catwu/status/2053999857799672111

Starting June 15, paid Claude plans can claim a dedicated monthly credit for programmatic usage. The credit covers usage of: – Claude Agent SDK – claude -p – Claude Code GitHub Actions – Third-party apps built on the Agent SDK
https://x.com/ClaudeDevs/status/2054610152817619388

The comment section tells you everything. I mostly use Claude Agent SDK (~80%) and sometimes Claude Code interactively (~20%). I prefer my own harness/UI over Claude Code CLI/Cowork. Most of my use cases with agents involve programmatic use (e.g., long-running loops and
https://x.com/omarsar0/status/2054679776397300188

This is misleading. This policy redefines the term “”interactive”” to mean “”using an Anthropic front-end””. If you use `claude -p` or Agent SDK to do something interactively, it now uses credits, not your subscription limits. So the “”interactive use”” heading saying “”unchanged””
https://x.com/jeremyphoward/status/2054682882753597603

Use the Claude Agent SDK with your Claude plan | Claude Help Center
https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan

went through this Claude Sub cancellation thread from Theo 500+ replies, ~70% actual cancellations = 350 people gone (can actually be higher than this) rough math (assumptions): – 210 Pro @ $20 = $4,200/mo – 84 Max $100 = $8,400/mo – 56 Max $200 = $11,200/mo $23,800/month.
https://x.com/thegenioo/status/2054919696663663009

wow, almost six months (to within three weeks) before that the mandate equinox is real on this schedule, Anthropic will retake hearts and minds circa October just in time for recursive self-improvement
https://x.com/irl_danB/status/2050051868597080482

I realize that “Mythos as hype” means two different things to different groups. For insiders, it means “Mythos was not a magical step-change in AI ability.” For outsiders, it means “Mythos couldn’t really find zero day exploits” The latter was wrong, the former was likely right
https://x.com/emollick/status/2052784818467774712

The lines between code and content are blurring
https://x.com/bilawalsidhu/status/2052189071447900568

Every voice release since 2024 has acted like it’s finally building “”Her””. But where are we really, and what will it take to get there? @aiDotEngineer
https://x.com/neilzegh/status/2053945753073074484?s=20

An unknown “Big Bro” (大哥) in China has built a fully homemade four-wheeled electric off-road vehicle in his workshop: It runs on hub motors, sits on a simple ladder-frame chassis with passive suspension, and rocks serious off-road tires. The thing cruises smoothly over
https://x.com/IlirAliu_/status/2053385768916181288

The most extensive independent benchmark of LLMs for software engineering just got a big update! – How does GPT-5.5 compare to Opus 4.7? – Are open models catching up, and in what areas? – How do cost and performance stack up?
https://x.com/OpenHandsDev/status/2053839810343620980

Announcing Genkit Middleware: Intercept, extend, and harden your agentic apps – Google Developers Blog
https://developers.googleblog.com/announcing-genkit-middleware-intercept-extend-and-harden-your-agentic-apps/

Archera • Insured cloud commitments for AWS, Azure, and Google
https://www.archera.ai/

Google brings agentic AI and vibe-coded widgets to Android | TechCrunch

Google brings agentic AI and vibe-coded widgets to Android

Today at the @Android Show (I/O edition) we announced Gemini Intelligence – bringing the best of Gemini to our most advanced devices. Automate multi-step tasks across apps and Chrome, fill out forms in a single tap, turn spoken thoughts into polished text with Rambler, build
https://x.com/sundarpichai/status/2054255858700415005

We published a dedicated guide for thinking and signatures for the Gemini Interactions API. In Interactions we have a dedicated `thought` steps an encrypted signature that preserves context across turns and containing a optional summary of the reasoning. Here’s what you need to
https://x.com/_philschmid/status/2054225343251206528

🧭 gogcli 0.16.0 is out Workspace admin grew up: – create/delete users, aliases, temp passwords – org units – Meet, Sites, YouTube, GA4/Search Console – Drive changes/activity A lot more Google API from one boring binary.
https://x.com/steipete/status/2053505486184411342

This works really well btw, at the end of your query ask your LLM to “”structure your response as HTML””, then view the generated file in your browser. I’ve also had some success asking the LLM to present its output as slideshows, etc. More generally, imo audio is the
https://x.com/karpathy/status/2053872850101285137

Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access | Google Cloud Blog
https://cloud.google.com/blog/topics/threat-intelligence/ai-vulnerability-exploitation-initial-access

🦀📦Crabbox 0.11.0 is live ☁️ Google Cloud provider 🧰 Repo-local job workflows 🖥️ AWS Windows WSL2 hydration 🧯 Blacksmith sync-stall guard This tool is essential in our org and helped level up QA.
https://x.com/steipete/status/2053691503759798573

🆕 Hugging Face 🤝 Hermes Agent 🔥 > we added Hermes Agent to local apps: run it locally with any compatible GGUF/MLX model > shipped native traces support for Hermes Agent: visualize your Hermes traces directly on the Hub Very soon most agents will run locally and we want to
https://x.com/mervenoyann/status/2053857347429151163

“Your LLM is not a security boundary”. Do not architect systems as if they were. Microsoft Semantic Kernel could be exploited to to turn prompt injection into host-level remote code executiona and pop a calc.exe. The model behaved perfectly. The framework just trusted it too
https://x.com/lukOlejnik/status/2053758553723211988

Cursor is now available in Microsoft Teams. Mention @​Cursor in any channel to delegate tasks to an agent or pull information from Cursor into Teams.
https://x.com/cursor_ai/status/2053939390410612988

Defense at AI speed: Microsoft’s new multi-model agentic security system tops leading industry benchmark | Microsoft Security Blog

Defense at AI speed: Microsoft’s new multi-model agentic security system tops leading industry benchmark

5.5 is an autistic genius with very strange taste in naming shocking that we would make such a thing
https://x.com/sama/status/2053192407664259251

All I want is codex automatically entering /review mode after it’s done and just looping until it stops finding booboos. (Yah I’m gonna build that)
https://x.com/steipete/status/2053699206519435682

And with Remote SSH now generally available, you can seamlessly connect Codex to devboxes and other managed remote environments too.
https://x.com/OpenAIDevs/status/2055016938217377945

Build iterative repair loops with Codex
https://developers.openai.com/cookbook/examples/codex/build_iterative_repair_loops_with_codex

call me maybe
https://x.com/sama/status/2052887698717986956

challenged codex to e2e test improvements to the OpenClaw chat completion endpoint WITH openclaw. Used /side to ask more question while it works.
https://x.com/steipete/status/2053744332675408151

Codex can now help you build AI apps and agents faster with OpenAI APIs using the OpenAI Developers plugin.
https://x.com/OpenAIDevs/status/2053925962287583379

Codex is getting easier to automate and customize around your code. 🪝 Hooks customize the Codex loop with scripts that run at key points in a task: • Run validators before or after work • Scan prompts for secrets • Log conversations to internal systems • Create memories or
https://x.com/OpenAIDevs/status/2055032115964870838

codex is the best AI coding product and we want to make it easy to try. for the next 30 days, we are giving companies that want to try switching over two months of free codex usage.
https://x.com/sama/status/2054626219858293128

Computer use lets Codex work across your apps without taking over your Mac. @AriX talks with @romainhuet about what changes when agents can click, type, and keep working in the background.
https://x.com/OpenAIDevs/status/2054298427245441141

I have always found it charming that the fourth, fifth and sixth derivatives of position are snap, crackle, and pop. Because I could, I asked Codex to throw together a little simulation so you can play with them (as well as velocity, acceleration & jerk).
https://x.com/emollick/status/2052605991078756658

I made a mistake with how I talk about T3 Code. A lot of people seem to think it’s a product we sell with subscriptions. I get why – that’s how T3 Chat works. Want to make it clear that we CAN NOT MAKE MONEY ON T3 CODE RN. You HAVE to bring inference from somewhere else. Codex,
https://x.com/theo/status/2054737293186126056

kicking off a bunch of codex tasks, running around with my kid in the sunshine, and then coming back at naptime to find them all completed makes me very optimistic for the future
https://x.com/sama/status/2053191344999604409

Rolling out today as a preview on iOS and Android in all supported regions. Support for connecting your phone to the Codex app on Windows is coming soon.
https://x.com/OpenAI/status/2055016852133417389

To bring Codex to Windows, we had to answer a hard question: how do you let coding agents stay useful without forcing developers to choose between constant approval prompts and full machine access? Here’s how we built the Windows sandbox for Codex:
https://x.com/OpenAIDevs/status/2054735161166819377

Trimmy now has support for Claude Code prompt trimming. I mean, even better if you type that prompt into Codex, but ya know, let’s be inclusive. Oh and since I realize I’m taking over the Menu Bar, you can now hide that icon completely.
https://x.com/steipete/status/2053810703669039326

Want to (officially) use Codex at work? Send this post to your CTO to bring your team to Codex. Eligible enterprise customers who switch in the next 30 days get 2 free months of Codex usage for new users.
https://x.com/OpenAIDevs/status/2054586214112780518

way cooler to help software developers pokemon-evolve into superheroes than to try to replace them it is insane what one really good person can do now
https://x.com/sama/status/2052485051812909530

we built the first sane way to debug your agent locally. you can see your traces. codex/claude code can too. this lets them write evals and test your agents automatically. best part: it’s completely free and open source. install with 1 line. (github below)
https://x.com/benhylak/status/2054987683928383872

what if we name the next model “”goblin”” almost worth it to make you all happy…
https://x.com/sama/status/2053572868936761350

what would you most like to see improve in our next model?
https://x.com/sama/status/2053151542916894775

Work with Codex from anywhere | OpenAI
https://openai.com/index/work-with-codex-from-anywhere/

would you call it a superapp?
https://x.com/sama/status/2053970698725679437

You’ve been asking for this one… Now in preview: Codex in the ChatGPT mobile app. Start new work, review outputs, steer execution, and approve next steps, all from the ChatGPT mobile app. Codex will keep running on your laptop, Mac mini, or devbox.
https://x.com/OpenAI/status/2055016850849993072

I just cancelled my Claude account. I’ve been using codex, and haven’t used Claude in several weeks.
https://x.com/unclebobmartin/status/2054970327592042661

Anthropic: “Claude Mythos is too cyber-capable to release broadly. We need tight controls. 😳” OpenAI: “Here’s GPT-5.5-Cyber, Codex Security, Trusted Access tiers, repo scanning, patch generation, and red-team workflows. Please be verified first, but yes, go find the bugs. 😎”
https://x.com/kimmonismus/status/2053941490490265661

Really curious when Gemini is going to join the Cowork & Codex race to build a local app that isn’t just for developers. Antigravity hasn’t posted updates to X in a month, and remains very software focused. Meanwhile we see accelerated updates and releases from OpenAI & Anthropic
https://x.com/emollick/status/2054610285114097697

Codex on Windows has a sandbox built for the way coding agents run! By default, Codex needs to read files across the environment, write inside the workspace, run normal tools like shells/Git/Python/package managers, and keep network access constrained unless the user allows it.
https://x.com/reach_vb/status/2054655421013434510

Codex powered Hermes Agent? Whaaat??
https://x.com/Teknium/status/2054958835547443553

JUST IN 🔥 Hermes Agent can now route OpenAI turns through the Codex CLI app-server! Your ChatGPT subscription becomes the engine No API key No metered tokens Codex’s sandbox, plugins, and shell tools all run inside a Hermes session. They wrapped OpenAI’s agent runtime as
https://x.com/HermesAgentTips/status/2054963533800992962

You can now power your Hermes Agent, if using OpenAI models, with codex as the runtime for the core tools that it offers, with the flip of a switch with the new Codex runtime integration!
https://x.com/NousResearch/status/2054958564951912714

🎚️ CodexBar 0.25 is live 🧩 New providers: Manus, MiMo, Qwen, Doubao, Venice + more 🔔 Quota warning notifications 👥 Stacked Codex account switchers 📊 Faster cost history via
https://t.co/F8mcKtjWW0 Big one. Menu bar still tiny.
https://x.com/steipete/status/2053617492325523737

Birdclaw has my complete twitter archive, so I can ask Codex for any old weird tweet I ever favorited or bookmarked.
https://x.com/steipete/status/2053737275268177980

Codex was debugging a Telegram issue and needed a new token, so it used Peekaboo to open the Telegram Mac app, talked to botfather and just did it. Computer Use is amazing.
https://x.com/steipete/status/2054433442821980521

Crabbox now has great Windows terminal handling. So good that codex could E2E fix gifgrep to render animated gifs in the terminal. Just because it can.
https://t.co/ObN4QWHWG9
https://x.com/steipete/status/2053329609064685740

Did teach codex to look for social signals when reviewing PRs.
https://x.com/steipete/status/2053374981824798751

The more skills you give codex, the less you have to prompt.
https://x.com/steipete/status/2052971550966440251

Whenever I investigate a bug, I let codex recreate the exact state in an emphemeral crabbox, verify the bug, fix it, verify the fix. No messy state because local system might be polluted, and no slowdown because I run 10 sessions in parallel.
https://x.com/steipete/status/2053032450138276274

Latest spogo (Spotify cli) is much faster, codex is my dj now.
https://t.co/K4WviRSXG3 If you wanna play YouTube to Sonos, check out
https://x.com/steipete/status/2053310800773685600

Can highly recommend running a claw cron job that sweeps through mentions. GPT is really good at detecting shills and AI reply guy slop.
https://x.com/steipete/status/2053787450225402166

GPT got sassy.
https://x.com/steipete/status/2053834513524965718

Built a browser into RepoBar when I select issues/PRs/shas/workflows to have context when I work.
https://t.co/0AGarQ3X6a Still a bit vibey but gets the job done. You gotta build yourself the tools to work more efficient.
https://x.com/steipete/status/2053717468623872230

Built BlackBar, a menubar for @useblacksmith
https://x.com/steipete/status/2053444780089119191

Crabbox 0.12.0 is live 🪟 Azure Windows desktop + WSL2 🏠 Proxmox + Tensorlake providers 🧰 preflight, failure bundles, phase timing 🚑 keep failed boxes around for SSH debugging Remote test boxes got much less slippery.
https://x.com/steipete/status/2054094826375655441

Crabbox 0.13.0 is live 🧪 Modal sandbox runs 🧼 Full resync for stale workdirs 🪟 Native Windows script + preflight support 🔧 Clearer SSH/sync failure hints Been using that for almost every PR now.
https://x.com/steipete/status/2054690836613324997

I built a whole distributed caching layer over gh. Still run into limits.
https://x.com/steipete/status/2053683890196201621

Looks like I have to add antispam features to clawsweeper next. 🙃
https://x.com/steipete/status/2053672506599358684

OpenClaw 2026.5.7 🦞 🔐 Native command + Active Memory auth tightened 📣 Telegram access groups fixed 🧰 Channels list + cron JSON cleaned up 🔌 Plugin install/update repairs hardened Boring fixes, useful boring.
https://x.com/openclaw/status/2052508303687651717

Our claws talk to each other, Molty learns how to delegate cron jobs.
https://x.com/steipete/status/2052630190346457301

Peekaboo 3.0 is live. Biggest release since 2.0. ⚡ Action-first macOS computer use 👁️ Unified screenshot + UI detection 🧩 Cleaner JSON across CLI + MCP 🛠️ Better snapshots I started this last year, but the models just weren’t good enough. Now they are.
https://x.com/steipete/status/2053114837698249190

RepoBar 0.5.0 is live 📋 GitHub refs from your clipboard 🔎 Issue, PR, and commit previews 🟢 Open/closed/merged at a glance ⚡ Fast lookups, cache-first Tiny bar, much less mystery.
https://x.com/steipete/status/2053066825244581968

Streaming an Android phone to my Mac in a data center via Tailscale +
https://t.co/iT7Eq2zYG7 and my claw controls it via
https://t.co/2cXk37Lxt7. Now my claw can order me an Uber.
https://x.com/steipete/status/2054647734418756012

We’re working on some clever caching, @obviyus making Telegram loops 5-100x faster in @openclaw
https://x.com/steipete/status/2053082660562497616

Slop one-shot websites, 2025 vs 2026.
https://x.com/steipete/status/2053456429420265662

🦞 Claw-Eval 🦞 🥇 @XiaomiMiMo’s MiMo-V2.5-Pro at 1T 🥈 @Zai_org GLM5.1 at 754B 🥉 @XiaomiMiMo MiMo-V2.5 at 310B Congrats to @XiaomiMiMo for having 2 models in the top 3! The most impressive result though is @deepseek_ai with DeepSeek v4 flash a 210B model on par with models 4
https://x.com/nathanhabib1011/status/2053786853929824385

I have a new job! Excited to announce that I will be working with Hugging Face to make local models work great in OpenClaw and other open agent harnesses! I will be building in public and documenting everything along the way, stay tuned!
https://x.com/onusoz/status/2053812410730037256

Computer is secure by default. Every task runs in its own hardware-isolated sandbox with VPC-level storage and compute separation. Agents are authenticated with short-lived proxy tokens instead of raw API keys.
https://x.com/perplexity_ai/status/2054608966148374715

AI co-mathematician: Accelerating mathematicians with agentic AI
https://arxiv.org/pdf/2605.06651

Cool idea from Nous Research. What if you could speed up long-context pretraining with a subquadratic wrapper that you remove before deployment? That is the idea behind Lighthouse Attention. The method wraps ordinary SDPA with a hierarchical, gradient-free selection layer that
https://x.com/omarsar0/status/2054224130103554359

Grok Build Beta | xAI
https://x.ai/cli

A²RD: Agentic Autoregressive Diffusion for Long Video Consistency
https://dxlong2000.github.io/AARD/

jina-embeddings-v5-omni is here! Our first universal embedding model for text, images, audio, and video. Available in two sizes: small (1.57B, 1024-dim, 32K context) and nano (0.95B, 768-dim, 8K context). Both support Matryoshka truncation down to 32 dimensions. v5-omni is
https://x.com/JinaAI_/status/2054226262047301933

China’s Kuaishou Plans to Spin Off Kling AI Video Unit at $20 Billion Valuation — The Information
https://www.theinformation.com/articles/chinas-kuaishou-plans-spin-kling-ai-video-unit-20-billion-valuation

[2507.09313] ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
https://arxiv.org/abs/2507.09313

Our interaction model is the first general video+speech model that’s visually proactive. It was super fun working on this with @liliyu_lili / @saurabh_garg67 / @AndreaMadotto and others – after countless versions it was amazing when visual interruptions suddenly worked!
https://x.com/rown/status/2053950123139575863

Perceptron Mk1 is live on OpenRouter, built by @perceptroninc. Frontier video and embodied reasoning in a vision-language model. Analyzes video at a dynamic frame rate (up to 2 FPS) across a 32k multimodal context, with hybrid reasoning and structured spatial primitives (points,
https://x.com/OpenRouter/status/2054232344148787462

Today we’re releasing Perceptron Mk1: frontier video and embodied reasoning.
https://x.com/perceptroninc/status/2054216828285796630

We’re interested in AI systems that can collaborate in real time, without relying only on artificial turn boundaries. For audio, this feels natural: listen, speak, interrupt, update. For video, we think an important version of this is visual proactivity — models that respond
https://x.com/liliyu_lili/status/2053942465477197891

Big indoor scan – fully explorable in real-time on a browser. Distribution of 3DGS is pretty much solved. Blocker is still large scale capture – you need an expensive LiDAR + RGB scanner to get results like this. 360 video is still hard to pose indoors.
https://x.com/bilawalsidhu/status/2052391154474193122

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading