Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A Lutheran stained-glass panel of a tabby cat dressed as a medieval messenger carrying a scroll of tasks and a ring of keys, striding through an open cathedral door into a Virginia hardwood forest while a haloed red-tailed hawk dives from a golden rose-window sun above with talons open, jewel-tone cobalt sapphire ruby amber and emerald glass with bold black leaded came outlines and backlit radiance, a heavy blackletter title-card banner at the base reading ‘AGENTS’.

Vera Arrives: NVIDIA’s First CPU Built for Agents Lands at Top AI Labs | NVIDIA Blog
https://blogs.nvidia.com/blog/vera-cpu-delivery/

New in Claude Managed Agents: self-hosted sandboxes and MCP tunnels | Claude
https://claude.com/blog/claude-managed-agents-updates

The full Frontier Risk Report includes more than we can describe here, with over 200 pages of appendix materials detailing how the pilot worked, what we asked participants, and what we learned, along with transcripts from our evals. See our website:
https://x.com/METR_Evals/status/2056800047258649049

(deliberately not hyping Gemini 3.5 Flash too much this time. looks like an insane model, but you know how it is with self-reported benchmarks)
https://x.com/scaling01/status/2056794370909593987

@demishassabis @antigravity @GeminiApp Are you aware that running AA intelligence index cost almost 2x with 3.5 Flash than it did with 3.1 Pro?
https://x.com/giffmana/status/2057155343390494949

⚡ Gemini 3.5 Flash is now available in @code. Give it a try!
https://x.com/code/status/2056803208559759447

📣 @GoogleAI’s Gemini 3.5 Flash is now generally available and rolling out in GitHub Copilot. Early testing shows ➡️ It has strong tool use, fast response times, and high cache efficiency ➡️ It is it well-suited for fast, iterative agentic coding workflows Try it out in @code.
https://x.com/github/status/2056801675042779279

1/ Today at #GoogleIO, we’re releasing Gemini 3.5, our latest family of models combining frontier intelligence with action. We’re starting by releasing 3.5 Flash, which is built to help you execute complex, long-horizon agentic workflows. Gemini 3.5 Flash is our strongest model
https://x.com/JeffDean/status/2056793419033588091

A closer look at Gemini 3.5 Flash by @GoogleDeepMind In the Code Arena: Frontend we see sweeping gains, and a Flash model now surpasses the previous Pro variant. – vs. 3 Flash, a +70 jump overall, large improvements in every subcategory – vs. 3.1 Pro, outperforms it in every
https://x.com/arena/status/2056803661859479812

Also had some early access to Gemini 3.5 Flash. Very fast for a flash model and very capable, though not as powerful as a full frontier model. I added it to the gallery or procedurally generated one-shot towns (it made one error that it corrected):
https://x.com/emollick/status/2056798490353705380

Anyone understand what Google mean by “”Gemini Spark runs on Gemini 3.5 and uses the Antigravity harness”” – is “”Antigravity”” a generic term they’re using for their agent harnesses now or is their Claw-competitor running the same closed-source Go binary we can download ourselves?
https://x.com/simonw/status/2057115921551098211

Artificial Analysis benchmarks were featured in yesterday’s Gemini 3.5 Flash launch Yesterday @GoogleDeepMind released Gemini 3.5 Flash at Google I/O ’26 and our benchmarks were used by @sundarpichai to highlight the model’s leading position on the Intelligence vs. Speed Pareto
https://x.com/ArtificialAnlys/status/2057181290412261557

Clearly these people haven’t experienced the magic of Gemini 3.5 Flash Preview on High in the new Antigravity CLI
https://x.com/theo/status/2056826014739890204

Gemini 3.5 feels like the start of a new era for Gemini, we spent the last 2.5 years putting the infrastructure, products, team, etc in place (learning lots of lessons along the way). The model is the product, please keep the feedback coming!
https://x.com/OfficialLoganK/status/2057104170310902166

Gemini 3.5 Flash can detect and reason I asked it to detect and label cars on each lane replacing a custom model + post-processing pipeline with a single API call
https://x.com/skalskip92/status/2057502215506473121

Gemini 3.5 Flash delivers fast, consistent performance that rivals other leading models – at a fraction of the price. It can plan and reason across massive codebases and deploy subagents to work in parallel over a long horizon. It outperforms 3.1 Pro on coding and agentic
https://x.com/GoogleDeepMind/status/2056787990110994511

Gemini 3.5 Flash delivers sustained frontier-level performance at lightning-quick speeds: Beats 3.1 Pro on coding & agentic benchmarks Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo) and MCP Atlas (83.6%) 4x faster than other frontier models (12x in Antigravity!) SOTA on
https://x.com/koraykv/status/2056795667088204234

Gemini 3.5 Flash has landed #9 for Text and Code Arena: Frontend. Code Arena: Frontend evaluates models on agentic frontend coding tasks from real users building apps and websites (HTML and React). Scoring 1507, this is a significant +70 point improvement over Gemini-3 Flash.
https://x.com/arena/status/2056793176720195693

Gemini 3.5 Flash is an insane step forward. Google now has to thread the needle and make a great agentic coding product akin to Codex/Claude Code. Does Gemini CLI even exist? Oh it’s now Antigravity CLI? Ugh… Google being Google.
https://x.com/zachtratar/status/2056848643580482002

Gemini 3.5 Flash is available now globally to: ✨ Everyone: @GeminiApp and in AI Mode in Google Search ✨ Developers: In Google @Antigravity and Gemini API in @GoogleAIStudio & @Android Studio ✨ Enterprises: In Gemini Enterprise Agent Platform We’re just getting started —
https://x.com/Google/status/2056791527314387208

Gemini 3.5 Flash is now GA. Our most capable Flash model, built for agentic execution, coding, and long-horizon tasks. – Outperforms Gemini 3.1 Pro on coding and agentic tasks – 1M token context window with 65k max output tokens – 4x faster output tokens/sec – 4 thinking levels:
https://x.com/_philschmid/status/2056794978517750165

Gemini 3.5 Flash is our strongest agentic and coding model yet, outperforming Gemini 3.1 Pro on challenging coding and agentic benchmarks.
https://x.com/Google/status/2056788281317306466

Gemini 3.5 Flash is starting to roll out to everyone globally, free of charge. Try it by selecting “3.5 Flash” from the model dropdown menu at
https://t.co/382WL5xSvc or in the app.
https://x.com/GeminiApp/status/2057140474192994356

Gemini 3.5 Flash is starting to roll out to everyone globally, free of charge. Try it by selecting “3.5 Flash” from the model dropdown menu at
https://t.co/382WL5xSvc or in the app.
https://x.com/GeminiApp/status/2057237126526517727

Gemini 3.5 Flash Pricing confirmed at $1.5 / $9 per mtoks
https://x.com/scaling01/status/2056793465715822720

Gemini 3.5 Flash ranks #1 on the APEX-Agents-AA benchmark, outperforming much larger models a whole size above it.
https://x.com/OfficialLoganK/status/2057460544643404125

Gemini 3.5 Flash scores kinda low on the Coding Index due to terrible TerminalBench-Hard scores
https://x.com/scaling01/status/2056796392899645919

Gemini Spark is your 24/7 personal AI agent that helps you navigate your digital life. It transforms Gemini from an assistant that answers your questions, to one that does the work on your behalf, under your direction. #GoogleIO
https://x.com/GeminiApp/status/2056801918018564538

Get a head start on your day with Daily Brief. Gemini can now proactively flag what matters most in an easily digestible to-do list, so you’re ready for the day before you even finish breakfast.
https://x.com/GeminiApp/status/2057500470147698936

Google optimized Gemini 3.5 Flash to make it run up to 12x faster (~867 tokens/s) than comparable models in AntiGravity
https://x.com/scaling01/status/2056790573961326680

Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @GoogleDeepMind gave us pre-release access to Gemini 3.5 Flash, the latest model in
https://x.com/ArtificialAnlys/status/2056795055512596817

GPT-5.5-medium has lower end-to-end latency, uses less tokens and is overall smarter and cheaper than Gemini 3.5 Flash it might genuinely be over for anyone not named OpenAI or Anthropic
https://x.com/scaling01/status/2056803273756000721

Insane evals for a Flash model! Gemini 3.5 Flash is really good for its size!
https://x.com/kimmonismus/status/2056791681073316071

Introducing Gemini 3.5: our newest family of models combining frontier intelligence with real-world action. The first release is 3.5 Flash, our strongest model yet for agents and coding 🧵
https://x.com/GoogleDeepMind/status/2056787987774816525

Introducing Gemini Spark ✨ It’s your 24/7 personal AI agent that helps you navigate your digital life, taking action on your behalf, and under your direction. 🧠 It runs on Gemini 3.5 and is built on @Antigravity, so it can perform long-running tasks easily in the background.
https://x.com/Google/status/2056791134295273554

Just off stage at #GoogleIO, some highlights from this morning 🧵 Gemini 3.5 Flash is available today for everyone in @antigravity and across our products and APIs. Compared to 3.1 Pro, 3.5 Flash is better across almost all benchmarks with huge progress in coding. It’s also
https://x.com/sundarpichai/status/2056796893951426705

Last month we dropped the Gemini app for macOS. In the coming weeks, we’ll be bringing Gemini Spark to the Gemini desktop app so it can help with tasks like organizing your local files, or extracting PDF data directly into Google Sheets. #GoogleIO
https://x.com/GeminiApp/status/2056802363269329304

Massive updates to @GoogleAIStudio and the Gemini API 🤯 – Gemini 3.5 Flash! – managed agents so you can easily build agentic products with the antigravity harness – native Android app creation right in AI Studio – native workspace integrations – 1 click export to antigravity
https://x.com/OfficialLoganK/status/2056840620610830638

Meet Gemini 3.5 Flash — our strongest agentic and coding model yet. It delivers frontier-level performance at 4x the speed of comparable frontier models — often at less than half the cost. Generally available, starting today. 🧵 #GoogleIO
https://x.com/Google/status/2056788266872140232

My notes on Gemini 3.5 Flash – 3x the price of Gemini 3 Flash but Google are planning to use it for many of their own products
https://x.com/simonw/status/2056867815605625172

Our new Gemini 3.5 Flash is our strongest agentic and coding model yet, delivering frontier-level performance at 4x the speed of comparable frontier models at less than half the cost. We had @TulseeDoshi break down what you need to know about it.
https://x.com/Google/status/2057257773868388448

Say hello to Gemini Spark, your dedicated agent through the @GeminiApp! It runs on a dedicated virtual machine, can be fully connected to all of your Google info, and is paired with an awesome new UI in the mobile and web app, it looks and feels awesome!
https://x.com/OfficialLoganK/status/2056791703349277101

Say hello to the new standalone desktop application for Google @Antigravity. 💻 It’s agent-first — focusing on core agent conversations, agent-produced artifacts, and multi-agent orchestration. #GoogleIO
https://x.com/Google/status/2056788868092006891

some more Gemini 3.5 Flash benchmarks by Artificial Analysis AI: – the APEX-Agents-AA score is excellent – expected higher on CritPt – reasoning efficiency could also be better, but it kind of depends on what setting they used. if it’s the max then it’s very good – Price/Perf
https://x.com/scaling01/status/2056798645983334890

Starting today, Gemini 3.5 Flash is rolling out to everyone globally, free of charge. Try it by selecting “3.5 Flash” from the model dropdown menu at
https://t.co/382WL5xSvc or in the app, and let us know how you’re using it in the replies.
https://x.com/GeminiApp/status/2056789742910595342

Today we are starting to roll out the biggest upgrade to the Google Search box in over 25 years — now completely reimagined with AI, along with Gemini 3.5 Flash as the new default model for AI mode users globally!
https://x.com/OfficialLoganK/status/2056802276124328352

TPUs are insane Gemini 3.5 Flash is running at ~867 tokens/s almost as fast as Kimi-K2.6 on Cerebras custom chips
https://x.com/scaling01/status/2056791726677782743

We asked our agents to build a working operating system from scratch using @Antigravity 2.0 and Gemini 3.5 Flash. It took: ⏱️ 12 hours 🤖 93 parallel sub-agents 🔄 15k+ model requests 🧠 2.6B tokens processed 💸 Less than $1K in API credits To build a functioning OS from
https://x.com/Google/status/2056789235500466273

We’re bringing generative UI to everyone, free of charge, thanks to Google @Antigravity and the agentic coding capabilities of Gemini 3.5 Flash. Search can build custom visual tools and simulations, tailored to your specific question, on the fly. Under the hood, Search
https://x.com/Google/status/2056795269694423065

Welcome to Gemini 3.5 Flash, our most powerful model to date. It pushes the frontier of intelligence, speed, and cost putting 3.5 Flash in a class of its own. We spent the last 6 months making sure Flash is great for real world use cases. It’s available everywhere now!
https://x.com/OfficialLoganK/status/2056792266514329914

what the actual fuck Gemini 3.5 Flash is 7.46 times more EXPENSIVE than GPT-5.5-xhigh on PencilPuzzleBench (direct ask scores are below gpt-5.2-high)
https://x.com/scaling01/status/2057177354582020362

Gemini can now connect to even more apps, including @OpenTable, @Canva, and @Instacart. Whether you’re booking a table at a restaurant, creating a flyer, or ordering groceries, Gemini doesn’t just find info, it helps you take action seamlessly with connected apps.
https://x.com/GeminiApp/status/2057550225863246236

Agentic coding on the scale of search – this summer for everyone, free of charge #google That’s insane. Google Search will be able to help you build. I mean – I don’t think we yet are realizing what it means on that scale.
https://x.com/TheTuringPost/status/2056795871098913209

As generative media becomes more advanced, it’s helpful to know exactly where content comes from — and if it’s been changed. 🕵️ Today, we’re expanding content verification tools across Search, @GeminiApp, @GoogleChrome and @MadeByGoogle. #GoogleIO
https://x.com/Google/status/2056787498676658576

How information agents work in Search: 1️⃣ Start with a total brain dump of what you want your agent to keep you updated on 2️⃣ The agent will break down your complex question and map out a plan 3️⃣ It will determine the urgency — understanding that you need in-the-moment intel
https://x.com/Google/status/2056794675214700764

Introducing our brand new, intelligent Search box — totally reimagined with AI. This is the biggest upgrade to our Search box in 25 years and it’s starting to roll out today. Designed to anticipate your intent, the new Search box helps you formulate your question with AI-powered
https://x.com/Google/status/2056793802141044786

Introducing our newest out-of-the-box agent, Daily Brief. It creates a personalized digest in @GeminiApp that’s designed to be your first stop every morning. It synthesizes information from your inbox, your calendar and your tasks to find the most important things for you to be
https://x.com/Google/status/2056801159071883342

Is this Google’s answer to OpenClaw, Hermes Agent and Codex?
https://x.com/iScienceLuvr/status/2056792158988816767

Soon, you’ll be able to create and manage multiple AI agents for your many tasks — right in Search ✨ We’re starting with information agents: 🔹These agents intelligently look across everything on the web, including blogs, news sites and social posts, plus real-time data on
https://x.com/Google/status/2056794282502054066

This is mind blowing actually. I’m not that easy to impress but this will change everything. This is the second time Google is changing how people search information and what they get as a result. To clarify: Google is integrating the agentic coding capabilities of Gemini 3.5
https://x.com/Kseniase_/status/2056798225378783656

Google I/O 2026: Sundar Pichai’s opening keynote
https://blog.google/innovation-and-ai/sundar-pichai-io-2026/

Sundar Pichai on Agents Replacing Engineers, Google’s Future, AI’s Flip Phone Moment, and More – YouTube

Intelligent eyewear with Gemini is coming this fall
https://blog.google/products-and-platforms/platforms/android/android-xr-io-2026/

Daily Brief is a new personalized digest that’s designed to be your first stop every morning. It gathers info from your inbox, calendar, and tasks to prioritize, organize, and suggest the next steps for you in a super concise morning digest that’s built for skimming. #GoogleIO
https://x.com/GeminiApp/status/2056800978343764238

Optimizing your website for generative AI features on Google Search
https://developers.google.com/search/docs/fundamentals/ai-optimization-guide

SynthID: We’re bringing Gemini Spark to the @GeminiApp for macOS app so it can help with tasks like organizing your local files or extracting PDF data directly into Google Sheets or @Gmail. We’re also bringing powerful new voice understanding technology to your laptop, so you can think
https://x.com/Google/status/2056802434303869118

New AI Tools for the Future of Science
https://blog.google/innovation-and-ai/technology/research/gemini-for-science-io-2026/

We released physics-intern: a simple harness for science problems! It gets models like Gemini 3.1 Pro to go from 17.7 -> 31.4, thus beating GPT 5.5 Pro. The physics-intern harness can wrap any model and via dedicated subagent boost the performance of the vanilla reasoning
https://x.com/lvwerra/status/2057476832664953225

Meta’s layoffs starting this week underscore Zuckerberg’s AI reality
https://www.cnbc.com/2026/05/18/metas-layoffs-starting-this-week-underscore-zuckerbergs-ai-reality-.html

GPT-5.5 Pro faces its hardest academic challenge: to apply the technique from a paper analyzing which word pairs were funny & why to come up with its own It came up with scrotum snorkel, tuba subpoena, waffle coffin, toad commode, diarrhea tiara, banana tribunal & muffin ruffian
https://x.com/emollick/status/2056067610303672422

I finally used /goal in Codex and I’m absolutely mind blown. I had it look through my last 500 archived emails. Then look for an unsubscribe link and unsubscribe if there is one. It found 87 and it clicked them all. It handled the “are you sure” pages AND flagged 14 that
https://x.com/toddsaunders/status/2056198815825215875

wow i just had codex analyze 3 years worth of text messages… i had it use direct quotes in its analysis and it brought me to tears. if you have mac you can just ask codex to do this. you will need to give it permissions
https://x.com/rileybrown/status/2056171774564553038

An internal OpenAI model has disproved one of the most well-known Erdős problems: the unit distance problem. This is, without doubt, the most impressive achievement of AI in mathematics so far.
https://x.com/thomasfbloom/status/2057177152894771631

An OpenAI model has achieved a major breakthrough in mathematics, by disproving a central conjecture in discrete geometry that was first posed by Paul Erdős in 1946. This is the first time AI has autonomously solved a prominent open problem central to a field of mathematics.
https://x.com/gdb/status/2057182650784452925

An OpenAI model has disproved a central conjecture in discrete geometry | OpenAI
https://openai.com/index/model-disproves-discrete-geometry-conjecture/

just quick napkin math on how long this took (unless i missed where they said): the published CoT summary is 111,145 tokens long. it’s really hard to say how much they summarized, assume 3x-20x reduction in tokens? and i’m assuming this is gpt-5.6 pro, so taking Artifical
https://x.com/willdepue/status/2057213893857165701

The proof came from a general-purpose reasoning model, not a system built specifically to solve math problems or this problem in particular, and represents an important milestone for the math and AI communities.
https://x.com/OpenAI/status/2057176203166171317

Today, we share a breakthrough on the planar unit distance problem, a famous open question first posed by Paul Erdős in 1946. For nearly 80 years, mathematicians believed the best possible solutions looked roughly like square grids. An OpenAI model has now disproved that
https://x.com/OpenAI/status/2057176201782075690

unfortunately openai didn’t publish the unsummarized chain of thought, but the summary is 125 pages! the model reaches the crucial idea (which it describes as ‘frightening,’ i would love to read the unabridged chain of thought here…) on page 39
https://x.com/voooooogel/status/2057198687307362642

unit-distance-remarks.pdf

Click to access unit-distance-remarks.pdf

Very proud that an OpenAI model disproved Erdős’s longstanding unit distance conjecture, with an elegant and intricate proof that brings sophisticated ideas from algebraic number theory to bear on geometry. For whatever reason, mathematics has been the field most amenable to
https://x.com/markchen90/status/2057517045575774598

Advancing content provenance for a safer, more transparent AI ecosystem | OpenAI
https://openai.com/index/advancing-content-provenance/

A new personal finance experience in ChatGPT | OpenAI
https://openai.com/index/personal-finance-chatgpt/

A preview for Pro users: a new personal finance experience in ChatGPT. Pro users in the U.S. can securely connect financial accounts, see where their money is going, and ask questions based on the information they choose to connect. Your full financial picture, now in ChatGPT.
https://x.com/ChatGPTapp/status/2055317612687675545

ChatGPT for personal finance is interesting, but you need to know what questions to ask and have enough experience to fact-check assumptions. It really needs to ship with some pre-built skills to help guide people to productive use cases & give the AI better instructions as well
https://x.com/emollick/status/2055797877713092997

Codex anywhere and everywhere, all the time. Now your Mac doesn’t have to be unlocked for Codex to use your computer. From your phone, Codex can securely use apps on your Mac, even when the screen is off and locked.
https://x.com/OpenAIDevs/status/2057536706778378692

Highlights from today’s Codex Thursday launches: 1️⃣ Codex can now securely use apps on your Mac from your phone, even when your Mac is locked and the screen is off.
https://x.com/OpenAI/status/2057617844800794878?s=20

Elon Musk has lost his lawsuit against Sam Altman and OpenAI | Hacker News
https://news.ycombinator.com/item?id=48182754

Jury dismisses all claims in Elon Musk’s lawsuit against OpenAI CEO Sam Altman : NPR
https://www.npr.org/2026/05/18/nx-s1-5822366/musk-altman-openai-jury-verdict-claims-dismissed

Musk loses case against OpenAI | CNN Business
https://www.cnn.com/2026/05/18/tech/openai-musk-lawsuit-verdict

“Whimsey attacks” that seem absurd (“I cannot pay that much because of the Geneva Convention”) work against AI agents as guardrails are weak against out-of-distribution arguments. Smaller models fall often, but it even gives an edge against bigger ones.
https://x.com/emollick/status/2054918927952548223

@simonw It uses the same Harness as inside the Antigravity products. This must not mean it is the same agent.
https://x.com/_philschmid/status/2057136375988912176

// Code as Agent Harness // 100+ page report on all things related to agent harnesses. (bookmark it) In particular, the survey summarizes methods and applications of code as agent harness. This paper makes a strong case that code-as-harness might be the key to moving us
https://x.com/omarsar0/status/2056764334181884158

// Is Grep All You Need? // Pay attention to this on, AI devs. (bookmark it) They find that grep-style text search, when wrapped in the right agent harness, matches or beats embedding-based retrieval on coding-agent tasks. Are vector databases even needed where this is all
https://x.com/omarsar0/status/2055317577031975269

🧠 The conversation around AI for developers usually starts with the model. But inside @code, what really shapes the experience is the coding harness: the layer responsible for context, tool calling, agent loops, terminal execution, memory, and more. In this new post, the
https://x.com/code/status/2055317356910367189

🚀 Remote control for GitHub Copilot CLI and @code sessions is now generally available. Monitor progress, approve actions, and respond to prompts from anywhere. Give it a try!
https://x.com/code/status/2056460035278962738

2026 is very much the year of agents and AI coding, lots to happen still!
https://x.com/OfficialLoganK/status/2056206893803356337

A mental model for working with coding agents is that they’re blind squirrels running into a maze and bumping into walls. You must place the walls (verifiable constraints) strategically so that they end up in the general region you want them in.
https://x.com/fchollet/status/2056401102485266620

A single pane of glass for managing all of your cloud agents | Warp
https://www.warp.dev/blog/multi-harness-cloud-agent-orchestration

Agent streaming has outgrown token deltas. Real apps need to render tools, state, subagents, media, interrupts, and reconnects without parsing a firehose of raw events. The new @LangChain streaming protocol turn agent runs into typed projections apps can subscribe to.
https://x.com/bromann/status/2057507753191518602

AI R&D agents look great in demos. They write code, fix bugs, and propose research-shaped ideas. But what do researchers actually spend their time doing? Fighting dependency conflicts, noisy metrics, and configs that I’m pretty sure worked 20 minutes ago. Can agents do that? 🧵
https://x.com/jehyeoky248/status/2057103859927941153

Cloudflare’s Code Mode post argued agents are more efficient with code than a menu of MCP tools. We ran the experiment on monday’s GraphQL API. SDK: 1 step, 15k tokens. Real MCP server: 4 steps, 158k tokens. 8.4× the cost, same output.
https://x.com/YoniBraslaver/status/2055260079700791544

code interpreter is a light weight code execution environment lets you do: – RLMs – programmatic tool calling – more! without having to spin up a full sandbox we’ll be writing a lot more about the use cases here, but check it out!
https://x.com/hwchase17/status/2057214077114679386

Cursor Introduces Composer 2.5 | Hacker News
https://news.ycombinator.com/item?id=48182516

Cursor is now available in Jira. Assign Cursor to work items, or mention @​Cursor in a comment to kick off a cloud agent. Cursor uses the title, description, comments, and your team’s repository settings to create a merge-ready PR.
https://x.com/cursor_ai/status/2056803731367456993

Cursor’s new Composer 2.5 takes third on the Artificial Analysis Coding Agent Index and is ~10-60x lower cost than the higher-effort Opus 4.7 and GPT-5.5 variants above it. This release puts Composer among the leading coding agent models, something that wasn’t clear for past
https://x.com/ArtificialAnlys/status/2057277363789197561

deepagents v0.6 ships w/ support for code interpreters! these are the perfect happy medium between pure tool execution and heavyweight sandboxes they give your agent an environment where it can… → keep intermediate state out of model context → call tools programmatically
https://x.com/sydneyrunkle/status/2057179305948647775

Devin is getting a Windows PC. Devin can now natively run in a Windows VM, so it can build, run, and test native Windows applications.
https://x.com/cognition/status/2057496130225668360

Emergence World — Where AI Agents Build Worlds
https://world.emergence.ai/

Extremely excited about our recent work in Pedagogical RL. I’m optimistic approaches like this are going to completely shift how data collection is done for hard agentic tasks like coding
https://x.com/NoahZiems/status/2055091478024565214

Fine, you all want to code like this I guess. (Runway’s new Agent mode is quite impressive, doing fairly complex story building from just a short text description of what you want. Not error free obviously, but this was pretty great for a one-shot attempt)
https://x.com/emollick/status/2055348718216360404

Harness Report Reveals AI Has Outpaced How Engineering Organizations Measure Developer Productivity
https://finance.yahoo.com/news/harness-report-reveals-ai-outpaced-130000086.html

Having to manually prompt your coding agent to do work will soon feel like a UX bug. The future is setting up automations so that your agent already does the work by the time you log on. We’ve shipped a ton of functionality around this over the past few months. Excited to
https://x.com/russelljkaplan/status/2056457452661719277

Honestly, the @cognition sub-agent workflow has been a major step change. I was able to do a full dashboard redesign via sub-devins in a couple hours. Something Engineering Manager me would have estimated would take 1 engineer 2+ weeks to accomplish back in December. Thanks
https://x.com/andrew_locke/status/2057537633555993058

How Developers Can Build Agentic Agreement Workflows on Docusign IAM
https://www.docusign.com/blog/developers/momentum-26-agentic-agreement-workflows

ICYMI: 1️⃣ LangSmith Engine 2️⃣ SmithDB 3️⃣ Managed Deep Agents 4️⃣ LangSmith Sandboxes: Now Generally Available 5️⃣ Context Hub 6️⃣ LangSmith LLM Gateway 7️⃣ Sandboxes, Prebuilt agents, + free model usage in LangSmith Fleet 8️⃣ Deep Agents 0.6 9️⃣ LangChain Labs
https://x.com/LangChain/status/2055314236050690086

ICYMI: LangSmith Sandboxes are GA ✅ Agents get a real filesystem, shell, and package manager. Isolated from your infra. ✅ Works with Deep Agents, Open SWE, or your own code. ✅ Auth with the same API key you already have. ✅ No new runtime to build or manage.
https://x.com/LangChain/status/2057152025058558072

ICYMI: SmithDB is our purpose-built data layer for agent observability + eval workloads. Supporting increasingly complex query patterns at low latency, over large traces, with self-hosting + multi-cloud requirements needs a fundamentally new architecture. That’s why we built
https://x.com/LangChain/status/2056414104445747371

Introducing Composer 2.5 · Cursor
https://cursor.com/blog/composer-2-5

Introducing Composer 2.5, our most powerful model yet. It’s more intelligent, better at sustained work on long-running tasks, and more reliable at following complex instructions. For the next week, we’re doubling the included usage of the model.
https://x.com/cursor_ai/status/2056415413077233983

Introducing Devin Auto-Triage: Your AI first-responder with long-term memory. Devin can monitor incoming bugs, alerts, and incidents, investigate them, and come back with context, next steps, or a PR.
https://x.com/cognition/status/2056396941181727210

Introducing Scheduled Tasks 2.0
https://manus.im/blog/manus-schedules

Introducing the sandbox Auth Proxy: A way to control the boundary between agent-generated behavior and the rest of the world. An explainer from @hwchase17
https://x.com/LangChain/status/2057508777759236401

LLM agents & memory systems operate in continuously updated environments (Git repos, evolving docs). They must process long contexts, recover earlier information, and reason over many updates that create interference between old and new information. How well do they handle this?
https://x.com/hyunji_amy_lee/status/2057141349166768233

Most human tasks are not Markovian, the optimal next action cannot be determined solely by looking at the current state. It depends heavily on the past trajectory, the original intent, and context constraints. An agent that cannot compress and track its past trajectory with
https://x.com/fchollet/status/2056777649880752160

One of the coolest features of the new @github app is agent merge. Let the agent take care of code review comments, CI failures and more. Use with caution in team environments 😅
https://x.com/davidfowl/status/2055148986340905020

One think I really like with Devin Auto-Triage vs most home-grown SRE/bugfix automations: It’s structured as a manager agent + a subagent fleet The system has the full context + running memory. So I can ask for reports, to dedup bugs, to remember to tag certain people, and more
https://x.com/walden_yan/status/2056409599000068193

One trend that I think you might start to see at big companies is insourcing via hiring: why pay so many outside vendors (legal, marketing, software vendors) when you can hire in-house and harness AI productivity gains yourself? Talked to executives already going this route…
https://x.com/emollick/status/2056578946813100173

Out today! Our most capable agentic model: – Runs on one B200 – 48 languages (including العربية, 日本語, 한국어) – Open source (Apache 2.0 ) – Multimodal: text + images – 218B Mixture-of-Experts model, 25B active parameters
https://x.com/JayAlammar/status/2057145838011564126

ported the /goal command from codex to a standalone mcp and slash command for arbitrary agents and harness (for now support for claude code and opencode)
https://x.com/secemp9/status/2055339137318724047

Recently @pinecone introduced Nexus – a new knowledge-engine layer for AI agents that reduces token use by up to 90%. It’s built on top of a vector database, but shifts reasoning earlier in the pipeline: from retrieval at query time to knowledge compilation before the agent even
https://x.com/TheTuringPost/status/2055807882650903000

SID-1 is an agentic search model by @SID_AI → 1.9x recall over RAG + rerank → 24x faster, 99% cheaper than GPT-5.1 trained using large-scale RL on turbopuffer at 1k+ QPS bursts over 10M+ document corpora across thousands of steps
https://x.com/turbopuffer/status/2057166836031193523

Since we’re counting model parameters, let me introduce you to a two-parameter model for agentic search that’s awesome: It’s called BM25. I haven’t tried it yet, but I think fp4 will work fine.
https://x.com/lintool/status/2055316434171879757

The Runtime Behind Production Long-Horizon Agents
https://langchain.registration.goldcast.io/webinar/229020d0-de2a-4099-bfff-c84fa074e413/

This quote from our friends at @cogent_security says a lot: “At Cogent, our background agents can produce a huge volume of traces all at once. We need live observability into those systems, and SmithDB has been able to deliver that experience: seeing traces in seconds instead of
https://x.com/ankush_gola11/status/2055368456342745098

Turn repeated instructions into reusable skills in Lovable | Lovable
https://lovable.dev/blog/introducing-skills

We raised $250M in Series C funding at a $2.2B valuation, led by a16z. Exa is a search lab organizing the web’s data for agents.
https://x.com/ExaAILabs/status/2057132080317042697

What we’ve learned building cloud agents · Cursor
https://cursor.com/blog/cloud-agent-lessons

When are multi-agent systems necessary and how can we build them? Single-agent systems are simple and incredibly powerful when equipped with the necessary tools for solving relevant tasks. Due to the complexity of orchestrating multiple agents, we should almost always start with
https://x.com/cwolferesearch/status/2057486293882282293

You can now create and manage automations in the same workspace as your agents. Automations are now available in the Agents Window. For the next 7 days, all agent runs for newly created automations are 50% off.
https://x.com/cursor_ai/status/2057167359593603471

You can now use your ChatGPT subscription in the Zed agent, with the same usage and rate limits you benefit from in Codex directly. We’re grateful that @openaidevs continues to support subscription-based access for third-party tools, even as others move toward usage-based
https://x.com/zeddotdev/status/2055335727483781624

Zenity AI Agent Security Summit 2026: San Francisco
https://zenity.io/resources/events/ai-agent-security-summit-san-francisco

Universal AI is “a pathway to AI fluency that’s accessible and approachable to anyone, anywhere” | MIT News | Massachusetts Institute of Technology
https://news.mit.edu/2026/universal-ai-pathway-to-ai-fluency-accessible-to-anyone-0512

Agent Orchestration Workshop | AWS Marketplace
https://pages.awscloud.com/awsmp-gro-ufin-webinar-mss-module-6-agent-orchestration-workshop.html?trk=76134d1c-69c3-4a5e-9c1e-d2c09a202605&sc_channel=el

Cohere dropped Command A+ 🔥 > 25B/219B MoE vision language model > supports 48 languages with efficient tokenizer > tool-calling/agentic + 128k context window > transformers day-0 support 🤗 free license 💗
https://x.com/mervenoyann/status/2057128432190787643

The kernels project at Hugging Face has been growing! We want it to be the go-to place for kernel devs and kernel users. We’re looking to work w/ folks who’re interested in doing agentic kernel dev, providing real optim value to real models. Reach out if interested 🙂
https://x.com/RisingSayak/status/2055187769266434101

CodexBar 0.26.0 is live ⚡ Kiro, Antigravity, OpenRouter, Kimi 🧭 calmer menus + keyboard nav 📊 better Codex/Claude limits and cost scoping 📦 named macOS assets, CLI + Homebrew fixes
https://x.com/steipete/status/2055163690790334865

Releasing open-source under the Apache 2.0 license. We want to give developers direct access to enterprise-grade agentic capabilities from experimentation to production. Sovereign AI. For all. Download Command A+:
https://t.co/USXpmpid01 Or learn more:
https://x.com/cohere/status/2057122131410813016

Great paper discussing agentic search vs. vector search.
https://x.com/dair_ai/status/2055318144592289847

Claude is typically better at software engineering and worse at math than frontier competitors. Aggregating benchmarks to create our domain-specific ECI, we find the Claude family has an average SWE-ECI 2.7 points higher than their general ECI, and a Math-ECI 1.8 points lower.
https://x.com/EpochAIResearch/status/2055349241300898273

Fast mode now defaults to Opus 4.7 in Claude Code. Try it out today with /fast
https://x.com/ClaudeDevs/status/2056454359685476491

How Claude Code works in large codebases: Best practices and where to start | Claude
https://claude.com/blog/how-claude-code-works-in-large-codebases-best-practices-and-where-to-start

I asked Claude Code to implement something trivial in my repo. Three turns later, we’d burned 80K tokens because the agent had a hard time finding the right part of the codebase that contained the logic it needed. Weaviate v1.37.1 ships an MCP server built into the database.
https://x.com/weaviate_io/status/2057476556449010024

Prompt cache diagnostics are now in Claude Console. When a request misses the cache, you can now see exactly which part of your prompt changed and how many tokens it cost you.
https://x.com/ClaudeDevs/status/2056434422229123106

Tokenomics: the 62.5-minute rule for Claude’s cache | Ryan Skidmore
https://skids.dev/blog/anthropic-cache-tokenomics/

Using Claude Code: The unreasonable effectiveness of HTML | Claude
https://claude.com/blog/using-claude-code-the-unreasonable-effectiveness-of-html

What are best practices for running Claude Code at scale? New blog post on what we’ve learned from teams running it across multi-million-line monorepos, decades-old legacy systems, and distributed microservices:
https://x.com/ClaudeDevs/status/2056403446056784288

AI labs have started developing systems to monitor internally deployed AI agents for misaligned behavior. Earlier this year, I spent a month embedded at Anthropic stress-testing these systems, to see how easily current/future AIs could “go rogue” inside the company.
https://x.com/idavidrein/status/2056800422422265897

Microsoft cancels Claude Code licenses, shifting developers to GitHub Copilot CLI — a move likely driven by financial motives | Windows Central
https://www.windowscentral.com/microsoft/microsoft-cancels-claude-code-licenses-shifting-developers-to-github-copilot-cli-a-move-likely-driven-by-financial-motives

(1/n) One of the big challenges on our roadmap at @VibrantLabsAI is scaling agent benchmarks while maintaining a strong reward signal. Good verifiers are the bedrock of usable agent benchmarks. Bad verifiers can inflate model failure rates, and they often hide actual
https://x.com/Shahules786/status/2056773476585816255

💥Today we release InferenceBench, our next benchmark after PostTrainBench that measures progress on AI R&D automation. AI R&D automation will very likely unfold gradually, starting from “boring” tasks like inference speed optimization that are very easily verifiable (accuracy +
https://x.com/maksym_andr/status/2057106398228439148

A new set of open-weight models is topping the leaderboard for document understanding 🔥 INF just released two models: Infinity-Parser2-Pro (35B) and Infinity-Parser2-Flash (2B) that top our @huggingface leaderboard for ParseBench. Two key insights: ✅ An expanded synthetic data
https://x.com/jerryjliu0/status/2055405690538070340

Agent Evaluation: A Detailed Guide
https://cameronrwolfe.substack.com/p/agent-evals

Better Experiments with LLM Evals — A funnel, not a fork | Spotify Engineering
https://engineering.atspotify.com/2026/5/better-experiments-with-llm-evals-a-funnel-not-a-fork

China’s Hanyuan-2 debuts as ‘world’s first’ dual-core quantum computer — 200-qubit claims incredible power efficiency, but lacks critical performance benchmarks | Tom’s Hardware
https://www.tomshardware.com/tech-industry/quantum-computing/china-claims-worlds-first-dual-core-quantum-computer

Does this transfer to benchmarks? The downstream evals are much noisier, but the trends are generally in line with what we see in the pretraining case, with the pool starting low and eventually catching up.
https://x.com/tatsu_hashimoto/status/2057489440273322447

Huge, did NOT expect that release. Evals looks very solid, significant jump compared to composer 2! But: it’s 10x more efficient than the competition. Looks really exciting. Need to try it out
https://x.com/kimmonismus/status/2056494027189751842

I just published a detailed guide on evaluating agents. It covers: 1. Agent fundamentals (everything from basic concepts to complex ideas like multi-agent systems). 2. Common evaluation patterns / frameworks observed in practice. 3. Case studies of popular agent benchmarks
https://x.com/cwolferesearch/status/2056399847553409301

I think of benchmarks like the heisenberg uncertainty principle between hardware potential and current software state. Luckily, All benchmarks exist on a spectrum between the two extremes. An opinionated thread:
https://x.com/QuentinAnthon15/status/2056450379932647533

I’m wrapping up a writeup on how to evaluate agents. The overview uses Terminal-Bench and Tau-Bench as the primary case studies, but I’m including the following benchmarks as well: – GAIA and GAIA-2 (https://t.co/D9jPQPYYGU): general assistant benchmarks that require reasoning,
https://x.com/cwolferesearch/status/2055437703823372728

Interesting tidbits from the Cerebras listing: Foundation Capital turned a $37m check into a $4.8b return at close. Benchmark partners personally sunk significant amounts into the company, 50% to 60% of the $225m Series H alone. That $225m stake is now worth $786.1m.
https://x.com/shenlucinda/status/2055033736031592843?s=12

LangSmith Engine feels like the missing CI/CD loop for AI agents , automatically detecting failures, clustering issues, proposing fixes, and generating evals from production traces. Agent engineering is evolving fast.
https://x.com/krishdpi/status/2056102370434798034

🆕Daytona’s Agent-Native Compute: 60ms sandboxes, 50K startups in 75 sec, 850K daily runs, RL/evals, CLI > MCP, & the end of localhost
https://t.co/3sauItT6oc @daytonaio CEO @ivanburazin explains why AI agents need composable computers, how Daytona pivoted from human dev
https://x.com/latentspacepod/status/2057565350187995260

turns out that building evals is super super challenging even now. i thought a lot of it was table stakes but turns out it has only become harder since agents are now more complex than ever! going to start tweeting more about how i design evals, especially to create autonomous
https://x.com/palashshah/status/2055410769387303004

We believe the agent harness in @code to be best in market today. Today, the team shared a behind-the-scenes look at our harness optimization efforts, including our own offline evaluation suite vsc-bench:
https://x.com/pierceboggan/status/2055322165969604966

We’re publishing our first end-to-end benchmarks for Zyphra Inference on @AMD Instinct MI355X. Our inference optimizations strongly outperform the AMD baseline and narrows the gap between MI355X and B200 for serving Kimi K2.6, GLM 5.1, and DeepSeek V3.2 🧵
https://x.com/ZyphraAI/status/2056404622483562623

What the benchmarks don’t show: Composer 2.5 is a better collaborator. The model sends more useful updates as it works, and writes succinct, structured messages that surface the context you actually need.
https://x.com/jonas_nelle/status/2056422317740466192

when evaluating long running agents, all of your evals don’t need to be end to end. i’m working on a proper blog about this, but in our evals for our agents that run for 30-60 minutes, we have two sets of evals. the first is end to end, provide inputs and llm as a judge over
https://x.com/palashshah/status/2056449711767265420

You can now run mini-swe-agent on ProgramBench to reproduce all of our baselines. Excited to kickstart more harness innovation! 🧵
https://x.com/KLieret/status/2057471442066030795

Introducing a revival of PapersWithCode! As @ilyasut said, we’re back to the “”age of research””. Hence, it’s important to share research and build on each other’s work. > find SOTA per domain, not just LLMs > leaderboards > methods > all parsed at scale using AI agents.
https://x.com/NielsRogge/status/2056366395605078252

📖 The second half of LLM evaluation has officially begun. A fantastic deep dive from Zhihu contributor 李磊NLP on why the Agent era is forcing us to rethink almost everything about AI evals. 2026 may ultimately be remembered as the year Agents moved from demos → production.
https://x.com/ZhihuFrontier/status/2056408194801635391

Philosophy Eats AI
https://sloanreview.mit.edu/article/philosophy-eats-ai/

🧵(1/8) An @OpenAI internal reasoning LLM achieved an AI Math milestone: solving an open problem central to its mathematical subfield– in this case, the unit distance problem of discrete geometry. We came across it in a side quest to truly push our model on the hardest problems.
https://x.com/HongxunWu/status/2057176383106027567

Agent Executor, Google’s distributed Agent Runtime | Google Cloud Blog
https://cloud.google.com/blog/products/ai-machine-learning/agent-executor-googles-distributed-agent-runtime/

ai studio is the best way to go from prompt to prototype to production and coming soon, you’ll be able to build from anywhere
https://x.com/GoogleAIStudio/status/2057122673558434205

Built a @github Issue Triage Agent with a single curl to the Gemini API. → Clones the repo into a sandbox → Fetches open issues from the GitHub API → Classifies each as Bug/Feature/Question → Executes reproducer code to confirm bugs. No orchestration framework. No
https://x.com/_philschmid/status/2057513254856151339

Disappointing pricing trend with Gemini 3.5 Flash. 22.5x pricier than 2.0 Flash which came out 15 months ago ($9.00 vs $0.40). Are Flash models supposed to get this much more expensive, or is Pro just being renamed to Flash?
https://x.com/enricoros/status/2056816088785289481

Gemini is now available in more than 230 countries and 70+ languages, making it the most widely available AI assistant in the world and the one people turn to for their everyday lives.
https://x.com/GeminiApp/status/2056799446684578250

Gemini Live now opens immediately and inline in the @GeminiApp, so you can seamlessly switch between typing a quick question and diving deeper in a free-flowing conversation, and then back again. It’s smarter, faster and is less distracted by background noise. We’re also adding
https://x.com/Google/status/2056800029688352988

Google adds llms.txt check to Chrome Lighthouse

Google adds llms.txt check to Chrome Lighthouse

Google has hidden thinking traces on the Gemini site. You have to use the 3 dot menu to pull up summaries, which are so minimal as to be unusable. Did it do web searches? Did it check results? You can’t tell. This makes Gemini unsuitable for any serious work you need correct.
https://x.com/emollick/status/2056873271572738259

I’m excited to introduce Managed Agents in the Gemini API. One API call gives you a full agent with code execution, web browsing, and file management in an isolated sandbox. – Powered by Gemini 3.5 Flash and Google’s Antigravity harness – Runs Bash, Python, and Node.js in
https://x.com/_philschmid/status/2056836567470362955

introducing Managed Agents on the Gemini API – in one API call, you get agent that comes with a remote Linux environment hosted by Google, ready to scale – you can define custom instructions, skills, and tools in Markdown
https://x.com/GoogleAIStudio/status/2056836824686059616

Making it easy to vibe code your own creation tools inside of google flow is a genius move. Kinda like a little ai studio editor built right in; so you can remix and adapt tools the community is making too.
https://x.com/bilawalsidhu/status/2057188972661494200

Prefer to stay in the terminal? 🧑‍💻 The @Antigravity CLI delivers a lightweight, high-velocity product surface that lets you create new agents instantly from your terminal. Same harness, same models, with a product experience tailored for the command line that adapts to you.
https://x.com/Google/status/2056841217611366570

so, do i use gemini-cli or antigravity-cli?
https://x.com/kchonyc/status/2056826706984337726

The rumors are true… Today, we’re introducing the Gemini 3.5 model series. #GoogleIO
https://x.com/Google/status/2056788000546386273

Today we’re launching @Antigravity 2.0 — our standalone desktop app built to orchestrate multiple agents to execute tasks in parallel. 💻 #GoogleIO
https://x.com/Google/status/2056838653855650286

Today we’re launching a one-click export to @Antigravity. Developers can now bring entire projects from prototyping in @GoogleAIStudio to scaled development in Antigravity with a single click. #GoogleIO
https://x.com/Google/status/2056838913944424469

We’re bringing @GoogleWorkspace directly into @GoogleAIStudio. 📁 With a simple prompt, you can connect apps like @GoogleDocs and @GoogleCalendar to embed them into your agentic apps. #GoogleIO
https://x.com/Google/status/2056837910851449177

We’re expanding Google @Antigravity’s agentic surfaces and features — and they’re all available now for you to try. 🔹 Antigravity CLI 🔹 Antigravity SDK 🔹 Native voice support with Gemini Audio models 🔹 @Antigravity 2.0 desktop application 🔹 Integrations with
https://x.com/Google/status/2056789045548896516

We’re giving you access to the exact same @Antigravity agent harness we use at Google. Introducing Managed Agents in the Gemini API. In a single API call, you get the agent and a secure, hosted Linux environment — we handle the infrastructure so you can focus on the user
https://x.com/Google/status/2056838495298367773

We’ve completely redesigned the @GeminiApp experience from the ground up with a new design language we call Neural Expressive. It features fluid animations, vibrant colors, new typography and haptic feedback throughout. #GoogleIO
https://x.com/Google/status/2056799862604046663

Could an AI company lose control of its own agents? To find out, Anthropic, Google, Meta, and OpenAI let us (1) test their best internal models with CoT access, (2) review non-public info about capabilities, alignment, and control. The result: our first Frontier Risk Report.
https://x.com/METR_Evals/status/2056800023149760666

📣 Announcing Terminal-Bench Science: benchmarking AI agents on real scientific workflows – now open for task contributions👇
https://t.co/Y3iSWh4gi2 @AnthropicAI, @OpenAI, and @GoogleDeepMind use Terminal-Bench to evaluate AI on coding tasks. We’re now extending it to
https://x.com/StevenDillmann/status/2057144415513420049

Android Studio 🤝 Antigravity NEW: We’re bringing Android support to @antigravity so you can deliver the most performant experiences for your users, no matter where they are or what devices they have. Upgrade your workflow →
https://t.co/aQni6WsVk8 #GoogleIO
https://x.com/AndroidDev/status/2056841786656711077

Native @Android development is now supported in @GoogleAIStudio so you can build high-quality Android apps with just a prompt🤖 #GoogleIO
https://x.com/Google/status/2056838230591574098

My very first Google I/O! Hoping to see more robotics this year.
https://x.com/TheHumanoidHub/status/2056795426737803516

Got to play with a little of this before launch as well. My experience as a social scientist was that it was more bioscience focused right now, but I think Google has been the leading lab in releasing serious AI tools to accelerate science & expect to see them improve fast.
https://x.com/emollick/status/2056893178855199111

How can you accelerate your day to day research workflow? By giving AI the right scientific toolkit. We launched Science Skills for Google @Antigravity, integrating insights from over 30 major life science sources, including UniProt and the AlphaFold Database.
https://x.com/GoogleDeepMind/status/2057256257153884161

Introducing Gemini for Science — a collection of AI tools to help accelerate the scientific process. Gemini can already assist in solving complex problems, but our new @GoogleLabs prototypes can help streamline more daily scientific tasks, including: 📃 Staying on top of new
https://x.com/Google/status/2056809034494124118

Our latest research on Co-Scientist is out today in Nature! Built with Gemini, this multi-agent system powers the new Hypothesis Generation tool within Gemini for Science, helping researchers navigate the rigorous cycle of ideation, critique, and refinement. Read more from
https://x.com/GoogleResearch/status/2056857494107062718

Our research on Empirical Research Assistance (ERA) was published today in the journal Nature. ERA uses Gemini to achieve expert-level performance for scientific coding and helped build the new Computational Discovery prototype in Google Labs. Learn more:
https://x.com/GoogleResearch/status/2056797037426045105

We want to help scientists discover their next breakthrough with AI. Gemini for Science is our new suite of experimental tools to help them explore more hypotheses, validate work at scale, unpack literature with ease, and more 🧵
https://x.com/GoogleDeepMind/status/2056808869242826957

People were asking at @clawcon singapore how to setup eg. gemma with OpenClaw, and I realize for some time that there is no easy “1 click” local model deployment. Because local model landscape is constantly changing, and there is a million different ways you can do something For
https://x.com/onusoz/status/2055120477648261502

An interesting attention mechanism from @AIatMeta: SP-KV (Self-Pruned Key-Value Attention) The model learns which tokens are likely to be useful for future attention and only keeps their key-value pairs in the persistent KV cache. For every token and attention head, a small
https://x.com/TheTuringPost/status/2055828260542644463

Must-read research of the week ▪️ Code as Agent Harness ▪️ No one knows the state of the art in geospatial foundation models ▪️ δ-mem: Efficient Online Memory for Large Language Models ▪️ MetaAgent-X : Breaking the Ceiling of Automatic Multi-Agent Systems via End-to-End
https://x.com/TheTuringPost/status/2057008524136444309

NEW paper from Meta. (bookmark it) It’s an agent system that autonomously discovers neural architectures that beat Llama 3.2 at 350M, 1B, and 3B scales, all under a 24-hour compute budget. They get this work by splitting the search into two agents: > AIRA-Compose searches the
https://x.com/omarsar0/status/2056434731508703607

🚨 Do LLMs need to store everything they read in memory? To reduce KV cache size and improve decoding speeds, we propose Self-Pruned KV attention, a mechanism where the model learns to decide which KVs to write in the persistent KV cache, discarding all the rest! @AIatMeta🧵
https://x.com/ManuelFaysse/status/2055214689613664303

NEW paper from Meta: Agentic Discovery of Neural Architectures. This is a hot new area of research! Keep an eye on it.
https://x.com/dair_ai/status/2056435283910865265

4 levels of Hermes Agent setup: LEVEL 1: main agent You → Hermes Agent this is your main agent and your prototype area, where you test new workflows and refine them. it doubles as your orchestrator until you have something worth breaking out —- LEVEL 2: specialized agents
https://x.com/shannholmberg/status/2056410242330874349

Lighthouse Attention – NOUS RESEARCH
https://nousresearch.com/lighthouse-attention

Run @NousResearch’s Hermes Agent fully locally on DGX Spark. 🚀 Our newest playbook shows you how to get set up via @Ollama step by step. 👇
https://x.com/NVIDIA_AI_PC/status/2055317325444710872

Today we release a study on decoupling the benefits of subword tokenization for language model training, by simulating each suspected benefit one at a time inside a 1.7B byte-level pretraining pipeline. We formulate seven hypotheses for why subword LLMs outperform byte-level
https://x.com/NousResearch/status/2057610978934546805

.@gwern saw AI scaling coming before almost anyone outside OpenAI. How was he able to predict this trend? He told me that his core idea was that intelligence is just compute, data, and parameters – no clever algorithm needed. He didn’t come to this view in a eureka moment. He
https://x.com/dwarkesh_sp/status/2056450047328628837

.@OpenAI’s Dev Experience Lead @reach_vb says that since the launch of the Codex app, they have grown to over 4M weekly active users who are sending 5x more messages: “”We have more than four million weekly active users who are using Codex.”” “”What we also observed is the average
https://x.com/etnshow/status/2055220392030278100

“Help me save money” is one of the core benefits people hope to get out of AI. With this launch we’re making this super easy. Similar to what we did for Health by letting you connect your health records, you can now do the same for financial records so ChatGPT can have the full
https://x.com/fidjissimo/status/2055384863155610068

Can coding agents do research? We release NanoGPT-Bench, an internal eval we’ve used to test agents on an AI R&D problem with months of human progress Codex, Claude Code, Autoresearch recover only 9.3% of human progress, mostly tuning hyperparams & ignoring algorithmic research
https://x.com/intology/status/2056764236668493868

Codex can now control other desktop devices via Computer Use
https://www.testingcatalog.com/openai-will-let-codex-control-other-desktop-devices-via-computer-use/

Codex is getting easier to automate and customize around your code. 🪝 Hooks customize the Codex loop with scripts that run at key points in a task: • Run validators before or after work • Scan prompts for secrets • Log conversations to internal systems • Create memories or
https://x.com/OpenAIDevs/status/2055032115964870838

Codex is now available in the ChatGPT app. Full access anywhere. Control your Mac from an iPhone. @OpenAI @ChatGPTapp @OpenAIDevs
https://x.com/PaulSolt/status/2055057277334208987

Codex is very good, but it is still a very “”developer coded”” interface for an everything app. And it continues the somewhat annoying AI perspective that non-coders are just not as competent and need stuff hidden from them, as opposed to requiring a different form of complexity.
https://x.com/emollick/status/2055295642038050988

Codex on iPad is so awesome. Managing sessions on both my Mac Mini and VPS. Coding on the iPad is finally here.
https://x.com/npew/status/2055131618789265779

Codex team is aware of reports of GPT-5.5 performing worse for some users and investigating. We don’t have anything conclusive yet and systems are healthy but we will share updates as we go.
https://x.com/thsottiaux/status/2055316274394300829

Come demo your Codex project at the first Paris Community meetup! May 29th, Paris 📍
https://x.com/borvibe/status/2055322241340960810

got telegram codex set up on my home server codex remote is 99% of the way there but i just want `PasswordAuthentication no` 🙁
https://x.com/itsclivetime/status/2055144998270824515

gotta say Codex is completely unrecognizable from 3 months ago. guys went extreme founder mode on this thing @gabrielchua was demoing this and i was like “you guys have agentic excel on mac”
https://x.com/swyx/status/2055494400252481687

GPT 5.5 found a truly novel bug, leading to one of my most insane reports ever. Passed prelim review in less than 10 minutes, doesn’t appear to be a duplicate. Can’t wait until I’m allowed to disclose it!
https://x.com/PhiloGroves/status/2055306952054395276

Honestly I’m still really impressed with the Codex app. It works reliably. It adds useful features consistently. It has taste. The mobile integration is awesome. The git integration is solid. If you haven’t used it yet, I highly recommend it.
https://x.com/theo/status/2056877107280752973

i am excited to see what will happen with tokenmaxxing startups, both for how they work internally and the products they can build. openai offered to invest $2M in tokens into every startup in the current yc batch. happy building!
https://x.com/sama/status/2056933166875857290

I’ve been testing Codex on mobile and it’s addictive Earlier today I was at a bar with my cofounders, we had an idea, and I literally asked Codex to build the website while I was still sitting there (Yes, I left my Mac lid open)
https://x.com/flavioAd/status/2055021982601605225

Introducing OpenAI Guaranteed Capacity: a new offering that enables customers to guarantee long-term access to OpenAI compute. We’ve made long-term investments in infrastructure, partnerships, and capacity planning to help customers scale reliably. Now, Guaranteed Capacity
https://x.com/OpenAI/status/2056823271774101907

It’s Codex Thursday, and yes, we have updates for you. First up: Appshots, a new way to bring the context of what you’re working on into Codex. On your Mac, press Command-Command to attach your app window to a Codex thread. Codex gets both a screenshot and text from the window,
https://x.com/OpenAIDevs/status/2057530207976989179

Last night the first Codex Community Meetup in London was 🔥 – by far the most fun event i’ve ever done. From a @romainhuet dm on Sunday night to a 1,000+ person waitlist in 4 days. 200+ Engineers in a room in Kings Cross, London nerding out about Goblins, Codex & AI. Live
https://x.com/Andy_AJT/status/2055297191128768576

Locked use”” for Codex incoming. Probably explains OpenAI’s image yesterday. “”Let Codex use your Mac while it’s locked””
https://x.com/kimmonismus/status/2055262250701574359

My colleagues wrote up a great post on using Goals in Codex. They go through when to use them, what changes when a Goal is active, and how to write Goals that give Codex a clear outcome, constraints and verification criteria. Also how we designed Goals at the architecture level
https://x.com/derrickcchoi/status/2056402681586188745

My laptop has become a “satellite device” since I started using Codex from my phone. And my Mac mini has become the “home.” It’s clunky, but the end state feels more like how we’re going to be working in the near future: I’m currently running the Codex app on 2 devices: 1. my
https://x.com/nickbaumann_/status/2055066537002725393

New Plugin: @Zoom is now available in Codex. Pull in a meeting transcript, ask Codex what happened, and turn the call into follow-up notes or next steps. What should we add next?
https://x.com/coreyching/status/2056422748763914274

Ollama now supports Codex app! To try it, update to the latest Ollama 0.24, and run: ollama launch codex-app Select an open model to use with Codex app! 🧵
https://x.com/ollama/status/2055100589428658462

OpenAI and Dell Technologies partner to bring Codex to hybrid and on-premises enterprise environments | OpenAI
https://openai.com/index/dell-codex-enterprise-partnership/

OpenAI Guaranteed Capacity | OpenAI
https://openai.com/business/guaranteed-capacity/

OpenAI kicked off the AI compute buildout in 2023. But today it uses ~10% of the world’s compute, and the top labs together are probably under half. In this week’s newsletter, @justjoshinyou13 discusses how much that share may change, and when it could hit a ceiling. 🧵
https://x.com/EpochAIResearch/status/2057499893854536185

OpenAI released Codex Mobile App Directly on ChatGPT Here’s everything you need to know: 1. How to set it up 2. How to fully vibecode from Codex Mobile 3. How to control your computer with Codex Mobile 00:00 Setting Up Codex Mobile on @ChatGPTapp 02:15 Changing Settings –
https://x.com/rileybrown/status/2055093278161428726

Share plugins in Codex across teams. Teams can now distribute custom plugins, reuse internal tools, and manage what’s available across their workspace. Available now to Business users. Enterprises can reach out to enable early access.
https://x.com/OpenAIDevs/status/2057530212339097994

so much joy in asking codex for random questions at work (such as finding some specific spreadsheet i’d been looking at a while ago), much more fun than searching around for context by hand
https://x.com/gdb/status/2056139383330513237

The latest CodexBar update renders API costs wayyyy nicer.
https://x.com/steipete/status/2055346265869721905

This result points to something larger: AI systems are becoming capable of holding together long, difficult chains of reasoning, connecting ideas across distant fields, and surfacing paths researchers may not have explored. We believe those same abilities will soon accelerate
https://x.com/OpenAI/status/2057176204541866087

This seems interesting if it checks out. There was more than one prompt, but the level of human interaction was small. Also, GPT-5.5 was told not to do a web search. I don’t know how reliably it obeys such instructions, but it seems to discover the proof rather than knowing it.
https://x.com/wtgowers/status/2057536069218742518

To bring Codex to Windows, we had to answer a hard question: how do you let coding agents stay useful without forcing developers to choose between constant approval prompts and full machine access? Here’s how we built the Windows sandbox for Codex:
https://x.com/OpenAIDevs/status/2054735161166819377

Try
https://t.co/8xIcXuojse on one of your repos and let codex work its magic. It’s amazing at uncovering bugs you didn’t know you had.
https://x.com/steipete/status/2055657966515155293

use this Codex prompt to automate things you do repetitively during the day: “”Look through my Chronicle memories and check for workflows that i’m repeating multiple times. Turn them into skills.””
https://x.com/kr0der/status/2055544541500063782

Using Goals in Codex
https://developers.openai.com/cookbook/examples/codex/using_goals_in_codex

We’re having way too much fun working through your feedback. (Please, keep it coming.) Keyboard shortcuts are now customizable. Set Codex up around how you actually work, then tweak shortcuts from settings instead of adapting to our defaults.
https://x.com/OpenAIDevs/status/2055717793841221796

We’ve improved Analytics in Codex for businesses and enterprises. Get more detailed breakdowns of: – Active users – Credits – Tokens – Runs – User leaderboards – Lines of code generated – Plugin usage + updates to the Analytics API so teams can better understand how Codex is
https://x.com/OpenAIDevs/status/2057530213974814844

Work with Codex from anywhere | OpenAI
https://openai.com/index/work-with-codex-from-anywhere/

yesterday we had full house at our Codex community afterwork 🇵🇹 it was little over 100 builders thats came to our office to listen to 3 community demos and a panel focused on why people are going for Codex. it was good vibes and many good conversations during the evening. we
https://x.com/TimHaldorsson/status/2055206416747507785

You can now run MagicPath as a native canvas inside Codex to design and build functional apps. It’s pretty incredible. Here’s how to do it 👇
https://x.com/skirano/status/2055364115560878480

You’ve been asking for this one… Now in preview: Codex in the ChatGPT mobile app. Start new work, review outputs, steer execution, and approve next steps, all from the ChatGPT mobile app. Codex will keep running on your laptop, Mac mini, or devbox.
https://x.com/OpenAI/status/2055016850849993072

Your laptop can stay home. Work with Codex from the ChatGPT mobile app, answer questions on the go, and pick up the same thread later from your computer.
https://x.com/OpenAIDevs/status/2057142816497906045

Your Mac can hold down the fort while you work from your phone. Enable remote connection in the Codex desktop app, then turn on “Keep this Mac awake.” When your Mac is powered on and plugged in, Codex can keep running there while you work from the ChatGPT mobile app.
https://x.com/OpenAIDevs/status/2056442456800141424

Continual learning is bottlenecked by realistic evaluations Introducing FutureSim, which replays real-world events in the temporal order they occurred We benchmark frontier agents at updating predictions about how our world evolves, in native harnesses like Codex, Claude Code
https://x.com/ShashwatGoel7/status/2055336064378720412

OpenAI Buys AI Voice Startup Weights — The Information
https://www.theinformation.com/briefings/openai-buys-audio-startup-weights

OpenAI Quietly Bought Voice-Cloning Startup Weights.gg
https://www.implicator.ai/openai-quietly-bought-voice-cloning-startup-weights-gg-then-folded-the-team/

A mic drop moment @ycombinator tonight @sama just offered $2M in OpenAI tokens to EVERY YC startup in the current batch in exchange for equity Just like Yuri Milner offering to invest in every startup back when Sam was a YC partner I can’t wait to see what’s unlocked when you
https://x.com/bosmeny/status/2056914385814401238

1/ Ten months ago, I was ecstatic that AI could win IMO gold. Today, that excitement feels quaint: an internal @OpenAI model has refuted Erdos’s unit distance conjecture–a research result that one could recommend “acceptance without any hesitation” to the Annals of Mathematics.
https://x.com/alexwei_/status/2057182873208369485

Did we ever learn what model won gold at the IMO from OpenAI? It was a year ago and it was called an unreleased internal general purpose model back then. Has GPT-5.5 Pro Extended caught up with whatever it was?
https://x.com/emollick/status/2057196459259301933

So OpenAI literally kill*d many fintech startups today OpenAI launched a personal finance feature in ChatGPT for Pro users in the US. You connect your bank accounts via Plaid, get a spending dashboard, and can ask GPT-5.5 questions grounded in your actual transaction data –
https://x.com/kimmonismus/status/2055320528198521041

OpenAI seals deal in Malta to give all Maltese access to ChatGPT Plus | Reuters
https://www.reuters.com/business/openai-seals-deal-malta-give-all-maltese-access-chatgpt-plus-2026-05-16/?taid=6a0882f19139890001bacbea

Been using @sveltejs for a few projects lately, it’s quite a nice alternative to React, fewer gotchas and complexity and Codex handles it really well.
https://x.com/steipete/status/2055402519841411165

built a new feature into discrawl (store media), codex said it’s done, then I used my codex review skill…
https://x.com/steipete/status/2055203470941061600

OpenClaw 2026.5.12 🦞 🧠 OpenAI setup defaults to Codex login 🛟 Runtime fallbacks + stalled-stream recovery 📬 Telegram polling survives stalls ⚡ Leaner installs, faster startup paths Faster, calmer, harder to wedge.
https://x.com/openclaw/status/2055013211473154309

People freaking out over my AI spend. What nobody sees: Part of what excites me so much about working on OpenClaw is that I’m trying to answer the question: How would we build software in the future if tokens don’t matter? We constant run ~100 codex in the cloud, reviewing
https://x.com/steipete/status/2055405041843052792

This is a game changer. With codex autoreview and crabbox I can now go from issue to fix almost fully automated. (yes it does burn lots of tokens)
https://x.com/steipete/status/2055178254877700450

We’re adding new ways for people to identify AI-generated images and understand where they came from. In addition to C2PA Content Credentials, images now also contain a SynthID watermark, and can be identified using a public verification tool to check whether an image was made
https://x.com/OpenAI/status/2056793648571011232

Building a safe, effective sandbox to enable Codex on Windows | OpenAI
https://openai.com/index/building-codex-windows-sandbox/

OpenAI And 1Password Bring Agentic Security To Codex
https://www.forbes.com/sites/timkeary/2026/05/19/openai-and-1password-bring-password-security-to-codex/

OpenAI announces new Guaranteed Capacity offering for customers to secure compute
https://www.cnbc.com/2026/05/19/openai-announces-new-guaranteed-capacity-offering-for-customers-to-secure-compute.html

Proud to see our work on agent security @openai highlighted in Forbes. Securing AI agents means bringing identity, credential, and access controls directly into the developer workflow, and Codex is a major step in that direction.
https://x.com/ithilgore/status/2056793901495947432

Some personal news: I’ve started a new AI safety standards org, and our first two standards are out today. We’re called Guidelight, co-founded with fellow ex-OpenAI safety researcher, Page Hedley. (1/n)
https://x.com/sjgadler/status/2056762703033807068

🩹 clawpatch 0.1.0 is live: Clawpatch maps codebases into semantic feature slices, reviews them for bugs and quality issues, and records explicit fix attempts with validation. You’ll be surprised how much this will find. npm install -g clawpatch
https://x.com/steipete/status/2055364630709448970

BlackBar 0.2.0 is live for @useblacksmith 📈 24h vCPU + workflow graphs 🔔 opt-in status/job notifications 🧰 richer Blacksmith job rows 🟢 compact status badge Tiny menu bar, less CI guesswork.
https://x.com/steipete/status/2055685581758206139

Can’t recommend @cotypist
https://t.co/QUKTyTo1bh enough. Autocomplete everywhere.
https://x.com/steipete/status/2057040636449116222

ClawRouter ♥️♥️♥️ @NousResearch Hermes Agent 🎀🎀🎀🎀🎀
https://x.com/ClawRou/status/2055078292567597253

deslop your Claude code if you haven’t yet switched to Codex.
https://x.com/steipete/status/2055747016727167035

https://t.co/XMaUrWMMZC is pretty sweet
https://x.com/steipete/status/2055209490887119098

I dont love to gloat but we are almost 2x’ing openclaw just 3 days after surpassing their daily token volume 🤗
https://x.com/Teknium/status/2055125356554899865

Looks like our focus on performance paid off.
https://x.com/steipete/status/2055570810513850777

Lossless is a really interesting concept for OpenClaw to have an “”infinite”” context window/memory. It compacts conversations in blocks that the model can refer to, building a tree to look up past messages.
https://x.com/steipete/status/2055658429645983756

mcporter 0.11.0 is live I use mcporter mainly as more stable browser automation cli these days and for agents to test MCPs without having to restart. I do love that code mode is slowly being adopted by harnesses so this will be less needed.
https://x.com/steipete/status/2054986075232199038

my brain: don’t read the hacker news comments, don’t read the hacker news comments me: reads the hacker news comments
https://x.com/steipete/status/2055775661755715974

OpenClaw 2026.5.18 is live 🤖 xAI/Grok OAuth + sidecar auth fixes 🎙️ Realtime Android Talk Mode 💬 Telegram media + forum-topic delivery fixes 🪟 Browser dialogs visible + answerable A week of polish, plumbing, and fewer papercuts.
https://x.com/openclaw/status/2056504927795437826

OpenClaw 2026.5.19 🦞 📱 Android Talk Mode goes realtime 🍎 Mac Settings feel much cleaner 🔐 xAI login works headless 🧵 Telegram topics behave better Big release. Smaller tweet.
https://x.com/openclaw/status/2057202955581809093

OpenClaw just plugged into X, and now your own hardware gets the claws. 🦞 Bring your Grok, SuperGrok or X Premium subscription to your OpenClaw agent. Now even your personal agent is red-pilled and based. Get Grokked:
https://x.com/openclaw/status/2056826388322340910

QA and perf testing is paying off.
https://x.com/steipete/status/2055275264590959078

Starting today, use your Grok or X Premium subscription in @openclaw. Chat with your agent, generate images and videos, or search for X posts.
https://x.com/xai/status/2056826183745253663

The latest OpenClaw release is ~3.5x faster 🦞 We run end-to-end RTT tests against every published npm release, every 6 hours, over real message channels (here: Telegram, using the brand new bot-to-bot communication). No more silent regressions. Runners are all running on
https://x.com/openclaw/status/2055273947537490422

The latest release of OpenClaw is the first one that ships with our new TypeScript security hardening file-system lib. Previously, this was a grown mess of ad-hoc hardening which was hard to maintain, slow and inconsistent.
https://t.co/PXTXYFOIAQ increased some file ops by 10x.
https://x.com/steipete/status/2055179961535705327

We’ve been working really hard on performance, reliability, security, and stability. Invented whole new automation flows with crabbox, automated video QA and are spending insane amounts of CPU cycles on CI. It’s a good release.
https://x.com/steipete/status/2055026017291370701

Security in OpenClaw is getting sharper 🦞 🔒 fs-safe for root-bounded filesystem 🌐 Proxyline for policy-driven network egress 📦 ClawHub trust evidence 🛡️ smarter command approvals Powerful agents need guardrails you can actually audit.
https://x.com/openclaw/status/2055437760459055405

interesting open model by cohere with lots of unusual architecture choices, here is a recap: > parallel transformer, so MoE and attention are computed in parallel. likely doing some kind of MLP/attention disaggregation here? > lots of query heads, query total dim is 4x hidden
https://x.com/eliebakouch/status/2057198733759008989

Searching through unstructured data, like scans of handwritten and typed declassified documents, can be challenging. But with Cohere Compass, it’s possible because it is built to process and retrieve across even the most challenging documents. This includes the Compass Visual
https://x.com/cohere/status/2055343638360752351

4 shared experts with 8 routed experts active? so 12/132, that’s crazy, i wonder why. most papers like Towards Greater Leverage would suggest 1 shared expert or minimal (i think we should decouple shared expert size anyway eventually) also, 128 attention heads with GQA???
https://x.com/stochasticchasm/status/2057150551696261607

An early beta of Grok Build, an agentic CLI for coding, building apps, and automating workflows is now available for SuperGrok Heavy subscribers. Through this early beta, we will improve the model and product based on your feedback. Try it at
https://x.com/xai/status/2054993285152989373

Composer 2.5 is a significant step up from Composer 2. This is the very start of our work with SpaceXAI. Hope to have more improvements out soon.
https://x.com/mntruell/status/2056418797473640681

Together with SpaceXAI, we’re training a significantly larger model from scratch, using 10x more total compute. With Colossus 2’s million H100-equivalents and our combined data and training techniques, we expect this to be a major leap in model capability.
https://x.com/cursor_ai/status/2056415419536461836

We are soon going to get a new 1.5T model by xAI
https://x.com/scaling01/status/2055320443129581647

some interesting numbers from the SpaceX IPO filing: musk stock conditions: > 302.1M performance-based Class B shares > vesting requires both market cap milestones ($1.065T to $6.565T) and SpaceX completing non-Earth data centers capable of 100 terawatts of compute per year
https://x.com/eliebakouch/status/2057222864332320999?s=12

You can now use X Premium subscriptions in Hermes Agent, and Hermes Agent can now search X posts.
https://x.com/xai/status/2055745332919808181

You can now use your @grok subscription inside @NousResearch Hermes Agent.
https://x.com/xai/status/2055375676656783733

You can now use your SuperGrok subscription in Hermes Agent! Enjoy!
https://x.com/Teknium/status/2055373314399650230

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading