Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A Byzantine gold-ground mosaic icon of a stylized flat-faced saint with six symmetrical arms, each hand holding a different small tool — scroll, key, quill, hammer, compass, envelope — a halo of concentric circuit-ring gyres behind the head and tiny hammered-gold clockwork birds orbiting the hands like couriers, warm candlelit glow on tesserae with visible grout, deep imperial purple and Tyrian crimson accents, the bold ivory Trajan capital title AGENTS set cleanly across the lower third, landscape 16:9, symmetrical iconic composition with generous negative space.

AI Startup Cognition Raises $1 Billion at $26 Billion Value – YouTube
https://www.youtube.com/watch?v=VuyOy5WN980

cognition is now the largest independent agent lab in the world. take the 200% utilization that everyone is hitting from this chart and run out the sales growth from this, i encourage you to go thru the exercise if you are new to investing a lot of you have read my cog
https://x.com/swyx/status/2059717021944926238

It is cliché at this point, but most people don’t realize how capable the current generation of AI systems in their harnesses really are (And, as opposed to previous times where non-lawyers or non-mathematicians were making these comments about law & math, now it is the experts)
https://x.com/emollick/status/2059431958447317381

Excited to share our most powerful new Claude Code feature: dynamic workflows! Mention “”workflow”” in a prompt and Claude will dynamically create an orchestration plan that it strictly follows, allowing you to confidently trust that every stage happens in the right order even
https://x.com/_catwu/status/2060054180379689074

The cost per accepted line of code varies by roughly 7x across model families.
https://x.com/cursor_ai/status/2060025070425395562

Google DeepMind’s Hassabis: AGI is 3 to 4 years away – Sherwood News
https://sherwood.news/tech/google-deepminds-hassabis-agi-is-3-to-4-years-away/

// Language Models Need Sleep // Let your agents “”sleep””, folks. On a serious note, this is a fascinating paper on getting the most from long-horizon agents. Here is the problem with agents today: Attention scales badly with context length, so long-horizon agents keep paying a
https://x.com/dair_ai/status/2059333792775745619

// Your Agents are Aging Too // Huh!? They need “”sleep,”” and now they are aging? Joke aside, great write-up on reliable agentic engineering. This new research introduces AgingBench, a longitudinal reliability benchmark. It organizes agent aging into four mechanisms, including
https://x.com/omarsar0/status/2059689897523642510

🧩 MCP does way more than you think Beyond tool lookups, MCP can connect Copilot to databases, APIs, and internal systems, plus power richer app-like experiences in chat. 🔗 Register:
https://x.com/code/status/2059666498285629707

15 AI agent workflows your team can run today
https://info.notion.so/resources/unlock-15-ai-agent-powered-workflows

5 patterns for building long-running AI Agents 1. Checkpoint-and-Resume → Save progress in batches (like every 50 documents) so the agent can recover from failures without restarting entire workflows 2. Delegated approval → Long-running agents can “freeze” mid-process,
https://x.com/TheTuringPost/status/2058240378718085230

Also extracted our image-logic into a separate library. Especially useful if you want to ensure small hacked images don’t explode your process. Rastermill – Portable image processing for Node agents. Uses Wasm+Rust to be fast.
https://x.com/steipete/status/2059423344961671290

As agent harnesses become more standardized, we’re going to see a lot more “managed agent services”
https://x.com/hwchase17/status/2060034741471199249

As agents use more context, input tokens have become the majority of price-equivalent token costs.
https://x.com/cursor_ai/status/2060025076947521984

autoreview is the most impactful skill I’ve added to my stack (next to
https://t.co/SEj2XRpaD1). It automatically reviews your code before landing a PR. Finds so many edge cases. Sometimes it runs for hours.
https://x.com/steipete/status/2059453909819654554

Build strong data foundations for agentic AI at scale | AWS Marketplace panel
https://pages.awscloud.com/awsmp-gim-yngd-webinar-aim-enterprise-ai-and-data-leader-panel-lt-panel-1.html?trk=9e986f3f-af52-49c8-ad21-bd3dd1f6187a&sc_channel=el

Cognition: The Devin is in the Details
https://www.swyx.io/cognition

computers roughly do two things.. show you something when you’re not asking, & do something when you are. both of these are about to radically change. ambient ai handles the first, agentic ai handles the second, & the seam between them is the new interface for almost all of
https://x.com/signulll/status/2057850735048458639

Deep Agents v0.6 brings Delta channels, reducing checkpoint storage by up to 100x for long-running agents, without sacrificing observability or resilience. Here’s a 200-turn coding agent session. Without Delta Channels: 5.3GB of checkpoint storage With Delta Channels: 129mb
https://x.com/LangChain/status/2059634226836746483

Dynamic workflows and adversarial code review was part of what made it possible to rewrite Bun in Rust in 6 days.
https://x.com/jarredsumner/status/2060050578026189172

Exploring Agent-Assisted Qualitative Analysis
https://www.sh-reya.com/blog/ai-qual-analysis/

Fleet agents can now securely write and run code. With computer use in LangSmith Fleet, agents get isolated execution environments. Analyze data, transform files, generate & write code, and run shell commands all within a secure virtual computer. Now in public beta.
https://x.com/LangChain/status/2059685293322858809

Folks: when you write skills, ask your agent to be token efficient, relax grammer. I see too many skills that write books in the skill description, and all that crap is loaded into every context. I wrote a skill that finds the worst offenders.
https://x.com/steipete/status/2058917897590673525

Gemini Managed Agents Dev Guide: 1 API call = Gemini 3.5 Flash + Antigravity Harness + remote Linux sandbox. No infra, no orchestration. – Antigravity quickstart (code/files/browsing) – Persistent multi-turn + streaming – Custom agents (AGENTS.md + mounts) – Ops:
https://x.com/_philschmid/status/2059263980913229989

Huge congrats to the Cognition team! If you tried Devin when it first came out, know it is a completely different product today and among the primary coding agents we use at Exa. Anytime I message the cog team at a crazy hour I get an instant answer. So well deserved.
https://x.com/nityasnotes/status/2059768072110776370

I always wanted a GitHub dashboard: See my repos, open Issues/PRs, what version I released last, how many commits since last release. So I built one for everyone.
https://x.com/steipete/status/2058381186884411473

I don’t think people realize how different programming feels when proactive automations are set up. On every important Slack channel inside Cognition, we have a Devin automation monitoring all posts – triaging, solving, routing, reviewing. Devin knows who works on which parts of
https://x.com/russelljkaplan/status/2057932036892197337

I would push back a little: because the models are so good & improving, they don’t have to be the product. But it is the model that is the prime mover. If they weren’t so generally capable, the harnesses & apps the labs build around them would be hard to build and wouldn’t work.
https://x.com/emollick/status/2057681650633322939

I’m refactoring an older part of the codebase (subagents) that touches a lot of code, and autoreview is running for 5h already and fixing tons of issues.
https://x.com/steipete/status/2058307930613518698

ICYMI: We introduced CoreWeave Sandboxes, now in public preview. It’s the execution layer for RL, agent tool use, and model evaluation, on your own CKS clusters or serverless through @wandb 👏 More from @deok_filho.
https://x.com/CoreWeave/status/2057852737073942634

Improving your agent has been a manual process of: ✅ Reading traces ✅ Looking for patterns ✅ Writing evals ✅ Creating fixes Now, LangSmith Engine runs that cycle for you.
https://x.com/LangChain/status/2059654417478012938

Introducing Repo2RLEnv Turn any repository into runnable, verifiable coding environments built from real PRs and commits for coding-agent evaluation or RL training > uv pip install repo2rlenv
https://x.com/adithya_s_k/status/2059991239890776269

It Twitter’s too busy for you, try
https://x.com/steipete/status/2058263360525730087

Late night sesh with the Hark agent team; I’m so excited for the potential here p.s. we have some exciting tomorrow morning @hark_labs
https://x.com/adcock_brett/status/2057324020085973264

Model-Harness-Task fit! it’s clear that RL post-training produces a model-harness fit via tool shapes and prompting as models are trained with the harness in the loop. Mentioned this in a previous LangChain blog, Cursor also has good content on this But there’s probably less
https://x.com/Vtrivedy10/status/2059712077925658717

More Devins in More Places | Cognition
https://cognition.ai/blog/series-d

New @latentspacepod Essay: why Agent Labs are clearly emerging in 2025 as a complement to Model Labs’ all becoming AI Cloud platforms.
https://x.com/swyx/status/1990886806250782876

New pet peeve: cli’s that install new skills onto my system without asking.
https://x.com/steipete/status/2058883349632934149

🆕The Age of Async Agents: Devin’s 7x PR growth, 80% AI commits, background agents, memory, testing, & Open-Inspect
https://t.co/x5Hw5S3egc @cognition cofounder + CPO @walden_yan and Open-Inspect creator @_colemurray explain why engineering is moving from local IDEs to cloud
https://x.com/latentspacepod/status/2060089484608459220

Out of every company I’ve seen, @Cloudflare has cracked the agent-platform of the future best. It’s currently not close. It addresses shortage issues organically. Most analysts are saying “”each agent needs its own computer”” (which is spiking forecasts). Cloudflare instead has
https://x.com/brandonjcarl/status/2059624598644109363

PSA: the IDE in Antigravity 2.0 is alive and well, we just landed an update to the UI which makes it more clear (see top right). Sorry for the confusion on this (pls keep feedback coming), we also just reset everyone’s weekly limits. Enjoy the weekend : )
https://x.com/OfficialLoganK/status/2057912550633947436

QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks “”We release Quest, a family of open models (ranging from 2B to 35B) that serve as general-purpose deep research agents designed to handle a wide range of long-horizon search tasks, with strong capabilities
https://x.com/iScienceLuvr/status/2059223911011930606

Robinhood is Now Open to Agents
https://robinhood.com/us/en/newsroom/robinhood-is-now-open-to-agents/

S tier AI products need model <> harness <> product symbiosis
https://x.com/dzhng/status/2057748510947082539

Say hello to open source deep research for your favorite agent harness. Our AI-Q agent skill packages the work of building a research pipeline into a portable skill. Drop it into your harness, and the agent delegates a research task to a local or hosted AI-Q server and gets back
https://x.com/NVIDIAAI/status/2057855521193881773

System scaling is the next real bottleneck in agentic AI. If you build agent orchestration layers, this is a clean map of where the engineering leverage actually sits. The labs own the model. You own the harness, and that is increasingly where agent quality is won or lost. The
https://x.com/dair_ai/status/2059294269698199929

The 2026-07-28 MCP Specification Release Candidate | Model Context Protocol Blog
https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/

The release candidate for MCP 2026-07-28 is out. The protocol is now stateless: no handshake, no session id, any request can hit any server instance. Plus extensions as first-class (MCP Apps, Tasks), auth hardening, and a proper deprecation policy so we don’t have to do this
https://x.com/dsp_/status/2057780712187580924

The stack around the model is becoming as important as the model itself
https://x.com/TheTuringPost/status/2058174163618005292

There’s security, and there’s clankers.
https://x.com/steipete/status/2058884046940225918

This is an idea I have been using for like 4 months now. Very easy to do with -p or Agent SDK. I doubt I will use CC for it, but great to see a native implementation of dynamic workflows. Agent-to-agent interactions are super effective, but also watch out for token use.
https://x.com/omarsar0/status/2060059612041171175

This seems more ideological than technical.
https://x.com/steipete/status/2058101804202602949

Today we’re bringing Cua Driver to Windows: background computer-use for any agent. Claude Code, Codex, or your own loop can drive real Windows apps through CLI or MCP while your desktop stays usable, with true multi synthetic pointer support. 1/6
https://x.com/trycua/status/2059688960838828391

Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.
https://x.com/serenaa_ge/status/2059308218564890875

We are quite short of compute, and that is going to result in compute becoming very expensive for complex agentic workflows even as single-turn chatbots get cheaper. So the richest companies & most pressing use cases will use AI agents & everyone else will be stuck with chatbots?
https://x.com/emollick/status/2057565824341127432

we just revamped the create_agent docs! the new agents page shows how to build a custom harness for your use case w/ create_agent as an easy entrypoint start w/ your prompt and tools, then add middleware to customize at any point in the loop
https://x.com/sydneyrunkle/status/2059280878694531280

We launched Context Hub as a way to manage skills, AGENTS.md files, and other context files an agent might need You can easily use it as a virtual filesystem in deepagents See this video for more info!
https://x.com/hwchase17/status/2059687279199924462

We spent a ton of time making worktrees actually work well for agent swarms and large repos. When you’re running 10s of agents that ship, you quickly realize you need worktree, but git’s defaults are brutal at this scale. Slow creation, every agent copying the whole repo, can’t
https://x.com/theskory/status/2059729539287167068

We used bots so far to enforce a 10PR per “person” limit. Great to see GitHub shipping that natively!
https://x.com/steipete/status/2057946259709628781

We’re excited to introduce Inherent, a lab designed from scratch to build AI agents that discover new knowledge. The coming era of machine-driven scientific inquiry demands a new kind of research institution and a new kind of AI. To achieve our mission, we live within the
https://x.com/inherent_labs/status/2060119235372752924

What do people use for SSO/SCIM/Endpoint Security in 2026. As we’re hiring people for the OpenClaw Foundation, I gotta level up.
https://x.com/steipete/status/2059421603268608302

With the Cursor SDK, you can build your own agents with Composer 2.5. It’s now available in Python and TypeScript. This long weekend, Composer usage is 90% off in the SDK. We’re excited to see what you build!
https://x.com/cursor_ai/status/2057913121558413770

You’d think they’d have addressed the already finnicky edit tool when making the parallel agents stuff more aggressive. Can’t imagine how many millions of dollars in tokens have been wasted on failed edits due to conflicting parallel agents.
https://x.com/theo/status/2060135394570797158

γ-World: Generative Multi-Agent World Modeling Beyond Two Players
https://research.nvidia.com/labs/sil/projects/gamma-world/

Announcing AA-WER Streaming, our new benchmark measuring streaming Speech to Text models on accuracy and latency for voice agent use cases. Pareto optimal models on this new benchmark include those from Cartesia, ElevenLabs, and Deepgram Streaming Speech to Text (STT) powers
https://x.com/ArtificialAnlys/status/2060021901234458958

Agent Judge: Solving Long-Context Evals for Production Agents — Judgment Labs
https://www.judgmentlabs.ai/blogs/agent-judge-solving-long-context-evaluations

Artificial Analysis and IBM Research are launching ITBench-AA, the first in a new series of benchmarks evaluating models on agentic enterprise IT tasks, starting with Site Reliability Engineering tasks where frontier models score below 50% ITBench-AA’s SRE tasks benchmark model
https://x.com/ArtificialAnlys/status/2059698327235805258

Interesting new SWE/agentic benchmark (DeepSWE) was released yesterday. 113 tasks across 91 repos in 5 languages. Here are interesting things I noticed: – The evaluation harness (mini-swe-agent) gives every model a single bash tool and the same SI. No vendor editing primitives.
https://x.com/_philschmid/status/2059564676569076021

Reasonix — DeepSeek-native AI coding agent for your terminal
https://esengine.github.io/DeepSeek-Reasonix/

RF-DETR just landed to @huggingface transformers 🥵🔥 sota real-time detection & segmentation models by @roboflow 💜 > play with our real-time demo > fine-tune the models on your use case with our tutorials (takes a toaster’s VRAM) > or just hand them to your agents 😄
https://x.com/mervenoyann/status/2059647988373373253

SHIPPED. Mistral Vibe is now the AI agent for long-horizon productivity and coding, and the home for Work mode, Code mode, the CLI, and a brand new VS Code extension. Let’s go… 🧵
https://x.com/mistralvibe/status/2059984963932499973

I’ve started experimenting with gBrain + Hermes Agent it’s a shared memory layer that sits underneath my Hermes Agent company. every specialist reads from the same brain before they do anything the architecture I’m currently testing: > inputs flow in: my ideas, strategy
https://x.com/shannholmberg/status/2057821004676956586

very belated but in retrospect i think @sama’s mythical “”build a business that gets better when models get better”” is basically what I called Agent Labs here. seeing a very direct correlation with model performance and agent lab revenue, discontinuity in Q4 2025 (clip from
https://x.com/swyx/status/2057119153337545096

OpenClaw’s dependency purge continues. Killed Sharp and Jimp. Replaced it with photon, a small WebAssembly that runs compiled Rust for image processing. 2MB vs 140MB.
https://x.com/steipete/status/2058922222790525272

Excited to dive into this – an open source agent designed with memory/continual learning in mind
https://x.com/hwchase17/status/2059487107144655356

To get Perplexity Computer and similar tools deeply embedded in enterprises, a continuous investment in security engineering is necessary. What’s interesting in the way we’re approaching it is putting these tools insde agentic sandboxes and having security workflows run
https://x.com/AravSrinivas/status/2057873563156402448

Fast, faster, Qwen. 🚀 Thrilled to see Qwen3.5 reaching a record-breaking 580 tps for agentic workloads on the TokenSpeed engine! This milestone wouldn’t be possible without our incredible partners. Huge thanks to @lightseekorg, @NVIDIAAI, the Mooncake team, and @tri_dao for
https://x.com/Alibaba_Qwen/status/2059674574397313277

Proud to see Qwen3.7‑Max debut at #4 in Code Arena, marking a significant milestone for Qwen in agentic web development. #AlibabaAI #Qwen
https://x.com/AlibabaGroup/status/2059317802935423028

Today’s best coding models from Qwen, DeepSeek, Minimax etc are trained on 1000s of concurrent RL environments to simulate SWE tasks in real git repos. Very excited to open source a tool that unlocks this capability for the whole AI community: point Repo2RLEnv at any GitHub
https://x.com/_lewtun/status/2059995216937886088

NEW paper worth reading. A full agentic workflow can be distilled into model weights and run at roughly 100x lower inference cost while preserving near-frontier task quality. The workflow includes multi-step LLM calls, tool invocations, intermediate scratchpads, and decision
https://x.com/dair_ai/status/2057846601843146760

Claude Opus 4.8 is now available in Windsurf and Devin CLI
https://x.com/cognition/status/2060050201990369662

Cursor Composer 2.5’s is 3-18x cheaper than Opus 4.7 in Claude Code (medium reasoning), and 5-32x cheaper than GPT-5.5 in Codex (medium) based on API pricing This low Cost per Task isn’t just driven by relatively low token pricing, it’s also driven by low relatively low token
https://x.com/ArtificialAnlys/status/2057914437156409577

Recently, I used dynamic workflows to catalogue all of our 100s of A/B test flags and find the ones rolled out to 0% or 100% so that we can quickly deprecate the stale ones. Instead of waiting for Claude Code to investigate each sequentially, dynamic workflows allowed Claude to
https://x.com/_catwu/status/2060054182447448387

two CTOP updates: 1. now supports Devin (in addition to Claude Code, Codex, OpenCode) 2. new CLI — you (or your agents) can run ctop ls, ctop search, ctop kill from the terminal
https://x.com/aakashadesara/status/2057809590616461399

Opus 4.8 is live. Benchmarks especially significant jump in Agentic coding, but more important: „Fast mode is available for Opus 4.8. It’s the same model at roughly 2.5x the speed, and we’ve made it three times cheaper than before.”
https://x.com/kimmonismus/status/2060044465385902436

Opus 4.8 is now supported in Hermes Agent ^_^
https://x.com/Teknium/status/2060054418821906652

The fact that tokens went from something no one even put in a budget line a year ago to an absolute requirement for coding now is the cause of handwringing, not that AI is not turning out to be useful No one knows who should get tokens, how much they should get & how to control
https://x.com/emollick/status/2059640930265686158

Gemini 3.5 Flash is a step forward for Google on speed and agentic capabilities but comes at a trade-off of being higher cost than prior models We have measured up to ~280 output tokens/sec, placing it on the speed/intelligence Pareto frontier and well ahead of Gemini 3 Flash.
https://x.com/ArtificialAnlys/status/2059316050391634302

Gemini 3.5 Flash ranks #1 on Automation Bench (from Zapier), beating every other frontier model at a much lower cost
https://x.com/OfficialLoganK/status/2057317567673594131

Thanks for all of the Antigravity feedback over the last couple of days, especially around the IDE. Our intention was never to remove the IDE support for developers, and we should have been clearer with that in the product from the beginning. We’ve made it clearer in 2.0 on how
https://x.com/_mohansolo/status/2057910616153882949

What if you can build an Agent with it own computer in a single api call? At my @Google I/O talk, I showed how to use Gemini Managed Agents and the new Interactions API to give your AI a secure, hosted Linux sandbox to execute code and manage its own memory.
https://x.com/_philschmid/status/2057833963633418426

Gemini 3.5 flash release is underwhelming for browser agents. Slight improvement in performance over Gemini 3.1 pro, at a small increase in total cost
https://x.com/Alezander907/status/2057686331380359566

Report: Microsoft tries to get back in the AI coding game with new model – Sherwood News
https://sherwood.news/tech/report-microsoft-tries-to-get-back-in-the-ai-coding-game-with-new-model/

The App Modernization Playbook | Microsoft Azure
https://info.microsoft.com/ww-landing-app-modernization-playbook.html

Train the AI’s playbook instead of retraining the AI itself Microsoft introduced a new open-source training method – SkillOpt It optimizes external skill documents, while the model’s internal weights stay frozen. Here’s the workflow: • The AI tries tasks • An optimizer
https://x.com/TheTuringPost/status/2059137971543269860

Almost everyone is building agent harness systems the wrong way. The default move: pick LangChain or LangGraph or the OpenAI Agents SDK, accept the loop, the tools, the memory, the orchestration, the policy engine, the credential store, the budget tracker, all of it, as one
https://x.com/ghumare64/status/2060072412868235587

Cloudsail: Instant Sandboxes for Coding Agents Create a new Cloudflare Sandbox for each task with a shell, Codex and GitHub access. Tokens are never exposed to the sandbox. Update your deps far away from your laptop. npm install -g cloudsail cs
https://x.com/cnakazawa/status/2057823910574588238

Codex computer use entirely driving iphone simulator to bug bash a feature it just built
https://x.com/JustinBleuel/status/2058228412158758950

Macro Evals for Agentic Systems
https://developers.openai.com/cookbook/examples/partners/macro_evals_for_agentic_systems/macro_evals_for_agentic_systems

Private MCP servers 🤝 OpenAI products Your team can keep MCP servers inside your network while ChatGPT, Codex, and the Responses API connect through outbound-only HTTPS. 🔗
https://x.com/OpenAIDevs/status/2059703536825565499

Secure MCP Tunnel | OpenAI API
https://developers.openai.com/api/docs/guides/secure-mcp-tunnels

🦀 The Rust frontend is officially merged into vLLM! As GPUs get faster, the frontend has become a real share of CPU time. The new Rust frontend is a drop-in alternative to the Python API server — same engine, same ZMQ boundary. Opt in with VLLM_USE_RUST_FRONTEND=1. Early
https://x.com/vllm_project/status/2059344804295942513

.@Microsoft has just open-sourced 2 useful tools: ▪️ RAMPART – a framework for stress-testing AI agents with repeatable attack and safety scenarios directly in CI. ▪️ Clarity – helps teams design the right system before they build it, saving results in a readable markdown
https://x.com/TheTuringPost/status/2057268273952264279

Cursor · The Cursor Developer Habits Report
https://cursor.com/insights

DeepSWE
https://deepswe.datacurve.ai/blog

Thanks to Julien Grok Build v0.1 now has its appropriate 256K context length in Hermes – sorry bout that!
https://x.com/Teknium/status/2057930638632812642

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading