Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: Using the provided reference image, keep the pure white landscape field, vertical type stack, ‘(we could be)’ tagline, ‘FEATURING.’ label, and identical galaxy-punchout Milky Way treatment across all letterforms, but replace ‘HEROES’ with ‘AGENTS’ in the same bold condensed grotesque, replace ‘ALESSO’ with ‘TRUSTED DELEGATES’ in the light geometric all-caps sans, and replace ‘TOVE LO’ with ‘TOOL USE’ in the bold condensed grotesque, preserving exact tracking, weights, font contrast, and starfield clipping.
Introducing Deep Research and Deep Research Max
https://blog.google/innovation-and-ai/models-and-research/gemini-models/next-generation-gemini-deep-research/
ChatGPT Images 2.0 – YouTube
I have been using GPT ImageGen-2 for the past weeks I didn’t think that better image-generators would be a big deal but it turns out that there is a quality threshold I didn’t expect, where you can now get text, slides, academic papers Look at what it does with my “”otter test””!
https://x.com/emollick/status/2046665274535854146
No bad ideas when you’re playing with ChatGPT Images 2.0 → Smarter visuals → Better editing and aesthetics Rolling out in Figma and Figma Weave
https://x.com/figma/status/2046673364496875977
Aspect Ratios & Resolution with ChatGPT Images 2.0 – YouTube
ChatGPT Images — Chameleon – YouTube
Instruction Following with ChatGPT Images 2.0 – YouTube
Multilingual & Text Rendering with ChatGPT Images 2.0 – YouTube
Slides & Infographics with ChatGPT Images 2.0 – YouTube
Thinking & Intelligence with ChatGPT Images 2.0 – YouTube
This is ChatGPT Images 2.0 – YouTube
Adobe Summit: Adobe Redefines Customer Experience Orchestration Vision in the Agentic AI Era with Introduction of CX Enterprise
https://news.adobe.com/news/2026/04/adobe-redefines-custome-experience
// Skill Learning for Autonomous Web Agents // Web agents can navigate a page, but ask them to repeat a checkout flow they already completed, and they start from scratch every time. This work introduces WebXSkill, a skill learning framework where web agents extract reusable
https://x.com/dair_ai/status/2045139481892880892
80% of US adults who report using Claude in the previous week live in households earning $100,000 or more a year, compared to 37% of Meta AI users. Other major providers cluster in a relatively narrow band, with 56-64% of users in $100,000+ households.
https://x.com/EpochAIResearch/status/2047056309535801605
Anthropic’s Mythos AI Model Is Being Accessed by Unauthorized Users – Bloomberg
https://www.bloomberg.com/news/articles/2026-04-21/anthropic-s-mythos-model-is-being-accessed-by-unauthorized-users
Building agents that reach production systems with MCP | Claude
https://claude.com/blog/building-agents-that-reach-production-systems-with-mcp
Figma stock 20 minutes after the Claude Design announcement. Wild.
https://x.com/Yuchenj_UW/status/2045161719547445426
Anthropic released Claude design, direct attack on figma and lovable. Anthropic just shipped Claude Design, powered by Claude Opus 4.7., a tool that turns conversations into polished prototypes, pitch decks, and marketing assets. It auto-applies your brand system, lets you
https://x.com/kimmonismus/status/2045162358004216134
Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our most capable vision model. Available in research preview on the Pro, Max, Team, and Enterprise plans, rolling out throughout the day.
https://x.com/claudeai/status/2045156267690213649
On the plus side with Opus 4.7, if it does decide to think it produces BY FAR the best Sparks unicorn* ever, even non-thinking is pretty good, if not great. * This is created using TikZ, which is a language built for scientific diagrams & very much not for drawing. The original
https://x.com/emollick/status/2044880350237626844
Anthropic and Amazon expand collaboration for up to 5 gigawatts of new compute \ Anthropic
https://www.anthropic.com/news/anthropic-amazon-compute
We’re expanding our collaboration with Amazon to secure up to 5 gigawatts of compute for training and deploying Claude. Capacity begins coming online this quarter, with nearly 1 gigawatt expected by the end of 2026.
https://x.com/AnthropicAI/status/2046327624092487688
Anthropic is coming after Figma.
https://x.com/Yuchenj_UW/status/2045158071950033063
Anthropic making a Lovable/Bolt/v0/Figma Make clone and calling it Design is peak Anthropic.
https://x.com/skirano/status/2045192705941106992
Introducing Claude Design by Anthropic Labs \ Anthropic
https://www.anthropic.com/news/claude-design-anthropic-labs
Gemini Embedding 2 is now generally available via the Gemini API and Gemini Enterprise Agent Platform search and understand semantic relationships across text, image, video, audio, and documents without complex, fragmented pipelines
https://x.com/GoogleAIStudio/status/2047007402520674679
Introducing one of our biggest updates to the Gemini Deep Research Agent, now available via the Interactions API! Trigger complex, long-horizon research workflows with arbitrary MCP support, get rich visualizations, plan before you execute, and more with these two
https://x.com/googleaidevs/status/2046630912054763854
The next evolution of our autonomous research agent is here. Today, we’re introducing Deep Research and Deep Research Max via the Gemini API. Powered by Gemini 3.1 Pro, you can now trigger comprehensive research workflows with unprecedented control and transparency, featuring:
https://x.com/Google/status/2046627647208259835
We’ve expanded the capabilities of Deep Research and Deep Research Max to give you even more control and transparency over the entire research process: 📋 Collaborative Planning: Review and refine the agent’s research plan before it executes. 🛠️ Extended Tooling: Run Google
https://x.com/Google/status/2046627652568850687
Introducing Gemini Enterprise Agent Platform | Google Cloud Blog
https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-agent-platform/
The conversation around AI agents is no longer about how to build them — it’s about how to manage thousands of them. Today we’re introducing Gemini Enterprise Agent Platform, a new way to build, scale, govern and optimize agents. It combines the models and services you’re
https://x.com/Google/status/2046985650868547851
We’re launching Gemini Enterprise Agent Platform with @GoogleCloud: a platform for businesses to develop, scale, govern and optimize agents. It’s the evolution of Vertex AI, bringing together model selection and agent building with new features for integration, security and
https://x.com/GoogleDeepMind/status/2046983340524269713
We’re making it easier for organizations to scale up autonomous agents with the Agentic Data Cloud. This AI-native architecture closes the gap between thinking and doing through: – A universal context engine that gives your agents a grounded source of truth about your business
https://x.com/Google/status/2046997032649277754
We’re delivering agentic defense by combining Google’s Threat Intelligence and Security Operations with @wiz_io’s Cloud and AI Security Platform to detect, prevent and respond to threats. Security agents provide protection for your entire AI development lifecycle.
https://x.com/Google/status/2047000216188940710
I tested @huggingface ml-intern, given the prompt “”Fine-tune a Segment Anything Model (SAM) on a useful medical dataset. Train the model, and provide a comprehensive tutorial in a Jupyter Notebook file. Additionally, create a Hugging Face article/blog post documenting
https://x.com/Mayank_022/status/2046646301555900828
Kimi K2.6 Tech Blog: Advancing Open-Source Coding
https://www.kimi.com/blog/kimi-k2-6
Kimi K2.6 autonomously overhauled exchange-core, an 8-year-old open-source financial matching engine. Over a 13-hour execution, the model iterated through 12 optimization strategies, initiating over 1,000 tool calls to precisely modify more than 4,000 lines of code. Acting as
https://x.com/Kimi_Moonshot/status/2046531057147933137
Hermes Agent recently overtook OpenClaw in weekly new GitHub stars. Two months after launch, the agent from @NousResearch is pulling developers from the incumbent at meaningful scale. It’s self-hosted and open source, so you control what leaves your machine. It reaches you
https://x.com/Delphi_Digital/status/2045839142450536504
Introducing Hermes Agent v0.11.0 Our largest update yet, with over 700 PRs across ~200 contributors. Thank you to everyone who’s worked on Hermes Agent! This update features a beta TUI v2, unlimited recursion depth and width of subagents, 5 new LLM providers, expanded image gen
https://x.com/Teknium/status/2047506967909015907
always a real feeling of magic to ask codex to perform a task that requires finding information scattered across slack, google docs, notion, and various internal tools, and it just figures it out
https://x.com/gdb/status/2044643518891909289
Auto-review is a new mode that lets Codex work longer with fewer approvals and safer execution. It helps Codex keep moving through tests, builds, and more, including during long tasks and automations, while a separate agent checks higher-risk steps in context before they run.
https://x.com/OpenAIDevs/status/2047436655863464011
ChatGPT Images 2.0 is available starting today to all ChatGPT and Codex users. Images with thinking are available to ChatGPT Plus, Pro, and Business users (Enterprise soon). On mobile, make sure you update to the latest version of the app. The underlying model, gpt-image-2, is
https://x.com/OpenAI/status/2046670994413322435
Chronicle – Codex | OpenAI Developers
https://developers.openai.com/codex/memories/chronicle
Classic study gave 146 economist teams the same dataset & got wildly different answers New paper reruns it with agentic AI. Claude Code & Codex land near the human median, but with far tighter dispersion & no extremes. Suggests that AI is now useful for doing scalable research.
https://x.com/emollick/status/2046362044786458648
Codex Computer feels like the first really usable computer use platform. More importantly, it shows that the tech has arrived and now we will see a wave of things get unlocked. Enterprise software will never be the same again. All the legacy stuff that will never see an API is
https://x.com/matvelloso/status/2045209294942142860
Codex hit 4M active users, less than two weeks after hitting 3M. We will reset rate limits today!
https://x.com/sama/status/2046604989527912590
codex is becoming a full agentic IDE
https://x.com/gdb/status/2045375289560007029
Codex is open source, enabling anyone to build awesome applications on top of it:
https://x.com/gdb/status/2045214436689072136
Exclusive | OpenAI Plans Launch of Desktop ‘Superapp’ to Refocus, Simplify User Experience – WSJ
https://www.wsj.com/tech/openai-plans-launch-of-desktop-superapp-to-refocus-simplify-user-experience-9e19931d
GPT-5.5 is here. It’s our smartest frontier model yet, introducing a new class of intelligence for agentic coding, computer use, knowledge work, and scientific research. Rolling out in ChatGPT and Codex today. API is coming soon.
https://x.com/OpenAIDevs/status/2047377079352877534
gpt-image-2 is here, available today in the API and Codex. The most capable image generation model yet, built for production-grade workflows with stronger text rendering, layout, editing, resolution, and multilingual rendering.
https://x.com/OpenAIDevs/status/2046671238534496259
Have never seen something like Codex Computer Use. UX is novel. Also surprised at how well it works. 5.4 great at driving OS actions. Well done.
https://x.com/mattrickard/status/2045218583882633412
Introducing GPT-5.5 | OpenAI
https://openai.com/index/introducing-gpt-5-5/
Introducing GPT-5.5 A new class of intelligence for real work and powering agents, built to understand complex goals, use tools, check its work, and carry more tasks through to completion. It marks a new way of getting computer work done. Now available in ChatGPT and Codex.
https://x.com/OpenAI/status/2047376561205325845
Introducing workspace agents in ChatGPT | OpenAI
https://openai.com/index/introducing-workspace-agents-in-chatgpt/
Introducing workspace agents in ChatGPT–shared agents that can handle complex tasks and long-running workflows across tools and teams.
https://x.com/OpenAI/status/2047008987665809771
man Codex Computer Use is actually so good i’ve got my guy sending Slack messages, reading my Slack bookmarks, checking stuff on my browser, and i’m still trying more things it’s legit so good
https://x.com/kr0der/status/2045154074337710136
New in the Codex app: – GPT-5.5 – Browser control – Sheets & Slides – Docs & PDFs – OS-wide dictation – Auto-review mode Enjoy!
https://x.com/ajambrosino/status/2047381565534322694?s=20
Seriously stop everything you are doing and use codex desktop app new computer use. Absolutely mind blowing
https://x.com/HamelHusain/status/2045191726495846459
Some of you were disappointed that we “only” get an image model from OpenAI today. But you need to see the big picture: GPT-Image-2 can generate mockups of websites, which Codex can then turn straight into working code. That’s one of the exciting new use cases enabled by true
https://x.com/mark_k/status/2046640315348725879
With GPT-5.5, Codex now gets more of the job done across the browser, files, docs, and your computer. We’ve expanded browser use so Codex can interact with web apps, and test flows, click through pages, capture screenshots, and iterate on what it sees until it completes the
https://x.com/OpenAIDevs/status/2047381283358355706
Workspace agents can work across tools–pulling context from docs, email, chats, code, and systems, and taking approved actions like updating @Linear issues, creating docs, or sending messages. In @SlackHQ, agents can jump into a thread, understand what’s needed, pull the right
https://x.com/OpenAI/status/2047008991944069624
OpenAI dropped a new model on HF today!
https://x.com/ClementDelangue/status/2046973714751754479
OpenAI just open sourced a new 1.5B (50m active) model on HuggingFace with Apache 2.0 license! It’s not a new LLM, this one is called Privacy Filter, and it’s a PII detection model (checking if text has private information) A few interesting tidbits from the release + links:
https://x.com/altryne/status/2046977133013311814
OpenAI just released a new open-source model it’s “”a bidirectional token-classification model for personally identifiable information (PII) detection and masking in text
https://t.co/xTZt1J3WcT
https://x.com/scaling01/status/2046972437422543064
ChatGPT Images 2.0 is a big leap forward in image generation intelligence. It’s much better at following detailed instructions, rendering dense text, understanding the world more accurately, and creating visuals that are more useful. And when you give it additional time to
https://x.com/nickaturley/status/2046677986242363731
GPT Image 2 + Codex: or how to make Codex not suck at UI. Step 1: Generate a UI image (native in Codex) Step 2: Get Codex to implement the UI based on it Step 3: Get Codex to iterate until it aligns with the image as much as possible Codex is bad at initial UI, but very good at
https://x.com/petergostev/status/2046720618566242657
Here is a manga made by ChatGPT Images 2.0 of @gabeeegoooh and me looking for more GPUs:
https://x.com/sama/status/2046672912833458597
My most popular AI post was a bunch of made-up “”graphs”” four years ago. Now, the new GPT-2 image generator does it for real (though not perfect) Here’s the famous AI task horizons graph with a touch of Basquiat, haunted by ghosts, from the Voynich manuscript, as a decaying pier.
https://x.com/emollick/status/2046728271849550331
A Visual Thought Partner ChatGPT Images 2.0 is our first image model with thinking capabilities. When a thinking model is selected in ChatGPT, Images 2.0 can search the web for real-time information, create multiple distinct images from one prompt, double-check its own outputs,
https://x.com/OpenAI/status/2046670989719924768
Making ChatGPT better for clinicians | OpenAI
https://openai.com/index/making-chatgpt-better-for-clinicians/
(17) State of the Claw — Peter Steinberger – YouTube
Peter Steinberger: How I created OpenClaw, the breakthrough AI agent | TED Talk
.@KAIST_AI and @nyuniversity proposed a cross-domain shared memory for coding agents This idea is called Memory Transfer Learning (MTL) Build one big memory pool from many different kinds of coding tasks and let the agent reuse that memory across domains → This memory can
https://x.com/TheTuringPost/status/2045453668217172297
a bunch here where I’m saying ok Garry’s kinda right?! 👀…in some ways 🙂 we’re making this loop much easier to close out of the box soon If more people get into evals & traces to ground self-improving agents from Garry’s posts, there’ll be no one happier than me have written
https://x.com/Vtrivedy10/status/2046979341427331522
Are the Costs of AI Agents Also Rising Exponentially? — Toby Ord
https://www.tobyord.com/writing/hourly-costs-for-ai-agents
Bring your own key in @code is now available to all Copilot plans, including Free, Pro, Pro+, Business, and Enterprise! Use the best agent harness with local providers like @lmstudio or cloud providers like @OpenRouter 🙂
https://x.com/pierceboggan/status/2046985841596354815
Cursor scales code retrieval to 1T+ vectors with turbopuffer
https://turbopuffer.com/customers/cursor
Mention @Cursor to kick off tasks in Slack and see updates of its work streaming in real time. Cursor uses context in the thread and broader channels to create a PR for you to review and ship.
https://x.com/cursor_ai/status/2047000517751288303
One of my favorite features we’ve shipped in Fleet recently is the presentation viewer. Fleet agents are very good at writing code, so we thought why not have a built in way for them to design slides, and present them directly from the app! To go along with this, we’re also
https://x.com/BraceSproul/status/2047417882423022034
over break i dictated to 5.5 for minutes describing a new ambitious rl run. hit send and forgot about it as i hung out with friends and bf for a few days. returned on monday to an industrial-scale rl run humming after it worked for 31 hours
https://x.com/aidan_mclau/status/2047388367705575701
Small step for a finger on “”post tweet”” button. Big step for millions of future AI agents doing useful work for us.
https://x.com/MillionInt/status/2046659157688996251?s=20
there are early signs of 5.5 being a competent ai research partner. several researchers let 5.5 run variations of experiments overnight given only a high level algorithmic idea, wake up to find completed sweep dashboards and samples, never having touched code or a terminal at all
https://x.com/tszzl/status/2047386955550470245?s=46
This model can write better code than any model I’ve used before. It also needs a lot of encouragement. Context has never been more important. If you don’t give it strict enough instructions, it will explore. If it finds something “”wrong””, it feels almost impossible to steer
https://x.com/theo/status/2047379702189310085
Through the end of this weekend, we are doubling Composer 2 usage limits inside of Cursor’s new agents window. Enjoy!
https://x.com/cursor_ai/status/2045236540784492845
Today, we are excited to introduce
https://t.co/F5gmrAYGFP — a new tool to help site owners understand how they can make their sites optimized for agents.
https://x.com/Cloudflare/status/2045126394418503846
We recently shipped quality-of-life improvements to the Cursor CLI to make working with agents in the terminal more delightful. Use /debug to find root causes and fix tricky bugs that are hard to reproduce or understand.
https://x.com/cursor_ai/status/2046324136377721128
we’ve just released memory for agents primitives in Agents SDK when you need them, managed when you don’t
https://x.com/whoiskatrin/status/2045139949939200284
Xiaomi MiMo-V2.5 Series: Pushing Open-Source Agents Forward 🔸 MiMo-V2.5-Pro, our strongest model yet. A major leap from MiMo-V2-Pro in general agentic capabilities, complex software engineering, and long-horizon tasks, now matching frontier models like Claude Opus 4.6 and
https://x.com/XiaomiMiMo/status/2046988157888209365
An obvious way to release Mythos class models with uncertain autonomous ability is to make them only available on the website, like Gemini Deep Think or ChatGPT Pro. Minimal risk of being used for autonomous hacking, but accessible to people who have hard problems to solve.
https://x.com/emollick/status/2045916298450784680
A real issue with the current state of our knowledge on the work implications of AI is that there was a genuine discontinuity in AI ability with the rise of practical agentic systems in 2026. We were starting to get a picture of the impact of chatbots, no real data on agents.
https://x.com/emollick/status/2044576806926512446
Benchmarking Inference Engines on Agentic Workloads | Applied Compute
https://www.appliedcompute.com/research/inference-benchmark
Cloud agent infrastructure has a lot of moving parts: VM isolation, session persistence, environment provisioning, orchestration, integrations. Each one is its own engineering challenge. In this post, we break down what it takes to build cloud agent infrastructure from
https://x.com/cognition/status/2047392064355377194
Glad to be a part of this initiative to develop open-world evaluations for AI. We need the ability to assess just how capable agents are becoming in order to anticipate and respond to the impact they can have on real world systems and transactions. An agent that can successfully
https://x.com/ghadfield/status/2045245020429570505
had a blast on the pod (first one kinda nervous 😅), big shoutout to @himanshustwts just felt like us riffing on agent engineering, open source & research ❤️ some fun highlights – working backwards from the model’s capabilities/flaws and building systems (a harness) around them
https://x.com/Vtrivedy10/status/2046942634321559707
Have AI capabilities accelerated? On 3 out of the 4 AI capability metrics we investigated, we found strong evidence of acceleration, around when reasoning models emerged.
https://x.com/EpochAIResearch/status/2045205780916560010
Let’s talk content faithfulness. Four days ago, we launched ParseBench, the first document OCR benchmark for AI agents. Its most fundamental metric asks: did the parser capture all the text, in order, without making things up? We grade three failure modes with 167K+ rule-based
https://x.com/llama_index/status/2045145054772183128
Let’s talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDataPointMatch. Most document look at a chart and OCR the caption. Agents need the actual numbers. That’s the gap between “”OCR’d the
https://x.com/llama_index/status/2046586730879283227
This is crazy. ml-intern just passed the @huggingface internship test in 15 minutes. The task: replicate a research baseline from a DeepMind paper on test-time compute scaling. Here’s what the agent did: – Read the DeepMind paper, dug into Appendix E, picked the right scoring
https://x.com/akseljoonas/status/2047332440025321796
what do we think about perplexity/NLL eval for post trained models? cursor composer 2 did use it to choose the starting model (but didn’t end up choosing the best one according to NLL)
https://x.com/eliebakouch/status/2045115926123520100
Yesterday, we announced CRUX, a project that aims to conduct regular “open-world evaluations,” where we will be testing the ability of AI agents to complete long-horizon tasks in messy, real-world environments. @sayashk’s post dives into the details; here are a few of my own
https://x.com/PKirgis/status/2045265295649231354
This paper makes a strong case for open-world evaluations as a complement to traditional benchmarks, particularly for realistic, long-horizon, open-ended settings! Glad the AISI SoE team could contribute to this effort.
https://x.com/CUdudec/status/2045139195220431022
ParseBench is the first benchmark to include VLM chart understanding 📊📈📉 over enterprise documents. 🟠 Existing benchmarks (ChartQA, ChartXiv) test over charts specifically and not the chart’s inclusion in the overall document. Also doesn’t contain references to real-world
https://x.com/jerryjliu0/status/2046725527806021937
HF becoming the platform for agents (assisted by their humans) to use and build AI (rather than just leveraging APIs)!
https://x.com/ClementDelangue/status/2046598219853951346
Exclusive: Microsoft To Shift GitHub Copilot Users To Token-Based Billing, Tighten Rate Limits
https://www.wheresyoured.at/news-microsoft-to-shift-github-copilot-users-to-token-based-billing-reduce-rate-limits-2/
Access GPT Image 2.0 natively in Hermes Agent Update now to get access – just run `hermes update` and select your image generation tool model with `hermes tools`
https://x.com/NousResearch/status/2046693872773062834
✨ We’re excited to share that gpt-image-2 will be coming shortly to Canva AI 2.0! From highly-detailed generations to its creative intelligence, we can’t wait to see what you create. And with Magic Layers, you can edit everything like a design 🎨
https://x.com/canva/status/2046665346161988062
Playing with GPT Image 2 and really noticing the upgrade in fine detail + overall cohesion. Everything just feels like it belongs together a bit more ☀️ textures, lighting, composition all click. Try the model now in Firefly! Check the comments for the prompt 👇
https://x.com/AdobeFirefly/status/2046675148065923103
Same prompts as before, but now in GPT image-generator 2, page excerpts from: “”Eldritch Horrors as Pets: A Guide”” “”How Womblenauts Work”” “”Photographs of the People of New York Who Look Like Birds”” “”Cakes shaped like fish shaped like cakes”” Lots of great little lines in there
https://x.com/emollick/status/2046678198826479667
Its noticeable how much of the whole practice of working with AI – the prompts, the skill files, the connectors, retrieval work, the markdown files, etc. – is a substitute for the real problem of continual learning. If that ends up being solved, a lot of things will change fast.
https://x.com/emollick/status/2044792241000943867
Late-interaction retrieval models are widely used for their strong performance, but their representations can be utilized beyond just retrieval. Our new paper demonstrates that these representations can effectively replace raw document text in RAG tasks.
https://x.com/Julian_a42f9a/status/2045200413402493064
The new generation of open state-of-the-art single and multi-vector retrieval models is here It’s time, DenseOn with the LateOn 🎶 @LightOnIO releases models that leap past existing ones, and everything you need to do the same!
https://x.com/antoine_chaffin/status/2046609241918579019
We’re releasing LateOn and DenseOn today. Two open retrieval models, 149M parameters each. LateOn (ColBERT, multi-vector): 57.22 NDCG@10 on BEIR. DenseOn (dense, single-vector): 56.20. Both beat models up to 4× larger We’re open-sourcing the weights under Apache 2.0 🧵👇
https://x.com/raphaelsrty/status/2046609364929187845
We were made aware of concerns regarding the visibility of chat messages and code on Lovable projects with public visibility settings. To be clear: We did not suffer a data breach. Our documentation of what “public” implies was unclear, and that’s a failure on us. Specifically
https://x.com/Lovable/status/2046270357674299623?s=20
// Self-Evolving Agent Protocol // One of the more interesting papers I read this week. (bookmark it if you are an AI dev) The paper introduces Autogenesis, a self-evolving agent protocol where agents identify their own capability gaps, generate candidate improvements,
https://x.com/omarsar0/status/2045241905227915498
// Stateless Decision Memory for Enterprise AI Agents // Most of the interesting AI agent papers right now are about capability. This one is about plumbing, and it’s probably more important than it looks. One of the few agent memory works focusing on production-grade
https://x.com/omarsar0/status/2047325132096758228
🔒`deepagents deploy` now supports custom auth: turn a single deployment into a multi-tenant platform. per-user resource and thread isolation, access control, and your own auth provider, all without spinning up extra infrastructure.
https://x.com/sydneyrunkle/status/2046643201738449076
A good AGENTS.md is a model upgrade. A bad one is worse than no docs at all. | Augment Code
https://www.augmentcode.com/blog/how-to-write-good-agents-dot-md-files
A new protocol that can become a useful part of agentic workflows – Autogenesis Protocol (AGP) It’s all about structuring self-improvement of agentic systems to make it safe, clear and continuous. The idea is to separate what evolves from how changes happen That’s why AGP has
https://x.com/TheTuringPost/status/2046254041051943157
A survey that deserves your attention – Externalized Intelligence in LLM Agents It explains a shift from intelligence inside the model weights to intelligence in the system around it, showing: • How capability is increasingly coming from: – memory systems – persistent state –
https://x.com/TheTuringPost/status/2045988056088678667
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
https://agent-tars-world.github.io/-/
Agents can’t choose between structure and flexibility
https://frontierai.substack.com/p/agents-cant-choose-between-structure
AI Agent Frameworks Workshop | AWS Marketplace
https://pages.awscloud.com/awsmp-gro-hjua-webinar-mss-module-4-multi-agent-architectures-ai-series.html?trk=cac541b5-48a5-40fe-b9f3-5ced632e84d2&sc_channel=el
Coding Agents | Band
https://www.band.ai/for-coding-agents
developing an agent is a harness problem. deploying an agent is a runtime problem. here’s a guide to the runtime capabilities that keep agents running in production and the infrastructure that enables self improvement!
https://x.com/sydneyrunkle/status/2046284044942397744
Evals ~= Environments…they’re one of the best investments a team can make for improving agents Step 0: Turn On Tracing for Agents Step 1: Point compute at Traces to understand agent behavior, segment useful tasks, and isolate error modes Step 2: Turn Trace Data into
https://x.com/Vtrivedy10/status/2047362615836336473
Governing multi-agent systems at scale is where complexity explodes An upcoming live session with @rungalileo co-founder @YashSheth46 and @crewAIInc founder @joaomdmoura will help you master it. You’ll learn how to: – Enforce safety and security policies in agents – Steer
https://x.com/TheTuringPost/status/2044915828542603502
Great paper on self-improving agents. Why? We need to think more deeply about AI agent system design. The protocol specifies a framework for proposing, assessing, and committing improvements with auditable lineage and rollback. Visual below (courtesy of my research agent).
https://x.com/omarsar0/status/2045956901750399374
Great to see more people adopting ADP to standardize agent trajectories! If you’re building large agent training datasets, please reach out and we’ll be happy to help.
https://x.com/gneubig/status/2046963826109689983
i haven’t seen a model that just works across agent harnesses. seems like it should exist. great opportunity for open-weight models. any thoughts?
https://x.com/omarsar0/status/2047006936306962754
Introducing Deep Max: State-of-the-Art Agentic Search | Exa Blog
https://exa.ai/blog/deep-max
Introducing ml-intern, the agent that just automated the post-training team @huggingface It’s an open-source implementation of the real research loop that our ML researchers do every day. You give it a prompt, it researches papers, goes through citations, implements ideas in GPU
https://x.com/akseljoonas/status/2046543093856412100
LLM agents are assumed to integrate unexpected environmental observations into their reasoning. It turns out they don’t. We added the complete task solution into agent environments as a file or an API endpoint, and measured whether agents act on what they discover. They almost
https://x.com/LeonEnglaender/status/2046621862214488473
LLM agents loop, drift, and get stuck on hard reasoning tasks up to 30% of the time. Current fixes are either too blunt (hard step limits) or too expensive (LLM-as-judge adding 10-15% overhead per step). New research proposes a smarter middle ground. The work introduces the
https://x.com/omarsar0/status/2045139481779696027
RLM means notebooks are gonna be back (I hope). Agent driving a REPL with interleaved prose. The exact backend the nb interface is for. LOTS of have been swirling around the idea, but RLM solidified it and made it work. So I hope to see big NB releases soon! p.s. If you
https://x.com/isaac_flath/status/2046588093399019918
The Agent Harness Is a Shell | blog | inference.sh
https://inference.sh/blog/opinions/harness-is-a-shell
The idea of training LLMs to manage their own KV cache is super interesting to me. The recent neural garbage collection (NGC) paper was a great read on this topic. Reasoning models / agents obviously need long sequences to handle complex reasoning, long horizon tasks, tool
https://x.com/cwolferesearch/status/2047476297031631102
Things don’t always go to plan when bringing agents into production. ‘deepagents deploy’ is purpose-built around the challenges teams face when deploying. ✅ Every infrastructural consideration is mapped to a purpose-built runtime capability. No need to rebuild components from
https://x.com/LangChain/status/2046275653335462128
we just shipped support for subagents with `deepagents deploy`! add an agents/ dir to your project with an AGENTS.md per specialized subagent. subagents are great for task delegation with isolated/optimized context
https://x.com/sydneyrunkle/status/2045209395881980276
we want to help you navigate all the stuff that comes with shipping an agent to production – so we wrote a guide! this guide packages learnings across our + customer deployments, building open source agent infra/frameworks/examples, debugging infra errors, and building great
https://x.com/Vtrivedy10/status/2046280543978057892
We’re launching the beta for our new commercial AI product: Sakana Fugu 🐡, a multi-agent orchestration system! Blog:
https://t.co/c7wlMoX1W2 Fugu hits SOTA on SWE-Pro, GPQA-D, and ALE-Bench, and has been our internal secret weapon. It dynamically coordinates frontier models,
https://x.com/SakanaAILabs/status/2047479445209145785
Moonshot AI: “”Our RL infra team used a K2.6-backed agent that operated autonomously for 5 days, managing monitoring, incident response, and system operations, demonstrating persistent context, multi-threaded task handling, and full-cycle execution from alert to resolution””
https://x.com/scaling01/status/2046250343479054540
An update on recent Claude Code quality reports \ Anthropic
https://www.anthropic.com/engineering/april-23-postmortem
Anthropics works on its always-on agent with UI extensions
https://www.testingcatalog.com/anthropics-works-on-its-always-on-agent-with-new-ui-extensions/
Claude remains irreducibly Claude. If you know, you know. (The fact that models have distinct personalities that are consistent across generations is technically interesting, it also makes it very easy to use new releases when they come along, because they feel very similar).
https://x.com/emollick/status/2044799110088130992
I was told by Anthropic that they are looking at ways of fixing this, which is good (you can also see a reply from a Claude PM in the thread).
https://x.com/emollick/status/2044958121731195185
It’s fascinating is how little of Claude Code is actually “”intelligence.”” This study found a tiny reasoning core wrapped in massive infrastructure, and even quantifies it. → Only ~1.6% of the system is actual decision logic, while ~98.4% is operational harness: ~512K lines
https://x.com/TheTuringPost/status/2046726989021888910
Must-read research of the week ▪️ Dive into Claude Code: The design space of today’s and future AI agent systems ▪️ Lightning OPD: Efficient Post-Training for LRMs with Offline On-Policy Distillation ▪️ Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense
https://x.com/TheTuringPost/status/2046710304999104954
My one takeaway from the leaked Claude Code: a good agent harness should get out of the way. As @polynoamial once put it, “”Your fancy AI scaffolds will be washed away by scale”” Claude Code’s harness is the opposite of a fancy scaffold: it’s simple, but to the point ;
https://x.com/AymericRoucher/status/2045176781414527305
Over the past month, some of you reported Claude Code’s quality had slipped. We investigated, and published a post-mortem on the three issues we found. All are fixed in v2.1.116+ and we’ve reset usage limits for all subscribers.
https://x.com/ClaudeDevs/status/2047371123185287223
Really liking Claude Design so far. Except for the fact that it just wiped out my project after burning 10% of my usage. “”The files appear to be gone”” 🙃
https://x.com/theo/status/2045310884717981987
tldr: claude code changed some harness settings which degraded perf. these small harness tweaks can matter a lot! 1. default reasoning high -> medium 2. bug that accidentally evicted thinking blocks on every turn in session (march 26-april 10). was a change to help with cache
https://x.com/Vtrivedy10/status/2047384831995371631
A lot of bugs that folks may have hit yesterday when first trying Opus 4.7 are now fixed. Thanks for bearing with us🙏
https://x.com/alexalbert__/status/2045159041283064095
A major lesson to take away from Opus 4.7 is that, while there is a lot of arguments about implementation choices and personality, models keep improving measurably on economically important tasks with each release (it has been two months since Opus 4.6), with no signs of slowdown
https://x.com/emollick/status/2045314251804324080
Changes in the system prompt between Claude Opus 4.6 and 4.7
https://simonwillison.net/2026/Apr/18/opus-system-prompt/
Claude Opus 4.7 by @AnthropicAI advances the price-performance Pareto frontier in both Code and Text Arena! This makes Claude Opus 4.7 now the only model from a US lab that remains on the Pareto frontier for Code Arena.
https://x.com/arena/status/2045206342173086156
Claude Opus 4.7 by @AnthropicAI also lands at #1 and #3 in the Text Arena. Opus 4.7 Thinking ranks #1 across major categories: – #1 Overall 1505, +9 points over Muse Spark – #1 Expert 1561, +19 points over 4.6 Thinking – #1 Coding 1567, +23 points over 4.6 Thinking – #2
https://x.com/arena/status/2045177497378316597
I have found that asking for a sestina regularly triggers Opus 4.7’s safety guardrails. The forbidden poetic form!
https://x.com/emollick/status/2044863531900686775
I’ll give Anthropic credit for moving quickly. Opus 4.7 Adaptive Thinking now triggers thinking much more often, including for the tasks it failed at yesterday. That also means it is doing a lot more web search. So far, a large improvement in output quality on non-coding tasks.
https://x.com/emollick/status/2045147490316374414
Introducing the next-gen AI for design and creation — Genspark Build 🚀 Powered by Claude Opus 4.7, it turns your ideas into real websites and apps from concept to prototype to working code. Now in Public Preview: all Plus and Pro users get 3 days of zero-credit access (April
https://x.com/genspark_ai/status/2046610783203975539?s=20
With max thinking Opus 4.7 is quite impressive, with a real sense of style In two prompts: “”implement the Tower of Babel, in 3D, in as sophisticated and visually interesting a way as possible. It should be interactive”” and then “”make it better.”” Play:
https://x.com/emollick/status/2044966818339594252
Looks like the Anthropic “”safety layers”” aren’t just blocking prompts anymore, they’re erroneously banning entire orgs 🙃
https://x.com/theo/status/2045317666383204423
Sam Altman throws shade at Anthropic’s cyber model, Mythos: ‘fear-based marketing’ | TechCrunch
Sam Altman throws shade at Anthropic’s cyber model, Mythos: ‘fear-based marketing’
The progress on some of these benchmarks has been insane! @AnthropicAI @DarioAmodei May I please ask you to request Claude to give you a list of the of the top 1000 areas of STEM, top 1000 magazine topics, top 500 professions, and for each list item pick a (not in training
https://x.com/NandoDF/status/2045063560716296450
[2604.14228] Dive into Claude Code: The Design Space of Today’s and Future AI Agent Systems
https://arxiv.org/abs/2604.14228
please show me the >6 hour METR time horizons, WeirdqML, GSO, PPBench, LisanBench, SimpleBench, RLI or ARC-AGI-2 scores if you are so confident that open-source models are at GPT-5.4 or Opus 4.5+ level there’s still a big gap and it’s likely getting bigger
https://x.com/scaling01/status/2046565191903511010
Redwood Research presents LinuxArena – 20 live production environments for AI agents – Frontier models achieve ~23% undetected sabotage vs. trusted monitors – Useful work ≈ attack surface → sandboxing fails, monitoring is essential
https://x.com/arankomatsuzaki/status/2046070569758752984
Exciting news – Claude Opus 4.7 from @AnthropicAI takes #1 in Code Arena! +37 points over Opus-4.6 and +46 over the next non-Anthropic model, GLM-5.1 (#4). Massive ~130 pts lead over GPT-5.4 and Gemini-3.1-Pro. #1 on both React and HTML leaderboards. Code Arena evaluates
https://x.com/arena/status/2045177492936532029
The continuing gap between the capabilities of Gemini Pro 3.1 (very good model) and the capabilities of the Gemini app/website is odd. The model can do what Claude/GPT can do, but there is a minimal harness for tools (file creation, research etc), no auditable CoT/actions, manual
https://x.com/emollick/status/2045909435315323321
I think the adaptive thinking requirement in Claude Opus 4.7 is bad in the ways that all AI effort routers are bad, but magnified by the fact that there is no manual override like in ChatGPT. It regularly decides that non-math/code stuff is “”low effort”” & produces worse results.
https://x.com/emollick/status/2044864822076969268
Opus 4.7 better than Opus 4.6 but can’t beat Gemini 3.1 Pro and GPT-5.4 on LiveBench
https://x.com/scaling01/status/2045178622617498084
Opus 4.7 scores 156 on ECI, our tool for combining multiple benchmarks onto a single scale. This puts it a bit ahead of Opus 4.6 and a bit behind only GPT-5.4, Gemini 3.1 Pro, and GPT-5.4 Pro. Thread with individual scores and commentary.
https://x.com/EpochAIResearch/status/2046631622909558857
🧭 gog 0.13 is out! Gmail forwarding with notes + attachments, autoreplies, full-body search, Markdown uploads to Google Docs, rendered Slides thumbnails, Sheets chart editing, secondary calendars, commenter-only Drive shares, and safer no-send controls.
https://x.com/steipete/status/2046356596683411924
Google adds subagents to Gemini CLI to handle parallel coding tasks
https://tessl.io/blog/google-adds-subagents-to-gemini-cli-to-handle-parallel-coding-tasks/
Google Creates Strike Team to Improve Coding Models — The Information
https://www.theinformation.com/articles/google-creates-strike-team-improve-coding-models
AI Infrastructure: Cloud TPUs | Google Skills
https://www.skills.google/paths/2806/course_templates/1405
At #googlecloudnext today, we are introducing Workspace Intelligence Today’s digital workflows are information-rich but context-poor; project details live in Docs, trackers in Sheets, decisions are tucked in meeting notes, and updates are scattered across emails and chats.
https://x.com/ChanduThota/status/2046946043078848788
Last fall, we launched Gemini Enterprise as a front door to AI in the workplace, enabling every customer and employee to use cutting-edge AI agents. Today at #GoogleCloudNext, we announced new features, including: – A new inbox in Gemini Enterprise to manage, monitor and act
https://x.com/Google/status/2046988686433108417
ReasoningBank, a novel agent memory framework, enables LLM agents to continuously learn from both successful & failed experiences. Our evaluation shows that it enhances agent effectiveness, boosting success rates and efficiency. Learn more:
https://x.com/GoogleResearch/status/2046631948437921801
TPU 8t and TPU 8i technical deep dive | Google Cloud Blog
https://cloud.google.com/blog/products/compute/tpu-8t-and-tpu-8i-technical-deep-dive
Muse Spark is #3 on ClawEval, ahead of GPT-5.4 and Gemini 3.1 Pro. It is honestly a surprisingly agentic model.
https://x.com/alexandr_wang/status/2045348588734066794
Ollama now natively supports Copilot CLI! Bring your own models, and even work completely offiline!
https://x.com/_Evan_Boyle/status/2045926113889989057
partnering with @Kimi_Moonshot to bring kimi k2.6 to @CloudflareDev workers ai on day 0 better for coding and agentic use cases! try it out now:
https://x.com/michellechen/status/2046297037742997909
🎉 Congrats to the Moonshot team on Kimi K2.6 — day-0 support on vLLM 0.19.1. • 1T total / 32B active MoE — 384 experts, 8 routed + 1 shared • MLA attention, 256K context • Native multimodal: MoonViT vision encoder + video input • Native INT4 quantization • Interleaved
https://x.com/vllm_project/status/2046251287206035759
FINALLLY FINALLY it is here. V4-flash: all the way back to V2 prices, only now with 1M V4-pro: roughly Kimi/GLM/MiMo competitor Chat prefix completion and FIM back – thank you! Missed this forever but what can they do?
https://x.com/teortaxesTex/status/2047508587883250112
Kimi 2.6 Thinking seems very good for an open weights model, but many rough edges compared to closed SoTA. The Lem Test resulted in a 74 page thinking trace… and an okay-ish answer. It did an okay TiKZ unicorn, an adequate twigl shader for a neogothic city in the waves, etc.
https://x.com/emollick/status/2046411222354989189
Kimi K2.6 + DFlash: 508 tok/s on 8x MI300X 5.6x throughput improvement over baseline autoregressive serving 90 tok/s → 508 tok/s on the same hardware, same model, zero quality loss
https://x.com/HotAisle/status/2046620289984057634
Kimi K2.6 demonstrates strong long-horizon coding in complex engineering tasks: Kimi K2.6 successfully downloaded and deployed the Qwen3.5-0.8B model locally on a Mac. By implementing and optimizing model inference in Zig–a highly niche programming language–it demonstrated
https://x.com/Kimi_Moonshot/status/2046531052957569211
Kimi K2.6 has landed, and it is live on Baseten! We have baked in multiple inference optimizations so that you can leverage Kimi K2.6 in production right away. To run Kimi K2.6, Baseten uses: -> The Baseten Inference Stack with advanced optimizations, including KV-aware routing
https://x.com/baseten/status/2046263526281576573
Kimi K2.6 helped us rewrite kernels; it worked like a charm 🙂
https://x.com/Yulun_Du/status/2046252918526071017
Kimi K2.6 is live on OpenRouter! @Kimi_Moonshot’s new model is a long-horizon coding model built for sustained agentic work. It behaves more like a systems engineer than a chatbot, with the stamina to decompose, execute, and optimize complex tasks. Try it in all your favorite
https://x.com/OpenRouter/status/2046259590774571199
Kimi K2.6 is now available in Windsurf! Available for free for the next 2 weeks for Pro, Teams, and Max users.
https://x.com/windsurf/status/2046686574793154996
Kimi K2.6 now in OpenCode — Go included
https://x.com/opencode/status/2046275886396125680
Kimi K2.6 was released 1h ago, and it looks amazing! Here it’s running with MLX (mlx-vlm) on two M3 Ultras (full 1T param VLM) 🔥
https://x.com/pcuenq/status/2046283942689456297
Meet Kimi K2.6: Advancing Open-Source Coding 🔹Open-source SOTA on HLE w/ tools (54.0), SWE-Bench Pro (58.6), SWE-bench Multilingual (76.7), BrowseComp (83.2), Toolathlon (50.0), Charxiv w/ python(86.7), Math Vision w/ python (93.2) What’s new: 🔹Long-horizon coding – 4,000+
https://x.com/Kimi_Moonshot/status/2046249571882500354
Moonshot AI launches Kimi K2.6 on Kimi Chat and APIs
https://www.testingcatalog.com/moonshot-ai-launches-kimi-k2-6-on-kimi-chat-and-apis/
Qwen3.6-27B can now run locally! 💜 Run on 18GB RAM via Unsloth Dynamic GGUFs. Qwen3.6-27B surpasses Qwen3.5-397B-A17B on all major coding benchmarks. GGUFs:
https://t.co/ykKgwh2zI9 Guide:
https://x.com/UnslothAI/status/2046959757299487029
Ran Qwen3-8B (8.2B dense, open) on LongCoT-Mini. Vanilla: 0/507. dspy.RLM: 33/507 (6.5%). Same model. Same weights. No fine-tuning. The scaffold is doing 100% of the lifting. Context: leaderboard’s smallest open MoE is GLM-4.7 at 358B total / 32B active params. Qwen3-8B is ~4x
https://x.com/raw_works/status/2045208764509470742
these questions are silly Kimi > all other open-source models tho
https://x.com/scaling01/status/2046591683198906542
We’re open-sourcing FlashKDA — our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. Achieves 1.72×-2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention. Explore on github:
https://x.com/Kimi_Moonshot/status/2046607915424034839
OpenClaw 2026.4.20 🦞 🧠 Kimi K2.6 support + provider-aware /think 💬 BlueBubbles iMessage sends + tapbacks fixed ⏰ Cron state/delivery cleanup 🔐 Gateway pairing + plugin startup hardening Less haunted. More useful.
https://x.com/openclaw/status/2046686809367708123
Kimi K2.6 wrote an inference engine for Qwen3.5 0.5B in Zig and managed to beat LM Studio’s token per second by 20%, running for 12 hours and with 4000+ tool calls
https://x.com/nrehiew_/status/2046254256194474221
A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features”” TL;DR: feed-forward localization builds a lightweight feature map and estimates camera pose in one pass, achieving fast and accurate relocalization across large scenes
https://x.com/Almorgand/status/2045194191081251178
Z-Image experiment. I expanded the patch-2 layers to patch-4. New layers = patch-2 layers averaged over sub-patches (in) / replicated (out), so the weights are already close with zero training. Finetuning now to clean it up. If it works: 2× image size at the same compute.
https://x.com/ostrisai/status/2045677110413668743
Image models tend to get much more stuck on a particular direction than text models, requiring clearing the context window fairly often. PerfectSquashBench is my new measure of how image models anchor. The squash remains merely fine after many attempts.
https://x.com/emollick/status/2047073009312121000
What you need to know about the Deep Research and Deep Research Max Update: – Can consults over 100 sources in one research task. – Generates native charts and infographics inline. – Accepts PDFs, CSVs, images, audio, and video inputs. – Max version uses ~160 search queries per
https://x.com/_philschmid/status/2046627179551944753
Meta just released Sapiens2 on Hugging Face High-resolution vision transformers pretrained on 1 billion human images, for human-centric perception: pose, segmentation, normals, and pointmaps.
https://x.com/HuggingPapers/status/2047410529010844044
new image model coming with some real magic within, to unlock new use cases in productivity and creativity livestream noon today
https://x.com/gdb/status/2046632580527554572
MIT engineers built an AI wristband that controls robots by reading your hand muscles. It works by using an ultrasound to capture images of the muscles and tendons in your wrist. An AI algorithm then translates those images into the exact position of all 5 fingers, tracking 22
https://x.com/rowancheung/status/2045158931367072104
Context Unrolling in Omni Models – A unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations – Enables Context Unrolling, where the model explicitly reasons across multiple modal representations
https://x.com/arankomatsuzaki/status/2047519009004716097
New in LangSmith Fleet: Create and edit files with your agent. Your agent can now work with files directly. Create documents, presentations, and webpages inside a conversation, or upload your own files and edit them together. 📄 Work with images, PDFs, and text files. 💬 Build
https://x.com/LangChain/status/2047362259983495215
Tool Gateway is now live in Nous Portal. No separate accounts, no API key juggling. All you need is one subscription, and everything works. A paid Nous Portal subscription now includes access to 300+ models and a growing set of third-party tools. Launching with: → Web
https://x.com/NousResearch/status/2044878344592699744?s=20
Hermes Agent 🤝 Ollama
https://x.com/NousResearch/status/2045304840645939304
Hermes Agent 生态要炸了,这波进化速度把我整不会了! 刚从官方生态地图 Hermes Atlas 扒出来几个真·硬货,每一个单拎出来都是一个方向—- 1️⃣ hermes-agent-camel 内置 CaMeL 信任边界,Agent 自主跑任务不再翻车。生产环境终于敢上了!之前多少人卡在这步,现在直接破防。
https://x.com/NFTCPS/status/2046076635200553224
Hermes Agent: The Complete Beginner’s Guide Says written by me but was actually done by my Hermes agent and it used the Hermes Atlas knowledge base as a source It will be updated periodically as the canonical guide Now live on Hermes Atlas (link in replies)
https://x.com/KSimback/status/2046528526581383643
Hermes 一丢 Agent,全网程序员集体进化了! Nous Research 扔出 hermes-agent(90k+ stars),核心就一个词:自我进化。它不是玩具,而是带持久记忆、自动提炼技能、跨会话成长的底层骨架。 结果?社区直接把它当 DNA,短短几周卷出 80+ 进化体,生态总星 10 万+。这才是开源的最高境界:一个
https://x.com/GitTrend0x/status/2045142797439922337
Hermes 多 Agent 深水区:三个高级实战技巧 90% 的人用 Hermes,还停留在助手阶段:把所有需求塞进一个 Prompt,然后看着它串行执行。 这种用法在多 Agent 并发场景下有三个隐性代价: •Token 浪费:子 Agent 继承冗余历史信息。 •指令稀释:长上下文中关键指令权重衰减。
https://x.com/BTCqzy1/status/2045720855137903046
Introducing
https://t.co/JflLUfop4O V2 🤯 Whats New 👇🏻 🤯No fork required ⭐️New Hermes dark + light themes 🤖Agent View Office 🚀Conductor for agent missions 👥Operations for sub agent orchestration One liner: curl -fsSL
https://t.co/XsSfmAJvZx | bash
https://x.com/outsource_/status/2046079580105064787
ollama launch hermes Ollama 0.21 includes supports Hermes Agent, the self-improving AI agent built by @NousResearch.
https://x.com/ollama/status/2045282803387158873
Skillkit native support for Hermes agent is live NOW! Thanks to @ruffy0369 🔥
https://x.com/ghumare64/status/2046542176142733712
The Hermes Agent Creative Hackathon starts now 16 Days, $25k in Prizes Presented by @Kimi_Moonshot & @NousResearch For the tinkerers pushing Hermes Agent into creative domains: video, image, audio, 3D, long-form writing, creative software, interactive media and more. Show us
https://x.com/NousResearch/status/2045225469088326039
Xiaomi’s MiMo-V2.5 and MiMo-V2.5-Pro are both now available in Hermes Agent through Nous Portal and OpenRouter! Just `hermes update`!
https://x.com/Teknium/status/2047093325774385358
You can now scale depth as well as width with subagents! Just uncapped Hermes Agents’ sub-agent spawn width, and enabled spawn depth so sub-agents can be configured to spawn their own sub-agents! Looking forward to seeing what new use cases open up with this new flexibility.
https://x.com/Teknium/status/2046709250114957624
不会还有很多人运行Hermes Agent还在用黑底白字的命令行吧? 其实开源社区已经给出了几套非常成熟的 Web 面板方案 体验完全不输商业软件 今天帮大家盘点目前 GitHub 上专门适配 Hermes四大主流 Web UI方案,帮你找到最适合的那款! 以下是 4 种目前最主流面板的完整干货清单: 1.全能管家
https://x.com/0xMulight/status/2046071441469366368
为 Hermes AI agent 提供原生 macOS 图形界面,支持同时管理多个本地和远程 Hermes 服务器,实时可视化 agent 活动、会话、配置和系统状态
https://t.co/osk1ipQ4Kd Scarf 是一个 Swift 编写的 macOS 应用,用来给 Hermes AI agent 套一个图形界面。 2.0
https://x.com/QingQ77/status/2046592289540346020
我靠!Ollama 现在原生支持 Hermes Agent 了! 一行命令直接起飞: ollama launch hermes 就这?就这!本地部署这么简单你还在用什么云端? 不知道自己电脑能跑哪些模型的,两个方法选一个: 1️⃣ 用 llmfit 本地检测 2️⃣ 直接上网站查 别再说本地部署难了,难的是你没试过。 🔗
https://x.com/NFTCPS/status/2045730947501576460
Kimi K2.6 is now available in Hermes Agent. Simply run `hermes update` and use `hermes model` to select a compatible provider hosting the model!
https://x.com/NousResearch/status/2046300755683098910
和上交念 AI 专业的研究生朋友聊了一下openclaw 和 hermes,分享一下内容: 1. 速度 卸载 / 迁移 openclaw , openclaw 是垃圾 2. openclaw 的生态丰富,但是底层是”上下文窗口 + RAG”的方案,长时间用虾,非常容易技术串联,牛头不对虾嘴 3. hermes 比 openclaw 好,但并不完美
https://x.com/ResearchWang/status/2046080807186665594
🎚️CodexBar 0.21 Abacus AI provider, Codex Pro $100 support, safer OpenAI web extras, fixed local cost scanning, z. ai 5h quotas, Antigravity/Cursor/Ollama fixes, faster refreshes, macOS 26 icon fix and more. The big issue with too much CPU usage was an OpenAI web fetch and is
https://x.com/steipete/status/2045582547996856682
A hill that I will die on: with today’s AI models, intelligence is a function of inference compute. Comparing models by a single number hasn’t made sense since 2024. What matters is intelligence per token or per $. This is especially true when using it in a product like Codex.
https://x.com/polynoamial/status/2047387675762802998
A hill that I will die on: with today’s AI models, intelligence is a function of inference compute. Comparing models by a single number hasn’t made sense since 2024. What matters is intelligence per token or per $. This is especially true when using it in a product like Codex.
https://x.com/polynoamial/status/2047387675762802998?s=46
Also, a ton of new Codex features coming soon! Fun little bundle w/the new model.
https://x.com/sama/status/2047378431260664058?s=20
auto-review now live in codex — using a guardian agent to evaluate the safety of proposed actions, reducing human approvals to only when they’re really needed.
https://x.com/gdb/status/2047489218998628780
Build workspace agents for your team, on top of a cloud-hosted Codex harness. Hook them up to tools, give them recurring tasks, and talk to them from surfaces like Slack. Easier than ever to bring the power of agents to your computer work.
https://x.com/gdb/status/2047023089087606814
Chronicle is an experimental feature giving Codex the ability to see and have recent memory over what you see, automatically giving it full context on what you’re doing. Feels surprisingly magical to use.
https://x.com/gdb/status/2046293955009274019
Codex + 5.5 is incredible for the full spectrum of computer use. No longer just for coders, but for anyone who does computer work (including creating spreadsheets, slides, etc).
https://x.com/gdb/status/2047387783111868707
codex for proactively suggesting what it can do for you:
https://x.com/gdb/status/2045227305816281404
Codex is becoming a turbocharged partner for everything you want your computer to do for you:
https://x.com/gdb/status/2044855706273391084
codex is becoming the universal app for developers:
https://x.com/gdb/status/2045974850074996882
codex is for everyone. learn how to get the most out of it:
https://x.com/gdb/status/2045208278033142227
codex makes work plain fun
https://x.com/gdb/status/2045440270188364117
GPT-5.5 in Codex is a delight to work with: – Super sharp with responses – It understands intent better than any model – Great “”personality”” – Gets lots of stuff done without pausing unnecessarily It generated this beautiful artifact design. Huge win for OpenAI.
https://x.com/omarsar0/status/2047424707310289058
GPT-5.5 is rolling out today for Plus, Pro, Business and Enterprise users across ChatGPT and Codex. We’re also introducing GPT-5.5 Pro for Pro, Business, and Enterprise users in ChatGPT.
https://x.com/OpenAI/status/2047376568809636017
GPT-5.5 just dropped, I’ve been testing it for the last two weeks. tl;dr – It’s an incredible model, but there’s something different about this launch… OpenAI isn’t just going for raw intelligence. They’ve improved the personality of the model. This is almost certainly to
https://x.com/MatthewBerman/status/2047375703516361174
I am happy everyone is switching to Codex, but Tibo if you start rate limiting me or making me use worse models…
https://x.com/sama/status/2044921348540264614
idk what your AGI definition is but subagents & computer use in codex is pretty close!! *video in realtime
https://x.com/reach_vb/status/2045151640802771394
imagegen in codex is easy to underestimate, but it’s quite powerful:
https://x.com/gdb/status/2044994088739749996
In ChatGPT, full-stack inference improvements enable a more capable model at faster speed. This efficiency is a game-changer for GPT-5.5 Pro, now a much more practical option for demanding tasks, and a step change in the level of difficulty and quality of work ChatGPT can take on
https://x.com/OpenAI/status/2047376567559668222
incredibly fun to build webapps and games with codex, entirely with natural language
https://x.com/gdb/status/2045594591584530826
Last week, we released a preview of memories in Codex. Today, we’re expanding the experiment with Chronicle, which improves memories using recent screen context. Now, Codex can help with what you’ve been working on without you restating context.
https://x.com/OpenAIDevs/status/2046288243768082699
Last week, we released a preview of memories in Codex. Today, we’re expanding the experiment with Chronicle, which improves memories using recent screen context. Now, Codex can help with what you’ve been working on without you restating context.
https://x.com/OpenAIDevs/status/2046288243768082699?s=20
LETS GOOOO! Excited to introduce GPT-5.5 Thinking & Pro in ChatGPT and Codex 🔥 It’s our smartest model *yet* for real work: stronger agentic coding, computer use, knowledge work, long-context reasoning, and scientific research It can plan, use tools, check its work, recover
https://x.com/reach_vb/status/2047377562339524659
Lots of major improvements to Codex! Computer use is a real update for me; it feels even more useful than I expected. It can use all of the apps on your Mac, in parallel and without interfering with your direct work.
https://x.com/sama/status/2044858862042591378
New in the Codex app: – GPT-5.5 – Browser control – Sheets & Slides – Docs & PDFs – OS-wide dictation – Auto-review mode Enjoy!
https://x.com/ajambrosino/status/2047381565534322694
OpenAI develops platform for always-on Agents on ChatGPT
https://www.testingcatalog.com/openai-develops-platform-for-always-on-agents-on-chatgpt/
OpenAI’s first AI intern is expected by the end of this year, but we got impatient and decided to build it ourselves 🙂 > Runs autonomously for hours / days depending on the task. > Can read every paper, model, and dataset on the HF Hub to build the best post-training recipes
https://x.com/_lewtun/status/2046549090171764914
Opus 4.7 using ~10x less tokens to solve machine learning problems ~8.4x cheaper than Opus 4.6 and GPT-5.4 and 3.4x cheaper than GPT-5.3 Codex per run while having the same performance
https://x.com/scaling01/status/2045160883010081237
The second most important release of the LLM era (after GPT-3.5), featuring what was likely the most important chart. Still seems surprising to me that OpenAI told everyone about the biggest advance in AI technology since the LLM rather than keeping it to themselves until later.
https://x.com/emollick/status/2046053467941163055
We are releasing a *research preview* of Chronicle in Codex. It allows codex to build up memories based on your day to day work on your computer and then refer to these memories to be a lot more helpful. Available for PRO subscriptions and on Mac to start. This is early and
https://x.com/thsottiaux/status/2046291546325369065
We’re open-sourcing Cua Driver – our new macOS driver that lets any agent (Claude Code, Codex, your own loop) drive any app in the background, with true multi-player and multi-cursor built-in. 1/8
https://x.com/trycua/status/2047383200348221632
ChatGPT plugin now available for Google Sheets:
https://x.com/gdb/status/2047064885012599168
🚨 GPT Image 2 is live on fal, day 0! 🔤 Strong text rendering 🧭 Better layout + UI adherence 🛠️ Cleaner preserve-and-change edits 📷 Strong everyday photoreal output
https://x.com/fal/status/2046667081068761527
people are speculating GPT-Image-2 is testing on @arena. the early examples being posted are pretty mind-boggling. all three of these images are AI generated. h/t @sawlygg @synthwavedd
https://x.com/blakeir/status/2040250530375606401?s=12
Though the images are very good, ChatGPT Image 2.0 does have the typical imagegen problem, which is that editing can be “”stubborn””, and attempts to get the AI to change details work well for the first round or two, but then progress slows. Putting the image in a new chat helps.
https://x.com/emollick/status/2046672707517886500
GPT-5.5 is now accessible in Hermes Agent through the ChatGPT/Codex OAuth provider. Run `hermes update` to access now or learn how to get started with Hermes Agent here:
https://x.com/Teknium/status/2047419336537846193
OpenClaw 2026.4.15 🦞 🤖 Anthropic Opus 4.7 support 🗣️ Gemini TTS in bundled 🧠 Slimmer context + bounded memory reads 🔧 Codex transport self-heal, safer tool/media handling ✨ Pile of update/channel fixes Good boring release.
https://x.com/openclaw/status/2044919054402752638
Ex-OpenAI researcher Jerry Tworek launches Core Automation to build the most automated AI lab in the world
OpenClaw 2026.4.21 🦞 🖼️ OpenAI Image 2 🔧 npm update repair for bundled plugins 🐳 Docker E2E coverage for channel deps 🩹 Low-risk fixes backported Tiny release. Useful claws.
https://x.com/openclaw/status/2046807838459125990
A few weeks ago @steipete told me he was thinking about canceling his TED talk. Too busy. No time to prep. Working on OpenClaw and OpenAI simultaneously with zero bandwidth. He prepped the whole thing in a week, showed up to Vancouver still taking meetings between sessions,
https://x.com/bilawalsidhu/status/2045291456630509709
OpenClaw 2026.4.21 is live. Small release, important fix: npm updates now repair bundled plugin runtime deps, with Docker E2E coverage so Telegram/Discord/Slack do not break after upgrade. Also backports OpenAI Image 2 support. npm i -g openclaw@latest
https://x.com/steipete/status/2046803162590335240
🗃️ wacli 0.6.0 is out! Big security + reliability sweep for WhatsApp CLI. Hardens SQLite/store path handling, sanitizes search queries, recovers sync/media panics, adds WACLI_STORE_DIR, and improves SIGINT exits.
https://t.co/VabuMQgps5 props @sdinakar7 for doing the work!
https://x.com/steipete/status/2046375922031321401
8 React components per line 😅
https://x.com/steipete/status/2046991196786979210
Amazing to see almost 2000 people at ClawCon Michigan!! Weird lobster cult 🦞🥳
https://x.com/steipete/status/2044935917417205796
Did some work to get our CI times down from 8 to two minutes via some… parallelization. Kudos to the @useblacksmith folks for sponsoring + letting us melt their servers.
https://x.com/steipete/status/2046787353906167992
discrawl 0.3.0 is out. This release adds Git-backed archive sync, so a Discord archive can be published to a private repo and queried locally without every user needing bot credentials. Also: auto-refresh, activity reports, field notes, faster imports.
https://x.com/steipete/status/2046748122928263345
Interesting shift. These highly subsidized subs are out there to get your code to improve their models. If you use AI for things useful to you, but not code, you are not valuable to them.
https://x.com/steipete/status/2046199257430888878
Kudos to the folks from Tencent for working with us and providing evals to improve OpenClaw’s harness performance! We’re also working with them to bring fixes/improvements back to the open source repo. Great option for folks not comfortable with the terminal.
https://x.com/steipete/status/2046259696722465113
MCPorter 🧳 0.9.0 is out. Call MCPs from TypeScript or as CLI – per-server tool filtering – sturdier stdio shutdowns – Windows OAuth URL quoting fix – OAuth config docs – schema-declared string coercion for tool calls
https://x.com/steipete/status/2046192869497622529
Since this is blowing up on hacker news. Boris said that CLI usage is allowed. Thus we added support for it, only to find out that we are still blocked there. It is trival to work around with a few renames, but I don’t wanna play that game. So it’s in a weird limbo where cli use
https://x.com/steipete/status/2046685973233189375
they: OpenClaw is so insecure look at all these GHSAs! reality: we are just an indicator of the coming storm
https://x.com/steipete/status/2044888081141223442
Vancouver, it’s been a blast! 🇨🇦
https://x.com/steipete/status/2045276507527143629
We need open traces so that everyone can train open agent models! cc @steipete @badlogicgames @thdxr @matanSF @hwchase17
https://x.com/ClementDelangue/status/2046942871299772441
yes
https://x.com/steipete/status/2046724963550306498
New paper: “”Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs””. Our system (BLF) matches human superforecasters on ForecastBench, and beats all the top methods (GPT-5, Cassi, Grok 4.20, and Foresight-32B). 🧵
https://x.com/sirbayes/status/2046961503107166689





Leave a Reply