Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A single continuous concentric ROYGBIV rainbow ribbon in Julio Le Parc style loops and branches through small arrow-tipped nodes and finally forms the bold word AGENTS, its letters built from the same layered violet-to-red spectrum bands with tiny clean white gaps at each crossover, centered on a pale off-white background with abundant negative space, flat matte screen-print finish, crisp printed edges, no shadow, no gradient, no texture.

Apple’s Messages app on iPhone now has a third-party AI agent – 9to5Mac
https://9to5mac.com/2026/06/04/apples-messages-app-on-iphone-now-has-a-third-party-ai-agent/

Uber reportedly now caps coding agents at $1,500/month per employee per tool – seems sensible to me, but it’s also an interesting hint at the value Uber thinks these tools are providing
https://x.com/simonw/status/2062143151184465964

Can we design legal agent verifiers that are up to 1,000x cheaper? Verifiers are LLM judges that check an agent’s work against rubric criteria: they’re used both in agent benchmarking and as reward signal in post-training. But verifiers can be a bottleneck at scale. For
https://x.com/harvey/status/2061866491033899371

Here Opus 4.8 built and play-tested a new RPG in Claude Code, including 3 PDF manuals and adventures, playtest notes, a website, and a playable solo adventure – then put it all on Netlify. No feedback from me at all.
https://x.com/emollick/status/2060045063275573723

In early May, the best superforecasters predicted that, by the end of the year, the longest METR 80% task horizons would reach 3-4 hours. In late May, Claude Mythos achieved that number.
https://x.com/emollick/status/2062235461364445204

I had Opus 4.8 in Claude Code write a sophisticated, if minor, academic paper from a archive of hundreds of de-identified research files from years ago I had to use GPT-5.5 Pro as a reviewer, it spotted one major error & some minor points. Opus corrected
https://x.com/emollick/status/2060098885561778341

AI research is a series of next-step decisions. We looked at sessions where a human researcher took a wrong turn, showed Claude the session up to that point, and asked it what to do next. Mythos Preview improved on humans 64% of the time–up from 22% in 2024.
https://x.com/AnthropicAI/status/2062568870872003021

Expanding Project Glasswing \ Anthropic
https://www.anthropic.com/news/expanding-project-glasswing

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever lived. In addition to being neat, it is an interesting test of combining research, code, design and stats for an AI.
https://x.com/emollick/status/2060165879908749490

Introducing Claude Opus 4.8 \ Anthropic
https://www.anthropic.com/news/claude-opus-4-8

Today, Anthropic engineers on average ship 8x as much code per quarter as they did compared to 2021-2025.
https://x.com/AnthropicAI/status/2062568864240836995

When AI builds itself \ Anthropic
https://www.anthropic.com/institute/recursive-self-improvement

Introducing Gemma 4 12B
https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/

Today we’re introducing Gemma 4 12B — our latest open model that brings advanced agentic reasoning, vision and audio directly to your laptop. It delivers performance nearing our larger Gemma models with a much smaller total memory footprint, while being small enough to run
https://x.com/Google/status/2062203526588088452

Exclusive: Microsoft is building a super app that combines coding, chat, and other Copilot AI tools | Fortune
https://fortune.com/2026/05/29/microsoft-working-on-super-app/

Exclusive: New screenshots of upcoming Copilot Super App
https://www.testingcatalog.com/exclusive-new-screenshots-of-upcoming-copilot-super-app/

Microsoft scout revealed „your always-on personal agent for work.” If “”AI”” was the Word of the Year in 2025, in 2026 it will be “”agents”” (always-on). Everything is agentic this year.
https://x.com/kimmonismus/status/2061875714933371220

This came as a surprise: Microsoft has unveiled handheld and desktop devices designed to control one’s agents. It reminds me of what I had expected from OpenAI’s hardware-standalone devices for controlling agents.
https://x.com/kimmonismus/status/2061860319547527191

MAI-Transcribe-1.5 is available at $6 per 1,000 minutes of audio via Microsoft Foundry.
https://x.com/ArtificialAnlys/status/2061878498609053909

Today we announced MAI-Thinking-1, a strong generalist and reasoning LLM built from the ground up without distilling third-party models. 97% on AIME 2025; 53% on SWE-Bench Pro; preferred by human raters over Sonnet 4.6 (blind side-by-side). Tech report:
https://x.com/asadovsky/status/2062008312603070891

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks. – It’s a
https://x.com/mustafasuleyman/status/2061880164498428188

MAI-Image-2.5 is here — now #3 on text-to-image and #2 on image-to-image Arena leaderboards, surpassing Nano Banana Pro. Leading image generation. Precise editing. Built for enterprise scale. It delivers strong performance on H100s, enabling deployment on existing
https://x.com/MicrosoftAI/status/2062240400299934143

Building a hill-climbing machine: Launching seven new MAI models | Microsoft AI
https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/

OpenClaw is on Microsoft now, with all necessary security precautions 🤍 Peter Steinberger introducing it himself. Excited for the whole @openclaw team: @steipete, @davemorin, @vincent_koc to name a few
https://x.com/TheTuringPost/status/2061870411571466666

It’s finally here. The official Hermes Desktop app. Available on all platforms.
https://x.com/Teknium/status/2061844602735538266

Introducing the Cosmos Coalition A new global initiative with NVIDIA and leading AI labs to build and open-source frontier world models for physical AI. Runway joins as a founding member, working alongside NVIDIA and a set of leading AI labs to build, share and accelerate world
https://x.com/runwayml/status/2061315089869721682

Jensen just launched NVIDIA Cosmos 3. Pitched as the first fully open omnimodel for physical AI: a mixture-of-transformers (reasoning + generation) with native vision reasoning and generation across text, image, video, sound, and action. Tops open-model leaderboards on
https://x.com/TheHumanoidHub/status/2061333253920080345

NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI | NVIDIA Newsroom
https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-3-the-open-frontier-foundation-model-for-physical-ai

NVIDIA Nemotron 3 Ultra Powers Faster, More Efficient Reasoning for Long-Running Agents | NVIDIA Technical Blog
https://developer.nvidia.com/blog/nvidia-nemotron-3-ultra-powers-faster-more-efficient-reasoning-for-long-running-agents/

Building apps has never been easier. With Sites, Codex can turn your work, ideas, and plans into an interactive website or app your team can explore, use, and share with a URL. Rolling out to Business and Enterprise plans, before expanding more broadly.
https://x.com/OpenAI/status/2061845949170045346

Codex for every role, tool, and workflow | OpenAI
https://openai.com/index/codex-for-every-role-tool-workflow/

Dreaming: Better memory for a more helpful ChatGPT | OpenAI
https://openai.com/index/chatgpt-memory-dreaming/

Watch me control my computer with just my voice. This is the future of operating systems. No hands. GPT-Realtime 2.0 is very, very underrated. Demo:
https://x.com/FarzaTV/status/2060865350036750847

👏👏 Introducing Qwen3.7-Plus — a multimodal agent model that unifies vision and language into one versatile agent foundation. ✅ Multimodal interactive hybrid agent: unified GUI & CLI operation across visual and text tasks ✅ Versatile coding agent & productivity assistant with
https://x.com/Alibaba_Qwen/status/2061506641120641494

Qwen
https://qwen.ai/blog?id=qwen3.7-plus

♾️ Continuity Work started in the app can be run locally or in the cloud, and regardless of where it runs, can be viewed and continued across clients from CLI to Mobile to Web. Mobile notifications when tasks complete, the ability to quickly check in on sessions and continue
https://x.com/lukehoban/status/2061905448287322243

❤️⚡Code-1-Flash is available in VS Code. Play with it and give us feedback.
https://x.com/mariorod1/status/2061914993550143513

A brand new W&B Weave is live! It watches production agents end to end, flags failure modes on its own, runs a full loop from inference to training, and blocks regressions. You can finally watch how your agent thinks across millions of traces instead of squinting at one. 🫡
https://x.com/wandb/status/2061894943203831996

Agent harness performance doesn’t scale reliably with raw test-time compute like tokens, tool calls, wall time, or cost. Researchers from @HIT_1920 found that the key scaling variable is Effective Feedback Compute (EFC) – feedback that is informative, valid, non-redundant, and
https://x.com/TheTuringPost/status/2060890500006302105

Agentic RL: Token-In, Token-Out Done Right
https://qgallouedec-tito.hf.space/

AI should earn its keep. Introducing the AI Productivity Guarantee. If Devin delivers less engineering value than you’re paying for, Cognition will fund your usage until it does, up to $10 million. It’s time for the AI industry to stop maximizing tokens and start maximizing
https://x.com/cognition/status/2062597242167628019

all agents in the future are going to need to write and execute code LangSmith Sandboxes are GA – try them out today
https://x.com/hwchase17/status/2061496556608504043

Browser progress. Now you can open “”Remote tabs”” that run in Cloudflare Browser Run instances. Right click the tab to get a shareable CDP URL where you can hand off to your agent and watch it do things on the website for you (like fill out a form).
https://x.com/BraydenWilmoth/status/2062180110208311558

Check out our technical blog for the Agent Arena methodology + a deep dive into how people delegate, correct, and steer agents:
https://x.com/arena/status/2062566769659912281

Computer-use agents are moving from the cloud to your local machine. Fast. When we launched Holo3 two months ago, the production feedback was clear: digital agents need to be blazing fast, cost-effective, and versatile. Today, we’re dropping Holo 3.1, engineered to run
https://x.com/hcompany_ai/status/2061815355341725925

Control your agents in production. | LaunchDarkly
https://launchdarkly.com/pa/platform/agent-control/

GitHub’s plan for Agents — Kyle Daigle, GitHub
https://www.latent.space/p/github

I am hooked on Dynamic Workflows! The idea of generating harnesses on the fly is so compelling that I reverse-engineered it for my agent orchestrator. And then I built a monitoring dashboard (as an HTML artifact) to track tasks, metrics, and reports. I can now use and monitor
https://x.com/omarsar0/status/2062553527730540611

Interrupt 2026 Recordings | The Agent Conference by LangChain
https://interrupt.langchain.com/recordings

Introducing Devin Desktop. Manage fleets of local and cloud agents from one surface. Plan, delegate, review, and ship without leaving your editor.
https://x.com/cognition/status/2061889596703551926

KV Cache re-use is the most important thing for agentic rollouts. We’ve integrated Mooncake Store into prime-rl with vLLM, you can now use it as a drop-in replacement for native CPU/Disk offloading, giving you cross-node prefix cache reuse to make your agents go brrr🚀
https://x.com/m_sirovatka/status/2061862853997465738

Managed Deep Agents keeps the project shape you already know: ↳ AGENTS.md, skills/, subagents/, + tools.json Context Hub gives your agent a managed place to retain and update this context across sessions, allowing agent definition to evolve over time.
https://x.com/LangChain/status/2061432934993674267

MiniMax M3 is now available on PBD TokenRouter. The first open-weights model to combine frontier coding and agent capabilities (59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 74.2% MCP Atlas), up to 1M tokens of context enabled by MiniMax Sparse Attention, and native
https://x.com/PBDTokenRouter/status/2061463048485838935

New work on Multi-Agent Computer Use (MACU). The future of computer-use agents lies in multi-agent systems that combine planning, coordination, and parallel execution. Paper:
https://t.co/TFHVDz6Th8 Webside + Code:
https://t.co/QPUFKgEacu MACU introduces a manager agent
https://x.com/rsalakhu/status/2062194674794668066

Standalone IDEs have about 6 months left to live. An interface for manually editing and refactoring doesn’t need to exist if you’re not manually editing and refactoring anymore. So what’s the right interface for a dev to be working in for 8h / day? Some parts are obvious: you
https://x.com/ScottWu46/status/2061998361373532187

Startup discovery platform @harmonic_ai rebuilt Scout, their AI platform using Deep Agents and LangSmith. Deep Agents: One frontier model + two tool sets (global company data and firm-specific context). Long-horizon execution and context window management handled out of the
https://x.com/LangChain/status/2062204592562073972

The AI agent bottleneck isn’t model performance — it’s permissions | VentureBeat
https://venturebeat.com/orchestration/the-ai-agent-bottleneck-isnt-model-performance-its-permissions

The events of the last 6 months in technology are arguable amongst the most important in human history The tools now increasingly exist for recursive self improvement of models & agents We are likely in very early lift off & exponential Largely unnoticed outside of tech
https://x.com/eladgil/status/2061129428084887593?s=20

Today I’m launching a new project called SynthTraces 🔥 It is a minimal codebase to generate synthetic coding agent session traces using Pi (from @badlogicgames) I wanted a large number of coding-agent traces, so I built a tiny harness where two models talk to each other: – an
https://x.com/julien_c/status/2062524414034423969

Verification is the hidden bottleneck for knowledge work agents, especially in legal AI — complex, long-horizon work is graded by rubrics with dozens of strict criteria. In new research with @langchain Labs, we study how to verify legal agents more efficiently, and show open
https://x.com/nikogrupen/status/2061866707988431039

We partnered with @FireworksAI_HQ to train open-source models for legal. Here’s what we found: 1) Hybrid legal agents can beat frontier models on quality and cost by routing selectively to a frontier advisor. We tested a hybrid setup where GLM 5.1 served as the primary worker,
https://x.com/harvey/status/2062218656420167785

We’re moving away from search as a web fetch tool call to search as codegen to be future proof in a world where code execution inside agent harnesses is the way to do almost all of our knowledge work. Doing this lets you compose multi-step primitives far more naturally and be
https://x.com/AravSrinivas/status/2061575845056278971

We’re presenting ParseBench at CVPR 2026 today. 🦙 Come learn why document understanding is an AGI-complete problem (an agent can’t act on a doc it can’t correctly read, and reading a real enterprise table is harder than it looks). The first doc-parsing benchmark built for AI
https://x.com/llama_index/status/2062525204262236266

Windsurf is now Devin Desktop | Devin
https://devin.ai/blog/windsurf-is-now-devin-desktop

With canvases, Cursor can create apps like dashboards, reports, and internal tools. Now you can publish a canvas and share it with your team via URL.
https://x.com/cursor_ai/status/2062611883249783083

What is AGI-hard – Latent.Space
https://www.latent.space/p/agi-hard

When it comes to observability for agents, regular application tracing doesn’t cut it. We need tools that understand the specific semantics of agents (multi-turn sessions, tool calls, long context, etc.). This is what we built the new Weave for. Agent-first observability and
https://x.com/neutralino1/status/2061949197851742525

Role-specific plugins in Codex are built around the work teams actually do. Plugins for Data Analytics, Creative Production, and Product Design give Codex the tools and context to create reports, creative directions, and prototypes. Built and used by OpenAI teams.
https://x.com/OpenAIDevs/status/2061888366791246071

Honestly if we put this demo in a fancy XR headset, we’d call it jarvis. The future of 3d doesn’t involve the death of autocad, maya and blender. It turns those tools into a shared canvas of collaboration with ai agents.
https://x.com/bilawalsidhu/status/2061450274011591084

Introducing Agent Arena: real-world agentic evals at scale. How do you evaluate agents doing actual work? We measure millions of live sessions where real users accomplish real tasks. On Arena, models now get web search, filesystem, and terminal tools to complete complex
https://x.com/arena/status/2062566749418233981

Shared my first trace from @NanoClaw_AI to @huggingface yesterday. Very cool! By default, all agents should store their traces on HF (in private) so that you can keep a history of them, analyze them,… & share them and post-train better models, harnesses and more. Excited
https://x.com/ClementDelangue/status/2062542713463980303

Jerry Liu built one of the most installed pieces of AI plumbing of the last three years. Then he sat down and told me the framework era he helped create is over. The agent harness ate the abstraction layer. The patterns @llama_index used to wrap (query rewriting, reasoning
https://x.com/ConorBronsdon/status/2062224321381323218

Be There for Every Customer With Meta Business Agent
https://about.fb.com/news/2026/06/meta-business-agent/

Having trouble connecting Hermes Agent Desktop to your remote instance? Check out this updated guide:
https://x.com/Teknium/status/2062170975949721612

Just pushed an update to help remote connecting with the Hermes Agent GUI over tailscale to function! Please update if you had any issues!
https://x.com/Teknium/status/2061984430370267210

The next evolution of Hermes Agent is here! Introducing Hermes Desktop: everything you love about Hermes, now native on your machine. First demoed in Jensen’s GTC keynote, it’s now in public preview.
https://x.com/NousResearch/status/2061843507417944552

The idea of OpenClaw is always that it should be yours. It’s modular and lean, only add what you need. Fewer skills, fewer tools = your agent can work more efficiently.
https://x.com/steipete/status/2061072753998856696

Been teaching codex to be my QA assistant. For every commit it creates a user-test scenario and uses webVNC (crabbox), computer/browser use (peekaboo/mcporter) to test OpenClaw like a user/QA person would. This runs in the background and opens PRs with fixes.
https://x.com/steipete/status/2061208638027395490

Introducing Search as Code, our new search architecture for AI agents. It writes Python that calls our search stack directly, instead of looping through function calls one at a time. Available in the Perplexity Agent API, and now default in Computer.
https://x.com/perplexity_ai/status/2061506359326384319

Today we’re announcing that hybrid agentic inference is coming to Perplexity Computer. Computer can split tasks between a local model running on your machine and frontier models in the cloud. This keeps private data on your device and maximizes token efficiency. Coming soon.
https://x.com/perplexity_ai/status/2061861293569765847

For two years the whole conversation was about context window size. Meanwhile the actual problem never moved: agents don’t remember anything between sessions. We kept patching it with RAG and manual context injection and calling that memory. HydraDB is going at the layer
https://x.com/kimmonismus/status/2061454202883432501

Mellum started with code completion. Mellum2 is built for more – handling both natural language and code. A 12B-parameter open-source LLM for routing, RAG, and sub-agents, optimized for ultra-low-latency inference. Now on @huggingface. Learn more:
https://x.com/jetbrains/status/2061444430884675791

2 tools for scientific figure generation and editing – Crafter and CraftEditor↓ They wrap image gen in a better agentic workflow: 1. Crafter – a multi-agent harness for scientific figures > makes targeted corrections > keeps structured memory of the figure > works with
https://x.com/TheTuringPost/status/2061883014410629400

Could we get a room-temp superconductor from bio? What about chips built from biological materials? For decades people have been predicting a material science revolution from biotech. So far it’s not panned out. But George Church, the godfather of modern synthetic biology,
https://x.com/dwarkesh_sp/status/2060398903347073188

Crafter A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
https://x.com/_akhaliq/status/2061835314599993392

Crafter: A multi-agent harness for scientific figure generation An owl, a fox, and a turtle collaborate to turn text, sketches, or masks into publication-ready charts, posters, and infographics. Everything exports to editable SVGs.
https://x.com/HuggingPapers/status/2061800325959324069

It’s a real honor to receive an honorary doctorate of science from @BrownUniversity . 😍
https://x.com/drfeifei/status/2059874169224147126

Mathematicians and scientists often peak in their 20s. Why? Maybe older scientists become stuck in their ways. Or maybe younger researchers feel free to be more creative. But @jacobkimmel’s hypothesis is that this isn’t because of social factors at all – it’s evolution:
https://x.com/dwarkesh_sp/status/2062203013679481179

We’re excited to introduce Inherent, a lab designed from scratch to build AI agents that discover new knowledge. The coming era of machine-driven scientific inquiry demands a new kind of research institution and a new kind of AI. To achieve our mission, we live within the
https://x.com/inherent_labs/status/2060119235372752924?s=20

.@MukilLoganathan’s Interrupt keynote on Sandboxes.
https://t.co/oddQOs0Q6O In 20 minutes, you’ll learn how to run agent code safely. Isolated from your runtime, with network controls, persistent state, and snapshot/restore when things go wrong.
https://x.com/LangChain/status/2061448130806116827

🤔It is time to rethink how we evaluate agent memory 🌍 As agents become longer horizon and more autonomous, memory is no longer just a module for storing past chats. 🛠️ It determines how agents track changing worlds, learn from past actions, revise outdated information, and
https://x.com/liuchen02938149/status/2061842528698311103

CUAs need to move beyond the prevailing single serial agent paradigm, and start being researched, evaluated, and deployed as multi-agent systems. We hope MACU provides useful insights and a reproducible foundation for future research. 💻 Code:
https://t.co/2tmxNMqHb1 📄 Paper:
https://x.com/kohjingyu/status/2062179533009178897

Grok Build 0.2.7 is now out, with /usage, /login, shared terminals across subagents, and improved image understanding See all updates at
https://x.com/xai/status/2060102590122385460

Grok STT and Grok TTS from @xai are now live on Vapi, the platform for enterprise voice AI. Build on Vapi to create custom voice agents that speak your customers’ language, capture the details that matter in regulated workflows, and sound noticeably more human on every call.
https://x.com/Vapi_AI/status/2062202760590762212

grok-build-0.1 is now available via the xAI API in public beta. This is the same model that powers the Grok Build CLI and excels at agentic coding. Priced at $1/m input and $2/m output, it’s extremely cost effective, intelligent, and fast.
https://x.com/xai/status/2060392249402552457

Why Video Agent models are next — Ethan He, xAI Grok Imagine
https://www.latent.space/p/video-agents

/goal and other fully automated AI agents are cool, but not a great model for the future of work with people. Instead you want your AI to know when to ask you GOOD questions, maybe because it is stuck, maybe because your taste matters, maybe because you would find it interesting.
https://x.com/emollick/status/2061192810422796321

Anthropic says 80% of its new production code is now authored by Claude — how your enterprise can keep up | VentureBeat
https://venturebeat.com/technology/anthropic-says-80-of-its-new-production-code-is-now-authored-by-claude-how-your-enterprise-can-keep-up

Big development – Anthropic is now advocating to build verification mechanisms to enable the option to pause AI development.
https://x.com/a_karvonen/status/2062572851916574730

Correction: Claude Opus 4’s ~3x average speedup dates to May 2025, not May 2024. This evaluation has only existed since September 2024, but we backtested it on earlier models: those from May 2024 showed no speedup whatsoever.
https://x.com/AnthropicAI/status/2062634151556292775

Had Claude Code build a snake game where the snake becomes aware it is in the game and then… stuff happens. Some impressive creative decisions by the AI (& also some very AI ones), I just gave a first prompt and some feedback on the game as it went.
https://x.com/emollick/status/2062039734453416361

Re Opus/sonnet: from what I understood they compare it to sonnet 4.6. only on SWE pro comparable to opus. If I’m correct the quote was „side by side with sonnet 4.6″
https://x.com/kimmonismus/status/2061918020843557110

We’ve updated /fork in Claude Code /fork now runs a background agent with your exact context (system prompt, tools, history, model) and prompt cache. The result gets returned to your session. /branch (the old /fork) still copies the transcript to a new session you drive.
https://x.com/ClaudeDevs/status/2061947411141169494

Claude really can roleplay an economist. I love this little comment Claude made after some robustness checks on the paper it wrote: “”On a 1-10 identification scale, I’d now put the paper at about 4.5 — better than the 3.5 I’d have given before these tests, but well short of
https://x.com/emollick/status/2060168513176658003

Running an AI-native engineering org | Claude
https://claude.com/blog/running-an-ai-native-engineering-org

Anthropic Opus 4.8 is new SOTA on ARC-AGI-3 Score: 1.5%, ~$10K ARC-AGI-3 analysis notes: * Opus 4.8 read the environment an abstraction *above* Opus 4.7, as objects & systems, not pictures * Opus 4.8 succeeded on early levels, but still committed to a wrong sub-goal
https://x.com/arcprize/status/2061512025638121516

Claude Opus 4.8: The System Card | Don’t Worry About the Vase
https://thezvi.wordpress.com/2026/05/29/claude-opus-4-8-the-system-card/

I had early access to Opus 4.8. Was impressed by it. Here is Opus 4.8’s one shot of “”create a visually interesting shader that can run in twigl, make it like an infinite city of neo-gothic towers partially drowned in a stormy ocean with large waves”” (this is all done with math)
https://x.com/emollick/status/2060042738637148470

Opus 4.8 Part 2: Model Welfare | Don’t Worry About the Vase
https://thezvi.wordpress.com/2026/06/01/opus-4-8-part-2-model-welfare/

Opus 4.8 vs MiniMax M3 tested both on default settings with the same prompt > Opus one shotted everything in 7 minutes > M3 needed an extra prompt to fix the “”break block”” feature and took 20+ minutes both got super close, judge both and lemme know which one looks better?
https://x.com/notjazii/status/2061407087293313210

The issue affected how Opus 4.8 requests were handled, causing the model to trigger more parallel tool calls than intended. It was unrelated to dynamic workflows.
https://x.com/ClaudeDevs/status/2061501790131265803

.@GoogleDeepMind’s Gemma 4 – 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes –model gemma4:12b-mlx Claude Code: ollama launch claude –model gemma4:12b-mlx and more 👇👇👇 (Note, this currently works via MLX)
https://x.com/ollama/status/2062250522598572345

Part of the work was rebuilding leaner and faster dependencies: –
https://t.co/QIuyWGhe3V – proxy layer –
https://t.co/c177W6KYqH – filesystem safety –
https://t.co/YW6iCQZX4V – Image engine in WASM –
https://t.co/45NHImVgb7 – Opus in WASM –
https://t.co/B3ELm7gF7x – PDF in WASM
https://x.com/steipete/status/2060133435423789092

Miso One is live: an open-weights voice model built to sound like a real person reading, with actual warmth and pacing where most TTS still goes flat. 8B params, free on GitHub, with one-shot voice cloning from a short sample at 110ms latency. Self-host it and your audio data
https://x.com/kimmonismus/status/2062210845308780639

We just published internal data on how much of Claude’s development is already being done by Claude: – Over 80% of all code merged into our codebase is now written by Claude – It’s been months since many researchers at Anthropic hand-wrote code – The typical Anthropic engineer
https://x.com/alexalbert__/status/2062580571214389510

We’ve added a CLI for Claude Platform to make every API endpoint runnable from your terminal. Call the Messages API, stand up Claude Managed Agents, pipe results straight into your shell. The ant CLI is well understood by coding agents (Claude Code) using the claude-api skill.
https://x.com/ClaudeDevs/status/2061877343078244459

We’ve reset 5-hour and weekly rate limits for all users on Pro and Max plans. We fixed an issue that caused some Claude Code sessions to spawn excessive parallel subagents, burning through usage faster than expected.
https://x.com/ClaudeDevs/status/2061501787769893055

5-Day AI Agents: Intensive Vibe Coding Course With Google | Kaggle
https://www.kaggle.com/competitions/5-day-ai-agents-intensive-vibecoding-course-with-google

getting started with managed agents in the gemini api: an inside look with @OfficialLoganK, @_philschmid, and @alihcevik
https://x.com/GoogleAIStudio/status/2061452967530701090

Together with @OfficialLoganK and @alihcevik we sat down and discussed the launch of Managed Agents in the Gemini API. We wanted to make building AI Agents much simpler. With just a single API call, you can now spin up an agent that reasons, writes and runs code, and manages
https://x.com/_philschmid/status/2061457703210197273

We believe AI can be a dedicated research partner to help discover the next breakthrough. Enter Co-Scientist: our latest Gemini-based multi-agent system that can generate, debate and evolve novel hypotheses for complex scientific problems 🧵
https://x.com/GoogleDeepMind/status/2061857539977842793

2-bit Gemma 4 12B GGUF, only 4.66 GB on disk, managed to cite 15 sites from a single prompt. Try this locally on >6GB RAM via Unsloth Studio. GitHub:
https://x.com/UnslothAI/status/2062470072179044447

Congrats to the @googlegemma team on the Gemma 4 12B launch 🎉 Day-0 support on vLLM is ready to go. It’s an encoder-free unified multimodal model — text, image, audio, and video all project straight into the LLM’s embedding space, no separate vision or audio towers. 256K
https://x.com/vllm_project/status/2062228047324201166

For the past years my research focus was on unifying models and training paradigms across modalities. Today I’m excited that we’re releasing our latest model aligned with this theme: Gemma 4 12B, a dense encoder-free model which processes raw text, image, and audio inputs! 1/
https://x.com/mtschannen/status/2062236357351579915

Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs. Google’s new model, Gemma 4 12B Unified supports image, audio and 256K context. You can run and train the model via Unsloth Studio. GGUF:
https://t.co/8cL321pVDh Guide:
https://x.com/UnslothAI/status/2062207258810053084

Meet Gemma 4 12B! A unified, encoder-free multimodal model designed to bring high-performance intelligence directly to your laptop, and released under an Apache 2.0 license. Bridging the gap between edge efficiency and advanced reasoning. Here is what’s new with Gemma 4 12B: 👇
https://x.com/googlegemma/status/2062202706882883696

Our new unified architecture allows Gemma 4 12B to process multimodal inputs natively. Here’s how ⬇️ Traditional models rely on separate encoders for images and audio. This adds latency and increases memory usage. So we streamlined this: 👁️ Vision: We took a novel approach to
https://x.com/Google/status/2062203532351090824

Today we’re shipping our biggest MLX-VLM release yet: v0.6.0 …and we are raising 💸 This one’s about turning your Apple devices into real local agent machines. From your desk to your pocket. What’s new: ⚡ Speculative decoding everywhere — Gemma 4 EAGLE3 + DFlash, Qwen
https://x.com/Prince_Canuma/status/2061541992790683726

We released Gemma 4 12B yesterday. Here is a visual guide that explains the full architecture. → How encoders typically connect modalities to LLMs → Why Gemma 4 removed the vision and audio encoders → How a single 12B model can handle text, images, and audio without
https://x.com/_philschmid/status/2062546814075609413

We’re launching Gemma 4 12B: Our unified, encoder-free model that brings powerful multimodal intelligence straight to your laptop 🚀 The model bridges the gap between our mobile E4B model and larger 26B MoE models, packaging frontier-class reasoning and native audio into a
https://x.com/googleaidevs/status/2062204432658386950

.@cassidoo is selling me on the GitHub Copilot app! Looks like GitHub has finally found its place in agentic development ❤️ #MSBuild
https://x.com/techgirl1908/status/2061870470237164018

Big paper on AI coding agents using Github & other data The auto-complete tools (Copilot) led to 2.2x more code, local agents like original Claude Code led to 7.4x, & current remote coding agents 17.3x(!) But human bottlenecks in coding means actual releases “”only”” went up 30%
https://x.com/emollick/status/2061659432233161023

Copilot CLI adds an experimental terminal UI with tabs, built-in rubber duck for feedback, prompt scheduling, and hands-free voice input. • Tabs let you view issues, PRs, and gists without leaving the CLI
https://x.com/GHchangelog/status/2061870684876272123

GitHub unveils the GitHub Copilot desktop app in preview, which introduces a new feature called canvases for bidirectional work between users and agents (@mariorod1 / The GitHub Blog) (Visit Techmeme dot com for the link and full context!)
https://x.com/Techmeme/status/2061875738694062419

I’ve been excited about many of the things we’ve shipped over the last year, across GitHub Copilot CLI, SDK, Cloud Agents, and more. But I am even more excited about what we’re delivering now with the GitHub Copilot App. This is the app that I use all day every day now. It is,
https://x.com/lukehoban/status/2061905434039246939

Introducing the GitHub Copilot app, the desktop home for agent-native software development on GitHub
https://x.com/pierceboggan/status/2061868635241828688

It’s time to move from renting intelligence to truly controlling your AI. Microsoft Frontier Tuning lets you take our models and make them uniquely your own, turning them from capable generalists to completely custom partners. It starts with reinforcement learning environments
https://x.com/mustafasuleyman/status/2062275417378041957

Just now from @satyanadella at Microsoft Build keynote : over 11,000 models are available in Microsoft Foundry Super proud to contribute to it – @huggingface provides 10,928 of those! Stay tuned for exciting announcements with new ways to Build with open models in Microsoft
https://x.com/jeffboudier/status/2061868927207244277

Microsoft also released MAI-Code-1-Flash It’s a 137B parameter MoE coding model with a context length of 256K tokens trained on over 10T tokens. At launch it will be available only in GitHub Copilot in Visual Studio Code. It seems to be stronger and more efficient than Claude
https://x.com/scaling01/status/2061891478176112794

Seven new models launching at Build: let’s go! Reasoning. Code. Image. Transcribe. Voice. Built from scratch on a clean data lineage, designed for efficiency, working seamlessly as a family of models Thread 🧵 #MSBuild
https://x.com/MicrosoftAI/status/2061887500541366489

Super interesting announcement: Project Solara. #MSBuild Microsoft’s thesis is that if agents become the primary way we interact with software, they should not be confined to laptops and smartphones. Project Solara is a platform for building agent-first devices. The company
https://x.com/TheTuringPost/status/2061865165734506683

The GitHub Copilot app is now available for more people, good time to try it I’m learning new things every day, for example the way it renders the plan with updates is absolutly beutiful!
https://x.com/OrenMe/status/2061873010664001605

Today’s news all comes down to this: we’re putting our relentless hill-climbing machine at your service. From launching top tier models to helping you make them your own, our commitment as a platform company is to keep you at the absolute frontier. For all the details:
https://x.com/mustafasuleyman/status/2061934667096596657

Microsoft leaked the training FLOPS for Claude Mythos based on their slide Claude Mythos used: 6.1*10^27 FLOPs (with 95% CI at 5.3*10^27 and 7.1*10^27, assuming 1 px measurement error)
https://x.com/scaling01/status/2061897540161728791

The 6.1e27 FLOP figure for Claude Mythos is not realistic. The Microsoft intern just threw some darts. Instead, let me throw some darts. This is my realistic estimate for Claude Mythos compute, total params, active params and training tokens! Mythos was likely a training run
https://x.com/scaling01/status/2061989029025853757

first ship on the new team: new docs for agents on Cloudflare! 🚀 really helps to visualize the breadth that @Cloudflare agent platform has to offer: Communication channels: – chat/slack/webhook agents – voice agents with voice – email agents with email service Core agent: –
https://x.com/thomasgauvin/status/2062512156076048447

🚨 MAI-Image-2.5 is now live on fal! 📸 Photorealistic images with natural lighting and accurate skin tones Refined text rendering for branding, packaging, and commercial design Text-to-image and image editing with precise, design-ready control
https://x.com/fal/status/2061920052664820199

Mai-1 thinking: Mid size model, 45b active parameter, MoE, side by side with sonnet 4.6 0 distillation „Microsoft’s first reasoning model”
https://x.com/kimmonismus/status/2061877528781025381

microsoft MAI tech report is a gold mine, one of the most transparent for a model at this scale. this model uses zero synthetic data or distillation from previous models. this means reasoning, agentic behavior, tool use are all learned fully during post-training with no cold
https://x.com/eliebakouch/status/2061965825037254947

This SkillOpt paper from Microsoft is a must-read! (bookmark it) I was a bit skeptical of the results reported in the paper when I shared it a few days ago. However, I managed to integrate it into my agent orchestrator and ran a few experiments. The results are mindblowing.
https://x.com/omarsar0/status/2062204469538881988

Three new @MicrosoftAI models now live on OpenRouter! Launching together: MAI-Image-2.5, MAI-Transcribe-1.5, and MAI-Voice-2. More on each below 🧵
https://x.com/OpenRouter/status/2061894672847671724

As token budgets take on a larger part of operating expenses over time, model routing is the inevitable conclusion. This is also one of the biggest areas of differentiation for the applied AI layer over time. By understanding the different work patterns in your domain, and
https://x.com/levie/status/2061974298760495132

Excited to see the use of GEPA-optimized LLM judges for data filtering in MAI-Thinking-1 model’s pre-training pipeline!
https://x.com/LakshyAAAgrawal/status/2062013650639241403

Give MAI-Code-1-Flash a try and let us know what you think!
https://x.com/pierceboggan/status/2062220583786709163

MAI-Thinking-1 is out! Excited to share what we are building and how climbing from scratch (no distillation) actually works: simple recipes, rigorous science, self-distillation, patience, and great infra. Check out our tech report has the full story of our RL climbs.
https://x.com/HannaHajishirzi/status/2061901432627044430

Super detailed tech report for MAI-Thinking-1, with a ton of info on all stages of the pipeline. I’m surprised so much of this info is released 🙂 Super long thread on my notes:
https://x.com/nrehiew_/status/2062013300196700395

this was an insanely good read, i think this is the most detailed report i’ve read at this scale in some aspects. i really hope MAI continues releasing those tech reports, thanks a lot to the team for this gift 🥹
https://x.com/eliebakouch/status/2062004670017486912

Today, Baseten and @MicrosoftAI are excited to announce that MAI-Thinking-1 is coming to Baseten. MAI-Thinking-1 is a model you can fine-tune without giving your data to the lab. Key characteristics include: → Clean data lineage, with zero distillation from third-party models
https://x.com/baseten/status/2061878701823066431

We’re excited to work with @Baseten to make MAI-Thinking-1 available to developers and enterprises.
https://x.com/MicrosoftAI/status/2061923309344756043

WOW microsoft new “”MAI Thinking 1″” model comes with a 109 page tech report that looks REALLY detailed, this is amazing
https://x.com/eliebakouch/status/2061877335960281459

Microsoft has released MAI-Transcribe-1.5: an exceptionally fast speech transcription model at a speed factor of ~276x, while still achieving 2.4% on AA-WER (#3), leading the accuracy-speed Pareto frontier MAI-Transcribe-1.5 is Microsoft AI (MAI)’s latest speech transcription
https://x.com/ArtificialAnlys/status/2061878491860324402

Microsoft introduces MAI-Thinking-1 It’s a 1T@35B parameter model pre-trained on 30T tokens with a maximum context length of 256k tokens using 8192 GB200 GPUs. Based on benchmarks it seems to be around GLM-5 level. Microsoft also released a comprehensive 109 pages tech-report:
https://x.com/scaling01/status/2061889624847343825

microsoft used gepa / dspy to tune the LLM judge prompt for quality scoring. @lateinteraction stays winning. from the mai-thinking-1 report
https://x.com/bj2rn/status/2061941109828301241

Today we’re announcing MAI-Thinking-1 with Microsoft and it will be available on Baseten soon. Microsoft built something genuinely different here: a commercial-grade thinking model trained on clean data with no distillation from third-party models and designed to be fine-tuned
https://x.com/tuhinone/status/2061879239817969756

big congrats to the microsoft AI team on MAI-Thinking-1! this is the kind of thoughtful post-training the field needs more of – focused on what actually matters to users excited to see a new frontier model in the race 😎
https://x.com/echen/status/2061907282607100075

Introducing MAI-Code-1-Flash A new coding model from Microsoft for fast, efficient assistance in everyday workflows Rolling out to @code developers in model picker and Auto now!
https://x.com/pierceboggan/status/2061877165810131297

It is difficult to know how good MAI-Thinking-1 is from the scores alone (like weirdly low GPQA & Terminal Bench 2.0) But Microsoft makes it really hard to try its models upon release (a general issue with many Microsoft AI products), so I dunno. Stats below Meta Spark, though.
https://x.com/emollick/status/2061907785768489127

MAI-Image-2.5 has officially released from @MicrosoftAI landing at #2 in the Image Edit Arena (Single-Image-Edit) with a score of 1401 and advances the Pareto frontier! This puts the model +10 pts over Nano Banana 2, Grok Imagine Image Quality and ChatGPT-Image-Latest-High
https://x.com/arena/status/2061887242579382660

MAI-Image-2.5 ranks #2 in the Image Edit Arena and advances the Pareto frontier. That means: at its price tier, no model scores higher on Arena. Congrats again to @MicrosoftAI on this release!
https://x.com/arena/status/2061894541888962712

Microsoft AI has announced their very own reasoning model, MAI-Thinking-1, they have a detailed tech report too! I really appreciate they’ve reported health evals: HealthBench Professional and MedXpertQA. These are both very solid benchmark tasks that I recommend people use.
https://x.com/iScienceLuvr/status/2061926066453962952

Awesome to have on you stage @steipete ! What an exciting time
https://x.com/mustafasuleyman/status/2061882830561628611

Such a privilege to work with Microsoft to bring claws to enterprises!
https://x.com/steipete/status/2061874084649025728

Major overhaul of the Hermes Dashboard It should now surface a complete management plane. Goal is to reduce or eliminate any needs to run a CLI command directly in your terminal. Let me know if we’re missing anything!
https://x.com/Teknium/status/2062315666439655499

You can use Hermes Desktop with Ollama using local or cloud models. Get started 👇👇👇
https://x.com/ollama/status/2062011585355551231

🚀 We’re excited to partner with @NVIDIARTXSpark pushing local AI agents forward on DGX Spark + RTX! This is exactly the direction explored in @Inferact’s hands-on #vLLM + #DGXSpark blog–serving large NVFP4 models locally on NVIDIA DGX Spark. vLLM is an ideal fit, bringing
https://x.com/vllm_project/status/2061530659160838549

We are proud to continue our collaboration with @nvidia with support for thier NVIDIA RTX Spark Laptop. Strengthening our support to support OpenShell and @Microsoft Security Primitives. Building ontop of our earlier work with NemoClaw and our existing fully-native Windows
https://x.com/openclaw/status/2061331260279054801?s=20

.@NVIDIA’s Cosmos 3 launched today… and guess who had early access? Agile Robots SE! They’ve been running it across their full portfolio: Thor single- and dual-arm, FR3 Duo. Focus? Simulation. Using Cosmos 3 as a neural simulator; a learned world model that generates
https://x.com/IlirAliu_/status/2061512207738012093

1/ NVIDIA just open-sourced Cosmos 3 at GTC Taipei! It’s the first fully open “”omnimodel”” for physical AI – one model that understands the real world, predicts what happens next, and generates the actions a robot should take. Weights, code, datasets. All open. And this is
https://x.com/kimmonismus/status/2061432501223162241

Breaking news: Cosmos 3 is here. They are attempting to do something completely new 🤯 Why is Physical AI much harder than building a chatbot? Understanding the world is not enough, robots need to predict it and act inside it. That’s the idea behind NVIDIA Cosmos 3: →
https://x.com/TheTuringPost/status/2061308942186414136

In case you missed this: NVIDIA shipped a text-to-image open weights model that looks seriously competitive 👀 (as part of its Cosmos 3 release)
https://x.com/victormustar/status/2061354267546427595?s=20

NVIDIA’s Cosmos 3 is what we’ve never seen before ‒ an Omnimodal World Model. It’s closing the loop for physical AI. All stack in one system: world and multimodal understanding, future generation, reasoning and action This is the next step in Jensen Huang’s AI progression:
https://x.com/TheTuringPost/status/2061474876083274238

NVIDIA’s Cosmos 3 lands at #1 among open weights models in both Text to Image and Image to Video on the Artificial Analysis Leaderboards! Cosmos 3 is a family of omnimodal world models for Physical AI from @nvidia, unifying language, image, video, audio and action in a single
https://x.com/ArtificialAnlys/status/2061494719998546206

NVIDIA’s Cosmos 3 lands at #1 among open weights models in both Text to Image and Image to Video on the Artificial Analysis Leaderboards! Cosmos 3 is a family of omnimodal world models for Physical AI from @nvidia, unifying language, image, video, audio and action in a single
https://x.com/ArtificialAnlys/status/2061494719998546206?s=20

is Hermes Agent ready for enterprises? NVIDIA built OpenShell, a runtime that wraps AI agents in the security IT teams need before they let anything touch sensitive systems. it plugs directly into Microsoft’s enterprise security stack. Hermes Agent now runs inside it. what this
https://x.com/shannholmberg/status/2061368566256189656

We have been working closely with @nvidia to ensure Hermes Agent works smoothly on their new @NVIDIARTXSpark superchip and integrates with the new OpenShell runtime, which connects Hermes to @Microsoft’s security primitives. Watch our feature in the big announcement at Computex:
https://x.com/NousResearch/status/2061323987804713083?s=20

@openclaw @NousResearch @LangChain As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, and post-training recipes. Available now on @huggingface →
https://x.com/NVIDIAAI/status/2062521383582646537

🚀 Day-0 support for NVIDIA Nemotron 3 Ultra on vLLM! Ready to be served with the latest vLLM stable release, the new open frontier reasoning model is built for long-running autonomous agents: 🧠 550B total / 55B active — Hybrid Transformer-Mamba MoE 📚 Up to 1M token context
https://x.com/vllm_project/status/2062574262163280172

420.2 tok/s on a 550B model. ⚡️ Nemotron-3-Ultra-550B-A55B reaches 420.2 tok/s powered by BLACKBOX AI Inference Engine. Blackbox now delivers the fastest inference in the industry, outperforming every other provider, including on smaller-parameter models. Check our blog in
https://x.com/blackboxai/status/2062546216949588001

Are you tired of waiting 17 minutes for an AI agent to finish a code change? As an agent’s context grows, standard transformer attention can turn long runs into a bottleneck. @NVIDIAAI Nemotron 3 Ultra addresses this with a hybrid architecture that replaces several
https://x.com/baseten/status/2062609272815685759

Big day for American open models… Nemotron 3 Ultra is now the strongest US open-weight model tested, while apparently serving 300+ tok/s 🤯 Comparable large DeepSeek/Kimi models are usually 50-100 tok/s btw
https://x.com/caspar_br/status/2061505720907182280

Introducing NVIDIA Nemotron 3 Ultra. A frontier smart open model built for long-running agents that need to plan, reason, use tools and keep working across complex coding, research and enterprise workflows. Up to 5x faster inference and up to 30% lower cost for agentic tasks.
https://x.com/nvidia/status/2062522316672667770

Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-latency multilingual speech recognition. AI natives can now build coding agents, deep research agents, and real-time voice systems on the AI
https://x.com/togethercompute/status/2062520009893576974

nemotron 3 is significantly less sparse than other models (~10% active vs ~3% for kimi K2/deepseek v4)
https://x.com/eliebakouch/status/2061607195268038777

Nemotron 3 Ultra (550B-A55B) is here – our strongest open-weight model and full training recipe to date. Heavy emphasis on real-world inference efficiency for long-context agentic workloads. Everything is open 🤗: base, post-trained, reward checkpoints, NVFP4 quantized
https://x.com/PavloMolchanov/status/2062538679470657727

Nemotron 3 Ultra is now the best open weight model on
https://t.co/EJXiSfWv2O 💚
https://x.com/ctnzr/status/2061483152741175757

Nemotron 3 Ultra was launched today, including a focus on low latency agentic performance. We tested it against peers under restricted turn-usage limits on Terminal-Bench v2.1 – @NVIDIA Nemotron 3 Ultra completes tasks at a much faster pace than peers due to its high inference
https://x.com/ArtificialAnlys/status/2062598349757567359

NVIDIA has just released Nemotron 3 Ultra, the new most intelligent US open weights model, with leading speed for its intelligence Nemotron 3 Ultra scores 47.7 on the Artificial Analysis Intelligence Index, well ahead of the next strongest US open weights models, Gemma 4 31B
https://x.com/ArtificialAnlys/status/2062527871529439438

NVIDIA just announced the release of Nemotron 3 Ultra in Jensen Huang’s Computex keynote: at 550B parameters (55B active), this is the largest Nemotron 3 model to date, and it is the most intelligent US open weights model We partnered with @nvidia to evaluate this model for
https://x.com/ArtificialAnlys/status/2061304911565144230?s=20

NVIDIA just released Nemotron 3 Ultra, a 550b-parameter agentic coding model with a 1m context window. It was built for token efficiency, and is up to 5x faster and 30% cheaper than other similar models. It’s the largest US open-weights model release ever. Free in Cline now!
https://x.com/cline/status/2062620668085297214

NVIDIA Nemotron 3 Ultra is here We have Day‑0 support for Nemotron 3 Ultra in prime-rl and Lab. Specialize Nemotron 3 Ultra for your use case.
https://x.com/PrimeIntellect/status/2062622550300275088

NVIDIA Nemotron 3 Ultra is on Fireworks, day zero. Nemotron Ultra is an open model for frontier reasoning and orchestration in long-running autonomous agents. Think use cases like coding agents, deep research, and complex enterprise workflows. Read on:
https://x.com/FireworksAI_HQ/status/2062568688201646321

NVIDIA’s Nemotron 3 Ultra is available on Ollama’s cloud! Try it 👇 Claude Code: ollama launch claude –model nemotron-3-ultra:cloud Hermes Agent: ollama launch hermes –model nemotron-3-ultra:cloud OpenClaw: ollama launch openclaw –model nemotron-3-ultra:cloud
https://x.com/ollama/status/2062591290743853291

Oh wow, they pre-trained Nemotron 3 Ultra in NVFP4 big update for estimating future model sizes and flops, especially for OpenAI models
https://x.com/scaling01/status/2062540298933219832

The @nvidia Nemotron 3 Ultra is live on CW Serverless Inference 🚀 Open, frontier-reasoning, built for long-running agents — 550B params (55B active), hybrid Transformer-Mamba MoE, up to 1M context. For orchestration, coding agents & deep research. No infra to manage.
https://x.com/wandb/status/2062577626242580896

We are excited to join Nvidia’s Nemotron Coalition of leading AI labs working together to advance open frontier foundation models. To celebrate we have partnered with @nvidia and @nebiustf to provide 2 free weeks of the new Nemotron 3 Ultra model on the Nous Portal!
https://x.com/NousResearch/status/2062554136625766409

With today’s launch of Nemotron 3 Ultra, @nvidia continues to expand its investment in open-source AI. Their flagship frontier-reasoning model, built for long-running autonomous agents, is available Day 0 on Modal. – 550B with 55B active parameters – Hybrid Transformer-Mamba MoE
https://x.com/modal/status/2062528720104227149

Nemotron 3.5 ASR is built for streaming multilingual speech recognition and voice agents. One 0.6B checkpoint. 40 language-locales. Sub-100ms latency. Cache-aware FastConformer carries context forward between chunks instead of repeatedly reprocessing overlapping audio. Try it:
https://x.com/togethercompute/status/2062520605102993436

Open speech models for real-time voice agents. @NVIDIAAI Nemotron 3.5 ASR is now available on fal. An open streaming speech recognition model supporting 40 language-locale combinations with ultra-low latency, native punctuation, and capitalization. Built for voice agents,
https://x.com/fal/status/2062521027020611933

Second big release from us today: Nemotron-3.5-ASR-Streaming! 🌎40 languages ⚡️80ms – 1s controllable latency 🔥240 – 2400 concurrent streams on 1xH100 🧱FastConformer Cache-Aware RNN-T architecture
https://x.com/PiotrZelasko/status/2062538923776290909

Build and launch apps to your team, using Codex:
https://x.com/gdb/status/2061988413105156128

Codex now has more than 5M weekly active users. But the bigger story is what people are using it for: not just writing code, but getting more work done across research, analysis, content, and operations. Our new report on how Codex is becoming a productivity tool for knowledge
https://x.com/OpenAINewsroom/status/2061834718224777579

Devin Desktop is our first product launch that’s fully agent-neutral. You can run your own custom background agents directly from the desktop app, Devin, or even Claude Code / Codex. Part of Cognition being the Independent Agent Lab is working well with all of the agents –
https://x.com/russelljkaplan/status/2061920322325205007

Every bug in @ChatGPTapp is getting fixed With the help of codex (and the rest of the lovely team and their codexes) along with a 7pm iced americano there will be zero bugs This is a formal request for tiny nits, error states, broken ui, etc The tinier the better!
https://x.com/JustinBleuel/status/2059820836362588426

Haven’t seen codex writing ad-hoc codemods before, but it just did for a bigger TypeScript migration. Impressed.
https://x.com/steipete/status/2061115471760441692

I do this with codex all the time. Ask it to review code for bugs and it will tell you all good, tell it there is a bug and it will LOOP AND LOOP and will find issues.
https://x.com/steipete/status/2060672154727825718

I told codex to use
https://t.co/oHS8ombQcW whenever I’m distracted and it needs my help to be unblocked, and ever once it a while I hear it talking to me, and it’s the coolest thing ever. (e.g. for releases, that needs npm and is 1Password-gated)
https://x.com/steipete/status/2061574752574283858

If you ever get tired of managing your Codex threads, just let Codex manage itself! Codex can now create threads, search them, organize them, pin the important ones, and spin up worktrees for parallel tasks.
https://x.com/guinnesschen/status/2060464235868836235

New in the LangSmith Sandboxes GA Release: Sandbox CLI ✅Build snapshots from Dockerfiles ✅Manage sandboxes ✅Open interactive consoles ✅Tunnel raw TCP ✅Use standard tools (ssh, scp, rsync, sftp) against a sandbox like any Linux box
https://x.com/LangChain/status/2062512156688466083

Nobody talks about how pleasant building with Codex feels
https://x.com/CarolMonroe/status/2060399106326204681

Now generally available, @OpenAI GPT-5.5, GPT-5.4, and Codex on Amazon Bedrock. Deploy frontier AI models with automatic scaling through Bedrock’s next-gen inference engine. With these offerings on Bedrock, customers can: ◽ Build autonomous agents that handle multi-step
https://x.com/awscloud/status/2061564484523524302

watching codex control my browser to do things it can’t do in the harness is a holy shit experience
https://x.com/Nickprince/status/2060769417919868973

We just launched Sites into Codex! Software creation was always about more than writing code. Sites in Codex fundamentally gives the power of end-to-end software creation to every user, no matter their technical fluency. These Sites are fully deployed to a URL, private to
https://x.com/TheRohanVarma/status/2061872164442403139

We just released the Codex Python SDK 🔥 You can now embed Codex directly into your Python apps and workflows! > Start threads > Run turns > Stream progress > Resume sessions > Pass images > Control sandbox access All whilst reusing your existing Codex auth. pip install
https://x.com/reach_vb/status/2061569472792572163

We’re making Codex more useful for your work by expanding plugins beyond individual tools. These plugins turn Codex into a specialist for a specific role with a single install, no coding required. Codex can access 62 popular apps and 110 skills for work across sales, data
https://x.com/OpenAI/status/2061887650391625870

Windows users, this one’s for you. Computer use now works on Windows, so Codex can take action on your Windows computer. And with Windows support for Codex in the ChatGPT mobile app, you can start, review, and steer tasks on the go while work continues on your Windows machine.
https://x.com/OpenAI/status/2060428604727771421

With GPT 5.5, /goal, autoreview and crabbox my prompts moved from ~30-60min to often 4-10h tasks and my confidence that it’s ready is much much higher. Yielding agents is a skill.
https://x.com/steipete/status/2060678430031597696

OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new way to build on Amazon Bedrock with OpenAI through the security, compliance, and governance workflows they already use. This is also the beginning of a broader expansion of OpenAI
https://x.com/OpenAI/status/2061564502160892138

OpenAI models and Codex, now in your AWS workflows. Build AI apps and software engineering workflows with OpenAI on Amazon Bedrock, using the AWS environments and controls your team already trusts.
https://x.com/OpenAIDevs/status/2061564710173224985

PSA: Codex now works with Amazon Bedrock! To configure, set: model_provider = “”amazon-bedrock”” and Codex can run local CLI, desktop, and IDE workflows against OpenAI models on Bedrock, using AWS-native auth/IAM. Both GPT-5.4 and GPT-5.5 are supported. Big unlock for teams
https://x.com/reach_vb/status/2061572961451094191

We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-5.5’s agentic coding and tool use together with stronger intelligence for drug discovery, analysis, design, and experimental workflows.
https://x.com/OpenAI/status/2062281977122996256

We’re taking steps to accelerate defensive progress in biology: – Launching Rosalind Biodefense to help trusted builders develop new biodefense and pandemic preparedness capabilities. – Expanding trusted access to GPT-Rosalind for select U.S. government and allied partners
https://x.com/OpenAI/status/2060376598642405492

“clanker” is not a slur. “vibe coding” is.
https://x.com/steipete/status/2060371944168358250

build the thing that builds the thing.
https://x.com/steipete/status/2060071399226495452

Couldn’t be more excited to have Vince on board. 🦞 Very few people understand the new ways, how software is built. He gets it.
https://x.com/steipete/status/2060306947035832628

Every claw release spins up hundreds of CI machines to QA test and eventually creates a ledger.
https://x.com/steipete/status/2059996262477222153

Finally got my visa sorted out and moving to San Francisco, just in time for MS Build and OpenClaw’s after hours!
https://x.com/steipete/status/2061031509088231640

Hit GitHub’s rate limit one too many times, so I built octopool: a Cloudflare Worker that pools your team’s PATs + GitHub App installations behind a shared read cache. Self-host on Cloudflare. Drop-in gh shim.
https://x.com/steipete/status/2059989719870558309

I smell a takedown in 3…2…1
https://x.com/steipete/status/2060294413377519808

It’s been great working with Omar to get observability and verifiable workspaces into OpenClaw.
https://x.com/steipete/status/2061877813053907083

No LLMs for finding bugs even?
https://x.com/steipete/status/2060358460831682895

We have over 1300 people on the waitlist for today’s OpenClaw event – will be livestreamed on Twitch and Discord tho!
https://x.com/steipete/status/2062307384018829768

We never had more npm downloads than this week on @openclaw – comined with Docker, GitHub, company-internal deployments and the numerous forks, real number is more in the 10-20 million downloads/week.
https://x.com/steipete/status/2062276065448669627

langsmith! ✅ Sandbox:
https://t.co/vaChlwHbrm ✅ Gateway:
https://t.co/UqWDeBFS2H ✅ Observability:
https://x.com/hwchase17/status/2062144718427857256

.@perplexity_ai made search code-driven and agent-controlled with a new architecture – Search as Code (SaC). The agent stops being stuck with one big monolithic search API, and generates Python code that composes search primitives for the task. SaC architecture includes 3 main
https://x.com/TheTuringPost/status/2061606295128641643

Congrats to the @Alibaba_Qwen team, the new Qwen-3.7 Plus is a best in class multimodal model, performing near SOTA as both a coding and general purpose agent across benchmarks. Free to try in Cline now! `npm i -g cline`
https://x.com/cline/status/2061580233778790439

[2606.03746] Qwen-Image-Flash: Beyond Objective Design
https://arxiv.org/abs/2606.03746

20 advanced RAG types to know in 2026 ▪️ Mindscape-Aware RAG (MiA-RAG) ▪️ Multi-step RAG with Hypergraph-based Memory (HGMem) ▪️ MegaRAG ▪️ Disco-RAG (discourse-aware) ▪️ Agentic RAG ▪️ A-RAG (with Hierarchical Retrieval Interfaces) ▪️ Predictive Prefetching RAG ▪️ SURE-RAG ▪️
https://x.com/TheTuringPost/status/2061085042391339199

Composer 2.5 is now available inside Grok Build. Composer 2.5 is a fast, highly intelligent model that excels on long-running tasks and following complex instructions.
https://x.com/xai/status/2061510464325206163

Grok @Imagine 1.5 Preview is here Try it today in the API:
https://x.com/grok/status/2062225080843747351?s=20

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading