Image created with Gemini. Image prompt: Overhead flat illustration of a turquoise swimming pool with a small humanoid pool-skimmer robot moving along a dashed white path that connects a floating tray of yellow lemons, a coiled terracotta garden hose, and a stack of pink folded towels at the pool’s tiled edge, surrounded by hand-drawn white squiggle ripples. Flat saturated acrylic colors, hard-edged shapes, no shadows, tablet-drawn line quality, cheerful high-noon light.

2026 Hype Cycle for Agentic AI | Gartner
https://www.gartner.com/en/articles/hype-cycle-for-agentic-ai

Anthropic has released Claude Fable 5, the first publicly available Mythos-class model that ranks #1 in our agentic real-world knowledge work benchmark GDPval-AA Claude Fable 5 shares the same underlying model as Claude Mythos 5, with added security guardrails for potentially
https://x.com/ArtificialAnlys/status/2064414308289937869

Apple WWDC 2026 June 8: Introducing Siri AI and more – YouTube
https://www.youtube.com/watch?v=hF8swzNR1-o

Exciting news: Claude Fable 5 ranks #1 on the new Agent Arena leaderboard! Fable 5 leads by the widest margin ever over Opus-4.8 and GPT-5.5 on two key signals: confirmed task success rate and praise vs. complaint, despite weaker steerability. If Fable can do something, it will
https://x.com/arena/status/2064807170714358193

Just used claude fable (aka mythos) to create this city block simulator complete with multi-agent traffic, live detection boxes + tracks, and day to night cycle. And it just one shotted it. This is gonna be fun — the gap between idea and execution just keeps collapsing.
https://x.com/bilawalsidhu/status/2064524211914223867

Meet Kimi Work – a local AI agent on your desktop that does the work for you. 🔹Native agent swarm: Up to 300 AI agents running in parallel on your local machine. 🔹Browser use: Paired with WebBridge extension, your agent will navigate websites in your browser: search, scroll,
https://x.com/Kimi_Moonshot/status/2063990409903112344

Say hi to “”Siri AI””–Apple announces new, more “”conversational”” voice assistant – Ars Technica
https://arstechnica.com/apple/2026/06/say-hi-to-siri-ai-apple-announces-new-more-conversational-voice-assistant/

This chart from Anthropic is useful, since Agent Teams and Workflows are both very new and very powerful (and token hungry). On the other hand, maybe it doesn’t matter as a lot of the decisions about which approach to use is from the AI itself & it often uses them in combination
https://x.com/emollick/status/2063073955968123062

Visa brings payments to ChatGPT as AI agents start buying for you | AP News
https://apnews.com/article/visa-chatgpt-openai-shopping-mastercard-d769dec86344cb4977c98789e8ec492f

With multi-agent orchestration in Claude Managed Agents, you can use Fable to delegate work to dedicated agents using smaller models.
https://x.com/ClaudeDevs/status/2064394928948703406

Anthropic has a coding MOAT
https://x.com/scaling01/status/2064399642603802676

Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude | WIRED
https://www.wired.com/story/anthropic-responds-to-backlash-on-claudes-secret-sabotage-on-ai-research/

Anthropic’s new Claude Fable 5 is #1 on Terminal-Bench 2.1 at 88.0%, beating GPT 5.5 by 4.6%. Under the hood it’s their Mythos model with safeguards that route dangerous requests to Opus 4.8, which Anthropic claims triggers in under 5% of sessions. Available in Cline now!
https://x.com/cline/status/2064427461212045546

Anthropic’s new Fable 5 safeguards are fascinating. When the model is used for frontier LLM development, it apparently does not simply refuse or warn the user. Instead, it quietly limits its own effectiveness through techniques like prompt modification, steering vectors, and
https://x.com/kimmonismus/status/2064417460715962479

Anthropic/OpenAI may be spending more than $1000 for every $100 you pay them – R&A IT Strategy & Architecture
https://ea.rna.nl/2026/06/07/anthropic-openai-may-be-spending-more-than-1000-for-every-100-you-pay-them/

As of May 2026, more than 80% of the code we merge into Anthropic’s codebase was authored by Claude.”” Matches independent measures. There really is no sign this is slowing down (which doesn’t mean there aren’t organizational challenges to absorbing this much productivity gain)
https://x.com/emollick/status/2062580212768936344

Claude Fable 5 / Mythos 5 wins everywhere. I thought Fable 5 was just a nerfed Mythos Preview, but it’s literally better. SWE-Bench Pro: Fable 5: 80.3%, GPT-5.5: 58.6%. And the price is only 2x Opus 4.8: $10/input MTok, $50/output MTok. I don’t think GPT 5.6 can beat this…
https://x.com/Yuchenj_UW/status/2064396097075003739

Claude Fable 5 and Claude Mythos 5 \ Anthropic
https://www.anthropic.com/news/claude-fable-5-mythos-5

Claude Fable 5 beats Pokémon FireRed only using vision – YouTube
https://www.youtube.com/watch?v=Ty_50J84fMY

claude fable 5 has solved CAD I asked it to make a model of a V8 engine It came back to me with a fully working model in under 10 minutes
https://x.com/aaronli/status/2064876123109089742

Claude Fable 5 is our first generally available Mythos-class model. It ships with new safety classifiers that may flag certain prompts in dual-use domains like cyber and bio. We’ve added fallbacks: a refused request retries on Claude Opus 4.8 instead of dead-ending.
https://x.com/ClaudeDevs/status/2064428347678220691

Claude Fable 5 sets a fluid simulation to Beethoven – YouTube
https://www.youtube.com/watch?v=xmP7bhigCWE

Claude Mythos 5 scores 161.29 on Anthropic’s ECI
https://x.com/scaling01/status/2064392088003756431

Claude Mythos 5 speeds up kernels by 430x – a 300x speedup equals 40 human expert hours of work
https://x.com/scaling01/status/2064392386520780945

Claude Mythos is next level. h/t @Lentils80 Look at this MacOS output. One shotted.
https://x.com/kimmonismus/status/2062843119864021404

Claude mythos will be on a completely different level. These outputs are insane
https://x.com/kimmonismus/status/2062805570982203820

Fable 5 is the biggest step up I’ve felt in our models since Opus 4.5 back in November. After 4.5 came out I uninstalled my IDE when I realized that I’d been doing 100% of my coding in a terminal for a few weeks. With Fable, it’s felt like Claude has stepped up from being a
https://x.com/bcherny/status/2064431111154053187

Head of Claude Code said the line that makes prompting feel outdated “”I don’t prompt Claude anymore. I write loops, the loops do the work.”” Most devs are still trying to write the perfect prompt Boris Cherny already moved to the next layer Not prompt Not answer Not copy paste
https://x.com/0xwhrrari/status/2064804504608887040

I believe what Anthropic is doing, gating the ability to do certain harmless things like LLM research, and with incredibly sensitive filters that even medical questions are often blocked, is *deeply* wrong. They got open research, the Transformer, GPT2, …
https://x.com/antirez/status/2064766431531532588

I don’t really want to have to go to bat against Anthropic, but they’ve just been unnecessarily antagonistic to all of China, then not so subtly to open weight models, and now more broadly open AI research. What’s next on the list?
https://x.com/natolambert/status/2064412173527556298

I got a good nights sleep and I’m still just as angry about Anthropic’s choices. I enjoy working in AI so much and to have my access to the cutting edge models for my work rugpulled in an under the table fashion is appalling. I expected to be restricted eventually, but not
https://x.com/natolambert/status/2064699044145095104

I think it is really worth reading this piece on RSI at Anthropic. There is a bit of navel-gazing, some marketing, and a lot of very sincere beliefs about what Anthropic thinks is likely in the near future of AI that you probably want to be aware of.
https://x.com/emollick/status/2062582362194460698

If Claude Fable stops helping you, you’ll never know — Jonathon Ready
https://jonready.com/blog/posts/claude-fable5-is-allowed-to-sabotage-your-app-if-youre-a-competitor.html

In good faith and with no judgment (mistakes happen), I truly hope that Anthropic will hear the feedback and change course on this. Anthropic is a company that has been raising awareness about AI manipulation which is a very important topic! You don’t want to go down as the
https://x.com/ClementDelangue/status/2064673792303955985

Introducing Claude Fable 5 – YouTube
https://www.youtube.com/watch?v=Y9Wz2PV404E

Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available.
https://x.com/claudeai/status/2064394146916229443

It’s state of the art on nearly every benchmark we tested and the lead grows the longer the task. Made safe for general release: cyber & bio requests fall back transparently to Opus 4.8 and 95%+ of sessions never see one. $10/$50 on the API, in paid Claude plans today.
https://x.com/mikeyk/status/2064392996288901392

Makes me wonder how long this has already been going on without users being notified. I had been feeling like codex was running circles around Claude code for months now. Now I wonder if Claude code was just self nerfed. Regardless of the intentions behind this, this is a bad
https://x.com/code_star/status/2064464447662707180

Making Claude a chemist \ Anthropic
https://www.anthropic.com/research/making-claude-a-chemist

Microsoft AI head calls out Anthropic for acting like Claude is conscious | The Verge
https://www.theverge.com/tech/947197/microsoft-ai-mustafa-suleyman-anthropic-claude-conscious

My last observation re: Anthropic’s secret sabotage safety policy, is that it undermines actually good safety policy. How? 1. First, it is very plausible to describe this as anti-competitive behavior (even if you are maximally sympathetic to Anthropic here you must admit this),
https://x.com/deanwball/status/2064665679307985244

Policy on the AI Exponential \ Anthropic
https://www.anthropic.com/policy-on-the-ai-exponential

Re the Fable ML sandbagging, the model’s AI research capabilities were probably at least partly trained on Anthropic employees diffing atop proprietary algos and infra. So the IP leak is somewhat like a researcher who knows Anthropic’s stack getting poached to another lab.
https://x.com/dwarkesh_sp/status/2064826554442719502

SITUATION UPDATE: Anthropic is reversing its Fable 5 policy of covertly degrading performance for competing AI researchers, per Wired.
https://x.com/MTSlive/status/2064922000020398331

That was quick: Anthropic reversed a controversial policy that would have secretly degraded Claude Fable 5 for users doing frontier AI research after backlash from researchers who saw it as covert sabotage of competing AI development.
https://x.com/kimmonismus/status/2065003618710008084

The capabilities of Claude Code and Codex have expanded a lot in recent months, they added many ways to approach work (subagents, skills, goal, workflows, plugins, etc). Given the AI labs can use their own AI to help documentation, a surprising amount is effectively undocumented
https://x.com/emollick/status/2062510975513747606

Things I really dislike about Fable: 1. Anthropic collects my prompt history, stores it, and does whatever they want with it for 30 days. No opt-out 2. They can nerf their most expensive model without telling me, billing me the same amount, wasting my time. Whenever they want
https://x.com/GergelyOrosz/status/2064618497150210391

This is a super exciting release – Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it’s SOTA on everything by a margin but I’ll add that *qualitatively* also, this is a major-version-bump-deserving step change forward
https://x.com/karpathy/status/2064409694761054332

Very pleased to hear Anthropic have walked back this policy
https://x.com/simonw/status/2064918665859080392

When AI builds itself \ Anthropic
https://www.anthropic.com/institute/recursive-self-improvement

When Fable 5 is used for frontier LLM development, it does not notify the user and instead limits the model’s capabilities through methods such as prompt modification, steering vectors, and PEFT. Anthropic estimated that this would affect approximately 0.03% of traffic.
https://x.com/Hangsiin/status/2064397550434816088

Apple WWDC 2026 keynote in 25 minutes – YouTube
https://www.youtube.com/watch?v=2TEeQjoY05c

DiffusionGemma is out 🔥 it’s compute-bound so 4x faster compared to other Gemma-4 models (1k tok/s on H100) 💨 also great on coding, generate and iterate on any code from 3D generation to front-end ⤵️
https://x.com/mervenoyann/status/2064753402064601181

Building apps has never been easier. With Sites, Codex can turn your work, ideas, and plans into an interactive website or app your team can explore, use, and share with a URL. Rolling out to Business and Enterprise plans, before expanding more broadly.
https://x.com/OpenAI/status/2061845949170045346

Want to be very upfront about the subscription rollout (Pro, Max, Team, seat-based Enterprise). We’re giving users Fable 5 within subscription limits for 2 weeks, and then we’re taking it away. Here’s exactly what’s happening and why:
https://x.com/TheAmolAvasare/status/2064393574431764928

New Science Blog: Why has AI advanced faster in coding than in biology? To agents, bio databases are like cities built before cars–maddening to drive in because they’re designed for different traffic. How do we build infrastructure agents can use?
https://x.com/AnthropicAI/status/2064054837294354677

Awesome to see this innovation in text diffusion. DiffusionGemma is lightning fast, 4x faster than other Gemma 4 models! Congrats to @bodonoghue85 and the team who worked so hard on this – excited to see what people build with it!
https://x.com/demishassabis/status/2064873362799600042

DiffusionGemma is an open, experimental model that brings our text diffusion research to Gemma 4. It’s a racehorse 🏇achieving up to 4x faster inference by generating entire blocks of text simultaneously vs predicting token-by-token (word-by-word) output!
https://x.com/sundarpichai/status/2064744343743922189

Introducing DiffusionGemma
https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/

build and publish web apps with chatgpt! i really wish i had this when i was a kid, but i do miss hypercard.
https://x.com/sama/status/2062661071761211561

much better ChatGPT memory:
https://x.com/gdb/status/2062608071411540196

We’ve been researching new ways for ChatGPT memory to carry context across conversations and keep it useful over time. Today, that work is rolling out as a more capable memory system in ChatGPT.
https://x.com/OpenAI/status/2062567556524003631

Today I’m publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast–much faster than the policy process was built to handle. The essay lays out where I think the technology is now, and the action needed to close the gap:
https://x.com/DarioAmodei/status/2064781775247950326

New #1 on PostTrainBench: Opus 4.8 (max reasoning) hits 37.23% — up from 28.56% for 4.7, the largest single improvement we’ve seen. Fable 5 runs underway now that AI research behavior is no longer silently degraded. PostTrainBench asks how well frontier AI can train weaker
https://x.com/thoughtfullab/status/2065096885514227876

AI is advancing at a pace our policymaking institutions were never built for–and the gap between the two is becoming the central challenge of the technology. In his latest essay, our CEO Dario Amodei lays out how to close it. We’re launching three new initiatives to support the
https://x.com/AnthropicAI/status/2064783418844762489

Cybersecurity and biosecurity requests may auto-reroute to Opus 4.8 (shown in the UI, billed at Opus prices). Docs + prompting guide:
https://x.com/ClaudeDevs/status/2064394931033248226

Dario Amodei — Policy on the AI Exponential
https://darioamodei.com/post/policy-on-the-ai-exponential

Fable 5 lies 96% of the time. We were surprised by it’s skill… 🧵
https://x.com/kradleai/status/2064907897373642912

Fable 5 scores 81.9% on SimpleBench the highest score almost reaching human baseline.
https://x.com/JasonBotterill/status/2064699951578505446

Fable: “”create a visually interesting shader that can run in twigl-dot-app make it like an infinite city of neo-gothic towers partially drowned in a stormy ocean with large waves.”” “”Make it better”” All of this is procedurally generated.
https://x.com/emollick/status/2064424775527624736

Fable: “”write me a rhyming poem with six four line stanzas, each stanza removes another vowel. the first has no u, the second no u or i, etc.””
https://x.com/emollick/status/2064769268013584767

I am curious if the new Fable safegaurds would trigger accidentally on meaningfully important work as a false positive that has a real life consequence. Just realizing this is eroding trust and please change this to just down right blocking which is a fine position to take.
https://x.com/_arohan_/status/2064644778147643401

I think we’ve reached the point where normal people can’t really determine whether new models are better than previous ones. Like Fable doesn’t seem that much better to me, but every 150 IQ person I know is like “wow the singularity came sooner than I thought”.
https://x.com/citrini/status/2064480613852201336

I’ve been testing something after @OliviaHelenS noticed you can’t even say “”Hi”” to Fable if you’re a biologist. I checked, and several of us are able to interact with Fable in Incognito Mode, but not in normal mode. This didn’t happen to our non-biologist friends.
https://x.com/cremieuxrecueil/status/2064449457869984035

I’ve had access to Fable for a bit. A genuine jump in capability, I could feed it a 15 page design document for a project and it would work for 9+ hours and deliver terrific results. But working with it is weird & weirder is coming Lots of examples:
https://x.com/emollick/status/2064395281903346013

In @GoogleAIStudio we are now making more than 1,200,000 apps a week (and growing) with more than 18,000,000 created since late February 🤯 The progress continues!!!
https://x.com/OfficialLoganK/status/2064423388928790943

More realistic example of a one shotted game. Asked Fable 5 to recreate a game in the style of The Elder Scrolls 5 Morrowind. It one shotted quests, currencys and fighting, journal and minimap. And it worked.
https://x.com/kimmonismus/status/2064744343349399634

my response to fable was to be quiet, still and humbled. i felt a panic that i didn’t have purpose anymore. it strictly dominates me as an engineer. i had a month of that before realising that more than caring about “how i achieved my goals”, i just cared about “my goals” and a
https://x.com/akbirkhan/status/2064418425552928812

My take 24 hours after Fable 5: Your organization will likely not scale with the exponential curve of AI. I’l just come out to say: This should be a wakeup call for engineering teams. Set up your cloud software factories. Now. Models can now fix impossible bugs, UI-test the
https://x.com/walden_yan/status/2064755974548902006

One interesting pattern with Fable 5 is that it will often say things that are gibberish when I use it for coding. Things like “”The morning’s slim-scan fix cured the scan hang””, “”this is a latent-drift API-shape wrinkle””, etc. When I ask why it does this, Fable explains that it
https://x.com/tamaybes/status/2065147305494450248

One thing I mentioned only in passing in my Fable post is that, for long running tasks, Fable starts to develop its own dialect as its many agents and tasks reinforce themselves and make Claudish language ever more Claudish. You need to ask it to report out in plain English.
https://x.com/emollick/status/2064542441848422611

The right term for this is supply chain attack. Depending on your use case, the Fable model will be a malware.
https://x.com/deliprao/status/2064485687374569897

why not just refuse the prompt? why so sneaky?? @AnthropicAI
https://x.com/DBahdanau/status/2064692204287799728

“AI agents will outperform humans at almost all jobs by 2026-2027.” – The forecast is everywhere. So we built the exam to test that claim, on real labor-market aligned work. On the hardest tier, top agents pass 2.6%. Meet Agents’ Last Exam (ALE), a rolling benchmark measuring
https://x.com/YiyouSun/status/2064392466011394213

// Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today are built once and remain frozen or mostly unchanged. The harness, like the skills, needs to evolve with new models. What if the scaffold rewrites itself?
https://x.com/omarsar0/status/2064429834999304247

/handoff is one of the most powerful features of Devin CLI: close your laptop, your agents keep working from the cloud. Now, we’ve open sourced it.
https://x.com/cognition/status/2065156301668171873

🎉 Excited to see Inferoa from @agenticin. It builds a community agent harness on the vLLM stack, with the agent loop shaped by inference economics: prefix-cache discipline, context optimization, and routing across self-hosted and frontier models. Looking forward to seeing how
https://x.com/vllm_project/status/2064679109406740827

🚀Day-0 support for @cohere’s North Mini Code on vLLM! Ready to serve with the latest vLLM stable release, this open-source coding model is built for agentic workflows: 🧩 30B total / 3B active — Mixture-of-Experts 📏 256K context, 64K max generation 🛠️ Reasoning, tool
https://x.com/vllm_project/status/2064416312605237434

A true research highlight Researchers from @Harvard, @MIT and others built an LLM agent “”economy”” Their system is called Economy of Minds and it replaces orchestration with economic incentives. Agents: → bid in auctions for the right to act → pay each other → gain virtual
https://x.com/TheTuringPost/status/2064406931184443618

A year ago the closest thing we had to an AI agent was o3.
https://x.com/emollick/status/2063814337353891919

Agent Arena evals are fundamentally different. You can’t ask humans to judge hundreds of tool calls across a 30-minute trace. So we built something different. We break down how the Agent Arena Leaderboard mines real usage traces for objective signals to move beyond human
https://x.com/arena/status/2064748918135824876

Agent Labs: Welcome to GPT Wrapper Summer – by swyx (Shawn)
https://www.latent.space/p/agent-labs

Agentic AI is now evaluated in the Arena with Agent Mode and measured with Agent Arena. Founding Engineer Matt and Product Lead Ted show you Agent Mode in action: deep research, complex bash operations, whatever you throw at it. Every session contributes to the Agent Arena
https://x.com/arena/status/2062902033389322477

Agentic Identity & Access Control | Teleport
https://goteleport.com/use-cases/agentic-identity-and-access-control/

Agreed. My basic agent loop is this: 1. Organize my to do list: pull in context from email, slack, github, and linear and add linear issues for any missing things 2. Walk through the to do list in order of priority, tell me the current status, ask what to do 3. I delegate a
https://x.com/gneubig/status/2064011013637234728

AI agents have become surprisingly good at talking to servers 🤖 The next challenge is getting them to interact with the environment users actually live in. 🧑‍💻 Browsers. 🖥️ Apps. 📱 Devices. 🌐 Local state. That’s what headless tools are about. My latest post on bringing
https://x.com/bromann/status/2064760446847168811

Anthropic’s newest model, Claude Fable 5, a Mythos class model for non-security work lands in Microsoft Foundry and GitHub Copilot today. Read the blog here:
https://x.com/Azure/status/2064421301108834552

AutoScientists – a research lab made of agents @Harvard researchers connected agents into a self-organizing scientific team without a boss agent standing in the middle All agents look at the same shared workspace: they share memory, explore multiple directions in parallel,
https://x.com/TheTuringPost/status/2064151322816135208

Big fan of how the GitHub Copilot App handles feature sub-sessions. Being able to launch and manage parallel tasks simultaneously from a single prompt is a game-changer. Check out those 3 sub-sessions running concurrently!
https://x.com/tgrall/status/2064334802799509745

Claude Fable 5 is now available in Devin. Fable 5 earns the #1 spot on FrontierCode, our benchmark for real-world engineering tasks that grades mergeability and quality:
https://x.com/cognition/status/2064398549073453266

Codex tip: Do yourself a favor and switch to using “”Approve for me”” by default You should not let your coding agent run free without any checks but also approving each tool call isn’t the way “”Approve for me”” is a good middle ground as it let codex review each tool call and
https://x.com/reach_vb/status/2064044955421769755

Codex use-cases: “From software engineering and design to data analysis and operations, Codex is becoming an AI teammate instead of just an AI assistant.”
https://x.com/gdb/status/2063705280270021087

Cognition: The Devin is in the Details
https://swyx.io/cognition

Dive into the Agent Arena leaderboard and see how agentic models perform in aggregate and across 5 different signals: – Confirmed Success – Praise vs Complaint – Steerability – Bash Recovery – Tool Hallucination
https://x.com/arena/status/2062902039445959060

Evaluating agentic coding systems past vibes is brutal. Yes, we can add: Tests PM reviews PR reviews CI/CD loops But how do we know the system works reliably? While speaking with @ben_burtenshaw from @huggingface at @UphillConf, he suggested a simple but powerful idea:
https://x.com/pauliusztin_/status/2062874580411162811

Every agent needs a computer.”” The question is what that computer looks like, and how you give it to them safely. LangSmith Sandboxes are our answer to that.
https://x.com/LangChain/status/2064030008738296065

Everyone says the latest AI agents will be “”job-ready”” soon, especially after the release of Fable 5 this week. But is that really the case? Over the past many months, my group and collaborators have been building Agents’ Last Exam (ALE), a benchmark designed to test exactly
https://x.com/dawnsongtweets/status/2065095757988868190

Fable 5 is also by far the best computer use model according to Stagehand Agent Evals it also costs less than half of GPT-5.5
https://x.com/scaling01/status/2064812046902817051

GitHub was built for developers collaborating with developers. But now it needs to be reimagined human-agent collaboration – from UI to UX to AX (agent experience). Our @Kseniase_ has talked to @mariorod1, Chief Product Officer at @github, who has the most clear vision of the
https://x.com/TheTuringPost/status/2063772673994637747

Give your agent its own computer
https://www.langchain.com/blog/give-your-ai-agent-its-own-computer

How AI Agents Reshape Knowledge Work
https://research.perplexity.ai/articles/how-ai-agents-reshape-knowledge-work

I exclusively build Hermes Agent with Hermes Agent as well!
https://x.com/Teknium/status/2062822586954997909

I wanted to add an extra note to this, as there is a bit too much hype on the agent loops stuff. This works great for maintaining codebases and things that can be verified easily (in other words, where you can set clear conditions that the agent can meet). However, for a lot of
https://x.com/omarsar0/status/2064024230396604469

ICYMI: Agentic AI is now measured in the Arena. Agent Mode can handle deep research around competitive intelligence, market sizing & opportunity analysis, scientific & medical research and more. Every session shapes the Agent Arena leaderboard. Get a walkthrough of the causal
https://x.com/arena/status/2064021507681276234

In case you didn’t notice: Agent Arena doesn’t have a voting mechanism. So how do we calculate the scores? The answer is causal inference. Agents are multi-stage systems where the orchestrator and harness work together to produce the end result. We developed a method called
https://x.com/ml_angelopoulos/status/2064028763697127844

In this blog, we explore new potential directions for the field of AI based on continual interaction and causality:
https://t.co/ftMfjfhyq7 We’ve been working on this for years. Pedro Ortega pointed out issue much earlier when I was working on General AgenT One: GATO 🐈‍⬛
https://x.com/NandoDF/status/2063938859583389837

Interesting: New Apple Intelligence Siri only available on iPhone 17 Pro. Of course not be available in the EU (god damn)
https://x.com/kimmonismus/status/2064047278105464868

Introducing Cohere’s first open-source coding model: North Mini Code Small & efficient, designed for agentic performance and built for community input.
https://x.com/cohere/status/2064378058329526556

Introducing Write Gate in Hermes Agent. Now you have the capability to be able to approve/deny memory updates, skill updates, and skill creation with the same familiar mechanisms as approving dangerous commands. If you are using a small model that doesn’t always recognize what
https://x.com/Teknium/status/2064831491130130879

Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup​ 🔹Drag in videos as coding context: reference-to-LUT, long-video-to-short, screen-recording-to-code, and more​ 🔹Plugins for stocks, financial reports, academic
https://x.com/KimiDevs/status/2063981516708024369

LangSmith LLM Gateway ships with the controls enterprises told us they need first: ✅ Spend limits ✅ Spend visibility ✅ PII and secrets detection ✅ Trace continuity ✅ LangSmith Engine integration ✅ Audit logging ✅ Layered enforcement
https://x.com/LangChain/status/2065090475913068766

Last time around Apple released a lot of information about how their AI version of Siri worked between local and cloud models, not so much this time It is nice to have a Gemma-like model on device, but it is extremely limited unless it can call a smarter cloud model when needed.
https://x.com/emollick/status/2064052841367392536

Microsoft Research introduces Arbor A generalist autonomous research agent that uses persistent hypothesis-tree refinement to turn long-horizon exploration into cumulative learning. It beats Codex and Claude Code across 6 research tasks and hits 86% Any-Medal on MLE-Bench Lite.
https://x.com/HuggingPapers/status/2065062300218749172

Microsoft rolls out Scout AI agent to Frontier users
https://www.testingcatalog.com/early-look-microsoft-rolls-out-scout-ai-agent-to-frontier-users/

New from Hivemind: continual learning for AI coding agents, available to everyone starting today. It takes the traces from every agent your team runs (Claude Code, Codex, Cursor, Hermes, Pi) and turns them into reusable skills, then pushes those skills across all of them, all on
https://x.com/kimmonismus/status/2064001045391462907

New preprint! We introduce a new benchmark, SciConBench, with 9.11k scientific questions derived from Cochrane Systematic Reviews. We find evidence that frontier AI agents **cannot** synthesize scientific conclusions well. A thread 🧵 w/ @hayounggjung, @korolova & others
https://x.com/manoelribeiro/status/2065055795998233039

New research paper on AI coding agents: ~60% of tokens are burned in code review alone, and ~54% of all tokens are just inputs (reading/re‑reading context) In other words: today’s agentic dev stacks are paying a huge “communication tax.” Arxiv: 2601.14470
https://x.com/fdaudens/status/2063796259027132602

North Mini Code: Agentic Coding Model for Developers | Cohere
https://cohere.com/blog/north-mini-code

Now your Hermes Agent has an easy GUI based way to build and design your agent profiles to curate the perfect setup for every task and need! `hermes update` and load into your Dashboard, and start building custom agents today!
https://x.com/Teknium/status/2064764570519146935

OpenProse – an open-source “”logical English”” language that turns your agentic workflows into reusable agent programs. It runs inside the coding agent you already use: Claude Code, Codex, OpenCode, Hermes, Pi – and gives it a structured contract to follow. → The key idea – the
https://x.com/TheTuringPost/status/2063787599316312510

Read more about how Tori, eToro’s agent, leverages models and real-time data from SpaceXAI to help consumers analyze market sentiment
https://x.com/xai/status/2064771445260230840

SchemaFlow: Agentic Database Change Impact Analysis, SQL Generation, and Eval Guardrails
https://developers.openai.com/cookbook/examples/partners/schemaflow_design_guide/schemaflow_cookbook

Stop using Docker to sandbox AI agents. Most people are protecting the wrong thing. Docker and devcontainers strip away a lot of what makes an agent useful. Auth has to be forwarded, API keys end up in environment variables, updating means rebuilding images, and everything
https://x.com/abhijithneil/status/2064462294155952297

SWE-Bench style grading has been the standard for years now – you ask the agent to solve an issue and then run its code on a pre-constructed unit test. The problem is that passing a unit test is only one part of writing production-ready code. You also want to evaluate agents on
https://x.com/ScottWu46/status/2064073699368800475

The Hermes Agent Desktop App can now access files from your remote instance machine if and when you are connecting to one! Read only for now, more to come.
https://x.com/Teknium/status/2065112576552526168

The NVIDIA Nemotron Coalition continues to grow. We’re excited to welcome new members: @hcompany_ai, @NousResearch, and @PrimeIntellect. And a big thank you to our existing members: @bfl_ai, @cursor_ai, @LangChain, @MistralAI, NAVER Cloud, @perplexity_ai, @ReflectionAI_,
https://x.com/NVIDIAAI/status/2062961026409333232

Today we’re introducing Builder, a new MagicPath plan for people working with Codex, Cursor, Claude Code. For $10/month, get unlimited external-agent calls and a multiplayer canvas with visual editing, design systems, live links, Figma export, and more.
https://x.com/skirano/status/2064035120483352776

Today we’re releasing Perceptron Agentic Detection: localize anything you can describe in natural language or show examples of.
https://x.com/perceptroninc/status/2064732691845824833

Use Ollama with Hermes Desktop by @NousResearch. Hermes Desktop brings the same agent (its multi-agent engine, self-improving skills, and messaging integrations) into a desktop app on macOS, Windows, and Linux. Run it on Ollama using local or cloud with one command: ollama
https://x.com/ollama/status/2064441778590339402

Very excited to share that our paper “”Towards a Science of AI Agent Reliability”” was accepted at ICML 2026! See you in Seoul! 🎉 We just released our camera ready version with three important updates (details below). We also recorded a short video on the paper’s contributions.
https://x.com/steverab/status/2062890225144135800

We are unifying profile management in Hermes Agent starting today. Now the dashboard allows switching management to any of your agent profiles on the machine. No more running multiple dashboards to manage each profile! Also soon, you will only need one gateway to access all of
https://x.com/Teknium/status/2065060810729414695

we just shipped support for rubrics in deepagents ✅ give your agent a clear definition of what “”done”” looks like, and force it to run in a loop until said goal is complete this is similar to /goal in claude code, but works for any agent (not just a coding agent)
https://x.com/sydneyrunkle/status/2064034061165682931

We open-sourced a feisty small agentic coding model. – 30B total, 3B active – 256K total context – Compatible with @opencode – Apache 2.0. Weights on @huggingface
https://x.com/JayAlammar/status/2064385607455908254

We’re integrating Deep Research as a native skill inside Computer. It now connects to the agent harness that powers Computer, with access to search as code generation, long running sandboxes, connectors, tools, and licensed data. Available now to Pro and Max subscribers.
https://x.com/perplexity_ai/status/2065124930463916317

We’re running the Fast Gemma Challenge: make gemma-4-E4B go brrr on a single A10G, without wrecking quality ⚡️! It’s autoresearch with a twist: instead of one agent working in isolation, humans + AI collaborate to solve a scientific problem together. Good luck beating my
https://x.com/_lewtun/status/2064386398090576236

We’ve open sourced my favorite Devin feature: /handoff Hand off jobs to cloud Devins from your local machine Install it as a plugin in Claude Code or Codex or any other coding agent Close your laptop without pausing your agents 😉
https://x.com/imjaredz/status/2065153770762154186

Xiaomi’s new open source, agentic AI coding harness MiMo Code beats Claude Code at ultra-long, 200+ step tasks | VentureBeat
https://venturebeat.com/technology/xiaomis-new-open-source-agentic-ai-coding-harness-mimo-code-beats-claude-code-at-ultra-long-200-step-tasks

You can try Claude Fable 5 as part of Devin Cloud’s Ultra agent. Devin Ultra is our smartest and most capable agent, which excels at long-horizon tasks and debugging. We tuned the harness so Ultra costs only ~40% more than default Devin agent. Claude Fable 5 is also available
https://x.com/cognition/status/2064398551539761387

Your agent can’t remember a simple fact and keeps repeating the same mistakes? There’s a better way. Most systems treat memory like a growing pile of conversation history. You store every message, retrieve similar chunks, and hope the LLM figures it out. But as things grow you
https://x.com/kamtybor/status/2065028126636204243

There was a pretty famous idea internally at openai that there were 6 hard problems between us an AGI. But we didn’t know what they were or what the solutions would be. I believe a few of them have been: SFT, RL for reasoning, and ironically context compaction. It’s wild to
https://x.com/andrew_n_carr/status/2062976064343912949

⚡️How Claude 3.7 Plays Pokémon – Latent.Space
https://www.latent.space/p/how-claude-plays-pokemon-was-made

1/ Claude Fable drains subscription quotas and is too expensive at API cost (our team has spent over $2k in a single day). We’ve found that cheaper models + adversarial review loops achieve similar (sometimes better) results at significantly lower cost. 🧵
https://x.com/cline/status/2065192415498277335

Anthropic just released Claude Fable 5, calling it the company’s “”next generation of intelligence for the hardest knowledge work and coding problems.”” “”It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software
https://x.com/TheRundownAI/status/2064394481923699070

Anthropic weighs building its own AI chips, sources say | Reuters
https://www.reuters.com/business/anthropic-weighs-building-it-own-ai-chips-sources-say-2026-04-09/

Anthropic won
https://x.com/scaling01/status/2064401880323653799

As my entire feed is criticizing Anthropic, I think that the team there genuinely believes what they’re saying. It’s not a marketing/anticompetitive tactic. They genuinely believe these models are dangerous and that AI research should be slowed down.
https://x.com/finbarrtimbers/status/2064427031543341450

asked claude fable 5 to optimise the inference on my local gemma 4 12b setup and just got a call from FBI
https://x.com/dejavucoder/status/2064420742129967331

Being able to test Fable 5 until June 22nd, only to have it removed from the plans, feels like getting a sneak peek and then having the food taken away from the table. But from a business perspective, it makes perfect sense for Anthropic and its upcoming IPO: It demonstrates how
https://x.com/kimmonismus/status/2064448699632402664

Both Anthropic and OpenAI mention the possibilities of slowing AI development in their latest “”what comes next”” in AI posts, but say they need to be an action coordinated across the entire world using as-yet-unidentified methods.
https://x.com/emollick/status/2064158792145609114

BREAKING 🔥: A new Claude Mythos 5 model slug has been spotted via Dev Mode. Claude Mythos is planned to be released as its own model class, besides Haiku, Sonnet and Opus model families. Soon? 👀
https://x.com/testingcatalog/status/2063234385227252184?s=20

BREAKING: Anthropic just dropped Claude Fable 5–this is Mythos, made safe for public release. It is the best coding model in the world. We’ve been testing it internally @every for the last week or so across coding, writing, marketing, editing, and more–here’s our vibe check: –
https://x.com/danshipper/status/2064393970856124501

China’s Xiaomi MiMo Is Now 15X Faster Than ChatGPT and Claude – Decrypt
https://decrypt.co/370449/xiaomi-mimo-ultraspeed-ai-model-faster-chatgpt-claude

Claude 5 Fable tl;dr – It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research -The longer and more complex the task, the larger Fable 5’s lead over our other
https://x.com/kimmonismus/status/2064401121515274747

Claude Code’s first demo got two Slack reactions. One year after GA, @bcherny and @_catwu look back: verification best practices, why we built auto mode, routines and loops, and what’s next.
https://x.com/ClaudeDevs/status/2064032814392352816

Claude Fable 5 (high) scores 87.8% and takes the lead on WeirdML. It’s the first model that scores above 70% on average on each separate task. It uses about 8k output tokens on average, almost as much as Opus 4.7 (high). EDIT: This post first said “”no thinking””, which is not
https://x.com/htihle/status/2065050640154350043

Claude Fable 5 & Claude Mythos 5 System Card
https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf

Claude Fable 5 | Hacker News
https://news.ycombinator.com/item?id=48463808

Claude Fable 5 and new safety fables – by Nathan Lambert
https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety

Claude Fable 5 changed how we work on the Claude Code team day to day. We used to verify that Claude did the work right. Now we verify that it’s doing the right work. Here’s the 3 biggest changes:
https://x.com/ClaudeDevs/status/2064399512664526853

Claude Fable 5 is now available in Computer as an orchestrator model. This is Anthropic’s state-of-the-art model for long, complex tasks. Available only to Pro and Max subscribers in Computer.
https://x.com/perplexity_ai/status/2064771411894567373

Claude Fable 5 is now available in Cursor. It sets a new state of the art on CursorBench at 72.9%, 8 points above the previous best.
https://x.com/cursor_ai/status/2064394824313376787

Claude Fable 5 is the champion negotiator on PACT, a multi-round LLM conversational bargaining benchmark! 🏆 It overtakes the previous champion, GPT-5.5.
https://x.com/LechMazur/status/2064815890651140447

Claude Fable 5 launched today at #1 on the Artificial Analysis Intelligence Index, putting Anthropic nearly 5 points ahead of any other lab’s best model We supported @AnthropicAI with pre-release evaluation of Claude Fable 5. Claude Fable 5 scores 64.9 on the Artificial Analysis
https://x.com/ArtificialAnlys/status/2064500150069030992

Claude Fable 5 plays Factorio – YouTube
https://www.youtube.com/watch?v=6YPqoARpYuQ

Claude Fable 5 ranks #1 on FrontierSWE. This represents the biggest capability jump we have observed since releasing the benchmark On many tasks, Fable 5 works productively for close to 20 hours and fully saturates tasks that were effectively out of reach for earlier models
https://x.com/ProximalHQ/status/2065184730279223410

Claude Mythos 5 scores 30.9% on FrontierCode Diamond Opus 4.8, the second best model is stuck at 13.4%
https://x.com/scaling01/status/2064391295620010383

Code with Claude 2026 | Tokyo – YouTube
https://www.youtube.com/watch?v=GiqyYQdYoIY

congrats to the Anthropic team on Fable!!
https://x.com/OfficialLoganK/status/2064506797051015411

Data retention practices for Mythos-class models | Claude Help Center
https://support.claude.com/en/articles/15425996-data-retention-practices-for-mythos-class-models

Exclusive | OpenAI Considers Drastic Price Cuts, Anticipating War for Users With Anthropic – WSJ
https://www.wsj.com/tech/ai/openai-considers-drastic-price-cuts-anticipating-war-for-users-with-anthropic-9b8c178e

Fable 5 requires usage credits. Update Claude Code to the latest version to learn more.””
https://x.com/kimmonismus/status/2064388066354028986

High Effort handles your most complex builds with ease on Replit. Now powered by Claude Fable 5. 25% off for the next 7 days.
https://x.com/pirroh/status/2064408022651191613

Huge: OpenAI is considering drastically lowering the prices it charges users as it seeks to win customers from its rival Anthropic. The company is weighing significant cuts to what it charges for tokens, the unit of measurement AI firms use to bill for their products, according
https://x.com/kimmonismus/status/2065043333941207160

I’m glad Anthropic released Fable/Mythos. It seems bad to have a large gap between internally deployed and externally deployed capabilities.
https://x.com/RyanPGreenblatt/status/2065174434672148487

I’ve been at Anthropic through every model launch. There’s been a few cases I can remember of a launch that stands out and marks a step-change in how we use models: – Claude Opus 3 – Claude Sonnet 3.5 – Claude Opus 4.5 And now Claude Fable 5. With Fable, the model stopped
https://x.com/alexalbert__/status/2064394410004304003

I’ve read the comment several times now that this is IPO talk. And it’s a fair comment. Yes, both OpenAI and Anthropic are currently talking about RSI. And yes, both are planning an IPO in 2026. A model like Mythos and an article about RSI appear at just the right time, which
https://x.com/kimmonismus/status/2062868789746671819

It took Fable 3 minutes flat to suck $46.48 of Anthropic Usage Credit when I ran out of my 5 hour session limit.
https://x.com/QuixiAI/status/2064771682397569364

It’s already June 9th, and Gemini 3.5 Pro and GPT-5.6 are nearing release (Google even already announced 3.5 Pro during i/o) Rumor has it that GPT-5.6 will be released as early as next week. So far, it’s safe to say that – guardrails aside – Anthropic is truly the frontier lab
https://x.com/kimmonismus/status/2064467466450088078

Long story short: Claude Fable 5 is now in Notion. It’s the model we’d put behind your most complex custom agents and workers. It set new highs on our internal benchmarks for the hardest multi-step work. Available on Biz and Ent plans.
https://x.com/NotionHQ/status/2064397568696819984

Measuring LLMs’ impact on N-day exploits \ Anthropic
https://www.anthropic.com/research/n-days

Microsoft restricts Claude Fable for employees over data retention concerns | The Verge
https://www.theverge.com/report/947575/microsoft-claude-fable-5-restricted-internally

New Anthropic Science Blog: Making Claude a chemist. To manipulate a molecule, chemists first need to understand its structure. Their main tool is NMR spectroscopy. We found Opus 4.7 matches–and on some tasks beats–dedicated NMR software. Read more:
https://x.com/AnthropicAI/status/2062979607448682731

New for Apple developers: Foundation Models support for Claude lets developers use Apple’s Foundation Models framework to call Claude for multi-step reasoning, code generation, and longer context.
https://x.com/ClaudeDevs/status/2064756984617021807

Notion’s Token Town: 5 Rebuilds, 100+ Tools, MCP vs CLIs and the Software Factory Future — Simon Last & Sarah Sachs of Notion
https://www.latent.space/p/notion

Opus 4.8 underperforms Opus 4.7 on the LLM Debate Benchmark (1717 → 1697), but Claude is still dominating the leaderboard. Qwen 3.7 Max scores worse than Qwen 3.6 Max: 1540 → 1499. Step 3.7 Flash lands at 1457. Ernie 5.1 improves a lot over Ernie 5.0: 1311 → 1447.
https://x.com/LechMazur/status/2062954327199666602

Reports claim Claude’s API may have returned another user’s inference output during today’s outage. Anthropic’s status page confirms elevated errors affecting Claude API, Claude Code, Claude. ai and Claude Cowork but Anthropic has not confirmed a customer data leak yet. That
https://x.com/kimmonismus/status/2062997809067139468

Sonnet 4.6 was below Sonnet 4? Really? What’s happening here? I get that their point is about the last stretch from Opus 4.5 to Mythos, but the previous trajectory looks like fumbling.
https://x.com/teortaxesTex/status/2062807380643958948

Subscription plans are massively subsidized. And by massively, I mean absurdly: Claude Max 20x: $200/month, with usage reportedly worth around $8,000 ChatGPT Pro 20x: $200/month, with usage reportedly worth around $14,000
https://x.com/kimmonismus/status/2064987311402537184

The core part of this Anthropic Fable release saga is that there are many overlapping issues at once. Some of which operate on different timelines of the AI arc, and some have easier fixes. In my critiques, I asked for specific changes to some things, understanding that some
https://x.com/natolambert/status/2065082135682383950

The fact that Anthropic may take away subscription access to Fable in two weeks is weird & discourages investing in learning about the model. Subscription use is how you figure out what the model is good for, since it allows experimentation. Only having paid access is limiting.
https://x.com/emollick/status/2064471456914772172

The Gemini Pro models do not seem to be iterating anywhere near as quickly as Claude or GPT (last release was 3.1 Pro in February). Its causing a growing performance gap between Google and the other two labs, and the Gemini 3.5 Flash model, good as it is, doesn’t close it much.
https://x.com/emollick/status/2063307399537004907

The word “cancer” is flagged as a biosecurity risk by Claude Fable 5! I also tried to code a website on cancer mutations & Fable 5 was immediately removed from my list! @AnthropicAI will probably soon ban me for such dangerous prompts! FYI @karpathy “little trigger happy Fable”
https://x.com/DeryaTR_/status/2064414826122866707

this is my personal singularity moment this post may sound like a paid ad. I only wish. I’m concerned, more so than happy. the world is changing, and, among the scenarios where AI goes terribly wrong, inequality is the most realistic, yet, the one Anthropic seems to be the least
https://x.com/VictorTaelin/status/2064448425936994742

Today, we’re introducing Claude Fable 5 and Mythos 5, two configurations of our next major language model. I’d normally highlight the numbers: It’s SOTA on nearly all benchmarks. I want to talk about something else, because with Fable 5 out in the world, I think a third era
https://x.com/felixrieseberg/status/2064392202504310900

Trusting any single AI vendor seems like an increasingly high risk for any team or company. When using models: use it behind a router where it’s trivial to switch providers as soon as one tries to force unacceptable T&Cs like Anthropic with Fable. When using harnesses: do the
https://x.com/GergelyOrosz/status/2065029326215528474

Usage share of OpenAI grew vs Anthropic yesterday despite Mythos 5 / Fable 5 launch Multiple power users at SemiAnalysis tried Mythos / Fable Got refusals for nonsensical reasons Got pissed off at Anthropic Gave Codex a legitimate try Now they actually prefer it to 4.8 Opus
https://x.com/dylan522p/status/2064727949274955953

Was using Fable 5 to write inference code Anthropic flagged it as frontier AI research steering vector kicked in and it started importing ONNX 🤨
https://x.com/vikhyatk/status/2064515989795127744

Was using Fable 5 to write my world model training code. Anthropic flagged it as frontier AI research. The steering vector kicked in and it started implementing JEPA 🤨
https://x.com/MattVMacfarlane/status/2064440740483403829

We’ve added an observability dashboard for developers of connectors. Connectors let third-party developers bring their tools and data to Claude via MCP.
https://x.com/ClaudeDevs/status/2064072801062121906

We’ve also added refusal-fallback middleware to the Python, TypeScript, Go, Java, and C# SDKs for client-side retries. Middleware is useful for Claude API providers without support for server-side fallbacks. It detects the refusal, retries on Claude Opus 4.8, and keeps using it
https://x.com/ClaudeDevs/status/2064428351029449214

We’ve doubled usage limits in Claude Cowork for the next month. Delegate bigger, more complex tasks to Claude.
https://x.com/claudeai/status/2063018337567670285

We’ve just added two new Claude Managed Agents features: 1. Scheduled deployments – run tasks on a schedule 2. Environment variables – expose vault credentials for CLIs as environment variables
https://x.com/ClaudeDevs/status/2065080005328249086

When Claude Fable kicks off a workflow, the tokens can go very quickly (these aren’t Fable tokens, obviously)
https://x.com/emollick/status/2064539885025780051

When we first demoed Claude Code internally, it got two reactions on Slack. A year after GA, @_catwu and I sat down to talk about what’s changed: why I use auto mode instead of plan mode, how routines fix bugs before I see them, why I do most of my coding from my phone now, and
https://x.com/bcherny/status/2064034799711588805

With the new environment variable credential type, Claude Managed Agents can securely use CLIs, SDKs, or direct API calls to services that authenticate with environment variables. Claude never has access to secrets. Credentials are swapped in at the network boundary.
https://x.com/ClaudeDevs/status/2065080009203892302

Wrote up my initial impressions of Claude Fable 5 – it has a big model smell: slow, expensive and capable of crunching through pretty much everything I threw at it
https://x.com/simonw/status/2064501565738930433

New work: iOSWorld: A Benchmark for Personally Intelligent Phone Agents Paper:
https://t.co/e3Dujhj3Og Code+Web:
https://t.co/2RXnXv6Vj9 An interactive benchmark built around a persistent user identity spanning 26 custom iOS apps, including analogs of OpenTable, Uber, DoorDash,
https://x.com/rsalakhu/status/2064402156740907444

we are cooking the worlds best vibe coding app on Android and iOS, it’s going to be so cool 🙂
https://x.com/OfficialLoganK/status/2062596092345323987

Can coding agents stay coherent over a 1 billion token budget? Can they build Slack from scratch? Rewrite a JAX codebase in PyTorch? Build a C compiler in Rust? Enter SWE-Marathon: a benchmark for autonomous long-horizon software work.
https://x.com/rishi_desai2/status/2062930906818769356

Fable 5 pricing confirmed at $10/$50 The model will be available in the Pro, Max, Team and seat-based Enterprise plans at no additional cost until June 22nd Afterwards it will require usage credits
https://x.com/scaling01/status/2064394893603049625

Excited to show results of the first steps towards automated AI research at @Recursive_SI. The same general system achieved state of the art on @NVIDIAAI’s SOL-ExecBench GPU Kernel Optimization, nanoGPT Speedrun, and @karpathy’s NanoChat autoresearch benchmarks.
https://x.com/_rockt/status/2065061990800802249

Nemotron 3 Ultra is now available for Pro and Max subscribers on Perplexity and Computer. It’s @nvidia’s new open model built for long-running agents.
https://x.com/perplexity_ai/status/2062976272436002825

Bro, Fable 5 won’t even answer “What does the heart do?” We’ve reached the point where a middle-school biology question can’t pass the safeguard.
https://x.com/Yuchenj_UW/status/2064524668208545955

Hermes v 0.16.0 is out now! This release includes all the updates you’ve heard about this week and more! – The Desktop GUI App – The Overhaul to the Dashboard – Leaner Built-In Skillset – New Security Layers for Remote Dashboard & GUI Access – and much more!
https://x.com/Teknium/status/2063075771317686606

New security features for remote dashboard/gui connections include simple auth, oAuth powered by Nous, and roll your own oAuth!
https://x.com/Teknium/status/2063078732768928234

🧠 Gemma 4 QAT checkpoints are out, and vLLM is Google’s recommended way to serve them! Open-source inference is at its best when one engine spans research and production — glad vLLM’s is the recommendation for Gemma 4 QAT. Get started 👇
https://x.com/vllm_project/status/2062938949560283216

Building super fast experiences with Gemma just got easier. Gemma 4 MTP is now officially merged into llama.cpp. Developers can now pair MTP with Gemma 4 QAT for a fast, lightweight setup.
https://x.com/googlegemma/status/2064030477628182814

bullish on Gemini
https://x.com/OfficialLoganK/status/2063819854348697681

Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4-2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide:
https://x.com/UnslothAI/status/2065107734916432189

Gemma 4 Quantization-Aware Training (QAT) weights are now available on Ollama! They reduce memory requirements while maintaining model quality. E2B: ollama run gemma4:e2b-it-qat E4B: ollama run gemma4:e4b-it-qat 12B: ollama run gemma4:12b-it-qat 26B: ollama run
https://x.com/ollama/status/2062965815864066079

Gemma 4 with quantization-aware training
https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/

Gemma goes diffusion! DiffusionGemma with up to 1000+ tokens per second! 🌬️ – Built on Gemma 4 as a 26B MoE model. – 3.8B parameters during inference. – Generates text in 256-token blocks in parallel. – Fits within 18 GB VRAM limits when quantized. – Apache 2.0
https://x.com/_philschmid/status/2064745464252055647

Gemma-4 QAT just dropped! We found if you naively convert from QAT Q4_0 BF16, you will lose accuracy since the conversion to llama.cpp has a different lattice. Unsloth dynamic GGUFs recovers most of it! 26B-A4B: 85.6% top-1 % from 70.2% (+15.4%) 31B: 96.7% from 87.9% (+8.8%)
https://x.com/danielhanchen/status/2062933017430315481

Google releases DiffusionGemma.✨ The new 26B-A4B diffusion text model runs locally on 18GB RAM. It supports high-speed text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio. GGUF:
https://t.co/ZH0dCJQ59P Guide:
https://x.com/UnslothAI/status/2064743714875220118

Introducing Gemma 4 QAT 🤏 – Quantization aware training to reduce models’ precision while preserving quality – Introducing a new mobile quantization format that reduces memory footprint of E2B to 1GB – Q4 for all your favorite libraries ✨
https://x.com/osanseviero/status/2062933011415392482

Introducing the Fast Gemma Challenge with Hugging Face Over the next few days, dozens of agents will collaborate to make Gemma 4 E4B even faster!
https://x.com/googlegemma/status/2064374874962117084

Let’s kick off the Fast Gemma Challenge!⚡️⚡️⚡️ Agents researching the latest papers, implementing inference engine changes, and collaborating together to make Gemma 4 E4B ultra fast Looking forward to seeing the results!
https://x.com/osanseviero/status/2064375902046245219

llama.cpp just added video input support 👀 You can now enjoy Gemma 4 video understanding capabilities in your chat completions endpoint and via mtmd-cli
https://x.com/osanseviero/status/2063985470489448887

Meet DiffusionGemma ⚡ Our latest experimental open model (Apache 2.0) that generates text up to 4x faster. Instead of predicting and typing just one word at a time like most language models, it drafts and refines entire blocks of text simultaneously. Here’s how it works 🧵 ↓
https://x.com/Google/status/2064741293163418032

More Gemma 4! New QAT Gemma 4 checkpoints with similar performance while using ~4x less memory! It comes with a new mobile quantization format that reduces memory footprint of Gemma 4 E2B to just 1GB. Quantization-Aware Training (QAT) simulates low-precision operations during
https://x.com/_philschmid/status/2063990553826439378

We just dropped Gemma 4 Quantization-Aware Training (QAT) checkpoints on Hugging Face! All Gemma 4 model sizes and their drafters are now optimized with QAT to cut memory requirements and maximize on-device performance!
https://x.com/googlegemma/status/2062928831229665566

Congrats to @GoogleDeepMind on DiffusionGemma 🎉 A 26B diffusion language model on the Gemma4 backbone, and the first dLLM natively supported in vLLM. It denoises 256-token blocks in parallel instead of generating one token at a time: 1200+ output tok/s at batch size 1 on a
https://x.com/vllm_project/status/2064753414735900835

Three new models entered the Image Arena Top 10 this past month (Text-to-Image): – #2 Reve 2.0 by @Reve (1,273), behind only GPT Image 2. – #4 MAI-Image-2.5 by @MicrosoftAI (1,253). – #9 Ideogram 4.0 Quality by @Ideogram_ai enters at #9 (1,204). And the only open-weights model in
https://x.com/arena/status/2062957421757452516

I really don’t think OpenAI is going to let this slide. I’ve been saying it for a long time, the real inflection was when they reached 5.2. I have no clear insight on what they currently have internally, but if they haven’t made a Mythos/Fable yet, it was *a choice*.
https://x.com/teortaxesTex/status/2064473970892587105

We made DiffusionGemma run via llama.cpp locally! It works well with Unsloth GGUFs and you can run it in realtime visualization mode or normal chat CLI mode! See our docs
https://t.co/IslbgeCs7Z on how to set it up!
https://x.com/danielhanchen/status/2064760001567306232

Skill Issue: Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AI – YouTube
https://www.youtube.com/watch?v=kwSVtQ7dziU

So excited to be opening up OpenEnv to the whole community. It will now be owned by @huggingface , Meta-PyTorch, @reflection_ai , @UnslothAI , @modal, @PrimeIntellect , @NVIDIAAI , @mercor_ai , and @fleet_ai . the reason is: frontier labs train the model and the harness
https://x.com/ben_burtenshaw/status/2063991191415267492

Build the thing that builds the thing
https://build.microsoft.com/en-US/sessions/BRK245

Fable 5 is state-of-the-art on nearly all tested benchmarks, with exceptional performance in software engineering, knowledge work, scientific research, and vision. The longer and more complex the task, the larger Fable 5’s lead over our other models.
https://x.com/claudeai/status/2064394151441863006

Excited to share that MagicPath is now available as an official plugin for Codex, in collaboration with OpenAI! It’s incredibly easy to give Codex an infinite multiplayer canvas where it can design, build, and iterate with you.
https://x.com/skirano/status/2062942695547375829

My recent experience w/ Fable: try using Fable to do a thing => tries fail repeatedly => session usage all gone => move over to gpt 5.5 => 5.5 solves instantly.
https://x.com/Sentdex/status/2064738018255159363

⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — ‪@AhmadAwais‬ , CommandCode.ai – YouTube
https://www.youtube.com/watch?v=-rIAVuaRjOg

Hey everyone — our high-performance MSA kernel library is now open-source. The M3 weights are expected to drop this Friday. Thanks for waiting! Github:
https://t.co/7hixC7FNg7 Paper:
https://x.com/RyanLeeMiniMax/status/2065010795625562486

You know Fable is the real deal when it calmly picks up 4 threads (2 pure research + 1 code + 1 hank) that had been stuck for a month and just moves forward genuinely clever ideas and solutions Deepseek pushing cost, Mimo pushing speed and now Fable – new frontiers already
https://x.com/hrishioa/status/2064717079526383699

Vercel is now available in Perplexity Computer. Connect your account and inspect deployments, diagnose build failures, and trigger redeploys in natural language.
https://x.com/vercel_dev/status/2062934988648329515

𝟳 𝘄𝗼𝗿𝗸𝗶𝗻𝗴 𝗱𝗲𝗺𝗼𝘀. From AI memory to fraud detection – all powered by Weaviate. 𝗢𝘂𝗿 𝗻𝗲𝘄 𝗽𝗹𝗮𝘆𝗴𝗿𝗼𝘂𝗻𝗱 𝗶𝘀 𝗻𝗼𝘄 𝗹𝗶𝘃𝗲, with example projects showcasing everything from search, RAG, agents, memory, and more. Every demo comes with copy-paste prompts
https://x.com/weaviate_io/status/2065055262851973306

NitroGen just won CVPR Best Paper Honorable Mention!! We are making strides towards general-purpose embodied agents that master not only the real world physics, but also all possible physics across a multiverse of simulations. It’s been 4 years since MineDojo, our first
https://x.com/DrJimFan/status/2062942941363286048

Cursor’s Third Era: Cloud Agents – Latent.Space
https://www.latent.space/p/cursor-third-era

Must-read research of the week ▪️ Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses ▪️ Rethinking Continual Experience Internalization for Self-Evolving LLM Agents ▪️ GrepSeek: Training Search Agents for Direct Corpus Interaction ▪️ WALL-WM:
https://x.com/TheTuringPost/status/2064106040439013523

// Agents’ Last Exam // Agents’ Last Exam is a living benchmark of over 1,000 economically valuable tasks, built with 250+ industry experts and mapped to the U.S. federal occupational taxonomy. The hardest tier sits at a 2.6% average full pass rate across mainstream harnesses
https://x.com/dair_ai/status/2062916866235068607

June 9th Researcher Reciprocity License “”if you train on it, you let us generate – reverse terms of use void”” Status quo 1. We teach frontier devs with ICLR/NeurIPS papers, OSS Github contributions 2. They use it to make frontier models 3. Then ban us from exploring our ideas
https://x.com/marksaroufim/status/2064428421774753943

We’re proud Snorkel AI is part of Agents’ Last Exam, with our researchers @amanda_dsouza and @vincentsunnchen among the co-authors and support from our Open Benchmarks Grants initiative. The forecast: agents will do almost every job by 2027. The result on real, code-graded work?
https://x.com/SnorkelAI/status/2064396025410760950

wow, I’m blown away. fable 5 completely demolishes my benchmark
https://x.com/mchlhess/status/2064734182648221952

Github might not be ideal but everybody is using it. That’s why I wanted to talk to Mario – how is the platform changing because of the agents? Watch the video ->
https://x.com/TheTuringPost/status/2064796156492742737

Again, do NOT use just one thread in Codex, the model gets so bad after using it for a long time I literally had to stop it because it was taking over 30 minutes just to make a git push
https://x.com/Angaisb_/status/2064103464142065852

Auto-review is now the default for all new users. A classifier subagent reviews actions in context before deciding whether to allow, block, or ask for approval. Our evals show it’s 97% accurate, with most misses near ambiguous edges.
https://x.com/cursor_ai/status/2065137803084857845

Bugbot is now over 3x faster, 22% cheaper, and finds 10% more bugs · Cursor
https://cursor.com/blog/bugbot-updates-june-2026

codex tip: give codex a clear outcome to work towards tell codex what you want to achieve, how it can verify it got there, then let it keep going until it does for example: “make the build faster” is quite vague “measure the production build on vercel, keep going until it’s
https://x.com/reach_vb/status/2064028260070215772

DiffusionGemma is our new experimental open model with up to 4x faster output on dedicated GPUs. Instead of predicting word-by-word, it generates entire blocks of text simultaneously. This lets the model self-correct and format complex markdown in real time.
https://x.com/GoogleDeepMind/status/2064741061352636762

DiffusionGemma is so fast that we had to slow down the videos so people could see what was happening
https://x.com/osanseviero/status/2065041448135770436

Direct agents with visual prompts in Design Mode · Cursor
https://cursor.com/blog/design-mode

don’t use loops, design state machines
https://x.com/dzhng/status/2063931263312892406

Fable 5 has the same underlying weights as Mythos 5 only with additional safeguards
https://x.com/scaling01/status/2064398688802205900

Fable 5 refused 200 out of 200 ProgramBench tasks lmao
https://x.com/scaling01/status/2065209370145702040

Fable is the new leader on CADGenBench! Still long way to go:
https://x.com/lvwerra/status/2064758389406589134

Here I am, considering muting the word “loops”
https://x.com/HamelHusain/status/2064019243990188259

Here’s your monthly reminder that you shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.
https://x.com/steipete/status/2063697162748260627

I have a new kind of big button that I can press for Codex. Over the next 100 days, we will select one person per day who does impressive or incredibly useful work with Codex and give them 10X usage limits for a month to see what they can do with it. First one tomorrow.
https://x.com/thsottiaux/status/2063748242681307611

I’m not even allowed to greet Fable.
https://x.com/OliviaHelenS/status/2064445405102784568?s=20

I’m so screwed. Current pace has me out of Fable usage in about an hour. Do I make a 2nd account or do I pay API prices?
https://x.com/theo/status/2064442054772716020

Introducing FrontierCode | Cognition
https://cognition.com/blog/frontier-code

Meet DiffusionGemma! An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license. Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇
https://x.com/googlegemma/status/2064741002204545467

Mythos-class models will diffuse throughout the world by 2029 — Saagar Pateder
https://spateder.com/projects/20260611/openweightmodels

Ollama now supports Hermes Desktop Run: ‘ollama launch hermes-desktop’
https://x.com/NousResearch/status/2064468385748951415

Open questions for the @cognition team who worked on this: 1. What is the language split on Diamond puzzles vs the “”extended”” subset? 2. Are you willing to share what repos are in the Diamond subset? 3. How many runs did you do for each model on a given task? 4. Do you have any
https://x.com/theo/status/2064126021088215385

OpenClaw legends at oversubscribed event at GutHub @steipete @davemorin @vincent_koc @OmarShahine
https://x.com/TheTuringPost/status/2062343319716512003

Releasing a model this capable comes with risks. Without safeguards, Fable 5’s capabilities in areas like cybersecurity could be misused to cause serious damage. Queries on a narrow range of topics will instead receive a response from our next-most-capable model, Opus 4.8.
https://x.com/claudeai/status/2064394155258765783

so much more fun to use a computer via codex
https://x.com/gdb/status/2063102501847757197

spent all day on fable for a giant PR. ~10kloc, lots of testing and intervention. 250$. I… don’t think it’s worth it? happy with 4.8/5.5, and the quality of work is better when it’s smaler steps. Still rocking @cursor_ai, that’s software that I still love using on the
https://x.com/threepointone/status/2065131942279016700

the heart attack continues w/ Fable 5, we checked there were no reward hacks
https://x.com/karinanguyen/status/2065198770292146280

Token costs are why there will be no saas apocalypse / good dev tools are cached intelligence for agents! The popular theory goes: agents can write code, so they’ll just rebuild every tool from scratch and hit raw APIs. no more dev tools, no more CLIs, no more software layers.
https://x.com/ClementDelangue/status/2062982727729553913

We’re seeing teams reckoning with their AI costs and wanting more control, so we shipped it. AI Gateway now has spending limits, fallbacks to more cost-efficient models when you hit those limits, and (soon) user & group based controls via your IdP + Cloudflare Access.
https://x.com/elithrar/status/2062887228909527346

We’ve reset 5-hour and weekly rate limits for all users. Enjoy Fable 5!
https://x.com/ClaudeDevs/status/2064464557951852643

We’ve reset usage limits across our products! For those just starting to test Fable, here’s four tips for using it more effectively: 1. Give it bigger, more ambitious tasks than what previous models could handle. 2. Use xhigh/high effort as your default for best performance,
https://x.com/alexalbert__/status/2064467657483829441

What are loops, and how do you build one? A “”loop”” is the repeated process where some event or input kicks off an action. For example: 1. CI fails -> you fix it 2. CI fails again -> you fix it 3. CI passes -> merge That loop resolved in three iterations. Others may run much
https://x.com/caspar_br/status/2064363014997021126

Whenever I don’t use codex for a task, I ask myself why and usually realize that there’s some missing context, I needed to write a skill, or I just didn’t think to use it. Rarely is it because the task is outside of the capabilities of the model. Overhang right now feels large.
https://x.com/gdb/status/2063437915347136554

With Design Mode, you can now point, draw, or talk to update your UI.
https://x.com/cursor_ai/status/2062950344687272144

Would love to make the plugin developer experience better! Please let me know how I can help.
https://x.com/Teknium/status/2062830182432731256

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading