About This Week’s Covers

An actual real life White Barn hand soap called Crisp Morning Air inspired this week’s cover. The names and smells of these soaps are hilarious and existential. I used the red font from Baudrillard’s Simulacra and Simulation, which I believe is Garamond, for the title text. These soaps are out of control.

Bonnie Tyler died this week on July 8, 2026. For the category covers, I used the theme of her iconic, dramatic 1983 MTV video for Total Eclipse of the Heart. The over-the-top composition of the video made these covers especially cheesy, not the prompt or the skill. I’ve included my favorites below:

This Week’s Humanities Selections

This week’s Humanities readings are samples from Shahrnush Parsipur, an Iranian author who died on July 3, 2026. The song of the week is the uniquely raspy It’s a Heartache by Bonnie Tyler.

“As we were walking along I was thinking about how many people had to drown so that the first human could learn to swim. Even so, there are still those who drown.”

“Her female fate was to bend, to become small, to fold; thus people would leave her alone and let her share herself in other ways. Love for her had to be of another kind. This love was a stranger to nighttime restlessness. It had no knowledge of trembling bodies and palpitating hearts. It was directed at reaching out rather than union. It worked according to a strategy that was opposite to that of the human male. Therefore, she could not laugh when she wished. She could not eat when she wished. She needed to draw a circle and put her genie—the evil and fire of her worldy desires—firmly under her control.”

“…join the family in the orchard where she had to tolerate the children who screamed all the time as they gorged themselves with cherries giving themselves diarrhea and eating yogurt at night as antidote.”

“She had not learned to be malicious. She only knew malice.”

This Week By The Numbers

Total Organized Headlines: 618

This Week’s AI News Overview Written By Me, Not AI (Adios, conciseness!)

This is my 145th week of organizing links, and I’m 400 links away from hitting 60,000. For the week ending July 10th, I organized 618 links, and 110 of them went into the executive summaries. I’m 11 weeks behind because I took the summer off to be with my family.

But first… this week’s moments of touching grass.

I don’t have a ton of visual evidence of touching grass this week. However, in a discrete sign of aging, I realized that I’ve been casually wearing my reading glasses on top of my sunglasses. No one pointed this out, which means I am surrounding myself with the right people 🙂

Form follows function… and hey, if you avoid mirrors… you never have to deal with fashion choices.

I’m continuing to train for my attempt to climb the Grand Traverse Peak with my daughter Rori in August. I climbed 360 floors on the StairMaster with a 15-pound pack without using my hands. That’s 3,863 feet, which is still about 1,200 feet shy of what I’ll be doing with Rori, but at least it’s building my cardio base.

On to the AI news of the week!
I’ve organized the news by the major frontier labs first: Anthropic, Meta, and OpenAI. The rest of the news is organized by category or company. Many of them are straightforward, so I’m just giving the headline and a link.

At the bottom of this summary is the big pile of the 110 links that fill out the balance of the summaries. This will be a mix of conversation and headlines to stick to the point.

First… take a look at this chart!!!

Anthropic
A week after Anthropic accused Alibaba of the largest distillation theft in history, Alibaba banned all of their employees from using Claude.
https://techcrunch.com/2026/07/04/alibaba-reportedly-bans-employees-from-using-claude-code/

Anthropic named former Fed chair Ben Bernanke to its independent trust.
https://www.cnbc.com/2026/07/09/anthropic-fed-chair-bernanke-independent-trust.html

Claude Cowork is now integrated with the mobile and web apps. I haven’t figured out how to use this yet, but it’s a big development if you can basically use your laptop as a little server.
https://claude.com/blog/cowork-web-mobile/

Anthropic has been renting significant computing space from SpaceX, and Elon Musk just a few months ago was body slamming Anthropic with personal attacks and hyperbolic fear-mongering. But now he’s praising Mythos and Fable, and he promises not to shut off access.
https://techcrunch.com/2026/07/09/elon-musk-praises-mythos-fable-promises-not-to-cut-off-anthropic/

Ethan Mollick, who is one of my favorite voices in AI discussion and a professor at the Wharton School, has been using Fable as a beta tester and concluded that it’s such a strong model that the limits are not what the model can do, but our ability to think about edge cases to push the model hard enough, because we simply assume that things can’t be done. Our imagination is the limit.

Anthropic published a paper that describes an internal storage space (subconscious) within Claude’s neural network, where it can store thoughts and references. It’s easy to anthropomorphize this, but in reality it’s just mappings and data. Because it’s not a human brain, it’s worth reading this report and trying to understand the differences. Anthropic calls it a global workspace. They have a five-minute explanation video that is worth watching (it’s over the top, below)…. Just like our brains use the subconscious to store and process things while we’re not aware of them, there’s evidence that Claude is doing a similar process, but I’d frame it more like a parenthetical coding clause.
https://www.anthropic.com/research/global-workspace

As you might expect, this has created a flurry of consternation. And fear not, Anthropic created a wildly lofty video…

In an tangential project, Anthropic is researching ways to cut off certain areas of understanding within a model, especially knowledge that could be used for good or bad purposes. This is part of a larger discipline called alignment, where you try to mold a personality of an intelligent model… that’s as flexible as possible for varying use cases… while not allowing the model to become a bad actor or be manipulated into providing dangerous information. This one is worth a skim. It’s an insight into how alignment is starting to evolve beyond just a guiding system prompt, but almost lobotomizing models for certain capabilities.
https://www.anthropic.com/research/off-switch-dual-use

Anthropic is launching a dashboard called Reflect that allows you to look back on how you use Claude. I have not found it yet.
https://www.anthropic.com/news/reflect-with-claude

Example of reflecting with Claude

Anthropic signed a $19 billion, 20-year data lease with a company called TeraWulf. Twenty years seems incredibly long in technology time.
https://finance.yahoo.com/technology/ai/articles/terawulf-shares-surge-19b-anthropic-161300778.html

The most AI frontier-lab-coded link of the week is an obnoxious blog post by Anthropic called “Inviting Hard Questions,” with a weird, diffused picture of a flower and some rocks. It’s a mixture between Deep Thoughts with Jack Handey and Ready Player One. I’m not sure how to take it seriously… but they sure do.
https://www.anthropic.com/news/hard-questions

The signal-to-noise ratio is largely noise, but there is a little bit of signal. Anthropic praises themselves for creating the “Anthropic public record”, which is a survey that asked 52,000 Americans their biggest hopes and fears about AI. They did another survey of 81,000 Claude users in 160 countries, and they did a bunch of focus groups.

On top of that, they’re studying how Claude is used with anonymized real-world data, which we saw last week with the usage report. Further, they’ve created the Anthropic Institute, a research division dedicated to figuring out what the biggest challenges are going to be from AI.

And now this blog post is asking the general public to send Anthropic their “hardest questions”, and they’re going to track and report the actions they’re taking to those questions. It seems like all the worst parts of a bad presidential town hall, where audience members ask questions that are carefully screened, combined with staying up all night reading random internet comments.

Not only did Anthropic publish a huge vapid blog post with a huge vapid YouTube video, there’s also a huge vapid website dedicated to it too. It’s This American Life with none of the humor and all of the pomp.

Zero out of ten. Would not recommend.

Meta
Meta announced that they are going to build their own chip and start production in September as part of an effort to double their computing capacity.
https://techcrunch.com/2026/07/09/metas-new-ai-chips-will-begin-production-in-september/

Meta also announced Muse Spark 1.1. Muse Spark 1.0 was the first major release from the Superintelligence Lab, launched in April. Version 1.1 focuses on agentic and coding use, and Meta says it can compete with GPT-5.5 and Opus 4.8. I’m not sure I believe that, but I’m sure it’s a great model.
https://dev.meta.ai/
https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/

Along with Muse Spark, there is an image model and a video model as part of the family. I have not tried either one. Supposedly the image model is an agentic image creation tool that uses refinement instead of just diffusion, which I think is what GPT and Gemini do as well.
https://ai.meta.com/blog/introducing-muse-image-muse-video-msl/
https://about.fb.com/news/2026/07/introducing-muse-image-meta-ai/

Perhaps bigger than anything else… Meta claims that Muse Spark 1.1 is the top legal model and is state-of-the-art on Harvey’s legal bench. AND… the best at taxes and medical benchmarks. This is just too good for me to believe, but I will keep my eye out over the next few weeks to see if this sticks.

That’s one trouble with any of these releases lately. Influencers all over the place are claiming state-of-the-art in cherry-picked benchmarks. For example, in a benchmark from last week, which I think was OpenAI’s latest model, they compared it to Mythos, but OpenAI had six hours to work on the problem and Mythos only had two. Absolutely unfair.

OpenAI
OpenAI launched a Chrome extension for GPT and has enabled significant developments with browser use: opening tabs, downloading files, opening up documents, dashboards, logging into places. This has been a wonderful trend for me personally, because not everything has an API, and sometimes robots are blocked. Now I can just ask GPT or Chrome or Claude to open up my browser and work for me. It’s often still very slow, but it’s getting much better and worth trying if you haven’t in a while.

OpenAI has been accused of trying to conceal evidence and lie about its technical capabilities. OpenAI allegedly hid 78 million log files that could have been handed over in discovery. That’s egregious.
https://arstechnica.com/tech-policy/2026/07/openai-faked-inability-to-search-training-data-hid-billions-of-logs-nyt-says/

Every week OpenAI and Anthropic seem to find a way to make their products more confusing and fractured. OpenAI has merged the Codex app into GPT. You have to switch back and forth and you lose your history when changing, and it’s easy to forget which one you’re using, and there’s not a lot of continuity when you switch between the mobile app, the Chrome or web browser version, and desktop app. I like the concept of putting everything into one place, but they’ve got a long way to go with the interface.

It’s now been a week since GPT-5.6 Sol came out (to beta tester reviews), and the word on the street is still that the biggest selling point is that it’s close to Fable’s performance, but significantly cheaper, and the trade-off is worth it. A cheaper second place is often a better pound-for-pound solution than the expensive first place.
https://openai.com/index/gpt-5-6/

For some reason OpenAI thinks a corny agentic pet is compelling. I am absolutely allergic to it and have never tried it. But you could pick a little persona to be your friend while you use ChatGPT. To me this is like the cherry-flavored vape of user experience features.

Over the last few weeks ChatGPT has been introducing Sites, which are hosted websites you can build in ChatGPT and then have an actual URL you can share with people. There’s only one privacy toggle right now, and it’s either public or private. I like it because it’s easier to share your output with someone without having to give them an HTML file. I don’t like it because you’re limited by having it hosted on GPT. So the scalability seems stunted out of the gate.

A trend across all of the models is low-latency voice features with extended reasoning and the ability to search the web. I feel like every week OpenAI releases a new voice model and they name it something else. The latest from OpenAI is ChatGPT Live. There’s a pretty clever commercial, that goes with it. That’s worth watching. The latency is getting really quick.
https://openai.com/index/introducing-gpt-live/

I’ve talked a lot lately about diffused disposable interfaces (slop) as the future of UI/UX… and this video above is a great example.

AI 2030
In April 2025, five authors released a paper called AI 2027 that attempted to guess what the future would hold as AI evolved over the next two years. They correctly predicted that by the end of 2025 there would be rising agentic use and that the data centers would get enormous (not too hard, but they got the nuances too).

They predicted that coding automation would take off in 2026, which is most definitely accurate. Coding is now moving faster than anyone predicted.

They also nailed the next prediction, which is in 2026, China would start to push back on computing reliance on U.S. chips. Just this past month, we’ve seen Alibaba banning the use of Claude and China attempting to lean into Huawei’s chips instead of Nvidia’s.

In 2027, the paper predicts reinforcement learning and self-improvement.

The paper also predicts China will steal an AI agent from the USA. While we haven’t seen an agent stolen yet, we have seen a lot of distillation of U.S. models.

The paper predicted that by 2027, the government and the U.S. frontier models will be working hand in hand, and we’re ahead of schedule on that one. All of the labs (OpenAI, Google, SpaceX, and Anthropic) are working with the government. Ironically, Anthropic’s Mythos and Project Glasswing might be the most impactful, despite Anthropic’s reluctance to partner on weapons.

Now there’s a new release from that same group of authors called: AI 2040 Plan A.

This is the positive vision of what they think can happen. Unfortunately, their first paper, AI 2027, predicted that AI would take over the world and destroy humanity (forgot to mention that above, sorry)… and now the authors are publishing Plan A as a roadmap to avoid this.

This is very much worth reading, both the original 2027 paper and the new one. Even if you’ve read the original 2027 paper, it’s worth going back to.

Plan A requires making a deal with China to slow down growth. It’s a little bit like a nuclear treaty, with a lot of trust built in. The paper includes an unsuccessful branch of the timeline where China attempts to cheat.

I don’t see Plan A happening. With the nuclear arms race, there were seismographs and other ways to detect testing, which led to the Fast Fourier transform. There is a spectacular YouTube video on this. I don’t know how we would have a corresponding measure of tracking AI cheating.

General News and Observations
Florent Daudens, formerly of Hugging Face, thinks that GPT-5.6 is the best daily driver, whereas Fable is a little more adventurous.

Ethan Mollick thinks that SpaceX and Meta have finally started to catch up in what he calls the near frontier category. They’re so good at being cheap that they approach the Pareto frontier, where you have the most powerful model at the best price, blended and weighted for overall value.

Over the past few weeks, we’ve talked a lot about smaller models being able to do as much as the previous versions of the strong models at low cost. This is going to be, in my mind, what drives mass AI adoption because the cost is going to keep coming down and the quality of the output is so good that organizations simply will have to adopt it to keep up (like getting a cell phone, email address, or an internet connection is hard to avoid now). And I don’t mean AI for writing essays and things….I mean agentic use, corporate use, operational efficiencies, things that are financially viable and compelling with tangible results for law, finance, medicine, etc.

Ethan Mollick says that OpenAI and Anthropic are so good at getting these weaker models to perform that they will box out third parties trying to create startup models because the big lab’s smaller versions of the frontier are simply too easy to stamp out on a regular basis…

And finally, Mollick expresses the same confusion I have with the product names. There’s ChatGPT, ChatGPT Codex, ChatGPT Work, Claude, Claude Code, and Claude Cowork. It’s an absolute jumble for anyone who doesn’t use these tools regularly, and even for those who do.

ByteDance
ByteDance released Seedream 5.0
https://seed.bytedance.com/en/blog/beyond-generation-it-understands-design-introducing-seedream-5-0-pro

China
China is looking to curb overseas access to their top models, and facing continued export controls of chips, DeepSeek announced that they will start to make their own chips.

Cohere
Cohere released a transcription tool that is the leading state-of-the-art open source model for Arabic speech recognition.

DoorDash
DoorDash published a lengthy blog post about how they have built an agentic code reviewer internally. It’s interesting if you’re into how corporations are using agents.
https://careersatdoordash.com/blog/how-we-learned-to-trust-our-ai-code-reviewer-at-doordash/

Google
Ethan Mollick points out that Google is slacking a little bit, and it’s been a long time since they released a competitive frontier model. Over the past few weeks, I’ve pointed out that Google is excelling in the open source side, like Gemma, and the medium-sized models that are cheap, like Gemini 2.5 Flash.

Every now and then Google comes out with some ridiculously bad idea, and this week it is a video creation tool that takes your photos and turns them into corny pastels and watercolors. It seems like a novelty, and I never like to dunk on a product team directly, but I don’t see the value of this.
https://blog.google/products-and-platforms/products/photos/video-remix/

Legal News
In more serious news, there are two major legal AI developments. One is that Harvey partnered with Artificial Analysis to launch a legal agentic benchmark. This is a big deal, and I highly recommend following it.
https://x.com/ArtificialAnlys/status/2074541975186165887

And second, a startup that I’ve never heard of named Norm AI has just hit a 1.2 billion valuation.
https://x.com/johnjnay/status/2074485345593245833

Microsoft
Speaking of corporations using mid-sized models, Microsoft is swapping out some of the expensive OpenAI and Anthropic models for their own MAI model, which is a strong multimodal model that Microsoft’s been iterating on over the past few months.
https://www.bloomberg.com/news/articles/2026-07-07/microsoft-replaces-openai-anthropic-with-own-ai-in-some-apps

Microsoft also came out with a quirky product called Ode Poetry. Satya Nadella announced it personally. It’s an AI companion designed to lift your mood through real poems written by people, with voices powered by AI.

I don’t mind the concept. However, the interface features an AI voice model that’s supposed to sound like an anthropologist named William Seegart. He has this syrupy slow, overly empathetic voice that, to me, reeks of insincerity. Because you know you’re talking to a robot, but it’s got this better version of HAL from 2001: A Space Odyssey.

It asks you how you feel, and then you tell it how you feel, and then it picks a poem and reads it out loud to you. It has too much of a Thomas Kinkade vibe for my taste. I love poetry, but I’m not sure I need a robot meditation voice, like evil English Teddy Ruxpin.

Try it yourself: https://odepoetry.ai/

WorldModels
There was a breakthrough performance in world models this week.

World models take generative/diffusion AI to the fourth dimension. If you think of text as one dimension, and images two dimensions, video generation is three…. A world model is when you can look around and control space and time, but it’s being diffused as you move your view.

Rather than having a predetermined world that’s rendered fully, like a video game engine, this is completely diffused in real time. In the past, this would mean all sorts of hallucinations and lack of continuity…

Then Google came out with Genie 3, and that was, in my mind, the first real flagship frontier world model with solid memory.

Fei-Fei Li’s lab, World Labs, came out with Marble. That was a text-to-world model. NVIDIA has those as well.

This is a new one because it’s a video gaming diffused world model where four people can play at the same time and control a video game that they each see on their computer in real time. It was trained on data from 10,000 hours of Rocket League matches. The team partnered with Epic Games, who owns Rocket League, to build this specialized world model.
https://mira-wm.com/
https://mira-wm.com/blog-post/

To my knowledge, none of the authors of the paper or the creators of the model have any previous connections with frontier labs, other than maybe a Microsoft connection or two.

NVIDIA
NVIDIA released Audex, an open-source mixture of experts model that can handle speech, sound, and reasoning all at the same time. It’s on Hugging Face, and it’s a unified multimodal model that can understand audio. So it’s not just speech-to-text; it’s understanding the sound itself, and it can also create sound.
https://x.com/_weiping/status/2074537900172050704

NVIDIA has partnered with Hugging Face to integrate its Isaac platform into Hugging Face’s robotics program. Isaac and GR00T 1.7 are incredibly powerful open source models from NVIDIA. I’ve featured them both several times over the years.

NVIDIA and Hugging Face have created a development platform that unifies the different parts of the pipeline so that people can create more quickly. The Isaac pipeline is a five-stage deployment process…from first using Isaac to simulate an environment to test a robot, you can then create the coding data. Then you move over to GR00T to build the logic policies to reason with humans and turn the training into multitask behaviors. Then you go through policies and simulation before you deploy in the real world with something called Jetson Thor.

I think Jim Fan at NVIDIA is the absolute coolest robotics mind out there, even though Figure and Tesla and Unitree and Boston Dynamics get a lot of the buzz because they’re building the humanoids. NVIDIA is open sourcing the brains and the training and the simulations, and I am a huge fan of Dr. Jim Fan.
https://blogs.nvidia.com/blog/hugging-face-lerobot-models-frameworks-open-robotics/?ncid=so-twit-128581-vt48&linkId=100000429635659

Twitter
Grok has not historically been a very strong model. Maybe just as a conversationalist in your car, Grok is fine. That makes sense if you have a Tesla and you just need to use it as a Siri or Alexa proxy. But as far as actual benchmarking, Grok is nowhere near the pack.

That said, Grok 4.5 just came out and it is a strong model at a low cost.
https://x.ai/news/grok-4-5

This idea of the Pareto frontier, that is where you’re going to see the open source models and Grok really start to kick in.

Grok used to be open source, but nowit is not. Ever since Grok 2.5, they have not open sourced it.

SpaceX has so much computing power that they’ve been renting tremendous amounts to Anthropic and others. And now that everything is unified under SpaceX, I predict we’re going to see some really powerful output from Grok over the next few months.

This Week’s Top Stories and All The Links That Go With Them – Clean Batch for Skimming

Anthropic

Alibaba: Alibaba bans employees from using Anthropic’s Claude Code tool (a week after being accused of distillation)

Alibaba Bans Employees From Using Claude — The Information
https://www.theinformation.com/briefings/alibaba-bans-employees-using-claude

Alibaba reportedly bans employees from using Claude Code | TechCrunch
https://techcrunch.com/2026/07/04/alibaba-reportedly-bans-employees-from-using-claude-code/

Bernanke: Anthropic adds Ben Bernanke to its long-term benefit trust

Anthropic names former Fed Chair Bernanke to its independent trust
https://www.cnbc.com/2026/07/09/anthropic-fed-chair-bernanke-independent-trust.html

Cowork: Claude Cowork works on mobile and web now

Claude Cowork is coming to mobile and web. Hand Claude a task at your desk and pick up the finished work from your phone. Close the laptop and Claude keeps going. Beta is rolling out over the next several weeks starting with the Max plan, with more plans to follow.”
https://x.com/claudeai/status/2074525815820169320

Claude Cowork on web and mobile: hand off work anywhere | Claude by Anthropic
https://claude.com/blog/cowork-web-mobile/

Elon Praise: Musk promises not to cut off Anthropic from SpaceX servers

Elon Musk praises Mythos/Fable, promises not to ‘cut off’ Anthropic | TechCrunch
https://techcrunch.com/2026/07/09/elon-musk-praises-mythos-fable-promises-not-to-cut-off-anthropic/

Fable: Ethan Mollick says the limit is us. Not Fable.

I can basically guarantee you are not being ambitious enough with the work you are assigning Fable Start asking for the maximum possible thing to figure out the far edges of what it can do. After that, you can decide where the system reaches it limits & revise requests downward”
https://x.com/emollick/status/2074137233607590328

Mechanistic Interpretability: Anthropic finds a hidden storage space inside Claude’s neural network – The J-Space

Interesting stuff. And the visualization at the end is worth trying:”
https://x.com/emollick/status/2074204151521648766

Anthropic research suggests that modern LLMs have access consciousness. Fascinating test with the J-space! We don’t yet have a convincing test for phenomenal consciousness, which is what most people intuitively understand consciousness to be.”
https://x.com/BorisMPower/status/2074201312531734567

Anthropic researchers found something unusual inside Claude. A small internal workspace that the model uses while solving certain problems. They call it the J-space, named after the Jacobian method they used to discover it. The J-space isn’t text. It’s not Claude’s”
https://x.com/LiorOnAI/status/2074198891990548940

Must-read research by Anthropic. Here is the simple explanation and why this is a big deal. We suspect LLMs perform “internal reasoning”. But little is known or do good methods exist to understand it. Anthropic claims that J-Space (which differs from chain-of-thought or”
https://x.com/omarsar0/status/2074264122330612223

The J-space lets us read, audit, and shape what Claude is actively thinking about—useful tools for keeping models trustworthy as they grow more capable. And it suggests surprising parallels between language models and our own minds. Read the full paper:”
https://x.com/AnthropicAI/status/2074185387577094398

A global workspace in language models \ Anthropic
https://www.anthropic.com/research/global-workspace

New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.”
https://x.com/AnthropicAI/status/2074185348142280912

Mechanistic Interpretability II: Anthropic is trying to modulize Claude so elements can be labotomized

An off switch for dual use knowledge in AI models \ Anthropic
https://www.anthropic.com/research/off-switch-dual-use

Open: Anthropic keeps OpenClaw in dark about lawsuit

We’re a very large customer of Anthropic and they still have yet to tell us about the lawsuit. I learned about it from a reporter, not our “partner.”
https://x.com/steipete/status/2074739318103629979

Reflect: Anthropic launches dashboard to track and reflect on Claude usage (I can’t find it yet)

A new way to reflect on how you use Claude \ Anthropic
https://www.anthropic.com/news/reflect-with-claude

TeraWulf: Anthropic signs $19 billion, 20-year data center lease with TeraWulf

TeraWulf shares surge on $19B Anthropic AI infrastructure lease deal
https://finance.yahoo.com/technology/ai/articles/terawulf-shares-surge-19b-anthropic-161300778.html

Vapid Blog Posts: In the most vapid obnoxious post of the year, Anthropic launches public initiative soliciting hard questions

Inviting hard questions \ Anthropic
https://www.anthropic.com/news/hard-questions

Meta

Chip: Meta to start producing in-house Iris AI chip in September

EXCLUSIVE: EXCLUSIVE Meta to put AI chip into production in September as it looks to double computing capacity, memo shows | Reuters
https://www.reuters.com/world/asia-pacific/meta-put-ai-chip-into-production-september-it-looks-double-computing-capacity-2026-07-09/

Meta’s new AI chips will begin production in September | TechCrunch
https://techcrunch.com/2026/07/09/metas-new-ai-chips-will-begin-production-in-september/

Coding: Can it be true: Meta’s new Muse Spark 1.1 rivals GPT-5.5 and Opus 4.8 on coding (doubt it without a harness)

Meta is releasing an agentic and coding model that can compete with GPT-5.5 and Opus 4.8? And on top of that, it is SOTA on HLE? Holy, that was not on my bingo card. Alexandr Wang cooked. What is happening today!”
https://x.com/kimmonismus/status/2075232528726708245

Muse Spark 1.1 just launched and it’s their most capable coding agent model yet. On Terminal-Bench 2.1 it scores 80.0%, in the same cluster as Opus 4.8 (82.7%) and GPT 5.5 (83.4%). Use it in Cline with the Meta API!”
https://x.com/cline/status/2075271057326719152

AI developers – we’ve published a technical guide covering how to get started with Muse Spark on Meta Model API. Muse Spark is a multimodal reasoning model built for agentic tasks, coding, computer use and long-context workflows. See what you can build 👉
https://x.com/MetaforDevs/status/2075268072022401526

Image and Video: Meta launches Muse Image and Muse Video

Introducing Muse Image: Image Generation Built for Your World
https://about.fb.com/news/2026/07/introducing-muse-image-meta-ai/

Muse Image works as an agent rather than a direct prompt-to-image model: it invokes tools, self-refines, improves with scaled test-time compute, and pairs with Muse Spark for collaborative media generation. 🧵👇”
https://x.com/AIatMeta/status/2074587864923250873

We’re launching Muse Image today!! 🎉 It’s an agentic image gen model that plans, writes code and uses search tools, and refines its own outputs in chain-of-thought. Image performance improves as we scale test-time compute with higher reasoning.
https://x.com/_tim_brooks/status/2074578008296628698

1/ releasing muse image today — the first image generation model from MSL. it’s agentic: pairs with muse spark to reason through your prompt, search the web, and plan before it generates. people get what they meant on the first try. live now in the Meta AI app.”
https://x.com/alexandr_wang/status/2074555909347369105

Introducing Muse Image and Muse Video, the first media generation models developed by Meta Superintelligence Labs. Muse Image is our most advanced image generation model yet. It follows instructions faithfully, edits with precision, composes from multiple references, and draws”
https://x.com/AIatMeta/status/2074577662840832382

Meta Muse Video just entered the Video Arena at #3. @AIatMeta’s new video model scored 1459 in the Text-to-Video Arena. It outperforms Alibaba’s HappyHorse 1.0 by +30pts and ranks ahead of Grok Imagine, Sora 2 Pro and Google Veo-3.1 models. Meta has now reached the video AI”
https://x.com/arena/status/2074591193783320851

1/ today, with the launch of muse image and the preview of muse video, we wanted to share more samples from the model alongside research details and eval results see some more samples from muse video below (and see thread for research blog!)”
https://x.com/alexandr_wang/status/2074579670712934428

Introducing Muse Image and Muse Video
https://ai.meta.com/blog/introducing-muse-image-muse-video-msl/

Meta launches Muse Image for AI image generation in its apps
https://www.testingcatalog.com/meta-launches-muse-image-across-its-apps-and-previews-muse-video/

Legal bench: Can it be true: Muse Spark 1.1 tops legal, tax, and medical AI benchmarks? (doubt it)

Muse Spark 1.1 is SOTA on Harvey’s Legal Bench, TaxEval, and MedScribe. It’s cool to see that our model outperforms even Fable in a few areas :)”
https://x.com/alexandr_wang/status/2075233663323947120

🏆 tell your local lawyer to try out muse spark 1.1″
https://x.com/alexandr_wang/status/2075270855832273137

Spark 1.1: Muse Spark 1.1

Today we are launching Muse Spark 1.1, an upgrade to muse spark 1 that greatly improves agentic, coding, multimodal, and computer use capabilities. We’re also launching the Meta Model API in public preview.
https://x.com/shengjia_zhao/status/2075220782465290620

Excited to share what we’ve been building at Meta Superintelligence Labs! Today we’re launching Muse Spark 1.1, our strongest model yet for complex agentic workflows — delivering massive gains in agents, computer use, coding, multimodal reasoning, and multi-agent orchestration.”
https://x.com/ren_hongyu/status/2075224643829711101

Introducing Muse Spark 1.1
https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/

OpenAI

Browser Use: ChatGPT browser use and Chrome extension

The new Chrome extension lets ChatGPT work alongside you on any task in the browser. Start new threads or resume existing ones from the extension, with access to ChatGPT desktop context like local files, projects, plugins, skills, and more. Install the ChatGPT extension or”
https://x.com/OpenAIDevs/status/2075276009902112976

The in-app browser is getting better at helping ChatGPT work across the web, with support for authenticated sites, multiple tabs, and file downloads. Open a document, dashboard, or Site and collaborate with ChatGPT with annotation mode, without leaving the desktop app.”
https://x.com/OpenAIDevs/status/2075292716737736919

Copyright: OpenAI accused of hiding ChatGPT logs in NYT copyright case

OpenAI may have made a fatal misstep in copyright fight with news orgs – Ars Technica
https://arstechnica.com/tech-policy/2026/07/openai-faked-inability-to-search-training-data-hid-billions-of-logs-nyt-says/

Federated Codex: OpenAI merges Codex and ChatGPT into confusing desktop app UX

We’re bringing Codex and ChatGPT together in one desktop app. The same powerful coding agent, now alongside ChatGPT Work, with new coding workflows, a new Chrome extension, revamped in-app browser, and faster Computer Use powered by GPT-5.6.”
https://x.com/OpenAIDevs/status/2075275868268789885

Today, Codex comes into ChatGPT with its own dedicated space, right next to the new Work agent, plus a lot more for developers! GPT-5.6 with Ultra and subagents for harder tasks, faster computer use, inline diff editing, PR review, and Sites to quickly deploy full-stack apps.”
https://x.com/romainhuet/status/2075286364476850430

GPT 5.6: OpenAI’s GPT-5.6 Sol cheaper than Fable 5 plus early feedback

OpenAI’s most anticipated model of the year is here, and it ranks #2 on Vals Index and Vals Multimodal Index. Although Fable 5 is still ahead on several benchmarks, GPT 5.6 is clearly in the same class, and is able to complete tasks like our CyberBench that Fable refuses.”
https://x.com/ValsAI/status/2075270642359029972

GPT-5.6 sets a new Terminal-Bench record at 91.9%. Priced the same as GPT 5.5 at $5/$30 per million tokens. With Fable getting pulled from Claude subscriptions and moving to API cost (~2x as expensive at $10/$50), we are glad to see @OpenAI’s commitment to accessibility.”
https://x.com/cline/status/2075278343927365991

GPT-5.6 Sol (max) scores similarly to Claude Fable 5 (max) in GDPval-AA v2, reflecting a similar ability to complete economically valuable tasks.”
https://x.com/ArtificialAnlys/status/2075268987550932998

GPT-5.6 Sol comes close second to Claude Fable 5 in the Artificial Analysis Intelligence Index at one third of the cost, and leads the Artificial Analysis Coding Agent Index in OpenAI’s Codex harness We supported @OpenAI with pre-release evaluation of GPT-5.6 Sol, Terra, and”
https://x.com/ArtificialAnlys/status/2075268970492657905

OpenAI just released GPT-5.6. But, instead of competing on benchmarks, they’re competing on cost curves. GPT-5.6 is doing something every frontier lab has been chasing: getting more work out of every token. It beats Fable 5 on coding-agent benchmarks while using less than half”
https://x.com/LiorOnAI/status/2075277748394967122

I was an early tester of GPT-5.6 Sol. I was asked to not share demos until after launch but it is a very good model. It is of similar ability, but quite different feel, than Fable. Fable wants to go off and do work on its own pace, Sol is faster but works with you in steps more.”
https://x.com/emollick/status/2074712677755035907

GPT-5.6 – ARC-AGI Results
https://arcprize.org/results/openai-gpt-5-6

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game It is the best model at orienting in a situation it’s never encountered”
https://x.com/arcprize/status/2075270869992264003

While scores at the top of ARC-AGI v1/2 are topping out at 95% there is still a lot the benchmark tell us about efficiency Not only does GPT-5.6 Sol get 92.5% on ARC-AGI-2 (sota) but it does so at 1 OOM less cost than GPT-5.5 Pro(!) GPT-5.5 Pro came out 3 months ago…”
https://x.com/GregKamradt/status/2075274981794300113

OpenAI says GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna”
https://x.com/scaling01/status/2075269113488789984

Quote from OpenAI on the livestream. ‘Already, Sol has been transforming our research program. As one example, GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna.’”
https://x.com/AndrewCurran_/status/2075273510331785658

GPT-5.6: Frontier intelligence that scales with your ambition | OpenAI
https://openai.com/index/gpt-5-6/

Not enough people are emotionally prepared for GPT-6″
https://x.com/scaling01/status/2075276735650648258

Previewing GPT-5.6 Sol: a next-generation model | OpenAI
https://openai.com/index/previewing-gpt-5-6-sol/

obviously the best model we have ever produced, but also one of the best blog posts we have ever produced:”
https://x.com/sama/status/2075266471316615436

Pets are CORNY: Pets are CORNY

There was a reason why I highlighted it, and I didn’t see it mentioned in other coverage. I was testing Pets, and here is what I discovered. It was a quick but telling test: Different Pets produced utterly different results. The idea, structure, language, what to address, and”
https://x.com/TheTuringPost/status/2075411271596335396

Sites: ChatGPT Sites

This was such a fun ship to be part of: ChatGPT Sites is now available on all paid plans. Give it an idea and it’ll design the site, wire up a database, connect your sources, and even handle auth. Keep building in ChatGPT, then publish a real, shareable URL. ChatGPT can just”
https://x.com/simpsoka/status/2075278935366287842

Sites is now in beta for Pro and Plus users, with EU and UK availability coming soon. During the beta, your plan includes Sites usage at no additional cost, subject to plan-specific limits.”
https://x.com/OpenAIDevs/status/2075337081304522853

Sites is now available in beta to Pro and Plus users. Build interactive apps with GPT‑5.6 and deploy them using Sites, with hosting, storage, and optional auth built in.”
https://x.com/OpenAIDevs/status/2075275892591591469

Voice: GPT-Live is better reasoning low latency voice

For questions that require web search, deeper reasoning, or more complex work, GPT-Live can delegate to our latest frontier model behind the scenes, and brings the result back into the conversation when it’s ready.”
https://x.com/OpenAI/status/2074907033577693636

We’ve reduced p95 latency by at least 25% across Realtime voice models through improved caching.”
https://x.com/OpenAIDevs/status/2074255420831735824

GPT-live (next-generation voice) launches today in ChatGPT. it feels magical and ‘real’. i have always preferred typing to talking to an AI, now i think that’s going to shift.”
https://x.com/sama/status/2074909079450050629

Introducing GPT-Live, a new generation of voice models for natural human-AI interaction. Rolling out in ChatGPT starting today. You’ll want to turn the sound on for this one.”
https://x.com/OpenAI/status/2074907025537224840

GPT Live is seriously good. GPT Live with interactive widgets is next level. After a few weeks with OpenAI’s new voice model, I turned voice on by default and there’s no going back.”
https://x.com/fdaudens/status/2075011510045233457

Introducing GPT-Live | OpenAI
https://openai.com/index/introducing-gpt-live/

what a good video”
https://x.com/sama/status/2075068286107316317

Work

Introducing ChatGPT Work, a new agent in ChatGPT powered by Codex and GPT-5.6. It can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work. It’s a whole new way to get work done.”
https://x.com/OpenAI/status/2075274271845404744

AGI

AI 2040 Paper: Will we get a 100 trillion parameter model by 2028?

AI 2040 also predicts 100T models in ~2028″
https://x.com/scaling01/status/2075296890325712944

Benchmarks

GPT 5.6 v. Fable: GPT 5.6 v. Fable

So who wins between GPT-5.6 Sol and Fable? After a few weeks of early access, GPT-5.6 is my daily driver. Fable feels more adventurous, focused and weirdly narrower. Sol feels more broadly useful and circumspect: closer to the brief, more reliable. Demo
https://x.com/fdaudens/status/2075285983054963081

Grok/Muse: Grok/Muse

It seems like SpaceX/Grok and Meta/Muse have started to keep pace in the near-frontier category while also introducing a new category of cheap, fast & closed specialized coding models. Both were tied with the Big Three at some point, fell behind, but may have started to return.”
https://x.com/emollick/status/2075254468912812206

Naming is Insane: Naming is Insane

I am *so confused* by ChatGPT v. ChatGPT Codex v. ChatGPT Work v. Claude v. Claude Code v. Claude Cowork right now!”
https://x.com/simonw/status/2075348941215006888

PotatoMonetBench: PotatoMonetBench

Big gains in PotatoMonetBench from Opus 4.5 to Fable: “There are two controls. One is slider, in which one side is labelled Maximum Potato and the other is Formalware, there are four positions. The other is a dial that goes from Monet to Drive-Thru.”
https://x.com/emollick/status/2074227397398794536

Smaller Models Gaining Ground: Smaller Models Gaining Ground

Anthropic and OpenAI will increasingly be able to serve weaker versions of frontier AI (Haiku, Instant) at very low cost And owning the full range of intelligence means they might provide “organizations of models” delivering lower cost/higher performance than any 3rd party could”
https://x.com/emollick/status/2074201073795874969

ByteDance

Seedance: Seedance

Quick test w/ Seedance. Video-to-video AI models are getting really good.”
https://x.com/bilawalsidhu/status/2074193958549275040

Seedream 5.0: Seedream 5.0

Another powerful new AI image model has arrived! ByteDance released Seedream 5.0 Pro, a model that claims to go beyond generation to “understand design”. Features include: – Improvements in text rendering, structure, alignment, and other design areas for strong infographics”
https://x.com/TheRundownAI/status/2074869830293786634

ByteDance debuts Seedream 5.0 Pro with advanced reasoning
https://www.testingcatalog.com/bytedance-debuts-seedream-5-0-pro-with-advanced-reasoning/

Seed News – ByteDance Seed Team
https://seed.bytedance.com/en/blog/beyond-generation-it-understands-design-introducing-seedream-5-0-pro

China

China Blocking Access to Top Models?: China Blocking Access to Top Models?

EXCLUSIVE: Beijing is looking at curbing overseas access to China’s top AI models, sources say | Reuters
https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/

DeepSeek Chips: DeepSeek Chips

Facing US export controls, China’s DeepSeek plans to make its own chips – Ars Technica
https://arstechnica.com/ai/2026/07/facing-us-export-controls-chinas-deepseek-plans-to-make-its-own-chips/

Cohere

Arabic Transcription: Arabic Transcription

We’ve built Cohere Transcribe Arabic, the world’s most accurate open-source model for Arabic speech recognition. Available under Apache 2.0″
https://x.com/cohere/status/2074499759616729149

DoorDash

Code Adoption: Code Adoption

How we learned to trust our AI code reviewer at DoorDash – DoorDash
https://careersatdoordash.com/blog/how-we-learned-to-trust-our-ai-code-reviewer-at-doordash/

Google

Falling behind: Is Google slacking?

Pretty long gap since the last dot on the green line…”
https://x.com/emollick/status/2075068836856913928

Photos Remix: Google Photos remix seems horrible

Create videos in seconds with Google Photos’ Video Remix
https://blog.google/products-and-platforms/products/photos/video-remix/

Law

Harvey: Harvey and Artificial Analysis launch legal agent benchmark

After our announcement last month, Artificial Analysis is now launching Harvey LAB-AA (Legal Agent Benchmark), our implementation of Harvey’s new agentic legal benchmark that evaluates language models on real-world legal work across 24 practice areas Harvey LAB-AA tests models”
https://x.com/ArtificialAnlys/status/2074541975186165887

Norm AI: Legal start-up Norm AI hits $1.2B valuation

We just raised a $120 million Series C at a $1.2 billion valuation, led by @khoslaventures, to pursue the full-stack approach to legal AI. Blackstone, Bain Capital, Craft Ventures, Coatue, Vanguard, New York Life, TIAA, Tony James, and Jeff Hammes (former Chairman of Kirkland &”
https://x.com/johnjnay/status/2074485345593245833

Microsoft

MAI: Microsoft is swapping OpenAI and Anthropic for MAI

Microsoft Replaces OpenAI, Anthropic With Own AI in Some Apps – Bloomberg
https://www.bloomberg.com/news/articles/2026-07-07/microsoft-replaces-openai-anthropic-with-own-ai-in-some-apps

Microsoft Ode: Ode is on demand humanities readings from Microsoft

Ode
https://odepoetry.ai/

Mira

Playable World Model: MIRA – A Playable World Model

I have been playing these world models/games since the first diffusion DOOM, and multiplayer at 20 FPS is pretty neat. “A dream of rocket league,” indeed.”
https://x.com/emollick/status/2074348274136346871

MIRA – Blog post
https://mira-wm.com/blog-post/

Introducing MIRA. A playable, multiplayer world model. A dream of Rocket League. Trained on 10k hours of data collected with publicly available bots, MIRA learns the dynamics of a four-player game. The model runs in real time at 20 fps, based on the keys you and the other”
https://x.com/gen_intuition/status/2074104524596457706

NVIDIA

Audex: NVIDIA Audex – an open audio-text model on Hugging Face

NVIDIA just released Audex on Hugging Face A unified audio-text MoE that handles speech, sound, and reasoning without regressing on text intelligence. 30B params. 3B active. 1M context.”
https://x.com/HuggingPapers/status/2074384562952749254

🚀 Introducing Audex, a unified audio-text LLM for text, speech, sound, and music 🚀 🏆 Audex delivers best-in-class performance among open models across: 🎧 Audio understanding 🗣️ Speech recognition and translation 🔊 Text-to-speech 🎵 General audio generation 🔄”
https://x.com/_weiping/status/2074537900172050704

Robots

Neo: Neo Fancy Hand

A stunning piece of engineering!! 1X unveils the new humanoid hand for NEO – 25 degrees of freedom: 22 fully actuated in the fingers and palm, plus 3 at the wrist. – The DoF are distributed anatomically rather than evenly, deliberately biased toward a thumb that genuinely”
https://x.com/TheHumanoidHub/status/2075264747419869294

Human, yet superhuman! 1X NEO with new-generation hands”
https://x.com/TheHumanoidHub/status/2075267374866186354

NVIDIA Huggingface: Nvidia Isaac joins Hugging Face’s LeRobot

19 million developers can now do more with frontier physical AI tools. 🤗 We’re expanding open robotics with @HuggingFace, bringing GR00T 1.7, an open VLA model for humanoid robots and Isaac Teleop directly into LeRobot. Read the blog ➡️ https://t.co/ElPM8aiYN6 #MACHINA2026″
https://x.com/NVIDIARobotics/status/2074380795855147072

NVIDIA and Hugging Face are bringing NVIDIA’s Isaac stack to LeRobot, Hugging Face’s open robotics library. Available now: – Isaac GR00T 1.7, a VLA foundation model for humanoids, claimed to be the first open and commercially viable one – Isaac Teleop, an open framework for”
https://x.com/TheHumanoidHub/status/2074866213818429882

Open-source robotics just leveled up. 🦾 @HuggingFace’s LeRobot 0.6 now includes NVIDIA Isaac GR00T 1.7 and Isaac Teleop, bringing the latest Isaac models and frameworks to the open robotics community. This step-by-step guide walks you through installation, data collection,”
https://x.com/NVIDIARobotics/status/2074390485251113317

SpaceX

Grok 4.5: Grok 4.5

Grok 4.5 from @SpaceXAI is live on OpenClaw. No OpenClaw update required, just connect your X Premium or SuperGrok subscription, select Grok 4.5 under the xAI provider, and use an Opus-class model that’s fast, low cost, and ready for agentic work.”
https://x.com/openclaw/status/2074973471977955556

SpaceXAI just released Grok 4.5, and it ranks #4 on GDPval-AA v2 with an Elo of 1543 – behind only the latest Claude releases from Anthropic on real-world agentic knowledge work tasks Grok 4.5 achieved this score at a cost of $0.49 per GDPval task to sit clearly on the Pareto”
https://x.com/ArtificialAnlys/status/2074942097158021371

Grok 4.5 is not quite at the level of Fable 5 or GPT-5.5, but it doesn’t need to be if you are using a combination of models in your agent orchestrator. 2x efficiency compared to leading models. Competitive with costs, too. Not bad. Definitely worth a look.”
https://x.com/omarsar0/status/2074921779718443155

xAI has just released Grok 4.5, beating Opus 4.8 on Terminal-Bench. Notably, Grok 4.5 is served at fast speeds w/ 2x greater token efficiency than other top models, while being priced ~5x cheaper than GPT. Meaning it delivers the highest intelligence per unit of time and cost.”
https://x.com/cline/status/2074928307427275108

Grok 4.5 is a genuine surprise success. Not only is it now playing in the same league as Claude and OpenAI’s GPT, it is also significantly cheaper. More token-efficient, less expensive to use, and still delivering outstanding performance: With this triad, it is a true”
https://x.com/kimmonismus/status/2074956721919869186

Introducing Grok 4.5 | SpaceXAI
https://x.ai/news/grok-4-5

it’s not a codemonkey, it’s a legitimate frontier model xAI is back in the game. What the hell does Cursor know?!”
https://x.com/teortaxesTex/status/2074923834075988419

We’ve partnered with SpaceXAI to train Grok 4.5. It’s our most powerful model yet and the first we’ve built for more than software engineering.”
https://x.com/cursor_ai/status/2074915744999969059

xAI is SpaceXAI: xAI is SpaceXAI

xAI Is Dead. Long Live SpaceXAI
https://gizmodo.com/xai-is-dead-long-live-spacexai-2000782034

We are now @SpaceXAI.”
https://x.com/SpaceXAI/status/2074214064746832060?s=20

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading