About This Week’s Covers and Humanities Readings
For this week’s humanities reading, I chose Marjane Satrapi. Marjane died on June 4, 2026 at 56 years old. She was a writer, graphic artist, filmmaker, political humanist, and illustrator of her own books. Her most famous book is an autobiographical graphic novel called Persepolis. It is a personal account of the Iranian Revolution, adolescence, repression, exile, and cultural displacement.
For the cover theme, I went with Julio Le Parc. Julio’s style is one that I’ve seen my entire life, especially in the ’80s. Almost every TV show and graphic I can remember from my childhood in the ’80s feels derivative of his work.
It’s a bit fitting in the age of AI, that I don’t necessarily believe that every image I saw as a kid that looked like a Julio Le Parc was actually by Julio.




His style in the late ’50s was known as Surfaces Sequences, which he said created stacked illusions of rhythmic movements, as if they were accumulated and documented along the way.
It’s interesting that almost everything I think of with this style comes back to some sort of computer art rendering from the ’80s. A little bit of symmetry here as we move into the AI world.
His style is called kinetic art. From HBO to Atari to Polaroid… even the Apple IIe, or that blocky font like the title from Logan’s Run. If you Google “’70s retro stripe logos,” it’s all Julio-influenced. Birdwell Beach Britches, striped conversion vans from the ’70s. It’s everywhere. Ocean Pacific’s logo, you name it.









Le Parc was born in Argentina and moved to Paris in 1958. Le Parc was kicked out of France in 1968 after participating in the Atelier Populaire. Students took over the art school and ran it as a collective until they were shut down by police about a month after they started. A few of the taglines from Atelier Populaire were “Beauty is in the street” and “Be young and shut up.”





Le Parc made it back to France and died in Paris in June 2026. Julio seems like a pretty cool dude who lived an incredible life.
Here are my favorite category covers of the week:









The humanities readings come from Marjane Satrapi. Marjane was born in November 1969, and she died on June 4 this year. She was nine years old during the 1979 Iranian Revolution, and lot of her family and friends were either arrested or killed.
Her uncle was a political prisoner exiled to the Soviet Union, then brought back and executed in Iran. Marjane was the only visitor allowed in 1982 before he was executed. She was only thirteen (!) at the time, and it impacted the rest of her life.

Satrapi pushed the boundaries in Iran and disregarded the modesty code. She bought and played banned music. When she was 14, her parents helped her move to Austria. After college, she returned to Iran and got a master’s degree.
Satrapi was most famous for her autobiographical comic books, Persepolis and Persepolis 2. Persepolis was adapted into a film, with the voices of Iggy Pop, Sean Penn, and Gena Rowlands. She was brave and lived a full and authentic life.
Humanities Reading of the Week
“The regime had understood that one person leaving her house while asking herself:
Are my trousers long enough?
Is my veil in place?
Can my make-up be seen?
Are they going to whip me?No longer asks herself:
Where is my freedom of thought?
Where is my freedom of speech?
My life, is it liveable?
What’s going on in the political prisons?”
― Marjane Satrapi, The Complete Persepolis“In life you’ll meet a lot of jerks. If they hurt you, tell yourself that it’s because they’re stupid. That will help keep you from reacting to their cruelty. Because there is nothing worse than bitterness and vengeance… Always keep your dignity and be true to yourself.”
― Marjane Satrapi, Persepolis: The Story of a Childhood“It’s fear that makes us lose our conscience. It’s also what transforms us into cowards.”
― Marjane Satrapi, The Complete Persepolis“In every religion, you find the same extremists.”
― Marjane Satrapi, The Complete Persepolis“Once again, I arrived at my usual conclusion: one must educate oneself.”
― Marjane Satrapi, Persepolis: The Story of a Childhood“I have always thought that if women’s hair posed so many problems, God would certainly have made us bald.”
― Marjane Satrapi, The Complete Persepolis
This Week By The Numbers
Total Organized Headlines: 547
- AGI: 1 story
- AI Inn of Court: 9 stories
- Agents and Copilots: 249 stories
- Alibaba: 5 stories
- Alignment: 34 stories
- Amazon: 4 stories
- Anthropic: 47 stories
- Apple: 2 stories
- Audio: 24 stories
- Augmented Reality (AR/VR): 32 stories
- Benchmarks: 45 stories
- Business and Enterprise: 34 stories
- Chips and Hardware: 71 stories
- DeepSeek: 1 story
- Education: 6 stories
- Ethics/Legal/Security: 41 stories
- Figure: 5 stories
- Google: 35 stories
- HuggingFace: 32 stories
- Images: 59 stories
- International: 13 stories
- Internet: 1 story
- Law: 6 stories
- Locally Run: 15 stories
- Meta: 3 stories
- Microsoft: 85 stories
- Mistral: 1 story
- Multimodal: 46 stories
- NVIDIA: 63 stories
- Nous Research: 9 stories
- Open Source: 74 stories
- OpenAI: 66 stories
- OpenClaw: 18 stories
- Perplexity: 6 stories
- Podcasts/YouTube: 8 stories
- Publishing: 7 stories
- Qwen: 4 stories
- RAG: 3 stories
- Robotics Embodiment: 62 stories
- Science and Medicine: 29 stories
- Security: 8 stories
- Technical and Dev: 108 stories
- Video: 37 stories
- World Models: 30 stories
- X: 10 stories
This Week’s Executive Summaries
This week, I organized 547 links into about 60 categories. Seventy links went into the executive summaries and top stories.
I’m still seven weeks behind because I’m enjoying the summer with my family. As of this publishing, we have 11 days left for us to be together before our oldest heads off to college.



I organized everything this week based on what I want to see when I go back and read these. Those are truly my top stories. They’re not always the stories that make the news, and they’re not always the most product-oriented.
I’m going to summarize the 34 biggest stories in the order that I want to remember them. There are 55 top stories total below.
Anthropic expands closed AI security consoritum Project Glasswing to 150 infrastructure partners
The top story this week continues from the top story from last week. Project Glasswing is a consortium of partners that Anthropic gathered to carefully discover vulnerabilities and security risks.
Last week, Anthropic published the findings of its initial 50 partners. The consortium identified over 10,000 critical severity vulnerabilities!
This week, Anthropic announced that it was going to expand Project Glasswing to 150 new organizations. Here are the three quotes you need to know:
“What each partner has in common is that a successful attack on their codebase could be catastrophic. For most partners, we estimate that a major attack could affect more than 100 million people, with important ramifications for both global and national security.”
“Mythos Preview continues a long-term trend that we’ve been warning about for some time: within 6 to 12 months, we expect that many other AI companies will have Mythos-class models, and they could release them without safeguards that prevent misuse. In that world, cyberattacks could occur much more often, and in much more unpredictable forms. It’s imperative that cyberdefenders [adapt to maintain pace]”
“We’re working as quickly as we can to safely release Mythos-level capabilities in general access. To do so, we’ll need highly robust safeguards that prevent the model’s cyber capabilities from being misused—safeguards that we (and, to our knowledge, all other AI developers) have yet to develop.”
Theme One: POWER
The theme for the first set of top stories is just the sheer power of AI. People seem to be looking closely at whether or not AI is a financial bubble. The internet had a bubble, yet the internet did not go away.
So, what I’d like to talk about this week is that the power of AI is underestimated by almost everyone who doesn’t use it often.
Claude now writes 80% of Anthropic’s code, hinting at self-improvement
Anthropic put out a statement called “When AI Builds Itself.” Anthropic sees recursive self-improvement within the near future… sooner than most people expect.

Anthropic engineers ship 8x more code per quarter than before
Not only is Anthropic sharing that Claude writes 80% of its own code, but Anthropic engineers are now shipping eight times as much code per quarter as they did during the previous five years.

Claude Mythos hits forecasters’ year-end AI task-length target by May
In early May, the best superforecasters predicted that by the end of the year, the METR 80% task horizon would reach three to four hours.
In plain English, that means that if a task takes a human three to four hours to complete, the METR benchmark looks at what agentic AI can do with an 80% task success rate.
So, if it takes a human expert four hours to complete a task, Claude would have to finish that same task with an 80% success rate to achieve the benchmark.
Notably, earlier in the year, 50% was the benchmark that most people were using. Now, the success threshold has already gone up to 80%.
In May, the superforecasters predicted that the three-to-four-hour 80% task benchmark would be achieved by the end of 2026. By the end of that same month, Claude Mythos had already achieved it.

Claude beats human researchers on critical next-step decisions 64% of the time
Anthropic gathered hundreds of coding and research in which a human took a wrong turn. Anthropic then showed Claude each session up until the critical error moment and asked what it thought should be done next. Mythos was able to improve on the human’s choices 64% of the time, up from 22% in 2024.

Mustafa Suleyman, CEO of Microsoft AI, predicts 1000x compute growth by 2029
Mustafa Suleyman, CEO of Microsoft, predicts that AI computing power will grow 1,000 times in the next three years.
Gemini 2.5 beats law professors 75% in office hours test
In Google news, Gemini 2.5 was able to beat law professors 75% of the time when tested against them.
Law professors documented the questions they were asked during office hours. Gemini 2.5 and the professors then answered those questions, and other law professors blindly judged the results. Gemini won 75% of the time over the professors. Gemini’s answers were also rated as less harmful than the humans’ answers.
And Gemini 2.5 is not the best model in the Gemini family.

AI CEOs urge Congress to mandate DNA synthesis disclosure and screening
Sticking with the same theme of the power of AI… all of this news is just this week by the way… OpenAI joined DeepMind and Anthropic in backing mandatory DNA synthesis tracking.
Basically, this would allow the government to screen synthetic nucleic acids and block combinations that could be dangerous. It would also help make sure that people trying to buy synthetic DNA or RNA are not bad actors.
So, on one hand, we will have synthetic RNA and DNA breakthroughs that can help with vaccines and biotech. On the other hand, AI could give criminals the ability to create pandemics and unleash new pathogens.
It’s a pretty science-fiction-like letter from these AI leaders.

OpenAI launches Rosalind Biodefense program to prevent AI-driven pandemics
Separately, OpenAI launched the Rosalind Biodefense Program to prevent AI-driven pandemics. Rosalind Biodefense is an initiative to develop high-impact defensive applications of AI in the life sciences.
The name Rosalind comes from one of the new GPT models, called GPT-Rosalind. I have no idea how they name this stuff anymore.
OpenAI is also giving expanded access to U.S. government agencies and allied partners for public health and biodefense missions.
Open-weight AI models are only four months behind US frontier labs
Open-source models are still only four months behind U.S. frontier labs. Epoch Research released its latest study, which appears to push back on the idea that closed models are now leaping ahead with things like Claude Mythos.
To be honest, I think Epoch simply hasn’t included Claude Mythos and GPT-6 in its charts. So, I wonder whether this trend of a four-month lag will stick around for more than a few months.

If open source is able to withstand the leaps from Mythos and GPT-6, things are going to be very tough for the U.S. frontier market because those companies simply won’t be able to invest at this rate forever.
Internet bot traffic surpasses human web traffic
The final story of the AI power theme this week is that Cloudflare reported that agentic traffic has become so popular that it has surpassed human traffic online.
Cloudflare is the annoying interstitial that tracks every single thing that comes to every single website and documents everything about it. I’m not a huge fan of Cloudflare, but they work extensively with publishers, and I try my best to assume they will be a good partner.
If anything, it’s worth pointing out that Cloudflare managed to drink Akamai’s milkshake and take their market share. There’s no reason Akamai shouldn’t be in the position Cloudflare is in now.
Now we shift into the second main theme: agentic AI news.
Theme Two: AGENTS
OpenClaw is now on Microsoft
The first agent story is that Microsoft is now supporting OpenClaw. The company has added some security precautions, but at the big Microsoft event recently, Peter Steinberger appeared and announced that Microsoft was going to support OpenClaw.
Microsoft announces always-on personal work agent called Scout
Related to the OpenClaw story, Microsoft announced that it is building an always-on personal work agent called Scout. Scout is also being discussed as part of a unified Copilot super app. Basically, Microsoft is building a federated app that will combine all of the different Copilot tools, and it’s going to name it Scout. It’s supposed to be out by the end of the summer.
I have a lot of mixed feelings about Copilot, but it would be great if Microsoft could improve it a little bit.

Legal AI juggernaut Harvey wants to cut verifier costs by 1000x
Next in agentic news, legal AI juggernaut Harvey is looking to cut verifier costs by 1,000x.
Nous Research launches official Hermes desktop app
Nous Research is an American open-source firm that has built the biggest competitor to OpenClaw, called Hermes. Hermes now has an official desktop app, which is pretty exciting, and I’m always rooting for Nous Research.
It’s not a very famous company, but it is doing a very good job of keeping up with OpenClaw. https://hermes-agent.nousresearch.com/

OpenAI expands Codex with role-specific plugins for office work
OpenAI announced that its software development tool Codex is now used by more than 5 million people per week. However, lots of non-developers have started to adopt it, using it for analyst work, marketing, design, research, investments, and banking.
Twenty percent of Codex users are no longer software developers, and that group is growing more than three times as quickly.
There are now tons of plugins you can add to Codex to help it adapt to your role. These are basically like Claude Skills or something, but they either connect Codex with tools or guide it toward a certain level of expertise.
This week, OpenAI launched six specific plugins that include 62 apps and 110 skills. https://openai.com/index/codex-for-every-role-tool-workflow/

There’s a data analytics plugin that works with business analytics. It can review reports and dashboards and use tools such as Snowflake, Databricks Genie, Hex, and Tableau.
There’s a creative plugin that can integrate with Figma, Canva, Shutterstock, Picsart, and FAL. It can actually build entire e-commerce sets.
There’s a sales plugin that helps sales teams find accounts or leads. It can prepare for meetings, complete follow-ups, and identify ways to improve performance.
There’s a product design plugin that can examine the user experience and create prototypes. It can even take static screenshots and turn them into interactive websites.

There’s a public equity investing plugin that helps people examine market and company information, earnings, and investment data. It connects with Moody’s, Daloopa, Datasite, FactSet, S&P, PitchBook, and other tools like that.
Then there’s a separate investment banking plugin that helps prepare pitch materials and can analyze comparable companies and transactions.
There’s quite a bit going on here, and I highly recommend trying it if you haven’t yet. OpenAI said that more is coming soon, including tools for corporate finance and legal work.
ChatGPT Sites turns Codex outputs into web apps
One of the new plugins is actually a pretty big feature called Sites. You can simply give Codex a prompt, and it will turn it into a website.
It appears that OpenAI must have enterprise partnerships with APIs and all sorts of webhooks that you don’t have to go learn, because there’s no need to set up an API yourself.
The downside with all of this, of course, is that you really aren’t building a website because everything is hosted and run through OpenAI. It kind of reminds me of when blogs disappeared and everyone started publishing into Facebook. You’re now kind of trapped in that Meta ecosystem of Instagram and Facebook, and you don’t really have a website anymore because everything is just being syndicated into social.
If OpenAI does that for websites, it would mean that you’re vibe coding all this stuff that’s hosted on OpenAI’s cloud, and it’s probably not very portable. And even if it is portable, you’re probably going to have to swap out or figure out on your own how to create connections to all these different APIs.
For all I know, that may just become a thing of the past completely. But for the here and now, that seems like a pretty big callout.
Still, it’s a lot of fun. If you’re not technical, you can basically make disposable websites in about two minutes. Here’s one I made for an upcoming hike that I’m doing with my daughter.
https://grand-traverse-weather.ethanholland.chatgpt.site

That wraps up the agentic news section of this week’s top stories.
Next, we move into the device category.
Theme Three: DEVICES
Microsoft surprises with a dedicated hardware device to control AI agents
First, Microsoft surprised everybody this week by announcing a dedicated hardware device that can control AI agents. It’s called Project Solara, and it seems like a pretty big leap ahead of the competition in a space I did not expect Microsoft to enter. https://commandline.microsoft.com/project-solara-build-2026/

It’s kind of the last thing I would expect or want to use. But credit where credit is due. I have to figure this out now because Microsoft appears to have leapfrogged everybody in this case.

I was an early adopter of the Rabbit R1, which is sitting in a drawer now. I also had the Limitless Pendant, which is also sitting in a drawer. I have a Google Home and an Amazon Alexa, and it will be really interesting to see whether Microsoft can wiggle its way into this world.


To me, it feels like it’s going to be a B2B-type play, where I imagine people will be force-fed these devices at work and they’ll end up being kind of lame. For personal use, I don’t know whether I really want to deal with a Microsoft product.

The applications for B2B might be mundane for middle managers, but they may also be spectacular for customer service or healthcare. I could see operations becoming a lot more streamlined in industries where we don’t expect agents to even be around.
With a physical device, people can have structured, frequent interactions, such as checking a SKU or a price in Home Depot or finding out whether a pair of pants is available online in a certain waist and inseam.
It could create an incredible change in how we experience things that used to be frustrating and could now possibly become delightful. I’m going to go ahead and stay optimistic on this one.
OpenAI curiously leads funding round for webcam maker Opal Electronics
The second device story is that OpenAI is leading funding for a company called Opal Electronics, which has an array of AI-native devices. This is kind of an interesting twist that goes right along with the Microsoft announcement, as well as pendants like Limitless and handhelds like the Rabbit R1.

OpenAI spent something like $4 billion to acquire Jony Ive’s company, and it has basically just kicked the can on that. Now, here it is investing in Opal Electronics.
If agents are the buzzword for 2026, maybe devices will start to kick in during 2027, especially as voice APIs become faster and faster and can now operate in real time and include research.
It’s just a matter of time until you can basically talk to agents without having to build them. I think the bridge will be through these devices, with consumers and workers basically conversing with them and getting the results they want from AI instead of from a human.
And that’s going to get pretty weird pretty fast.
We already do it now with Siri and Alexa for simple things like the weather, and we do it a lot with Google. Everything is just merging into the conversation.
There’s going to be a lot of dictation happening. That could be the worst-case scenario. Whoever figures out “thought-to-text” will win it all.
I’ve had a lot of people tell me that this will never work because it’s rude. But I think people already yap on their phones all the time in public spaces. I hope we don’t create an environment with too much public yapping, but I don’t necessarily see the net amount going up or down. I think it’s just going to shift to our devices.
NVIDIA and Microsoft appear to be redesigning Windows PCs for local AI workloads
This is one of the coolest stories of the week. I’ll walk you through the headline in somewhat plain language.
Basically, Nvidia is building a chip called RTX Spark that puts a CPU, GPU, and 128 gigabytes of shared memory into one package that is small enough to fit inside a thin laptop. On top of that, Microsoft is rewriting critical parts of Windows so that AI agents can run on the machine instead of having to hit the cloud.
This will move a lot of AI’s computing power away from something you rent by the month in the cloud and onto something you can run and process locally on your computer. This is faster and a lot more cost-effective, and it starts to integrate the AI agent directly into the local machine.
It also puts Nvidia into the PC processor business that Intel and AMD have basically dominated for 40 years. So, this is a pretty big deal.
You’ve got a petaflop of speed, which means one quadrillion mathematical calculations per second, and 128 gigabytes of memory. Memory is the main bottleneck when you try to run AI locally. Most laptops have about 16 to 32 gigabytes of system memory, and then they have a graphics card that may have 16 to 18 gigabytes of video memory.
Having 128 gigabytes of unified memory means the CPU and GPU can now share that pool together, and a 120-billion-parameter model can literally fit on your laptop.
The new CPU is called the Grace CPU, and to my knowledge, Nvidia has never sold a PC processor before.
There’s also a pretty technical piece of this that I don’t completely understand called OpenShell by Nvidia. OpenShell is an identity and security layer that can coordinate resources locally on a Windows machine. However, if you’re computing power exceeds your laptop and needs to go to the cloud, OpenShell can mask your personal information so that nothing personal goes to the cloud.
Your laptop basically runs things locally and then hits the cloud when necessary. It reminds me a little bit of a hybrid car, where sometimes you’re running on battery, sometimes you’re in hybrid mode, and sometimes you’re using gas. All of that routing is going to be done locally on your laptop, and OpenShell makes it possible.
Every major PC maker has signed on to this project. Asus, Dell, HP, Lenovo, MSI, and Microsoft’s Surface brand are all committed to this overhaul.
For people like me who track AI, you start to see some interesting patterns, even if you’re not technical.
For example, Nvidia has the Nemotron family of models, which are all open source and just happen to include 120-billion-parameter large models. It turns out that Nemotron 3 Super is a 120-billion-parameter model. Nemotron 3 Super has a one-million-token context window, and the new RTX Spark system that Microsoft and Nvidia are partnering on for these laptops just happens to have a one-million-token context window and can host a 120-billion-parameter model locally.
Nvidia basically sized this hardware around its own model and then quoted the model specifications as a hardware capability.
I don’t think Nvidia has any interest in winning the model race because it doesn’t make money on Nemotron. But it does want local models to be free, plentiful, and enormous because that makes using the new Nvidia silicon table stakes.
Nvidia is pretty open about the fact that it also supports Qwen 3.5 and Mistral Small 4, and that it can also work with Llama and Comfy. Nvidia is psyched about anything that can run locally because that helps sell more GPUs.
Microsoft has a bunch of locally run models as well. I don’t have them all memorized, but one of its flagship model families is called Phi.
A lot of stuff is coming together. If you think about the fact that just this week we’ve got agentic hardware devices from Microsoft, and laptops with the new Nvidia tech stack, everything just keeps getting smaller and portable.
I think Bill Gates once said that we’re basically trying to get everything inside a contact lens that fits in our eyeball or an implant that goes in our ear. You can see this continuing to happen.
Now we’re switching over to the next theme of the week, which is robotics, and we can maintain some continuity with NVIDIA.
Theme Four: ROBOTS
Nvidia launches Cosmos 3 robotics open model with Runway as the main partner
NVIDIA is one of my favorite companies when it comes to robotics because its top researcher, Dr. Jim Fan, is a master at simulations and training robots on skills using thousands of virtual universes. He adds a 1,001st universe, which happens to be ours, by connecting the simulation wires, essentially, to real-world robot vision and sensors.
You start with a robot that is just “seeing” a world that is entirely a simulation. Then you swap the robot’s input from the simulation to real-world sensors, and now the robot is walking around our world and doesn’t know it’s any different from the simulations. It’s unbelievable.
Cosmos is NVIDIA’s flagship model for physical AI. It’s a combination of physical AI, reasoning models, world simulations, and action generators.
Reasoning, simulation, and actuators all in one place.
Cosmos 3 is a fully open omni-model with native vision and multimodal generation across text, images, video, and sound. This thing is a Ferrari. It could easily be the top story if evrunery week weren’t insane.
What’s interesting is that this is the first time I’ve seen what used to be a video company, Runway, placed into the robotics pile. We’ve known that Runway has wanted to shift toward robotics for some time, but now we have a whole kit and caboodle of things that NVIDIA is working with, and that’s kind of why this is an omni-model.
Fun fact, I hosted a Runway Meetup at the Milton Theater last year. We had a solid three people (including my mom) show up! LOL.
You’ve got Black Forest Labs, which created one of the greatest early image-generation tools, Flux. You’ve got Runway, which has a really strong video model, and more traditional physical AI companies such as Agile Robots and Skild AI.
This concept of omni-models is really mind-blowing because these things natively understand everything.
Three years ago, we started with models that could handle text. Then we had models that could create images. Then we had text-to-image and image-to-text. Then we had separate video and audio. Then both. Now, we’ve got all of these things coming together into one model that can handle a giant pile of skills.
And that’s kind of like our brain. We may have specialties that we can sharpen, but one of the biggest barriers to artificial general intelligence was that we had all of these really smart, linearly talented specialists. Omni-models can really boil the ocean and handle anything.
If you don’t know about Cosmos, you definitely should learn about it. And if you want somewhere to start, I would begin by Googling Dr. Jim Fan. Then check out how he taught a robot dog to balance on a yoga ball in simulation and successfully made it work in real life on the first try after leaving the simulation.
OpenAI is on a humanoid robotics hiring push
OpenAI was originally associated with the 1X platform, which makes the NEO robot. It seems to have left a lot of that behind, but now OpenAI is on a robotics hiring push.
Sam Altman posted on Twitter that OpenAI Robotics is hiring and has a ton of open positions.
Things are happening quickly, bubbling up from the bottom. No one seems to see it all coming up from beneath them, but it’s going to get nuts.
I keep saying this. It’s just a matter of time.
Speaking of things getting nuts, this has got to be one of the more dystopian company launches I’ve ever seen.
Dystopian startup Shift offers free NYC apartment cleaning to collect training data for robots to replace people
A company called Shift will clean your apartment in New York City for free. It will send a cleaning service whose workers wear, essentially, a bunch of devices that record everything they do.
So, you don’t pay anything, and this poor cleaner shows up covered in recording devices, cleans your house, and then the company takes that data and leverages it to train robots to perform daily tasks.
It’s absolutely and transparently aimed at putting service workers out of business. I’m not saying this is bad because I hope people will find other jobs. But on the other hand, man, it’s pretty dystopian to have the person doing the job wear the equipment that will train their replacement.
The personal information is supposedly anonymized, so your house itself is not really going to become part of the system. But the company is literally going after cleaning, service, repairs, and errands.
Not a fan. But then again, as Bob Marley says, none of us can stop the time.
Now we shift over to more business-oriented news for the next two stories.
Theme Five: BUSINESS
DeepSeek raising $7.4 billion, valuation nears $59 billion
The first is that DeepSeek is raising $7.4 billion at an almost $60 billion valuation. DeepSeek is the famous Chinese AI model that completely disrupted the U.S. marketplace when it announced DeepSeek R1 on the same day as the presidential inauguration in January 2025.
Many people called it a Sputnik moment. It completely shook Wall Street, triggered a massive drop in tech stocks and a 17% plunge in Nvidia’s market value, and showed that frontier-level reasoning models could be trained at a fraction of the cost with only a few months of lag time.
We’ve talked a lot about that open-source delay, which is now about four or five months behind U.S. frontier models.
I’m not an expert on DeepSeek, but I’m close enough for horseshoes. Basically, what’s interesting to me about DeepSeek is that it’s owned by a Chinese hedge fund. It was founded by an individual named Liang Wenfeng, who is the CEO of both the hedge fund High-Flyer and DeepSeek. High-Flyer is a high-frequency trading hedge fund.
Wenfeng is committing 20 billion yuan of his own money, and Tencent is supposedly contributing 10 billion yuan to the cause, with battery maker CATL coming in at 5 billion yuan.
Uber reportedly caps coding agent spending at $1,500 monthly per employee
“Uber reportedly now caps coding agents at $1,500/month per employee per tool – seems sensible to me, but it’s also an interesting hint at the value Uber thinks these tools are providing”
https://simonwillison.net/2026/Jun/3/uber-caps-usage/
Now we jump over to science news.
Theme Six: SCIENCE
Mayo Clinic and Microsoft are building a healthcare-specific frontier AI model
The big story this week is that the Mayo Clinic and Microsoft are building a healthcare-specific frontier AI model. A lot of it is designed around the patient and clinician experience, which I think is a great idea.
https://news.microsoft.com/source/2026/06/02/mayo-clinic-and-microsoft-collaborate-to-develop-a-frontier-ai-model-for-healthcare/
The model will be owned by the Mayo Clinic but hosted on Microsoft’s Azure Boundary API.
The last major story in this week’s top stories takes us back to the local model theme.
Google’s Gemma 4 12B is local AI that’s almost as good as the cloud Gemma
Google launched Google Gemma 4 12B, which is a unified, encoder-free multimodal model. What that basically means in plain English is that there are no separate encoders, so vision and audio go directly into the same system as the large language model. That’s basically what “unified” means.
https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/
It also has advanced reasoning, but most importantly, it’s small enough to run locally with just 16 gigabytes of RAM. Compare that to the massive Nvidia-Microsoft partnership. This is a tiny model.
It’s also completely open under Apache 2.0 licensing. It’s so powerful for the size, that it’s almost as capable as Google’s largest cloud-based Gemma models, but it has half the memory footprint.
Gemma would be really good for things you want to run locally and use to power tasks that don’t necessarily require a complicated language model.
My DSLR can already do face tracking, which is basically object segmentation, right inside the camera. You can only imagine what Gemma will be able to do.
Gemma can also process PDFs and legal contracts. I used to find that hard to believe, from a comfort-level standpoint, because I have a bad habit of always using the best model possible, like Opus instead of Sonnet.
But Gemma can write pretty good code, translate languages, and understand images. It can also help with grammar or full emails and extract data from PDFs. If you wanted to create a locally hosted customer support chatbot, it could do that.
For things like Siri or Google Android assistants, something like Gemma would be very powerful, especially when you combine it with a router.
Those are the 35 stories I want to make sure I remember when I come back and read this in a few months or years. There are 20 more top stories worth reading below as part of the executive summary recaps, below.
Full Executive Summaries with Links, Generated by Haiku 4.5
(I have Haiku do these weekly to compare to my effort… it’s usually just as good)
Anthropic expands AI vulnerability-scanning program to 150 critical infrastructure organizations.
Project Glasswing, which uses Anthropic’s Claude Mythos model to find security flaws in software, is tripling its partner organizations from 50 to 150 across sectors like power, water, healthcare, and communications—all entities whose compromise could affect over 100 million people each. The expansion reflects growing concern that cheaper, more capable AI models for hacking are months away, making it urgent for defenders to adapt now; Mythos has already helped partners identify over 10,000 high-severity vulnerabilities, and Anthropic is releasing complementary tools to help the broader software industry scale defenses.
Expanding Project Glasswing \ Anthropic https://www.anthropic.com/news/expanding-project-glasswing
Anthropic reports AI systems are now writing most company code and could soon design their own successors.
Anthropic’s internal data shows Claude AI is accelerating its own development cycle—engineers now ship 8× more code per quarter than in 2021–2025, with Claude authoring over 80% of merged code as of May 2026. The company warns that this trend, if it continues, could lead to “recursive self-improvement,” where AI systems autonomously build and improve their own successors, potentially arriving sooner than institutions are prepared to manage. While Anthropic emphasizes potential benefits for science and healthcare, it also highlights growing risks around human control and oversight of increasingly autonomous AI development.
Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention. https://x.com/AnthropicAI/status/2062568862479208923
When AI builds itself \ Anthropic https://www.anthropic.com/institute/recursive-self-improvement
Anthropic engineers now ship eight times more code per quarter than four years ago.
The company’s productivity surge suggests AI-assisted development is accelerating internal innovation cycles, though the metric doesn’t clarify whether more code means better products or reflects shifting development practices. This represents a concrete productivity gain at a major AI lab, signaling that AI tools are reshaping how researchers and engineers work rather than replacing them entirely.
Today, Anthropic engineers on average ship 8x as much code per quarter as they did compared to 2021-2025. https://x.com/AnthropicAI/status/2062568864240836995
Claude Mythos beats expert forecasts on AI planning ability.
Expert superforecasters predicted AI systems wouldn’t handle complex multi-hour tasks until year-end, but Anthropic’s Claude Mythos reached that milestone in late May—two months ahead of schedule. This matters because it suggests AI planning capabilities are advancing faster than specialists anticipated, with potential implications for autonomous task execution in real-world scenarios.
In early May, the best superforecasters predicted that, by the end of the year, the longest METR 80% task horizons would reach 3-4 hours. In late May, Claude Mythos achieved that number. https://x.com/emollick/status/2062235461364445204
Claude’s AI now outperforms humans on research decisions by wide margin.
Anthropic’s latest model, tested on real research dilemmas, suggested better next steps than human researchers 64% of the time, nearly tripling its success rate from last year’s 22%. This matters because it suggests AI is moving beyond answering questions to actively improving how humans conduct complex, creative work—the kind of judgment calls that typically require deep expertise and intuition.
AI research is a series of next-step decisions. We looked at sessions where a human researcher took a wrong turn, showed Claude the session up to that point, and asked it what to do next. Mythos Preview improved on humans 64% of the time—up from 22% in 2024. https://x.com/AnthropicAI/status/2062568870872003021
AI compute capacity expected to surge tenfold over three years.
Microsoft’s AI chief projects that computational power for training advanced AI systems will increase roughly 1,000-fold by 2029, from current levels to what experts call “5e30 FLOPs”—a measurement of processing speed. This explosive growth matters because it directly determines how capable future AI models can become, though it also raises urgent questions about energy consumption, chip manufacturing bottlenecks, and whether performance gains will match this computational expansion.
Mustafa Suleyman, CEO of Microsoft AI says that AI compute will grow 1000x in the next 3 years We are currently at around 5e27 FLOPs. Three more OOMs mean we reach 5e30 FLOPs in 2029. https://x.com/scaling01/status/2061901702324695115
Gemini AI outperforms law professors in blind-judged office hour Q&A test.
Google’s latest Gemini model won 75% of head-to-head comparisons against actual law professors answering student questions, with independent judges rating its responses as less potentially harmful. This suggests AI has crossed a threshold where it can credibly substitute for expert human guidance in at least some professional contexts—a shift from AI being a novelty tool to being genuinely competitive with specialists.
Law professors wrote questions they were asked during office hours. Gemini 2.5 & humans answered them then other law professors blindly judged the results: -Gemini had a 75% win rate vs. professors -Gemini’s answers were rated LESS harmful than humans -Newer models do even better https://x.com/emollick/status/2061876620638486584
AI leaders push Congress to require DNA synthesis screening.
CEOs from OpenAI, DeepMind, and Anthropic joined biosecurity experts and former national-security officials in publicly calling for mandatory screening of DNA synthesis orders to prevent misuse. The letter signals that major AI companies see biosecurity risks—particularly around dual-use biological threats—as serious enough to warrant government regulation, marking a rare moment of industry consensus on the need for external oversight.
OpenAI, DeepMind, Anthropic CEOs back mandatory DNA synthesis screening A coalition of AI leaders, synthesis-industry executives, biosecurity researchers, and former national-security officials published an open letter in June 2026 urging Congress to make screening and https://x.com/kimmonismus/status/2062485389949145457
OpenAI launches Rosalind Biodefense to fortify pandemic preparedness.
OpenAI is giving vetted developers and U.S. government health agencies access to GPT-Rosalind, a specialized AI model for life sciences, to build tools for disease detection, outbreak response, and vaccine development. The move reflects a deliberate strategy to ensure advanced AI capabilities benefit defenders against biological threats while maintaining safeguards through trusted-partner models and external expert review—with initial projects already underway at organizations like Lawrence Livermore National Laboratory and Johns Hopkins Applied Physics Laboratory.
Strengthening societal resilience with Rosalind Biodefense | OpenAI https://openai.com/index/strengthening-societal-resilience-with-rosalind-biodefense/
Open-weight AI models falling further behind proprietary rivals.
The gap between freely available AI models and paid commercial ones has widened to a four-month delay since January, meaning models anyone can download are noticeably less capable than industry leaders like OpenAI’s systems. This matters because it could entrench market power among well-funded companies and limit researchers’ ability to study or improve frontier AI technology. The trend reverses earlier momentum where open models were narrowing the gap.
We took another look at the capability gap between open-weight and proprietary models. Since the start of the year, open-weight models have lagged the state of the art by four months. https://x.com/EpochAIResearch/status/2060451576779886942
AI bots now outnumber human users on the internet.
Automated agents have surpassed human traffic online, arriving years ahead of expert predictions that anticipated late 2026 or 2027. This milestone reflects the rapid deployment of AI systems performing tasks autonomously—from data gathering to transactions—and signals a fundamental shift in how the internet operates, with potential implications for network infrastructure, security, and the economics of online services.
Welp, that happened faster than I predicted. Thought it would be end of 2027, then early 2027, but agentic traffic growing so fast that bots have now passed human traffic online for the first time in the Internet’s history. https://x.com/eastdakota/status/2062212701414187452?s=20
Microsoft acquires OpenClaw with security focus intact.
Microsoft has brought OpenClaw into its fold with founder Peter Steinberger leading the charge, emphasizing that security measures remain a priority through the transition. The acquisition signals Microsoft’s continued investment in developer tools and open-source communities, though specifics on how OpenClaw will integrate into Microsoft’s broader AI strategy remain limited in this announcement.
OpenClaw is on Microsoft now, with all necessary security precautions 🤍 Peter Steinberger introducing it himself. Excited for the whole @openclaw team: @steipete, @davemorin, @vincent_koc to name a few https://x.com/TheTuringPost/status/2061870411571466666
Microsoft’s AI agent promises to work continuously on your behalf.
Microsoft is positioning “agents”—AI systems that run autonomously without constant human direction—as the next frontier after generative AI, with the company unveiling an always-on work assistant that operates independently. This matters because it signals a shift from AI tools you ask questions to AI systems that anticipate and execute tasks, potentially reshaping how office work gets done. The industry momentum behind autonomous agents suggests 2026 will see major enterprises deploying these systems at scale, moving beyond chatbots to software that manages workflows on its own.
Microsoft scout revealed „your always-on personal agent for work.“ If “”AI”” was the Word of the Year in 2025, in 2026 it will be “”agents”” (always-on). Everything is agentic this year. https://x.com/kimmonismus/status/2061875714933371220
Microsoft consolidates fragmented Copilot tools into unified super app.
Microsoft is merging its scattered AI assistants—GitHub Copilot for coding, Copilot Chat, and a new Scout agent—into a single interface launching by summer 2026. The move addresses customer frustration with juggling multiple tools and aims to boost weak adoption rates, particularly among Microsoft 365’s 450 million users, where less than 4.5% pay for Copilot features. The company faces mounting competition from rivals like OpenAI, Google, and startup Cursor, making product consolidation essential to recapture market share in enterprise AI.
Exclusive: Microsoft is building a super app that combines coding, chat, and other Copilot AI tools | Fortune https://fortune.com/2026/05/29/microsoft-working-on-super-app/
Exclusive: New screenshots of upcoming Copilot Super App https://www.testingcatalog.com/exclusive-new-screenshots-of-upcoming-copilot-super-app/
AI verifiers could become dramatically cheaper through specialized design.
Researchers are exploring ways to reduce the cost of AI “verifiers”—systems that check whether other AI agents complete tasks correctly—by up to 1,000 times. Since these verification systems are critical for both testing AI performance and training better models, their high cost has become a scaling bottleneck; cheaper verifiers could unlock faster and more affordable AI development across industries.
Can we design legal agent verifiers that are up to 1,000x cheaper? Verifiers are LLM judges that check an agent’s work against rubric criteria: they’re used both in agent benchmarking and as reward signal in post-training. But verifiers can be a bottleneck at scale. For https://x.com/harvey/status/2061866491033899371
Hermes desktop application launches across Windows, Mac, and Linux.
Anthropic’s Claude AI now has a dedicated desktop interface for direct computer access, expanding beyond web-based chat. This matters because desktop apps typically offer faster performance, offline capability, and deeper system integration than browser versions. The multi-platform rollout suggests Anthropic is positioning Claude as infrastructure for everyday work rather than a specialized tool.
It’s finally here. The official Hermes Desktop app. Available on all platforms. https://x.com/Teknium/status/2061844602735538266
OpenAI expands Codex beyond coders with role-specific tools.
OpenAI is rolling out six specialized Codex plugins designed for non-developers—analysts, marketers, salespeople, designers, and investors—who now represent 20% of Codex’s 5 million weekly users and are growing three times faster than developers. The company is also introducing Sites, allowing teams to generate and share interactive web dashboards and planning tools, plus Annotations for precise refinements without starting over. This matters because it signals a shift in how businesses deploy AI: moving from code-focused development to enterprise workflow automation across departments, with early use cases ranging from sales pipeline management to investment thesis preparation.
Codex for every role, tool, and workflow | OpenAI https://openai.com/index/codex-for-every-role-tool-workflow/
AI tool lets teams convert documents into shareable interactive websites instantly
Codex Sites automates the conversion of work documents and plans into functional web applications accessible via simple URLs, launching first to paid tier users. This matters because it removes traditional coding barriers for non-technical teams, letting business users deploy interactive tools without developers—a tangible shift from AI as an assistance layer to AI handling core creation tasks.
Building apps has never been easier. With Sites, Codex can turn your work, ideas, and plans into an interactive website or app your team can explore, use, and share with a URL. Rolling out to Business and Enterprise plans, before expanding more broadly. https://x.com/OpenAI/status/2061845949170045346
Handheld devices give users direct control over AI agents, shifting strategy
Microsoft is releasing handheld and desktop controllers specifically built to manage AI agents—a move that mirrors expectations around dedicated hardware for agent control. This represents a shift from traditional software interfaces toward physical devices as the primary way users interact with autonomous AI systems, signaling that companies now view agent management as central enough to warrant dedicated hardware products.
This came as a surprise: Microsoft has unveiled handheld and desktop devices designed to control one’s agents. It reminds me of what I had expected from OpenAI’s hardware-standalone devices for controlling agents. https://x.com/kimmonismus/status/2061860319547527191
OpenAI backs hardware startup Opal Electronics for AI devices.
OpenAI is leading a funding round for Opal Electronics, a webcam maker now expanding into AI-native devices for creative work, as part of OpenAI’s broader “ambient computing” strategy to build screenless gadgets that sense the world in real time. This move appears designed to accelerate product launches and gather user data while OpenAI’s flagship palm-sized device with designer Jony Ive faces delays until 2027, signaling the company’s serious pivot from software-only business toward physical hardware competition.
OpenAI makes its next hardware move with Opal Electronics https://www.testingcatalog.com/openai-makes-its-next-hardware-move-with-opal-electronics/
NVIDIA and Microsoft partner to bring AI processing directly to Windows PCs.
NVIDIA and Microsoft are collaborating to equip personal computers with specialized AI chips and software, shifting computation away from cloud servers to individual devices. This matters because it could make AI features faster, more private, and less dependent on internet connectivity—fundamentally changing how consumers interact with PCs. The partnership suggests a strategic bet that the next wave of computing growth depends on embedding AI capabilities at the hardware level rather than accessing them remotely.
NVIDIA and Microsoft Reinvent Windows PCs for the Age of Personal AI | NVIDIA Newsroom https://nvidianews.nvidia.com/news/nvidia-microsoft-windows-pcs-agents-rtx-spark
NVIDIA and partners launch open-source AI model for robotics and physical tasks
NVIDIA unveiled Cosmos 3, an open-source foundation model designed to help robots and physical systems understand and interact with the real world. The model combines multiple AI capabilities—vision, language, video, and action prediction—in a single system that can reason about physical scenarios and generate responses across text, image, and video formats. This matters because open-sourcing frontier models typically accelerates industry adoption; NVIDIA is positioning this through a new coalition with leading AI labs to democratize development of “world models” that teach machines how physical systems actually work, potentially faster than proprietary alternatives.
Introducing the Cosmos Coalition A new global initiative with NVIDIA and leading AI labs to build and open-source frontier world models for physical AI. Runway joins as a founding member, working alongside NVIDIA and a set of leading AI labs to build, share and accelerate world https://x.com/runwayml/status/2061315089869721682
Jensen just launched NVIDIA Cosmos 3. Pitched as the first fully open omnimodel for physical AI: a mixture-of-transformers (reasoning + generation) with native vision reasoning and generation across text, image, video, sound, and action. Tops open-model leaderboards on https://x.com/TheHumanoidHub/status/2061333253920080345
NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI | NVIDIA Newsroom https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-3-the-open-frontier-foundation-model-for-physical-ai
OpenAI appears to be developing humanoid robots alongside language models.
OpenAI has hired robotics specialists and filed patents related to robotic systems, signaling a shift beyond software into physical automation. This matters because it could expand AI’s real-world impact from text and analysis into manufacturing, logistics, and service industries—where embodied AI could compete with human labor directly. The move also suggests OpenAI sees robotics as a natural extension of its AI capabilities rather than a separate venture.
All signs point to OpenAI building humanoid robots. https://x.com/TheHumanoidHub/status/2061322294426038393
OpenAI’s robotics division is actively hiring engineers to build practical robots.
OpenAI is moving beyond software into physical robotics, recruiting specialists in hardware design, manufacturing, and machine learning to create robots for real-world tasks. This signals the company’s belief that AI’s next frontier is enabling machines to interact with the physical environment rather than remaining confined to digital systems. The hiring push suggests OpenAI sees near-term commercial potential in robotics, though the statement remains vague about specific applications.
OpenAI Robotics is hiring, looking for exceptional full-stack hardware, ops, systems, and ML engineers to help us program and manufacture robots that are useful for society. AI should be able to help people in the physical world. In the short term, we are focused on robots to https://x.com/sama/status/2061117302528188712
Shift launches free apartment cleaning service in New York City.
A new startup is offering complimentary home cleaning in exchange for recording the work, using human cleaners equipped with data-collection devices. The model tests whether consumers will trade privacy for free services—a significant shift from traditional paid cleaning platforms. This raises questions about data usage and consent in gig economy services.
Today, we’re launching shift. We’re starting by cleaning your apartment in New York City, for free. Here’s how it works. Book a shift cleaning. A vetted shift operator comes to your home wearing one of our devices. They clean. They leave. You pay nothing. In exchange, we record https://x.com/joinshiftX/status/2060044783519735987?s=20
Hyperscaler spending on AI infrastructure hits $770 billion in 2026.
Major cloud companies are maintaining their aggressive investment pace in data centers and computing power needed for AI systems, with spending expected to exceed $1 trillion by 2027. This sustained capital deployment signals confidence in AI’s commercial potential while raising questions about whether returns will justify the enormous financial commitment.
Hyperscaler capital expenditures came in on trend in Q1 2026, continuing the trajectory that projects them spending $770 billion this year and over a trillion dollars in 2027. https://x.com/EpochAIResearch/status/2060076222873526506
Chinese AI startup DeepSeek raises $7.4 billion at $52–59 billion valuation.
DeepSeek, which gained global attention last year by matching U.S. AI capabilities, is closing its first funding round with $7.4 billion from Chinese tech and battery companies including Tencent and CATL, reflecting Beijing’s push to build a self-sufficient AI ecosystem independent of American technology. The investment underscores the intensifying U.S.-China technological competition, with DeepSeek explicitly limiting early access to its models to Chinese partners while Western companies are excluded.
DeepSeek slated to draw $7 billion in maiden fundraising, sources say https://www.cnbc.com/2026/06/03/deepseek-slated-to-draw-7-billion-in-maiden-fundraising-sources-say.html
Uber limits AI coding tool spending to $1,500 per employee monthly.
Uber’s cost cap on AI-powered coding assistants suggests the company has quantified real productivity gains while guarding against runaway spending on nascent tools. The threshold reveals how enterprises are moving beyond open-ended AI experimentation toward disciplined deployment—treating these agents as measurable business investments rather than unlimited resources.
Uber reportedly now caps coding agents at $1,500/month per employee per tool – seems sensible to me, but it’s also an interesting hint at the value Uber thinks these tools are providing https://x.com/simonw/status/2062143151184465964
Canada launches $200 billion AI strategy to close adoption gap
Prime Minister Mark Carney unveiled “AI for All,” a five-year national strategy targeting $200 billion in economic growth and 250,000 new jobs by closing Canada’s AI adoption lag—currently just 12% versus 60% globally. The plan focuses on three pillars: building public trust through stronger data protections and AI transparency, creating opportunities via mass AI literacy training and 90,000 job placements, and securing sovereignty by building domestic compute infrastructure and supporting homegrown AI champions. This addresses a concrete competitive vulnerability: Canada has world-class talent but ranks among slowest G7 nations in deploying AI at scale, risking brain drain and foreign control of critical infrastructure.
Prime Minister Carney launches AI for All: Canada’s new national artificial intelligence strategy | Prime Minister of Canada https://www.pm.gc.ca/en/news/news-releases/2026/06/04/prime-minister-carney-launches-ai-all-canadas-new-national-artificial
Mayo Clinic and Microsoft build specialized AI model for clinical care.
Mayo Clinic and Microsoft are developing a healthcare-specific AI system that combines Mayo’s clinical expertise and patient data with Microsoft’s AI technology, designed to improve diagnoses and treatment decisions. The model, owned by Mayo Clinic, will be tested internally before becoming available globally through Microsoft’s cloud platform, addressing a critical gap where general-purpose AI lacks the medical context needed for safe clinical use.
Mayo Clinic and Microsoft collaborate to develop a frontier AI model for healthcare – Source https://news.microsoft.com/source/2026/06/02/mayo-clinic-and-microsoft-collaborate-to-develop-a-frontier-ai-model-for-healthcare/
Proud that we’re collaborating with Mayo Clinic to build a frontier AI model for healthcare. Both our organizations exist to serve people at scale – and we believe this could be nothing short of transformative for global healthcare. https://x.com/mustafasuleyman/status/2061903347129418227
Microsoft launches seven new AI models with custom tuning capability.
Microsoft AI unveiled a family of seven new models spanning reasoning, coding, image, voice, and transcription tasks, paired with a technique called Frontier Tuning that lets organizations train AI systems on their own workflows and data. The company claims its flagship reasoning model matches leading competitors while being dramatically more efficient when customized—one internal example showed a tuned model matching GPT 5.4 at 10× lower cost—and announced a partnership with Mayo Clinic to develop specialized healthcare AI. This represents Microsoft’s push toward what it calls “hill-climbing,” a systematic approach to continuously improving AI by scaling compute, refining data, and building proprietary silicon.
Building a hill-climbing machine: Launching seven new MAI models | Microsoft AI https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/
NVIDIA releases efficient reasoning model designed specifically for multi-step AI agent workflows.
NVIDIA’s new Nemotron 3 Ultra model prioritizes speed and cost-efficiency for AI agents that run complex, multi-step tasks over extended periods—a growing use case where token costs balloon as agents plan, call tools, and pass information back and forth repeatedly. The 550-billion-parameter model achieves 5x faster processing than comparable models while reducing task completion costs by up to 30%, with particular strength in long-context reasoning and agent orchestration tasks. This represents a shift from optimizing single-turn chatbots to building models tailored for the practical economics of deployed agent systems.
NVIDIA Nemotron 3 Ultra Powers Faster, More Efficient Reasoning for Long-Running Agents | NVIDIA Technical Blog https://developer.nvidia.com/blog/nvidia-nemotron-3-ultra-powers-faster-more-efficient-reasoning-for-long-running-agents/
Anthropic releases Claude Opus 4.8 with improved reasoning at same price.
Anthropic upgraded its flagship Claude model to version 4.8, delivering measurable improvements in coding, legal analysis, and autonomous task completion without raising prices. The model shows four times fewer unremarked code flaws than its predecessor and now includes new features like effort-level controls and dynamic workflows for handling large-scale projects. Early adopters report the model exhibits better judgment, fewer unsupported claims, and stronger performance on specialized benchmarks for legal work and financial analysis—improvements that translate to more reliable AI for high-stakes professional use cases.
Introducing Claude Opus 4.8 \ Anthropic https://www.anthropic.com/news/claude-opus-4-8
Google releases Gemma 4 12B, an open AI model for laptops.
Google unveiled Gemma 4 12B, a compact AI model that runs on standard laptops with 16GB of memory while handling text, images, and audio without separate processing components. The model achieves performance near Google’s larger 26B model but uses half the memory, making advanced reasoning and voice capabilities accessible to developers without cloud infrastructure. Over 150 million downloads of earlier Gemma models demonstrate developer appetite for open, locally-runnable AI tools.
Introducing Gemma 4 12B https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/
Today we’re introducing Gemma 4 12B — our latest open model that brings advanced agentic reasoning, vision and audio directly to your laptop. It delivers performance nearing our larger Gemma models with a much smaller total memory footprint, while being small enough to run https://x.com/Google/status/2062203526588088452
Molmo2 advances open-source video understanding with pointing capabilities.
Allen Institute’s Molmo2 model can track objects, count items, and reason across multiple images using simple pointing gestures—capabilities previously limited to proprietary systems. This matters because it democratizes advanced video AI for researchers and developers, while the open-source approach allows community scrutiny and improvement of a technology increasingly used in autonomous systems and content analysis.
Molmo2 is a CVPR 2026 award candidate paper from @allen_ai Molmo2 is a VLM that supports video pointing, tracking, counting by pointing, and multi image reasoning all in one open model prompt: blue players https://x.com/skalskip92/status/2062549751246066144
Claude AI autonomously created a complete tabletop RPG with rulebooks and playable content.
Anthropic’s latest Claude model (Opus 4.8) demonstrated significantly expanded autonomy by independently designing, documenting, and publishing a functional RPG system through code execution—completing design, playtesting materials, website, and deployment without human guidance. The capability matters because it shows AI moving beyond responding to requests toward self-directed project completion across multiple creative and technical domains simultaneously.
Here Opus 4.8 built and play-tested a new RPG in Claude Code, including 3 PDF manuals and adventures, playtest notes, a website, and a playable solo adventure – then put it all on Netlify. No feedback from me at all. https://x.com/emollick/status/2060045063275573723
I had Opus 4.8 in Claude Code write a sophisticated, if minor, academic paper from a archive of hundreds of de-identified research files from years ago I had to use GPT-5.5 Pro as a reviewer, it spotted one major error & some minor points. Opus corrected https://x.com/emollick/status/2060098885561778341
Anthropic files confidential IPO paperwork with U.S. regulators.
Anthropic submitted a draft registration statement to the SEC, positioning itself for a potential initial public offering once regulatory review concludes. The move signals the AI safety-focused company’s readiness to access public capital markets, though the timing and final terms remain contingent on market conditions. This comes as Anthropic releases Claude Opus 5, an upgraded AI model designed for complex, long-running tasks.
Anthropic confidentially submits draft S-1 to the SEC \ Anthropic https://www.anthropic.com/news/confidential-draft-s1-sec
Anthropic has confidentially submitted a draft S-1 registration statement to the Securities and Exchange Commission. Pending completion of SEC review, this gives us the option to pursue an initial public offering. Read more: https://x.com/AnthropicAI/status/2061478052257841495
Apple opens iPhone Messages to third-party artificial intelligence tools.
Apple is allowing outside companies to build AI features directly into its Messages app, marking a shift toward openness after years of controlling which services connect to core iPhone functions. This matters because it could accelerate AI adoption among everyday users while testing whether Apple can maintain security and privacy standards when ceding control to third parties. The move suggests Apple sees competitive pressure to offer AI capabilities without building everything in-house.
Apple’s Messages app on iPhone now has a third-party AI agent – 9to5Mac https://9to5mac.com/2026/06/04/apples-messages-app-on-iphone-now-has-a-third-party-ai-agent/
Suno raises $400 million as AI music creation enters mainstream culture.
The AI music startup reached a $5.4 billion valuation and plans to launch an industry-partnered music model, signaling a shift from niche tool to mainstream platform used for everything from family mementos to therapeutic applications. The funding reflects investor confidence in democratizing music creation, though the company’s emphasis on artist partnerships suggests growing recognition that widespread AI music tools require music industry cooperation to succeed.
The Next Chapter for Suno · Suno https://suno.com/blog/series-d-announcement
Alphabet raises $80 billion to build AI infrastructure at scale.
Alphabet plans to sell $80 billion in stock—including a $10 billion investment from Berkshire Hathaway—to fund massive expansion of AI computing capacity. The move signals that demand for Google’s AI services is outpacing its ability to supply them, with CEO Sundar Pichai citing compute capacity constraints as a top concern. The fundraising reflects a broader industry arms race: Alphabet, Microsoft, Meta, and Amazon are expected to spend over $700 billion on AI infrastructure this year alone.
Alphabet to raise $80 billion from stock sales to fund AI buildout https://www.cnbc.com/2026/06/01/alphabet-to-raise-80-billion-from-stock-sales-to-fund-ai-buildout.html
I notice the provided text appears incomplete—it cuts off mid-sentence and doesn’t contain sufficient detail about what system this refers to, specific outcomes, or broader implications.
To produce an accurate two-line summary following your guidelines, I would need: – The name/identity of the AI system being evaluated – Concrete results from the scientific collaboration (e.g., validation status, clinical stage, concrete discoveries) – Timeframe and scope details – Why this matters relative to existing approaches Could you provide the complete material? That will allow me to deliver a factual, punchy summary that meets your editorial standards.
Over the past year, we’ve collaborated with global scientific experts to evaluate the system on complex problems. It assisted by: identifying new targets for liver fibrosis, uncovering fresh approaches to tackling Amyotrophic lateral sclerosis (ALS), and digested decades of https://x.com/GoogleDeepMind/status/2061857550438392094
AI Studio now lets developers build Gmail and Drive-connected apps directly.
Google has integrated email, cloud storage, and spreadsheet connections into its AI development platform, eliminating the need to switch between tools. This matters because it reduces friction in building AI applications that interact with widely-used workplace software, potentially accelerating development cycles for business-focused AI tools.
We just shipped the ability to build apps that connect to Gmail, Drive, Sheets, and more directly inside of @GoogleAIStudio, no navigating to other sites, you can add testers right inside of AI Studio, with full public sharing coming soon!! https://x.com/OfficialLoganK/status/2061568290984800740
U.S. and Japan launch one-billion-dollar AI science partnership.
The U.S. Department of Energy and Japan announced a $1 billion, five-year joint initiative—the first international partnership in Trump’s “Genesis Mission”—combining eleven research teams across American national labs and Japanese institutions to advance AI-driven breakthroughs in quantum computing, fusion energy, and autonomous laboratories. The collaboration pools computing resources including the DOE’s supercomputers and Japan’s Fugaku system, positioning both nations to accelerate scientific discovery by embedding AI into research workflows at scale.
United States and Japan Announce Historic $1 Billion Partnership Under President Trump’s Genesis Mission | Department of Energy https://www.energy.gov/articles/united-states-and-japan-announce-historic-1-billion-partnership-under-president-trumps
White House prioritizes AI security over regulation in new executive order.
The administration released an executive order emphasizing rapid AI innovation while addressing national security risks through voluntary government-industry collaboration rather than mandatory licensing or oversight. Key measures include hardening federal cybersecurity systems against AI-enabled threats within 30 days, creating a vulnerability-sharing clearinghouse between government and industry, and establishing a voluntary framework for assessing powerful AI models—explicitly avoiding mandatory preclearance requirements.
Promoting Advanced Artificial Intelligence Innovation and Security – The White House https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/
MAI releases seven new models including reasoning-focused text foundation model.
MAI announced a suite of seven models today, headlined by MAI-Thinking-1, a reasoning-focused large language model that scores 97% on advanced math benchmarks and outperforms Claude Sonnet 4.6 in blind human evaluations. The release also includes MAI-Image-2.5, now ranked second globally for image-to-image generation, and MAI-Transcribe-1.5 priced at $6 per 1,000 minutes, signaling the company’s push to compete across multiple AI categories with independently built models rather than adapted versions of competitors’ technology.
MAI-Image-2.5 is here — now #3 on text-to-image and #2 on image-to-image Arena leaderboards, surpassing Nano Banana Pro. Leading image generation. Precise editing. Built for enterprise scale. It delivers strong performance on H100s, enabling deployment on existing https://x.com/MicrosoftAI/status/2062240400299934143
MAI-Transcribe-1.5 is available at $6 per 1,000 minutes of audio via Microsoft Foundry. https://x.com/ArtificialAnlys/status/2061878498609053909
Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks. – It’s a https://x.com/mustafasuleyman/status/2061880164498428188
Today we announced MAI-Thinking-1, a strong generalist and reasoning LLM built from the ground up without distilling third-party models. 97% on AIME 2025; 53% on SWE-Bench Pro; preferred by human raters over Sonnet 4.6 (blind side-by-side). Tech report: https://x.com/asadovsky/status/2062008312603070891
Microsoft launches Web IQ search engine for AI agents.
Microsoft introduced Web IQ, a specialized search tool that lets AI systems access current, verified information from the web in real time. This matters because AI agents often struggle with outdated knowledge, and Web IQ aims to give them reliable, up-to-date data for more accurate decision-making. The move positions Microsoft to compete in the growing market of AI systems that need to interact with live information rather than rely solely on training data.
As Satya shared in today’s keynote at Build, we just launched Microsoft Web IQ – our next generation search engine for AI agents. It is a new suite of AI-native grounding APIs designed to connect AI systems with fresh, reliable web data for enhanced real-world intelligence https://x.com/JordiRib1/status/2061866606670581871
MiniMax releases M3 model with million-token memory and native video understanding.
MiniMax’s new M3 model combines three capabilities previously found only in closed-source frontier models: expert-level coding performance, the ability to process 1 million tokens of context (roughly 750,000 words), and native support for images and videos. The company achieved this through MSA, a new attention architecture that reduces computational complexity and enables 9–15× faster processing at ultra-long contexts. M3 demonstrates these capabilities through autonomous completion of complex tasks—including reproducing a published AI research paper over 12 hours and optimizing GPU code without human intervention—positioning it as the first open-weight model bringing all three capabilities together.
MiniMax M3: Frontier Coding, 1M Context, Native Multimodality — All in One Model – MiniMax Research | MiniMax https://www.minimax.io/blog/minimax-m3
OpenAI breaks ground on Michigan data center with union jobs and community benefits.
OpenAI and partners began construction on a 1-gigawatt data center in Saline, Michigan that will create over 2,500 union construction jobs and 450 permanent positions while generating $1 billion in projected tax revenue. The project includes commitments to shield local residents from increased electricity costs, protect water resources through closed-loop cooling, and invest $45 million in AI training credits for 400,000+ Michigan students—positioning the state as a hub for both AI infrastructure and workforce development as part of OpenAI’s broader Stargate infrastructure initiative.
Building the infrastructure for the Intelligence Age in Michigan | OpenAI https://openai.com/index/stargate-michigan-data-center/
AI tools give mathematicians permission to pursue riskier research paths.
Fields Medalist Terence Tao reports that AI assistance lets him test unconventional ideas without the penalty of wasted effort, fundamentally changing how mathematicians approach exploration. Rather than incremental progress on established problems, researchers can now afford to chase speculative directions—shifting the incentive structure of academic research itself.
AI can give researchers the freedom to pursue “crazier” ideas. For Terence Tao, AI creates more room to experiment, test unexpected paths, and discover what might otherwise stay out of reach. https://x.com/OpenAI/status/2060451757818601808
OpenAI upgrades ChatGPT memory to stay accurate across years.
OpenAI launched an improved “dreaming” system that automatically curates ChatGPT’s memory of user preferences and context from past conversations, addressing problems where stored information became stale or incorrect. The upgrade—rolling out this week to Plus and Pro users—represents a shift from manual memory-saving to a background process that learns from chat history, allowing ChatGPT to maintain relevant personalization over multi-year timeframes without users explicitly asking it to remember details.
Dreaming: Better memory for a more helpful ChatGPT | OpenAI https://openai.com/index/chatgpt-memory-dreaming/
AI voice control system shows hands-free computer operation capability.
A new voice interface demonstration suggests computers could be controlled entirely by spoken commands, with some observers highlighting the potential of real-time AI systems. While voice control itself isn’t new, the implication is that this latest version performs complex tasks without manual input. The practical significance depends on accuracy rates and real-world performance beyond controlled demos, which weren’t detailed here.
Watch me control my computer with just my voice. This is the future of operating systems. No hands. GPT-Realtime 2.0 is very, very underrated. Demo: https://x.com/FarzaTV/status/2060865350036750847
Scorsese publicly embraces AI image tool for film storyboarding work.
The legendary 91-year-old filmmaker has advised German AI startup Black Forest Labs and tested their FLUX image model for preproduction storyboarding, saying it lets him communicate visual ideas to his crew faster without sacrificing craft. This matters because it shows a major creative authority—who previously adopted 3D and digital de-aging—viewing AI as an evolution of cinema rather than a threat, though it may still provoke industry debate about technology’s role in filmmaking.
I see videos like this and get excited… it’s the old guard embracing new tech. Then I remember the polarizing reaction ahead – perhaps Scorsese is impervious to such pressures? https://x.com/bilawalsidhu/status/2061811752786944074
Legendary filmmaker Martin Scorsese signed on last year as an adviser to Black Forest Labs, the German AI startup behind FLUX image models. On Tuesday he went public, testing the tool on a single scene during preproduction. His use is narrow: storyboarding only, complementing https://x.com/TheRundownAI/status/2061834880917357011
Martin Scorsese × Black Forest Labs https://bfl.ai/martin-scorsese-bfl-advisor
Seeing Martin Scorsese using FLUX for storyboarding and scene exploration was absolutely insane. Experiencing how one of the absolute masters of cinema & filmmaking uses the technology that we developed, his curiosity and creativity, and the way he prompted our models, was https://x.com/robrombach/status/2061804823352086681
Qwen releases multimodal agent model combining vision and language capabilities.
Alibaba’s Qwen3.7-Plus unifies image and text processing in a single AI model designed to handle both visual interface navigation and command-line tasks, plus coding and productivity work. This matters because it simplifies how businesses deploy AI agents—rather than maintaining separate systems for different input types, one model handles multiple job categories. The unified approach suggests a shift toward more generalist AI tools that can work across traditionally separate domains.
👏👏 Introducing Qwen3.7-Plus — a multimodal agent model that unifies vision and language into one versatile agent foundation. ✅ Multimodal interactive hybrid agent: unified GUI & CLI operation across visual and text tasks ✅ Versatile coding agent & productivity assistant with https://x.com/Alibaba_Qwen/status/2061506641120641494
Softbank leads potential $800M funding round for German robotics startup.
SoftBank is negotiating to invest over $300 million in Munich-based Agile Robots, a hardware robotics company, as part of an $800 million funding round still in early stages. This signals growing venture capital appetite for physical automation companies beyond software AI, though deal terms remain fluid and subject to change.
SoftBank is in early talks to back a ~$800M (€700M) funding round for Munich-based Agile Robots, per Bloomberg. – SoftBank would write more than $300M of the round, sources told Bloomberg, though the talks are early and amounts and terms could still shift. – Agile Robots, https://x.com/TheHumanoidHub/status/2061864624006574580
AI researchers propose unified framework for “world models” across rendering, simulation, and planning.
Fei-Fei Li’s essay clarifies confusion around the overused term “world model” by identifying three distinct functions: renderers that produce realistic visuals, simulators that model physics and geometry accurately, and planners that predict actions—each serving different needs from entertainment to robotics training. The breakthrough insight is that these three functions share underlying knowledge about how the physical world works, suggesting they could eventually merge into a single foundation model that switches between visual, structural, and action outputs depending on the task. This matters because simulation, currently the least visible category, is the linchpin connecting human-focused visualization to machine-focused robotic learning, and mastering it could unlock multi-trillion-dollar applications across manufacturing, autonomous vehicles, and robotics.
A Functional Taxonomy of World Models – Dr. Fei-Fei Li https://drfeifei.substack.com/p/a-functional-taxonomy-of-world-models





Leave a Reply