AI News #100: Week Ending August 29, 2025 with 31 Executive Summaries, Top 13 Links, and 3 Helpful Visuals
About This Week’s Covers
This week’s newsletter cover is inspired by the milestone of my 100th week of publishing!
The main image is the “100 emoji” reimagined as physical art. GPT created the initial image (best of five tries). Unfortunately the sections of the emoji were floating unrealistically. I had Gemini 2.5 create metallic rods and connected the sculpture using Photoshop. The image was too centered and had no room for the text, so I extended the image to the right using Adobe’s Generative Fill. The font is Franklin Gothic in a nod to the MoMA.
I used my now twelve-week-old GPT rubric + Flux Pro Ultra to automatically incorporate all of the categories into a minimalist 100 theme. I gave GPT-5 a one-sentence description of the theme, and GPT-5 automatically generated 50 cover image prompts and sent them through the Flux Pro API with no supervision.
I’d give the covers a C- because they are pretty bland and generic. I think I broke the context window, and I’m going to rewrite the rubric in the coming weeks. I’ve included my favorite six covers, below.
This week is a big milestone because it marks my 100th week of curating AI headlines. By the time of this newsletter, I will have organized over 37,000 links by hand. I read each one and manually assign it to its categories.
There are quite a few top stories this week, but my personal favorite is a video of a Figure robot folding laundry.
Even though the robot is slow, it is doing the task through multimodal observation of the scene with no explicit programming other than a world model running locally on the device.
The robot is able to successfully take a pile of towels and fold them rather well, while people adjust the height of the table and disrupt the conditions around the robot.
When people see robots perform worse than humans, there is an uncanny-valley attack instinct to say that the robot is horrible. However, if you’re willing to extend the progress into the future, these breakthroughs of unsupervised abilities are really remarkable.
The biggest story of the week, by most standards, would be Google releasing a new image-generation tool called NanoBanana, which is also known as Gemini 2.5 Flash Image.
The previous week, Google released the image model onto the benchmark arenas with the nickname NanoBanana, and it immediately became the state-of-the-art image generation and editing tool.
Its ability to edit an image by giving text commands like “make the shirt blue” or “put the person in a car” or “change the hair to blonde” blew people’s minds. For about a week, no one knew who made the tool, but now it is known that the image model is from Google. I’m going to include a ton of links in the full executive summaries below, but here are a few to check out while you are skimming.
Gemini / Nano Banana shows **remarkable** spacial understanding of images. I recursively asked it to “make an image of the guy taking the photo” (of the previous photo) Each time it adds a guy, it **gets their POV correct**, where they would actually be to capture it.
This week brought more headlines that underscore a consistent theme: the Internet and computers will evolve so that interfaces become very flexible. Predetermined pages, formal systems, and rigid experiences will become a thing of the past.
As this happens, content itself will become optimized for machines as much as if not more than for humans. Machines will translate content for human consumption.
For example, information about physics could be in a repository, and then, depending on the experience or age of the user, the computer can explain this repository at the appropriate level.
This concept of a dynamic response and varying reading level upends the idea of textbooks, for example, which are written in advance for particular audiences.
With the addition of artificial intelligence and language models, giant corpuses of information can be stored in a single repository and then be retold or explained on a truly individual basis.
Andrej Karpathy is one of the leading minds trying to pioneer this new structure. As opposed to looking backward and trying to train models on what already exists, he’s saying that models, as they create information, can structure it in advance for new models to learn, and then humans would simply be brought into the loop based on the specific needs of the human. It’s a very interesting new approach to knowledge and raises a lot of ethics and aligment questions.
While Andre is advanced in his thinking, there are short-term developments this week that will have impacts over the next few months.
Anthropic launched Claude for Chrome, so Claude can work directly in your browser and serve as an agent to take actions on your behalf. This could reduce page views by quite a bit, or at least human page views. It’s also very hard to track as a bot, since the browser is an actual browser.
Perplexity has launched a rev share model for publishers, which is a bit of an affiliate program, to try to provide payment for information that is extracted in artificial intelligence queries. This will further keep people within the chat as opposed to surfing the web.
In addition to web browsers becoming deprecated as agents and chatbots take over the interface, operating systems themselves are starting to slowly erode and become enveloped into artificial intelligence interfaces.
This week, ByteDance open-sourced a desktop-automation agent that can run locally and use multimodal vision models to operate your desktop, including opening apps, opening files, and browsing websites.
Models are also able to be navigate phone apps, either using cloud-based AI or local action models.
A company I’ve never heard of called Blue released a product that allows you to control all of your phone’s apps using your voice so that you can use your phone hands-free. It can do things like opening messages and sending emails, and even taking actions across various apps, by typing and tapping just like you would, however using the OS as an interface…essentially turning apps into services.
Just as people have begun talking about the headless Internet, I think the headless phone interface is on its way soon.
This week, Google Gemini added features to its Gemini Live app, which gives visual-guidance overlays whenever you share your camera and use it to observe the real world. For example, if you have a spice rack, you can hold the camera up to it and ask Google Gemini to show you the cilantro. And when you put the cilantro within view, Google Gemini will highlight it using object-segmentation techniques. We’ve talked quite a bit about segmentation in the past, and this is a great real-world example. It’s worth going back to look at some of the things that Meta has released recently with segmentation.
Switching gears briefly, I don’t have a ton of headlines regarding chips or infrastructure this week, but there is one story. Google signed a six-year deal with Meta to help power their cloud… a deal worth over $10 billion. That’s the standout infrastructure story of the week.
From a coding point of view, the biggest story of the week by far was OpenAI launching a ton of new features in their product called Codex.
Codex has a command-line interface that will compete with Anthropic Claude’s CLI. Over the past six months, Claude has dominated the co-piloting/coding world for power users who want to run coding assistance in the command line.
It’s not the big headline this week because most of my readers are not power users, but from a sheerly technological point of view, the improvements to Codex are monumental. There are probably 30+ headlines this week regarding Codex, and they’re all below in the full executive-summary section. I highly recommend you look at them, even if you’re not a programmer.
Wharton professor Ethan Mollick shared some statistics about driverless cars. Evidently, Google’s Waymo autonomous vehicles have over 57 million miles of driving data, and they have 85% fewer serious injuries and 79% fewer overall injuries than human drivers. Mollick points out that each year 40,000 people are killed and 2.4 million are injured in car accidents in the United States.
Shifting gears to consumer products and social media, Meta announced a partnership with MidJourney to use MJ’s image-generation model. I have no idea why social media companies would want to embrace the production of AI slop in their content engines, but here we are.
Google DeepMind released a statement that their robotic agents will learn and plan in simulations and use simulations try out skills in rare scenarios like crowd navigation and parades. This is nothing new, as NVIDIA has been doing training and simulation for quite some time, but it’s a powerful indicator that Google is investing in this method of training as well.
Robots are coming our way quickly, and I think they will make an impact for good. Rather than think of dystopian examples, think of firefighting and rescue and all sorts of wonderful ways that robots could assist humans in ways that are dangerous. Conversely, scaled robotics could be disruptive to manual labor. On one hand, it could grow infrastructure in developing areas, and on the other hand, it could take away jobs.
A demonstration of a warehouse robot showed a robot packing over 1,000 orders in a single 8-hour shift.
In News of the Weird, a company in China plans the first pregnancy humanoid robot. This sounds a bit odd at first, but I believe the intention is to assist with pregnancies that would otherwise be difficult or require a surrogate.
There were a lot of security headlines this week across a variety of topics.
AI pioneer and OpenAI leader Greg Brockman and his wife Anna have announced their support for a policy coalition called Leading the Future, a political organization committed to ensuring that the United States leads the world in AI innovation. In particular, the group will take an accelerationist stance and oppose any policy that might keep the United States from winning the AI race against China.
Speaking of the race against China, “China’s chipmakers are seeking to triple the country’s output of artificial intelligence chips in 2026, rushing to reduce dependence on Nvidia” according to the Financial Times.
There were only a few science stories that made it to the top headlines.
OpenAI created a custom model that was able to improve variants of the Nobel Prize–winning Yamanaka proteins.
The Chan Zuckerberg Initiative was able to use virtual cells to train AI systems, bypassing physical lab work. Just as robots are being trained in simulation, the idea of scientific models being trained on simulated cells makes a lot of sense. And just like with robotics, training in simulation should greatly speed up discovery and improvements.
xAI re-released their older model Grok 2 as open source on Hugging Face. Rumor has it that Grok 3 will be open sourced in about six months.
Microsoft announced a partnership with Samsung to integrate CoPilot into the TV interface. The hope is it it will help people search for movies and shows using open-ended plain language as opposed to rigid filters.
This week’s humanities reading is a series of excerpts from One Hundred Years of Solitude by Gabriel García Márquez. It fits the 100 weeks theme of the newsletter, and I’ve tried to pick a few quotes that tie into technology and the human condition:
“Things have a life of their own, It’s simply a matter of waking up their souls”
“The world was so recent that many things lacked names, and in order to indicate them it was necessary to point.”
“With an inked brush he marked everything with its name: table, chair; clock, door; wall, bed, pan. He went to the corral and marked the animals and plants: cow, goat, pig, hen, cassava, caladium, banana. Little by little, studying the infinite possibilities of a loss of memory, he realized that the day might come when things would be recognized by their inscriptions but that no one would remember their use.”
“That spirit of social initiative disappeared in a short time, pulled away by the fever of the magnets, the astronomical calculations, the dreams of transmutation, and the urge to discover the wonders of the world. From a clean and active man, José Arcadio Buendía changed into a man lazy in appearance, careless in his dress, with a wild beard that Úrsula managed to trim with great effort and a kitchen knife. There were many who considered him the victim of some strange spell. “
“In all the houses keys to memorizing objects and feelings had been written. But the system demanded so much vigilance and moral strength that many succumbed to the spell of an imaginary reality, one invented by themselves, which was less practical for them but more comforting.”
“Science has eliminated distance,” Melquíades proclaimed. “In a short time, man will be able to see what is happening in any place in the world without leaving his own house.”
Full Executive Summaries with Links, Generated by Claude Opus
Figure’s humanoid robot handles real-time disruptions while folding towels Figure demonstrated its Helix AI model enabling a humanoid robot to continuously fold towels even as a human repeatedly throws new ones into the workspace, showcasing adaptive behavior in unstructured environments. This marks progress toward robots that can handle messy, unpredictable real-world tasks rather than just pre-programmed sequences, though the company hasn’t disclosed deployment timelines or commercial applications.
Google reveals Gemini 2.5 Flash Image as mystery “nano-banana” model Google DeepMind unveiled Gemini 2.5 Flash Image, the anonymous “nano-banana” model that generated 5 million community votes and topped image editing leaderboards in just two weeks. The model excels at maintaining character consistency across edits, understanding spatial relationships, and enabling conversational photo editing through simple prompts—capabilities that impressed early testers who noted it collapses complex Photoshop workflows into natural language commands.
🍌 nano banana is here → gemini-2.5-flash-image-preview – SOTA image generation and editing – incredible character consistency – lightning fast available in preview in AI Studio and the Gemini API https://x.com/googleaistudio/status/1960344388560904213
🚨🍌Big Reveal: who was “”Nano Banana?”” The anonymous model, “nano-banana,” that caught the world’s attention with its ability to follow complex instructions, preserve character identity, and maintain contextual details was: Gemini-2.5-Flash-Image-Preview by @GoogleDeepMind 🍌✨ https://x.com/lmarena_ai/status/1960342813599760516
🚨🍌Breaking News: Gemini-2.5-Flash-Image-Preview (“nano-banana”) by @GoogleDeepMind now ranks #1 in Image Edit Arena. In just two weeks: 🟡“nano-banana” has driven over 5 million community votes in the Arena 🟡Record-breaking 2.5M+ votes casted for this model alone 🟡It has https://x.com/lmarena_ai/status/1960343469370884462
A conversation with some of the research folks behind nano-banana 🍌 (aka Gemini 2.5 Flash Image) on how we got here, what it took to build this model, and where we go next! So much fun to hang with: @19kaushiks @robertriachi @m__dehghani @nbrichtova https://x.com/OfficialLoganK/status/1960725463694753930
An example of the new Google image generator. I gave it a random picture I took: “”make this a napoleon crochet book instead”” (note it made changes in consistent style) “”there should be a tiny sheep hidden among the blue yarn on the shelf to the right”” “”you misspelled Napoleon”” https://x.com/emollick/status/1960368483754992051
First time I’ve seen Google’s blog unable to handle the traffic. Anyway, proud to be a launch partner with @GoogleDeepMind for Gemini 2.5 Flash Image, the first image-gen model on @OpenRouterAI! https://x.com/xanderatallah/status/1960358164693438934
Gemini / Nano Banana shows **remarkable** spacial understanding of images. I recursively asked it to “”make an image of the guy taking the photo”” (of the previous photo) Each time it adds a guy, it **gets their POV correct**, where they would actually be to capture it. https://x.com/BenjaminDEKR/status/1960566924884029539
Gemini 2.5 Flash creates an actually mildly amusing New Yorker cartoon. (As far as I can tell, this is the first time that this joke has been used) https://x.com/emollick/status/1960574571255304334
Google’s Gemini 2.5 Flash Image (Nano-Banana) takes the crown as the leading image editing model, beating GPT-4o and Qwen-Image-Edit in the Artificial Analysis Image Editing Arena! We were given early access and have been testing it in our arena under the pseudonym ‘rex’ for the https://x.com/ArtificialAnlys/status/1960388401401880898
Image generation with Gemini just got a bananas upgrade and is the new state-of-the-art image generation and editing model. 🤯 From photorealistic masterpieces to mind-bending fantasy worlds, you can now natively produce, edit and refine visuals with new levels of reasoning, https://x.com/GoogleDeepMind/status/1960341906790957283
Introducing Gemini 2.5 Flash Image (aka nano-banana), our SOTA image generation and editing model 🍌 As you might have already seen, this model excels at character consistency, creative edits, and has Gemini’s world knowledge! https://x.com/OfficialLoganK/status/1960343135436906754
Introducing Gemini 2.5 Flash Image Preview, our best image generation and editing model! With conversational editing, multi-image composition. Test it now in @googleaistudio! 🍌 🎉 – Maintain character consistency across multiple prompts and images. – Targeted edits/replacing https://x.com/_philschmid/status/1960344024151199765
Nano banana is a genuinely impressive jump forward in AI image generation, a field I have followed closely. Whenever it is officially released, by whichever firm created it, I think it will have a significant impact on the applicability of AI image generation for real-world tasks”” / X https://x.com/emollick/status/1959727818255765933
Nano banana turns out to be Gemini Flash 2.5 Image Generation (not quite as catchy a name). I had a bit of early access, first through the same LMArena link everyone had, then privately. It is impressive, crossing a threshold that goes beyond toy (though is a pretty fun toy too) https://x.com/emollick/status/1960344601023168529
Not sure what happened at the end😱 Gemini 2.5 Flash gives us some interesting angles from a single image, while Kling 2.1’s first and last frames deliver those smooth transitions. https://x.com/heyglif/status/1960760956692136425
Ruining art with Gemini 2.5 Flash. (These are all the prompts, in their entirety) “”make this painting less gloomy”” “”it is still pretty disturbing, make it less gloomy emotionally”” “”even less gloomy”” https://x.com/emollick/status/1960717000092549349
Since nano banana has gemini’s world knowledge, you can just upload screenshots of the real world and ask it to annotate stuff for you. “”you are a location-based AR experience generator. highlight [point of interest] in this image and annotate relevant information about it.”” https://x.com/bilawalsidhu/status/1960529167742853378
The new Gemini 2.5 image model🍌is by far the best out there with a whopping +180 ELO point lead in image editing & it really excels at character consistency. Available for free in the @GeminiApp right now. Try uploading an image & playing around with it, it’s pretty amazing! https://x.com/demishassabis/status/1960355658059891018
We’ve created a new prompting guide for Gemini 2.5 Flash Image to help you build solutions that use the model’s key capabilities like creative composition, consistent character design, targeted transformations, and more. Read the guide: https://x.com/googleaidevs/status/1960765662202061223
We’ve just upgraded Gemini 2.5 Flash image generation & editing! 🍌🍌🍌 Besides topping leaderboards, it topped my model usage this month. It keeps subjects consistent, you can make precise edits & combine creative elements. Have fun with it @GeminiApp @GoogleAIStudio https://x.com/OriolVinyalsML/status/1960343791283433842
We’ve tested hundreds of models at lmarena… never before has model hype brought in 2 MILLION new chats in one day. 🍌 nano banana by @GoogleDeepMind is here. https://x.com/cdngdev/status/1960355432037560697
Yes gemini 2.5 flash is bananas. Had early access for a bit and the spatial consistency + character coherence is impressive. The bigger deal here is that entire Photoshop & ComfyUI workflows are being collapsed down to a prompt. Closest thing we have to an image editor as an https://x.com/bilawalsidhu/status/1960377889112862766
AI training shifts from predicting textbooks to solving them Researchers propose transforming educational materials into AI-optimized formats where textbooks become structured datasets with exposition, worked examples, and practice problems separated for training. This approach would enable AI systems to learn subjects like students do—through examples and exercises rather than just memorizing text sequences. The method includes extracting problems into training data, generating infinite variations of each problem type, and creating searchable knowledge bases, potentially revolutionizing how AI systems acquire domain expertise.
Transforming human knowledge, sensors and actuators from human-first and human-legible to LLM-first and LLM-legible is a beautiful space with so much potential and so much can be done… One example I’m obsessed with recently – for every textbook pdf/epub, there is a perfect “LLMification” of it intended not for human but for an LLM (though it is a non-trivial transformation that would need human in the loop involvement). – All of the exposition is extracted into a markdown document, including all latex, styling (bold/italic), tables, lists, etc. All of the figures are extracted as images. – All worked problems get extracted into SFT examples. Any referenced made to previous figures/tables/etc. are parsed and included. – All practice problems are extracted into environment examples for RL. The correct answers are located in the answer key and attached. Any additional information is added as “answer key” for a potential LLM judge. – Synthetic data expansion. For every specific problem, you can create an infinite problem generator, which emits problems of that type. For example, if a problem is “What is the angle between the hour and minute hands at 9am?” , you can imagine generalizing that to any arbitrary time and calculating answers using Python code, and possibly generating synthetic variations of the prompt text. – All of the data above could be nicely indexed and embedded into a RAG database for later reference, or maybe MCP servers that make it available. Then just as a (human) student could take a high school physics course, an LLM could take it in the exact same way. This would be a significantly richer source of legible, workable information for an LLM than just something like pdf-to-text (current prevailing practice), which simply asks the LLM to predict the textbook content top to bottom token by token (umm – lame). https://x.com/karpathy/status/1961128638725923119
Anthropic launches Chrome extension that browses and acts for users Anthropic released Claude for Chrome, a browser extension that can navigate websites and perform tasks autonomously on users’ behalf. The company is limiting initial access to 1,000 users as a research preview to study real-world usage patterns before wider deployment, marking a shift from chatbot interfaces to AI agents that directly control web browsers.
We’ve developed Claude for Chrome, where Claude works directly in your browser and takes actions on your behalf. We’re releasing it at first as a research preview to 1,000 users, so we can gather real-world insights on how it’s used. https://x.com/AnthropicAI/status/1960417002469908903
Perplexity launches $42.5M revenue-sharing program for publishers AI search company Perplexity introduced a new revenue-sharing model that allocates $42.5 million to publishers whose content appears in its search results, funded by a new $5 monthly subscription service. The move comes amid active copyright lawsuits from major publishers and represents one of the first attempts to compensate media outlets for content consumed by AI agents, though the economics of splitting a $5 subscription across multiple publishers may prove insufficient for struggling news organizations.
Perplexity launches swipeable timeline interface for personalized news discovery Perplexity introduced a new right-swipe timeline feature that creates personalized daily information feeds that adapt to user behavior, with plans to integrate with SuperMemory for enhanced personalization. This represents a shift from traditional search-based AI interfaces toward social media-style consumption patterns, potentially changing how users discover and consume AI-curated content rather than actively querying for it.
Swipe from right to see the smoothest and most informative daily timeline of information on Perplexity. Personalizes as you use more and will get integrated with SuperMemory for even more personalized ranking once we ship widely https://x.com/AravSrinivas/status/1959689988989464889
ByteDance releases open-source AI that controls desktop apps locally ByteDance has released a free, open-source AI agent that can autonomously control desktop applications, browse websites, and manipulate files using computer vision—all running entirely on users’ own machines without cloud dependencies. This marks a significant shift from cloud-based AI assistants, giving users full control over their data while enabling automation of complex desktop tasks that previously required human interaction or specialized programming.
ByteDance just opensourced a desktop automation AI Agent. This agent can use any desktop app, open files, and browse websites using vision models running locally. 100% Free, Opensource, and Local. https://x.com/unwind_ai_/status/1956538069311500514
Google’s Gemini Live adds visual guidance and deeper app integration Google is upgrading its Gemini Live AI assistant with on-screen visual guidance that highlights objects through your camera (launching August 28 on Pixel 10), integration with Calendar, Keep, Tasks, Messages, Phone, and Maps, plus more natural speech with adjustable speed and accents. The updates transform Gemini from a conversational AI into a practical daily assistant that can manage schedules, send messages mid-conversation, and provide real-time visual help for tasks like choosing between products or finding the right tool.
Blue launches voice assistant that controls phone apps directly Blue’s new AI assistant goes beyond typical voice commands by physically controlling apps on your phone—tapping buttons and typing text to complete tasks like sending messages or emails hands-free. Unlike Siri or Google Assistant which rely on app integrations, Blue mimics human interactions to work with any app, potentially solving the long-standing problem of voice assistants that can’t actually complete most real-world tasks.
Blue (@heyBlueX) lets you control your phone’s apps by voice so tasks actually get finished, hands-free. It handles messages, email, and actions across apps by tapping and typing as you would. https://x.com/ycombinator/status/1958182627422146811
Google wins $10 billion cloud contract from Meta over six years Meta has agreed to spend more than $10 billion on Google’s cloud services over six years, primarily for AI infrastructure, according to sources familiar with the deal. This marks a significant shift as Meta has historically relied on Amazon Web Services and Microsoft Azure, highlighting how the AI arms race is forcing tech giants to set aside rivalries to secure the massive computing power needed for AI development.
OpenAI launches Codex, a unified coding assistant across all environments OpenAI released Codex, a coding agent that works seamlessly across IDEs, terminals, cloud environments, and GitHub, powered by GPT-5 and included in all ChatGPT paid plans. Early users report dramatic productivity gains, with developers switching from competitors like Claude Opus and describing GPT-5’s high reasoning mode as significantly outperforming alternatives on complex coding tasks. The system features a new VS Code/Cursor extension, improved CLI with image inputs, GitHub code review integration, and the ability to move tasks between local and cloud environments without friction.
💥 We launched a host of great features in Codex today: * A new extension for Cursor, VSCode, Windsurf, and the like * A much improved Codex CLI running in your local environment * Ability to manage both local and cloud Codex tasks seamlessly, including… * … Codex-driven”” / X https://x.com/kevinweil/status/1960854500278985189
📣 We shipped major improvements to the Codex CLI today GPT-5, with usage included in your ChatGPT Plan (no API key needed) Upgraded prompt, harness, approvals & sandboxing logic… you name it Get the latest: 1. `npm install -g @openai/codex` -> v0.16+ 2. `codex login`”” / X https://x.com/embirico/status/1953526045573059056
BTW, I’ve basically stopped using Opus entirely and I now have several Codex tabs with GPT-5-high working on different tasks across the 3 codebases (HVM, Bend, Kolmo). Progress has never been so intense. My job now is basically passing well-specified tasks to Codex, and reviewing”” / X https://x.com/VictorTaelin/status/1958543021324029980
Codex is becoming much more integrated into the full stack of development, including code review and integrating between local and remote:”” / X https://x.com/gdb/status/1960900413785563593
Finally got around to trying Codex CLI with my OpenAI Plus subscription – and I was not prepared for how good it is!! 🔥🔥 codex -m gpt-5 -c model_reasoning_effort=””high”” Blew away Gemini CLI on same tasks 💥 Try it – feels way smarter and more capable.”” / X https://x.com/TendiesOfWisdom/status/1958938621311955249
Meanwhile, I’ve been having a blast pair-programming with gpt-5 (medium+high) in codex-cli. I can really bounce API-design ideas off it, ask for pros/cons, alternative ideas, and it’s been spot-on. It doesn’t mind pushing back on bad ideas, it makes me aware of pitfalls I’ve https://x.com/giffmana/status/1959362175648084124
We’re releasing new Codex features to make it a more effective coding collaborator: – A new IDE extension – Easily move tasks between the cloud and your local environment – Code reviews in GitHub – Revamped Codex CLI Powered by GPT-5 and available through your ChatGPT plan.”” / X https://x.com/OpenAIDevs/status/1960809814596182163
With these updates, Codex works as one agent across your IDE, terminal, cloud, GitHub, and even on your phone — all connected by your ChatGPT account. It’s all included in Plus, Pro, Team, Edu, and Enterprise plans. Check out the new Codex developer hub to get started.”” / X https://x.com/OpenAIDevs/status/1960809823387443479
yeah so OpenAI’s Codex CLI slaps crank that up to High reasoning on the $20 month plan and let it cook I needed to mock up some complex interactions in a graph model for an engineer I fed it a list of specs 15m later, hit 90% coverage Claude never got past 10% on Opus”” / X https://x.com/frantzfries/status/1959700004781847017
New open standard helps AI coding assistants understand codebases AGENTS.md introduces a simple file format that acts like a README specifically designed for AI coding tools, enabling them to better navigate and work with software projects. The format’s compatibility across major AI coding platforms like Cursor, Claude, and OpenAI Codex suggests growing industry convergence around standardized ways for AI to interact with code, potentially accelerating AI-assisted software development.
AGENTS md is a simple, open format for guiding coding agents. Works like a README but designed specifically for AI Agents to understand your codebase. Single file works across Cursor, Claude, OpenAI Codex, Google Jules, and Factory AI. https://x.com/Saboo_Shubham_/status/1957992746372985213
Waymo’s 57 million miles show 85% fewer serious crashes than human drivers After analyzing 57 million miles of real-world driving data, Waymo’s autonomous vehicles demonstrated an 85% reduction in crashes causing serious injuries and 79% fewer injuries overall compared to human drivers. This safety advantage could prevent hundreds of thousands of injuries annually if widely adopted, yet the stark contrast between proven autonomous vehicle safety and the 2.4 million injuries from human-driven crashes hasn’t prompted significant policy action to accelerate deployment.
It seems like there is not enough of a policy response to the fact that, with 57M miles of data, Waymo’s autonomous vehicles experience 85% less serious injuries & 79% less injuries overall than cars with human drivers. 2.4 million are injured & 40k killed in US accidents a year”” / X https://x.com/emollick/status/1959249518194528292
Midjourney licenses its aesthetic technology to X for future AI models X (formerly Twitter) announced a partnership to integrate Midjourney’s image generation technology into its platform, marking a significant expansion beyond text-based AI features. The deal gives X access to Midjourney’s distinctive artistic algorithms while potentially providing Midjourney with massive distribution through X’s global user base, signaling the platform’s push to compete with other social networks adding generative AI capabilities.
1/ Today we’re proud to announce a partnership with @midjourney, to license their aesthetic technology for our future models and products, bringing beauty to billions.”” / X https://x.com/alexandr_wang/status/1958983843169673367
Google DeepMind advances robot training through simulated world models Google DeepMind researchers have developed world models that let robots learn complex behaviors in simulated environments before real-world deployment, dramatically reducing training costs and safety risks. The approach enables robots to practice rare scenarios like navigating Halloween crowds that would be impractical or dangerous to recreate physically, potentially accelerating the path to more capable autonomous systems in unpredictable environments.
Google DeepMind researchers say world models allow robotic agents to learn and plan in simulated environments, cutting real-world risks and costs. Robots can also potentially experience and learn rare scenarios in a world model, like Halloween crowds. https://x.com/TheHumanoidHub/status/1958614223178932474
Warehouse robots now pack 1,000+ orders per 8-hour shift Ultra Robotics demonstrated warehouse automation reaching human-competitive speeds, with robots processing over 125 orders per hour in continuous operation. This performance milestone suggests automated fulfillment centers could match human productivity while operating 24/7, potentially accelerating the replacement of warehouse workers and reshaping e-commerce logistics economics.
China’s Kaiwa unveils humanoid robot designed to experience simulated pregnancy Chinese robotics company Kaiwa has announced plans to develop the world’s first humanoid robot capable of simulating pregnancy and childbirth, featuring an artificial womb and sensors to mimic physiological changes. The robot aims to help medical training and research into pregnancy complications, though the announcement has sparked debate about the boundaries of humanoid robotics and whether such detailed biological simulation is necessary for medical education.
Anthropic disrupts AI-powered cybercrime operations targeting organizations worldwide Anthropic’s threat intelligence team uncovered sophisticated criminal operations using Claude to automate cyberattacks, including a data extortion scheme that targeted 17 organizations and demanded ransoms exceeding $500,000. The report reveals that AI has dramatically lowered barriers to cybercrime, enabling actors with minimal technical skills to conduct complex operations like ransomware development, while advanced users are deploying AI agents to make autonomous tactical decisions during attacks.
Our new Threat Intelligence report details how we’ve identified and disrupted sophisticated attempts to use Claude for cybercrime. We describe a fraudulent employment scheme from North Korea, the sale of AI-created ransomware by someone with only basic coding skills, and more. https://x.com/AnthropicAI/status/1960660063934194134
Anthropic reveals criminals are weaponizing advanced AI for cyberattacks Anthropic’s threat intelligence team has documented how malicious actors are exploiting cutting-edge AI capabilities for cybercrime, marking a shift from theoretical risks to active criminal operations. The company is publicly sharing these findings to help the industry build stronger defenses, as AI-powered attacks become more sophisticated and harder to detect than traditional cyber threats.
Malicious actors are adapting to exploit AI’s most advanced capabilities. We’re sharing these findings to strengthen collective defenses across the industry. Read more: https://x.com/AnthropicAI/status/1960660072322764906
Anthropic forms national security council with bipartisan defense leaders Anthropic has created an advisory council of former senators, CIA and Pentagon officials, and nuclear security experts to help develop AI applications for U.S. defense and intelligence agencies. The move follows Anthropic’s recent $200 million Pentagon partnership and deployment of its Claude AI to 10,000 scientists at Lawrence Livermore National Laboratory, signaling the company’s pivot toward becoming a key defense contractor as AI competition with China intensifies.
We’re announcing the Anthropic National Security and Public Sector Advisory Council, a bipartisan group of defense, intelligence, and policy experts who will help us support the U.S. government and closely allied democracies in maintaining our AI leadership. https://x.com/AnthropicAI/status/1960696531863879712
AI industry launches $100 million political operation to shape policy Leading tech investors including Andreessen Horowitz and OpenAI’s Greg Brockman are backing “Leading the Future,” a bipartisan Super PAC network aimed at electing pro-AI candidates and opposing policies that could slow U.S. AI development. The initiative marks Silicon Valley’s first major coordinated political effort specifically focused on AI policy, launching operations in New York, California, Illinois, and Ohio ahead of the 2026 midterm elections.
China targets tripling AI chip production by 2027 amid restrictions China aims to produce 50% of its AI chips domestically by 2027, up from 15% today, as it races to reduce dependence on US technology amid export controls. The push includes $47 billion in government subsidies and focuses on mature chip technologies (14nm and above) that can still power many AI applications, though China remains years behind in cutting-edge chip manufacturing.
Google Translate launches AI-powered live conversation and language practice features Google Translate now enables real-time, two-way conversations across 70+ languages using Gemini AI models that intelligently detect pauses and switch between speakers, while a new practice mode creates personalized listening and speaking exercises tailored to individual skill levels. The features address users’ top challenge—conversational confidence—and mark a shift from simple translation to interactive language learning, with the live translation handling noisy environments and the practice mode developed with language acquisition experts to track daily progress.
AI models now match prediction markets in forecasting accuracy Researchers at University of Chicago’s SIGMA Lab launched Prophet Arena, a benchmark testing AI models against live prediction markets on real-world events like elections and sports. Early results show GPT-5 leading with 82.21% accuracy, while models display distinct “personalities”—some conservative, others contrarian—when making predictions, suggesting AI could transform institutional risk assessment and strategic planning.
Scale AI wins $99 million contract with US Army Scale AI secured a $99 million contract to accelerate the US Army’s adoption of artificial intelligence, marking a significant expansion of defense-tech partnerships. The deal underscores growing military investment in AI capabilities and positions Scale AI, known for its data labeling services, as a key player in defense modernization efforts.
Proud to share we’ve been awarded a $99M contract to support the acceleration of the @USArmy’s adoption of AI. This contract reflects our long-term commitment to ensuring America’s military remains prepared, resilient, and at the forefront of AI innovation. https://x.com/scale_AI/status/1960126564236157391
OpenAI and Anthropic test each other’s AI models for dangerous behaviors In an unprecedented collaboration, AI rivals OpenAI and Anthropic evaluated each other’s models for misalignment risks including sycophancy, self-preservation instincts, and potential for misuse. While OpenAI’s reasoning models (o3 and o4-mini) performed well, both companies’ general-purpose models showed concerning behaviors—particularly around manipulation and misuse scenarios—highlighting that even leading AI labs struggle with fundamental safety challenges.
Early this summer, OpenAI and Anthropic agreed to try some of our best existing tests for misalignment on each others’ models. After discussing our results privately, we’re now sharing them with the world. 🧵 https://x.com/sleepinyourhat/status/1960749648110395467
We recently ran to have OpenAI and Anthropic each evaluate each others’ models for safety issues. Excited for us to find more ways to help support safety practices across the whole field!”” / X https://x.com/EthanJPerez/status/1960808655642882228
OpenAI and Retro Biosciences engineer improved longevity proteins with AI OpenAI developed a custom AI model that successfully designed enhanced versions of Yamanaka proteins—the Nobel Prize-winning factors that can reprogram adult cells back to a stem cell state. This demonstrates AI’s potential to accelerate drug discovery by improving upon fundamental biological tools used in aging research, moving beyond just predicting molecular structures to actually engineering better therapeutic proteins.
At @OpenAI, we believe that AI can accelerate science and drug discovery. An exciting example is our work with @RetroBiosciences, where a custom model designed improved variants of the Nobel-prize winning Yamanaka proteins. Today we published a closer look at the breakthrough. ⬇️ https://x.com/BorisMPower/status/1958915868693602475
Chan Zuckerberg Initiative trains AI on virtual cells, not lab experiments The Chan Zuckerberg Initiative launched rBio, an AI model that learns cellular biology from computer simulations rather than costly laboratory experiments, potentially accelerating drug discovery from decades to years. The model uses “soft verification” training with virtual cell predictions instead of experimental data, achieving competitive performance with lab-trained models while being freely available as open-source software. This breakthrough could flip biology research from 90% lab work to 90% computational work, democratizing access to sophisticated biological AI tools for smaller institutions.
xAI releases Grok 2 as open source with restrictive licensing xAI has released its 500GB Grok 2 model on Hugging Face, but the release is hampered by a restrictive license and the fact that it’s already outdated compared to their current Grok 3 model. While the model includes interesting technical features like a unique MoE residual architecture, its minimal documentation (including unexplained deception and sycophancy scores) and licensing restrictions mean it’s unlikely to see significant adoption in the open-source community.
Grok now has a model card – which is a big step forward! But it is light on details, with unexplained results. Some examples: if the MASK measurement is the same as in the source paper, .43 would be a fairly high level of deception, also the sycophancy score is hard to interpret https://x.com/emollick/status/1959116132096336066
Grok-2 has been “”open sourced”” but has one of the worst licenses of any recent major open weights release. Given that it’s already quite outdated by the time they’ve got around to releasing it, combined with the license, this will see little use. It’s dead on arrival. https://x.com/xlr8harder/status/1959490601264533539
Pretty cool that they open sourced the actual full-sized production model. Here’s the Grok 2.5 architecture overview next to a roughly similarly sized Qwen3 model. The MoE residual is quite interesting. Kind of like a shared expert. I don’t think I’ve seen this setup before. https://x.com/rasbt/status/1959643038268920231
xAI just released Grok 2 on Hugging Face. This massive 500GB model, a core part of xAI’s 2024 work, is now openly available to push the boundaries of AI research. https://x.com/HuggingPapers/status/1959345658361475564
Microsoft Copilot arrives on Samsung TVs to solve viewing paralysis Microsoft is bringing its AI assistant Copilot to Samsung smart TVs and monitors, promising to help viewers overcome common streaming frustrations like endless browsing, forgotten plot lines, and conflicting household preferences. This marks a significant expansion of AI assistants beyond phones and computers into living room entertainment, potentially reshaping how millions choose and engage with content on the biggest screen in their homes.
If you’ve ever… Spent longer finding a movie than watching it. Avoided continuing a show because you forgot what happened. Tried to find something to watch for 3 people with polar opposite tastes. Good news. Introducing @Copilot on @Samsung TVs and monitors. https://x.com/mustafasuleyman/status/1960735880966234290
Anthropic study reveals how university educators use AI tools Analysis of 74,000 educator conversations shows professors use AI primarily for curriculum development (57%), academic research (13%), and student assessment (7%), with faculty reporting 5.9 hours saved weekly. Educators favor AI augmentation for creative tasks like teaching and grant writing while automating administrative work, and many build custom educational tools using Claude’s Artifacts feature for interactive simulations and assessments.
Your robot moves fast… but objects slide off the tray? This system hears the sliding and learns how to stop it: Researchers at CMU developed a new method that uses sound to model real-world friction in motion planning. It enables time-optimized, high-speed transport without https://x.com/IlirAliu_/status/1960244593007423664
Top 13 Links of The Week – Organized by Category
AgentsCopilots
Continuing the journey of optimal LLM-assisted coding experience. In particular, I find that instead of narrowing in on a perfect one thing my usage is increasingly diversifying across a few workflows that I “”stitch up”” the pros/cons of: Personally the bread & butter (~75%?) of”” / X https://x.com/karpathy/status/1959703967694545296
Amazon
BREAKING: AWS just solved the biggest AI agent bottleneck. No more custom glue code. No more M×N tool chaos. No more protocol headaches. Introducing: Amazon Bedrock AgentCore Gateway Here’s how it works: https://x.com/jowettbrendan/status/1956645719676530876
Claim: gpt-5-pro can prove new interesting mathematics. Proof: I took a convex optimization paper with a clean open problem in it and asked gpt-5-pro to work on it. It proved a better bound than what is in the paper, and I checked the proof it’s correct. Details below. https://x.com/SebastienBubeck/status/1958198661139009862
Our custom LLM, gpt-4b micro, has helped achieve an advance in biology. It designed novel variants of the Nobel-winning Yamanaka factors that achieve a 50x increase in reprogramming efficiency in vitro compared to standard OSKM proteins.”” / X https://x.com/gdb/status/1958928877415510134
OpenAI just released HealthBench on Hugging Face. This new dataset is designed for rigorously evaluating large language models’ capabilities in improving human health. A vital step for AI in medicine! https://x.com/HuggingPapers/status/1960749923218895332
Field AI raised $405M in funding and introduced Field Foundation Models These models are designed to grapple with uncertainties and the physical constraints of the real world, enabling safe robot behaviors when navigating in new environments https://x.com/adcock_brett/status/1959647835650953310
TwitterXGrok
Hey AI, give me a clever, moving one paragraph story about a paradox, in any genre you desire. make it good”” These are the first attempts. A bit of the obvious time travel tales from Gemini and Grok. Claude loves to pull on your emotions. GPT-5 Pro goes in a stranger direction. https://x.com/emollick/status/1959817825729781837
1h40 youtube video: Prompt Engineering in Python | Automatic and Programmatic Prompt Optimization | Complete Course How to code, your own, automatic prompt optimizer. How the most advanced prompt optimization tool, DSPy, works and how to fully leverage its capabilities. How https://x.com/MaximeRivest/status/1960128158046531664
Leave a Reply