About This Week’s Covers

This week’s cover references OpenAI’s Operator, which allows ChatGPT to browse the web like a human. The image is a twist on the classic Sade album Diamond Life, home of the single “Smooth Operator.” Sade is replaced with a robot.

The rest of the covers were created using Claude 3.5 and the Ideogram API, with the theme “‘a sophisticated but deceptive charmer who moves through high society. Use a striking red background lighting’ + category name.”

This Week’s Executive Summaries

Quotes Of the Week

Andrej Karpathy: “Projects like OpenAI’s Operator are to the digital world as Humanoid robots are to the physical world. One general setting (monitor keyboard and mouse, or human body) that can in principle gradually perform arbitrarily general tasks, via an I/O interface originally designed for humans. In both cases, it leads to a gradually mixed autonomy world, where humans become high-level supervisors of low-level automation.”
Karpathy

Morgan Stanley: “Humanoid Hype? Barely 6 months following the publication of our Humanoid Blue Paper, we are receiving more incoming client requests to discuss humanoids than the entirety of our auto OEM, dealer and supplier coverage combined.”
Adcock_brett 

Dr. Jim Fan of the NVIDIA Robotics Lab: “Whether you like it or not, the future of AI will not be canned genies controlled by a “safety panel”. The future of AI is democratization. Every internet rando will run not just o1, but o8, o9 on their toaster laptop. It’s the tide of history that we should surf on, not swim against.”
DrJimFan  

Sam Altman: Advancing AI may require “changes to the social contract. The entire structure of society will be up for debate and reconfiguration.” 
TFTC21

Ethan Mollick: “Most important implications of Operator: 1) General purpose web agents aren’t there yet, but seem more workable than expected (Operator is quite good) 2) Companies aren’t thinking enough about how to market to the preferences of agents 3) Security is going to get very weird fast” 
Emollick

DeepSeek

DeepSeek Releases Multimodal Image Model: Challenges Industry Leaders in Image Generation
DeepSeek has released Janus-Pro-7B, an open-source system that can both understand and create images. Early tests show it performing better than popular image generators like DALL-E 3 and Stable Diffusion on key benchmarks. What sets the system apart is that it processes visual information through separate pathways while using a single unified brain (which I believe is the key to true multimodality).
Rowancheung | nrehiew_

Open Source Community Launches Project to Recreate DeepSeek R1
HuggingFace has announced “Open-R1,” a community project to recreate DeepSeek’s recently released AI reasoning system from scratch. DeepSeek made waves last week by releasing their R1 model, which matches or exceeds the performance of similar systems like OpenAI’s, but kept some key details under wraps. While DeepSeek shared their AI model publicly, they didn’t release the training data or code needed to fully reproduce their results. HuggingFace’s project aims to fill these gaps by reverse-engineering the missing pieces and sharing them with the public. The goal is to help researchers and companies build better AI systems that can tackle complex reasoning tasks like math problems and coding challenges, while making the entire process transparent to the wider AI community.
huggingface

Operator and Agents

Perplexity Launches AI Assistant to Handle Everyday Digital Tasks
Perplexity released a digital assistant that helps with online activities, combining web browsing capabilities with task execution. The app, only available on the Google Play Store, claims to handle everything from restaurant bookings to reminders, with the ability to maintain context across multiple requests. It’s multimodal, so users can engage through text or their device’s camera.
Perplexity_ai 

A few examples of OpenAI’s Operator browser-using automation:
“Scheduling an appointment with my barber after looking at my Google Calendar schedule/availability Note that in this demo, ChatGPT Operator pinged me that I needed to sign in to Google to check my calendar I tried a second time, and my login was saved session-to-session”
rowancheung 

“Finding a top-rated dog walker in Vancouver BC This is no easy task, so I wanted to test how well ChatGPT Operator could handle it To my surprise, I got 3 really solid options at the end” 
rowancheung

“Booking a one-way flight from Zurich to Vienna using the Booking integration This one required a bit of back and forth, with ChatGPT Operator pinging me and asking for my flight preference and having me take control of entering payment details”
Rowancheung 

Leaping AI Aims to Automate Call Centers with Self-Learning Voice Agents
Y Combinator startup Leaping AI looks to automate customer service with AI voice agents that can handle entire call centers, achieving 70% automation rates while maintaining 90% customer satisfaction. The system is already managing thousands of daily calls for beta clients. In a standout case, German wine merchant Hawesko replaced over 100 human agents with AI voices that independently handle customer support calls, including complex tasks like wine purchasing recommendations. According to Leaping AI, the AI agents continuously improve their performance through automatic analysis and self-adjustment.
ycombinator

Meta AI Adds Memory and Personalization Features – On Their Way to Agents
Meta AI can now remember user preferences shared in private chats and offer personalized recommendations based on users’ social media activity. The assistant learns details like dietary preferences and interests during conversations, then uses this information along with Facebook profile data and content viewing patterns to provide tailored suggestions. Currently available in the US and Canada on Facebook, Messenger, WhatsApp, and Instagram, the memory feature works only in one-on-one conversations with the AI and allows users to delete stored information at any time (wink).
fb

Alibaba’s Mobile AI Agent Can Use A Phone and Learns From Past Actions
Alibaba launched Mobile-Agent-E, an AI system that helps users complete complex smartphone tasks by learning from its previous experiences. The system uses a manager-worker structure where a central planner breaks down tasks while specialized agents handle visuals, actions, error checking, and memory. Early tests show it performs 22% better than existing mobile AI assistants at handling multi-step tasks across different apps.  This seems like pretty big news. Still waiting for Apple to make their move!
_akhaliq | huggingface

OpenAI + US Government 

OpenAI Launches ChatGPT Gov for Federal Agencies
OpenAI introduced ChatGPT Gov, allowing U.S. government agencies to deploy ChatGPT Enterprise features within secure Microsoft Azure cloud environments. Over 90,000 users across 3,500 government agencies have already used ChatGPT, with early adopters including the Air Force Research Laboratory and Los Alamos National Laboratory. Pennsylvania state employees reported saving nearly two hours per day on routine tasks using the system, while Minnesota’s translation office leverages it to serve multilingual communities.
openai

OpenAI Partners with U.S. National Labs to Advance Scientific Research
OpenAI is providing its advanced AI reasoning models to scientists at Los Alamos, Lawrence Livermore, and Sandia National Laboratories. The partnership will deploy OpenAI’s technology on Venado, a specialized NVIDIA supercomputer at Los Alamos, to support research in materials science, renewable energy, astrophysics, and nuclear security. The collaboration aims to strengthen U.S. tech leadership by helping scientists tackle challenges in disease prevention, cybersecurity, energy infrastructure, and threat detection. 
Openai 

Anthropic

Anthropic CEO: AI Could Help Double Human Lifespan by 2030
Anthropic CEO Dario Amodei predicts AI will accelerate biological research so significantly that human lifespans could double within five years. Speaking at the World Economic Forum, Amodei explained that AI systems are approaching PhD-level capabilities in mathematics, programming, and biology, potentially compressing 100 years of scientific progress into a decade. While optimistic about AI’s impact on disease treatment and healthcare, he acknowledged that real-world factors like clinical trials and regulations will affect the pace of implementation.
observer

Anthropic CEO: AI Export Controls Still Critical Despite DeepSeek’s Advances
Anthropic CEO Dario Amodei argues that recent AI advances by Chinese company DeepSeek strengthen the case for U.S. chip export controls to China. While DeepSeek demonstrated strong performance with their V3 and R1 models, Amodei explains their achievements reflect expected industry cost reductions rather than a fundamental breakthrough. He notes that DeepSeek’s reported 50,000-chip infrastructure cost around $1 billion – comparable to U.S. AI labs. Looking ahead to 2026-2027, Amodei emphasizes that developing superintelligent AI will require millions of chips and tens of billions of dollars. He contends that export controls remain vital in determining whether the future will be a “bipolar world” with both U.S. and China having advanced AI capabilities, or a “unipolar world” where the U.S. and allies maintain a strategic advantage.  This is at odds with the “open source call to action” from Dr. Jim Fan at NVIDIA, at the top of this week’s executive summaries.
darioamodei

Top 74 Links of The Week – Organized by Category

Agents and Copilots

“This Qwen2.5-VL looks exciting! – strong and general vision capabilities – agentic features to support computer/phone use – long video understanding & capturing events – visualize localization – generated structured outputs Opens both base and instruct models in 3 sizes: 3B, 

“🚨 Anthropic CEO Dario Amodei plans to build a “virtual collaborator” this year – an AI agent that operates on your computer or at work – it will handle tasks like writing and testing code, communicating with coworkers on platforms like Slack or Google Docs 

“You don’t need to pay $200 for AI. We’re launching Open Operator – an open source reference project that shows how easy it is to add web browsing capabilities to your existing AI tool. It’s early, slow, and might not work everywhere. But it’s free and open source! 🔗👇 

“🤖 Just watched an AI agent autonomously browse a news site, find & summarize Oscar nominations! Open-source alternative to pricey AI tools with @BrowserUse @LangChainAI @Gradio with GPT-40 model. (can’t wait to have the option to use open-source ones) 

“ByteDance dropped UI-TARS, an open vision-language model outperforming Claude computer-use and GPT-4o on 10+ GUI agent benchmarks It processes screenshots as input and performs human-like interactions 

“We’ve released the SOTA open multimodal model, Qwen2.5-VL! It shows significant improvements across various aspects compared to the previous version. I’m thrilled to have made contributions to the Agent section of Qwen2.5-VL. The journey continues—see you tomorrow! 

“Announcing Qwen2.5-VL Cookbooks! 🧑‍🍳A collection of notebooks showcasing use cases of Qwen2.5-VL, include local model and API. Examples include Compute use, Spatial Understanding, Document Parsing, Mobile Agent, OCR, Universal Recognition, Video Understanding. 

Major misunderstanding about AI infrastructure investments: Much of those billions are going into infrastructure for *inference*, not training. Running AI assistant services for billions of people requires a lot of compute. Once you put video understanding, reasoning, large-scale memory, and other capabilities in AI systems, inference costs are going to increase. The only real question is whether users will be willing to pay enough (directly or not) to justify the capex and opex.

“Introducing /extract The era of writing web scrapers is over. Write a prompt, get web data. That’s it. Now in open beta with 500K free tokens 🔥 

AGI (Artificial General Intelligence)

“13 months later: one more “there’s an important missing perspective” tweet*, this time on the R1 conversation. R1 and o1 are incredibly impressive models, but you gotta admit that they’re not really *that* big of a leap relative to many timelines. We’re just baking in CoT into” / X

“The broader picture of the last month is not DeepSeek or StarGate or whatever but that all the technical & investment signals are now pointing to an expectation of continued rapid acceleration of AI capabilities. And we still lack a clear articulation about what that looks like.” / X

“I wonder if worrying about beating out other players in open AI models only really makes sense if you think that the end goal is to produce a closed AGI model that gives you advantages over everyone else. Otherwise what is the value of investing in your own model vs using one?” / X

Audio

Riffusion’s free AI music platform could be the Spotify of the future | VentureBeat

“We’ve raised a $180M Series C to give every AI agent a voice. The past year has been about building the foundations of AI audio – now we’re focused on making speech the standard for how we interact with technology. 

Ethics/Legal/Security

Law School Now Requires Students To Get Artificial Intelligence Certification – Above the Law

“visited @Helion_Energy today. the machine is making rapid progress (and the scale is nuts)–it feels like walking through a sci-fi movie!” / X

“Just fyi, @deepseek_ai collects your IP, keystroke patterns, device info, etc etc, and stores it in China, where all that data is vulnerable to arbitrary requisition from the 🇨🇳 State. From their own privacy policy: 

Trump revokes Biden executive order on addressing AI risks | Reuters

Multimodality

“This Qwen2.5-VL looks exciting! – strong and general vision capabilities – agentic features to support computer/phone use – long video understanding & capturing events – visualize localization – generated structured outputs Opens both base and instruct models in 3 sizes: 3B, 

“Announcing Qwen2.5-VL Cookbooks! 🧑‍🍳A collection of notebooks showcasing use cases of Qwen2.5-VL, include local model and API. Examples include Compute use, Spatial Understanding, Document Parsing, Mobile Agent, OCR, Universal Recognition, Video Understanding. 

Relightable Full-body Gaussian Codec Avatars

OpenAI

“Been playing with the new Operator for a little bit before launch and it is both very much still an experiment and also a good indicator of where things are going. It goes on the web and does things for you. Still many rough edges but here is an example of using it for shopping. 

“NEW VIDEO: Hands-on review of OpenAI Operator, their first AI agent that can control your browser. As OpenAI moves away from open source, Chinese labs like DeepSeek and ByteDance are becoming the new champions of open AI development, and doing it with access to a fraction of the 

Exclusive | OpenAI in Talks for Huge Investment Round Valuing It at Up to $300 Billion – WSJ

“The new o1 canvas is quite impressive: “create an interactive tool that visually shows me how correlation works, and why correlation alone is not a great descriptor of the underlying data in many cases. make it accessible to nonmath people and highly interactive and engaging” 

“Canvas update: today we’re rolling out a few highly-requested updates to canvas in ChatGPT. ✅Canvas now works with OpenAI o1—Select o1 from the model picker and use the toolbox icon or the “/canvas” command ✅Canvas can render HTML & React code” / X

“BREAKING: OpenAI has begun using the official DeepSeek R1 API to generate synthetic data for distilling and training its next-generation models. Sources say OpenAI employees have been discouraging public use of the R1 API, likely to divert traffic and prioritize their internal 

“ok we heard y’all. *plus tier will get 100 o3-mini queries per DAY (!) *we will bring operator to plus tier as soon as we can *our next agent will launch with availability in the plus tier enjoy 😊” / X

“big. beautiful. buildings. 

“A revolution can be neither made nor stopped. The only thing that can be done is for one of several of its children to give it a direction by dint of victories. -Napoleon” / X

“6. Researching a good birthday gift for my mom based on what she likes Similar to the Reddit block, ChatGPT Operator couldn’t access NYTimes, so it pivoted and found another site. Really neat. Also cool to see it compare and find the best price across the web for me, too 

“7. Booking a one-time house cleaner for my home through the Thumbtack integration based on my budget ChatGPT Operator came back to me with four highly rated options within my price range 

“8. Finding the best/cheapest health insurance coverage in Switzerland This was interesting since most prices are not publicly available and are gated behind a meeting ChatGPT Operator did what it could, and presented me with a good blog for me to read further 

Open Source

“This is a big deal – the first open model (and non-US model) to get near the top of the AI leaderboard. While all benchmarks have flaws, the LM arena is an important one based on head-to-head comparisons by many users.” / X

“@PalmerLuckey > The $5M number is bogus. It is pushed by a Chinese hedge fund to slow investment in American AI startups Palmer, I offer you a way out. Look at this string, talk to your ML friends, and admit you have been misled. Otherwise, reveal that you’re an innumerate jingoistic fraud. 

“DeepSeek is a really good model, but it is not generally a better model than o1 or Claude. But since it is both free & getting a ton of attention, I think a lot of people who were using free “mini” models are being exposed to what a early 2025 reasoner AI can do & are surprised” / X

“I don’t have too too much to add on top of this earlier post on V3 and I think it applies to R1 too (which is the more recent, thinking equivalent). I will say that Deep Learning has a legendary ravenous appetite for compute, like no other algorithm that has ever been developed” / X

“those who think RL use less compute don’t know RL at all 😅 SFT: human generates data and machine learns RL: machine generates data and machine learns” / X

“Just fyi, @deepseek_ai collects your IP, keystroke patterns, device info, etc etc, and stores it in China, where all that data is vulnerable to arbitrary requisition from the 🇨🇳 State. From their own privacy policy: 

“DeepSeek showed us in just 4 days: – Open-source AI is only <6 months behind closed AI – China is leading the open-source AI race (was not on my bingo card) – we are entering the LLM RL golden era – distilled models are powerful, we’ll have highly intelligent AI running locally” / X

“An obvious, “we are so back” moment in the AI circle somehow turned into “it’s so over” in mainstream. > unbelievable shortsightedness > the power of o1 in the palm of every coder’s hand to study, explore, and iterate upon > ideas compound > the rate of compounding accelerates” / X

“Announcing Qwen2.5-VL Cookbooks! 🧑‍🍳A collection of notebooks showcasing use cases of Qwen2.5-VL, include local model and API. Examples include Compute use, Spatial Understanding, Document Parsing, Mobile Agent, OCR, Universal Recognition, Video Understanding. 

Grok-3 model from xAI spotted ahead of anticipated release

“These four points on DeepSeek seem very likely correct and important to understand about the economics of building AI models and what DeepSeek actually did. . 

“🚀 The open source community is unstoppable: 4M total downloads for DeepSeek models on @huggingface, with 3.2M coming from the +600 models created by the community. That’s 30% more than yesterday! 

“Pretty cool that with the new Qwen 2.5 models you can ask questions / generate using a reasonably sized code-base as context, all running on a laptop with mlx-lm. The 7B runs pretty fast on an M4 Max using the mlx-lm code base (~16k lines) as context: 

“This is the clearest breakdown you’ll find about DeepSeek R1, from how it works to why it crushed math problems, and how @huggingface plans to reproduce it openly 

“What do DeepSeek and open-source AI mean for journalism? Plenty of insights in this interview with @_KarenHao, plus some snippets by yours truly! 

Qwen2.5-Max: Exploring the Intelligence of Large-scale MoE Model | Qwen

“Everyone: DeepSeek just appeared out of nowhere! 😱 Me: – DeepSeek Coder in 2023 – MoE in Feb – Math in Feb – VL in March – V2 in May – Coder V2 in June – Prover in August – V2.5 in September – VL 2 in December – V3 in December They’ve consistently shipped for 1+ years 😁” / X

“Yes, DeepSeek R1’s release is impressive. But the real story is what happened in just 7 days after: Original release: 8 models, 540K downloads. Just the beginning… The community turned those open-weight models into +550 NEW models on @huggingface. Total downloads? 2.5M—nearly 

“Perplexity with DeepSeek R1 is really cool, but I wish I could inspect the chain of thought in full — it passes by so quickly and I can’t go back to look at it. The uncensored CoT is the coolest part about R1 and this would be a simple UX change. Anyone else feel the same? 

“One of the most overlooked contributing factors to the success of DeepSeek-R1 is the Mixture-of-Experts (MoE) base model from which it is derived–DeepSeek-v3… The recently-proposed DeepSeek LLMs–including DeepSeek-R1, DeepSeek-v3 and more–have made waves within LLM research for 

“The main two implications of DeepSeek are going to be (1) an acceleration of AI model development until they hit a wall and (2) increased likelihood that there is no wall in the near future” / X

DeepSeek displaces ChatGPT as the App Store’s top app | TechCrunch

“Two days ago we released Qwen2.5-Max, and today we update the price of its API: Input tokens: $1.6 / million tokens Output tokens: $6.4 / million tokens We hope this change provides you some help and we’d love to hear your feedback about our new model! 👉🏻” / X

U.S. Navy bans use of DeepSeek AI: ‘Imperative’ to avoid using

“You can now run inference directly on Hugging Face model pages – powered by Together AI! 

“R1 be like: this is, without any exaggeration, the most impressive LLM philosophizing I’ve seen ever. It not just surpasses Opus. It dissects Opus, like a professor would a student’s essay. 

Welcome to Inference Providers on the Hub 🔥

“TinyZero reproduction of R1-Zero “experience the Ahah moment yourself for < $30″ Given a base model, the RL finetuning can be relatively very cheap and quite accessible.” / X

Publishing

Copyright and Artificial Intelligence, Part 2 Copyrightability Report

“Really good to see the Copyright Office moving towards understanding the full extent of what AI tools can offer to the industry. “The use of AI tools to assist rather than stand in for human creativity does not affect the availability of copyright protection for the output,”” / X

“Next big thing for brands: knowing what brands agents prefer. If you ask for stock prices, Claude with Computer Use goes to Yahoo Finance while Operator does a Bing search Operator loves buying from the top search result on Bing. Claude has direct preferences like 1-800-Flowers 

“CNN gears up for major digital pivot, including: – Cutting 200 TV jobs while hiring 200 for tech roles – Ramping up vertical video (50-100 daily) “biggest steps to overhaul the company, steering it away from its reliance on traditional TV and trying to cash in on digital 

“Most important implications of Operator: 1) General purpose web agents aren’t there yet, but seem more workable than expected(Operator is quite good) 2) Companies aren’t thinking enough about how to market to the preferences of agents 3) Security is going to get very weird fast” / X

“Introducing Perplexity Assistant. Assistant uses reasoning, search, and apps to help with daily tasks ranging from simple questions to multi-app actions. You can book dinner, find a forgotten song, call a ride, draft emails, set reminders, and more. Available on Play Store. 

“This is a huge victory for A.I. Filmmakers across the world. The Copyright Office has officially declared what I’ve been preaching for so long. Work that has been manipulated and transformed by humans is indeed copyrightable, even if the tools used were rooted in A.I. . This is a” / X

“🤖 Just watched an AI agent autonomously browse a news site, find &amp; summarize Oscar nominations! Open-source alternative to pricey AI tools with @BrowserUse @LangChainAI @Gradio with GPT-40 model. (can’t wait to have the option to use open-source ones) 

Oscar nominations: AI-enhanced films score nods for acting, editing

Robotics and Embodiment

WATCH: PPB unveils robotic dog for ‘dangerous tactical incidents’

“Unitree is preparing for the RoboCup Humanoid League 2025 (July) in Salvador, Brazil. The soccer upgrade of the Unitree G1 features a 180° depth camera and additional two DoF in the neck. The development kit includes an RL framework integrating Isaac Sim or MuJoCo, along with 

“JUST IN: 🇨🇳 China opens training base for humanoid robots in Shanghai. 

Video

“**NEW PIKA 2.1 MODEL OUT NOW** ————— “WHAT ONCE BURNED” In the flicker of the fire, burning stories in the flames of time, They linger through the ages, neither alive nor expired. For now they rest, yet still they grieve. Should their memories ever leave? They 

https://x.com/Oranguerillatan/status/1883923157007913248

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading