About This Week’s Covers
The category cover theme for this week’s AI newsletter is The White Lotus Season 3. I had Claude 3.7 automatically write prompts for 40 categories and send them to Black Forest Labs’ Flux Ultra as a batch. One Python command generates all the images. Keep in mind that if I simply wanted to create a great image with AI, I could iterate through prompts and achieve solid results within a few runs. This is purposefully done the hard way.
Every week, my category images are first tries, no redos. They are complicated and even include their own text in the prompt. For example, MidJourney and DALL·E 3 simply can’t handle them (total disaster). Ideogram and Flux are the two best for this type of prompt. The Thai tapestry theme based on the White Lotus open credits turned out pretty well. With a few refinements in the prompts and touch-ups in Photoshop, these would be solid.
Here are my favorite six of the forty category covers this week:

This Week By The Numbers
Total Organized Headlines: 453
- AGI: 16 stories
- Agents and Copilots: 198 stories
- Anthropic: 44 stories
- Apple: 3 stories
- Audio: 15 stories
- Augmented Reality (AR/VR): 26 stories
- Autonomous Vehicles: 2 stories
- Business and Enterprise: 32 stories
- Chips and Hardware: 59 stories
- Education: 2 stories
- Ethics/Legal/Security: 19 stories
- Figure: 10 stories
- Google: 34 stories
- Images: 18 stories
- International: 28 stories
- Llama: 6 stories
- Locally Run: 16 stories
- Meta: 8 stories
- Microsoft: 1 story
- Mistral: 16 stories
- Multimodal: 36 stories
- NVIDIA: 49 stories
- Open Source: 50 stories
- OpenAI: 52 stories
- Perplexity: 13 stories
- Podcasts/YouTube: 2 stories
- Publishing: 6 stories
- RAG: 5 stories
- Robotics Embodiment: 60 stories
- Science and Medicine: 15 stories
- Technical and Dev: 21 stories
- Video: 22 stories
- X: 11 stories
This Week’s Executive Summaries
This week was especially busy because Nvidia had their big conference which included a lot of product announcements. There was also a lot of agent news across the board. I organized over 400 headlines and 21 of them merited executive summaries.
I’m noticing a theme of Nvidia trying to ensure that they remain the computing chip of choice by giving away a ton of software and AI models that are optimized for their chips. It makes sense that Nvidia would open source as much as they can to get a lock on the ecosystem so people are, for lack of a better term, addicted to their chips.
Nvidia released two open models that allow AI agents to reason. Nvidia also released two major simulators for robotic training in virtual reality. They also introduced a new foundation model for robots to understand their environment through vision, including interacting with things they’ve never seen before. Nvidia also released a data platform to help companies maximize the integration between their data and AI tools through a special Nvidia data architecture. Lastly, Nvidia released two open-sourced multilingual speech-to-text and text-to-speech models, which also support translation.
Anthropic has given its AI the ability to search the web, which is a big improvement; however, it’s a little bit disappointing that you can’t paste in a web address and tell it to browse or refer to it. Seems a bit limited.
Manu, the extensively hyped Chinese reasoning agent, has entered its third week of attention with many people getting accepted into the beta group and coming back with stories of incredible capabilities. For example, this week someone ran a professional-level stock analysis in one hour that would’ve taken about two weeks for a human to accomplish.
A new study came out that shows that AI agents are doubling their abilities every seven months. This is kind of a new Moore’s Law for agents. It implies that within five years, AI agents will be able to quickly accomplish things that used to take humans weeks.
OpenAI has continued evolving their agent API with a new SDK that allows anyone to integrate their service with language models through structured interfaces. This is potentially a huge deal to allow plain language interaction with agents that can complete complicated transactional tasks. It’s a direct counterpoint to Anthropic’s MCP service. It’s only going to be a matter of weeks until we see integrations where agents can go out and purchase things effectively or make reservations, etc.
OpenAI’s chief product officer predicts that this year will mark the moment that AI permanently outperforms humanity at coding.
xAI has purchased a video foundation model company, which implies that it’s going after video companies like Runway and Sora, while also potentially looking towards building world models for robot embodiment and autonomous vehicles.
Payment processing provider Stripe has created an incredible template to show how companies can give discrete documentation to AI models through a markdown document. I highly recommend opening the document and reading how well Stripe has articulated what they can do so that a prompt engineer can simply build tools integrating with Stripe without knowing any technical details. It’s really a pivotal moment that is going unsung.
Meta’s open-source model Llama has hit 1 billion downloads.
OpenAI has released three advanced audio models, which include text-to-speech and speech-to-text with better transcription than the previous leader, Whisper. The text-to-speech is impressive as well because it can be manipulated and adjusted for inflections and tone, depending on the use case. The goal is to give agents voices for interaction with people.
There are a few more executive summaries in addition to these and also 15 separate compelling visuals of the week to check out. For the real gluttons for punishment, I also selected the remaining top 50 headlines of the week at the bottom. There’s no sign of slowing down.
Anthropic’s Claude Can Search The Web
Claude now searches the internet to provide timely information with direct source citations. This combines Claude’s knowledge base with real-time data, improving accuracy for queries requiring current information. Publisher fears come true, as users receive sources in conversation format rather than separate search results. Web search is currently available to paid US users. The feature aims to help sales teams analyze industry trends, financial analysts assess market data, researchers build stronger proposals, and shoppers compare products across multiple sources.
Claude can now search the web \ Anthropic https://www.anthropic.com/news/web-search
“Web search is now available in claude dot ai. Claude can finally search the internet! https://x.com/alexalbert__/status/1902765482727645667?s=46
NVIDIA Introduces Open Reasoning AI Models for Building AI Agents
NVIDIA launched the Llama Nemotron family of open reasoning models designed to help developers build AI agents that can work independently or in teams to solve complex tasks. Post-trained by NVIDIA, these models improve multistep math, coding, reasoning and decision-making with up to 20% better accuracy and 5x faster inference than competing open models. Available in Nano, Super, and Ultra sizes as NVIDIA NIM microservices, the models are optimized for different deployment scenarios from single PCs to multi-GPU servers. Major companies including Microsoft, SAP, ServiceNow, Accenture, and Deloitte are already integrating these models to enhance their enterprise AI offerings. NVIDIA is also releasing supporting tools including the AI-Q Blueprint and AgentIQ toolkit to help businesses connect knowledge to AI agents that can perceive, reason, and act.
NVIDIA Launches Family of Open Reasoning AI Models for Developers and Enterprises to Build Agentic AI Platforms | NVIDIA Newsroom https://nvidianews.nvidia.com/news/nvidia-launches-family-of-open-reasoning-ai-models-for-developers-and-enterprises-to-build-agentic-ai-platforms
“NVIDIA has announced their first reasoning models, a new family of open weights Llama Nemotron models: Nano (8B), Super (49B) and Ultra (249B) From our early testing, @nvidia’s Nemotron Super 49B scores 64% on GPQA Diamond in reasoning mode and 54% in non-reasoning mode – we https://x.com/ArtificialAnlys/status/1902386178206429434
NVIDIA Releases Two Major Simulation Tools for AI and Robotics
NVIDIA Launches Cosmos-Transfer for Advanced World Simulation NVIDIA released Cosmos-Transfer, a world generation model that creates simulations based on multiple spatial inputs like segmentation, depth, and edge data. The system allows users to assign different weights to various inputs across locations, enabling highly controllable simulations. Applications include robotics training and autonomous vehicle data enrichment, with NVIDIA demonstrating real-time generation capabilities using their GB200 NVL72 hardware. NVIDIA Partners with DeepMind and Disney on “Newton” Physics Engine NVIDIA, Google DeepMind, and Disney Research introduced Newton, an open-source physics engine for robotics simulation built on NVIDIA Warp. The engine offers GPU acceleration, works with existing frameworks like MuJoCo, and includes differentiable physics for more accurate simulations. Disney is already implementing the technology to develop entertainment robots, including Star Wars-inspired BDX droids, while the system helps address the gap between simulation and real-world robot performance.
“Jensen just announced Newton, an open-source physics engine for robotics simulation developed by NVIDIA, Google DeepMind, and Disney Research. ⦿ Built on NVIDIA Warp, it offers GPU-accelerated performance and compatibility with frameworks like MuJoCo and Isaac Lab. ⦿ Features https://x.com/TheHumanoidHub/status/1902091038342484239
“Nvidia just released Cosmos-Transfer1 on Hugging Face Conditional World Generation with Adaptive Multimodal Control https://x.com/_akhaliq/status/1902187161841000938
“21/ @sundarpichai: “I’m really excited about the next phase of our partnership (with NVIDIA) as we work together on agentic AI, robotics and bringing the benefits of AI to more people around the world.”” / X https://x.com/AtomSilverman/status/1902087731100209624
NVIDIA Introduces GR00T N1 Foundation Model for Humanoid Robots
NVIDIA released GR00T N1, an open foundation model designed to simplify humanoid robot development. This cross-embodiment AI system processes images and language to perform manipulation tasks across different environments and within varying robot forms. Trained on extensive humanoid datasets and synthetic data, GR00T N1 enables robots like the Fourier GR-1 to grasp objects, transfer items between arms, and complete multi-step tasks. The model features a dual architecture combining a vision-language system for reasoning with a diffusion transformer that translates plans into physical movements. NVIDIA’s training approach combines internet-scale video data with synthetic simulations (generating 750,000 trajectories in just 11 hours) and real robot demonstrations, improving performance by 40% compared to using only real-world data. The 2B parameter version available today is the first in a planned series of customizable models.
Accelerate Generalist Humanoid Robot Development with NVIDIA Isaac GR00T N1 | NVIDIA Technical Blog https://developer.nvidia.com/blog/accelerate-generalist-humanoid-robot-development-with-nvidia-isaac-gr00t-n1/
“Nvidia just released GR00T-N1-2B on Hugging Face open foundation model for generalized humanoid robot reasoning and skills https://x.com/_akhaliq/status/1902124817228194289
“Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand – with only 2B parameters, N1 learns from the most diverse physical action dataset https://x.com/DrJimFan/status/1902117478496616642
“Nvidia dropping a 2B open foundation model for humanoid robots on the hub!🔥 https://x.com/reach_vb/status/1902120408742080558
NVIDIA Isaac GR00T N1: An Open Foundation Model for Humanoid Robots | Research https://research.nvidia.com/publication/2025-03_nvidia-isaac-gr00t-n1-open-foundation-model-humanoid-robots
“NVIDIA announced GR00T N1, the first fully customizable open-source humanoid robot foundation model, designed to advance general-purpose robotics. ⦿ A dual-system AI inspired by human cognition: System-1 handles fast, intuitive actions, while System-2 enables methodical https://x.com/TheHumanoidHub/status/1902083623207211507
“NEO Gamma vacuuming the floor at NVIDIA GTC Would you rather have a vacuum cleaning bot or a versatile humanoid that can use your manual vacuum and other cleaning tools in addition to tackling many other home chores? You’ll have to wait for the latter. https://x.com/TheHumanoidHub/status/1902129279905001738
“Jim Fan at NVIDIA GTC: “This year is the year for humanoid robots.” Hyped for Jensen’s keynote tomorrow! [10am PT] https://x.com/TheHumanoidHub/status/1901738197333516555
“At GTC 2025, Jensen Huang made several AI and robotics reveals: —Next-gen GPUs: Blackwell Ultra (2025), Vera Rubin (2026), and Feynman (2028) —GR00T N1 open AI for humanoids —DGX Spark & DGX Station personal AI computers —Newton robotics physics engine https://x.com/rowancheung/status/1902249965780406666
Nvidia Partners with Storage Companies to Integrate Large Data Sets with AI Systems
NVIDIA announced a new system called the AI Data Platform that helps companies use their stored information more effectively with AI. Leading storage companies like Dell, IBM, and HP are using this design to build integrated and ‘smarter’ storage systems. These optimized architectures can quickly answer complex questions about company data without requiring technical expertise. The technology uses NVIDIA’s specialized computer chips to process information faster while using less electricity. This means businesses can get useful insights from their data almost instantly, whether that data is in documents, images, or videos. Companies plan to start offering these improved storage solutions this month. The trend here is that in NVIDIA is trying to become the dialtone for everything computing.
NVIDIA and Storage Industry Leaders Unveil New Class of Enterprise Infrastructure for the Age of AI | NVIDIA Newsroom https://nvidianews.nvidia.com/news/nvidia-and-storage-industry-leaders-unveil-new-class-of-enterprise-infrastructure-for-the-age-of-ai?ncid=ref-inpa-350226
“13/ NVIDIA AI Data Platform: New reference architecture to transform enterprise data systems from file storage to knowledge serving via AI query agents collaborating with data storage leaders to create a platform that serves knowledge via AI query agents rather than just files” / X https://x.com/AtomSilverman/status/1902087716462072234
NVIDIA Releases Two Open Source Multilingual Speech and Translation Models
NVIDIA has open-sourced Canary 1B and 180M Flash, high-performance speech recognition and translation models that rank second on the Open ASR Leaderboard. The models achieve over 1,000 times real-time processing speed and come in two sizes (880M and 180M parameters) optimized for on-device use. They support English, German, French, and Spanish with capabilities for word-level and segment-level timestamps. The models demonstrate improved accuracy with fewer hallucinations and are available under a CC-BY license that permits commercial applications.
“NEW: Nvidia just open sourced Canary 1B & 180M Flash – multilingual speech recognition AND translation models 🔥 > Second on the Open ASR Leaderboard > Achieves greater than 1000 RTF 🤯 > 880M & 180M sizes – perfect for on-device > Supports word-level & segment-level timestamps https://x.com/reach_vb/status/1902730989811413250
AI Agent Tool, Manus, Contininues Hype Cycle into Third Week
Manus, the now viral agent combining research capabilities with computer operation skills, has continuted demonstrating impressive performance on complex tasks. In testing, it completed what would typically be two weeks of professional-level stock analysis in approximately one hour. The China/Singapore-based company reports using Claude Sonnet 3.5 for execution alongside a specially post-trained Qwen model for planning. Founder Yichao suggests that developing agentic capabilities may be more about alignment than foundational technology. Early users highlight both the tool’s functional capabilities and streamlined user experience, with examples including the ability to generate 3D games through simple prompts.
“Got access and it’s true… Manus is the most impressive AI tool I’ve ever tried. – The agentic capabilities are mind-blowing, redefining what’s possible. – The UX is what so many others promised… but this time it just works. prompt: “code a threejs game where you control a https://x.com/victormustar/status/1898505307896131708
“Manus, the new AI product that everyone’s talking about, is worth the hype. This is the AI agent we were promised. Deep Research+Operator+Computer Use+Lovable+memory. Asked it to “Do a professional analysis of Tesla stock ” and it did ~2wks of professional-level work in ~1hr! https://x.com/deedydas/status/1898444603071795378
AI Agents Are Doubling Their Capabilities Every 7 Months
AI Agents Getting Better at Complex Tasks, Still Fall Short on Reliability A recent study shows AI systems are improving at handling longer tasks, with agent capabilities doubling approximately every 7 months – similar to a “Moore’s Law for AI agents.” The research found that successful AI runs cost less than 10% of what human software engineers would charge for the same work. Despite this progress, these systems aren’t consistently reliable yet. Researchers predict that within five years, AI agents could independently complete software projects that currently take humans days or weeks to finish.
“A new paper shows that AI agents are improving rapidly at long tasks – but they aren’t reliable yet. That being said this feels significant: “more than 80% of successful runs cost less than 10% of what it would cost for a human [L4 software engineer] to perform the same task.” https://x.com/emollick/status/1902443733158609005
“When will AI systems be able to carry out long projects independently? In new research, we find a kind of “Moore’s Law for AI agents”: the length of tasks that AIs can do is doubling about every 7 months. https://x.com/METR_Evals/status/1902384481111322929
Measuring AI Ability to Complete Long Tasks – METR https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
Open AI’s Agent API and SDK Mark The Start of A New Era
OpenAI introduced an Agents SDK and Responses API that streamlines AI workflow automation across tools. The platform combines chat capabilities with web search, file management, and computer interaction tools. In a demonstration, an AI agent located an invoice, analyzed billing data from a spreadsheet, and processed payment through Stripe. The system allows multiple agents to work together while monitoring performance. This development threatens many AI agent startups by eliminating the need for complex prompt engineering, custom orchestration, and extensive debugging. The move positions OpenAI as a comprehensive AI development platform but locks developers into their ecosystem, contrasting with Anthropic’s more open approach with Claude’s Open Model Context Protocol.
“Multi-Agent workflows are the future of AI. OpenAI released new Agent APIs today, and Box built an Agent that combines documents from Box and web search tools to generate answers. Enterprise devs can grab sample code from our GitHub repo to customize with their data. https://x.com/levie/status/1899522600172507617
“With Agents SDK, @OpenAI has definitely stirred up the orchestration landscape where @crewAIInc and @LangChainAI have fought hard for the pole position. But the disruption doesn’t end at agent orchestration.👇” / X https://x.com/BetterSayAJ/status/1899708901299286300
“🔷 OpenAI Drops Powerful New Tools for AI Agents! 🚀 @OpenAI 🔹 OpenAI introduces new tools & APIs to streamline logic, orchestration & interactions. 🔹 Key Features: ✅ Responses API – Chat & tool use combined for seamless interactions ✅ Built-in Tools – Web search, file https://x.com/MervinPraison/status/1899528674845610432
“Great to see @AgentOpsAi trending on github after their partnership with @OpenAI was announced with the launch of the Agents SDK @AgentOpsAi is a Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks https://x.com/atulkumarzz/status/1900436286928003255
“🤖 Agents SDK—our new open-source SDK for orchestrating multi-agent workflows, improving upon Swarm. Configure agents with built-in tools, hand off tasks, add safety guardrails, and visualize execution traces for debugging and optimizing performance. https://x.com/OpenAIDevs/status/1899531857143972051
“@OpenAI Just Wiped Out 1,000 AI Agent Startups with One Announcement The AI agent landscape just changed overnight. OpenAI unveiled its Agents SDK & Responses API, making AI-powered workflow automation ridiculously simple. The demo? A Stripe AI agent that: → Found an invoice https://x.com/adugbovictory/status/1899749168383770921
“Build a multi-agent AI research team using OpenAI Agents SDK in less than 100 lines of Python Code. It combines specialized agents that plan research, search the web, and compile reports with proper citations. 100% Opensource Code with step-by-step tutorials. https://x.com/unwind_ai_/status/1901815794670960947
“@OpenAI launched its Agents SDK yesterday, and I played around with it by creating a sales playbook agent! Here’s what it can do: – Listen to sales calls and provide real-time guidance – Recommend responses to objections based on your sales playbooks (1/2) https://x.com/danteocualesjr/status/1899942964526498042
“OpenAI released new tools for building custom enterprise agents, including a Responses API with tools for web search, computer use, and file search There’s also open-source Agents SDK to orchestrate single and multi-agent systems https://x.com/adcock_brett/status/1901303264924016784
“Today, @OpenAI released the agents SDK and computer use via their API. 3h later I built an agent with it that applies to hundreds of jobs for me 🤯🤯🤯 It was honestly not that hard and I did not write a single line of code, this was all vibecoded with Cursor. The year of https://x.com/know_u_self/status/1899567846591484359
API Reference – OpenAI API https://platform.openai.com/docs/api-reference/realtime-client-events/session
“Today, @OpenAI launched their Agents SDK developer framework. Until now, agents have mostly been chat-oriented (chat in, chat out), but agents will increasingly be action-oriented (data in, action out). We, @stripe, would love to show you a few financial agents we’ve built! ⤵️” / X https://x.com/jeff_weinstein/status/1899543525198614996
“We’ve been working with OpenAI for the past few weeks testing the Agents SDK and are launching as a tracing provider. At @AgentOpsAI, we will be sharing best practices for getting started with the SDK. Here is everything you need to get started (🧵): https://x.com/AgentOpsAI/status/1899870669397205125
“OpenAI just launched the Agents SDK – a simple yet powerful toolkit for building AI apps that can actually do things in the real world! I summarized everything you need to know about it and how it is going to be a game changer for anyone building agents. https://x.com/AtomSilverman/status/1899511053601698073
AI Code Models Expected to Surpass Top Human Programmers Next Year
OpenAI’s Chief Product Officer Kevin Weil predicts that 2025 will mark the moment AI permanently outperforms humans at programming. Weil points to AI’s rapidly improving competitive coding abilities, forecasting that an AI system will reach the number one position in coding competitions next year. He frames this shift as democratizing software creation, suggesting it will enable anyone to build whatever they imagine regardless of technical background.
“”This is the year that AI gets better than humans at programming forever. And there’s no going back.” OpenAI CPO Kevin Weil highlights the rapid improvement of AI models in competitive coding, predicting one will reach the number 1 spot in 2025. He says AI surpassing humans in https://x.com/vitrupo/status/1901258363364897016
xAI Acquires Video AI Startup Hotshot
Elon Musk’s xAI has acquired Hotshot, a startup that developed three video foundation models over the past two years. The acquisition positions xAI to compete directly with established video AI companies like Runway and Sora. This move expands Musk’s AI strategy beyond text-based systems (where xAI already competes with OpenAI and Google) and image generation into comprehensive video capabilities. The acquisition suggests xAI is building toward a complete multimodal AI platform that can handle text, images, and video generation.
“xAI acquired an AI video startup. Elon about to start gunning for Runway, Sora, Veo too. Damn. Isn’t he already gunning for OpenAI, Google and Perplexity with deep search and deep think? Bro also dropping a Midjourney class image generator while gunning for Character AI and” / X https://x.com/bilawalsidhu/status/1901723862720778267
Stripe Adds Machine-Readable Documentation for AI Tools
Stripe has integrated its documentation with large language models by adding /llms.txt and Markdown formats to its developer resources. This update allows developers to easily import Stripe’s technical information directly into AI systems. The new machine-readable format enables more efficient transfer of payment processing knowledge into developers’ preferred AI tools, streamlining implementation processes.
Stripe Developers on X: “We’ve added /llms.txt and Markdown to @stripe docs: https://t.co/mG9FIaXLTS Use the .md pages to quickly move Stripe knowledge into your LLM of choice. 📄 https://t.co/qkxhlLtHcL” / X https://x.com/StripeDev/status/1901662810897101256
Google Launches Gemini Robotics to Bridge AI and Physical World
Google DeepMind has introduced Gemini Robotics, a model built on Gemini 2.0 that can handle tasks it hasn’t encountered during training. This AI system adds physical actions as an output capability, allowing direct robot control. The company also unveiled Gemini Robotics-ER, which provides advanced spatial understanding for roboticists. As part of this initiative, Google is partnering with Apptronik to develop humanoid robots powered by Gemini 2.0, while working with selected testers to refine the technology. These models aim to enable robots to perform a wider range of real-world tasks than previously possible.
“Gemini Robotics model is able to generalize to tasks it has never seen before in training. https://x.com/TheHumanoidHub/status/1900262977582096576
“Meet Gemini Robotics: our latest AI models designed for a new generation of helpful robots. 🤖 Based on Gemini 2.0, they bring capabilities such as better reasoning, interactivity, dexterity and generalization into the physical world. 🧵 https://x.com/GoogleDeepMind/status/1899839624068907335
Perplexity Launches Model Context Protocol
Perplexity has released its Model Context Protocol (MCP) server for Sonar, enabling AI assistants to perform real-time web searches. This integration allows Claude and other AI systems to access current information from across the internet, delivering more timely and accurate responses. Developers can now implement this capability through the Perplexity API, giving their preferred AI applications access to up-to-date information.
“The Perplexity API Model Context Protocol (MCP) is now available. We’ve built an MCP server for Sonar, giving AI assistants real-time, web search research capabilities. Powered by Perplexity, Claude can now search the web and deliver real-time and accurate insights on demand. https://x.com/perplexity_ai/status/1899849114583765356
“Perplexity API now supports MCP. You can use it give any of your favorite AIs real time information, eg: Claude. https://x.com/AravSrinivas/status/1899850017546129445
AskPerplexity can now ingest videos and offer explanations
I don’t have a lot of context for this one, but it looks like you can tag perplexity in a tweet and have it watch a video and tell you what is in it. If so, that’s pretty cool!
“AskPerplexity can now ingest videos and offer explanations! https://x.com/AravSrinivas/status/1901840001866023146
Meta’s Llama AI Models Hit 1 Billion Downloads
Meta announced that its Llama AI models have reached the milestone of 1 billion downloads. The company credits the open source AI community for contributing significantly to Llama’s development and improvement over the past two years. Meta shared a recap highlighting collaborative achievements and expressed gratitude for the partnership and community enthusiasm that continues to enhance Llama’s capabilities.
“Llama has surpassed 1 billion downloads 🚀 And we would not be where we are today without the contributions of the entire open source AI community – thanks to each of you who helped make Llama better. We threw together a short recap of some of the amazing things we have built https://x.com/Ahmad_Al_Dahle/status/1902001873781194956
Creative Industries Unite Against AI Copyright Exemption Push
OpenAI and Google are urging the US Administration to remove copyright protections for AI training data, prompting strong opposition from entertainment industry professionals. In a formal response to the Administration’s AI Action Plan, representatives from film, television, music and other creative fields argue that America’s global AI leadership shouldn’t undermine the $229 billion creative sector that supports 2.3 million jobs. They emphasize that tech companies valued in the trillions could simply license content legally rather than seeking special exemptions. The group warns this issue extends beyond entertainment to all knowledge industries, potentially threatening America’s intellectual property leadership across multiple fields.
Hollywood’s response to AI Action Plan https://s3.documentcloud.org/documents/25590860/hollywoods-response-to-ai-action-plan.pdf
OpenAI Launches Advanced Audio Models for Voice Applications
OpenAI has introduced three audio models in its API to enhance voice interaction. Two new speech-to-text models outperform OpenAI’s previous Whisper model, improving transcription accuracy in challenging conditions like accents and noisy environments. The text-to-speech model allows developers to control not just content but delivery style and emotion with instructions like “talk like a sympathetic customer service agent.” These models are fully integrated with OpenAI’s Agents SDK, enabling developers to build voice agents with minimal code. The release builds on OpenAI’s recent focus on agent capabilities, extending their utility beyond text-based interactions to more natural spoken communication.
Introducing next-generation audio models in the API | OpenAI https://openai.com/index/introducing-our-next-generation-audio-models/
“🔊 Three new audio models for you today! * A new text to speech model that gives you control over timing and emotion—not just what to say, but how to say it * Two speech to text models that meaningfully outperform Whisper All available in the API + fully integrated into our” / X https://x.com/kevinweil/status/1902769861484335437
“We’re also holding a radio contest. 📻 Tweet out your https://t.co/1kI0L2V1jW TTS creations (hit “share”). The top three most creative ones will win a Teenage Engineering OB-4. Keep it to ~30 seconds, and be creative with voice, delivery, pronunciation, or even tone changes” / X https://x.com/OpenAIDevs/status/1902773659497885936
“Three new state-of-the-art audio models in the API: 🗣️ Two speech-to-text models—outperforming Whisper 💬 A new TTS model—you can instruct it *how* to speak 🤖 And the Agents SDK now supports audio, making it easy to build voice agents. Try TTS now at https://x.com/OpenAIDevs/status/1902773579323674710
Open Source Speech Model Outperforms Major Tech Companies
OPEN SOURCE SPEECH MODEL OUTPERFORMS MAJOR TECH COMPANIES Kokoro-82M has become the highest-ranked open weights text-to-speech model in the Artificial Analysis Speech Arena, scoring 1086 ELO points. This places it ahead of solutions from major tech companies while trailing only proprietary models from OpenAI, Cartesia, and ElevenLabs. The 82M-parameter model uses StyleTTS 2 architecture, supports multiple languages including English, French, and Hindi, and only required $1,000 and several hundred hours of audio to train. At $0.63 per million characters when run on Replicate, it costs approximately 100 times less than some leading proprietary alternatives.
“Kokoro-82M v1.0 is now the leading open weights Text to Speech Model in Artificial Analysis Speech Arena! Kokoro currently scores a Speech Arena ELO of 1086, just behind proprietary speech models from OpenAI, Cartesia, and ElevenLabs. For the first time, this puts an open https://x.com/ArtificialAnlys/status/1902762871106441703
Adobe Introduces Horribly Named “Experience Platform Agent Orchestrator”
Adobe’s Experience Platform Agent Orchestrator deploys specialized AI agents that handle specific marketing tasks such as audience targeting, content creation, and site optimization. These agents operate within Adobe’s enterprise applications and can integrate with third-party systems, streamlining workflow automation across different marketing platforms. Buzzword bingo unlocked! But you get the idea.
Adobe Experience Platform Agent Orchestrator is designed to unlock your full potential, making personalization at scale easier than ever. https://x.com/AdobeExpCloud/status/1902042855520305191
Stability.ai Creates 3D Videos From 2D Images
Stability.ai released Stable Virtual Camera, a diffusion model that transforms still images into 3D videos with realistic depth and camera movement. The system, currently in research preview, can generate videos following user-defined camera paths or choose from 14 preset camera movements including 360° rotations, spirals, and dolly zooms. Unlike traditional approaches, Stable Virtual Camera doesn’t require complex 3D reconstruction techniques and can work with as few as one or as many as 32 input images.
Introducing Stable Virtual Camera: Multi-View Video Generation with 3D Camera Control — Stability AI https://stability.ai/news/introducing-stable-virtual-camera-multi-view-video-generation-with-3d-camera-control
15 AI Visuals and Charts: Week Ending March 21, 2025
“An introduction to the Model Context Protocol (MCP), a new standard enabling AI agents to interact with diverse APIs in a unified way. I created the first version for a commercetools MCP server. https://x.com/lgiavedoni/status/1899778198642069916
“🔴Excited to share my TED AI talk! How we advance AI into an era of superintelligence, with true *practical* impact in our world. ❓I ask: For all the groundbreaking milestones we’ve reached, all the intelligence benchmarks we’ve shattered, for the $53B invested in generative https://x.com/stephzhan/status/1902377745335947580
NVIDIA Isaac GR00T N1 for Complex Manipulation Tasks – YouTube https://www.youtube.com/watch?v=H2wcr8UYsnI
NVIDIA Isaac GR00T N1: An Open Foundation Model for Humanoid Robots – YouTube https://www.youtube.com/watch?v=m1CH-mgpdYg
“Stability AI unveiled Stable Virtual Camera, a new diffusion model It transforms single images into 3D videos with 14 dynamic camera paths Currently in research preview under a non-commercial license! https://x.com/rowancheung/status/1902250078913360345
“google’s ai is finally based again prompt: “edit the image to add massive flames coming out of the front two nozzles of the flame thrower, having nice layers of glow, cinematic lighting, billowing fiery tendrils. turn the person in the background into a storm trooper.” https://x.com/bilawalsidhu/status/1901731029800608031
“Gosh y’all know I love 3d gaussian splats — but I love it even more when you line up your exterior & interior 3d scan and make a killer transition like this https://x.com/bilawalsidhu/status/1901156838483104013
“the level of scene understanding the new gemini model exhibits is wild. you’d need multiple controlnets to get this result, and now it’s just a text prompt. prompt: edit this image to create a 3d wireframe representation of every unique object and subject in this scene. it https://x.com/bilawalsidhu/status/1901413678202892369
Thoughts on Claude Code / AI workflows https://docs.google.com/presentation/d/1XRIflchTrZR2aqvxd7dROu8-76mf5N4Priwk8l9wNns/edit?slide=id.g3412c9db504_2_108#slide=id.g3412c9db504_2_108
“Photoshopping will never be the same. Gemini 2.0 Flash in a @Gradio app = 🤯 https://x.com/fdaudens/status/1901704690598842864
“I find playing with Gemini multimodal image generation to be really fun. Took a pic: “turn the bottles into a Saturn V complete with tiny ground crew. Add a neon sign to the cups saying ‘moon’ with an up arrow” “Make the rocket out of legos. Make the crew ducklings on stilts” https://x.com/emollick/status/1901370982557794658
“Using Gemini Flash Experimental to ruin art by adding ice cream. https://x.com/emollick/status/1900056829683462234
GTC March 2025 Keynote with NVIDIA CEO Jensen Huang – YouTube https://www.youtube.com/watch?v=_waPvOwL9Z8&t=1885s
DiffPortrait360: Consistent Portrait Diffusion for 360 View Synthesis https://freedomgu.github.io/DiffPortrait360/
“Watch for a 14min demo of me using Manus for the 1st time. It’s *shockingly* good. Now imagine this in 2-3 years when: – it has >180 IQ – never stops working – is 10x faster – and runs in swarms by the 1000s AGI is coming – expect rapid progress. https://x.com/mckaywrigley/status/1898756745545252866
Top 46 Links of The Week – Organized by Category
AGI
“When working with LLMs I am used to starting “New Conversation” for each request. But there is also the polar opposite approach of keeping one giant conversation going forever. The standard approach can still choose to use a Memory tool to write things down in between” / X https://x.com/karpathy/status/1902737525900525657
“I just created 4 weeks on content in 2 minutes with Manus! This is the closest I’ve felt to AGI. Manus creates separate docus with each 𝕏 post/thread saved in as drafts. Final step: copy over to Typefully or any post scheduler and automate your social media growth. https://x.com/Lyle_AI/status/1898538952186851663
“Kevin Weil, OpenAI’s CPO, says that the next obvious step for AGI beyond the digital world is robotics and real-world impact. https://x.com/TheHumanoidHub/status/1901544115000742364
““I believe now is the right time to start preparing for AGI” The same warnings are now appearing with increasing frequency from smart outside observers of the AI industry, like @kevinroose (below) & Ezra Klein. I think ignoring the possibility they are right is a real mistake. https://x.com/emollick/status/1900575976284660146
AgentsCopilots
“The updated Google Deep Research is really good. It is much more obviously smart and agentic in its research progression, while still casting the widest net of any of the Deep Research tools. It is becoming clear that more powerful models leads to better agents for research. https://x.com/emollick/status/1900377760322785550
“22/ Visa is using AI agents with the profiler feature to streamline cybersecurity and automate phishing email analysis.” / X https://x.com/AtomSilverman/status/1902087732874395817
“✨ AppFolio’s copilot, Realm-X – powered by LangGraph and LangSmith – saves property managers over 10 hours per week ✨ Realm-X Assistant is an AI copilot designed to streamline property managers’ daily tasks. It offers a conversational interface that empowers users to https://x.com/LangChainAI/status/1901699861264900112
“🤓 I don’t understand why MCP (Model Context Protocol) took four months to get attention from the community, but I am glad it finally took off. When it was released last November, I immediately saw its value for AI agents. That’s why I built an MCP server with ORKL right away to https://x.com/fr0gger_/status/1899588470835921040
Europe, Meet Your Newest Assistant: Meta AI | Meta https://about.fb.com/news/2025/03/europe-meet-your-newest-assistant-meta-ai/
“Created Agentic AI book writer using OpenAI’s latest Agents SDK. (Coded by Cursor + Claude 3.7) Got 4 AI agents waiting to serve me. – book analyst – ghost writer – proof reader – graphics designer Entire book content was written in 337.69 seconds, automated. Awesome! https://x.com/seree/status/1900962604266475826
“The long-term agentic capabilities of Claude 3.7 are severely slept on. It’s capable of doing tasks out of the box that were impossible with massive multi-agent systems even 3 months ago. It enables an entirely new class of products. See for yourself tomorrow 👀” / X https://x.com/mathemagic1an/status/1901869700222693647
“18/ @normativeaiannounced $48 million in funding to continue building the Legal Engineering Automation Platform (LEAP), their proprietary system for creating AI Agents with legal and regulatory domain expertise. https://x.com/AtomSilverman/status/1900661690880192717
“Today we’re releasing the #1 interface for MCP Servers. Everything you already love about Cursor, across every app on your computer. What Apple Intelligence & Copilot could have been. 1) Screen grounding: @ any service in-context, and write just by hitting tab. 2) Scroll 👇 https://x.com/PimDeWitte/status/1898023432937521592
“I built a Deep Research AI Agent team using OpenAI Agents SDK and Firecrawl. It combines multiple AI agents that can autonomously search the web, extract content, and generate detailed reports with enhanced analysis. 100% Opensource Code with step-by-step tutorials. https://x.com/Saboo_Shubham_/status/1901460343655948330
“🦜🤖LangManus You had to know it was coming! This community effort attempts to replicate Manus using the LangStack (LangChain + LangGraph) Still early innings – but check it out here! https://x.com/hwchase17/status/1902774800860451116
“manus is the craziest AI agent! it was asked to: • Plan 2-month family trip • Australia → NZ → Argentina → Antarctica watch it self-assign tasks, browse the web, research, and then creates a stunning itinerary with stays, budget & a food guide!🤯 https://x.com/LamarDealMaker/status/1898454061277458498
“Build your first agent using OpenAI’s Agents SDK and AgentOpsAI” / X https://x.com/n_sri_laasya/status/1900129244443009468
“YC batch overview (which is a good summary of early stage in general right now): -agent everything or “everythings agent” -barbell of round dynamics, either party round with angels and small feeler checks from big funds or the most outrageous round ever at seed (one co raised” / X https://x.com/NWischoff/status/1900206344462098876
“the big 4 & big 3 consultancies are getting unbundled — into an army of premium ai consultants social media unbundled media & entertainment — gave a youtuber more power than a hollywood studio now ai is doing it to knowledge work” / X https://x.com/bilawalsidhu/status/1899974041005355452
“Why? As model capabilities change (hello reasoning and LM assistants), benchmarks need to follow! The leaderboard is slowly becoming obsolete; we feel it could encourage people to hill climb in irrelevant directions. So this is the end! (hold your breath & count to 10)” / X https://x.com/clefourrier/status/1900280341887091071
“Tools like @v0 and @Replit really are so great for making disposable tools that you just wouldn’t bother making before. I’m reading Tiny Experiments from @neuranne and asked v0 to make be a pact tracker and this is the, well, v0. https://x.com/atallahfrost/status/1899699565860708480
“Measuring AI Ability to Complete Long Tasks A new analysis from METR suggests within 5 years, AI systems may be capable of automating many software tasks that currently take humans a month. “To quantify the capabilities of AI systems in terms of human capabilities, we propose a https://x.com/iScienceLuvr/status/1902244871785549909
“I used @ManusAI_HQ to help me find an apartment in London (moving there soon). I’M HONESTLY IMPRESSED. I’m not a London expert, so check for yourself if the results make sense. Research time: ~ 10 min. Prompt and results in 🧵 https://x.com/hkproj/status/1898394847678980167
“Manus AI is AGI for me!!! I added a PDF of the Class 12th Physics ‘Semiconductors’ NCERT chapter and asked Manus to make a website around it. It created a fully interactive website: breaking the chapter into several parts, using animations, simulations, and visual explanations https://x.com/BugNinza/status/1899067812548903179
Audio
“Introducing #Ray2 Flash—3x faster, 3x cheaper new model. Flash brings Ray’s frontier production-ready Text-to-Video, Image-to-Video, audio, and control capabilities with high quality and speed to all subscribers—so you can create more, faster, and without limits. Available now. https://x.com/LumaLabsAI/status/1898056614684381218
“Alright, this pretty WILD! Orpheus 3B – high quality, emotive Text to Speech – Apache 2.0 licensed! 🔥 > Zero shot voice cloning > Natural & emotive speech > Controllable intonation > Trained on 100K hours of audio > Input AND output streaming > 100ms latency 🤯 > Easy to https://x.com/reach_vb/status/1902445501427114043
EthicsLegalSecurity
“I wrote a quick new post on “Digital Hygiene”. Basically there are some no-brainer decisions you can make in your life to dramatically improve the privacy and security of your computing and this post goes over some of them. Blog post link in the reply, but copy pasting below https://x.com/karpathy/status/1902046003567718810
“id like to publicly state i dont buy the whole ‘there will be a ton of new jobs’ thing for normal people. there will be many new jobs but not for normal people” / X https://x.com/nearcyan/status/1901932030386127224
“There was controversy over the way the big AI companies did image generation a year ago. Now Grok can make pictures of famous people, Gemini can remove people or watermarks from photos, etc. Is it a shift in responsibility from firms to users? A lack of negative impact? Apathy?” / X https://x.com/emollick/status/1901798161846595851
“🤯 Gemma 3’s image analysis blew me away! Tested 2 ways to extract airplane registration numbers from photos with 12B model: 1️⃣ @Gradio app w/API link (underrated feature IMO) + ZeroGPU infra on @huggingface in Google Colab. Fast & free. 2️⃣ @lmstudio server + local processing https://x.com/fdaudens/status/1900285203135987943
OpenAI
The court rejects Elon’s latest attempt to slow OpenAI down | OpenAI https://openai.com/index/court-rejects-elon/
OpenSource
“Want to build useful newsroom tools with AI? We’re launching a @huggingface x Journalism Slack channel where journalists turn AI concepts into real newsroom solutions. Inside the community: ✅ Build open-source AI tools for journalism ✅ Get direct help from the community ✅ https://x.com/fdaudens/status/1901995208868245661
“It’s pretty wild what you can build with AI. A client wanted a website chatbot that could answer questions, take reservations, and upsell live. Instead of waiting for a dev, I built a full app in just 1 week using @grok and @lovable_dev https://x.com/cjohan247/status/1899109274535489800
“New RL Method thats better than GRPO! 🤯@ByteDanceOSS released a new open source RL method that outperforms GRPO. DAPO or Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO) achieves 50 points on the AIME 2024 benchmark with 50% fewer training steps. TL;DR: 🏆 50% https://x.com/_philschmid/status/1902258522059866504
Perplexity
“We’re excited to announce major upgrades to our Sonar models, delivering superior performance at lower costs. Recent benchmarks reveal that Sonar and Sonar Pro outperform competitors, all while remaining significantly more affordable. https://x.com/Perplexity_ai/status/1902756765843755503
Publishing
“Effortlessly Extract Appointments from Files and Websites 🎉 Every semester, my kids’ school sends PDFs packed with important dates—manually copying them to Outlook was exhausting! 🤯 So, as a non-developer, I built a Minimum Lovable Product (MLP) using @lovable_dev (amazing https://x.com/_DBrugger/status/1899530386524266697
Robotics
“Chinese automaker Xpeng Motors may invest up to $13.8 billion in humanoid robots. He Xiaopeng: “XPeng Motors has been deeply involved in the humanoid robotics industry for five years and may continue for another 20 years, investing an additional 50 billion yuan or even https://x.com/TheHumanoidHub/status/1900219602413768849
“Like clockwork, another Chinese humanoid drops. Beijing-based NOETIX Robotics has unveiled the N2, a 3’7” tall humanoid weighing 20 kg (44 lb), with 18 DOF for the whole body and an NVIDIA Jetson installed. The price will start at ¥39,900 ($5,500). https://x.com/TheHumanoidHub/status/1900674346294854118
“Elon: ⦿ There will be a massive number of robots in 10 years ⦿ You’ll be able to put snap-on parts on Optimus to change its look ⦿ 20% chance of killer robots annihilating humanity, 80% chance we’ll have an age of abundance ⦿ Whoever controls the AI chips wins the AI race https://x.com/TheHumanoidHub/status/1902252391396995138
ScienceMedicine
“UCSF created a BCI that enabled a paralyzed man to control a robotic arm using only his thoughts—for 7 months straight! It detects neural signals and converts them into commands—with an AI model—for the robotic arm to execute https://x.com/adcock_brett/status/1901303433644106062
“The math here is that a little under 50% of global GDP is human labor Humanoids are synthetic humans; it is the biggest TAM of our lifetime by a long shot Humanoids will build other humanoids who will then go out and do construction, manufacturing, help in the home Perhaps” / X https://x.com/adcock_brett/status/1900197901583999374
“Dear community, For the last 2 years, we’ve evaluated over 13K models with the Open LLM Leaderboard, using our research cluster to provide open, fair and reproducible evaluations to all. However, all good things come to an end: the leaderboard is officially retiring! https://x.com/clefourrier/status/1900280339613860057
“12/ Manus AI conducted online research, generated Python code, validated all results, and published a detailed report in 47 minutes. Prompt: “Calculate the optimal Hohmann transfer orbit for a spacecraft traveling from Earth to Mars”+ more https://x.com/AtomSilverman/status/1901701424658157639
“Manus AI is freaking insane and I have not used anything like this before. Disclaimer: I am not paid by Manus AI to write this, I just feel lucky to have gotten access. For the given prompt, Manus AI conducted online research, generated Python code, validated all results, and https://x.com/ai_for_success/status/1898393871698301208
TechPapers
“I suspect that a lot of “AI training” in companies and schools has become obsolete in the last few months As models get larger, the prompting tricks that used to be useful are no longer good; reasoners don’t play well with Chain-of-Thought; hallucination rates have dropped, etc.” / X https://x.com/emollick/status/1901011090068197484
“Thinking for longer (e.g. o1) is only one of many axes of test-time compute. In a new @Google_AI paper, we instead focus on scaling the search axis. By just randomly sampling 200x & self-verifying, Gemini 1.5 ➡️ o1 performance. The secret: self-verification is easier at scale! https://x.com/ericzhao28/status/1901704339229732874





Leave a Reply