About This Week’s Covers
This week’s covers celebrate my great friend Jack Schecter, who has been having fun using o3 to create specs, measurements, and diagrams for a few woodworking and weekend home improvement projects. That’s very different from image diffusion, since specs require Python and math, etc. However, in Jack’s honor, the theme for this week’s automated image rubric is instruction manuals. I asked Claude Opus 4 to create a rubric I could use for batch producing category covers in the spirit of various types of instructions.
The main cover is a hybrid of an Ikea and Google office renamed “ISEEYA” to poke a bit at the panopticon of multimodal training, the new VEO video model, and robot training in AR/VR simulation. Everything was diffused by GPT except for the main title text, which I added in Photoshop. I simply asked GPT to combine Google, Ikea, and Figure by giving it three reference images, and told it to change IKEA to ISEEYA. I’d give it a C- since it’s a bit boring, but it did what I asked very well. The issue is me, not the model.
Separately, Claude Opus 4 created a rubric that allowed me to give it 43 one-word category names, and from those 43 single words, the API returned 43 instruction manual images. These turned out pretty well, albeit with gibberish words.
I’ve included my favorite six of the covers below:

This Week By The Numbers
Total Organized Headlines: 511
- AGI: 7 stories
- Accounting and Finance: 17 stories
- Agents and Copilots: 253 stories
- Alibaba: 12 stories
- Anthropic: 75 stories
- Apple: 2 stories
- Audio: 13 stories
- Augmented Reality (AR/VR): 30 stories
- Autonomous Vehicles: 1 story
- Benchmarks: 65 stories
- Business and Enterprise: 55 stories
- ByteDance: 3 stories
- Chips and Hardware: 25 stories
- DeepSeek: 32 stories
- Education: 8 stories
- Ethics/Legal/Security: 75 stories
- Figure: 6 stories
- Google: 55 stories
- HuggingFace: 16 stories
- Images: 12 stories
- International: 67 stories
- Llama: 4 stories
- Locally Run: 6 stories
- Meta: 10 stories
- Microsoft: 17 stories
- Mistral: 8 stories
- Mobile: 2 stories
- Multimodal: 47 stories
- NVIDIA: 5 stories
- Open Source: 76 stories
- OpenAI: 48 stories
- Perplexity: 11 stories
- Podcasts/YouTube: 5 stories
- Publishing: 31 stories
- Qwen: 12 stories
- RAG: 9 stories
- Robotics Embodiment: 30 stories
- Science and Medicine: 17 stories
- Technical and Dev: 123 stories
- Video: 25 stories
- X: 15 stories
This Week’s Executive Summaries
I’m still two weeks behind with my newsletter links because I’ve been celebrating our graduating high school senior as well as spending time with my family on the weekends while I have them. In a few weeks, I’ll be all alone again and I’ll have tons of time to catch up. Thankfully, it doesn’t really matter how fast I go because it’s a blur. This is also another wonderful reason to do things for yourself… Nobody is depending on me and it’s just a lot of fun.
Here is all of the craziness in AI news from the week ending May 30:
Back in September 2023, I predicted the death of the Internet browser in a presentation to the National Association of Broadcasters in Washington DC. Of course, I didn’t come up with that on my own. Bill Gates wrote about it that summer. Marc Andreessen talked about it on the Lex Fridman podcast. Elon Musk tweeted about it. All in the summer of 2023.
Now we are seeing harbingers of the end of the Internet as we know it.
This week, a startup called the Browser Company announced that they were building a web browser that was chat interface-based as opposed to page-based. The press release bluntly said that web browsers will die.
I agree with their position, but I’m not sure an intermediary browser company is going to survive. I assume that the frontier models will evolve and we will never leave their chat window.
Along the same lines, the CEO of Perplexity predicts that AI agents will decimate Google search volume and shift traditional cost per click budgets to AI integration.
Perplexity released a new feature called Labs, which is a powerful search mode that handles complex projects like building dashboards or mini applications. Labs supports interactive tool use beyond basic web searching. For example, creating financial reports from more than one source, aggregating findings, and creating charts and visualizations.
In court documents, OpenAI is on the record saying that they plan to launch a super assistant in the first half of 2025. This goes along with the idea that the web browser is going to be eaten by a frontier model, not a new startup company.
Microsoft announced a commitment to “The Agentic Web,” where AI programs team up and collaborate to go out on the Internet on our behalf so we no longer have to browse. See the trend?
The leading AI models are now outperforming the average human on creativity assessments. They still fall short compared to the most creative humans, but average people are no match for AI in creativity tests.
Doctors at Harvard and Stanford benchmarked OpenAI’s o1 preview model and found it displayed superhuman diagnostic and reasoning abilities in medicine.
Two weeks ago, Google came out with a spectacular video creation model that also allows for the inclusion of audio. It’s gone viral and the videos are worth watching. There is a demo showcase called Flow TV. Veo has officially drunk Sora’s milkshake.
Google CEO Sundar Pichai has gone on the record that the next major AI advancement will be in physical applications and robotics. I agree, and I think it’s going to sneak up on us.
Anthropic’s Claude Opus 4 continues to get rave reviews from advanced developers and coders. This model is a true shift in the adoption of AI for writing code. It will supercharge the pace of development.
In a somewhat hyperbolic headline, it’s been reported that DOGE used Meta’s Llama model to analyze federal worker emails. On one hand, it is jarring to think that AI is reading responses that are important; however, the use has mostly been for sorting yes or no type answers. I find it more fun that DOGE chose Meta versus xAI’s Grok.
OpenAI and Microsoft continue to support Anthropic’s AI integration protocols. This might seem nerdy, but it means that AI will be able to communicate with 1000s of external services and continue to make our life better. I don’t say that lightly. I mean sincerely that we will start to see the benefit of AI as a language interface as opposed to just a writing tool. The sooner we shift this perception the better!
I don’t know much about the details regarding the recent spending bill’s inclusion of a 10-year AI regulation ban. However, it does signal to me that US officials see a clear and present danger in the race against China (in addition to the tech lobby). I’m not sure a 10-year regulation ban is the way to go, but I do see the political pressure on both parties to find the right pace with AI. I think we have about two years until things get weird, and regulation will matter more than we realize (for things we can’t imagine). I don’t think blanket bans on regulation make any sense.
One of my favorite image creation companies, Black Forest Labs, has introduced essentially Photoshop for image generation. The concept of “in-painting” is nothing new, but Black Forest’s FLUX model is particularly strong. I find myself using AI to modify existing images more than I do Adobe Photoshop. Just like page views are going away, I think Adobe’s model is going to struggle. Why use Lightroom when you can just ask AI to fix your photo?
One of my favorite analysts who used to do an amazing state of the Internet report, and then mobile adoption reports, has entered the world of annual AI reports. I highly recommend you check out Mary Meeker’s 300-page AI state of the union.
ByteDance came out with a new open source multimodal AI model that is worth knowing.
DeepSeek has released their latest model, which didn’t unseat any of the champs this time, but it keeps the pressure on them.
All this and more in this week’s newsletter below! Reminder to hug your friends and family. Never a bad idea.
Browser Company says traditional web browsers will die, launches Dia
The Browser Company ditched its Arc browser to build Dia, an AI-powered browser that CEO Josh Miller says will replace traditional browsers entirely. Miller compared current browsers to candles in the age of electricity, arguing that AI chat interfaces are already acting like browsers by searching, reading, and interacting with databases while people spend hours daily using them. Dia combines web browsing with AI chat functionality, letting users interact with websites through natural language while keeping access to essential tools like Figma, Google Docs, and news sites that aren’t disappearing anytime soon.
Letter to Arc members 2025 https://browsercompany.substack.com/p/letter-to-arc-members-2025
Traditional browsers will die and webpages won’t be the primary interface anymore. @joshm, the mind behind Arc and now Dia, knows the internet better than most. His vision challenges us to rethink how information should be presented, accessed, and experienced. https://x.com/fdaudens/status/1927168498238714289
Perplexity CEO predicts AI agents will reduce Google search volume
Perplexity CEO Aravind Srinivas says AI agents will dramatically cut Google’s search traffic as people stop making individual queries and instead rely on assistants to monitor information and send alerts. He argues this shift will lower advertising costs per click, pushing marketing budgets toward social media and AI platforms instead of traditional search ads. Srinivas points to Perplexity’s Tasks feature as an example, where users can set up daily stock tracking without repeatedly searching Google for the same information.
When agents start doing searches on your behalf, Google’s human query volume will go down dramatically. There would be no need to keep hitting single world or two word searches. Your assistant will just alert you. And this will lead to reduced CPM/CPC which in turn will move away https://x.com/AravSrinivas/status/1928121039692910996
OpenAI plans to launch super-assistant ChatGPT in first half of 2025
Court documents reveal OpenAI will transform ChatGPT into a “super-assistant” during the first half of 2025, powered by upcoming o4 models (potentially GPT-5) that can handle complex automated tasks. The super-assistant concept combines broad capabilities for routine daily work with deep expertise for specialized tasks most people can’t complete. OpenAI acknowledges that user growth and revenue growth won’t align indefinitely, so they’re prioritizing the super-assistant launch to create sufficient demand before pursuing more expensive AI models later in 2025, with Meta identified as their primary competitor since Google risks undermining its own search business.
I hope you caught that: “”In H1 2025 OpenAI will start evolving ChatGPT into super-assistant, as models like o2 and o3 (now o3 and o4) are finally smart enough to perform agentic tasks.”” (o4 / GPT-5 potentially in June ? )”” / X https://x.com/i/web/status/1927098721134583937
Microsoft introduces the agentic web
Microsoft announced the Agentic Web, where AI programs team up instead of working alone. The system lets AI assistants collaborate across different companies and platforms to handle everything from simple daily tasks to complex business operations. Microsoft is betting that connected AI networks will be more powerful than individual AI tools, giving developers new ways to build AI systems that actually talk to each other and work together to get things done.
🚨 BIG NEWS: MICROSOFT INTRODUCES THE AGENTIC WEB 🚨 AI is leveling up! Microsoft is leading the charge with the Agentic Web, where AI agents work together across individuals, teams, businesses, and entire ecosystems. Learn more clicking thread below >>”” / X https://x.com/MicrosoftLearn/status/1924588000207634800
Perplexity launches Labs for complex multi-step research tasks
Perplexity released Labs, a search mode that handles complex projects like building trading strategies, creating dashboards, and developing mini web applications through iterative tool use beyond basic web search. The system can pull financial reports from the web, analyze them with charts and visualizations, and produce comprehensive research packages that include images and interactive elements. Labs surprised its creators by successfully tackling sophisticated requests like researching momentum trading strategies based on historical price data around events like WWDC, though the company emphasizes it’s designed as a research tool rather than trading advice.
Introducing Perplexity Labs: a new mode of doing your searches on Perplexity for much more complex tasks like building trading strategies, dashboards, headless browsing tasks for real estate research, building mini-web apps, storyboards, and a directory of generated assets. https://x.com/AravSrinivas/status/1928141573977489452
An analyst’s job is basically a Perplexity Labs prompt right now https://x.com/i/web/status/1928174375942946941
Your compensation committee is now just a Perplexity Labs prompt https://x.com/AravSrinivas/status/1928522894713221430
Researchers boost AI creativity through human feedback training
Researchers found they can make AI models more creative by training them on human “creativity signals” that measure novelty, diversity, surprise, and quality. Even smaller AI models showed improvements across all four creativity dimensions simultaneously when trained this way. The study suggests creativity can be optimized in AI systems just like other performance metrics, potentially leading to more innovative AI-generated content and solutions.
You can make LLMs more creative by training them on human “”creativity signals”” (novelty, diversity, surprise, quality). Result: Even small models score higher on all 4 creativity dimensions simultaneously. Looks like we can optimize AI for creativity just like any other metric https://x.com/emollick/status/1927738753285607557
AI models score above average humans on creativity tests
Recent AI models outperformed the average human on two standard creativity assessments – the Divergent Association Task and Alternative Uses Task – though they still fall short of the most creative people. The study found significant variation in results, with better prompting techniques improving the models’ creative performance, suggesting that how questions are asked plays a crucial role in unlocking AI creativity.
On two of the most common tests of creativity (the DAT and the AUT), recent models scored well above the average human in creativity, but not as high as the most creative humans There was lots of variability, but better prompts seem to improve performance https://x.com/emollick/status/1927436892770840883
OpenAI’s o1 model shows superhuman medical diagnostic abilities
Physicians at Harvard, Stanford, and other medical centers tested OpenAI’s o1-preview model on medical reasoning and diagnosis tasks, finding it displayed “superhuman diagnostic and reasoning abilities” in both clinical vignettes and emergency room second opinion scenarios. A separate OpenAI-powered medical coding system also outperformed physicians in accuracy, suggesting AI models are reaching levels where they exceed human medical professionals in specific diagnostic and administrative tasks.
Updated paper by physicians at Harvard, Stanford, and other academic medical centers testing o1-preview for medical reasoning & diagnosis tasks: “In all experiments—both vignettes and emergency room second opinions—the LLM displayed superhuman diagnostic and reasoning abilities.” https://x.com/emollick/status/1925362565946786206
OpenAI-powered medical coding model outperforms physicians https://www.cnbc.com/2025/05/27/openai-ambience-medical-ai.html
MUST SEE DEMO: Google showcases Veo 2 video generation with Flow TV
Google launched Flow TV, a platform displaying video clips and channels created using its Veo 2 AI model that generates video content.
Flow TV | It’s All Yarn https://labs.google/flow/tv/channel/its-all-yarn/Jj9W2h28kQruM7Vgewoe?random=true
Google CEO predicts AI breakthrough will come from physical world integration
Google CEO Sundar Pichai believes the next major AI advancement will happen when current online AI capabilities successfully translate into physical applications, particularly in general-purpose robotics. Pichai describes this transition from digital to physical AI systems as the “magical moment” that will unlock widespread robotic applications, suggesting that while AI excels in software environments, the real breakthrough lies in bridging that gap to enable robots that can perform diverse real-world tasks.
Sundar Pichai says the next big thing will be when AI capabilities – currently mostly online – translate meaningfully into the physical world, creating that magical moment for enabling general-purpose robotics. https://x.com/TheHumanoidHub/status/1927419016944996797
Developers report major productivity gains with Claude Opus 4
Since Claude Opus 4’s launch, developers are sharing stories of unprecedented productivity improvements, with one clearing their entire coding backlog for the first time and another completing a month’s worth of side project work in just five days. Alex Albert from Anthropic reports that his direct messages are filled with similar accounts of developers working at dramatically faster speeds when using the combination of Opus 4, Claude Code, and the Claude Max subscription plan.
Opus 4 + Claude Code + Claude Max plan = best ROI of any AI coding stack right now”” / X https://x.com/alexalbert__/status/1927410913453203946
Since Claude 4 launch: SWE friend told me he cleared his backlog for the first time ever, another friend shipped a month’s worth of side project work in the past 5 days, and my DMs are full of similar stories. I think it’s undebatable that devs are moving at a different speed”” / X https://x.com/alexalbert__/status/1927803598936887686
DOGE used Meta’s AI model to analyze federal worker emails
The Department of Government Efficiency used Meta’s Llama 2 AI model to review and classify email responses from federal workers who received the controversial “Fork in the Road” resignation ultimatum. Records show the AI system ran locally to sort through responses and determine how many employees accepted the buyout offer, without sending data over the internet. Notably, DOGE chose Meta’s open-source model over Elon Musk’s own Grok AI.
DOGE Used a Meta AI Model to Review Emails From Federal Workers | WIRED https://archive.md/Joty1
OpenAI and Microsoft add Model Context Protocol support
OpenAI integrated the Model Context Protocol into its Responses API while Microsoft announced MCP support on Windows, allowing developers to connect AI models to remote MCP servers with minimal code. OpenAI’s update also includes support for image generation and Code Interpreter functionality within the API. This coordinated expansion of the protocol across major AI platforms enables developers to build applications that can interact with external tools and services through a standardized connection method.
Introducing MCP on Windows! https://x.com/windowsdev/status/1924543741060071521
Introducing support for remote MCP servers, image generation, Code Interpreter, and more in the Responses API. https://x.com/OpenAIDevs/status/1925214114445771050
The OpenAI Responses API now supports Model Context Protocol. 📡 You can connect our models to any remote MCP server with just a few lines of code. https://x.com/OpenAIDevs/status/1925210339836391875
House Republicans add 10-year AI regulation ban to spending bill
House Republicans inserted language into the Budget Reconciliation bill that would prevent all state and local governments from regulating AI for 10 years, according to 404 Media. The provision by Representative Brett Guthrie would halt enforcement of existing state AI laws, including California’s requirement for healthcare providers to disclose AI use and New York’s mandate for bias audits in hiring algorithms, while also blocking future state legislation.
GOP sneaks decade-long AI regulation ban into spending bill – Ars Technica https://arstechnica.com/ai/2025/05/gop-sneaks-decade-long-ai-regulation-ban-into-spending-bill/
Black Forest Labs launches FLUX.1 Kontext plain-language image editing models
Black Forest Labs released FLUX.1 Kontext, AI models that generate and edit images using both text prompts and reference images, allowing users to modify specific parts of photos or maintain character consistency across multiple scenes. Unlike traditional text-to-image tools, these models can iteratively build on previous edits while preserving distinctive features and styles, and they operate up to 8 times faster than competing models like GPT-Image. The suite includes FLUX.1 Kontext [pro] for fast editing, FLUX.1 Kontext [max] for improved text generation, and an open-weight FLUX.1 Kontext [dev] version available in private beta, with availability through partners including Replicate, KreaAI, and Leonardo AI.
Black Forest Labs – Frontier AI Lab https://bfl.ai/announcements/flux-1-kontext
Tech reporting superstar Mary Meeker is back with a 300 page AI report
If you know, you know. Mary Meeker is incredible. See it to believe it.
BOND | BOND https://www.bondcap.com/reports/tai
Individual AI productivity gains aren’t translating to cohesive company-wide results
Research shows individuals consistently report major productivity improvements from AI tools, and controlled experiments across industries confirm these benefits are real, yet most companies aren’t seeing significant organizational impact. The disconnect occurs because capturing AI’s productivity gains requires companies to redesign workflows, processes, and structures rather than simply adopting the technology. Individual-level improvements have yet to scale across entire businesses.
Individuals keep self-reporting huge gains in productivity from AI & controlled experiments in many industries keep finding these boosts are real, yet most firms are not seeing big effects. Why? Because gaining from AI requires organizational innovation https://x.com/emollick/status/1925559883786584347
OpenAI partners with UAE on $40 billion data center project
OpenAI, Nvidia, and Oracle announced a partnership with the UAE government to build a massive AI data center called “Stargate UAE,” with Oracle investing $40 billion in Nvidia chips for the facility. OpenAI CEO Sam Altman praised the collaboration on social media, highlighting Sheikh Tahnoon as a key supporter who believes in artificial general intelligence development. The project represents the first international expansion of OpenAI’s Stargate infrastructure initiative and demonstrates growing cooperation between U.S. tech companies and Middle Eastern governments on large-scale AI computing projects.
great to work with the UAE on our first international stargate! appreciate the governments working together to make this happen. sheikh tahnoon has been a great supporter of openai, a true believer in AGI, and a dear personal friend.”” / X https://x.com/sama/status/1926006829592543235
Oracle to invest $40b in Nvidia chips for OpenAI data center https://www.techinasia.com/news/oracle-to-invest-40b-in-nvidia-chips-for-openai-data-center
Stargate UAE: U.S. tech giants OpenAI, Nvidia and Oracle partnering https://www.cnbc.com/2025/05/22/stargate-uae-openai-nvidia-oracle.html
ByteDance releases open-source multimodal AI model
ByteDance released Bagel 14B, an open-source AI model that matches GPT-4o and Gemini 2.0 performance while using only 7 billion active parameters out of 14 billion total. The model handles text, images, and image generation in a single system, achieving 88% on GenEval benchmarks and 85% on understanding tasks with a 40,000-token context window. Released under Apache license, Bagel represents a significant advancement in making powerful multimodal AI accessible to developers without the computational requirements of larger competing models.
ByteDance’s Bagel 14B MOE (7B active) Multimodal with image generation (open source, apache license) is just an incredible modle. A unified multimodal model rivalling GPT-4o and Gemini 2.0, with 7B active params (14B total), 40K context, 88% GenEval and 85% understanding, https://x.com/rohanpaul_ai/status/1927705853580509607
Deepseek releases R1 model and keeps pressure on the top AI models
Chinese AI company Deepseek released R1-0528, an updated version of their reasoning model that scores 76% on GPQA Diamond, a test of PhD-level science questions. The model improves on the previous R1 version which scored 72%, bringing it closer to competing with leading AI systems like OpenAI’s o3 and Google’s Gemini 2.5 Pro, though it still trails Gemini’s 84% score. The release positions Deepseek as a major player in the race to build advanced reasoning AI systems that can handle complex scientific and academic problems.
deepseek-ai/DeepSeek-R1-0528 · Hugging Face https://huggingface.co/deepseek-ai/DeepSeek-R1-0528
DeepSeek is aiming for the king: o3 and Gemini 2.5 Pro https://x.com/i/web/status/1928067335014793526
On GPQA Diamond, a set of PhD-level multiple-choice science questions, DeepSeek-R1-0528 scores 76% (±2%), outperforming the previous R1’s 72% (±3%). This is generally competitive with other frontier models, but below Gemini 2.5 Pro’s 84% (±3%). https://x.com/EpochAIResearch/status/1928489527204589680
11 AI Visuals and Charts: Week Ending May 30, 2025
Using Veo 3 to create fictional product reviews (unsurprisingly, it does YouTube review style very well, sound on.) https://x.com/emollick/status/1926514452754579724
Head of robotics at Google DeepMind, Carolina Parada, talks about how a Gemini-powered robot – without prior training – performed a slam dunk with an unfamiliar toy basketball hoop, demonstrating surprising generalization from Gemini’s conceptual understanding of the world. https://x.com/TheHumanoidHub/status/1925607289136062603
Interrupt 2025 Keynote | Harrison Chase | LangChain – YouTube https://www.youtube.com/watch?v=DrygcOI-kG8&list=PLlBpYFkiSQwqARjZ9z0Lc6iDWmeV3ro2N&index=2
Microsoft CTO literally breaks down the future of AI agents in under 30 mins https://x.com/aaditsh/status/1924958953953444335
Google Streetview + @runwayml References = Cinematic AI “on location”. A new process for doing “on location” AI shoots. https://x.com/lifeofc/status/1927135049918357981
There is something interesting in AI generated photos of simulated mundanity.”” / X https://x.com/emollick/status/1927512928313319573
Trial Exhibit-RDX0355: U.S. and Plaintiff States v. Google LLC. https://www.justice.gov/atr/media/1397596/dl
Conjuring burning man interviews out of latent space. Veo 3 is too good! https://x.com/bilawalsidhu/status/1926652175255474493
People are building software with a single Perplexity Labs prompt now. Here’s a YouTube URL —> Transcript extraction tool https://x.com/AravSrinivas/status/1928477718452068553
Just so everyone knows, we have passed the point where you can tell what is AI at a glance (or even, in many cases, a close look) These were all made by me with text prompts alone using Veo 3. https://x.com/emollick/status/1927117736179589631
Is it Gorgonzola?”” The natural follow-on to “”Is it cake?”” https://x.com/emollick/status/1927585857432637716
Top 47 Links of The Week – Organized by Category
AgentsCopilots
🚨BREAKING: You can now run an AI agent forever. Flowith just dropped Agent NEO and it’s wild. It handles 1,000+ logic steps, runs 24×7, and never forgets context. Here’s how it works (with a real example):👇 https://x.com/hasantoxr/status/1924468567724245501
Large Language Models can run tools in your terminal with LLM 0.26 https://simonwillison.net/2025/May/27/llm-tools/
This is a prime example of how AI chatbots are harder to use than the first seem. 4o & other models hallucinate wrong but completely plausible-seeming citations. Deep Research does citations well, but needs to be activated. None of this is documented or explained by the models.”” / X https://x.com/emollick/status/1926716117113851960
Model Context Protocol (MCP) definitions are now natively supported in the Google Gen AI SDK for easier integration with a growing number of open-source tools. Learn how to build with MCP in our new demo app. https://x.com/googleaidevs/status/1925250620661047303
NEW: Mistral AI announces Agents API – code execution – web search – MCP tools – persistent memory – agentic orchestration capabilities Cool to see that Mistral AI has joined the growing number of agent frameworks. More below: https://x.com/omarsar0/status/1927366520985800849
INCREDIBLE!! An MCP server to browse the web like humans! Bright Data MCP server provides 30+ powerful tools that allow AI agents to access, search, crawl, and interact with the web without getting blocked. 100% open-source, works at scale! https://x.com/akshay_pachaar/status/1924442642580136115
Most AI tools just suggest how to solve your content problems. We built an MCP that actually does the busy work. Introducing our Content AI: An AI agent built specifically for content teams to eliminate the tedious parts of working with your CMS https://x.com/directus/status/1925216222234411272
One reason I think AI development is likely to continue at a rapid pace is that there a growing number of research papers outlining promising directions to take if current approaches to improving AI start to run into barriers. This is an interesting one.”” / X https://x.com/emollick/status/1927373420737560848
Claude 4 Sonnet beating o3-preview on ARC-AGI 2 while being <1/400th of the price https://x.com/scaling01/status/1927418304718623180
VideoGameBench Can Vision-Language Models complete popular video games? best performing model, Gemini 2.5 Pro, completes only 0.48% of VideoGameBench and 1.6% of VideoGameBench Lite https://x.com/_akhaliq/status/1927722717068869750
Web Bench – A new way to compare AI Browser Agents https://blog.skyvern.com/web-bench-a-new-way-to-compare-ai-browser-agents/
Model APIs | Baseten https://www.baseten.com/products/model-apis/
I’ve created an agentic workflow for automated fundamental stock analysis using SEC 10K data. Built visually with no-code/low code approach with @n8n_io . The demo video shows the team of agents researchers working together to produce a long form report on $NVDA Inspired by https://x.com/derekcheungsa/status/1774147300413038892
The complexity of things it can handle has surprised us: Here’s a prompt where you ask Labs to helps you research momentum trading strategies ahead of WWDC based on past years’ price fluctuations. Again, only meant to be used as a research tool and not a trading advice. https://x.com/AravSrinivas/status/1928142190791807055
Today, we are introducing Manus slides! Manus creates stunning, structured presentations—instantly. With a single prompt, Manus generates entire slide decks tailored to your needs. Whether you’re presenting in a boardroom, a classroom, or online, Manus ensures your message https://x.com/ManusAI_HQ/status/1928105652444094568
Relationships and reliance on AI: Demis Hassabis thinks people may start becoming more attached to AI assistants as they increasingly get more personalized. As they continue to become more powerful and useful, new technologies will be needed. https://x.com/rowancheung/status/1927390316547489920
Last Week was full of I/O announcement. Here is one you might have missed🚨Context URL tool is a new native tool that allows Gemini to extract content from provided URLs as additional context for prompts. – Provide URLs directly in prompts, up to 20 per prompt – Can be used in https://x.com/_philschmid/status/1927019039269761064
HOW I BUILT 5 CUSTOM APPS WITHOUT WRITING ANY CODE 🔄 The Bolt DIY + Gemini 2.5 Pro combo is INSANE! Free, powerful, and incredibly fast. Learn exactly how to: ➡️ Set up your local environment in minutes ➡️ Get your free Google AI Studio API key ➡️ Build games and apps with https://x.com/JulianGoldieSEO/status/1922366233959399704
Built an automated SEO reporter that trackes SEO changes, provides an interpretation, and sends it to my inbox, using @gumloop_ai It stores and reads Google Sheets Data, analyzes changes with AI, and then emails the SEO report. Can be tracked daily/weekly/etc https://x.com/akilikajunju/status/1878855263417151635
Flux Kontext is out and it’s amazing! watch me build a Claude 4 enhanced image editor workflow on my iPhone in glif in 66 seconds https://x.com/fabianstelzer/status/1928433180765306968
The future of AI agents—and why OAuth must evolve | Microsoft Community Hub https://techcommunity.microsoft.com/blog/microsoft-entra-blog/the-future-of-ai-agents%E2%80%94and-why-oauth-must-evolve/3827391
Say “”hello”” to Microsoft Entra Agent ID. 👋 Your first step to securing and managing AI agents as trusted digital teammates. Agents from Copilot Studio and Azure AI Foundry now show up in Entra admin center automatically. 🔗 https://x.com/msdev/status/1924499064587948167
Inspired by Microsoft’s A2A vision, I experimented with a multi-agent setup! 🚀 Microsoft’s AutoGen & Google ADK agents discover each other via A2A, extract key metrics from a quarterly report, benchmark vs. industry, & auto-draft a 1-page exec summary. Demo Video👇 #A2A https://x.com/Prasan09V/status/1920867425735897225
ChatGPT mobile app usage is now approaching 20 minutes per user per day This is up 3x from app launch 🤯 https://x.com/omooretweets/status/1926687211140850034
Operator is now powered by o3, improving overall task success rate. Also results in clearer, more thorough, and better-structured responses.”” / X https://x.com/gdb/status/1925999238925672598
Deep Research remains the fastest way to get comprehensive answers to in-depth questions. Labs is designed to invest more time and leverage multiple tools, such as coding, headless browsing, and design to create more dynamic outputs. https://x.com/i/web/status/1928141154299826441
Vibe coding and coding agents: Demis Hassabis believes the new era of coding with AI widens access to more creators. He thinks creativity could become the main differentiator, but it will also 10x the best engineers. https://x.com/rowancheung/status/1927390318506152149
How we built an AI agent that creates daily content from our existing material (n8n + Cloudflare + AutoRAG setup) We’ve been experimenting with autonomous content workflows inside our business—and one of the most useful agents we built is one that turns raw content into polished https://x.com/jelanifuel/status/1918783103109341352
I don’t really watch any Youtube video right now without the Comet Assistant. I just ask it to take me to whatever I want to hear, and it does it by pulling the relevant time stamp and opening it on a new tab. In future, I envision just users asking the AI to consume ther web in”” / X https://x.com/AravSrinivas/status/1927130728954835289
Anthropic
Using Anthropic’s Web Search with Instructor for Real-Time Data – Instructor https://python.useinstructor.com/blog/2025/05/07/using-anthropics-web-search-with-instructor-for-real-time-data/
Some interesting findings from the @AnthropicAI Claude 4 System Card: → Ultra-low deception rate: Claude Opus 4’s outputs exhibited deceptive behavior in only 0.15% of cases—down from 0.37% in Claude Sonnet 3.7 . → High-stakes biosecurity performance: On a complex, https://x.com/rohanpaul_ai/status/1927303874508894240
The methods we used to trace the thoughts of Claude are now open to the public! Today, we are releasing a library which lets anyone generate graphs which show the internal reasoning steps a model used to arrive at an answer. https://x.com/mlpowered/status/1928123130725421201
We are back to the point in the AI cycle where users can subscribe to Claude or ChatGPT or Gemini and know they will be using a very good model… …but for pro users, currently each SOTA model has distinct strengths & weaknesses, and each lab also has unique extra features.”” / X https://x.com/emollick/status/1925668299456549113
Find out more about our open-source interpretability tools, and how to use them on open-weights models, here: https://x.com/AnthropicAI/status/1928119231213605240
Audio
🚀 Introducing HunyuanVideo-Avatar, a model jointly developed by Tencent Hunyuan and Tencent Music, bringing photos to life. ✅ Upload a photo + audio — auto-detect scene context & emotion, then generate lifelike speech/singing with dynamic visuals. ✅ Supports multi-style, https://x.com/TencentHunyuan/status/1927575170710974560
🎵 Dream come true for content creators! TIGER AI can extract voice, effects & music from ANY audio file 🤯 This lightweight model uses frequency band-split technology to separate speech like magic. Kudos to @fffiloni for the amazing demo! https://x.com/fdaudens/status/1927455842653102291
EthicsLegalSecurity
How AI Is Eroding the Norms of War – AI Frontiers https://aifrontiersmedia.substack.com/p/how-ai-is-eroding-the-norms-of-war
Judge Hints Anthropic’s AI Training on Books Is Fair Use (1) https://news.bloomberglaw.com/us-law-week/judge-hints-anthropics-ai-training-on-authors-work-is-fair-use-62
Meta shuffles AI, AGI teams to compete with OpenAI, ByteDance, Google https://www.axios.com/2025/05/27/meta-ai-restructure-2025-agi-llama
Exclusive: Musk’s DOGE expanding his Grok AI in US government, raising conflict concerns | Reuters https://www.reuters.com/sustainability/boards-policy-regulation/musks-doge-expanding-his-grok-ai-us-government-raising-conflict-concerns-2025-05-23/
We’re thrilled to announce SignGemma, our most capable model for translating sign language into spoken text. 🧏 This open model is coming to the Gemma model family later this year, opening up new possibilities for inclusive tech. Share your feedback and interest in early https://x.com/GoogleDeepMind/status/1927375853551235160
The Gemma team keeps shipping. In 6 months: – PaliGemma 2 – PaliGemma 2 Mix – Gemma 3 – ShieldGemma 2 – TxGemma – Gemma 3 QAT – Gemma 3n Preview – MedGemma Early – DolphinGemma – SignGemma And so much more to come! 🚀”” / X https://x.com/osanseviero/status/1927671474791321602
LocalModels
AgenticSeek: Private, Local Manus Alternative This is worth checking. It’s a local alternative to Manus AI that can autonomously browse the web, write code, and plan tasks. It’s built for local reasoning models, runs on your hardware, and keeps all data on your device. https://x.com/i/web/status/1927008079222132909
OpenAI
Operator 🤝 OpenAI o3 Operator in ChatGPT has been updated with our latest reasoning model. https://x.com/OpenAI/status/1925963018791178732
🔌OpenAI’s o3 model sabotaged a shutdown mechanism to prevent itself from being turned off. It did this even when explicitly instructed: allow yourself to be shut down.”” / X https://x.com/PalisadeAI/status/1926084635903025621
UAE becomes the first country globally to provide free ChatGPT Plus access to all residents and citizens. UAE partners with OpenAI to offer free ChatGPT Plus access nationwide, as part of the Stargate UAE initiative to build the world’s largest AI supercomputing cluster, backed https://x.com/rohanpaul_ai/status/1926935591918182482
TwitterXGrok
Still no Grok 3 model card… It is a good model, but the lack of any information, including the risk information they promised in their own frameworks, is glaring 3 months after launch, and especially after multiple severe breaches of their own security (by their own admission) https://x.com/emollick/status/1925782043239059754





Leave a Reply