AI News #134: Week Ending April 24, 2026 with 49 Executive Summaries
About This Week’s Covers
This week’s cover theme is a song that is special to me because my daughter and I used it as a way to stay inspired throughout her first year at the University of Colorado Boulder. Alesso’s “Heroes (We Could Be)” was our anthem.
Rori and I were 1,600 miles apart, and I only saw her for a handful of days over the past nine months. However, when Rori was in high school, she danced competitively six days a week until 9 p.m., sometimes later, and I never really saw her. So despite her being 1,600 miles away, we FaceTimed daily, sometimes even more than once a day! It was awesome.
As she sent me her adventures, whether academic, exploring the Rocky Mountains, or taking road trips, we always used the Alesso “Heroes (We Could Be)” song as our psych-up music. Rori is a photographer for the student paper, and on her last weekend before coming home, she photographed Alesso at Red Rocks, FaceTimed me, and we sang “Heroes (We Could Be)” together.
So for this week’s cover, I decided to take the Alesso cover and transform it into the AI newsletter cover. I did that with Photoshop by hand, recreating all the elements and fonts, as well as the punched-out text with the galaxy in the background.
There are a lot of websites that can help you find the right fonts, and in the case of the Heroes font, I was able to track down Akzidenz-Grotesk Medium Condensed. However, for the Alesso logo font, the consensus on the internet was incorrect and said that it was a font called Alien League. Luckily, Alesso’s brand team, Acid and Marble, posted a great overview of the brand identity and included the actual font, which is RBNo2 Light by René Bieder. https://www.behance.net/gallery/32016489/ALESSO-Identity
I love when people share how they do things, and that’s one of the main reasons I overdo it with my own newsletter. As a photographer, beyond just the ability to compose a shot, a lot of it has to do with the gear. It could be the lens, the f-stop, the shutter speed, the ISO… things that are helpful to teach people who are coming up.
For the category covers, I decided to use my Claude skill, and I simply asked it to swap out Alesso’s name with the category name, leave “We Could Be,” and then, for the subtitle, put something that would be aspirational for that category to accomplish. Then I let Claude partner with the Gemini API tool to recreate the images.
I was pretty impressed with how well it did with the fonts, and how well Gemini was able to recreate the punched-out text with the galaxy, considering it had to recreate it for each particular letter. If you really think through how diffusion works, it’s pretty amazing.
I also thought Claude’s prompts were pretty creative, considering I did not guide it other than the overarching idea. I’ve put a few of my favorites below.
Humanities Reading for The Week
We go hide away in daylight We go undercover, wait out the sun Got a secret side in plain sight Where the streets are empty, that’s where we run
Every day people do everyday things, but I can’t be one of them I know you hear me now, we are a different kind, we can do anything We could be heroes, me and you
Anybody’s got the power They don’t see it ’cause they don’t understand Spinnin’ round and round for hours You and me, we got the world in our hands Every day people do everyday things, but I can’t be one of them I know you hear me now, we are a different kind, we can do anything We could be heroes, me and you
All we’re looking for is love and a little light Love and a little light We could be heroes, me and you -Alesso Heroes (We could be)
For the week ending April 24, I organized 588 headlines into 49 categories. 138 links informed 50 executive summaries. I’ve tried to organize the first set as the most important, but all 50 stories are worth knowing this week.
Here’s the deep dive into these stories, with links and examples.
Top Stories
Benchmarks
Microsoft AI CEO: Frontier AI computing power is set to grow one-thousandfold by 2028 Mustafa Suleyman is the CEO of Microsoft AI. This week he tweeted that since he began working on AI in 2010, the amount of computing power for the best AI models has grown by a factor of one trillion. He estimates that computing power will increase by 1,000 times by the end of 2028. So that’s 1,000 times the existing one-trillion-times increase from 2010.
This type of exponential growth is what makes the trajectory of the future hard to comprehend or predict. I agree with the criticisms that AI is a bubble and that much of the tech scene is self-indulgent, trying to grow their company’s valuation. However, even if we pull away from this cynical lens, technology’s improvement happens regardless of any individual company’s ambition… and the arc has been and will continue to be steep.
“Since I began work on AI in 2010, training compute for frontier models has grown by one trillion times. Now we’re looking at something like another thousand-fold growth in effective compute by the end of 2028. 1000x the existing 1,000,000,000,000x. Extraordinary stuff.” https://x.com/mustafasuleyman/status/2046989133676257284
Google
Google says new TPU8t chips can scale to one-million-chip clusters(!!!) In related news this week, Google announced that they can scale their architecture to a million chips within one cluster. Google’s chips are called TPUs, which stands for Tensor Processing Units. Google’s TPUs have made the news a lot more in the past few weeks and appear to be chipping away at NVIDIA’s dominance.
White House and Silicon Valley accuse China of industrial-scale US AI model theft Anthropic, Google, OpenAI, and the White House have accused China of stealing from American frontier models. There’s a lot to unpack here, and the irony in the room is that, of course, the American frontier models are often accused of stealing from all of human content in order to make their models smarter.
Almost like Russian nesting dolls, these U.S. frontier companies are accusing China of using U.S. models to train open-source models that don’t have access to the same computing power or resources. The technical term for how this is happening is called distillation, and that’s where you create a small model to act like a student, trained by a bigger model that acts as a mentor or tutor.
Chinese model companies are using the U.S. company APIs and learning by querying andn mimicking the major models. For example, Anthropic claims it saw more than 150,000 exchanges between DeepSeek and Claude talking about alignment. MiniMax had more than 13 million exchanges. Kimi spoke with Claude 3.4 million times.
Scott Galloway has theorized that if China were trying to disrupt American innovation and economics, the best way to do it would be to release free models as soon as possible after America invests in building the best closed models. So, for example, if Anthropic invests billions of dollars to make the latest version of Opus, and then three months later a free version comes out from a Chinese model that’s open source… for consumers or corporations, it’s a lot cheaper to use a free model that’s essentially as good.
Most corporate enterprise AI systems use what’s called a wrapper that allows secure usage of artificial intelligence via API. This wrapper can swap out the engine to adjust the price of any query. For example, if you need the greatest model for an important complex task, you’re probably going to use Opus 4.7 or GPT-5.5. But this costs a lot of money per query. If all you’re doing is something simple like understanding a simple instuction for an operational agent, you could just use an open model that’s every cheap, like Qwen, Gemma, DeepSeek, or MiniMax.
The pricing competition between open source and frontier models is a real issue when it comes to enterprise AI adoption. For example, I believe Uber was using Qwen for quite some time.
Beyond the economics and intellectual property, there’s the AI arms race in general, especially when it comes to security. Once AI agents start talking to each other directly, whoever has the most powerful agent is potentially able to hack weaker agent’s systems.
Whether or not it’s a merited fear, there is a lot of talk around essentially an arms race with China for computing dominance, whether that’s quantum or artificial intelligence. China has demonstrated that open source is a scalable way to overwhelm frontier models in the United States. And now U.S. frontier model companies are partnering with the government to frame this as a national security issue.
There have been many predictions of how artificial intelligence will play out. Almost all of them have a phase where the U.S. government starts to work directly with the frontier models to close down access in the name of security. And you can see this starting to happen now, whether it’s the arguments between Anthropic and the Department of Defense, this new alignment around protectionism, or Anthropic limiting availability of it’s best model Mythos. Despite the heated rhetoric from the U.S. directed at Anthropic, Anthropic is still very much at the table.
Once the strongest closed models become unreleased to the public and protected by the U.S. government, we’re in a new phase of artificial intelligence, especially the overlap between private- and public-sector partnerships.
Anthropic launches Claude Design, taking direct aim at Figma Switching gears completely, Anthropic introduced Claude Design this week, which directly goes after Figma’s design tools. Claude can now build realistic prototypes, wireframes, mockups, pitch decks, presentations, and marketing collateral, all within a single AI chat interface. The multimodality of AI shifting from just chat to images, videos, and layout is finally happening at a consumer level.
People are going to start to get used to the concept of chat becoming a secondary feature, as chat is now an interface for output that could be any type of media… even website design and creation (with coding tools).
Anthropic exec Mike Krieger left Figma’s board this week after reports of an incoming launch of a competing product. Now, Claude Design is live. How it works: describe the design and Claude Opus 4.7 builds the first version. Refine with inline comments, direct edits, or https://x.com/TheRundownAI/status/2045176722476208454
Anthropic released Claude design, direct attack on figma and lovable. Anthropic just shipped Claude Design, powered by Claude Opus 4.7., a tool that turns conversations into polished prototypes, pitch decks, and marketing assets. It auto-applies your brand system, lets you https://x.com/kimmonismus/status/2045162358004216134
Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our most capable vision model. Available in research preview on the Pro, Max, Team, and Enterprise plans, rolling out throughout the day. https://x.com/claudeai/status/2045156267690213649
OpenAI
OpenAI and Google swap image-generation benchmark lead, then GPT-Image-2 surges ahead Last week, OpenAI released the latest version of ChatGPT Image. It is now the strongest image tool on the imaging benchmarks. I went through the specs last week, but this week we have some more examples of feats of strength, and they’re worth looking at below.
Arena Trends: Text-to-Image, Jan 2026 – Apr 2026 For most of the year, @GoogleDeepMind and @OpenAI traded the top spot within a tight margin – GPT-Image vs. Nano Banana – with the rest of the field clustered below 1,200. Today, GPT-Image-2 breaks away with a score of 1,512, 242 https://x.com/arena/status/2046690103515648061
A Visual Thought Partner ChatGPT Images 2.0 is our first image model with thinking capabilities. When a thinking model is selected in ChatGPT, Images 2.0 can search the web for real-time information, create multiple distinct images from one prompt, double-check its own outputs, https://x.com/OpenAI/status/2046670989719924768
ChatGPT Images 2.0 is a big leap forward in image generation intelligence. It’s much better at following detailed instructions, rendering dense text, understanding the world more accurately, and creating visuals that are more useful. And when you give it additional time to https://x.com/nickaturley/status/2046677986242363731
ChatGPT Images 2.0 is available starting today to all ChatGPT and Codex users. Images with thinking are available to ChatGPT Plus, Pro, and Business users (Enterprise soon). On mobile, make sure you update to the latest version of the app. The underlying model, gpt-image-2, is https://x.com/OpenAI/status/2046670994413322435
Exciting news – GPT-Image-2 by @OpenAI has claimed the #1 spot across all Image Arena leaderboards! A clean sweep with a record-breaking +242 point lead in Text-to-Image – the largest gap we’ve seen to date. – #1 Text-to-Image (1512), +242 over #2 (Nano-banana-2 with web-search https://x.com/arena/status/2046670703311884548
GPT-Image-2 takes #1 in every single Text-to-Image category — all 7 of them. Surpassing the next leading model (Nano-banana-2 with web-search) across the board. Here’s the drill-down on improvements vs. its predecessor, GPT-Image-1.5-High-Fidelity: – #1 Product, Branding & https://x.com/arena/status/2046670705958551938
GPT-ImageGen-2 did this in one shot, with just the prompt “”turn all of Tennyson’s Ulysses into a comic, across as many pages as needed. make it great, include the full text”” 10 pages, though it did use what seems to be the ImageGen-2 ‘s preferred “”spackled drawing”” style 1/ https://x.com/emollick/status/2046843402021380556
I have been using GPT ImageGen-2 for the past weeks I didn’t think that better image-generators would be a big deal but it turns out that there is a quality threshold I didn’t expect, where you can now get text, slides, academic papers Look at what it does with my “”otter test””! https://x.com/emollick/status/2046665274535854146
Introducing ChatGPT Images 2.0 A state-of-the-art image model that can take on complex visual tasks and produce precise, immediately usable visuals, with sharper editing, richer layouts, and thinking-level intelligence. Video made with ChatGPT Images https://x.com/OpenAI/status/2046670977145372771
My most popular AI post was a bunch of made-up “”graphs”” four years ago. Now, the new GPT-2 image generator does it for real (though not perfect) Here’s the famous AI task horizons graph with a touch of Basquiat, haunted by ghosts, from the Voynich manuscript, as a decaying pier. https://x.com/emollick/status/2046728271849550331
No bad ideas when you’re playing with ChatGPT Images 2.0 → Smarter visuals → Better editing and aesthetics Rolling out in Figma and Figma Weave https://x.com/figma/status/2046673364496875977
This wasn’t the case with previous image generators, but the LLM you select has a huge effect on GPT-imagegen-2 output. GPT-5.4 Thinking and GPT-5.4 Pro will produce much better images, especially for complex things. This is, of course, not intuitive or explained anywhere. https://x.com/emollick/status/2046960756608868533
Anthropic
Anthropic commits $100 billion to Amazon for 5GW expansion to train Claude Anthropic committed $1 billion to Amazon to reserve enough computing power to train its next version of Claude.
We’re expanding our collaboration with Amazon to secure up to 5 gigawatts of compute for training and deploying Claude. Capacity begins coming online this quarter, with nearly 1 gigawatt expected by the end of 2026. https://x.com/AnthropicAI/status/2046327624092487688
OpenAI’s Codex surges to 4 million users as new features let it control your computer, browse the web, and work through complex tasks on its own OpenAI’s programming tool, Codex, is a competitor to Claude Code. Codex now has 4 million users and has released the ability to allow Codex to control your computer and browse the web, so you can leave your computer and let it run.
always a real feeling of magic to ask codex to perform a task that requires finding information scattered across slack, google docs, notion, and various internal tools, and it just figures it out https://x.com/gdb/status/2044643518891909289
Auto-review is a new mode that lets Codex work longer with fewer approvals and safer execution. It helps Codex keep moving through tests, builds, and more, including during long tasks and automations, while a separate agent checks higher-risk steps in context before they run. https://x.com/OpenAIDevs/status/2047436655863464011
Codex Computer feels like the first really usable computer use platform. More importantly, it shows that the tech has arrived and now we will see a wave of things get unlocked. Enterprise software will never be the same again. All the legacy stuff that will never see an API is https://x.com/matvelloso/status/2045209294942142860
GPT Image 2 + Codex: or how to make Codex not suck at UI. Step 1: Generate a UI image (native in Codex) Step 2: Get Codex to implement the UI based on it Step 3: Get Codex to iterate until it aligns with the image as much as possible Codex is bad at initial UI, but very good at https://x.com/petergostev/status/2046720618566242657
man Codex Computer Use is actually so good i’ve got my guy sending Slack messages, reading my Slack bookmarks, checking stuff on my browser, and i’m still trying more things it’s legit so good https://x.com/kr0der/status/2045154074337710136
Some of you were disappointed that we “only” get an image model from OpenAI today. But you need to see the big picture: GPT-Image-2 can generate mockups of websites, which Codex can then turn straight into working code. That’s one of the exciting new use cases enabled by true https://x.com/mark_k/status/2046640315348725879
With GPT-5.5, Codex now gets more of the job done across the browser, files, docs, and your computer. We’ve expanded browser use so Codex can interact with web apps, and test flows, click through pages, capture screenshots, and iterate on what it sees until it completes the https://x.com/OpenAIDevs/status/2047381283358355706
OpenAI’s GPT-5.5 reclaims the top spot in AI rankings, with a new model built to handle complex, multi-step work with less human guidance OpenAI also released its latest GPT model, GPT-5.5, which is neck and neck with Opus on several benchmarks.
Huch like Anthropic has with its Figma competitor, OpenAI is integrating its tools (chat, code, imagery, computer use), so users can generate images and designs/layouts within Codex. We can see symmetry between the OpenAI and Anthropic product lines as multimodality (chat, video, imagery, code, and audio) starts to be integrated across all the apps and APIs . Introducing GPT-5.5 | OpenAI https://openai.com/index/introducing-gpt-5-5/
GPT-5.5 is here. It’s our smartest frontier model yet, introducing a new class of intelligence for agentic coding, computer use, knowledge work, and scientific research. Rolling out in ChatGPT and Codex today. API is coming soon. https://x.com/OpenAIDevs/status/2047377079352877534
GPT-5.5 takes OpenAI back to the clear number one in AI. OpenAI’s new model tops the Artificial Analysis Intelligence Index by 3 points, breaking a three-way tie with Anthropic and Google OpenAI gave us pre-release access to test all five reasoning effort levels: xhigh, high, https://x.com/ArtificialAnlys/status/2047378419282034920
GPT-5.5, not fully saturating the TikZ unicorn test yet but getting awfully close … (yes this is actual TikZ code, I personally find it so unbelievable that I’m putting the code below for anyone to verify for themself) https://x.com/sebastienbubeck/status/2047383628922167390?s=46
gpt-image-2 is here, available today in the API and Codex. The most capable image generation model yet, built for production-grade workflows with stronger text rendering, layout, editing, resolution, and multilingual rendering. https://x.com/OpenAIDevs/status/2046671238534496259
I’ve been an early tester of GPT-5.5, and it destroyed the “”GPT Plays Pokémon FireRed”” benchmark. GPT-5.4 never finished the game, it got stuck in a loop, reloading the last save and retrying the final rival fight over and over. GPT-5.5 not only beat it on the first try, but did https://x.com/clad3815/status/2047392779006013833?s=12
Introducing GPT-5.5 A new class of intelligence for real work and powering agents, built to understand complex goals, use tools, check its work, and carry more tasks through to completion. It marks a new way of getting computer work done. Now available in ChatGPT and Codex. https://x.com/OpenAI/status/2047376561205325845
Anthropic
MCP downloads triple to 300M monthly as agents reach production A popular agent connector tool, MCP, has reached 300 million monthly downloads. This is a sign that new agents are popping up everywhere, even if we don’t know about them.
Google rebrands its AI developer platform as Gemini Enterprise Agent Platform, adding tools to build, monitor, and control fleets of business AI agents Google has completely rebranded its Gemini developer platform to be the Gemini Enterprise Agent Platform.
The conversation around AI agents is no longer about how to build them — it’s about how to manage thousands of them. Today we’re introducing Gemini Enterprise Agent Platform, a new way to build, scale, govern and optimize agents. It combines the models and services you’re https://x.com/Google/status/2046985650868547851
We’re launching Gemini Enterprise Agent Platform with @GoogleCloud: a platform for businesses to develop, scale, govern and optimize agents. It’s the evolution of Vertex AI, bringing together model selection and agent building with new features for integration, security and https://x.com/GoogleDeepMind/status/2046983340524269713
We’re making it easier for organizations to scale up autonomous agents with the Agentic Data Cloud. This AI-native architecture closes the gap between thinking and doing through: – A universal context engine that gives your agents a grounded source of truth about your business https://x.com/Google/status/2046997032649277754
Google Cloud and Wiz launch AI-powered security tools that hunt threats, protect AI agents, and flag unauthorized AI apps across organizations Google released an AI-powered security tool that you can set to automatically look for threats across an organization. You can just leave it on, and it basically crawls and scours all of the systems, keeping an eye out for any kind of risks.
We’re delivering agentic defense by combining Google’s Threat Intelligence and Security Operations with @wiz_io’s Cloud and AI Security Platform to detect, prevent and respond to threats. Security agents provide protection for your entire AI development lifecycle. https://x.com/Google/status/2047000216188940710
HuggingFace
An AI agent completed a full machine learning research project on its own, from finding data to training a model to publishing the results Hugging Face is an open-source code repository, and they’ve released an agent called Intern, which can take on fairly complicated tasks. One person had Hugging Face’s Intern fine-tune a Segment Anything model (Meta’s image understanding tool) on a medical dataset and then provide a tutorial as a blog post, and it did it all on its own.
I tested @huggingface ml-intern, given the prompt “”Fine-tune a Segment Anything Model (SAM) on a useful medical dataset. Train the model, and provide a comprehensive tutorial in a Jupyter Notebook file. Additionally, create a Hugging Face article/blog post documenting https://x.com/Mayank_022/status/2046646301555900828
Op-Eds
Wharton’s Ethan Mollick argues that studying literature and art builds something AI can’t generate: the judgment to know what’s actually goodANDAI is erasing a basic assumption: that the things we encounter were made by someone who devoted real time and effort to them Wharton’s Ethan Mollick posted that we’re reaching a point where it’s going to be hard to understand whether anything we see or read is the result of a lifetime of learning or simply artificial intelligence prompting. The idea of a thesis being created by an AI research agent means that we have to constantly be aware of whether this is a document someone went to school for years to produce or simply content (even if it’s viable) that was prompted or agentically produced.
Ethan argues that this is where a humanities degree will become helpful, because in a world where anyone can generate a tsunami of slop, understanding taste and quality is going to be a critically human skill.
An advantage of the humanities — of reading & seeing a wide range of great works from many perspectives & cultures — is that you develop your sense of taste. In a world where anyone can produce a flood of writing & visual output is for cheap, that has never been more important. https://x.com/emollick/status/2046774096730411360
One thing thing about AI, for better and worse, is that “”everything around me is somebody’s life work”” is no longer a true assumption going forward. https://x.com/emollick/status/2045318277958709540
Publishing
AI-generated track tops global iTunes chart ?! As if on cue, this week an AI-generated song took the No. 1 spot on the iTunes global charts (!), and a new report found that AI-generated music now represents 44% of all newly uploaded music.
Salesforce rebuilds its entire platform so AI agents can run it without anyone ever logging in or clicking through menus Salesforce announced that it has rebuilt its entire platform so that AI agents can run on it without logging in or clicking menus. Basically, anything that can be used on Salesforce now has connectors to enable agents to access and use it.
OpenAI launches free clinical ChatGPT for verified US healthcare providers OpenAI launched a free clinical version of ChatGPT for healthcare providers. It has been built explicitly to help with medical research and documentation and drive higher-quality patient care.
Allen Institute’s OlmoEarth reads satellite images to map forests, spot wildfires, and track land changes with almost no setup required AllenAI has become my favorite unknown AI lab. They’re a nonprofit, and I’m not sure if everyone would say they’re a frontier lab. However, they have come out with the neatest tools for audio and visual multimodal AI. This week, the Allen Institute released OLMo Earth, which can be trained to understand satellite images, map forests and wildfires, and track the way land changes over time.
Allen’s models are extremely good at visual understanding. They have one model that can watch a video of someone cooking food, understand all of the ingredients and amounts, and actually back out a recipe simply by watching someone cook.
Open-source Hermes Agent surpasses OpenClaw in developer adoption Another important, rather unknown, company to follow in the United States is Nous Research. They have an open-source agent called Hermes that is actually surpassing OpenCLAW in volume. It’s difficult for me to believe, but that’s what I’m seeing this week. In particular, Hermes overtook OpenCLAW in weekly GitHub stars. There tends to be a lot of marketing out there about Hermes, so I have to parse it, but even when you extract the signal from the noise, it’s worth knowing that Hermes is a big competitor with OpenCLAW if you’re following those types of strong, fully autonomous, locally run, agents.
Hermes Agent recently overtook OpenClaw in weekly new GitHub stars. Two months after launch, the agent from @NousResearch is pulling developers from the incumbent at meaningful scale. It’s self-hosted and open source, so you control what leaves your machine. It reaches you https://x.com/Delphi_Digital/status/2045839142450536504
Introducing Hermes Agent v0.11.0 Our largest update yet, with over 700 PRs across ~200 contributors. Thank you to everyone who’s worked on Hermes Agent! This update features a beta TUI v2, unlimited recursion depth and width of subagents, 5 new LLM providers, expanded image gen https://x.com/Teknium/status/2047506967909015907
Meta
Meta plans to cut 10% of staff in May restructuring Meta announced plans to cut 10% of its staff. While this sounds like (and is) a huge number, I’m pretty sure this just takes them back to 2024 levels after quite a bit of hiring during and post-pandemic.
Palantir goes viral by turning political ideology into a business strategy with a manifesto Defense company Palantir went viral when its CEO published a manifesto, which was basically his political ideology turned into a business strategy. It’s a little surreal to read, simply because it reads out of a movie, but I don’t think there’s anything too surprising from a defense contractor’s point of view. It’s worth skimming if you have a moment.
Introducing one of our biggest updates to the Gemini Deep Research Agent, now available via the Interactions API! Trigger complex, long-horizon research workflows with arbitrary MCP support, get rich visualizations, plan before you execute, and more with these two https://x.com/googleaidevs/status/2046630912054763854
Introducing our biggest upgrades to the Deep Research API yet… including Deep Research Max (our SOTA system), MCP support, Native charts & infographics, planning mode, full tool support (including Google tools), full multi-modal input support, & real-time progress streaming! https://x.com/OfficialLoganK/status/2046628030777631000
The next evolution of our autonomous research agent is here. Today, we’re introducing Deep Research and Deep Research Max via the Gemini API. Powered by Gemini 3.1 Pro, you can now trigger comprehensive research workflows with unprecedented control and transparency, featuring: https://x.com/Google/status/2046627647208259835
We’ve expanded the capabilities of Deep Research and Deep Research Max to give you even more control and transparency over the entire research process: 📋 Collaborative Planning: Review and refine the agent’s research plan before it executes. 🛠️ Extended Tooling: Run Google https://x.com/Google/status/2046627652568850687
DeepSeek V4 Pro launches as largest open model with frontier level benchmarks DeepSeek is once again the open-source king and it’s competitive with frontier models 1st or 2nd place on 12/22 benchmarks https://x.com/scaling01/status/2047512176856899985
Deepseek V4 Pro is the biggest open model ever with 1.6T total 49B active, trained on 33T tokens, 1M context, with 2 new attention mechanisms, Muon, mHC, open source kernels, FP4 QAT, MIT license and with one of the best tech repot of the year https://x.com/eliebakouch/status/2047519300399837677
Google
Google releases unified Gemini Embedding 2 across all media types (text, image, video, audio, and documents) A few weeks ago, Google announced an embedding tool that can understand all types of media. It can understand the semantic relevance between images, videos, and audio. So, for example, if you had a sad song, it could match the emotion of that sad song with a sad movie, a sad story, or a sad picture.
This type of softer emotional connection between different types of media was never possible before. In the past, embedding could only work within the same type of media. So a sad poem could be matched to another sad poem, or a joyful song could be matched to another joyful song. Now, with this new Google tool, emotional or semantic search and association of a feelings can be trained and queried across any type of media set.
Gemini Embedding 2 is now generally available via the Gemini API and Gemini Enterprise Agent Platform search and understand semantic relationships across text, image, video, audio, and documents without complex, fragmented pipelines https://x.com/GoogleAIStudio/status/2047007402520674679
Anthropic’s Opus 4.7 is crushing the unicorn drawing benchmark On the plus side with Opus 4.7, if it does decide to think it produces BY FAR the best Sparks unicorn* ever, even non-thinking is pretty good, if not great. * This is created using TikZ, which is a language built for scientific diagrams & very much not for drawing. The original https://x.com/emollick/status/2044880350237626844
Benchmarks
Claude agents replicate economists’ study with tighter consistency Classic study gave 146 economist teams the same dataset & got wildly different answers New paper reruns it with agentic AI. Claude Code & Codex land near the human median, but with far tighter dispersion & no extremes. Suggests that AI is now useful for doing scalable research. https://x.com/emollick/status/2046362044786458648
Cursor
Cursor strikes xAI compute pact amid $50B funding round… lots of TBDs in the small print The Path Forward for AI Startups A lot of founders are messaging each other after the SpaceXAI <> Cursor “IPO-deferred acquisition”. Common discussion topic: what is the future for independent startups? Must ~everyone ultimately be acquired by a frontier lab or go extinct? https://x.com/russelljkaplan/status/2047077659985981616
SpaceXAI and @cursor_ai are now working closely together to create the world’s best coding and knowledge work AI. The combination of Cursor’s leading product and distribution to expert software engineers with SpaceX’s million H100 equivalent Colossus training supercomputer will https://x.com/SpaceX/status/2046713419978453374
SpaceXAI and @cursor_ai are now working closely together to create the world’s best coding and knowledge work AI. The combination of Cursor’s leading product and distribution to expert software engineers with SpaceX’s million H100 equivalent Colossus training supercomputer will https://x.com/SpaceX/status/2046713419978453374?s=20a
The structure of the deal is pretty interesting here. I think what’s happening is: 1. xAI is having trouble training a SOTA coding model (hence cofounder departures), bunch of idle GPUs 2. Cursor doesn’t have capital to blow on a $5B training run to compete with Codex/Claude 3. https://x.com/0xrwu/status/2046721359263285478
Figure
Figure’s humanoid robot output suggests 2,500-unit annual run rate Figure shows a serious humanoid production ramp. Brett’s graph doesn’t have a y-axis. Assuming June 2023 = 1 unit and the graph is linear: 2023: 2 units 2024: 26 2025: 79 2026 (till Apr 21st): 312 Annualized run-rate based on the past 21 days of production: 2,589/year. https://x.com/TheHumanoidHub/status/2047041923387695504
Moonshot open-sources Kimi K2.6 with state-of-the-art coding benchmarks Kimi K2.6 Tech Blog: Advancing Open-Source Coding https://www.kimi.com/blog/kimi-k2-6
Kimi K2.5 widens gap between the US and China in open weights model intelligence. The leading US open weights model remains OpenAI’s gpt-oss-120b, which has now been eclipsed by an ever-growing list of open weights releases from China. https://x.com/ArtificialAnlys/status/2016250140219343163?s=20
Kimi K2.6 autonomously overhauled exchange-core, an 8-year-old open-source financial matching engine. Over a 13-hour execution, the model iterated through 12 optimization strategies, initiating over 1,000 tool calls to precisely modify more than 4,000 lines of code. Acting as https://x.com/Kimi_Moonshot/status/2046531057147933137
NVIDIA
NVIDIA CEO Jensen Huang’s Dwarkesh podcast interview is packed with wild moments Also, somehow everyone missed that Jensen Huang all but called Dario Amodei’s mindset a loser’s mindset https://x.com/TheTuringPost/status/2046585887400604116
Jensen on the famous story about Larry Ellison and Elon Musk begging him for GPUs over dinner: “”That never happened. We absolutely had dinner, and it was a wonderful dinner. At no time did they beg for GPUs. They just had to place an order.”” Jensen says Nvidia’s allocation https://x.com/dwarkesh_sp/status/2044989230112506351
Nvidia has locked up many years of scarce components – almost a hundred billion dollars in purchase commitments. Is this Nvidia’s big moat? A competitor might design a great accelerator, but they don’t have Jensen’s LTAs with SK Hynix, TSMC, etc. Jensen: “If our next several https://x.com/dwarkesh_sp/status/2044808033411223559
The most viral moment of the @dwarkesh_sp x Jensen Huang conversation was widely misunderstood It wasn’t really just about China, TPUs, or even Nvidia’s moat. It was about a much deeper disagreement over what it means for America to “”win”” in AI https://x.com/TheTuringPost/status/2046366547619270665
OpenAI
Stargate data center buildout slips behind schedule in Texas In 2025, OpenAI announced Stargate, a $500 billion data center initiative. We surveyed all 7 US sites and found visible development at each. There’s a long road ahead, but the project appears on track to reach 9+ GW by 2029—comparable to New York City’s peak power demand. 🧵 https://x.com/EpochAIResearch/status/2045258390147088764
We now estimate that only about 0.3 GW of total facility power is operational for Stargate Abilene, not 0.6 GW. We have moved the 0.6 GW milestone to late May and the 1.2 GW milestone from Q3 to Q4 2026, but both are uncertain. More about this change and our methodology in 🧵 https://x.com/EpochAIResearch/status/2047442515608162481
Hyatt rolls out OpenAI tools to employees worldwide “Hyatt’s innovative approach with OpenAI reflects how Hyatt is elevating its use of technology and enhancing human connections. The company is making artificial intelligence broadly accessible to its employees, enabling teams to spend less time on manual” / X https://x.com/TheRealAdamG/status/2046262564158333211
OpenAI just open sourced a new 1.5B (50m active) model on HuggingFace with Apache 2.0 license! It’s not a new LLM, this one is called Privacy Filter, and it’s a PII detection model (checking if text has private information) A few interesting tidbits from the release + links: https://x.com/altryne/status/2046977133013311814
This story has now been updated with more details. Three leaders departed from OpenAI today: – Kevin Weil, VP of OpenAI for Science – Srinivas Narayanan, CTO of B2B Applications – Bill Peebles, Head of Sora https://x.com/zeffmax/status/2045248266384838800?s=46
Today is my last day at OpenAI, as OpenAI for Science is being decentralized into other research teams. It’s been a mind-expanding two years, from Chief Product Officer to joining the research team and starting OpenAI for Science. Accelerating science will be one of the most https://x.com/kevinweil/status/2045230426210648348?s=20
Workspace agents can work across tools—pulling context from docs, email, chats, code, and systems, and taking approved actions like updating @Linear issues, creating docs, or sending messages. In @SlackHQ, agents can jump into a thread, understand what’s needed, pull the right https://x.com/OpenAI/status/2047008991944069624
Alibaba’s Qwen3.6-27B beats much larger model on coding 🚀 Meet Qwen3.6-27B, our latest dense, open-source model, packing flagship-level coding power! Yes, 27B, and Qwen3.6-27B punches way above its weight. 👇 What’s new: 🧠 Outstanding agentic coding — surpasses Qwen3.5-397B-A17B across all major coding benchmarks 💡 Strong https://x.com/Alibaba_Qwen/status/2046939764428009914
Post-trained Qwen model hits Pareto-optimal accuracy-cost frontier We’ve post trained a model on top of Qwen that achieves Pareto optimality on accuracy-cost curves. Unlike our previous post trained models, this model has been trained to be good at search and tool calls simultaneously, allowing us to unify the tool call router and https://x.com/AravSrinivas/status/2047019688920756504
Qwen 3.6-Max-Preview solves AIME-2026 #15 after like 30 minutes of thinking, but on first try. Preview or not, it’s more baked than DeepSeek-Expert. Other tests validate this impression. It doesn’t screw up. Alibaba Qwen is, after all, a frontier lab. https://x.com/teortaxesTex/status/2046166258853269990
Robotics
US military swaps $75K robot dogs for $3K coyote rovers To prevent dangerous airbase bird strikes, the US military last year replaced $75k Boston Dynamics’ quadrupeds with $3k rovers attached with plastic coyote decoys. Making security robots look as human as possible will serve as a similar psychological visual deterrent. https://x.com/TheHumanoidHub/status/2044860410420183489
Tesla
Tesla files five Optimus patents focused on robotic hand mechanics Tesla published five new patents yesterday, all for the Optimus robotic hand and arm: 1. Covers cable routing through the wrist joint. The key innovation is that cables are arranged in a lateral stack on the forearm side and transition to a vertical stack on the hand side, https://x.com/TheHumanoidHub/status/2045270641210011767
Tesla pivots Fremont factory to mass-produce one million Optimus robots Elon on Q1 call, confirming what he posted last week: “”AI5 chip will (first) go into Optimus and the data center. It’s looking like we’ll be able to achieve unsupervised robotaxi with AI4, which is far better than human safety, so AI5 is certainly not immediately needed in https://x.com/TheHumanoidHub/status/2047075601312633316
Elon on Q1 earnings call: – Optimus v3 is functional, but there are some aesthetic elements to be worked on. It may be unveiled in the middle of 2026. – We’re a little hesitant to show it off because we find our competitors do a frame-by-frame analysis and copy everything we https://x.com/TheHumanoidHub/status/2047070791234375765
Tesla in Q1 2026 Shareholder Update: – Preparations for our first large-scale Optimus factory will begin shortly in Q2. – First-gen line designed for 1M robots/year will replace Model S/X lines in Fremont Factory, second-gen line is being prepared at Giga Texas (long-term https://x.com/TheHumanoidHub/status/204704948987438309
Automated Executive Summaries with The Same Links, Generated by Claude Haiku 4.5 (I test it every week to see how it does)
AI training power is set to multiply a thousandfold by 2028. Frontier AI models have already grown a trillion times more powerful since 2010, and researchers expect another thousand-fold jump in computing capacity by 2028—a pace that would dramatically accelerate AI capabilities and raise questions about resource consumption, competitive dynamics, and the feasibility of sustaining such exponential growth.
Since I began work on AI in 2010, training compute for frontier models has grown by one trillion times. Now we’re looking at something like another thousand-fold growth in effective compute by the end of 2028. 1000x the existing 1,000,000,000,000x. Extraordinary stuff. https://x.com/mustafasuleyman/status/2046989133676257284
Google’s AI chip clusters now scale to one million processors in single system. Google announced it can now operate a million of its custom AI training chips (TPU8t) within a single interconnected cluster, a significant leap in computational scale that enables larger and more complex AI model training. This matters because it removes a major technical bottleneck—previously, scaling beyond certain thresholds required splitting work across separate systems, adding complexity and latency. The achievement signals Google can now support proportionally larger AI models than competitors constrained by smaller unified clusters.
White House accuses China of stealing American AI models at scale. The U.S. government released a policy memo on Wednesday alleging that Chinese entities are conducting “industrial-scale” distillation campaigns—using automated queries to extract capabilities from frontier AI models without permission—and committed to sharing intelligence with American companies and exploring accountability measures. The technique is legal gray area: it doesn’t require stealing code, just feeding thousands of carefully crafted questions to a model and training a cheaper rival on the responses, a method Anthropic documented in 24,000 fraudulent accounts generating 16 million exchanges with Claude. The memo arrives weeks before a planned Trump-Xi summit, positioning AI protection as both national security policy and negotiating leverage, though enforcement challenges remain steep since distillation happens over the internet rather than at physical borders.
Anthropic launches Claude Design to let users build prototypes and presentations through conversation. Anthropic released Claude Design, a new tool powered by its latest vision model that generates polished designs, prototypes, and pitch decks from text descriptions—letting teams skip traditional design software workflows. The product directly competes with Figma and similar design tools by automating visual creation, automatically applying brand systems, and seamlessly handing off designs to code implementation, addressing a real friction point where non-designers struggle to visualize and share ideas quickly.
Anthropic exec Mike Krieger left Figma’s board this week after reports of an incoming launch of a competing product. Now, Claude Design is live. How it works: describe the design and Claude Opus 4.7 builds the first version. Refine with inline comments, direct edits, or https://x.com/TheRundownAI/status/2045176722476208454
Anthropic released Claude design, direct attack on figma and lovable. Anthropic just shipped Claude Design, powered by Claude Opus 4.7., a tool that turns conversations into polished prototypes, pitch decks, and marketing assets. It auto-applies your brand system, lets you https://x.com/kimmonismus/status/2045162358004216134
Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our most capable vision model. Available in research preview on the Pro, Max, Team, and Enterprise plans, rolling out throughout the day. https://x.com/claudeai/status/2045156267690213649
OpenAI’s GPT-Image-2 dominates text-to-image rankings with record margin. OpenAI launched ChatGPT Images 2.0, an image generation model that combines reasoning capabilities with improved instruction-following, text rendering, and real-time web search. The model achieved the #1 position across all competitive benchmarks with a 242-point lead—the largest gap ever recorded—outperforming competitors at tasks including product design, complex layouts, and text-heavy visuals. The upgrade is now available to ChatGPT users, with reasoning features limited to paid tiers.
Arena Trends: Text-to-Image, Jan 2026 – Apr 2026 For most of the year, @GoogleDeepMind and @OpenAI traded the top spot within a tight margin – GPT-Image vs. Nano Banana – with the rest of the field clustered below 1,200. Today, GPT-Image-2 breaks away with a score of 1,512, 242 https://x.com/arena/status/2046690103515648061
A Visual Thought Partner ChatGPT Images 2.0 is our first image model with thinking capabilities. When a thinking model is selected in ChatGPT, Images 2.0 can search the web for real-time information, create multiple distinct images from one prompt, double-check its own outputs, https://x.com/OpenAI/status/2046670989719924768
ChatGPT Images 2.0 is a big leap forward in image generation intelligence. It’s much better at following detailed instructions, rendering dense text, understanding the world more accurately, and creating visuals that are more useful. And when you give it additional time to https://x.com/nickaturley/status/2046677986242363731
ChatGPT Images 2.0 is available starting today to all ChatGPT and Codex users. Images with thinking are available to ChatGPT Plus, Pro, and Business users (Enterprise soon). On mobile, make sure you update to the latest version of the app. The underlying model, gpt-image-2, is https://x.com/OpenAI/status/2046670994413322435
Exciting news – GPT-Image-2 by @OpenAI has claimed the #1 spot across all Image Arena leaderboards! A clean sweep with a record-breaking +242 point lead in Text-to-Image – the largest gap we’ve seen to date. – #1 Text-to-Image (1512), +242 over #2 (Nano-banana-2 with web-search https://x.com/arena/status/2046670703311884548
GPT-Image-2 takes #1 in every single Text-to-Image category — all 7 of them. Surpassing the next leading model (Nano-banana-2 with web-search) across the board. Here’s the drill-down on improvements vs. its predecessor, GPT-Image-1.5-High-Fidelity: – #1 Product, Branding & https://x.com/arena/status/2046670705958551938
GPT-ImageGen-2 did this in one shot, with just the prompt “”turn all of Tennyson’s Ulysses into a comic, across as many pages as needed. make it great, include the full text”” 10 pages, though it did use what seems to be the ImageGen-2 ‘s preferred “”spackled drawing”” style 1/ https://x.com/emollick/status/2046843402021380556
I have been using GPT ImageGen-2 for the past weeks I didn’t think that better image-generators would be a big deal but it turns out that there is a quality threshold I didn’t expect, where you can now get text, slides, academic papers Look at what it does with my “”otter test””! https://x.com/emollick/status/2046665274535854146
Introducing ChatGPT Images 2.0 A state-of-the-art image model that can take on complex visual tasks and produce precise, immediately usable visuals, with sharper editing, richer layouts, and thinking-level intelligence. Video made with ChatGPT Images https://x.com/OpenAI/status/2046670977145372771
My most popular AI post was a bunch of made-up “”graphs”” four years ago. Now, the new GPT-2 image generator does it for real (though not perfect) Here’s the famous AI task horizons graph with a touch of Basquiat, haunted by ghosts, from the Voynich manuscript, as a decaying pier. https://x.com/emollick/status/2046728271849550331
No bad ideas when you’re playing with ChatGPT Images 2.0 → Smarter visuals → Better editing and aesthetics Rolling out in Figma and Figma Weave https://x.com/figma/status/2046673364496875977
This wasn’t the case with previous image generators, but the LLM you select has a huge effect on GPT-imagegen-2 output. GPT-5.4 Thinking and GPT-5.4 Pro will produce much better images, especially for complex things. This is, of course, not intuitive or explained anywhere. https://x.com/emollick/status/2046960756608868533
Anthropic secures five gigawatts of Amazon computing power through 2036 Anthropic and Amazon expanded their partnership with a $100 billion, decade-long commitment to build AI infrastructure, with the first major capacity additions arriving within months to handle surging demand for Claude. The deal reflects real constraints: Claude’s revenue run rate jumped from $9 billion to $30 billion in just months, straining systems during peak usage. This infrastructure play is distinctive because it locks in custom chip technology across multiple generations, positioning Amazon’s homemade processors as critical to frontier AI development rather than generic cloud compute.
We’re expanding our collaboration with Amazon to secure up to 5 gigawatts of compute for training and deploying Claude. Capacity begins coming online this quarter, with nearly 1 gigawatt expected by the end of 2026. https://x.com/AnthropicAI/status/2046327624092487688
Anthropic’s Claude chatbot company hits $1 trillion valuation, surpassing OpenAI. Anthropic reached a $1 trillion valuation on secondary share markets, exceeding OpenAI’s $880 billion, driven by its Claude Code tool’s rapid developer adoption and major partnerships with Amazon and Palantir. The company’s annualized revenue jumped from $9 billion to $39 billion in just four months, signaling that investor demand now exceeds available shares—a scarcity that inflates valuations beyond traditional metrics. This marks a significant competitive shift, as Anthropic was valued at only $380 billion three months prior.
OpenAI’s Codex gains screen-reading memory to reduce repetitive context. Chronicle, a new opt-in feature for Codex on macOS, lets the AI assistant watch your screen and remember what you’re working on, eliminating the need to repeatedly explain your context. The feature aims to boost productivity by helping Codex understand ongoing projects and identify relevant tools automatically, though it comes with security trade-offs—unencrypted local storage and potential risks from screen capture—requiring users to pause it during sensitive work.
always a real feeling of magic to ask codex to perform a task that requires finding information scattered across slack, google docs, notion, and various internal tools, and it just figures it out https://x.com/gdb/status/2044643518891909289
Auto-review is a new mode that lets Codex work longer with fewer approvals and safer execution. It helps Codex keep moving through tests, builds, and more, including during long tasks and automations, while a separate agent checks higher-risk steps in context before they run. https://x.com/OpenAIDevs/status/2047436655863464011
Codex Computer feels like the first really usable computer use platform. More importantly, it shows that the tech has arrived and now we will see a wave of things get unlocked. Enterprise software will never be the same again. All the legacy stuff that will never see an API is https://x.com/matvelloso/status/2045209294942142860
GPT Image 2 + Codex: or how to make Codex not suck at UI. Step 1: Generate a UI image (native in Codex) Step 2: Get Codex to implement the UI based on it Step 3: Get Codex to iterate until it aligns with the image as much as possible Codex is bad at initial UI, but very good at https://x.com/petergostev/status/2046720618566242657
man Codex Computer Use is actually so good i’ve got my guy sending Slack messages, reading my Slack bookmarks, checking stuff on my browser, and i’m still trying more things it’s legit so good https://x.com/kr0der/status/2045154074337710136
Some of you were disappointed that we “only” get an image model from OpenAI today. But you need to see the big picture: GPT-Image-2 can generate mockups of websites, which Codex can then turn straight into working code. That’s one of the exciting new use cases enabled by true https://x.com/mark_k/status/2046640315348725879
With GPT-5.5, Codex now gets more of the job done across the browser, files, docs, and your computer. We’ve expanded browser use so Codex can interact with web apps, and test flows, click through pages, capture screenshots, and iterate on what it sees until it completes the https://x.com/OpenAIDevs/status/2047381283358355706
OpenAI releases GPT-5.5, claiming clear lead in AI intelligence. OpenAI launched GPT-5.5, a new AI model designed to work more autonomously on complex tasks like coding, research, and data analysis without requiring step-by-step human direction. The model reportedly ranks first on OpenAI’s intelligence benchmark by a three-point margin, matches the speed of its predecessor while performing significantly better, and is rolling out today to ChatGPT and Codex users, with API access coming soon. The release is notable for emphasizing efficiency and agentic capability—the ability to plan and execute multi-step work independently—rather than raw scale alone.
GPT-5.5 is here. It’s our smartest frontier model yet, introducing a new class of intelligence for agentic coding, computer use, knowledge work, and scientific research. Rolling out in ChatGPT and Codex today. API is coming soon. https://x.com/OpenAIDevs/status/2047377079352877534
GPT-5.5 takes OpenAI back to the clear number one in AI. OpenAI’s new model tops the Artificial Analysis Intelligence Index by 3 points, breaking a three-way tie with Anthropic and Google OpenAI gave us pre-release access to test all five reasoning effort levels: xhigh, high, https://x.com/ArtificialAnlys/status/2047378419282034920
GPT-5.5, not fully saturating the TikZ unicorn test yet but getting awfully close … (yes this is actual TikZ code, I personally find it so unbelievable that I’m putting the code below for anyone to verify for themself) https://x.com/sebastienbubeck/status/2047383628922167390?s=46
gpt-image-2 is here, available today in the API and Codex. The most capable image generation model yet, built for production-grade workflows with stronger text rendering, layout, editing, resolution, and multilingual rendering. https://x.com/OpenAIDevs/status/2046671238534496259
I’ve been an early tester of GPT-5.5, and it destroyed the “”GPT Plays Pokémon FireRed”” benchmark. GPT-5.4 never finished the game, it got stuck in a loop, reloading the last save and retrying the final rival fight over and over. GPT-5.5 not only beat it on the first try, but did https://x.com/clad3815/status/2047392779006013833?s=12
Introducing GPT-5.5 A new class of intelligence for real work and powering agents, built to understand complex goals, use tools, check its work, and carry more tasks through to completion. It marks a new way of getting computer work done. Now available in ChatGPT and Codex. https://x.com/OpenAI/status/2047376561205325845
Claude upgrades agent platform with production-ready integration protocol. Anthropic released detailed guidance for building AI agents that connect to cloud systems using MCP (Model Context Protocol), a standardized framework now handling millions of daily connections. The protocol matters because production agents increasingly operate in cloud environments where they need secure, scalable access to enterprise data and services—MCP solves this by providing a common layer that works across different AI platforms and deployment types, rather than requiring custom integration work for each agent-service combination. Downloads of MCP SDKs have tripled to 300 million monthly since January, reflecting broad adoption across major AI platforms and enterprises building production agents.
Google launches enterprise platform to build and govern AI agents at scale. Google is consolidating its AI development tools into a new Gemini Enterprise Agent Platform designed to let companies build, deploy, and control multiple AI agents across their operations. The platform addresses a shift from managing individual AI tasks to delegating business outcomes autonomously, offering governance controls, security guardrails, and integration with over 200 AI models. Real-world customers like Comcast, PayPal, and L’Oréal are already using it to handle complex workflows—from customer service to financial operations—with persistent memory and multi-agent coordination capabilities.
The conversation around AI agents is no longer about how to build them — it’s about how to manage thousands of them. Today we’re introducing Gemini Enterprise Agent Platform, a new way to build, scale, govern and optimize agents. It combines the models and services you’re https://x.com/Google/status/2046985650868547851
We’re launching Gemini Enterprise Agent Platform with @GoogleCloud: a platform for businesses to develop, scale, govern and optimize agents. It’s the evolution of Vertex AI, bringing together model selection and agent building with new features for integration, security and https://x.com/GoogleDeepMind/status/2046983340524269713
We’re making it easier for organizations to scale up autonomous agents with the Agentic Data Cloud. This AI-native architecture closes the gap between thinking and doing through: – A universal context engine that gives your agents a grounded source of truth about your business https://x.com/Google/status/2046997032649277754
Google and Wiz deploy AI agents to automate enterprise cybersecurity threat detection. Google Cloud introduced three new AI agents for security operations that automate threat hunting, detection creation, and threat analysis—reducing manual work from 30 minutes to 60 seconds per alert. The partnership with Wiz extends protection across multiple cloud platforms and AI development environments, addressing the growing reality that attackers are using AI to accelerate attacks; one security firm found threat handoff times have compressed from eight hours to 22 seconds over three years.
We’re delivering agentic defense by combining Google’s Threat Intelligence and Security Operations with @wiz_io’s Cloud and AI Security Platform to detect, prevent and respond to threats. Security agents provide protection for your entire AI development lifecycle. https://x.com/Google/status/2047000216188940710
Hugging Face intern creates fine-tuned medical image segmentation model. An AI intern successfully adapted Segment Anything Model (SAM)—a general-purpose image recognition tool—for medical imaging tasks, then documented the process in a tutorial and blog post. This matters because SAM is typically trained on everyday photos, but medical imaging requires specialized adaptation; the public tutorial accelerates how quickly other developers can apply similar techniques to healthcare datasets, potentially speeding up adoption of AI in clinical workflows.
I tested @huggingface ml-intern, given the prompt “”Fine-tune a Segment Anything Model (SAM) on a useful medical dataset. Train the model, and provide a comprehensive tutorial in a Jupyter Notebook file. Additionally, create a Hugging Face article/blog post documenting https://x.com/Mayank_022/status/2046646301555900828
AI tools democratize content creation, but human judgment becomes scarcer. As artificial intelligence makes it cheap and easy for anyone to generate endless writing and images, the ability to recognize quality—something developed through deep engagement with great works across cultures—has become a rare and valuable skill. This shift mirrors historical moments when abundance of supply made discernment the real bottleneck, not access to tools.
An advantage of the humanities — of reading & seeing a wide range of great works from many perspectives & cultures — is that you develop your sense of taste. In a world where anyone can produce a flood of writing & visual output is for cheap, that has never been more important. https://x.com/emollick/status/2046774096730411360
AI systems now can replicate what took humans years to master. Machine learning models trained on vast datasets can now produce outputs—writing, images, code, analysis—that previously represented years of specialized human effort. This democratizes creation but raises urgent questions about attribution, compensation, and the value of expertise. The shift fundamentally changes how society thinks about skill acquisition and creative work’s economic worth.
One thing thing about AI, for better and worse, is that “”everything around me is somebody’s life work”” is no longer a true assumption going forward. https://x.com/emollick/status/2045318277958709540
AI-generated music now dominates new uploads on streaming platforms. Nearly 75,000 AI tracks are uploaded daily to Deezer, representing 44% of all new music, while an AI song simultaneously claimed the top spot on iTunes’ global charts. Though AI tracks account for only 1–3% of actual streams on Deezer, the platform found that 85% of those streams were fraudulent, highlighting how AI music is being weaponized to game the system rather than reach genuine listeners. The dual developments underscore a critical moment: AI music production is accelerating faster than the industry’s defenses, with one survey showing 80% of listeners want AI music clearly labeled and 52% believe AI tracks should be excluded from main charts.
Salesforce makes its entire platform accessible to AI agents without browser interfaces. Salesforce has rebuilt its 25-year-old platform to let AI agents directly access all customer data, workflows, and business logic through APIs and commands rather than forcing users to click through interfaces. The move, called Headless 360, reflects a shift toward “agentic enterprises” where AI handles routine work like case updates and approvals within conversations on Slack and other channels. The company is backing this with new monitoring tools to ensure agents behave reliably in production, addressing a core enterprise concern about deploying AI systems at scale.
OpenAI releases free ChatGPT version specifically designed for clinicians OpenAI launched ChatGPT for Clinicians, a free AI tool tailored for medical documentation, research, and clinical decision support, available to verified U.S. physicians, nurse practitioners, physician assistants, and pharmacists. The move addresses widespread burnout from administrative tasks: physician AI adoption jumped from 48% to 72% in a year, with clinician usage of ChatGPT doubling annually. Testing by physicians rated 99.6% of responses as safe and accurate, and the tool outperformed both competing AI systems and human physicians on benchmark tasks, though OpenAI emphasizes it supplements rather than replaces clinical judgment.
AI2 releases Earth-observation embeddings for satellite analysis Allen Institute for AI has made OlmoEarth embeddings—compressed numerical summaries of satellite imagery—available through its Studio platform, allowing researchers and developers to search, classify, and analyze Earth observation data without labeled training datasets. The embeddings demonstrated strong performance on tasks like land-cover mapping (achieving 0.84 accuracy from just 60 labeled pixels) and similarity search, with all code and model weights open-sourced so outputs are fully transparent and reproducible.
Open source agent Hermes overtakes rival with developer momentum. Hermes Agent, an open-source AI tool from Nous Research, has surpassed OpenClaw in weekly GitHub adoption just two months after launch—a rare shift in developer preference driven by its self-hosted design that keeps data on users’ machines. The latest version 0.11.0 brings substantial improvements including a redesigned interface, expanded AI model support, and deeper automation capabilities across 200+ contributors, signaling sustained community confidence in the project.
Hermes Agent recently overtook OpenClaw in weekly new GitHub stars. Two months after launch, the agent from @NousResearch is pulling developers from the incumbent at meaningful scale. It’s self-hosted and open source, so you control what leaves your machine. It reaches you https://x.com/Delphi_Digital/status/2045839142450536504
Introducing Hermes Agent v0.11.0 Our largest update yet, with over 700 PRs across ~200 contributors. Thank you to everyone who’s worked on Hermes Agent! This update features a beta TUI v2, unlimited recursion depth and width of subagents, 5 new LLM providers, expanded image gen https://x.com/Teknium/status/2047506967909015907
Meta aims to cut ten percent of workforce by May. Meta is laying off roughly 10,000 employees as part of CEO Mark Zuckerberg’s “Year of Efficiency” drive to reduce costs and streamline operations. This matters because it signals how AI investment priorities are reshaping big tech’s organizational structure—companies are consolidating teams and cutting roles deemed less critical to their AI and infrastructure ambitions. The timing and scale suggest cost discipline is becoming as central to AI strategy as development spending itself.
Palantir’s manifesto resonates by repackaging AI governance concerns as bold contrarianism. CEO Alex Karp positioned Palantir as a principled counterweight to Silicon Valley’s supposed recklessness, arguing the company should prioritize American national security and democratic values in AI deployment. The manifesto gained traction not because it introduced novel ideas—responsible AI governance, transparency, and human oversight are standard industry talking points—but because it framed these concepts as defiant stances against imaginary opponents, appealing to audiences hungry for moral clarity in an uncertain AI landscape.
Google launches Deep Research Max, an AI agent for autonomous enterprise analysis. Google released two upgraded autonomous research agents—Deep Research and Deep Research Max—built on its Gemini 3.1 Pro model, designed to automate complex data analysis for finance, life sciences, and market research. The key distinction is Deep Research Max uses extended reasoning to produce comprehensive, heavily cited reports from blended web and proprietary data sources, while the faster version suits real-time user interfaces. New capabilities include native chart generation, secure connection to private data via Model Context Protocol, and researcher-guided planning before execution—moving beyond simple web search into specialized professional workflows.
Introducing one of our biggest updates to the Gemini Deep Research Agent, now available via the Interactions API! Trigger complex, long-horizon research workflows with arbitrary MCP support, get rich visualizations, plan before you execute, and more with these two https://x.com/googleaidevs/status/2046630912054763854
Introducing our biggest upgrades to the Deep Research API yet… including Deep Research Max (our SOTA system), MCP support, Native charts & infographics, planning mode, full tool support (including Google tools), full multi-modal input support, & real-time progress streaming! https://x.com/OfficialLoganK/status/2046628030777631000
The next evolution of our autonomous research agent is here. Today, we’re introducing Deep Research and Deep Research Max via the Gemini API. Powered by Gemini 3.1 Pro, you can now trigger comprehensive research workflows with unprecedented control and transparency, featuring: https://x.com/Google/status/2046627647208259835
We’ve expanded the capabilities of Deep Research and Deep Research Max to give you even more control and transparency over the entire research process: 📋 Collaborative Planning: Review and refine the agent’s research plan before it executes. 🛠️ Extended Tooling: Run Google https://x.com/Google/status/2046627652568850687
Google brings AI email summaries to workplace Gmail accounts. Google is expanding its AI Overviews feature—previously limited to search results and consumer accounts—into workplace Gmail, allowing employees to ask natural language questions and receive instant summaries pulled from multiple emails without opening them individually. The feature will roll out across business, enterprise, and education Workspace plans, making AI-powered email filtering a default option for millions of office workers. This marks a significant shift in how workplace communication tools operate, automating the reading and synthesis of email conversations at scale.
Google open-sources DESIGN.md format for AI design consistency. Google Labs released an open specification for DESIGN.md, a file format that lets designers document their design rules and brand guidelines in a standardized way. Rather than AI systems guessing at design intent, DESIGN.md allows AI agents to understand the reasoning behind design choices and generate interfaces that match brand standards while meeting accessibility requirements. The move enables designers to reuse design systems across projects and different tools, addressing a practical workflow problem that affects how AI assists in design work.
DeepSeek releases massive open model matching paid frontier AI systems. DeepSeek’s new V4 Pro model—the largest open-source AI yet with 1.6 trillion parameters—ranks first or second on nearly half of standard performance benchmarks, directly competing with paid systems like OpenAI’s offerings. The MIT-licensed release includes novel efficiency techniques that reduce computational requirements, making cutting-edge AI capabilities available for free to researchers and developers rather than locked behind proprietary paywalls.
Deepseek V4 Pro is the biggest open model ever with 1.6T total 49B active, trained on 33T tokens, 1M context, with 2 new attention mechanisms, Muon, mHC, open source kernels, FP4 QAT, MIT license and with one of the best tech repot of the year https://x.com/eliebakouch/status/2047519300399837677
Gemini’s new embedding model handles five data types in one system. Google released Gemini Embedding 2, which processes text, images, videos, audio, and documents together to find semantic connections—eliminating the need for separate specialized tools. This matters because companies currently juggle multiple systems to search and understand different media types, so consolidation reduces complexity and cost. The model is available now through Google’s API and enterprise platform.
Gemini Embedding 2 is now generally available via the Gemini API and Gemini Enterprise Agent Platform search and understand semantic relationships across text, image, video, audio, and documents without complex, fragmented pipelines https://x.com/GoogleAIStudio/status/2047007402520674679
Adobe launches CX Enterprise, an end-to-end AI agent system for marketing. Adobe unveiled CX Enterprise, a new platform that deploys AI agents to automate customer marketing workflows—from content creation to personalized outreach—while maintaining human oversight and brand consistency. The system integrates with tools from AWS, Google Cloud, Microsoft, OpenAI and others, allowing businesses to build custom AI workflows using reusable “agent skills” rather than starting from scratch. This represents Adobe’s shift from selling isolated AI features to offering a cohesive agent-based operating system designed specifically for how marketing teams work.
Anthropic’s Mythos AI model has suffered unauthorized access breaches. Anthropic discovered that its unreleased Mythos AI model was being accessed by people without authorization, according to Bloomberg reporting. The breach is significant because it highlights security vulnerabilities in AI development pipelines before models reach public release—a stage typically assumed to be controlled. The incident underscores growing risks as AI capabilities advance and more organizations handle increasingly powerful systems during development phases.
I appreciate you sharing this material, but I’m unable to produce the requested summary. The provided text is a fragment discussing Opus 4.7’s performance on some task involving TikZ diagrams, but it lacks essential context: what Opus 4.7 is, what problem it solves, why this matters to business or society, and what the concrete evidence or results are. To write an accurate two-line summary, I’d need: – Clear explanation of what happened (a product release? a capability benchmark?) – Why this is significant or distinctive – Specific, verifiable results or metrics Could you provide the complete article or fuller context?
On the plus side with Opus 4.7, if it does decide to think it produces BY FAR the best Sparks unicorn* ever, even non-thinking is pretty good, if not great. * This is created using TikZ, which is a language built for scientific diagrams & very much not for drawing. The original https://x.com/emollick/status/2044880350237626844
AI economists produce consistent answers where humans wildly disagreed. A classic experiment where 146 human economists analyzed identical data produced dramatically different conclusions. When AI systems like Claude repeated the task, they clustered near the middle with far less variation and no extreme outliers—suggesting AI could standardize analysis in fields where human judgment creates unpredictable scatter. This hints at a practical use case beyond general capabilities: using AI to reduce noise in scalable research workflows. — Claude users skew wealthier than competitors’ audiences. About 80% of Americans using Claude weekly earn over $100,000 annually, nearly double the 37% figure for Meta AI users and notably higher than competitors clustering at 56–64%. The gap signals either different marketing reach, product positioning that appeals more to high-income users, or adoption patterns tied to where people encounter AI tools—a meaningful difference in who’s shaping early AI experience and feedback.
Classic study gave 146 economist teams the same dataset & got wildly different answers New paper reruns it with agentic AI. Claude Code & Codex land near the human median, but with far tighter dispersion & no extremes. Suggests that AI is now useful for doing scalable research. https://x.com/emollick/status/2046362044786458648
80% of US adults who report using Claude in the previous week live in households earning $100,000 or more a year, compared to 37% of Meta AI users. Other major providers cluster in a relatively narrow band, with 56–64% of users in $100,000+ households. https://x.com/EpochAIResearch/status/2047056309535801605
AI coding startup Cursor reaches fifty billion dollar valuation on strong revenue growth. Cursor, an AI-coding assistant, is raising at least $2 billion at a $50 billion valuation—nearly double its value from six months ago—while forecasting over $6 billion in annualized revenue by end of 2026. The company recently achieved profitability on enterprise sales by building its own proprietary model, reducing dependence on rivals like Anthropic’s Claude; a separate study of 500 developer teams found that better AI models don’t replace existing work but instead drive a 44 percent increase in usage and shift developers toward more complex tasks like code review and architecture rather than basic coding.
# The Path Forward for AI Startups A lot of founders are messaging each other after the SpaceXAI <> Cursor “IPO-deferred acquisition”. Common discussion topic: what is the future for independent startups? Must ~everyone ultimately be acquired by a frontier lab or go extinct? https://x.com/russelljkaplan/status/2047077659985981616
SpaceXAI and @cursor_ai are now working closely together to create the world’s best coding and knowledge work AI. The combination of Cursor’s leading product and distribution to expert software engineers with SpaceX’s million H100 equivalent Colossus training supercomputer will https://x.com/SpaceX/status/2046713419978453374
SpaceXAI and @cursor_ai are now working closely together to create the world’s best coding and knowledge work AI. The combination of Cursor’s leading product and distribution to expert software engineers with SpaceX’s million H100 equivalent Colossus training supercomputer will https://x.com/SpaceX/status/2046713419978453374?s=20a
The structure of the deal is pretty interesting here. I think what’s happening is: 1. xAI is having trouble training a SOTA coding model (hence cofounder departures), bunch of idle GPUs 2. Cursor doesn’t have capital to blow on a $5B training run to compete with Codex/Claude 3. https://x.com/0xrwu/status/2046721359263285478
Humanoid robot production at Figure AI accelerates dramatically toward thousands annually. Figure AI’s humanoid robot manufacturing has ramped from 2 units in 2023 to an annualized production rate of 2,589 units by April 2026, suggesting the company is moving from prototype phase to scaled production. This trajectory—growing 1,200-fold in three years—represents a shift from lab demonstrations to commercial manufacturing, though the viability of sustained demand at these volumes remains unproven.
Figure shows a serious humanoid production ramp. Brett’s graph doesn’t have a y-axis. Assuming June 2023 = 1 unit and the graph is linear: 2023: 2 units 2024: 26 2025: 79 2026 (till Apr 21st): 312 Annualized run-rate based on the past 21 days of production: 2,589/year. https://x.com/TheHumanoidHub/status/2047041923387695504
Meta poaches key researchers from Google’s advanced AI division. Meta hired several senior scientists from Google’s DeepMind-adjacent team, signaling intensifying competition for AI talent at the frontier. This matters because top researchers are concentrated at a handful of labs, and their movement affects which companies lead in cutting-edge breakthroughs. The hiring reflects Meta’s strategic push to compete directly with Google and OpenAI in developing next-generation AI systems.
Meta’s internal surveillance tool raises employee privacy questions. Meta has deployed an AI monitoring system to track staff activity, according to Business Insider, prompting concerns from workers about privacy and autonomy in the workplace. The development is notable because it highlights how AI tools designed for external business purposes are increasingly being turned inward to monitor employees—a practice that differs from Meta’s public-facing AI initiatives and raises questions about workplace trust and data protection standards.
Chinese AI model Kimi K2.6 demonstrates autonomous coding prowess on complex tasks. Kimi K2.6, a Chinese AI model, completed multi-hour engineering projects independently—including rewriting a financial matching engine’s 4,000+ lines of code to achieve 185% throughput gains and optimizing machine learning inference across multiple programming languages. The model’s ability to execute extended, multi-step coding tasks with minimal human intervention represents a capability gap that widens China’s position in open-source AI development, where Chinese models now outperform leading U.S. alternatives on certain benchmarks.
Kimi K2.5 widens gap between the US and China in open weights model intelligence. The leading US open weights model remains OpenAI’s gpt-oss-120b, which has now been eclipsed by an ever-growing list of open weights releases from China. https://x.com/ArtificialAnlys/status/2016250140219343163?s=20
Kimi K2.6 autonomously overhauled exchange-core, an 8-year-old open-source financial matching engine. Over a 13-hour execution, the model iterated through 12 optimization strategies, initiating over 1,000 tool calls to precisely modify more than 4,000 lines of code. Acting as https://x.com/Kimi_Moonshot/status/2046531057147933137
Nvidia’s supply dominance, not chip design, is its real competitive moat. Jensen Huang revealed that Nvidia’s true advantage isn’t superior engineering but rather decades-long contracts securing scarce components from suppliers like TSMC and SK Hynix—worth nearly $100 billion in commitments. This creates a barrier competitors can’t easily replicate: even if a rival designs a better accelerator, they can’t obtain the materials to manufacture it at scale, fundamentally shifting how AI infrastructure competition works.
Jensen on the famous story about Larry Ellison and Elon Musk begging him for GPUs over dinner: “”That never happened. We absolutely had dinner, and it was a wonderful dinner. At no time did they beg for GPUs. They just had to place an order.”” Jensen says Nvidia’s allocation https://x.com/dwarkesh_sp/status/2044989230112506351
Nvidia has locked up many years of scarce components – almost a hundred billion dollars in purchase commitments. Is this Nvidia’s big moat? A competitor might design a great accelerator, but they don’t have Jensen’s LTAs with SK Hynix, TSMC, etc. Jensen: “If our next several https://x.com/dwarkesh_sp/status/2044808033411223559
The most viral moment of the @dwarkesh_sp x Jensen Huang conversation was widely misunderstood It wasn’t really just about China, TPUs, or even Nvidia’s moat. It was about a much deeper disagreement over what it means for America to “”win”” in AI https://x.com/TheTuringPost/status/2046366547619270665
OpenAI’s Stargate data center project advances across seven U.S. sites. OpenAI’s $500 billion Stargate initiative is moving from announcement to reality, with visible construction underway at all seven planned locations and one Texas facility already delivering 0.3 gigawatts of AI computing capacity. By 2029, the project is projected to reach over 9 gigawatts—equal to New York City’s peak power demand and sufficient to power as many GPUs as existed globally at the end of 2025. The scale of this infrastructure buildout matters because it signals how the AI industry is racing to secure physical capacity for training and running advanced models, while developers are experimenting with onsite power generation and water-efficient cooling to sidestep grid constraints.
In 2025, OpenAI announced Stargate, a $500 billion data center initiative. We surveyed all 7 US sites and found visible development at each. There’s a long road ahead, but the project appears on track to reach 9+ GW by 2029—comparable to New York City’s peak power demand. 🧵 https://x.com/EpochAIResearch/status/2045258390147088764
We now estimate that only about 0.3 GW of total facility power is operational for Stargate Abilene, not 0.6 GW. We have moved the 0.6 GW milestone to late May and the 1.2 GW milestone from Q3 to Q4 2026, but both are uncertain. More about this change and our methodology in 🧵 https://x.com/EpochAIResearch/status/2047442515608162481
Hyatt partners with OpenAI to automate routine employee tasks. Hyatt is deploying AI tools across its workforce to handle manual work, freeing staff to focus on guest interactions—a shift that shows how hospitality companies are using AI not to replace workers but to redeploy them toward higher-value customer service. The move is notable because it treats AI as a productivity tool for frontline employees rather than a cost-cutting mechanism.
“Hyatt’s innovative approach with OpenAI reflects how Hyatt is elevating its use of technology and enhancing human connections. The company is making artificial intelligence broadly accessible to its employees, enabling teams to spend less time on manual” / X https://x.com/TheRealAdamG/status/2046262564158333211
OpenAI open sources PII detection model under permissive Apache license OpenAI released a 1.5 billion-parameter model called Privacy Filter that automatically detects and masks personally identifiable information like names and addresses in text. The model is freely available under an Apache 2.0 license on HuggingFace, making it accessible to developers building privacy safeguards into applications. This represents a shift toward open-sourcing specialized safety tools rather than keeping them proprietary.
OpenAI just open sourced a new 1.5B (50m active) model on HuggingFace with Apache 2.0 license! It’s not a new LLM, this one is called Privacy Filter, and it’s a PII detection model (checking if text has private information) A few interesting tidbits from the release + links: https://x.com/altryne/status/2046977133013311814
OpenAI is building a desktop application that consolidates its AI tools into one interface, aiming to streamline how users access ChatGPT and related services. This matters because it signals a shift from browser-based access toward a dedicated platform—similar to how Microsoft moved Office online—and suggests OpenAI sees integration and convenience as competitive advantages against rivals like Google and Anthropic. The move also indicates OpenAI believes its current multi-product approach feels fragmented enough to warrant simplification. nan
OpenAI shuts down ambitious research projects, loses three key leaders. OpenAI is consolidating around enterprise AI by cutting experimental initiatives like Sora (which cost $1 million daily) and its science research group, prompting departures of three executives including the architects of video AI and the science platform. The moves signal a strategic pivot away from speculative moonshots toward commercial products, though departing researchers argue that breakthrough innovation requires space outside mainstream roadmaps.
This story has now been updated with more details. Three leaders departed from OpenAI today: – Kevin Weil, VP of OpenAI for Science – Srinivas Narayanan, CTO of B2B Applications – Bill Peebles, Head of Sora https://x.com/zeffmax/status/2045248266384838800?s=46
Today is my last day at OpenAI, as OpenAI for Science is being decentralized into other research teams. It’s been a mind-expanding two years, from Chief Product Officer to joining the research team and starting OpenAI for Science. Accelerating science will be one of the most https://x.com/kevinweil/status/2045230426210648348?s=20
OpenAI launches shared AI agents for team workflows and approvals. OpenAI introduced workspace agents—AI assistants that teams can build once and share to handle recurring business tasks like lead qualification, report generation, and vendor screening. The agents run continuously in the cloud, integrate with tools like Slack and external apps, and operate under organizational controls requiring approval for sensitive actions, addressing a gap where AI has helped individuals but not cross-team workflows that depend on shared context and handoffs.
Workspace agents can work across tools—pulling context from docs, email, chats, code, and systems, and taking approved actions like updating @Linear issues, creating docs, or sending messages. In @SlackHQ, agents can jump into a thread, understand what’s needed, pull the right https://x.com/OpenAI/status/2047008991944069624
Developer releases unsupervised AI agent that triggers massive open-source movement. Peter Steinberger’s OpenClaw represents a genuine shift from chatbots to autonomous agents—software that can act independently on the internet rather than just respond to questions. The project’s rapid open-source adoption suggests the field views this as a watershed moment in AI capability, though Steinberger’s own framing (“the lobster is loose”) hints at concerns about once-deployed autonomous systems being difficult to contain or reverse.
Alibaba’s compact AI model outperforms much larger competitor in coding tasks. Alibaba released Qwen3.6-27B, a smaller open-source AI model that unexpectedly beats its own larger 397-billion-parameter version at writing and debugging code—a significant efficiency gain suggesting that model size alone doesn’t determine coding ability. This matters because it could reduce the computational cost and energy use of deploying AI coding assistants, potentially making advanced AI tools more accessible to smaller organizations.
🚀 Meet Qwen3.6-27B, our latest dense, open-source model, packing flagship-level coding power! Yes, 27B, and Qwen3.6-27B punches way above its weight. 👇 What’s new: 🧠 Outstanding agentic coding — surpasses Qwen3.5-397B-A17B across all major coding benchmarks 💡 Strong https://x.com/Alibaba_Qwen/status/2046939764428009914
Alibaba’s Qwen model solves elite math problems that stump most AI systems. Alibaba’s latest Qwen model successfully solved a notoriously difficult math competition problem (AIME-2026 #15) on its first attempt after extended reasoning, outperforming comparable systems from competitors like DeepSeek. This demonstrates that Alibaba has established itself as a frontier AI lab capable of competing with leading developers on complex problem-solving tasks, marking a shift in the competitive landscape beyond the usual suspects.
We’ve post trained a model on top of Qwen that achieves Pareto optimality on accuracy-cost curves. Unlike our previous post trained models, this model has been trained to be good at search and tool calls simultaneously, allowing us to unify the tool call router and https://x.com/AravSrinivas/status/2047019688920756504
Qwen 3.6-Max-Preview solves AIME-2026 #15 after like 30 minutes of thinking, but on first try. Preview or not, it’s more baked than DeepSeek-Expert. Other tests validate this impression. It doesn’t screw up. Alibaba Qwen is, after all, a frontier lab. https://x.com/teortaxesTex/status/2046166258853269990
Military trades expensive robot dogs for cheaper decoy rovers. The US Air Force ditched $75,000 Boston Dynamics quadrupeds in favor of $3,000 rovers fitted with plastic coyote decoys to prevent bird strikes at airbases, achieving the same deterrent effect at a fraction of the cost. This shift highlights a practical tension in defense procurement: expensive humanoid robotics often lose to simpler, cheaper solutions when the actual job—in this case, visual deterrence—doesn’t require advanced capabilities.
To prevent dangerous airbase bird strikes, the US military last year replaced $75k Boston Dynamics’ quadrupeds with $3k rovers attached with plastic coyote decoys. Making security robots look as human as possible will serve as a similar psychological visual deterrent. https://x.com/TheHumanoidHub/status/2044860410420183489
Tesla’s Optimus robot hand patents reveal wiring breakthrough. Tesla disclosed five patents for its Optimus humanoid robot focusing on the hand and arm, with the most significant innovation addressing a longstanding robotics challenge: efficiently routing cables through the wrist joint by stacking them laterally in the forearm and vertically in the hand. This cable-routing solution matters because it enables more compact, flexible wrist designs that could improve dexterity and durability in Tesla’s manufacturing robots, distinguishing this from general robotics progress by tackling a specific mechanical bottleneck in humanoid design.
Tesla published five new patents yesterday, all for the Optimus robotic hand and arm: 1. Covers cable routing through the wrist joint. The key innovation is that cables are arranged in a lateral stack on the forearm side and transition to a vertical stack on the hand side, https://x.com/TheHumanoidHub/status/2045270641210011767
Tesla accelerates humanoid robot production with Optimus factory launch. Tesla announced preparations to build its first large-scale Optimus humanoid robot factory starting Q2 2026, with an initial production line targeting 1 million units annually at its Fremont facility. The company is pairing this manufacturing push with expanded AI infrastructure, including the Cortex 2 data center at Giga Texas for robot training, and plans to unveil Optimus v3 in mid-2026. While Tesla believes its current AI4 chip enables safer autonomous robotaxis than human drivers, the move signals confidence that humanoid robots represent a significant near-term business opportunity beyond vehicles.
Elon on Q1 call, confirming what he posted last week: “”AI5 chip will (first) go into Optimus and the data center. It’s looking like we’ll be able to achieve unsupervised robotaxi with AI4, which is far better than human safety, so AI5 is certainly not immediately needed in https://x.com/TheHumanoidHub/status/2047075601312633316
Elon on Q1 earnings call: – Optimus v3 is functional, but there are some aesthetic elements to be worked on. It may be unveiled in the middle of 2026. – We’re a little hesitant to show it off because we find our competitors do a frame-by-frame analysis and copy everything we https://x.com/TheHumanoidHub/status/2047070791234375765
Tesla in Q1 2026 Shareholder Update: – Preparations for our first large-scale Optimus factory will begin shortly in Q2. – First-gen line designed for 1M robots/year will replace Model S/X lines in Fremont Factory, second-gen line is being prepared at Giga Texas (long-term https://x.com/TheHumanoidHub/status/2047049489874383097
Leave a Reply