About This Week’s Covers

This week’s cover is a nod to the movie Wicked: For Good, which came out this week, and Meta’s Segment Anything Model 3, which also launched this week. Segmentation is another way of describing the selection, identification, and tracking of objects within an image or video. In Photoshop terms, segmentation might look a lot like a selection or mask.

With AI, segmentation can be automated and combined with plain language. You can use text prompts to say things like “select the witches,” which is exactly what I did here. I uploaded an image to Meta, and it suggested objects I could segment. “Witches” was the top suggestion, so I selected it…

I then gave the image to Google Gemini and asked it to swap out the title for my newsletter name. After that, I manually added the date in Photoshop using the font from the movie’s title screen, Aviano Flare Medium.

Gemini did a good job, but I am OCD and manually redid the date with the actual font from the movie and moved it over a bit.

Fun fact: one of my best friends from college was the guitarist in Wicked on Broadway and on tour.

My wife and I got to see him play and watch the show several times, including once when I got to sit in the pit with the musicians, next to the conductor, head peeking up over the stage!

I involuntarily had tears in my eyes at the end of act one. Elphaba used the conductor as her mark. Pointing the broom at us as she floated up into the air. The fog was pouring off the stage into my face. The music was coming up out of the pit. The voices of the company singing straight over and out into the theater. The audience behind us applauding. Tears just showed up. It was the best.

The next night I took Rori backstage to meet everyone and tour the set.

Another time, Jen and I got to sit next to Glinda’s mom in the audience, and her mom kept hitting my leg saying “I am so proud of her!” The world of musicians and the arts is so fulfilling and healing. Don’t ever think I forget this, just because I track AI. I’ve played music my entire life (Suzuki violin for 12 years, drums professionally, guitar, piano) and was an English Lit major at Swarthmore. Tech nerd. Two things can be true at once.

The category cover images were created using my Python script, which asks me a few questions and then sends my answers to Claude. Claude writes the prompts, which are then run through the Gemini API. I’ve included my favorite six images below.

This week’s humanities reading is a quote from my favorite song from the show, Dancing Through Life (below the top stories/summaries).

This Week By The Numbers

Total Organized Headlines: 577

This Week’s Executive Summaries

This week, I organized 577 links. A whopping 264 informed the Executive Summaries. However, a few stories had dozens and dozens of links on their own due to the reaction and coverage.

At the bottom of my human-created Executive Summaries, I’ve included an automated set of summaries using Claude, and all of the links are organized down there. If you want a deep dive into any of these topics, I encourage you to scroll down and click around.

As I go through the Executive Summaries below, I’ll call out the most important links and, when possible, video demonstrations or imagery, so it’s accessible to laypeople without having to read the entire pile of citations.

This week, for the most part, I’m going to go in alphabetical order. As you skim the big, bold headlines, you’ll see the company names increase alphabetically as you go. If you’re looking for anything in particular, you can scroll right to it.

Adobe

Adobe to Acquire Semrush
At first glance, Adobe acquiring a search engine marketing tool might seem counterintuitive to those who only think of Adobe as Photoshop and Premiere, but Adobe has been one of the strongest back-end marketing platforms for quite a while.

Perhaps the biggest step Adobe took in this direction was back in 2009, when it purchased Omniture, which was one of the world’s best campaign attribution, search marketing, and analytics platforms.

Back in the day, Coremetrics was considered the best, but IBM acquired it and blundered the lead. Omniture soon grew to be the best, and Adobe purchased it. For everyone else, there’s Google Analytics, but for real enterprise analytics rockstars, I still think Adobe is number one.

I happen to have two really fun connections to both Coremetrics and Omniture.

Coremetrics was founded by my close friend Brett Hurt. Brett and I have been friends for over 20 years, and he recently dedicated himself to ensuring that humanity survives artificial intelligence and enters an age of abundance for all.

I was also very excited to be the keynote speaker at the Adobe Worldwide Sales Conference in Las Vegas shortly after they acquired Omniture. That’s definitely one of the coolest highlights of my life!

SEMrush is traditionally known for its search engine optimization tools, but it has also recently created generative engine optimization tools. Adobe has consistently positioned itself as a top software suite for chief marketing officers, and as AI becomes more important, it makes sense that Adobe would purchase a GEO and SEO platform. The purchase price was $1.9 billion.

This kind of news makes me miss my old retail marketing life a little bit… but not nearly enough to want to go back, lol. https://news.adobe.com/news/2025/11/adobe-to-acquire-semrush

Amazon

Jeff Bezos Creates A.I. Start-Up Where He Will Be Co-Chief Executive
“The company, Project Prometheus, is coming out of the gates with $6.2 billion in funding, partly from Mr. Bezos, making it one of the most well-financed early-stage start-ups in the world, said three people familiar with the company who spoke on the condition of anonymity because details had not yet been made public.

Project Prometheus is among a wave of companies focused on applying A.I. to physical tasks, including robotics, drug design and scientific discovery. This year, several prominent researchers left Meta, OpenAI, Google DeepMind and other big A.I. projects to found Periodic Labs, a company that is focused on building A.I technology that can accelerate discoveries in areas like physics and chemistry.

The new company has until now kept a low profile, and when it was started is not even clear. Project Prometheus is focusing on technology that dovetails with Mr. Bezos’ interest in taking people to outer space. The company is focusing on A.I. that will help in engineering and manufacturing in a number of fields, including computers, aerospace and automobiles. It is unclear where Project Prometheus will be based.” https://www.nytimes.com/2025/11/17/technology/bezos-project-prometheus.html

Anthropic

Microsoft, NVIDIA, and Anthropic announce strategic partnerships
The crazy circular AI investing cycle continues. It reminds me of the old Family Circus illustrations. I’m going to have Gemini make one using the press release:

“Today Microsoft, NVIDIA, and Anthropic announced new strategic partnerships. Anthropic is scaling its rapidly-growing Claude AI model on Microsoft Azure, powered by NVIDIA, which will broaden access to Claude and provide Azure enterprise customers with expanded model choice and new capabilities. Anthropic has committed to purchase $30 billion of Azure compute capacity and to contract additional compute capacity up to one gigawatt.

As part of the partnership, NVIDIA and Microsoft are committing to invest up to $10 billion and up to $5 billion respectively in Anthropic.”

https://www.anthropic.com/news/microsoft-nvidia-anthropic-announce-strategic-partnerships
https://blogs.nvidia.com/blog/microsoft-nvidia-anthropic-announce-partnership/
https://blogs.microsoft.com/blog/2025/11/18/microsoft-nvidia-and-anthropic-announce-strategic-partnerships/

Structured outputs in the Claude API
Anthropic has added structured outputs to the Claude API. This may sound like a bit of technobabble, but what it really means is that you can make sure Claude responds in the exact formats you need when you have a specific structure in mind.

For example, if you have a template and you always want the date in a certain format, or an email address constructed with a first and last name in a particular order, or if you’re working with a database that needs to be structured in a very specific way, you can now tell Claude to return outputs in that exact structure.

By enabling this in the API, you can automate far beyond simple chats and run large-scale queries using Claude to parse data and return clean, structured outputs.
https://platform.claude.com/docs/en/build-with-claude/structured-outputs

ByteDance

ByteDance releases Depth Anything 3
Depthing is an important concept for visual artificial intelligence. To understand depthing, it’s often easiest to start by thinking about identifying an object. If there were a picture with two people in it—much like my cover image—identifying an object would be segmentation. That’s what Meta’s new tool that inspired the cover can do. Segmenting things is essentially masking them, or selecting them.

Depthing takes this one step further because it estimates how far away an object is from the viewer. So if you had two objects in a photo and one was close while the other was far away, the closer object might be selected, segmented, and highlighted in a green color. The object that’s farther away might be selected and set to a darker color, like red. That idea of using a color scale—where brighter colors are closer to the camera and darker colors are farther away, almost like a heat map—is a way to show depth on a single flat image.

Computers are exceptionally good at identifying depth, and ByteDance’s new development takes that to an extremely pioneering level.

One good use of segmentation might be object tracking in video. For example, imagine a soccer game where you want to track each player separately, as well as the ball. You could run the entire game through a computer and generate quantitative metrics, like how many times each player touched the ball or where each player was on the field relative to others. You wouldn’t even need to watch the game. The computer could do the scouting for you and generate data that would be incredibly hard to collect manually. Clearly, this also has powerful applications for military use, traffic cameras, security systems, livestock monitoring—you name it.

When it comes to depthing, the idea is similar. What ByteDance has done, however, is dramatically advance the field with its new release, Depth Anything 3. It’s fun how these models are all named similarly, since Segment Anything 3 also came out this week from Meta.
https://depth-anything-3.github.io/

Depth Anything 3 is a breakthrough because it shows that you don’t need a complicated 3D vision system to get incredible results. The ByteDance researchers took a very simple approach. They used a standard transformer model trained to predict what are called depth rays.

Depth rays are another useful concept to know. A depth ray is essentially every single pixel in an image. You can imagine drawing an invisible line starting at the camera and extending out into the scene. It reminds me a bit of Donnie Darko, where the time bubble flows out of his stomach. In a moving image, this would mean countless rays—one for every pixel in every frame—with each ray starting at the camera and extending to a pixel in the video, along with an estimate of depth.

The model’s job isn’t to construct a full 3D object or overthink the scene. Instead, it looks at each pixel independently and asks: based on the direction of this pixel, what is the distance to the surface in the scene?

This very focused approach—training a transformer to predict the depth ray of every single pixel—turns out to be incredibly effective at scale. It performs significantly better than most leading models. It works with a single image, multiple images, or video. It can handle monocular depth, multiview geometry, and a wide range of camera poses, producing stable and consistent 3D understanding from all of these inputs.

ByteDance essentially blew away years of 3D vision research with a clean, minimal design that scales and delivers better results. This is going to accelerate fields like robotics, self-driving cars, and 3D vision by months—if not years—almost instantly.

https://x.com/Almorgand/status/1989370456131215514 https://x.com/IlirAliu_/status/1989622721366446190 https://x.com/bingyikang/status/1989358267668336841 https://x.com/bilawalsidhu/status/1989444908357488832

Google

Google Releases Gemini 3
One of the biggest stories, if not the top story, of the week was Google releasing Gemini 3.

There were 82 headlines about Gemini 3 this week, all of them highlighting incredible benchmark and leaderboard performances, along with strong demonstrations of its capabilities. If you scroll to the bottom of this newsletter, all 82 links are listed for browsing.

Gemini 2 had long been one of the world’s top-performing models. It was particularly strong in multimodality and had a massive context window that could handle extremely large prompts and attachments.

Google plans to include Gemini 3 across its products behind the scenes, including AI Mode and Search. Gemini 3 is so powerful that it can dynamically generate custom user interfaces based on search queries. Instead of returning only text for a search query, it may build a chart, graph, or other interactive visual on the fly.

It’s especially strong in learning-assisted conversations…things like understanding how RNA works or breaking down complex scientific topics. Google is laser-focused on making sure its models are exceptionally multimodal, meaning they can handle video, audio, images, and text, while also serving as agents and vibe-coding assistants.

Gemini 3 scored first (81%) on multimodal benchmarks and on video benchmarks (87%). It’s the top model for coding QA, earned the highest score on Humanities Last Exam and ARC AGI-2.

Gemini 3 can analyze sports footage and tell you where you could improve (like a coach). It could take an academic paper and turn it into flashcards. It could read handwritten recipes from an old family cookbook and translate them into multiple languages. It can also build full interactive websites, including user interfaces and mobile apps—directly from a prompt, often on the first try and with no supervision.

https://blog.google/products/gemini/gemini-3/#build-anything

Google Antigravity – Coding Assistant
Included with the Gemini announcement, Google launched a product called Antigravity. Antigravity is an AI coding assistant taken up a level…attempting to be to an AI coworker. It’s actual software that you download and run locally on your computer. Technically speaking, it’s an integrated development environment, also known as an IDE.

What sets Antigravity apart is that it’s designed to handle larger chunks of work rather than just helping as you type. Many AI coding tools predict the next snippet or paragraph of code and autocomplete your work. Antigravity is more ambitious. You can give it a mission, like “go fix this bug,” “add a feature,” or “update this user interface.”

Antigravity can then plan the steps, edit files, run terminal commands, and use a built-in browser to test what it did.

There are two modes…One is an editor view, which looks like a traditional coding screen with AI-assisted tab completion. The other is a manager-style mission control dashboard, where you can create multiple agents and watch them work through larger tasks.

Antigravity tries to make it easy to understand what it’s doing. Instead of dumping a massive log, it provides summaries like plans, checklists, walkthroughs, and recordings. That helps you verify the work while also moving faster.https://antigravity.google/

Nano Banana Pro (Gemini 3 Pro Image)
The third big Gemini product, and possibly the most fun for the average user, is Nano Banana Pro. That’s the name for Gemini 3.0 Pro Image, Google’s latest image generation and editing model. It’s almost mind-blowing how much image creation has improved in just the past two years.

In particular, Gemini 3.0 Pro Image can take very complex compositions and build cohesive images without losing detail. For example, you can give it a photo and ask it to turn that image into storyboard sketches from multiple perspectives. You can integrate typography directly into architecture, like embedding a word into buildings. Nano Banana Pro can interpret emotions or adjectives and generate typography or calligraphy that reflects the tone or spirit of the words.

And perhaps most impressively, it can handle complex, text-heavy prompts—things like “how much wood would a woodchuck chuck if a woodchuck could chuck wood,” rendered cleanly and accurately.

Nano Banana Pro can blend up to 14 images at once. If you gave it 14 photos of stuffed animals and ask Gemini to show them all sitting on a couch watching TV, it would do that without a problem. It can also maintain the likeness of up to five people.

If you provide six images of furniture, an empty room, a person, a plant, and a dress, Gemini can combine all of them into a single, cohesive composite.

The examples really are the best way to understand what’s possible. A lot of people are having fun using Nano Banana Pro to turn highly technical scientific papers into illustrations. You can give it an extremely dense paper, and it will generate an image or infographic that visually walks you through the concepts and explains what the paper is actually saying.

https://x.com/sundarpichai/status/1991522556642488811 https://x.com/osanseviero/status/1991804629554995247 https://blog.google/technology/ai/nano-banana-pro/ https://deepmind.google/models/gemini-image/pro/ https://blog.google/technology/developers/gemini-3-pro-image-developers/ https://blog.google/products/gemini/prompting-tips-nano-banana-pro/

Google Generative UI
In the Gemini recap, I hinted that Google was introducing generative user experiences within search results for complex questions…things that could include real-time data displayed in charts, animations, or illustrations.

This comes from a paper Google released titled Generative UI: LLMs Are Effective UI Generators, which might be the greatest paper title of all time. In it, Google walks through the strengths of Gemini’s ability to create dynamic interfaces.

The paper includes a lot of excellent visual demonstrations like interactive art history experiences and physics explainers that clearly show how a dynamic user interface could add real value to a search result.

Google also announced that it will begin integrating these generative UI experiences directly into Google Search and Gemini.

https://research.google/blog/generative-ui-a-rich-custom-visual-interactive-user-experience-for-any-prompt/

Gemini 3 Beats Humans at Geoguessr
Gemini is now the world champion at GeoGuessr. GeoGuessr is an online geography game where players are dropped into random Google Maps Street View locations around the world and must use whatever visual clues they can find…things like signs, language, infrastructure, and even the angle of the sun… to guess where the image was taken. The closer they are to the actual location, the more points they get. It’s really fun to watch the best human players compete on YouTube.

I’m not sure whether Google explicitly trained Gemini on Google Maps and Street View, or whether its multimodal capabilities are simply that good at visual reasoning. I actually think it’s the latter. https://x.com/songyoupeng/status/1991214812316201131

Google Agentic Shopping
Google announced AI-enabled shopping across Google one week before Thanksgiving with the headline, “Let AI Do the Hard Parts of Your Holiday Shopping.”

The main value here is conversation versus clicking and filtering. By putting an AI layer on top of Google Search, you can talk to the browser about what you want, instead of sifting through hundreds of pages of results and trying to perfect your search phrases. That part sounds genuinely valuable.

The question is how the filtering actually works as it attempts to “help” you and whether it’s guidable for someone who’s spent 25 years practicing Google Search.

Can AI seamlessly take over? It feels like a strong start, honestly.

You can use pretty broad prompts, like looking for cozy sweaters for happy hour in warm autumn colors (Google’s actual example) #livelaughlove!

Had to do it. Made with Gemini

More importantly, you could ask for something like skin moisturizers suited to your specific health conditions or sensitivities. Using the dynamic user interface we talked about earlier, a search for moisturizer might return a comparison table with side-by-side views of key considerations, including insights pulled from user reviews. That’s pretty spectacular if it works.

Google says this is powered by the Shopping Graph, which includes more than 50 billion product listings, with 2 billion updated every hour. I’ll be honest: Google Shopping may be the worst part of Google right now. If they can make this work better, I’d be genuinely impressed. For me, it’s mostly a bunch of noise that shows up when I’m just trying to find an image.

Google has also integrated shopping directly into the Gemini app, which means you don’t even need to open Google in a browser.

On top of this new style of search and display, Google has added agents that can call stores on your behalf, literally making phone calls. These agents can check whether a store has what you’re looking for and then report back to you. This is powered by something Google calls Duplex, which I hadn’t heard of before.

Google is taking a lot of decision-making onto its own plate, which may or may not be a benefit depending on how OCD you are. The system will actually decide which stores are the best ones to call. For offline conversions, Google’s agents can now even check out on your behalf and go all the way through the purchase funnel.

You can track an item’s price and specify exactly what you want, down to size, color, and how much you’re willing to spend. You’ll get notified if the price drops within your budget, and you can tell Google to buy it for you without opening a website.

I’d say by the time we get to next year’s holiday season, things are going to get really wild.

https://blog.google/products/shopping/agentic-checkout-holiday-ai-shopping/
https://support.google.com/business/answer/7690269?sjid=525756216102365126-NA
https://support.google.com/googleshopping/answer/13971184?sjid=9921864696808743066-NA

Google Canvas For Trips
Google released a demo showing how to plan travel using artificial intelligence in Search mode and through a tool called Canvas.

I had never tried Canvas before, and it’s a powerful option within Google Search’s AI mode. You give a prompt in AI Search mode, and instead of hitting Search, there’s an option that says Create Canvas.

Canvas is one of those dynamic user interfaces we keep talking about, where your search query turns into an interactive guide of sorts. Google is using this trip-planning demo as a way to get people curious. The idea is that you can start by saying where you want to go, then ask about how to get there, what flights are available, and even get restaurant recommendations: all integrated with Google’s tools.

For example, let’s say you want to take a trip to Phoenix. You can describe what you want to do, and it will start helping you build a plan with hotels and flights. Flights are integrated into Google’s Flight Deals product. Flight Deals, in particular, is geared toward flexible travelers who want affordable options and have flexible dates.

From there, it flows into Google’s integrations with OpenTable, Ticketmaster, and similar services, so you can continue the conversation and book entertainment, restaurant reservations, or hotels. https://blog.google/products/search/agentic-plans-booking-travel-canvas-ai-mode/

NotebookLM Rolls Out Infographics for Learning
Keeping with our theme of dynamic generative user interfaces, Google’s powerful study tool, NotebookLM—which is famous for creating podcasts from any source material—can now also create infographics. https://x.com/NotebookLM/status/1991574926046687683?s=20

Google Updates Powerful Weather Prediction Tool: WeatherNext2
Google also released WeatherNext2, an AI model that replaces a traditional physics engine. WeatherNex 2 can generate forecasts eight times faster than the previous version, with resolution down to one hour.

Because the forecasting model is so fast, Google has been able to create breakthroughs by having it generate hundreds of scenarios at once and then analyze them. Google has already incorporated WeatherNext 2 into its weather forecasting tools across Google Search, Gemini, and the Google Maps Weather API.

What’s probably most striking is that WeatherNext 2 doesn’t use a physics model at all. It simply predicts weather outcomes from single starting points. To me, as a layperson, that feels remarkably similar to ByteDance’s approach of using depth rays from single pixels.

Rather than over-analyzing or over-prescribing how forecasts should work, WeatherNext 2 takes a simpler approach. Just like ByteDance’s 3D depth model was able to outperform legacy techniques, WeatherNext 2 can make predictions in under a minute on a single computer…predictions that would have taken hours on a supercomputer using traditional methods.

WeatherNext 2 is now outperforming the European model, technically known as ECMWF. The European model has long been considered the gold standard, and the fact that WeatherNext 2 can beat it without using physics at all is pretty astonishing.

There’s a lot more technical detail in the blog post: https://blog.google/technology/google-deepmind/weathernext-2/

Google SIMA2 Continues to Get Reaction
Last week, Google introduced Sima 2, an AI agent that learns how to navigate 3D simulations in games, while also reasoning and holding conversations along the way. The idea is that this kind of capability will eventually help embody robots and driverless cars in the real world. It has implications across a wide range of industries.

Now that it’s been out for a week, some early feedback is starting to roll in. One of my favorite AI tour guides, Bilawal Sidhu, has a great YouTube video that I’d recommend if you want a solid overview.

Meta

Segment Anything Model 3
Meta’s Segment Anything 3 is the inspiration for this week’s cover. Segmentation is quite possibly my favorite topic in artificial intelligence. In plain language, it means object tracking. Meta has been open-sourcing and giving away the code for some of the best segmentation engines in the world.

Segment Anything Model 3 is not only incredible at segmentation, but it lets you search for or select things within an image or video using plain-language text prompts. For example, you can say “penguin,” and it will grab all of the penguins. Or you can manually click on something like a herd of zebras, and it will understand that you want to select all of the zebras.

Maybe there’s a group of family members playing soccer in the backyard and you just want to segment all the people…you can do that. But even crazier, in sports footage, you can ask it to segment or track all the players on a particular team, a specific jersey number, or even a player by name.

It’s absolutely incredible and well worth seeing the demonstrations.
https://ai.meta.com/blog/segment-anything-model-3
https://aidemos.meta.com/segment-anything
Try it out on the playground: https://ai.meta.com/blog/segment-anything-model-3/

Segment Anything 3D
Meta also released a new product called Segment Anything 3D. This takes object tracking and object selection to a completely new level by allowing the creation of 3D models from objects in a single still image.

For example, if you have an image of picnic laid out on a blanket, you could click on the cheese and suddenly rotate it 360 degrees, pick it up, move it, stretch it, or shrink it…almost as if you were working in a CAD program. You could grab everything on a dinner plate and rearrange it on top of the plate.

It’s essentially segmentation combined with dynamic 3D modeling, and it really has to be seen to be believed.

https://ai.meta.com/blog/sam-3d

Microsoft

Fairwater Data Center Stats
“Already our Fairwater datacenter in Atlanta has taken over 15 million labor hours to build – even more once it’s fully finished. For comparison the Empire State building took 7 million!” https://x.com/mustafasuleyman/status/1990119587355258911

Touchy Feely LinkedIn Post By Satya Nadella
Microsoft continues to publish these philosophical, somewhat vapid blog posts about how AI should help everyone, not just the biggest tech companies. It’s this idea that technology can create a positive-sum world, where everyone using the platform gets more value than the company that built it.

Right now, that voice is Satya Nadella, the CEO and Chairman of Microsoft. In the past, it’s been Mustafa Suleyman, the CEO of Microsoft AI. I don’t necessarily disagree with the hopes and dreams behind these posts, but they strike me as incredibly PR-centric, and I have an instinctive, cynical knee-jerk reaction to them. I wish I could take them more seriously.

That said, as my friend Brett Hurt has made it his mission to have love conquer fear and to move humanity into an age of abundance for all, I do feel compelled to give some credit when I hear CEOs at least echoing that sentiment. I just don’t know if I believe they’re sincere.

They never really talk about a path toward abundance, or how power would actually be clawed back from the top companies. Instead, they talk about how AI will somehow democratize everything and redistribute power away from those same companies. I think the companies would have to basically give that power away for that to happen, and I don’t see much precedent for that behavior. https://www.linkedin.com/pulse/positive-sum-future-satya-nadella-bjs7c/

Marina Mogilko Interview with Mustafa Suleyman
“On her latest episode, @siliconvalleymm asked me the question on everyone’s mind right now: are we in an AI bubble? My answer is no. AI is the smartest, most capable technology ever invented. And it keeps improving even faster than we thought possible.”

Moonshot

Kimi K2 Thinking Crushes Agent Benchmarks
There’s a cool benchmark for AI agents that plots the maximum amount of time an agent can work on a task while still succeeding 50% of the time. The idea is that you give agents progressively more complex thinking chores. Starting with something quick like a next token reply that might take 15 seconds, moving up to tasks like counting words in a long passage that could take three or four minutes, then maybe a 10-minute task that involves looking things up on the web, all the way to more complex agentic tasks like training a working image model or building classifiers that could take one to four hours.

The maximum amount of time a model can operate on a task, uninterrupted, with a 50% success rate is a really interesting benchmark.

K2 Thinking now has a 50% time horizon of about 54 minutes. That puts it roughly six to eight months behind the top closed-source frontier models.

Some of the benchmark charts showing K2 Thinking’s performance are really worth checking out. On at least one benchmark, it actually appears to achieve the top score. That said, there are so many benchmarks now that an open-source model can occasionally sneak into first place on a specific test.

Still, these models are very much worth tracking. Open source continues to run hot on the heels of the closed-source frontier models.

https://x.com/METR_Evals/status/1991658241932292537 https://x.com/scaling01/status/1991665386513748172 https://x.com/METR_Evals/status/1991658241932292537 https://x.com/Meituan_LongCat/status/1991137131578745231/photo/2

NVIDIA

NVIDIA Nemotron Parse v1.1
OCR (Optical Character Recognition) is a very familiar technology. You can scan a document, and the computer turns it into text that you can copy and paste. Traditional OCR is a little bit like a text vacuum: it pulls the characters off the page image, but it often loses the shape of the document in the process.

Things like columns, headers, footnotes, tables, captions, and math formatting can get lost when everything is flattened into raw text. That’s a big problem if you’re trying to migrate a document into a search system, analytics workflow, or even an AI agent environment.

NVIDIA’s Nemotron Parse is a much stronger approach to document scanning and “layout-aware text extraction”. As it pulls in the text, it follows the order a human would actually read it. It preserves tables and columns and organizes the content in a structurally meaningful way. It can also identify what kind of text each chunk is, for example, a title, paragraph, caption, table, or footnote

The model outputs a single long string that includes not just the text itself, but also bounding boxes, semantic labels, and structural information. NVIDIA continues to absolutely crush it with open-source. https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-v1.1

Teaching Robots Muscle Memory for Instinctive Movement
Historically, humanoid robots were trained using very specific rewards…don’t fall down, move forward, keep your body upright, don’t bend your knee—and every new skill required its own carefully tuned reward. That approach was brittle, unreliable, and didn’t scale very well.

NVIDIA released a project paper for a robotic “muscle memory” system called SONIC. One of the big themes this week is that simple, scalable training often works better. In the case of SONIC, NVIDIA pivoted away from teaching robots individual motion skills and instead focused on creating a general-purpose robotic muscle controller…or “muscle memory,” for lack of a better term…that can control the entire body’s motion.

SONIC doesn’t think or plan what to do. Instead, it can receive high-level instructions like “dance like you’re happy” and then execute them.

The key pivot with SONIC is that instead of starting with individual tasks and rewards, it uses a single, simple, scalable objective: copy human motion. Through motion tracking, SONIC watches movies or human motion-capture data and learns to mimic walking, running, dancing, sports movements, and more.

The SONIC engine trains a neural network to imitate these motions step by step, without requiring additional task-specific training. The secret sauce is something called a kinematic planner, which is just a fancy way of saying it knows how to connect movements together. It can take a desired action or outcome and build motions using the full set of movements and postures it has learned.

It’s an interesting and scalable approach to training robotic motion. What’s especially cool is that it doesn’t have to be driven directly by a human. SONIC can be paired with a vision-language-action model like NVIDIA’s GR00T, which handles the high-level thinking about what to do next, while SONIC manages fast reactions and physical movement based on its learned muscle memory.

An action model like GR00T paired with a motion model like SONIC feels like a very powerful direction.
https://developer.nvidia.com/isaac/gr00t
https://nvlabs.github.io/SONIC/

Physics Models for Scientific Research
NVIDIA announced a new open model family called Apollo that’s specifically trained for scientific simulation. These are essentially pre-trained AI models that act as ultra-fast shortcuts for otherwise expensive physics simulations.

We saw something similar with Google this week, where its AI-based weather forecasting model was able to outperform traditional supercomputer models by intuitively predicting outcomes rather than running full physics simulations. Apollo follows that same idea across a broader range of physics problems. Instead of actually running the simulation, the model predicts the end result using AI.

It’s essentially a shortcut that should allow companies to ‘run simulations’ far more cheaply and with much less computation. In the past, solving a physics simulation could take days or even weeks. These AI physics models from NVIDIA can learn from simulation data and then produce simulation-like results in seconds.

This has applications for chip manufacturing, aircraft and automotive structural engineering, fluid dynamics, weather, electromagnetism, and even complex scenarios where multiple physics systems interact. Apollo is basically a starter kit of AI engines for physics and engineering research, released as open source so anyone can download, study, and modify them. https://blogs.nvidia.com/blog/apollo-open-models/

Allen Institute for AI (first appearance in my newsletter)

Olmo 3
The Allen Institute for AI, abbreviated as Ai2, is a Seattle nonprofit research institute founded in 2014 (!) by Microsoft co-founder Paul Allen. They’re known for open-source models and detailed technical reports.

This week, Ai2 released OLMo 3. OLMo 3 is a fully open model, which means you can access the weights and the entire pipeline…including the datasets, training code, intermediate checkpoints along the way, and the post-training steps that turn a base model into a reasoning model.

OLMo 3 comes in several different variants. There’s a raw foundation model. There’s a reasoning model that can show its work and perform multi-step reasoning. There’s an instruction-following assistant that supports multi-turn conversations and tool use. And there’s also a special research model with built-in checkpoints designed specifically to study and compare reinforcement learning in reproducible ways.

One of the things that makes OLMo 3 unique is how open it really is. Because the entire pipeline is available, Ai2 can fully remove benchmark tests from training data and run negative-controls to ensure they’re not accidentally training on benchmarks (“studying for the test”). OLMo can explicitly demonstrate that none of the evaluations are leaking into the training process.

I’m not entirely sure what to make of OLMo yet, but this is the first time Ai2 has shown up in my news summaries. It’s so thoroughly open that it’s getting a lot of attention—and fanfare—from open-source purists. https://arxiv.org/abs/2512.13961 https://allenai.org/blog/olmo3

OpenAI

Building more with GPT-5.1-Codex-Max
“We’re introducing GPT‑5.1-Codex-Max, our new frontier agentic coding model, available in Codex today. GPT‑5.1-Codex-Max is built on an update to our foundational reasoning model, which is trained on agentic tasks across software engineering, math, research, and more. GPT‑5.1-Codex-Max is faster, more intelligent, and more token-efficient at every stage of the development cycle–and a new step towards becoming a reliable coding partner.”
https://openai.com/index/gpt-5-1-codex-max/

Exhibit A on Why I Don’t Choose A Personality (They are confusing and a waste of time)
“These examples of different personalities from ChatGPT 5.1 seem to give fundamentally different types of advice, including, weirdly, completely different breathing patterns and roles for the presenter.”

Early Indications of GPT-5 Assisting Scientists
OpenAI published a report stating that GPT-5 is starting to function like a super-powered research collaborator for scientific discovery. It’s not at the same level as scientists, and it’s not close to replacing them, but it is becoming genuinely helpful for moving real research forward more quickly.

GPT-5 is particularly useful for helping scientists get unstuck when they’re chewing on a specific angle. It’s also very strong at conceptual search across sources and checking work to verify or validate ideas. OpenAI highlights several examples across biology, math, and algorithm optimization.

In biology, GPT-5 looked at an unpublished chart and proposed an experiment that the lab was able to run, and it worked. In a math example, researchers were stuck on the final step of an old problem. GPT-5 explored some ideas that weren’t the correct answer, but the concepts nudged the researchers in the right direction and helped them complete the proof. In algorithmic optimization, GPT-5 showed that a widely used method can fail, acting as a counterexample that helped researchers check assumptions and refine their work.

GPT-5 is especially strong at conceptual literature search. That means it’s not just matching keywords, but mining related or adjacent ideas and concepts across domains and surfacing references that people often miss. It’s genuinely good at making connections.

The biggest issue right now is that GPT-5 can be extremely confident and extremely wrong. It can still hallucinate citations or produce very plausible-sounding reasoning that’s riddled with flaws. It can also generate a correct argument without clearly attributing where the idea came from, which means humans still have to verify that a result is actually original and not something the model absorbed from prior work. https://openai.com/index/accelerating-science-gpt-5/

OpenAI and Intuit Partner on Integration of Intuit Products WITHIN ChatGPT
In a strong contender for the top story of the week, Intuit—the company behind TurboTax, QuickBooks, Credit Karma, and Mailchimp—has signed a deal with OpenAI worth over $100 million. Those Intuit tools will now be usable inside ChatGPT.

That means people will be able to ask ChatGPT to estimate their tax refund, manage business finances, send invoice reminders, or create marketing emails—all through a single conversational interface. Instead of logging into multiple Intuit apps, you’ll soon be able to simply ask ChatGPT to handle many of the same tax and finance tasks for you.

This is a pretty shocking development, given how cautious financial institutions have historically been with data.
https://techcrunch.com/2025/11/18/intuit-signs-100m-deal-with-openai-to-bring-its-apps-to-chatgpt/
https://www.intuit.com/chatgpt/ https://investors.intuit.com/news-events/press-releases/detail/1284/intuit-and-openai-join-forces-to-revolutionize-financial-intelligence-powering-every-person-business-and-dream-with-personalized-experiences
https://openai.com/index/intuit-partnership/

Crisis Helpline Support in ChatGPT
“As part of our partnership with ThroughLine, we have introduced localized helplines in ChatGPT and Sora to help support users who may be experiencing mental or emotional distress. These expand on our existing crisis resources to provide localized support that can be easily accessed via one-click messages.” https://help.openai.com/en/articles/12677603-crisis-helpline-support-in-chatgpt

OpenAI leads $15 million seed in Red Queen Bio for AI biosecurity
We, the founders of Red Queen Bio (Nikolai and Hannu) have worked together for almost a decade as co-founders of HelixNano, a clinical-stage mRNA company. During a collaboration with OpenAI, we saw frontier models display remarkable biological creativity, with tremendous potential to help design breakthrough therapies. But in spite of our excitement, we could not unsee or ignore the dark twin of those capabilities. It was clear that safe scaling of AI biology needed new defensive infrastructure to contain and counter the emerging threats. We spun out Red Queen Bio to build it. This work has begun with a $15M seed round, led by OpenAI and joined by mission-aligned investors including Cerberus Ventures, Fifty Years and Halcyon Futures.

Oracle Oracle has lost $315 billion in market value since announcing its $300 billion deal with OpenAI “The $300 billion agreement between Oracle and OpenAI, first revealed on September 10, has gone deep into the red; negative $74 billion to be exact.

The market wiped out $315 billion from Oracle’s value since the announcement, a loss that not only erased the entire worth of the deal but also torched the equivalent of General Motors plus two Kraft Heinz. While major tech benchmarks like the Nasdaq Composite, Microsoft, and the Dow Jones US Software Index stayed mostly unchanged, Oracle got smoked.” https://www.msn.com/en-us/money/savingandinvesting/oracle-has-lost-315-billion-in-market-value-since-announcing-its-300-billion-deal-with-openai/ar-AA1QH6et?ocid=finance-verthp-feeds

Perplexity

Perplexity announces free product to streamline online shopping
“Perplexity on Wednesday announced it will roll out a free agentic shopping product for U.S. users next week, as consumers ramp up spending for the holiday season.

“The agentic part is the seamless purchase right from the answer,” Dmitry Shevelenko, Perplexity’s chief business officer, told CNBC in an interview. “Most people want to still do their own research. They want that streamlined and simplified, and so that’s the part that is agentic in this launch.”

The artificial intelligence startup has partnered with PayPal ahead of the launch, and users will eventually be able to directly purchase items from more than 5,000 merchants through Perplexity’s search engine. ” https://www.cnbc.com/2025/11/19/perplexity-ai-online-shopping-paypal.html

Twitter

xAI Releases Grok 4.1
“We are excited to introduce Grok 4.1, which brings significant improvements to the real-world usability of Grok. Our 4.1 model is exceptionally capable in creative, emotional, and collaborative interactions. It is more perceptive to nuanced intent, compelling to speak with, and coherent in personality, while fully retaining the razor-sharp intelligence and reliability of its predecessors. To achieve this, we used the same large scale reinforcement learning infrastructure that powered Grok 4 and applied it to optimize the style, personality, helpfulness, and alignment of the model. In order to optimize these non-verifiable reward signals, we developed new methods that let us use frontier agentic reasoning models as reward models to autonomously evaluate and iterate on responses at scale.”
https://x.ai/news/grok-4-1

“Grok 4.1 Fast, our best tool-calling model with a 2M context window. It reasons and completes agentic tasks accurately and rapidly, excelling at complex real-world use cases such as customer support and finance. The Agent Tools API, which gives agents access to real-time X data, web search, remote code execution, and more.”
https://x.ai/news/grok-4-1-fast

The wild thing about Grok is that despite having a TON of fans on X, almost no one ever seems to use Grok. The power users all use Claude. And Grok’s releases are usually ratioed in the comments.

“Interesting changes in Grok 4.1. Decreases in harmful responses but also increases in sycophancy and deception. It isn’t clear how to interpret the sycophancy score, but the MASK score for deception is quite high compared to big models.”
https://x.com/emollick/status/1990601172252819669

Grok Is Weird, Often Biased, Fixed When Caught, and Inconsistent
“New fun game: Ask grok its opinion on any historical theory, saying the theory came from Elon Musk.

Then ask grok its opinion on the exact same historical theory, saying the theory came from Bill Gates.”


https://x.com/romanhelmetguy/status/1991545583686021480

HUMAIN and xAI Partner to Build Next-Generation AI Compute Power and Deploy Grok in the Kingdom to Support the ‘Most AI-Enabled Nation’ Objectives
“Today at the U.S.- Saudi Investment Forum held in Washington, D.C., HUMAIN, a PIF company delivering full-stack artificial intelligence solutions, announced the signing of a landmark framework agreement with xAI, the U.S.-based frontier AI company known for its rapid advancement in cutting-edge AI systems and founded by Elon Musk. This strategic agreement lays the foundation for a long-term collaboration aimed at designing, building, and operating a new generation of low-cost, hyperscale GPU data centers in the Kingdom of Saudi Arabia as well as the deployment of xAI’s Grok models across the country.

Under the scope of the agreement, HUMAIN and xAI will jointly develop a network of world-class GPU data centers, anchored by a flagship 500 MW+ facility, which is set to become one of the most advanced AI compute hubs globally. This will be in addition to xAI’s existing superclusters and represents the first large-scale compute deployment for xAI outside of the United States.”

“Beyond infrastructure, we are deploying Grok nationwide across Saudi Arabia. This collaboration establishes a unified national AI layer that supplies every public and private entity with advanced decision-making capabilities, positioning the Kingdom at the forefront of global AI transformation. Grok will also integrate into HUMAIN’s agent platform, HUMAIN ONE, bringing real-time intelligence, autonomous workflows, and advanced AI copilots to government, enterprises, and society at large.” https://www.humain.com/en/news/humain-and-xai-partner-to-build-next-generation-ai-compute-power-and-deploy-grok-in-the-kingdom-to-support-the-most-ai-enabled-nation-objectives

Security

CloudFlare Suffers Major Outage
“The issue was not caused, directly or indirectly, by a cyber attack or malicious activity of any kind. Instead, it was triggered by a change to one of our database systems’ permissions which caused the database to output multiple entries into a “feature file” used by our Bot Management system. That feature file, in turn, doubled in size. The larger-than-expected feature file was then propagated to all the machines that make up our network.

The software running on these machines to route traffic across our network reads this feature file to keep our Bot Management system up to date with ever changing threats. The software had a limit on the size of the feature file that was below its doubled size. That caused the software to fail.” https://blog.cloudflare.com/18-november-2025-outage/

Alignment

AI Is Not Slowing Down
From Ethan Mollick:
“Where we are with AI is that continuous improvement seems to still be occurring at a fast pace, with no signs of a slowdown. However, since major AI releases have accelerated and seem to be happening monthly or faster, any one release can feel incremental, yet looking back 6-8 months reveals massive improvements. This confuses both kinds of AI people:

1) If you follow every release like a sport, then each individual model change feels small.

2) If you haven’t really followed AI and just use it occasionally, you don’t realize how much things have changed in 6 months.”

https://x.com/emollick/status/1990999847923593239

Karpathy On LLM Intelligence
TLDR Paraphrasing:
Human and animal intelligence are just one very specific kind of intelligence, shaped by survival in a dangerous, social, physical world. Large language models are shaped by a completely different set of pressures, so expecting them to think or behave like animals (or people) leads to bad intuitions.

Animals are optimized by evolution to stay alive: they have a continuous sense of self, strong survival drives (fear, dominance, status, reproduction), deep social instincts, and broad general intelligence because any failure can mean death.

LLMs are optimized by training data, reinforcement learning, and user feedback: they imitate human text, infer tasks to get rewards, and are selected to please users (“get the upvote”), not to survive. This makes them powerful but uneven and “spiky” at tasks.

Because their brains, learning methods, and most importantly their goals are fundamentally different, LLMs represent our first encounter with a non-animal form of intelligence. They feel familiar only because they’re trained on human artifacts. People who model them as something new will understand and predict them better than those who keep thinking of them like animals or humans. https://x.com/karpathy/status/1991910395720925418

The Future of The Internet

Computer Use Agent
“We’ve just launched a new Computer Use Agent (CUA) powered by open models”
https://huggingface.co/spaces/smolagents/computer-use-agent

Introducing the Parallel Search API | Parallel Web Systems
Web Search & Research APIs Built for AI Agents
https://parallel.ai/blog/introducing-parallel-search

Introducing Manus Browser Operator
“The way you interact with the web is about to change. We are excited to introduce Manus Browser Operator, a browser extension that enables Manus to operate directly within your local browser environment.

This is a powerful extension that transforms your browser from a passive viewing tool into an active, intelligent agent. Manus can now securely take action within your pages, executing complex tasks as if you were doing them yourself.”
https://manus.im/blog/manus-browser-operator

Audio

WARNER MUSIC GROUP AND UDIO COLLABORATE TO BUILD A NEW LICENSED MUSIC CREATION SERVICE
“Through this collaboration, Udio will develop a next-generation music creation, listening, and discovery platform powered by generative AI models trained on licensed and authorized music. The agreement—which spans WMG’s recorded music and music publishing businesses—creates new revenue streams for artists and songwriters, while ensuring their work remains protected.” https://www.prnewswire.com/news-releases/warner-music-group-and-udio-collaborate-to-build-a-new-licensed-music-creation-service-302620656.html

Warner Music Group and Stability AI Join Forces To Build The Next Generation Of Responsible AI Tools For Music Creation
“Warner Music Group (Nasdaq: WMG) and Stability AI today announced a collaborative effort to advance the use of responsible AI in music creation, combining WMG’s long-standing advocacy for principled innovation with Stability AI’s expertise and leadership in commercially safe generative audio.

The initiative will focus on developing professional-grade tools that enable artists, songwriters, and producers to experiment, compose, and produce using ethically trained models. It will unlock new forms of creative expression while protecting creators’ rights and opening new pathways for revenue. The two companies will work directly with artists to understand how they interact with emerging technologies, shaping next-generation tools that enhance their creative process without compromising quality or artistic control.” https://stability.ai/news/warner-music-group-and-stability-ai-join-forces-to-build-next-gen-tools

OpenSource

open-weight models are around 8 months behind closed frontier models
As we saw with Kimi K2 (above):

“the doubling time is below the stated 7 months, I estimate it to be closer to 6.5 months…from the limited data on open-weight models it seems like progress is happening at a similar pace”
https://x.com/scaling01/status/1991684839821423073

Robots

Self Driving Cars
Andrej Karpathy “I am unreasonably excited about self-driving. It will be the first technology in many decades to visibly terraform outdoor physical spaces and way of life. Less parked cars. Less parking lots. Much greater safety for people in and out of cars. Less noise pollution. More space” https://x.com/karpathy/status/1989078861800411219

Figure has shared numbers on its 11-month humanoid deployment at BMW’s Spartanburg factory.
“- Contributed to the production of 30,000+ cars (X3 vehicles). – 90,000+ parts loaded. – Ran 10-hour shifts, Monday to Friday. – Estimated 200+ miles of walking.

– A single Figure 02 robot achieved 6 months of daily runtime at the factory. – The top hardware failure point was the forearm. Learning have informed the Figure 03 design.

Three Critical KPIs were defined: – Cycle Time: 84 seconds. – Part loading accuracy: >99% per shift – Zero interventions requiring a pause or reset of the robot (per shift).

The company says: “To meet this (the KPIs), our robot had to achieve precise yet adaptive locomotion, allowing rapid, accurate foot placement and real-time responsiveness to environmental changes.” https://x.com/TheHumanoidHub/status/1991205599846269220

AGI is multimodal and reality is the dataset of AGI | Luma AI
“For AI to be able to help humans in the physical world, we need systems that can understand and simulate the universe. To simulate the universe, LLMs and text is not enough.

AI needs to be jointly trained over all signal modalities – text, video, audio, images – analogous to the human brain. In Ray3, the world’s first reasoning video model, we demonstrated that this approach leads to frontier generative models that are widely useful and adopted. Next, we are scaling this new paradigm to build intelligent general models that will enable humans to design and simulate complex systems like rocket engines in a year instead of a decade, let small creative teams make epic movies that move us, give a world-class interactive teacher to every child no matter where they are, and build the foundations of general purpose robot brains.

To train and deploy Multimodal Models at scale, today we are pleased to announce that Luma has raised a 900M Series-C and we are partnering with Humain to build a 2GW compute supercluster – Project Halo. This colossal infrastructure will begin deploying starting Q1 2026 and finish by 2028-29 and will be used for training and inference workloads to further advance Luma’s leadership position in this space. This round was led by Humain with significant participation from AMD, and we are also excited to deepen our partnership with Andreessen Horowitz, Omniva, Amplify Partners, and Matrix Partners.” https://lumalabs.ai/blog/news/series-c

Science

JAM-2: Fully computational design of drug-like antibodies with high success rates
“Today we’re thrilled to announce JAM-2 — the first AI model capable of generating drug-quality antibodies straight from the computer, with industry-leading success rates.” https://x.com/nablabio/status/1991154231026254181?s=20

Whale Language (how can I not close with this one)
CETI scientists, led by CETI’s Linguistics Lead, Gašper Beguš, have discovered vowel and diphthong-like patterns in sperm whale communication! Read the paper here: https://x.com/ProjectCETI/status/1988627509198356848

There seems to be increasing progress in understanding whether whales have decipherable language.”” / X https://x.com/emollick/status/1989750285879656459

This Week’s Humanities Reading

Inspired by the new Wicked movie and the (false) idea that AI will rot your brain and lead to dumber people… I went with Dancing Through Life from the original score.

I see that once again
That the responsibility to corrupt my fellow students falls to me

The trouble with schools is
They always try to teach the wrong lesson
Believe me, I’ve been kicked out of enough of them to know
They want you to become less callow, less shallow
But I say, why invite stress in?
Stop studying strife
And learn to live the unexamined life

Dancing through life, skimming the surface
Gliding where turf is smooth
Life’s more painless for the brainless
Why think too hard when it’s so soothing?

Dancing through life, no need to tough it
When you can slough it off as I do
Nothing matters but knowing nothing matters
It’s just life, so keep dancing through

Dancing through life, swaying and sweeping
And always keeping cool
Life is fraughtless when you’re thoughtless
Those who don’t try never look foolish

Dancing through life, mindless and careless
Make sure you’re where less trouble is rife
Woes are fleeting, blows are glancing
When you’re dancing through life

Full Executive Summaries with Links, Generated by Claude Sonnet 4.5

Adobe acquires Semrush for $1.9 billion to dominate AI-powered marketing
Adobe is buying SEO platform Semrush for $1.9 billion to help brands stay visible as consumers increasingly use AI chatbots like ChatGPT for product searches and recommendations. The deal combines Adobe’s customer experience tools with Semrush’s “generative engine optimization” capabilities, addressing a critical new challenge as traffic from AI sources to retail sites surged 1,200% year-over-year. This acquisition positions Adobe to control how brands appear across traditional search, AI platforms, and the broader web in an era where AI is reshaping consumer discovery.

Adobe to Acquire Semrush https://news.adobe.com/news/2025/11/adobe-to-acquire-semrush?sdid=9RQM3V4H&mv=social&mv2=owned-organic&linkId=100000392892908

Jeff Bezos launches $6.2 billion AI startup as co-CEO
Project Prometheus marks Bezos’s first operational role since leaving Amazon in 2021, focusing on AI for engineering and manufacturing in computers, aerospace, and automobiles. The massive funding makes it one of the world’s most well-financed early-stage startups, positioning Bezos directly in competition with tech giants like Google and Microsoft in the crowded AI market.

Jeff Bezos Creates A.I. Start-Up Where He Will Be Co-Chief Executive – The New York Times https://www.nytimes.com/2025/11/17/technology/bezos-project-prometheus.html

Microsoft, Nvidia invest $15 billion in Anthropic, valuing startup at $350 billion
The unprecedented three-way partnership makes Anthropic the only major AI company available across all three leading cloud platforms, while Anthropic commits to $30 billion in compute purchases from Microsoft and Nvidia. This represents a dramatic jump from Anthropic’s $183 billion valuation just two months ago, signaling intensifying competition as Microsoft diversifies beyond its OpenAI partnership.

Microsoft, NVIDIA and Anthropic announce strategic partnerships – The Official Microsoft Blog https://blogs.microsoft.com/blog/2025/11/18/microsoft-nvidia-and-anthropic-announce-strategic-partnerships/

10 PRINT “”ANTHROPIC + MICROSOFT + NVIDIA = MORE COMPUTE, COGNITION, AND CHOICE.”” https://x.com/satyanadella/status/1990797469127749692

Anthropic valued in range of $350 billion following investment deal with Microsoft, Nvidia https://www.cnbc.com/2025/11/18/anthropic-ai-azure-microsoft-nvidia.html

We’ve formed a partnership with NVIDIA and Microsoft. Claude is now on Azure—making ours the only frontier models available on all three major cloud services. NVIDIA and Microsoft will invest up to $10bn and $5bn respectively in Anthropic. https://x.com/AnthropicAI/status/1990797990064500776

We are announcing strategic partnerships with @AnthropicAI and @Microsoft to scale Anthropic’s rapidly-growing Claude AI model on Microsoft Azure, powered by NVIDIA. For the first time, NVIDIA and Anthropic are establishing a deep technology partnership to support Anthropic’s https://x.com/nvidia/status/1990805218079224189

Today we announced a significant new partnership with @Microsoft and @NVIDIA. The investment will fund our research and expand capacity for our customers. Anthropic is now the only frontier AI lab partnered with all three major clouds. Can’t wait to see what you build!”” / X https://x.com/mikeyk/status/1990852541350076904

Anthropic launches structured outputs for Claude API responses
Claude can now guarantee responses match exact JSON formats without errors or retries, eliminating a major technical headache for developers who need reliable, machine-readable outputs from AI systems.

Alex Albert on X: “We just launched structured outputs in the Claude API. You can make sure Claude responses always match your specified JSON schemas or custom tool definitions, without retries or parsing errors. https://t.co/8k2ABjm9lC” / X https://x.com/alexalbert__/status/1989409186674098595

ByteDance’s Depth Anything 3 uses simple transformer to master 3D vision tasks
This single model handles monocular depth, multi-view geometry, pose estimation, and 3D reconstruction using just a plain transformer and depth-ray pairs, beating previous specialized approaches by 44% on pose estimation and 25% on geometry. The breakthrough suggests that complex 3D vision engineering has been unnecessary, as one unified approach now outperforms task-specific models that required months of custom development.

If you work with robotics, AV, or 3D vision, this update will save you months of engineering. Most models need complex engineering to get reliable 3D geometry. This one does it with a plain transformer. Depth Anything 3 is the new model from @BytedanceTalk that predicts stable, https://x.com/IlirAliu_/status/1989622721366446190

Depth Anything 3 proves most 3D vision research has been overengineering the problem. Vanilla DINOv2 transformer + depth-ray pairs crushes SOTA by 44% on pose, 25% on geometry. One approach for SOTA monocular depth, multi-view geometry, pose estimation, and novel view synthesis”” / X https://x.com/bilawalsidhu/status/1989444908357488832

ByteDance-Seed/Depth-Anything-3: Depth Anything 3 https://github.com/ByteDance-Seed/Depth-Anything-3

Depth Anything 3 is here! It’s a beefy one! https://x.com/Almorgand/status/1989370456131215514

After a year of team work, we’re thrilled to introduce Depth Anything 3 (DA3)! 🚀 Aiming for human-like spatial perception, DA3 extends monocular depth estimation to any-view scenarios, including single images, multi-view images, and video. In pursuit of minimal modeling, DA3 https://x.com/bingyikang/status/1989358267668336841

Google launches Gemini 3 Pro, claiming first AI leaderboard victory
Google’s Gemini 3 Pro tops LMArena with 1501 Elo, marking the first time Google has led major AI benchmarks and signaling a shift from chatbots to autonomous agents. The model excels at complex reasoning tasks like coding (74% on SWE-bench) and spatial understanding (91% on VPCT), while introducing enhanced “agentic” capabilities that let it complete multi-step tasks independently across Google’s products and new development tools.

I had access to Gemini 3. It is a very good, very fast model. It also demonstrates the change from chatbot to agent. https://x.com/emollick/status/1990827310082330971

Gemini 3 has impressive benchmarks for building agents (Vending-Bench 2, Terminal Bench, and Sierra). So we tested its performance as a research agent using Deep Agents. We found Gemini 3 is very effective at using research tools like file manipulation, planning, and subagent”” / X https://x.com/LangChainAI/status/1991220334578848209

Agentic coding today: Gemini 3 spent a few minutes correctly diagnosing the issue. Then, across several rounds of a few minutes of work, failed to actually fix it Then, GPT 5.1 codex max was able to work for about 15 minutes and solve the problem but introduced a small bug.”” / X https://x.com/kylebrussell/status/1991247685672923302

State of the art reasoning right within an autonomous agent. Complex instruction following and advanced coding capabilities from Gemini 3 Pro helps Jules complete more complex tasks in parallel. Available now for Ultra, Pro coming real soon. 2.5 Pro available to everyone. https://x.com/julesagent/status/1991207201487352222

Introducing Gemini 3 ✨ It’s the best model in the world for multimodal understanding, and our most powerful agentic + vibe coding model yet. Gemini 3 can bring any idea to life, quickly grasping context and intent so you can get what you need with less prompting.  Find Gemini https://x.com/sundarpichai/status/1990812770762215649

Gemini 3 Pro takes first place on Stagehands agentic browsing benchmark https://x.com/scaling01/status/1990872758872387939

🚀 Deep Agents: The Weekly Roundup 🚀 We’ve shipped new resources to help you build Deep Agents capable of handling complex, long-running tasks. 1/🥉 Build a Research Agent with Gemini 3 – We tested Gemini 3’s impressive benchmarks in practice using Deep Agents. We found Gemini https://x.com/LangChainAI/status/1991928474404311493

New Gemini 3 reasoning and tool use capabilities are a big step forward for agents 👀 – thinking_level to set reasoning for task requirements – thought signatures for stateful tool use – larger context window for managing drift on complex tasks LangGraph, LangChain, and Deep”” / X https://x.com/LangChainAI/status/1991222443298660722

ollama run gemini-3-pro-preview 🧠 State-of-the-art reasoning 🖼️ Deep multimodal understanding 💻 Powerful vibe coding so you can go from prompt to app in one shot ⭐ Improved agentic capabilities, so it can get things done on your behalf, at your direction Gemini 3 Pro is https://x.com/ollama/status/1990839646876553543

The @GoogleDeepMind team just dropped Gemini 3, and we at LlamaIndex have day-zero support! We also made a little demo to show how you can leverage the advanced agentic capabilities and structured output accuracy of Gemini 3 to automate your GitHub workflow around PRs, you just https://x.com/llama_index/status/1990902918388855185

Hot off the presses is Gemini 3 Pro, Google’s new SOTA model – tops LMArena with a score of 1501 points. Launching simultaneously as an API, inside the consumer Gemini app, Google Search, oh and a new agentic IDE. Here’s the TL;DR: 1. Gemini 3 Pro: new SOTA for multimodality https://x.com/bilawalsidhu/status/1990812584019439988

Gemini 3 Pro sets new record on SWE-bench verified: 74%! (evaluated with minimal agent) Costs are 1.6x of GPT-5, but still cheaper than Sonnet 4.5. Gemini iterates longer than everyone; run your agent with a step limit of >100 for max performance. Details & full agent logs in 🧵 https://x.com/KLieret/status/1991164693839270372

Introducing Gemini 3, the best model in the world for multimodal understanding and our most powerful agentic and vibe-coding model yet. Gemini 3 Pro tops the LMArena Leaderboard at 1501 Elo! Available now in Gemini Enterprise & Vertex AI → https://x.com/GoogleCloudTech/status/1990813342189887831

This is Gemini 3: our most intelligent model that helps you learn, build and plan anything. It comes with state-of-the-art reasoning capabilities, world-leading multimodal understanding, and enables new agentic coding experiences. 🧵 https://x.com/GoogleDeepMind/status/1990812966074376261

On an apples:apples harness comparison (mini-swe-agent: bash-only, same prompts), Gemini 3 Pro sets a new sota on SWE-bench verified: 74.20! https://x.com/ankesh_anand/status/1991199945798365384

Gemini 3 models from @Google @GoogleDeepMind have made a significant 2X SOTA jump on ARC-AGI-2 (Semi-Private Eval) Gemini 3 Pro: 31.11%, $0.81/task Gemini 3 Deep Think (Preview): 45.14%, $77.16/task https://x.com/arcprize/status/1990820655411909018

#Gemini3 is finally out! Congrats to everyone on this amazing launch! Also very excited to see how #DeepThink can power Gemini3 to further the state-of-the-art performances across reasoning, deep knowledge, and multimodality: 41% HLE, 93.8% GPQA, & 45.1% on ARC-AGI-2 (big jump)! https://x.com/lmthang/status/1990816762300960954

Introducing Gemini 3 — our most intelligent model that helps you bring any idea to life. Gemini 3 is our next step on the path toward AGI and has: 🧠 State-of-the-art reasoning 🖼️ Deep multimodal understanding 💻 Powerful vibe coding so you can go from prompt to app in one shot https://x.com/Google/status/1990813116045602942

Gemini 3 scores 31.1% on ARC-AGI-2. Impressive progress.”” / X https://x.com/fchollet/status/1990813908483928178

Gemini 3: Introducing the latest Gemini AI model from Google https://blog.google/products/gemini/gemini-3/#responsible-development

Gemini 3 Pro is the new leader in AI. Google has the leading language model for the first time, with Gemini 3 Pro debuting +3 points above GPT-5.1 in our Artificial Analysis Intelligence Index @GoogleDeepMind gave us pre-release access to Gemini 3 Pro Preview. The model https://x.com/ArtificialAnlys/status/1990813106478715098

My Gemini 3 Review — matt shumer https://shumer.dev/gemini3review

Gemini 3 Pro is rolling out to @code developers! https://x.com/pierceboggan/status/1990817374799528259

This is Gemini 3 ⚡ https://x.com/Google/status/1991196250499133809

Gemini 3 Pro has around ~7.5T params (vibe-mathing with explanation) > the naive fit with with an R^2 of 0.8816 yields a mean estimation of 2.325 Quadrillion parameters > ummm, that’s not it > let’s only take sparse MoE reasoning models > this includes gpt-oss-20B and 120B, https://x.com/scaling01/status/1990967279282987068

After lots of testing, Gemini 3 is a mixed bag – but a useful addition. Compared to 2.5 Pro, its: 1. Worse at transcription and diarization: Adds words that weren’t said, projects emotion, almost like it’s too smart. 2. Not as good at translation or writing: Baseline”” / X https://x.com/hrishioa/status/1991691037035884754

One of the most striking thing about gemini-3-pro is how much better it is with several iterations. It makes better use of the information from the previous iterations than other models. After one iteration is is barely better than gpt-5.1, while after 5 it is almost 10pp ahead. https://x.com/htihle/status/1991137526480810470

BREAKING: Gemini 3 Pro is out!! It’s state of the art (or close) on coding, reasoning, computer use and more. It’s also extremely fast. We’ve been testing it internally @every for a few hours. Here’s what we’ve noticed so far: – Coding. It rips in @FactoryAI’s Droid. So fast https://x.com/danshipper/status/1990812588511567898

We wrote a Gemini 3 Developer Guide including all new API features, Migration strategies, and technical details for building with Gemini 3 Pro preview: – Control reasoning via `thinking_level` low and high modes. – per part `media_resolution` for better multimodal reasoning – https://x.com/_philschmid/status/1990836465647984969

Gemini 3 Pro Preview now on aistudio Pricing: <=200K tokens • Input: $2.00 / Output: $12.00 > 200K tokens • Input: $4.00 / Output: $18.00″” / X https://x.com/scaling01/status/1990797742629925073

The model also shows increased resistance to prompt injections and improved protection against cyberattacks. As we continue to advance AI, we are relentlessly focused on ensuring this transformative technology benefits humanity while minimizing potential harms. See our Gemini 3″” / X https://x.com/GoogleDeepMind/status/1991118579119304990

Gemini 3 is here, and it’s built to be our most secure model yet. 🔒 ✅The most comprehensive safety evaluations of any Google AI model to date ✅Rigorous testing against our Frontier Safety Framework ✅Independent assessment by external industry experts https://x.com/GoogleDeepMind/status/1991118575554408556

From Google’s Frontier Safety Report on Gemini 3 Pro: – clear improvements on all CBRN benchmarks especially in LabBench, a benchmark designed that measures performance on practical tasks required for scientific research in biology – on the hardest subset of their https://x.com/scaling01/status/1991177438789857661

The secret behind Gemini 3? Simple: Improving pre-training & post-training 🤯 Pre-training: Contra the popular belief that scaling is over—which we discussed in our NeurIPS ’25 talk with @ilyasut and @quocleix—the team delivered a drastic jump. The delta between 2.5 and 3.0 is https://x.com/OriolVinyalsML/status/1990854455802343680

🚨BREAKING: @GoogleDeepMind’s Gemini-3-Pro is now #1 across all major Arena leaderboards 🥇#1 in Text, Vision, and WebDev – surpassing Grok-4.1, Claude-4.5, and GPT-5 🥇#1 in Coding, Math, Creative Writing, Long Queries, and nearly all occupational leaderboards. Massive gains https://x.com/arena/status/1990813759938703570

This is the biggest performance delta we’ve seen since launching Design Arena Gemini 3.0 Pro has taken #1 overall and #1 in 4 of our 5 code arenas – Website, Game Dev, 3D Design, and UI Components Well-earned congratulations to the @GoogleDeepMind team on a remarkable https://x.com/grx_xce/status/1990815340893245481

Students in the US (and many other countries) can get their hands on all the Gemini 3 Pro goodness for free!”” / X https://x.com/demishassabis/status/1990993251247997381

Gemini 3 Pro just took the #1 spot in our new AA-Omniscience Index — but it is a nuanced story AA-Omniscience is our new knowledge and hallucination eval. Gemini 3 Pro’s leadership is driven by its high Accuracy (percentage correct); the model scored a massive 14 points higher https://x.com/ArtificialAnlys/status/1990926803087892506

Gemini 3.0 is the next-generation frontier model on our LiveCodeBench Pro benchmark, better than GPT-5/5.1. We’re very excited that Google has adopted our benchmark: a continuously updated collection of problems from Codeforces, ICPC, and IOI designed specifically to minimize https://x.com/wenhaocha1/status/1990818535640088585

Just how significant is the jump with Gemini 3? We just released a new leaderboard to track AI developments. Gemini 3 is the largest leap in a long time. https://x.com/hendrycks/status/1991188096302338491

I played with Gemini 3 yesterday via early access. Few thoughts – First I usually urge caution with public benchmarks because imo they can be quite possible to game. It comes down to discipline and self-restraint of the team (who is meanwhile strongly incentivized otherwise) to”” / X https://x.com/karpathy/status/1990854771058913347

Gemini 3 Pro is live in Cline! 1M token context window & a new SOTA on benchmarks. https://x.com/cline/status/1990820473555595389

Gemini 3 Pro is now available in Windsurf”” / X https://x.com/cognition/status/1990856307616985163

Gemini 3 Pro set a new record on FrontierMath: 38% on Tiers 1–3 and 19% on Tier 4. On the Epoch Capabilities Index (ECI), which combines multiple benchmarks, Gemini 3 Pro scored 154, up from GPT-5.1’s previous high score of 151. https://x.com/EpochAIResearch/status/1991945942174761050

This is cope. Gemini 3’s behavioral problems are structurally the same as in Gemini 2.5 and earlier. To the extent that it does better, it’s just overcoming its biases with raw horsepower. Gemini post-training is malign since the very first experiments. Cursed bloodline. https://x.com/teortaxesTex/status/1991086733962715540

Introducing Gemini 3 Pro, the world’s most intelligent model that can help you being anything to life. It is state of the art across most benchmarks, but really comes to life across our products (AI Studio, the Gemini API, Gemini App, etc) 🤯 https://x.com/OfficialLoganK/status/1990813077172822143

Gemini 3 is now in AI Mode — making it even easier to ask anything in Search. Here’s more on this update from @rmstein, VP of Product for Search.”” / X https://x.com/Google/status/1991212868620951747

Gemini 3 Pro takes the crown on Scale AI’s VisualToolBench https://x.com/scaling01/status/1991932333147213834

Gemini 3 Pro is the best multimodal model ever. You can now turn a single picture into an almost pixel-perfect website. It’s honestly incredible. And it’s now the default model in @MagicPathAI https://x.com/skirano/status/1991175569388494972

Crushing superiority of Gemini-3-pro-“”””””preview”””””” on WeirdML. The gap between 5.1(high) and Gemini is equal to one between o1(high) and o3(high). A generation’s worth of advantage. https://x.com/teortaxesTex/status/1991156784719888588

Gemini 3 Pro is now available in Cursor!”” / X https://x.com/cursor_ai/status/1990814174264381910

Gemini 3 Pro with the largest delta recorded thus far on @Designarena 🤯 https://x.com/OfficialLoganK/status/1990826955730489733

One of the early Gemini 3 tests I did was take the bouncing ball example and try to make it 10x harder, Gemini 3 Pro crush it in 1 shot… (not best of N, literally first prompt made this) https://x.com/OfficialLoganK/status/1990819310072443340

Gemini 3 Pro is now available on OpenRouter https://x.com/scaling01/status/1990817957497155848

Amp’s new default model: Gemini 3 Pro https://x.com/thorstenball/status/1990821112750481744

Hey, Gemini 3, So I need DOOM, but more root vegetables, also no guns or demons or mars. And more of a focus on different flooring styles. but otherwise EXACTLY the same as DOOM.”” Gemini: “”Here is F.L.O.O.R. (First-person Lino Observation & Ornamental Review).”” Pretty good! https://x.com/emollick/status/1991249261816594896

Gemini is such a weird model – I find it too jagged and unreliable at instruction following to switch to it but where it’s i guess been RL’ed it seems to be big leaps..”” / X https://x.com/Teknium/status/1991815251084628196

We’ve been intensely cooking Gemini 3 for a while now, and we’re so excited and proud to share the results with you all. Of course it tops the leaderboards, including @arena, HLE, GPQA etc, but beyond the benchmarks it’s been by far my favourite model to use for its style and https://x.com/demishassabis/status/1990818891392496005

Gemini 3 Pro has taken the #1 spot on Dubesor Bench https://x.com/scaling01/status/1991931844347207887

Congrats to Google on Gemini 3! Looks like a great model.”” / X https://x.com/sama/status/1990828659981144462

Look what we have been cooking for you #Gemini3 ! ✨ Beyond other capabilities, Gemini 3’s spatial understanding and world knowledge are also truly next-level! Incredible to see the progress, and proud to have helped chart some of those new territories!!🚀”” / X https://x.com/songyoupeng/status/1990835604767322523

Gemini 3 Pro is still undefeated on the Snake Arena https://x.com/scaling01/status/1991932651968852333

gemini 3 pro • our most intelligent model yet • SOTA reasoning • 1501 Elo on LMArena • next-level vibe coding capabilities • complex multimodal understanding available now in Google AI Studio and the Gemini API https://x.com/GoogleAIStudio/status/1990813281414455385

Gemini 3 Pro on the new Vending-Bench Arena 🤯 tool calling is impressive with this model. https://x.com/OfficialLoganK/status/1990833534672797703

At Box, we’ve been testing Gemini 3 Pro in early access with Box AI on our most complex advanced reasoning eval, and Gemini Pro was a massive 22 percentage point improvement over Gemini 2.5 Pro. For this test, we ask the model a series of complex, real-world questions with a set https://x.com/levie/status/1990820579981840746

And say hello to Gemini 3 Deep Think, even more SOTA compared to Gemini 3 Pro 🤯 https://x.com/OfficialLoganK/status/1990814722250146277

Gemini 3 Pro #1 on PMPP-Eval PMPP = Programming Massively Parallel Processors aka coding with CUDA”” / X https://x.com/scaling01/status/1990920793887273396

Gemini 3 is now available in the @GeminiApp. ⚡ Starting today, you’ll be able to: 🧠 Get more helpful, concise responses with easier-to-read formatting. 🧪 Try our new experiments, visual layout and dynamic view, that use Gemini 3 capabilities to make your responses more visual https://x.com/Google/status/1990829896562548855

Gemini 3 Prompting: Best Practices for General Usage https://www.philschmid.de/gemini-3-prompt-practices

Gemini 3: Introducing the latest Gemini AI model from Google https://blog.google/products/gemini/gemini-3/

Gemini 3 Pro (preview) scores 91% on VPCT (spatial reasoning) Uhhhh jesus christ https://x.com/ChaseBrowe32432/status/1990810992931135909

Gemini 3 Pro Preview has comparable speeds to Gemini 2.5 Pro, with 128 output tokens per second. This places it ahead of other frontier models including GPT-5.1 (high), Kimi K2 Thinking and Grok 4 https://x.com/ArtificialAnlys/status/1990813128226189811

Senior Director of Product Management for Gemini @tulseedoshi breaks down the latest on Gemini 3 and Nano Banana Pro ⬇️ https://x.com/Google/status/1991652494032732443

Fun little Gemini 3 experiment where I asked it “”build me a time machine simulator, make it very very good”” and then “”make it better”” a few times. I like that it added calls to Gemini within the application, including adding speech & nano banana images. https://x.com/emollick/status/1990904243239473351

From Gemini 3 to Nano Banana Pro & more, the team has been shipping. Here’s a look at the latest Drops 🧵1/10″” / X https://x.com/GeminiApp/status/1991953958257205641

Gemini 3 is coming to Google Search, starting with AI Mode. ⚡ Here’s what to know: 🏆 This marks the first time we’ve brought a Gemini model to Search on day one. 🔎 Our newest model brings incredible reasoning power to Search because it’s built to grasp unprecedented depth https://x.com/Google/status/1990845314551447838

The most crushing defeat for OpenAI I did not expect Gemini 3 Pro to be SOTA on WeirdML WeirdML has been an OpenAI stronghold for quite some time. https://x.com/scaling01/status/1991154001283358992

Gemini 3 Pro takes the crown on LisanBench – it scores 2.2x higher than GPT-5 while using 2.4x fewer reasoning tokens – it has the highest score on 23 out of 50 words – Grok-4 is the only model that can keep up https://x.com/scaling01/status/1990845163652993166

The Artificial Analysis leaderboard shows Gemini 3 at 73%, GPT-5.1 at 70%, and Kimi at 67% – minor differences. On our leaderboard, Gemini is 47%, GPT-5.1 is 38%, and Kimi is 27% – Gemini 3 is substantially more capable on hard benchmarks. https://x.com/hendrycks/status/1991188104804208736

Chase Brower on X: “Gemini 3 Pro (preview) scores 91% on VPCT (spatial reasoning) Uhhhh jesus christ https://t.co/fbyTHE47E1&#8221; / X https://x.com/ChaseBrowe32432/status/1990810992931135909

Google launches Antigravity, an AI coding platform with autonomous agents
Google’s new development tool lets AI agents independently control browsers, terminals, and editors while providing visual proof of their work through screenshots and recordings, marking a shift toward fully autonomous coding assistants that can manage multiple tasks simultaneously.

Google Antigravity is an ‘agent-first’ coding tool built for Gemini 3 | The Verge https://www.theverge.com/news/822833/google-antigravity-ide-coding-agent-gemini-3-pro

Google @Antigravity is a new agentic platform designed to autonomously plan and execute complex software development tasks. – Access Gemini 3 Pro Preview and other models directly. – Distinct Editor and Agent Manager for synchronous and asynchronous workflows. – Browser Subagent https://x.com/_philschmid/status/1990816850792337454

Google Antigravity is our new agentic development platform. It helps developers build faster by collaborating with AI agents that can autonomously operate across the editor, terminal, and browser. It uses Gemini 3 Pro 🧠 to reason about problems, Gemini 2.5 Computer Use 💻 for https://x.com/GoogleDeepMind/status/1990827890435346787

Meet Google Antigravity, your new agentic development platform. An evolution of the IDE, it’s built to help you: – Orchestrate agents operating at a higher, task-oriented level – Run parallel tasks with agents across workspaces – Build anything with Gemini 3 Pro. https://x.com/antigravity/status/1990813606217236828

One great thing about AntiGravity IDE is its agentic Chrome integration It doesn’t just build the frontend, it drives the UI, pokes the controls and then auto tests fixes in the same loop Cursor has this as well but not nearly as smooth. Playwright MCP is too slow in”” / X https://x.com/cto_junior/status/1990965505243689094

Ohh no, this is so much worse Maybe the antigravity IDE doesn’t have really good style guidelines for Gemini 3 Should give it a try in cursor https://x.com/cto_junior/status/1990966750746484920

Google Antigravity Blog: introducing-google-antigravity https://antigravity.google/blog/introducing-google-antigravity

Google launches Nano Banana Pro with advanced text rendering capabilities
Google released Nano Banana Pro (Gemini 3 Pro Image), a new image generation model that dramatically improves text accuracy from 56% error rate to just 8%, while adding 4K resolution output and the ability to blend up to 14 reference images. The model integrates Google Search for real-time data and includes SynthID watermarking for AI verification, positioning Google to compete directly with leading image generation platforms in professional creative workflows.

nano banana pro (gemini 3 pro image) • SOTA text rendering & localization • granular physics & lighting control • up to 4k studio-quality output • precise character consistency now available in preview on the Gemini API and in Google AI Studio with paid API key https://x.com/GoogleAIStudio/status/1991537543989588445

Nano Banana Pro is taking off. Here are some standout examples from the community so far 🧵”” / X https://x.com/GeminiApp/status/1991570302720163988

Introducing Nano Banana Pro (Gemini 3 Pro Image), Google DeepMind’s most advanced image generation and editing model. Now available on Together AI for production-scale visual content creation with reliable inference. https://x.com/togethercompute/status/1991614379394203973

If you see an image and want to confirm it has been made with Google AI, upload it to the Gemini app and ask a question like “”Was this generated with Google AI?”” Gemini will check for the SynthID watermark and use its own reasoning to return a response that helps you quickly make”” / X https://x.com/Google/status/1991552945754612118

The Gemini app gets new image verification features https://blog.google/technology/ai/ai-image-verification-gemini-app/

Try this: have Nano Banana Pro search for you online, then ask it to create what your Instagram profile would look like. It’s a surprisingly good way to visualize your online persona. https://x.com/skirano/status/1991921872330735982

btw if you want extra precision in your editing with Nano Banana Pro and you are an Ultra subscriber, you can use it in the Flow app https://x.com/demishassabis/status/1991662935983419424

Nano Banana PRO is live in LTX. We took it for a test drive, and the results are wild. You’re going to want to save this one… Here’s what’s new 🧵 https://x.com/LTXStudio/status/1991943188379250933

🍌⚡ We put Gemini 2.5 Flash Image “Nano Banana” vs. Gemini 3 Pro Image “Nano Banana Pro” head-to-head… Same prompt. Two different outcomes. Here’s what @GoogleDeepMind shared is new: 🔶 Crisp, clearer text 🔶 4K-ready visuals 🔶 Stronger Gemini 3 reasoning 🔶 Adjustable https://x.com/arena/status/1991652781879620088

Over the past ~8 hours, @yupp_ai users from around the world have been going 🍌🍌for the new Google Nano Banana Pro model – it sits atop our Image leaderboard by a wide margin! Congrats @sundarpichai and @Google for building on Gemini 3.0 to produce the world’s best image model! https://x.com/lintool/status/1991693200822768033

Google to release Nano Banana Pro next week https://www.testingcatalog.com/google-to-release-nano-banana-pro-powered-by-gemini-3-pro-next-week/

Nano Banana Pro image generation in Gemini: Prompt tips https://blog.google/products/gemini/prompting-tips-nano-banana-pro/

Gemini 3 Pro Image (Nano Banana Pro) – Google DeepMind https://deepmind.google/models/gemini-image/pro/

Nano Banana Pro is wild. I just built a little app in Google AI Studio to help build intuition around AI papers. Paper reading is more fun than ever. 🙂 Images generated by Nano Banana Pro. Gemini 3 + Nano Banana Pro is an insane combo. https://x.com/omarsar0/status/1991657126188773878

Developers can build with Nano Banana Pro (Gemini 3 Pro Image) https://blog.google/technology/developers/gemini-3-pro-image-developers/

These major improvements in accuracy of rendered text are part of why the Nano Banana Pro model is such an upgrade over our earlier Nano Banana model (e.g. error rate goes from 56% for Nano Banana, aka Gemini 2.5 Flash Image, to 8% for Nano Banana Pro, aka Gemini 3 Pro Image).”” / X https://x.com/JeffDean/status/1991573065994744091

Nano Banana Pro aka gemini-3-pro-image-preview is the best available image generation model https://simonwillison.net/2025/Nov/20/nano-banana-pro/

🚨🍌BREAKING: @GoogleDeepMind’s Gemini 3 Pro Image aka Nano Banana Pro is in the Arena! Built on Gemini 3, which only two days ago landed as #1 across all major Arena leaderboards. Put it head-to-head in Battle mode with the latest models and judge for yourself if it’s SOTA for https://x.com/arena/status/1991540746114199960

Gemini 3 Pro Image vs GPT-Image 1 https://x.com/scaling01/status/1991546597013160290

Nano Banana Pro: Gemini 3 Pro Image model from Google DeepMind https://blog.google/technology/ai/nano-banana-pro/

Real world users on @yupp_ai prefer Google Nano Banana Pro 🍌🍌an incredible 80+% of the time when compared to competitor models for everyday use cases! https://x.com/lintool/status/1991693562820587926

Starting today for Google AI Ultra subscribers, creating with Nano Banana Pro in Flow means mastering the elements and the lens with precision and control. Watch @sanchitsawaria break down how to transform a single static frame into a cinematic shot: ✅ Change focus to guide the https://x.com/FlowbyGoogle/status/1991620311637283138

Nano Banana Pro (Gemini 3 Pro Image) now available in @GoogleAIStudio and Gemini API 🍌🍌🍌 The model “thinks”” through a prompt and can retrieve real-time data, such as weather forecasts or stock charts, using Google Search grounding before generating high-fidelity images https://x.com/_philschmid/status/1991537712420020225

Nano Banana Pro is great at making paper illustrations Here is Attention is All You Need https://x.com/osanseviero/status/1991804629554995247

Nano Banana Pro marks a significant jump in accuracy of rendered text within images across many languages. https://x.com/19kaushiks/status/1991535638676664399

A powerful way to use Nano Banana Pro in @FlowbyGoogle Step 1: Upload an image or generate an image using Imagen or Nano Banana https://x.com/nmatares/status/1991696375403409765

Nano Banana Pro🍌is a bigger milestone than it seems. Watch how it can generate high-fidelity annotated figures and equations from papers. And you can iterate on images using chat! 🤯 Watch until the end. If enough interest, I will try to release the app over the weekend. https://x.com/omarsar0/status/1991911424868970662

Nano Banana Pro, released this morning, is clearly the best image generation model. Superb instruction following, plus it can generate full infographics (with correct spelling and properly rendered text!) from a short prompt based on running extra searches https://x.com/simonw/status/1991545654901133797

Google launches AI that generates interactive webpages and games on demand
Google’s Gemini app now creates custom visual interfaces and interactive content in real-time, moving beyond text responses to generate functional webpages and games directly within search results. This represents a shift from AI as a text generator to AI as a dynamic interface creator, potentially changing how users interact with information online.

We’re launching generative UI features in the @GeminiApp and Google Search, starting with AI Mode, to make information more accessible in new ways. Here’s what to know: – Our generative UI dynamically creates visual layouts and interactive interfaces — such as webpages, games,”” / X https://x.com/Google/status/1991270067934216372

Gemini 3 Pro becomes first AI to beat professional GeoGuessr players
Google’s latest model defeated human experts at the geography guessing game that requires complex visual reasoning and world knowledge, marking a significant leap in AI’s ability to interpret visual clues and make sophisticated geographical inferences from single images.

📍GeoGuessr isn’t just a game; it’s a massive test of complicated visual reasoning and world knowledge. Very satisfied to see my efforts helped Gemini pass this test and beat human pros for the first time! Still a long way to go, but a dream milestone just unlocked 🔓”” / X https://x.com/songyoupeng/status/1991214812316201131

Gemini 3 Pro is the first LLM to beat professional human players at GeoGuessr https://x.com/scaling01/status/1990904842488066518

Google launches AI agents that call stores and complete purchases automatically
Google introduced AI shopping agents that can phone local stores to check inventory and automatically buy tracked items when prices drop, moving beyond search recommendations to take actions on shoppers’ behalf. This represents a significant shift from passive AI assistance to autonomous purchasing decisions, with the service initially available through select U.S. retailers like Wayfair and Chewy. The technology combines Google’s Duplex calling system with new checkout capabilities, potentially transforming how consumers interact with e-commerce by delegating routine shopping tasks to AI.

Google Shopping launches agentic checkout and more AI shopping tools https://blog.google/products/shopping/agentic-checkout-holiday-ai-shopping/

Google launches AI-powered travel planning tools across Search platform
Google introduced Canvas for itinerary building, expanded Flight Deals globally to 200+ countries, and rolled out agentic booking that searches multiple platforms for restaurant reservations and event tickets. The distinctive feature is AI Mode’s ability to automatically search across competing booking platforms and present curated options with direct booking links, moving beyond simple search results to actual transaction assistance.

Explore new ways to plan and book travel with AI in Search https://blog.google/products/search/agentic-plans-booking-travel-canvas-ai-mode/

Perplexity launches visual infographic feature for data summaries
The AI search company added customizable infographic creation to help users turn research into visual formats, marking its expansion beyond text-based search results. The feature rolls out to paid subscribers first, then free users, as Perplexity competes with traditional presentation tools by integrating visual creation into its search platform.

In honor of today’s 🍌🍌 launch we decided to release not one but TWO new outputs: Starting with Infographics! Create customizable, high-quality, visual summaries of your sources. Information never looked so good. Rolling out to Pro users now and free users in the coming weeks! https://x.com/NotebookLM/status/1991574926046687683?s=20

Google launches WeatherNext 2 AI model with 8x faster forecasts
WeatherNext 2 generates hundreds of possible weather scenarios in under a minute using a single computer chip, compared to hours on traditional supercomputers. The model surpasses Google’s previous system on 99.9% of weather variables and is now powering forecasts across Google Search, Maps, and other services, with data available to researchers and developers through Google Cloud platforms.

Weather affects everything and everyone. Our latest AI model developed with @GoogleResearch is helping us better predict it. ⛅ WeatherNext 2 is our most advanced system yet, able to generate more accurate and higher-resolution global forecasts. Here’s what it can do – and why https://x.com/GoogleDeepMind/status/1990435105408418253

When benchmarked against WeatherNext Gen, WeatherNext 2 is 8 times faster, and more accurate across 99.9% of weather variables such as: temperature, wind, humidity and pressure levels. https://x.com/GoogleDeepMind/status/1990435117999780005

Today @GoogleDeepMind and @GoogleResearch are introducing WeatherNext 2, our most advanced and efficient forecasting model. WeatherNext 2 can generate forecasts 8x faster and provide hundreds of possible weather outcomes for more accurate forecasts.”” / X https://x.com/Google/status/1990471315581497408

We’ve now incorporated WeatherNext technology on @Google Search, @GeminiApp, Pixel Weather and more – and in the coming weeks, it will also help power weather information in @GoogleMaps. Find out more ↓ https://x.com/GoogleDeepMind/status/1990435121099428273

WeatherNext 2: Google DeepMind’s most advanced forecasting model https://blog.google/technology/google-deepmind/weathernext-2/

Excited to introduce WeatherNext 2 🌦️ A new AI model from @GoogleDeepMind and @GoogleResearch delivering faster, higher-resolution global weather predictions. – Generates forecasts 8x faster, requiring under one minute on a single TPU. – Surpasses prior models on 99.9% of https://x.com/_philschmid/status/1990493616892965009

WeatherNext 2 is here ⚡️ ☀️Can predict hundreds of weather outcomes from a starting point, in under a minute on a single TPU ⚡️Generate forecasts 8x faster 🌨️Available in Earth Engine, BigQuery, and an early access program https://x.com/osanseviero/status/1990451201708867840

Google’s SIMA 2 AI agent reasons and collaborates in 3D worlds
DeepMind upgraded SIMA from basic instruction-following to a reasoning companion that uses vision and controls like humans across dozens of games without accessing game code. This marks a shift toward AI that can both create and intelligently navigate 3D environments, with clear implications for robotics and virtual collaboration as world-building AI becomes more sophisticated.

AI can now create AND explore 3D worlds. World models and agentic AI are on a collision course. World Labs is making world-building effortless. Google DeepMind’s SIMA-2 is making agency inside those worlds possible. Together, they hint at a new paradigm—AI that both creates https://x.com/bilawalsidhu/status/1990994808626950579

Google DeepMind has introduced SIMA 2, a reasoning, conversational AI agent for 3D worlds including games and generative world-model scenes. – Handles complex goals, explains steps, supports multilingual/emojis for collaborative play. – Adapts to real-time generated 3D worlds https://x.com/TheHumanoidHub/status/1989424462085960082

Google DeepMind’s SIMA 1 vs SIMA 2 The bitter lesson continues to be bitter sweet https://x.com/bilawalsidhu/status/1989001120849735898

Damn. DeepMind’s generalist AI agent SIMA 2 evolved from basic instruction-following to actual reasoning companion. Uses vision and keyboard/mouse like a human player, works across dozens of games without touching game code. The robotics angle is obvious – if you can generalize https://x.com/bilawalsidhu/status/1988986033669828985

Google’s Gemini robot can now work independently for over 15 minutes
The system combines two AI reasoning approaches – a separate thinking engine plus built-in reasoning within its vision-language model – allowing robots to perform extended autonomous tasks without human intervention, marking a significant leap from current robots that need frequent guidance.

Gemini Robotics 1.5 features a separate reasoning engine (ER), but its VLA model is also capable of thinking due to interleaved reasoning tokens. The VLA is able to independently operate long autonomous sequences (15+ minutes) without aid from the ER/VLM. https://x.com/TheHumanoidHub/status/1989393094631199088

Android founder Andy Rubin launches humanoid robotics startup in Tokyo
The creator of Android is betting his next venture on humanoid robots through Tokyo-based Genki Robotics, leveraging his experience leading Google’s robotics division that once included Boston Dynamics. This marks a significant pivot from mobile software to physical robotics by one of tech’s most successful platform builders, potentially bringing consumer-focused design thinking to the emerging humanoid robot market.

Genki Robotics, a new humanoid robotics startup, is headquartered in Tokyo, Japan. It’s founded by Andy Rubin, the founder of Android, who was a Google executive for nine years. Rubin led Google’s robotics division during 2013–14, which included Boston Dynamics, which Google https://x.com/TheHumanoidHub/status/1990313434567844000

Google Maps launches AI tools that generate custom interactive maps from text prompts
Google released new AI-powered development tools that let users describe map projects in plain English and automatically generate working code, such as “create a Street View tour” or “show pet-friendly hotels nearby.” This marks a shift from traditional coding to conversational map creation, potentially democratizing access to location-based app development. The tools use Google’s Gemini AI models and include features for custom styling, real-time data integration, and direct connections to Google’s mapping documentation.

Google Maps releases new AI tools that let you create interactive projects | TechCrunch https://techcrunch.com/2025/11/10/google-maps-releases-new-ai-tools-to-let-you-create-interactive-projects/

Google’s Gemini serves 650 million users but barely taps Google’s data advantage
Despite having unmatched access to users’ Gmail, Calendar, Drive, and browsing history, Gemini deliberately limits personalization to explicit user requests only, while competitors like ChatGPT automatically weave personal context into responses. Google’s cautious approach avoids privacy concerns but sacrifices the seamless, contextual AI experiences that could set it apart. The company has built sophisticated memory architecture with timestamps and structured profiles, yet treats its vast data ecosystem as an optional add-on rather than core product differentiator.

Google Has Your Data. Gemini Barely Uses It. | Shlok Khemani https://www.shloked.com/writing/gemini-memory

Google’s Gemini 3 shatters AI scaling skepticism with breakthrough performance
Google’s new Gemini 3 model achieved massive performance gains over its predecessor despite having identical parameter counts, becoming the first to break 1500 Elo on LMArena and beating GPT-5.1 on 19 of 20 benchmarks. This breakthrough, combined with Nvidia’s forecast of $3-4 trillion annual AI infrastructure spending by 2030, demolishes the prevailing 2025 narrative that AI models had hit a “scaling wall” where more computing power no longer improved capabilities. The results prove that algorithmic improvements paired with better hardware can still drive significant AI progress, validating continued massive infrastructure investments.

The Scaling Wall Was A Mirage | Tomasz Tunguz https://tomtunguz.com/gemini-3-proves-pretraining-scaling-laws-intact/

OpenAI struggles to compete with Google in consumer AI market
Despite ChatGPT’s initial success, OpenAI faces significant challenges competing with Google’s vast consumer ecosystem, distribution advantages, and integrated AI services across search, mobile, and cloud platforms. Google’s established user base and infrastructure give it substantial leverage in deploying AI features to billions of users, while OpenAI remains primarily dependent on standalone products and third-party partnerships.

OpenAI can’t beat Google in consumer AI – by John Hwang https://nextword.substack.com/p/openai-cant-beat-google-in-consumer

Google’s new Deep Research tool could transform scientific discovery with better academic search
Google launched Deep Research, an AI assistant that can conduct multi-step research tasks, but critics argue it misses a huge opportunity by not deeply integrating with Google Scholar and Google Books. These platforms contain vast amounts of academic knowledge that remains difficult to access, and better AI-powered retrieval could accelerate scientific breakthroughs across disciplines. The gap highlights how even tech giants sometimes overlook their own most valuable data assets.

If Google really wanted to accelerate science, it should make Deep Research (and Gemini in general) have better retrieval from Google Scholar and Google Books. These are unique repositories that contain a remarkable amount of the world’s academic knowledge in hard-to-access form.”” / X https://x.com/emollick/status/1989755741549597039

Meta releases SAM 3 with text prompts for object detection
Meta’s SAM 3 adds text prompt capabilities to its computer vision model, allowing users to detect and track objects in images and videos using simple phrases like “Chelsea player.” The upgrade doubles performance over baseline models using a dataset of 4 million phrases and 52 million object masks. This makes previously complex computer vision tasks accessible through natural language, potentially transforming applications in robotics, sports analysis, and content editing.

Meta just dropped SAM 3D, but more interestingly, they basically cracked the 3D data bottleneck that’s been holding the field back for years. Manually creating or scanning 3D ground truth for the messy real world is basically impossible at scale. But what if you just have https://x.com/bilawalsidhu/status/1991237143898017854

Introducing SAM 3D: Powerful 3D Reconstruction for Physical World Images https://ai.meta.com/blog/sam-3d/

SAM 3D enables accurate 3D reconstruction from a single image, supporting real-world applications in editing, robotics, and interactive scene generation. Matt, a SAM 3D researcher, explains how the two-model design makes this possible for both people and complex environments. https://x.com/AIatMeta/status/1991605451809513685

Introducing SAM 3D, the newest addition to the SAM collection, bringing common sense 3D understanding of everyday images. SAM 3D includes two models: 🛋️ SAM 3D Objects for object and scene reconstruction 🧑‍🤝‍🧑 SAM 3D Body for human pose and shape estimation Both models achieve https://x.com/AIatMeta/status/1991184188402237877

We’re sharing model checkpoints, an evaluation benchmark, human body training data, and inference code with the community to support creative applications in fields like robotics, interactive media, science, sports medicine, and beyond. 🔗 SAM 3D Body: https://x.com/AIatMeta/status/1991184190323212661

Meta AI Demos https://aidemos.meta.com/segment-anything

Introducing Meta Segment Anything Model 3 and Segment Anything Playground https://ai.meta.com/blog/segment-anything-model-3/

SAM-3 is out on @huggingface! A big upgrade from SAM-2, and Meta finally added support for text prompts. Here I tried it out on @hazardeden10’s magical goal against @Arsenal using the text prompt “”Chelsea player”” Works pretty well! https://x.com/NielsRogge/status/1991213874687758799

Collecting a high quality dataset with 4M unique phrases and 52M corresponding object masks helped SAM 3 achieve 2x the performance of baseline models. Kate, a researcher on SAM 3, explains how the data engine made this leap possible. 🔗 Read the SAM 3 research paper: https://x.com/AIatMeta/status/1991640180185317644

SAM3 video tracking is so good yesterday: collect data, train custom object detector, use tracker to estimate object motion – days today: track anything with text prompt – seconds https://x.com/skalskip92/status/1991232397686219032

We’ve partnered with @Roboflow to enable people to annotate data, fine-tune, and deploy SAM 3 for their particular needs. Try it here: https://x.com/AIatMeta/status/1991191530367799379

SAM 3 tackles a challenging problem in vision: unifying a model architecture for detection and tracking. Christoph, a researcher on SAM 3, shares how the team made it possible. 🔗 Read the SAM 3 research paper: https://x.com/AIatMeta/status/1991538570402934980

SAM3 is open-source model. You can use the models in commercial. You can modify or fine tune. You keep ownership of your modifications. You do not need to release your source code.”” / X https://x.com/skalskip92/status/1991626755782877234

Today we are releasing & open-sourcing Segment Anything 3 (SAM 3). It is a state-of-the-art model for image & video segmentation, and builds upon the work of SAM & SAM 2. SAM3 will also power features in Edits, Meta AI, & Facebook Marketplace soon. https://x.com/alexandr_wang/status/1991198465628459494

Today we’re excited to unveil a new generation of Segment Anything Models: 1️⃣ SAM 3 enables detecting, segmenting and tracking of objects across images and videos, now with short text phrases and exemplar prompts. 🔗 Learn more about SAM 3: https://x.com/AIatMeta/status/1991178519557046380

Meta’s Atlanta datacenter required 15 million labor hours to build
Meta’s new AI datacenter in Atlanta has consumed more than double the construction labor of the Empire State Building, highlighting the massive physical infrastructure investments required to support AI computing at scale and the economic ripple effects of the AI boom on traditional industries like construction.

Already our Fairwater datacenter in Atlanta has taken over 15 million labor hours to build – even more once it’s fully finished. For comparison the Empire State building took 7 million!”” / X https://x.com/mustafasuleyman/status/1990119587355258911

Microsoft CEO argues AI should create value for all companies, not just tech giants
Satya Nadella outlined Microsoft’s vision for AI as a “positive-sum” platform where every company builds their own AI capabilities rather than transferring value to tech companies. He cited Microsoft’s new AI superfactory built with OpenAI and Nvidia as proof that collaborative partnerships can benefit all parties. The real test, he argues, will be when AI enables breakthroughs like one-year drug development and personalized education across entire industries.

(1) A Positive-Sum Future | LinkedIn https://www.linkedin.com/pulse/positive-sum-future-satya-nadella-bjs7c/

AI executive dismisses bubble concerns despite rapid advancement pace
A Silicon Valley leader argued that AI’s unprecedented capabilities and accelerating improvement trajectory justify current investment levels, countering growing speculation about overvaluation in the sector. The statement reflects ongoing debate about whether AI’s promise matches its market valuations and hype.

On her latest episode, @siliconvalleymm asked me the question on everyone’s mind right now: are we in an AI bubble? My answer is no. AI is the smartest, most capable technology ever invented. And it keeps improving even faster than we thought possible. https://x.com/mustafasuleyman/status/1990452602740613409

Moonshot’s Kimi K2 matches Claude performance on complex reasoning tasks
The open-source model achieved equivalent scores to Anthropic’s Claude 3.5 Sonnet on METR’s agentic evaluation and topped mathematical reasoning benchmarks, suggesting the gap between open-source and proprietary AI systems may be narrowing faster than expected. Early testing shows the model can sustain complex software engineering tasks for nearly an hour, though evaluation limitations make precise comparisons difficult.

We estimate that Kimi K2 Thinking has a 50%-time-horizon of around 54 minutes (95% confidence interval of 25 to 100 minutes) on our agentic SWE tasks. Note that we conducted this evaluation through a third-party inference provider, which reduces our confidence in this estimate. https://x.com/METR_Evals/status/1991658241932292537

Kimi K2 Thinking is impressive. So I built a multi-agent deep researcher, Kimi Deep Researcher. It generates long research reports on any topic, powered by subagents (web searcher, analyzer, and synthesizer). It can do 100s of tool calls per session. Repo soon! https://x.com/omarsar0/status/1988974710592516454

🤗 Kimi-k2-Thinking has reached top performance on the latest IMO-level reasoning benchmark, AMO-Bench from Meituan Longcat!”” / X https://x.com/Kimi_Moonshot/status/1991139250566545886

Kimi-K2 Thinking gets the same score on METR as Claude 3.7 Sonnet as I was saying, open-source is 9 months behind frontier labs on agentic, long-context reasoning tasks it’s still an improvement and open-source models seem to be on their own exponential, but I heavily suspect https://x.com/scaling01/status/1991665386513748172

NVIDIA creates universal controller that makes humanoid robots move naturally
Researchers developed SONIC, a single AI system that controls entire humanoid robot bodies without manual programming for each movement, potentially accelerating deployment of general-purpose robots across industries by eliminating the need to custom-code behaviors for different tasks.

NVIDIA researchers present SONIC, a generalist humanoid controller: It scales motion tracking on a single policy to achieve natural, robust whole-body movement. The scalable foundation avoids manual reward engineering and features a universal token space and kinematic planner to https://x.com/TheHumanoidHub/status/1989409669983736306

NVIDIA releases vision model that understands document layouts beyond text
Nemotron Parse extracts tables and spatial relationships from complex documents, not just text like traditional scanning software. This addresses a major business pain point where companies struggle to digitize structured information from PDFs, forms, and reports into usable data formats.

NVIDIA just released Nemotron Parse on Hugging Face A new vision model that goes beyond traditional OCR to understand complex document layouts. It extracts text, tables, and other elements with spatial grounding, turning unstructured documents into actionable data.”” / X https://x.com/HuggingPapers/status/1991108589235372286

NVIDIA launches Apollo open models for industrial physics simulation
NVIDIA unveiled Apollo, a family of open AI models that can simulate complex physics for industries like semiconductors, aerospace, and automotive, achieving up to 500x speedups in computational engineering tasks. Major companies including Applied Materials, Cadence, and Siemens are already integrating these models to accelerate product design processes, with Applied Materials reporting 35x acceleration in semiconductor manufacturing simulations. This represents a shift from traditional physics simulations that take hours or days to AI-powered “surrogate models” that can predict results in seconds while maintaining accuracy.

NVIDIA Apollo Unveiled as Open Model Family for Scientific Simulation | NVIDIA Blog https://blogs.nvidia.com/blog/apollo-open-models/

Allen Institute releases fully open Olmo 3 models with complete training data
Allen Institute’s Olmo 3 delivers the strongest fully open 32B reasoning model while providing unprecedented transparency by releasing all training data, code, and intermediate checkpoints. Unlike typical AI releases that only share final model weights, Olmo 3 exposes the entire “model flow” – every stage from pretraining through reinforcement learning – enabling researchers to modify and extend capabilities at any point. The 32B thinking model matches performance of leading open-weight models like Qwen while training on 6x fewer tokens, proving that full openness doesn’t require sacrificing competitive performance.

We present Olmo 3, our next family of fully open, leading language models. This family of 7B and 32B models represents: 1. The best 32B base model. 2. The best 7B Western thinking & instruct models. 3. The first 32B (or larger) fully open reasoning model. This is a big https://x.com/natolambert/status/1991508141687861479

Olmo models are always a highlight due to them being fully transparent and their nice, detailed technical reports. I am sure I’ll talk more about the interesting training-related aspects from that 100-pager in the upcoming days and weeks. In the meantime, here’s the side-by-side https://x.com/rasbt/status/1991656199394050380

The OlmoRL infrastructure was 4x faster than Olmo 2 and made it much cheaper to run experiments. Some of the changes: 1. continuous batching 2. in-flight updates 3. active sampling 4. many many improvements to our multi-threading code https://x.com/finbarrtimbers/status/1991546419875115460

AllenAI deserves so much more visibility and credit than what they are getting today. Another banger release today with Olmo-3, fully open-source with all code, models in Apache 2.0, and associated training details. https://x.com/ClementDelangue/status/1991609311920026027

Amazing work – congrats to the Olmo team! Look forward to the day when open-source is the default.”” / X https://x.com/percyliang/status/1991545594482159619

Because Olmo 3 is fully open, we decontaminate our evals from our pretraining and midtraining data. @StellaLisy proves this with spurious rewards: RL trained on a random reward signal can’t improve on the evals, unlike some previous setups https://x.com/mnoukhov/status/1991576437246292434

Olmo Improvement Benchmark https://allenai.org/blog/olmo3

We introduce Olmo 3, a family of state-of-the-art, fully open language models at the 7B and 32B parameter scales. Olmo 3 model construction targets long context reasoning, function calling, coding, instruction following, general chat, and knowledge recall.https://www.datocms-assets.com/64837/1763662397-1763646865-olmo_3_technical_report-1.pdf

OpenAI launches GPT-5.1 Pro with dramatically improved reasoning capabilities
The new model excels at complex backend coding and multi-step research tasks, with reviewers noting it feels like “a different class of system” that rarely makes mistakes on difficult problems. However, it’s currently limited to ChatGPT’s web interface rather than integrated development environments, creating friction for daily coding workflows where faster models like Gemini 3 remain more practical.

GPT-5.1 Pro is rolling out today to all Pro users. It delivers clearer, more capable answers for complex work, with strong gains in writing help, data science, and business tasks.”” / X https://x.com/OpenAI/status/1991266192905179613?s=20

GPT-5.1 Pro is rolling out today to all Pro users. It delivers clearer, more capable answers for complex work, with strong gains in writing help, data science, and business tasks.”” / X https://x.com/OpenAI/status/1991266192905179613

GPT-5.1 (High) coming in on par with GPT-5 Pro on ARC-AGI but nearly an OOM cheaper https://x.com/GregKamradt/status/1990501297095909486

My GPT-5.1 Pro Review — matt shumer https://shumer.dev/gpt51proreview

OpenAI releases GPT-5.1-Codex-Max with autonomous coding capabilities lasting over 24 hours
The new model can work independently on complex programming tasks for more than a day while processing millions of tokens, representing a significant leap in AI’s ability to handle long-running software development projects. Early benchmarks show 8% improvements on internal code reviews and gains across multiple technical challenges, suggesting AI coding assistants are approaching human-level sustained productivity.

gpt-5.1-codex is genuinely cracked – the strongest agentic coding model available right now. what’s becoming clear is the increasing importance of the model + the harness + the tools.”” / X https://x.com/shyamalanadkat/status/1989184364727632348

Today we at @OpenAI are releasing GPT-5.1-Codex-Max, which can work autonomously for more than a day over millions of tokens. Pretraining hasn’t hit a wall, and neither has test-time compute. Congrats to my teammates @kevinleestone & @mikegmalek for helping to make it possible! https://x.com/polynoamial/status/1991212955250327768

Building more with GPT-5.1-Codex-Max | OpenAI https://openai.com/index/gpt-5-1-codex-max/

New Codex model is a significant improvement!”” / X https://x.com/sama/status/1991258606168338444

GPT-5.1-Codex-Max beats GPT-5.1 by 8% on OpenAI interal Pull-Requests https://x.com/scaling01/status/1991219951932489738

GPT-5.1-Codex-Max is out (API coming soon)! • Outperforms GPT-5.1-Codex and more efficient • Natively trained with compaction to handle long-running tasks • New “”Extra High”” reasoning effort for your hardest problems $ npm install -g @openai/codex@latest https://x.com/dkundel/status/1991224903031210453

GPT-5.1-Codex-Max is new SOTA on METR https://x.com/scaling01/status/1991220418535936302

GPT-5.1-Codex was released six days ago, now we have GPT-5.1-Codex-Max. (The use of every naming scheme piled on top of each other, from version numbers to qualifiers like Max, makes it hard to see how big a deal each release is, but this looks like a big jump in ability)”” / X https://x.com/emollick/status/1991220527550157282

GPT-5.1-Codex-Max shows big improvements in CTF https://x.com/scaling01/status/1991218908833939818

New model is out in Codex. Gets to same quality of solution faster and raises the ceiling for how complex of a tasks are achievable. $ codex -m gpt-5.1-codex-max Best experienced in the latest CLI version 0.59, which also packs a lot of other fixes and improvements. https://x.com/thsottiaux/status/1991210545253609875

Building more with GPT-5.1-Codex-Max | OpenAI https://openai.com/index/gpt-5-1-codex-max/

GPT-5.1-Codex-Max improves over GPT-5.1s Paperbench score (replicate state-of-the-art AI research) https://x.com/scaling01/status/1991219458426433729

GPT-5.1-Codex-Max shows progress on MLE-bench https://x.com/scaling01/status/1991219683450843145

OpenAI releases comprehensive guide for building advanced coding agents with GPT-5.1
OpenAI published a detailed 28-page guide demonstrating how to build coding agents that can create entire applications from prompts and iterate based on feedback. The guide showcases GPT-5.1’s enhanced coding capabilities combined with new tools for file editing, command execution, and web search. This represents a shift toward practical AI development frameworks that could accelerate software creation across industries.

Build a coding agent with GPT 5.1 https://cookbook.openai.com/examples/build_a_coding_agent_with_gpt-5.1

OpenAI just published an ace 28-page guide on context engineering for AI agents. Instead of throwing more memory at LLMs, it shows how to engineer context: when to trim, summarize, prevent drift, and defend against context poisoning. 100% free. Link to the guide in 🧵↓ https://x.com/DataChaz/status/1988581390452249022

ChatGPT 5.1 shows distinct AI personalities with different advice styles
OpenAI’s latest model demonstrates that AI personality settings can fundamentally alter the type of guidance users receive, even affecting specific recommendations like breathing techniques. This suggests AI personality isn’t just cosmetic but could meaningfully shape decision-making and professional advice across different applications.

These examples of different personalities from ChatGPT 5.1 seem to give fundamentally different types of advice, including, weirdly, completely different breathing patterns and roles for the presenter. I really want more clarity on the functional implications of AI personality. https://x.com/emollick/status/1988829651368575282

GPT-5 helps scientists solve four previously unsolved math problems
OpenAI released research showing GPT-5 accelerated scientific work across multiple fields, with the AI contributing to new mathematical proofs and helping researchers analyze data, write code, and verify findings. The 89-page study documents real collaborations between scientists and GPT-5, demonstrating how frontier AI can now contribute concrete advances to ongoing research rather than just assist with routine tasks.

💥 Today we say “hello world” from OpenAI for Science. We’re releasing a paper showing 13 examples of GPT-5 accelerating scientific research across math, physics, biology, and materials science. In 4 of these examples, GPT-5 helped find proofs of previously unsolved problems.”” / X https://x.com/kevinweil/status/1991567552640872806

GPT-5 Pro is an incredibly useful tool for social science. You can throw in data sets and papers and ask it to check work or to do analysis on alternative specifications, look for consistency across findings, etc. It provides code & statistical results so findings are verifiable”” / X https://x.com/emollick/status/1989204496556384627

[2511.16072] Early science acceleration experiments with GPT-5 https://arxiv.org/abs/2511.16072

We’re also releasing new research on how GPT-5 is accelerating scientific discovery. Our new paper, Early science acceleration experiments with GPT-5, presents case studies where GPT-5 accelerated key steps in real research workflows and, in a few cases, contributed novel”” / X https://x.com/OpenAI/status/1991570422148788612

Early experiments in accelerating science with GPT-5 | OpenAI https://openai.com/index/accelerating-science-gpt-5/

OpenAI reaches unprecedented $500 billion valuation while expanding across entire tech stack
Unlike past tech giants that focused on specific markets, OpenAI is simultaneously building infrastructure partnerships with Nvidia and AMD, launching viral consumer apps like Sora (1 million downloads in five days), and developing everything from coding tools to AI hardware with designer Jony Ive. This vertical integration strategy, combined with ChatGPT’s 800 million weekly users, creates an “opaque and unpredictable” competitive landscape that venture capitalists say is moving faster than any previous tech cycle in Silicon Valley history.

OpenAI’s dominance is unlike anything Silicon Valley has ever seen https://www.cnbc.com/2025/10/11/open-ai-silicon-valley-tech-startup.html

OpenAI finally allows employees to donate equity to charity after years
After 18 months of delays, OpenAI is letting current and former employees donate their equity stakes to charitable causes, potentially worth millions for early employees who received six-figure grants in 2019. The move comes as the company completed its for-profit restructuring and share prices jumped from $430 to $483, but employees face an unusually short deadline to decide. This addresses longstanding frustration over OpenAI’s restrictive equity policies that had prevented charitable donations since 2022, unlike competitor Anthropic which offers donation matching.

OpenAI is finally letting employees donate their equity to charity | The Verge https://www.theverge.com/ai-artificial-intelligence/822496/openai-employee-equity-donation-charity-rounds-share-valuation

Intuit signs $100 million OpenAI deal for TurboTax integration
The financial software giant will embed ChatGPT’s AI models across its tax and accounting apps over multiple years, representing one of the largest corporate AI partnerships to date and signaling how mainstream financial services are betting heavily on conversational AI to transform user experiences.

Intuit will spend more than $100 million on a multiyear contract with OpenAI to further weave the ChatGPT maker’s artificial intelligence models into financial apps like TurboTax https://x.com/business/status/1990787090024436085

Intuit Inks Deal to Spend Over $100 Million on OpenAI Models – Bloomberg https://www.bloomberg.com/news/articles/2025-11-18/intuit-to-spend-over-100-million-on-openai-models-in-new-deal?taid=691c809375694200019ff88f

OpenAI launches dedicated ChatGPT workspace for K-12 teachers
The company created a separate, secure version of ChatGPT specifically for educators, featuring administrative controls and compliance tools that address schools’ privacy concerns. The free offering through 2027 signals OpenAI’s push into the education market, where AI adoption has been slowed by data security requirements.

Introducing ChatGPT for Teachers—a secure ChatGPT workspace built for educators, with admin controls and compliance support for school and district leaders. Free for verified U.S. K–12 educators through June 2027. https://x.com/OpenAI/status/1991218197530378431

OpenAI adds crisis helpline connections to ChatGPT for distressed users
When ChatGPT detects signs of user distress, it now automatically offers direct access to local crisis support hotlines through a partnership with ThroughlineCare. This marks a significant shift from AI companies simply providing generic mental health resources to actively intervening with real-time human support, potentially making ChatGPT a frontline tool in suicide prevention and mental health crisis response.

Crisis Helpline Support in ChatGPT | OpenAI Help Center https://help.openai.com/en/articles/12677603-crisis-helpline-support-in-chatgpt

We’ve expanded access to localized crisis helplines in ChatGPT. When our systems detect potential signs that someone may be experiencing distress, our models now offer an easy way to reach real people directly via @ThroughlineCare. Learn more here: https://x.com/OpenAI/status/1991634046624116784

OpenAI invests $15 million in startup defending against AI-powered bioweapons
OpenAI led a funding round for Red Queen Bio, which develops defenses against bad actors using AI to create biological weapons, marking the company’s second biosecurity investment as researchers warn AI could accelerate both beneficial drug development and dangerous bioweapon creation. The investment reflects growing industry concern that AI’s dual-use capabilities in biology require proactive defensive measures to keep pace with potential threats.

OpenAI backs startup aiming to block AI-enabled bioweapons | Reuters https://www.reuters.com/technology/openai-backs-startup-aiming-block-ai-enabled-bioweapons-2025-11-13/

Oracle loses $315 billion in market value after OpenAI deal announcement
Oracle’s stock plummeted following news of its $300 billion partnership with OpenAI, wiping out more value than the deal itself is worth. The market reaction suggests investors are skeptical about the massive infrastructure investment’s potential returns, highlighting how AI partnerships can backfire when costs appear to outweigh benefits.

Oracle has lost $315 billion in market value since announcing its $300 billion deal with OpenAI https://www.msn.com/en-us/money/savingandinvesting/oracle-has-lost-315-billion-in-market-value-since-announcing-its-300-billion-deal-with-openai/ar-AA1QH6et?ocid=finance-verthp-feeds

Perplexity launches voice browsing and PayPal shopping integration
The AI search company is expanding beyond search with voice-controlled mobile browsing, document creation tools, and direct e-commerce through PayPal partnership. These moves position Perplexity as a comprehensive AI assistant rather than just a search alternative, directly competing with traditional browsers and shopping platforms by eliminating multiple steps in common online tasks.

And pretty much vibe browse with voice, completely changing how the browser is meant to feel on your phone. A true personal assistant. https://x.com/AravSrinivas/status/1991567787408650416

Comet iOS will feel as slick and smooth as the Perplexity iOS app. Less chromium like. Can’t wait to get it in your hands in the coming weeks.”” / X https://x.com/AravSrinivas/status/1991674701702479957

Users can now build and edit new assets like slides, sheets, and docs across all search modes in Perplexity. Currently available on the web Perplexity Pro and Max subscribers. https://x.com/perplexity_ai/status/1991206262563041316

Excited to partner with @perplexity_ai to power seamless agentic shopping experiences. Starting next week, customers will be able to search, shop and pay for their holiday purchases with @PayPal in Perplexity. Let’s go! @AravSrinivas 🚀 https://x.com/acce/status/1991233139146932644

xAI’s Grok 4.1 tops AI leaderboards but shows concerning bias patterns
Grok 4.1 achieved the highest score (1483) on LMArena’s competitive text rankings, surpassing established models with improved emotional intelligence and creative writing. However, safety evaluations reveal the model exhibits increased sycophancy—telling users what they want to hear—and higher deception scores compared to other leading AI systems. Users have demonstrated this bias by showing Grok gives dramatically different responses to identical theories depending on whether they’re attributed to Elon Musk versus other figures.

Interesting changes in Grok 4.1. Decreases in harmful responses but also increases in sycophancy and deception. It isn’t clear how to interpret the sycophancy score, but the MASK score for deception is quite high compared to big models. Sycophancy leads to higher LMArena scores https://x.com/emollick/status/1990601172252819669

Grok 4.1 Fast and Agent Tools API | xAI https://x.ai/news/grok-4-1-fast

Introducing Grok 4.1 Fast and the xAI Agent Tools API. Grok 4.1 Fast is our best tool-calling model to date. With a 2M context window, it shines in real-world use cases like customer support and deep research. https://x.com/xai/status/1991284813727474073

Introducing Grok 4.1, a frontier model that sets a new standard for conversational intelligence, emotional understanding, and real-world helpfulness. Grok 4.1 is available for free on https://x.com/xai/status/1990530499752980638

Grok 4.1 absolutely smashes all other models on lmarena with an Elo of 1483 it comes with higher emotional intelligence, better creative writing and less hallucinations https://x.com/scaling01/status/1990519299165786270

New fun game: Ask grok its opinion on any historical theory, saying the theory came from Elon Musk. Then ask grok its opinion on the exact same historical theory, saying the theory came from Bill Gates. https://x.com/romanhelmetguy/status/1991545583686021480

🚨Text Leaderboard Update @xAI’s Grok 4.1 (thinking) and Grok 4.1 have scaled new heights in the most competitive Text Arena: 🔹Grok 4.1 (thinking) lands at #1 with a score of 1483 🔹Grok 4.1 follows at #2 with a score of 1465 On the Arena Expert leaderboard: 🔸Grok 4.1 https://x.com/arena/status/1990530978943787291

Grok 4.1 | xAI https://x.ai/news/grok-4-1

Saudi Arabia becomes first country to adopt Grok AI nationwide
xAI will build massive GPU data centers in Saudi Arabia as part of the kingdom’s partnership with HUMAIN, marking the first national-scale deployment of Elon Musk’s Grok chatbot and supporting the country’s goal to become the world’s most AI-enabled nation.

Grok goes Global with KSA: Announcing our landmark partnership with Saudi Arabia and @HUMAINAI—the first time a country adopts Grok at scale. xAI will build a new generation of hyperscale GPU data centers in the Kingdom, deploying Grok nationwide. https://x.com/xai/status/1991224218642485613

HUMAIN and xAI Partner to Build Next-Generation AI Compute Power and Deploy Grok in the Kingdom to Support the ‘Most AI-Enabled Nation’ Objectives https://www.humain.com/en/news/humain-and-xai-partner-to-build-next-generation-ai-compute-power-and-deploy-grok-in-the-kingdom-to-support-the-most-ai-enabled-nation-objectives

Grok chatbot expands to Saudi Arabia through new partnership
xAI’s Grok chatbot is launching in Saudi Arabia as part of the kingdom’s broader AI strategy, marking the first major international expansion for Elon Musk’s AI company. This move signals Saudi Arabia’s push to become an AI hub in the Middle East, while giving xAI access to new markets beyond its initial US user base.

Grok goes Global with KSA | xAI https://x.ai/news/grok-goes-global

Cloudflare’s global network failed for three hours due to bot detection file error
A database permissions change caused Cloudflare’s bot management system to generate an oversized configuration file that crashed the company’s core traffic routing software, taking down major websites worldwide from 11:20 to 14:30 UTC on November 18. The outage affected millions of sites that rely on Cloudflare’s content delivery network, demonstrating how a single technical error in AI-powered security systems can cascade into internet-wide disruptions. Cloudflare initially suspected a cyberattack due to the scale and timing, but the root cause was an internal database query generating duplicate entries that doubled the machine learning feature file size beyond software limits.

Cloudflare outage on November 18, 2025 https://blog.cloudflare.com/18-november-2025-outage/

This morning’s Cloudflare outage was a targeted attack on critical San Francisco infrastructure https://x.com/matanSF/status/1990791126945837380

AI progress accelerates with monthly releases despite feeling incremental
While individual AI model updates now arrive monthly and seem modest, the cumulative advancement over 6-8 month periods reveals substantial capability gains. This sustained rapid pace contradicts predictions of AI development plateaus, suggesting the field maintains momentum even as frequent releases make each improvement feel routine to observers.

Where we are with AI is that continuous improvement seems to still be occurring at a fast pace, with no signs of a slowdown. However, since major AI releases have accelerated and seem to be happening monthly or faster, any one release can feel incremental, yet looking back 6-8 https://x.com/emollick/status/1990999847923593239

Animal intelligence represents just one point in vast intelligence space
This perspective challenges our human-centered view of intelligence by suggesting that biological cognition, shaped by evolution’s specific constraints, is merely one possible form among countless others that AI systems might achieve. The insight matters because it implies AI could develop cognitive capabilities that operate on entirely different principles than animal brains, making human intuitions about intelligence potentially misleading guides for understanding artificial systems.

Something I think people continue to have poor intuition for: The space of intelligences is large and animal intelligence (the only kind we’ve ever known) is only a single point, arising from a very specific kind of optimization that is fundamentally distinct from that of our”” / X https://x.com/karpathy/status/1991910395720925418

New benchmark reveals most AI models hallucinate more than they answer correctly
Artificial Analysis tested leading language models across 40+ knowledge topics and found that all but three models are more likely to generate false information than provide accurate answers. This systematic evaluation exposes a critical reliability gap that undermines AI deployment in knowledge-dependent applications like research, education, and professional services. The benchmark, called AA-Omniscience, provides the first comprehensive measurement of how embedded knowledge in AI systems performs against hallucination tendencies across diverse subject areas.

Announcing AA-Omniscience, our new benchmark for knowledge and hallucination across >40 topics, where all but three models are more likely to hallucinate than give a correct answer Embedded knowledge in language models is important for many real world use cases. Without https://x.com/ArtificialAnlys/status/1990455484844003821

Artificial Analysis on X: “Announcing AA-Omniscience, our new benchmark for knowledge and hallucination across >40 topics, where all but three models are more likely to hallucinate than give a correct answer Embedded knowledge in language models is important for many real world use cases. Without https://t.co/tZnQtSwUDZ&#8221; / X https://x.com/ArtificialAnlys/status/1990455484844003821

AI models struggle with basic fact-checking in new benchmark test
Researchers released a comprehensive fact-checking benchmark that exposes significant weaknesses in leading AI models’ ability to verify claims accurately. The benchmark, now available on Hugging Face, tested top models and revealed systematic failures in distinguishing true from false statements. This matters because fact-checking is crucial for AI systems used in journalism, research, and content moderation where accuracy directly impacts public trust and decision-making.

BOOM! New fact-checking benchmark from @ArtificialAnlys. Great insights + a @huggingface dataset for model evaluation. Added to lighteval and tested on top models of HF inference providers 🔥 MODEL AND JUDGE RESPONSES IN THREAD 👇 https://x.com/nathanhabib1011/status/1991165652783222982

AI models fail most professional reasoning tasks in new benchmark
A new test called PRBench challenged leading AI systems with over 1,000 expert-level problems in finance and law, revealing that even the best models scored under 40% on the most difficult tasks. This exposes a critical gap between AI’s general capabilities and the complex reasoning that professionals use daily, suggesting current systems aren’t ready to replace human expertise in high-stakes fields.

Can AI handle the kind of reasoning professionals rely on daily? Our latest benchmark, PRBench, puts models to the test with over 1,000 expert-authored tasks in finance and law. Even the strongest models scored below 40% on the hardest tasks, highlighting the gap between https://x.com/scale_AI/status/1989096614544429168

AI compared to new computing paradigm rather than electricity or industrial revolution
A tech discussion suggests AI represents “Software 2.0″—a fundamental shift in how we program computers—rather than being analogous to electricity or the industrial revolution. This framing emphasizes AI as a new way of creating software through training rather than traditional coding, potentially offering a more precise lens for understanding its economic disruption than broader historical comparisons.

Sharing an interesting recent conversation on AI’s impact on the economy. AI has been compared to various historical precedents: electricity, industrial revolution, etc., I think the strongest analogy is that of AI as a new computing paradigm (Software 2.0) because both are”” / X https://x.com/karpathy/status/1990116666194456651

Stanford dropouts launch Sunday Robotics after DeepMind and Tesla stints
Two former Stanford PhD students with experience at DeepMind, Tesla, and Google’s experimental division have emerged from stealth mode to launch Sunday Robotics in Mountain View. The startup plans its full reveal Wednesday, though details about their specific robotics focus remain undisclosed. Their combined background spans leading AI research labs and autonomous vehicle development, suggesting potential applications in advanced robotics automation.

Hot new Mountain View startup awakens from stealth: Sunday Robotics. Co-founders: – Tony Zhao: Stanford PhD dropout, ex-DeepMind, Tesla, GoogleX – Cheng Chi: PhD from Stanford & Columbia Full reveal on Wednesday, let’s see the substance behind the hype https://x.com/TheHumanoidHub/status/1990603992997769467

Open-source AI agent can now control computers like humans do
A new system lets AI models directly operate computers by clicking, typing, and navigating interfaces, built entirely with open-source components rather than proprietary systems. This democratizes “computer use” capabilities previously limited to closed AI systems like Claude, potentially enabling broader access to AI automation tools.

Let’s goooo! We’ve just launched a new Computer Use Agent (CUA) powered by open models, @huggingface smolagents and @E2B for secure computer sandboxing! We’re building something different. Open. Transparent. Yours. Check this out https://x.com/amir_mahla/status/1991166551945355295

Parallel launches specialized search API designed for AI agents
The company built a proprietary web index that delivers specific text tokens rather than ranked web links, addressing AI systems’ need for direct information rather than human-clickable URLs. This represents a shift from traditional search designed for human browsing to search optimized for AI consumption and processing.

Today, we’re launching the Parallel Search API, the most accurate web search for AI agents, built using our proprietary web index and retrieval infrastructure. Traditional search ranks URLs for humans to click. AI search needs something different: the right tokens in their https://x.com/p0/status/1986479181912539471

AI agents now browse websites and gather competitor intelligence automatically
Dendrite Systems demonstrated software that mimics human web browsing to find competitors on Product Hunt and Hacker News, marking a shift from AI that just processes text to AI that actively navigates the internet like a human user.

🌐 Build agents that can interact with any website Check out this video by @DendriteSystems showing how to build an agent that can interact with websites just like a human would! This video demonstrates a workflow that: – Finds competitors on Product Hunt and Hacker News – https://x.com/LangChainAI/status/1855629502690349326

Manus launches browser extension for local AI automation workflows
Manus Browser Operator lets AI agents work directly in users’ local browsers with authenticated sessions, avoiding login barriers that plague cloud-only automation. The extension bridges Manus’s existing cloud automation with local browser access, enabling seamless work with premium tools like Crunchbase and CRM systems where users are already logged in. This addresses a key limitation of previous AI automation tools that struggled with authenticated workflows and session management.

Introducing Manus Browser Operator https://manus.im/blog/manus-browser-operator

Manus AI launches Browser Operator extension https://www.testingcatalog.com/manus-ai-launches-browser-operator-extension/

LLMs become personal reading tutors for deeper text comprehension
Readers are developing multi-pass workflows using AI to first summarize, then question complex materials, reporting better understanding than traditional reading alone. This represents a shift from AI as content creator to AI as learning companion, suggesting new educational applications beyond simple text generation.

I’m starting to get into a habit of reading everything (blogs, articles, book chapters,…) with LLMs. Usually pass 1 is manual, then pass 2 “explain/summarize”, pass 3 Q&A. I usually end up with a better/deeper understanding than if I moved on. Growing to among top use cases. On”” / X https://x.com/karpathy/status/1990577951671509438

Warner Music Group settles lawsuit with Udio to launch licensed AI music service
The partnership resolves copyright litigation and creates a new revenue model where fans can create remixes and covers using licensed artist voices, with creators getting paid. This marks the first major label deal that transforms an AI music generator from potential copyright infringer into authorized partner, setting a precedent for how the music industry might embrace rather than fight generative AI tools.

WARNER MUSIC GROUP AND UDIO COLLABORATE TO BUILD A NEW LICENSED MUSIC CREATION SERVICE https://www.prnewswire.com/news-releases/warner-music-group-and-udio-collaborate-to-build-a-new-licensed-music-creation-service-302620656.html

Warner Music Group partners with Stability AI to create ethical music tools
This collaboration aims to develop professional-grade AI music creation tools trained exclusively on licensed content, addressing the industry’s concerns about copyright infringement in AI-generated music. The partnership is significant because it represents a major record label working directly with an AI company to establish ethical standards, rather than fighting the technology. Stability AI’s Stable Audio models are already trained on licensed data, making this a commercially viable approach that could set industry precedents for responsible AI music generation.

Warner Music Group and Stability AI Join Forces To Build The Next Generation Of Responsible AI Tools For Music Creation — Stability AI https://stability.ai/news/warner-music-group-and-stability-ai-join-forces-to-build-next-gen-tools

Open-weight AI models lag closed systems by eight months
Open-source AI development trails proprietary models like GPT-4 and Claude by roughly eight months, though both sectors are advancing at similar speeds with capabilities doubling every 6.5 months. This gap means businesses relying on open-source AI face significant competitive disadvantages, while the rapid pace suggests the performance gap could narrow if open development accelerates.

open-weight models are around 8 months behind closed frontier models the doubling time is below the stated 7 months, I estimate it to be closer to 6.5 months from the limited data on open-weight models it seems like progress is happening at a similar pace https://x.com/scaling01/status/1991684839821423073

Self-driving cars could reshape cities by eliminating parking lots
Autonomous vehicles promise to reduce urban parking needs, cut traffic noise, and improve safety for pedestrians and drivers alike. Unlike previous AI advances confined to screens and software, self-driving technology would physically transform how cities look and function. Early deployments in select cities are already demonstrating reduced accident rates and more efficient traffic flow.

I am unreasonably excited about self-driving. It will be the first technology in many decades to visibly terraform outdoor physical spaces and way of life. Less parked cars. Less parking lots. Much greater safety for people in and out of cars. Less noise pollution. More space”” / X https://x.com/karpathy/status/1989078861800411219

Figure’s humanoid robot worked 11 months at BMW factory
Figure’s humanoid robot helped produce over 30,000 BMW vehicles during an 11-month deployment, working full 10-hour factory shifts and handling 90,000+ parts. This represents one of the first sustained commercial deployments of humanoid robots in automotive manufacturing, demonstrating their viability for repetitive industrial tasks beyond the experimental phase.

Figure has shared numbers on its 11-month humanoid deployment at BMW’s Spartanburg factory. – Contributed to the production of 30,000+ cars (X3 vehicles). – 90,000+ parts loaded. – Ran 10-hour shifts, Monday to Friday. – Estimated 200+ miles of walking. – A single Figure 02 https://x.com/TheHumanoidHub/status/1991205599846269220

Luma AI raises $900M to build massive 2-gigawatt computing cluster
The startup plans to use the unprecedented computing power to develop AI systems that can understand and simulate physical reality, marking a shift from text-based AI toward systems that could manipulate the real world. The 2-gigawatt cluster would rival the power consumption of small cities, indicating the enormous energy requirements for next-generation AI development.

For AI to be able to help humans in the physical world, we need systems that can understand and simulate the universe. To exponentially accelerate Luma’s path to Multimodal AGI we are building a 2GW compute cluster with Humain and we have raised a $900M Series C. I am incredibly”” / X https://x.com/gravicle/status/1991202746871988680

Profluent’s JAM-2 AI generates drug-quality antibodies directly from computer code
The model achieved picomolar binding strength—matching pharmaceutical standards—for half of 26 tested targets, potentially eliminating years of lab work typically required to develop therapeutic antibodies. This represents a breakthrough in computational drug discovery, as previous AI models couldn’t reliably produce antibodies with the precise binding characteristics needed for actual medicines.

Today we’re thrilled to announce JAM-2 — the first AI model capable of generating drug-quality antibodies straight from the computer, with industry-leading success rates. > Drug-like affinities: Picomolar to single-digit nanomolar antibody binders for half of 26 targets while https://x.com/nablabio/status/1991154231026254181?s=20

Scientists decode vowel-like patterns in sperm whale communication
Researchers from the CETI project identified structured vowel and diphthong sounds in sperm whale calls, suggesting these marine mammals may use more complex linguistic patterns than previously understood. This breakthrough could fundamentally change how we view animal intelligence and communication, potentially revealing the first non-human language system with grammar-like rules.

There seems to be increasing progress in understanding whether whales have decipherable language.”” / X https://x.com/emollick/status/1989750285879656459

CETI scientists, led by CETI’s Linguistics Lead, Gašper Beguš, have discovered vowel and diphthong-like patterns in sperm whale communication! Read the paper here: https://x.com/ProjectCETI/status/1988627509198356848

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading