Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A luminous Lutheran stained-glass panel shaped like a film reel unspooling across the composition, each glass frame depicting a sequential moment of a red-tailed hawk seizing an outdoor cat while a bald athletic minister in running gear watches from the central rose-window frame, rendered with bold black leaded cames, jewel-tone cobalt sapphire ruby amber and emerald glass backlit like a cathedral window, with a heavy blackletter banner reading VIDEO integrated into the base.

Big improvements in quality and capabilities from the Omni team, excited to see how we translate this to progress for embodied AGI
https://x.com/jparkerholder/status/2056789448554062232

By default Omni has a bit of a professional “”look””, but can also create “”normie”” videos with the right prompting.
https://x.com/shlomifruchter/status/2056858151987884087

Creating, remixing, and editing a video is easier than ever with Gemini Omni. It offers a fluid, conversational way to create and edit. Just upload a video from your camera roll and ask Gemini to make changes.
https://x.com/GeminiApp/status/2057159933934907825

Edit your own videos with Gemini Omni with just a conversation. 🎥 Prompt the changes you want to see to reimagine the action, change the point of view, or adjust the lighting over multiple turns. Every instruction builds on the last, so your characters stay consistent, the
https://x.com/Google/status/2056786888930062369

Gemini Omni can create anything from any input, starting with video. 🪄 This means you can combine images, audio, video and text as input and generate high-quality videos. Or use drawings to create in a way that matches your vision. #GoogleIO
https://x.com/Google/status/2056786781992071172

Gemini Omni combines an intuitive understanding of physics with Gemini’s real-world knowledge and reasoning. 🌐 Now, the stories and outputs you create won’t just look photorealistic — they’ll behave like the real world. #GoogleIO
https://x.com/Google/status/2056786589175677089

Gemini Omni Flash is rolling out starting today. Here’s where you can find it: 🔹 Today: Google AI Plus, Pro and Ultra subscribers globally in the @GeminiApp and @FlowbyGoogle . 🔹Rolling out starting this week, for no cost: @YouTube Shorts and the YouTube Create app.
https://x.com/Google/status/2056789307856462061

Gemini Omni Flash: > a recording from a capsule on the london eye, a jerky zoom into something in the distance and then refocusing (with a bit of back and forth) (no timestamp or dialog) Note the world knowledge of London’s landscape, and the way the video is gently moving like
https://x.com/fofrAI/status/2056789242274259242

Gemini Omni is a major leap in world understanding & multimodal editing! It can take photos, video & audio and build entirely new scenes. Over time it’ll be able to handle any input & any output – starting w/ video You can even give it your own videos & iterate on your ideas:
https://x.com/demishassabis/status/2056831486251380783

Gemini Omni is coming to the Gemini app for paid subscribers today. It lets you bring your ideas to life using any combination of text, images, and video inputs. Just open up Gemini, attach a video from your camera roll, and change it around. It’s that simple. #GoogleIO
https://x.com/GeminiApp/status/2056800579159216202

Gemini Omni is our new AI model that can create anything from any input, starting with video. 🪄 Hear from @DemisHassabis on how you can mix text, audio, and images to generate and edit high-quality videos just by having a conversation. #GoogleIO
https://x.com/Google/status/2057180052979409172

Gemini Omni is so fun – insanely great at editing videos!
https://x.com/joshwoodward/status/2056827449556845051

Gemini Omni is wild
https://x.com/osanseviero/status/2056863263305105424

Gemini Omni was bigger news than Gemini 3.5 Flash
https://x.com/scaling01/status/2057143531622334678

Gemini Omni: “”a dramatic reading of Death by Water from the Wasteland by a man eating garlic bread while balanced on a unicycle on a small platform over a churning sea of tomato sauce in which, at the center, sites a meatball with bright blue eyes wearing a top hat””
https://x.com/emollick/status/2056791733619376315

I had early Gemini Omni access: “”sea otter in a pilot’s uniform explains why Spirit Airlines went bankrupt to a river otter who is distracted by their laptop while they are in a hot air balloon over NYC. in the next balloon over, william shakespeare fights a robot made of pizza””
https://x.com/emollick/status/2056788122369712148

Introducing Gemini Omni 🔮…….. Omni is our new model that can create anything from any input — starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon!
https://x.com/OfficialLoganK/status/2056787874260164628

Meet Gemini Omni — our new AI model that can create anything from any input, starting with video. 🪄 #GoogleIO
https://x.com/Google/status/2056786395067552140

Nano Banana for video is here! Google has long touted that Gemini is natively multi-modal in & out — but Omni is the first glimpse into the power of that paradigm applied to creation. Toss in a video and do multi-turn edits. Toss in audio and get reactive visuals. It’s kinda
https://x.com/bilawalsidhu/status/2056790381514076449

Starting today, Gemini Omni is rolling out to all Google AI Plus, Pro and Ultra subscribers globally at
https://t.co/382WL5xSvc and in the app. Soon, we will support more output formats, like image and audio. Give Gemini Omni a try and share your creations in the replies. 👇
https://x.com/GeminiApp/status/2056814117047132301

The Odyssey and the Iliad get so many movie treatments but the sequel, the Roman Aeneid, is entirely ignored. Here is a teaser trailer from one prompt to Gemini Omni. The first pass made all the flags Danish(?) but Omni is capable of editing video, so I asked for their removal.
https://x.com/emollick/status/2056855332127711387

The real „wow” moment is Gemini Omni. A world model towards AGI. It can create anything from any input. This is insane.
https://x.com/kimmonismus/status/2056802929957568881

Trying Gemini Omni on real life #google
https://x.com/TheTuringPost/status/2057167259877916679

We’re dropping Gemini Omni: our first step towards a model that can create anything from anything – starting with video. It combines Gemini’s intelligence with our generative media systems – representing a leap forward in world understanding, multimodality, and editing 🧵
https://x.com/GoogleDeepMind/status/2056786446636212467

A new paper from @ylecun, @NadavTimor and others: “”On Training in Imagination”” The main question of this study is: Given imperfect world models, imperfect reward models, noisy labels, and limited budget, how should we train most efficiently? The researchers analyze two sources
https://x.com/TheTuringPost/status/2056182805412098431

Big ass 3d gaussian splats in traditional GIS software makes me very happy
https://x.com/bilawalsidhu/status/2054955174649532680

Honestly reality capture is really fun. You should try it. Use your phone, drone, DSLR and now 360 camera and easily turn video footage into a photorealistic 3d model of the real world. Full playlist on NeRFs & Gaussian Splatting:
https://x.com/bilawalsidhu/status/2056503463526224116

I had to 3D reconstruct it!
https://x.com/Almorgand/status/2056681309854986515

Running around a 3d scan with a cartoon avatar. Photorealistic environment + stylized characters. The front flip absolutely makes this video lol
https://x.com/bilawalsidhu/status/2055722453431697834

Scaling laws hit 3D: 10B params, 70% memory save… A large-scale feed-forward model that performs 3D reconstruction on both static and dynamic scenes while exploring transfer of learned geometric representations. The work scales models to 10B parameters and training data to
https://x.com/IlirAliu_/status/2056645067285188974

Starchild-1: The First Real-Time Multimodal World Model
https://odyssey.ml/introducing-starchild-1

Static scans provide context but the world is dynamic. Provide RGB+D imagery and get articulated simulation ready 3d assets back. Am very bullish on this approach reaching a quality threshold for production 3d use cases.
https://x.com/bilawalsidhu/status/2056056016861528316

This thing just rips at 3d reconstruction of otherwise impossible scenarios — fpv flying through windows, people flying through the sky; I mean damn!
https://x.com/bilawalsidhu/status/2056715017702011162

UniCorrn: Unified Correspondence Transformer Across 2D and 3D”” TL;DR: one shared Transformer unifies 2D-2D, 2D-3D, and 3D-3D matching, outperforming specialized methods across correspondence tasks
https://x.com/Almorgand/status/2054984265746555351

谢赛宁团队+ Adobe research +ANU 的 RAEv2 出啦, 搞视觉重建/扩散模型/原生多模态/语义对齐/图像生成/世界模型的同学们值得关注: RAE 是非常重要的一个研究,改进了重建和引导生成,同时保留了Representation space的全局语义。这使得它能够取代VAE为理解和生成提供一个更优雅的 unified
https://x.com/recatm/status/2057456332861567359

Last year on 60 minutes we talked about the potential to ground our world models with Street View data. Thanks to hard work from @poolio and others this is now possible–Genie can simulate worlds based on real locations, and you can try it too!! Check out the blog for more info
https://x.com/jparkerholder/status/2056798252264018232

Real-world models are here! Stoked to share how we’re bringing real-world locations to life by integrating Street View into Genie. Try it now at
https://t.co/j6c1N38tRS and read the blog for more info:
https://x.com/poolio/status/2056796361987850705

That explains this startup.
https://x.com/emollick/status/2055829444397351110

For those saying “”the tomato sauce blood from the sword wound that flying Shakespeare inflicted on the pizza robot while the otters discussed Spirit Airlines wasn’t thick enough”” or whatever… this was state of the art in July 2025 (2 years) for “”an otter using wifi on a plane””
https://x.com/emollick/status/2056884715366310053

LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
https://jjihwan.github.io/projects/LiteFrame/

We’re adding new ways for people to identify AI-generated images and understand where they came from. In addition to C2PA Content Credentials, images now also contain a SynthID watermark, and can be identified using a public verification tool to check whether an image was made
https://x.com/OpenAI/status/2056793648571011232

Meet LongCat-Video-Avatar 1.5🐱–our upgraded, open-source digital human framework. Built for real production, not just short demos. What’s New: 🔹 Upgraded Audio Encoder: Replaces Wav2Vec2 with Whisper-Large, yielding significantly smoother and more natural lip dynamics. 🔹
https://x.com/Meituan_LongCat/status/2057494106889486646

Aleph 2.0 is here. Now you can edit a single frame in your video, preview the change and then Aleph 2.0 carries that edit across the rest of your video. Try it now in the new Edit Studio on web at the link below.
https://x.com/runwayml/status/2057530497597600169

AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation”” TL;DR: flow-map distillation enables video diffusion models that scale gracefully across arbitrary sampling step budgets while preserving high few-step generation quality
https://x.com/Almorgand/status/2055358506295799913

Today, we’re incredibly excited to release the new Edit Studio inside Runway. Edit Studio is a brand-new way to create, edit, and ship your content with Runway. We’re equally thrilled to release the long-awaited Aleph 2.0 as the first model to be experienced within Edit Studio.
https://x.com/iamneubert/status/2057535909524824226

I know the upcoming film version of the Odyssey is controversial, so I whipped together a completely accurate version that I think will be happily accepted by everyone as the most definitive version since Homer’s original, if not more so.
https://x.com/emollick/status/2056100170123644940

Runway started by helping filmmakers — now it wants to beat Google at AI | TechCrunch

Runway started by helping filmmakers — now it wants to beat Google at AI

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading