Image created with Ideogram 3.0. Image prompt: Lower-East-Side street-corner photograph reminiscent of a late-80s album cover: weathered red-brick tenement with exterior fire-escapes, canvas awning shading racks of vintage clothes; above the awning, a hand-painted board reads ‘Video SPORTSWEAR’; a hanging blade sign in cursive script reads ‘Video Boutique’; a stack of VHS tapes titled ‘Video Classics’ leans against a stereo; warm golden-hour light, subtle 35mm film grain, muted yet punchy color palette, gritty NYC vibe.

Gemini Live camera and screen sharing in @GeminiApp is available on @Android and rolling out to iOS, starting today. https://x.com/Google/status/1924876301573239061

Google is bringing real-time AI camera sharing to Search | The Verge https://www.theverge.com/news/670597/google-search-live-ai-mode-gemini-ios

In Flow, AI can help make clips from prompts, build them into scenes and then save your ingredients – such as characters, locations, objects or styles – all in one place. ↓ https://x.com/GoogleDeepMind/status/1924896542848090276

Check out Veo 3 🔥🔥🔥 sound on 🔊”” / X https://x.com/_tim_brooks/status/1924895946967810234

From capturing real-world physics – like the noise and movement of water, or the look and sound of walking in snow – to lip syncing, Veo 3 is great at understanding what you want. You can tell a short story in your prompt, and the model gives you back a clip that brings it to https://x.com/GoogleDeepMind/status/1924893531300077675

Say goodbye to the silent era of video generation: Introducing Veo 3 — with native audio generation. 🗣️ Quality is up from Veo 2, and now you can add dialogue between characters, sound effects and background noise. Veo 3 is available now in the @GeminiApp for Google AI Ultra https://x.com/Google/status/1924893837295546851

Veo 3 is available today for Ultra subscribers in the United States in the @GeminiApp. Find out more about where you can use it ↓ https://x.com/GoogleDeepMind/status/1924893533787332996

Veo 3, our SOTA video generation model, has native audio generation and is absolutely mindblowing. For filmmakers + creatives, we’re combining the best of Veo, Imagen and Gemini into a new filmmaking tool called Flow. Ready today for Google AI Pro and Ultra plan subscribers. https://x.com/sundarpichai/status/1924909490081825195

Veo 3: “”a big broadway musical about garlic bread, with elaborate costumes and a sondheim-like vibe”” https://x.com/emollick/status/1925065546082484418

Veo 3: “”a scene from an unnerving 1970s childrens show with live action puppets and Lovecraftian overtones singing a song”” https://x.com/emollick/status/1925047195738218505

Video, meet audio. 🎥🤝🔊 With Veo 3, our new state-of-the-art generative video model, you can add soundtracks to clips you make. Create talking characters, include sound effects, and more while developing videos in a range of cinematic styles. 🧵 https://x.com/GoogleDeepMind/status/1924893528062140417

I thought my “”otter on a plane using Wifi”” benchmark was already done, but Veo 3 adds higher quality… and sound Here is “”an otter on a plane using wifi on their phone, the flight attendant asks them “”do you want a drink ?”” and the otter nods”” (One of the first set of 4 videos) https://x.com/emollick/status/1925018308182524391

Google Beam: Be there from anywhere with our breakthrough communication technology. https://starline.google/

Alibaba’s Wan dropped Wan2.1-VACE, a unified AI for video creation and editing Available in 1.3B, 14B sizes, the model can handle reference-to-video generation, video-to-video editing, and masked video-to-video editing Open-sourced under Apache 2.0 https://x.com/adcock_brett/status/1924133827095498952

Google is shipping 3D video conferencing tech this year. Project Starline is now Google Beam, with HP being the first OEM to bring it to market. It literally feels like a portal. Having tried this tech while I was at Google for regular meetings — lemme tell you your recall is https://x.com/bilawalsidhu/status/1924889726542348699

Hi Google Beam👋! What started as Project Starline has evolved into a revolutionary 3D video communications platform. The combination of our AI video model and light field display allows you to make eye contact and read subtle cues as if you were face-to-face 🤯 https://x.com/GoogleAI/status/1924880505146847454

What if you could take a 2D video call and make it feel like you’re really there? Google Beam, our new AI-first video communication platform, does just that — using a state-of-the-art AI video model to transform 2D video streams into a realistic 3D experience. #GoogleIO https://x.com/Google/status/1924875328037466302

Google presents LightLab Controlling Light Sources in Images with Diffusion Models https://x.com/_akhaliq/status/1923135902827642901

EVA: Expressive Virtual Avatars from Multi-view Videos https://vcai.mpi-inf.mpg.de/projects/EVA/

New sota open-source depth estimation: Marigold IID 🌼 > normal maps, depth maps of scenes & faces > get albedo (true color) and BRDF (texture) maps of scenes, they even release a depth-to-3D printer format demo 😮 link to all models and demos on the next one ⤵️ https://x.com/mervenoyann/status/1923318140965990814

🤖 From this week’s issue: Google introduced Veo 3 and Imagen 4, and a new tool for filmmaking called Flow. https://x.com/dl_weekly/status/1925904865164689539

Flow is available for Google AI Pro and Ultra plan subscribers in the US, with more countries coming soon. Try it here ↓ https://x.com/GoogleDeepMind/status/1924896551496716667

Get into the zone with Flow. 🎬 It combines the best of our most advanced models Veo, Imagen and Gemini into 1️⃣ master filmmaking tool – helping you weave cinematic clips, dynamic scenes, and compelling narratives into stories with consistent results. https://x.com/GoogleDeepMind/status/1924896540138586528

Introducing Flow: a new type of AI filmmaking tool that combines the best of Veo, Imagen and Gemini — built with and for creatives. Flow helps you maintain character and visual consistency from one clip to the next. See how emerging filmmakers are using it 🎥 https://x.com/Google/status/1924896843441336440

Google Flow is the closest thing I’ve seen to a multimodal AI studio for creatives. And it’s available today. It feels like a generative camera and soundstage, where you can “capture” all the shots that you need — and feel confident you have everything to put it together in https://x.com/bilawalsidhu/status/1924901783664787942

Real-time speech translation directly in Google Meet matches your tone and pattern so you can have free-flowing conversations across languages Launching now for subscribers. ¡Es mágico! https://x.com/sundarpichai/status/1924909694524805567

Announcing Veo 3, Imagen 4, and Lyria 2 on Vertex AI | Google Cloud Blog https://cloud.google.com/blog/products/ai-machine-learning/announcing-veo-3-imagen-4-and-lyria-2-on-vertex-ai

I really want to try out Veo3. Really really bad. But the reality is, I will probably only run a dozen or so tests through it and move on, so I cannot justify a subscription of this amount. I have visited this screen a dozen or so times the past few days. https://x.com/ostrisai/status/1925917357731410313

It’s official — Veo 3 and Imagen 4 is here and available starting today. > Veo 3 is not only higher quality video with support for subject and style references, but it can *natively* generate audio (sound effects, music AND dialogue!) > Imagen 4 similarly now crushes it at https://x.com/bilawalsidhu/status/1924897257629089855

New Google AI Ultra subscription tier will give you access to Gemini 2.5 Pro Deep Think, Veo 3 and Project Mariner”” / X https://x.com/scaling01/status/1924891236109799838

non-human intelligence comes in peace the dialogue, lip movement, environmental audio — all perfectly synced — all from one prompt what should i prompt with google veo 3 next? https://x.com/bilawalsidhu/status/1924931758556082677

This matches what I am seeing, the model is a huge leap in video creation and is good at direction following, but the two most common failure modes are that it adds nonsense “”subtitles”” to videos and that some videos lack sound. Wouldn’t be a big deal but credits are not returned”” / X https://x.com/emollick/status/1925305547651190787

We’re excited to shape the future of Flow with AI filmmakers like @HenryDaubrez who used it to create Electric Pink: a short video exploring a pink-haired superhero crafting his dream adventure using his childhood inspirations. 📽️✨↓ https://x.com/GoogleDeepMind/status/1924896549248594225

Gemini Diffusion: diffusion-based LLM, much faster than autoregressive LLMs Gemini 2.5 Pro Deep Think: doubles o3’s score on 2025 USAMO math competition Imagen 4: can spell Veo 3: native audio generation, characters can speak Google is back. Artificial Pichai Intelligence.”” / X https://x.com/Yuchenj_UW/status/1924896740068753825

Brandcast 2025: YouTube’s Upfronts highlights – YouTube Blog https://blog.youtube/news-and-events/brandcast-2025/

CGS-GAN 3D Consistent Gaussian Splatting GANs for High Resolution Human Head Synthesishttps://fraunhoferhhi.github.io/cgs-gan/

DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
https://ziqiaopeng.github.io/dualtalk/

Earlier this month, we released Gen-4 References, our most general and flexible image generation model yet. It became one of our most popular releases ever, with new use cases and workflows being discovered every minute. Today, we’re making it available via the Runway API, https://x.com/runwayml/status/1923440921464442887

Gen-4 References API is out! It’s time to build”” / X https://x.com/c_valenzuelab/status/1923441791665058270

Here is a new workflow for Gen-4 References: Element extraction and composition. 1) Extract the chair from image 1 and put it on a flat chroma green background color 2) Show me a front angle photo of the striped couch from image 1 on a flat chroma green background color 3) https://x.com/c_valenzuelab/status/1924596075568222654

dont you dare community note this this is my only coping mechanism left”” (pretty funny poke at Veo) / X https://x.com/nearcyan/status/1924963816359788640

It was the week of video generation at @huggingface, on top of many new LLMs, VLMs and more! Let’s have a wrap 🌯 LLMs 💬 > Alibaba Qwen released WorldPM-72B, new World Preference Model trained with 15M preference samples (OS) > II-Medical-8B, new LLM for medical reasoning that https://x.com/mervenoyann/status/1924430139242283172

✨ All in One, Wan for All✨ We are excited to introduce our latest model to our talented community creators: Wan2.1-VACE, All-in-One Video Creation and Editing model. Model size: 1.3B, 14B License: Apache-2.0 📌 Wan2.1-VACE provides solutions for various tasks, including https://x.com/Alibaba_Wan/status/1922655324919779604

Bilibili dropped AniSORA on Hugging Face – anime video generation model capable of making manga, tuber, mad-style parodies and more! – Apache 2.0 licensed! 🔥 https://x.com/reach_vb/status/1924425789774123316

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading