Image created with Ideogram 3.0. Image prompt: Lower-East-Side street-corner photograph reminiscent of a late-80s album cover: weathered red-brick tenement with exterior fire-escapes, canvas awning shading racks of vintage clothes; above the awning, a hand-painted board reads ‘Multimodality SPORTSWEAR’; a hanging blade sign in cursive script reads ‘Multimodality Boutique’; a collage of tapes, polaroids, and comics labeled ‘Multimodality’ covers a brick panel; warm golden-hour light, subtle 35mm film grain, muted yet punchy color palette, gritty NYC vibe.
Gemini Live camera and screen sharing in @GeminiApp is available on @Android and rolling out to iOS, starting today. https://x.com/Google/status/1924876301573239061
Google is bringing real-time AI camera sharing to Search | The Verge https://www.theverge.com/news/670597/google-search-live-ai-mode-gemini-ios
Gemini 2.5 can now organize vast amounts of multimodal information, reason about everything it sees, and write code to simulate anything. ↓ #GoogleIO https://x.com/GoogleDeepMind/status/1924878250255516126
Last year, we introduced Project Astra: a research prototype exploring capabilities for a universal AI assistant. 🤝 We’ve been making it even better with improved voice output, memory and computer control – so it can be more personalized and proactive. Take a look ↓ #GoogleIO https://x.com/GoogleDeepMind/status/1924883244459425797
Check out Veo 3 🔥🔥🔥 sound on 🔊”” / X https://x.com/_tim_brooks/status/1924895946967810234
From capturing real-world physics – like the noise and movement of water, or the look and sound of walking in snow – to lip syncing, Veo 3 is great at understanding what you want. You can tell a short story in your prompt, and the model gives you back a clip that brings it to https://x.com/GoogleDeepMind/status/1924893531300077675
Say goodbye to the silent era of video generation: Introducing Veo 3 — with native audio generation. 🗣️ Quality is up from Veo 2, and now you can add dialogue between characters, sound effects and background noise. Veo 3 is available now in the @GeminiApp for Google AI Ultra https://x.com/Google/status/1924893837295546851
Veo 3 is available today for Ultra subscribers in the United States in the @GeminiApp. Find out more about where you can use it ↓ https://x.com/GoogleDeepMind/status/1924893533787332996
Veo 3, our SOTA video generation model, has native audio generation and is absolutely mindblowing. For filmmakers + creatives, we’re combining the best of Veo, Imagen and Gemini into a new filmmaking tool called Flow. Ready today for Google AI Pro and Ultra plan subscribers. https://x.com/sundarpichai/status/1924909490081825195
Veo 3: “”a big broadway musical about garlic bread, with elaborate costumes and a sondheim-like vibe”” https://x.com/emollick/status/1925065546082484418
Veo 3: “”a scene from an unnerving 1970s childrens show with live action puppets and Lovecraftian overtones singing a song”” https://x.com/emollick/status/1925047195738218505
Video, meet audio. 🎥🤝🔊 With Veo 3, our new state-of-the-art generative video model, you can add soundtracks to clips you make. Create talking characters, include sound effects, and more while developing videos in a range of cinematic styles. 🧵 https://x.com/GoogleDeepMind/status/1924893528062140417
AI learns how vision and sound are connected, without human intervention | MIT News | Massachusetts Institute of Technology https://news.mit.edu/2025/ai-learns-how-vision-and-sound-are-connected-without-human-intervention-0522
UAE launches Arabic language AI model as Gulf race gathers pace | Reuters https://www.reuters.com/world/middle-east/uae-launches-arabic-language-ai-model-gulf-race-gathers-pace-2025-05-21/
Elon on Optimus in today’s CNBC interview ⦿ Currently, Optimus is being trained from demonstrations collected by humans wearing mocap suits with cameras on their heads – performing primitive tasks such as opening doors, picking up objects, and dancing. This is needed to https://x.com/TheHumanoidHub/status/1924981814311133509
Optimus can now learn from first-person video. Many new skills are emerging that can be instructed via natural language. Next step: expand this to learning from third-person videos (random internet videos) and push reliability via self-play (Reinforcement learning). https://x.com/TheHumanoidHub/status/1925057174092579253
Exclusive: Google Sees Smart Glasses as the ‘Next Frontier’ for AI. And It’s Not Working Alone – CNET https://www.cnet.com/tech/computing/exclusive-google-sees-xr-smart-glasses-as-the-ultimate-use-for-ai-with-warby-parker-samsung-and-xreal-on-deck/#ftag=CAD590a51e
Releasing the OpenAI to Z Challenge — using o3/o4 mini and GPT 4.1 models to discover previously unknown archaeological sites:”” / X https://x.com/gdb/status/1923105670464782516
Stability AI open-sourced Stable Audio Open Small, a text-to-audio AI —341M-parameters —Generates 11s of audio, including drum loops, foley, riffs, and textures —Optimized for Arm-based consumer devices https://x.com/adcock_brett/status/1924133939376996539
Today we’re open-sourcing Stable Audio Open Small, a 341M-parameter text-to-audio model optimized to run entirely on @Arm CPUs. This means 99% of smartphones can now generate music-production samples in seconds, right on-device with no internet required. Built for fast, https://x.com/StabilityAI/status/1922675163411497094
Enterprise Document AI & OCR | Mistral AI https://mistral.ai/solutions/document-ai
o3, show me a photo of the most stereotypical X and LinkedIn feeds as seen on a mobile device. Really lean into it.”” https://x.com/emollick/status/1922866235307356423
Salesforce just dropped BLIP3-o on Hugging Face A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset https://x.com/_akhaliq/status/1923001183804764391
tis the year of any-to-any/omni models BAGEL by @BytedanceTalk 7B native multimodal model that understands and generates both image + text outperforms leading VLMs like Qwen 2.5-VL 👏 and has Apache 2.0 license 😱 https://x.com/mervenoyann/status/1925218434964472067
Google Meet is getting real-time speech translation | TechCrunch https://techcrunch.com/2025/05/20/google-meet-is-getting-real-time-speech-translation/
Gemma keeps delivering! I’m very excited to share with you the most recent LMArena results for the Gemma 3 family💥 Gemma 3 stays as the best open model that can run on a single GPU. And stay tuned, more to come! https://x.com/osanseviero/status/1923159046900548014
Google Flow is the closest thing I’ve seen to a multimodal AI studio for creatives. And it’s available today. It feels like a generative camera and soundstage, where you can “capture” all the shots that you need — and feel confident you have everything to put it together in https://x.com/bilawalsidhu/status/1924901783664787942
Announcing Veo 3, Imagen 4, and Lyria 2 on Vertex AI | Google Cloud Blog https://cloud.google.com/blog/products/ai-machine-learning/announcing-veo-3-imagen-4-and-lyria-2-on-vertex-ai
I really want to try out Veo3. Really really bad. But the reality is, I will probably only run a dozen or so tests through it and move on, so I cannot justify a subscription of this amount. I have visited this screen a dozen or so times the past few days. https://x.com/ostrisai/status/1925917357731410313
It’s official — Veo 3 and Imagen 4 is here and available starting today. > Veo 3 is not only higher quality video with support for subject and style references, but it can *natively* generate audio (sound effects, music AND dialogue!) > Imagen 4 similarly now crushes it at https://x.com/bilawalsidhu/status/1924897257629089855
New Google AI Ultra subscription tier will give you access to Gemini 2.5 Pro Deep Think, Veo 3 and Project Mariner”” / X https://x.com/scaling01/status/1924891236109799838
non-human intelligence comes in peace the dialogue, lip movement, environmental audio — all perfectly synced — all from one prompt what should i prompt with google veo 3 next? https://x.com/bilawalsidhu/status/1924931758556082677
This matches what I am seeing, the model is a huge leap in video creation and is good at direction following, but the two most common failure modes are that it adds nonsense “”subtitles”” to videos and that some videos lack sound. Wouldn’t be a big deal but credits are not returned”” / X https://x.com/emollick/status/1925305547651190787
We’re excited to shape the future of Flow with AI filmmakers like @HenryDaubrez who used it to create Electric Pink: a short video exploring a pink-haired superhero crafting his dream adventure using his childhood inspirations. 📽️✨↓ https://x.com/GoogleDeepMind/status/1924896549248594225
Gemini Diffusion: diffusion-based LLM, much faster than autoregressive LLMs Gemini 2.5 Pro Deep Think: doubles o3’s score on 2025 USAMO math competition Imagen 4: can spell Veo 3: native audio generation, characters can speak Google is back. Artificial Pichai Intelligence.”” / X https://x.com/Yuchenj_UW/status/1924896740068753825
Multimodal model support is here in 0.7! Ollama now supports multimodal models via its new engine. Cool vision models to try👇 – Llama 4 Scout & Maverick – Gemma 3 – Qwen 2.5 VL – Mistral Small 3.1 and more 😍 Blog post 🧵👇 https://x.com/ollama/status/1923139667563528347
Meet Document AI, our end-to-end document processing solution powered by the world’s best OCR model! https://x.com/MistralAI/status/1925577532595696116
Everyone has this👇 technology that we were all was blown away by a year ago, for free. I even screenshotted the problem from the video to use in this demo to test it out.. Yet a weakness is the GPT-4o vision system, which still has trouble with details like where my line was. https://x.com/emollick/status/1924304791636725855
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset Author’s Explanation: https://x.com/TheAITimeline/status/1924232118755824119
Emerging Properties in Unified Multimodal Pretraining https://arxiv.org/pdf/2505.14683
MMLongBench Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly https://x.com/_akhaliq/status/1924477810430624017
New video of AGIBOT’s Lingxi X2 robot showcasing dynamic movements, fall recovery, quiet operation, and vision-based perception/planning. https://x.com/TheHumanoidHub/status/1923259033160384676
X2C: A Dataset Featuring Nuanced Facial Expressions
for Realistic Humanoid Imitation https://lipzh5.github.io/X2CNet/
Elon, November 2023: “Optimus will figure out how to do things by watching videos.” https://x.com/TheHumanoidHub/status/1925067964027699466
Seeing blood clots before they strike | The University of Tokyo https://www.u-tokyo.ac.jp/focus/en/press/z0508_00409.html
CGS-GAN 3D Consistent Gaussian Splatting GANs for High Resolution Human Head Synthesishttps://fraunhoferhhi.github.io/cgs-gan/
DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
https://ziqiaopeng.github.io/dualtalk/
Grok can now generate charts. Currently works in browser, will come to more platforms in the coming days. https://x.com/512×512/status/1923893565664657418




