Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A symmetrical Byzantine mosaic icon of a six-winged golden seraph centered in a vast gold-tessera apse, each hammered wing tipped with a different sensory relic — an eye, an ear-trumpet, a lyre, an unfurled scroll, a painter’s brush, a speaking mouth — encircled by a concentric halo of clockwork circuit-rings, lit by warm candle-glow with deep imperial purple and Tyrian crimson borders, the title ‘MULTIMODALITY’ rendered in bold ivory-gold Trajan capitals across the lower third, painterly tactile surface grain, 16:9 full-bleed.
ByteDance just open-sourced one of the most capable multimodal models out there. BAGEL does image generation, editing, style transfer, and visual understanding – all in a single 7B parameter model. Apache 2.0 licensed! One model. No switching between specialized tools. Amazing
https://x.com/kimmonismus/status/2060050186076815792
I think people don’t realize why Gemini Omni is different than other video AIs. It is fully multimodal, so it can edit video natively, too I took the famous “”train “” movie from 1896 & made it a bullet train, LEGO, added a time traveler, a centipede, muppets… (see reflections?)
https://x.com/emollick/status/2057874739817808223
Omni is pretty nuts. It is NOT seedance. Any input in/out. It’s more than nano banana for video – it’s quite literally industrial light & magic. Now effectively reduced to an insanely realistic AR filter that you can apply on demand to anyone’s footage.
https://x.com/bilawalsidhu/status/2057300479340695960
Google just turned Street View into a video game. The mother lode of ground level data — 280 billion real world panoramas, now playable in real time. Here’s everything you need to know in 7 mins: 00:00 Genie 3 Grounded In Reality 00:44 Real-time Demos! 03:48 The Bigger
https://x.com/bilawalsidhu/status/2057262850209419553
Project Genie 🤝 @GoogleMaps Street View You can now take real U.S. places and transform them into new, interactive worlds. 🌍
https://x.com/GoogleDeepMind/status/2057842131142590512
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
https://research.nvidia.com/labs/lpr/locate-anything/
Rice and Baylor join BrainGate to develop brain-computer interfaces for people with paralysis
https://www.news-medical.net/news/20260528/Rice-and-Baylor-join-BrainGate-to-develop-brain-computer-interfaces-for-people-with-paralysis.aspx
Big milestone today. But still very early for self-driving software. When I was working on Autopilot at Tesla, we had an internal infra project called “”Operation Vacation”” – the goal being that the entire Autopilot engineering team could take a vacation and come back to a better
https://x.com/russelljkaplan/status/2059696925096776113
Handover event of last Tesla Model S & X: Franz: Next time you come back and stand in this place, it will be full of robots. Elon: I guess next time, we’ll be handing over the first Optimus… well, Optimus will just hand itself over. Hrushikesh: Fremont built the most cars of
https://x.com/TheHumanoidHub/status/2057907084411519278
Tesla’s Fremont line for the planned 1 million unit/year Optimus. At first glance it looks like a render, but those people walking around seem real.
https://x.com/TheHumanoidHub/status/2057930556374200479
Sonic 3.5 is now the #1 text to speech model on the @ArtificialAnlys leaderboard! You no longer have to trade off quality and latency – Sonic 3.5 also has the fastest time to first audio at 82ms end to end. See full benchmark results 👇
https://x.com/cartesia/status/2057880195403800633
Granola — The AI Notepad for back-to-back meetings
https://www.granola.ai/?via=adops-tldr-tech&dub_id=itx1yUOnpPaKx3A8
Gave google omni a sketched camera path and asked it to generate drone POV footage.
https://x.com/bilawalsidhu/status/2059419767417487718
Google just revealed Omni, personalized cross-device intelligence, and Spark agents at I/O 2025. I sat down with CEO Sundar Pichai to figure out what comes next: 1:46 Omni: “”Nano Banana for video”” 4:59 The future of YouTube 7:04 Advice for AI skeptics 9:33 Why your mom should
https://x.com/rowancheung/status/2057491344697012384
Nano Banana for video is here 🍌🎥 Gemini Omni is our new AI model that makes creating and editing videos as easy as having a conversation. Here’s how it works ↓
https://x.com/Google/status/2057881884219035752
We came across a really interesting tool that fixes a pretty common problem inside AI video workflows: extending AI cinematic scenes with seamless continuity. Using Omni models, you can take the last frame from an existing video clip and prompt something like: “show me what
https://x.com/CuriousRefuge/status/2057920807389806699
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini 🚀 Today, we’re sharing the @GoogleDeepMind white paper for GE 2, our first native multimodal embedding model. Whether it’s text, audio, video, or image, GE 2 provides a unified representation of the input.
https://x.com/mseyed/status/2059504005387284629
Google has the only true Omni model, but the elements aren’t hooked up. It appears it can take in & output audio, images. video, songs, text, code, etc. But right now each type of output is separate. When you can access the model directly, blending modes, a lot becomes possible.
https://x.com/emollick/status/2059774997535584325
Improving AI labels for viewers and creators – YouTube Blog
https://blog.youtube/news-and-events/improving-ai-labels-viewers-creators/
🎉 Congrats to @StepFun_ai on releasing Step-3.7-Flash, with day-0 support in vLLM. – 198B sparse MoE vision-language model, ~11B active params per token, native image + text input – 256K context window for long docs, multi-file repos, and dense visual interfaces – FP8 and NVFP4
https://x.com/vllm_project/status/2060155953715323288
MAI-Image-2.5 launches at No. 3 on Arena | Microsoft AI
https://microsoft.ai/news/mai-image-2-5-launches-at-no-3-on-arena-ai/
Meet MAI-Image-2.5 – ranked third on the @arena text-to-image leaderboard. It’s another great advance in quality. And with Build just a week away, there is much more to come. Learn more here:
https://x.com/MicrosoftAI/status/2059344061358563838
Meet MAI-Image-2.5 – ranked third on the @arena text-to-image leaderboard. It’s another great advance in quality. And with Build just a week away, there’s much more to come from the @MicrosoftAI team. I can’t wait.
https://x.com/mustafasuleyman/status/2059346031167570299
You can now transcribe meetings in real time using Codex and ask Codex questions about meetings as they’re happening! I updated my new Codex Meeting Recorder skill to use GPT Realtime Whisper. Tell Codex to use the skill, and it will start transcription and show it in the
https://x.com/_simonsmith/status/2059626873479422250
Exciting news, MAI-Image-2.5 (Preview) from @MicrosoftAI debuts at #3 in the Text-to-Image Arena with a score of 1,254 — a +72 point improvement over MAI-Image-2. A top 5 arena previously held only by @GoogleDeepMind and @OpenAI has a new lab in the mix. Congrats to the
https://x.com/arena/status/2059346024632820146
Elon Musk’s Neuralink reveals new robotic system for BCI brain implant surgeries | MobiHealthNews
https://www.mobihealthnews.com/news/elon-musks-neuralink-reveals-new-robotic-system-bci-brain-implant-surgeries





Leave a Reply