Image created with OpenAI GPT-Image-1. Image prompt: TikTok LIVE phone-screen POV, floating hearts & spinning album art, HorrorJump sudden flashlight flare and shaky-cam, featuring play-button PIP and timeline scrubber; soft-glow studio lighting, photoreal 8k

Ring Video Descriptions deliver real-time, Gen AI descriptions of what’s happening https://www.aboutamazon.com/news/devices/ring-video-descriptions-gen-ai

NO MORE SEO VideoPrism by @GoogleDeepMind is 🔥 it’s a versatile video encoder that can be plugged into text encoder or LLMs the authors first train a CLIP-like video-text model, then distill video encoder in masked manner to VideoPrism 😮 all models with A2.0 license on @huggingface 🤗 https://x.com/mervenoyann/status/1937572802896200181
https://research.google/blog/videoprism-a-foundational-visual-encoder-for-video-understanding/

Gemma 3n is out, with day-0 MLX support 👏 https://x.com/awnihannun/status/1938283694416077116

Introducing Gemma 3n: The developer guide – Google Developers Blog https://developers.googleblog.com/en/introducing-gemma-3n-developer-guide/

We’re fully releasing Gemma 3n, which brings powerful multimodal AI capabilities to edge devices. 🛠️ Here’s a snapshot of its innovations 🧵 https://x.com/GoogleDeepMind/status/1938278533517746686

We’ve taken community feedback very seriously, and that’s why for Gemma 3n launch we’re so proud to partner with so many in this amazing ecosystem Thanks to @huggingface, @ollama, @Prince_Canuma for MLX, @UnslothAI, @ggerganov llama.cpp/GGUFs, @NVIDIAAIDev, @kaggle,”” / X https://x.com/osanseviero/status/1938349897503412553

Veo 3 for Developers – Paige Bailey – YouTube https://www.youtube.com/watch?v=hlcAZ2lX_ZI

11/ TL;DR MJ has dropped a decent model with unlimited video gen for $60/month. if you’re all about the MJ aesthetics, doing abstract generations (sans text), and don’t mind using MMAudio or similar to add audio in post, this is a great addition to the toolkit. https://x.com/bilawalsidhu/status/1935528281031462945

MidJourney Video 1/ First off. It’s fun to just click through your MJ catalog and see it come to life. – No text to video; only image to video for now – Works with MJ or uploaded imagery – Can choose high/low motion + auto or custom prompt – Can extend clips 4x – SD output at 24 fps; no upscaling https://x.com/bilawalsidhu/status/1935527429768163481

MidJourney Video 4/ Pretty good with motion graphics and AR visuals. Lack of text rendition (a weakness for MJ in general) comes through. Would not recommend for titles. But you can still get some beautiful abstract visuals (e.g head locked AR gen on the left, multi-monitor generations on right) https://x.com/bilawalsidhu/status/1935527747725709668

MidJourney Video 6/ Muzzle flashes look pretty good. But I had a very hard time getting shell casings to work properly. Let me know if you find a good prompt to achieve this, because I couldn’t. https://x.com/bilawalsidhu/status/1935527929569755261

MidJourney Video 7/ MJ video nails that high end unreal engine “”rendered”” look. Fisheye lens distortion and sweeping camera move is nice. But notice how wonky all the cars in the scene look. As it stands, I don’t think we’ll be pulling any 3d objects or scenes out of this video model. https://x.com/bilawalsidhu/status/1935527993830686966

MidJourney Video 8/ Some generations get this weird “”unsharp mask”” look (a technique for sharpening) as the generation progresses. Lmk if you spot it too. https://x.com/bilawalsidhu/status/1935528060612395421

MidJourney Video 10/ MJ video does seem like an amazing tool to make abstract visual elements you composite elsewhere. I hope they remove the extend duration limit (20s / 4 times max) because it could be an amazing tool for screensavers, music videos and concert visuals. https://x.com/bilawalsidhu/status/1935528210588213670

MidJourney Video 2/ Fast generation time. Works well for that wide angle vlogging style. You can extend any clip 4x. Two great examples below. Of course, it’s begging for dialogue. Sure, you can add the facial performance in post – but it won’t look half as good. Veo 3 has spoiled us here. https://x.com/bilawalsidhu/status/1935527555404271877

MidJourney Video 3/ MJ video does okay in my handshake test (homie on the left really went in hard lmao) Physics is a weakness of this model — doesn’t matter if it’s soft-body or rigid-body subject matter. Might get slightly better as user ratings roll in, but still far behind the SOTA. https://x.com/bilawalsidhu/status/1935527672484179979

MidJourney Video 5/ The dinosaur test comes next. Movement looks decent, but the rest of the physics in the scene are all over the place. The slipping tanks in the background reminds me a bit of Sora. Relative scale and relative motion is pretty wonky. https://x.com/bilawalsidhu/status/1935527810564767970

MidJourney Video 9/ Testing fluid simulations here. Not only is it pretty far from SOTA, sometimes I get generations with this stop motion-like choppy FPS look (e.g. wine glass on right). https://x.com/bilawalsidhu/status/1935528134113407408

The new Hailou 02 AI video model really does seem to have made huge strides in the “”gymnastics problem”” where fast flipping motions lead to distortion Here are the first three results of the “”a man in elaborate robes does a backflip while holding two pool noodles”” (a hard test!) https://x.com/emollick/status/1936091679850705019

カプセルインタフェース:視聴覚や動き,力加減を伝え,全身リアル体験を実現 – CapsuleInterface: Full-Body Experience via Senses and Motion – YouTube https://www.youtube.com/watch?v=8a46Uap367k

Today we’re introducing you to the future of video. The world’s first Creative Operating System, we call it the HeyGen Video Agent. Upload a doc, some footage, or even just a sentence. It analyzes your input. Finds the story. Writes the script with taste. Selects the shots https://x.com/joshua_xu_/status/1938252187941122091

Hunyuan3D-2.1 passed my in-the-wild test 🤯 insanely good model! https://x.com/mervenoyann/status/1937161670444589215

3DGH: 3D Head Generation with Composable Hair and Face https://c-he.github.io/projects/3dgh/

Meta Held Deal Talks With Startup Runway in AI Recruiting Push – Bloomberg https://www.bloomberg.com/news/articles/2025-06-23/meta-held-deal-talks-with-startup-runway-in-ai-recruiting-push?embedded-checkout=true

Day 2/5 of #MiniMaxWeek: Introducing Hailuo 02, World-Class Quality, Record-Breaking Cost Efficiency 🎥 – Best-in-class instruction following – Handles extreme physics (yes, it does acrobatics 🤹) – Native 1080p https://x.com/MiniMax__AI/status/1935026724468871550

So here’s the app I created for the @lovable_dev Weekend AI showdown. It’s a video editor, fully built using Lovable + @AnthropicAI. Here is the link: https://x.com/Ahoo_Ahuu/status/1934624238071353400

Phew, it’s been a while. timm 1.0.16 released today to provide the image encoder for Gemma 3n. Additions kept on stacking so I haven’t had chance to finalize a release since the last one to support SigLIP-2 backbones. Lots of stuff in there: * Gemma 3n encoder (via a”” / X https://x.com/wightmanr/status/1938311403934519807

RT @GoogleDeepMind: We’re fully releasing Gemma 3n, which brings powerful multimodal AI capabilities to edge devices. 🛠️ Here’s a snapshot…”” / X https://x.com/slashML/status/1938394979727999455

RT @osanseviero: I’m so excited to announce Gemma 3n is here! 🎉 🔊Multimodal (text/audio/image/video) understanding 🤯Runs with as little as…”” / X https://x.com/algo_diver/status/1938374626910060782

RT @reach_vb: Google COOKED yet again – Multimodal Gemma3n 4B and 2B now available in Transformers, vLLM, MLX AND Llama.cpp 🤯 The model ca…”” / X https://x.com/reach_vb/status/1938476208330866751

RT @simonw: I’m really impressed by the new Gemma 3n I tried a 7.5GB model from Ollama and a 15GB model through mlx-vlm – they seem very c…”” / X https://x.com/osanseviero/status/1938581225452486911

RT @UnslothAI: Run Gemma 3n locally with our Dynamic GGUFs!✨ @Google’s Gemma 3n supports audio, vision, video & text and the 4B model fits…”” / X https://x.com/osanseviero/status/1938307534840074522

KerasHub lets you use any Hugging Face checkpoint for all top models like Llama, Gemma, Mistral, etc… Run your workflows in JAX, PyTorch, TensorFlow – inference, LoRA fine-tuning, large-scale training from scratch Blog post: https://x.com/fchollet/status/1938208330062655678

Inside Disney’s Campaign to Protect Darth Vader from AI – Bloomberg https://www.bloomberg.com/news/newsletters/2025-06-22/inside-disney-s-campaign-to-protect-darth-vader-from-ai

Researchers introduced STORM, a text-video model that trims video input to one-eighth the usual size yet still yields state-of-the-art scores. STORM inserts mamba layers between a SigLIP vision encoder and a Qwen2-VL language model: the mamba layers aggregate information across https://x.com/DeepLearningAI/status/1936438967391453522

Disney and Universal filed a lawsuit against image generation company Midjourney, accusing it of training its models on their copyrighted content and reproducing it without permission. The studios claim Midjourney system generated unauthorized images of characters like https://x.com/DeepLearningAI/status/1937314755066171580

The lifelike feel of @midjourney ‘s videos are in a class of their own. Some beautiful clips 👇 https://x.com/rohanpaul_ai/status/1936646300130308291

Most of the value in AI video won’t be captured by the creation tools. It’ll accrue to the platforms — X, YT, IG, etc. — where that content is distributed, ranked, and monetized. Can’t imagine it playing out any other way.”” / X https://x.com/bilawalsidhu/status/1935854180310434255

Vision-Language Models (VLMs) struggle with genuine deep comprehension of long-form video content because existing benchmarks often test shallow details or rely on low-quality, automatically generated questions. This paper introduces Movie Facts and Fibs (MF^2), a new benchmark https://x.com/rohanpaul_ai/status/1937466483275366761

Excited to share the release of VideoPrism! 🎥 📏Generate video embeddings 👀Useful for classifiers, video retrieval, and localization 🔧Adaptable for your tasks Model: https://x.com/osanseviero/status/1937560015348597124

ASMR Shiba has something to say 🐾 https://x.com/fdaudens/status/1937491856650277300

Introducing Vibemotion The first ever AI that turns a single prompt into stunning motion graphics and videos in minutes @vibemotion_ai https://x.com/adithya_s_k/status/1937530085109825630

Summer’s here, and the waves are calling! 🌊 With SurfSurf Effect, surf anytime, anywhere—no limits, just pure summer fun. 🌞 Don’t leave your fur friend behind — let’s hit the waves and ride into adventure! 🏄‍♂️💥 #surfsurf #klingeffects #klingai https://x.com/Kling_ai/status/1937393240225063042

I think the work we are doing at Runway is part of a new foundation for an entirely new media landscape. Just as the camera transformed how we capture reality, AI is transforming how we create it. The models and technical capabilities we’ve built are just the beginning – they’re”” / X https://x.com/c_valenzuelab/status/1937615643731272177

Midjourney’s new animation features continue to be compelling to play with because they really do let you make things that don’t feel like standard AI videos. Here I made some vast and strange machines. https://x.com/emollick/status/1935775887607447687

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading