Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Photorealistic wide shot of six freestanding Ionic limestone columns on a sunny campus quad with a complete classical entablature spanning the top, the word VIDEO carved in large Roman serif letters in the center architrave, and a detailed bas-relief frieze below showing a continuous film reel with frame perforations and sequential motion imagery carved into the limestone, late afternoon golden hour lighting, red brick buildings and green lawn in background, crisp architectural photography.

Okay coolest Adobe Sneaks of the year award goes to Project Light Touch. Adobe calls this “”spatial lighting mode”” — interactively move your light source around within a 3D volume and voila — your image is accurately relit. They’re probably using ML to infer a PBR + depth map https://x.com/bilawalsidhu/status/1983982560054296843

NVIDIA World Simulation with Video Foundation Models for Physical AI https://huggingface.co/papers/2511.00062

Meta, Google, Apple – they’re all building AI replicas that capture your face, expressions, movements, personality. This goes way beyond Face ID. They’re basically creating a version of you that knows you better than you know yourself. The fidelity is remarkable too. We went https://x.com/bilawalsidhu/status/1985398951407722901

Sora: “Tiktok style high energy video explainer about the spinning columns of penguins in the sky. The pillar has always been there.” We live in a time of strange wonders (not the penguin pillar. That has always been there) https://x.com/emollick/status/1983718548326195236

Coca-Cola | Holidays are Coming, Fantastical :90 – YouTube https://www.youtube.com/watch?v=eoXX905YK6M

Veo 3.1 now has a Camera Adjustment Feature, allowing you to change the angle and movement of a previously generated video. Taking it out for a test spin, here’s our “”Test”” video, in the thread we’ll check out how the feature does! https://x.com/TheoMediaAI/status/1986104791454388289

Sora: “”that infamous dramatic Oscar winning scene where the lead keeps getting hit by the boom mic but nobody notices”” https://x.com/emollick/status/1985923709786603845

Chaining ffmpeg with a Browser Agent https://100x.bot/a/chaining-ffmpeg-with-browser-agent

(32) PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & Editing – YouTube https://www.youtube.com/watch?v=4hFybgTk4kE

Instant Skinned Gaussian Avatars for Web, Mobile and VR Applications TL;DR: animates GS by leveraging parallel splat-wise processing to dynamically follow the underlying skinned mesh in real time while preserving high visual fidelity. https://x.com/Almorgand/status/1985377664526323886

PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & Editing https://antoniooroz.github.io/PercHead/

SAM 2++: Tracking Anything at Any Granularity”” TL;DR: unifies video tracking across masks, boxes & points; uses task-specific prompts, a unified decoder, and a task-adaptive memory to track at any granularity. Backed by a new large-scale dataset. https://x.com/Almorgand/status/1986112315050369103

ByteDance released BindWeave Subject-Consistent Video Generation via Cross-Modal Integration https://x.com/_akhaliq/status/1986058046876070109

AI-video to robot transfer Generate AI video with Google VEO -> reconstruct the 3D motion -> train sim-to-real https://x.com/TheHumanoidHub/status/1985802136568123421

Vidu Q2 launches at #8 on the Artificial Analysis Text to Video Leaderboard, surpassing standard Sora 2 and Wan 2.5! It is also one of the first models to support video generation with multiple reference images, enabling more controllable results by using multiple angles of the https://x.com/ArtificialAnlys/status/1985781760236630305

The Sora app is now available on Android in: Canada Japan Korea Taiwan Thailand US Vietnam https://x.com/soraofficialapp/status/1985766320194142540

Yes, this is the official Sora handle. https://x.com/soraofficialapp/status/1985849973830046152

Happy Halloween from Sora and the monsters of Monster Manor. Created using characters, now available in the Sora app. https://x.com/OpenAI/status/1984318204374892798

Veo 3.1’s limit of 8 seconds through Flow and Gemini is a disadvantage compared to Sora 2. Especially as the models get better at breaking up a 20 second or more generation into multiple scenes and shots on their own, seeing 8 second clips becomes the visual equivalent to “”delve”””” / X https://x.com/emollick/status/1985727061428678701

Adobe Delivers New AI Innovations, Assistants and Models Across Creative Cloud to Empower Creative Professionals https://news.adobe.com/news/2025/10/adobe-max-2025-creative-cloud

Introducing Canva’s Creative Operating System https://www.canva.com/newsroom/news/creative-operating-system/

Matrix Oscar Winner: “”It Was a Prototype Disguised as a Blockbuster”” I spoke to John Gaeta, the man who built bullet time. While Hollywood thought they were making a sci-fi trilogy, his team was running a research project in plain sight. We covered: – How they raided Berkeley https://x.com/bilawalsidhu/status/1984356297937133831

Qwen3-VL Accuracy Differences on Ollama vs MLX Video: https://x.com/andrejusb/status/1985612661447331981

Introducing Odyssey-2: instant, interactive AI video https://odyssey.ml/introducing-odyssey-2

The results are in. LTX-2 is now ranked #3 video model on @ArtificialAnlys Video Arena. No surprise here. LTX-2 delivers on quality and speed. Huge credit to the LTX team here at @lightricks for making it happen. https://x.com/LTXStudio/status/1986442720534016449

New, extremely challenging visual reasoning benchmark “”MIRA”” where current models fail…great resource to research reasoning with imgs/video🌌 https://x.com/Muennighoff/status/1986519726823211129

MotionStream Real-Time Video Generation with Interactive Motion Controls model runs in real time on a single NVIDIA H100 GPU (29 FPS, 0.4s Latency) https://x.com/_akhaliq/status/1986054085766750630

A fun use of Sora 2 is summoning what feel like half-remembered content because they draw on cliches and presumed training data. Prompt: “”That really weird twist from the 1990s British science fiction show”” https://x.com/emollick/status/1985730430113394919

we are launching the ability to buy extra gens in sora today. we are doing this for two main reasons: first, we have been quite amazed by how much our power users want to use sora, and the economics are currently completely unsustainable. we thought 30 free gens/day would be”” / X https://x.com/billpeeb/status/1984011952155455596

here’s the bf16 video for reference, close enough right? i guess this is due to the reduced quantization block size in NVFP4 (16) vs MXFP4 (32) so ideally the former preserves the details more but just a speculation. https://x.com/mrsiipa/status/1986123806357020865

Introducing Cambrian-S it’s a position, a dataset, a benchmark, and a model but above all, it represents our first steps toward exploring spatial supersensing in video. 🧶 https://x.com/sainingxie/status/1986685042332958925

One interesting potential side effect of the nature of AI generated music & video is that it may make good critics good at making stuff. These systems have high variance, so the best approach is making many songs/videos and selecting the best. Critics are practiced at selection”” / X https://x.com/emollick/status/1985139697849413849

We discuss their papers showing that model diffing is unexpectedly easy when fine-tuning in a narrow domain, and on finding and fixing flaws with crosscoders, a sparse autoencoder based approach Video: https://x.com/NeelNanda5/status/1986590670631674217

We present MotionStream — real-time, long-duration video generation that you can interactively control just by dragging your mouse. All videos here are raw, real-time screen captures without any post-processing. Model runs on a single H100 at 29 FPS and 0.4s latency. https://x.com/xxunhuang/status/1985806498811789738

you might ask — isn’t this just a data or scaling problem? partly, yes. that’s why we’re building the new Cambrian-S video MLLM family. we want to push the limits of the current paradigm. we think data and scaling are essential for supersensing (just not sufficient.) the core https://x.com/sainingxie/status/1986685054559342916

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading