Image created with Flux Pro v1.1 Ultra. Image prompt: Video, cinema clapboard built from bands of small bananas with a camera lens ringed by tiny bananas, crisp reflections, photorealistic, editorial, minimal, high detail, 3:2 landscape

USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning https://bytedance.github.io/USO/

If you think Apple is not doing much in AI, you’re getting blindsided by the chatbot hype and not paying enough attention! They just released FastVLM and MobileCLIP2 on Huggingface. The models are up to 85x faster and 3.4x smaller than previous work, enabling real-time vision language model (VLM) applications! It can even do live video captioning 100% locally in your browser 🤯🤯🤯 https://x.com/ClementDelangue/status/1962526559115358645

🗺️ ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling”” TL;DR: high-fidelity 3D humans across a wide range of poses, capturing both skeletal structure and surface details; separates internal skeleton from the external surface, (1/3) https://x.com/Almorgand/status/1962581481055797586

We connect the autoregressive pipeline of LLMs with streaming video perception. Introducing AUSM: Autoregressive Universal Video Segmentation Model. A step toward unified, scalable video perception — inspired by how LLMs unified NLP. 📝 https://x.com/miran_heo/status/1962649613590302776

Notebook LM Rolling out NEW audio overview formats:
(Default) Deep Dive: a thorough examination of your sources
Brief: 1-2 minute, bite-sized overviews
Critique: an expert review, offering constructive feedback on your material
Debate: a thoughtful debate between two hosts https://x.com/NotebookLM/status/1962949985546187120

Why Runway is eyeing the robotics industry for future revenue growth | TechCrunch https://techcrunch.com/2025/09/01/why-runway-is-eyeing-the-robotics-industry-for-future-revenue-growth/

Nano Banana + Veo 3 https://x.com/dev_valladares/status/1961621010144247858

How do we generate videos on the scale of minutes, without drifting or forgetting about the historical context? We introduce Mixture of Contexts. Every minute-long video below is the direct output of our model in a single pass, with no post-processing, stitching, or editing. 1/4 https://x.com/GordonWetzstein/status/1963583050744250879

Finally…an AI video editor that just works!! Edit any videos or cut the best moments directly from YouTube link from just a simple English prompt. This is insane! https://x.com/Saboo_Shubham_/status/1962891766232739919

People ask how I get such clean 3D scans with a DSLR — and I must admit it’s a bit of a dark art. But this new $5K PortalCam changes everything. LiDAR precision + SLAM speed + 3D Gaussian Splat fidelity in a device anyone can use. The use cases are wild. Let me show you 🧵 https://x.com/bilawalsidhu/status/1963337887027707987

90’s SGI computers were a vibe and half too. Reality Engine systems were like $250-750K and I wanted one so bad as a kid! These beautiful beasts powered VFX in everything from Jurassic Park and Terminator 2 to The Matrix and Lord of the Rings. https://x.com/bilawalsidhu/status/1962755877481349170

Pixie: Physics from Pixels”” TL;DR: NeRF, GS w/ physics; neural network mapping pretrained visual features (i.e., CLIP) to dense material fields of physical properties in a single forward pass, enabling real‑time physics simulations. https://x.com/Almorgand/status/1961076683093524561

Lipsync Studio https://higgsfield.ai/create/speech

I wrote about the era of Mass Intelligence. GPT-5 and Google’s Nano Banana are examples of how advanced AI is now making their way to far more users, at scale, as both performance and efficiency keep improving. We are going to see a lot of weird things happening, all at once. https://x.com/emollick/status/1961169796491329653

TikTok owner ByteDance sets valuation at over $330 billion in planned buyback – The Japan Times https://www.japantimes.co.jp/business/2025/08/28/tech/tiktok-bytedance-valuation-330-billion/

HunyuanWorld-Voyager is here and fully open-source! The world’s first ultra-long-range world model with native 3D reconstruction, redefining AI-driven spatial intelligence for VR, gaming, and simulations. ✅Direct 3D Output: Exports point cloud videos to 3D formats without tools https://x.com/TencentHunyuan/status/1962741518797836708

Entire startups have raised more venture capital on the backs of Adobe video edits than actual products. Insane if you think about it. After Effects might be the most valuable VC fundraising tool ever invented.”” / X https://x.com/bilawalsidhu/status/1962915517326332086

The utter disrespect for CGI & VFX continues to baffle me. Do these Hollywood heavyweights make these remarks because it’s an easy PR win? Or do they genuinely think modern cinema and TV would be anywhere close to where it is without CGI assisted storytelling?”” / X https://x.com/bilawalsidhu/status/1962583444158062641

Netflix House Philadelphia Opens Nov. 12; the Dallas Location Arrives Dec. 11; How to Buy Tickets – Netflix Tudum https://www.netflix.com/tudum/articles/netflix-house

MiniCPM-V 4.5 achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro, and strong open-source models like Qwen2.5-VL 72B powered https://x.com/_akhaliq/status/1963587749400727980

California tech startup once worth $1 billion shuts down https://www.sfgate.com/tech/article/flip-startup-shuts-down-billion-21022297.php

Filmmaker Henry Daubrez joins Google Labs team to work on Flow https://blog.google/technology/google-labs/flow-resident-filmmaker/

Real-time AI video is more for content creation than gaming in the short term. This is a crude example of performance capture. Once we pipe tracked camera poses from your iPhone directly to the model, every creator gets an interactive movie set. https://x.com/bilawalsidhu/status/1961159080761831475

🚀Introducing Wan2.2-S2V — a 14B parameter model designed for film-grade, audio-driven human animation. 🎬Going beyond basic talking heads to deliver professional-level quality for film, TV, and digital content. And it’s open-source! ✨ Key features: 🔹 Long-video dynamic https://x.com/Alibaba_Wan/status/1960350593660367303

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading