Image created with gemini-3.1-flash-image-preview with claude-sonnet-4-5. Image prompt: Using the provided reference image, maintain the exact left-third close-crop composition with deep blue-purple cinematic lighting and misty atmospheric smoke bleeding rightward, but replace the central figure with a film director slumped in their chair, face lit by cold monitor glow, glitter scattered across their shoulders like film grain, expression weighted with creative exhaustion, right two-thirds open hazy space with ‘video’ in thin lowercase white Helvetica Neue Light, same melancholic post-party emotional register and HBO prestige drama aesthetic.

People are asking what’s the difference between Falcon Perception and SAM3, so here’s my opinion: SAM3:
https://t.co/KVRbuHm8H1 Falcon Perception:
https://t.co/QDgMlOBvDH First, sam3 does “”promptable concept segmentation””: simple noun phrases (like “”yellow bus””, “”red apple””) +
https://x.com/dahou_yasser/status/2041474094252933195

Today we’re releasing WildDet3D–an open model for monocular 3D object detection in the wild. It works with text, clicks, or 2D boxes, and on zero-shot evals it nearly doubles the best prior scores. 🧵
https://x.com/allen_ai/status/2041545111151022094

kays on X: “I noticed there wasn’t anything like this out there, so I wrote a tiny visual blog for those wanting to introduce themselves to Dynamic Gaussian Splatting and their current methods 🖼️ Feel free to check out, these are some of the visuals taken from it https://t.co/6W2qx2yI1K” / X
https://x.com/pabloadaw/status/2041650303804555278

We’re excited to be rolling out two model updates today! Marble 1.1: Improves lighting and contrast, with a major reduction in visual artifacts. Marble 1.1-Plus: Our new model built for scale. Create larger, more complex environments than ever before.
https://x.com/theworldlabs/status/2041554646561677701

I showed you SAM 3 all week. This is a 0.6B model that outperforms it. Falcon Perception. Type “”detect the plane”” and it segments every plane in the frame. Pixel-accurate masks from natural language. Fighter jets. Fire. Crowds. All on a MacBook via MLX. No cloud.
https://x.com/MaziyarPanahi/status/2040776481673281936

agents that make explainer videos > agents that summarize PDFs
https://x.com/lucatac0/status/2041018088913608923

Jeanne on X: “I’ve combined Manim @NousResearch’s Hermes Agent skill + @yifan_zhang_’s Math Code. Math Code executes the proof on a problem called Jordan’s Lemma and Hermes Agent with @claudeai Sonnet 3.7 directs Math Code, writes a script, gets Manim to render an explanatory video. https://t.co/qOsmOpvPlS” / X
https://x.com/prompterminal/status/2040982307377381583

Nous Research on X: “Introducing the Manim skill for Hermes Agent. Manim is an engine for creating precise programmatic animations for mathematical and technical explainers, made famous by the @3blue1brown channel. https://t.co/nyNeNthhZB” / X
https://x.com/NousResearch/status/2040931043658567916

PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models”” TL;DR: diffusion pipeline for scalable generation of photorealistic human data with 3D annotations
https://x.com/Almorgand/status/2040096997470843366

StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision”” TL;DR: adapts a pretrained 3D-aware transformer to stereo vision with a training-free pipeline, achieving SOTA performance on KITTI
https://x.com/Almorgand/status/2041569246883332385

Subtle hand held camera shake will convince people a 3d game render is real life footage. No generative ai required.
https://x.com/bilawalsidhu/status/2041643400433201384

PoseDreamer: Scalable Photorealistic Human Data Generation with Diffusion Models
https://prosperolo.github.io/posedreamer/

We always need more visuals! Checkout this on for dynamic gaussian splatting
https://x.com/Almorgand/status/2041773431524302968

The next version of @OpenClaw comes with native video generation. To start, I added support for the following companies: – Alibaba – BytePlus – fal – Google – MiniMax – OpenAI – Qwen – Together – xAI
https://x.com/steipete/status/2040928953653744003

📝Summarize 0.13 is out! 🎞️ Local video slides (–slides) 🤖 More model backends (GitHub Copilot) 🧠 Better GPT-5.4 support 📺 Better media handling (HLS detection.m3u8) It graduated from my tap to official homebrew formula! 🍺 brew install summarize
https://x.com/steipete/status/2041669438882087180

Robots can now reconstruct 3D scenes in real time from a single RGB camera. [📍 Projects page + paper] No depth sensor. No retraining. 30 FPS. Researchers at the Imperial College London introduced KV-Tracker, a training-free method that makes heavy models like π³ and Depth
https://x.com/IlirAliu_/status/2041062366025031787

Researchers just taught a robot to play tennis. From just clips of a few amateur players performing basic forehands, backhands, and shuffles… …a robot learned one of the fastest, most coordinated physical skills there is. Insane!
https://x.com/rowancheung/status/2040085788256506190

Robotics pre-training *from scratch* has been a heretical idea for the last two years. That “there’s no internet of robotics data” has led to two prevailing conclusions: 1) we need to use pretrained model backbones and 2) we need to scale robotics data. The first conclusion in
https://x.com/xiao_ted/status/2041547335935853025

PixVerse C1 is live–our first model built for film production. Coherent action, storyboard-to-video, ref-guided consistency. 1080p, 15s, native audio. Available on PixVerse Web and API Platform. RT+Follow+Reply=300Creds(72H ONLY)
https://x.com/PixVerse_/status/2041536108660740162

Seedance 2.0 is now on Runway. Use text, image, video or audio as inputs to generate stunning multi-shot video sequences with full sound design and dialogue. Available now on Unlimited plans and Enterprise accounts outside of the US. Get started now at the link below.
https://x.com/runwayml/status/2041517519664463940

Nvidia’s answer to Tesla’s data advantage in self-driving Ali Kani, who has been at NVIDIA Automotive for almost 8 years, explains ↓ Watch the full video to see a test of Nvidia’s driving system on real streets and explore how they plan to bring self-driving to every car:
https://x.com/TheTuringPost/status/2041089313388343530

Text to Video Leaderboard – Top AI Video Models
https://artificialanalysis.ai/video/leaderboard/text-to-video

back in 2020 when i was writing blogs, i was always inspired by the 3blue1brown animations. it’s great to see that agents can now write manim code for anything by prompting. now that we have this, every blog should have these little animations.
https://x.com/casper_hansen_/status/2041046264758858081

We’ve added a new pseudonymous video model to our Text to Video and Image to Video Arenas.’HappyHorse-1.0′ is currently landing in the #1 spot for Text and Image to Video (No Audio) and the #2 spot for Text and Image to Video (With Audio). Further details coming soon. Example
https://x.com/ArtificialAnlys/status/2041591989083500933

VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward”” TL;DR: latent geometry-guided RL aligns video diffusion models for consistent 4D scene structure, improving camera stability and cross-view coherence
https://x.com/Almorgand/status/2039772881505149093

We’ve been studying what it takes to get NVFP4 & MXFP8 deliver good speedups on modern flow models for image & video gen. on B200 🕵️‍♂️ Today, I’m excited to share those findings! Bringing some cool recipes through Diffusers and TorchAO with `torch.compile` 🔥 Hop in ⬇️
https://x.com/RisingSayak/status/2042597708402430290

OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation”” TL;DR: panoramic video generation framework enabling long-horizon, consistent scene exploration with trajectory control and refinement
https://x.com/Almorgand/status/2041919499079725085

Researchers at Netflix just released a new AI model It erases objects from video, then rewrites the physics of the entire scene as if that object never existed It’s called VOID (Video Object and Interaction Deletion) Current inpainting tools simply paint over the gap left by
https://x.com/rowancheung/status/2041507881858826404

VOID: Video Object and Interaction Deletion
https://void-model.github.io/

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading