Image created with gemini-3.1-flash-image-preview with claude-sonnet-4-5. Image prompt: Using the provided reference image, maintain the exact left-third close-crop composition with deep blue-purple cinematic lighting and misty atmospheric smoke bleeding rightward, but replace the central figure with a film director slumped in their chair, face lit by cold monitor glow, glitter scattered across their shoulders like film grain, expression weighted with creative exhaustion, right two-thirds open hazy space with ‘video’ in thin lowercase white Helvetica Neue Light, same melancholic post-party emotional register and HBO prestige drama aesthetic.
People are asking what’s the difference between Falcon Perception and SAM3, so here’s my opinion: SAM3:
https://t.co/KVRbuHm8H1 Falcon Perception:
https://t.co/QDgMlOBvDH First, sam3 does “”promptable concept segmentation””: simple noun phrases (like “”yellow bus””, “”red apple””) +
https://x.com/dahou_yasser/status/2041474094252933195
Today we’re releasing WildDet3D–an open model for monocular 3D object detection in the wild. It works with text, clicks, or 2D boxes, and on zero-shot evals it nearly doubles the best prior scores. 🧵
https://x.com/allen_ai/status/2041545111151022094
kays on X: “I noticed there wasn’t anything like this out there, so I wrote a tiny visual blog for those wanting to introduce themselves to Dynamic Gaussian Splatting and their current methods 🖼️ Feel free to check out, these are some of the visuals taken from it https://t.co/6W2qx2yI1K” / X
https://x.com/pabloadaw/status/2041650303804555278
We’re excited to be rolling out two model updates today! Marble 1.1: Improves lighting and contrast, with a major reduction in visual artifacts. Marble 1.1-Plus: Our new model built for scale. Create larger, more complex environments than ever before.
https://x.com/theworldlabs/status/2041554646561677701
I showed you SAM 3 all week. This is a 0.6B model that outperforms it. Falcon Perception. Type “”detect the plane”” and it segments every plane in the frame. Pixel-accurate masks from natural language. Fighter jets. Fire. Crowds. All on a MacBook via MLX. No cloud.
https://x.com/MaziyarPanahi/status/2040776481673281936
agents that make explainer videos > agents that summarize PDFs
https://x.com/lucatac0/status/2041018088913608923
Jeanne on X: “I’ve combined Manim @NousResearch’s Hermes Agent skill + @yifan_zhang_’s Math Code. Math Code executes the proof on a problem called Jordan’s Lemma and Hermes Agent with @claudeai Sonnet 3.7 directs Math Code, writes a script, gets Manim to render an explanatory video. https://t.co/qOsmOpvPlS” / X
https://x.com/prompterminal/status/2040982307377381583
Nous Research on X: “Introducing the Manim skill for Hermes Agent. Manim is an engine for creating precise programmatic animations for mathematical and technical explainers, made famous by the @3blue1brown channel. https://t.co/nyNeNthhZB” / X
https://x.com/NousResearch/status/2040931043658567916
PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models”” TL;DR: diffusion pipeline for scalable generation of photorealistic human data with 3D annotations
https://x.com/Almorgand/status/2040096997470843366
StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision”” TL;DR: adapts a pretrained 3D-aware transformer to stereo vision with a training-free pipeline, achieving SOTA performance on KITTI
https://x.com/Almorgand/status/2041569246883332385
Subtle hand held camera shake will convince people a 3d game render is real life footage. No generative ai required.
https://x.com/bilawalsidhu/status/2041643400433201384
PoseDreamer: Scalable Photorealistic Human Data Generation with Diffusion Models
https://prosperolo.github.io/posedreamer/
We always need more visuals! Checkout this on for dynamic gaussian splatting
https://x.com/Almorgand/status/2041773431524302968
The next version of @OpenClaw comes with native video generation. To start, I added support for the following companies: – Alibaba – BytePlus – fal – Google – MiniMax – OpenAI – Qwen – Together – xAI
https://x.com/steipete/status/2040928953653744003
📝Summarize 0.13 is out! 🎞️ Local video slides (–slides) 🤖 More model backends (GitHub Copilot) 🧠 Better GPT-5.4 support 📺 Better media handling (HLS detection.m3u8) It graduated from my tap to official homebrew formula! 🍺 brew install summarize
https://x.com/steipete/status/2041669438882087180
Robots can now reconstruct 3D scenes in real time from a single RGB camera. [📍 Projects page + paper] No depth sensor. No retraining. 30 FPS. Researchers at the Imperial College London introduced KV-Tracker, a training-free method that makes heavy models like π³ and Depth
https://x.com/IlirAliu_/status/2041062366025031787
Researchers just taught a robot to play tennis. From just clips of a few amateur players performing basic forehands, backhands, and shuffles… …a robot learned one of the fastest, most coordinated physical skills there is. Insane!
https://x.com/rowancheung/status/2040085788256506190
Robotics pre-training *from scratch* has been a heretical idea for the last two years. That “there’s no internet of robotics data” has led to two prevailing conclusions: 1) we need to use pretrained model backbones and 2) we need to scale robotics data. The first conclusion in
https://x.com/xiao_ted/status/2041547335935853025
PixVerse C1 is live–our first model built for film production. Coherent action, storyboard-to-video, ref-guided consistency. 1080p, 15s, native audio. Available on PixVerse Web and API Platform. RT+Follow+Reply=300Creds(72H ONLY)
https://x.com/PixVerse_/status/2041536108660740162
Seedance 2.0 is now on Runway. Use text, image, video or audio as inputs to generate stunning multi-shot video sequences with full sound design and dialogue. Available now on Unlimited plans and Enterprise accounts outside of the US. Get started now at the link below.
https://x.com/runwayml/status/2041517519664463940
Nvidia’s answer to Tesla’s data advantage in self-driving Ali Kani, who has been at NVIDIA Automotive for almost 8 years, explains ↓ Watch the full video to see a test of Nvidia’s driving system on real streets and explore how they plan to bring self-driving to every car:
https://x.com/TheTuringPost/status/2041089313388343530
Text to Video Leaderboard – Top AI Video Models
https://artificialanalysis.ai/video/leaderboard/text-to-video
back in 2020 when i was writing blogs, i was always inspired by the 3blue1brown animations. it’s great to see that agents can now write manim code for anything by prompting. now that we have this, every blog should have these little animations.
https://x.com/casper_hansen_/status/2041046264758858081
We’ve added a new pseudonymous video model to our Text to Video and Image to Video Arenas.’HappyHorse-1.0′ is currently landing in the #1 spot for Text and Image to Video (No Audio) and the #2 spot for Text and Image to Video (With Audio). Further details coming soon. Example
https://x.com/ArtificialAnlys/status/2041591989083500933
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward”” TL;DR: latent geometry-guided RL aligns video diffusion models for consistent 4D scene structure, improving camera stability and cross-view coherence
https://x.com/Almorgand/status/2039772881505149093
We’ve been studying what it takes to get NVFP4 & MXFP8 deliver good speedups on modern flow models for image & video gen. on B200 🕵️♂️ Today, I’m excited to share those findings! Bringing some cool recipes through Diffusers and TorchAO with `torch.compile` 🔥 Hop in ⬇️
https://x.com/RisingSayak/status/2042597708402430290
OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation”” TL;DR: panoramic video generation framework enabling long-horizon, consistent scene exploration with trajectory control and refinement
https://x.com/Almorgand/status/2041919499079725085
Researchers at Netflix just released a new AI model It erases objects from video, then rewrites the physics of the entire scene as if that object never existed It’s called VOID (Video Object and Interaction Deletion) Current inpainting tools simply paint over the gap left by
https://x.com/rowancheung/status/2041507881858826404
VOID: Video Object and Interaction Deletion
https://void-model.github.io/





Leave a Reply