Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A 16:9 landscape gallery poster in the style of Julio Le Parc, with the bold letters ‘AR/VR’ centered and built from precise concentric ROYGBIV rainbow bands, violet outermost stepping inward through blue, green, yellow, orange to red, where two circular letter openings expand outward into a symmetrical pair of nested rainbow-banded lenses evoking a VR headset portal. Flat matte finish, crisp screen-printed edges, clean off-white background with abundant negative space, no shadows or gradients.

Introducing the Cosmos Coalition A new global initiative with NVIDIA and leading AI labs to build and open-source frontier world models for physical AI. Runway joins as a founding member, working alongside NVIDIA and a set of leading AI labs to build, share and accelerate world
https://x.com/runwayml/status/2061315089869721682

Jensen just launched NVIDIA Cosmos 3. Pitched as the first fully open omnimodel for physical AI: a mixture-of-transformers (reasoning + generation) with native vision reasoning and generation across text, image, video, sound, and action. Tops open-model leaderboards on
https://x.com/TheHumanoidHub/status/2061333253920080345

NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI | NVIDIA Newsroom
https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-3-the-open-frontier-foundation-model-for-physical-ai

Honestly if we put this demo in a fancy XR headset, we’d call it jarvis. The future of 3d doesn’t involve the death of autocad, maya and blender. It turns those tools into a shared canvas of collaboration with ai agents.
https://x.com/bilawalsidhu/status/2061450274011591084

Fucking cool. Giving AR portal vibes – like watching volumetric video with full 6dof head tracking. Maybe a future version of omni will be converting 2d to 3d video for our immersive displays and glasses.
https://x.com/bilawalsidhu/status/2060943911363690875

i like big splats and i cannot lie. you can now compress, tile and stream city scale 3d gaussian splats — glTF has an official 3DGS extension now too. this is what the future of google earth looks like. no more broccoli trees. no more melted powerlines. immaculate ground
https://x.com/bilawalsidhu/status/2060518632547877359

Omni’s take on rendering the camera path, then the *actual* earth studio render below by tatsuya. I reckon google could turn these into real spatial benchmarks for their ai video models.
https://x.com/bilawalsidhu/status/2060886445770870905

.@NVIDIA’s Cosmos 3 launched today… and guess who had early access? Agile Robots SE! They’ve been running it across their full portfolio: Thor single- and dual-arm, FR3 Duo. Focus? Simulation. Using Cosmos 3 as a neural simulator; a learned world model that generates
https://x.com/IlirAliu_/status/2061512207738012093

1/ NVIDIA just open-sourced Cosmos 3 at GTC Taipei! It’s the first fully open “”omnimodel”” for physical AI – one model that understands the real world, predicts what happens next, and generates the actions a robot should take. Weights, code, datasets. All open. And this is
https://x.com/kimmonismus/status/2061432501223162241

Breaking news: Cosmos 3 is here. They are attempting to do something completely new 🤯 Why is Physical AI much harder than building a chatbot? Understanding the world is not enough, robots need to predict it and act inside it. That’s the idea behind NVIDIA Cosmos 3: →
https://x.com/TheTuringPost/status/2061308942186414136

In case you missed this: NVIDIA shipped a text-to-image open weights model that looks seriously competitive 👀 (as part of its Cosmos 3 release)
https://x.com/victormustar/status/2061354267546427595?s=20

NVIDIA’s Cosmos 3 is what we’ve never seen before ‒ an Omnimodal World Model. It’s closing the loop for physical AI. All stack in one system: world and multimodal understanding, future generation, reasoning and action This is the next step in Jensen Huang’s AI progression:
https://x.com/TheTuringPost/status/2061474876083274238

NVIDIA’s Cosmos 3 lands at #1 among open weights models in both Text to Image and Image to Video on the Artificial Analysis Leaderboards! Cosmos 3 is a family of omnimodal world models for Physical AI from @nvidia, unifying language, image, video, audio and action in a single
https://x.com/ArtificialAnlys/status/2061494719998546206

NVIDIA’s Cosmos 3 lands at #1 among open weights models in both Text to Image and Image to Video on the Artificial Analysis Leaderboards! Cosmos 3 is a family of omnimodal world models for Physical AI from @nvidia, unifying language, image, video, audio and action in a single
https://x.com/ArtificialAnlys/status/2061494719998546206?s=20

A month later, I’m still thinking about that day in Augsburg. We co-hosted the KUKA × @VisComp1 Simulation Event… and honestly, it was one of those events where you leave with your head full. The room was packed with customers, partners, and simulation engineers from across
https://x.com/IlirAliu_/status/2061741801405374760

A system that automatically creates thousands of varied humanoid loco-manipulation demonstrations from one single teleoperated example: How? Via interleaved locomotion planning, manipulation planning, and skill adaptation. It features a new nine-task humanoid
https://x.com/IlirAliu_/status/2060637097120215196

AGIBOT has unveiled AGILE (AgiBot Generative Intelligent Locomotion Engine), a perception-control foundation model for whole-body humanoid locomotion. – It fuses visual perception, balance, and motion planning end-to-end, replacing the traditional split between “”seeing”” and
https://x.com/TheHumanoidHub/status/2060411921451680114

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners”” TL;DR: learns temporally consistent pixel-level representations from videos using linear in-context supervision from depth and motion cues
https://x.com/Almorgand/status/2060386938612232547

Helix4D: Complex 4D Mesh Generation”” TL;DR: generates temporally consistent 4D meshes under extreme topology changes like melting, shattering, and transparency using flow matching and temporal attention chains
https://x.com/Almorgand/status/2060031819471348115

Hey! A new vision encoder for robotics is in town 👀🤖 Instead of using models trained on static images (CLIP, SigLIP, DINO), we bake the dynamics-awareness directly into perception. It transfers well everywhere and boosts real-world OOD success by +22.5% Check it out👇
https://x.com/jbhuang0604/status/2061840469966090308

Larus went ham with this one! Love the synced highlighting on the camera path, something I wanted to try myself. Makes me think these could end up as spatial reasoning benchmarks for ai video models, esp in cities with existing 3d data as ground truth.
https://x.com/bilawalsidhu/status/2060373038982459741

Most robot vision systems take weeks to set up… or you do it in 6 minutes. Not 6 days. Not 6 hours. 6 minutes… data recording + model training included. I’ve been working closely with the team at Lentil Robotics, and what they’ve built is genuinely different from anything
https://x.com/IlirAliu_/status/2061862045541109967

Noble Machines’ Moby robot hauling a 50 lb crate. The whole-body control (WBC) is the physical foundation of the learning stack. It absorbs payload variation, contact response, and center-of-mass shifts in real time at the control layer and enabled cleaner data-focused learning
https://x.com/TheHumanoidHub/status/2059859238462353733

Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs
https://www.latent.space/p/andon

Robots can now reconstruct 3D scenes in real time from a single RGB camera. [📍 Projects page + paper] No depth sensor. No retraining. 30 FPS. Researchers at the Imperial College London introduced KV-Tracker, a training-free method that makes heavy models like π³ and Depth
https://x.com/IlirAliu_/status/2059906536240013690

THE LiDAR odometry package your robot needs. Most localization stacks assume the environment will cooperate. Distinct geometry. Clear point cloud differentiation. That assumption fails the moment you deploy in a warehouse, hospital, or industrial facility. GenZ-ICP is built to
https://x.com/IlirAliu_/status/2060783800788160674

Triangle Splatting SLAM”” TL;DR: the first dense RGB-D SLAM system built on differentiable triangle splatting, enabling real-time tracking, mesh reconstruction, collision checking, and scene editing from a unified representation.
https://x.com/Almorgand/status/2061839159438835881

We treat 3d scanning like a tech demo, but it’s actually spatial memory capture. Damn near teleportation. A few hundred photos of my parent’s old home, and now it’s immortalized forever. 3d gaussian splat made w/ reality capture + litchfeld.
https://x.com/bilawalsidhu/status/2061134940813611505

When Does LeJEPA Learn a World Model? It’s a new paper from @ylecun together with @klindt_david and @randall_balestr, that brings LeJEPA theory closer to the real goal of JEPA-style models. → LeJEPA learns a World Model when • The world’s latent variables are Gaussian •
https://x.com/TheTuringPost/status/2060153308392857933

A practical step forward for real-world manipulation: an open-source world model that replaces rigid action chunking with event-grounded prediction. It anchors planning and control to actual physical moments (reach → grasp → contact → place), giving robots more natural
https://x.com/IlirAliu_/status/2060422419479974312

The first open-source unified world model for scalable robot manipulation: 5B-parameter open-source unified video-action world model that combines policy and world modeling to generate robot actions, predict future visuals, and evaluate task progress from observations, language,
https://x.com/IlirAliu_/status/2061870964519076312

Well that escalated quickly 😂 Insane camera path generation harry!
https://x.com/bilawalsidhu/status/2061886480847450588

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading