Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Top-down isometric PS1-style pixel art scene showing a city block with three visible stacked layers: solid chunky sprites on street level, translucent neon holographic UI layer floating above, and fantasy virtual elements in upper layer, all casting colored shadows through each other with dithered textures and saturated primary colors against dark asphalt, 32-bit graphics aesthetic with CRT glow
NitroGen: A Foundation Model for Generalist Gaming Agents”” TL;DR: vision-to-action model trained on 40k+ hrs of gameplay across 1,000+ games, mapping raw pixels to gamepad actions for generalist agents.”” https://x.com/Almorgand/status/2011847937899589672
Virtual avatars have become insanely accessible. Japan invented the VTuber trend in the 2010s, but the barriers to creation have been nuts — you needed a mocap suit, beefy pc to run unreal engine, and often multiple operators. Now all you need is the phone in your pocket and a”” https://x.com/bilawalsidhu/status/2013974187405082694
ICo3D: An Interactive Conversational 3D Virtual Human https://ico3d.github.io/
It’s really cool to see “legacy” players in geospatial like ESRI adopt 3D gaussian splatting in ArcGIS. Radiance fields are v. complementary to point clouds and textured 3D meshes — exceedingly human readable and much closer to reality than anything else.”” https://x.com/bilawalsidhu/status/2012954912766746958
Mixing 3d models with real world scans (made w/ 3d gaussian splatting) is a powerful combination. And when you do it all inside an unbiased, physically accurate 3d rendered like Octane – the fidelity is truly next level.”” https://x.com/bilawalsidhu/status/2013677348474769615
This is some quietly impressive work on making video world models actually controllable in 4D space. VerseCrafter lets you take an input image, use something like Blender to animate the 3D camera path and object trajectories, then uses that to condition generation. Scribbling in”” https://x.com/bilawalsidhu/status/2011922881597292811
What if sim and reality were one? This system keeps them in sync… always! 👉 Project & Paper ⬇️ > uses 3D Gaussians and particles to predict and correct object states > creates a real-time digital twin that stays synced with the real world > enables rich policy training using”” https://x.com/IlirAliu_/status/2013535990283989158
From text to assembled objects, end to end. The idea is simple but powerful: You describe an object in words. AI designs it in 3D. A robot figures out how to build it from real parts. Text to robotic assembly of multicomponent objects is introduced by a team from MIT and”” https://x.com/IlirAliu_/status/2012239250939347083
ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and TEst-time Generative Adaptation https://kim-youwang.github.io/elite
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control”” TL;DR: 4D-aware video world model with explicit camera + object motion control via geometric representations, enabling high-fidelity, controllable video synthesis.”” https://x.com/Almorgand/status/2014043511964815574
VeloDepth: Video Depth Propagation for Consistent, Efficient Depth”” TL;DR: video depth pipeline warping/refining deep features for fast, temporally stable depth prediction across frames.”” https://x.com/Almorgand/status/2012206027815395804
Humans usually guess what robots feel. This system actually lets you feel it. This is an end-to-end high resolution touch teleoperation demo built with Samsung. Real tactile signals from a robot hand were streamed directly into a human fingertip display and a glove in real”” https://x.com/IlirAliu_/status/2012601631347470741
Excited for this! I think it is by far the most interesting approach of any BCI effort.”” https://x.com/sama/status/2013403329322242261
There are many ideas on how to build better world models, and one of the recent ones comes from @Princeton with Web World Models (WWMs). ▪️ The key principle: separate rules from imagination. WWM builds around these two pieces: 1. The physical layer is handled by code. It’s”” https://x.com/TheTuringPost/status/2013016473514717330
Elon Musk says Tesla’s restarted Dojo3 will be for ‘space-based AI compute’ | TechCrunch https://techcrunch.com/2026/01/20/elon-musk-says-teslas-restarted-dojo3-will-be-for-space-based-ai-compute/
Hyundai Motor Group has appointed Milan Kovac, the former Head of Engineering for Tesla Optimus, as an outside director of Boston Dynamics (a strategic board seat). From informed sources: Milan is not joining Boston Dynamics in a full-time capacity as an executive or tech lead.”” https://x.com/TheHumanoidHub/status/2012241287131644250
Instead of generating more AI slop””… Tesla’s building world models for Optimus, for evaluation and closed-loop RL.”” https://x.com/TheHumanoidHub/status/2011632089708576793
Jason Calacanis recently got to see Optimus 3. He says: “”Nobody will remember that Tesla ever made a car. They will only remember the Optimus and that he is going to make a billion of those.”””” https://x.com/TheHumanoidHub/status/2011640366089585045
Tesla has officially updated their mission statement to: ‘Building a World of Amazing Abundance.’”” https://x.com/TheHumanoidHub/status/2013716027306320257
The video Elon shared is AI generated, but the input image (starting frame) looks like a real picture. Could that be a sneaky reveal of Tesla Optimus 3?”” https://x.com/TheHumanoidHub/status/2011983766458417608





Leave a Reply