Every week, I organize 400 to 700 links into roughly 60 categories as part of my ongoing effort to learn about AI. This is my personal notebook, which I enjoy sharing with friends… a hobby and a labor of love, rather than a commercial publication or product.
If you arrived here through a search or shared link, this page collects the links I found for Augmented Reality (AR/VR) for the week ending July 24, 2026.
As part of my learning process, I like to automate the category covers. It gives me a chance to learn Python and APIs.
This week’s cover prompt was written using Claude Opus 4.7, and the image was generated using Gemini 3.1 Flash Image Preview.
Category cover image prompt:
A giant glossy chrome VR headset floating in a deep cosmic purple-black sky with its two lenses glowing as portals pouring flowing rainbow ribbons of magenta, orange, yellow, lime, turquoise and violet outward in sweeping arcs, surrounded by starbursts and glitter sparkle under a golden mothership beam from above, with the title AR/VR arced across the top in huge fat rounded 1970s funk bubble letters filled with chrome and glitter and stacked multicolor drop shadows, 1970s psychedelic Afrofuturist concert poster style, one bold central subject, balanced composition with breathing room, maximum saturation.
This Week in Augmented Reality (AR/VR) News
Here’s a quick AI-generated summary by Claude Sonnet 5.5, based on the headlines and excerpts accompanying this week’s links:
- 3D scene capture keeps getting faster: Two research posts from the same account describe reconstructing scenes quickly from loose photos. One builds Gaussian Splat scenes from unordered captures, the other builds metric-scale indoor 3D from sparse panoramas. Bilawal Sidhu praised a scan that fuses aerial, ground, indoor and outdoor data.
- Captured scenes as robot training data: Researchers from Amazon FAR, Berkeley, Stanford and CMU scanned real rooms with an iPhone and rebuilt them as Gaussian Splatting scenes. Per the post, a robot trained on zero real-world data then walked, picked up boxes and followed multi-step instructions in the real world.
- Coding agents now make 3D content: Ethan Mollick said Codex used computer control to install Blender, which he had never used, and make an animated 3D otter, needing only one click to grant Windows permission. Sidhu also called making 3D motion graphics with coding agents absurdly fun.
This summary was generated by Claude Sonnet 5.5 to help you explore the links below. Rest assured, I select, organize, and check the links by hand in Google Sheets, and write the introduction and personal commentary in The Main Newsletters myself each week as a labor of love.
This week's links related to Augmented Reality (AR/VR)
The new 4-step Cosmos 3 Super models generate images and video up to 25x faster than the originals, and still rank among the best open-weight models on @ArtificialAnlys. 🥇 #1 for image-to-video (no audio) 🥈 #2 for text-to-image Try them on @huggingface:”
https://x.com/NVIDIAAI/status/2079949373069197658
NVIDIA’s Cosmos3 Edge is out! it watches videos streams & understands the mechanics/physics in them 🔥 it can reason in words, images, or next action prediction. physical AI reasoning, on the edge. try it on @huggingface (or on your edge device) ▶️
https://x.com/HuggingApps/status/2079923165157859362
Introducing Cosmos 3 Edge
https://huggingface.co/blog/nvidia/cosmos3edge
Such an immaculate 3d scan – fusing aerial/ground, indoor/outdoor results in the most magical reconstructions”
https://x.com/bilawalsidhu/status/2079356672963715184
Computer use with Codex is really impressive: “I’ve never used Blender, download and install it using your computer control and then use it to make an otter in 3D and turn it into a little animation” I only had to click once to give Windows permission to install Blender. Neat!”
https://x.com/emollick/status/2078318882473796029
Making 3d motion graphics with coding agents is absurdly fun. Breaking down geospatial tech for my next video.”
https://x.com/bilawalsidhu/status/2080146952524550329
Trained on zero real-world data. Learned to walk, pick up boxes, and follow multi-step instructions… in the REAL world. ( 📌 Paper below) Researchers from Amazon FAR, Berkeley, Stanford, and CMU scanned real rooms with an iPhone, rebuilt them as 3D Gaussian Splatting scenes,”
https://x.com/IlirAliu_/status/2079476975391985836
The latest Tesla iOS app decompile shows clear signs that Tesla is actively building home robot support for Optimus. – The app will request explicit permission to collect video and spatial data while Optimus operates in the home – Users will be able to view and filter”
https://x.com/TheHumanoidHub/status/2079827794490696169
MAI-Image-2.5-Pro launches today in Foundry for preview. It’s our highest-fidelity, professional-grade image model. It’s for super high quality imagery, detailed editing, precise in-image text rendering. It joins our image family of models so builders can pick the point on the”
https://x.com/mustafasuleyman/status/2080336466660724998
Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes” TLDR: feed-forward network that reconstructs complete metric-scale indoor 3D scenes from sparse, unordered panoramic images using learned covisibility-based reference selection and geometry-aware multi-task prediction”
https://x.com/Almorgand/status/2080359178804048141
Immediate 3D Gaussian Splat Reconstruction of Unordered Input with Global Consistency” TL;DR: enables immediate 3D Gaussian Splatting reconstruction from unordered image captures using fast matching, co-visibility graphs, loop closure, and scalable global optimization.”
https://x.com/Almorgand/status/2079269361253032104
Fast Foundation Stereo: When Foundation Models Meet Efficient Stereo Matching” TL;DR: distills stereo foundation models into an efficient real-time stereo matching framework, preserving high accuracy while dramatically reducing inference cost.”
https://x.com/Almorgand/status/2080354313428193577





Leave a Reply