Image created with Gemini. Image prompt: A flat acrylic-style collage of dozens of slightly offset rectangular photo panels forming one backyard swimming pool scene, turquoise water with hand-drawn white squiggle ripples, terracotta tile edge and lawn green grass, high noon light with no shadows, hard-edged shapes and saturated flat color, visible mismatched seams between panels, wide open composition with generous white space.

Apple Beat Google to It! Hyper-Realistic Maps at WWDC 2026 — Gaussian Splats & the ongoing 3D maps land grab to own the substrate that connects the physical and digital world. Hopping on with @RadianceFields to discuss.
https://x.com/bilawalsidhu/status/2064146894930989138

holy crap! apple just beat google to the punch — 3d gaussian splatting is coming to apple maps. these 3d scenes are made from oblique aerial imagery. but unlike blobby photogrammetry — no more broccoli trees, no more melted powerlines — ground level detail that actually holds
https://x.com/bilawalsidhu/status/2064057313057439795

ICYMI google’s been shipping gaussian splats in maps for a while — but tucked inside immersive view and indoor only. I used it heavily in NYC. Meanwhile apple’s bringing radiance fields to 300+ cities this fall, built from aerial imagery. Surprised they didn’t drop ground-level
https://x.com/bilawalsidhu/status/2064187152930365502

Your app can now search the web for images. Web search in the Responses API now supports image results in addition to text results, so you can build apps that surface products, places, visual references, and source links for inspiration.
https://x.com/OpenAIDevs/status/2064395155688616153

Turn data and comparisons into charts, directly in ChatGPT. Available now on mobile and web.
https://x.com/ChatGPTapp/status/2064018770839113769?s=20

Apple’s gaussian splat maps are rolling out on the developer beta! I hope they make a legit vision pro app – would be dope to drop into a city with the homies embodied as persona avatars.
https://x.com/bilawalsidhu/status/2064494805023911966

🎉 Meet vLLM-Omni v0.22.0, a major upgrade for omnimodal world models and production-grade multimodal serving. 🌍 Day-0 @NVIDIAAI Cosmos 3 world models: text, image, audio, video, and action, in and out. 🤖 Robot serving: DreamZero + OpenPI realtime API. 🎙️ Production TTS:
https://x.com/vllm_project/status/2064013506882703421

Today we published a technical blog post about Ideogram 4.0 — our goal is to enable more innovation and creativity. It’s a 9.3B Diffusion Transformer trained from scratch, paired with a frozen 8B VLM as text encoder. The nf4 checkpoint runs on a 24GB consumer GPU. Thread 🧵
https://x.com/ideogram_ai/status/2062956373957292281

In the Image Arena: open-weight Text-to-Image has a clear leader, with a tight race directly behind it: – #1 Ideogram-4.0 Quality has set the pace this week with a score of 1204. @ideogram_ai – #2 Hunyuan Image 3.0 by @TencentHunyuan with a score of 1151, just +1 pt ahead of
https://x.com/arena/status/2062997992777609534

Three new models entered the Image Arena Top 10 this past month (Text-to-Image): – #2 Reve 2.0 by @Reve (1,273), behind only GPT Image 2. – #4 MAI-Image-2.5 by @MicrosoftAI (1,253). – #9 Ideogram 4.0 Quality by @Ideogram_ai enters at #9 (1,204). And the only open-weights model in
https://x.com/arena/status/2062957421757452516

1X has hired Samarth Sinha to lead its new World Models research group. Samarth was previously a founding researcher at Luma AI, scaling large multimodal models. The mandate: build the next generation of foundation world models for NEO humanoid, trained on large-scale,
https://x.com/TheHumanoidHub/status/2062595716946784758

A new v0 robotics benchmark by independent researcher that tests robot foundation models on four tasks: Requiring spatial reasoning, geometric understanding, occlusion handling, and visuospatial planning. Using a low-cost open-source SO-101 robotic arm. The tasks progress in
https://x.com/IlirAliu_/status/2063894548821016989

A zero-shot video-language reward model. Trained on over 1M trajectories from 21 robot embodiments. It predicts: Frame-level task progress and generalizes zero-shot to unseen tasks, scenes, and robots, yielding 2.4-4.5x better success rates in: Online/offline RL, data
https://x.com/IlirAliu_/status/2063168037964877993

Boston Dynamics taught Atlas a Rabona kick. Learned from human mocap, retargeted to Atlas, trained in sim via RL, deployed zero-shot to the real robot. Soccer skills demand whole-body coordination, and similar recipe transfers to warehouse work.
https://x.com/TheHumanoidHub/status/2062607675339497680

If you can teach a robot a new skill in under an hour, you’ve just made your whole lab more flexible overnight. Pharma and biotech need automation that’s flexible enough for unstructured work… yet auditable enough to trust in regulated environments. That combination doesn’t
https://x.com/IlirAliu_/status/2062958380239511894

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading