Image created with OpenAI GPT-Image-1. Image prompt: 1966 Kodachrome photo-look, thin white frame, forest-green title band in upper left with stacked yellow/white serif text reading “MULTIMODALITY”, thin vertical scratch along film scene featuring a collage scrapbook page taped to the pen; gentle film grain, overcast daylight

Bloomberg: where are your robots 😂😂 https://x.com/adcock_brett/status/1930997923539828884

Brett Adcock says the Figure 03 robot is 93% cheaper than Figure 02; the company is well-capitalized to support sufficient training and inference compute, produce hundreds of thousands of robots, and hire the best people; the current focus is to keep the team small and fast. https://x.com/TheHumanoidHub/status/1931263196217909682

Brett Adcock says the latest autonomous demo of Figure 02 is fully end-to-end and uses a single neural network – camera frames in, actions out. “You cannot code your way out of this problem.” https://x.com/TheHumanoidHub/status/1931039140512145724

Figure CEO Brett Adcock says his robots will share a single brain. When one learns something new, they all get smarter. Want an employee or a home assistant? You’ll pick the one that learns from everyone’s mistakes. This is how the flywheel spins. And why he believes the first https://x.com/vitrupo/status/1931001200604037145

fyi, we just posted a deep-dive write-up on the 60min logistics video Full behind-the-scenes on the AI work powering the latest Helix release https://x.com/adcock_brett/status/1932192198025773371

Here’s my Bloomberg interview: → Figure’s business model → Best robot AI models today → Robots in the home and workforce Interview link in comment below https://x.com/adcock_brett/status/1932071569633022445

Millions. There is a potential to shipping millions of robots doing just stuff like this A little under half of GDP is human labor And our humanoid robots are just synthetic humans who can work longer and ultimately faster / more accurate”” / X https://x.com/adcock_brett/status/1931886869316538830

There’s no way to code your way out of this problem Every bag is different; Every pile of packages is different This is however perfect for neural networks”” / X https://x.com/adcock_brett/status/1932280240170250319

This could be a winner-take-all industry Whoever builds the smartest & cheapest robot wins More robots = lower cost = more training data = smarter Helix No one wants the dumb robot in their home or workplace”” / X https://x.com/adcock_brett/status/1931091232912212078

“Stair mode” in NEO’s RL controller engages stereo RGB vision to infer the height of the floor around it, combining this with proprioceptive history to anticipate each step’s height and plan precise, stable foot placement. https://x.com/TheHumanoidHub/status/1932869701774028831

1X announces Redwood AI Redwood is a vision-language transformer model that empowers NEO to perform end-to-end mobile manipulation tasks – retrieving objects, opening doors, and navigating complex home environments. Trained on a large dataset of teleoperated and autonomous https://x.com/TheHumanoidHub/status/1932481396335128821

1X announces their latest reinforcement learning (RL) controller, which unlocks NEO’s full-body mobility for home environments, enabling Redwood AI (1X’s in-house AI model) to interact with the physical world more naturally and broadly. The unified controller supports walking https://x.com/TheHumanoidHub/status/1932864588648964459

NEOs sighted in the natural world. It’s a teaser for 1X updates dropping throughout the week. https://x.com/TheHumanoidHub/status/1932115480342593867

Redwood NEO’s AI https://x.com/1x_tech/status/1932474830840082498

Something new soon. Something NEO soon. Stay tuned.”” / X https://x.com/TheHumanoidHub/status/1931849987744530923

Apple (AAPL) Targets Spring 2026 for Release of Delayed Siri AI Upgrade – Bloomberg https://www.bloomberg.com/news/articles/2025-06-12/apple-targets-spring-2026-for-release-of-delayed-siri-ai-upgrade?srnd=undefined&sref=9hGJlFio&embedded-checkout=true

Apple doesn’t report benchmarks for their AIs, reporting on an ill-documented head-to-head evaluation But even by their standards, Apple’s latest on device models are mostly worse than the open Gemma 3-4B from Google or Qwen 3-4B And their server LLM is similar to Llama 4 Scout https://x.com/emollick/status/1932420903515590997

Our vision is for AI that uses world models to adapt in new and dynamic environments and efficiently learn new skills. We’re sharing V-JEPA 2, a new world model with state-of-the-art performance in visual understanding and prediction. V-JEPA 2 is a 1.2 billion-parameter model, https://x.com/AIatMeta/status/1932808881627148450

Google Search AI Mode now offers data visualization and charts https://blog.google/products/search/ai-mode-data-visualization/

Introducing the V-JEPA 2 world model and new benchmarks for physical reasoning https://ai.meta.com/blog/v-jepa-2-world-model-benchmarks/

It’s true. The Meta offers for the “”superintelligence”” team are actually insane. If you work at the big AI labs, Zuck is personally negotiating $10M+/yr in cold hard liquid money. I’ve never seen anything like it.”” / X https://x.com/deedydas/status/1932828204575961477

Meta’s Mark Zuckerberg Creating New Superintelligence AI Team – Bloomberg https://www.bloomberg.com/news/articles/2025-06-10/zuckerberg-recruits-new-superintelligence-ai-group-at-meta?embedded-checkout=true

ChatGPT voice is getting really good — https://x.com/gdb/status/1931456650336141752

Wow, new expressive voice in ⁦⁦@ChatGPTapp⁩ doesn’t just talk, it performs. Feels less like an AI and more like a human friend. Nice work ⁦@OpenAI⁩ team. 🎤🎶🚀 https://x.com/shaunralston/status/1931361225046405233

Haven’t tried the updated Advanced Voice that was recently launched to all paid users in ChatGPT? Then take a listen below. Prompt: Wish me an awkward happy birthday. https://x.com/OpenAI/status/1932166285447856130

The new ChatGPT Advanced Voice Mode is super interesting – lots of deliberate use of disfluencies (nervous laughs, ums & ahs) and vocal changes make it feel much more human than the previous version Really shows the possibilities from multimodal voice vs most AI’s text-to-speech”” / X https://x.com/emollick/status/1931557886947205629

This clip from Figure shows what’s so exciting about end to end applications. Handling a messy, unstructured feed of objects like this would be very difficult to scale with traditional methods https://x.com/chris_j_paxton/status/1932072973847941505

This is 🤯 Figure 02 autonomously sorting and scanning packages, including deformable ones. The speed and dexterity are amazing. https://x.com/TheHumanoidHub/status/1930706769061564921

Uncut hour-long footage of Figure 02 autonomously transferring and flattening packages for a scanner down the line. The robot is using Figure’s Helix model, a generalist VLA that now incorporates upgrades in temporal memory and force feedback. https://x.com/TheHumanoidHub/status/1931394946768249324

Welcome to the most boring video we’ve ever posted Here’s 60 minutes of our humanoid robot solving logistics, powered by our Helix neural network https://x.com/adcock_brett/status/1931391783306678515

I finally built PodPixel using @Replit 🎉 An web app that transcribes podcasts & pulls out all links/resources with context. Just use the search or drop a URL, and find those links. Try yourself. https://x.com/designworkplan/status/1928756748153659509

Apple Intelligence gets even more powerful with new capabilities across Apple devices – Apple https://www.apple.com/newsroom/2025/06/apple-intelligence-gets-even-more-powerful-with-new-capabilities-across-apple-devices/

This is a very thought provoking interview with my former student. I do think AI personas (esp multimodal and real time) may be addictive and seem better than humans – but so is heroin (albeit heroin has less useful applications than AI).”” / X https://x.com/sirbayes/status/1932155427703431647

Agility’s Digit executes a multi-step task autonomously from a natural language command. “”Bring me the ingredients to make pasta.”” https://x.com/TheHumanoidHub/status/1930387671626690985

Vision Transformers have high computational costs. Existing token reduction methods like pruning and merging are exclusive, causing significant information loss and needing post-training to recover performance. This paper presents Token Transforming, a unified many-to-many https://x.com/rohanpaul_ai/status/1932718446648918269

World first: Breakthrough AI powered Brain-Computer Interface Enables Real-Time Speech for ALS Patient → A 45-year-old man with ALS can now produce expressive speech and melody using a brain-computer interface (BCI) that translates brain signals into audio in 10 milliseconds. https://x.com/rohanpaul_ai/status/1933094038816858372

Voice cloning is now trivially easy with open source tools, while live avatar videos of real people are easy with proprietary tools & a variety of open source tools are getting there. Very limited time to adjust legal & financial safeguards to new ways of authenticating people”” / X https://x.com/emollick/status/1931364236304830675

We just launched our biggest update yet. Meet Higgsfield Speak — the fastest way to make motion-driven talking videos. Pick a style, choose an avatar, type a script. We do the rest — cinematic motion, voice, emotion. Comment Speak to get the full guide + promo code in the DM. https://x.com/higgsfield_ai/status/1930686472845455417

Figure update coming in the next 1-2 hours…standby”” / X https://x.com/adcock_brett/status/1931384282993537028

Figure’s Helix model can also perform a human–robot handover. “Helix’s single neural network policy produces the appropriate response based on what it sees. The system can be taught new context-dependent behaviors with only a handful of demonstrations.” https://x.com/TheHumanoidHub/status/1931534295619023323

Humanoid flips box (fun meme spoof of Figure celebration) 😂 https://x.com/adcock_brett/status/1931850724343964116

If I asked Figure team to lose commercialization time to be at an event they might quit: https://x.com/adcock_brett/status/1931219135175770271

Image of the next Figure robot model https://x.com/adcock_brett/status/1932821044789919766

Is this working Dan? (Figure response to heckler) https://x.com/adcock_brett/status/1930693311771332853

It really feels like general robotics is within reach One robot for every human Congrats to Louis who led the project!”” / X https://x.com/adcock_brett/status/1931509884484567323

New jobs are now posted at Figure, this will be a monster year for us Join us to build next-generation physical AI, humanoid robots, and all the subsystems that power them: > AI, Training Infra > AI, Large Scale Training > AI, Large Scale Model Evals > AI, Reinforcement https://x.com/adcock_brett/status/1881792213514158565

This was a few months ago…Helix is now showing massive improvements in logistics – can’t wait to show you what’s new https://x.com/adcock_brett/status/1930655461105426604

You there Dan? In the meantime, this is fully autonomous – powered by Helix 🧬 The AI policy learns to flip each package barcode-side down and even flattens puffy packages just like a human would”” / X https://x.com/adcock_brett/status/1930854226529251565

If I could tell my younger self one thing after 20 years of building tech: move fast Speed is the ultimate advantage…the ultimate moat I used to chase perfect – every launch, every feature. But perfect slows you down. It blocks feedback, delays learning, kills momentum”” / X https://x.com/adcock_brett/status/1933226344156221746

Used gemini 2.5 pro to build a shot counter for myself + write an after effects script to create this AR style HUD overlay. Footage captured w/ my meta rayban glasses. Insane how much better 2.5 pro is at media understanding vs 1.5 pro (when I last tried this). Using both video https://x.com/bilawalsidhu/status/1931030893772017703

1X Robotics introduced Redwood, an AI merging language control, locomotion, and whole body manipulation The neural net enables 1X Neo to do more chores at home, including opening doors and picking up never-before-seen objects in unfamiliar locations https://x.com/rowancheung/status/1932694752534684082

Paper page – GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents https://huggingface.co/papers/2506.03143

apple really put “”the last in the group chat to get the joke”” in an ad about apple intelligence”” / X https://x.com/swyx/status/1932137205268688983

Apple WWDC 2025 > What users wanted: Siri that actually works > What users got: “You’ll immediately notice how the playback controls refract the environment. Sidebars and toolbars reflect the depth of your workspace and offer a subtle hint of the content.” I wanted more. But”” / X https://x.com/bilawalsidhu/status/1932168211963007179

Apple’s “spatial scenes” remind me of Facebook 3d photos from 2018. Take any photo and use AI to give it real depth and parallax. Glad Apple is starting to think beyond stereo photo/video for the Vision Pro; 6dof media needs to be a first class citizen. https://x.com/bilawalsidhu/status/1932286185285750791

Apple’s Visual Intelligence was showcased with a familiar demo for anyone following recent developer conferences: More ways to buy stuff, more swiftly, powered by AI. https://x.com/TechCrunch/status/1932147112164069608

I am a graphics programmer, and here’s my feedback on Apple’s Liquid Glass beta. The idea is cool, but it’s difficult to work with from a UX perspective. Let’s start with the main problems: 1 – Low Contrast: It’s clearly not readable, but there are many different ways to fix it. https://x.com/XorDev/status/1932429551256101328

Interesting to see Apple double down on conventional UIs while ignoring AI when the goal of the big AI firms is to make it so that you just talk to AI to get whatever you want done, without touching a UI.”” / X https://x.com/emollick/status/1932225668487463374

lmfao Apple models sound so 2010ish”” / X https://x.com/cto_junior/status/1932128352036605962

New iOS feels like a junior designer discovered the gradient tool, and are now using it EVERYWHERE. I’ve been there, that was me once.”” / X https://x.com/dzhng/status/1932135452569714863

RT @fkasummer: apple is about to have their windows vista moment”” / X https://x.com/zacharynado/status/1932259455368102098

Updates to Apple’s On-Device and Server Foundation Language Models – Apple Machine Learning Research https://machinelearning.apple.com/research/apple-foundation-models-2025-updates

What could happen at Apple’s WWDC 2025? See latest rumors https://www.usatoday.com/story/tech/2025/06/04/apple-wwdc-2025-rumors/84017268007/

Windows Vista walked so iOS 26 could run.”” / X https://x.com/skirano/status/1932145646963704199

WWDC: Apple opens its AI to developers but keeps its broader ambitions modest | Reuters https://www.reuters.com/business/wwdc-apple-faces-ai-regulatory-challenges-it-woos-software-developers-2025-06-09/

Edge AI Innovation: Real-Time Pose Detection | Dell https://www.dell.com/en-us/blog/edge-ai-innovation-real-time-pose-detection/

1X CEO Bernt Bornich explained what a world model is on the latest ‘NVIDIA AI Podcast.’ https://x.com/TheHumanoidHub/status/1930762366775693338

alignhuman.github.io https://alignhuman.github.io/

Seedance https://seed.bytedance.com/en/seedance

SyncTalk++ https://ziqiaopeng.github.io/synctalk++/

Reimagining TTS with LLM-Powered Audio Generation | Bland AI https://www.bland.ai/blogs/new-tts-announcement

RT @freddy_alfonso_: 🚨 NotebookLM Dethroned?! 🚨 Meet vui: The new open-source dialogue generation model. 💪100M Params, 40k hours audio!…”” / X https://x.com/_akhaliq/status/1932149790747525396

OpenAI to continue working with Scale AI after Meta deal | Reuters https://www.reuters.com/technology/openai-continue-working-with-scale-ai-after-meta-deal-2025-06-13/

Richard Sutton argues that AI must move beyond human-generated static data into the “Era of Experience,” where agents learn through continuous interaction with the world. This will require building upon RL with better algorithms capable of continual learning and meta-learning. https://x.com/TheHumanoidHub/status/1931969449688719439

🎉CVPR, hosted annually by IEEE, is the leading event in computer vision. Renowned for its high quality and affordability, it provides exceptional value to students, researchers, and industry professionals. 👤The CVPR event features Pengfei Wan, Head of Kling Video Generation https://x.com/Kling_ai/status/1932464913018147291

Congratulations to our excellent multimodal team! 🥳 new blog and model release 🚀 A state of the art CLIP model from just data curation. https://x.com/code_star/status/1932438873399033943

Dolphin: new OCR model by @BytedanceTalk with MIT license 🐬 the model first detects element in the layout (table, formula etc) and then parses each element in parallel for generation ⤵️ model and demo is on @huggingface Hub 🤗 https://x.com/mervenoyann/status/1933465022857982394

RT @nickhjiang: Vision transformers have high-norm outliers that hurt performance and distort attention. While prior work removed them by r…”” / X https://x.com/TimDarcet/status/1932707025718247935

RT @ZyphraAI: Zyphra is expanding! Join our growing team in Palo Alto. We have multiple roles open across multimodal foundation models, RL…”” / X https://x.com/QuentinAnthon15/status/1932128395594510598

Vision-language-action models suffer from high inference latency and discontinuities between action chunks. Real-time chunking (RTC), new research from @physical_int applies an inference-time freezing and inpainting scheme to ensure smooth asynchronous action execution. ⚙️ The https://x.com/rohanpaul_ai/status/1932382707574505763

DesignBench provides a benchmark for multimodal LLMs evaluating front-end engineering across popular frameworks and tasks like generation, edit, and repair. Methods 🔧: → DesignBench contains 900 real-world webpage samples for HTML/CSS, React, Vue, and Angular frameworks. → https://x.com/rohanpaul_ai/status/1932279554954940445

Dating ancient manuscripts using radiocarbon and AI-based writing style analysis | PLOS One https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0323185

[2411.12915] VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge https://arxiv.org/abs/2411.12915

CVPR 2025: VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge – YouTube https://www.youtube.com/watch?v=_Z2KMfDXkwY&list=PLaU7MWI8yG9Uy8P_3K5R4_H6HZ5AsYsxD&index=5&t=4s

Graph Neural Network Combining Event Stream and Periodic Aggregation for Low-Latency Event-based Vision https://openaccess.thecvf.com/content/CVPR2025/papers/Dampfhoffer_Graph_Neural_Network_Combining_Event_Stream_and_Periodic_Aggregation_for_CVPR_2025_paper.pdf

Introducing Manus video generation. Manus transforms your prompts into complete stories—structured, sequenced, and ready to watch. With a single prompt, Manus plans each scene, crafts the visuals, and animates your vision. From storyboard creation to concept visualization—your https://x.com/ManusAI_HQ/status/1929913745503072551

Also illustrates how far Siri has fallen behind. The gap between it and ChatGPT Advanced Voice Mode is vast (as is the gap between Siri and Gemini Voice, which is not quite as advanced as ChatGPT) The usual fast follower approach may fail as people come to trust “”their”” chatbot.”” / X https://x.com/emollick/status/1931914916341944456

Vision-language-action (VLA) models in robotics often suffer from latency and jerky transitions, struggling to act smoothly while thinking ahead. A new paper from Physical Intelligence introduces Real-Time Chunking (RTC) – a method that lets robots plan the next actions while https://x.com/TheHumanoidHub/status/1932158191502279021

Scaling laws drive smarter humanoid robots. Figure 02’s accuracy in the package handling task improved from 88.2% to 94.4% after the training data was scaled up by 6x. https://x.com/TheHumanoidHub/status/1931398514359353393

Figure AI CEO on Building General Purpose Robots – YouTube https://www.youtube.com/watch?v=zObe3aOz5fw

Introducing V-JEPA 2, a new world model with state-of-the-art performance in visual understanding and prediction. V-JEPA 2 can enable zero-shot planning in robots—allowing them to plan and execute tasks in unfamiliar environments. Download V-JEPA 2 and read our research paper https://x.com/AIatMeta/status/1932923002276229390

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading