Meta announced Muse Spark in Voice Mode and Meta Glasses
https://www.testingcatalog.com/meta-to-release-muse-spark-in-voice-mode-and-meta-glasses/

Today we’re introducing Meta AI Voice Conversations powered by Muse Spark that let you talk naturally to Meta AI (interrupt, switch topics, or swap languages), and as you talk, Meta AI can generate images and pull up recommendations from Reels, maps, and more. We’re also bringing
https://x.com/MetaNewsroom/status/2054205287515484397

we launched some muse spark updates yesterday, including muse spark voice and live AI w your camera in Meta AI app + muse spark rolling out to glasses 😎 check them out!
https://x.com/alexandr_wang/status/2054588354914832439

Seeing the demos come together over the last week has been awesome — so many things that previously required a special-purpose model (e.g. real-time translation, event detection in video) turn out to be zero-shot instruction following once you have a general-purpose model with
https://x.com/johnschulman2/status/2053940940885332028

Interaction Models: A Scalable Approach to Human-AI Collaboration – Thinking Machines Lab
https://thinkingmachines.ai/blog/interaction-models/

Interaction Models: A Scalable Approach to Human-AI Collaboration – Thinking Machines Lab
https://thinkingmachines.ai/blog/interaction-models/

People talk, listen, watch, think, and collaborate at the same time, in real time. We’ve designed an AI that works with people the same way. We share our approach, early results, and a quick look at our model in action.
https://x.com/thinkymachines/status/2053938892152435174

Sharing our work on full-duplex multimodal models — real-time interaction that’s natural and intuitive without compromising on intelligence. We started Thinky in part to differentially advance capabilities for human-AI collaboration, which are underemphasized relative to
https://x.com/johnschulman2/status/2053940452789981426

thinking machines is using SGLang btw
https://x.com/eliebakouch/status/2053982248253190180

Thinking Machines know how to surprise. Those simultaneous abilities (not only translation but also creating graph while replying to a question) are pretty remarkable. Can’t wait to try it out and also learn how much it costs to use
https://x.com/TheTuringPost/status/2053975565179253010

Thinking Machines on X: “People talk, listen, watch, think, and collaborate at the same time, in real time. We’ve designed an AI that works with people the same way. We share our approach, early results, and a quick look at our model in action. https://t.co/AFJZ5kH7Ku https://t.co/uxl1InS6Ay” / X
https://x.com/thinkymachines/status/2053938892152435174

Thinky’s secret plan: 1: Increase Human<->AI bandwidth 2: Raise ceiling of human+AI intelligence 3: Help humans continue as main-characters in the new world We are at Step 1. Interaction Models are great real-time collaborative tools for humans. Here’s a preview:
https://x.com/soumithchintala/status/2053940215505645938

Very cool announcement from Thinky! The model looks nice (they go into some reasonable amount of detail), and reading some parts of the blog you can definitely see that the infea guys had a lot of fun there!
https://x.com/giffmana/status/2053953584300003405

After the Iran war started the US asked satellite firms to delay (then pause) war zone imagery. Why let adversaries use commercial US assets for targeting? Now Iran has released their own imagery. WaPo geolocated them – exposing the full extent of damage to US military sites.
https://x.com/bilawalsidhu/status/2052227041710219421

Human visual positioning system. In the era of ubiquitous maps, photo is equal to location.
https://x.com/bilawalsidhu/status/2053217307661406540

I love these projects because every city is already publishing itself, often unwittingly. Rodin published every towed car in SF in real time. SF noticed and shut the window in 3.5 hours. So many websites and services we use every day are effectively sensors to read world state.
https://x.com/bilawalsidhu/status/2054351915127857642

Love this. Old school matte paintings are a perfectly good way to augment reality. Much larger field of view than holding up your smartphone too.
https://x.com/bilawalsidhu/status/2053538357733384544

Normalize visualizing music with 3d gaussian splats
https://x.com/bilawalsidhu/status/2053588182109704427

Open source real-time world models legit feel like puppeteering reality itself. Now @reactorworld is putting it in everyone’s hands at sub 50 ms latency — type a prompt, change the scene, and control what unfolds as it generates.
https://x.com/bilawalsidhu/status/2052453815342014753

Synthetic Aperture Radar is cool af — a change detection machine that rips through cloud cover, day or night. And ICEYE knows how to show it.
https://x.com/bilawalsidhu/status/2052929763455504887

Using AI is about to feel like screen sharing with a really smart and attentive genius. Maybe this is the year computer use AI starts feeling magical?
https://x.com/bilawalsidhu/status/2054260705872740583

The lines between code and content are blurring
https://x.com/bilawalsidhu/status/2052189071447900568

Meta silently dropped Sapiens2 last week 🔥 a family of high-res models trained on 1B human images > for pose estimation, body-part segmentation, surface normals, pointmaps (sota) > 6 sizes: 0.1B → 5B params (all ViT patch 16) > high-res: 1024×768 and 4K
https://x.com/mervenoyann/status/2054187884417102319

Perceptron Mk1 shocks with highly performant video analysis AI model 80-90% cheaper than Anthropic, OpenAI & Google | VentureBeat
https://venturebeat.com/technology/perceptron-mk1-shocks-with-highly-performant-video-analysis-ai-model-80-90-cheaper-than-anthropic-openai-and-google

[2507.09313] ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
https://arxiv.org/abs/2507.09313

Our interaction model is the first general video+speech model that’s visually proactive. It was super fun working on this with @liliyu_lili / @saurabh_garg67 / @AndreaMadotto and others – after countless versions it was amazing when visual interruptions suddenly worked!
https://x.com/rown/status/2053950123139575863

Perceptron Mk1 is live on OpenRouter, built by @perceptroninc. Frontier video and embodied reasoning in a vision-language model. Analyzes video at a dynamic frame rate (up to 2 FPS) across a 32k multimodal context, with hybrid reasoning and structured spatial primitives (points,
https://x.com/OpenRouter/status/2054232344148787462

Today we’re releasing Perceptron Mk1: frontier video and embodied reasoning.
https://x.com/perceptroninc/status/2054216828285796630

We’re interested in AI systems that can collaborate in real time, without relying only on artificial turn boundaries. For audio, this feels natural: listen, speak, interrupt, update. For video, we think an important version of this is visual proactivity — models that respond
https://x.com/liliyu_lili/status/2053942465477197891

Big indoor scan – fully explorable in real-time on a browser. Distribution of 3DGS is pretty much solved. Blocker is still large scale capture – you need an expensive LiDAR + RGB scanner to get results like this. 360 video is still hard to pose indoors.
https://x.com/bilawalsidhu/status/2052391154474193122

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading