Image created with Flux Pro v1.1 Ultra. Image prompt: Giant “100” as pure white negative‑space cutout dominating the frame; minimalist poster style; clean waveform and spectrogram bands threading through the zeros; midnight blue backdrop; high contrast, crisp edges, soft studio light, no other text, no logos

Blue (@heyBlueX) lets you control your phone’s apps by voice so tasks actually get finished, hands-free. It handles messages, email, and actions across apps by tapping and typing as you would. https://x.com/ycombinator/status/1958182627422146811

Google Translate adds live translation and language learning https://blog.google/products/translate/language-learning-live-translate/

If you’ve ever… Spent longer finding a movie than watching it. Avoided continuing a show because you forgot what happened. Tried to find something to watch for 3 people with polar opposite tastes. Good news. Introducing @Copilot on @Samsung TVs and monitors. https://x.com/mustafasuleyman/status/1960735880966234290

Introducing gpt-realtime and Realtime API updates for production voice agents | OpenAI https://openai.com/index/introducing-gpt-realtime/#image-input

“”Huge Realtime API release today! Details below, but TLDR: – GA (out of beta) – better instruction following, naturalness, audio – MCP support – new voices – SIP (telephony) support – new WebRTC APIs and video support Demos: https://x.com/juberti/status/1961116594211364942

Introducing gpt-realtime and Realtime API updates for production voice agents | OpenAI https://openai.com/index/introducing-gpt-realtime/#additional-capabilities

Introducing gpt-realtime and Realtime API updates for production voice agents | OpenAI https://openai.com/index/introducing-gpt-realtime/#remote-mcp-server-support

The Realtime API is officially out of beta and ready for your production voice agents! We’re also introducing gpt-realtime—our most advanced speech-to-speech model yet—plus new voices and API capabilities: 🔌 Remote MCPs 🖼️ Image input 📞 SIP phone calling ♻️ Reusable prompts https://x.com/OpenAIDevs/status/1961124915719053589

Voice is the OG modality. So excited for image inputs, function calling & MCP support in the Realtime API GA! `gpt-realtime` is a lot more natural and expressive, and every time a SOTA voice model is released, you know what I gotta do… Here is the new voice Marin, on https://x.com/swyx/status/1961124194789499233

Apple in talks to use Google’s Gemini AI to power revamped Siri, Bloomberg News reports | Reuters https://www.reuters.com/business/apple-talks-use-googles-gemini-ai-power-revamped-siri-bloomberg-news-reports-2025-08-22/

Harvard dropouts to launch ‘always on’ AI smart glasses that listen and record every conversation | TechCrunch https://techcrunch.com/2025/08/20/harvard-dropouts-to-launch-always-on-ai-smart-glasses-that-listen-and-record-every-conversation/

Testing Wan 2.2 S2V with my musician character and a Smells Like Teen Spirit cover. It is pretty good, but not perfect. The background music may have a large effect. Curious if it would be better to preprocess the audio better or finetune the model to handle music better. https://x.com/ostrisai/status/1960907113821298877

Eleven v3 (alpha), now available in the API | ElevenLabs https://elevenlabs.io/blog/eleven-v3-alpha-now-available-in-the-api

Today we’re announcing the open-source release of HunyuanVideo-Foley, our new end-to-end Text-Video-to-Audio (TV2A) framework for generating high-fidelity audio.🚀 This tool empowers creators in video production, filmmaking, and game development to generate professional-grade https://x.com/TencentHunyuan/status/1960920482779423211

🚨Text Leaderboard Update: A new model provider, @MicrosoftAI has broken into the Top 15 this week! 💠MAI-1-preview by @MicrosoftAI debuts at #13. Congrats to the Microsoft AI team! As the Text Arena is one of the most competitive races, breaking into the Top 15 is no small https://x.com/lmarena_ai/status/1961112908026593557

At Microsoft we have a bold vision for applied AI—responsible, reliable, and filled with personality and expertise. The launch of MAI-Voice-1 and MAI-1-preview, our first in-house models, are just the beginning.”” / X https://x.com/yusuf_i_mehdi/status/1961112928230461615

Excited to share our first @MicrosoftAI in-house models: MAI-Voice-1 and MAI-1-preview. Details and how you can test below, with lots more to come⬇️ https://x.com/mustafasuleyman/status/1961111770422186452

Two in-house models in support of our mission   | Microsoft AI https://microsoft.ai/news/two-new-in-house-models/

microsoft is dropping (still uploading) VibeVoice-1.5B model on @huggingface! i love the multi-speaker conversational audio feature for podcasts! https://x.com/MaziyarPanahi/status/1959994276198351145

microsoft/VibeVoice-1.5B · Hugging Face https://huggingface.co/microsoft/VibeVoice-1.5B

VibeVoice A Frontier Open-Source Text-to-Speech Model https://x.com/_akhaliq/status/1960106923191140373

VibeVoice is a framework from @MSFTResearch for generating expressive, long-form, multi-speaker audio conversations. Create podcasts from text. MIT licensed🔥 Synthesize speech up to 90 minutes long with 4 distinct speakers 🤯 https://x.com/Gradio/status/1960023019239133503

A smarter way to talk to your TV: Microsoft Copilot launches on Samsung TVs and monitors   | Microsoft Copilot Blog https://www.microsoft.com/en-us/microsoft-copilot/blog/2025/08/27/a-smarter-way-to-talk-to-your-tv-microsoft-copilot-launches-on-samsung-tvs-and-monitors/

Introducing gpt-realtime — our best speech-to-speech model for developers, and updates to the Realtime API https://x.com/OpenAI/status/1961110295486808394

Some notes on the gpt-realtime release it replaces chained STT→LLM→TTS with a single speech-in/speech-out model (lower latency, richer nuance) – huge imo 🔥 On benchmarks (vs GPT4o-realtime): > scores 82.8% vs 65.6% on BigBench (reasoning) > 30.5% vs 20.6% on MultiChallenge”” / X https://x.com/reach_vb/status/1961140618295394579

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading