Image created with Flux Pro v1.1 Ultra. Image prompt: Assembly instruction diagram for a vintage stereo receiver with analog dials and wood case, component layout style, 1970s aesthetic, wood grain brown and brushed silver colors, aged cream paper texture, “AUDIO” in retro serif font as manual title, hand-drawn style annotations, vacuum tube placements shown

🚀 Introducing HunyuanVideo-Avatar, a model jointly developed by Tencent Hunyuan and Tencent Music, bringing photos to life. ✅ Upload a photo + audio — auto-detect scene context & emotion, then generate lifelike speech/singing with dynamic visuals. ✅ Supports multi-style, https://x.com/TencentHunyuan/status/1927575170710974560

🎵 Dream come true for content creators! TIGER AI can extract voice, effects & music from ANY audio file 🤯 This lightweight model uses frequency band-split technology to separate speech like magic. Kudos to @fffiloni for the amazing demo! https://x.com/fdaudens/status/1927455842653102291

Built this application 100% with vibe coding using @v0 , @cursor_ai , and 21st from @serafimcloud Meet your new productivity companion for deep focus, ambient sounds, and a beautiful interface. ☕️ https://x.com/birobirobirodev/status/1921720561358573934

Introducing EVI 3 • Hume AI https://www.hume.ai/blog/introducing-evi-3

Launched http://Audiomemo.ai
– Your Clarity Companion https://x.com/Makerealcents/status/1922798540490764717

We’re rolling out voice mode in beta on mobile. Try starting a voice conversation and asking Claude to summarize your calendar or search your docs. https://x.com/AnthropicAI/status/1927463559836877214

bad AI music prompt: “rap about a robot riding a unicorn” good AI music prompt: “Bulgarian 1950s folklore techno with tickled goat skin drums, polyrhythmic throat shouts and glockenspiel arpeggios, cathedral reverb sampled through a 70s landline””” / X https://x.com/fabianstelzer/status/1927649423657521608

wow, @kyutai_labs has been cooking with this one 🔥 Talk to your AI voice clone with just a 10 sec sample. Super low latency + open-source soon. Interesting finding: they suggest that coupling STT+TTS with an LLM actually yields better reasoning and function-calling capabilities https://x.com/fdaudens/status/1925977675606155561

I built a NotebookLM clone using N8N and ChatGPT. It pulls in news via RSS, then generates the script and audio dialog with ChatGPT and text-to-speech. I even did intro music with Udio! https://x.com/markdoylecto/status/1873376687133729205

🚀 We’re open sourcing Chatterbox – our state-of-the-art Voice Cloning model that includes text-to-speech and voice conversion! In recent testing, 63.75% of listeners preferred Chatterbox over ElevenLabs. Not only is it free and open source (MIT license), it’s demonstrably https://www.resemble.ai/chatterbox/

[2505.17862v1] Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities https://arxiv.org/abs/2505.17862v1

Learn how to easily build ChatGPT powered @telegram bot that incorporates speech-to-text and text-to-speech for natural conversations.🚀 Built visually with no-code using @n8n_io and @OpenAI APIs See in the demo the app in action including the ability to be used for learning https://x.com/derekcheungsa/status/1751327704631197915

Unitree hosted a one-hour livestream demonstrating the G1 robot performing fighting moves in response to voice commands. https://x.com/TheHumanoidHub/status/1925811114660548653

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading