“Happy to share you my latest hack for @ElevenLabsDevs hackathon 🤩 Gamepal is the old-time gamer hotline but powered by modern technologies! Built with @lovable_dev @ElevenLabsDevs @ExaAILabs đź«¶ Try it out 👇 https://x.com/picsoung/status/1893811252075368513
Text-to-Speech Generator with 450+ AI Voices https://podcastle.ai/ai-voices
Sesame on X: “At Sesame, we believe in a future where computers are lifelike. Today we are unveiling an early glimpse of our expressive voice technology, highlighting our focus on lifelike interactions and our vision for all-day wearable voice companions. https://t.co/Edp8V8urgC https://t.co/Mc5nWnBJZM” / X
https://x.com/sesame/status/1895159087010324615
“FINALLY! Generate a full song with lyrics in < 20 seconds! ⚡ 🔥 DiffRhythm is ⟡ just out ⟡ an open weights end-to-end full song generation model that generate 1-2min songs in just a few seconds 🏎️💨 Give it a reference + lyrics and get a song back! Sound on! 🔊 ▶️ https://x.com/multimodalart/status/1896862125659988322
“@elevenlabsio Hackathon – Summary In 12 hours, our 2-person (one alsmost non technical) team built Neighbour – a voice-powered assistant for seniors, leveraging Eleven Labs technology. The app is built with Next.js, using Supabase as the database. Lovable handled the coding. • https://x.com/im_the_kk/status/1893742399685378426
DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion https://aslp-lab.github.io/DiffRhythm.github.io/
“This could be huge for AI storytelling. I’ve tried so many AI lipsync tools and have never been satisfied. This is my first try with @hedra_labs Character-3 and it’s blown me away. It was so easy. I uploaded a track from @SunoMusic and an image from @midjourney and it just https://x.com/TomLikesRobots/status/1898009257598980587
[2503.01183] DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion https://arxiv.org/abs/2503.01183
“At Sesame, we believe in a future where computers are lifelike. Today we are unveiling an early glimpse of our expressive voice technology, highlighting our focus on lifelike interactions and our vision for all-day wearable voice companions. https://x.com/sesame/status/1895159087010324615
“This is wild, open suno/udio is here DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion video by @multimodalart https://x.com/_akhaliq/status/1896938481911542002
“The new AI voice from Sesame really is a powerful illustration of where AI is going. This is all real-time, from my browser. Excellent use of disfluencies, pauses, even intakes of breathe really make this seem like a human, though bits of uncanniness remain, at least for now. https://x.com/emollick/status/1896757383566950466
“Introducing Scribe — the most accurate Speech to Text model. It has the highest accuracy on benchmarks, outperforming previous state-of-the-art models such as Gemini 2.0 and OpenAI Whisper v3. It’s now the leading model for English, Spanish, Italian, and many more. With support https://x.com/elevenlabsio/status/1894821477230485570
“Real time voice mode now works on Mac app too! Chat while you’re working! It will keep listening to you in the background so you can have it on and keep working or doing whatever you were doing. Update the Mac App. And the shortcut is Cmd + Shift + M. https://x.com/AravSrinivas/status/1897408183620264028
“🚀 Phi-4 just dropped & it’s now #1 on the Open ASR Leaderboard! 🏆🔥 But you can make it even better! 🎯 Fine-tune it for your specific use case—whether it’s handling noisy audio, improving accuracy in low-resource languages, or custom domain adaptation! Try it out with this” / X https://x.com/Tu7uruu/status/1895161283743490548
“Continuing from last week’s post on the rise of the Voice Stack, there’s an area that today’s voice-based systems often struggle with: Voice Activity Detection (VAD) and the turn-taking paradigm of communication. When communicating with a text-based chatbot, the turns are clear:” / X https://x.com/AndrewYNg/status/1897776017873465635
Stability AI and Arm Bring On-Device Generative Audio to Smartphones — Stability AI https://stability.ai/news/stability-ai-and-arm-bring-on-device-generative-audio-to-smartphones
“Introducing Poe Apps: a new, easy way to create and use visual interfaces into any combination of the 100+ text, image, video, and audio models on Poe. (1/5) https://x.com/poe_platform/status/1894435707814637741
“🔥 Phi-4-multimodal-instruct just landed on the Hugging Face Speech Recognition Leaderboard — and it’s ranked #1 with an impressive 6.14% WER! Big moves in ASR — check out how it compares to Whisper, Canary & more đź‘€ https://x.com/Tu7uruu/status/1896948226743558530
“Today, we’re releasing Octave: the first LLM built for text-to-speech. 🎨Design any voice with a prompt 🎬 Give acting instructions to control emotion and delivery (sarcasm, whispering, etc.) 🛠️Produce long-form content on our Creator Studio Unlike traditional TTS that just https://x.com/hume_ai/status/1894833497824481593
“BOOM! Phi 4 Multimodal (MIT licensed) – the new king of the Open ASR Leaderboard đź’Ą Beats Nvidia Canary, OpenAI Whisper and more 🤩 Bonus: the model can do much more – speech summarisation, diarization and doubles up as an Audio LM too! https://x.com/reach_vb/status/1897014754943910266




