We’ve made a step change quality improvement in our voice changer model, now available in playground and API. It’s now state of the art, the style transfer abilities are amazing. Record and turn your voice into an amazing character in one click. / X https://x.com/krandiash/status/1892725226498359365

Today, we’re excited to announce a beta release of Zonos, a highly expressive TTS model with high fidelity voice cloning. We release both transformer and SSM-hybrid models under an Apache 2.0 license. Zonos performs well vs leading TTS providers in quality and expressiveness. https://x.com/ZyphraAI/status/1888996367923888341

This video has been entirely generated with AI (face and voice) on the last version on Argil AI 🤯 I have known @LaodisOfficial for years, and for the first time, we have reached the point where the difference between going into a studio and using AI is no different. Do you https://x.com/BrivaelLp/status/1890311661241749821

Spotify Opens Up Support for ElevenLabs Audiobook Content — Spotify https://newsroom.spotify.com/2025-02-20/spotify-opens-up-support-for-elevenlabs-audiobook-content/

Learn everything there is to know about LlamaParse in this comprehensive video! In this video, @mesudarshan covers: ➡️ Multiple parsing modes ➡️ Using parsing instructions to improve quality ➡️ Output formats available ➡️ Parsing audio and images ➡️ JSON mode ➡️ Using it all in https://x.com/llama_index/status/1890499579214491967
stepfun-ai/

Improvements will happen rapidly and almost daily according to the team. There is also a Grok-powered voice app coming too — about a week away! / X https://x.com/omarsar0/status/1891715813956108699

@WholeMarsBlog Voice mode is still a little patchy, so probably launches in about a week, but it’s awesome / X https://x.com/elonmusk/status/1891676673898119254

Step-Audio a 130 billion parameter multimodal LLM that is responsible for understanding and generating human speech https://x.com/_akhaliq/status/1891528348590833834

No system card for Grok 3 yet, so no perspectives on risk mitigation. This is especially key for voice, and is why labs have been slow with full multimodal, you can imitate anyone’s voice, and also the AI tended to take your voice and repeat it back to you. From 4o system card: https://x.com/emollick/status/1891745345392058496

BREAKING: Grok’s voice unveiled https://x.com/teslaownersSV/status/1891719294469222495

📱 Turn any text into a podcast instantly! Transform articles, papers & blogs into audio content using open-source AI (deepseek-r1) + kokoro TTS. Think NotebookLM but fully open source 🎧 Nice work @ngxson! https://x.com/fdaudens/status/1891690883176604053

Step-Audio-Chat · Hugging Face https://huggingface.co/stepfun-ai/Step-Audio-Chat

chat, is this for real? A 132B parameter end to end Speech LM??? voice in, voice out 🤯 APACHE 2.0 LICENSED?? https://x.com/reach_vb/status/1891517368603492697

Organizational life is about to get much weirder. This paper creates an early form of meeting delegates, where you send an AI to a meeting on your behalf, and it uses your voice and knowledge to advance your agenda A lot of old organizational methods need to be rethought for AI https://x.com/emollick/status/1891527817826828565

“We’ve made a step change quality improvement in our voice changer model, now available in playground and API. It’s now state of the art, the style transfer abilities are amazing. Record and turn your voice into an amazing character in one click.” / X https://x.com/krandiash/status/1892725226498359365

“Today, we’re excited to announce a beta release of Zonos, a highly expressive TTS model with high fidelity voice cloning. We release both transformer and SSM-hybrid models under an Apache 2.0 license. Zonos performs well vs leading TTS providers in quality and expressiveness. https://x.com/ZyphraAI/status/1888996367923888341

“BREAKING: Grok’s voice unveiled https://x.com/teslaownersSV/status/1891719294469222495

“This video has been entirely generated with AI (face and voice) on the last version on Argil AI 🤯 I have known @LaodisOfficial for years, and for the first time, we have reached the point where the difference between going into a studio and using AI is no different. Do you https://x.com/BrivaelLp/status/1890311661241749821

“📱 Turn any text into a podcast instantly! Transform articles, papers & blogs into audio content using open-source AI (deepseek-r1) + kokoro TTS. Think NotebookLM but fully open source 🎧 Nice work @ngxson! https://x.com/fdaudens/status/1891690883176604053

Spotify Opens Up Support for ElevenLabs Audiobook Content — Spotify https://newsroom.spotify.com/2025-02-20/spotify-opens-up-support-for-elevenlabs-audiobook-content/

“Learn everything there is to know about LlamaParse in this comprehensive video! In this video, @mesudarshan covers: ➡️ Multiple parsing modes ➡️ Using parsing instructions to improve quality ➡️ Output formats available ➡️ Parsing audio and images ➡️ JSON mode ➡️ Using it all in https://x.com/llama_index/status/1890499579214491967

stepfun-ai/Step-Audio-Chat · Hugging Face https://huggingface.co/stepfun-ai/Step-Audio-Chat

“chat, is this for real? A 132B parameter end to end Speech LM??? voice in, voice out 🤯 APACHE 2.0 LICENSED?? https://x.com/reach_vb/status/1891517368603492697

“Improvements will happen rapidly and almost daily according to the team. There is also a Grok-powered voice app coming too — about a week away!” / X https://x.com/omarsar0/status/1891715813956108699

“@WholeMarsBlog Voice mode is still a little patchy, so probably launches in about a week, but it’s awesome” / X https://x.com/elonmusk/status/1891676673898119254

“Step-Audio a 130 billion parameter multimodal LLM that is responsible for understanding and generating human speech https://x.com/_akhaliq/status/1891528348590833834

“Organizational life is about to get much weirder. This paper creates an early form of meeting delegates, where you send an AI to a meeting on your behalf, and it uses your voice and knowledge to advance your agenda A lot of old organizational methods need to be rethought for AI https://x.com/emollick/status/1891527817826828565

“No system card for Grok 3 yet, so no perspectives on risk mitigation. This is especially key for voice, and is why labs have been slow with full multimodal, you can imitate anyone’s voice, and also the AI tended to take your voice and repeat it back to you. From 4o system card: https://x.com/emollick/status/1891745345392058496

“Audiobox Aesthetics is a model for unified automatic quality assessment for speech, music and sound. Try the demo on @huggingface ➡️ https://x.com/AIatMeta/status/1893009390980170001

“Nothing better than user love!! Congrats @HeyGen_Official https://x.com/saranormous/status/1892945844376047679

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading