Image created with Flux Pro Ultra. Image prompt: A Minecraft screenshot with blocky note blocks, jukeboxes, and sound wave patterns built from colored blocks rising into the sky, with “AUDIO” written in pixelated Minecraft font across the top
“MusiCoT, a novel chain-of-thought prompting technique, enhances high-fidelity music creation by aligning AI with human musical thought processes. Generate lyrics and music seamlessly across 10 major languages https://x.com/rohanpaul_ai/status/1909583479743856897
ibm-granite/granite-speech-3.2-8b · Hugging Face https://huggingface.co/ibm-granite/granite-speech-3.2-8b
“The WhatsApp MCP server can now send and receive images, videos, and voice notes Combine it with the new ElevenLabs MCP server to give it superpowers — using AI to transcribe the voice notes and send audio messages with 3,000+ voices https://x.com/LukeHarries_/status/1909303780941640041
“Introducing the official ElevenLabs MCP server. Give Claude and Cursor access to the entire ElevenLabs AI audio platform via simple text prompts. You can even spin up voice agents to perform outbound calls for you — like ordering pizza. https://x.com/elevenlabsio/status/1909300782673101265
Introducing Amazon Nova Sonic: Human-like voice conversations for generative AI applications | AWS News Blog https://aws.amazon.com/blogs/aws/introducing-amazon-nova-sonic-human-like-voice-conversations-for-generative-ai-applications/
Amazon’s Nova Sonic foundation model understands voice in a whole new way https://www.aboutamazon.com/news/innovation-at-amazon/nova-sonic-voice-speech-foundation-model
“Amazon launched Nova Sonic speech-to-speech AI for human-like interactions —Outperforms OpenAI’s voice models with ~ 80% less cost —4.2% word error rate across languages — 46.7% better accuracy than GPT-4o for noisy environments —On Amazon Bedrock https://x.com/rowancheung/status/1909845011551633891
Hugging Face and Cloudflare Partner to Make Real-Time Speech and Video Seamless with FastRTC https://huggingface.co/blog/fastrtc-cloudflare
Google’s NotebookLM can now find its own sources | The Verge https://www.theverge.com/news/642490/google-notebooklm-discover-sources-ai-audio-overviews
“Two studies from OpenAI and MIT Media Lab examined how ChatGPT may influence users’ emotions. In a controlled trial with nearly 1,000 participants, researchers found that ChatGPT voice conversations correlated with reduced loneliness and more emotionally expressive https://x.com/DeepLearningAI/status/1909308766756954578
“Google getting ready to ship: Veo-2 Gemini 2.0 Flash live (audio/video chat) Gemini 2.5 Flash preview” / X https://x.com/scaling01/status/1909904138013417878
“Github 👨🔧: TTS Towards Human-Sounding Speech ——- → Leverages a Llama-3b LLM backbone for speech synthesis, showcasing emergent capabilities. → Produces highly natural speech output, focusing on realistic intonation, emotion, and rhythm. → Supports zero-shot voice https://x.com/rohanpaul_ai/status/1909126492971536685
“Google made headlines by making Deep Research available on Gemini 2.5 Pro Exp The move enables Gemini to create superior research reports over rivals Also includes new audio overviews to turn reports into podcast-like conversations! https://x.com/rowancheung/status/1909845062466298148
“And we’re bringing even more helpful AI capabilities to the Workspace tools you use every day, including: 🎧 New audio generation features in Docs ✏️ Help me refine — your personal writing coach in Docs 🎥 High-quality, original video clips in Vids, powered by Veo 2 📊” / X https://x.com/Google/status/1910081783351427166
“Midjourney released V7, the first major update to its image model in almost a year. It includes: —Improved generation quality —Better prompt adherence —A faster and voice-capable Draft Mode to iterate on ideas —Currently in alpha testing phase https://x.com/rowancheung/status/1909154260912124294




