Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A centered golden clockwork nightingale of hammered gold and silver filigree perched on a metal bough, its open beak releasing concentric spiral sound-rings made of glittering tesserae across a Byzantine gold-ground mosaic apse, small flattened stylized mosaic saints in the margins with hands cupped to their ears, warm candlelit sacral glow, burnished gold with deep imperial purple and Tyrian crimson accents, the bold ivory Trajan-capital title ‘AUDIO’ rendered across the lower third with clean separation, painterly tactile mosaic grain, 16:9 full-bleed.
Introducing Music v2, our groundbreaking new music model
https://elevenlabs.io/blog/introducing-music-v2
Announcing AA-WER Streaming, our new benchmark measuring streaming Speech to Text models on accuracy and latency for voice agent use cases. Pareto optimal models on this new benchmark include those from Cartesia, ElevenLabs, and Deepgram Streaming Speech to Text (STT) powers
https://x.com/ArtificialAnlys/status/2060021901234458958
Today we are launching Music v2. Better vocals, instrumentation, and arrangement across every genre, improved multilingual support plus capabilities that weren’t possible before.
https://x.com/ElevenLabs/status/2059312414198235642?s=20
WavFlow: Audio Generation in Waveform Space”” TL;DR: generates high-fidelity audio directly in raw waveform space using flow matching, eliminating the need for latent compression while matching SOTA audio generation quality
https://x.com/Almorgand/status/2057801315028271352
Sonic 3.5 is now the #1 text to speech model on the @ArtificialAnlys leaderboard! You no longer have to trade off quality and latency – Sonic 3.5 also has the fastest time to first audio at 82ms end to end. See full benchmark results 👇
https://x.com/cartesia/status/2057880195403800633
Granola — The AI Notepad for back-to-back meetings
https://www.granola.ai/?via=adops-tldr-tech&dub_id=itx1yUOnpPaKx3A8





Leave a Reply