Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A wide off-white cover with a single horizontal rainbow soundwave rippling symmetrically outward in nested concentric ROYGBIV bands — red at the core rising through orange, yellow, green, blue to violet at the crest — and the word ‘AUDIO’ centered within the swell, its bold letters constructed from the identical layered-spectrum line-work, flat matte finish, crisp printed edges, generous negative space, Julio Le Parc kinetic op-art style.

The Next Chapter for Suno · Suno
https://suno.com/blog/series-d-announcement

Introducing Gemma 4 12B
https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/

Today we’re introducing Gemma 4 12B — our latest open model that brings advanced agentic reasoning, vision and audio directly to your laptop. It delivers performance nearing our larger Gemma models with a much smaller total memory footprint, while being small enough to run
https://x.com/Google/status/2062203526588088452

Watch me control my computer with just my voice. This is the future of operating systems. No hands. GPT-Realtime 2.0 is very, very underrated. Demo:
https://x.com/FarzaTV/status/2060865350036750847

Miso One is live: an open-weights voice model built to sound like a real person reading, with actual warmth and pacing where most TTS still goes flat. 8B params, free on GitHub, with one-shot voice cloning from a short sample at 110ms latency. Self-host it and your audio data
https://x.com/kimmonismus/status/2062210845308780639

Google just released Magenta RealTime 2 on Hugging Face The only open-weights model for real-time continuous music generation on device. Steer it with text, audio, or MIDI at ~200ms latency.
https://x.com/HuggingPapers/status/2062260306039259236

2-bit Gemma 4 12B GGUF, only 4.66 GB on disk, managed to cite 15 sites from a single prompt. Try this locally on >6GB RAM via Unsloth Studio. GitHub:
https://x.com/UnslothAI/status/2062470072179044447

Congrats to the @googlegemma team on the Gemma 4 12B launch 🎉 Day-0 support on vLLM is ready to go. It’s an encoder-free unified multimodal model — text, image, audio, and video all project straight into the LLM’s embedding space, no separate vision or audio towers. 256K
https://x.com/vllm_project/status/2062228047324201166

For the past years my research focus was on unifying models and training paradigms across modalities. Today I’m excited that we’re releasing our latest model aligned with this theme: Gemma 4 12B, a dense encoder-free model which processes raw text, image, and audio inputs! 1/
https://x.com/mtschannen/status/2062236357351579915

Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs. Google’s new model, Gemma 4 12B Unified supports image, audio and 256K context. You can run and train the model via Unsloth Studio. GGUF:
https://t.co/8cL321pVDh Guide:
https://x.com/UnslothAI/status/2062207258810053084

Meet Gemma 4 12B! A unified, encoder-free multimodal model designed to bring high-performance intelligence directly to your laptop, and released under an Apache 2.0 license. Bridging the gap between edge efficiency and advanced reasoning. Here is what’s new with Gemma 4 12B: 👇
https://x.com/googlegemma/status/2062202706882883696

Our new unified architecture allows Gemma 4 12B to process multimodal inputs natively. Here’s how ⬇️ Traditional models rely on separate encoders for images and audio. This adds latency and increases memory usage. So we streamlined this: 👁️ Vision: We took a novel approach to
https://x.com/Google/status/2062203532351090824

Today we’re shipping our biggest MLX-VLM release yet: v0.6.0 …and we are raising 💸 This one’s about turning your Apple devices into real local agent machines. From your desk to your pocket. What’s new: ⚡ Speculative decoding everywhere — Gemma 4 EAGLE3 + DFlash, Qwen
https://x.com/Prince_Canuma/status/2061541992790683726

We released Gemma 4 12B yesterday. Here is a visual guide that explains the full architecture. → How encoders typically connect modalities to LLMs → Why Gemma 4 removed the vision and audio encoders → How a single 12B model can handle text, images, and audio without
https://x.com/_philschmid/status/2062546814075609413

We’re launching Gemma 4 12B: Our unified, encoder-free model that brings powerful multimodal intelligence straight to your laptop 🚀 The model bridges the gap between our mobile E4B model and larger 26B MoE models, packaging frontier-class reasoning and native audio into a
https://x.com/googleaidevs/status/2062204432658386950

first ship on the new team: new docs for agents on Cloudflare! 🚀 really helps to visualize the breadth that @Cloudflare agent platform has to offer: Communication channels: – chat/slack/webhook agents – voice agents with voice – email agents with email service Core agent: –
https://x.com/thomasgauvin/status/2062512156076048447

Microsoft has released MAI-Transcribe-1.5: an exceptionally fast speech transcription model at a speed factor of ~276x, while still achieving 2.4% on AA-WER (#3), leading the accuracy-speed Pareto frontier MAI-Transcribe-1.5 is Microsoft AI (MAI)’s latest speech transcription
https://x.com/ArtificialAnlys/status/2061878491860324402

Microsoft introduces MAI-Thinking-1 It’s a 1T@35B parameter model pre-trained on 30T tokens with a maximum context length of 256k tokens using 8192 GB200 GPUs. Based on benchmarks it seems to be around GLM-5 level. Microsoft also released a comprehensive 109 pages tech-report:
https://x.com/scaling01/status/2061889624847343825

microsoft used gepa / dspy to tune the LLM judge prompt for quality scoring. @lateinteraction stays winning. from the mai-thinking-1 report
https://x.com/bj2rn/status/2061941109828301241

Today we’re announcing MAI-Thinking-1 with Microsoft and it will be available on Baseten soon. Microsoft built something genuinely different here: a commercial-grade thinking model trained on clean data with no distillation from third-party models and designed to be fine-tuned
https://x.com/tuhinone/status/2061879239817969756

Nemotron 3.5 ASR is built for streaming multilingual speech recognition and voice agents. One 0.6B checkpoint. 40 language-locales. Sub-100ms latency. Cache-aware FastConformer carries context forward between chunks instead of repeatedly reprocessing overlapping audio. Try it:
https://x.com/togethercompute/status/2062520605102993436

Open speech models for real-time voice agents. @NVIDIAAI Nemotron 3.5 ASR is now available on fal. An open streaming speech recognition model supporting 40 language-locale combinations with ultra-low latency, native punctuation, and capitalization. Built for voice agents,
https://x.com/fal/status/2062521027020611933

Second big release from us today: Nemotron-3.5-ASR-Streaming! 🌎40 languages ⚡️80ms – 1s controllable latency 🔥240 – 2400 concurrent streams on 1xH100 🧱FastConformer Cache-Aware RNN-T architecture
https://x.com/PiotrZelasko/status/2062538923776290909

OpenAI just dropped a completely new kind of model gpt-realtime-translate takes in speech audio from any language and outputs speech in your target language LLMs are great, but you need specialized models for specialized use cases We’re running this on our smart glasses
https://x.com/caydengineer/status/2060426641701269917

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading