Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Minimalist white marble museum gallery with single spotlight illuminating a central pedestal displaying vintage microphone, camera lens, and speaker arranged as unused artifacts, dramatic shadows, architectural emptiness, cold luxury aesthetic, pristine isolation, text overlay ‘MULTIMODALITY’ in bold white sans-serif.

Amazon has launched a new speech-to-speech model, Nova Sonic 2.0, which ranks #2 on our Artificial Analysis Big Bench Audio Speech Reasoning benchmark! The new model achieves a reasoning accuracy score of 87.1% on Big Bench Audio, placing second overall behind Google’s Gemini https://x.com/ArtificialAnlys/status/1995950101068763393

🚨BREAKING: Text Leaderboard Update: A new open source model has landed on the leaderboard! Mistral-Large-3 lands at #6 among open models and #28 overall on the Text leaderboard. Mistral 3 is the next generation of Mistral AI models and their most capable model family to date. https://x.com/arena/status/1995877395510051253

Introducing Mistral 3 | Mistral AI https://mistral.ai/news/mistral-3

Introducing Mistral Code | Mistral AI https://mistral.ai/news/mistral-code

Introducing the Mistral 3 family of models: Frontier intelligence at all sizes. Apache 2.0. Details in 🧵 https://x.com/MistralAI/status/1995872766177018340

Magistral | Mistral AI https://mistral.ai/news/magistral

Mistral Small 3 | Mistral AI https://mistral.ai/news/mistral-small-3

Mistral Small 3.1 | Mistral AI https://mistral.ai/news/mistral-small-3-1

Voxtral | Mistral AI https://mistral.ai/news/voxtral

Aerial intelligence with prompts. Moondream segments pools, tennis courts, and even solar panels, with pixel-perfect accuracy. https://x.com/moondreamai/status/1997058204589871395

Moondream’s new segmentation just dropped. Prompt: “dirty laundry items on the bed.” Moondream: pixel-perfect + actually understands the scene. SAM 3: grabs the floor. https://x.com/moondreamai/status/1996001944838832501

Open-Vocabulary Image Segmentation | Moondream
https://moondream.ai/skills/segment

We’ve always had leading document OCR. Today we’re excited to showcase our infrastructure for letting you build document agents 📑🤖 Our latest release lets you easily build, edit, and deploy a multi-step agentic document workflow directly within LlamaCloud. 1️⃣ Start with https://x.com/jerryjliu0/status/1996349988205637773

Hand tracking gets all the attention in XR, but the real unlock is what you feel when your finger actually touches something. This prototype shows how a single haptic thimble on your fingertip can change the whole experience. Not a glove. Not a full suit. Just a small actuator https://x.com/IlirAliu_/status/1996151860600644008

I really like this idea of uncovering more of a scene as you move the camera. There’s a new aesthetic to it. It’s the first time AI video has felt flexible and truly controllable. The model zero-shots all the items you describe and pairs them with great motion. Whisper Thunder is https://x.com/c_valenzuelab/status/1995539870983266493

This print appears to be happening in midair… a new method of 3D printing: It combines direct ink writing with up-conversion particles-assisted photopolymerization. With this method of printing ceramics, the need for support structures is eliminated. Paper: https://x.com/IlirAliu_/status/1995569849884639713

Wow.. AI assisted keyframe interpolation in the latest release of cascadeur. Define two keyframes manually and get believable animation in between. You can also mix regular, AI and physics-based interpolation – oh and it’s all generated locally: https://x.com/bilawalsidhu/status/1994944476142604435

Visualizing my re-designed living room in 3D Nano banana -> World Labs -> WebAR (threejs) Each of these gaussian splats are only 1.5 mb in size (.spz file type) https://x.com/XRarchitect/status/1995541338335678801

More inference workloads now mix autoregressive and diffusion models in a single pipeline to process and generate multiple modalities – text, image, audio, and video. Today we’re releasing vLLM-Omni: an open-source framework that extends vLLM’s easy, fast, and cost-efficient”” / X https://x.com/vllm_project/status/1995566791234629989

Wispr Flow | Effortless Voice Dictation https://wisprflow.ai/

Gemini 3 Pro is the frontier of multimodal AI, delivering SOTA performance across document, screen, spatial, and video understanding. Read our deep dive on how we’ve pushed our core capabilities to power hero use cases across: + Docs: “”derender”” complex docs into structured https://x.com/googleaidevs/status/1996973083467333736

Google out here building the Borg cube for real https://x.com/bilawalsidhu/status/1995650915785986491

Happy to share that the @GoogleDeepMind Gemini team is starting a new research team in Singapore! This new team will be focused on advanced reasoning, LLM/RL and improving bleeding edge SOTA models such as Gemini, Gemini Deep Think and beyond. 🔥 This team will be led by yours https://x.com/YiTayML/status/1996640869584445882

I was in Singapore earlier this year to visit the office, and this is going to be a very-high impact part of the Gemini team! If you’re interested in working on Gemini and want to be in Singapore working with awesome people like @YiTayML and @quocleix, see below ⬇️”” / X https://x.com/JeffDean/status/1996644208854388983

Opera rolls out Gemini-powered AI features across its browsers – 9to5Mac https://9to5mac.com/2025/12/01/opera-browsers-get-google-gemini-integration/

Our Gemini 3 Vibe Code hackathon started!, Build applications using the new Gemini 3 Pro model with a price pool of $500k. 🤯 > Top 50 winners receive $10,000 in Gemini API credits each. > Access Gemini 3 Pro Preview directly in Google AI Studio. > Leverage advanced reasoning”” / X https://x.com/_philschmid/status/1996990062836244732

Take an early look at how Google Gemini projects will work – Android Authority https://www.androidauthority.com/google-gemini-projects-2-3620950/

Today, we’re rolling out an updated Deep Think mode available in the Gemini app for Google AI Ultra subscribers. Here’s what you need to know: — Gemini 3 Deep Think mode pushes the boundaries of intelligence even further, delivering meaningful improvement in reasoning https://x.com/GoogleAI/status/1996657213390155927

Today, we’re rolling out an updated Deep Think mode available in the Gemini app for Google AI Ultra subscribers. Here’s what you need to know: — Gemini 3 Deep Think mode pushes the boundaries of intelligence even further, delivering meaningful improvement in reasoning https://x.com/GoogleAI/status/1996657213390155927?s=20

Ultra users, ready to try Gemini 3 Deep Think mode? Here’s how: 1) Select ‘Deep Think’ in the prompt bar 2) Select ‘Thinking’ from the model drop down 3) Type your prompt & submit”” / X https://x.com/GeminiApp/status/1996670867770953894

We’re hiring research scientists & student researchers at Google DeepMind. DM or email me if you’re interested! I’ll be at NeurIPS this week. Happy to chat in person!”” / X https://x.com/RuiqiGao/status/1995572419218796567

We’re pushing the boundaries of intelligence even further with Gemini 3 Deep Think. 🧠 This mode meaningfully improves reasoning capabilities by exploring many hypotheses simultaneously to solve problems. Here’s how it coded a simulated dominoes game from a single prompt ⬇️ https://x.com/GoogleDeepMind/status/1996658401233842624

With state-of-the-art reasoning, richer visuals, and deeper interactivity, Gemini 3 is more intuitive, more powerful, and more personalized. Start exploring at https://x.com/GeminiApp/status/1995534313044238347

Depth Anything 3 can reconstruct this FPV video in just a few seconds on a A100 🤯 It was not long ago that I used to let agisoft metashape chug all night on a 3d scan, and here we are https://x.com/bilawalsidhu/status/1996354738078752987

Microsoft just released VibeVoice-Realtime-0.5B https://x.com/_akhaliq/status/1996602953885499466

VibeVoice https://microsoft.github.io/VibeVoice/

And Mistral Large 3, a frontier class open source MoE. https://x.com/MistralAI/status/1995872771516354828

🎉 Congratulations to the Mistral team on launching the Mistral 3 family! We’re proud to share that @MistralAI, @NVIDIAAIDev, @RedHat_AI, and vLLM worked closely together to deliver full Day-0 support for the entire Mistral 3 lineup. This collaboration enabled: • NVFP4 https://x.com/vllm_project/status/1995890057224618154

Europe still has one frontier model maker that can generally keep pace with Chinese open weights models, though no reasoner for Mistral 3 yet means they are behind the curve of actual performance – DeepSeek r1 got 71.5% on GPQA Diamond (& 1-shot, not 5-shot) back in January. https://x.com/emollick/status/1996068920596594932

I want to especially thank @MistralAI for releasing the base models for Mistral 3. Fewer companies are sharing base models and this opens many use cases from custom instruct to non-instruct cases”” / X https://x.com/QuixiAI/status/1996272948378804326

Meet the Ministral 3 models from @MistralAI! – 3B, 8B, and 14B models – Instruct, reasoning, and base variants – Supports tool use and vision input – Open-weights, Apache 2.0 licensed https://x.com/lmstudio/status/1995908228526604451

Mistral 3 is now available on Ollama v0.13.1 (currently in pre-release on GitHub). 14B: ollama run ministral-3:14b 8B: ollama run ministral-3:8b 3B: ollama run ministral-3:3b Please update to the latest Ollama. https://x.com/ollama/status/1995885696360566885

Mistral releases Ministral 3, their new reasoning and instruct models! 🔥 Ministral 3 comes in 3B, 8B, and 14B with vision support and best-in-class performance. Run the 14B models locally with 24GB RAM. Guide + Notebook: https://x.com/UnslothAI/status/1995874975631503479

NEW: @MistralAI released a fantastic family of multimodal models, Ministral 3. You can fine-tune them for free on Colab using TRL ⚡️, supporting both SFT and GRPO https://x.com/SergioPaniego/status/1996257877871509896

NEW: @MistralAI releases Mistral 3, a family of multimodal models, including three start-of-the-art dense models (3B, 8B, and 14B) and Mistral Large 3 (675B, 41B active). All Apache 2.0! 🤗 Surprisingly, the 3B is small enough to run 100% locally in your browser on WebGPU! 🤯 https://x.com/xenovacom/status/1995879338583945635

Run Mistral Large 3 on Ollama’s cloud: ollama run mistral-large-3:675b-cloud”” / X https://x.com/ollama/status/1996682858933768691

Super nice to see Mistral Large 3 as the #1 OSS model for coding on lmarena 🥳😎🙌 And the spoiler alert! 👀👀”” / X https://x.com/sophiamyang/status/1996587296666128398

Support for running Mistral Large 3 locally will be available in Ollama soon.”” / X https://x.com/ollama/status/1996683156817416667

The Bert-Nebulon Alpha Stealth model is live now as @MistralAI’s new Mistral Large 3! Try the full release now on OpenRouter: https://x.com/OpenRouterAI/status/1995904288560988617

The world’s best small models–Ministral 3 (14B, 8B, 3B), each released with base, instruct and reasoning versions. https://x.com/MistralAI/status/1995872768601325836

Mistral Large 3 debuts as the #1 open source coding model on the @arena leaderboard. We’d love for you to try it! More on coding in a few days… 👀 https://x.com/MistralAI/status/1996580307336638951

Mistral AI raises 1.7B€ to accelerate technological progress with AI | Mistral AI https://mistral.ai/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai

Elicit now understands figures! Elicit is the first AI tool that can systematically parse, interpret, and extract data from figures across thousands of papers. That includes Kaplan-Meier curves, heatmaps, reaction schemes, and microscopy images. Figures contain critical https://x.com/elicitorg/status/1995926919783862369

We’re building out an applied research team to push SOTA on document understanding using LLMs/VLMs and other emerging techniques 📈📑 We’re on a mission to understand and orchestrate the most complex document types, from PDFs to Excel. You’re responsible for research, evals, and https://x.com/jerryjliu0/status/1997048645817192638

Laying the Foundations for Visual Intelligence–Our $300M Series B | Black Forest Labs https://bfl.ai/blog/our-300m-series-b

.@Stanford researchers showed what happens when you shrink a multimodal model. They look specifically at how reducing the size of the LLM inside a multimodal model affects the model’s overall abilities. ➡️ The part that suffers most is vision. And perception really collapses. https://x.com/TheTuringPost/status/1994548273387032753

VPS or visual positioning system… this is how AR glasses and robots will understand where they are in the real world and know where to go. It’s also how you can spatially annotate reality with cm level accuracy – all using a machine readable model of the world. https://x.com/bilawalsidhu/status/1995959232714473966

Document understanding is a huge use case for VLMs, but historically there’s been no single “”good”” benchmark to measure progress here (unlike SWE-bench for coding). This past week I did a deep dive into OlmOCR-Bench, a recent document OCR benchmark that is a huge step in the https://x.com/jerryjliu0/status/1996668513562644823

with transformers v5 RC comes `any-to-any` pipeline and a new model class: AutoModelForMultimodalLM 👏 these unlock models that take in 2+ inputs and 2+ outputs, like Gemma3n (all modalities to text) and Qwen3-Omni (all modalities to text+audio) docs on the next one 🙌🏻 https://x.com/mervenoyann/status/1996908863673737450

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading