“AskPerplexity can now ingest videos and offer explanations! https://x.com/AravSrinivas/status/1901840001866023146

SmolDocling is based on SmolVLM, trained on pages and Docling transcriptions (new DocTags format to capture the elements and locality better)
https://x.com/mervenoyann/status/1901668064602579150

“Announcing @MistralAI Small 3.1: multimodal, multilingual, Apache 2.0, the best model in its weight class. 💻 Lightweight: Runs on a single RTX 4090 or a Mac with 32GB RAM, perfect for on-device applications. 🗣️ Fast-Response Conversations: Ideal for virtual assistants and other https://x.com/sophiamyang/status/1901675671815901688

“🤯 Gemma 3’s image analysis blew me away! Tested 2 ways to extract airplane registration numbers from photos with 12B model: 1️⃣ @Gradio app w/API link (underrated feature IMO) + ZeroGPU infra on @huggingface in Google Colab. Fast & free. 2️⃣ @lmstudio server + local processing https://x.com/fdaudens/status/1900285203135987943

“Introducing the Gemma package, a minimalistic library to use and fine-tune Gemma 🔥 Including docs on: – Fine-tuning – Sharding – LoRA – PEFT – Multimodality – Tokenization !pip install gemma https://x.com/osanseviero/status/1902456220876787763

“Fresh: SmolDocling 🔥 state-of-the-art open-source lightning-fast OCR 📑 read single document in 0.35 seconds (on A100) using 0.5GB VRAM it’s a 256M model that beats every model (even ones 27x larger, including Qwen2.5VL 🤯) https://x.com/mervenoyann/status/1901668060257190186

“Microsoft launched Phi-4-multimodal, a high-performing open weights model capable of processing text, images, and speech simultaneously. Built on a transformer architecture with 5.6 billion parameters, the model delivers impressive results in speech transcription, image https://x.com/DeepLearningAI/status/1902372844421460458

“Introducing #Ray2 Flash—3x faster, 3x cheaper new model. Flash brings Ray’s frontier production-ready Text-to-Video, Image-to-Video, audio, and control capabilities with high quality and speed to all subscribers—so you can create more, faster, and without limits. Available now. https://x.com/LumaLabsAI/status/1898056614684381218

“Kokoro-82M v1.0 is now the leading open weights Text to Speech Model in Artificial Analysis Speech Arena! Kokoro currently scores a Speech Arena ELO of 1086, just behind proprietary speech models from OpenAI, Cartesia, and ElevenLabs. For the first time, this puts an open https://x.com/ArtificialAnlys/status/1902762871106441703

“Tweet out your https://t.co/fwYIeucwEn TTS creations (hit “share”) to @OpenAIDevs. The top three most creative ones will win a Teenage Engineering OB-4. Keep it to ~30 seconds, and be creative with voice, delivery, pronunciation, or even tone changes mid-script. Ends Friday!” / X https://x.com/kevinweil/status/1902769865254903888

“🚀 Sesame Labs’ CSM-1b takes #3 on @huggingface! This open-source beast packs 1M hours of training, real-time synthesis, voice cloning & emotional intelligence – all Apache 2.0 licensed. The open-source AI race is on fire! 🔥 https://x.com/fdaudens/status/1900504351472259534

Introducing next-generation audio models in the API | OpenAI https://openai.com/index/introducing-our-next-generation-audio-models/

“Alright, this pretty WILD! Orpheus 3B – high quality, emotive Text to Speech – Apache 2.0 licensed! 🔥 > Zero shot voice cloning > Natural & emotive speech > Controllable intonation > Trained on 100K hours of audio > Input AND output streaming > 100ms latency 🤯 > Easy to https://x.com/reach_vb/status/1902445501427114043

“19/ Companies are already using the tooling from Oracle + Nvidia – Real-world implementations already in use by Soley Therapeutics (drug discovery) – Pipefy (business process automation) – DeweyVision (media data processing) – SoundHound (voice AI)” / X https://x.com/AtomSilverman/status/1902087726985572390

New Gemini features: Canvas and Audio Overview https://blog.google/products/gemini/gemini-collaboration-features/

“We’re also holding a radio contest. 📻 Tweet out your https://t.co/1kI0L2V1jW TTS creations (hit “share”). The top three most creative ones will win a Teenage Engineering OB-4. Keep it to ~30 seconds, and be creative with voice, delivery, pronunciation, or even tone changes” / X https://x.com/OpenAIDevs/status/1902773659497885936

“Luma Labs released Ray2 Flash, a 3x faster and more affordable version of its video generation model It supports text-to-video and image-to-video with audio and advanced control options Now available to all Luma subscribers, without any limits! https://x.com/adcock_brett/status/1901303510555152869

“Three new state-of-the-art audio models in the API: 🗣️ Two speech-to-text models—outperforming Whisper 💬 A new TTS model—you can instruct it *how* to speak 🤖 And the Agents SDK now supports audio, making it easy to build voice agents. Try TTS now at https://x.com/OpenAIDevs/status/1902773579323674710

“@OpenAI MOAR AUDIO – LETSGOOO! – don’t even care if it’s open weights or not, the community follows what OAI/ big labs do!” / X https://x.com/reach_vb/status/1902741809010295197

“Lots of new audio stuff today: – ASR: gpt-4o-transcribe with SoTA performance – TTS: gpt-4o-mini-tts with playground at https://x.com/juberti/status/1902771172615524791?s=46

“🔊 Three new audio models for you today! * A new text to speech model that gives you control over timing and emotion—not just what to say, but how to say it * Two speech to text models that meaningfully outperform Whisper All available in the API + fully integrated into our” / X https://x.com/kevinweil/status/1902769861484335437

“The amount of capability overhang in current AI systems is hard to overstate, even in narrow areas like vision & image creation. If AI development stopped today (and no indication that is happening), we have a couple decades of figuring out how to integrate it into work & life.” / X https://x.com/emollick/status/1902020093229346906

“Mistral have released Mistral Small 3.1, adding image input and a 128k token context window to Mistral Small 3 Key results and info: ➤ @MistralAI Small 3.1 scores an Artificial Analysis Intelligence Index of 35, in line with Mistral 3 and other models such as GPT-4o mini and https://x.com/ArtificialAnlys/status/1902017023917666351

ImaginTalk https://imagintalk.github.io/

Mistral Small 3.1 | Mistral AI https://mistral.ai/news/mistral-small-3-1

“🚀 Just dropped: SmolDocling-256M brings powerful OCR to your local machine! Only 256M in size but packs full document processing power. https://x.com/fdaudens/status/1901755195505160678

“Nice video on @MistralAI Small 3.1 from @1littlecoder 🫶 https://x.com/sophiamyang/status/1902038297620443612

“🔥 Mistral-Small-3.1 (24B) is already #3 on @huggingface after just 1 day! Brings SOTA vision + 128k context while fitting on a 32GB MacBook. Crushes benchmarks in reasoning, multilingual & visual tasks. Apache 2.0 licensed 🚀 https://x.com/fdaudens/status/1902111100503572865

“@MistralAI @huggingface Available on @MistralAI La Plateforme `mistral-small-latest` : https://x.com/sophiamyang/status/1901677125918134276

“Introducing: ShieldGemma 2 – a 4B model for image safety classification 🛡️ 👀Use as input filter for VLMs ❌or for blocking dangerous image generation outputs 🦺Great for production deployments Blog: https://x.com/osanseviero/status/1901764379328037047

“Effortlessly Extract Appointments from Files and Websites 🎉 Every semester, my kids’ school sends PDFs packed with important dates—manually copying them to Outlook was exhausting! 🤯 So, as a non-developer, I built a Minimum Lovable Product (MLP) using @lovable_dev (amazing https://x.com/_DBrugger/status/1899530386524266697

“We’ve just unveiled ERNIE 4.5 & X1! 🚀 As a deep-thinking reasoning model with multimodal capabilities, ERNIE X1 delivers performance on par with DeepSeek R1 at only half the price. Meanwhile, ERNIE 4.5 is our latest foundation model and new-generation native multimodal model. https://x.com/Baidu_Inc/status/1901089355890036897

Llama4 is probably coming next month, multi modal, long context : r/LocalLLaMA https://www.reddit.com/r/LocalLLaMA/comments/1jes8ue/llama4_is_probably_coming_next_month_multi_modal/

“We’ve just unveiled ERNIE 4.5 & X1! 🚀 As a deep-thinking reasoning model with multimodal capabilities, ERNIE X1 delivers performance on par with DeepSeek R1 at only half the price. Meanwhile, ERNIE 4.5 is our latest foundation model and new-generation native multimodal model. https://x.com/Baidu_Inc/status/1901089355890036897

“Mistral AI released Small 3.1, a SOTA multilingual and multimodal LLM —24B (can run on a laptop) —128k token context window —Outperforms Gemma 3 and GPT-4o Mini on most benchmarks —Inference speed of 150 tokens/sec —Open-source under Apache 2.0 license https://x.com/rowancheung/status/1901887637465809285

“Introducing Mistral Small 3.1. Multimodal, Apache 2.0, outperforms Gemma 3 and GPT 4o-mini. https://x.com/MistralAI/status/1901668499832918151

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading