“Audiobox Aesthetics is a model for unified automatic quality assessment for speech, music and sound. Try the demo on @huggingface ➡️ https://x.com/AIatMeta/status/1893009390980170001
“@bradlightcap huge congrats, open models can help it go even higher: https://x.com/_akhaliq/status/1892600666276671710
“Qwen2.5-VL Technical Report just dropped https://x.com/_akhaliq/status/1892433462910501170
“Qwen2.5-VL Technical Report was just released https://x.com/arankomatsuzaki/status/1892422884049768473
“Announcing Evo 2: The largest publicly available, AI model for biology to date, capable of understanding and designing genetic code across all three domains of life. https://x.com/arcinstitute/status/1892248139333091577
“🚀 Day 1 of #OpenSourceWeek: FlashMLA Honored to share FlashMLA – our efficient MLA decoding kernel for Hopper GPUs, optimized for variable-length sequences and now in production. ✅ BF16 support ✅ Paged KV cache (block size 64) ⚡ 3000 GB/s memory-bound & 580 TFLOPS” / X https://x.com/deepseek_ai/status/1893836827574030466
“Wow! Just tried it with various documents, including a handwritten letter from Claes Oldenburg to Ellen H. Johnson. Impressive results! Will definitely add this to the AI toolkit for journalists https://x.com/fdaudens/status/1894745220094349495
“Tools like LangSmith let you quickly evaluate new models and there a lot of new models these days 🙃” / X https://x.com/hwchase17/status/1894107965839356302
“Extracting structured data from unstructured documents is a huge use case for our customers. We’ve just made it a lot simpler with LlamaExtract, now in public beta! LlamaExtract enables you to: ➡️ Define and customize schemas for data extraction, either programmatically or in https://x.com/llama_index/status/1895164615010722233
ModelSpace/GemmaX2-28-2B-v0.1 · Hugging Face https://huggingface.co/ModelSpace/GemmaX2-28-2B-v0.1
“Deespeek released their MLA implementaion, here is how it works 💡 Multi-head Latent Attention (MLA) speeds up LLM inference and reduces memory needs. It uses “low-rank joint compression” to shrink the Key-Value (KV) to reduce memory usage by up to 93.3% and improve throughput https://x.com/_philschmid/status/1894017216640901302
“🎉 In addition to our Qwen2.5-VL report release, we’re thrilled to offer AWQ quantized models in 3B, 7B, and 72B sizes. Experience enhanced performance with these optimized versions! 🔗 Get started here: https://x.com/Alibaba_Qwen/status/1892576743904670022
“🚀 We released the tech report of Qwen2.5VL ( https://x.com/Alibaba_Qwen/status/1892576737848160538
“As always, vLLM going at lightning fast speed here and integrating EP ⚡️ https://x.com/reach_vb/status/1894266500271223021
“OlmOCR is a new drop by @allen_ai to parse any PDF 📝🤝 I have fed one of my old master’s notes and it did a great job 💗 It is based on Qwen2VL-7B and works out of the box with transformers, has Apache 2.0 license 🔥 https://x.com/mervenoyann/status/1894422823646409090
olmOCR https://simonwillison.net/2025/Feb/26/olmocr/
“Ollama’s JavaScript library is updated to v0.5.14 – improved header configuration for clients – browser compatibility fixes for platform detection 👇👇👇 https://x.com/ollama/status/1894119195253903508
“🚀 Day 3 of #OpenSourceWeek: DeepGEMM Introducing DeepGEMM – an FP8 GEMM library that supports both dense and MoE GEMMs, powering V3/R1 training and inference. ⚡ Up to 1350+ FP8 TFLOPS on Hopper GPUs ✅ No heavy dependency, as clean as a tutorial ✅ Fully Just-In-Time compiled” / X https://x.com/deepseek_ai/status/1894553164235640933
“Loving these new drops from @arcee_ai and they are Apache 2.0! Respect!” / X https://x.com/cognitivecompai/status/1892648693691551839
“🧰 New LangChain Python Integrations! We’ve added 17 new integration packages this month! Check out the whole list of integration packages here: https://x.com/LangChainAI/status/1894398108517241284
“LLMs are automating data ETL end-to-end – and it starts with structured extraction. We’re excited to announce the launch of LlamaExtract 🧑🔬🤖: a GenAI-native extraction agent that adapts the latest models to offer accurate structured extraction over large amounts of complex https://x.com/jerryjliu0/status/1895179354960994591
“I built a Deepseek R1 RAG Reasoning Agent running locally on my computer. It’s an Agentic RAG reasoning agent that can think, reason and fall back to web search if needed. 100% Opensource code with step-by-step tutorial. https://x.com/Saboo_Shubham_/status/1890966230133342318
“Efficient Triton based implementation for DeepSeek’s Native Sparse Attention! 🔥 https://x.com/reach_vb/status/1893021796577714346
DeepSeek to open source parts of online services code | TechCrunch https://techcrunch.com/2025/02/21/deepseek-to-open-source-parts-of-online-services-code/
““It turns out ingenuity [and] focus on data quality is in some ways more impactful than just burning money.” Very good read about @cohere in @the_logic” / X https://x.com/fdaudens/status/1892663693650853936
“DeepSeek just dropped DeepEP – a high-performance communication library for Mixture-of-Experts (MoE) and expert parallelism, featuring FP8 support, asymmetric-domain bandwidth forwarding (NVLink to RDMA), low-latency pure RDMA kernels (163 µs dispatch), hook-based https://x.com/reach_vb/status/1894262653440184603
“Github 👨🔧: 🎨 Refly is an open-source AI-native creation engine. Its intuitive free-form canvas interface combines multi-threaded dialogues, AI knowledge base integration, chrome extension clip & save, contextual memory, intelligent search, WYSIWYG AI editor and more, empowering https://x.com/rohanpaul_ai/status/1893077150657585158
“🤔 Think deeply about the future of Qwen. As a sneak peek into our upcoming QwQ-Max release, this version offers a glimpse of its enhanced capabilities, with ongoing refinements and an official Apache 2.0-licensed open-source launch of QwQ-Max and Qwen2.5-Max planned soon. Stay” / X https://x.com/huybery/status/1894131290246631523
“LETS GOOOOO – QwQ & Qwen 2.5 Max OPENSOURCE SOON! 🔥🔥🔥 “Very soon, we are about to release the official version of QwQ-Max, and we will open-weight both QwQ-Max and Qwen2.5-Max under the license of Apache 2.0! Furthermore, we will also provide smaller variants, e.g., QwQ-32B,” / X https://x.com/reach_vb/status/1894133551173701972
olmOCR – Open-Source OCR for Accurate Document Conversion https://olmocr.allenai.org/blog
“P2L is all open-source! Paper: https://x.com/lmarena_ai/status/1894767022791438490
“DeepSeek #2 OSS release! MoE kernels, expert parallelism, FP8 for both training and inference!” / X https://x.com/danielhanchen/status/1894212351932731581
“Emilia-Large 200K+ Hours of Open-Source Speech Data the largest TTS pretraining datasets With 200K+ hours of multilingual speech data, fully open-source https://x.com/_akhaliq/status/1895136683756245489
“DualPipe – DeepSeek’s 4th release this week! Reduces pipeline bubbles when compared to 1F1B pipelining (1 forward 1 backward) and ZB1P (Zero bubble pipeline parallelism) ZB1P is in PyTorch: https://x.com/danielhanchen/status/1894935737315008540
“DeepSeek is accelerating the launch of its next reasoner AI Reports suggest they planned to release R2 in early May but now want it out ASAP This comes after R1 went viral with its efficient yet powerful reasoning capabilities—matching industry leaders https://x.com/rowancheung/status/1894691541676888458
“@Alibaba_Qwen You guys are the best! So so looking forward to QwQ and Qwen 2.5 Max on Hugging Face! 😍 As always, we’re here if/ when you need us! – Thanks a lot for your commitment to Open Source and Science! 🫡” / X https://x.com/reach_vb/status/1894133865742004391
Tencent releases new AI model, says replies faster than DeepSeek-R1 | Reuters https://www.reuters.com/technology/artificial-intelligence/tencent-releases-new-ai-model-says-replies-faster-than-deepseek-r1-2025-02-27/
“<think>…</think> QwQ-Max-Preview Qwen Chat: https://t.co/FBpr7zfQY6 Blog: https://t.co/Zta7ezuIiC 🤔 Today we release “Thinking (QwQ)” in Qwen Chat, backed by our QwQ-Max-Preview, which is a reasoning model based on Qwen2.5-Max. This model is still for preview. It is highly https://x.com/Alibaba_Qwen/status/1894130603513319842
“🚀 Day 2 of #OpenSourceWeek: DeepEP Excited to introduce DeepEP – the first open-source EP communication library for MoE model training and inference. ✅ Efficient and optimized all-to-all communication ✅ Both intranode and internode support with NVLink and RDMA ✅” / X https://x.com/deepseek_ai/status/1894211757604049133
“Perplexity open-sourced an unbiased, post-trained version of DeepSeek R1 For context, some users found R1 refused to engage with CCP-censored topics They created a diverse dataset set of 1K+ examples to “uncensor” R1 while retaining its math, reasoning https://x.com/adcock_brett/status/1893708481015796180
“🚨 Off-Peak Discounts Alert! Starting today, enjoy off-peak discounts on the DeepSeek API Platform from 16:30–00:30 UTC daily: 🔹 DeepSeek-V3 at 50% off 🔹 DeepSeek-R1 at a massive 75% off Maximize your resources smarter — save more during these high-value hours! https://x.com/deepseek_ai/status/1894710448676884671
“We’re excited to release Command R7B Arabic – a compact open-weights AI model optimized to deliver state-of-the-art Arabic language capabilities to enterprises in the MENA region. https://x.com/cohere/status/1895186668841509355
“Day 3 of DeepSeek’s OSS release! DeepGEMM – FP8 matrix mult kernels for dense and MoEs with fine-grained scaling!! Also all kernels are JIT compiled at runtime! Super cool! I talked about DeepSeek’s FP8 methods here: https://x.com/danielhanchen/status/1894554391140864240
“I agree: “The model is the product”! IMO, if you don’t learn to train your own models (based on open-source), you won’t have a successful product long-term!” / X https://x.com/ClementDelangue/status/1894831939305218559
“DeepSeek open sourced FlashMLA – up to 3000 GB/s (memory-bound) and 580 TFLOPS in computation-bound on H800s 🔥 > BF16 support > Paged KV cache (block size 64) > Optimised for variable sequence length Day 1 of their Open Source Week! 🐳 https://x.com/reach_vb/status/1893904274825875755
“It is available today on our platform as well as accessible on @HuggingFace and @Ollama.” / X https://x.com/cohere/status/1895186677360140614
“METR is running a pilot field experiment to measure how AI tools affect open source developer productivity. If you’re an open source developer who wants to make $150/hour to work on issues of your own choosing – consider expressing interest 👇 https://x.com/METR_Evals/status/1894257205680967907
“Find more in our blog: https://x.com/cohere/status/1895186678438076477
“Love that DeepSeek is building on FlashAttention-3 code, this is why OSS can move so fast ❤️ FA3 recently enabled MLA as well, thanks to my student @tedzadouri. If you want MLA prefill & decode with full features (arbitrary page size, sliding window, rotary…), check out FA3!” / X https://x.com/tri_dao/status/1893874966661157130
“DeepSeek’s first OSS package release of the week!! Optimized multi latent attention CUDA kernels! Currently BF16 for now – maybe FP8 in the future? For a refresh of what MLA is – I explained it a bit in my DeepSeek V3 analysis tweet: https://x.com/danielhanchen/status/1893847594247377271
“💡AI is transforming every industry, creating unprecedented efficiencies and enabling entirely new classes of products. At Together AI, we believe the future of AI is open source, and we have built a cloud company for this AI-first world by combining state-of-the-art open source” / X https://x.com/togethercompute/status/1892609241212715045
“Clever test of AI reasoning ability adds the option “none of these” to the common MMLU benchmark, forcing the AI to consider options rather than just picking the best. The result is a big drop in accuracy for most models, though Reasoners (o3 and DeepSeek) hold up much better. https://x.com/emollick/status/1892613509805900191
“🚀 Day 5 of #OpenSourceWeek: 3FS, Thruster for All DeepSeek Data Access Fire-Flyer File System (3FS) – a parallel file system that utilizes the full bandwidth of modern SSDs and RDMA networks. ⚡ 6.6 TiB/s aggregate read throughput in a 180-node cluster ⚡ 3.66 TiB/min” / X https://x.com/deepseek_ai/status/1895279409185390655
“it’s surprising how many people were giving me grief over my bullishness on Grok” / X https://x.com/teortaxesTex/status/1892509598612877503
“xAI unveiled Grok-3, its new AI model with reasoning capabilities, achieving SoTA performance across math, science, and coding On launch, they also revealed that the model scored #1 on Chatbot Arena above all other models Impressive speed https://x.com/adcock_brett/status/1893708244100530495
“now here’s something I expect Elon to beat Dario on: organic reaction to being called a retard. big Grok W. (also, smart that it has a clue about X’s peculiarities) https://x.com/teortaxesTex/status/1894049060308336799
xAI’s Grok 3 is available for free to everyone ‘for a short time’ https://www.engadget.com/ai/xais-grok-3-is-available-for-free-to-everyone-for-a-short-time-130031943.html
Grok 3 Beta — The Age of Reasoning Agents https://x.ai/blog/grok-3
“Grok 3 API with 1M context coming thank me later 🫡” / X https://x.com/teortaxesTex/status/1894733474239582465
“Grok 3 is the first (and only) model to solve this non-riddle. https://x.com/emollick/status/1894526521835946353
“The chain-of-thought in the reasoning model for Grok 3 is very relatable. https://x.com/emollick/status/1892818440731185156
“grok 3 build a endless runner style game where a hugging face collects GPUs https://x.com/_akhaliq/status/1893847291221426578
“Another reason I would love a model card for Grok 3 (and it would be helpful for the team to release one)” / X https://x.com/emollick/status/1892443277485359546
“Grok 3 is a very good model but it is hard to imagine that it is not eclipsed by coming releases from the other labs. There just isn’t enough of an edge there The question is whether X is aiming for a marathon or a sprint, and if they can keep the speed of improvement as high” / X https://x.com/emollick/status/1892732503464579502
“Grok 3 on how to become emperor of Rome (both summary and full version) https://x.com/emollick/status/1892463906016170322
“this is what using grok 3 feels like 😭 https://x.com/bilawalsidhu/status/1892392477987930185
Grok 3 appears to have briefly censored unflattering mentions of Trump and Musk | TechCrunch https://techcrunch.com/2025/02/23/grok-3-appears-to-have-briefly-censored-unflattering-mentions-of-trump-and-musk/
“I will respond to Grok 3 the only way I can imagine, which is to begin a campaign of sabotage against heavy water production in southern Norway.” / X https://x.com/_aidan_clark_/status/1893810160662945918
“this is huge, Hugging Face Inference providers now supports over 8 different providers including sambanova, together, fireworks, hyperbolic, replicate, fal and more and close to 100 models https://x.com/_akhaliq/status/1892628229871030602
“HOLY SHITT, Microsoft dropped an open-source Multimodal (supports Audio, Vision and Text) Phi 4 – MIT licensed! 🔥 > Beats Gemini 2.0 Flash, GPT4o, Whisper, SeamlessM4T v2 > Models on Hugging Face hub, integrated with/ Transformers! Phi-4-Multimodal: > Modalities: Integrates https://x.com/reach_vb/status/1894989136353738882
“We just crossed 2,000 organizations who upgraded to @huggingface Enterprise including Mercedes, Nvidia, Deutsche Telekom, Synopsys, ServiceNow, Fidelity, Chegg, Perplexity, Procter and Gamble, Meta, Mistral, AMD, Snowflake, Shopify, Jasper, Writer, IBM, J&J, Credit Karma, Adyen, https://x.com/ClementDelangue/status/1894778750631371149
“ok this is fine “` import matplotlib.pyplot as plt # Data and styling (sorted by pass@1 scores) models = [‘o3-mini(high)’, ‘Grok-3 mini Reasoning’, ‘o1 (medium)’, ‘o3-mini(medium)’, ‘Grok-3 Reasoning Beta’, ‘Deepseek-R1’, ‘DeepSeek-R1-D-Llama-70B’, ‘Gemini-2” / X https://x.com/teortaxesTex/status/1892487275948224793
“@InceptionAILabs If you got curious by the is diffusion large language model but realized it’s not open source: it’s your lucky day 🍀 LLaDA 8B – an open source apache 2 large diffusion language model is also just out! ✨ competitive with LLaMA 3 8B https://x.com/multimodalart/status/1895046839159668876
“@karpathy It shares the #1 large diffusion language with an open source one ⚡️, LLaDA, Large Language Diffusion Model (8B params), competitive with LLaMA 3 – weights were just out yesterday too” / X https://x.com/multimodalart/status/1895039220722319532
“With this and LLaDA-8B I am not sure how you can ignore diffusion language models as being a real possibility for the future of language models. Maybe I am being hyperbolic but I wouldn’t be surprised if GPT-5 or GPT-6 is a diffusion model! Anyway yeah I’m quite bullish on” / X https://x.com/iScienceLuvr/status/1895078017548046751
“Miquel and team COOKED: SmolVLM2 – Apache 2.0 licensed VideoLMs (ranging from 2.2B to 256M) – can even run on a FREE colab🔥 Beats models 5x it’s size – runs on an iPhone! Trained with memory efficiency as a focus you can pass extremely long videos with little VRAM required🤗 https://x.com/reach_vb/status/1892578169615523909
“Check-out SmolVLM2 with day-zero support for MLX + MLX Swift to run locally on any Apple device. Small yet powerful enough to do fast *video* understanding on an iPhone! https://x.com/awnihannun/status/1892594913893556707
“Introducing Perplexity’s new voice mode. Ask any question. Hear real-time answers. Update your iOS app to start using. Coming soon to Android and Mac app. https://x.com/perplexity_ai/status/1894788583770509505
“we just dropped SmolVLM2: world’s smollest video models in 256M, 500M and 2.2B ⏯️🤗 we also release the following 🔥 > an iPhone app (runs on 500M model in MLX) > integration with VLC for segmentation of descriptions (2.2B) > a highlights extractor (2.2B) https://x.com/mervenoyann/status/1892576290181382153
“Physical AI is a civilizational technology. In a few years, intelligent robots will be as many as iPhones. I’d love to see your coolest open-source robotics project: model, simulation, hardware, you name it! Reply with links, and I’ll pick one winner for the NVIDIA GTC Golden https://x.com/DrJimFan/status/1892980857255928292
Vevo Therapeutics Open Sources Tahoe-100M, the World’s Largest Single-Cell Dataset, as the Inaugural Contribution to Arc Institute’s New Virtual Cell Atlas | Arc Institute https://arcinstitute.org/news/news/arc-vevo
“Meta presents: SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution Achieves 41.0% solve rate on SWE-bench Verified with Llama 3 https://x.com/arankomatsuzaki/status/1894596772804350016
“Cool fact: Llama 1 was trained with a batch size of ~4M tokens for 1.4 trillions tokens while DeepSeek was trained with a batch size of ~60M tokens for 14 trillion tokens” / X https://x.com/andrew_n_carr/status/1892403826717569120
“💪Open source models like DeepSeek-R1 and Meta’s Llama have emerged as formidable alternatives to proprietary solutions, marking a decisive shift in the AI landscape. Together AI has established itself as the definitive platform powering this transformation, delivering the” / X https://x.com/togethercompute/status/1892609242957582505
“Meta just dropped SWE-RL Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution Trained on top of Llama 3, our resulting reasoning model, Llama3-SWE-RL-70B, achieves a 41.0% solve rate on SWE-bench Verified — a human-verified collection of real-world https://x.com/_akhaliq/status/1894584315352076608
“Meta PARTNR dataset and code ⬇️ https://x.com/AIatMeta/status/1894524604900938078
Empowering innovation: The next generation of the Phi family | Microsoft Azure Blog https://azure.microsoft.com/en-us/blog/empowering-innovation-the-next-generation-of-the-phi-family/
Microsoft Has Dropped Some AI Data Center Leases, TD Cowen Says – Bloomberg https://archive.md/WP5rT
Microsoft releases new Phi models optimized for multimodal processing, efficiency – SiliconANGLE https://siliconangle.com/2025/02/26/microsoft-releases-new-phi-models-optimized-multimodal-processing-efficiency/
“🚀 Day 0: Warming up for #OpenSourceWeek! We’re a tiny team @deepseek_ai exploring AGI. Starting next week, we’ll be open-sourcing 5 repos, sharing our small but sincere progress with full transparency. These humble building blocks in our online service have been documented,” / X https://x.com/deepseek_ai/status/1892786555494019098
“AlphaMaze: Teaching a 1.5B LLM to think visually and solve ARC-AGI like puzzles! 🤯 Powered by DeepSeek R1 1.5B + GRPO All with Apache licensed checkpoints and dataset 🤗 https://x.com/reach_vb/status/1892999150255440012
“Phi 4 Multimodal Instruct: https://x.com/reach_vb/status/1894991935456124941
“How are Vision Language Models trained? 🖼️ @Alibaba_Qwen 2.5-VL demonstrates state-of-the-art performance through dynamic resolution processing, absolute time encoding for videos, and a redesigned Vision Transformer in 3B/7B/72B variants for edge and cloud deployment. 👀 https://x.com/_philschmid/status/1892506190656999925
“🚀 Day 0: Warming up for #OpenSourceWeek! We’re a tiny team @deepseek_ai exploring AGI. Starting next week, we’ll be open-sourcing 5 repos, sharing our small but sincere progress with full transparency. These humble building blocks in our online service have been documented,” / X https://x.com/deepseek_ai/status/1892786555494019098
SmolVLM2: Bringing Video Understanding to Every Device https://huggingface.co/blog/smolvlm2





One response to “Open Source: AI News Week Ending 02/28/2025”
[…] Open Source: 96 stories […]