“Or impact of the fact that closed-source players use open-source AI (currently dominated by Chinese players) like open-source datasets? The countries or companies that win open-source AI will have massive power and influence on the future of AI.” / X
https://x.com/ClementDelangue/status/1877767382120255792

“MiniMax-01 coder mode is now available in ai-gradio pip install –upgrade “ai-gradio[minimax]” import gradio as gr import ai_gradio demo = gr.load( name=’minimax:MiniMax-Text-01′, src=ai_gradio.registry, coder=True ) demo.launch()”
https://x.com/_akhaliq/status/1880059318785176043

“🥳 Transformers 4.48.0 is out. ModernBert + new models added #HuggingFace”
https://x.com/fdaudens/status/1877772241884065827

“We’re excited to introduce Tasks! For the first time, ChatGPT can manage tasks asynchronously on your behalf—whether it’s a one-time request or an ongoing routine. Here are my favorite use cases: 1/ ChatGPT checks stock price every morning!”
https://x.com/karinanguyen_/status/1879270529066262733

“🥳🎉Announcing @MistralAI new model: Codestral 25.01 – new SOTA coding model, #1 on LMSYS! – Lightweight, fast, and proficient in over 80 programming languages, – Optimized for low-latency, high-frequency usecases – 2x faster than the previous version – Supports tasks such as”
https://x.com/sophiamyang/status/1878902888434479204

kyutai/helium-1-preview-2b · Hugging Face
https://huggingface.co/kyutai/helium-1-preview-2b

“TabPFNv2 models are out, and they’re immensely performant 🤯 Get started with 2 LoC ⬇️”
https://x.com/mervenoyann/status/1877725448823525665

“Model page:”
https://x.com/ollama/status/1879216139542434055

“Ollama v0.5.5 is here! Lots more quality of life updates and fixes as we begin to transition to Ollama’s new engine.”
https://x.com/ollama/status/1879212554435911978

“New Day, New Embedding Model! @llama_index releases “vdr-2b-multi-v1” a 2B multimodal, multilingual embedding model designed for document retrieval without OCR or complex data processing. 👀 TL;DR: 🧠 Built on @Alibaba_Qwen 2-VL 2B with Matryoshka Representation Learning 🌐
https://x.com/_philschmid/status/1877778889566494843

“a little teaser: with the amazing @AymericRoucher we are cooking vision support for smolagents 🤠 soon you’ll be able to use APIs like gpt-4o as well as a huge variety of @huggingface transformers vision LMs 🤝 lfg 🎉
https://x.com/mervenoyann/status/1879947783442202666

“That’s a APACHE 2.0 Licensed STREAMING multimodal 8B model beating GPT4o 💥
https://x.com/reach_vb/status/1879124653626765381

(1) Moondream: how does a tiny vision model slap so hard? — Vikhyat Korrapati – YouTube

“🚀 LlamaIndex’s first open-source model is here! vdr-2b-multi-v1: A groundbreaking multilingual visual document retrieval model that’s 3x faster. Search across 5 languages without OCR, backed by 500k training samples. Try it:
https://x.com/fdaudens/status/1877811907454808236

“Awesome new app for local LLMs that runs on iPhone, iPad, Mac, etc built with MLX Swift. Also code is open source and MIT licensed!
https://x.com/awnihannun/status/1878843809460875593

“Introducing OuteTTS 0.3 1B & 500M 🔥 > Zero shot voice cloning > Multilingual (en, jp, ko, zh, fr, de) > Trained on 20,000 hours of audio > Powered by OLMo-1B & Qwen 2.5 0.5B > Speed & emotion control > Powered by HF grants 🤗 Open science ftw!
https://x.com/reach_vb/status/1879647151145590905

“New open Omni model released! 👀@OpenBMB MiniCPM-o 2.6 is a new 8B parameters, any-to-any multimodal model that can understand vision, speech, and language and runs on edge devices like phones and tablets. TL;DR: 🧠 8B total parameters (SigLip-400M + Whisper-300M + ChatTTS-200M
https://x.com/_philschmid/status/1879163439559389307

“We’ve released a new multilingual, open-source visual embedding model and training set on Huggingface! vdr-2b-multi-v1 produces single dense vectors for visual document retrieval, enabling efficient large-scale systems. Key features: ➡️ Trained on 5 languages (IT, ES, EN, FR,
https://x.com/llama_index/status/1877778352087699962

“🎬 You can now create AI videos with transparent backgrounds! Game-changer for VFX creators – and it’s all open-source! My favorite examples:” / X
https://x.com/fdaudens/status/1877361506368573817

“@Alibaba_Qwen exciting, qwen is also available to try out in anychat:
https://x.com/_akhaliq/status/1877478548874699153

“Qwen COOKED yet again: commercially permissive Process Reward Models (72B + 7B)🔥 Along with a detailed tech report!
https://x.com/reach_vb/status/1879089267139588276

Purr-fectly informed | Mistral AI | Frontier AI in your hands
https://mistral.ai/news/mistral-afp/

Codestral 25.01 | Mistral AI | Frontier AI in your hands
https://mistral.ai/news/codestral-2501/

“@natolambert “Tier 2 country” holy shit, they keep doing it > Poland they don’t even care about fanatical proven loyalty. Urge to go back to my Tier 3 country and finetune DeepSeek V3 on smuggled GPUs.” / X
https://x.com/teortaxesTex/status/1877730865003786384

“Today, we’re launching early access for North! Our all-in-one secure AI workspace platform combines LLMs, search, and agents into an intuitive interface that effortlessly integrates AI into your daily work to achieve peak productivity.
https://x.com/cohere/status/1877335657908949189

AFP and Mistral AI announce global partnership to enhance AI responses with reliable news content | AFP.com
https://www.afp.com/en/agency/inside-afp/press-release/afp-and-mistral-ai-announce-global-partnership-enhance-ai-responses

Trying out QvQ—Qwen’s new visual reasoning model
https://simonwillison.net/2024/Dec/24/qvq/

“The new Grok app is so fast and smooth. At this rate of shipping and quality, the @xai team will leapfrog everyone.” / X
https://x.com/amasad/status/1879418028989063552

“Phi-4 (4-bit) in @lmstudio on an M4 max is quite fast and quite good:
https://x.com/awnihannun/status/1878564132125085794

“LLMQuoter Enhances RAG capabilities using a “quote-first-then-answer” strategy. Adopts Llama-3B and finetunes with LoRA on a 15K sample subset of HotPotQA to enhance RAG by identifying key quotes before passing them to reasoning models. “This workflow reduces cognitive
https://x.com/omarsar0/status/1878820053933855147

“🦙 Llama 3.3 70B is now available on Together AI for free! The new 70B model delivers similar capabilities to the much larger Llama 3.1 405B model, with improved reasoning 🤔, math ➕➖, and instruction-following 🧠. Explore this model and unleash your creativity today! 🎨
https://x.com/togethercompute/status/1879231968434684254

“Today we’re excited to share that our work on SeamlessM4T from Meta FAIR was published in the latest issue of @Nature ➡️
https://x.com/AIatMeta/status/1879593558728188045

Mark Zuckerberg gave Meta’s Llama team the OK to train on copyrighted works, filing claims | TechCrunch

Mark Zuckerberg gave Meta’s Llama team the OK to train on copyrighted works, filing claims

“🚨 MiniCPM-o 2.6, an 8B multimodal LLM, outperforms GPT-4o, Gemini 1.5 Pro, and Sonnet in visual (70.2 avg on OpenCompass, 1.8M pixel – 1344 x 1344), speech (bilingual realtime, beats GPT-4o-realtime in ASR/STT), with 75% fewer vision tokens, supporting 30+ languages 🔥
https://x.com/reach_vb/status/1879119335354220734

“looks pretty dope – 8B Instruct from @intern_lm – beats Qwen 2.5 7B/ Llama 3.1 8B AND APACHE 2.0 licensed🔥 Trained on 4 Trillion tokens (much much less than the later) & supports o1 like `Thinking` mode! Model checkpoints on the Hub + compatible w/ Transformers! 🤗
https://x.com/reach_vb/status/1879538728710164786

“OpenBioLLM-8B and OpenBioLLM-70B are new fine-tuned Llama models developed by Saama to streamline tasks that can accelerate clinical trials, opening up new possibilities in personalized medicine. More details in their research paper ➡️
https://x.com/AIatMeta/status/1880338816491499737

“Using the DINOv2 open source model from Meta FAIR, @virgosvs developed EndoDINO, a foundation model that delivers SOTA performance across a range of GI endoscopy tasks ➡️
https://x.com/AIatMeta/status/1878887608681742625

Announcing Open-Source SAEs for Llama 3.3 70B and Llama 3.1 8B – Goodfire Papers
https://www.goodfire.ai/blog/sae-open-source-announcement/

“What a beginning to this year in open ML 🤠 Let’s unwrap! Multimodal 🖼️ > ByteDance released SA2VA: a family of vision LMs that can take image, video, text and visual prompts > moondream2 is out with new capabilities like outputting structured data and gaze detection! @vikhyatk
https://x.com/mervenoyann/status/1877728986341515663

“there’s a new multimodal retrieval model in town 🤠 @llama_index released vdr-2b-multi-v1 > uses 70% less image tokens, yet outperforming other dse-qwen2 based models > 3x faster inference with less VRAM 💨 > shrinkable with matryoshka 🪆 > can do cross-lingual retrieval!
https://x.com/mervenoyann/status/1878757447294500968

“Wait WTF, @MiniMaxAI_ dropped MiniMax-Text 01, 456B parameters (45.9B active) beats DeepSeek v3 with FOUR MILLION context length – commercially permissive! 🔥 > On Hugging Face Hub & works w/ Transformers (custom code) 💥
https://x.com/reach_vb/status/1879232524066726039

“So, a 7B model just outperformed OpenAI’s O1 in math.
https://x.com/fdaudens/status/1877518980027388068

“Step-by-step visual reasoning that’s both accurate and lightning-fast. LlamaV-o1 introduces a comprehensive framework for step-by-step visual reasoning in LLMs, with a new benchmark, evaluation metric, and curriculum learning approach. —– 🤔 Original Problem: → Current
https://x.com/rohanpaul_ai/status/1879813276827304431

“Huge. UC Berkeley just released a $450 open-source reasoning model that matches o1. Sky-T1-32B-Preview is a fully open-source model designed for reasoning and coding tasks. Achieves 82.4% on Math500 and 86.3% on LiveCodeBench-Easy. It includes training data, code, and model
https://x.com/LiorOnAI/status/1878876546066506157

“🎉 Introducing DeepSeek App! 💡 Powered by world-class DeepSeek-V3 🆓 FREE to use with seamless interaction 📱 Now officially available on App Store & Google Play & Major Android markets 🔗Download now:
https://x.com/deepseek_ai/status/1879465495788917166

“LETS GOO! @kyutai_labs drops Helium-1 Preview, 2B-parameter multilingual base LLM targeting edge and mobile devices beats Qwen 2.5 1.5B – CC-BY licensed 🔥 > outperforming/ comparable to Owen 1.5B, Gemma 2B & Llama 3B > trained on 2.5T tokens with a 4096 context size >
https://x.com/reach_vb/status/1878860650560025011

DeepSeek
“DeepSeek-V3, the company’s latest open LLM, surpasses Llama 3.1 405B and GPT-4o on key benchmarks, especially in coding and math tasks. Using a mixture-of-experts architecture with 671 billion parameters, only 37 billion are active at once, DeepSeek V3 was trained at a low cost” / X

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading