the Hugging Face logo on a chalkboard in a classroom
“Introducing Spellcaster: AI Grammarly, for documentation proof reading Spellcaster is an open-source CLI tool that uses LLM agents to improve your codebase’s documentation. It scans for: * Grammar errors * Spelling mistakes * Bad code examples Here’s how it works (🧵):
“Open-source models beating closed models will become more and more common. Scaling has diminishing returns. The best solution will not have the largest scale but best approach or data. Especially with test-time compute, you do not need the best model to have the best solution.” / X
The Most Capable Open Source AI Model Yet Could Supercharge AI Agents | WIRED
“🎉🎉 CrewAI version v0.63.6 is out! 🎉🎉 🤖 New LLM class to interact with LLMs 🧠 Support to custom memory interfaces 🦾 Adding support to o1 family models 🎤 Updated logs format 📃 Docs updated 🪲 Bugs fixed 💇♀️ Oh and small logo update 🙂 RT please? 🙏⚡️ we move fast!” / X
“Colossal-AI now makes it possible to reduce model training costs by 30% with one single line of code! Achieved by upgrading mixed precision training that now supports BF16 (O2) + FP8 (O1).” / X
“New Open source SoTA Multimodal (Vision) Language model dropped – Molmo Outperformed Claude 3.5 Sonnet, GPT4V, Gemini 1.5 Pro – using 1000x less data. 🤯 🗣️ Novel dataset (PixMo) with detailed human-spoken image captions 🧠 Architecture: Vision encoder + LLM 🔓 Open weights,
“Meet Molmo: a family of open, state-of-the-art multimodal AI models. Our best model outperforms proprietary systems, using 1000x less data. Molmo doesn’t just understand multimodal data—it acts on it, enabling rich interactions in both the physical and virtual worlds. Try it
“$IBM and NASA are partnering to develop an open-source AI model for weather and climate analysis. Hugging Face CEO @ClementDelangue discusses:
“Moshi, the speech-based AI assistant from Kyutai is now open source and comes with a detailed technical paper.” / X
Announcing Pixtral 12B | Mistral AI | Frontier AI in your hands
“Here is an open-source, self-hosted AI starter kit. This template will bootstrap a fully-featured low-code development environment to build AI applications:
[2409.16693v1] CaBRNet, an open-source library for developing and evaluating Case-Based Reasoning Models
Apple
Siri May Not Get Its Apple Intelligence Update Until January 2025
Gemma
DataGemma-FullPaper
Hugging Face
“We just crossed 1,000,000 free public models on Hugging Face! That’s the ones the media covers like Llama, Gemma, Phi, Flux, Mistral, Phi, Starcoder, Qwen, Stable diffusion, Grok, Whisper, Olmo, Command, Zephyr, OpenELM, Jamba, Yi but also 999,984 others. Why? Because contrary” / X
“@derekm00r3 @polynoamial @thomaspower @OpenAI OpenAI and Google-DeepMind may have clammed up. But open source has never been healthier. Just look at the number of projects on Github. HuggingFace just reached 1 million models. The entire world runs on Linux (except for a few desktops and phones). Hardly signs that open source” / X
“Open Dataset release by @OpenAI! 👀 OpenAI just released a Multilingual Massive Multitask Language Understanding (MMMLU) dataset on @huggingface! 🌍 MMLU test set available in 14 languages, including Arabic, German, Spanish, French,…. 🧠 Covers 57 categories from elementary to
GRIN_MoE.pdf · microsoft/GRIN-MoE at main
Meta/Llama
“With Llama 3.2 we released our first-ever lightweight Llama models: 1B & 3B. These models empower developers to build personalized, on-device agentic applications with capabilities like summarization, tool use and RAG where data never leaves the device.
“The really MASSIVE week continues with 🚀 Llama 3.2 Release – Introduces 1B and 3B text models for edge devices, 11B and 90B vision models – All models support 128K token context – 1B/3B outperform Gemma 2 2.6B and Phi 3.5-mini on key tasks – 11B/90B vision models competitive
Llama 3.2 goes small and multimodal · Ollama Blog
Llama 3.2
meta-llama (Meta Llama)
“The lightweight Llama 3.2 models shipping today include support for @Arm, @MediaTek & @Qualcomm to enable the developer community to start building impactful mobile applications from day one.
“I just pulled the numbers on vision-language benchmarks for Llama-3.2-11B (vision). Surprisingly, the open-source community at large isn’t behind in the lightweight model class! Pixtral, Qwen2-VL, Molmo, and InternVL2 all stand strong. OSS AI models have never been stronger. The
“At Connect, we announced Hyperscape technology, bringing photorealistic spaces into the metaverse with your mobile phone (scanning not yet available). We’ve launched a demo app so you can experience these high-fidelity spaces yourself!
“📣 Introducing Llama 3.2: Lightweight models for edge devices, vision models and more! What’s new? • Llama 3.2 1B & 3B models deliver state-of-the-art capabilities for their class for several on-device use cases — with support for @Arm, @MediaTek & @Qualcomm on day one. •
Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
meta-llama/llama-stack: Composable building blocks to build Llama Apps
“A few technical insights on our lightweight Llama 3.2 1B & 3B models. 🦙🧵” / X
“My analysis of Llama 3.2: 1. New 1B and 3B text only LLMs 9 trillion tokens 2. New 11B and 90B vision multimodal models 3. 128K context length 4. 1B and 3B used some distillation from 8B and 70B 5. VLM 6 billion img, text pairs 6. CLIP MLP GeLU + cross attention Long analysis:
“🚀 Big news! We’re thrilled to announce the launch of Llama 3.2 Vision Models & Llama Stack on Together AI. 🎉 Free access to Llama 3.2 Vision Model for developers to build and innovate with open source AI.
“These lightweight Llama models were pretrained on up to 9 trillion tokens. One of the keys for Llama 1B & 3B however was using pruning & distillation to build smaller and more performant models informed by powerful teacher models. Pruning enabled us to reduce the size of extant
“We’re seeing exciting performance from these new models, with results that outperform Gemma 2 2.6B and Phi 3.5-mini models on a range of tasks even at smaller sizes!
“Zuck’s AI strategy in a nutshell: – Free frontier model for devs – Undercut closed-source rivals – Crowdsource best use cases for multimodal AI – Skip cloud wars; instead monetize biz/creator agents – Harvest fresh data/content for Meta ecosystem – Profit Zuck’s XR strategy in
“Llama 3.2 Multimodal benchmarks MMMU for 11B 60.3 vs Claude Haiku 50.2 MMMU for 90B 60.3 vs GPT 4o mini 59.4 90B looks extremely powerful!
“Llama 3.2 1B in 4-bit runs at ~60 toks/sec with MLX Swift on my iPhone 15 pro. It’s quite good and easily runs on-device:
“Llama 3.2 is out with a major update: it can now process images. Key highlights: • 11B and 90B vision models • Small 1B and 3B text models for mobile devices” / X
“Llama can now see and run on your phone!👀🖼️ Llama 3.2 released with Multimodal support in Llama Vision and tiny llamas for on-device usage. 10 new llama released by @AIatMeta from 1B text only to 90B Multimodal (text+image) 🚀 But with EU restrictions 🇪🇺 TL;DR: 🖼️ Llama 3.2
“Llama 3.2 features 11B & 90B models, our first multimodal Llama models with support for vision tasks. These models can take in both image and text prompts to deeply understand and reason on inputs.
“2. Meta AI can now ‘see’ images! Similar to ChatGPT, you can now share photos and have Meta AI reply to any photo in chat. But where Meta is going a step further is by allowing users to actually edit photos, like removing an object, adding a hat, or changing backgrounds, etc.,
“Our vision models required an entirely new architecture to support image reasoning. This was accomplished by training a set of adapter weights that integrate the pre-trained image encoder into the pre-trained language model.” / X
Qwen
“Alibaba’s Qwen 2.5 7B coder is reportedly as good as GPT-4 at coding. For reference, Qwen is more than 300 times cheaper at $0.09 per million tokens. Running on a local machine, it’s free — another big win for open-source!
“We have GPT-4 for coding at home! I looked up @OpenAI GPT-4 0613 results for various benchmarks and compared them with @Alibaba_Qwen 2.5 7B coder. 👀 > 15 months after the release of GPT-0613, we have an open LLM under Apache 2.0, which performs just as well. 🤯 > GPT-4 pricing
“GPT-4 for coding at home! Qwen 2.5 Coder 7B outperforms other @OpenAI GPT-4 0613 and open LLMs < 33B, including @BigCodeProject StartCoder, @MistralAI Codestral, or Deepseek, and is released under Apache 2.0. 🤯 Details: 🚀 Three model sizes: 1.5B, 7B, and 32B (coming soon) up





Leave a Reply