OLMo 2: The best fully open language model to date | Ai2
Tülu 3 is a leading instruction following model family, offering fully open-source data, code, and recipes designed to serve as a comprehensive guide for modern post-training techniques.
“I’ve spent the last two years scouring all available resources on RLHF specifically and post training broadly. Today, with the help of a totally cracked team, we bring you the fruits of that labor — Tülu 3, an entirely open frontier model post training recipe. We beat Llama 3.1
Ai2 releases new language models competitive with Meta’s Llama | TechCrunch
DocumindHQ/documind: Open-source platform for extracting structured data from documents using AI.
As Cohere and Writer mine the ‘Live AI’ arena, Pathway joins the pack with a $10M round | TechCrunch
“🚀 Introducing Hymba-1.5B: a new hybrid architecture for efficient small language models! ✅ Outperforms Llama, Qwen, and SmolLM2 with 6-12x less training ✅ Massive reductions in KV cache size & good throughput boost ✅ Combines Mamba & Attention in a Hybrid Parallel
“🚀 Mind-blown by SmolVLM – a tiny but mighty vision language model! ✨ Key specs: – 2.25B parameters – Only 5GB GPU RAM needed – Apache 2.0 license – Fine-tunable on Google Colab free tier #AI #MachineLearning
Databricks
Databricks nears multibillion funding round at $55 billion valuation
DeepSeek
“DeepSeek launched DeepSeek-R1, a model family that demonstrates competitive performance in reasoning-intensive tasks while offering transparent reasoning steps—a departure from how OpenAI handles reasoning tokens in o1. Learn more in #TheBatch:
Hugging Face
SmolVLM – a Hugging Face Space by HuggingFaceTB
ColPali 🤝 Vespa – Visual Retrieval – a Hugging Face Space by vespa-engine
Tulu 3 Datasets – a allenai Collection
showlab/ShowUI-2B · Hugging Face
AI Video Composer – a Hugging Face Space by huggingface-projects
🧠 Reasoning Models – a zh-ai-community Collection
bluesky-community (Bluesky Community)
“HuggingFace Inference Endpoints now supports deploying llama.cpp-powered instances on CPU servers too This is a first step towards a wider low-cost cloud LLM availability, especially with new AI-friendly instruction sets on the rise. More info:
Mistral
“New study shows LLMs outperform neuroscience experts at predicting experimental results in advance of experiments (86% vs 63% accuracy). They use a fine-tuned Mistral 7B but other models worked too. Suggests LLMs can integrate scientific knowledge at scale to support research
Qwen
Alibaba’s Qwen with Questions reasoning model beats o1-preview | VentureBeat
“🔥 Battle for the top reasoning LLM intensifies! The QwQ-32B-Preview is a very good reasoning LLM. Full video of my tests here:
“Alibaba researchers unveil Marco-o1, an LLM with advanced reasoning capabilities
Alibaba releases an ‘open’ challenger to OpenAI’s o1 reasoning model — OODAloop
Extending the Context Length to 1M Tokens! | Qwen
“QwenVL-Flux is ABSURDLY cool ⚡️ It’s kind of like an IP Adapter, with QwenVL as the image encoder, adding multiple high quality features: – Image Variation 🎨 – Image Blending 🔄 – Style Transfer 🎯 Kudos to @kuer5ord for training it! ▶️ Play at
“🎨 Excited to announce that Qwen2VL-Flux demo is now ready for testing! This is a lightweight variation of the full model, perfect for quick image transformations with optional text guidance.
QwQ: Reflect Deeply on the Boundaries of the Unknown | Qwen
QwQ-32B-Preview – a Hugging Face Space by Qwen
“QwQ whats this? (model looks like normal qwen2, so no RL techniques are happening at inference time) https://x.com/gazorp5/status/1861883506055606567





Leave a Reply