OLMo 2: The best fully open language model to date | Ai2

Tülu 3 is a leading instruction following model family, offering fully open-source data, code, and recipes designed to serve as a comprehensive guide for modern post-training techniques.

“I’ve spent the last two years scouring all available resources on RLHF specifically and post training broadly. Today, with the help of a totally cracked team, we bring you the fruits of that labor — Tülu 3, an entirely open frontier model post training recipe. We beat Llama 3.1 

Ai2 releases new language models competitive with Meta’s Llama | TechCrunch

DocumindHQ/documind: Open-source platform for extracting structured data from documents using AI.

As Cohere and Writer mine the ‘Live AI’ arena, Pathway joins the pack with a $10M round | TechCrunch

“🚀 Introducing Hymba-1.5B: a new hybrid architecture for efficient small language models! ✅ Outperforms Llama, Qwen, and SmolLM2 with 6-12x less training ✅ Massive reductions in KV cache size & good throughput boost ✅ Combines Mamba & Attention in a Hybrid Parallel 

“🚀 Mind-blown by SmolVLM – a tiny but mighty vision language model! ✨ Key specs: – 2.25B parameters – Only 5GB GPU RAM needed – Apache 2.0 license – Fine-tunable on Google Colab free tier #AI #MachineLearning 

Databricks

Databricks nears multibillion funding round at $55 billion valuation

DeepSeek

“DeepSeek launched DeepSeek-R1, a model family that demonstrates competitive performance in reasoning-intensive tasks while offering transparent reasoning steps—a departure from how OpenAI handles reasoning tokens in o1. Learn more in #TheBatch: 

Hugging Face 

SmolVLM – a Hugging Face Space by HuggingFaceTB

ColPali 🤝 Vespa – Visual Retrieval – a Hugging Face Space by vespa-engine

Tulu 3 Datasets – a allenai Collection

showlab/ShowUI-2B · Hugging Face

AI Video Composer – a Hugging Face Space by huggingface-projects

🧠 Reasoning Models – a zh-ai-community Collection

bluesky-community (Bluesky Community)

“HuggingFace Inference Endpoints now supports deploying llama.cpp-powered instances on CPU servers too This is a first step towards a wider low-cost cloud LLM availability, especially with new AI-friendly instruction sets on the rise. More info: 

Mistral

“New study shows LLMs outperform neuroscience experts at predicting experimental results in advance of experiments (86% vs 63% accuracy). They use a fine-tuned Mistral 7B but other models worked too. Suggests LLMs can integrate scientific knowledge at scale to support research 

Qwen

Alibaba’s Qwen with Questions reasoning model beats o1-preview | VentureBeat

“🔥 Battle for the top reasoning LLM intensifies! The QwQ-32B-Preview is a very good reasoning LLM. Full video of my tests here: 

“Alibaba researchers unveil Marco-o1, an LLM with advanced reasoning capabilities 

Alibaba releases an ‘open’ challenger to OpenAI’s o1 reasoning model — OODAloop

Extending the Context Length to 1M Tokens! | Qwen

“QwenVL-Flux is ABSURDLY cool ⚡️ It’s kind of like an IP Adapter, with QwenVL as the image encoder, adding multiple high quality features: – Image Variation 🎨 – Image Blending 🔄 – Style Transfer 🎯 Kudos to @kuer5ord for training it! ▶️ Play at 

“🎨 Excited to announce that Qwen2VL-Flux demo is now ready for testing! This is a lightweight variation of the full model, perfect for quick image transformations with optional text guidance. 

QwQ: Reflect Deeply on the Boundaries of the Unknown | Qwen

QwQ-32B-Preview – a Hugging Face Space by Qwen

“QwQ whats this? (model looks like normal qwen2, so no RL techniques are happening at inference time) https://x.com/gazorp5/status/1861883506055606567

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading