“ModernBERT ftw! @answerdotai & @LightOnIO killing it!! 🔥 > ModernBERT-base: 22 layers, 149M params > ModernBERT-large: 28 layers, 395M params > 2 trillion tokens of English and code data. > Up to 8,192 tokens, ideal for processing long documents > RoPE for long-context support”
https://x.com/reach_vb/status/1869791808030708054
“Falcon 3 models were released a few hours ago! Huggingface Link:”
https://x.com/scaling01/status/1869007562034544939
“Introducing CerebrasCoder! An open-source app that generates websites with Llama3.3-70b from @CerebrasSystems as fast as you can type. 100% free and open-source.
https://x.com/stevekrouse/status/1869029269646835716
Bringing Grok to Everyone
https://x.ai/blog/grok-1212
“🔍 Economists make the case in a statement published on @mozilla: Open-source AI isn’t just good tech—it’s smart economics. See their key arguments visualized: #OpenSourceAI #AIEconomics
https://x.com/fdaudens/status/1868752892041347327
[2412.13061v1] VidTok: A Versatile and Open-Source Video Tokenizer
https://arxiv.org/abs/2412.13061v1
Databricks
Databricks to Hit $62 Billion Valuation in New Funding Round
https://finance.yahoo.com/news/databricks-hit-62-billion-valuation-153858326.html
“Data and AI platform developer @databricks has raised an eye-popping $10 billion in a new round of funding that boosts the company’s valuation to $62 billion, the company said today.
https://x.com/CRN/status/1869077611353092200
Meta/Llama
“Just 10 days after o1’s public debut, we’re thrilled to unveil the open-source version of the groundbreaking technique behind its success: scaling test-time compute đź§ đź’ˇ By giving models more “time to think,” LLaMA 1B outperforms LLaMA 8B in math—beating a model 8x its size.
https://x.com/ClementDelangue/status/1868740932251844806
“How we implemented test-time computing for open models to solve complex math problems like @OpenAI o1. đź‘€ Test-time compute methods use dynamic inference strategies to have LLMs “think longer” on harder problems, e.g. difficult math problems. By scaling test-time compute,
https://x.com/_philschmid/status/1868919520741445797
“We outperform Llama 70B with Llama 3B on hard math by scaling test-time compute 🔥 How? By combining step-wise reward models with tree search algorithms 🙂 We show that smol models can match or exceed the performance of their much larger siblings when given enough “time to
https://x.com/_lewtun/status/1868703456602865880
“🔬 Mind-blowing: Small language models (1B-3B params) outperform 8B and 70B models on math problems when given “time to think”! New research from @HuggingFace shows how test-time compute scaling can make tiny models mighty. We’re open sourcing the full recipe and sharing a
https://x.com/fdaudens/status/1868764037669925154
“🚀 Just extracted chart data from a PDF & recreated it in under 1 min! Llama Vision on @HuggingChat is a game-changer for data visualization work. No more manual data entry headaches! 📊✨ #DataScience #ProductivityHack
https://x.com/fdaudens/status/1869488842606035346
Meta launches Llama 3.3, shrinking powerful 405B open model | VentureBeat
Meta launches open source Llama 3.3, shrinking powerful bigger model into smaller size
“As we wrap up 2024, we’re sharing an update on our progress with Llama and the impact it’s having around the world. Read the full update here ➡️
https://x.com/AIatMeta/status/1869775975917257037
Phi
Introducing Phi-4: Microsoft’s Newest Small Language Model Specializing in Complex Reasoning | Microsoft Community Hub




