“Our latest technical report is here! The Qwen2.5 series has been widely recognized and loved, and we’re thrilled to see technology enabling more wonderful things to happen. Enjoy it!”
https://x.com/huybery/status/1869952907677991200

Finally, a Replacement for BERT: Introducing ModernBERT
https://huggingface.co/blog/modernbert

“We’ve been building LOTUS at Stanford and Berkeley to make LLM-powered data processing fast, easy and declarative. LOTUS is an open-source query engine that makes programming as easy as writing Pandas and optimizes your programs for up to 400x speedups. To celebrate the
https://x.com/lianapatel_/status/1869454886351389001

“🚀 Introducing DeepSeek-V3! Biggest leap forward yet: ⚡ 60 tokens/second (3x faster than V2!) 💪 Enhanced capabilities 🛠 API compatibility intact 🌍 Fully open-source models & papers 🐋 1/n
https://x.com/deepseek_ai/status/1872242657348710721

“DISTILLED REASONING CAPABILITIES FROM DEEPSEEK R1 🤯
https://x.com/reach_vb/status/1872246633649553556

[2412.08905] Phi-4 Technical Report
https://arxiv.org/abs/2412.08905

“They have 2 types of RL rewards. Verifiers (code, math) and standard model based RM. Importantly the model based RM is trained COT style GRPO from deepseek math used here
https://x.com/nrehiew_/status/1872318217395572895

“Open Science FTW! That’s a fully open weight model, that you can deploy wherever you want, however you want mogging closed source APIs!
https://x.com/reach_vb/status/1871281525075140846

“🎄Happy holidays and we wish you enjoy this year. Before moving to 2025, Qwen has the last gift for you, which is QVQ! 🎉 This may be the first open-weight model for visual reasoning. It is called QVQ, where V stands for vision. It just reads an image and an instruction, starts
https://x.com/Alibaba_Qwen/status/1871602879972405626

“Qwen2.5 advances LLM capabilities through expanded pre-training data (18T tokens), sophisticated post-training techniques, and efficient architecture optimizations for enhanced performance across scales. Solution/Methods in this Paper ⚡: → Pre-training data expanded from 7T”
https://x.com/rohanpaul_ai/status/1870940304314102271

So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer – The GitHub Blog

So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer

“🔥 While everyone’s focused on big LLMs, encoder models quietly power 1B+ monthly downloads – 3x more than all LLMs combined! And after 6 years, these AI workhorses got a major upgrade: Meet ModernBERT. 2x faster, 8K context, runs on consumer GPUs 🚀 #AI #MachineLearning”
https://x.com/fdaudens/status/1869790286643380356

“Not your weights, not your brain! Starting today, you can run your private GGUFs from the Hugging Face hub directly in @ollama! 🔥 You asked, we delivered! Works out of the box, all you need to do is add your Ollama SSH key to your profile, and that’s it! ⚡ Run private”
https://x.com/reach_vb/status/1871287195702833214

“You can now get on-demand GH200/ H200/ H100 directly via your @huggingface account, powered by @LambdaAPI 🔥 What’s your excuse to not learn/ practice this holidays? 🤗”
https://x.com/reach_vb/status/1871159830024790319

Serverless LoRA Inference
https://docs.together.ai/docs/lora-inference

GuanjieChen/Skip-DiT · Hugging Face
https://huggingface.co/GuanjieChen/Skip-DiT

DeepSeek
“Holy f-ck! They also dropped the Instruct model on the Hub – that’s literally the same model that runs on DeepSeek Chat! 🔥 That’s the best open weight LLM right now and second best on AiderBench (after o1) Now we wait for the model card! 🫡

https://x.com/reach_vb/status/1872193295876813128
“🌌 Open-source spirit + Longtermism to inclusive AGI 🌟 DeepSeek’s mission is unwavering. We’re thrilled to share our progress with the community and see the gap between open and closed models narrowing. 🚀 This is just the beginning! Look forward to multimodal support and” / X

https://x.com/deepseek_ai/status/1872242666265801105
Meta/Llama

“Through experimentation @LinkedIn found EON-8B, a domain-adapted version of Llama 3.1 8B, to be 75x and 6x cost effective in comparison to GPT-4 and GPT-4o respectively. More on their domain-adapted foundation model work ➡️
https://x.com/AIatMeta/status/1871337917786026056

“Yes they do train on R1. You know, I’m used to it but it’s bizarre how their papers increasingly differ from the competition. Llama 3, Qwen 2.5 are just “more dakka bruh”. This is… like 10 papers in one, a condensed textbook from 2 years into the future of open source AGIs.
https://x.com/teortaxesTex/status/1872250466987545056

The future of AI: Built with Llama
https://ai.meta.com/blog/future-of-ai-built-with-llama/

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading