Bringing Open-Source Models to Spreadsheets 🚀 | by Florent Daudens | Generative AI in the Newsroom
https://generative-ai-newsroom.com/bringing-open-source-models-to-spreadsheets-c440fc4818b4
QVQ: To See the World with Wisdom | Qwen
https://qwenlm.github.io/blog/qvq-72b-preview/
Nvidia to open-source Run:ai, the software it acquired for $700M to help companies manage GPUs for AI | VentureBeat
“Top 25 AI models in 2024 on Hugging Face 🔥 @bfl_ml w/ Flux.1-dev & Flux.1-schnell – current SoTA open Text to Image models @AIatMeta w/ Llama 3.X series (1B to 70B) 🦙- competitive LLMs across sizes @StabilityAI w/ SD 3.5 Medium & Large @GoogleAI Gemma 7B 💎 @xai w/ grok-1
https://x.com/reach_vb/status/1873689441404965334
DeepSeek
Tom Dörr on X: “DeepSeek-V3: A 671B parameter MoE language model with 37B parameters activated per token, trained on 14.8 trillion tokens, using Multi-head Latent Attention and Multi-Token Prediction, and supported by FP8 mixed precision training https://t.co/cWfYHheqQM” / X – https://x.com/tom_doerr/status/1874031396013879744
Deepseek: The Quiet Giant Leading China’s AI Race
https://www.chinatalk.media/p/deepseek-ceo-interview-with-chinas
DeepSeek-V3, ultra-large open-source AI, outperforms Llama and Qwen on launch | VentureBeat
DeepSeek-V3, ultra-large open-source AI, outperforms Llama and Qwen on launch
“Just added DeepSeek-Engineer on Github 🐋 Wanted to test the API, so I created a quick coding assistant that can read, create, and diff edit files using structured outputs. It’s very simple and minimal, and a good foundation if you want to learn how coding assistants work!
https://x.com/skirano/status/1872382787422163214
vLLM on X: “⬆️pip install -U vLLM You can now run DeepSeek-V3 on latest vLLM many different ways: 💰 Tensor parallelism on 8xH200 or MI300x, or TP16 on IB connected nodes: `–tensor-parallel-size` 🌐 Pipeline parallelism (!) across two 8xH100 or any collection of machines without high speed” / X – https://x.com/vllm_project/status/1872453508127130017
“It’s funny that SOTA models like deepseek will gladly output they are trained by OAI (1.5 years after gpt-4) but maybe not so funny… how hard could it be to steer the model in post training to not admit it is trained by OAI/GPT-4? It’s obvious the model will output this when
https://x.com/abacaj/status/1872523867077407188
“Exciting News from Chatbot Arena❤️🔥 @OpenAI’s o1 rises to joint #1 (+24 points from o1-preview) and @deepseek_ai DeepSeek-V3 secures #7, now the best and the only open model in the top-10! o1 Highlights: – Achieves the highest score under style control – #1 across all domains
https://x.com/lmarena_ai/status/1873695386323566638
“Just published in my Newsletter today. Exploring DeepSeek-V3 Technical Report – They Just Changed The Game of AI Model Training (link in comment and bio – consider subscribing, its FREE and I write here daily )
https://x.com/rohanpaul_ai/status/1872754473715855773
“DeepSeek COOKED! It is the only open-weight (commercially permissive) model in the top 10. More so, it is a model that has been post-trained on ~5K H800 GPU hours (0.01M $) *only* – now imagine what a better post-trained model can do! 🔥
https://x.com/reach_vb/status/1873751949289705628
“Other interesting information: 1. 14.8 trillion tokens of pretraining data 2. They find that distillation from their reasoning model DeepSeek-R1 is extremely helpful for coding (LiveCodeBench) and math (MATH-500). But this ends up causing the model to produce very long” / X
https://x.com/madiator/status/1872505935832474018
“⬆️pip install -U vLLM You can now run DeepSeek-V3 on latest vLLM many different ways: 💰 Tensor parallelism on 8xH200 or MI300x, or TP16 on IB connected nodes: `–tensor-parallel-size` 🌐 Pipeline parallelism (!) across two 8xH100 or any collection of machines without high speed” / X
https://x.com/vllm_project/status/1872453508127130017
“🚀 Introducing DeepSeek-V3! Biggest leap forward yet: ⚡ 60 tokens/second (3x faster than V2!) 💪 Enhanced capabilities 🛠 API compatibility intact 🌍 Fully open-source models & papers 🐋 1/n
https://x.com/deepseek_ai/status/1872242657348710721
“Cool things from DeepSeek v3’s paper: 1. Float8 uses E4M3 for forward & backward – no E5M2 2. Every 4th FP8 accumulate adds to master FP32 accum 3. Latent Attention stores C cache not KV cache 4. No MoE loss balancing – dynamic biases instead More details: 1. FP8: First large




