“Hugging Face has quietly become the biggest AI app store with 400,000 total apps, 2,000 new apps created every day, getting visited 2.5M times every week! Now you can search through any of them with AI or categories. The future of AI will be distributed, have fun everyone!
https://x.com/ClementDelangue/status/1886861567326650526
“llama.cpp RTX 50 benchmarks on the *Distilled* DeepSeek models” / X
https://x.com/ggerganov/status/1885426263243862263
“🚀 Announcing the all new @MistralAI Le Chat – your ultimate AI sidekick for life & work, now live on mobile! So many exciting features 🧵:
https://x.com/sophiamyang/status/1887517050697842899
“OpenAI’s Deep Research alternative by @huggingface They replicated it with an open-source version that already scores 54% on the same GAIA benchmark on which OpenAI’s scored 67%. → The GAIA benchmark includes complex, multi-step queries requiring multimodal understanding,
https://x.com/rohanpaul_ai/status/1886947545890611591
microsoft/phi-4 · Hugging Face
https://huggingface.co/microsoft/phi-4
“Deepseek V3 beaten! @allen_ai released Tülu 3 405B, an open-source post-training model built on Llama 3.1 405B. It outperforms Deepseek V3 the base model behind Deepseek R1 and is on par with @OpenAI GPT-4o. 👀 Tülu 3 405B is not a Reasoning model like R1 or o1, but team shared
https://x.com/_philschmid/status/1885253101214404813
“Best way to start the week – install llama.vscode More than 1000 installs so far!
https://x.com/ggerganov/status/1886313165710917968
“this is what a happy llama.cpp user looks like
https://x.com/ggerganov/status/1886493193518100727
“I find a lot of the discussion about Deep Research to be based on assumptions & a history of AI failing at citations Yet the expert researchers who’ve tried to use it find it very impressive. The criticism is like assuming you can still use 6 fingers on a hand to spot AI images
https://x.com/emollick/status/1886870485654512094
“idk if this is news but R1-low-mid-high are coming soon it seems
https://x.com/teortaxesTex/status/1887403244046995458
Why everyone is freaking out about DeepSeek | The Verge
https://www.theverge.com/ai-artificial-intelligence/598846/deepseek-big-tech-ai-industry-nvidia-impac
“The @CohereForAI team is hiring a research executive partner to make our model launches spark joy ✨, … drive the many cross-institutional collaborations which define our work 🌎 and shape the future of what’s possible with research at the frontier. 🚀
https://x.com/sarahookr/status/1885073573116612741
“More information and confirmations, Deepseek! @SemiAnalysis_* released a thorough report confirming that the $6M training figure is misleading and that it has access to thousands of GPUs. TL;DR; (free version): ✅ $6M “training cost” is misleading – excludes infrastructure” / X
https://x.com/_philschmid/status/1885264300450754594
“Nice example of the continuing jagged frontier of AI. Deep Research is really impressive in many fields, but apparently not great at math research. The only way to figure this out is trial & error (and building intuition). That can make using these systems confusing to start.” / X
https://x.com/emollick/status/1886931823717949723
“671-billion-parameter DeepSeek-R1 model at up to 3,872 tokens per second
https://x.com/_akhaliq/status/1885150800256680409
“Is the news about DeepSeek as big a deal as everyone is making out? Yes, it’s Sputnik. It is Sputnik 2.0. You know that story about how NASA spent a million dollars designing a pen that could write in space and the Russians brought a pencil? That just happened again.
https://x.com/JonathanRoss321/status/1886784302513377379
“Wow, someone just released a notebook to train a reasoning LLM with the new RL algorithm from DeepSeek, GRPO. In <2 hours, you can transform a very small model, Qwen 0.5 (500 million parameters) into a tiny math reasoning machine.
https://x.com/LiorOnAI/status/1886850811378196685
“Chinese startup DeepSeek dropped an open-source image-generation model, Janus-Pro It tops DALL-E 3 and StabIe Diffusion in image quality and accuracy benchmarks Another big move for open source, regardless of the R1 controversy
https://x.com/adcock_brett/status/1886098038520787198
“I didn’t know that DeepSeek-R1 has been downloaded 1.2 million times from HF already. DeepSeek-R1 was launched on Jan 20. That’s crazy!
https://x.com/omarsar0/status/1887259405579649411
“Chat, we getting a $56 Million open source LLM on Hugging Face! 🔥 🇪🇺/ acc
https://x.com/reach_vb/status/1886510019140612607
Who is Liang Wenfeng? DeepSeek founder comes from AI investing | TechCrunch
Who is Liang Wenfeng? DeepSeek founder comes from AI investing
“🔥 Qwen2.5-Max is now ranked #7 in the Chatbot Arena, surpassing DeepSeek V3, o1-mini and Claude-3.5-Sonnet. It is ranked 1st in math and coding, and 2nd in hard prompts. 👉🏻 Try Qwen2.5-Max here:
https://x.com/Alibaba_Qwen/status/1886485743998279944
Hugging Face researchers are trying to build a more open version of DeepSeek’s AI ‘reasoning’ model | TechCrunch
Hugging Face researchers are trying to build a more open version of DeepSeek’s AI ‘reasoning’ model
Alibaba’s Qwen team releases AI models that can control PCs and phones | TechCrunch
Alibaba’s Qwen team releases AI models that can control PCs and phones
“Heading to Paris with @Huggingface team for the AI Action Summit! Looking forward to great discussions on open-source AI. Plus we will be co-hosting a party at Station F! Want to meet up and connect with the team? DM me about open-source or grabbing coffee – always eager to
https://x.com/fdaudens/status/1886429525409509379
“🚀 Big updates are here on Qwen Chat ! Visit
https://x.com/Alibaba_Qwen/status/1886105723047973138
“This work is led by @TairanHe99 in collaboration with CMU lab @GuanyaShi We open-source everything: paper & code! Check it out:
https://x.com/DrJimFan/status/1886824977191854327
Open-source DeepResearch – Freeing our search agents
https://huggingface.co/blog/open-deep-research
“If you like MLX and you like Rust, check-out mlx-rs. Comes with examples of text generation with Mistral and MNIST training:
https://x.com/awnihannun/status/1886846423905575330
“Surprised we haven’t seen more about Deepseek r1-zero (no one seems to host it?) Unlike r1, which was trained to “think” in a readable, kinda charming way, r1-zero is the self-trained reasoner that had the *aha moment* about math & produces “thoughts” that are not human readable
https://x.com/emollick/status/1885150199279972717
“Free Deepseek R1 on OpenRouter 🤔
https://x.com/rohanpaul_ai/status/1885675708824879391
“🎧 Wow, that finale in Planet Money’s latest on DeepSeek! Must-listen episode feat. @lvwerra breaking down AI’s democratization in the most fun, accessible way possible 🤖✨
https://x.com/fdaudens/status/1885858178396463128
“deepseek R1 at 44 Output Tokens per Second on @replicate is super fast developers can get started easily with ai-gradio pip install –upgrade “ai-gradio[replicate]” import gradio as gr import ai_gradio gr.load( name=’replicate:deepseek-ai/deepseek-r1′, src=ai_gradio.registry,
https://x.com/_akhaliq/status/1885385810419044623
“Beautiful piece by @huggingface on Open-R1’s effort to replicate the DeepSeek-R1 pipeline and dataset. Some of the Key Takeaways – Some of the Evaluation results closely match DeepSeek’s benchmarks, but handling its massive 6,000+ token responses remains a challenge. →
https://x.com/rohanpaul_ai/status/1886180331675422728
DeepSeek temporarily suspends API service top-ups | Reuters
https://www.reuters.com/technology/artificial-intelligence/deepseek-temporarily-suspends-api-service-top-ups-2025-02-06/
“agents will actually work once tool calling gets built into reasoning models deepseek r1 + web search is a v0 of this. its also why the reasoning traces are *very* important (and great to read)” / X
https://x.com/SullyOmarr/status/1884351883307085972
DeepSeek R1 struggles with its identity – and more • The Register
https://www.theregister.com/2025/01/27/deepseek_r1_identity/
“We saw this with DeepSeek v3 when it was fine-tuned on reasoning data. Looks like reasoning data really hurts instruction following” / X
https://x.com/nrehiew_/status/1885392655489663271
“Introducing DeepSeek R1 Trend Finder 🔎 It gets posts from key influencers using @firecrawl_dev and the X API, finds any trends with R1, then pings you in slack. Check it out:
https://x.com/ericciarla/status/1886456419240829193
“open Sourced the r1 chat template! – easy integration with @aisdk! – works with deepseek r1 providers like @GroqInc! repo here →
https://x.com/zaidmukaddam/status/1885041197711781898
“No reasoning model consistently solves this puzzle, but DeepSeek’s thought here was insane (it also got it wrong): “A young boy who has been in a car accident is rushed to the emergency room. Upon seeing him, the surgeon says, “I can operate on this boy!” How is this possible?”
https://x.com/emollick/status/1886271713970704506
“The lack of good benchmarks is pretty glaring with DeepSeek and other reasoning models. These models are clearly good at some things & are extremely mediocre at others (DeepSeek is not good at multi-turn conversation, for example). MMLU and standard scores don’t help with that.” / X
https://x.com/emollick/status/1885159401150984370
DeepSeek might have a trademark problem in the US | TechCrunch
Taiwan bans government departments from using DeepSeek AI | Reuters
https://www.reuters.com/technology/taiwan-bans-government-departments-using-deepseek-ai-2025-02-03/
One very interesting observation about DeepSeek-R1: few-shot prompting actually degrades its performance. I’m not 100% certain of the cause, but the lack of good few-shot learning performance is likely due to the fact that R1 is trained to adhere to a strict format. During RL, we apply a formatting loss that rewards the model for outputting a long CoT reasoning trace followed by a final answer / summary. Few-shot examples don’t naturally fit into this reasoning format.
https://x.com/cwolferesearch/status/1886810699726275041
“Expect the next UX battle in AI to be over what “thinking tokens” to show DeepSeek found that producing less pleasant (mixed languages, etc) thoughts might get better outcomes, but people didn’t like them. They trained r1 to show nicer thoughts, which makes DeepSeek nice to use
https://x.com/emollick/status/1885747742472896897
“People don’t talk enough about a giant DeepSeek achievement over most US models – it actually has a reasonable name.” / X
https://x.com/emollick/status/1884847466857873837
“S1 Simple Scaling: SFT Qwen 32B Outperforms O1 preview on Math by a whooping 27% 🔥 > s1-32B outperformed o1-preview on competition math questions (MATH and AIME24) by up to 27% > with budget forcing, s1-32B improved from 50% to 57% on AIME24 > created a small dataset, s1K,
https://x.com/reach_vb/status/1886378372483186885
“IBM released Granite-Vision-3.1-2B, a small vision LM with impressive performance on different tasks 😮🔥 it comes with transformers and vLLM support from the get-go 💗 you can run it in Colab T4, so I built a notebook to put it to test, find it on the next one ⤵️
https://x.com/mervenoyann/status/1887521464292614382
“Let’s gooo – DeepSeek just released an OFFICIAL DEMO for DeepSeek VL2 Small. TL;DR It’s really powerful at OCR, text extraction and chat use-cases! 🔥
https://x.com/reach_vb/status/1887094223469515121
“Hugging Face announces SmolLM2 When Smol Goes Big — Data-Centric Training of a Small Language Model
https://x.com/_akhaliq/status/1887371050628903065
“🚀 Install any @huggingface Spaces app (transcription, OCR, TTS, AI scraping) as a PWA on your desktop! Transform web apps into local tools – no coding needed.” / X
https://x.com/fdaudens/status/1887243422009786821
“🎯 Kokoro TTS just hit v1.0! 🚀 Small but mighty: 82M parameters, runs locally, speaks multiple languages. The best part? It’s Apache 2.0 licensed! This could unlock so many possibilities ✨
https://x.com/fdaudens/status/188496764787324135
“@MistralAI Performs competitively with open weight models three times its size and with proprietary GPT4o-mini model across Code, Math, General knowledge and Instruction following benchmarks.
https://x.com/sophiamyang/status/1884973000820142266
“1/10 Mistral just released Mistral Small 3 (24B). How does it compare to the models Mistral benchmarks it against [1]? We found that Mistral Small 3 is competitive on GPQA Diamond but does worse on MATH Level 5, ranking below Qwen 2.5 32B and GPT-4o Mini in math reasoning.
https://x.com/EpochAIResearch/status/1885117404755235158
“This week in open AI was 🔥 Let’s recap! 🤗 LLMs 💬 > Huge: @allen_ai released new Tülu models that outperform DeepSeek R1 using Reinforcement Learning with Verifiable Reward (RLVR) based on Llama 3.1 405B 🔥 > @MistralAI is back to open-source with their “small” 24B models
https://x.com/mervenoyann/status/1885389118328242589
“The latest @MistralAI model, Mistral-Small-3-24B, is now live in the Arena! A compact yet powerful model: 81% MMLU, $0.1 (input/M tokens), and Apache 2.0 license. Check out its performance and compare at
https://x.com/lmarena_ai/status/1885372975035408840
“@MistralAI launches Mistral Small 3, a 24B parameter open weights model with similar intelligence to GPT-4o mini Key details: ➤ Artificial Analysis Quality Index jump from 61 to 72, compared to the Sep ‘24 version of Mistral Small ➤ 24B parameter dense model ➤ Apache 2.0
https://x.com/ArtificialAnlys/status/1885328803331006743
“Mistral Small 3 is available on Mistral, @togethercompute and @FireworksAI_HQ at launch, with Mistral’s first party API offering the lowest price.
https://x.com/ArtificialAnlys/status/1885328807269458398
“DeepSeek r1 is exciting but misses OpenAI’s test-time scaling plot and needs lots of data. We introduce s1 reproducing o1-preview scaling & performance with just 1K samples & a simple test-time intervention. 📜
https://x.com/Muennighoff/status/1886405528777073134
“LMAO Qwen 2.5 VL can perform Computer Use, out of the box, taking on OpenAI Operator HEAD ON! 🐐
https://x.com/reach_vb/status/1883961488320389376
“s1 Simple test-time scaling After supervised finetuning the Qwen2.5- 32B-Instruct language model on s1K and equipping it with budget forcing, our model s1-32B exceeds o1-preview on competition math questions by up to 27% (MATH and AIME24). Further, scaling s1-32B with budget
https://x.com/_akhaliq/status/1886244987551052061
“After DeepSeek and Kimi, there’s new agentic multimodal model from china that almost beats Claude Sonnet 3.5, Gemini-2 Flash & GPT-4o. Qwen 2.5 VL lets you create $200/month OpenAI Operator like agent for Computer and mobile phones. And it’s 100% Opensource. Let that sink in.
https://x.com/Saboo_Shubham_/status/1884080762863177902
“🤖📊 LangGraph: Dynamic Agent Workflows A powerful LangChain extension that supercharges AI agent systems with cyclical workflows. Build sophisticated multi-agent LLM applications with dynamic state management. Learn more about building advanced AI agents with LangGraph:
https://x.com/LangChainAI/status/1885372475057344881
“Dropping a clone of Deep Research in 12 hours 🔭 An AI Agent that reasons large amounts of web data extracted with @firecrawl_dev Open source. Powered by the @aisdk
https://x.com/nickscamara_/status/1886287956291338689
“Introducing Open Deep Research 🔭 An open source AI Agent that reasons large amounts of web data extracted with @firecrawl_dev Open source. Powered by the @aisdk
https://x.com/nickscamara_/status/1886459999905521912
“I built a Meme Generator web AI Agent with browser use that can automatically generate a meme by operating the browser. It works with DeepSeek, Claude Sonnet 3.5 and OpenAI GPT-4o. 100% Opensource Code with step-by-step tutorial.
https://x.com/Saboo_Shubham_/status/1884443229124522080
“📊 R1 just built its own download dashboard! Some fresh stats: +6M downloads for 800+ derivative models vs 2M for originals. Watch the numbers grow. 🚀
https://x.com/fdaudens/status/1886133195306930423
“i work with AI agents and their developers all day every day. here’s how i believe DeepSeek will fundamentally change how we see AI agents 🧵
https://x.com/braelyn_ai/status/1884002410391330958
“Actually serious here. R1 works in a very brute force try all approaches way and so I see approaches that I would never have thought of or edge cases that I would have forgotten about.” / X
https://x.com/nrehiew_/status/1885343197372531139
“Fluid (vercel.com/fluid) tl;DR: ◆ Serverless DX, Servers efficiency ◆ Full @nodejs & Python runtimes (more soon) ◆ `waitUntil`, streaming, cold start prevention… ◆ Even faster and more efficient free hobby tier ◆ One click to enable & zero-config experience 👇
https://x.com/rauchg/status/1886855760321372365
“Happy to release SWE Arena, your vibe coding platform! SWE Arena supports real-time code execution and rendering, covering various frontier LLMs & VLMs! We actually had this idea two years ago inside @BigCodeProject with @ArjunGuha and @dan_fried. However, there wasn’t much tech” / X
https://x.com/terryyuezhuo/status/1886450697497120891
“SWE Arena looks amazing for vibe coding Arena supports real-time code execution and rendering, covering various frontier LLMs & VLMs
https://x.com/_akhaliq/status/1886452970520293864
“My new favorite bookmark: like history on @huggingface — perfect way to spot what’s catching fire in the AI community 🔥
https://x.com/fdaudens/status/1884771012476047417
“We should make a @raycastapp extension for Hugging Face Inference Providers!” / X
https://x.com/reach_vb/status/1885337481639346641
“TIGER-Lab replaced answers in SFT with critiques. They claim superior performance in reasoning tasks without any
https://x.com/maximelabonne/status/1885291354852393216
stepfun-ai/GOT-OCR-2.0-hf · Hugging Face
https://huggingface.co/stepfun-ai/GOT-OCR-2.0-hf
vidore/colqwen2-v0.1 · Hugging Face
https://huggingface.co/vidore/colqwen2-v0.1
S1: The $6 R1 Competitor? Tim Kellogg shares his notes on a new paper, s1: Simple test-time scaling, which describes an inference-scaling model fine-tuned on top of Qwen2.5-32B-Instruct for just $6 – the cost for 26 minutes on 16 NVIDIA H100 GPUs.
https://simonwillison.net/2025/Feb/5/s1-the-6-r1-competitor/
S1: The $6 R1 Competitor? – Tim Kellogg
https://timkellogg.me/blog/2025/02/03/s1




