“This is a big deal – the first open model (and non-US model) to get near the top of the AI leaderboard. While all benchmarks have flaws, the LM arena is an important one based on head-to-head comparisons by many users.” / X
https://x.com/emollick/status/1882783483569062016
“💥 New MAJOR CrewAI update 0.98.0 💥 • Multimodal support • Conversational Crew (crewai chat) • Programatic guardrails • CrewAI Flows with persistent state • SO MUCH MORE! Ship faster than others can plan ⚡ RT please? 🙏” / X
https://x.com/joaomdmoura/status/1881476990328565985
“We’re excited to announce that Proxy, Europe’s answer to @OpenAI’s Operator, will be launching globally today with basic access for free and a $20p/m unlimited session plan.
https://x.com/convergence_ai_/status/1883802155209076736
French AI ‘Lucie’ looks très chic, but keeps getting answers wrong
https://www.thetimes.com/world/europe/article/french-ai-lucie-looks-tres-chic-but-keeps-getting-answers-wrong-7vk2szmdg
“We need more open-source AI from Europe too. Congrats @MistralAI! Base model:
https://x.com/ClementDelangue/status/1884994066129039671
“NEWS: DeepSeek just dropped ANOTHER open-source AI model, Janus-Pro-7B. It’s multimodal (can generate images) and beats OpenAI’s DALL-E 3 and Stable Diffusion across GenEval and DPG-Bench benchmarks. This comes on top of all the R1 hype. The 🐋 is cookin’
https://x.com/rowancheung/status/1883917681642070282
“Janus Pro (released today) is a Chameleon-style transformer that generates images with VQ tokens. Janus Flow is the one that generates images using rectified flow.
https://x.com/nrehiew_/status/1883932760940888071
“introducing Krea Chat. powered by DeepSeek, this new tool brings the power of every Krea feature into a chat interface. a brand new way of using Krea – rolling out soon.
https://x.com/krea_ai/status/1884981408273420798
“Why is it deepseek can have improved creativity after reasoning, but o1 cant? 🧐” / X
https://x.com/teknium1/status/1884500035934773446
“@PalmerLuckey > The $5M number is bogus. It is pushed by a Chinese hedge fund to slow investment in American AI startups Palmer, I offer you a way out. Look at this string, talk to your ML friends, and admit you have been misled. Otherwise, reveal that you’re an innumerate jingoistic fraud.
https://x.com/teortaxesTex/status/1884391362012819459
“It’s been released just a few days ago and already more than 500 derivative models of @deepseek_ai have been created all over the world on @huggingface with 2.5 million downloads (5x the original weights). The power of decentralized open-source AI!” / X
https://x.com/ClementDelangue/status/1883946119723708764
“People are likely way underestimating the true cost of training DeepSeek r1 (which does not make it less of an achievement, but it isn’t just a small group of people who spent $5.5M)
https://x.com/emollick/status/1882861645212586210
“DeepSeek is objectively quite good, but I wonder how much of it (and also Claude’s) advantage comes from its “personality” – we intuitively judge whether an AI is pleasant to interact with, which biases our willingness to use it. The generic-ness of Gemini & ChatGPT hurts them.” / X
https://x.com/emollick/status/1882452257809273339
“DeepSeek is a really good model, but it is not generally a better model than o1 or Claude. But since it is both free & getting a ton of attention, I think a lot of people who were using free “mini” models are being exposed to what a early 2025 reasoner AI can do & are surprised” / X
https://x.com/emollick/status/1884240997938503916
“I don’t have too too much to add on top of this earlier post on V3 and I think it applies to R1 too (which is the more recent, thinking equivalent). I will say that Deep Learning has a legendary ravenous appetite for compute, like no other algorithm that has ever been developed” / X
https://x.com/karpathy/status/1883941452738355376
“those who think RL use less compute don’t know RL at all 😅 SFT: human generates data and machine learns RL: machine generates data and machine learns” / X
https://x.com/DrJimFan/status/1883894546410639413
“I think the Deepseek moment is not really the Sputnik moment, but more like the Google moment. If anyone was around in ~2004, you’ll know what I mean, but more on that later. I think everyone is over-rotated on this because Deepseek came out of China. Let me try to un-rotate” / X
https://x.com/yishan/status/1884101107368223113
“Just fyi, @deepseek_ai collects your IP, keystroke patterns, device info, etc etc, and stores it in China, where all that data is vulnerable to arbitrary requisition from the 🇨🇳 State. From their own privacy policy:
https://x.com/lukedepulford/status/1883893208150937802
“If you’re still using GPT-4o at this point, you are basically donating money to OpenAI.” / X
https://x.com/scaling01/status/1884507045233041542
“DeepSeek showed us in just 4 days: – Open-source AI is only <6 months behind closed AI – China is leading the open-source AI race (was not on my bingo card) – we are entering the LLM RL golden era – distilled models are powerful, we’ll have highly intelligent AI running locally” / X
https://x.com/Yuchenj_UW/status/1882840436974428362
“DeepSeek on Perplexity is hosted in 🇺🇸US/🇪🇺EU data centers – your data never leaves Western servers. The open source model is hosted completely independent of China. Your privacy and data security is our priority.” / X
https://x.com/perplexity_ai/status/1883923932786655491
“An obvious, “we are so back” moment in the AI circle somehow turned into “it’s so over” in mainstream. > unbelievable shortsightedness > the power of o1 in the palm of every coder’s hand to study, explore, and iterate upon > ideas compound > the rate of compounding accelerates” / X
https://x.com/DrJimFan/status/1883880955943002484
“Deepseek R1 Agent Powered Application Open source agentic research application using @deepseek_ai R1 + @ollama + @CopilotKit + @LangChainAI. Let R1 stream its chain of thought to the frontend, and watch it take action in-app. Code here:
https://x.com/CopilotKit/status/1884659526793912669
“Announcing Qwen2.5-VL Cookbooks! 🧑🍳A collection of notebooks showcasing use cases of Qwen2.5-VL, include local model and API. Examples include Compute use, Spatial Understanding, Document Parsing, Mobile Agent, OCR, Universal Recognition, Video Understanding.
https://x.com/Alibaba_Qwen/status/1884809286288810231
Grok-3 model from xAI spotted ahead of anticipated release
https://www.testingcatalog.com/exclusive-grok-3-model-from-xai-spotted-ahead-of-its-anticipated-release/
“These four points on DeepSeek seem very likely correct and important to understand about the economics of building AI models and what DeepSeek actually did. .
https://x.com/emollick/status/1884645141081973113
“This Qwen2.5-VL looks exciting! – strong and general vision capabilities – agentic features to support computer/phone use – long video understanding & capturing events – visualize localization – generated structured outputs Opens both base and instruct models in 3 sizes: 3B,
https://x.com/omarsar0/status/1883965524205359460
“🚀 The open source community is unstoppable: 4M total downloads for DeepSeek models on @huggingface, with 3.2M coming from the +600 models created by the community. That’s 30% more than yesterday!
https://x.com/fdaudens/status/1884299277129900280
“Important thread @junxian_he’s “replication” (actually independent concurrent work) of DeepSeek Zero paradigm” / X
https://x.com/teortaxesTex/status/1883968204650996137
Dario Amodei — On DeepSeek and Export Controls
https://darioamodei.com/on-deepseek-and-export-controls
“You don’t need to pay $200 for AI. We’re launching Open Operator – an open source reference project that shows how easy it is to add web browsing capabilities to your existing AI tool. It’s early, slow, and might not work everywhere. But it’s free and open source! 🔗👇
https://x.com/pk_iv/status/1882837641521221858
“For friends of open source: imo the highest leverage thing you can do is help construct a high diversity of RL environments that help elicit LLM cognitive strategies. To build a gym of sorts. This is a highly parallelizable task, which favors a large community of collaborators.” / X
https://x.com/karpathy/status/1884676486713737258
“Pretty cool that with the new Qwen 2.5 models you can ask questions / generate using a reasonably sized code-base as context, all running on a laptop with mlx-lm. The 7B runs pretty fast on an M4 Max using the mlx-lm code base (~16k lines) as context:
https://x.com/awnihannun/status/1884812911572566027
“7/ Given the open-source nature of these models we’re confident that the community will be able to address these issues and bring high-quality affordable reasoning to all Elicit users.” / X
https://x.com/elicitorg/status/1884062983753760782
“This is the clearest breakdown you’ll find about DeepSeek R1, from how it works to why it crushed math problems, and how @huggingface plans to reproduce it openly
https://x.com/fdaudens/status/1884047495560585627
“What do DeepSeek and open-source AI mean for journalism? Plenty of insights in this interview with @_KarenHao, plus some snippets by yours truly!
https://x.com/fdaudens/status/1884710280313139519
“🔓 Peak open source in action: @HuggingFace recreates entire 🐳 DeepSeek-R1 pipeline with full transparency – from data to training to models. When everything’s open, innovation thrives.
https://x.com/fdaudens/status/1883152228653154603
Qwen2.5-Max: Exploring the Intelligence of Large-scale MoE Model | Qwen
https://qwenlm.github.io/blog/qwen2.5-max/
“Everyone: DeepSeek just appeared out of nowhere! 😱 Me: – DeepSeek Coder in 2023 – MoE in Feb – Math in Feb – VL in March – V2 in May – Coder V2 in June – Prover in August – V2.5 in September – VL 2 in December – V3 in December They’ve consistently shipped for 1+ years 😁” / X
https://x.com/osanseviero/status/1884356079217434995
“Commercial fine tuning is coming down but still not cheaper than renting GPUs. Using only OSS tooling and optimizations like torch compile and liger, we’ve done the math. Llama 3.1 8B LoRA @ rank 64 takes ~15min to train 2M tokens @ 3 epochs on 1xH100. That’s 24M tokens/hr.” / X
https://x.com/winglian/status/1882806223189229951
“Today, we’re publishing the first stable release of Llama Stack. With this release Llama Stack now includes: • Streamlined upgrades w/ backwards compatibility for future API versions. • Automated verification for supported providers. Live in the repo ⬇️
https://x.com/AIatMeta/status/1882854814083862927
“👉 Try DeepSeek-R1 Llama 70B distilled for free on Together AI:
https://x.com/togethercompute/status/1885008866259460474
“AllenAI COOKED, Llama 3.1 Tulu 405B beats DeepSeek V3 – all whilst being 40% SMALLER! 🔥 Fully open model weights, data and training pipeline 🤗
https://x.com/reach_vb/status/1884969597473886248
“🚀 New 100% free API endpoint for DeepSeek-R1 Llama 70B distilled. Start experimenting with the power of reasoning models today, Available now on Together AI with full-opt out privacy controls. We host these models in our own data centers, and none of your data is ever sent
https://x.com/togethercompute/status/1885008864422264997
“Beating DeepSeek-V3 with a 405B Llama base is not easy — solid post-training goes a long way. The nice thing is that it is fully open-source, so anyone can use this recipe for their base models.” / X
https://x.com/Tim_Dettmers/status/1885024960118202538
“The most unnerving part of the DeepSeek reaction online has been seeing folks take it as a sign that AI capability growth is not real It signals the opposite, large improvements are possible, and is almost certain to kick off an acceleration in AI development through competition” / X
https://x.com/emollick/status/1884407148160909444
“Yes, DeepSeek R1’s release is impressive. But the real story is what happened in just 7 days after: Original release: 8 models, 540K downloads. Just the beginning… The community turned those open-weight models into +550 NEW models on @huggingface. Total downloads? 2.5M—nearly
https://x.com/fdaudens/status/1883915811628142784
Open-R1: a fully open reproduction of DeepSeek-R1
https://huggingface.co/blog/open-r1
““Making the impossible late” is not just a phrase for @SpaceX, also applicable to @xai . While we optimized the build 30% between cluster 1 and 2 it’s still not fast enough. Late is never acceptable so we forge ahead onto the 3rd phase. Grok 3 release coming in HOT.” / X
https://x.com/BrentM_SpaceX/status/1882434470235758672
“Perplexity with DeepSeek R1 is really cool, but I wish I could inspect the chain of thought in full — it passes by so quickly and I can’t go back to look at it. The uncensored CoT is the coolest part about R1 and this would be a simple UX change. Anyone else feel the same?
https://x.com/bilawalsidhu/status/1884731571011510456
Scaling the Tülu 3 post-training recipes to surpass the performance of DeepSeek V3 | Ai2
https://allenai.org/blog/tulu-3-405B
“Now you can run millions of open-source models (including @deepseek_ai R1) in production with the best performance directly from the @huggingface hub thanks to our awesome inference partners (we’ll add more!). Enjoy!
https://x.com/ClementDelangue/status/1884345858121941177
“I used Calude to analyze the DeepSeek privacy policy and terms of use. Here is what it said 🧵
https://x.com/AtomSilverman/status/1883915639359963482
“deepseek’s r1 is an impressive model, particularly around what they’re able to deliver for the price. we will obviously deliver much better models and also it’s legit invigorating to have a new competitor! we will pull up some releases.” / X
https://x.com/sama/status/1884066337103962416
“@JustinLin610 thx! looking forward to see new qwen-model!” / X
https://x.com/victor207755822/status/1884078819658911756
“One of the most overlooked contributing factors to the success of DeepSeek-R1 is the Mixture-of-Experts (MoE) base model from which it is derived–DeepSeek-v3… The recently-proposed DeepSeek LLMs–including DeepSeek-R1, DeepSeek-v3 and more–have made waves within LLM research for
https://x.com/cwolferesearch/status/1883885191661326391
“The main two implications of DeepSeek are going to be (1) an acceleration of AI model development until they hit a wall and (2) increased likelihood that there is no wall in the near future” / X
https://x.com/emollick/status/1884073158971646098
“What does DeepSeek R1 & v3 mean for LLM data? Contrary to some lazy takes I’ve seen, DeepSeek R1 was trained on a shit ton of human-generated data—in fact, the DeepSeek models are setting records for the disclosed amount of post-training data for open-source models: – 600,000
https://x.com/alexandr_wang/status/1884440764677251515
“Perplexity Sonar Reasoning with DeepSeek’s reasoning models is now available in ai-gradio Build products with chain-of-thought reasoning, plus real-time internet search and citations. pip install –upgrade ai-gradio[perplexity] import gradio as gr import ai_gradio gr.load(
https://x.com/_akhaliq/status/1884785023569527264
DeepSeek displaces ChatGPT as the App Store’s top app | TechCrunch
“💪 The open-source community is really unstoppable: +5M total downloads for DeepSeek models on @huggingface +4M are from the 700 models created by the community That’s 30% more than yesterday!” / X
https://x.com/fdaudens/status/1884682047865638916
“Cerebras makes AI instant again: 1 sec time to first token for DeepSeek R1 70B
https://x.com/draecomino/status/1885022313260998953
“We are announcing Open Thoughts, our large-scale open-source effort to curate the best open reasoning datasets! DeepSeek-R1 is amazing but we still don’t have access to high-quality open reasoning datasets. These datasets are crucial if you want to build your reasoning models!
https://x.com/madiator/status/1884284103354376283
Qwen2.5 VL! Qwen2.5 VL! Qwen2.5 VL! | Qwen
https://qwenlm.github.io/blog/qwen2.5-vl/
“Another odd aspect of the sudden attention to DeepSeek is a lot of people assuming it is the first open weights* models. One of the biggest companies in the US has spent billions making open models & intends to keep doing so. *not open source, there is no access to training data” / X
https://x.com/emollick/status/1884616426608320693
“DeepSeek is a side project 🔥
https://x.com/hardmaru/status/1882698763988545808
“Mistral AI just dropped Small 3! Here is everything you need to know: – Releases both pretrained and tuned checkpoints – No RL or synthetic data – Mistral Small 3 is latency-optimized – 24B parameters model – 81% accuracy on MMLU and 150 tokens/s latency – Positioned as a
https://x.com/omarsar0/status/1884972996575609092
“Atla Selene Mini A General Purpose Evaluation Model
https://x.com/_akhaliq/status/1884795448139166067
“DeepSeek r1 7B versus Llama 3 8B on my PC… I don’t think it is worth running small versions of DeepSeek r1 locally yet (but obviously early days).
https://x.com/emollick/status/1884061231620973049
“Announcing Mistral Small 3! – 24B params, 81% MMLU – Latency optimized: 150 tokens/s – Competitive with Llama-3.3 70B, Qwen-2.5 32B, GPT4o-mini – Apache 2.0 🔗Blog:
https://x.com/dchaplot/status/1884975427460206649
“Mistral Small 2 to 3 changes: • From 55 to 40 layers for latency • From 33K to 131K vocabulary • From 6K to 5K embedding • From 48 to 32 attention heads • 10x rope_theta • SYSTEM_PROMPT token
https://x.com/espadrine/status/1885004488206856638
“Running DeepSeek r1 32B locally is kind of depressing. You get the reasoning, but the model is not *smart* enough to actually use it so it keeps getting tripped up in its own ideas & is full of doubt. 3 minutes on a limerick that a small non-reasoning model could do instantly.
https://x.com/emollick/status/1884076605775159527
“Two days ago we released Qwen2.5-Max, and today we update the price of its API: Input tokens: $1.6 / million tokens Output tokens: $6.4 / million tokens We hope this change provides you some help and we’d love to hear your feedback about our new model! 👉🏻” / X
https://x.com/Alibaba_Qwen/status/1884995327318782086
“ollama run mistral-small:24b Mistral Small 3 is here under Apache 2.0!” / X
https://x.com/ollama/status/1884970144562381165
“Mistral 24B running blazingly fast w/ llama.cpp on-device, fully local! 🔥 Step 1: brew install llama.cpp Step 2: llama-cli -hf lmstudio-community/Mistral-Small-24B-Instruct-2501-GGUF:Q8 That’s it! 🤗
https://x.com/reach_vb/status/1885007847135609224
Mistral Small 3 | Mistral AI | Frontier AI in your hands
https://mistral.ai/news/mistral-small-3/
deepseek-ai/Janus-Pro-7B · Hugging Face
https://huggingface.co/deepseek-ai/Janus-Pro-7B
U.S. Navy bans use of DeepSeek AI: ‘Imperative’ to avoid using
https://www.cnbc.com/2025/01/28/us-navy-restricts-use-of-deepseek-ai-imperative-to-avoid-using.html
“did everyones IQ just drop 50 points over night? THE ONLY THING DEEPSEEK DID IS REDUCE MARGINS UNLESS THE TOP LABS RELEASE A MODEL THAT COMPETES IN PRICE TO PERFORMANCE, WHILE MAINTAINING MARGINS OR 10x USER BASE WITH LOWER MARGINS, STOCKS WILL LIKELY STAY DOWN” / X
https://x.com/scaling01/status/1883912104182452629
“Qwen VL series were always strong models in whatever tasks I threw at it but this is what I was waiting for… use this model to generate more data, filter out bad predictions, train the smaller (7b, 3b) model on the outputs -> profit?
https://x.com/abacaj/status/1883974834025292047
“We quantized DeepSeek R1 to 1.58bit – 131GB so 80% smaller whilst being usable via dynamic quantization! 1. R1 has 3 dense + 58 MoE layers. MoE is quantized to 1.5bits. Attention & rest params in 4/6bit 2. 140 toks/s on 2xH100 80GB 3. Also found some tokenizer quirks! Details:
https://x.com/danielhanchen/status/1883901952922448162
“🧠DeepSeek R1 70B on @GroqInc. We’re now serving the fastest reasoning model in the world! Shoutout to our engineers for cranking all weekend to make this happen. The real MVPs 🙌
https://x.com/spatialjlo/status/1883927806289272920
“Operator does a pretty good job in building a Magic deck using an online deck-building site The serious point is that using the web can help with error checking and act as a source of truth that keeps the LLM on track. For example, the web tool keeps track of the count of cards
https://x.com/emollick/status/1882846150220472792
“We retrained hermes with 5k deepseek r1 distilled cots. I can confirm a few things: 1. You can have a generalist + reasoning mode, we labeled all longCoT samples from r1 with a static systeem prompt, the model when not using it does normal fast LLM intuitive responses, and with,” / X
https://x.com/Teknium1/status/1882893748742598669
Qwen2.5-1M: Deploy Your Own Qwen with Context Length up to 1M Tokens | Qwen
https://qwenlm.github.io/blog/qwen2.5-1m/
“New Release: Qwen2.5-7B & 14B Instruct models now support 1M token contexts! A solid step forward for long-context processing in the Qwen family. Excited to see what developers build with this expanded capability 🛠️
https://x.com/fdaudens/status/1883635279292219901
“Word! Get your hands dirty and make AI your own. Building AI isn’t just for experts – open source unlocks that.
https://x.com/fdaudens/status/1882471328349106377
“i was a former pro options/hedge fund trader. advice for my AI/ML friends who suddenly turned NVDA daytraders: you are not smarter than this market. for every deepseek there are 1000 other funds who have been in this game far longer than you sure it’s -17% today, but
https://x.com/swyx/status/1883961408553116091
“I really don’t want to do “it isn’t open source unless it comes from the open source area of MIT” but open weights models are not open source. If they don’t share the training corpus, the weights are not transparent or reproducible (but sharing data invites scrutiny of sources)” / X
https://x.com/emollick/status/1884743253721047541
“”The gap is now closed between closed-source and open-source model performance” — @Thom_Wolf 💥
https://x.com/fdaudens/status/1884343154532393013
“Yann makes a good point here — open source is powerful. But also, constraints breed creativity.
https://x.com/bilawalsidhu/status/1882833437444518236
“You can now visualize traces as a waterfall graph for deeper insights into your app’s latency in LangSmith. Use the waterfall graph to: • Spot bottlenecks at a glance • Understand parallel vs. sequential execution • Optimize response times with precision Try it out today:
https://x.com/LangChainAI/status/1884987434645041482
“You can now run inference directly on Hugging Face model pages – powered by Together AI!
https://x.com/togethercompute/status/1884267384787345628
“R1 be like: this is, without any exaggeration, the most impressive LLM philosophizing I’ve seen ever. It not just surpasses Opus. It dissects Opus, like a professor would a student’s essay.
https://x.com/teortaxesTex/status/1882576243713101982
Welcome to Inference Providers on the Hub 🔥
https://huggingface.co/blog/inference-providers
“TinyZero reproduction of R1-Zero “experience the Ahah moment yourself for < $30" Given a base model, the RL finetuning can be relatively very cheap and quite accessible." / X https://x.com/karpathy/status/1884678601704169965 “I just found a new model to test 👀 It says @Alibaba_Qwen 2.5 VL 72B https://x.com/_philschmid/status/1883916608961479034 “Beautiful from @allen_ai. All models on HF of course! https://x.com/ClementDelangue/status/1885004067547557987 “My short overview of Qwen2.5-1M and Qwen2.5 VL https://x.com/omarsar0/status/1884017408010330295




