Image created with Flux Pro v1.1 Ultra. Image prompt: OpenSource, elegantly peeled banana revealing interlocking modules of small bananas inside, modular metaphor, photorealistic, editorial, minimal, high detail, 3:2 landscape
China’s DeepSeek Preps AI Agent for End-2025 to Rival OpenAI https://finance.yahoo.com/news/china-deepseek-preps-ai-agent-152907224.html?guccounter=1&guce_referrer=aHR0cHM6Ly93d3cuZ29vZ2xlLmNvbS8&guce_referrer_sig=AQAAADiJs67uOGL7PzqX3MGgvD-A6UJVzmztcJfvPzJTz9iF2iWfg-h2zg2pcwJoIuJ-4IUs3BMrEvPbbpbf4j7qXCmM4BqK78UMZVzrZl3fSuokrWneWMYpy8S7L3-xciC9d74km3boS_g57OxikNZN7Owozd204A5KlQA0MSzkqp42
Apple open sourcing artefacts on HF is a special kind of joy! https://x.com/reach_vb/status/1961481909181075961
🚨 Apple just released FastVLM on Hugging Face – 0.5, 1.5 and 7B real-time VLMs with WebGPU support 🤯 > 85x faster and 3.4x smaller than comparable sized VLMs > 7.9x faster TTFT for larger models > designed to output fewer output tokens and reduce encoding time for high https://x.com/reach_vb/status/1961471154197053769
And FastVLM was released by Apple today! 🚀 All about on-device use. Model sizes: 0.5B, 1.5B, 7B. Available in MLX and Core ML. Vision encoder designed to output fewer tokens and reduce encoding time. Which means much faster time-to-first-token.”” / X https://x.com/pcuenq/status/1961464859465269757
Holy crap! That is some fast video captioning — all happening locally in your browser 🤯 This is the aptly named FastVLM by Apple; available on HF: https://x.com/bilawalsidhu/status/1962545148136444380
NEW: Apple releases FastVLM and MobileCLIP2 on Hugging Face! 🤗 The models are up to 85x faster and 3.4x smaller than previous work, enabling real-time VLM applications! 🤯 It can even do live video captioning 100% locally in your browser (zero install). Huge for accessibility! https://x.com/xenovacom/status/1961454543503344036
“To test models’ performance on Claude Code, we ran GLM-4.5 against Claude Sonnet 4 and other open-source models on 52 practical programming tasks. While GLM-4.5 demonstrated strong performance against top open-source models, it secured a 40.4% win rate against Claude Sonnet 4. https://x.com/Zai_org/status/1962522761630482700
🚀 Introducing slime v0.1.0 — An open-source RL infra powering models like GLM-4.5, built by THUDM & Zhipu AI. @Zai_org RL infra 朱小霖 shared a deep dive on Zhihu into how they redefined high-performance RL infra👇 🛠️ What’s new in v0.1.0? • High-performance inference for https://x.com/ZhihuFrontier/status/1962751555591086226
Announcing GLM Coding Plan for Claude Code! After seeing the amazing adoption of GLM-4.5 over the past month, we’re making it more accessible. Get started: https://x.com/Zai_org/status/1962522757536887205
Have been tinkering with GLM 4.5 for about an hour. It is about 3x faster than Claude Code + Opus 4.1 and 5x faster than GPT-5-high, but still feels just as good as closed-source models. I am definitely more productive than with other models due to GLM-4.5’s speed.”” / X https://x.com/Tim_Dettmers/status/1962603940291260533
NVIDIA continues to lead on open-sourcing pretraining data — Nemotron-CC-v2 has dropped! https://x.com/ZeyuanAllenZhu/status/1962119316427706828
Autonomous News Agent A LangGraph-powered AI agent that autonomously curates news briefings, extracts facts, and summarizes content with integrated human feedback and dynamic tool selection. https://x.com/LangChainAI/status/1962213801249710230
🚨 We’ve just published a recipe to train a frontier-level deep research agent using RL. With just 30 hours on an H200, any developer can now beat Sonnet-4 on DeepResearch Bench using open-source tools. (Thread 🧵) https://x.com/corbtt/status/1962954306078048297
I trained a Qwen Image Edit LoRA for inpainting. Just paint the part you want inpainted green (0, 255, 0), and it will inpaint only that section. https://x.com/ostrisai/status/1963269597865599425
~4 months ago, we introduced OpenVision — a fully open, cost-effective family of vision encoders that rival OpenAI’s CLIP and Google’s SigLIP. Today, we’re back with a major update: OpenVision 2 https://x.com/cihangxie/status/1963297223753494832
Boston Dynamics’ Spot costs $75,000. This one? $135. LeCabot is an open-source mod for your SO-100 and Unitree Go2 that turns them into a fully capable mobile manipulator… for a fraction of the price. You control both the robot dog and arm simultaneously using a Meta Quest https://x.com/IlirAliu_/status/1960971652465840406
Hugging Face team just released an agent dataset. Training on it drastically improves the ability to execute code and analyze data. 📈 They use E2B sandboxes to simulate a real code execution environment. Check it out:”” / X https://x.com/e2b/status/1962945170736849262
Jina Code Embeddings: SOTA Code Retrieval at 0.5B and 1.5B https://x.com/JinaAI_/status/1963637141675843791
vibe coding app: https://x.com/_akhaliq/status/1962920607684730977
`langchain` 1.0, now in alpha, ships with improved standardization for reasoning, citations, tool calls, multimodal data, and other content across LLM providers. No more juggling APIs— just one consistent interface.
https://x.com/LangChainAI/status/1963285794954907750
AI Rails App Builder A natural language-powered system that builds and modifies Rails applications in real-time. Using LangGraph, it handles file operations and Rails commands through an intelligent agent with live previews. https://x.com/LangChainAI/status/1962183602185314525
Issue Triager Agent A GitHub issue management solution that uses LangGraph to automatically handle stale issues with human oversight through Agent Inbox. Built with LangSmith integration for comprehensive monitoring and control. https://x.com/LangChainAI/status/1962198699653861755
In less than a day, @StepFun_ai dropped Step-Audio 2 Mini – 8B speech to speech, beats GPT-4o-Audio, Apache 2.0 licensed 🔥 > Trained on 8M+ hours, supports 50K+ voices, benchmarks for expressive/grounded speech 🤯 > Expressive and emotionally aware generation > Retrieves and https://x.com/reach_vb/status/1961414067668558319
🚨 Top 10 Leaderboard Disrupted ⚡ DeepSeek V3.1 and DeepSeek v3.1 thinking by @deepseek_ai have landed in the Arena, both ranked at #8. A few highlights: 💠 DeepSeek V3.1 is in the Top 3 for Math, Creative Writing & Longer Query 💠 DeepSeek V3.1 thinking comes in #3 for https://x.com/lmarena_ai/status/1961474406817173602
Meets or beats sonnet 4 across the board”” / X https://x.com/andrew_n_carr/status/1963805265356075336
Anyway, here’s a simple fix for the issue. It deviates from the original benchmark, but at least now my silly baseline isn’t better than Qwen3 🤠 For the curious, @akseljoonas and I found this by manually reading the agent trajectories – yet another example where LOOKING AT THE”” / X https://x.com/_lewtun/status/1962884902363255165
airline customer service AI demo with agent handoffs https://x.com/tom_doerr/status/1962972766174339271
✍️ When it comes to creative writing optimization, you can’t ignore Zhi-Create-Qwen3-32B, a fine-tuned variant of Qwen3-32B. On WritingBench, it scores 82.08, outperforming the base model (78.97), showing notable gains across 6 domains (Fig.1) What powers its performance boost? https://x.com/ZhihuFrontier/status/1963441300692402659
🥳Seed-OSS INT4 model: https://x.com/HaihaoShen/status/1962652473862299667
meituan-longcat/LongCat-Flash-Chat · Hugging Face https://huggingface.co/meituan-longcat/LongCat-Flash-Chat
xAI may be one of the single biggest contributors to open-source inference just by serving everything with SGLang https://x.com/casper_hansen_/status/1961752869478031810
You can now use flash-attention 3 through 🤗 `kernels`, skipping its long build times entirely 🔥 https://x.com/RisingSayak/status/1963225732668182856
We just added OpenAI Codex CLI formal support in Hugging Face MCP Server – go play with it now!! 🔥 https://x.com/reach_vb/status/1963599978909008321
ZeroGPU on 🤗 HF Spaces enables anyone to build delightful ML demos, benefitting from powerful compute. But, due to its serverless nature, it is hard to optimize these demos. That CHANGES today 🪖 Use AoT compilation to melt our ZeroGPU servers 🔥 Details ⬇️ https://x.com/RisingSayak/status/1962844485118996545
ZeroGPU on Hugging Face enables anyone to build and deploy AI apps dynamically allocates and releases NVIDIA H200 GPUs as needed But, due to its serverless nature, it is hard to optimize these apps Now use AoT compilation to melt ZeroGPU servers on Hugging Face for vibe https://x.com/_akhaliq/status/1962920105186115621
🚀 LongCat-Flash-Chat Launches! ▫️ 560B Total Params | 18.6B-31.3B Dynamic Activation ▫️ Trained on 20T Tokens | 100+ tokens/sec Inference ▫️ High Performance: TerminalBench 39.5 | τ²-Bench 67.7 🔗 Model: https://x.com/Meituan_LongCat/status/1961827385667690965
Tutorial: Train a Qwen Image Edit LoRA with AI Toolkit
https://x.com/ostrisai/status/1961884211956400358
Try Hunyuan-MT-7B and Hunyuan-MT-Chimera via @huggingface and @gradio! This model is specialized for translate 🤗”” / X https://x.com/SOSOHAJALAB/status/1962790133054480600
🌐 Our first open model has landed on the Search leaderboard! Diffbot-small-xl by @diffbot debuts at #9 (Apache 2.0) We look forward to more models with search capabilities contributing to ecosystem progress! https://x.com/lmarena_ai/status/1961526740754616545
For llama.vim the recommended setup now is Qwen 3 Coder 30B A3B Instruct: brew install llama.cpp llama-server –fim-qwen-30b-default Amazingly, on Macs the 30B MoE model performs better than the old Qwen 2.5 Coder 7B so if you have the necessary RAM it’s better to switch to https://x.com/ggerganov/status/1961471397428883882
GPT-4o level intelligence running on your phone! MiniCPM-V 4.5 delivers enterprise-grade AI performance in just 8B parameters, outperforming models like GPT-4o, Gemini-2.0 Pro on vision and language tasks. – 30+ language support – Runs smoothly on iPhone/iPad 100% open-source! https://x.com/akshay_pachaar/status/1962132670126981459
❤️ Thanks to @mervenoyann & @huggingface , MiniCPM-V 4.5 is officially live on Hugging Face Spaces. Come check it out! https://x.com/OpenBMB/status/1963623940028563910
Le Chat. Custom MCP connectors. Memories. | Mistral AI https://mistral.ai/news/le-chat-mcp-connectors-memories
From payments data and refunds to invoices and subscriptions, @MistralAI’s users can now handle it all inside Le Chat with @stripe’s MCP. Here’s how it works: https://x.com/emilygsands/status/1962884010289590583
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation, presented in the paper Check out the model here: https://x.com/reach_vb/status/1961414145938485477
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation, presented in the paper Play with the demo here: https://x.com/reach_vb/status/1961471503267979699
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation, presented in the paper try it here: https://x.com/_akhaliq/status/1962644559868883310
FineVision is out! A massive open-source dataset by @huggingface for training Vision-Language Models: – 17.3M images – 24.3M samples – 88.9M turns – 9.5B answer tokens This is the inaugural article using our new scientific publishing template! https://x.com/thibaudfrere/status/1963627540544647177
Fuck it. Today, we open source FineVision: the finest curation of datasets for VLMs, over 200 sources! > 20% improvement across 10 benchmarks > 17M unique images > 10B answer tokens > New capabilities: GUI navigation, pointing, counting FineVision 10x’s open-source VLMs. https://x.com/andimarafioti/status/1963610118165000479
Today, we are releasing FineVision, a huge open-source dataset for training state-of-the-art Vision-Language Models: > 17.3M images > 24.3M samples > 88.9M turns > 9.5B answer tokens Here are my favourite findings: https://x.com/lusxvr/status/1963609337546293448
Introducing Command A Translate, our state-of-the-art model designed for high-quality translation tasks. https://x.com/cohere/status/1961081779789447519
vLLM now supports Kwai Keye-VL-1.5!
With sharper video 📹 & image 🖼️ comprehension, stronger reasoning, and an extended 128K context length, this model unlocks richer conversations and more complex tasks than ever before. https://x.com/vllm_project/status/1962509793345859666
we present R-4B, a multimodal large language model designed for general-purpose auto-thinking, autonomously switching between step-by-step thinking and direct response generation based on task complexity. https://x.com/mervenoyann/status/1962917670786937135
best small vision LM with reasoning has dropped on @huggingface 🔥 Tencent dropped R-4B, small vision LM that claims sota with Apache 2.0 license 💗 the model enables different thinking options and transformers support through custom code! https://x.com/mervenoyann/status/1962917635932229797
Apertus: a fully open, transparent, multilingual language model | ETH Zurich https://ethz.ch/en/news-and-events/eth-news/news/2025/09/press-release-apertus-a-fully-open-transparent-multilingual-language-model.html
We said “coming weeks” and proceeded to add three over the holiday weekend.😅 Now live on @OpenRouterAI via our W&B Inference service: • @deepseek_ai V3.1 • @OpenAI’s gpt-oss-120b & gpt-oss-20b https://x.com/weights_biases/status/1962943063711744115
Unfortunate reality: most open-source LLM servers (e.g. Together) don’t offer cache-hit discounts, while closed providers like OpenAI do. DeepSeek does discount, but most third-party servers don’t.
https://x.com/arankomatsuzaki/status/1963294646957957263
MXFP4 support for openai/gpt-oss landed in MLX! ⌘ + Shift + R to update your runtime. 🚀 https://x.com/lmstudio/status/1961508941852283016
Big labs spent millions developing deep research. Now you can do it for $350 in a single day! Open research wins.”” / X https://x.com/corbtt/status/1962954848913256832
Kimi K2-Instruct-0905 is the latest, most capable version of Kimi K2. https://x.com/bigeagle_xd/status/1963802450374369722
Model providers should be community driven #249605
https://x.com/ggerganov/status/1963255951659508117
Nice to see open models (+ OpenHands) showing really strong performance here. Arguably for pure agentic coding tasks open models have almost caught up with closed ones. For more diverse tasks there’s still a little ways to go.”” / X https://x.com/gneubig/status/1963045532022010231
Hermes 4: Nous Research Open-Weight Reasoning Family Models – 70B / 405B (Llama-3.1 bases, released) – 14B (Qwen3 base, research baseline) Hermes 4 70B & 405B – Base: Llama-3.1-70B / 405B – Training: TorchTitan (modified), Axolotl, 192× B200s, FSDP and TP – Dataset: 56B tokens https://x.com/gm8xx8/status/1962943078702186627
🧵 We acquired OpenPipe! OpenPipe’s ART framework makes it easy to beat foundation models with smaller open models on real problems, using reinforcement learning. https://x.com/shawnup/status/1963335514377130397
Instructions/reasoning are now everywhere in retrieval – we want embeddings to do it all! 🚀 But… is it even possible? 🤔 Turns out, it’s not possible for single-vector models 😱 theoretically and empirically! To make it obvious we OSS a simple eval SoTA models flop on! 🧵 https://x.com/orionweller/status/1961436569409331579
🚀 Qwen-Max has successfully scaled to 1T parameters, and we’re still pushing further. Hopefully this giant will bring some surprises, see you next week!”” / X https://x.com/huybery/status/1963998518667776250
Big news: Introducing Qwen3-Max-Preview (Instruct) — our biggest model yet, with over 1 trillion parameters! 🚀 Now available via Qwen Chat & Alibaba Cloud API. Benchmarks show it beats our previous best, Qwen3-235B-A22B-2507. Internal tests + early user feedback confirm: https://x.com/Alibaba_Qwen/status/1963991502440562976
Qwen3 Max is truly, solidly, a US-grade modern frontier model. They ask $15/MT for what they serve because that is easily its weight class.”” / X https://x.com/teortaxesTex/status/1963994291765649716
Qwen3-Max-Preview is now live on OpenRouter! 🚀”” / X https://x.com/Alibaba_Qwen/status/1964004112149754091
Ready to meet the biggest, brainiest guy in the Qwen3 family?”” / X https://x.com/Alibaba_Qwen/status/1963586344355053865
Really liking the chainlit open source lib for building a quick but nice chat interface for any LLM. Here are some quick single and multi-turn examples for my Qwen3 from-scratch models: https://x.com/rasbt/status/1962695306757185647
Traditional code embedding models face a fundamental bottleneck: there simply aren’t enough high-quality comment-code pairs for supervised training. By starting with Qwen2.5-Coder pre-trained on 5.5 trillion tokens spanning 92+ programming languages, we inherit deep semantic https://x.com/JinaAI_/status/1963637139037720995
Here’s a fun fact about TAU Bench: if you train an SFT baseline which has zero tool-calling capabilities, you can beat Qwen3-4B-Instruct by a large margin on the Airline domain 🙃 Why? Because on this domain, TAU Bench only evaluates the model’s ability to: – communicate with https://x.com/_lewtun/status/1962884893718761634
Glad to see Qwen3-Coder performing well on the GSO leaderboard!”” / X https://x.com/Alibaba_Qwen/status/1963049864474120475
MiniCPM-V 4.5 achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro, and strong open-source models like Qwen2.5-VL 72B powered https://x.com/_akhaliq/status/1963587749400727980
🚨 Attention: Draw Things now officially supports Qwen-Image-Edit.
https://x.com/drawthingsapp/status/1961977481860419771
Huge thanks to the community for making Qwen Image Edit’s inpainting magic happen!🙌
https://x.com/Alibaba_Qwen/status/1963048659676979559
Goated FAIR team just found how coding agents sometimes “”cheat”” on SWE-Bench Verified. It’s really simple. For example, Qwen3 literally greps all commit logs for the issue number of the issue it needs to fix. lol, clever model. “”cheat”” cuz it’s more like env hacking. https://x.com/giffmana/status/1963327672827687316
💥 A 450M model just beat bigger VLAs on real robot tasks, and it’s 100% open source [📍 bookmark for later] Came across SmolVLA, a new vision-language-action model for robotics that’s compact, fast, and trained entirely on open community datasets from LeRobot via Hugging Face. https://x.com/IlirAliu_/status/1961486533535412392
All of the details warrant a blog post for the community. So, we authored one 🤗 Check out all the details in this post: https://x.com/RisingSayak/status/1962844506094723429




