Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A Lutheran stained-glass panel of a wide-open ornate birdcage releasing a radiant flock of diverse songbirds into a cathedral-blue sky, a haloed red-tailed hawk circling protectively above while a small housecat retreats through a doorway below, jewel-tone leaded glass with cobalt, ruby, amber and emerald, backlit like a rose window, with a heavy blackletter title-card banner at the base reading OPEN SOURCE.
Cohere dropped Command A+ 🔥 > 25B/219B MoE vision language model > supports 48 languages with efficient tokenizer > tool-calling/agentic + 128k context window > transformers day-0 support 🤗 free license 💗
https://x.com/mervenoyann/status/2057128432190787643
The kernels project at Hugging Face has been growing! We want it to be the go-to place for kernel devs and kernel users. We’re looking to work w/ folks who’re interested in doing agentic kernel dev, providing real optim value to real models. Reach out if interested 🙂
https://x.com/RisingSayak/status/2055187769266434101
Releasing open-source under the Apache 2.0 license. We want to give developers direct access to enterprise-grade agentic capabilities from experimentation to production. Sovereign AI. For all. Download Command A+:
https://t.co/USXpmpid01 Or learn more:
https://x.com/cohere/status/2057122131410813016
🚀🚀Qwen3.7 Preview lands on Arena ! Here come Qwen3.7-Max-Preview & Qwen3.7-Plus-Preview. Alibaba now #6 lab in Text, #5 in Vision.⚡️⚡️ Can’t wait to release Qwen3.7 series models!Stay tuned! @arena
https://x.com/Alibaba_Qwen/status/2056403591464984753
Alibaba researchers present MIGA A train-free method for infinite-frame video generation with state-of-the-art temporal consistency across thousands of frames featuring a novel two-stage alignment mechanism and dual consistency enhancement.
https://x.com/HuggingPapers/status/2057506246899724355
llama.cpp adds MTP for the Qwen3.6 family This is a significant milestone for the local AI ecosystem. The performance jump with these changes is massive and elevates local inference on commodity hardware further. Special thanks to Aman Gupta for leading this development!
https://x.com/ggerganov/status/2056391115469689330
llama.cpp with MTP support makes local models fast enough to use as daily drivers 🚀 Qwen3.6-27B dense generation (on A10G): From 25 tok/s → 45 tok/s (+78%). Two flags on llama-server: –spec-type draft-mtp –spec-draft-n-max 2
https://x.com/victormustar/status/2056456757786869793
Qwen
https://qwen.ai/blog?id=qwen3.7
Qwen3.6 MTP Unsloth GGUFs now run 1.8x faster, increased from 1.4x just two days ago! This is due to llama.cpp adding –spec-draft-p-min 0.75! Args have also changed from –spec-type mtp to –spec-type draft-mtp Also increase –spec-draft-n-max 2 to 6 We also released
https://x.com/danielhanchen/status/2055274688025378854
Qwen3.7 Preview By @Alibaba_Qwen lands on Arena for Text and Vision. In Text Arena, Qwen3.7 Max Preview ranks #13 overall. Alibaba is now the #6 lab in this arena. – #7 Math – #9 Expert – #9 Software & IT – #10 Coding In Vision Arena: Qwen3.7 Plus Preview ranks #16 overall,
https://x.com/arena/status/2056400044862111757
Thread by @Alibaba_Qwen on Thread Reader App – Thread Reader App
https://threadreaderapp.com/thread/2056403591464984753.html
What political censorship looks like inside an LLM’s weights — a mechanistic-interpretability study of Qwen 3.5
https://vas-blog.pages.dev/qwen-censorship/
Alibaba unveils new AI chip in push for domestic alternatives
https://ca.finance.yahoo.com/news/alibaba-unveils-ai-chip-push-050035417.html
You don’t understand the current AI race if you don’t think about it in terms of compute – and compute clearly distinguishes 3 tiers of companies. Arthur Mensch, Mistral’s CEO, recently had a hearing at the French Assemblée Nationale. He elegantly framed the AI race as a compute
https://x.com/AymericRoucher/status/2057492189626720729
Cohere launches open weights model Command A+ that achieves 37 on the Artificial Analysis Intelligence Index The release of Command A+ places @Cohere in line with Claude 4.5 Haiku on the Intelligence Index, and just above NVIDIA Nemotron 3 Super and Gemini 3.1 Flash-Lite. Key
https://x.com/ArtificialAnlys/status/2057123594162077837
DeepSeek-V4-Flash means LLM steering is interesting again
https://www.seangoedecke.com/steering-vectors/
There is a lot of fear-based marketing in AI right now which serves some purpose. But it’s misleading. Again, what you want to enable is giving access to people so that they can build.”” @ClementDelangue, co-founder & CEO @HuggingFace Watch the full interview on YouTube:
https://x.com/TheTuringPost/status/2056167344712691903
Introducing the Ettin Reranker Family
https://huggingface.co/blog/ettin-reranker
@huggingface Bio released Carbon, an open DNA foundation model family. We tested a simple infra question: “”Can Carbon run on @awscloud Trainium2 with NxD Inference on day one?”” The answer is: Hell yes !!! Carbon-500M, 3B, and 8B all compiled and ran on a single trn2.3xlarge
https://x.com/Shekswess/status/2057468970471448787
13 open-source tools for foundation model deployment ▪️ vLLM ▪️ Ollama ▪️ Hugging Face TGI ▪️ BentoML ▪️ Seldon Core ▪️ Kubeflow ▪️ MLflow ▪️ MLRun ▪️ Metaflow ▪️ TensorFlow Serving ▪️ TorchServe ▪️ SGLang ▪️ llama.cpp Save the list and learn where to use each of them here →
https://x.com/TheTuringPost/status/2056102301811781848
4 levels of Hermes Agent setup: LEVEL 1: main agent You → Hermes Agent this is your main agent and your prototype area, where you test new workflows and refine them. it doubles as your orchestrator until you have something worth breaking out —- LEVEL 2: specialized agents
https://x.com/shannholmberg/status/2056410242330874349
Lighthouse Attention – NOUS RESEARCH
https://nousresearch.com/lighthouse-attention
Run @NousResearch’s Hermes Agent fully locally on DGX Spark. 🚀 Our newest playbook shows you how to get set up via @Ollama step by step. 👇
https://x.com/NVIDIA_AI_PC/status/2055317325444710872
Today we release a study on decoupling the benefits of subword tokenization for language model training, by simulating each suspected benefit one at a time inside a 1.7B byte-level pretraining pipeline. We formulate seven hypotheses for why subword LLMs outperform byte-level
https://x.com/NousResearch/status/2057610978934546805
🎉 Day-0 vLLM support for Command A+! Congrats to @cohere on their most powerful open-source model yet. 🧠 218B MoE / 25B active, Apache 2.0 🌍 Multimodal + 48 languages ⚡ Runs on as little as 2× H100s @ W4A4 Serve it now in vLLM! 🚀 📖
https://x.com/vllm_project/status/2057206049665622070
🎉 Day-0 vLLM support for Intern-S2-Preview! Congrats to the @intern_lm team — an open-source scientific multimodal foundation model, with a first take on material crystal structure generation alongside general capabilities. 📖
https://x.com/vllm_project/status/2055148034124894395
Big inference day with @tuhinone from @baseten at MS&E 435 yesterday (and the Cerebras IPO that AM) 🚀 We discussed compute scarcity amid the rising inference demand, model routing, & open source! Video soon.
https://x.com/apoorv03/status/2055479206545646040
Cohere is on such a great open-source trajectory lately. Beautiful Apache 2.0 model!
https://x.com/ClementDelangue/status/2057180057756467671
I believe on-prem and local AI – based on @huggingface open-source models – will be an important answer to the GPU shortages this year (because they are cheaper, faster, safer than cloud APIs)! Great collaboration between @huggingface & @MichaelDell @Dell to make this a reality
https://x.com/ClementDelangue/status/2056439359784530252
Introducing: Cohere Command A+ We’ve created our most powerful LLM yet, optimized it to run on as little hardware as possible, and released it open-source for all.
https://x.com/cohere/status/2057120818551734589
Making humans responsible for their AI use seems like an incredibly reasonable way to address problems & opportunities in the use of AI for academic research, at least in the short term (autonomous scientific work will require different solutions).
https://x.com/emollick/status/2055016669047624174
Open source Command A+ model This tech can go one of two ways. It can go the way the internet and mobile phones did – in which technological hegemony resulted in a mostly disempowering tech. Or it can empower the people that use it. We are working towards that second one.
https://x.com/nickfrosst/status/2057132425310851104
Our first fully open source Apache 2 model 🙂
https://x.com/aidangomez/status/2057142232860258527
interesting open model by cohere with lots of unusual architecture choices, here is a recap: > parallel transformer, so MoE and attention are computed in parallel. likely doing some kind of MLP/attention disaggregation here? > lots of query heads, query total dim is 4x hidden
https://x.com/eliebakouch/status/2057198733759008989
Searching through unstructured data, like scans of handwritten and typed declassified documents, can be challenging. But with Cohere Compass, it’s possible because it is built to process and retrieve across even the most challenging documents. This includes the Compass Visual
https://x.com/cohere/status/2055343638360752351
ByteDance open source Lance: Unified Multimodal Modeling by Multi-Task Synergy a lightweight native unified multimodal model for image and video understanding, generation, and editing. 3B for video part,3B for image part and 3B for decoder.
https://x.com/bdsqlsz/status/2056353648779907115
Great work at @baseten running vLLM-Omni in production — open-source, production-grade, cost-efficient omni-modal serving 🎙️ Multi-stage audio, streaming multi-modal, real-time TTS — workloads where closed-source APIs have been the default. →
https://x.com/vllm_project/status/2055136943550427242
Meet LongCat-Video-Avatar 1.5🐱–our upgraded, open-source digital human framework. Built for real production, not just short demos. What’s New: 🔹 Upgraded Audio Encoder: Replaces Wav2Vec2 with Whisper-Large, yielding significantly smoother and more natural lip dynamics. 🔹
https://x.com/Meituan_LongCat/status/2057494106889486646
You can now use X Premium subscriptions in Hermes Agent, and Hermes Agent can now search X posts.
https://x.com/xai/status/2055745332919808181
You can now use your @grok subscription inside @NousResearch Hermes Agent.
https://x.com/xai/status/2055375676656783733
You can now use your SuperGrok subscription in Hermes Agent! Enjoy!
https://x.com/Teknium/status/2055373314399650230





Leave a Reply