Image created with Gemini. Image prompt: A flat overhead view of a turquoise swimming pool assembled from a grid of slightly offset rectangular photo panels, each panel drawn with a different style of hand-drawn white ripple squiggles, surrounded by terracotta pink and hot yellow towels and small open toolboxes on lawn green tiles, flat saturated acrylic color with hard edges and no shadows under bright noon light.

Meet Kimi Work – a local AI agent on your desktop that does the work for you. 🔹Native agent swarm: Up to 300 AI agents running in parallel on your local machine. 🔹Browser use: Paired with WebBridge extension, your agent will navigate websites in your browser: search, scroll,
https://x.com/Kimi_Moonshot/status/2063990409903112344

DiffusionGemma is out 🔥 it’s compute-bound so 4x faster compared to other Gemma-4 models (1k tok/s on H100) 💨 also great on coding, generate and iterate on any code from 3D generation to front-end ⤵️
https://x.com/mervenoyann/status/2064753402064601181

Awesome to see this innovation in text diffusion. DiffusionGemma is lightning fast, 4x faster than other Gemma 4 models! Congrats to @bodonoghue85 and the team who worked so hard on this – excited to see what people build with it!
https://x.com/demishassabis/status/2064873362799600042

DiffusionGemma is an open, experimental model that brings our text diffusion research to Gemma 4. It’s a racehorse 🏇achieving up to 4x faster inference by generating entire blocks of text simultaneously vs predicting token-by-token (word-by-word) output!
https://x.com/sundarpichai/status/2064744343743922189

Introducing DiffusionGemma
https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/

Five labs, five minds: building a multi-model finance drama on small models
https://huggingface.co/blog/build-small-hackathon/thousand-token-wood-sim-v2

🚀Day-0 support for @cohere’s North Mini Code on vLLM! Ready to serve with the latest vLLM stable release, this open-source coding model is built for agentic workflows: 🧩 30B total / 3B active — Mixture-of-Experts 📏 256K context, 64K max generation 🛠️ Reasoning, tool
https://x.com/vllm_project/status/2064416312605237434

I exclusively build Hermes Agent with Hermes Agent as well!
https://x.com/Teknium/status/2062822586954997909

Introducing Cohere’s first open-source coding model: North Mini Code Small & efficient, designed for agentic performance and built for community input.
https://x.com/cohere/status/2064378058329526556

Introducing Write Gate in Hermes Agent. Now you have the capability to be able to approve/deny memory updates, skill updates, and skill creation with the same familiar mechanisms as approving dangerous commands. If you are using a small model that doesn’t always recognize what
https://x.com/Teknium/status/2064831491130130879

Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup​ 🔹Drag in videos as coding context: reference-to-LUT, long-video-to-short, screen-recording-to-code, and more​ 🔹Plugins for stocks, financial reports, academic
https://x.com/KimiDevs/status/2063981516708024369

Last time around Apple released a lot of information about how their AI version of Siri worked between local and cloud models, not so much this time It is nice to have a Gemma-like model on device, but it is extremely limited unless it can call a smarter cloud model when needed.
https://x.com/emollick/status/2064052841367392536

North Mini Code: Agentic Coding Model for Developers | Cohere
https://cohere.com/blog/north-mini-code

Now your Hermes Agent has an easy GUI based way to build and design your agent profiles to curate the perfect setup for every task and need! `hermes update` and load into your Dashboard, and start building custom agents today!
https://x.com/Teknium/status/2064764570519146935

The Hermes Agent Desktop App can now access files from your remote instance machine if and when you are connecting to one! Read only for now, more to come.
https://x.com/Teknium/status/2065112576552526168

The NVIDIA Nemotron Coalition continues to grow. We’re excited to welcome new members: @hcompany_ai, @NousResearch, and @PrimeIntellect. And a big thank you to our existing members: @bfl_ai, @cursor_ai, @LangChain, @MistralAI, NAVER Cloud, @perplexity_ai, @ReflectionAI_,
https://x.com/NVIDIAAI/status/2062961026409333232

Use Ollama with Hermes Desktop by @NousResearch. Hermes Desktop brings the same agent (its multi-agent engine, self-improving skills, and messaging integrations) into a desktop app on macOS, Windows, and Linux. Run it on Ollama using local or cloud with one command: ollama
https://x.com/ollama/status/2064441778590339402

We are unifying profile management in Hermes Agent starting today. Now the dashboard allows switching management to any of your agent profiles on the machine. No more running multiple dashboards to manage each profile! Also soon, you will only need one gateway to access all of
https://x.com/Teknium/status/2065060810729414695

We open-sourced a feisty small agentic coding model. – 30B total, 3B active – 256K total context – Compatible with @opencode – Apache 2.0. Weights on @huggingface
https://x.com/JayAlammar/status/2064385607455908254

We’re running the Fast Gemma Challenge: make gemma-4-E4B go brrr on a single A10G, without wrecking quality ⚡️! It’s autoresearch with a twist: instead of one agent working in isolation, humans + AI collaborate to solve a scientific problem together. Good luck beating my
https://x.com/_lewtun/status/2064386398090576236

Super excited to announce that @arcee_ai is the first major American AI lab to replace AWS S3 with Hugging Face for ALL their models and datasets, public AND private 🔥🔥🔥 Multi-million $ partnership to support American open-source AI, let’s go!
https://x.com/ClementDelangue/status/2064323874049679643

Kimi desktop version is here! Available for macOS (Apple Silicon) and Windows. Try it 👉
https://x.com/crystalsssup/status/2063992904209842215

We are releasing our first quantized checkpoints for the Qwen3.5 series of models, co-designed jointly with our inference engine to achieve maximum possible performance on Apple hardware Starting from 0.8B, 2B and 4B models
https://x.com/norpadon/status/2064040631479976240

Can coding agents stay coherent over a 1 billion token budget? Can they build Slack from scratch? Rewrite a JAX codebase in PyTorch? Build a C compiler in Rust? Enter SWE-Marathon: a benchmark for autonomous long-horizon software work.
https://x.com/rishi_desai2/status/2062930906818769356

Also, a lot depends on Chinese labs continuing to ship open weights models. If they stop, the frontier falls further and further behind to those who want to use local/fine-tuned models. I think this is possible because open weights may not be a good business model as costs rise.
https://x.com/emollick/status/2062746121055695077

Moonshot AI eyes US$30 billion valuation as China’s AI race intensifies | South China Morning Post
https://www.scmp.com/tech/article/3356348/moonshot-ai-eyes-us30-billion-valuation-chinas-ai-race-intensifies

The core problem with open weights is that the business model of frontier open weights AI does not look like open source, as there are very few cases where you can make money from closed ancillary services, and they are very expensive to make relative to any potential revenue.
https://x.com/emollick/status/2064700407838818694

We believe openness drives innovation. Both fp8 and nf4 checkpoints are in our repo, with the nf4 variant fitting on a single 24 GB GPU. Huggingface:
https://t.co/SNEh1yZsq8 Github:
https://t.co/mYuEfS2L9n Blog:
https://x.com/ideogram_ai/status/2062956472489922584

Hermes v 0.16.0 is out now! This release includes all the updates you’ve heard about this week and more! – The Desktop GUI App – The Overhaul to the Dashboard – Leaner Built-In Skillset – New Security Layers for Remote Dashboard & GUI Access – and much more!
https://x.com/Teknium/status/2063075771317686606

New security features for remote dashboard/gui connections include simple auth, oAuth powered by Nous, and roll your own oAuth!
https://x.com/Teknium/status/2063078732768928234

🧠 Gemma 4 QAT checkpoints are out, and vLLM is Google’s recommended way to serve them! Open-source inference is at its best when one engine spans research and production — glad vLLM’s is the recommendation for Gemma 4 QAT. Get started 👇
https://x.com/vllm_project/status/2062938949560283216

Building super fast experiences with Gemma just got easier. Gemma 4 MTP is now officially merged into llama.cpp. Developers can now pair MTP with Gemma 4 QAT for a fast, lightweight setup.
https://x.com/googlegemma/status/2064030477628182814

bullish on Gemini
https://x.com/OfficialLoganK/status/2063819854348697681

Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4-2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide:
https://x.com/UnslothAI/status/2065107734916432189

Gemma 4 Quantization-Aware Training (QAT) weights are now available on Ollama! They reduce memory requirements while maintaining model quality. E2B: ollama run gemma4:e2b-it-qat E4B: ollama run gemma4:e4b-it-qat 12B: ollama run gemma4:12b-it-qat 26B: ollama run
https://x.com/ollama/status/2062965815864066079

Gemma 4 with quantization-aware training
https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/

Gemma goes diffusion! DiffusionGemma with up to 1000+ tokens per second! 🌬️ – Built on Gemma 4 as a 26B MoE model. – 3.8B parameters during inference. – Generates text in 256-token blocks in parallel. – Fits within 18 GB VRAM limits when quantized. – Apache 2.0
https://x.com/_philschmid/status/2064745464252055647

Gemma-4 QAT just dropped! We found if you naively convert from QAT Q4_0 BF16, you will lose accuracy since the conversion to llama.cpp has a different lattice. Unsloth dynamic GGUFs recovers most of it! 26B-A4B: 85.6% top-1 % from 70.2% (+15.4%) 31B: 96.7% from 87.9% (+8.8%)
https://x.com/danielhanchen/status/2062933017430315481

Google releases DiffusionGemma.✨ The new 26B-A4B diffusion text model runs locally on 18GB RAM. It supports high-speed text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio. GGUF:
https://t.co/ZH0dCJQ59P Guide:
https://x.com/UnslothAI/status/2064743714875220118

Introducing Gemma 4 QAT 🤏 – Quantization aware training to reduce models’ precision while preserving quality – Introducing a new mobile quantization format that reduces memory footprint of E2B to 1GB – Q4 for all your favorite libraries ✨
https://x.com/osanseviero/status/2062933011415392482

Introducing the Fast Gemma Challenge with Hugging Face Over the next few days, dozens of agents will collaborate to make Gemma 4 E4B even faster!
https://x.com/googlegemma/status/2064374874962117084

Let’s kick off the Fast Gemma Challenge!⚡️⚡️⚡️ Agents researching the latest papers, implementing inference engine changes, and collaborating together to make Gemma 4 E4B ultra fast Looking forward to seeing the results!
https://x.com/osanseviero/status/2064375902046245219

Meet DiffusionGemma ⚡ Our latest experimental open model (Apache 2.0) that generates text up to 4x faster. Instead of predicting and typing just one word at a time like most language models, it drafts and refines entire blocks of text simultaneously. Here’s how it works 🧵 ↓
https://x.com/Google/status/2064741293163418032

More Gemma 4! New QAT Gemma 4 checkpoints with similar performance while using ~4x less memory! It comes with a new mobile quantization format that reduces memory footprint of Gemma 4 E2B to just 1GB. Quantization-Aware Training (QAT) simulates low-precision operations during
https://x.com/_philschmid/status/2063990553826439378

We just dropped Gemma 4 Quantization-Aware Training (QAT) checkpoints on Hugging Face! All Gemma 4 model sizes and their drafters are now optimized with QAT to cut memory requirements and maximize on-device performance!
https://x.com/googlegemma/status/2062928831229665566

Congrats to @GoogleDeepMind on DiffusionGemma 🎉 A 26B diffusion language model on the Gemma4 backbone, and the first dLLM natively supported in vLLM. It denoises 256-token blocks in parallel instead of generating one token at a time: 1200+ output tok/s at batch size 1 on a
https://x.com/vllm_project/status/2064753414735900835

this model is the opposite of mythos. Its small, cost effective, apache 2.0, and locally deployable. This is the way LLMs should go. small, open source, transparent and sovereign vs large, expensive, proprietary and hegemonic
https://x.com/nickfrosst/status/2064396337404096809

We made DiffusionGemma run via llama.cpp locally! It works well with Unsloth GGUFs and you can run it in realtime visualization mode or normal chat CLI mode! See our docs
https://t.co/IslbgeCs7Z on how to set it up!
https://x.com/danielhanchen/status/2064760001567306232

⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — ‪@AhmadAwais‬ , CommandCode.ai – YouTube
https://www.youtube.com/watch?v=-rIAVuaRjOg

Sovereign AI for all.
https://x.com/cohere/status/2064414912768618898

this is the biggest wake-up call to protect and nourish open source AI if you don’t build out sovereign and independent models+infra closed labs will patronize you to an insulting degree
https://x.com/rasdani_/status/2064409800641859747

ETH Zurich just open-sourced their entire 2026 robot learning course. Not a MOOC. The actual course. Slides, lecture recordings, coding assignments, GitHub repo. The curriculum goes from imitation learning and RL all the way to Vision-Language-Action models and foundation
https://x.com/IlirAliu_/status/2064617567893721457

DiffusionGemma is our new experimental open model with up to 4x faster output on dedicated GPUs. Instead of predicting word-by-word, it generates entire blocks of text simultaneously. This lets the model self-correct and format complex markdown in real time.
https://x.com/GoogleDeepMind/status/2064741061352636762

DiffusionGemma is so fast that we had to slow down the videos so people could see what was happening
https://x.com/osanseviero/status/2065041448135770436

Meet DiffusionGemma! An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license. Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇
https://x.com/googlegemma/status/2064741002204545467

Ollama now supports Hermes Desktop Run: ‘ollama launch hermes-desktop’
https://x.com/NousResearch/status/2064468385748951415

Profiling in PyTorch (Part 2): From nn.Linear to a Fused MLP
https://huggingface.co/blog/torch-mlp-fusion

We are open-sourcing solutions from this system so that the community can check them and build upon them:
https://t.co/Ed1ldf05lz More details in our blog:
https://x.com/_rockt/status/2065061993271202171

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading