Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: A charming 80s suburban house at dusk with a traditional Chinese moon gate as the front entrance, decorated with both Halloween jack-o-lanterns and red paper lanterns, autumn leaves scattered on the lawn, trick-or-treaters in costumes approaching with orange plastic buckets, warm golden light spilling from the circular doorway, film photography aesthetic with slight grain and rich fall colors

Many people are confused by Minimax’s recent return to full attention – especially since it was the first large-scale pivot toward hybrid linear attention – and by Kimi’s later adoption of hybrid linear variants (as well as earlier attempts by Qwen3-Next, or Qwen3.5). I actually”” / X https://x.com/SonglinYang4/status/1984021551914926514

ollama run qwen3-vl Ollama’s engine now supports all the Qwen 3 VL models locally. 2B to 235B parameter sizes. The smaller models work exceptionally well for their size. The latest version of Ollama v0.12.7 is needed! Give it a try! 👇👇👇 https://x.com/ollama/status/1983683646864126155

Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer A new paper from Weizmann Institute of Science, getting reconstructions that are not complete nonsense from ONLY 15 min of data. We previously demonstrated SOTA for 1 hr of data with MindEye2. This https://x.com/iScienceLuvr/status/1984195725253804449

PewDiePie in 2025: – built a 10×4090 rig – runs Llama 70B, gpt-oss-120B & Qwen 245B locally via vLLM – built a custom web UI (chat, RAG, search, TTS) – ran protein-folding simulations for charity – created an AI “council”, a swarm of 64 models – now fine-tuning his own model https://x.com/Yuchenj_UW/status/1984309989134254493

A great deep-dive on On-Policy Distillation — an efficient way to post-train smaller LLMs with dense, on-policy feedback. Excited to see Qwen featured in the experiments, showcasing strong math-reasoning gains and continual-learning recovery. Excellent work by @thinkymachines 👏”” / X https://x.com/Alibaba_Qwen/status/1983053298447069275

Always visualize. We’ve caught so many interesting patterns / phenomenon In particular in either multimodal or MoE modeling by inspecting / debugging visually. @kilian_maciej ran this exercise of the Qwen3 MoE series and some very interesting patterns emerged”” / X https://x.com/AkshatS07/status/1982629716495663521

IBM Granite team released Granite 4 Nano models 1B variant outperforms Qwen3-1.7B with fewer params on a mix of tasks from math to coding 👏 https://x.com/mervenoyann/status/1983192115577503974

Qwen 3 Max Thinking has released Should also be up on VB shortly! https://x.com/legit_api/status/1984284268412191216

Qwen3-VL models are now live in LM Studio! 🎉🚀 A powerful collection of vision-language models. Happy Halloween! 🎃👻 https://x.com/lmstudio/status/1984330903880155154

We dive deep into every part of the stack that we didn’t build. Here’s a deep-dive into (potentially undisclosed?) MoE depth wise up-cycling that Qwen3 does. I have a lot of Qwen folks that follow me, would anyone like to clarify :)”” / X https://x.com/ArmenAgha/status/1982613142321746130

Very nice blog post from Thinky (@_kevinlu et al) about on-policy distillation for LLMs — we published this idea back in 2023 and it is *publicly* known to be successfully applied to Gemma 2 & 3, and Qwen3-Thinking (and probably many closed frontier models)! The idea behind”” / X https://x.com/agarwl_/status/1982880080482140372

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading