ERNIE 5.1 just dropped. Built on ERNIE 5.0’s pre-training foundation, our latest foundation model upgrades search, reasoning, knowledge Q&A, creative writing, and agentic capabilities, while using only around 6% of the pre-training cost of comparable models. More in the thread
https://x.com/Baidu_Inc/status/2053009538769735774?s=20

[2605.10730] Qwen-Image-2.0 Technical Report
https://arxiv.org/abs/2605.10730

Exciting: local ML is (finally) going mainstream 🔥 – new GGUF uploads on HF nearly doubled in 2 months – smaller models (like gemma 4 and qwen 27b) starting to be really good and run on a lot of hardwaers – people forking and vibing over llama.cpp (MTP / DS4 / turboquant…)
https://x.com/victormustar/status/2053780086596288781

GB 200s change how one does the prefill and decode disaggregation when serving large MoEs like Qwen. We’ve published details of our stack quantifying the throughput benefits compared to serving on Hoppers.
https://x.com/AravSrinivas/status/2054206802133504234

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading