Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Black and white photograph of dramatic cumulus clouds forming dragon-like shape rising upward against dark sky, faint city skyline at bottom edge, bold sans-serif text ZHIPU AI as title card in lower third, high contrast cinematic composition, film grain texture, contemplative skyward perspective
GLM-4.6V (from @Zai_org) is the real deal. It also sounds like Sonnet. It punches pretty close to Sonnet 4 on coding tasks & visual understanding. This is the first OSS vision model that can really critique designs at a useful enough level. It’s only been a few days since we https://x.com/hrishioa/status/1998636234806341873
GLM-4.6V Series is here🚀 – GLM-4.6V (106B): flagship vision-language model with 128K context – GLM-4.6V-Flash (9B): ultra-fast, lightweight version for local and low-latency workloads First-ever native Function Calling in the GLM vision model family Weights: https://x.com/Zai_org/status/1998003287216517345
We believe every phone can become an AI phone. Here’s what we’ve built, and with the power of open source, you can build even better. https://x.com/Zai_org/status/1999118116086051034
Zhipu AI just released GLM-4.6V on Hugging Face This new multimodal model achieves SOTA visual understanding, features native function calling for agents, and handles 128k context for documents. Perception to action! https://x.com/HuggingPapers/status/1998373902595301589
GLM-4.6V can read my horrendous hand writing and explain the math correctly Really loving this model, how well it does tool calling, how many languages it knows and its visual accuracy. https://x.com/0xSero/status/1998328482930073887
GLM-4.6V is out. This is new vision language model from @Zai_org – it’s a MOE with 12B active parameters and 106B total. – there’s a leaner variant with 9B – context lengths are 128k – it has native multimodal function calling Should be perfect for agentic tasks like browser https://x.com/ben_burtenshaw/status/1998019922664865881
GLM-4.6V just dropped on Hugging Face https://x.com/_akhaliq/status/1998052965597241647
GLM-4.6V: Open Source Multimodal Models with Native Tool Use https://z.ai/blog/glm-4.6v
🎉Congrats to the @Zai_org team on the launch of GLM-4.6V and GLM-4.6V-Flash — with day-0 serving support in vLLM Recipes for teams who want to run them on their own GPUs. GLM-4.6V focuses on high-quality multimodal reasoning with long context and native tool/function calling, https://x.com/vllm_project/status/1998019338033680574
GLM-4.6V has day zero support on MLX-VLM 🚀 Quants uploading to the hub. PS: Install from source because they changed vision model type https://x.com/Prince_Canuma/status/1998024143212851571
I tested multimodal tool calling with GLM-4.6V on HuggingChat, it works very well! https://x.com/mervenoyann/status/1998405366313345295
interesting how some benchmark doesn’t seems to get huge boost between glm-4.6V and the flash version (which is ONLY 9B dense compare to 106B A12B MoE)”” / X https://x.com/eliebakouch/status/1998015034979389563





Leave a Reply