Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Black and white cinematic photograph of multiple distinct cloud types (puffy cumulus, wispy cirrus, flat stratus) converging and merging into a single unified cloud formation at center frame, high contrast film grain, bold sans-serif title card reading MULTIMODALITY at bottom third, dramatic skyward perspective, rich grayscale gradients showing textural differences between cloud types as they blend together.

GLM-4.6V (from @Zai_org) is the real deal. It also sounds like Sonnet. It punches pretty close to Sonnet 4 on coding tasks & visual understanding. This is the first OSS vision model that can really critique designs at a useful enough level. It’s only been a few days since we https://x.com/hrishioa/status/1998636234806341873

GLM-4.6V Series is here🚀 – GLM-4.6V (106B): flagship vision-language model with 128K context – GLM-4.6V-Flash (9B): ultra-fast, lightweight version for local and low-latency workloads First-ever native Function Calling in the GLM vision model family Weights: https://x.com/Zai_org/status/1998003287216517345

We believe every phone can become an AI phone. Here’s what we’ve built, and with the power of open source, you can build even better. https://x.com/Zai_org/status/1999118116086051034

Zhipu AI just released GLM-4.6V on Hugging Face This new multimodal model achieves SOTA visual understanding, features native function calling for agents, and handles 128k context for documents. Perception to action! https://x.com/HuggingPapers/status/1998373902595301589

GLM-4.6V can read my horrendous hand writing and explain the math correctly Really loving this model, how well it does tool calling, how many languages it knows and its visual accuracy. https://x.com/0xSero/status/1998328482930073887

GLM-4.6V is out. This is new vision language model from @Zai_org – it’s a MOE with 12B active parameters and 106B total. – there’s a leaner variant with 9B – context lengths are 128k – it has native multimodal function calling Should be perfect for agentic tasks like browser https://x.com/ben_burtenshaw/status/1998019922664865881

GLM-4.6V just dropped on Hugging Face https://x.com/_akhaliq/status/1998052965597241647

GLM-4.6V: Open Source Multimodal Models with Native Tool Use https://z.ai/blog/glm-4.6v

Gemini 3 Pro: the frontier of vision AI https://blog.google/innovation-and-ai/technology/developers-tools/gemini-3-pro-vision/

We’ve developed the FACTS Benchmark Suite with @GoogleResearch. 📊 It’s the industry’s first comprehensive test evaluating LLM factuality across four dimensions: internal model knowledge, web search, grounding, and multimodal inputs. https://x.com/GoogleDeepMind/status/1998831084277313539

Native Multimodal Function Calling is finally here. 👁️⚡️ GLM-4.6V (106B) and Flash (9B) from @Zai_org just landed on @ZenMuxAI . This is a massive leap for Agentic workflows: – No OCR Detour: Pass images directly as function args. – 128k Context: Handles massive docs & long https://x.com/ZenMuxAI/status/1998018534736343495

🎉Congrats to the @Zai_org team on the launch of GLM-4.6V and GLM-4.6V-Flash — with day-0 serving support in vLLM Recipes for teams who want to run them on their own GPUs. GLM-4.6V focuses on high-quality multimodal reasoning with long context and native tool/function calling, https://x.com/vllm_project/status/1998019338033680574

GLM-4.6V has day zero support on MLX-VLM 🚀 Quants uploading to the hub. PS: Install from source because they changed vision model type https://x.com/Prince_Canuma/status/1998024143212851571

I tested multimodal tool calling with GLM-4.6V on HuggingChat, it works very well! https://x.com/mervenoyann/status/1998405366313345295

interesting how some benchmark doesn’t seems to get huge boost between glm-4.6V and the flash version (which is ONLY 9B dense compare to 106B A12B MoE)”” / X https://x.com/eliebakouch/status/1998015034979389563

Gemini 2.5 Text-to-Speech model updates https://blog.google/innovation-and-ai/technology/developers-tools/gemini-2-5-text-to-speech/

Gemini 3 Pro continues to be SOTA on most multi-modal benchmarks and use cases! https://x.com/OfficialLoganK/status/1997003665433838026

We just updated our suite of Gemini TTS models 🗣️, they now come with: – Richer tone versatility and stricter adherence to style prompts – Smarter context-aware speed adjustments and better instruction following – Consistent character voices in multi-speaker scenarios”” / X https://x.com/OfficialLoganK/status/1998884687457173580

If you thought AI couldn’t see well before, look again 👀✨ The future is looking very sharp.”” / X https://x.com/songyoupeng/status/1997072574778601516

Yep, the point we wanted to make here is that GPT-5.2’s vision is better, not pe… | Hacker News https://news.ycombinator.com/item?id=46235267

𝐈𝐦𝐩𝐥𝐞𝐦𝐞𝐧𝐭𝐢𝐧𝐠 Qdrant’s 𝐒𝐞𝐦𝐚𝐧𝐭𝐢𝐜 𝐒𝐞𝐚𝐫𝐜𝐡 𝐚𝐭 𝐒𝐜𝐚𝐥𝐞: 𝐅𝐫𝐨𝐦 100𝐊 𝐈𝐦𝐚𝐠𝐞𝐬 𝐭𝐨 𝐌𝐞𝐚𝐧𝐢𝐧𝐠-𝐀𝐰𝐚𝐫𝐞 𝐑𝐞𝐭𝐫𝐢𝐞𝐯𝐚𝐥 Shared by DV Suresh Dasari Amazing breakdown of how semantic search was implemented for a catalog of 100,000+ carpet https://x.com/qdrant_engine/status/1998302093736583429

GigaTIME: Scaling tumor microenvironment modeling using virtual population generated by multimodal AI – Microsoft Research https://www.microsoft.com/en-us/research/blog/gigatime-scaling-tumor-microenvironment-modeling-using-virtual-population-generated-by-multimodal-ai/

🚀 Qwen3-Omni-Flash just got a massive upgrade (2025-12-01 version) ! What’s improved: 🎙️ Enhanced multi-turn video/audio understanding – conversations flow naturally ✨ Customize your AI’s personality through system prompts (think roleplay scenarios!) 🗣️ Smarter language https://x.com/Alibaba_Qwen/status/1998776328586477672

Bioinspired robot: Fly – Roll – Walk – Crawl [Paper ⬇️] Multi-Modal Mobility Morphobot (M4), a revolutionary robot inspired by nature’s most adaptable creatures. ✅ Capable of multiple forms of movement: flying, rolling, crawling, and more. ✅ Features adaptive appendages that https://x.com/IlirAliu_/status/1997226282120106338

Diffusion and flow models work in robotics because they can model complex action distributions. But what if… the real reason has nothing to do with generative modeling at all? A new study puts this to the test. Key findings that stood out: ✅ Regression matches flow models https://x.com/IlirAliu_/status/1997960545907990573

EngineAI CEO survives a powerful kick from EngineAI’s new T800 humanoid. https://x.com/TheHumanoidHub/status/1997290717908291617

EngineAI CEO Takes a Kick from T800 Robot to Settle CGI Debates | Humanoids Daily https://www.humanoidsdaily.com/news/engineai-ceo-takes-a-kick-from-t800-robot-to-settle-cgi-debates

Founders in robotics and industrial AI, please… • realise your pilot problem is a customer selection problem • realise your sales problem is a value communication problem • realise your churn problem is a deployment system problem • realise your pricing problem is an ROI”” / X https://x.com/IlirAliu_/status/1998031566195339566

Hardware Production 🤝 Robotics But these motors… are CRAZY synced: CNC motor synchronization helps make sure that different motors in a machine work together smoothly so that the machine can cut and shape materials with perfect accuracy. This is really important for making https://x.com/IlirAliu_/status/1997381376342155315

Robot policies fail on the hard parts of manipulation. The moment contact, friction, or force uncertainty shows up, the success rate drops fast. CR DAgger shows a very different path. You take a pre trained policy. You let a human correct it in the real world for a short https://x.com/IlirAliu_/status/1996871611392069708

Releasing jina-VLM: our new 2B vision language model achieves SOTA on multilingual visual question answering and document understanding among open 2B-scale VLMs. https://x.com/JinaAI_/status/1997926488843190481

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading