Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: A 1961 Ferrari 250 GT California Spyder in Rosso Corsa red positioned at an elegant intersection where flowing ribbons of golden light, soft blue sound waveforms, and translucent geometric text patterns converge and blend around the vehicle, cinematic automotive photography with warm studio lighting, polished chrome reflections, clean minimal background with subtle depth of field, premium rendering style
Today, we’re one step closer to AI as an operating system. A computer you can talk to, that can see what you see, and take action – all with your permission, all more intuitive than ever. Vision now GA globally + more on today’s @Windows blog: https://x.com/mustafasuleyman/status/1978808627008847997
ByteDance just released Sa2VA on Hugging Face. This MLLM marries SAM2 with LLaVA for dense grounded understanding of images & videos, offering SOTA performance in segmentation, grounding, and QA. https://x.com/HuggingPapers/status/1978745567258829153
Exclusive: First real samples of Veo 3.1 generated videos https://www.testingcatalog.com/first-real-samples-of-veo-3-1-generated-videos/#google_vignette
Neuralink’s patient-8, Nick Wray, fed himself for the first time since being paralyzed by ALS. He accomplished this by controlling a robotic arm purely with his thoughts, using the Telepathy implant. https://x.com/TheHumanoidHub/status/1976891521363591588
📍Reconstructing 3D scenes from photos shouldn’t destroy the colors. But most methods still fail when lighting or exposure changes: leaving blown-out windows, flat walls, or strange shadows. Neural Exposure Fields (NExF) fixes that. ✅ Learns an exposure value for every 3D https://x.com/IlirAliu_/status/1976703320359051682
Surfer 2 is here. 🏄🏄 Our new Cross-Platform Computer-Use Agent exceeds state-of-the-art on the 4 main benchmarks: WebVoyager, AndroidWorld, WebArena and OSWorld. 🖥️🌐📱 Find out more at https://x.com/hcompany_ai/status/1978935436111229098
The 2025 BEHAVIOR Challenge, launched by @drfeifei’s team at Stanford, is a global competition to train robots in household tasks, testing reasoning, navigation, and manipulation in simulated homes. ⦿ 50 tasks, 1,000 activities (e.g., cooking, cleaning) ⦿ 10,000 expert demos https://x.com/TheHumanoidHub/status/1976355634510737626
💃MPMAvatar: Learning 3D Gaussian Avatars with Accurate and Robust Physics-Based Dynamics”” TL;DR: photorealistic 3D human avatars from multi-view videos with realistic, robust dynamics; tailored Material Point Method for garments + 3DGS for rendering https://x.com/Almorgand/status/1978068062340337881
The ROI for spatial intelligence is very clear in public sector applications”” / X https://x.com/bilawalsidhu/status/1977767505369411818
Google’s Gemma AI model helps discover new potential cancer therapy pathway https://blog.google/technology/ai/google-gemma-ai-cancer-therapy-discovery/
Using AI to identify genetic variants in tumors with DeepSomatic https://research.google/blog/using-ai-to-identify-genetic-variants-in-tumors-with-deepsomatic/
Congrats to the team from Google and Yale on the release of C2S-ScaleGemma weights today! Building on some of our graph learning expertise, I’m impressed by how far we’ve come with advancing both science & math with foundation models, and look forward to more!”” / X https://x.com/mirrokni/status/1978640091292553509
An exciting milestone for AI in science: Our C2S-Scale 27B foundation model, built with @Yale and based on Gemma, generated a novel hypothesis about cancer cellular behavior, which scientists experimentally validated in living cells. With more preclinical and clinical tests,”” / X https://x.com/sundarpichai/status/1978507110477332582
Introducing… Cell2Sentence Scale 27B🤏 Based on Gemma, it’s an open model that generated hypotheses about cancer cellular behavior. In collaboration with Yale, we confirmed the predictions with experimental validation in living cells Super excited about this one 🤯”” / X https://x.com/osanseviero/status/1978508266939281804
Introducing our 2025 US GP Livery 🤩✨ #McLaren | @geminiapp https://x.com/McLarenF1/status/1978566601495433588
Ming-UniVision: Joint Image Understanding and Generation via a Unified Continuous Tokenizer | INCLUSION AI https://inclusionai.github.io/blog/mingtok/
make nanochat multimodal for < $10! this evening, i trained nanochatVL: via a projection model (llava-style) between SigLIP ViT and @karpathy nanochat to extend its understanding to images it’s a huge wip rn, but have a few promising results! now i can finally sleep https://x.com/_rajanagarwal/status/1978376536152785368
Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields”” TL;DR: geometry-grounded features add structure yet hurt pose estimation. Visual-only semantics (e.g., CLIP/DINO) remain more versatile. Framework: SPINE for inversion without initial guess. https://x.com/Almorgand/status/1976638982818873555
🚀 PaddleOCR-VL is here! Introducing PaddleOCR-VL (0.9B) — the ultra-compact Vision-Language model that reaches SOTA accuracy across text, tables, formulas, charts & handwriting. Breaking the limits of document parsing!🌍 Powered by: • NaViT dynamic vision encoder • ERNIE https://x.com/PaddlePaddle/status/1978809999263781290
icymi there’s a two new Nanonets OCR models, Nanonets-OCR2-3B and Nanonets-OCR2-1.5B-exp 🙌🏻 this model can handle forms (checkboxes), recognize watermarks, describes images, charts in docs and more! it even handles flowcharts 🤯 it’s multilingual and Apache-2.0 licensed 🙏🏻 https://x.com/mervenoyann/status/1978837720353927415
chatgpt for plumbers:”” / X https://x.com/gdb/status/1977105278077485375
1/ That came as a surprise to me: “AI’s integration and impact varies by industry, according to the survey; Plumbers were the most likely to say AI has helped their business grow; cleaners were the “biggest adopters of AI”; while electricians had “the highest satisfaction rates” https://t.co/NHu9ngRAa2″ / X
https://x.com/kimmonismus/status/1976932982746497380
Qwen3-VL has already become one of the most popular multimodal models supported by vLLM – Try it out!”” / X https://x.com/rogerw0108/status/1978158856611024913
Excited to announce the launch of Qwen3-VL-Flash on Alibaba Cloud Model Studio! 🚀 A powerful new vision-language model that combines reasoning and non-reasoning modes, outperforming open-source Qwen3-VL-30B-A3B and Qwen2.5-72B with faster responses, stronger capabilities, and https://x.com/Alibaba_Qwen/status/1978841775411503304
PhysHSI enables humanoid robots to perform diverse, natural humanoid-scene interactions and loco-manipulation. Trained in simulation with RL and adversarial motion priors (AMP) on retargeted, object-annotated MoCap data, it uses LiDAR-camera fusion for robust sim-to-real. https://x.com/TheHumanoidHub/status/1978154439597769160
Training robots to move like humans just got a lot easier. [📍GitHub] Most motion imitation codebases break somewhere: Unstable, incomplete, or missing key details. That’s why the team behind DeepMimic, AMP, and ASE just released MimicKit, a new open framework that finally https://x.com/IlirAliu_/status/1976550447109320917
Most robot policies fail because… they don’t know where to look or what to focus on. PEEK fixes that, using vision-language models (VLMs) to guide any visuomotor policy. ✅ Adds VLM-generated overlays showing “where” and “what” directly on training images ✅ Works with ACT, https://x.com/IlirAliu_/status/1977049479980228649
D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI https://worv-ai.github.io/d2e/




