Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A 16:9 flat matte op-art cover in the style of Julio Le Parc, the word IMAGES lettered large and centered with each letter built from concentric ROYGBIV rainbow bands — violet outside stepping inward to red — a single continuous rainbow ribbon extending from the letters to draw a clean rectangular picture-frame border around the composition, on a pale warm off-white background with generous negative space, crisp screen-printed edges, no shadows, no gradients, no glow.

Introducing Gemma 4 12B
https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/

Today we’re introducing Gemma 4 12B — our latest open model that brings advanced agentic reasoning, vision and audio directly to your laptop. It delivers performance nearing our larger Gemma models with a much smaller total memory footprint, while being small enough to run
https://x.com/Google/status/2062203526588088452

Legendary filmmaker Martin Scorsese signed on last year as an adviser to Black Forest Labs, the German AI startup behind FLUX image models. On Tuesday he went public, testing the tool on a single scene during preproduction. His use is narrow: storyboarding only, complementing
https://x.com/TheRundownAI/status/2061834880917357011

MAI-Image-2.5 is here — now #3 on text-to-image and #2 on image-to-image Arena leaderboards, surpassing Nano Banana Pro. Leading image generation. Precise editing. Built for enterprise scale. It delivers strong performance on H100s, enabling deployment on existing
https://x.com/MicrosoftAI/status/2062240400299934143

Building a hill-climbing machine: Launching seven new MAI models | Microsoft AI
https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/

Molmo2 is a CVPR 2026 award candidate paper from @allen_ai Molmo2 is a VLM that supports video pointing, tracking, counting by pointing, and multi image reasoning all in one open model prompt: blue players
https://x.com/skalskip92/status/2062549751246066144

Introducing the Cosmos Coalition A new global initiative with NVIDIA and leading AI labs to build and open-source frontier world models for physical AI. Runway joins as a founding member, working alongside NVIDIA and a set of leading AI labs to build, share and accelerate world
https://x.com/runwayml/status/2061315089869721682

Jensen just launched NVIDIA Cosmos 3. Pitched as the first fully open omnimodel for physical AI: a mixture-of-transformers (reasoning + generation) with native vision reasoning and generation across text, image, video, sound, and action. Tops open-model leaderboards on
https://x.com/TheHumanoidHub/status/2061333253920080345

NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI | NVIDIA Newsroom
https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-3-the-open-frontier-foundation-model-for-physical-ai

New open model Ideogram-4.0-Quality has landed at #8 in the Text-to-Image Arena. This makes the new model by @ideogram_ai the #1 open model in that arena! Scoring 1204, this open model approaches the performance of Nano Banana Pro. Congrats to the @ideogram_ai team on this
https://x.com/arena/status/2062203346996605116

2-bit Gemma 4 12B GGUF, only 4.66 GB on disk, managed to cite 15 sites from a single prompt. Try this locally on >6GB RAM via Unsloth Studio. GitHub:
https://x.com/UnslothAI/status/2062470072179044447

Congrats to the @googlegemma team on the Gemma 4 12B launch 🎉 Day-0 support on vLLM is ready to go. It’s an encoder-free unified multimodal model — text, image, audio, and video all project straight into the LLM’s embedding space, no separate vision or audio towers. 256K
https://x.com/vllm_project/status/2062228047324201166

For the past years my research focus was on unifying models and training paradigms across modalities. Today I’m excited that we’re releasing our latest model aligned with this theme: Gemma 4 12B, a dense encoder-free model which processes raw text, image, and audio inputs! 1/
https://x.com/mtschannen/status/2062236357351579915

Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs. Google’s new model, Gemma 4 12B Unified supports image, audio and 256K context. You can run and train the model via Unsloth Studio. GGUF:
https://t.co/8cL321pVDh Guide:
https://x.com/UnslothAI/status/2062207258810053084

Meet Gemma 4 12B! A unified, encoder-free multimodal model designed to bring high-performance intelligence directly to your laptop, and released under an Apache 2.0 license. Bridging the gap between edge efficiency and advanced reasoning. Here is what’s new with Gemma 4 12B: 👇
https://x.com/googlegemma/status/2062202706882883696

Our new unified architecture allows Gemma 4 12B to process multimodal inputs natively. Here’s how ⬇️ Traditional models rely on separate encoders for images and audio. This adds latency and increases memory usage. So we streamlined this: 👁️ Vision: We took a novel approach to
https://x.com/Google/status/2062203532351090824

Today we’re shipping our biggest MLX-VLM release yet: v0.6.0 …and we are raising 💸 This one’s about turning your Apple devices into real local agent machines. From your desk to your pocket. What’s new: ⚡ Speculative decoding everywhere — Gemma 4 EAGLE3 + DFlash, Qwen
https://x.com/Prince_Canuma/status/2061541992790683726

We released Gemma 4 12B yesterday. Here is a visual guide that explains the full architecture. → How encoders typically connect modalities to LLMs → Why Gemma 4 removed the vision and audio encoders → How a single 12B model can handle text, images, and audio without
https://x.com/_philschmid/status/2062546814075609413

We’re launching Gemma 4 12B: Our unified, encoder-free model that brings powerful multimodal intelligence straight to your laptop 🚀 The model bridges the gap between our mobile E4B model and larger 26B MoE models, packaging frontier-class reasoning and native audio into a
https://x.com/googleaidevs/status/2062204432658386950

PrismML — Introducing 1-bit and Ternary Bonsai Image 4B: Image Generation for Local Devices
https://prismml.com/news/bonsai-image-4b

The Layout Bet – Reve Blog
https://blog.reve.com/posts/the-layout-bet/

Today, we’re launching Reve 2.0, the best 4K image model in the world. We invented a new way to generate and edit any image using precise layouts. For the first time, it’s possible to create images you can touch.
https://x.com/reve/status/2062260665121919101

Ideogram 4 not only revamped their model to the best they built yet, but also they flipped from closed to open weights! With a 9B parameter model, I can’t wait to see what the community will do with it Try it now:
https://x.com/multimodalart/status/2062210597148930139

Ideogram 4.0 is now live on fal! SOTA open weights model for realism, text rendering and artistic generation Typography that reads clean on posters, packaging and brand work Lightning fast with multiple quality options for every workflow
https://x.com/fal/status/2062202673361780873

Ideogram just released their latest and best v4 image model open weights State of the art and open weights go well together 🤗 Model:
https://t.co/DUcL7BBH7D Demo:
https://x.com/huggingface/status/2062206083914158287

Introducing Ideogram 4.0: the best open image model in the world. Think it. Make it. Own it. Download the weights, fine-tune on your own data, and run it on your hardware. Live on every Ideogram plan and the API today.
https://x.com/ideogram_ai/status/2062202208700313872

big congrats to the microsoft AI team on MAI-Thinking-1! this is the kind of thoughtful post-training the field needs more of – focused on what actually matters to users excited to see a new frontier model in the race 😎
https://x.com/echen/status/2061907282607100075

Introducing MAI-Code-1-Flash A new coding model from Microsoft for fast, efficient assistance in everyday workflows Rolling out to @code developers in model picker and Auto now!
https://x.com/pierceboggan/status/2061877165810131297

It is difficult to know how good MAI-Thinking-1 is from the scores alone (like weirdly low GPQA & Terminal Bench 2.0) But Microsoft makes it really hard to try its models upon release (a general issue with many Microsoft AI products), so I dunno. Stats below Meta Spark, though.
https://x.com/emollick/status/2061907785768489127

MAI-Image-2.5 has officially released from @MicrosoftAI landing at #2 in the Image Edit Arena (Single-Image-Edit) with a score of 1401 and advances the Pareto frontier! This puts the model +10 pts over Nano Banana 2, Grok Imagine Image Quality and ChatGPT-Image-Latest-High
https://x.com/arena/status/2061887242579382660

MAI-Image-2.5 ranks #2 in the Image Edit Arena and advances the Pareto frontier. That means: at its price tier, no model scores higher on Arena. Congrats again to @MicrosoftAI on this release!
https://x.com/arena/status/2061894541888962712

Microsoft AI has announced their very own reasoning model, MAI-Thinking-1, they have a detailed tech report too! I really appreciate they’ve reported health evals: HealthBench Professional and MedXpertQA. These are both very solid benchmark tasks that I recommend people use.
https://x.com/iScienceLuvr/status/2061926066453962952

.@NVIDIA’s Cosmos 3 launched today… and guess who had early access? Agile Robots SE! They’ve been running it across their full portfolio: Thor single- and dual-arm, FR3 Duo. Focus? Simulation. Using Cosmos 3 as a neural simulator; a learned world model that generates
https://x.com/IlirAliu_/status/2061512207738012093

1/ NVIDIA just open-sourced Cosmos 3 at GTC Taipei! It’s the first fully open “”omnimodel”” for physical AI – one model that understands the real world, predicts what happens next, and generates the actions a robot should take. Weights, code, datasets. All open. And this is
https://x.com/kimmonismus/status/2061432501223162241

Breaking news: Cosmos 3 is here. They are attempting to do something completely new 🤯 Why is Physical AI much harder than building a chatbot? Understanding the world is not enough, robots need to predict it and act inside it. That’s the idea behind NVIDIA Cosmos 3: →
https://x.com/TheTuringPost/status/2061308942186414136

In case you missed this: NVIDIA shipped a text-to-image open weights model that looks seriously competitive 👀 (as part of its Cosmos 3 release)
https://x.com/victormustar/status/2061354267546427595?s=20

NVIDIA’s Cosmos 3 is what we’ve never seen before ‒ an Omnimodal World Model. It’s closing the loop for physical AI. All stack in one system: world and multimodal understanding, future generation, reasoning and action This is the next step in Jensen Huang’s AI progression:
https://x.com/TheTuringPost/status/2061474876083274238

NVIDIA’s Cosmos 3 lands at #1 among open weights models in both Text to Image and Image to Video on the Artificial Analysis Leaderboards! Cosmos 3 is a family of omnimodal world models for Physical AI from @nvidia, unifying language, image, video, audio and action in a single
https://x.com/ArtificialAnlys/status/2061494719998546206

NVIDIA’s Cosmos 3 lands at #1 among open weights models in both Text to Image and Image to Video on the Artificial Analysis Leaderboards! Cosmos 3 is a family of omnimodal world models for Physical AI from @nvidia, unifying language, image, video, audio and action in a single
https://x.com/ArtificialAnlys/status/2061494719998546206?s=20

[2606.03746] Qwen-Image-Flash: Beyond Objective Design
https://arxiv.org/abs/2606.03746

A month later, I’m still thinking about that day in Augsburg. We co-hosted the KUKA × @VisComp1 Simulation Event… and honestly, it was one of those events where you leave with your head full. The room was packed with customers, partners, and simulation engineers from across
https://x.com/IlirAliu_/status/2061741801405374760

A system that automatically creates thousands of varied humanoid loco-manipulation demonstrations from one single teleoperated example: How? Via interleaved locomotion planning, manipulation planning, and skill adaptation. It features a new nine-task humanoid
https://x.com/IlirAliu_/status/2060637097120215196

AGIBOT has unveiled AGILE (AgiBot Generative Intelligent Locomotion Engine), a perception-control foundation model for whole-body humanoid locomotion. – It fuses visual perception, balance, and motion planning end-to-end, replacing the traditional split between “”seeing”” and
https://x.com/TheHumanoidHub/status/2060411921451680114

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners”” TL;DR: learns temporally consistent pixel-level representations from videos using linear in-context supervision from depth and motion cues
https://x.com/Almorgand/status/2060386938612232547

Helix4D: Complex 4D Mesh Generation”” TL;DR: generates temporally consistent 4D meshes under extreme topology changes like melting, shattering, and transparency using flow matching and temporal attention chains
https://x.com/Almorgand/status/2060031819471348115

Hey! A new vision encoder for robotics is in town 👀🤖 Instead of using models trained on static images (CLIP, SigLIP, DINO), we bake the dynamics-awareness directly into perception. It transfers well everywhere and boosts real-world OOD success by +22.5% Check it out👇
https://x.com/jbhuang0604/status/2061840469966090308

Larus went ham with this one! Love the synced highlighting on the camera path, something I wanted to try myself. Makes me think these could end up as spatial reasoning benchmarks for ai video models, esp in cities with existing 3d data as ground truth.
https://x.com/bilawalsidhu/status/2060373038982459741

Most robot vision systems take weeks to set up… or you do it in 6 minutes. Not 6 days. Not 6 hours. 6 minutes… data recording + model training included. I’ve been working closely with the team at Lentil Robotics, and what they’ve built is genuinely different from anything
https://x.com/IlirAliu_/status/2061862045541109967

Noble Machines’ Moby robot hauling a 50 lb crate. The whole-body control (WBC) is the physical foundation of the learning stack. It absorbs payload variation, contact response, and center-of-mass shifts in real time at the control layer and enabled cleaner data-focused learning
https://x.com/TheHumanoidHub/status/2059859238462353733

Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs
https://www.latent.space/p/andon

Robots can now reconstruct 3D scenes in real time from a single RGB camera. [📍 Projects page + paper] No depth sensor. No retraining. 30 FPS. Researchers at the Imperial College London introduced KV-Tracker, a training-free method that makes heavy models like π³ and Depth
https://x.com/IlirAliu_/status/2059906536240013690

THE LiDAR odometry package your robot needs. Most localization stacks assume the environment will cooperate. Distinct geometry. Clear point cloud differentiation. That assumption fails the moment you deploy in a warehouse, hospital, or industrial facility. GenZ-ICP is built to
https://x.com/IlirAliu_/status/2060783800788160674

Triangle Splatting SLAM”” TL;DR: the first dense RGB-D SLAM system built on differentiable triangle splatting, enabling real-time tracking, mesh reconstruction, collision checking, and scene editing from a unified representation.
https://x.com/Almorgand/status/2061839159438835881

We treat 3d scanning like a tech demo, but it’s actually spatial memory capture. Damn near teleportation. A few hundred photos of my parent’s old home, and now it’s immortalized forever. 3d gaussian splat made w/ reality capture + litchfeld.
https://x.com/bilawalsidhu/status/2061134940813611505

When Does LeJEPA Learn a World Model? It’s a new paper from @ylecun together with @klindt_david and @randall_balestr, that brings LeJEPA theory closer to the real goal of JEPA-style models. → LeJEPA learns a World Model when • The world’s latent variables are Gaussian •
https://x.com/TheTuringPost/status/2060153308392857933

A practical step forward for real-world manipulation: an open-source world model that replaces rigid action chunking with event-grounded prediction. It anchors planning and control to actual physical moments (reach → grasp → contact → place), giving robots more natural
https://x.com/IlirAliu_/status/2060422419479974312

The first open-source unified world model for scalable robot manipulation: 5B-parameter open-source unified video-action world model that combines policy and world modeling to generate robot actions, predict future visuals, and evaluate task progress from observations, language,
https://x.com/IlirAliu_/status/2061870964519076312

Composer 2.5 is now available inside Grok Build. Composer 2.5 is a fast, highly intelligent model that excels on long-running tasks and following complex instructions.
https://x.com/xai/status/2061510464325206163

Grok @Imagine 1.5 Preview is here Try it today in the API:
https://x.com/grok/status/2062225080843747351?s=20

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading