Image created with gemini-2.5-flash-image with claude-sonnet-4-5-20250929. Image prompt: A cinematic photograph of a large canvas on an easel displaying a half-finished AI-generated artwork blending photorealistic mountains with abstract colorful swirls in deep blues, rich reds, and slate greys. Two elegant lit birthday candles sit directly on the canvas surface, their warm flames illuminating the digital brushstrokes with high contrast lighting against a dark background.

SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views”” TL;DR: feed-forward framework for 3DGS from sparse unposed views; predicts Gaussians + poses, enforces geometry via reprojection, SOTA novel view synthesis, even in extreme settings.”” https://x.com/Almorgand/status/1970910944948781195

Nvidia just released Lyra on Hugging Face Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation TL;DR: Feed-forward 3D and 4D scene generation from a single image/video trained with synthetic data generated by a camera-controlled video diffusion model https://x.com/_akhaliq/status/1970949464606245139

🍌 @GeminiApp just passed 5 billion images in less than a month. What a ride, still going! Latest trend: retro selfies of you holding a baby version of you. Can’t make this stuff up!”” / X https://x.com/joshwoodward/status/1970894369562796420

Meta scrapping Unity to build their own game engine (Horizon Engine) is really interesting. I doubt it has as much to do with the Unity tax and more so to allow them to vertically integrate with all their own layers of ~SOTA AI starting with gaussian splatting”” / X https://x.com/nearcyan/status/1968475789021852075

Wow. Qwen Image Edit now has native support for ControlNet (depth maps, edge maps, keypoint maps etc)”” / X https://x.com/bilawalsidhu/status/1970193454505541755

I noticed nano banana has changed PowerPoint. People who are good at using it (& have good imaginations) can come up with genuinely funny or interesting images that make presentations more compelling and coherent – no more pasted-in cartoons A new skill for the presenting class”” / X https://x.com/emollick/status/1968519431908180429

Today, we’re officially launching Wan2.5-Preview! It’s set to reshape the future of visual generation with a new architecture and powerful features. • Architectural Features: Native Multimodality, Deep Alignment ∘ Native Multimodal Architecture: Adopts a new, unified framework”” / X https://x.com/Alibaba_Wan/status/1970697244740591917

Apple presents Manzano: Simple & scalable unified multimodal LLM • Hybrid vision tokenizer (continuous ↔ discrete) cuts task conflict • SOTA on text-rich benchmarks, competitive in gen vs GPT-4o/Nano Banana • One model for both understanding & generation • Joint recipe: https://x.com/arankomatsuzaki/status/1969974676802990478

Introducing: Hyperscape Capture 📷 Last year we showed the world’s highest quality Gaussian Splatting, and the first time GS was viewable in VR. Now, capture your own Hyperscapes, directly from your Quest headset in only 5 minutes of walking around. https://x.com/JonathonLuiten/status/1968474776793403734

Meshcapade can now pull apart both 3D camera tracking + human pose estimation data. Effectively like Wonder Dynamics / Autodesk Flow Studio at this point. Useful AI model for 3D artists and as an input into video-to-video workflows. https://x.com/bilawalsidhu/status/1969458783480135711

I don’t want much. I just want nano banana for video. And i’m sure I’ll get soon :-)”” / X https://x.com/bilawalsidhu/status/1968732880244228490

Decart has dropped their first version of Nano Banana for video. Image editing is fun, but I really love instruction-based video editing. Feels like going from photoshop —> blender + after effects. https://x.com/bilawalsidhu/status/1968923439369994586

GenExam: The first multidisciplinary text-to-image exam is now on Hugging Face This new benchmark challenges T2I models with 1,000 rigorous, exam-style prompts across 10 subjects. It comes with ground-truth images and detailed scoring for semantic correctness and visual https://x.com/HuggingPapers/status/1968527551703433595

LLM-JEPA: much worse than image JEPAs imo I read through the paper and it feels super useless because they are constrained to paired data like “”text <-> SQL””. It’s not generalizable to arbitrary data, it just adds a term in the loss so that the embeddings of the “”SQL and text”””” / X https://x.com/scaling01/status/1969410266304545066

Microsoft introduces Latent Zoning Network (LZN) A unified principle for generative modeling, representation learning, and classification. LZN uses a shared Gaussian latent space and modular encoders/decoders to tackle all three core ML problems at once! https://x.com/HuggingPapers/status/1970218823140687885

Manzano is a multimodal LLM that unifies image understanding and generation. It uses a shared ViT with two adapters: continuous embeddings for image-to-text and discrete FSQ tokens (64K) for text-to-image, both in the same semantic space to reduce task conflict. A single https://x.com/gm8xx8/status/1969974517024923936

Kimi Here is the page link and prompt: link: https://t.co/RChtmDlqF5 prompt: “”Design a minimalistic brand website in the style of Le Labo. The layout resembles a laboratory logbook: black-and-white color scheme, lots of white space, typewriter-style fonts, and product labels that look https://x.com/crystalsssup/status/1971183638734832004

We are building “Open Source Nano Banana for Video” – here is open source demo v0.1 We are open sourcing Lucy Edit, the first foundation model for text-guided video editing! Get the model on @huggingface 🤗, API on @FAL, and nodes on @ComfyUI 🧵 https://x.com/DecartAI/status/1968769793567207528

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading