Google is adding digital watermarks to images edited with Magic Editor AI | TechCrunch https://techcrunch.com/2025/02/06/google-is-adding-digital-watermarks-to-images-edited-with-magic-editor-ai/

New creative updates to help advertisers generate lifestyle imagery https://blog.google/products/ads-commerce/new-creative-updates-advertisers-generate-lifestyle/

“Is Noise Conditioning Necessary for Denoising Generative Models? “Motivated by research on blind image denoising, we investigate a variety of denoising-based generative models in the absence of noise conditioning. To our surprise, most models exhibit graceful degradation, and in https://x.com/iScienceLuvr/status/1892053059221717486

“Large Language Diffusion Models Presents LLaDA, a 8B diffusion LM, trained entirely from scratch, rivaling LLaMA3 8B in performance despite being trained on 7x fewer tokens (2T tokens). https://x.com/arankomatsuzaki/status/1891343406334693879

“PaliGemma 2 Mix is here 💥 @GoogleDeepMind PaliGemma 2 mix is an open, multi-task vision-language model that can judge your outfit while counting all the red balls in an image. 🖼️ Trained for Question Answering, Captioning, OCR, Detection, Segmentation and more. ⬇️ https://x.com/_philschmid/status/1892258568176320927

“Let’s goo! Quite psyched to onboard @nebiusaistudio @novita_labs & @hyperbolic_labs over to Hugging Face Hub Inference Providers! 🔥 Starting today you can access SoTA VLMs (Qwen 72B VLM), Text to Image models like Flux AND LLMs like DeepSeek R1 with even more providers directly https://x.com/reach_vb/status/1891909914412433839

“The challenge in text-to-video generation lies in achieving high visual quality at high resolutions due to immense computational costs. Single-stage diffusion models require substantial parameters and function evaluations, making high-resolution video generation inefficient. https://x.com/rohanpaul_ai/status/1892179215631499459

“in what sense is this diffusion? I see no SDE, no probability flow, no noise. not every iterative sampling method is diffusion! this paper is genuinely impressive but it’s a new thing I don’t see how I would port diffusion intuitions over to it.” / X https://x.com/gallabytes/status/1891356261582557438

[2502.08391v1] ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification https://arxiv.org/abs/2502.08391v1

“As I wait for the SigLIP2 OpenCLIP and timm checkpoints to upload, a few comments. Like v1, these are very strong image encoders, but even better. They should become a go to ViT encoder for many tasks. Yet, even when we had ViT encoders from SigLIP, DFN CLIP, DataComp CLIP,” / X https://x.com/wightmanr/status/1892981509461540867

“Diffusion Models without Classifier-free Guidance Directly learning the modified score from classifier-free guidance during training, leading to faster convergence and eliminating the need for two model forward passes during inference. Achieves new SOTA FID on ImageNet 256×256 https://x.com/iScienceLuvr/status/1891847953087619147

“Current diffusion-based video generation is computationally and memory intensive, making it inaccessible on smartphones. This paper introduces On-device Sora. It is a pioneering solution to enable diffusion-based text-to-video generation directly on mobile devices. On-device https://x.com/rohanpaul_ai/status/1891408704496652446

“Arena Explorer is now ready for 🎨 Text-to-image & 🌐 Web Dev! You can dig into the user-submitted prompts for both features, and go even deeper into sub-categories. #OpenSource 🎨 Here are the top image generation prompt categories: 🔹 Art Imagery 🔹 Fashion Photography 🔹 https://x.com/lmarena_ai/status/1892617996251889878

“Amazing new open-source text-to-video model just dropped. Image-to-video generation models struggle with synchronizing motion and maintaining realism. Existing diffusion-based models generate high-quality frames but fail in complex physics and logical sequences. Step-Video-T2V https://x.com/rohanpaul_ai/status/1891413622867288100

[2502.12154v1] Diffusion Models without Classifier-free Guidance https://arxiv.org/abs/2502.12154v1

Large Language Diffusion Models https://ml-gsai.github.io/LLaDA-demo/

“Testing Topaz Starlight, their new diffusion-based video upscaling tech to breathe new life into old videos. Candidate #1: My VFX showreel from 2002. The results are…. interesting 🫣 Input vs. Output. More examples below (1/3) https://x.com/bilawalsidhu/status/1890499502899142927

“Image-to-video generation offers limited control over synchronizing camera and object movements in cinematic shots. Current methods lack intuitive and precise control over these intertwined movements, hindering creative expression. This paper proposes MotionCanvas, a system https://x.com/rohanpaul_ai/status/1891406891047104699

“I’m waiting for the checkpoints, but a comeback of diffusion models in language wasn’t something I expected. https://x.com/maximelabonne/status/1891472925603090497

“Large Language Diffusion Models (LLaDA) Proposes a diffusion-based approach that can match or beat leading autoregressive LLMs in many tasks. If true, this could open a new path for large-scale language modeling beyond autoregression. More on the paper: Questioning https://x.com/omarsar0/status/1891568386494300252

“RelaCtrl Relevance-Guided Efficient Control for Diffusion Transformers https://x.com/_akhaliq/status/1892779115847012854

[2502.10294v1] QMaxViT-Unet+: A Query-Based MaxViT-Unet with Edge Enhancement for Scribble-Supervised Segmentation of Medical Images https://arxiv.org/abs/2502.10294v1

“Large Language Diffusion Models introduce LLaDA, a diffusion model with an unprecedented 8B scale, trained entirely from scratch, rivaling LLaMA3 8B in performance. A text generation method different from the traditional left-to-right approach Prompt: Explain what artificial https://x.com/_akhaliq/status/1891487050936815693

“Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation https://x.com/_akhaliq/status/1892781056777936950

“This is unpleasantly phrased, but also wrong? If “boomer prompts” means using full sentences as a prompt rather than lists of tags (what it meant in the Midjourney community), then boomer prompting is actually quite effective for reasoners & matches OpenAI’s guidelines.” / X https://x.com/emollick/status/1890186641777864942

“Want strong SSL, but not the complexity of DINOv2? CAPI: Cluster and Predict Latents Patches for Improved Masked Image Modeling. https://x.com/TimDarcet/status/1890389871543419255

“Adobe just launched AI video generator. key features: – generate motion graphics (new in AI world) – text/image to video – Shot size control (close up, wide shot.. ) – Camera angle control (low angle, aerial shot ..) – camera motion control – keyframe link in comment https://x.com/EHuanglu/status/1889726675862098159

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading