“MCP Claude that have full control on ChatGPT 4o to generate full storyboard in Ghibli style ! All automatic I am doing nothing at all, we live a pretty crazy time @AnthropicAI @OpenAI https://x.com/OdinLovis/status/1905424459185750117

Ghibli effect: ChatGPT usage hits record after rollout of viral feature | Reuters https://www.reuters.com/technology/artificial-intelligence/ghibli-effect-chatgpt-usage-hits-record-after-rollout-viral-feature-2025-04-01/

“very crazy first week for images in chatgpt – over 130M users have generated 700M+ (!) images since last tuesday India is now our fastest growing chatgpt market 💪🇮🇳 the range of visual creativity has been extremely inspiring we appreciate your patience as we try to serve” / X https://x.com/bradlightcap/status/1907810330018726042

“believe it or not we put a lot of thought into the initial examples we show when we introduce new technology” / X https://x.com/sama/status/1905069374035411209

“🔍 Just discovered CountGD – a powerful app that counts objects using text prompts, visual examples, or both! There’s even a notebook to analyze image batches https://x.com/fdaudens/status/1906787856770625566

“The ChatGPT action figure trend is wild — but turning those stills into video really brings them to life. That last hint of uncanniness you can pixel peep in a photo just vanishes in motion. Suddenly, they feel like real objects! https://x.com/bilawalsidhu/status/1906460664148713546

“Pretty wild that 4o learned to estimate depth as an emergent property. Can also generate normal maps and PBR shaders. The depth maps aren’t super accurate right now but holy shit — general purpose multimodal reasoning might take us far.” / X https://x.com/bilawalsidhu/status/1905452746884432110

“Everyone saying ChatGPT’s new image model is “just style transfer” completely misses the point of multimodality. This is closer to an AI equipped with ControlNets, LoRAs, IP adapters — and a graphic designer’s mind. It’s an expert compositor with world knowledge. A creative https://x.com/bilawalsidhu/status/1905299900838875526

“EasyControl just dropped on Hugging Face Adding Efficient and Flexible Control for Diffusion Transformer Free chatgpt style ghibli image generation with easy control https://x.com/_akhaliq/status/1907087568404894136

“People calling this ghibli slop are missing the point. A comfyui workflow got collapsed into a mf-ing text prompt. It’s so energizing to see people riff off each other’s creativity. It’s like seeing a photoshop tennis match play out at insane scale — except everyone can join the” / X https://x.com/bilawalsidhu/status/1905107166719017012

“Ideogram 3.0 launched with new text rendering and graphic design capabilities —Creates complex layouts, logos, and typography —Outperforms Google’s Imagen 3, Flux Pro 1.1, and Recraft V3 —Style References to control generations —Available to free users https://x.com/rowancheung/status/1905130195738075507

“New 4o prompt: Kaiju Conflict “Show this exact scene five minutes after two kaiju monsters emerge and start fighting with each other.” Drop your own kaiju showdowns below! https://x.com/bilawalsidhu/status/1906395112990531926

“Gradio just crossed 1,000,000 monthly developers using it in March! In my opinion, it’s become one of the most important tools for democratization of AI that powered massively impactful apps like LMArena, Stable diffusion, Illusion Diffusion, InstantID, Kokoro, and thousands of https://x.com/ClementDelangue/status/1907122608782688766

“Standard image tokenization allocates coding capacity uniformly, regardless of regional complexity. This is inefficient. This paper introduces TokenSet, representing images as unordered token sets for dynamic capacity allocation based on semantics. It proposes a dual https://x.com/rohanpaul_ai/status/1905781707326005484

“Multimodal LLMs (MLLMs) struggle to accurately detect and categorize student errors in math problems involving both text and images. MATHAGENT solves this by decomposing error detection into three phases handled by specialized agents, improving error step identification accuracy https://x.com/rohanpaul_ai/status/1906896549277319328

“LVLMs visual understanding behaviors are underexplored. This paper introduces a heatmap visualization method for Large Vision-Language Models to reveal relevant image regions for open-ended question Answering. 📌 Token selection pinpoints image-relevant words in free-form https://x.com/rohanpaul_ai/status/1905829266715213908

“Gen-4 is so much fun. I have a series of Midjourney generations in this miniature diorama style that no other model has been able to animate. They freeze all the figures or if they do move totally change the style. Gen-4 interprets beautifully in terms of movement and https://x.com/TomLikesRobots/status/1906847002257760761

“The new image gen in ChatGPT is now out to 100% of free users!” / X https://x.com/kevinweil/status/1906834158405726574

“chatgpt image gen now rolled out to all free users!” / X https://x.com/sama/status/1906867488320843823

“can yall please chill on generating images this is insane our team needs sleep” / X https://x.com/sama/status/1906210479695126886

“Creative agencies are facing one of the biggest disruptions with such high-quality AI generated images starting from one basic product photo. https://x.com/rohanpaul_ai/status/1905777892849856765

“Kevin Frans and colleagues at @UCBerkeley introduced a new way to speed up image generation with diffusion models. Their “shortcut” method trains models to take larger noise-removal steps—the equivalent of multiple smaller ones—without losing output quality. Unlike established https://x.com/DeepLearningAI/status/1906768474816295165

“Creating detailed prompts needed by Text-to-Image models from simple user input is difficult and inefficient. TIPO (Text to Image with Text Presampling for Prompt Optimization) uses a light-weight model to refine simple prompts into detailed ones. Usually, if you just type a https://x.com/rohanpaul_ai/status/1905679791929655415

“UniCombine, a Diffusion Transformer framework for unified multi-conditional image generation. 📌 It uses Conditional Multi-Modal Diffusion Transformer Attention. This allows for unified handling of diverse conditions. 📌 Pre-trained Condition Low-Rank Adaptation modules enable https://x.com/rohanpaul_ai/status/1905797557638365472

[2502.01385v1] Detecting Backdoor Samples in Contrastive Language Image Pretraining https://arxiv.org/abs/2502.01385v1

New AI Innovation in Industry-Leading Adobe Premiere Pro Empowers Video Pros to Generate, Edit and Search Footage at Lightning Speed https://news.adobe.com/news/2025/04/new-ai-innovation-in-industry

“Higgsfield is super cool. Right now it takes a still image and you animate the camera—Ken Burns on steroids. Can’t wait till we can upload videos and add motion in post. Eventually, edit existing camera moves too. Genuinely useful for tons of creators. https://x.com/bilawalsidhu/status/1907572513661399152

[2504.00784v1] CellVTA: Enhancing Vision Foundation Models for Accurate Cell Segmentation and Classification https://arxiv.org/abs/2504.00784v1

“prompt: sam altman as a cricket player in anime style https://x.com/sama/status/1907531078581272931

A Unified Image-Dense Annotation Generation Model for Underwater Scenes
https://hongklin.github.io/TIDE/

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading