Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Cinematic 80s suburban front porch at twilight with costumed trick-or-treaters holding various devices—Walkman, Polaroid camera, walkie-talkie—as glowing translucent ribbons of sound waves, light trails, and speech bubbles swirl together above them in the crisp autumn air, fall leaves scattered on steps, warm porch light, nostalgic film grain

A weird gap in the Google AI line-up is between the Gemini research tools and NotebookLM. Gemini lets you do Deep Research, but only a subset of other NotebookLM features, while NotebookLM won’t let you trigger a deep research report or do other types of AI interactions.”” / X https://x.com/emollick/status/1983611024113856610

A next-gen visual model trained on structured JSON for precise, controllable generation. 💪 FIBO is a text-to-image model that transforms prompts into JSON schemas, enabling predictable visuals at scale. Trained on extended, structured captions—often 1,000+ words—FIBO https://x.com/bria_ai_/status/1983564638697517549

IBM just released Granite 4.0 Nano, a family of four tiny open weights models (1B and 350M) focused on efficiency with strong performance for their size Granite 4.0 H 1B and Granite 4.0 1B score 14 and 13 on the Artificial Analysis Intelligence Index, while the 350M variants https://x.com/ArtificialAnlys/status/1983611955668775411

open-source OCR models are super cheap to run and privacy first 🤝 BUT there’s a ton of new models out there: DeepSeek-OCR, Nanonets, PaddleOCR, how do you pick them? 🤯 don’t worry though, @huggingface got you covered! 🫡🧶 https://x.com/mervenoyann/status/1980685830411931885

Emu3.5: Native Multimodal Models are World Learners The Emu series of models have been very interesting in the multimodal space. This one seems to add a diffusion image generation mode despite being trained with next-token prediction. Also claims to be comparable to Nano Banana https://x.com/iScienceLuvr/status/1984190340279234888

Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer A new paper from Weizmann Institute of Science, getting reconstructions that are not complete nonsense from ONLY 15 min of data. We previously demonstrated SOTA for 1 hr of data with MindEye2. This https://x.com/iScienceLuvr/status/1984195725253804449

NVIDIA Just Released 8M Sample Open Dataset + OCR Tooling on @huggingface – 3x larger than v1 (just 2 months ago!) – Image/video QA, reasoning, multilingual OCR – Commercial-ready (CC-BY-4.0) @NVIDIAAI is one of the few major AI labs releasing datasets 🤗 https://x.com/vanstriendaniel/status/1983238971644608924

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading