Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: Using the provided reference image, keep the pure white landscape background, vertical type stack, galaxy-punchout starfield treatment, and exact font hierarchy, but replace ‘HEROES’ with ‘MULTIMODALITY’ in bold condensed grotesque all-caps with Milky Way texture clipped inside, keep ‘(we could be)’ unchanged beneath it, replace ‘ALESSO’ with ‘THE SYNESTHETE’ in light geometric all-caps galaxy-punchout, keep ‘FEATURING.’ unchanged, and replace ‘TOVE LO’ with ‘CROSS-MODAL FUSION’ in bold condensed grotesque all-caps galaxy-punchout, matching letter tracking and austere album-cover composition exactly.
Palantir for tracking your uber eats and fedex deliveries
https://x.com/bilawalsidhu/status/2046620539955872215
Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis | Ai2
https://allenai.org/blog/olmoearth-embeddings
A Visual Thought Partner ChatGPT Images 2.0 is our first image model with thinking capabilities. When a thinking model is selected in ChatGPT, Images 2.0 can search the web for real-time information, create multiple distinct images from one prompt, double-check its own outputs,
https://x.com/OpenAI/status/2046670989719924768
Claude Opus 4.7 from @AnthropicAI takes #1 in Vision & Document Arena! In Document Arena: Opus 4.7 lands +4 points over Opus-4.6 and +45 over the next non-Anthropic model, GPT-5.4 (#6). This is huge ~70 pts lead over Muse Spark and Gemini-3.1-Pro. Real world research work like
https://x.com/arena/status/2046224760657658239
TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
https://gdm-tipsv2.github.io/
vision🍌 is here
https://t.co/MjL0xUt82H if you got into computer vision the way I did, starting with pixel-level labeling tasks like segmentation, edges, depth, or surface normals, you’ll probably feel the same seeing these results — something big has quietly shifted, and it’s
https://x.com/sainingxie/status/2047339789926429166
A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features”” TL;DR: feed-forward localization builds a lightweight feature map and estimates camera pose in one pass, achieving fast and accurate relocalization across large scenes
https://x.com/Almorgand/status/2045194191081251178
Z-Image experiment. I expanded the patch-2 layers to patch-4. New layers = patch-2 layers averaged over sub-patches (in) / replicated (out), so the weights are already close with zero training. Finetuning now to clean it up. If it works: 2× image size at the same compute.
https://x.com/ostrisai/status/2045677110413668743
Image models tend to get much more stuck on a particular direction than text models, requiring clearing the context window fairly often. PerfectSquashBench is my new measure of how image models anchor. The squash remains merely fine after many attempts.
https://x.com/emollick/status/2047073009312121000
What you need to know about the Deep Research and Deep Research Max Update: – Can consults over 100 sources in one research task. – Generates native charts and infographics inline. – Accepts PDFs, CSVs, images, audio, and video inputs. – Max version uses ~160 search queries per
https://x.com/_philschmid/status/2046627179551944753
Meta just released Sapiens2 on Hugging Face High-resolution vision transformers pretrained on 1 billion human images, for human-centric perception: pose, segmentation, normals, and pointmaps.
https://x.com/HuggingPapers/status/2047410529010844044
new image model coming with some real magic within, to unlock new use cases in productivity and creativity livestream noon today
https://x.com/gdb/status/2046632580527554572
MIT engineers built an AI wristband that controls robots by reading your hand muscles. It works by using an ultrasound to capture images of the muscles and tendons in your wrist. An AI algorithm then translates those images into the exact position of all 5 fingers, tracking 22
https://x.com/rowancheung/status/2045158931367072104
Context Unrolling in Omni Models – A unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations – Enables Context Unrolling, where the model explicitly reasons across multiple modal representations
https://x.com/arankomatsuzaki/status/2047519009004716097
New in LangSmith Fleet: Create and edit files with your agent. Your agent can now work with files directly. Create documents, presentations, and webpages inside a conversation, or upload your own files and edit them together. 📄 Work with images, PDFs, and text files. 💬 Build
https://x.com/LangChain/status/2047362259983495215
Yay, finally! Introducing Vision Banana🍌 from @GoogleDeepMind, our unified model that outperforms SoTA specialist models on various vision tasks! By treating 2D/3D vision tasks as image generation, we unlock a new foundation for CV. Project page:
https://t.co/GQgRi6mWwC (1/5)
https://x.com/songyoupeng/status/2047312019976785944
Building a Fast Multilingual OCR Model with Synthetic Data
https://huggingface.co/blog/nvidia/nemotron-ocr-v2
VLM Performance:Qwen3.6-27B is natively multimodal, supporting both vision-language thinking and non-thinking modes in a single unified checkpoint — the same as Qwen3.6-35B-A3B. It handles images and video alongside text, enabling multimodal reasoning, document understanding,
https://x.com/Alibaba_Qwen/status/2046939788184547610





Leave a Reply