Image created with gemini-2.5-flash-image with claude-sonnet-4-5-20250929. Image prompt: A solemn courtroom witness stand in a historic London legal chamber with dark oak paneling and leaded windows, where visual art, musical notation, and luminous text simultaneously emanate from the carved stone platform, converging in golden light above it. Cinematic photograph with regal lighting, emphasizing the dignified integration of multiple forms of testimony into one trusted source of truth.

OpenAI Ramps Up Robotics Work in Race Toward AGI | WIRED https://www.wired.com/story/openai-ramps-up-robotics-work-in-race-toward-agi/

the proof is in the ice cream (GPT telling the difference between ice cream and dogs”” / X https://x.com/gdb/status/1967038157586649118

So it looks like Claude got there first: an actually smart phone assistant that can take complex requests that involve both common sense and complicated constraints. It is still beta feeling though & I found I needed to use the bigger Opus model as Sonnet was not smart enough. https://x.com/emollick/status/1966170169367232556

Introducing SpatialVID: A massive new video dataset for 3D spatial intelligence Crucial for training next-gen models, it features over 7,000 hours of diverse, in-the-wild video with dense annotations like camera poses, depth maps, and dynamic masks. https://x.com/HuggingPapers/status/1967260292569845885

🤖🌐 ParserGPT: Smart Web Scraping Transform messy websites into clean CSV data using LLMs and deterministic rules. Powered by LangChain and LangGraph, ParserGPT learns website structures once to enable efficient, repeated data extraction. Learn more about ParserGPT: https://x.com/LangChainAI/status/1967257030756028505

🖇️ IBM just released a tiny document VLM, Granite-Docling-258M. – converts PDFs into structured text formats like HTML or Markdown while preserving layout – accurately recognizes and format equations including inline math, handle tables, code blocks, and charts, support https://x.com/rohanpaul_ai/status/1968561354987442246

ever wondered how the text search on your phone image gallery works? @AIatMeta released MetaCLIP2, and we’ve added it to @huggingface transformers 🔥 it’s a multilingual model that can understand image + text! find the notebook for text-to-image search on the next ⤵️ https://x.com/mervenoyann/status/1966544046744011242

Excited to release a preview of Moondream 3. A 9B param, 2B active MoE vision language model that makes no compromises; offering state-of-the-art visual reasoning while still retaining an efficient and deployment-friendly form factor. https://x.com/vikhyatk/status/1968800178640429496

Say hi to https://x.com/interaction/status/1965093198482866317

NavFoM: Embodied Navigation Foundation Model • Trained on 8M samples across quadrupeds, drones, wheeled robots & vehicles • Handles VLN, ObjNav, tracking, & driving in one unified model • Outperforms on each domain + real-world deployment https://x.com/arankomatsuzaki/status/1967806725387588069

How Reka Speech does efficient timestamped transcription: 1. Encode audio with a 300M encoder → feed into a 500M backbone LM 2. During prefilling, offload the query/key embeddings to CPU (QK cache) 3. Generate transcript autoregressively (KV cache stays on GPU) 4. Pull the QK https://x.com/artetxem/status/1968027334033682727

Introducing Reka Speech: an efficient and accurate transcription & translation model. 🗣️ On modern GPUs (e.g. H100), it runs 8x–35x faster than existing solutions for batch processing. https://x.com/RekaAILabs/status/1967989101111722272

(13) Made On YouTube 2025: Auto-Dubbing – YouTube https://www.youtube.com/watch?v=8W3noE2Uxag

me and @TodePond spent the last couple months building a really good whiteboard agent, out now in 4.0 ! check it out with npm create tldraw it has a really good readme that we spent a lot of time on !”” / X https://x.com/max__drake/status/1968764136419975599

Optimize work and maximize outcomes with the next generation of AI Companion for Zoom Workplace | Zoom – Zoom https://news.zoom.com/ai-companion-3-0-and-zoom-workplace/

transformers @huggingface comes with Kosmos2.5 model by @MicrosoftAI 🤩 to celebrate this, we built a notebook for fine-tuning with OCR with detection 🔥 you can also try the model out of the box with the demo 🫡 https://x.com/mervenoyann/status/1966487632659005667

tldraw canvas agent starter kit is dropping today https://x.com/tldraw/status/1968655029247648229

Introducing Magistral Small 1.2 & Magistral Medium 1.2, minor updates to our Magistral 1.1 models! – Multimodality: Now equipped with a vision encoder, these models handle both text and images seamlessly. – Performance Boost: 15% improvements on math and coding benchmarks such https://x.com/MistralAI/status/1968670593412190381

Magistral Medium 1.2 is out Multimodality: Now equipped with a vision encoder, these models handle both text and images seamlessly Performance Boost: 15% improvements on math and coding benchmarks such as AIME 24/25 and LiveCodeBench v5/v6 default in anycoder under Magistral https://x.com/_akhaliq/status/1968708201236381858

> demo (for OCR with boxes and conversion to markdown): https://x.com/mervenoyann/status/1966488556831977672

IBM just released small swiss army knife for the document models: granite-docling-258M 🔥 not only a document converter but also can do document question answering, understand multiple languages 🤯 with Apache 2.0 license 👏 https://x.com/mervenoyann/status/1968316714577502712

ByteDance unveils SAIL-VL2, a SOTA vision-language foundation model. It achieves comprehensive multimodal understanding and reasoning, outperforming at 2B & 8B scales. https://x.com/HuggingPapers/status/1968588429433913714

here’s a notebook tutorial to walk you through how to do this step-by-step 💗 GH → https://x.com/mervenoyann/status/1966544570436424074

Beautiful open smol moe for vision language task Weight are here if you want to try: https://x.com/eliebakouch/status/1968809452640825650

1/ Introducing Isaac 0.1 — our first perceptive-language model. 2B params, open weights. Matches or beats models significantly larger on core perception. We are pushing the efficient frontier for physical AI. https://x.com/perceptroninc/status/1968365052270150077

Excited to introduce Isaac – our first open-weights model which excels at localization and visual understanding.”” / X https://x.com/kilian_maciej/status/1968396992104874452

One of the cooler results, is Isaac 0.1 out-performs or matches YOLO models trained on thousands of samples with a dozen in-context samples for very specific detection tasks. https://x.com/ArmenAgha/status/1968378019753627753

“People who are serious about robot learning should build their own hardware,” says NVIDIA’s embodied AI research co-lead. This is likely a general statement, not a hint at NVIDIA’s plans, but it would be awesome if NVIDIA designed and made its own robot hardware. https://x.com/TheHumanoidHub/status/1966216768290222552

Sergey Levine says making a robotics foundation model is more like the Apollo program than it is like a science experiment. https://x.com/TheHumanoidHub/status/1967960666217566694

Unitree first open-source world-model on @huggingface! UnifoLM-WMA-0 is Unitree‘s first open-source world-model–action architecture spanning multiple types of robotic embodiments, designed specifically for general-purpose robot learning. Its core component is a world-model https://x.com/ClementDelangue/status/1968001710770520135

What if humanoid robots could learn new tasks just by watching people play? ❗️That’s the idea: Instead of relying on expensive teleoperation data, MimicDroid trains robots using regular human play videos. This makes it cheaper, faster, and more scalable for robots to adapt to https://x.com/IlirAliu_/status/1968216390155841582

(1/N) How close are we to enabling robots to solve the long-horizon, complex tasks that matter in everyday life? 🚨 We are thrilled to invite you to join the 1st BEHAVIOR Challenge @NeurIPS 2025, submission deadline: 11/15. 🏆 Prizes: 🥇 $1,000 🥈 $500 🥉 $300 https://x.com/drfeifei/status/1962971299246178664

Paper2Agent brings research papers ‘to life.’ This open tool from @Stanford transforms static papers into interactive AI assistants that can explain and apply their methods. It builds on the MCP and works in 2 layers: – Paper2MCP: Extracts the paper’s methods and code into an https://x.com/TheTuringPost/status/1968829219858956774

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading