“Aya can see now!” / X https://x.com/aidangomez/status/1896946200135495708

“proj: https://x.com/arankomatsuzaki/status/1896371410123309432

“Even more releases today! We’ve been hard at work with the great folks at @hedra_labs getting ready for this launch. Check out Character-3, deployed and scaled on @togethercompute!” / X https://x.com/realDanFu/status/1897757302243156440

“i think the droids will use Gibberlink to communicate with each other” / X https://x.com/ggerganov/status/1896615325116035244

“@MistralAI Check out our blog post: https://x.com/sophiamyang/status/1897716142401060867

“@MistralAI Multilingual capability 🔥 https://x.com/sophiamyang/status/1897715804042338327

The OS for Human-AI Interaction https://www.tavus.io/

“📽️ AI content creation just took a huge leap forward. We’re teaming up with @hedra_labs to bring you Character-3, the most powerful omnimodal AI model in production—now live inside Hedra Studio. No complicated tools. No long workflows. Just AI-powered content creation.” / X https://x.com/togethercompute/status/1897756209069138116

(1) Aran Komatsuzaki on X: “Nvidia presents: Token-Efficient Long Video Understanding for Multimodal LLMs SotA results across various long video understanding benchmarks while reducing the computation costs by up to 8x and the decoding latency by 2.4-2.9x for the fixed numbers of input frames https://t.co/0M6MQliA11” / X
https://x.com/arankomatsuzaki/status/1897854511814770733

“Announcing: Agentic Document Extraction! PDF files represent information visually – via layout, charts, graphs, etc. – and are more than just text. Unlike traditional OCR and most PDF-to-text approaches, which focus on extracting the text, an agentic approach lets us break a https://x.com/AndrewYNg/status/1895183929977843970

“Here is another Gibberlink experiment: Two AI agents autonomously encrypt their audio chat (video by Anton Pidkuiko) https://x.com/ggerganov/status/1896592079997788300

“Microsoft released the most powerful vision language action model this week 🔥 MAGMA-8B can operate in both physical and digital world: embodied robots, web automation and more! 🤯 https://x.com/mervenoyann/status/1895497146344026184

“@MistralAI love using Mistral OCR to extract math equations from pdfs: https://x.com/sophiamyang/status/1897715242936713364

Mistral OCR | Mistral AI https://mistral.ai/en/news/mistral-ocr

“Human: “Put the Ketchup away, where you think it belongs” (Ketchup was not present in the training set) https://x.com/adcock_brett/status/1897399420595134526

“@MistralAI An example of the OCR model extracting text as well as imagery from a given PDF into a markdown file: https://x.com/sophiamyang/status/1897713540506824954

“Mistral OCR is nice and fast but other models outperform it on document processing. We did a comprehensive benchmark on Mistral OCR and compared it against a comprehensive set of different LLM/LVM-powered parsing techniques – these direct parsing using gemini https://x.com/jerryjliu0/status/1898037050185859395

“Huge VLM release from @CohereForAI is just in 🔥 Aya-Vision is a new VLM family based on SigLIP and Aya, and it outperforms many larger models 🤩 > 8B and 32B models covering 23 languages and two new benchmark dataset 🔥 > supported by @huggingface transformers from get-go! 🤗 https://x.com/mervenoyann/status/1896924022438588768

Aya Vision: Expanding the worlds AI can see https://cohere.com/blog/aya-vision

“TALKPLAY Multimodal Music Recommendation with Large Language Models https://x.com/_akhaliq/status/1895532477823013144

“new sota for multilingual vision 🙂 Aya-Vision-32B https://x.com/nickfrosst/status/1896948730622075051

“I have been impressed by GPT-4.5’s vision ability. It can differentiate and count much better than any other model. It even spotted the butterfly. https://x.com/emollick/status/1895211249656570258

“”We do not describe the world we see, we see the world we can describe.” René Descartes Very proud to release Aya Vision 🌿 today, which expands the worlds AI can see. We pushed very hard to build something efficient, accessible and global. This is an important step forward.” / X https://x.com/sarahookr/status/1896953483913498722

“Announcing @MistralAI OCR –  the world’s best document understanding API. 🔍 State-of-the-art understanding of complex documents 🌍 Natively multilingual and multimodal ⚡ Fastest in its category 📄 Doc-as-prompt, structured output 🔒 Available for on-prem deployment https://x.com/sophiamyang/status/1897713370029068381

“HAIC Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models https://x.com/_akhaliq/status/1896396258165871089

“Mistral released a SOTA multimodal OCR https://x.com/scaling01/status/1897695665871872427

“Nice video by @Sam_Witteveen on @MistralAI OCR 🔥 https://x.com/sophiamyang/status/1898059704351277297

AI models make precise copies of cuneiform characters https://phys.org/news/2025-03-ai-precise-cuneiform-characters.html

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading