object segmentation. a sunny day with blue skies. an apple orchard. all of the apples are starkly bold primary colors. a sign reads “Multimodality” –chaos 50 –ar 4:3 –style raw –personalize 9zxyhz8
“Multimodal Table Understanding Introduces Table-LLaVa 7B, a multimodal LLM for multimodal table understanding. Competitive with GPT-4V and significantly outperforms existing MLLMs on multiple benchmarks. They also develop a large-scale dataset MMTab, covering table images,
Trint (@TrintHQ) / X
Trint’s speech-to-text platform makes audio and video searchable, editable and shareable.
“MMWorld Towards Multi-discipline Multi-faceted World Model Evaluation in Videos Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of “world models” — interpreting and reasoning about complex real-world dynamics. To assess these abilities,
“Researchers from Microsoft, Google, MIT, and Oxford released new research presenting DenseAV. It’s an AI algorithm that learns the meaning of words and sounds by watching videos, and can help better understand animal communication and discover language.
“Depth Anything V2 This work presents Depth Anything V2. Without pursuing fancy techniques, we aim to reveal crucial findings to pave the way towards building a powerful monocular depth estimation model. Notably, compared with V1, this version produces much finer and more
“TikTok presents Depth Anything V2 Trained from 595K synthetic labeled images and 62M+ real unlabeled images, providing the most capable monocular depth estimation model proj:
SEO is dead: AI will rescan the internet and make duplicate content worthless. – Ethan B. Holland – https://ethanbholland.com/2024/07/05/seo-is-dead-ai-will-rescan-the-internet-and-make-duplicate-content-worthless/
“What If We Recaption Billions of Web Images with LLaMA-3? Web-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance model training across various
“What If We Recaption Billions of Web Images with LLaMA-3 ? – Finetunes a LLaVA-1.5 and recaptions ~1.3B images from the DataComp-1B dataset – Opensources the resulting dataset data:
“”What If We Recaption Billions of Web Images with LLaMA-3?”🤯 And the results confirm that this enhanced dataset, Recap-DataComp-1B generated this way, offers substantial benefits in training advanced vision-language models. For discriminative models like CLIP, we observe
[2406.08478] What If We Recaption Billions of Web Images with LLaMA-3?
“DenseAV is an algorithm capable of discovering the meaning of language and locations of sounds just by watching unlabeled videos. DenseAV is completely unsupervised and never sees text during its training. Learn more:
OpenAI
(must see) “sharing a peek of the 4o demo we did at @openai’s New York tech week reception! met so many cool ny founders and builders – huge props to the team for making this happen!
(must see) Real time multimodal audio dialog demo with translations

Heads up! You’ve scrolled to the end of this category. There may have been just one or two links (above), so go back up and double check to be sure you didn’t quickly scroll down past it.
Be Sure To Read This Week’s Main Post:
This week’s executive overview and top links are here:
AI News #37: Week Ending 06/14/2024 with Executive Summary and Top 7 Must-Read Links
The post you just read is an deep dive extension of my weekly newsletter, This Week In AI, an executive summary of the top things to know in AI. Each week, I create an accessible overview for laypeople to feel confident they are conversant with the week’s AI developments. I include a curated list of must-click links of the week, to offer everyone a hands-on opportunity to explore the most intriguing updates in artificial intelligence across various categories, including robotics, imagery, video, AR/VR, science, ethics, and more. Beyond the overview, I post these topic-based deeper dives (below). If you haven’t read this week’s overview, I recommend starting there.
- Agents/Copilots
- Amazon
- Apple
- Artificial General Intelligence (AGI)
- Augmented and Virtual Reality (AR/VR)
- Autonomous Vehicles
- AI Audio
- Business and Enterprise AI
- Chips and Hardware
- Consumer Products
- Education
- Ethics/Legal Security
- Images/Photos
- International AI News
- Locally Run AI Models
- Mobile
- Meta
- Microsoft
- OpenAI
- Open Source
- Podcasts/YouTube
- Publishing and News
- Retrieval-Augmented Generation (RAG) News
- Robots and Embodiment
- Science and Medicine
- Video
- Vision/Multimodality
- X/Twitter/Grok
- Tech and Development
Credits/Sources

Most of these weekly links come from just a few prolific oversharing sources. Please follow them, as they work hard to find the news each week and they make it a lot easier for me to compile.
- Robert Scoble: https://x.com/Scobleizer
- Ethan Mollick: https://www.linkedin.com/in/emollick/
- Alan Thompson: https://lifearchitect.ai/
- Theoretically Media: https://www.youtube.com/@TheoreticallyMedia
- The Rundown: https://www.therundown.ai/
- Bilawal Sidhu: https://twitter.com/bilawalsidhu/
- TLDR: https://tldr.tech/ai
- Jeremiah Owyang: https://twitter.com/jowyang
- Nick St. Pierre: https://twitter.com/nickfloats
- Dr. Jim Fan: https://twitter.com/DrJimFan
- All About AI: https://www.youtube.com/@AllAboutAI
- Marshall Kirkpatrick: https://aitimetoimpact.com/
- AI News (Smol Talk): https://buttondown.email/ainews/archive/
- Andrej Karpathy: https://x.com/karpathy
- Brett Adcock: https://x.com/adcock_brett
- Florent Daudens: https://x.com/fdaudens
- Ate-a-Pi: https://x.com/8teAPi
- Francesco Marconi: https://x.com/fpmarconi
- Charlie Beckett: https://x.com/CharlieBeckett
For previous issues, please visit the archives!

Thanks for reading!





Leave a Reply