Learn everything there is to know about LlamaParse in this comprehensive video! In this video, @mesudarshan covers: ➡️ Multiple parsing modes ➡️ Using parsing instructions to improve quality ➡️ Output formats available ➡️ Parsing audio and images ➡️ JSON mode ➡️ Using it all in https://x.com/llama_index/status/1890499579214491967
stepfun-ai/

Step-Audio a 130 billion parameter multimodal LLM that is responsible for understanding and generating human speech https://x.com/_akhaliq/status/1891528348590833834

“Microsoft silently updated OmniParser on the hub 👀 60% faster than v1 – sub-second latency on a 4090! “OmniParser is a general screen parsing tool, which interprets/converts UI screenshot to structured format, to improve existing LLM based UI agent.” Bonus: you can try it https://x.com/reach_vb/status/1891467489030082875

Google Lens updates: iOS features and AI Overviews https://blog.google/products/google-lens/lens-on-ios-ai-overviews/

“Variational Rectified Flow Matching New paper from Apple; introduces a framework that enhances classic rectified flow matching by modeling multi-modal velocity vector-fields. “variational rectified flow matching introduces a latent variable that permits to disentangle https://x.com/iScienceLuvr/status/1890306105043218450

“We’ve made a step change quality improvement in our voice changer model, now available in playground and API. It’s now state of the art, the style transfer abilities are amazing. Record and turn your voice into an amazing character in one click.” / X https://x.com/krandiash/status/1892725226498359365

“Today, we’re excited to announce a beta release of Zonos, a highly expressive TTS model with high fidelity voice cloning. We release both transformer and SSM-hybrid models under an Apache 2.0 license. Zonos performs well vs leading TTS providers in quality and expressiveness. https://x.com/ZyphraAI/status/1888996367923888341

“BREAKING: Grok’s voice unveiled https://x.com/teslaownersSV/status/1891719294469222495

“This video has been entirely generated with AI (face and voice) on the last version on Argil AI 🤯 I have known @LaodisOfficial for years, and for the first time, we have reached the point where the difference between going into a studio and using AI is no different. Do you https://x.com/BrivaelLp/status/1890311661241749821

“📱 Turn any text into a podcast instantly! Transform articles, papers & blogs into audio content using open-source AI (deepseek-r1) + kokoro TTS. Think NotebookLM but fully open source 🎧 Nice work @ngxson! https://x.com/fdaudens/status/1891690883176604053

Spotify Opens Up Support for ElevenLabs Audiobook Content — Spotify https://newsroom.spotify.com/2025-02-20/spotify-opens-up-support-for-elevenlabs-audiobook-content/

“Learn everything there is to know about LlamaParse in this comprehensive video! In this video, @mesudarshan covers: ➡️ Multiple parsing modes ➡️ Using parsing instructions to improve quality ➡️ Output formats available ➡️ Parsing audio and images ➡️ JSON mode ➡️ Using it all in https://x.com/llama_index/status/1890499579214491967

stepfun-ai/Step-Audio-Chat · Hugging Face https://huggingface.co/stepfun-ai/Step-Audio-Chat

“chat, is this for real? A 132B parameter end to end Speech LM??? voice in, voice out 🤯 APACHE 2.0 LICENSED?? https://x.com/reach_vb/status/1891517368603492697

“Improvements will happen rapidly and almost daily according to the team. There is also a Grok-powered voice app coming too — about a week away!” / X https://x.com/omarsar0/status/1891715813956108699

“@WholeMarsBlog Voice mode is still a little patchy, so probably launches in about a week, but it’s awesome” / X https://x.com/elonmusk/status/1891676673898119254

“Step-Audio a 130 billion parameter multimodal LLM that is responsible for understanding and generating human speech https://x.com/_akhaliq/status/1891528348590833834

“Organizational life is about to get much weirder. This paper creates an early form of meeting delegates, where you send an AI to a meeting on your behalf, and it uses your voice and knowledge to advance your agenda A lot of old organizational methods need to be rethought for AI https://x.com/emollick/status/1891527817826828565

“No system card for Grok 3 yet, so no perspectives on risk mitigation. This is especially key for voice, and is why labs have been slow with full multimodal, you can imitate anyone’s voice, and also the AI tended to take your voice and repeat it back to you. From 4o system card: https://x.com/emollick/status/1891745345392058496

“Google presents: SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features Opensources model ckpts with four sizes from 86M to 1B https://x.com/arankomatsuzaki/status/1892777324715634971

“DeepMind just released new set of PaliGemma 2 Checkpoints tailored specifically for use like Optical Character Recognition, Captioning and more! 🔥 They come in sizes ranging from 3B to 28B – all open weights AND work with Transformers! 💥 https://x.com/reach_vb/status/1892269416831717648

“Great example of the jagged frontier- even smart AIs have weaknesses at basic tasks. Reading clocks is hard because AI vision systems are crude & calendar facts require vision & math The best model, Gemini, gets only 22% of clocks rights, while o1 gets 80% of calendar questions. https://x.com/emollick/status/1890460620262134176

“Google just released PaliGemma 2 Mix: new versatile instruction vision language models 🔥 > Three new models: 3B, 10B, 28B with res 224, 448 💙 > Can do vision language tasks with open-ended prompts, understand documents, and segment or detect anything 🤯 https://x.com/mervenoyann/status/1892267634923593760

Google Workspace Updates: Scroll through live captions and translated captions in Google Meet https://workspaceupdates.googleblog.com/2025/02/google-meet-caption-history.html

[2502.08391v1] ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification https://arxiv.org/abs/2502.08391v1

“Audiobox Aesthetics is a model for unified automatic quality assessment for speech, music and sound. Try the demo on @huggingface ➡️ https://x.com/AIatMeta/status/1893009390980170001

“AlphaMaze Enhancing Large Language Models’ Spatial Intelligence via GRPO https://x.com/_akhaliq/status/1892774257425264784

Announcing Gaze Detection. https://moondream.ai/blog/announcing-gaze-detection

“👀” / X https://x.com/fdaudens/status/1890410351507783941

“🔊 Meet Kokoro Web – Free, ML speech synthesis on your compturer, that’ll make you ditch paid services! 28 natural voices, unlimited generations, and WebGPU acceleration. Perfect for journalists and content creators. Test it with full articles—sounds amazingly human! 🎯🎙️” / X https://x.com/fdaudens/status/1890465947078590775

“SigLIP 2 blog from @mervenoyann and team to learn more and try it out: https://x.com/_philschmid/status/1892873071419113506

“we put out an update to chatgpt (4o). it is pretty good. it is soon going to get much better, team is cooking.” / X https://x.com/sama/status/1890816782836904000

Helix: A Vision-Language-Action Model for Generalist Humanoid Control https://www.figure.ai/news/helix

“One of the best Vision-Language Encoder got an Update! @GoogleDeepMind releases SigLIP 2! SigLIP 2 merges captioning pretraining, self-supervised learning, and online data curation, and outperforms its previous version in 10+ tasks, with support for flexible resolutions and https://x.com/_philschmid/status/1892869075266662632

“SigLIP 2 is new version of SigLIP: best open-source multimodal encoders by @GoogleDeepMind, now on @huggingface 🤗 What’s new? > Improvements from new masked loss, self-distillation and dense features (better localization) > Dynamic resolution with Naflex (better OCR) https://x.com/mervenoyann/status/1892869097227989071

“The challenge in creating omni-modal models is their performance gap compared to specialized single modality models. Existing omni-modal models also lack balanced performance and efficient training. This paper introduces Ola, an Omni-modal Language Model, using progressive https://x.com/rohanpaul_ai/status/1891408499898261902

“3/ AgentStack from @AgentOpsAI v0.3.4  – CLI command to create custom tools for any framework  – CLI command to undo AgentStack actions  – improvements to @firecrawl_dev – vision tools with @AnthropicAI  – SQL tool added to core tools @braelyn_ai https://x.com/AtomSilverman/status/1890534557817974900

“I like this, but found o1 Pro (which wasn’t tested) did surprisingly well on the sample puzzles (and I do wonder how much the limitations of vision play in). https://x.com/emollick/status/1890199820956315971

“Tool for generating vision AI code from prompts https://x.com/tom_doerr/status/1890064224619041073

“Your weekly recap of open AI is here, and it’s packed with models! 👀 Multimodal > @opengvlab released InternVideo 2.5 Chat models, new video LMs with long context > AIDC released Ovis2 model family along with Ovis dataset, new vision LMs in different sizes (1B, 2B, 4B, 8B, https://x.com/mervenoyann/status/1890484105311060329

“Building efficient Vision-Language Models without pre-trained vision encoders is difficult. This paper addresses the difficulty of learning visual perception from scratch within a unified model and minimizing vision-language interference. Proposes EVEv2.0. EVEv2.0 is an https://x.com/rohanpaul_ai/status/1892178341744054730

“ByteDance presents Phantom Subject-consistent video generation via cross-modal alignment https://x.com/_akhaliq/status/1892073250974216476

“On the heels of Humanity’s Last Exam, @scale_AI & @ai_risks have released a new very-hard reasoning eval: EnigmaEval: 1,184 multimodal puzzles so hard they take groups of humans many hours to days to solve. All top models score 0 on the Hard set, and <10% on the Normal set 🧵 https://x.com/alexandr_wang/status/1891208692751638939?s=46

“Is computer vision “solved”? Not yet Current models score 0% on ZeroBench 🧵1/6 https://x.com/jrobertsai/status/1891506671056261413?s=46

“Meta presents: Intuitive physics understanding emerges from self-supervised pretraining on natural videos Predicting outcomes in a rep space leads to physics understanding, unlike video prediction in pixel space and MLLMs, which reason through text. https://x.com/arankomatsuzaki/status/1891721882065391692

“Meta’s AI can now read your mind with 80% accuracy, no implants needed! Using non-invasive brain activity recordings (MEG & EEG) from 35 participants as they typed, Meta trained an AI to reconstruct sentences. This is life-changing for people with brain lesions. AI was https://x.com/Yuchenj_UW/status/1891560033193955795

One response to “Multimodal: AI News Week Ending 02/21/2025”

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading