Image created with Flux Pro v1.1 Ultra. Image prompt: Multimodality, trio of motifs in small bananas: page icon (text), waveform (audio), camera frame (image), evenly spaced, photorealistic, editorial, minimal, high detail, 3:2 landscape
UI-TARS-2 Technical Report Advancing GUI Agent with Multi-Turn Reinforcement Learning https://x.com/_akhaliq/status/1963229296236937443
Can AI agents reliably navigate the web? Does the choice of agent scaffold affect web browsing ability? To answer these questions, we added Online Mind2Web, a web browsing benchmark, to the Holistic Agent Leaderboard (HAL). We evaluated 9 models (including GPT-5 and Sonnet 4) https://x.com/sayashk/status/1963343022252315112
We can finally share UI-TARS-2🥳🥳 — a native GUI agent trained with multi-turn agent RL ⚡️⚡️Key highlights (all-in-one model!): 💻Computer Use: 47.5 OSWorld · 50.6 WindowsAgentArena 📱Phone Use: 73.3 AndroidWorld 🛜Browser Use: 88.2% Online-Mind2Web 🎮Gameplay: ~60% human https://x.com/TsingYoga/status/1963629621326614940
🏓🤖 Our humanoid robot can now rally over 100 consecutive shots against a human in real table tennis — fully autonomous, sub-second reaction, human-like strikes. https://x.com/ZhiSu22/status/1961244573658673222
HITTER https://humanoid-table-tennis.github.io/
Humanoid robots playing table tennis fully autonomously. The ‘HITTER’ system combines a model-based planner with a reinforcement learning (RL) whole-body controller. It is fully autonomous but relies on an external sensing system. A 9-camera OptiTrack motion capture setup https://x.com/TheHumanoidHub/status/1961338417628979237
Google’s on a roll. That’s a lot of performance for that tiny size! I just embedded 1.4 million documents in ~80 mins on my M2 Max for free. Would’ve been ~$200 with the text-embedding-3-large, with worse quality.”” / X https://x.com/rishdotblog/status/1963805087014502497
Duolingo is facing an existential crisis as Google Translate rolls out features to tutor users—and even handle live translation as a bonus | Fortune https://fortune.com/2025/08/27/duolingo-existential-crisis-ai-google-translate-language-learning-live-translation/
If you think Apple is not doing much in AI, you’re getting blindsided by the chatbot hype and not paying enough attention! They just released FastVLM and MobileCLIP2 on Huggingface. The models are up to 85x faster and 3.4x smaller than previous work, enabling real-time vision language model (VLM) applications! It can even do live video captioning 100% locally in your browser 🤯🤯🤯 https://x.com/ClementDelangue/status/1962526559115358645
AI stethoscope can detect three heart conditions in 15 seconds – BHF https://www.bhf.org.uk/what-we-do/news-from-the-bhf/news-archive/2025/august/ai-stethoscope-can-detect-three-heart-conditions-in-15-seconds
🗺️ ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling”” TL;DR: high-fidelity 3D humans across a wide range of poses, capturing both skeletal structure and surface details; separates internal skeleton from the external surface, (1/3) https://x.com/Almorgand/status/1962581481055797586
We connect the autoregressive pipeline of LLMs with streaming video perception. Introducing AUSM: Autoregressive Universal Video Segmentation Model. A step toward unified, scalable video perception — inspired by how LLMs unified NLP. 📝 https://x.com/miran_heo/status/1962649613590302776
Notebook LM Rolling out NEW audio overview formats:
(Default) Deep Dive: a thorough examination of your sources
Brief: 1-2 minute, bite-sized overviews
Critique: an expert review, offering constructive feedback on your material
Debate: a thoughtful debate between two hosts https://x.com/NotebookLM/status/1962949985546187120
Really excited about this new AI research that’s pushing the boundaries of what’s currently possible in astrophysics. 🌌”” / X https://x.com/sundarpichai/status/1963668228481159371
Using AI to advance our understanding of fundamental physics is the dream. Excited to see our latest AI model ‘Deep Loop Shaping’ help @LIGO and @Caltech detect the gravitational waves of intermediate-mass black holes better! Published in @ScienceMagazine”” / X https://x.com/demishassabis/status/1963795824854335528
We’re helping to unlock the mysteries of the universe with AI. 🌌 Our novel Deep Loop Shaping method published in @ScienceMagazine could help astronomers observe more events like collisions and mergers of black holes in greater detail, and gather more data about rare space https://x.com/GoogleDeepMind/status/1963664018515849285
Ha. The fact that this works is just great. https://x.com/emollick/status/1961304185661518161
Free AI Sound Effect Generator | Add Sound Effects to Video & Audio | ElevenLabs https://elevenlabs.io/sound-effects
Introducing Lovable Voice Mode Turn your ideas into reality without touching your keyboard. https://x.com/lovable_dev/status/1963255845900484632
Vibevoice from @MicrosoftAI is #1 trending on HF for the past few days! This is Frontier Open-Source Text-to-Speech Model VibeVoice designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. A core innovation of VibeVoice https://x.com/ClementDelangue/status/1963537036616323388
~4 months ago, we introduced OpenVision — a fully open, cost-effective family of vision encoders that rival OpenAI’s CLIP and Google’s SigLIP. Today, we’re back with a major update: OpenVision 2 https://x.com/cihangxie/status/1963297223753494832
Researchers are teaching robots to walk on Mars from the sand of New Mexico | Newsroom | Oregon State University https://news.oregonstate.edu/news/researchers-are-teaching-robots-walk-mars-sand-new-mexico
Interested in building and benchmarking deep research systems? Excited to introduce DeepScholar-Bench, a live benchmark for generative research synthesis, from our team at Stanford and Berkeley! 🏆Live Leaderboard https://x.com/lianapatel_/status/1961487232331911651
What if we could use the entire planet’s atmosphere as a sensor? DARPA: “let’s blow stuff up in New Mexico and see what happens.” *detonates test bombs* DARPA: “wait… why are we detecting a SpaceX rocket launch 1000+ miles away?” it worked TOO well. AtmoSense has cracked https://x.com/bilawalsidhu/status/1960905069638877536
AI co-pilot boosts noninvasive brain-computer interface by interpreting user intent | EurekAlert! https://www.eurekalert.org/news-releases/1096148
Smart ring rivalry heats up: Ultrahuman sues Oura over patent claims https://www.msn.com/en-us/money/other/smart-ring-rivalry-heats-up-ultrahuman-sues-oura-over-patent-claims/ar-AA1L2RLo?apiversion=v2&noservercache=1&domshim=1&renderwebcomponents=1&wcseo=1&batchservertelemetry=1&noservertelemetry=1
this repo is wild. 100+ production-ready AI Agents, RAG, Multi-Agent teams, Voice Agents, MCP, and LLM apps with step-by-step tutorials. 100% free and open source by @Saboo_Shubham_ 👏 link in next post 👀 https://x.com/MakerThrive/status/1962661273335742780
90% success rate in unseen environments. No new data, no fine-tuning. Autonomously. Most robots need retraining to work in new places. What if they didn’t? Robot Utility Models (RUMs) learn once and work anywhere… zero-shot. A team from NYU and Hello Robot built a set of https://x.com/IlirAliu_/status/1961692920836215229
A robot that sees the terrain and predicts its own future… up to 5 seconds ahead? This is real. ❗️Best Systems Paper finalist at #RSS2025 The team introduces a perceptive Forward Dynamics Model that helps legged robots safely navigate rough, complex environments: no manual https://x.com/IlirAliu_/status/1962569938805141861
People ask how I get such clean 3D scans with a DSLR — and I must admit it’s a bit of a dark art. But this new $5K PortalCam changes everything. LiDAR precision + SLAM speed + 3D Gaussian Splat fidelity in a device anyone can use. The use cases are wild. Let me show you 🧵 https://x.com/bilawalsidhu/status/1963337887027707987
AHELM: A Holistic Evaluation of Audio-Language Models “”we introduce AHELM, a benchmark that aggregates various datasets — including 2 new synthetic audio-text datasets called PARADE, which evaluates the ALMs on avoiding stereotypes, and CoRe-Bench, which measures reasoning over https://x.com/iScienceLuvr/status/1962799344001917360
In less than a day, @StepFun_ai dropped Step-Audio 2 Mini – 8B speech to speech, beats GPT-4o-Audio, Apache 2.0 licensed 🔥 > Trained on 8M+ hours, supports 50K+ voices, benchmarks for expressive/grounded speech 🤯 > Expressive and emotionally aware generation > Retrieves and https://x.com/reach_vb/status/1961414067668558319
I have a new terrible test of AI ability: “”create an execute the most annoying functional CAPTCHA in the world. Really go all out”” First off, Gemini 2.5 Pro Deep Think: Got the assignment, actually funny. Love the first line. I pasted the SVG it created below the image. https://x.com/emollick/status/1961648878286946329
Now you can grep a PDF (and any document) Introducing SemTools – simple parsing and semantic search for the command line https://x.com/LoganMarkewich/status/1961448960184520945
Introducing EmbeddingGemma🎉 🔥With only 308M params, this is the top open model under 500M 🌏Trained on 100+ languages 🪆Flexible embeddings (768 to 128 dims) with Matryoshka 🤗Works with your favorite open tools 🤏Runs with as little as 200MB https://x.com/osanseviero/status/1963635281032040914
Excited to share our first @MicrosoftAI in-house models: MAI-Voice-1 and MAI-1-preview. Details and how you can test below, with lots more to come⬇️ https://x.com/mustafasuleyman/status/1961111770422186452
8. VibeVoice It’s a long-form speech synthesis model using next-token diffusion for continuous data generation. – Generates up to 90 minutes of speech involving 4 speakers in a 64K token window, delivering high-fidelity, multi-speaker dialogue synthesis. – A novel tokenizer https://x.com/TheTuringPost/status/1962850737777684595
Researchers pioneer optical generative models, ushering in a new era of sustainable generative AI https://phys.org/news/2025-08-optical-generative-ushering-era-sustainable.html
Computer vision (understanding reality) 🤝 Computer graphics (generating reality) https://x.com/bilawalsidhu/status/1962517172267384853
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation, presented in the paper Check out the model here: https://x.com/reach_vb/status/1961414145938485477
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation, presented in the paper Play with the demo here: https://x.com/reach_vb/status/1961471503267979699
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation, presented in the paper try it here: https://x.com/_akhaliq/status/1962644559868883310
FineVision is out! A massive open-source dataset by @huggingface for training Vision-Language Models: – 17.3M images – 24.3M samples – 88.9M turns – 9.5B answer tokens This is the inaugural article using our new scientific publishing template! https://x.com/thibaudfrere/status/1963627540544647177
Fuck it. Today, we open source FineVision: the finest curation of datasets for VLMs, over 200 sources! > 20% improvement across 10 benchmarks > 17M unique images > 10B answer tokens > New capabilities: GUI navigation, pointing, counting FineVision 10x’s open-source VLMs. https://x.com/andimarafioti/status/1963610118165000479
Today, we are releasing FineVision, a huge open-source dataset for training state-of-the-art Vision-Language Models: > 17.3M images > 24.3M samples > 88.9M turns > 9.5B answer tokens Here are my favourite findings: https://x.com/lusxvr/status/1963609337546293448
Introducing POINTS-Reader — a vision-language model for end-to-end document conversion, delivering SOTA performance on OmniDocBench with blazing-fast throughput. https://x.com/ZhihuFrontier/status/1963192346222432750
Introducing Command A Translate, our state-of-the-art model designed for high-quality translation tasks. https://x.com/cohere/status/1961081779789447519
vLLM now supports Kwai Keye-VL-1.5!
With sharper video 📹 & image 🖼️ comprehension, stronger reasoning, and an extended 128K context length, this model unlocks richer conversations and more complex tasks than ever before. https://x.com/vllm_project/status/1962509793345859666
we present R-4B, a multimodal large language model designed for general-purpose auto-thinking, autonomously switching between step-by-step thinking and direct response generation based on task complexity. https://x.com/mervenoyann/status/1962917670786937135
best small vision LM with reasoning has dropped on @huggingface 🔥 Tencent dropped R-4B, small vision LM that claims sota with Apache 2.0 license 💗 the model enables different thinking options and transformers support through custom code! https://x.com/mervenoyann/status/1962917635932229797
Tencent open sources two high-performing translation models https://the-decoder.com/tencent-open-sources-two-high-performing-translation-models/
Tencent released Hunyuan-MT-7B and Hunyuan-MT-Chimera The Hunyuan Translation Model comprises a translation model, Hunyuan-MT-7B, and an ensemble model, Hunyuan-MT-Chimera. The translation model is used to translate source text into the target language, while the ensemble model https://x.com/_akhaliq/status/1962644501605835140
Apertus: a fully open, transparent, multilingual language model | ETH Zurich https://ethz.ch/en/news-and-events/eth-news/news/2025/09/press-release-apertus-a-fully-open-transparent-multilingual-language-model.html
ROSE: Remove Objects with Side Effects in Videos”” TL;DR: Diffusion transformer; five common cases: shadows, reflections, light, translucency and mirror as video side effects to remove https://x.com/Almorgand/status/1962846321372471755
HunyuanWorld-Voyager is here and fully open-source! The world’s first ultra-long-range world model with native 3D reconstruction, redefining AI-driven spatial intelligence for VR, gaming, and simulations. ✅Direct 3D Output: Exports point cloud videos to 3D formats without tools https://x.com/TencentHunyuan/status/1962741518797836708
Introducing gpt-realtime — our best speech-to-speech model for developers, and updates to the Realtime API https://x.com/OpenAI/status/1961110295486808394
Last week we released @OpenAIDevs gpt-realtime our latest speech-to-speech model. Prompting S2S models can be extremely powerful. To get the most out of the model we compiled a series of prompting tips of what we saw work with early customers. To present them here’s Cedar our https://x.com/dkundel/status/1962916750632353826
MiniCPM-V 4.5 achieves an average score of 77.0 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT-4o-latest, Gemini-2.0 Pro, and strong open-source models like Qwen2.5-VL 72B powered https://x.com/_akhaliq/status/1963587749400727980
abs: https://x.com/iScienceLuvr/status/1962800402409365590
code: https://x.com/iScienceLuvr/status/1962798182964113547
website: https://x.com/iScienceLuvr/status/1962799346292007272




