Image created with OpenAI GPT-Image-1. Image prompt: rich crimson, bright ivory, deep navy Independence-Day palette, vibrant, celebratory, wholesome, authentic, photorealistic desert campfire under stars with tiny flags in sand scene featuring a row of world flags behind a large US flag centerpiece; natural lighting, subtle film grain, high detail
Apple just released a Sage Mixtral 8x7b fine-tune w/ Apache license on the hub 👀 Uses State-Action Chains (SAC) to enhance dialogue generation by incorporating latent variables for emotional states and conversational strategies. Key comparisons: > SAC vs. standard LM”” / X https://x.com/reach_vb/status/1939970610702028899
Huawei announces open-sourcing of Pangu models to accelerate AI application, value creation https://www.ecns.cn/news/sci-tech/2025-07-01/detail-ihesxvny3991876.shtml
🚀 Meet Qwen-TTS – now live via the Qwen API ! Trained on millions of hours of speech, it delivers ultra-natural, expressive audio with smart prosody, pacing, and emotion. 🗣️ Supports 3 Chinese dialects: Beijing, Shanghai, Sichuan 🎙️ 7 bilingual voices: Cherry, Ethan, Chelsie, https://x.com/Alibaba_Qwen/status/1939553252166836457
Tencent released Hunyuan-A13B, a new open-source hybrid reasoning model It nears or matches models like o1 and DeepSeek R1 on major benchmarks, while remaining efficient enough to run on a single GPU Also includes “”fast and slow”” modes to adjust efficiency levels https://x.com/rowancheung/status/1939601169271197973
The real star of the show is the (Baidu) 21B A3B, ~30% smaller than Qwen3 30B A3B and better on most benchmarks! 🔥 https://x.com/reach_vb/status/1939584854045466791
Meet Qwen-VLo, your AI creative engine: • Concept-to-Polish: Turn rough sketches or text prompts into high-res visuals • On-the-Fly Edits: Refine product shots, adjust layouts or styles with simple commands • Global-Ready: Generate image in multiple languages • Progressive https://x.com/Alibaba_Qwen/status/1938604105909600466
Meet Jan-nano, a 4B model that outscores DeepSeek-v3-671B using MCP. It’s built on Qwen3-4B with DAPO fine-tuning, it handles: – real-time web search – deep research Model + GGUF: https://x.com/menloresearch/status/1934809407604576559
Announcing the Open Source Release of the ERNIE 4.5 Model Family | ERNIE Blog https://yiyan.baidu.com/blog/posts/ernie4.5/
Baidu just released the weights for multiple ERNIE 4.5 variants including multimodal models https://x.com/scaling01/status/1939509144903422131
MASSIVE release from Baidu – Ernie 4.5 VLMs & LLMs, Models beat DeepSeek v3, Qwen 235B and competitive to OpenAI O1 (for VLM) – Apache 2.0 licensed 💥 https://x.com/reach_vb/status/1939569283111235645
The ERNIE 4.5 series is now officially open source. This family of models includes 10 variants—from MoE models with 47B and 3B active parameters, the largest having 424B total parameters, to a 0.3B dense model—all available now to the global AI community for open research and https://x.com/Baidu_Inc/status/1939724778157511126
Wait, that’s JUST a 3B multimodal model understanding and generation model AND apache 2.0 licensed 🔥 https://x.com/reach_vb/status/1939627598830559644
RT @togethercompute: Announcing DeepSWE 🤖: our fully open-sourced, SOTA software engineering agent trained purely with RL on top of Qwen3-3…”” / X https://x.com/tri_dao/status/1940765882227347585
Xiaomi’s first AI-powered eyewear brings smartphone firm into ‘war of hundreds of glasses’ | South China Morning Post https://www.scmp.com/tech/big-tech/article/3315917/xiaomis-first-ai-powered-eyewear-brings-smartphone-firm-war-hundreds-glasses
More humanoid robots from China! Beijing-based industrial robot maker ROKAE Robotics has unveiled two humanoid robot models. https://x.com/TheHumanoidHub/status/1937926026514088103
New video of CL-3 humanoid by Chinese company LimX Dynamics. https://x.com/TheHumanoidHub/status/1940431073827328010
AI avatars in China just proved they are ace influencers https://www.cnbc.com/2025/06/19/ai-humans-in-china-just-proved-they-are-better-influencers.html
Kyutai released their Streaming Text to Speech model, ~2B param model, ultra low latency (220ms), CC-BY-4.0 license 🔥 Trained on 2.5 Million Hours of audio, it can serve up to 32 users w/ less than 350ms latency on a SINGLE L40 🤯 Incredible release by kyutai folks, go check https://x.com/reach_vb/status/1940786922546249961
LLMs can synthesize many programs, but how should we search among them? New from @SakanaAILabs – AB-MCTS frames code generation as an adaptive tree search, guided by external feedback. Beats baselines on synthesis benchmarks including ARC-AGI.”” / X https://x.com/ndea/status/1940166177424384354
A bit later, but Huawei did open source their 72B MoE. Now they’re in the proud club «did Scout better than Meta». DeepSeek ethos spreads… Also, gitcode claims 15T tokens vs 13 in the tech report. Another detail: MoGE is their original load balancing solution. interesting https://x.com/teortaxesTex/status/1940341153754382688
Introducing Mistral Small 3.2, a small update to Mistral Small 3.1 to improve: – Instruction following: Small 3.2 is better at following precise instructions – Repetition errors: Small 3.2 produces less infinite generations or repetitive answers – Function calling: Small https://x.com/MistralAI/status/1936093325116781016
METR results for DeepSeek V3 and R1 kinda sucks, huh? https://x.com/scaling01/status/1939770925781487779
It’s fun that Deepseek is being served at lower latencies and the same cost by a number of companies. A few firms have implemented Deepseek’s high rank / wide EP inference set up with higher efficiency too. As such traffic has left Deepseek’s direct API to other providers”” / X https://x.com/dylan522p/status/1940872241753039319
RT @ivanfioravanti: DeepSeek-R1-0528-5bit on MLX pushing M3 Ultra 512GB to its limits! 501GB used mem visibile on mactop in the video! Con…”” / X https://x.com/awnihannun/status/1940067135054913892
RT @tngtech: Today we release DeepSeek-TNG R1T2 Chimera. This new Chimera is a Tri-Mind Assembly-of-Experts model with three parents, nam…”” / X https://x.com/swyx/status/1940660469733511388
Policy throttles silicon; DeepSeek’s R2 waits in the queue. Washington’s export clamp has dried up fresh H20 supply, so those chips cannot be freed for training the larger R2. Engineers completed an initial R2 pass, yet CEO Liang Wenfeng says reasoning and coding lag. Teams https://x.com/rohanpaul_ai/status/1939242685828927512
DAMN! DeepSeek R1T2 – 200% faster than R1-0528 & 20% faster than R1 🔥 Significantly better than R1 on GPQA & AIME 24 made via Assembly of Experts w/ DS V3, R1 & R1-0528 MIT license – available on Hugging Face 🤗 https://x.com/reach_vb/status/1940536684061643239




