Image created with OpenAI GPT-Image-1. Image prompt: rich crimson, bright ivory, deep navy Independence-Day palette, vibrant, celebratory, wholesome, authentic, photorealistic desert campfire under stars with tiny flags in sand scene featuring a row of world flags behind a large US flag centerpiece; natural lighting, subtle film grain, high detail

Apple just released a Sage Mixtral 8x7b fine-tune w/ Apache license on the hub 👀 Uses State-Action Chains (SAC) to enhance dialogue generation by incorporating latent variables for emotional states and conversational strategies. Key comparisons: > SAC vs. standard LM”” / X https://x.com/reach_vb/status/1939970610702028899

Huawei announces open-sourcing of Pangu models to accelerate AI application, value creation https://www.ecns.cn/news/sci-tech/2025-07-01/detail-ihesxvny3991876.shtml

🚀 Meet Qwen-TTS – now live via the Qwen API ! Trained on millions of hours of speech, it delivers ultra-natural, expressive audio with smart prosody, pacing, and emotion. 🗣️ Supports 3 Chinese dialects: Beijing, Shanghai, Sichuan 🎙️ 7 bilingual voices: Cherry, Ethan, Chelsie, https://x.com/Alibaba_Qwen/status/1939553252166836457

Tencent released Hunyuan-A13B, a new open-source hybrid reasoning model It nears or matches models like o1 and DeepSeek R1 on major benchmarks, while remaining efficient enough to run on a single GPU Also includes “”fast and slow”” modes to adjust efficiency levels https://x.com/rowancheung/status/1939601169271197973

The real star of the show is the (Baidu) 21B A3B, ~30% smaller than Qwen3 30B A3B and better on most benchmarks! 🔥 https://x.com/reach_vb/status/1939584854045466791

Meet Qwen-VLo, your AI creative engine: • Concept-to-Polish: Turn rough sketches or text prompts into high-res visuals • On-the-Fly Edits: Refine product shots, adjust layouts or styles with simple commands • Global-Ready: Generate image in multiple languages • Progressive https://x.com/Alibaba_Qwen/status/1938604105909600466

Meet Jan-nano, a 4B model that outscores DeepSeek-v3-671B using MCP. It’s built on Qwen3-4B with DAPO fine-tuning, it handles: – real-time web search – deep research Model + GGUF: https://x.com/menloresearch/status/1934809407604576559

Announcing the Open Source Release of the ERNIE 4.5 Model Family | ERNIE Blog https://yiyan.baidu.com/blog/posts/ernie4.5/

Baidu just released the weights for multiple ERNIE 4.5 variants including multimodal models https://x.com/scaling01/status/1939509144903422131

MASSIVE release from Baidu – Ernie 4.5 VLMs & LLMs, Models beat DeepSeek v3, Qwen 235B and competitive to OpenAI O1 (for VLM) – Apache 2.0 licensed 💥 https://x.com/reach_vb/status/1939569283111235645

The ERNIE 4.5 series is now officially open source. This family of models includes 10 variants—from MoE models with 47B and 3B active parameters, the largest having 424B total parameters, to a 0.3B dense model—all available now to the global AI community for open research and https://x.com/Baidu_Inc/status/1939724778157511126

Wait, that’s JUST a 3B multimodal model understanding and generation model AND apache 2.0 licensed 🔥 https://x.com/reach_vb/status/1939627598830559644

RT @togethercompute: Announcing DeepSWE 🤖: our fully open-sourced, SOTA software engineering agent trained purely with RL on top of Qwen3-3…”” / X https://x.com/tri_dao/status/1940765882227347585

Xiaomi’s first AI-powered eyewear brings smartphone firm into ‘war of hundreds of glasses’ | South China Morning Post https://www.scmp.com/tech/big-tech/article/3315917/xiaomis-first-ai-powered-eyewear-brings-smartphone-firm-war-hundreds-glasses

More humanoid robots from China! Beijing-based industrial robot maker ROKAE Robotics has unveiled two humanoid robot models. https://x.com/TheHumanoidHub/status/1937926026514088103

New video of CL-3 humanoid by Chinese company LimX Dynamics. https://x.com/TheHumanoidHub/status/1940431073827328010

AI avatars in China just proved they are ace influencers https://www.cnbc.com/2025/06/19/ai-humans-in-china-just-proved-they-are-better-influencers.html

Kyutai released their Streaming Text to Speech model, ~2B param model, ultra low latency (220ms), CC-BY-4.0 license 🔥 Trained on 2.5 Million Hours of audio, it can serve up to 32 users w/ less than 350ms latency on a SINGLE L40 🤯 Incredible release by kyutai folks, go check https://x.com/reach_vb/status/1940786922546249961

LLMs can synthesize many programs, but how should we search among them? New from @SakanaAILabs – AB-MCTS frames code generation as an adaptive tree search, guided by external feedback. Beats baselines on synthesis benchmarks including ARC-AGI.”” / X https://x.com/ndea/status/1940166177424384354

A bit later, but Huawei did open source their 72B MoE. Now they’re in the proud club «did Scout better than Meta». DeepSeek ethos spreads… Also, gitcode claims 15T tokens vs 13 in the tech report. Another detail: MoGE is their original load balancing solution. interesting https://x.com/teortaxesTex/status/1940341153754382688

Introducing Mistral Small 3.2, a small update to Mistral Small 3.1 to improve: – Instruction following: Small 3.2 is better at following precise instructions – Repetition errors: Small 3.2 produces less infinite generations or repetitive answers – Function calling: Small https://x.com/MistralAI/status/1936093325116781016

METR results for DeepSeek V3 and R1 kinda sucks, huh? https://x.com/scaling01/status/1939770925781487779

It’s fun that Deepseek is being served at lower latencies and the same cost by a number of companies. A few firms have implemented Deepseek’s high rank / wide EP inference set up with higher efficiency too. As such traffic has left Deepseek’s direct API to other providers”” / X https://x.com/dylan522p/status/1940872241753039319

RT @ivanfioravanti: DeepSeek-R1-0528-5bit on MLX pushing M3 Ultra 512GB to its limits! 501GB used mem visibile on mactop in the video! Con…”” / X https://x.com/awnihannun/status/1940067135054913892

RT @tngtech: Today we release DeepSeek-TNG R1T2 Chimera. This new Chimera is a Tri-Mind Assembly-of-Experts model with three parents, nam…”” / X https://x.com/swyx/status/1940660469733511388

Policy throttles silicon; DeepSeek’s R2 waits in the queue. Washington’s export clamp has dried up fresh H20 supply, so those chips cannot be freed for training the larger R2. Engineers completed an initial R2 pass, yet CEO Liang Wenfeng says reasoning and coding lag. Teams https://x.com/rohanpaul_ai/status/1939242685828927512

DAMN! DeepSeek R1T2 – 200% faster than R1-0528 & 20% faster than R1 🔥 Significantly better than R1 on GPQA & AIME 24 made via Assembly of Experts w/ DS V3, R1 & R1-0528 MIT license – available on Hugging Face 🤗 https://x.com/reach_vb/status/1940536684061643239

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading