Image created with Flux Pro v1.1 Ultra. Image prompt: photorealistic still image of a middle-aged man standing behind a woman, woman covering part of her face with her hand, man looking over her shoulder, both illuminated with warm stadium jumbotron lighting, natural skin tones, subtle lens flare, shallow depth of field, exact color temperature of a live event projection, both wearing shirts with the open padlock open-source symbol, cinematic realism –no text, captions, watermarks

After DeepSeek R1, there’s new Claude 4 level model from China that outperforms DeepSeek v3, Qwen and OpenAI GPT-4.1 Meet Kimi k2 – 1 trillion parameter model purpose-built for agentic workflows with native MCP integration. 100% Opensource and FREE to try. Let that sink in. https://x.com/Saboo_Shubham_/status/1943694224584818808

RT @carlothinks: Ex-Alibaba CTO just made the boldest claim about AI & global power: “China is building the future of AI, not Silicon Vall…”” / X https://x.com/glennko/status/1950642750916792580

There is now a path for China to surpass the U.S. in AI. Even though the U.S. is still ahead, China has tremendous momentum with its vibrant open-weights model ecosystem and aggressive moves in semiconductor design and manufacturing. In the startup world, we know momentum”” / X https://x.com/AndrewYNg/status/1950941108000964654

This week’s letter from Andrew Ng in The Batch asks a blunt question: Can surging performance from China’s open-weights models and home-grown chips let it overtake the U.S. in AI? He lays out the data behind China’s momentum, explains why Washington’s new action plan is helpful”” / X https://x.com/DeepLearningAI/status/1951354901843288546

RT @OpenRouterAI: Qwen3 Coder has now passed Grok 4 in the Programming prompt rankings Tied with Kimi! https://x.com/huybery/status/1949270432567460309

RT @Alibaba_Wan: 🚀 Introducing Wan2.2: The World’s First Open-Source MoE-Architecture Video Generation Model with Cinematic Control! 🔥 Key…”” / X https://x.com/ClementDelangue/status/1949832988834873417

Pierre and team really cooked with this vision language model (VLM)! Excited for you to try it out! 111B open parameters”” / X https://x.com/JayAlammar/status/1950931480349143259

RT @1vnzh: Command A Vision – SOTA enterprisemaxx multimodal model – Outperforms GPT 4.1, Llama 4 Maverick, and Mistral Medium 3 in enterpr…”” / X https://x.com/aidangomez/status/1950927454383616343

RT @nickfrosst: cohere vision model 🙂 weights on huggingface https://x.com/andrew_n_carr/status/1951068402090647608

And @Microsoft!”” / X https://x.com/Yoshua_Bengio/status/1951270687957553235

EU AI Act: General-Purpose AI Code of Practice · Final Version
https://code-of-practice.ai/?section=safety-security

Yoshua Bengio on X: “I’ve been thrilled to see the support for the Safety & Security Chapter of the Code of Practice. Most frontier AI companies have now signed on to it: @AnthropicAI, @Google, @MistralAI, @OpenAI, @xAI Why this is important: 🧵 1/6″ / X
https://x.com/Yoshua_Bengio/status/1951263044056588677

Releasing Open Weights for FLUX.1 Krea https://www.krea.ai/blog/flux-krea-open-source-release

RT @bfl_ml: Today we are releasing FLUX.1 Krea [dev] – a new state-of-the-art open-weights FLUX model, built for photorealism. Developed…”” / X https://x.com/multimodalart/status/1950923544998658557

Suddenly there are tons more weird LLM arena models – cuttlefish, kraken, etc. I just hope we are not going to see a repeat of the Llama 4 incident, where different versions of the same model are being tuned to max out the arena score https://x.com/emollick/status/1949671630390665231

Lisan al Gaib on X: “horizon-alpha is by OpenAI, but extremely weak on LisanBench, even though it had 5 trials per word instead of just 1 it gets beaten by qwen3-30b-a3b you can tell it’s a very small model https://t.co/Lv6mrNbB9k” / X
https://x.com/scaling01/status/1950730582104604964

not even remotely close to o3-mini level”” / X https://x.com/scaling01/status/1950730792251891948

RT @Teknium1: Looks like OpenAI’s been using Nous’ YaRN and kaiokendev’s rope scaling for context length extension all along – of course ne…”” / X https://x.com/jeremyphoward/status/1951368366943510739

Especially notable given Zuckerberg’s note that Meta will not necessarily open source future models. US companies are still doing great small open models, but, aside from whatever OpenAI releases, it appears that frontier open weights will mean Chinese models (& maybe Mistral).”” / X https://x.com/emollick/status/1950610040945004957

I’m now using Qwen3-Coder in Claude Code. Works with any model actually, but this is surely the best one currently. There are a bunch of proxies on GitHub that make this possible, but none worked well enough for me, so I implemented this myself using LiteLLM. Guide in comments: https://x.com/WolframRvnwlf/status/1948046368213176624

Kimi K2 – On-par with Claude 4, but 80% cheaper!! I connected Kimi K2 to Claude Code to get a sense of real performance (Kimi Code!) Overall findings: 1. Exceptional coding capability 2. Cost only 20% of Claude 4 (Huge!) 2. Only downside is API is a bit slow 🧵 Below is some https://x.com/jasonzhou1993/status/1944320164889284947

Qwen3 Coder has a 5.32% diff edit failure rate in Cline, based on real-world data. It’s right alongside Claude Sonnet 4 and Kimi K2 as an excellent model in performing diff edits. Not bad for an open-source model at $0.30/$1.20 per million input/output tokens 👀 https://x.com/cline/status/1949973297455599998

Introducing FlowMaker 🌊🤖 A fully open-source, low-code way of building custom agent workflows. Build agents via a drag and drop interface, run it directly in the app, and also directly export it into a deeply custom workflow backed by @llama_index.TS. It’s a fantastic visual https://x.com/jerryjliu0/status/1948797112789205111

Deep agents with qwen3-coder!”” / X https://x.com/hwchase17/status/1951072092625240203

The small-sized Qwen3-Coder is here! This is a little gift prepared for local users🎁, it’s extremely fast, and it has basic agentic coding capabilities! Flash Coding!”” / X https://x.com/huybery/status/1950925963979796877

These new open source models (GLM, Kimi) continue to be odd. Great stats, some solid performances, but also fail tests that DeepSeek & smaller closed models have beaten for months. https://x.com/emollick/status/1949844122119840084

What started as an open source experiment to push models to their limits became something unexpected. 2.7M developers. Inbound from Fortune 100 companies. A $32M bet on the future of coding. The story behind it all, and why our long-term bet is on open-source: https://x.com/cline/status/1951005843417358427

(Since I am on a benchmark theme today) The ARC team does well keeping AI labs honest about their benchmarks, including showing that Qwen’s big ARC-AGI performance doesn’t replicate But ARC-AGI also has a strong philosophy of what AI should do. We need other benchmarking efforts”” / X https://x.com/emollick/status/1948476524027404733

RT @ggerganov: AMD teams contributing to the llama.cpp codebase. Great support from the community with the review process. Exciting to see…”” / X https://x.com/ggerganov/status/1950047168280060125

Step3 benchmarks at last. The first «DeepSeek-like» that’s strongly multimodal (Ernie disappointed). It’s very different from V3, too – another in-house attention, the logic around inference economics. A big release. https://x.com/teortaxesTex/status/1951008169989382218

RT @abidlabs: Why should you pay for an experiment tracking library? Excited to introduce a new 💯open-source library from @HuggingFace: Tr…”” / X https://x.com/_akhaliq/status/1950617338136383605

Official release of Wan2.2 an open-source text-to-video and image-to-video model! https://x.com/scaling01/status/1949828474878746924

AMD teams contributing to the llama.cpp codebase. Great support from the community with the review process. Exciting to see this open-source collaboration!”” / X https://x.com/ggerganov/status/1949907603942691027

Llama-8b 1.2M sequence length training is now possible on a 1x H200 gpu with ALST + FA3 + Liger-Kernel. That’s 2.4x longer than with 1x H100. Ready to run recipes: https://x.com/StasBekman/status/1950232169227624751

ollama run qwen3-coder”” / X https://x.com/ollama/status/1951147035895480356

ollama run qwen3:30b https://x.com/ollama/status/1950291777216262259

horizon-alpha is by OpenAI, but extremely weak on LisanBench, even though it had 5 trials per word instead of just 1 it gets beaten by qwen3-30b-a3b you can tell it’s a very small model https://x.com/scaling01/status/1950730582104604964

RT @MistralAI: In our continued commitment to open-science, we are releasing the Voxtral Technical Report: https://x.com/GuillaumeLample/status/1950855212677075122

StepFun open sources some of their inference infra https://x.com/teortaxesTex/status/1950127131754651655

There still does not exist an open model that consistently beats R1-0528 on hard coding. We’re on a plateau. All these models used similar compute, have similar active param count, near-identical architecture, and probably are trained on similar data. New ideas needed. https://x.com/teortaxesTex/status/1951200161805312297

GLM-4.5 models are the latest addition to chinese frontier open-source models Blog: https://x.com/scaling01/status/1949825490488795275

If your org has a policy against using open weights models from China you’re at a significant competitive disadvantage.”” / X https://x.com/corbtt/status/1950334347971874943

.@JanLiphardt is a Professor of Bioengineering at Stanford and the founder of OpenMind. Jan is building an open, decentralized ecosystem for real-world robotics and AI. We took Iris (the robot) for a walk across the street from their office and had a chat about the vision behind https://x.com/TheHumanoidHub/status/1948814860219003316

China dropped these open-source models in July: – GLM-4.5 – GLM-4.5-Air – Wan-2.2 – Qwen3 Coder – Qwen3-235B-A22B-Thinking-2507 – Qwen3-235B-A22B-2507 – Kimi K2 Meanwhile: – OpenAI still hasn’t released the open-source model – Anthropic is doing stricter rate limits – Meta might”” / X https://x.com/Yuchenj_UW/status/1950034092457939072

For anyone in AI, coding or research: this model is insane! Solving complex math and 256K-context easily: Alibaba dropped their most advanced AI brain yet: Qwen3 235B A22B Thinking 2507. It’s built for pure reasoning power. Here’s why this is wow: ✅ Solves harder problems. https://x.com/IlirAliu_/status/1949156307950309686

RT @AlibabaGroup: Qwen3 is recognized as the #1 open model in the Arena🏆! Ranked #3 overall and #1 in Coding, Hard Prompts, and Math – a re…”” / X https://x.com/lmarena_ai/status/1951328014140129551

Qwen/Qwen3-235B-A22B-Thinking-2507 · Hugging Face https://huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507

GSPO is the most impressive Alibaba Qwen research paper to date, I think. They’ve started publishing strong stuff just months ago. New 235B-Thinking truly is comparable to R1-0528. I am more optimistic about Qwen than ever.”” / X https://x.com/teortaxesTex/status/1949601984308207781

The GSPO paper by @Alibaba_Qwen is already the third most popular one on @huggingface for the month of July. I suspect this will have a massive impact on the field! https://x.com/ClementDelangue/status/1949934196148895799

If you’re a researcher or engineer releasing open science papers & open models and datasets, I bow to you 🙇🙇🙇 From what I’m hearing, doing so, especially in US big tech, often means fighting your manager and colleagues, going through countless legal meetings, threatening to”” / X https://x.com/ClementDelangue/status/1950927952641749194

supervision, the open-source library I created 2 years ago, is crossing 30,000 stars on GitHub! thank you to everyone who helped me build this project! it took us 4,000+ commits, 1,000+ PRs and 100+ contributors to do it. link: https://x.com/skalskip92/status/1949857474862866659

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading