Image created with gemini-3.1-flash-image-preview. Image prompt: 1960s Pink Panther cartoon cel of the pink panther tiptoeing across a flat lime-green background, casually tossing oversized rolled blueprint scrolls marked with a bold black code-bracket symbol, an open empty padlocked briefcase left behind him, loose confident black ink outlines and minimalist flat composition with generous negative space, large playful hand-lettered title text reading ‘Open Source’ at the top.

BREAKING – OFFICIAL RESULTS: GPT-5.6 Sol by @OpenAI is 1st overall on Design Arena with an Elo of 1353. This puts GPT-5.6 Sol above Claude Fable 5 by @AnthropicAI and in the same performance band as GLM 5.2 by @Zai_org on frontend design. This is an 18-position and 60-point Elo”
https://x.com/DesignArena/status/2076391367446860249

ByteDance just released UniVR-34B on Hugging Face The first model to learn complex reasoning, physical dynamics, and long-term planning directly from visual demonstrations — no text chains needed.”
https://x.com/HuggingPapers/status/2076513044340097501

DeepSeek reportedly in talks to raise $1.5B, then IPO | TechCrunch
https://techcrunch.com/2026/07/14/deepseek-reportedly-in-talks-to-raise-1-5b-then-ipo/

I assume Google escapes this trap, but this is what happened to Meta with Llama 4 and xAI post Grok 4. Only company to have escaped the “disappointing next giant model trap” without a major setback to their lead was OpenAI, with Orion/GPT-4.5.”
https://x.com/emollick/status/2077849021150888408

👀 This 7B Model Beat a 72B One by Learning Where to Look Zhihu contributor Hsing @onehsing shared the story behind OmniAgent, an ICML 2026 paper developed by researchers from CUHK, SJTU, NTU, and Alibaba’s Qwen team. On LVBench, OmniAgent-7B scored 50.5, beating Qwen2.5-VL-72B”
https://x.com/ZhihuFrontier/status/2076962763394695225

OpenClaw v2026.7.1 is out, with 3,063 contributions from 532 contributors! • Web UI and onboarding overhauls • Major iOS, Android and MacOS app work • Telegram, Slack, Discord, and iMessage improvements • GPT-5.6, Muse Spark 1.1, and MORE”
https://x.com/openclaw/status/2076900503259414944

The weekly ClawCast is going live in Discord at 11:30am PT! Join us to hear from @sodio on why Openclaw might be the most misunderstood open source project. Join discord!”
https://x.com/openclaw/status/2077455353730756992

Streaming some open source in a few! 🙌”
https://x.com/steipete/status/2077088803500818682

Newest AI models to explore ↓ Frontier / Commercial GPT-5.6 (Sol, Terra, Luna) GPT-Live Claude Fable 5 Claude Mythos 5 Muse Spark 1.1 Muse Image Muse Video Grok 4.5 Research / Open Gemma 4 InternVLA-A1.5 NVIDIA Audex SenseNova-Vision RynnWorld-4D Vidu S1 AlayaWorld”
https://x.com/TheTuringPost/status/2077203833999229056

My benchmark where I have AIs create one file procedurally-generated harbor towns through history in one shot now has GPT-5.6 Pro, Fable, Kimi K3, and Inkling. You can play with all the simulations:
https://t.co/Bsby6xoiEF I think they are surprisingly indicative.”
https://x.com/emollick/status/2077840214223982975

Post-Kimi K3 and open weights models getting closer to the frontier again, I wonder if Anthropic and OpenAI will be allowed to increase their release cadence by the government. Mythos came out in April (before Opus 4.7) which means Fable 5 is already an “older” model.”
https://x.com/emollick/status/2077844979305594892

Kimi K3 on my shader test: “create a visually interesting shader that can run in twigl-dot-app make it like an infinite city of neo-gothic towers partially drowned in a stormy ocean with large waves.” “Make it better” Very good model, not Sol Max or Fable, but great open weights”
https://x.com/emollick/status/2077783731691995348

Good lord! This is quite the showing by Kimi. Seems we’re gonna have Fable 5 on the sub for a while longer.”
https://x.com/bilawalsidhu/status/2077908938297655722

NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval
https://huggingface.co/blog/nvidia/nemotron-3-embed-wins-rteb

Perplexity has the best (both on cost and performance) deep and wide research harness in Computer. One of the contributing factors is strong internal evals and benchmarks. Today, we’re open-sourcing WANDR, the benchmark we use internally for measuring research capabilities.”
https://x.com/AravSrinivas/status/2077105849638728118

today we’re open-sourcing an eval/RL environment for measuring agentic search performance. importantly, these environments were synthesized from production traces, offering a real-world distribution, with weak human supervision. internally, we’ve been using these RL environments”
https://x.com/denisyarats/status/2077117794869805145

I wrote about open models and how to set up a coding agent based on GLM-5.2, for training two small LMs. This article comes after our recent update to @Dell Enterprise Hub, where we added to the catalog the new GLM-5.2-FP8. Read it here:
https://x.com/juanjucm/status/2076714987569963508

Here is Ternary Bonsai 27B running an end-to-end agentic workflow locally with Hermes on an NVIDIA GeForce RTX 5090 GPU. The model reasons, calls tools, reads outputs, modifies files, and surfaces insights – all on consumer hardware, while all private files, intermediate states,”
https://x.com/PrismML/status/2077084899904024918

12 free courses to master LLMs ▪️ Cohere LLM University ▪️ Hugging Face LLM Course ▪️ Hugging Face AI Agents Course ▪️ Google / Kaggle 5-Day Gen AI Intensive ▪️ DeepLearning. AI Short Courses ▪️ Hugging Face Context Course ▪️ Google / Kaggle 5-Day AI Agents Intensive ▪️”
https://x.com/TheTuringPost/status/2076430688317026619

Kimi K3 seems really good, closest to the frontier yet, but also wow does the model/harness love to loop back over and over again over tasks tweaking and changing things at max level. Anyhow, examples incoming, when they finish.”
https://x.com/emollick/status/2077770187521069152

I think OpenRouter is not a good measure of actual model usage in a world of agentic tools (not that I doubt that Chinese open weights model usage is up, but this could also look like a graph of usage shifting to Codex/Code/Cowork). We really need better data indicators for AI!”
https://x.com/emollick/status/2076769431120756834

Cursor, Copilot, Pi, and OpenCode tracing: Now in LangSmith. Full session observability, no extra instrumentation. ✅Identify, group, and query any coding-agent trace with the same stable keys, regardless of which agent produced it ✅See the full run tree: turns, model calls,”
https://x.com/LangChain/status/2077076144248021236

You can now access and activate your banked resets in Hermes Agent directly with /usage reset when using a codex/openai subscription”
https://x.com/Teknium/status/2077006948223090777

Prior to today, tool calls would work in parallel only if all were safe to parallelize. Now you gain big speedups for parallel tool calls so long as any subset are parallelizable. Check it out”
https://x.com/Teknium/status/2077132644979200150

We built a tracing plugin for every Codex session into LangSmith. Now every turn (tool calls, token usage, subagent threads) lands in LangSmith as a real trace you can dig into. Two config blocks and one flag, and it’s live.”
https://x.com/LangChain/status/2077045458917052492

> For me, agentic performance on spreadsheet tasks is the whole of Shortcut existence. So I’m going to beat competitors by obsessing over the harness. custom harness is the only way you will beat the labs agentic experience heres how to build one:”
https://x.com/hwchase17/status/2076784403414651035

OpenWiki Brains: Proactive Memory for AI Agents
https://www.langchain.com/blog/introducing-openwiki-brains-general-purpose-wiki-memory-for-agents

We’ve open-sourced Grok Build and have reset usage limits for all users. Open sourcing Grok Build allows anyone to support making a reliable and robust harness. Check out our code, including the Git repo for the Grok Build CLI.”
https://x.com/SpaceXAI/status/2077494535387828644

3 recaps to help you stay on track with what’s most important in AI in 2026 ▪️ AI Agents in 2026: Local, Physical, Responsible AI – Skill Engineering – OpenClaw and Hermes Agent – VLAs – Recursive self-improvement ▪️ AI Concepts and Techniques in 2026: – Conditional Memory -“
https://x.com/TheTuringPost/status/2076104540513001532

What are the best models you can run on your @NVIDIAAI DGX Spark? ✨ Mid-July 2026 Edition 1× DGX Spark • ⁠Qwen 3.6 35b NVFP4 — 256k ctx, 81 tok/s • ⁠Qwen 3.6 27b NVFP4 — 256k ctx, 33 tok/s 2× DGX Sparks ← sweet spot! • DeepSeek v4 Flash — 1M ctx, 60 tok/s • MiMo-V2.5″
https://x.com/MiaAI_lab/status/2076951362407944622

So I guess it is time to wonder: how does pre-clearance work for open weights models? No model card yet from Kimi K3 but maybe at weight release in a couple weeks, yet open models are easy to jailbreak. Do open models claiming to be Mythos/Sol level (K3 is not yet there, but”
https://x.com/emollick/status/2077912902032392661

Happy to see a new open weights model, but, so far, Inkling is pretty rough in my tests, not close to frontier Chinese open weights models. As one example, here is it failing to pass the Lem Test (done by every frontier model since DeepSeek r1/Sonnet 3.5).”
https://x.com/emollick/status/2077593908540850491

open source AI is…happening? 🙂 Anthropic & OpenAI better hope AT&T is the exception, not the rule”
https://x.com/amir/status/2090515013635305683

One thing I am kind of surprised by is that full multi-modal (any-any) models have not become a bigger deal. It seems Google is the only Lab releasing these, OpenAI uses selective multimodal capabilities, Anthropic famously has no multimodal output & open weights models are mixed”
https://x.com/emollick/status/2076772061427401162

For the first time since Claude Code came out, I moved one of my actual work pipelines to @pidotdev & open-weight models. And after a weekend of fighting with it, I prefer the report it made. 36 pages vs 21 from Claude, more information-dense, prose I liked more, and pennies”
https://x.com/TheZachMueller/status/2076746035758502275

Simulate realistic touch at massive scale​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​! {📌 with open code & sim integration released} A scalable tactile sensor simulator integrated into the Genesis physics engine that enables high-throughput parallel training of dexterous”
https://x.com/IlirAliu_/status/2077301186646946020

Mini 6 dof Arm. 3D printed planetary gearboxs & more… [📍GitHub link below ] A mini 6-axis arm driven by stepper motors with custom 3D printed split ring planetary gearboxs and an inverted belt differential wrist with custom bearings, driven by low-cost stepper motors and”
https://x.com/IlirAliu_/status/2075851609767051545

We’re open sourcing WANDR. WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer.”
https://x.com/perplexity_ai/status/2077099503723946121

Google announces Gemma 4 optimized for the Pixel 10’s TPU
https://9to5google.com/2026/07/14/pixel-10-gemma-4/

We’ve just released the 1-bit & 4-bit version of Hy3, a flagship-scale 295B model that can be served on a single GPU. 👌 Run Hy3 with llama.cpp, enable MTP, and experience powerful intelligence on dramatically lower hardware.🚀🚀🚀 Can’t wait to see what you build. #Hy3 #Hy”
https://x.com/TencentHunyuan/status/2076953120765280284

Peter Corke won the highest honor in robotics. His entire robotics curriculum is free online. 📌 Has been for 10 years. 👇️ You visit his website and find 200+ free lessons, an open-source toolbox used by researchers worldwide, and a full textbook that runs in Google Colab.”
https://x.com/IlirAliu_/status/2076004351311577348

More NVFP4 dynamic quants! We made them for all Gemma-4 sizes (E2B, E4B, 12B, 26B-A4B, 31B) – they’re all W4A4 + FP8 KV cache calibrated + W8A8 for attention / important layers. We also made Qwen3.5-122B-A10B and GLM-4.7-Flash NVFP4 ones as well – 397B and others will come soon!”
https://x.com/danielhanchen/status/2077072556537020914

🤗 MOSS-VL-Realtime is now open source on @huggingface . The 11B model family supports text, single and multiple images, single and multiple videos, and interleaved visual-text inputs in Chinese and English.@MosiAI_Official Highlights: 🏗️ Cross-Attention architecture”
https://x.com/Open_MOSS/status/2076993673552879790

🤗 MOSS-VL-Realtime is now open source on @huggingface . Built for real-time visual understanding over continuous video streams: 🧠 11B vision-language model 📜 Apache-2.0 license 💬 Ask questions at any point in a video stream 👀 Keeps watching while generating a response 🔄”
https://x.com/MosiAI_Official/status/2076989390191202577

Big unlock for open-source AI inference: Hugging Face Transformers models can now run in vLLM at native speed, often matching or beating hand-written implementations. Until now, every new architecture often needed to be built twice: – Once in Transformers for training and”
https://x.com/ClementDelangue/status/2076763231788339669

Model Routing Is Simple. Until It Isn’t.
https://huggingface.co/blog/ibm-research/model-routing-is-simple-until-it-isnt

There’s a major gap in this otherwise compelling vision. Lots of interesting signals recently: token budget caps, intelligence sovereignty, increasing competition (Meta/Grok/GLM) that drives the commoditization of intelligence. All seem to point toward a more benevolent”
https://x.com/ysu_nlp/status/2076481232117067894

A note of caution: I will say that when doing some complex statistical auditing of some of my prior academic work, Kimi K3 Max messed up in a bunch of ways, including misapplying statistics and applying some stuff badly. A bit from GPT-5.6 Pro critiquing K3 (which I agree with):”
https://x.com/emollick/status/2077869293031624793

Kimi K3 Tech Blog: Open Frontier Intelligence
https://www.kimi.ai/blog/kimi-k3

Kimi K3 cannot write a good murder mystery (though neither can any other model). That remains the jaggedest of frontiers. They both make things too obvious (the letter) and too obscure, and cannot foreshadow to save their artificial lives.”
https://x.com/emollick/status/2077951790868238616

Announcing LlamaCoder v4 – generate apps in 1 prompt! • Rebuilt it around GLM 5.2 & its strengths • Migrated to using Base UI w/ @shadcn • Improved parsing, planning, & design of apps • New WebAssembly-based renderer 100% free, open source, and powered by @togethercompute.”
https://x.com/nutlope/status/2076722464671793184

GLM-5.3-Flash: Frontier Intelligence, Flash Cost
https://z.ai/blog/glm-5.3-flash

.@satyanadella’s Reverse Information Paradox is real. What @satyanadella’s calling for already exists: open models. Over 9M+ developers have used @ollama to access open models and keep their competitive edge in-house. You can’t be locked out of a model you own. It’s your”
https://x.com/mchiang0610/status/2076736707471556755

Welcome @ATT to open models!”
https://x.com/ollama/status/2090601698402447748

For the last few months, the conversation around open models has centered on cost and performance optimization. But Satya highlights something more existential: open models aren’t just an optimization. They’re the foundation of a new software flywheel for every organization,”
https://x.com/jmorgan/status/2076750580052369896

A high-quality haptic control for robots as a weekend project? All for under $600. [📍 Save Github & arXiv for later] A team built DOGlove, an open-source haptic glove for robot teleoperation that you can assemble yourself for under $600. It is built for real control, not”
https://x.com/IlirAliu_/status/2076951255256047993

Inkling: Our Open-Weights Model – Thinking Machines Lab
https://thinkingmachines.ai/news/introducing-inkling/

Open-weight models surge to 29% of volume, price per token flattens – Vercel
https://vercel.com/blog/ai-gateway-production-index-july-2026

The State of Open Source AI — v1.0.1 · July 2026
https://stateofopensource.ai/

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading