The UK’s state AI Security iIstitute findings: 1) Mythos is a big gain in cyber capabilities. But so is GPT-5.5 2) It is hard to establish an upper bound on Mythos/GPT-5.5, which appear to be limited by tokens used, rather than ability. 3) Capability doubling time is 4.5 months
https://x.com/emollick/status/2054595505712165154
Kimi K2.6 is now open-weight #1 on Finance Agent Benchmark V2.
https://x.com/Kimi_Moonshot/status/2054803169994272819
Meet Kimi Web Bridge – Kimi’s browser extension. Agent can now interact with websites like a human: search, scroll, click, type and complete tasks. Supports Kimi Code CLI, Claude Code, Cursor, Codex, Hermes, and more. Available now on
https://t.co/sUqDpi0HQr and the Chrome Web
https://x.com/Kimi_Moonshot/status/2054918374837322140
Seeing the demos come together over the last week has been awesome — so many things that previously required a special-purpose model (e.g. real-time translation, event detection in video) turn out to be zero-shot instruction following once you have a general-purpose model with
https://x.com/johnschulman2/status/2053940940885332028
Interaction Models: A Scalable Approach to Human-AI Collaboration – Thinking Machines Lab
https://thinkingmachines.ai/blog/interaction-models/
Interaction Models: A Scalable Approach to Human-AI Collaboration – Thinking Machines Lab
https://thinkingmachines.ai/blog/interaction-models/
People talk, listen, watch, think, and collaborate at the same time, in real time. We’ve designed an AI that works with people the same way. We share our approach, early results, and a quick look at our model in action.
https://x.com/thinkymachines/status/2053938892152435174
Sharing our work on full-duplex multimodal models — real-time interaction that’s natural and intuitive without compromising on intelligence. We started Thinky in part to differentially advance capabilities for human-AI collaboration, which are underemphasized relative to
https://x.com/johnschulman2/status/2053940452789981426
thinking machines is using SGLang btw
https://x.com/eliebakouch/status/2053982248253190180
Thinking Machines know how to surprise. Those simultaneous abilities (not only translation but also creating graph while replying to a question) are pretty remarkable. Can’t wait to try it out and also learn how much it costs to use
https://x.com/TheTuringPost/status/2053975565179253010
Thinking Machines on X: “People talk, listen, watch, think, and collaborate at the same time, in real time. We’ve designed an AI that works with people the same way. We share our approach, early results, and a quick look at our model in action. https://t.co/AFJZ5kH7Ku https://t.co/uxl1InS6Ay” / X
https://x.com/thinkymachines/status/2053938892152435174
Thinky’s secret plan: 1: Increase Human<->AI bandwidth 2: Raise ceiling of human+AI intelligence 3: Help humans continue as main-characters in the new world We are at Step 1. Interaction Models are great real-time collaborative tools for humans. Here’s a preview:
https://x.com/soumithchintala/status/2053940215505645938
Very cool announcement from Thinky! The model looks nice (they go into some reasonable amount of detail), and reading some parts of the blog you can definitely see that the infea guys had a lot of fun there!
https://x.com/giffmana/status/2053953584300003405
ERNIE 5.1 just dropped. Built on ERNIE 5.0’s pre-training foundation, our latest foundation model upgrades search, reasoning, knowledge Q&A, creative writing, and agentic capabilities, while using only around 6% of the pre-training cost of comparable models. More in the thread
https://x.com/Baidu_Inc/status/2053009538769735774?s=20
DeepSeek V4 Flash is ~90% cheaper than GPT 5.4 Mini and ~70% cheaper than Gemini 3.1 Flash Lite For devs pushing ~500M tok/month, this is the difference between: GPT 5.4 Mini: ~$394/mo Gemini 3.1 Flash Lite: ~$131/mo DeepSeek V4 Flash: ~$71/mo … roughly ~$3,900/dev/yr back
https://x.com/masondrxy/status/2053855842076942555
We Tested DeepSeek V4 Pro and Flash Against Claude Opus 4.7 and Kimi K2.6
https://blog.kilo.ai/p/we-tested-deepseek-v4-pro-and-flash
PSA, we are hiring for our DevX team to kick start our India presence 🇮🇳! If you want to help builders get the most out of Gemini in India, please ping me via DM or email. India is our largest market from an AI Studio user pov, very excited to visit later this year as well!!
https://x.com/OfficialLoganK/status/2052199234607182077
[2605.10730] Qwen-Image-2.0 Technical Report
https://arxiv.org/abs/2605.10730
Exciting: local ML is (finally) going mainstream 🔥 – new GGUF uploads on HF nearly doubled in 2 months – smaller models (like gemma 4 and qwen 27b) starting to be really good and run on a lot of hardwaers – people forking and vibing over llama.cpp (MTP / DS4 / turboquant…)
https://x.com/victormustar/status/2053780086596288781
GB 200s change how one does the prefill and decode disaggregation when serving large MoEs like Qwen. We’ve published details of our stack quantifying the throughput benefits compared to serving on Hoppers.
https://x.com/AravSrinivas/status/2054206802133504234





Leave a Reply