Image created with gemini-2.5-flash-image with claude-sonnet-4-5-20250929. Image prompt: A sleek robotic hand with chrome and matte black finish carefully placing a second lit candle onto a modern minimalist birthday cake on a slate grey surface, vibrant confetti in deep blue and rich red scattered around, cinematic lighting with high contrast creating dramatic shadows, celebratory yet sophisticated atmosphere, close-up composition emphasizing the precision of the gesture.

A SOTA moment to me: Kimi’s OK Computer generate this website in just one shot > It designed a very beautiful site, all images were AI-generated, and when you click, the sidebar expands. > Inside the sidebar, there’s a handwritten letter, it really feels like a website made by a https://x.com/crystalsssup/status/1971133240619757794

Say hi to OK Computer, Kimi’s agent mode 🤖🎸 Your AI product & engineering team, all in one. ✨ From chat → multi-page websites, mobile first designs, editable slides ✨ From up to 1 million rows of data → interactive dashboards ✨ Agency: self-scopes, surveys & designs ✨ https://x.com/Kimi_Moonshot/status/1971078467560276160

Grok 4 Fast | xAI https://x.ai/news/grok-4-fast

Grok Code Fast 1 | xAI https://x.ai/news/grok-code-fast-1

in my testing so far, i found grok-4-fast is to be fast af like 2x-3x fast (higher throughput, you can check exact numbers on openrouter) and takes much less reasoning time. it’s significantly weaker at instruction following than gpt-5-mini. the task i tested on was a large”” / X https://x.com/dejavucoder/status/1969383391029313598

The Grok-4-fast journey has been incredible—kicking off right after the Grok 4 launch in July. None of it happens without the absolute GOAT @s_tworkowski our incredibly talented teammates @LiTianleli @mycharmspace , and the unwavering backing from @Yuhu_ai_ . This kind of”” / X https://x.com/ShuyangGao62860/status/1969240703080546376

The new Grok 4 Fast seems to do well with creative coding challenges (“”create a visually interesting shader that can run in twigl, make it like the ocean in a storm””, “”make a futuristic starship panel for p5js””) for a small model but not quite as great in creative language tests https://x.com/emollick/status/1969477042203771150

With the new Grok 4 Fast, the price/performance curve for AI shifted again. I updated my chart to reflect I also think GPQA Diamond is likely maxed out (the tests themselves have errors, making it impossible to get to 100%), I am going to need to do this with a harder benchmark. https://x.com/emollick/status/1969845283816161726

Not to take away from Grok 4 Fast (which seems like a very good model) or from Artificial Analysis (one of the few organizations doing independent benchmarking), but the Intelligence Index is an average of pretty saturated benchmarks (aside from HLE), we really need better ones.”” / X https://x.com/emollick/status/1969270709361733942

Veo is a more general reasoner than you might think. Check out this super cool paper on “”Video models are zero-shot learners and reasoners”” from my colleagues at @GoogleDeepMind. https://x.com/tkipf/status/1971063116734841248

Nova Act extension: Build and test AI agents without leaving your IDE https://labs.amazon.science/blog/nova-act-extension-build-and-test-ai-agents-without-leaving-your-ide

We’ve been building agents for a while and Claude Code SDK has unlocked capabilities that were previously impossible This is an enterprise grade Business Analyst Agent we’ve built out on top of the SDK https://x.com/ponnappa/status/1968710157409374241

I find it unimaginably based that the OAI Evals team keeps making benchmarks finding that Claude is better and publishing it anyway. they are 3 for 3 this year in acknowledging specifically how much Claude is better at tasks OAI care about. there is no sarcasm here folks. this https://x.com/swyx/status/1971404125553242253

NEW: Anthropic web search ✨ OpenRouter now uses the native web engines for OpenAI and Anthropic models by default For all other models, our custom web search will be used, powered by @ExaAILabs Configurable! 👇 https://x.com/OpenRouterAI/status/1968360919488151911

This next level! You can now combine Browser Use with Code Execution and have Gemini 2.5 control our browser via UI controls and write dynamic Javascript to extract data from pages or do other things! 🤯 Below is an example how Gemini 2.5 uses javascript tool and writes a https://x.com/_philschmid/status/1968685597519654994

Announcing our public preview of Chrome DevTools MCP! Experience the full power of DevTools in your AI coding agent → https://x.com/ChromiumDev/status/1970505063064825994

Chrome DevTools (MCP) for your AI agent  |  Blog  |  Chrome for Developers https://developer.chrome.com/blog/chrome-devtools-mcp

Browser Use 0.7.8 can write Javascript Code🤯 This was never possible. Browser agents used to be limited by 🎯Clicking on raw coordinates 🔂Selecting from a list of buttons 🐁Complex mouse movements Today, we combine coding agents with browser agents. Now we can interact with https://x.com/gregpr07/status/1968453999914590212

OpenAI might also be developing AI glasses, a voice recorder, and a pin | The Verge https://www.theverge.com/news/781854/openai-chatgpt-hardware-rumors-smart-speaker-glasses-pin

GPT-5 is the best model for code quality out there 2 years ago, we created the world’s hardest software design quiz. Only 5 questions, multiple choice. Yet only about 3% of software engineers get them. The average score is somewhere between 2 and 3. Supposedly brilliant models https://x.com/jimmykoppel/status/1968683689421701413

Measuring the performance of our models on real-world tasks | OpenAI https://openai.com/index/gdpval/

The implications of OpenAI’s plan to rent $450 billion worth of servers before the end of this decade are 🤯 https://x.com/amir/status/1969043037805228388

Announcing strategic partnership with @nvidia for millions of GPUs — about as much compute as they’ve shipped in 2025 in total — and an investment up to $100B as these GPUs are deployed: https://x.com/gdb/status/1970173243350008201

AI should do more than just answer questions; it should anticipate your needs and help you reach your goals. That’s what we’re beginning to build, starting with ChatGPT Pulse (rolling out now to Pro, with goal of making it available to everyone over time): https://x.com/fidjissimo/status/1971258542578663829

Introducing ChatGPT Pulse | OpenAI https://openai.com/index/introducing-chatgpt-pulse/

Now in preview: ChatGPT Pulse This is a new experience where ChatGPT can proactively deliver personalized daily updates from your chats, feedback, and connected apps like your calendar. Rolling out to Pro users on mobile today. https://x.com/OpenAI/status/1971259652684878019

Today we are launching my favorite feature of ChatGPT so far, called Pulse. It is initially available to Pro subscribers. Pulse works for you overnight, and keeps thinking about your interests, your connected data, your recent chats, and more. Every morning, you get a”” / X https://x.com/sama/status/1971297661748953263

For both $NVDA and OpenAI, the $100B $NVDA investment is perfect: 1. For OAI, the biggest question was how they were going to raise the future +$300B as the valuation is already very high, and the cash burn for the next few years is projected to be crazy. On top of it, there is”” / X https://x.com/rihardjarc/status/1970170005858726278

Two years ago. No model had surpassed GPT-4 & it wasn’t clear that was possible. Now you can get better than GPT-4 level performance on open weights models running on consumer hardware, and the state of the art in LLMs is cheaper & faster and very much more capable than GPT-4.”” / X https://x.com/emollick/status/1970790843868213361

We’ve released a large-scale study on how people are using ChatGPT. Consumer adoption has broadened beyond early-user groups, and lots of economic value is being created through both personal and professional use: https://x.com/gdb/status/1969953507215302836

Introducing Perplexity Email Assistant: an agent that behaves as your personal/executive assistant on your email client (Gmail, Outlook); scheduling meetings, prioritizing emails, and drafting replies for you. Available to all Perplexity Max subscribers from today! https://x.com/AravSrinivas/status/1970165878751973560

Introducing Perplexity Email Assistant. Now anyone can have a personal assistant in their email that schedules meetings, drafts replies, and labels priorities. Perplexity Email Assistant is now available on Gmail and Outlook for all Perplexity Max subscribers. https://x.com/perplexity_ai/status/1970165704826716618

Introducing Perplexity Search API We’ve built a search index of billions of webpages to provide real-time, quality information from the web. Now developers have access to the full power of our index, providing the most accurate results in milliseconds. https://x.com/perplexity_ai/status/1971274917401461236

Perplexity Search API: Providing direct search results in milliseconds for grounding LLMs and agents with real-time information from the web. This is an effort that began more than two years ago: to build our own search index. So much progress in a short period of time. We look”” / X https://x.com/AravSrinivas/status/1971275716357656987

Read more about Perplexity Enterprise Max on our blog: https://x.com/perplexity_ai/status/1968707015389364335

We’re excited to announce Perplexity Enterprise Max. Get unlimited Labs queries, 10x file uploads, premium security features for your org, and access to Comet Max Assistant. Enterprise Max is our most powerful tier for enterprise teams who are looking to get more work done. https://x.com/perplexity_ai/status/1968707003175641098

Qwen3-VL is finally released and open-sourced, available in both Thinking and Instruct versions! This time, we’ve placed special emphasis on strengthening Visual Agent and Visual Coding, which are crucial steps toward building a true Digital Agent 🚀”” / X https://x.com/huybery/status/1970650821747712209

🎙️ Meet Qwen3-TTS-Flash — the new text-to-speech model that’s redefining voice AI! Demo: https://x.com/Alibaba_Qwen/status/1970163551676592430

🔥 Qwen-Image-Edit-2509 IS LIVE — and it’s a GAME CHANGER. 🔥 We didn’t just upgrade it. We rebuilt it for creators, designers, and AI tinkerers who demand pixel-perfect control. ✅ Multi-Image Editing? YES. Drag in “person + product” or “person + scene” — it blends them like https://x.com/Alibaba_Qwen/status/1970189775467647266

🚀 Introducing Qwen3-LiveTranslate-Flash — Real‑Time Multimodal Interpretation — See It, Hear It, Speak It! 🌐 Wide language coverage — Understands 18 languages & 6 dialects, speaks 10 languages. 👁️ Vision‑Enhanced Comprehension — Reads lips, gestures, on‑screen text and https://x.com/Alibaba_Qwen/status/1970565641594867973

🚀 Introducing Qwen3-Omni — the first natively end-to-end omni-modal AI unifying text, image, audio & video in one model — no modality trade-offs! 🏆 SOTA on 22/36 audio & AV benchmarks 🌍 119L text / 19L speech in / 10L speech out ⚡ 211ms latency | 🎧 30-min audio https://x.com/Alibaba_Qwen/status/1970181599133344172

🚀 We’re thrilled to unveil Qwen3-VL — the most powerful vision-language model in the Qwen series yet! 🔥 The flagship model Qwen3-VL-235B-A22B is now open-sourced and available in both Instruct and Thinking versions: ✅ Instruct outperforms Gemini 2.5 Pro on key vision https://x.com/Alibaba_Qwen/status/1970594923503391182

🛡️ Meet Qwen3Guard — the Qwen3-based safety moderation model series built for global, real-time AI safety! 🌍 Supports 119 languages and dialects ✅ 3 sizes available: 0.6B, 4B, 8B ⚡ Low-latency, Real-time streaming detection with Qwen3Guard-Stream 📝 Robust Full-context safety https://x.com/Alibaba_Qwen/status/1970510193537753397

Alibaba Qwen officially achieves frontier lab status LFG https://x.com/zephyr_z9/status/1970587657421156622

Alibaba released Qwen3-Next-80B-A3B in Base, Instruct, and Thinking variants under an open-weights Apache 2.0 license, targeting faster long-context inference. The 80-billion-parameter mixture-of-experts design swaps most vanilla attention layers for Gated DeltaNet ones and the https://x.com/DeepLearningAI/status/1970254860416131146

Announcing the open-source release of Qwen3-VL! A powerful vision-language model that can operate GUIs, code https://t.co/ww8tsXcd1u charts from mockups, and recognize “”everything”” from daily life to specialized fields. Highlights: 🔹 Precise event location in videos up to 2 https://x.com/Ali_TongyiLab/status/1970665194390220864

NEW: Qwen 235B A22B Vision Language Model is OUTT! Apache 2.0 licensed and upto 1 Million context length 🤯 https://x.com/reach_vb/status/1970589927134937309

Qwen https://qwen.ai/blog?id=1675c295dc29dd31073e5b3f72876e9d684e41c6&from=research.research-list

Qwen https://qwen.ai/blog?id=241398b9cd6353de490b0f82806c7848c5d2777d&from=research.latest-advancements-list

Qwen https://qwen.ai/blog?id=99f0335c4ad9ff6153e517418d48535ab6d8afef&from=research.latest-advancements-list

Qwen https://qwen.ai/blog?id=b2de6ae8555599bf3b87eec55a285cdf496b78e4&from=research.latest-advancements-list

Qwen https://qwen.ai/blog?id=f0bbad0677edf58ba93d80a1e12ce458f7a80548&from=research.research-list

Qwen https://qwen.ai/blog?id=f50261eff44dfc0dcbade2baf1b527692bdca4cd&from=research.research-list

Qwen https://qwen.ai/blog?id=fdfbaf2907a36b7659a470c77fb135e381302028&from=research.research-list

Qwen3 VL might be the best multimodal (vision) model on the planet”” / X https://x.com/scaling01/status/1970591728433283354

Qwen3-Omni is new sota any-to-any model🔥 everything you have to know ⤵️ > a 30B MoE model with 3B active params, comes in three variants: instruct, thinking and captioner 🤩 thinking is for reasoning and captioner is for robust speech generation 🗣️ > it understands everything https://x.com/mervenoyann/status/1970444546216444022

Qwen3-Omni Technical Report A unified multimodal model that matches same-size Qwen text-only and vision-only baselines while pushing audio and audio-visual SOTA. Key technical details below: https://x.com/omarsar0/status/1970502225379381662

We’re excited to announce the upgrade of Qwen3-Coder, and the upgraded API `qwen3-coder-plus` is now available on Alibaba Cloud Model Studio with major improvements: 💻 Enhanced terminal task capabilities and better performance on Terminal Bench (w/ Qwen Code / Claude Code) 🏆 https://x.com/Alibaba_Qwen/status/1970582211993927774

🚨 New Models Update! 🔥 Qwen3 coming in hot into the Arena with three different models: 🔹Qwen3-VL-235b-a22b-thinking for Text & Vision 🔹Qwen3-VL-235b-a22b-instruct for Text & Vision 🔹Qwen3-Max-2025-9-23 for Text Check out the thread to learn more about them and get https://x.com/arena/status/1970920636957831611

Qwen just released Qwen3Guard-Gen-8B on Hugging Face This new safety moderation model offers three-tiered severity classification and multilingual support for AI content. https://x.com/HuggingPapers/status/1970504452466413639

Wow. Qwen Image Edit now has native support for ControlNet (depth maps, edge maps, keypoint maps etc)”” / X https://x.com/bilawalsidhu/status/1970193454505541755

Skild AI’s omni-bodied brain, trained on 100,000 diverse simulated robots for 1000 years, enables remarkable real-world adaptability. In-context adaptation allows the brain to discern the robot form and adapt to extreme changes like chopped limbs or walking on stilts. https://x.com/TheHumanoidHub/status/1970981739200909811

Gemini Robotics 1.5 brings AI agents into the physical world – Google DeepMind https://deepmind.google/discover/blog/gemini-robotics-15-brings-ai-agents-into-the-physical-world/

New Gemini Robotics 1.5 models will enable robots to better reason, plan ahead, use digital tools like Search, and transfer learning from one kind of robot to another. Our next big step towards general-purpose robots that are truly helpful — you can see how the robot reasons as https://x.com/sundarpichai/status/1971244716046872577

Talk to robots! Today we’re releasing our SOTA Gemini Robotics 1.5 model showing the power of using our multimodal Gemini models as a base, so it can understand & reason about the physical world. Robotics will be massive in the future – super excited by our pioneering work here!”” / X https://x.com/demishassabis/status/1971292365592854602

We’re making robots more capable than ever in the physical world. 🤖 Gemini Robotics 1.5 is a levelled up agentic system that can reason better, plan ahead, use digital tools such as @Google Search, interact with humans and much more. Here’s how it works 🧵 https://x.com/GoogleDeepMind/status/1971243947792925005

Introducing our Agentic Leaderboards. These new leaderboards test AI agents in real-world, high-complexity environments, setting a new standard for completing end-to-end digital tasks. https://x.com/scale_AI/status/1969416303015301128

Everyone says AI learns just like humans do – by picking up patterns and styles. But here’s what they’re missing: Can you watch every YouTube video ever made? Read every book on earth? That’s the real difference. It’s like facial recognition – sure you can spot a face in a https://x.com/bilawalsidhu/status/1969819287738307003

Introducing the Data Commons Model Context Protocol (MCP) Server: Streamlining Public Data Access for AI Developers – Google Developers Blog https://developers.googleblog.com/en/datacommonsmcp/

Satya Nadella is haunted at the prospect of Microsoft not surviving the AI era | The Verge https://www.theverge.com/tech/780946/microsoft-satya-nadella-town-hall-comments-ai-era-notepad

👏 Super proud of the @weaviate_io labs team! The Weaviate Query Agent is now in GA! Check out this demo notebook and some exciting customer stories: 📄 https://x.com/bobvanluijt/status/1968609785416196347

A good rule of thumb is that generalist AI models are less likely to be useful when being “a friendly assistant” is unhelpful. It is why there is opportunity (for now, at least) for new approaches in coaching & teaching (some adversarial elements) and digital twins & simulation.”” / X https://x.com/emollick/status/1969454060462874671

AI agents built into cloud apps will always be sub-par compared to state of the art agents that you can use with Obsidian. This is because your Obsidian data is in your control, in plain text formats that are ideal for LLMs to process. You can choose to run any of the https://x.com/kepano/status/1969035924152738238

AI Engineer Code Summit https://apply.ai.engineer/

Code World Model: producing code by imagining the effect of executing instructions and planning instructions that produce the desired effect.”” / X https://x.com/ylecun/status/1970967341052854748

Deploy your own AI vibe coding platform — in one click! https://blog.cloudflare.com/deploy-your-own-ai-vibe-coding-platform/

if the agent ingests anything, then it’s permissions should be dropped to the level of the author of that information”” I like that lot, it’s a very succinct way of explaining the problem here Anyone who can author text that gets into your agent can control what that agent does”” / X https://x.com/simonw/status/1969247680762429784

MetaBuddy boosted their user engagement by 3x and cut trainer analysis time by 60% – using Weaviate’s Query Agent. 𝗕𝘂𝘁 𝗳𝗶𝗿𝘀𝘁, 𝘄𝗵𝗮𝘁 𝗶𝘀 𝗠𝗲𝘁𝗮𝗕𝘂𝗱𝗱𝘆? MetaBuddy is a platform that connects fitness enthusiasts through interactive wellness experiences, events, and https://x.com/weaviate_io/status/1968691524318761165

ml infra is really hard. great job to everyone who worked on the debug and writeup.”” / X https://x.com/itsclivetime/status/1968534889151742437

Rocket.new | Build Web & Mobile Apps 10x Faster Without Code https://www.rocket.new/

RPG A Repository Planning Graph for Unified and Scalable Codebase Generation https://x.com/_akhaliq/status/1969976136232022289

We’re launching GLM Coding Plans with @Zai_org for Cline users. $3/month gets you 120 prompts per 5-hour cycle with GLM-4.5. $15/month gets you 600 prompts. Both plans give you frontier-level coding AI at a fraction of typical subscription costs. (details below) https://x.com/cline/status/1968820438156640490

100+ MCP servers now available in a container. Just pull the Docker image to use your favourite MCP server. 100% Opensource. https://x.com/Saboo_Shubham_/status/1968724555188281575

Metastone unveils MCP-AgentBench A new benchmark evaluating real-world language agent performance with MCP-mediated tools. It features 33 live servers & 188 tools to rigorously test agent capabilities beyond traditional metrics. https://x.com/HuggingPapers/status/1969864853238985001

Did you see that the Agent Research Environment is MCP compatible? -> using any MCP tools with any agent is now completely trivial! Check it out! We’ve used an LLM agent to 1) move a robot arm remotely 2) depending on real time web search results! 😀 How to in thread ^^ https://x.com/clefourrier/status/1970394602592182627

Factory are one of the most formidable teams in the AI Coding space and I think the Droid concept strikes exactly the right balance between concerns of engineer replacement and the ill defined expectations of agents. while there is some way to go on SWEBench still, its p clear https://x.com/swyx/status/1971310686585356295

6 months, 25 million revenue agents & 3 trillion tokens later… Rox is now globally available 🌎 Just as coding agents 10x’d engineering, revenue agents 10x customer work. With Rox, humans are evolving to orchestrators while agents manage the end-to-end customer lifecycle. https://x.com/rox_ai/status/1967973771681337578

Cohere adds $100M in second close to latest round as it scales security-first enterprise AI https://cohere.com/blog/september-2025-funding-round

Notion 3.0 with Agents is out today! It’s the first Knowledge Work Agent in the world. It works with Notion databases. It can do multi-step actions and autonomous work up to 20+ mins, and a brand-new memory system (using Notion pages and databases! 👌) With 3.0, Notion isn’t https://x.com/ivanhzhao/status/1968761820241609063

The best agents for software development are becoming the best agents for everything. Droids are the best software development agents in the world, reaching #1 on Terminal-Bench. We have raised $50M from NEA, Sequoia Capital, J.P. Morgan, Nvidia, Abstract Ventures, and other https://x.com/FactoryAI/status/1971271085653156054

ZapConnect 2025 | Zapier’s annual conference https://zapier.com/zapconnect

Cohere hits $7B valuation a month after its last raise, partners with AMD | TechCrunch https://techcrunch.com/2025/09/24/cohere-hits-7b-valuation-a-month-after-its-last-raise-partners-with-amd/

Many people think LLMs are non-deterministic. This is often not true! You just need 3 lines of code to make your LLM deterministic LLMs (as any PyTorch model) are non-deterministic only when they include certain operations or when using multiple GPUs Try the code yourself https://x.com/gabriberton/status/1968559505966350705

out of all my AI PhD friends on the job market this year, the ones that did the best (by far) write kernels. companies are paying out the wazoo for this skillset models approaching human-level performance in kernel-writing over the next year looks pretty unlikely”” / X https://x.com/jxmnop/status/1970498857386541137

First everything was an agent, now everything is superintelligence. Any vaguely defined term becomes marketing jargon.”” / X https://x.com/emollick/status/1968811667111731512

TorchAO Quantized Models and Quantization Recipes Now Available on HuggingFace Hub – PyTorch https://pytorch.org/blog/torchao-quantized-models-and-quantization-recipes-now-available-on-huggingface-hub/

Weaviate is ISO 27001 compliant! 🎉 But what does that actually mean for you? ISO 27001:2022 is the international standard for information security management systems. It requires organizations to systematically manage information security risks through continuous monitoring, https://x.com/weaviate_io/status/1970912361381843104

Our CEO @jdroege sat down with @axios to talk about where AI is headed and how we make AI work for businesses and governments. https://x.com/scale_AI/status/1968730759285317907

🚀 ARE: scaling up agent environments and evaluations Everyone talks about RL envs so we built one we actually use. In the second half of AI, evals & envs are the bottleneck. Today we OSS it all: Meta Agent Research Environment + GAIA-2 (code, demo, evals). 🔗Links👇 https://x.com/ThomasScialom/status/1970122143993037170

Expanding model choice in Microsoft 365 Copilot | Microsoft 365 Blog https://www.microsoft.com/en-us/microsoft-365/blog/2025/09/24/expanding-model-choice-in-microsoft-365-copilot/

We’re testing something 🧪 GitHub Copilot CLI is now in public preview. Build, debug, and deploy with GitHub Copilot coding agent without leaving your terminal. Try it out and tell us what you think 👇 https://x.com/github/status/1971295695853306059

Kimi Infra team dropped K2 Vendor Verifier > You can visually see the difference in tool call accuracy across providers on OpenRouter. https://x.com/crystalsssup/status/1971158566343184511

LIMI: Less Is More for Agency • Argues agentic AI doesn’t need more data, just better data • 78 curated demos → 73.5% on AgencyBench (beats models trained on 10k samples) • Outperforms SOTA (Kimi-K2: 24.1%, DeepSeek: 11.9%, Qwen3: 27.5%, GLM-4.5: 45.1%) • Establishes Agency https://x.com/arankomatsuzaki/status/1970328242688246160

me and @TodePond spent the last couple months building a really good whiteboard agent, out now in 4.0 ! check it out with npm create tldraw it has a really good readme that we spent a lot of time on !”” / X https://x.com/max__drake/status/1968764136419975599

🌟Our latest LangChain Academy course – Deep Agents with LangGraph – is now live!🌟 Many agents today follow the same simple pattern: run in a loop, call tools. That architecture works well enough, but it breaks down as tasks get more complex. Today, companies of all sizes – https://x.com/LangChainAI/status/1968708505201951029

After gpt-oss, the latest gpt-5-codex is the second model to be Responses API only! Makes sense, since Responses is objectively a better choice than Chat Completions for Agentic use-cases”” / X https://x.com/reach_vb/status/1970585119900528964

tldraw canvas agent starter kit is dropping today https://x.com/tldraw/status/1968655029247648229

AI has officially beaten me at the ICPC World Finals. It reminds me of a rare ICPC skill: being able to quickly read a teammate’s code and spot bugs. This skill takes years to train, and explains why AI often makes coding slower (see arXiv:2507.09089). No matter how strong AI”” / X https://x.com/ZeyuanAllenZhu/status/1968568919482089764

Cross-Agent Privilege Escalation: When Agents Free Each Other https://simonwillison.net/2025/Sep/24/cross-agent-privilege-escalation/#atom-everything

First set of @aisdk agents docs have shipped! – foundations – workflow patterns – loop control – `Agent` class https://x.com/nicoalbanese10/status/1968398677686301134

Introducing Parallel Thinking for Reka Research! Instead of one line of reasoning, we explore multiple paths in parallel, then resolve the best answer. Big accuracy gains on Research-Eval (+4.2) and SimpleQA (+3.5). Now live in the Reka Research API! https://x.com/RekaAILabs/status/1971241107322540194

Turso is an incredible technical feat. A Rust rewrite of sqlite, with an async-first architecture, incoming support for concurrent writes, vector search, and browser / wasm support out of the box. I think this has a very good chance of being a foundational piece of https://x.com/rauchg/status/1969515038512926823

Tool Calls Are Expensive And Finite https://www.reillywood.com/blog/tool-calls-are-expensive-and-finite/

AI agency breakthrough: Less data, more power New research with LIMI shows AI agents can achieve 73.5% on benchmarks, outperforming SOTA models by 50%+ using only 78 carefully chosen samples. The “”Agency Efficiency Principle”” is here. https://x.com/HuggingPapers/status/1970400645871185942

Code generators can write functions, but fail with full repositories. It’s because natural language is ill-suited for software structures. @Microsoft introduced the Repository Planning Graph (RPG), a blueprint that links abstract project goals to clear code structures: • https://x.com/TheTuringPost/status/1970068577509327197

Grok 4 Training Resource Footprint https://epochai.substack.com/p/grok-4-training-resource-footprint

Grok-4 Fast SVG test – not great https://x.com/scaling01/status/1969426187312156790

Team integrated @grok-4-fast-reasoning and @grok-coding-fast-1 this weekend. Fast as F…! More like Mach 4, not Grok 4. Congrats @xai team! 🇺🇸 https://x.com/ssankar/status/1970292424917574061

China’s Alibaba just dropped a Python framework for building multi-agent apps. AgentScope lets you build AI agents visually with MCP tools, memory, rag, and reasoning capabilities. Works with any LLM and supports real-time steering. 100% Opensource. https://x.com/Saboo_Shubham_/status/1967274908742025356

Agents that load dynamic MCP tools risk security and quality issues: • Prompt injection • Unreliable tool calls • Unexpected changes • Wasted tokens 𝚖𝚌𝚙-𝚝𝚘-𝚊𝚒-𝚜𝚍𝚔 generates static tools you control so they stay stable and predictable. https://x.com/vercel/status/1968416108018548766

Claude is pretty funny: “Give me 10 brilliant ideas for a science fiction short short story, pick the most brilliant and execute it terribly” It picked: “People start receiving Amazon packages from parallel universes where they made different life choices.” And it was terrible https://x.com/emollick/status/1969185633018024057

Clockwise MCP | The MCP Server for Time https://www.getclockwise.com/mcp

Factory AI CEO @matanSF on why their agents (droids) outperform: Most agents are locked to a single model; Claude code for Claude, Codex for GPT. “We built ours to be fully model-agnostic. Like an engineer fluent in many languages vs just one, they’re more adaptable and https://x.com/tbpn/status/1971322883315314995

Hey Claude: “”Sentient crabs have emerged from the depths and are causing the apocalypse. But work inside my company must go on. Create a powerpoint for the most prosaic and boring meeting that only somewhat alludes to the Crabpocalypse outside.”” (It went with pretty dark humor) https://x.com/emollick/status/1970215922427142323

It’s interesting how “”better at code”” has become the defining goal of almost every AI lab over the last twelve months I think Claude Code getting a bunch of people onto $200/month plans proved that code is one of the most economically valuable applications of this technology”” / X https://x.com/simonw/status/1970147806225854531

More terse than Claude. Works perfectly with Cline’s thinking slider — max it out and the model thinks exactly as much as needed. Full details: https://x.com/cline/status/1970619811853148550

New on the Claude Developer Platform: tool helpers in beta for the Python and Typescript SDKs. Tool helpers simplify tool creation and execution with: – Automatic input validation – A tool runner for automated tool handling in conversations. https://x.com/alexalbert__/status/1968721888487829661

The Claude Code SDK now supports custom tools and hooks directly in code. Additionally, we’ve refreshed all our docs with complete references and 10 new guides on how to utilize the SDK. https://x.com/trq212/status/1966586970458542297

Tri Dao says Claude Code makes him 1.5x more productive and that it’s quite helpful at writing Triton kernels https://x.com/scaling01/status/1970146206203416666

China’s Alibaba just dropped an opensource 30B agentic LLM that outperforms Claude 4 Sonnet, DeepSeek v3.1, Kimi k2 on a range of agentic search benchmarks. Only 3B parameters are activated per token. 100% open-source. https://x.com/unwind_ai_/status/1969053988143477186

New, very needed benchmark from @scale_AI: SWE-Bench Pro Includes: – Multi-file edits – 100+ lines changed on average – Complex dependencies across large codebases Current top model scores: – GPT-5: 23.3% – Claude Opus 4.1: 22.7% – Others drop further (<15%)”” / X https://x.com/alexandr_wang/status/1969805196462358919

Results so far No single model dominates: GPT-5 “high” reasoning leads on tough tasks but collapses on time-critical ones. Claude-4 Sonnet balances speed vs accuracy but at higher cost. Open-source models (like Kimi-K2) show promise in adaptability. Scaling curves plateau, https://x.com/omarsar0/status/1970147904087322661

Claude Sonnet 4 and Opus 4.1 are now available in Microsoft 365 Copilot, bringing Claude’s advanced reasoning capabilities to millions of enterprise users. Read more: https://x.com/AnthropicAI/status/1970907112831328296

@OpenAI Interesting: 1. Linear progress across OpenAI generations (GPT-4o, o3, GPT-5) 2. Claude Opus 4.1 is on top, nearing industry expert, much better than GPT-5 high. Thanks for acknowledging competitors. https://x.com/Yuchenj_UW/status/1971254164069212231

Bring your design context straight into @code with the @figma MCP server! Go from idea to product while staying in the flow. https://x.com/code/status/1970621943821861217

We’re excited to partner with Figma as one of their MCP client partners! https://x.com/allhands_ai/status/1970955961293795831

it’s quite incredible how bad Sonnet 4 is at long-context retrieval Grok-4 > GPT-5 ~ Gemini 2.5 Pro > Claude 4 Sonnet https://x.com/scaling01/status/1970661469667660100

Language Models that Think and Chat Better Proposes a simple RL recipe to improve small open models (eg, 8B) that rivals GPT-4o and Claude 3.7 Sonnet (thinking). Pay attention to this one, AI devs! Here are my notes: https://x.com/omarsar0/status/1971215698140819516

Voice Assist in Action | Try Our AI Voice Assistant | CallRail https://www.callrail.com/voice-assist-in-action

Last week we found an issue with SWE-Bench, allowing agents to cheat by looking at future commits. Instead of celebrating the SWE-Bench Devs for quickly fixing the issue and being transparent, the HN crowd is dunking on them and drawing wildly inaccurate conclusions about”” / X https://x.com/TacoCohen/status/1966421688846778561

What are popular AI coding benchmarks actually measuring? – nilenso blog https://blog.nilenso.com/blog/2025/09/25/swe-benchmarks/

Email, scheduling, rescheduling, time zones— @AprilAssistant can handle it all. Built by @Neha_Suresh_M and @vedhsaka in just 50 days, April is the AI voice agent making productivity real. Thousands of paying users already rely on it every day. https://x.com/ycombinator/status/1969090360023695462

Notion 3.0: Agent is here. Before, Notion was your tool. With 3.0, Notion is your AI teammate. It’s the most advanced knowledge work agent. It handles 20+ minute multi-step actions, works with databases, and has a state-of-the-art memory system. https://x.com/NotionHQ/status/1968734667533611052

Knowledge graph agents might not be ready for prime time, but they are promising. This paper introduces ARK-V1, a lightweight agent that helps LLMs answer questions by actively walking through a knowledge graph instead of relying only on memorized text. Here are my notes: https://x.com/omarsar0/status/1970497643324555664

Paper2Agent brings research papers ‘to life.’ This open tool from @Stanford transforms static papers into interactive AI assistants that can explain and apply their methods. It builds on the MCP and works in 2 layers: – Paper2MCP: Extracts the paper’s methods and code into an https://x.com/TheTuringPost/status/1968829219858956774

Meet Agent Fill, our agentic form filler. Built for finance & ops teams that don’t want to waste time filling out PDFs. Available today in alpha. https://x.com/RampLabs/status/1960799796522049721

Google introduces Test-Time Diffusion Deep Researcher Don’t sleep on diffusion models. Test-Time Diffusion Deep Researcher (TTD-DR) is a deep research agent that models research writing as a diffusion process. Instead of static reasoning or bolted-on tools, the system drafts https://x.com/omarsar0/status/1970864565710921891

How are developers using AI? Inside Google’s 2025 DORA report https://blog.google/technology/developers/dora-report-2025/

We did Google I/O back to back years with NotebookLM and both nights before the show I literally didn’t sleep. The first year of I/O (2023) we were going to live demo Notebook (Project Tailwind) and announce it to the world. Google Labs was fairly unknown and this was our one https://x.com/raizamrtn/status/1968508322329575452

We’re now moving beyond models that react to single instructions and creating systems that can truly tackle problems in a general way – on the path towards solving AGI in the physical world. Developers can now use Gemini Robotics-ER 1.5 via the Gemini API in @GoogleAIStudio. https://x.com/GoogleDeepMind/status/1971243970953879643

Ollama now has a web search API and MCP server! ⚡️ Augment local and cloud models with the latest content to improve accuracy 🔧 Build your own search agent 🔍 Directly plugs into existing MCP clients like @OpenAI Codex, @cline, Goose (@jack) and more! Let’s go!!!! 🧵👇 https://x.com/ollama/status/1971085470785319349

(🧵) Today, we release Meta Code World Model (CWM), a 32-billion-parameter dense LLM that enables novel research on improving code generation through agentic reasoning and planning with world models. https://x.com/syhw/status/1970960837721653409

CWM: An Open-Weights LLM for Research on Code Generation with World Models | Research – AI at Meta https://ai.meta.com/research/publications/cwm-an-open-weights-llm-for-research-on-code-generation-with-world-models/

New from Meta FAIR: Code World Model (CWM), a 32B-parameter research model designed to explore how world models can transform code generation and reasoning about code. We believe in advancing research in world modeling and are sharing CWM under a research license to help empower https://x.com/AIatMeta/status/1970963571753222319

new research from Meta FAIR: Code World Model (CWM), a 32B research model we encourage the research community to research this open-weight model! pass@1 evals, for the curious: 65.8 % on SWE-bench Verified 68.6 % on LiveCodeBench 96.6 % on Math-500 76.0 % on AIME 2024 🧵 https://x.com/alexandr_wang/status/1970973317227225433

Get ready to level up 🎮 Gaming Copilot (Beta) is rolling out now on PCs, with mobile coming next month. Voice mode + screen awareness mean help is just a quick question away – no pausing required – whether you need a tip to beat that level or to remember who that NPC is. https://x.com/mustafasuleyman/status/1968709610602184794

our team has been hard at work to improve the code search experience within GitHub Copilot to give you faster, more accurate code results! we wrote this blog to share more about our Copilot-Embedding model, and specifically details on the training techniques if you like posts https://x.com/pierceboggan/status/1970950784251724007

Two big updates for GitHub Copilot coding agents today. 🚀 Coding Agent is now generally available. And we’ve introduced a new Copilot CLI as part of the GitHub Copilot product suite! https://x.com/lukehoban/status/1971391939858792584

Use Copilot Vision at your own risk 🙁 Used it on my phone this morning to estimate nutrition of my breakfast and turns out I’ve been consuming an insane amount of sugar in my acai every day… Sometimes ignorance is bliss”” / X https://x.com/mustafasuleyman/status/1969053618088329247

📢 New Model(s) Drop: GPT-5 Codex Low, Medium and High are now on Yupp! @OpenAI’s frontier coding models, with adaptive thinking effort and lower token usage. We explored their capabilities with some tough coding prompts: https://x.com/yupp_ai/status/1970617312559669685

Codex CLI 0.39 was released by @OpenAI BIG new feature: Codex CLI now includes /review for automated code reviews! GPT-5-Codex will investigate and find critical bugs in your code. It’s like having another team member always available. https://x.com/mark_k/status/1968934227149291535

Excited to share that OpenAI’s GPT-5-Codex is now live in Windsurf! We’re making it free to paid users for a limited time. We can’t wait to hear what you build with it! https://x.com/windsurf/status/1970549712551100523

GPT-5-Codex is Live in Cline. OpenAI’s agent-optimized version of GPT-5: > 400K context window > adaptive reasoning: 93% fewer tokens on simple tasks, 2x more on complex ones > variable thinking that scales with task complexity > $1.25/$10 per million tokens Built for coding https://x.com/cline/status/1970619799119241709

GPT-5-Codex is live in the Responses API. If you use the Codex CLI via API key, you can now also use GPT-5-Codex.”” / X https://x.com/OpenAIDevs/status/1970535239048159237

GPT-5-Codex is now available in Cursor. Let us know your thoughts!”” / X https://x.com/cursor_ai/status/1970540811168473250

gpt-5-codex is now in the API:”” / X https://x.com/gdb/status/1970631954887565823

GPT-5-Codex is now rolling out to @code developers! https://x.com/pierceboggan/status/1970572801267638421

GPT-5-Codex is relentless and runs until whatever you give it is really done. https://x.com/steipete/status/1969054373385801896

GPT-5-Codex, meet Droid. Optimized for agentic coding and tuned in Factory in collaboration with @OpenAI, we find GPT-5-Codex to be a strong daily-driver. Particular strengths: – Long-running tasks – Autonomous pull requests – Quick questions with adaptive reasoning Available https://x.com/FactoryAI/status/1970549069996302846

GPT-5-Codex, optimized for agentic coding, is rolling out to @code now! Try it out and let us know what you think. https://x.com/code/status/1970579099472056350

I am really liking the new gpt-5-codex. this is its attempt at one shotting minecraft in three.js https://x.com/JasonBotterill3/status/1969730846417629277

I have a Codex project, a Deep Research run, a ChatGPT agent task, & a GPT-5 Pro task all going at the same time. Is this that agent managing thing that some people keep saying is supposed to be the future of work? The current AI interfaces aren’t really up to the task.”” / X https://x.com/emollick/status/1969980811504939443

OpenAI solved the “”make sure your code actually runs”” reward with gpt-5 codex and it shows”” / X https://x.com/andrew_n_carr/status/1969784664912179533

OpenAI tests ChatGPT Agent upgrades powered by new models https://www.testingcatalog.com/openai-tests-chatgpt-agent-upgrades-powered-by-new-alpha-models/

when chatgpt said moondream wasn’t a frontier model, i took it personally”” / X https://x.com/vikhyatk/status/1968811248381784167

Why we built the Responses API https://developers.openai.com/blog/responses-api/

wow! GPT-5-Codex in Cursor just gave me the best reasoning I have ever seen in an LLM. the fact that it was able to recover and continue reasoning multiple times is mindblowing. I’ve seen many models say “but wait” repeatedly and still struggle to find the actual cause. https://x.com/_overment/status/1970630704489803857

Agent run times aren’t everything. I gave the same high level task to Sonnet 4 and GPT-5-Codex. Sonnet completed the task across 7 files in ~6 minutes. Codex is so far at 8 files but 32 minutes in still…. The Sonnet output was perfectly acceptable.”” / X https://x.com/zachtratar/status/1970625784500130065

Aviro (@aviro_ai) makes enterprise AI agents continuously upskill to deliver on complex tasks. Their runtime layer, Cortex, helped their in-house deep research agent beat OpenAI’s by 70% on enterprise search and top Microsoft’s Deep Research benchmark. https://x.com/ycombinator/status/1968691488222503194

DeepSeek’s updated V3.1 Terminus ties with gpt-oss-120b (high) as the most intelligent open weights model and offers increased instruction following and long context reasoning capabilities 🧠 Our benchmarking results indicate DeepSeek V3.1 Terminus shows a greater intelligence https://x.com/ArtificialAnlys/status/1971114096008495501

Sam Altman says ChatGPT will stop talking about suicide with teens | The Verge https://www.theverge.com/ai-artificial-intelligence/779053/sam-altman-says-chatgpt-will-stop-talking-about-suicide-with-teens

there’s this guy on tiktok who video calls chatgpt, shows it objects, asks “how much gram,” then weighs them, and if chatgpt is wrong he puts his phone in the fridge where chatgpt is forced to chat with the condiments. roko’s basilisk is definitely making a note of this man. https://x.com/paularambles/status/1971234855309672467

Shipped. For instance, you can see that gpt-oss-120b is 196 GB right from the “Files” tab https://x.com/mishig25/status/1968598133543256151

The Illusion of Readiness (Health AI) • GPT-5 & peers ace med benchmarks—but stress tests reveal fragility • Guess answers w/o images, flip under trivial prompt tweaks • Fabricate “reasoning” that sounds right but isn’t • Leaderboard wins ≠ real-world readiness https://x.com/arankomatsuzaki/status/1970684893966516477

🚀 LongCat-Flash-Thinking: Smarter reasoning, leaner costs! 🏆 Performance: SOTA open-source models on Logic/Math/Coding/Agent tasks 📊 Efficiency: 64.5% fewer tokens to hit top-tier accuracy on AIME25 with native tool use, agent-friendly ⚙️ Infrastructure: Async RL achieves a https://x.com/Meituan_LongCat/status/1969823529760874935

Excited to introduce: Gamma 3.0 A generational leap for the world’s most popular AI presentation tool. Two major changes: 1. Gamma Agent – with one prompt, you can make sweeping edits across the presentation. a. Say ‘make it more visual’ and it will scan each slide for data https://x.com/thisisgrantlee/status/1967943621782413388

.@Alibaba_Qwen shipping velocity is unmatched Avg 3.5 releases per month, or almost 1 release per week And the majority are open-weights models. Image credt @Smol_AI https://x.com/awnihannun/status/1970839682503348623

[23 Sept 2025] Alibaba Yunqi: 7 models released in 4 days (Qwen3-Max, Qwen3-Omni, Qwen3-VL) and $52B roadmap congrats @Alibaba_Qwen ! https://x.com/Smol_AI/status/1970842828512088486

📢 New Model(s) Drop: Qwen3 VL 235B A22B Instruct & Thinking are now on Yupp! These latest models are @Alibaba_Qwen’s most powerful vision-language models yet. https://x.com/yupp_ai/status/1970640795259851079

🚀 Qwen3-Max is here—no preview, just power! Qwen Chat: https://x.com/Alibaba_Qwen/status/1970599097297183035

🚀 Your Personal AI Travel Designer Is Here! 🍁 Stop wasting hours planning trips. Qwen Chat Travel Planner crafts complete, day-by-day itineraries tailored JUST for you — powered by Amap, Fliggy APIs and Search. ✅ Recommends perfect hotels & transport routes ✅ Builds https://x.com/Alibaba_Qwen/status/1970554287202935159

🚨 Top 10 Open Model Leaderboard Update New open models have entered the Text Arena, and the top 10 rankings by provider have shifted for September! 🔹Qwen-3-235b-a22b-instruct from @Alibaba_Qwen holds the crown at #1 🏆 🔹Longcat-flash-chat from @Meituan_LongCat makes a strong https://x.com/arena/status/1968705194868535749

ChatGPT – Qwen Aug–sep 2025 Timeline (interactive) https://chatgpt.com/canvas/shared/68d3972d363881918f24524394a87d87

Four new releases from Qwen https://simonwillison.net/2025/Sep/22/qwen/#atom-everything

Just enabled full cudagraphs by default on @vllm_project! This change should offer a huge improvement for low latency workloads on small models and efficient MoEs For Qwen3-30B-A3B-FP8 on H100 at bs=10 1024/128, I was able to see a speedup of 47% 🔥 https://x.com/mgoin_/status/1970601094142439761

qwen3-coder-plus is now available on Anycoder Enhanced terminal task capabilities and better performance on Terminal Bench (w/ Qwen Code / Claude Code) SWE-Bench performance up to 69.6 Safer code generation available as Qwen3-Coder-Plus-2025-09-23 https://x.com/_akhaliq/status/1970595669896503462

Qwen3-Max The Tau bench score is insane”” / X https://x.com/scaling01/status/1970599394337587671

Qwen3-Omni-30B-A3B: – instruct – thinking and -captioner https://x.com/scaling01/status/1970182151019659493

So far, Qwen3-Max seems impressive for a non-reasoning model, doing a good job at a lot of my weird tests that even some reasoners struggle with. https://x.com/emollick/status/1970847381966180685

the new Qwen3-Max is now available in anycoder as default as Qwen3-Max-2025-09-23 https://x.com/_akhaliq/status/1970618469344235677

Try the new Qwen models in the Arena!”” / X https://x.com/Alibaba_Qwen/status/1971097727477088717

Qwen (Qwen) https://huggingface.co/Qwen

Inference Providers @huggingface powered by @novita_labs supports Qwen3-VL, the bleeding-edge vision LM 🔥 the model is quite large (22B active 235B total params) so this makes it super easy to try 💚 https://x.com/mervenoyann/status/1971168938848551021

🚨 New Models Alert: WebDev 💻 GPT-5-Codex and Qwen3-Coder-Plus are both now available on WebDev Arena! In the WebDev Arena, you can test out all the best frontier AI coding models on web development tasks. Vote for your preferred response and see how they stack up on the https://x.com/arena/status/1970962780225507775

Most grasping methods fail outside clean lab settings. Open-loop breaks under noise. Closed-loop fails in clutter… Grasp-MPC, from TUM and collaborators, combines model-based MPC with data-driven value functions for robust 6DoF closed-loop grasping… even on moving or cluttered https://x.com/IlirAliu_/status/1970197071123652990

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading