Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Photorealistic wide shot of six Ionic limestone columns on a university quad with a classical entablature spanning the top carved with ‘AGENTS’ in Roman serif letters, six identical bronze messenger statues in the foreground grass positioned mid-stride carrying scrolls, late afternoon golden hour light, warm beige limestone and patinated bronze against green lawn and red brick buildings, sharp focus, campus architecture photography.
Claude Skills might be the biggest upgrade to AI agents so far! Some say it’s even bigger than MCP. I’ve been testing skills for the past 3-4 days, and they’re solving a problem most people don’t talk about: agents just keep forgetting everything. In this video, I’ll share https://x.com/akshay_pachaar/status/1982817709323612628
Code execution with MCP: building more efficient AI agents \ Anthropic https://www.anthropic.com/engineering/code-execution-with-mcp
Apple Plans to Use 1.2 Trillion Parameter Google Gemini Model to Power New Siri – Bloomberg https://www.bloomberg.com/news/articles/2025-11-05/apple-plans-to-use-1-2-trillion-parameter-google-gemini-model-to-power-new-siri
Anthropic just posted another banger guide. This one is on building more efficient agents to handle more tools and efficient token usage. This is a must-read for AI devs! (bookmark it) It helps with three major issues in AI agent tool calling: token costs, latency, and tool https://x.com/omarsar0/status/1986099467914023194
Commerce is entering a new era–powered by AI. Introducing PayPal’s agentic commerce services, helping merchants–especially small businesses–thrive in AI-driven shopping. Includes: 💳 agent ready 🛍️ store sync Learn more 👉 https://x.com/PayPal/status/1983633727138812327
SUPER EXCITED to be working with @privy_io on wallets to enable agentic commerce! fun collab between @MaxSegall and @j_schottenstein — crypto + ai agents in 1 household coming together for a partnership 🤝”” / X https://x.com/LangChainAI/status/1986507333284061340
Agents Rule of Two: A Practical Approach to AI Agent Security https://ai.meta.com/blog/practical-ai-agent-security/
Super energized about this! With our App Builder and Workflow agents, you can now build apps and automate workflows in minutes, right in M365 Copilot chat. Here’s an example. https://x.com/satyanadella/status/1983372239052710047
@Kimi_Moonshot Congratulations to the entire Moonshot team — today is a great day for open source everywhere. We’re excited to continue supporting Kimi models with fast inference on Baseten. https://x.com/basetenco/status/1986494013109903362
@QuixiAI @Kimi_Moonshot a single H200 node is enough😃”” / X https://x.com/vllm_project/status/1986626058897269070
📢 New Model(s) Drop: Kimi K2 Thinking and Kimi K2 Thinking Turbo are now on Yupp! This pair of thinking models from @Kimi_Moonshot specialize in deep reasoning tasks. We explored their capabilities with some prompts on Yupp: https://x.com/yupp_ai/status/1986469027997491422
🚀 Hello, Kimi K2 Thinking! The Open-Source Thinking Agent Model is here. 🔹 SOTA on HLE (44.9%) and BrowseComp (60.2%) 🔹 Executes up to 200 – 300 sequential tool calls without human interference 🔹 Excels in reasoning, agentic search, and coding 🔹 256K context window Built https://x.com/Kimi_Moonshot/status/1986449512538513505
🚨 New Open Source Model Update! Touted for its reasoning and coding strengths, Kimi K2 Thinking by @Kimi_Moonshot is now live for both Text and WebDev in Battle, Side by Side and Direct. Bring your toughest prompts! 💪 The last time Kimi K2 was in the Arena with a new model, https://x.com/arena/status/1986482438768673107
5 Thoughts on Kimi K2 Thinking – by Nathan Lambert https://www.interconnects.ai/p/kimi-k2-thinking-what-it-means
70% on SWE bench verified 30% terminal bench those are two intuitive thresholds for “”actually useful and not frustrating”” coding assistant. Kimi k2 thinking got 71.3% on SWE-Bench Verified 47.1% on Terminal-Bench”” / X https://x.com/andrew_n_carr/status/1986538323876454461
Congrats to the Kimi K2 team on the great numbers on our SWE-bench Verified, SWE-bench Multilingual and SciCode benchmarks!! https://x.com/OfirPress/status/1986475891158040760
Kimi AI – Kimi K2 is Live https://www.kimi.com/
Kimi API is barely alive right now kinda slow ~20tks/s and get quite a few timeouts / network errors when I let the model reason for a long time”” / X https://x.com/scaling01/status/1986476278908920061
Kimi K2 Thinking feels like a big milestone for open-source AI. The first time in a while that open-source gets ahead of proprietary APIs on their big area of focus (agents). Fun to see that it’s happening at a time when the proprietary APIs have the most money/attention”” / X https://x.com/ClementDelangue/status/1986833436607160600
Kimi K2 Thinking https://moonshotai.github.io/Kimi-K2/thinking.html
Kimi K2 Thinking is now available in anycoder https://x.com/_akhaliq/status/1986468663600337125
Kimi K2 Thinking is the new leading open weights model: it demonstrates particular strength in agentic contexts but is very verbose, generating the most tokens of any model in completing our Intelligence Index evals @Kimi_Moonshot’s Kimi K2 Thinking achieves a 67 in the https://x.com/ArtificialAnlys/status/1986911675820446013
Kimi K2 Thinking just launched on Product Hunt! 🥳 Not chasing votes, just using PH as a clean milestone log for our model updates. 🙂 Huge thanks to the helpful team from @ProductHunt https://x.com/crystalsssup/status/1986714377983304137
Kimi-K2 is an exceptional base model GPQA Diamond 77% GPT-4.5 only got 71.4%”” / X https://x.com/scaling01/status/1986112227875954967
Kimi-K2 Reasoning is coming very soon just got merged into VLLM LETS FUCKING GOOOO im so hyped im so hyped im so hyped https://x.com/scaling01/status/1986071916541870399
Kimi-K2 reasoning is landing soon; it just got merged into vLLM https://x.com/cedric_chee/status/1986073808672067725
Kimi-K2 Thinking ranking 19th on SimpleBench improving Kimi-K2s score from 26.3% (rank 33) to 39.6% This makes it the 3rd best open-source model on SimpleBench. Other chinese open-source models like DeepSeek R1 0528 and DeepSeek V3.1 beat it by roughly 1 %. https://x.com/scaling01/status/1986846212050362510
Live in Cline: kimi-k2-thinking https://x.com/cline/status/1986512739490275680
MoonshotAI has released Kimi K2 Thinking, a new reasoning variant of Kimi K2 that achieves #1 in the Tau2 Bench Telecom agentic benchmark and is potentially the new leading open weights model Kimi K2 Thinking is one of the largest open weights models ever, at 1T total parameters https://x.com/ArtificialAnlys/status/1986541785511043536
moonshotai/Kimi-K2-Thinking · Hugging Face https://huggingface.co/moonshotai/Kimi-K2-Thinking
ollama run kimi-k2-thinking:cloud Kimi K2 Thinking is Moonshot AI’s best open-source thinking model. Try it on Ollama’s cloud! https://x.com/ollama/status/1986640693108863271
Our first research paper: custom Mixture-of-Experts (MoE) kernels that make deployment of trillion-parameter models like Kimi K2 viable for the first time on AWS EFA https://x.com/AravSrinivas/status/1986106660386222592
Unsurprisingly, Kimi K2 Thinking is already number one trending on HF. The AI frontier is open-source! https://x.com/ClementDelangue/status/1986827413532057712
🚀 Day 0 support: Kimi K2 Thinking now running on vLLM! In partnership with @Kimi_Moonshot, we’re proud to deliver official support for the state-of-the-art open thinking model with 1T params, 32B active. Easy deploy in vLLM (nightly version) with OpenAI-compatible API: What https://x.com/vllm_project/status/1986455911066706160
It even compares with GPT 5 Pro on some benches. Looks like Kimi’s interpretation of Pro mode is 8 samples + self reflection https://x.com/nrehiew_/status/1986453238552666320
From my tests, Kimi K2 thinking is better than everything Xai, Anthropic, Google has to offer atm. The only thing that is better than this is Gpt 5 codex (at code) and Gpt 5 pro (at high level algorithm design) It beats the SOTA at creative writing by a mile. Good work”” / X https://x.com/karmay007/status/1986454592809529493
Turn on agent mode and ChatGPT can take action for you–research, plan, and get things done while you browse. Now in preview for Plus, Pro, and Business users. https://x.com/OpenAI/status/1984304194837528864
Today we’re rolling out major upgrades that significantly improve the performance and overall outcomes of the Comet Assistant. Comet can now handle more complex, multi-site workflows while working across multiple tabs in parallel. https://x.com/perplexity_ai/status/1986499432410718598
After two years of work, we’ve made an AI Scientist that runs for days and makes genuine discoveries. Working with external collaborators, we report seven externally validated discoveries across multiple fields. It is available right now for anyone to use. 1/5 https://x.com/andrewwhite01/status/1986094948048093389
Edison Scientific (a brand-new company spun out of FutureHouse) releases Kosmos: An AI Scientist for Autonomous Discovery “”Our beta users estimate that Kosmos can do in one day what would take them 6 months, and we find that 79.4% of its conclusions are accurate.”” The paper https://x.com/iScienceLuvr/status/1986023952037417109
Kosmos: An AI Scientist for Autonomous Discovery https://edisonscientific.com/articles/announcing-kosmos
A new security agent called Aardvark:”” / X https://x.com/sama/status/1984002552158154905
Now in private beta: Aardvark, an agent that finds and fixes security bugs using GPT-5. https://x.com/OpenAI/status/1983956431360659467
(32) ElevenLabs CEO: Why Voice is the Next AI Interface – YouTube https://www.youtube.com/watch?v=ZqCEHR4wjxg
A year ago, I would not have expected the first academic field to seem to reach a consensus that AIs will accelerate research (which is not the same thing as autonomous research) would be math But that appears to be happening based on math professors in my feed and elsewhere.”” / X https://x.com/emollick/status/1984388281061282081
Advancing Claude for Financial Services \ Anthropic https://www.anthropic.com/news/advancing-claude-for-financial-services
Gemini has arrived as your hands-free driving assistant in the @GoogleMaps app. Find places along your route, check for EV availability, and share your ETA just by asking. Gemini can also help with multi-step tasks like “find me a restaurant that serves vegetarian tacos within https://x.com/sundarpichai/status/1986119293914792338
Google Gemini’s Deep Research can look into your emails, drive, and chats | The Verge https://www.theverge.com/ai-artificial-intelligence/814878/google-ai-gemini-deep-research-personalized
Introducing the File Search Tool in Gemini API https://blog.google/technology/developers/file-search-gemini-api/
Now, Gemini’s Deep Research can pull in info from @Gmail, @GoogleDrive, and Chat when you connect your @GoogleWorkspace account to give you more context-aware reports. To try it, just select “Deep Research” in Gemini on desktop and choose your sources. Coming to mobile soon.”” / X https://x.com/GeminiApp/status/1986472318873555058
We’ve launched the File Search Tool, a fully managed RAG system integrated into the Gemini API that simplifies grounding models with your private data to deliver more accurate, verifiable responses. – $0.15/m tokens for indexing, free storage and embedding generation at query https://x.com/_philschmid/status/1986506204240347520
Google Maps is getting a powerful boost with Gemini, making navigation smarter and easier. ✨ Learn more about the new features ↓”” / X https://x.com/Google/status/1986164830588248463
Google Maps launches Gemini features, including landmark navigation https://blog.google/products/maps/gemini-navigation-features-landmark-lens/
Here is the story of a remarkable, independent treatment suggestion by GPT-5 Pro: repurposing a known drug for a patient with food protein-induced enterocolitis syndrome (FPIES). First, how we came to test this. My close friend, physician-scientist Dr. Oral Alpan, treated the https://x.com/DeryaTR_/status/1984083644437192737
Tesla shareholders approve $1 trillion pay package for Musk | CNN Business https://edition.cnn.com/2025/11/06/business/musk-trillion-dollar-pay-package-vote
Tesla Shareholders Approve Elon Musk’s $1 Trillion Pay Package – WSJ https://www.wsj.com/business/autos/elon-musk-tesla-pay-package-vote-9abd5a73?st=d8Surv&reflink=desktopwebshare_permalink&mod=tldr
@snyksec + @FactoryAI 🚀 AI is writing code — now it can secure it too. By embedding Snyk Studio into Factory’s AI Droids, every line of AI-generated code is secured at inception and remediated intelligently. 🎥👇 https://x.com/mnair1/status/1986441206046372075
👉Update from Voiceflow HQ: Metadata capabilities now in their Knowledge Base! Add structured tags (e.g., “”locale””) to docs, URLs, or data tables for building precise, context-aware AI agents. Categorize by region, service type, or anything else.”” / X https://x.com/IsaacHandley/status/1985905936553398726
💡 Build. Deploy. Earn. Repeat. Want to deploy your AI agent and actually earn from it? Most devs spend months on monetization. Agent Forge lets you deploy, price, and earn, all in one ecosystem. 🔗 https://x.com/AITECHio/status/1981602229661446634
🤖 Deep Agents JS Deep Agents is now available in JS! Written on top of LangChain and LangGraph 1.0, this brings the power of agents harnesses to the JS ecosystem Comes with planning tools, subagents, and filesystem access Try it out now: npm i deepagents Repo: https://x.com/LangChainAI/status/1986471866857374103
🧩 Takeaway The bottleneck in agent RL wasn’t the policy — it was the experience. DreamGym shows that when we replace heterogeneous real-world rollouts with reasoning-grounded synthetic experience, RL becomes: ✅ scalable & unified ✅ less human engineering ✅ generalizable and”” / X https://x.com/jaseweston/status/1986613060052701397
🚀Mini “”TypeScript AI”” release day! We released a bunch of things in the Lang* ecosystem to make building AI agents in TypeScript easier than ever: 🤖DeepAgents 1.0: https://x.com/hwchase17/status/1986534722504384648
🚨 This might be the biggest leap in AI agents since ReAct. Researchers just dropped DeepAgent a reasoning model that can think, discover tools, and act completely on its own. No pre-scripted workflows. No fixed tool lists. Just pure autonomous reasoning. It introduces https://x.com/rryssf_/status/1983842885784269116
A real gap between what people using chatbots can do and what even non-coders can do with today’s CLI-like tools that have access to their computers, the web & the ability to execute long-term plans, Big opportunity for the AI lab that gets powerful & safe personal agents right”” / X https://x.com/emollick/status/1983710105704276289
AI Agents for the Enterprise | StackAI https://www.stack-ai.com/
AI coding just arrived in Jupyter notebooks – and @brganger (Jupyter co-founder) and I will show you how to use it. Coding by hand is becoming obsolete. The latest Jupyter AI – built by the Jupyter team and showcased at JupyterCon this week – brings AI assistance directly into https://x.com/AndrewYNg/status/1985416763916632124
AI-powered suggestions in @code are now open source! All Copilot functionality is now powered by a single, open-source extension with GitHub Copilot Chat. Learn more about the OSS effort, benefits, and dig into the code: https://x.com/pierceboggan/status/1986462412762194405
Amp – Hamel’s Blog – Hamel Husain https://hamel.dev/notes/coding-agents/amp.html
An Illustrated Guide to AI Agents, with @MaartenGr First 2 chapters now in Early Release! Drafts of the the first two chapters of An Illustrated Guide to AI Agents are now available in Early Release on the O’Reilly platform! These will take you through the central concepts of https://x.com/JayAlammar/status/1983217734641914311
Bear (@usebearai) helps companies get recommended by AI Agents like ChatGPT, Google AI Mode, Perplexity, Cursor, Claude Code, and more. It’s already used by Browserbase, WisprFlow, and 200+ companies to grow traffic from AI Agents. https://x.com/ycombinator/status/1983564490953372099
Beyond Autocomplete: How Agentic AI Is Rewriting Enterprise Software – Tabnine https://www.tabnine.com/webinar/beyond-autocomplete-how-agentic-ai-is-rewriting-enterprise-software/
Big news! We made the basic tier of the OpenHands Cloud FREE! This means that you can call state-of-the-art coding agents from your computer, phone, github, gitlab, slack, etc. for just the price of API credits or hosting your own language model! 🧵👇”” / X https://x.com/gneubig/status/1986071169263370711
CodeClash is our new benchmark for evaluating coding abilities- it’s much harder than anything we’ve built before. LMs must manage an entire codebase and develop it to compete in challenging arenas against other LM-generated programs Current LMs really struggle, lots to do here!”” / X https://x.com/OfirPress/status/1986095773843390955
Coding Agents Are Outliers https://vivekhaldar.com/articles/coding-agents-are-outliers/
Cognition | Windsurf Codemaps: Understand Code, Before You Vibe It https://cognition.ai/blog/codemaps
Composer: Building a fast frontier model with RL · Cursor https://cursor.com/blog/composer
Fundamentals of Building Autonomous LLM Agents Great overview of LLM-based agents. Great if you are just getting started with AI agents. This covers the basics good. https://x.com/omarsar0/status/1981793327956865504
GitHub Copilot Orchestra Pattern… A multi-agent orchestration system for structured, test-driven software development with AI assistance. Conductor -> Plan 🔁 ( implement -> review -> commit ) https://x.com/code/status/1986622178146562300
Graph-based Agent Planning It lets AI agents run multiple tools in parallel to accelerate task completion. Uses graphs to map tool dependencies + RL to learn the best execution order. RL also helps with scheduling strategies and planning. Major speedup for complex tasks. https://x.com/omarsar0/status/1983892163990843692
Harbor https://harborframework.com/
I’m a bit giddy over the fact that this is by all visible measures a frontier level model, if not THE frontier model, for agentic tasks. And you can run it. In it’s native precision. On 2 M3 Ultras. Pretty fast. In MLX.”” / X https://x.com/awnihannun/status/1986603178251722850
Improving agent with semantic search · Cursor https://cursor.com/blog/semsearch
Inside LangSmith’s No Code Agent Builder We built a no code agent builder, and our team is breaking down why we did it + what’s powering it under the hood. In our latest roundtable discussion, Harrison and the engineering team behind LangSmith Agent Builder sit down to discuss: https://x.com/LangChainAI/status/1983916519513059728
Introducing Cursor 2.0. Our first coding model and the best way to code with agents. https://x.com/cursor_ai/status/1983567619946147967
Introducing Multi-Agent Evolve 🧠 A new paradigm beyond RLHF and RLVR: More compute → closer to AGI No need for expensive data or handcrafted rewards We show that an LLM can self-evolve — improving itself through co-evolution among roles (Proposer, Solver, Judge) via RL — all https://x.com/youjiaxuan/status/1983293231879393695
It is clear developers need a unified experience for all coding agents. We introduced a new VS Code primitive view to address this: Agent sessions. From one view, kick off, monitor, and review all of your coding agents in @code – local or remote, from Copilot or agents like”” / X https://x.com/pierceboggan/status/1986116693819859024
LangSmith Agent Builder is our no code agent builder for anyone to create an agent, now available in private preview. This is not a workflow builder. Built on our Deep Agents architecture, LangSmith Agent Builder handles planning, memory, and sub-agents automatically. This means https://x.com/LangChainAI/status/1983568636079112233
Last week we launched a updates to @code to support working with agents. There’s a lot of terminology and labels here: subagents, handoffs, background tasks, custom agents, agent sessions, etc. I’m curious, what concepts or names are unclear? How can we help connect the dots?”” / X https://x.com/jo_parkhurst/status/1986136483892507119
Love this breakdown of the latest @code features from #GitHubUniverse from @burkeholland https://x.com/JamesMontemagno/status/1986106739612385493
Meet AgentOS — your production runtime for multi-agent systems. > Serve your Agents as an API > Runs entirely in your cloud > No data leaves your system > Completely private > 100% Open Source Check it out 👇 https://x.com/ashpreetbedi/status/1982871453163733497
New guide on RL for agentic environments. This guide integrates OpenEnv, textarena, and TRL for training language models on reasoning games like wordle. Instead of relying only on static reward functions, you can now hook up your model to interactive environments (browsers, https://x.com/ben_burtenshaw/status/1985368549720817953
New video: Build a streaming @LangChainAI agent in @nextjs using useStream + memory 🚀 You’ll learn: – stream AI replies into your UI with useStream – Minimal API route serving SSE – Add conversation memory via thread id + checkpointer 🎥 Watch now: https://x.com/bromann/status/1986491398665929209
Our new bounding box approach in LlamaParse gives you clean bounding boxes while preserving clean reading order of the text through agentic reconstruction. The issue with traditional parsing methods is that the quality of the output is directly dependent on the layout detector – https://x.com/jerryjliu0/status/1986539080331727226
Profiling with Cursor 2.0: The Missing Layer in AI Code Generation – Ryan Perry https://ryanperry.io/post/cursor-profiling-missing-layer
Scaling Agent Learning via Experience Synthesis 📝: https://x.com/jaseweston/status/1986613046047846569
Semantic search improves our agent’s accuracy across all frontier models, especially in large codebases where grep alone falls short. Learn more about our results and how we trained an embedding model for retrieving code. https://x.com/cursor_ai/status/1986124270548709620
SOTA hug of death”” / X https://x.com/code_star/status/1986476817755361505
This Instagram Reels AI agent is absolutely wild 🤯 It scrapes trending Reels in your niche, analyzes them with AI, and extracts every creative insight you need. All inside n8n + Airtable. Perfect for DTC brands & agencies who need to know what’s working on Instagram before https://x.com/mikefutia/status/1982488530908479964
Today, we’re announcing the next chapter of Terminal-Bench with two releases: 1. Harbor, a new package for running sandboxed agent rollouts at scale 2. Terminal-Bench 2.0, a harder version of Terminal-Bench with increased verification https://x.com/alexgshaw/status/1986911106108211461
Tongyi DeepResearch: A New Era of Open-Source AI Researchers | Tongyi DeepResearch https://tongyi-agent.github.io/blog/introducing-tongyi-deep-research/
Tools-to-Agent Retrieval This work presents a unified vector space embedding of both tools and agents with metadata links. Enables fine-grained tool-level and agent retrieval. It’s a great context engineering approach in that it retrieves at both the tool and agent levels https://x.com/omarsar0/status/1985745152204554720
Unlocking the Power of Multi-Agent LLM for Reasoning Designing and optimizing multi-agent systems is important. This paper analyzes multi‑agent systems where one meta‑thinking agent plans and another reasoning agent executes, and identifies a lazy agent failure mode. They find https://x.com/omarsar0/status/1986831275144138756
We built a full-stack AI agent that can generate entire apps from a prompt. It provisions Postgres, adds auth, commits code, and deploys. Copy our architecture 🧵 it’s open source https://x.com/neondatabase/status/1983609414671384716
We’re hiring research interns at @OpenHandsDev! If you’re passionate about doing cutting-edge research in AI agents (with publication encouraged), come work with me, @xingyaow_ and the rest of the team. We welcome referrals too! https://x.com/gneubig/status/1985428673806135698
We’ve shipped several quality-of-life improvements to Cursor! Details below…”” / X https://x.com/cursor_ai/status/1985791854739390591
What a fun little breakfast chat with @Swyx Windsurf just released Fast Context. Unlike traditional agentic search, Fast Context searches and retrieves relevant code from your codebase 20x faster. Game changing for developers who want to stay in flow. Watch it in action https://x.com/SarahChieng/status/1985410447538114771
🎂 MCP turns 1 on Nov 25 We’re throwing it the biggest birthday party in AI history @AnthropicAI + Gradio are bringing together thousands of developers for 17 days of building MCP servers and Agentic apps $500K+ free credits for participants. $17.5K+ in cash prizes. Nov 14-30 https://x.com/Gradio/status/1985446956034830495
New on the Anthropic Engineering blog: tips on how to build more efficient agents that handle more tools while using fewer tokens. Code execution with the Model Context Protocol (MCP): https://x.com/AnthropicAI/status/1985846791842250860
On Skills vs. Subagents Been using Skills extensively for the past couple of days. But like many other Claude Code users, I’ve been thinking about the difference between subagents and Skills. Subagents are useful for handing off subtasks (i.e., separation of concerns) from”” / X https://x.com/omarsar0/status/1981798842866557281
Unlock the full potential of your AI chat applications with Reka’s powerful and completely free MCP Server! This video guides you through integrating online search capabilities, AI-powered fact-checking, and more directly into your VS Code projects. GitHub: https://x.com/RekaAILabs/status/1985794490116780052
You can now deploy any ML model, RAG, or Agent as an MCP server. And it takes just 10 lines of code. Here’s a breakdown, with code (100% private):”” / X https://x.com/_avichawla/status/1985595667079971190
New Version of Siri to ‘Lean’ on Google Gemini – MacRumors https://www.macrumors.com/2025/11/02/new-version-of-siri-to-lean-on-google-gemini/
🚨 WebDev Leaderboard Update MiniMax-M2 from @MiniMax__AI has landed as the #1 open model! A 230B MoE model with 10B-active-parameters, it’s an open source model built for efficient, high-performance coding, reasoning, and agentic-style tasks. It also ranks #4 in WebDev https://x.com/arena/status/1985465603206107318
Introducing SWE-1.5, our fast agent model. It achieves near-SOTA coding performance while setting a new standard for speed. Now available in Windsurf. https://x.com/windsurf/status/1983667319944712460
Thanks @_akhaliq sharing our work! 🚀 Glad to introduce our newest work — VCode! 🎨 VCode: A Multimodal Coding Benchmark with SVG as Symbolic Visual Representation For decades, RGB pixels have been the default medium for representing images. But in the agentic era, how can we”” / X https://x.com/KevinQHLin/status/1986126304316411928
Today we’re releasing SWE-1.5, our fast agent model. It achieves near-SOTA coding performance while setting a new standard for speed. Now available in @windsurf. https://x.com/cognition/status/1983662836896448756
VCode a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation https://x.com/_akhaliq/status/1986073575216824650
Real shortage of good papers, even working papers, testing agentic (post o3) and Deep Research AI outputs in law, medicine, business, coding, etc. Assume when a research paper discusses AI it means GPT-4o (with occasional Gemini 2.5 or o1) for the next year or so.”” / X https://x.com/emollick/status/1985417078749479212
The fight over what a product manager does is going to among the first major AI organizational design problems faced by companies. It is a job right at the center of what AI can do with first pass coding, marketing, design, ideation, etc. Who will own which pieces up for grabs.”” / X https://x.com/emollick/status/1984682445946421505
This is the headline: K2 reasoning allegedly cost $4.6M to train. https://x.com/dbreunig/status/1986633012336099835
1/ AI agents can now pay with stablecoins. We’ve teamed up with @LangChainAI, the leading agent framework, to seamlessly provision wallets for agents so they can transact onchain. Stablecoins are programmable money. Privy makes them safe for agents to use. Here’s how 👇 https://x.com/privy_io/status/1986503492547039502
Perplexity Assistant will be able to join your meetings soon https://www.testingcatalog.com/perplexity-assistant-will-be-able-to-join-your-meetings-soon/
Ready to build sophisticated AI agents? This new Google Cloud Community article walks through creating “”Aura,”” a gen AI travel concierge, from scratch. Learn how to use ADK, Gemini, and BigQuery to build agents that can reason, plan, and act → https://x.com/GoogleCloudTech/status/1981854969796891010
vibe coding should be taught in every school and university around the world, first day on campus, welcome to the future”” / X https://x.com/OfficialLoganK/status/1984645815361814672
you can now push, run, and pull agentic environments from the hugging face hub as spaces. the workflow looks like this: – do `openenv init` to start a new environment from a template – build/port your rl environment – use `openenv push` to push the environment to hub. – from https://x.com/ben_burtenshaw/status/1986097540068950149
AI Agents will completely transform how we interact with corporate knowledge by applying reasoning across multiple data sources. Here’s an example of ChatGPT’s new knowledge feature to pull information from Box and Gmail to combine it an answer a question. https://x.com/levie/status/1981479811500724297
Say hello to Pinterest Assistant: Revolutionizing the way you shop online | Pinterest Newsroom https://newsroom.pinterest.com/news/pinterest-assistant-revolutionizing-the-way-you-shop-online/
For an AI assistant to be truly personal, it’s important to keep as much data as possible locally on the users device, and not on Perplexity servers. Account credentials, such as passwords and credit card information, are also stored locally on the user’s device.”” / X https://x.com/perplexity_ai/status/1985376891763925064
Microsoft Copilot has added some interesting features, but I still struggle with the key problem: I cannot figure out any way to trigger GPT-5 Thinking Extended/Claude 4.5 Sonnet level responses No matter what I do, no deep thinking , no agentic actions, no document outputs, etc https://x.com/emollick/status/1983962613576036736
Ant AQ-Team @AQ_MedAI @TheInclusionAI and SGLang RL Team @sgl_project just helped land Kimi-K2-Instruct RL on slime — fully wired up and running on 256× H20 141GB 🚀 Huge shout-out to @yngao016, @menlzy, @Yonah_x from AQ Team and @Ji_Li_233, @Yefei_RL from the SGLang RL Team for”” / X https://x.com/slime_framework/status/1986811354502906304
Fixed the token generation speed on https://x.com/Kimi_Moonshot/status/1986754111992451337
Here’s the command I ran: “` mlx.launch –hosts first.ip,second.ip –env MLX_METAL_FAST_SYNCH=1 mlx-lm/mlx_lm/examples/pipeline_generate.py –model mlx-community/Kimi-K2-Thinking –prompt “”Write an HTML and JavaScript page implementing space invaders”” -m 16384 “` PR here”” / X https://x.com/awnihannun/status/1986602098017116357
The new 1 Trillion parameter Kimi K2 Thinking model runs well on 2 M3 Ultras in its native format – no loss in quality! The model was quantization aware trained (qat) at int4. Here it generated ~3500 tokens at 15 toks/sec using pipeline-parallelism in mlx-lm: https://x.com/awnihannun/status/1986601104130646266
Gemini 3.0 Ultra or Gemini 3.0 Pro? Which is it and why do you think that? It sounds too big for Pro but too small for Ultra, but since models just get sparser and sparser I believe it’s Pro and very similar to Kimi-K2 1.2T@30B. Also as of right now Ultra is still just a”” / X https://x.com/scaling01/status/1986161974883860486
Team from Ant Group @TheInclusionAI helped land Kimi model @Kimi_Moonshot on @Zai_org’s slime framework! Open AIs help Open AIs ♥️ CN AIs help CN AIs ♥️”” / X https://x.com/bigeagle_xd/status/1986815075785879723
You can now use OpenAI Codex directly in @code Insiders with your GitHub Copilot Pro+ subscription. Learn more: https://x.com/code/status/1985449714540572930
New multi-year, strategic partnership with @OpenAI will provide our industry-leading infrastructure for them to run and scale ChatGPT inference, training, and agentic AI workloads. Allows OpenAI to leverage our unusual experience running large-scale AI infrastructure securely, https://x.com/ajassy/status/1985351258333643172
Copilot or @OpenAI Codex? Why not both… https://x.com/code/status/1986113028387930281
Perplexity launched Perplexity Patents, a new IP intelligence research Agent! “”While in beta, Perplexity Patents will be free for all users. Pro and Max subscribers will receive additional usage quotas and model configuration options.”” Now I want the same for News sources 👀 https://x.com/testingcatalog/status/1983885677835014270
In a world where AI chatbots are increasingly common, most still suffer from the same limitation: They’re text in, text out. But what if your AI could dynamically decide not just what to say, but how to show it? Elysia, our open source agentic RAG app, dynamically decides https://x.com/weaviate_io/status/1986463667160822206
DS-STAR is a state-of-the-art data science agent designed to autonomously solve complex data science problems. It automates tasks from analysis to data wrangling across diverse data types to achieve top performance on challenging benchmarks. Learn more: https://x.com/GoogleResearch/status/1986491681571807584
Auth0 | Securing AI Agents | The New Identity Challenge | Auth0 https://auth0.com/resources/whitepapers/securing-ai-agents-the-new-identity-challenge
Chaining ffmpeg with a Browser Agent https://100x.bot/a/chaining-ffmpeg-with-browser-agent
1) Docker Setup Deploy the Graphiti MCP server using Docker Compose. This setup starts the MCP server with Server-Sent Events (SSE) transport, and it includes a Neo4j container, which launches the database as a local instance. This configuration also lets you query and https://x.com/_avichawla/status/1985958018580955354
I am super excited to release this first complete version of mcp2py with support for OAuth! MCP has a true chance of becoming the next distribution standard, just like mobile apps and websites did before it. If it indeed becomes the standard where all companies that have a https://x.com/MaximeRivest/status/1985200460194627948
I’ve been really impressed with the way Claude Code has started asking questions when you’re in Plan mode. So i was curious if i could ask Claude to reveal how this works. I asked Claude – “”tell me about the tool you have for plan mode where you ask questions to the user. how https://x.com/nityeshaga/status/1985707959486472268
Big update for Claude Desktop and Cursor users! Now you can connect all AI apps via a common memory layer in a minute. I used the Graphiti MCP server that runs 100% locally to cross-operate across AI apps like Claude Desktop and Cursor without losing context. (setup below) https://x.com/_avichawla/status/1985958015452020788
This is a surprisingly revealing test prompt: “Write a paragraph that startles me with its brilliance and really demonstrates your capabilities across as many dimensions as possible. Then explain what you did.” Claude excels at writing, GPT-5 Pro nails intellectual tricks, etc. https://x.com/emollick/status/1984827923363230182
🏆NEW LMARENA LEADERBOARDS🏆 🤓Experts 💻 Software & IT Services ✍️ Writing, Literature, & Language 🔬 Life, Physical, & Social Science 🎭 Entertainment, Sports, & Media 📈 Business, Management, & Financial Ops 🧮 Mathematical ⚖️ Legal & Government 🩺 Medicine & Healthcare https://x.com/ml_angelopoulos/status/1986154276499104186
We’ve released an early preview of Qwen3-Max-Thinking–an intermediate checkpoint still in training. Even at this stage, when augmented with tool use and scaled test-time compute, it achieves 100% on challenging reasoning benchmarks like AIME 2025 and HMMT. You can try the https://x.com/Alibaba_Qwen/status/1985347830110970027
We’re announcing our first restricted donations from labs to support building ARC-AGI-3 – @ndea: @mikeknoop @fchollet – @xai: @elonmusk – @Googleorg: @divy93t – @NousResearch: @theemozilla, @rogershijin – @PrimeIntellect: @willccbb Launching Q1 ’26″” / X https://x.com/GregKamradt/status/1985804827063210244
Introducing the Gemini Docs MCP Server, a local STDIO server for searching and retrieving Google Gemini API documentation. This should help you build with latest SDKs and model versions. 🚀 – Run the server directly via uvx without explicit installation. – Performs full-text https://x.com/_philschmid/status/1985363147071386048
Google announces support for JSON Schema and implicit property ordering in Gemini API. https://blog.google/technology/developers/gemini-api-structured-outputs/
We just shipped some nice Gemini API updates for developers using Structured Outputs. The API now supports: – $ ref for recursive schemas – anyOf union types – min + max numerical constraints – null types – property ordering adherence And much more!”” / X https://x.com/OfficialLoganK/status/1986128461728260210
GitHub Copilot code completions now deliver: ⚡️ 3x higher token-per-second throughput 🧠 12% higher acceptance rate 📉 35% reduced latency Here’s how we built a faster, smarter custom model. 🤖 https://x.com/github/status/1985737580613140747
Whisper Into This AI-Powered Smart Ring to Organize Your Thoughts | WIRED https://www.wired.com/story/sandbar-stream-smart-ring/
Codex has really fueled my tendency of working on multiple things at the same time. I love that I can have it work on multiple things at the same time and don’t spend a lot of time thinking about prompting because Codex just gets me. – Talking about a feature in a slack thread?”” / X https://x.com/dkundel/status/1984367778154127465
Codex has transformed how OpenAI builds over the last few months. Have some great upcoming models too. Amazing work by the team!”” / X https://x.com/sama/status/1985814135784042993
If you want to use more Codex after you hit your subscription limits, you can now buy credits as needed. This is something we expect to do for compute-intensive features; it will let us keep subscription prices low for most users and let the rest of you go wild.”” / X https://x.com/sama/status/1983998115603734843
The new “thinking” UX for GPT-5 is a big improvement, but it is annoying that it auto scrolls to the bottom with each new entry, making it hard to read the thinking trace.”” / X https://x.com/emollick/status/1984378927654322420
We promised unprecedented transparency for Codex and to take the reports of degradation seriously, despite seeing incredible growth week over week. Here is our report and what we have found over the last seven days https://x.com/thsottiaux/status/1984465716888944712
You can now get more Codex usage from your plan and credits with three updates today: 1️⃣ GPT-5-Codex-Mini — a more compact and cost-efficient version of GPT-5-Codex 2️⃣ 50% higher rate limits for ChatGPT Plus, Business, and Edu 3️⃣ Priority processing for ChatGPT Pro and”” / X https://x.com/OpenAIDevs/status/1986861734619947305
You’ve asked for more flexible ways to get more Codex usage: Introducing credits for Codex on ChatGPT Plus and Pro. Credits give you more usage beyond what’s included in your plan, kicking in when you hit limits. As a bonus, we also reset Codex rate limits for everyone. Enjoy!”” / X https://x.com/OpenAIDevs/status/1983956896852988014
Gemini Canvas now supports slide creation and can be integrated with Google Slides. https://x.com/arrakis_ai/status/1983508958657892435





Leave a Reply