Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: 1961 Ferrari 250 GT California Spyder in a premium garage, empty driver seat with steering wheel turning autonomously, subtle holographic navigation waypoints floating in windshield, cinematic lighting with warm highlights on Rosso Corsa paint and chrome, elegant reflections on polished floor, soft depth of field, studio automotive photography style
Equipping agents for the real world with Agent Skills \ Anthropic https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills
Today we introduced Gemini Enterprise, built with our most advanced Gemini models. It allows you to chat with your company’s documents, data and apps as well as build and deploy AI agents, all grounded in your information and context. Have a look at how it helps you build an https://x.com/sundarpichai/status/1976338416611578298
Cranston AI (@cranston_ai) does your company’s bookkeeping & taxes with AI. Their agents pull in context from across the business and, after human review, file a full corporate tax return with the IRS. https://x.com/ycombinator/status/1975591950255358411
Today, we’re one step closer to AI as an operating system. A computer you can talk to, that can see what you see, and take action – all with your permission, all more intuitive than ever. Vision now GA globally + more on today’s @Windows blog: https://x.com/mustafasuleyman/status/1978808627008847997
Customize Claude Code with plugins \ Anthropic https://www.anthropic.com/news/claude-code-plugins
Claude Skills are awesome, maybe a bigger deal than MCP https://simonwillison.net/2025/Oct/16/claude-skills/
Claude Skills: Customize AI for your workflows \ Anthropic https://www.anthropic.com/news/skills
Today we’re introducing Skills in claude dot ai, Claude Code, and the API. Skills let you package specialized knowledge into reusable capabilities that Claude loads on demand as agents tackle more complex tasks. Here’s how they work and why they matter for the future of agents: https://x.com/alexalbert__/status/1978877498411880550
I am not going to lie. I see a lot of potential in the Skills feature that Anthropic just dropped! Just tested with Claude Code. It leads to sharper and precise outputs. It’s structured context engineering to power CC with specialized capabilities, leveraging the filesystem. https://x.com/omarsar0/status/1978919087137804567
Claude and your productivity platforms \ Anthropic https://www.anthropic.com/news/productivity-platforms
Preparing for AI’s economic impact: exploring policy responses \ Anthropic https://www.anthropic.com/research/economic-policy-responses
Claude can now create and use files \ Anthropic https://www.anthropic.com/news/create-files
Big progress on this important benchmark (but still weird artifacts). https://x.com/emollick/status/1976702663330038205
Walmart is moving very fast in AI. Amazon still seems to block ChatGPT agents from even visiting its site. Interesting reversal in agentic commerce, a lesson learned from e-commerce where Amazon moved fast?”” / X https://x.com/emollick/status/1978130496207888717
ChatGPT instant checkout for Walmart:”” / X https://x.com/gdb/status/1978123494870196228
Walmart teams up with OpenAI to allow purchases in ChatGPT https://www.cnbc.com/2025/10/14/walmart-openai-chatgpt-shopping.html
welcome @Walmart to instant checkout 🤝”” / X https://x.com/bradlightcap/status/1978116720171643127
More Articles Are Now Created by AI Than Humans https://graphite.io/five-percent/more-articles-are-now-created-by-ai-than-humans
Scientific frontiers of agentic AI – Amazon Science https://www.amazon.science/blog/scientific-frontiers-of-agentic-ai
Introducing: The fastest web agent in the world⚡[ bu 1.0 ] We built a special LLM that reduces the latency by 6x while keeping the same performance. The agents can now take 20 steps per minute. How many can a human do? 👨💻 ⚡Go give it a try. It feels crazy⚡ Read the comment https://x.com/gregpr07/status/1976153187167195370
Dari (@use_dari) built the easiest way to build reliable browser agents. Dari learns a workflow once, then caches the deterministic DOM steps. They also handle 2FA via text codes, TOTP, or email. Congrats on the launch, @avyvar and @benhong03! https://x.com/ycombinator/status/1975985451308818885
We just solved authentication for AI Agents. Announcing 1Password + Browserbase, enabling secure agentic password autofill for your browser agents. Available exclusively on Director dot ai and Browserbase. Full post below. https://x.com/browserbase/status/1975964196329386278
So this is a first for me. I just had a pretty big refactoring session with codex-cli, and eventually it started going completely off the rails. It became very dumb, made bad mistakes, only followed half my instructions and made up the other half, misused tools in ways so stupid https://x.com/giffmana/status/1978560796238991422
The Claude app is now a surprisingly good personal assistant. For whatever reason, Sonnet 4.5 is much better at working with Gmail/Google Calendar than other AI, even though most other models can access your email, they seem to do minimal work while Claude goes much deeper/wider”” / X https://x.com/emollick/status/1978101986357662156
Surfer 2 is here. 🏄🏄 Our new Cross-Platform Computer-Use Agent exceeds state-of-the-art on the 4 main benchmarks: WebVoyager, AndroidWorld, WebArena and OSWorld. 🖥️🌐📱 Find out more at https://x.com/hcompany_ai/status/1978935436111229098
AI can be confusing. How do you teach people that asking default GPT-5 a question and following up by asking for links to its sources will result in hallucinated cites while asking GPT-5 Thinking to answer a question and provide sources will get you accurate citations & links?”” / X https://x.com/emollick/status/1976308653339689206
RE: the agent/workflow debate Agents and workflows are a spectrum. A system can be more or less ‘agentic’. A pure ‘agent’ is too volatile to be sent to production – you need a bit of determinism to rein it in. https://x.com/mattpocockuk/status/1975847938396983790
Built an Agent Builder in @v0 Image and text models, javascript, export to code, conditionals, and API requests are all supported. Zero code written. https://x.com/max_leiter/status/1976002883775758482
We’re entering a new era—where AI is built into the tools people already use every day. With today’s updates, every Windows 11 PC becomes an AI PC, with Copilot at the center, ready to help you think, create, and act. And it all starts with “Hey Copilot.” https://x.com/yusuf_i_mehdi/status/1978808604200259785
This is starting to turn into a real pattern: search, code execution, and the ability to recursively call sub-agents are the Big 3 tools that can make almost any agent more effective.”” / X https://x.com/corbtt/status/1978595347833188591
Introducing SWE-grep and SWE-grep-mini: Cognition’s model family for fast agentic search at >2,800 TPS. Surface the right files to your coding agent 20x faster. Now rolling out gradually to Windsurf users via the Fast Context subagent – or try it in our new playground! https://x.com/cognition/status/1978867021669413252
super excited to release SWE-grep and SWE-grep-mini! SWE-grep-mini achieves extreme inference speeds of >2,800 TPS: 20x faster than Haiku 4.5 while beating Sonnet 4.5, Opus 4.1 & GPT-5 on our CodeSearch eval our vision: make agentic search as fast as embedding search. the https://x.com/silasalberti/status/1978871477605929229
ok so let me explain why subagents kill long context Like you can spend $500m building 100 million context models, and they would be 1) slow, 2) expensive to use, 3) have huge context rot. O(n) is the lower bound. Cog’s approach is something you learn in day 1 of @CS50 – https://x.com/swyx/status/1978874342743343254
In Pydantic AI 1.1.0 you can now use @PrefectIO to orchestrate your agents. Loved this collaboration!”” / X https://x.com/AAAzzam/status/1978611295243981273
Excited to announce @codegen is now live on @clickup ! Horizontal software is the ideal playground for agents – and we’ve designed a tight integration. Watch @codegen hop across your notes, tasks, whiteboards, etc. and deliver ready-to-ship code ⚡️ No CS degree required. https://x.com/mathemagic1an/status/1978563833275744364
Is this what happens when AI assistants mediate all our interactions in the physical & digital world? https://x.com/bilawalsidhu/status/1976348541401563459
Is there a good guide to using the AI CLI tools for non-coders? Either for non-coding uses (data analysis, automating work) or vibe-prototyping? I don’t want to have to write one myself, but there are a ton of capabilities in computer use with files that would be useful for many”” / X https://x.com/emollick/status/1976413226050212107
Introducing Scenes: your new way to tell stories on c.ai https://blog.character.ai/introducing-scenes-your-new-way-to-tell-stories-on-c-ai/
Chrombot ( https://x.com/Eval_Engine/status/1958899009822892197
We have a new state-of-the-art result on TheAgentCompany from Shanghai AI lab: MUSE + Gemini 2.5, solving 41.1% of the real-world inspired tasks. The new method is based on “”learning on the job””, a memory-based method. https://x.com/gneubig/status/1978564697499574761
Announcing my new course: Agentic AI! Building AI agents is one of the most in-demand skills in the job market. This course, available now at https://x.com/AndrewYNg/status/1975614372799283423
Why America Builds AI Girlfriends and China Makes AI Boyfriends https://www.chinatalk.media/p/why-america-builds-ai-girlfriends
Introducing the first-ever Google Agents Development Kit (ADK) Community Call. Come meet the team, learn about our AI Agents roadmap and ask any question you may have. Happening next Wednesday at 9:30-10:30am PT (link below) https://x.com/Saboo_Shubham_/status/1976334618715619573
DevDay AMA on Reddit tomorrow with some of the people behind the ships— AgentKit Apps SDK Sora 2 in the API GPT-5 Pro in the API Codex and more. Tomorrow, 11 AM PT. https://x.com/OpenAI/status/1976057496168169810
Codex is so good, and is going to get so amazing. I am having a hard time imagining what creating software at the end of 2026 is going to look like.”” / X https://x.com/sama/status/1977177505971736675
We just dropped our Open Agent Builder example app 👀 A 100% open source n8n style workflow builder powered by Firecrawl, @LangChainAI, @convex_dev, @ClerkDev and more. Check it out!”” / X https://x.com/firecrawl_dev/status/1978878728827478289
Cua ❤️ open-weight models Highly requested by the Discord community, we tested Moondream3 and Salesforce GTA-1 for UI grounding in computer-use agents 1/3 https://x.com/trycua/status/1976001242901119401
Agents don’t just chat — they act by fetching data, sending messages, calling APIs, and updating records, which makes securing them a whole new challenge. Our latest blog post breaks down how to implement authentication and authorization for agents. 🔒In this blog post, learn https://x.com/LangChainAI/status/1978121116867567644
Why does RL work for enhancing agentic reasoning? This paper studies what actually works when using RL to improve tool-using LLM agents, across three axes: data, algorithm, and reasoning mode. Instead of chasing bigger models or fancy algorithms, the authors find that real, https://x.com/omarsar0/status/1978112328974692692
[2509.25140] ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory https://arxiv.org/abs/2509.25140
RIP fine-tuning ☠️ This new Stanford paper just killed it. It’s called ‘Agentic Context Engineering (ACE)’ and it proves you can make models smarter without touching a single weight. Instead of retraining, ACE evolves the context itself. The model writes, reflects, and edits https://x.com/rryssf_/status/1976269613072843063
Is ACE the next Context Engineering Technique? ACE (Agentic Context Engineering) is a new framework that beats current state-of-the-art optimizers like GEPA by treating context as an evolving, structured space of accumulated knowledge. What is ACE? ACE treats context as an https://x.com/_philschmid/status/1977618096383721725
🚀 Introducing 𝐀𝐠𝐞𝐧𝐭 𝐒3, the most advanced computer-use agent, now 𝐚𝐩𝐩𝐫𝐨𝐚𝐜𝐡𝐢𝐧𝐠 𝐡𝐮𝐦𝐚𝐧-𝐥𝐞𝐯𝐞𝐥 𝐩𝐞𝐫𝐟𝐨𝐫𝐦𝐚𝐧𝐜𝐞🧠💻 Just one year ago, Agent S scored ~20% on OSWorld: SOTA then, but far from human 72%. Today, Agent S3 reaches 6̳9̳.̳9̳%̳ (⬆10% over https://x.com/xwang_lk/status/1973914981591838841
💃New Multi-Agent RL Method: WaltzRL💃 📝: https://x.com/jaseweston/status/1978185306999341256
we raised a $13m seed round from 120+ of Silicon Valley’s top investors for @mastra_ai, the leading TypeScript agent framework https://x.com/calcsam/status/1976346378013147359
[2510.08558] Agent Learning via Early Experience https://arxiv.org/abs/2510.08558
🔥Introducing #AgentFlow, a new trainable agentic system where a team of agents learns to plan and use tools in the flow of a task. 🌐 https://x.com/lupantech/status/1976016000345919803
Is your LLM-based multi-agent system actually coordinating? That’s the question behind this paper. They use information theory to tell the difference between a pile of chatbots and a true collective intelligence. They introduce a clean measurement loop. First, test if the https://x.com/omarsar0/status/1977784668323008641
Readers responded with both surprise and agreement last week when I wrote that the single biggest predictor of how rapidly a team makes progress building an AI agent lay in their ability to drive a disciplined process for evals (measuring the system’s performance) and error”” / X https://x.com/AndrewYNg/status/1978867684537438628
Open Agent Builder coming soon 👀 https://x.com/Dev__Digest/status/1976308510347673652
Open-sourcing retrieve-dspy! 💻🚀 While developing Search Mode for Weaviate’s Query Agent, we dove into the literature. It was amazing, and overwhelming, to see how many different takes on Compound Retrieval Systems there are! 📚 From perspectives on Reranking, such as to https://x.com/CShorten30/status/1978567334424932523
The Claude Code SDK is now the Claude Agent SDK Why? Because we realized the Claude Code agent harness is useful for much more than coding. In fact, we’re moving to using it to power most of our own agent loops at Anthropic. https://x.com/trq212/status/1975243734083379662
New on the Anthropic Engineering Blog: Our tips for developers on using Agent Skills, a new way to extend Claude’s capabilities with instruction folders, scripts, and resources: https://x.com/AnthropicAI/status/1978896757489594404
Using this new agent in ClickUp you can totally leapfrog the IDE, Cursor, Claude Code… Codegen is now available in ClickUp. “”Hey @codegen, fix this bug”” – me “”PR submitted & ready for review!”” – Codegen “”Sweet, ship it”” – me We’ve always believed that the future of AI value https://x.com/DJ_CURFEW/status/1978568748794794476
Here’s your comprehensive guide to understanding and building Model Context Protocol (MCP) Servers for C# developers. It’s a great hands-on resource for exploring MCP and AI assistant integration in .NET environments. 📚 https://x.com/dotnet/status/1960029410649649626
GitHub Copilot + MCP = 🤯 Assign an issue → Copilot fetches data from a remote MCP server → Builds a green-themed HTML file 🎨💻 Autonomous dev workflows are here. This is next-level automation for your dev. 🔗 https://x.com/dotnet/status/1960096101333131310
We’re excited to announce that KumoRFM now supports Model Context Protocol (MCP), the open standard that connects AI agents with external tools and data. Explore more in our blog: https://x.com/Kumo_ai_team/status/1963320126843032026
ICYMI: Oracle Database now includes a Model Context Protocol (MCP) Server made available via Oracle SQLcl. Using any MCP-supported client, your favorite AI assistant and LLM, you can securely connect and interact with your database! https://x.com/OracleDevs/status/1961213462433980436
Announcing Open Agent Builder – A @firecrawl_dev powered n8n-style workflow builder example app Build AI agent workflows with a visual canvas by connecting Firecrawl, LLMs, logic nodes, and MCPs, then deploy as an API. Fork the repo and build your own workflow app today 👇 https://x.com/CalebPeffer/status/1978852506286571737
🚀 Product Highlight: @Ask_Rube (by @composiohq) — The universal MCP connector for AI agents. Access 500+ integrations like Gmail, GitHub, Slack, and Notion through one secure, managed endpoint. Highlights: • 500+ Production-Ready Integrations — Gmail, GitHub, Slack, Notion & https://x.com/MCP_Community/status/1959956426434265538
France’s @alpic_ai, a #cloud platform built specifically for Model Context Protocol (#MCP), announced today it has raised €5.1 million in #preSeed funding to build infrastructure for #AI agents to interact with the digital world 🇫🇷 🤖 https://x.com/EU_Startups/status/1963532351629136370
more side by side comparisons of @AnthropicAI Claude Haiku 4.5 vs Sonnet 4.5 in @windsurf here’s Haiku vs Sonnet on @flavioAd’s hexagon test the thing about fast models is that you can iterate with human guidance >2x in the time that it takes the slower model to do 1 round. https://x.com/swyx/status/1978578087752372452
Claude Skills | Hacker News https://news.ycombinator.com/item?id=45607117
📢 New Model Drop: Claude Haiku 4.5 is now on Yupp! @AnthropicAI’s latest small model, engineered for performance, speed and low cost. We investigated this new model with some prompts: https://x.com/yupp_ai/status/1978602716961415192
Today we’re introducing Claude Code Plugins in public beta. Plugins allow you to install and share curated collections of slash commands, agents, MCP servers, and hooks directly within Claude Code. https://x.com/claudeai/status/1976332881409737124
Claude Haiku 4.5 System Card https://assets.anthropic.com/m/99128ddd009bdcb/original/
Cerebras took Cognition’s latest code retrieval model, put it on the wafer, and it’s now running circles around Claude and Cursor. It’s now live in production for Windsurf users.”” / X https://x.com/draecomino/status/1978898418354561225
Claude’s unwillingness to continue conversations after the context window is full is very frustrating. I am okay losing early context if I am working interactively on a project & making progress, I am not okay with suddenly being cut off and being forced to start a new chat. https://x.com/emollick/status/1977604962772512800
Introducing Claude Haiku 4.5 \ Anthropic https://www.anthropic.com/news/claude-haiku-4-5
Anthropic launches their first Haiku model in 11 months – Claude 4.5 Haiku jumps 35 points in Artificial Analysis Intelligence Index to become relevant again Claude 4.5 Haiku is 3x cheaper per token than Claude 4.5 Sonnet. Running the Artificial Analysis Intelligence Index costs https://x.com/ArtificialAnlys/status/1978661658290790612
Learn how to build a Model Context Protocol (MCP) server that enables #LLM access to network devices—safely and securely. 🤖 See how it works (and get the code!) in the latest #AI Break series’ blog by @Kareem_isk. https://x.com/LearningatCisco/status/1960087189699657861
🚨 Text Leaderboard Update Community votes are in, and @anthropicAI’s Claude Haiku 4.5 ranks #22! It has quickly become one of the best value models on the most competitive leaderboard. It delivers a solid punch at a fraction of the cost of its bigger siblings. ⚡️ A few https://x.com/arena/status/1978966289248063885
Claude. make the most recursive, self-referential presentation you can imagine. seriously go big with this, don’t just run with your first idea, revise it multiple times (ironic instruction, yes)”” “”come on, i said recursive & self-referential. improve it. (see what i did?)”” https://x.com/emollick/status/1977927117212971165
🎉 𝗠𝗖𝗣 𝗶𝘀 𝗡𝗼𝘄 𝗚𝗔 𝗶𝗻 𝗩𝗶𝘀𝘂𝗮𝗹 𝗦𝘁𝘂𝗱𝗶𝗼! 🧠🛠️ Model-Context Protocol (MCP) is now generally available in @VisualStudio! Build smarter, more context-aware apps with full support and tooling baked right into your favorite IDE. It’s time to bring AI-native https://x.com/dotnetdevs_io/status/1963300959443923040
Claude now connects to Microsoft 365. Claude can search for information in SharePoint, OneDrive, Outlook and Teams, providing tailored responses seamlessly. https://x.com/AnthropicAI/status/1978864348236779675
MCP often gets criticized for its propensity to bloat context with “tool overload”. I think this criticism is unfounded. The reality is, you should never have more than a few MCP servers active in your context at a given time. And if a single MCP server has too many tools -“” / X https://x.com/tadasayy/status/1978170863192346660
Ran Haiku-4.5 against my NYT Connections Eval with DSPY project and the results are in! – Baseline score of 64% – Optimized score of 71% – Complete in only 25 minutes – Total cost $11 This means Haiku 4.5 is the fastest model I’ve tested so far (ignoring Haiku 3.5 which did https://x.com/pdrmnvd/status/1978570006863790299
Anthropic and Salesforce expand partnership to bring Claude to regulated industries \ Anthropic https://www.anthropic.com/news/salesforce-anthropic-expanded-partnership
We’re expanding our partnership with @Salesforce. Claude is now a preferred model in Agentforce for regulated industries. We’re deepening Claude’s integration with Slack, and Salesforce is rolling out Claude Code for its global engineering organization. https://x.com/AnthropicAI/status/1978125047270154567
Introducing the world’s first Portfolio Construction MCP💠 A groundbreaking Model Context Protocol server that brings Wall Street-grade portfolio optimization / re-balancing directly to AI agents, wallets, and other crypto projects. https://x.com/ChainAware/status/1958573602422358292
AI Agents can now see and test the code they generated. Chrome DevTools MCP gives AI agents live Chrome instances to analyze performance traces, inspect DOM, and debug issues in real-time. Works with Cursor, Gemini CLI, Claude Code and more. 100% Opensource. https://x.com/Saboo_Shubham_/status/1973937860626792868
Salesforce AI Research introduces MCP-Universe: the first benchmark to truly test LLM agents in real-world scenarios with live Model Context Protocol servers. https://x.com/HuggingPapers/status/1959347736429674567
Introducing NotebookLM for arXiv papers 🚀 Transform dense AI research into an engaging conversation With context across thousands of related papers, it captures motivations, draws connections to SOTA, and explains key insights like a professor who’s read the entire field https://x.com/askalphaxiv/status/1978466642146545838
We tested Search Mode from Weaviate’s Query Agent on five popular Information Retrieval benchmarks — BEIR, LoTTe, EnronQA, WixQA, and BRIGHT! 📊 Of these benchmarks, we found the largest relative improvement from Search Mode over Hybrid Search on BRIGHT! ⚖️🚀 BRIGHT from https://x.com/CShorten30/status/1978107101936230745
GPQA Diamond and 𝜏²-Bench Telecom (an agentic benchmark requiring models to act in a customer service role) both show outsized performance for GPT-5 and o3 compared to GPT-4.1, but while the reasoning models cost >10x to run GPQA, in 𝜏²’s customer service environment they cost https://x.com/ArtificialAnlys/status/1978561356401111051
📣New paper: Rigorous AI agent evaluation is much harder than it seems. For the last year, we have been working on infrastructure for fair agent evaluations on challenging benchmarks. Today, we release a paper that condenses our insights from 20,000+ agent rollouts on 9 https://x.com/sayashk/status/1978565190057869344
Reasoning models are expensive to run with traditional benchmarks, but often get cheaper in agentic workflows as they get to answers in fewer turns Through 2025 we’ve seen test-time compute drive up the cost of frontier intelligence, but with agentic workflows there’s a key https://x.com/ArtificialAnlys/status/1978561353792344302
Saw that DGX Spark vs Mac Mini M4 Pro benchmark plot making the rounds (looks like it came from @lmsysorg). Thought I’d share a few notes as someone who actually uses a Mac Mini M4 Pro and has been tempted by the DGX Spark. First of all, I really like the Mac Mini. It’s https://x.com/rasbt/status/1978608882156269755
7. Small models can punch above their weight. Their best 4B-parameter model, trained with this recipe (real data + diverse RL + GRPO-TCR), beats 14B–32B models on tough benchmarks like AIME25 and GPQA-Diamond. Smart data and tuning trump raw size. Really good paper for AI devs”” / X https://x.com/omarsar0/status/1978112412743258361
The return of the physicists: “”CMT-Benchmark: A benchmark for condensed matter theory built by expert researchers.”” https://x.com/SuryaGanguli/status/1977740051108036817
This work is incredibly close to my heart. So grateful to our team @GoogleDeepMind @GoogleResearch & @Yale collaborators. It’s been a hard road and a hard-won milestone, and this is just the beginning. The path from a finding like this to the clinic is very long, requiring”” / X https://x.com/AziziShekoofeh/status/1978601041814777914
“An exciting milestone for AI in science: Our C2S-Scale 27B foundation model, built with @Yale and based on Gemma, generated a novel hypothesis about cancer cellular behavior, which scientists experimentally validated in living cells. With more preclinical and clinical tests,” / X
https://x.com/sundarpichai/status/1978507110477332582
Introducing AgentKit—build, deploy, and optimize agentic workflows. 💬 ChatKit: Embeddable, customizable chat UI 👷 Agent Builder: WYSIWYG workflow creator 🛤️ Guardrails: Safety screening for inputs/outputs ⚖️ Evals: Datasets, trace grading, auto-prompt optimization https://x.com/OpenAIDevs/status/1975269388195631492
No more “memory full”!! Out for Plus and Pro users now. This has been one of our biggest pieces of feedback to date… what else do you want to see from memory in ChatGPT?”” / X https://x.com/ChristinaHartW/status/1978617445809352821
Clouded Judgement 10.10.25 – The ChatGPT App Store Moment https://cloudedjudgement.substack.com/p/clouded-judgement-101025-the-chatgpt
ChatGPT Apps are very powerful and can include full-fledged applications:”” / X https://x.com/gdb/status/1978508510359855119
What is ChatGPT Go? | OpenAI Help Center https://help.openai.com/en/articles/11989085-what-is-chatgpt-go
very nice vercel integration with chatgpt:”” / X https://x.com/gdb/status/1976560187445305815
Excited to release new repo: nanochat! (it’s among the most unhinged I’ve written). Unlike my earlier similar repo nanoGPT which only covered pretraining, nanochat is a minimal, from scratch, full-stack training/inference pipeline of a simple ChatGPT clone in a single, https://x.com/karpathy/status/1977755427569111362
OAI Workforce Blueprint [Oct 2025] https://cdn.openai.com/global-affairs/f319686f-cf21-4b8e-b8bc-84dd9bbfb999/oai-workforce-blueprint-oct-2025.pdf
The best ChatGPT that $100 can buy — now on your own k8s/cloud! Run @karpathy’s nanochat (pretrain → finetune → eval → serve) with SkyPilot 🚀 💰 Complete training for ~$100 ⚡ Deploy web UI for ~$2–3/hr 🌐 Run on any cloud or k8s with SkyPilot https://x.com/skypilot_org/status/1978273387903410412
You can now ship @nextjs apps to @chatgptapp https://x.com/rauchg/status/1976444029652160909
ChatGPT can now automatically manage your saved memories—no more “memory full.” You can also search and sort memories by recency, and choose which to re-prioritize in settings. Rolling out to Plus and Pro users on the web globally starting today. https://x.com/OpenAI/status/1978608684088643709
Claude Code subagents are all you need. Some will complain on # of tokens. However, the output this spits out will save you days. The code quality is mindblowing! Agentic search works exceptionally well. The subagents run in parallel. ChatGPT’s deep research is no match! https://x.com/omarsar0/status/1978235329237668214
OpenAI develops shared ChatGPT prompts feature for Teams https://www.testingcatalog.com/openai-develops-shared-chatgpt-prompts-feature-for-teams/
Why There Hasn’t Been a ChatGPT Moment Yet in Manufacturing https://theshearforce.substack.com/p/why-there-hasnt-been-a-chatgpt-moment
chatgpt for slack:”” / X https://x.com/gdb/status/1977835678202535984
Stop scrolling! @karpathy just turned 100 dollars of GPU time into your own ChatGPT: Train a small chat model end to end in about 4 hours on a single 8×H100 box, then chat with it in a web UI. Clean code. One script. Full pipeline. ✅ One command speedrun from empty box to https://x.com/IlirAliu_/status/1977995281212780598
BOOM: We’ve just re-launched HuggingChat v2 💬 – 115 open source models in a single interface is stronger than ChatGPT 🔥 Introducing: HuggingChat Omni 💫 > Select the best model for every prompt automatically 🚀 > Automatic model selection for your queries > 115 models https://x.com/reach_vb/status/1978854312647307426
gpt-5 pro for searching scientific literature:”” / X https://x.com/gdb/status/1977190029727518973
Introducing the most advanced open-source chat template built on top of AI SDK. → Agents (Orchestration, handoffs & context guardrails) → Artifacts (Canvas, Charts) → Rate limits (Messages, tool permissions) Link ⬇️🧵 https://x.com/pontusab/status/1976304838200983572
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution “”we introduce BigCodeArena, an open human evaluation platform for code generation backed by a comprehensive and on-the-fly execution environment. Built on top of Chatbot Arena, BigCodeArena https://x.com/iScienceLuvr/status/1977694597603291492
GLM-4.6 is now live on BigCodeArena. Shout-out to @qinkai1028 and the whole @Zai_org team for this great model!”” / X https://x.com/terryyuezhuo/status/1978554496058851650




