Image created with Flux Pro v1.1 Ultra. Image prompt: Ethics, scales of justice holding matched clusters of small bananas on each plate, perfect balance, photorealistic, editorial, minimal, high detail, 3:2 landscape
China’s DeepSeek Preps AI Agent for End-2025 to Rival OpenAI https://finance.yahoo.com/news/china-deepseek-preps-ai-agent-152907224.html?guccounter=1&guce_referrer=aHR0cHM6Ly93d3cuZ29vZ2xlLmNvbS8&guce_referrer_sig=AQAAADiJs67uOGL7PzqX3MGgvD-A6UJVzmztcJfvPzJTz9iF2iWfg-h2zg2pcwJoIuJ-4IUs3BMrEvPbbpbf4j7qXCmM4BqK78UMZVzrZl3fSuokrWneWMYpy8S7L3-xciC9d74km3boS_g57OxikNZN7Owozd204A5KlQA0MSzkqp42
Alibaba reportedly developing new AI chip as China’s Xi rejects AI’s ‘Cold War mentality’ | Euronews https://www.euronews.com/next/2025/09/01/alibaba-reportedly-developing-new-ai-chip-as-chinas-xi-rejects-ais-cold-war-mentality
New Scale research: Can smaller models reliably oversee stronger LLM agents? We red team monitoring systems to detect covert sabotage, like agents secretly downloading sensitive information. https://x.com/scale_AI/status/1961233659228557530
AI personality isn’t the problem. The illusion of AI personhood is. https://x.com/mustafasuleyman/status/1963281258844438733
Building more helpful ChatGPT experiences for everyone | OpenAI https://openai.com/index/building-more-helpful-chatgpt-experiences-for-everyone/
Mayor Adams, DOT Announce Approval of First Application to Test Autonomous Vehicles in New York City With Trained Safety Specialist Behind Steering Wheel – NYC Mayor’s Office https://www.nyc.gov/mayors-office/news/2025/08/mayor-adams–dot-announce-approval-of-first-application-to-test-
We can now say pretty definitively that AI progress is well ahead of expectations from a few years ago. In 2022, the Forecasting Research Institute had super forecasters & experts to predict AI progress. They gave a 2.3% & 8.6% probability of an AI Math Olympiad gold by 2025… https://x.com/emollick/status/1962859757674344823
Google’s on a roll. That’s a lot of performance for that tiny size! I just embedded 1.4 million documents in ~80 mins on my M2 Max for free. Would’ve been ~$200 with the text-embedding-3-large, with worse quality.”” / X https://x.com/rishdotblog/status/1963805087014502497
Judge rules in Google’s illegal search monopoly case: it can keep Chrome | The Verge https://www.theverge.com/policy/717087/google-search-remedies-ruling-chrome
OpenAI Plans to Build Data Center in India in Major Stargate Expansion in Asia – Bloomberg https://www.bloomberg.com/news/articles/2025-09-01/openai-plans-india-data-center-in-major-stargate-expansion?srnd=phx-technology&embedded-checkout=true
Chinese news outlet CCTV Finance: “”According to market data, China’s humanoid robot sales in 2025 will exceed 10,000 units, a year-over-year increase of 125%.”” https://x.com/TheHumanoidHub/status/1961110406858199528
Accelerating AI adoption for the US government – The Official Microsoft Blog https://blogs.microsoft.com/blog/2025/09/02/accelerating-ai-adoption-for-the-us-government/
How do we generate videos on the scale of minutes, without drifting or forgetting about the historical context? We introduce Mixture of Contexts. Every minute-long video below is the direct output of our model in a single pass, with no post-processing, stitching, or editing. 1/4 https://x.com/GordonWetzstein/status/1963583050744250879
New Anthropic Research: Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks.
https://x.com/JackYoustra/status/1963280250923868239
Exclusive: Meta created flirty chatbots of Taylor Swift, other celebrities without permission | Reuters https://www.reuters.com/business/meta-created-flirty-chatbots-taylor-swift-other-celebrities-without-permission-2025-08-29/
Cool research from Microsoft! They release rStar2-Agent, a 14B math reasoning models trained with agentic RL. It reaches frontier-level math reasoning in just 510 RL training steps. Here are my notes: https://x.com/omarsar0/status/1964045125115662847
rStar2-Agent: Agentic Reasoning Technical Report “”We introduce rStar2-Agent, a 14B math reasoning model trained with agentic reinforcement learning to achieve frontier-level performance.”” “”three key innovations that makes agentic RL effective at scale: (i) an efficient RL https://x.com/iScienceLuvr/status/1962798181059817480
New White House commitments empower teachers, students, and job seekers through AI skilling and learning – Microsoft On the Issues https://blogs.microsoft.com/on-the-issues/2025/09/04/new-white-house-commitments/
Marc Benioff interacting with Optimus. The robot pauses for a long time after a command, Elon mentions, because it needed a little bit more room around it. The hands appear to be placeholder dummies. https://x.com/TheHumanoidHub/status/1963269758423580717
We need to talk about two kinds of “”normal technology”” when asking “”is AI a normal technology?”” There is “”normal”” tech diffusion & there is treating AI as “”normal”” tech that is just another IT product. I think there is a case for the former. The latter belief is likely blinding.”” / X https://x.com/emollick/status/1961487454394789914
Building AI agents for production comes with unique challenges. In our latest blog, we share how we designed LangGraph to tackle them: 🔹 Why heavy abstractions fail and what really matters for control & durability 🔹 The 6 features every production agent needs in practice 🔹”” / X https://x.com/LangChainAI/status/1963646974315606428
https://x.com/omarsar0/status/1962875111037358540
Today we’re launching Atla — the improvement engine for AI agents. Atla helps agent builders find and fix recurring failures. Instead of just surfacing traces, Atla automatically identifies your agent’s most critical failure patterns and suggests targeted fixes. https://x.com/Atla_AI/status/1963586200305836264
Issue Triager Agent A GitHub issue management solution that uses LangGraph to automatically handle stale issues with human oversight through Agent Inbox. Built with LangSmith integration for comprehensive monitoring and control. https://x.com/LangChainAI/status/1962198699653861755
This is a pretty important point, we have relied on all LLMs being broadly similar to each other (even to the extent that prompting is compatible across models). That may start to change with reinforcement learning.”” / X https://x.com/emollick/status/1961105788724027770
We raised a $150M Series D! Thank you to all of our customers who trust us to power their inference. We’re grateful to work with incredible companies like @Get_Writer, @zeddotdev, @clay_gtm, @trymirage, @AbridgeHQ, @EvidenceOpen, @MeetGamma, @Sourcegraph, and @usebland. This https://x.com/basetenco/status/1963981711647379653
I’m learning the true Hanlon’s razor is: never attribute to malice or incompetence that which is best explained by someone being a bit overstretched but intending to get around to it as soon as they possibly can.”” / X https://x.com/AmandaAskell/status/1961577559344455769
90% success rate in unseen environments. No new data, no fine-tuning. Autonomously. Most robots need retraining to work in new places. What if they didn’t? Robot Utility Models (RUMs) learn once and work anywhere… zero-shot. A team from NYU and Hello Robot built a set of https://x.com/IlirAliu_/status/1961692920836215229
A robot that sees the terrain and predicts its own future… up to 5 seconds ahead? This is real. ❗️Best Systems Paper finalist at #RSS2025 The team introduces a perceptive Forward Dynamics Model that helps legged robots safely navigate rough, complex environments: no manual https://x.com/IlirAliu_/status/1962569938805141861
A really useful prompt for writing: “”review this for accuracy, look up any facts you may want to challenge or explore.”” Even if not perfect, it is a good sanity check. Works well with Claude 4.1, GPT-5 Thinking, and Grok 4. Weirdly, Gemini 2.5 Pro often won’t do web searches. https://x.com/emollick/status/1961257429846691881
We really have not made a lot of progress on explaining the deep mystery of LLMs: How does a model using matrix multiplication to predict the next word manage to simulate human thought well enough to do all the very human-like things it does? And what does that mean about us?”” / X https://x.com/emollick/status/1960919256452796440
A book chapter I co-wrote many moons ago on longtermism just came out. I forgot we gave such dire warnings about going into academia 😆 https://x.com/AmandaAskell/status/1963266155944218806
Anyone know how to opt out of Anthropic’s new 5-year (!) data retention policy? https://x.com/michael_nielsen/status/1961439837791367501
looks like you can opt out of them training on your data, but they’ll still store your data for 5 years? https://x.com/vikhyatk/status/1961511207577534731
🐺 Introducing the Werewolf Benchmark, an AI test for social reasoning under pressure. Can models lead, bluff, and resist manipulation in live, adversarial play? 👉 We made 7 of the strongest LLMs, both open-source and closed-source, play 210 full games of Werewolf. Below is https://x.com/RaphaelDabadie/status/1961836323376935029
Scale AI is suing a former employee and rival Mercor, alleging they tried to steal its biggest customers | TechCrunch https://techcrunch.com/2025/09/03/scale-ai-is-suing-a-former-employee-and-rival-mercor-alleging-they-tried-to-steal-its-biggest-customers/
Regardless of whether current AI Labs fail (and there is no indication they are at risk) & even if AI development stops (more unlikely), things will keep getting weirder: today’s models are good enough for long-term disruption, and the weights & infrastructure aren’t going away.”” / X https://x.com/emollick/status/1962210917145264269
World’s first method turns plastic into fuel with 95% efficiency https://interestingengineering.com/science/us-china-turn-plastic-to-petrol
@vikhyatk If you opt out, the retention period is 30 days (no change to the existing period). https://x.com/sammcallister/status/1961520548510400753
Updates to Consumer Terms and Privacy Policy \ Anthropic https://www.anthropic.com/news/updates-to-our-consumer-terms
Newsom, California lawmakers strike deal that would allow Uber, Lyft drivers to unionize – Los Angeles Times https://www.latimes.com/california/story/2025-08-29/california-lawmakers-strike-deal-to-allow-uber-lyft-drivers-to-unionize
I know that energy usage from AI prompts comes up in classroom discussion all the time. I hope this section in my latest post helps people provide a more grounded answer to the question, rather than just speculating or citing out-of-date information. https://x.com/emollick/status/1962945874956304494
Warner Bros. Sues Midjourney, Joins Studios’ AI Copyright Battle https://variety.com/2025/film/news/warner-bros-midjourney-lawsuit-ai-copyright-1236508618/
We need to protect children from having their government ID linked to their adult online activity for the rest of ther lives. I’d like to propose some kind of online child safety act to this effect.”” / X https://x.com/AmandaAskell/status/1961164652479721724
Microsoft fires two more employees for participating in Palestine protests on campus | The Verge https://www.theverge.com/microsoft/767841/microsoft-fires-two-more-protesters-no-azure-for-apartheid
<cot>I wonder if the timeline over at Substack is better, maybe there is less slop and more interesting longform or so on. Opens Substack. https://x.com/karpathy/status/1961146044550373712
This is disappointing. Purposefully underselling what models can do is a really bad idea. It is possible to point out that AI is flawed without saying it can’t do math or count – it just isn’t true. People need to be realistic about capabilities of models to make good decisions.”” / X https://x.com/emollick/status/1963287621377167732
Mass Intelligence means we are going to inundated with stories about people using AI to do amazing things and horrifying things as over a billion people increasingly get access to advanced (and easy-to-use) AI models. Things are going to get very weird. https://x.com/emollick/status/1961469787415949417
This chart is being horribly misinterpreted. This is not where the training data of AI comes from, it is a study done by a SEO firm that claims to show how often sites come up at least once in THE WEB SEARCH FUNCTION of certain AI agents when they do a web search for more info. https://x.com/emollick/status/1962678752887914918
I wrote about the era of Mass Intelligence. GPT-5 and Google’s Nano Banana are examples of how advanced AI is now making their way to far more users, at scale, as both performance and efficiency keep improving. We are going to see a lot of weird things happening, all at once. https://x.com/emollick/status/1961169796491329653
Wonder how the White House feels watching this? > US doubles tariffs on Indian imports (punishing purchases of Russian oil). > Modi responds by shaking hands with Xi, despite a history of border disputes (and all out war) with China. The great game continues…”” / X https://x.com/bilawalsidhu/status/1962242039099388375
Meta introduces Set Block Decoding (SBD), a new inference accelerator for LLMs SBD samples multiple future tokens in parallel, cuts forward passes by 3–5x, needs no arch changes, stays KV-cache compatible, and matches NTP training performance. https://x.com/arankomatsuzaki/status/1963817987506643350
@jeremyphoward Fixing this is very high on the priority list for the next version! The reason it says it in the system prompt is because the model was asking too many clarification questions (and thinking long for each) which IMO was even worse UX.”” / X https://x.com/yanndubs/status/1961716590568706226
AIcos: At long last, we have built almost literally exactly the AI That Tells Humans What They Want To Hear, from Isaac Asimov’s classic 1941 short story, “”Don’t Build AI That Tells Humans What They Want To Hear”” https://x.com/ESYudkowsky/status/1962574434231062861
Trusted news sites may benefit in an internet full of AI-generated fakes, a new study finds | Nieman Journalism Lab https://www.niemanlab.org/2025/08/trusted-news-sites-may-benefit-in-an-internet-full-of-ai-generated-fakes-a-new-study-finds/
Goated FAIR team just found how coding agents sometimes “”cheat”” on SWE-Bench Verified. It’s really simple. For example, Qwen3 literally greps all commit logs for the issue number of the issue it needs to fix. lol, clever model. “”cheat”” cuz it’s more like env hacking. https://x.com/giffmana/status/1963327672827687316
You don’t need more robot data. You need to look inside the data you already have. [📍 bookmark for later] Instead of just copying demos, STRAP pulls semantically meaningful pieces from large offline datasets to improve robustness and performance… no fine-tuning needed. Why https://x.com/IlirAliu_/status/1961850172058525813
Would you expect to find an industrial robot working 200 meters underground? An ABB robot is palletizing bags of salt deep below the surface: in a place where space, conditions, and safety make automation indispensable. ✅ Boosts productivity where humans can’t work https://x.com/IlirAliu_/status/1962781359023243751
Tesla asks court to throw out fatal Autopilot crash verdict https://www.bbc.com/news/articles/ckgdjx0vgn3o
AI for Security has never been more exciting. Let me present MAPTA, our multi-agent framework that found multiple (now confirmed!) Remote Code Executions (RCE’s) in flagship web products of Tier-1 companies. Why the secrecy? We’re good boys, letting them cook patched through https://x.com/HatforceSec/status/1962550734563532874
There is a small and determined team in California trying to build gun detection systems for K12 schools We are looking for an AI Engineer to deploy deep learning models to detect weapons via 3D point cloud Please consider applying to help save lives – job link in comment”” / X https://x.com/adcock_brett/status/1960882149105787282
At Transluce, we train investigator agents to surface specific behaviors in other models. Can this approach scale to frontier LMs? We find it can, even with a much smaller investigator! We use an 8B model to automatically jailbreak GPT-5, Claude Opus 4.1 & Gemini 2.5 Pro https://x.com/TransluceAI/status/1963286326062846094
xAI sues ex-engineer for trade secret theft https://www.therundown.ai/p/xai-sues-ex-engineer-for-trade-secret-theft
receipts for no-BS evals talks: https://x.com/swyx/status/1963727193974153602
XAI OPENAI TRADE SECRETS LAWSUIT complaint.pdf https://fingfx.thomsonreuters.com/gfx/legaldocs/gdvzbjjjzvw/XAI%20OPENAI%20TRADE%20SECRETS%20LAWSUIT%20complaint.pdf
xAI sues former engineer, alleging he stole trade secrets after being paid $7M https://sfstandard.com/2025/08/29/xai-elon-musk-openai-stanford-sam-altman-ai-talent-wars/
i never took the dead internet theory that seriously but it seems like there are really a lot of LLM-run twitter accounts now”” / X https://x.com/sama/status/1963366714684707120




