Image created with Flux Pro v1.1 Ultra. Image prompt: Giant “100” as pure white negative‑space cutout dominating the frame; minimalist poster style; balanced scales made of circuit traces centered in the zeros; emerald backdrop; high contrast, crisp edges, soft studio light, no other text, no logos
Transforming human knowledge, sensors and actuators from human-first and human-legible to LLM-first and LLM-legible is a beautiful space with so much potential and so much can be done…
One example I’m obsessed with recently – for every textbook pdf/epub, there is a perfect “LLMification” of it intended not for human but for an LLM (though it is a non-trivial transformation that would need human in the loop involvement).
– All of the exposition is extracted into a markdown document, including all latex, styling (bold/italic), tables, lists, etc. All of the figures are extracted as images.
– All worked problems get extracted into SFT examples. Any referenced made to previous figures/tables/etc. are parsed and included.
– All practice problems are extracted into environment examples for RL. The correct answers are located in the answer key and attached. Any additional information is added as “answer key” for a potential LLM judge.
– Synthetic data expansion. For every specific problem, you can create an infinite problem generator, which emits problems of that type. For example, if a problem is “What is the angle between the hour and minute hands at 9am?” , you can imagine generalizing that to any arbitrary time and calculating answers using Python code, and possibly generating synthetic variations of the prompt text.
– All of the data above could be nicely indexed and embedded into a RAG database for later reference, or maybe MCP servers that make it available.
Then just as a (human) student could take a high school physics course, an LLM could take it in the exact same way. This would be a significantly richer source of legible, workable information for an LLM than just something like pdf-to-text (current prevailing practice), which simply asks the LLM to predict the textbook content top to bottom token by token (umm – lame). https://x.com/karpathy/status/1961128638725923119
Malicious actors are adapting to exploit AI’s most advanced capabilities. We’re sharing these findings to strengthen collective defenses across the industry. Read more: https://x.com/AnthropicAI/status/1960660072322764906
Watch Jacob Klein and Alex Moix from Anthropic’s Threat Intelligence team discuss what Anthropic is doing to disrupt AI cybercrime: https://x.com/AnthropicAI/status/1960660074948354501
It seems like there is not enough of a policy response to the fact that, with 57M miles of data, Waymo’s autonomous vehicles experience 85% less serious injuries & 79% less injuries overall than cars with human drivers. 2.4 million are injured & 40k killed in US accidents a year”” / X https://x.com/emollick/status/1959249518194528292
AI Now Matches Prediction Markets in Forecasting Real Events, Study Finds https://finance.yahoo.com/news/ai-now-matches-prediction-markets-230346352.html
Google scores six-year Meta cloud deal worth over $10 billion https://www.cnbc.com/2025/08/21/google-scores-six-year-meta-cloud-deal-worth-over-10-billion.html
Gemini for Government – brings together the best of Google’s AI-optimized & accredited commercial cloud, SOTA Gemini models, and agentic solutions to support the missions of government agencies! All for less than $0.50 per agency : )”” / X https://x.com/OfficialLoganK/status/1958549753148408045
Congratulations to the @scale_AI team on a major $99M contract with the @USArmy!”” / X https://x.com/alexandr_wang/status/1960195704275743035
Proud to share we’ve been awarded a $99M contract to support the acceleration of the @USArmy’s adoption of AI. This contract reflects our long-term commitment to ensuring America’s military remains prepared, resilient, and at the forefront of AI innovation. https://x.com/scale_AI/status/1960126564236157391
Scale AI and DoD Expand Partnership to Advance Army R&D | Scale https://scale.com/blog/scale-ai-dod-expand-army-rd-partnership
My wife Anna and I are supporting @LeadingFutureAI because we believe that AI can massively improve quality of life for every person (and every animal!). We believe the goal of AI policy should be to unlock this outcome. That means taking a balanced view, which we think of as”” / X https://x.com/gdb/status/1960022650228793440
Perplexity Is Launching a New Revenue-Share Model for Publishers – WSJ https://www.wsj.com/business/media/perplexity-ai-search-publisher-revenue-507987e5?gaa_at=eafs&gaa_n=ASWzDAh7IXUez7FN0pV8tS6RJPJiyucYbQ7TqVVq5v5uIQCmptaIEy7EmDeq7YP8qA%3D%3D&gaa_ts=68acf00e&gaa_sig=9LlQZ1VfgdWZeHollVleg5iI8GhXdOYPZrqyfcFsJu2MvpLol1OCDH19znUZBcg-D6zi8CV0qZGZxh6WAHnF-g%3D%3D
Perplexity’s $42.5M publisher peace offering https://www.therundown.ai/p/perplexitys-42-5m-publisher-peace-offering
Swipe from right to see the smoothest and most informative daily timeline of information on Perplexity. Personalizes as you use more and will get integrated with SuperMemory for even more personalized ranking once we ship widely https://x.com/AravSrinivas/status/1959689988989464889
China’s Kaiwa plans world’s first pregnancy humanoid robot https://interestingengineering.com/innovation/china-worlds-first-pregnancy-humanoid-robot
Detecting and countering misuse of AI: August 2025 \ Anthropic https://www.anthropic.com/news/detecting-countering-misuse-aug-2025
Introducing the Anthropic National Security and Public Sector Advisory Council \ Anthropic https://www.anthropic.com/news/introducing-the-anthropic-national-security-and-public-sector-advisory-council
Our new Threat Intelligence report details how we’ve identified and disrupted sophisticated attempts to use Claude for cybercrime. We describe a fraudulent employment scheme from North Korea, the sale of AI-created ransomware by someone with only basic coding skills, and more. https://x.com/AnthropicAI/status/1960660063934194134
We’re announcing the Anthropic National Security and Public Sector Advisory Council, a bipartisan group of defense, intelligence, and policy experts who will help us support the U.S. government and closely allied democracies in maintaining our AI leadership. https://x.com/AnthropicAI/status/1960696531863879712
Early this summer, OpenAI and Anthropic agreed to try some of our best existing tests for misalignment on each others’ models. After discussing our results privately, we’re now sharing them with the world. 🧵 https://x.com/sleepinyourhat/status/1960749648110395467
Findings from a Pilot Anthropic – OpenAI Alignment Evaluation Exercise https://alignment.anthropic.com/2025/openai-findings/
We recently ran to have OpenAI and Anthropic each evaluate each others’ models for safety issues. Excited for us to find more ways to help support safety practices across the whole field!”” / X https://x.com/EthanJPerez/status/1960808655642882228
China seeks to triple output of AI chips in race with the US – Google Search https://www.google.com/search?q=China+seeks+to+triple+output+of+AI+chips+in+race+with+the+US&sourceid=chrome&ie=UTF-8
Exclusive | Silicon Valley Launches Pro-AI PACs to Defend Industry in Midterm Elections – WSJ https://www.wsj.com/politics/silicon-valley-launches-pro-ai-pacs-to-defend-industry-in-midterm-elections-287905b3
Grok 2 from @xai has just been released on @huggingface: https://x.com/ClementDelangue/status/1959356467959439464
Grok-2 has been “”open sourced”” but has one of the worst licenses of any recent major open weights release. Given that it’s already quite outdated by the time they’ve got around to releasing it, combined with the license, this will see little use. It’s dead on arrival. https://x.com/xlr8harder/status/1959490601264533539
Pretty cool that they open sourced the actual full-sized production model. Here’s the Grok 2.5 architecture overview next to a roughly similarly sized Qwen3 model. The MoE residual is quite interesting. Kind of like a shared expert. I don’t think I’ve seen this setup before. https://x.com/rasbt/status/1959643038268920231
The @xAI Grok 2.5 model, which was our best model last year, is now open source. Grok 3 will be made open source in about 6 months. https://x.com/elonmusk/status/1959379349322313920
xAI just released Grok 2 on Hugging Face. This massive 500GB model, a core part of xAI’s 2024 work, is now openly available to push the boundaries of AI research. https://x.com/HuggingPapers/status/1959345658361475564
xai-org/grok-2 · Hugging Face https://huggingface.co/xai-org/grok-2
Grok now has a model card – which is a big step forward! But it is light on details, with unexplained results. Some examples: if the MASK measurement is the same as in the source paper, .43 would be a fairly high level of deception, also the sycophancy score is hard to interpret https://x.com/emollick/status/1959116132096336066
MIT report misunderstood: Shadow AI economy booms while headlines cry failure https://venturebeat.com/ai/mit-report-misunderstood-shadow-ai-economy-booms-while-headlines-cry-failure
Agentic Browser Security: Indirect Prompt Injection in Perplexity Comet | Brave https://brave.com/blog/comet-prompt-injection/
A Teen Was Suicidal. ChatGPT Was the Friend He Confided In. – The New York Times https://www.nytimes.com/2025/08/26/technology/chatgpt-openai-suicide.html
Some of the quotes from ChatGPT in this piece are quite eye-opening. OpenAI’s safeguards clearly failed here. https://x.com/lefthanddraft/status/1960340188145787005
Why is no one talking about this? This is why I don’t use an AI browser You can literally get prompt injected and your bank account drained by doomscrolling on reddit: https://x.com/zack_overflow/status/1959308058200551721
YouTube secretly used AI to edit people’s videos. The results could bend reality https://www.bbc.com/future/article/20250822-youtube-is-using-ai-to-edit-videos-without-permission
Building AI that answers with confidence requires two things: fresh context and transparent execution. Static agents miss both. We paired @tavilyai (real-time web) with W&B Weave (tracing, evals, ops) to ship research agents that are accurate, current and trustworthy. https://x.com/weave_wb/status/1960428416236445931
📈 Process reward strikes back 🚨 I think it is obvious that eventually we need to rely on stepwise judges instead of final outcome rewards. As tasks get longer (or even endless), it is unreasonable to push up/down all steps involved. Here we show you can obtain stepwise labels”” / X https://x.com/tesatory/status/1960533462672400724
🪜Introducing: StepWiser🦉 📝: https://x.com/jaseweston/status/1960529697055355037
In era of pretraining, what mattered was internet text. You’d primarily want a large, diverse, high quality collection of internet documents to learn from. In era of supervised finetuning, it was conversations. Contract workers are hired to create answers for questions, a bit”” / X https://x.com/karpathy/status/1960803117689397543
Harvard dropouts to launch ‘always on’ AI smart glasses that listen and record every conversation | TechCrunch https://techcrunch.com/2025/08/20/harvard-dropouts-to-launch-always-on-ai-smart-glasses-that-listen-and-record-every-conversation/
Has anyone seen a copy of the Project NANDA report on 95% of AI pilots failing? It has been reported about widely, but I haven’t been able to actually read it, though I filled out a form to get access. I want to understand what they found and how they measured it. Any links?”” / X https://x.com/emollick/status/1958598242884854077
Okay, got the report. I would read it yourself. I am not sure how generalizable the findings are based on the methodology (52 interviews, convenience sampled, failed apparently means no sustained P&L impact within six months but no coding explanation). https://x.com/emollick/status/1958602041367961972
AI is already transforming the global economy, and America should lead in advancing a policy governing AI expansion that drives innovation and broad economic growth for working people across the country. LTF and its affiliated organizations will promote policies that unlock the”” / X https://x.com/LeadingFutureAI/status/1959956536370790436
Melania Trump announces nationwide AI contest for students https://thehill.com/homenews/5470552-melania-trump-ai-school-challenge/
We are underinvesting in areas where research suggests AI could be good enough, right now, to make an impact Controlled studies find that current LLMs can be very useful in medicine & education, especially in disadvantaged areas, but more work is needed to deliver on the promise”” / X https://x.com/emollick/status/1960809095281434881
An interesting discussion from a scholar of literature on some of the odd weak points of GPT-5’s figurative writing ability, and what this might tell us about the problems of AI-driven evaluation of AI writing (and other uses of LLMs as a judge) when training models.”” / X https://x.com/emollick/status/1960445234875392090
Great Western Railway Battery Train Sets New World Record with 200-Mile Journey | Rail Industry Connect | Connecting you with rail industry insight and best practice https://www.railindustryconnect.co.uk/great-western-railway-battery-train-sets-new-world-record-with-200-mile-journey/
AI persuades best by overwhelming people with information instead of using psychological tricks https://the-decoder.com/ai-persuades-best-by-overwhelming-people-with-information-instead-of-using-psychological-tricks/
AI’s value is precisely because it’s something so different from humans. Never tired, infinitely patient, able to process more data than a human mind ever could. This is what benefits humanity. Not an AI that claims to feel shame, jealousy, fear + so on. 📝 https://x.com/mustafasuleyman/status/1958582940524597396
Richard Sutton says the AI industry has “”lost its way”” by ignoring core principles of intelligence https://the-decoder.com/richard-sutton-says-the-ai-industry-has-lost-its-way-by-ignoring-core-principles-of-intelligence/
Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence https://digitaleconomy.stanford.edu/wp-content/uploads/2025/08/Canaries_BrynjolfssonChandarChen.pdf
There has not been a lot of evidence so far that AI is actually impacting jobs. This paper suggests this may be starting, showing a decrease in entry-level job openings relative to experienced ones & in the jobs where AI is mostly automates, not augments, work (like coding).”” / X https://x.com/emollick/status/1960340823972647393
10,000 prompts is the new 10,000 hours”” / X https://x.com/reidhoffman/status/1960392913130541551
We have data on the environmental impact per AI prompt: Gemini: 0.00024 kWh & 0.26 mL water ChatGPT: 0.0003 kWh & 0.38 mL …the same energy as one Google search in 2008 & 6 drops of water. Seems to be improving, too: Google reports a 33x drop in energy use per prompt in a year. https://x.com/emollick/status/1958742855876227348
Gen Z wants to have their AI cake and eat it, too: KPMG intern survey reveals a generation that wants to have things both ways | Fortune https://fortune.com/2025/08/20/gen-z-careers-new-hires-ai-mentorship-work-ethic/
Japanese city proposes two-hour daily limit for smartphone use – The Japan Times https://www.japantimes.co.jp/news/2025/08/22/japan/society/japan-city-proposes-two-hour-daily-smartphone-limit/
One interesting way to understand or fact-check claims is to ask an AI to make a simulation of a process and see if it makes sense. “”Gemini 2.5 (with Canvas on), create a simulation showing this feedback loop”’ https://x.com/emollick/status/1959363417858351125
Google pays $30M to settle lawsuit over children’s YouTube data | TechCrunch https://techcrunch.com/2025/08/19/google-pays-30m-to-settle-lawsuit-over-childrens-youtube-data/
at least we can all agree, this is the worst style of art ever made https://x.com/onionweigher/status/1960210909629944048
A big milestone for Hermes. We did a lot of work to make a frontier level openmodel that does not dictate what expression you can elicit from the model. Super strong at math, coding, STEM, and creativity. Model Weights: https://x.com/Teknium1/status/1960420619620901135
Hermes 4 – Nous Research https://hermes4.nousresearch.com/
Hermes 4 technical breakdown: ▫️ Open Source LLM ▫️ Fine-tune of Llama 3.1 ▫️ 405B & 70B params ▫️ Hybrid reasoning ▫️ Trained on 3.5 million reasoning samples ▫️ Trained using 192 NVIDIA B200 GPUs ▫️ Uncensored ▫️ Steerable, aligned to the user ▫️ Creativity enhanced (like”” / X https://x.com/vectro/status/1960734604601569560
Nous Research presents Hermes 4, our latest line of hybrid reasoning models. https://x.com/NousResearch/status/1960416954457710982
Fourth model launch of the day 🔥 – introducing Hermes 4, from @NousResearch Hermes 4 is trained for steerability and lower refusal rates, topping RefusalBench and beating Grok 4 https://x.com/OpenRouterAI/status/1960436262923592065
1/8 🧵 GPT-5’s storytelling problems reveal a deeper AI safety issue. I’ve been testing its creative writing capabilities, and the results are concerning – not just for literature, but for AI development more broadly. 🚨”” / X https://x.com/ChristophHeilig/status/1960358655745724438
GPT-5 says ‘I don’t know’. Love this, thank you. https://x.com/koltregaskes/status/1957474061153436094
It’s rare for competitors to collaborate. Yet that’s exactly what OpenAI and @AnthropicAI just did—by testing each other’s models with our respective internal safety and alignment evaluations. Today, we’re publishing the results. Frontier AI companies will inevitably compete on”” / X https://x.com/woj_zaremba/status/1960757419245818343
OpenAI plans a new build with Oracle that would add 4.5 gigawatts of data-center capacity, an outgrowth of their “Stargate” program. The Wall Street Journal reported OpenAI will pay Oracle $30 billion annually. The plan follows a 1.2-gigawatt site in Abilene, Texas. Selection https://x.com/DeepLearningAI/status/1960900145421177053
Musk v. OpenAI just got messier https://tech.therundown.ai/p/musk-brings-zuck-into-openai-drama
OpenAI lawyers question Meta’s role in Elon Musk’s $97B takeover bid | TechCrunch https://techcrunch.com/2025/08/21/openai-lawyers-question-metas-role-in-elon-musks-97b-takeover-bid/
🎙️ In this episode, I talk with @BrendahNjiru, founder of @HomyRobotics, where she’s building emotionally intelligent humanoid robots for senior living: With a background in neuroscience and Alzheimer’s research at Cornell, Brendah brings a rare scientific depth to robotics. We https://x.com/IlirAliu_/status/1958515933259178236
Tracking drones in the wild is one of the hardest problems in computer vision. 🚁🔥 Small, fast-moving objects. Changing backgrounds. Real-time constraints… Most models break under these conditions. That’s why this open-source project from @chesterzelaya, extended by https://x.com/IlirAliu_/status/1959310919147622741
Classic deep state Washington thinking around tech is focused purely on *control* and *risk* and has a lack of understanding of technology/developer ecosystems work. As @DavidSacks says: for the American AI stack to win, we need to maximize marketshare. This means maximizing”” / X https://x.com/sriramk/status/1961072926561550366
i was literally shocked by the huge llm infra diff in US vs China and GPU vs TPU… I was chatting with senior folks about how global batch aux loss is hard under the ppvp constrain, as basically you have to do all f all b to get the fi correct and that’s challenging for peak”” / X https://x.com/JingyuanLiu123/status/1959093411283443726
I saw a photo of a new U.S. reshoring factory. The #1 robot brand powering Trump’s “Make America Great Again” revival isn’t… American. It’s Japanese. That sent me down a rabbit hole…. What I found is a 40-year silent invasion that reshaped American factories while almost https://x.com/IlirAliu_/status/1959610913134133549
🧵 TIL Distributed Infra The main reason to add PP is when DP comms (optimizer/state sync) are fine but TP can’t scale further due to bandwidth/latency or memory/geometry limits. Prefer TP+DP on TPUs / NVLink islands.”” / X https://x.com/mr_besher/status/1959215227972505960
DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization DuPO generates annotation-free feedback via a generalized duality, addressing RLVR’s reliance on costly labels and dual learning’s limitation to strictly invertible tasks. The idea is simple: split https://x.com/gm8xx8/status/1959926238065127724
turns out training a model specifically for translation outperforms all other models, including gpt5. https://x.com/nickfrosst/status/1961093091554713686
This is illustrative: 1) When you get an instant AI answer, it is from a small model, which are weak models, especially at math. 2) Non-reasoning models, like the one powering AI overview, only “think” as they write, they make mistakes & then back justify them as they write more.”” / X https://x.com/emollick/status/1959609082031063139
Docent, our tool for analyzing complex AI behaviors, is now in public alpha! It helps scalably answer questions about agent behavior, like “is my model reward hacking” or “where does it violate instructions.” Today, anyone can get started with just a few lines of code! https://x.com/TransluceAI/status/1960411239919837654
It appears that the marginal energy used by a standard prompt from a modern LLM is relatively established at this point, roughly 0.0003 kWh (8-10 seconds of streaming Netflix) Water is more complicated (.25mL to 5mL+), depending on definitions. Training resources are less clear.”” / X https://x.com/emollick/status/1959989512228208785
Elon could start by making X the best matchmaker on the planet. Put Grok to work to build connections in the real world. Not pull us away from it into the arms of an AI companion.”” / X https://x.com/bilawalsidhu/status/1958570141119037766
xai-vs-apple-and-openai.pdf https://s3.documentcloud.org/documents/26073662/xai-vs-apple-and-openai.pdf




