Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: A pristine 1961 Ferrari 250 GT California Spyder in Rosso Corsa red positioned in a modern minimalist studio with floating holographic technical blueprints and circuit board wireframes emerging from its body, one side of the car transitioning to translucent schematic overlay, warm golden studio lighting mixing with cool blue holographic projections, polished chrome reflections, cinematic automotive photography with subtle depth of field, elegant and technical aesthetic.

Cranston AI (@cranston_ai) does your company’s bookkeeping & taxes with AI. Their agents pull in context from across the business and, after human review, file a full corporate tax return with the IRS. https://x.com/ycombinator/status/1975591950255358411

Big progress on this important benchmark (but still weird artifacts). https://x.com/emollick/status/1976702663330038205

an interesting forecasting benchmark: at current trends, we’re one year away from models matching the performance of superforecasters (and gpt-4.5 is sota on this benchmark!)”” / X https://x.com/gdb/status/1976139319787364408

Over 1.3 quadrillion tokens a month across Google, so much progress : ) so much more to go! https://x.com/OfficialLoganK/status/1976359039581012127

GPT-5 and Gemini 2.5 Pro just achieved gold medal performance in the International Olympiad of Astronomy and Astrophysics (IOAA). AI is now world class at cutting edge physics. https://x.com/deedydas/status/1977029236390285608

I don’t think people have updated enough on the capability gain in LLMs, which (despite being bad at math a year ago) now dominate hard STEM contests: The International Math Olympiad, the International Olympiad on Astronomy & Astrophysics, International Informatics Olympiad… https://x.com/emollick/status/1977460160197956089

🚨 🎬 Video Arena Disrupted! @Openai’s Sora 2 and Sora 2 Pro have landed on the Text-to-Video leaderboard. 🏆 Sora 2 Pro is the first to tie rank with Veo 3 variants for #1. 🥉 Sora 2 comes in at #3, pushing the non-audio variants of Veo 3 into 5th! Video models with audio https://x.com/arena/status/1978149396996051007

OpenAI’s Head of Sora @billpeeb says a stunning 70% of Sora’s nearly 2 million weekly active users are creating content. https://x.com/tbpn/status/1976759087456305191

This matches what the GDPEval paper found. Experts should try using AI a couple times on any task, and then resort to doing it themselves (with appropriate minor AI assistance) if they can’t get AI to work for them. You still save time overall, even when AI fails on some cases. https://x.com/emollick/status/1977874249214779558

Surfer 2 is here. 🏄🏄 Our new Cross-Platform Computer-Use Agent exceeds state-of-the-art on the 4 main benchmarks: WebVoyager, AndroidWorld, WebArena and OSWorld. 🖥️🌐📱 Find out more at https://x.com/hcompany_ai/status/1978935436111229098

Google’s Gemini 2.5 Native Audio Thinking is the new leading Speech to Speech model per our Artificial Analysis Big Bench Audio benchmark The new model achieves a score of 92% on Big Bench Audio, the highest result recorded by Artificial Analysis to date. This not only places it https://x.com/ArtificialAnlys/status/1977720537519636756

Princeton’s HAL leaderboard has AssistantBench, and o3 beats GPT-5 med on it to obtain 38.8% accuracy! Models have definitely been getting much better at answering personal-assistant questions, but there’s still more room to grow. https://x.com/OfirPress/status/1978925179876020247

Defining and evaluating political bias in LLMs | OpenAI https://openai.com/index/defining-and-evaluating-political-bias-in-llms/

The 2025 BEHAVIOR Challenge, launched by @drfeifei’s team at Stanford, is a global competition to train robots in household tasks, testing reasoning, navigation, and manipulation in simulated homes. ⦿ 50 tasks, 1,000 activities (e.g., cooking, cleaning) ⦿ 10,000 expert demos https://x.com/TheHumanoidHub/status/1976355634510737626

Elon Musk’s xAI joins race to build ‘world models’ to power video games with artificial intelligence https://www.afr.com/technology/musk-s-xai-joins-race-to-build-world-models-to-power-video-games-20251012-p5n1wj

We have a new state-of-the-art result on TheAgentCompany from Shanghai AI lab: MUSE + Gemini 2.5, solving 41.1% of the real-world inspired tasks. The new method is based on “”learning on the job””, a memory-based method. https://x.com/gneubig/status/1978564697499574761

Why does RL work for enhancing agentic reasoning? This paper studies what actually works when using RL to improve tool-using LLM agents, across three axes: data, algorithm, and reasoning mode. Instead of chasing bigger models or fancy algorithms, the authors find that real, https://x.com/omarsar0/status/1978112328974692692

[2509.25140] ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory https://arxiv.org/abs/2509.25140

RIP fine-tuning ☠️ This new Stanford paper just killed it. It’s called ‘Agentic Context Engineering (ACE)’ and it proves you can make models smarter without touching a single weight. Instead of retraining, ACE evolves the context itself. The model writes, reflects, and edits https://x.com/rryssf_/status/1976269613072843063

Is ACE the next Context Engineering Technique? ACE (Agentic Context Engineering) is a new framework that beats current state-of-the-art optimizers like GEPA by treating context as an evolving, structured space of accumulated knowledge. What is ACE? ACE treats context as an https://x.com/_philschmid/status/1977618096383721725

🚀 Introducing 𝐀𝐠𝐞𝐧𝐭 𝐒3, the most advanced computer-use agent, now 𝐚𝐩𝐩𝐫𝐨𝐚𝐜𝐡𝐢𝐧𝐠 𝐡𝐮𝐦𝐚𝐧-𝐥𝐞𝐯𝐞𝐥 𝐩𝐞𝐫𝐟𝐨𝐫𝐦𝐚𝐧𝐜𝐞🧠💻 Just one year ago, Agent S scored ~20% on OSWorld: SOTA then, but far from human 72%. Today, Agent S3 reaches 6̳9̳.̳9̳%̳ (⬆10% over https://x.com/xwang_lk/status/1973914981591838841

💃New Multi-Agent RL Method: WaltzRL💃 📝: https://x.com/jaseweston/status/1978185306999341256

we raised a $13m seed round from 120+ of Silicon Valley’s top investors for @mastra_ai, the leading TypeScript agent framework https://x.com/calcsam/status/1976346378013147359

[2510.08558] Agent Learning via Early Experience https://arxiv.org/abs/2510.08558

🔥Introducing #AgentFlow, a new trainable agentic system where a team of agents learns to plan and use tools in the flow of a task. 🌐 https://x.com/lupantech/status/1976016000345919803

Is your LLM-based multi-agent system actually coordinating? That’s the question behind this paper. They use information theory to tell the difference between a pile of chatbots and a true collective intelligence. They introduce a clean measurement loop. First, test if the https://x.com/omarsar0/status/1977784668323008641

Readers responded with both surprise and agreement last week when I wrote that the single biggest predictor of how rapidly a team makes progress building an AI agent lay in their ability to drive a disciplined process for evals (measuring the system’s performance) and error”” / X https://x.com/AndrewYNg/status/1978867684537438628

Open Agent Builder coming soon 👀 https://x.com/Dev__Digest/status/1976308510347673652

Open-sourcing retrieve-dspy! 💻🚀 While developing Search Mode for Weaviate’s Query Agent, we dove into the literature. It was amazing, and overwhelming, to see how many different takes on Compound Retrieval Systems there are! 📚 From perspectives on Reranking, such as to https://x.com/CShorten30/status/1978567334424932523

Ran Haiku-4.5 against my NYT Connections Eval with DSPY project and the results are in! – Baseline score of 64% – Optimized score of 71% – Complete in only 25 minutes – Total cost $11 This means Haiku 4.5 is the fastest model I’ve tested so far (ignoring Haiku 3.5 which did https://x.com/pdrmnvd/status/1978570006863790299

Salesforce AI Research introduces MCP-Universe: the first benchmark to truly test LLM agents in real-world scenarios with live Model Context Protocol servers. https://x.com/HuggingPapers/status/1959347736429674567

We tested Search Mode from Weaviate’s Query Agent on five popular Information Retrieval benchmarks — BEIR, LoTTe, EnronQA, WixQA, and BRIGHT! 📊 Of these benchmarks, we found the largest relative improvement from Search Mode over Hybrid Search on BRIGHT! ⚖️🚀 BRIGHT from https://x.com/CShorten30/status/1978107101936230745

GPQA Diamond and 𝜏²-Bench Telecom (an agentic benchmark requiring models to act in a customer service role) both show outsized performance for GPT-5 and o3 compared to GPT-4.1, but while the reasoning models cost >10x to run GPQA, in 𝜏²’s customer service environment they cost https://x.com/ArtificialAnlys/status/1978561356401111051

📣New paper: Rigorous AI agent evaluation is much harder than it seems. For the last year, we have been working on infrastructure for fair agent evaluations on challenging benchmarks. Today, we release a paper that condenses our insights from 20,000+ agent rollouts on 9 https://x.com/sayashk/status/1978565190057869344

Reasoning models are expensive to run with traditional benchmarks, but often get cheaper in agentic workflows as they get to answers in fewer turns Through 2025 we’ve seen test-time compute drive up the cost of frontier intelligence, but with agentic workflows there’s a key https://x.com/ArtificialAnlys/status/1978561353792344302

Saw that DGX Spark vs Mac Mini M4 Pro benchmark plot making the rounds (looks like it came from @lmsysorg). Thought I’d share a few notes as someone who actually uses a Mac Mini M4 Pro and has been tempted by the DGX Spark. First of all, I really like the Mac Mini. It’s https://x.com/rasbt/status/1978608882156269755

7. Small models can punch above their weight. Their best 4B-parameter model, trained with this recipe (real data + diverse RL + GRPO-TCR), beats 14B–32B models on tough benchmarks like AIME25 and GPQA-Diamond. Smart data and tuning trump raw size. Really good paper for AI devs”” / X https://x.com/omarsar0/status/1978112412743258361

The return of the physicists: “”CMT-Benchmark: A benchmark for condensed matter theory built by expert researchers.”” https://x.com/SuryaGanguli/status/1977740051108036817

Concern and excitement about AI around the world | Pew Research Center https://www.pewresearch.org/global/2025/10/15/concern-and-excitement-about-ai/

Introducing MAI-Image-1, debuting in the top 10 on LMArena | Microsoft AI https://microsoft.ai/news/introducing-mai-image-1-debuting-in-the-top-10-on-lmarena/

Tiny Recursion Model (TRM) results on ARC-AGI – ARC-AGI-1: 40%, $1.76/task – ARC-AGI-2: 6.2%, $2.10/task Thank you to @jm_alexia for contributing TRM, a well written, open source, and thorough research to the community based on the HRM from @makingAGI https://x.com/arcprize/status/1978872651180577060

Banger paper from Meta and collaborators. This paper is one of the best deep dives yet on how reinforcement learning (RL) actually scales for LLMs. The team ran over 400,000 GPU hours of experiments to find a predictable scaling pattern and a stable recipe (ScaleRL) that https://x.com/omarsar0/status/1978865039529689257

Meet our third @MicrosoftAI model: MAI-Image-1 #9 on LMArena, striking an impressive balance of generation speed and quality Excited to keep refining + climbing the leaderboard from here! We’re just getting started. https://x.com/mustafasuleyman/status/1977827977338716626

One of the most fun parts of OpenAI is watching people here level up so fast and do such excellent work. We are operating at a high level across many different disciplines and many of the people doing it have never done it before, and joined us at the beginning of their career.”” / X https://x.com/sama/status/1976799538523292027

AI is apparently already accelerating science. Measuring academic publications of authors: “we find that productivity among GenAI users rose by 15 percent in 2023 relative to non-users and further increased to 36 percent in 2024” and the quality of publications also went up. https://x.com/emollick/status/1977073589406122443

[2510.13786] The Art of Scaling Reinforcement Learning Compute for LLMs https://arxiv.org/abs/2510.13786

Really enjoying @1a3orn ‘s blog. Great technical explainer content, and thoughtful skepticism about AI risk. You know how @slatestarcodex said he finds about one great new blog a year? I fear I’ve just found mine. https://x.com/dwarkesh_sp/status/1977995499744735677

🚀 As reinforcement learning advances LLM reasoning (think O1-style breakthroughs), cost has become the key bottleneck. To solve this, @TencentHunyuan Reasoning & Pretrain team introduced a new RL approach — scaling reasoning without human-labeled data. Let’s see Hunyuan https://x.com/ZhihuFrontier/status/1977684644100468911

LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings https://arxiv.org/pdf/2510.08338

[2510.06105] Moloch’s Bargain: Emergent Misalignment When LLMs Compete for Audiences https://arxiv.org/abs/2510.06105

GitHub repo: https://x.com/karpathy/status/1977755430093980034

Mamba3 just silently dropped on ICLR🤯 A faster, longer-context, and more scalable LLM architecture than Transformers A few years ago, some researchers started rethinking sequence modeling from a different angle. Instead of stacking more attention layers, they went back to an https://x.com/JundeMorsenWu/status/1977664753011916859

I find it funny that a day after I post a paper, it is reposted by accounts that frame the paper as “”This is mind-blowing🤯”” or “”AI just destroyed X”” or some other bombastic phrase. On one hand, AI is really doing amazing stuff, on the other, pace your use of superlatives!”” / X https://x.com/emollick/status/1978155771880742995

Tandem Training for Language Models “”we pursue methods that encourage models to produce solutions that remain intelligible to weaker collaborators.”” “”we introduce tandem training for language models, a reinforcement learning (RL) paradigm in which rollout tokens are https://x.com/iScienceLuvr/status/1978794773747314765

My notes on nanochat, including links to the training data it uses https://x.com/simonw/status/1977867015818997883

Recursive Language Models | Alex L. Zhang https://alexzhang13.github.io/blog/2025/rlm/

1/9 Introducing LOTION (Low-precision optimization via stochastic-noise smoothing), a principled alternative to Quantization-Aware Training (QAT) that explicitly smooths the quantized loss surface while preserving all global minima of the true quantized loss. Details below: https://x.com/ShamKakade6/status/1978531483909259302

nanochat day 1: sharing everything we have: – we’ve got an org on the hub to share resources and discuss learning – we’ve trained a tokenizer and published it on the hub. – integrated base training with trackio for free logging. curves! if you’re also working on this. join the https://x.com/ben_burtenshaw/status/1978062142709326053

This is great work, the type where you can’t wait to try it once you hit the end of the paper. Well done @a1z”” / X https://x.com/dbreunig/status/1978873161841066464

A lot of problems with AI discourse are because “”being good at AI”” (called Theory of Mind in this paper) is a skill that seems to be independent of “”being great at your job”” So you have amazing experts who gain from AI, and others who do not, and they don’t understand each other https://x.com/emollick/status/1977581271019856099

I asked Nick if an analogy to Github repos would help illustrate the advantages of sexual recombination over asexual cloning or lateral gene transfer. Recombination is like a normal pull request – you have an organized diff at the same site as the previous functionality, and https://x.com/dwarkesh_sp/status/1976714084822450466

tpuf ANN v3 can search 100 billion vectors with a p99 of 200ms simplicity scales [v3 in beta, unfiltered search, 1024D, k=10, 92% recall] https://x.com/turbopuffer/status/1978173877571441135

The Art of Scaling Reinforcement Learning Compute for LLMs “”We present the first large-scale systematic study, amounting to more than 400,000 GPU-hours, that defines a principled framework for analyzing and predicting RL scaling in LLMs.”” “”we propose a best-practice recipe, https://x.com/iScienceLuvr/status/1978793969384624226

Sneak peak from a paper about scaling RL compute for LLMs: probably the most compute-expensive paper I’ve worked on, but hoping that others can run experiments cheaply for the science of scaling RL. Coincidentally, this is similar motivation to what we had for the NeurIPS best https://x.com/agarwl_/status/1978528133184725131

This is the most impressive plot I’ve seen all year: – Scaling RL not only works, but can be predicted from experiments run with 1/2 the target compute – PipelineRL crushes conventional RL pipelines in terms of compute efficiency – Many small details matter for stability & https://x.com/_lewtun/status/1978826407376458125

QeRL: Beyond Efficiency — Quantization-enhanced Reinforcement Learning for LLMs “”this is the first framework to enable RL training of a 32B LLM on a single H100 80GB GPU”” “”QeRL addresses these issues by combining NVFP4 quantization with Low-Rank Adaptation (LoRA), accelerating https://x.com/iScienceLuvr/status/1978046212621373719

Not All Bits Are Equal: What We Learned From 1700 Experiments on Memory-Optimal Reasoning Given a fixed memory budget, how should you allocate across model weights, KV cache, and test-time compute to maximize accuracy in reasoning models? For example: would you choose a 32B, https://x.com/DimitrisPapail/status/1978108550854382052

I feel fortunate working on problems this hard w/ such great operators A few things I’ve learned: 1. working on hard/challenging companies makes so many things easier 2. hardware is so much more enjoyable than pure play SW, you’ll never go back 3. measure projects in deep”” / X https://x.com/adcock_brett/status/1978498395166867571

Diffusion Transformers with Representation Autoencoders “”Most DiTs continue to rely on the original VAE encoder”” “”In this work, we explore replacing the VAE with pretrained representation encoders (e.g., DINO, SigLIP, MAE) paired with trained decoders, forming what we term https://x.com/iScienceLuvr/status/1978053094769615296

Diffusion Transformers with Representation Autoencoders https://rae-dit.github.io/

🚀We officially release Ring-1T, the open-source trillion-parameter thinking model built on the Ling 2.0 architecture. Ring-1T achieves silver-level IMO reasoning through pure natural language reasoning. → 1 T total / 50 B active params · 128 K context window → Reinforced by https://x.com/AntLingAGI/status/1977767599657345027

It should be mandatory to run DSPy with GEPA as a baseline in your paper. If you can’t outperform that, why are you publishing results?”” / X https://x.com/casper_hansen_/status/1977668375783596286

Nanonets released a new version of their SoTA OCR model 🔥 Supports LaTeX, Multilingual, Complex tables and much more Works out of the box with transformers, vllm and all major runners! 🤗 https://x.com/reach_vb/status/1978061301399052485

In an AI-native world, multi-tenancy is foundational. Most vector databases treat multi-tenancy as an afterthought. Weaviate built it into the core architecture, scaling to millions of tenants while optimizing costs and performance. Key innovations: • One shard per tenant = https://x.com/weaviate_io/status/1978112245436453044

This is the first project of @rikiyatakehi’s internship: sane, tiny ColBERT baselines for the ModernEncoder era. … and the 17M (0.017B) model beats every single model on the LongEmbed long context leaderboard 🤯”” / X https://x.com/bclavie/status/1978854449062793335

@TimDarcet (Which means, at least for me, its nonsensical to compare method that builds on foundation model vs model from scratch) (And if they want to compare, they should do it will total resources used)”” / X https://x.com/cloneofsimo/status/1978639683677806592

Lots of folks have been asking for a gist or simple notebook to try out RLMs. While we work on some more exciting experiments, here’s a self-contained, minimal version I quickly put together for people to build on top of. Happy hacking 🙂 https://x.com/a1zhang/status/1978948676287340753

I don’t know what labs are doing to these poor LLMs during RL but they are mortally terrified of exceptions, in any infinitesimally likely case. Exceptions are a normal part of life and healthy dev process. Sign my LLM welfare petition for improved rewards in cases of exceptions.”” / X https://x.com/karpathy/status/1976077806443569355

Artisanal shims for the bitter lesson age – nilenso blog https://blog.nilenso.com/blog/2025/10/14/bitter-lesson-applied-ai/

Finally, Python 3.14 lets you disable GIL! It’s a big deal because earlier, even if you wrote multi-threaded code, Python could only run one thread at a time, giving no performance benefit. But now, Python can run your multi-threaded code in parallel. And uv fully supports it! https://x.com/_avichawla/status/1977985594103140710

Btw the ratio of bandwidth-to-flops (dense fp16) on the DGX spark is not that outlandish. It’s actually quite comparable to server-grade machines (e.g. H100, B200). Maybe it feels unusual because it’s unusual for a consumer-grade machine. https://x.com/awnihannun/status/1978595876303167575

DGX spark explainer (And I still can’t wait to get one to run MLX on it!) https://x.com/awnihannun/status/1978544161688068225

Nick Lane has a theory about the evolution of life which explains *why* life is the way that it is – all the contingent mechanisms you’re just supposed to take as givens in biology class. He thinks early life was continuous with the spontaneous chemistry of deepsea hydrothermal https://x.com/dwarkesh_sp/status/1976697329043570772

Hybrid Reinforcement (HERO): When Reward Is Sparse, It’s Better to Be Dense 🦸‍♂️ 💪 📝: https://x.com/jaseweston/status/1977756142571864539

automatic CI tests are now live for the RL environments on the hub right now it’s just a simple integration test, but soon you’ll be able to run hosted debug evals as part of the CI on your environments lots of things coming to improve the quality assurance of RL environments”” / X https://x.com/johannes_hage/status/1978577811393884394

LLMs are getting better at character-level text manipulation | Tom Burkert https://blog.burkert.me/posts/llm_evolution_character_manipulation/

Pioneering GenAI for product development | Atlassian https://www.atlassian.com/whitepapers/pioneering-gen-ai-for-product-development

In my blog post on latents for generative modelling, I pointed out that representation learning and reconstruction are two separate tasks (§6.3), which autoencoders try to solve simultaneously. Separating them makes sense. It opens up a lot of possibilities, as this work shows!”” / X https://x.com/sedielem/status/1978143596701249733

1/8 Second Order Optimizers like SOAP and Muon have shown impressive performance on LLM optimization. But are we fully utilizing the potential of second order information? New work: we show that a full second order optimizer is much better than existing optimizers in terms of https://x.com/ShamKakade6/status/1978147672105353543

Are Large Reasoning Models Interruptible? “”In this work, we challenge the frozen world assumption and evaluate LRM robustness under two realistic dynamic scenarios: interruptions, which test the quality of the model’s partial outputs on a limited budget, and dynamic context, https://x.com/iScienceLuvr/status/1978044847216095361

Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity | Notion https://simonucl.notion.site/verbalized-sampling

Why are embeddings so cheap? – by Piotr Mazurek https://www.tensoreconomics.com/p/why-are-embeddings-so-cheap

Thanks AK for sharing. Code is available at https://x.com/yukangchen_/status/1978146373745639894

🧠 Besides this, here are some internal Q&A insights. 💡 Q1: How do you convert pretrain corpus into RL tasks? Two schemes tested: (a) Use LLM rewriting to normalize structure, then treat each line as a segment. (b) Use the NLP classic NLTK for segmentation. Results? Almost no”” / X https://x.com/ZhihuFrontier/status/1977688143005634992

Dr.LLM: Dynamic Layer Routing in LLMs Neat technique to reduce computation in LLMs while improving accuracy. Routers increase accuracy while reducing layers by roughly 3 to 11 per query. My notes below: https://x.com/omarsar0/status/1978829550709866766

Inflight updates + continuous batching are key infra updates needed for RL at scale.”” / X https://x.com/finbarrtimbers/status/1977738036403200088

now that memory will never be full. how else can we make memory better ?”” / X https://x.com/_samirism/status/1978621143172165915

skills/document-skills at main · anthropics/skills https://github.com/anthropics/skills/tree/main/document-skills

Imagine your LLM inference automatically getting faster in production (by up to 400%!) 🆕Enter: ATLAS–a not so traditional speculator that adapts to your workload as it evolves. The more you use it, the better it performs. https://x.com/togethercompute/status/1978210662095475097

BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution “”we introduce BigCodeArena, an open human evaluation platform for code generation backed by a comprehensive and on-the-fly execution environment. Built on top of Chatbot Arena, BigCodeArena https://x.com/iScienceLuvr/status/1977694597603291492

GLM-4.6 is now live on BigCodeArena. Shout-out to @qinkai1028 and the whole @Zai_org team for this great model!”” / X https://x.com/terryyuezhuo/status/1978554496058851650

This paper shows that you can predict actual purchase intent (90% accuracy) by asking an LLM to impersonate a customer with a demographic profile, giving it a product & having it give its impressions, which another AI rates. No fine-tuning or training & beats classic ML methods. https://x.com/emollick/status/1976656964622123141

🫧 Is AI a bubble? – by Azeem Azhar and Nathan Warren https://www.exponentialview.co/p/is-ai-a-bubble

Its staggering to see people use dinov2 over and over to fit Imagenet, and choose to not account the pretraining compute in the overall considerations. https://x.com/cloneofsimo/status/1978326330132283891

nanochat d32, i.e. the depth 32 version that I specced for $1000, up from $100 has finished training after ~33 hours, and looks good. All the metrics go up quite a bit across pretraining, SFT and RL. CORE score of 0.31 is now well above GPT-2 at ~0.26. GSM8K went ~8% -> ~20%, https://x.com/karpathy/status/1978615547945521655

One More (Small) Thing: Introducing mxbai-colbert-edge-v0 17M and 32M. They are are the result of an easily reproducible way to train ColBERT models from scratch. They’re strong, too: the 17M variant would rank first on the LongEmbed leaderboard for models under 1B parameters. https://x.com/mixedbreadai/status/1978853869557055492

Cognition | Introducing SWE-grep and SWE-grep-mini: RL for Multi-Turn, Fast Context Retrieval https://cognition.ai/blog/swe-grep

⬆️ LLMs’ forecasting abilities are steadily improving. GPT-4 (released March 2023) achieved a difficulty-adjusted Brier score of 0.131. Nearly two years later, GPT-4.5 (released Feb 2025) scored 0.101—a substantial improvement. A linear extrapolation of state-of-the-art LLM https://x.com/Research_FRI/status/1975909516777537614

day 3 on nanochat and it’s getting integrated! – we now have weights for the small and large models with 20 and 32 layers. Karpathy shared his own checkpoints on the hub. – I’ve got demos for both weights on the hub as space. – we have a PR on transformers to integrate NanoChat, https://x.com/ben_burtenshaw/status/1978832914952401081

The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections https://arxiv.org/pdf/2510.09023

The Turing Test for video … 😅”” / X https://x.com/demishassabis/status/1978644313824534954

StreamingVLM Real-Time Understanding for Infinite Video Streams https://x.com/_akhaliq/status/1977757009572237678

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading