Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Cinematic view inside a futuristic orbital station lab with a glowing spherical AI processor core floating at center, surrounded by suspended circuit boards and holographic technical displays, deep space visible through large viewport, cool blue and green lighting with dramatic rim lights, high-tech military aesthetic inspired by Ender’s Game, photorealistic science fiction rendering.
GPT-5, Claude, Kimi, and Gemini: “”I can travel back in time to any time before 1500 and change only one thing, what is the single thing you would change, nothing obvious.”” https://x.com/emollick/status/1987355374928769395
Introducing Nested Learning: A new ML paradigm for continual learning https://research.google/blog/introducing-nested-learning-a-new-ml-paradigm-for-continual-learning/
“I finally reached human-level performance (85%) on ARC-AGI v1 for under $10k and within 12 hours. I use the same multi-agent collaboration with evolutionary test-time compute, now powered by GPT-5 pro with lower parallelism. https://x.com/jerber888/status/1987982067116777521
@OpenAI’s GPT-5.1 delivers a solid upgrade from GPT-5 for agentic coding. We’ve noticed that the model is more steerable, overthinks less, and is better at frontend design. The model is also faster on most tasks because it dynamically adjusts reasoning depth based on the”” / X https://x.com/cognition/status/1989081722353529178
We’ve been testing Box AI with GPT-5.1 for the past week to compare it to GPT-5 for enterprise content use-cases. It’s a very strong upgrade from GPT-5. It’s super fast, performing ~2X (or more) faster on our tests on long documents (30,000+ tokens); and we saw an 8 percentage point gain in data extraction from our most our most challenging documents (across 1,000+ data fields) from a variety of content types. https://x.com/levie/status/1989051715207983511
GPT-5 on Sudoku-Bench 🧩 Since releasing Sudoku-Bench in May 2025, when no LLM could solve a classic 9×9 puzzle, we’ve been evaluating the latest generation of models. GPT-5 now leads our leaderboard with 33% puzzles solved–approximately 2x the previous leader–and is the first https://x.com/SakanaAILabs/status/1988080410392404021
GPT-5.1: A smarter, more conversational ChatGPT | OpenAI https://openai.com/index/gpt-5-1/
GPT-5.1 isn’t “GPT-5 but faster.” In our evals of the model, we found it’s the highest-precision model we’ve ever tested for code-related tasks like code review. Less noise, more fixes, reviews that read like patches again. https://x.com/coderabbitai/status/1989035006774354387
GPT-5.1 is a great new model that we think people are going to like more than 5. But with 800M+ people using ChatGPT, one default personality won’t work for everyone. We launched new preset personalities so people can make ChatGPT their own. https://x.com/fidjissimo/status/1988683216681889887
Moving beyond one-size-fits-all – Fidji Simo https://fidjisimo.substack.com/p/moving-beyond-one-size-fits-all
Baidu just dropped an open-source multimodal AI that it claims beats GPT-5 and Gemini | VentureBeat https://venturebeat.com/ai/baidu-just-dropped-an-open-source-multimodal-ai-that-it-claims-beats-gpt-5
AI progress and recommendations | OpenAI https://openai.com/index/ai-progress-and-recommendations/
This is an important one, I think. AI progress and recommendations: https://x.com/sama/status/1987232631680053745
Super original work! What if you could match non-identical objects? (just accepted to WACV’26 congrats!)”” / X https://x.com/Almorgand/status/1988240870986953120
🚨 Video Arena leaderboard update! 🎬 There is a new model provider in the Video Arena Vidu Q2 Turbo and Vidu Q2 Pro by @ViduAI_official have just made their debut with strong initial performance, both landing in the Top 10 for Image-to-Video: 🔹Vidu Q2 Turbo lands #6 with a https://x.com/arena/status/1989056583872180298
I keep coming back to GDPval, there is a lot in that paper that sheds light on the coming impact of AI on knowledge work, especially as agentic work starts to become a real thing, replacing the back-and-forth cyborg/centaur prompting we have used for years https://x.com/emollick/status/1988088613125714402
The Next Stage of AI Coding Evaluation Is Here https://news.lmarena.ai/code-arena/
As AIs get smarter & more useful, our benchmarks become less useful. Measuring general knowledge or coding ability gives us only a glimpse into what an AI model can do. Anyone who wants to use AI seriously for real work will need to assess it themselves. https://x.com/emollick/status/1988440050716279110
Most models: think → tool call → think → tool call K2 Thinking: keeps tool calls inside the reasoning trace so multi-step workflows don’t drift. We’ll show how Moonshot post-trained for agentic tool calling and demo complex workflows running in one model call.”” / X https://x.com/togethercompute/status/1988009780149878904
It turns out that Kimi K2 Thinking is also a beast at deep research. It can run 200-300 tool requests for impressive multi-agent capabilities. Would you like to see a code example of it?”” / X https://x.com/omarsar0/status/1987912692099682399
Kimi K2 Thinking is impressive. So I built a multi-agent deep researcher, Kimi Deep Researcher. It generates long research reports on any topic, powered by subagents (web searcher, analyzer, and synthesizer). It can do 100s of tool calls per session. Repo soon! https://x.com/omarsar0/status/1988974710592516454
These are pretty impressive benchmarks from a Chinese open weights model. Especially big is the agentic capability, which has generally lagged in the open weights models. Be interesting to see independent confirmation soon, I found K2 a solid, but kind of weird, model to use.”” / X https://x.com/emollick/status/1986452925418270871
🚀 Hello, Kimi K2 Thinking! The Open-Source Thinking Agent Model is here. 🔹 SOTA on HLE (44.9%) and BrowseComp (60.2%) 🔹 Executes up to 200 – 300 sequential tool calls without human interference 🔹 Excels in reasoning, agentic search, and coding 🔹 256K context window Built https://x.com/Kimi_Moonshot/status/1986449512538513505
🚀We’re going live with @Kimi_Moonshot on Nov 19 for a technical deep dive on Kimi K2 Thinking Learn about the 1T parameter MoE that allows your AI agent to make 300 tool calls in one run. Register: https://x.com/togethercompute/status/1988009777247510564
from Kimi AMA: – K3 will likely use KDA or some other hybrid attention mechanism – Kimi-K2 will get vision https://x.com/scaling01/status/1987916859400659011
I wonder if part of what makes Kimi K2 Thinking impressive is that it produces a lot more thinking tokens for even minor & non-technical queries than any model I have used. This is the thinking trace for “”write me a really good sentence about cheese”” it is 1,595 tokens long! https://x.com/emollick/status/1987286609713107261
Try Kimi-K2-Thinking now on Together AI https://x.com/togethercompute/status/1988011880443470217
I’m sorry Kimi bros The problem is and was 100% the OpenRouter API and it’s starting to piss me off that long reasoning always breaks Just use Kimi API for now and not OpenRouter if you have requests that take a lot of reasoning tokens. Simpler requests work fine with”” / X https://x.com/scaling01/status/1987938809628291168
since testing Kimi-K2 Thinking I have become very wary of providers on OpenRouter might switch to original provider APIs only they need to do quality testing for every model and provider”” / X https://x.com/scaling01/status/1988399213563236810
Kimi K2 Thinking passes the Lem Test the first time, very few models have done so Just like Kimi K2, however, this remains a very weird & interesting model in a way that is hard to benchmark. Its writing is often very good but sometimes doesn’t hold up under close investigation https://x.com/emollick/status/1986552301922738651
Thanks everyone for testing Kimi K2 Thinking and sharing benchmark results! We’ve noticed that benchmark outcomes can vary across providers. Some third-party endpoints show substantial accuracy drops (e.g., 20+ pp), which has negatively affected scores on reasoning-heavy tasks”” / X https://x.com/Kimi_Moonshot/status/1987892275092025635
Kimi AMA on K2 Thinking: 1. $4.6M training cost is not an official number 2. Trained on H800s (nerfed H100s) 3. KDA (Kimi Delta Attention) hybrids with NoPE MLA perform better than full MLA with RoPE 4. Muon scales well to 1T parameters. “there are tens of optimizers and”” / X https://x.com/Yuchenj_UW/status/1987940704929395187
Test out Kimi K2 Thinking vs. all the frontier models for yourself at: https://x.com/arena/status/1987947224173781185
Testing Kimi K-2 has reminded me of how insane it is that firms picking AIs are treating them as fungible based on benchmarks Kimi & Grok & Claude & every other model have strengths, quirks & weaknesses that can make a big difference in aggregate Develop your own benchmarks!”” / X https://x.com/emollick/status/1986604851770360213
In our new Expert and Occupational leaderboards: The previous, non-thinking Kimi K2 is ranked #7 for Hard Prompts, particularly excelling in the ‘Legal & Government’ category under the ‘Occupational’ leaderboard, while falling behind in ‘Instruction Following’. Kimi K2 Thinking https://x.com/arena/status/1987947222299013630
k2 vision is happening. this is not a drill. https://x.com/code_star/status/1987917177417289794
Whenever people ask me, “Is Muon optimizer just hype?” I need to show them this. Muon isn’t just verified and used in Kimi; other frontier labs like OpenAI are using it and its variants. It’s also in PyTorch stable now! https://x.com/Yuchenj_UW/status/1987955443420065816
Latest LisanBench results for Kimi-K2 Thinking Kimi-K2 Thinking is the best open-source model and 7th best model overall, right between GPT-5 and GPT-5-Mini Raw Scores: Glicko-2 ratings – better indicator of relative strength Kimi-K2 Thinking managed to set new high-scores https://x.com/scaling01/status/1987952884927934966
🚨 Leaderboard Update! Kimi K2 Thinking by @Kimi_Moonshot has landed on the Text leaderboard as the #2 open source model (MIT modified), tied for #7 overall. These are real-world results. With only a six-point difference with @Zai_org ‘s GLM 4.6, the competition is tight. Kimi https://x.com/arena/status/1987947219224526902
A must-read paper → Fundamentals of Building Autonomous LLM Agents Reviews the core cognitive subsystems that make up autonomous LLM-powered agents, including: – Perception – Reasoning & planning: CoT, MCTS, ReAct, Tree-of-Thought (ToT) techniques – Long- & short-term memory – https://x.com/TheTuringPost/status/1984686406430871892
Another banger whitepaper from Google. This time, they discuss context engineering and how to build effective memory for AI agents. Highly recommended read for AI devs. (bookmark it) I think this is an excellent intro on how to think about memory for AI agents. kaggle. https://x.com/omarsar0/status/1989081828678893837
New Video: What matters right now in mechanistic interpretability? A lot has changed in AI and interp! The priorities have moved on, frontier models are WAY more interesting now I discuss the new big picture, my vision for the field, common mistakes and promising directions https://x.com/NeelNanda5/status/1989297683140354267
Building our own inference platform was by no means easy, but would have taken significantly longer if not for @modal”” / X https://x.com/ArmenAgha/status/1988763002674508160
Since I’m really not into benchmaxxing, I’ve been underselling the evals but: we’re SOTA on anything non-code (*including* math). https://x.com/Dorialexander/status/1987977993440936433
benchmark designers should “train on the test set” to expose exploitable non-visual shortcuts”” / X https://x.com/sainingxie/status/1988019293926080611
New SWE/ML Leaderboard just dropped, like WeirdML but with a human baseline Turns out, all LLMs slow you down compared to human experts in ML/HPC optimization tasks (measured by runtime) https://x.com/scaling01/status/1989338806575903109
We’re proud to support @arcprize’s mission to build rigorous interactive benchmarks that measure generalized intelligence https://x.com/NousResearch/status/1988733248693027053
Sudoku-Bench https://pub.sakana.ai/sudoku-gpt5/
💥HipKittens, from Stanford ML, can 2x performance of AMD kernels vs RoCM’s composable kernels baseline. Kills baseline and up to 2x in some test.”” / X https://x.com/qubitium/status/1988389379984027742
Excited to share this piece from @VentureBeat spotlighting how Baseten is redefining the AI infrastructure game: “Baseten takes on hyperscalers with new AI training platform that lets you own your model weights.” Thanks VentureBeat! Read full article https://x.com/basetenco/status/1987943307532476746
AI has been built on one vendor’s stack for too long. AMD’s GPUs now offer state-of-the-art peak compute and memory bandwidth — but the lack of mature software / the “CUDA moat” keeps that power locked away. Time to break it and ride into our multi-silicon future. 🌊 It’s been a https://x.com/simran_s_arora/status/1988320513052324127
It’s official: SkyPilot integrates natively with @wandb! With SkyPilot + W&B, you can: 🚀 Launch and scale experiments on any cloud or K8s 🔄 Recover from GPU failures and seamlessly keep tracking experiments 🤝 Collaborate with your team – share SSH access and training logs https://x.com/skypilot_org/status/1989377870469501106
Fast Kernels on AMD hardware!! Great work from @simran_s_arora @_williamhu @Drewwad and team on HipKittens ( https://x.com/AnushElangovan/status/1988393252555493739
14 days. 2.2X faster inference. @Modular_AI + AMD Instinct MI355X GPU = state-of-the-art AI performance. A great example of what’s possible when cutting-edge GPUs meet next-gen AI software. 🔗 https://x.com/AMD/status/1987898172484567238
Build times for gigawatt-scale data centers | Epoch AI https://epoch.ai/data-insights/data-centers-buildout-speeds
Baseten used @nvidia Dynamo to double inference speed for long-context code generation and increased throughput by 1.6x. Dynamo simplifies multi-node inference on Kubernetes, helping us scale deployments while reducing costs. Read the full blog post below👇”” / X https://x.com/basetenco/status/1989058852789317717
Scaling Large MoE Models with Wide Expert Parallelism on NVL72 Rack Scale Systems | NVIDIA Technical Blog https://developer.nvidia.com/blog/scaling-large-moe-models-with-wide-expert-parallelism-on-nvl72-rack-scale-systems/
A new addition to the ERNIE open-source model family is here! Meet ERNIE-4.5-VL-28B-A3B-Thinking, our lightweight multimodal reasoning model. > 3B active parameters with enhanced semantic alignment between visual and language modalities > Outperforming Gemini-2.5-Pro and https://x.com/Baidu_Inc/status/1988182106359411178
The 4.5 trillion dollar elephant in the room https://stevenadler.substack.com/p/the-45-trillion-dollar-elephant-in
Great new capability in Databricks powered by our AI research team! We trained a document parsing system that delivers leading quality at 3-5x lower cost and outperforms leading VLMs like GPT-5 and Claude. This is critical to connect AI to so many kinds of data. https://x.com/matei_zaharia/status/1988325177193885885
You can now get more Codex usage from your plan and credits with three updates today: 1️⃣ GPT-5-Codex-Mini — a more compact and cost-efficient version of GPT-5-Codex 2️⃣ 50% higher rate limits for ChatGPT Plus, Business, and Edu 3️⃣ Priority processing for ChatGPT Pro and”” / X https://x.com/OpenAIDevs/status/1986861734619947305?s=20
GPT-5.1 is now live in Warp. It’s much faster (40% faster task completion on a subset of SWE-bench Verified) without compromising quality. GPT-5.1 is available to all Warp users, and is now the default model for all new users. https://x.com/warpdotdev/status/1989049715837829326
OpenAI’s $1 Trillion Infrastructure Spend | Tomasz Tunguz https://tomtunguz.com/openai-hardware-spending-2025-2035/
everyone complained that the GPT5.1 release yesterday had no benchmarks. now you have them. note minor regressions in AIME and Taubench, which increases confidence that this is not benchmarkmaxxing i think more generally model comms for a consumer AI model lab has to be split https://x.com/swyx/status/1989047883639980141
Anthropic to Outpace OpenAI in Server Efficiency, Internal Projections Show — The Information https://www.theinformation.com/articles/anthropic-projects-cost-advantage-openai
Multiverse Computing – Quantum AI software revolution. https://multiversecomputing.com/
for the first time 5.1 instant uses adaptive reasoning when responding. shipping this led the team down some pretty fun ML rabbit holes! if you’re an RL nerd, please reach out 🙂 I’ll be at neurips looking for folks to join the Science of Posttraining team”” / X https://x.com/allisontam_/status/1989138927970848936
We’re introducing Project AELLA, in partnership with @inference_net & @wyndlabs_ai AELLA is an open initiative to make 100M scientific papers accessible via LLM made structured summaries. Available now: – Dataset of 100K summaries – 2 fine-tuned LLMs – 3d visualizer 👇 https://x.com/laion_ai/status/1988330466706157818
Breaking: we release a fully synthetic generalist dataset for pretraining, SYNTH and two new SOTA reasoning models exclusively trained on it. Despite having seen only 200 billion tokens, Baguettotron is currently best-in-class in its size range. https://x.com/Dorialexander/status/1987930819021635964
In camera-ready, we include preliminary scaling experiments using Magistral, a near-frontier pure RLVR model, and the conclusion remains consistent. I’m also curious: if we scale RLVR compute to 10-1000× of Magistral, would it actually produce new knowledge beyond pretraining? https://x.com/YangYue_THU/status/1987716984524730604
This is a very cool paper. It first shows that first-gen college students don’t realize a lot of unwritten rules that lead to success (the value of internships, student clubs, letters from professors). But giving them access to an LLM for guidance significantly closes the gap.”” / X https://x.com/emollick/status/1987399908924498032
SkyPilot v0.10.5 is released! 🚀 • 18x more efficient managed jobs • Robust and fast SkyPilot API server at scale • Improved usability of Python SDK & admin policy and more. https://x.com/skypilot_org/status/1989083081953931284
[2511.06385] From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies https://arxiv.org/abs/2511.06385
From Demonstrations to Safe Deployment: Path-Consistent SafetyFiltering for Diffusion Policies https://tum-lsy.github.io/pacs/
PEFT has reached 20k GitHub stars ⭐⭐⭐⭐⭐ Right on time for the new PEFT release, v0.18.0. And it’s not a small one at that, with lots of new PEFT methods and other improvements. Details in the 🧵 https://x.com/BenjaminBossan/status/1988993386729390191
LangGraph. CrewAI. Agno. Which one to pick? The good news is that this will not matter soon! Finally, we have a full picture of how the industry is solving this with just three open protocols that work across ALL frameworks. It’s not about picking the best framework. https://x.com/_avichawla/status/1989228336997101946
We heard you like Jevons – by a16z New Media – a16z https://www.a16z.news/p/jevons-or-bust
All the great breakthroughs in science are, at their core, compression. They take a complex mess of observations and say, “”it’s all just this simple rule””. Symbolic compression, specifically. Because the rule is always symbolic — usually expressed as mathematical equations. If”” / X https://x.com/fchollet/status/1989340153114976598
Turns out you can communicate across containers via 63-bits of available space in a shared lock you acquire on /proc/self/ns/time that all processes have access to. No networking required. The post has a demo of a chat app communicating across unprivileged containers. https://x.com/eatonphil/status/1988616517609541872
Check out Hermes 4 70b & 405b models from Nous Research. Now available for you through both the Nous Research and Cline providers. (Nous Research brand booklet linked below) https://x.com/cline/status/1989432694867193988
What happens if we let models discover what training data is best… and when to see it? My proposal for getting beyond the data wall: 1/n https://x.com/joemelko/status/1987715636861251667
Excited to share our latest work on untangling language models by training them with extremely sparse weights! We can isolate tiny circuits inside the model responsible for various simple behaviors and understand them unprecedentedly well. https://x.com/nabla_theta/status/1989043939374924251
Today we’re releasing a full suite of models and a dataset: 1️⃣ Omnilingual ASR: A suite of ASR models ranging from 300M to 7B parameters, supporting 1600+ languages 2️⃣ Omnilingual w2v 2.0: a 7B-parameter multilingual speech representation model that can be leveraged for other”” / X https://x.com/AIatMeta/status/1987957744138416389
Paper: https://x.com/Azaliamirh/status/1988753594531868682
RF-DETR paper is finally on arXiv – real time detection with DINOv2 backbone – runs neural architecture search (NAS) over about 6000 architecture variants – uses weight sharing across all configs – first real-time segmentation DETR to break past top YOLO results ↓ more https://x.com/skalskip92/status/1989004912609411133
Baseten takes on hyperscalers with new AI training platform that lets you own your model weights | VentureBeat https://venturebeat.com/ai/baseten-takes-on-hyperscalers-with-new-ai-training-platform-that-lets-you
I don’t fully understand this paper (it’s a lot of numerical analysis and optimization stuff) but from what I can tell the gist is: 1. RL update imposes an implicit “”KL leash”” that prevents the policy from diverging too far away from the original model 2. when starting with a”” / X https://x.com/iScienceLuvr/status/1988756370867564689
dataset: https://x.com/_akhaliq/status/1987989916974829809
Introducing Lightning Grasp, a high-performance procedural grasp synthesis algorithm that generates thousands of dexterous grasps in seconds, across diverse hands and challenging objects. ⚡️ 10-100x faster than sota. Paper: https://x.com/zhaohengyin/status/1988318037804806431
📢📢 We know that in post-training, RL tends to generalize much better than SFT; on the other hand, we know both are nothing but ΔW on top of a base model weight. 🤔Why not open the blackbox of the model, and check how their ΔWs are different? 💡It turns out that RL takes a”” / X https://x.com/tydsh/status/1989049095575728156
Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following – Introduces a new benchmark with over 1,600 prompts and expert-curated rubrics to evaluate the ability to follow complex, multi-turn instructions – Introduces a novel post-training https://x.com/iScienceLuvr/status/1989274582822592634
[2511.07418] Lightning Grasp: High Performance Procedural Grasp Synthesis with Contact Fields https://arxiv.org/abs/2511.07418
SAEs assume that model representations are static. But LLM features drift & evolve across context! @EkdeepL @can_rager @sumedh_hrs’s new paper introduces a neuroscience-inspired method to capture these dynamic representations, plus some cool results & demos:”” / X https://x.com/GoodfireAI/status/1989010394380485083
Black-Box On-Policy Distillation of Large Language Models https://x.com/_akhaliq/status/1989341114760126965
Labeling data is one of the highest leverage things someone in a position of leadership can do”” / X https://x.com/model_mechanic/status/1987945123439931785
LightGlueStick: a Fast and Robust Glue for Joint Point-Line Matching TL;DR: SuperPoint backbone; Attentional Line Message Passing (ALMP) architecture which exposes the connectivity of the lines to the network https://x.com/Almorgand/status/1988649405449400663
Update: Dojo now supports the OpenEnv spec. https://x.com/chakra_ai/status/1989377867965513880
Common Ground between AI 2027 & AI as Normal Technology https://asteriskmag.substack.com/p/common-ground-between-ai-2027-and
Cut the Cost of Complexity Report 2025 | IBM https://www.ibm.com/reports/cost-of-complexity
Inline terminal output coming to @code. It will auto expand when the command fails, otherwise can be opened manually. Currently it only shows when the command has finished, you can turn this experimental feature on via `””https://t.co/4yOdCnWVjf.terminal.outputLocation””: “”none””` https://x.com/Tyriar/status/1989439441971396952
@itsclivetime Dynamic mixed precision seems to be the best path. Whatever on average optimizes for the least amount of energy required to flip the least number of gates to achieve the right answer.”” / X https://x.com/elonmusk/status/1987994042937036805
Even G2 and R1. A new quietly extraordinary power for your everyday life. #EvenRealities #EvenG2 https://x.com/EvenRealities/status/1988629357036753060?s=20
Feed the Beast | Derek Larson https://www.dtlarson.com/feed-the-beast
Today I decided to replace the KL penalty with some yolo crazy approach, which worked. When looking at it closely it is the standard KL penalty with a minor but very important change that assures a property I tried to obtain months ago without success. Today is a good day.”” / X https://x.com/francoisfleuret/status/1988364427675189640
How to Train an LLM: Part 1 – Omkaar Kamath https://omkaark.com/posts/llm-1b-1.html
Python inspect + GEPA is a truly wild combination”” / X https://x.com/JoshPurtell/status/1988025269006069845
🚨We converted pretrained LLMs into looped LLMs that can crank up performance by looping for more iterations. Our looped models surpass the performance of the pretrained models we started out with, showing that existing models benefit from increased computational depth. 📜1/9 https://x.com/micahgoldblum/status/1988265009508655528
first time i see a plot like this with a linear, not log-scale, x axis idk how this will play out, but super cool to see @dorialexander testing a different paradigm”” / X https://x.com/lateinteraction/status/1988016952451735772
The 2025 Production AI Stack Report | Temporal https://temporal.io/pages/ai-production-stack-report
Track your W&B runs and system metrics without ever leaving your terminal! Real-time, offline metrics delivered where you actually work. Try it now by upgrading to the latest version of the SDK. Let us know what you think! https://x.com/wandb/status/1988401739872301137
🤫 Something’s been brewing in stealth. Our SDK team’s side project, codenamed W&B LEET, is being unleashed. We are releasing a full Terminal UI (TUI) for live, interactive W&B monitoring right in your terminal. No browser, no internet, no problem. https://x.com/wandb/status/1988401253156876418
[whitepaper] Camunda Compared to Alternatives | Camunda https://page.camunda.com/wp-camunda-compared-to-alternatives-guide
baidu/ERNIE-4.5-VL-28B-A3B-Thinking · Hugging Face https://huggingface.co/baidu/ERNIE-4.5-VL-28B-A3B-Thinking
builddotai/Egocentric-10K · Datasets at Hugging Face https://huggingface.co/datasets/builddotai/Egocentric-10K
playing around with cudagraphs: they do bring nice speedups if we keep their limitations aside for a moment https://x.com/maharshii/status/1989375005231362428
one of the most exciting data / pretraining releases in quite a while this is a roadmap towards the “cognitive core” the x-axis here is log-scale (!!)”” / X https://x.com/willccbb/status/1987998615785402785
New paper! Language has rich, multiscale temporal structure, but sparse autoencoders assume features are *static* directions in activations. To address this, we propose Temporal Feature Analysis: a predictive coding protocol that models dynamics in LLM activations! (1/14) https://x.com/EkdeepL/status/1989009095953895756
At AI Dev 25 x NYC, @ozenhati (Head of Developer Relations, @GroqInc) showed how compound AI systems can build deep-research agents with a single API call. She walked through how agents choose tools, reason over results, and loop until they reach an answer — and why latency https://x.com/DeepLearningAI/status/1989431887224275433
Can LMs learn to faithfully describe their internal features and mechanisms? In our new paper led by Research Fellow @belindazli, we find that they can–and that models explain themselves better than other models do. https://x.com/TransluceAI/status/1989395421236793374
We did a deep dive into how to evaluate & benchmark LLMs. Read our recent blog to get up to speed on: ⚪️The 5 principles of good LLM benchmarking ⚪️Identifying LLM capabilities and limitations ⚪️Types of model evaluation methods Full rundown: https://x.com/togethercompute/status/1987949723106557975
The State of AI: Global Survey 2025 | McKinsey https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
So power is the core determinant of where AI data centers are built. Other factors like latency matter surprisingly little — it takes >100× longer to generate model responses than transmit data from Texas to Tokyo. Even serving LLMs from the Moon may not be a big latency issue!”” / X https://x.com/EpochAIResearch/status/1987944164286447895
Where does this power come from? Usually a mix of on-site fossil fuel generation and interconnection to the grid. E.g. Stargate Abilene will start off with on-site natural gas, then connect to the grid to access Texas’ abundant renewable power.”” / X https://x.com/EpochAIResearch/status/1987944176089231428
I keep warning that so many of our systems are still built around the assumption that quality writing and analysis are costly and therefore meaningful signals. Our systems are very much not ready for the revelation that this is no longer true, as this planning objection AI shows https://x.com/emollick/status/1987659629128479170
.@ylecun’s new paper lays out a full theory for JEPAs and turns it into a practical method – LeJEPA Together with @randall_balestr, they highlight 2 key ideas behind JEPAs: – The “”ideal”” form of JEPA embeddings is an isotropic Gaussian – A new SIGReg objective pushes the https://x.com/TheTuringPost/status/1989039076302049701
Most trackers lose sight of an object once it changes shape… [👇Code & Dataset] an apple turns into slices, a caterpillar into a butterfly, and the model just gives up. Researchers at Cornell built a new system called Track Any State that does something different: it follows https://x.com/IlirAliu_/status/1988319369160781978
📣 Announcing MUSI: 1st Multimodal Spatial Intelligence Workshop @ICCVConference! 🎙️All-star keynotes: @sainingxie, @ManlingLi_, @RanjayKrishna, @yuewang314, and @QianqianWang5 – plus a panel on the future of the field! 🗓 Oct 20, 1pm-5:30pm HST 🔗 https://x.com/songyoupeng/status/1975811164765643058
Nvidia presents TiDAR Think in Diffusion, Talk in Autoregression https://x.com/_akhaliq/status/1988963077690438097
🤖 From this week’s issue: A technical blog post explaining how NVIDIA TensorRT-LLM’s Wide Expert Parallelism efficiently scales large Mixture-of-Experts models on GB200 NVL72 systems, achieving significant performance and cost improvements. https://x.com/dl_weekly/status/1987913458654786008
We releasing a large update to 📄FinePDFs! – 350B+ highly education tokens in 69 languages, with incredible perf 🚀 – 69 edu classifiers, powered by ModernBert and mmBERT – 300k+ EDU annotations for each of 69 languages from Qwen3-235B https://x.com/HKydlicek/status/1988328336469459449
New friend @ZenMuxAI is now shipping GLM-4.6!”” / X https://x.com/Zai_org/status/1989005078926143810





Leave a Reply