Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A Byzantine gold-ground mosaic icon of a hammered-gold motherboard as sacred reliquary, its circuit traces rendered as illuminated-manuscript filigree with enameled chip-cabochons, a small golden clockwork bird perched on the heat-sink bough beneath a circuit-ring halo, warm candlelit glow on tesserae with deep purple and Tyrian crimson accents, the bold Trajan-capital word TECH in gilded ivory letters across the lower third, symmetrical centered composition, painterly tactile surface grain, 16:9.
Claude Opus 4.8 takes the lead on the Artificial Analysis Intelligence Index at 61.4, with Anthropic retaking the #1 spot on GDPval-AA and advancing in terminal use and scientific reasoning To reach the leading position on the Intelligence Index, @Anthropic made large
https://x.com/ArtificialAnlys/status/2060117582120976868
It “”feels like the first smart model in a long while”” due to this
https://x.com/zephyr_z9/status/2060077152729694586
The cost per accepted line of code varies by roughly 7x across model families.
https://x.com/cursor_ai/status/2060025070425395562
Announcing AA-WER Streaming, our new benchmark measuring streaming Speech to Text models on accuracy and latency for voice agent use cases. Pareto optimal models on this new benchmark include those from Cartesia, ElevenLabs, and Deepgram Streaming Speech to Text (STT) powers
https://x.com/ArtificialAnlys/status/2060021901234458958
Agent Judge: Solving Long-Context Evals for Production Agents — Judgment Labs
https://www.judgmentlabs.ai/blogs/agent-judge-solving-long-context-evaluations
Artificial Analysis and IBM Research are launching ITBench-AA, the first in a new series of benchmarks evaluating models on agentic enterprise IT tasks, starting with Site Reliability Engineering tasks where frontier models score below 50% ITBench-AA’s SRE tasks benchmark model
https://x.com/ArtificialAnlys/status/2059698327235805258
Interesting new SWE/agentic benchmark (DeepSWE) was released yesterday. 113 tasks across 91 repos in 5 languages. Here are interesting things I noticed: – The evaluation harness (mini-swe-agent) gives every model a single bash tool and the same SI. No vendor editing primitives.
https://x.com/_philschmid/status/2059564676569076021
NEW paper worth reading. A full agentic workflow can be distilled into model weights and run at roughly 100x lower inference cost while preserving near-frontier task quality. The workflow includes multi-step LLM calls, tool invocations, intermediate scratchpads, and decision
https://x.com/dair_ai/status/2057846601843146760
Opus 4.8 is live. Benchmarks especially significant jump in Agentic coding, but more important: „Fast mode is available for Opus 4.8. It’s the same model at roughly 2.5x the speed, and we’ve made it three times cheaper than before.”
https://x.com/kimmonismus/status/2060044465385902436
Anthropic just launched Claude Opus 4.8, and it is the new leader on our GDPval-AA benchmark for agentic real-world work tasks Opus 4.8 scored 1890 on GDPval-AA at launch with its ‘max’ effort setting, +137 points from Opus 4.7 and +121 points ahead of the next-best model,
https://x.com/ArtificialAnlys/status/2060042848268083411
Anthropic says Opus 4.8 ranks 1st on FrontierSWE
https://x.com/scaling01/status/2060046440563388838
Anthropic to introduce AI Fluency scorecard in Claude
https://www.testingcatalog.com/anthropic-to-introduce-personal-ai-fluency-scorecard-in-claude/
We have, as far as I can tell, no good tests of the productivity impact of the autonomous coding tools that appeared starting in December 2025. Every paper out there is from prior to the Claude Code/Codex revolution. A huge gap in our knowledge about what is happening in coding.
https://x.com/emollick/status/2059118330472972331
Qwen3.7 Max (20250517) debuts at #4 in Code Arena: Frontend – the top-ranked Chinese lab on the board, surpassing GLM-5.1 and is now on par with Claude Opus 4.6 on agentic web development tasks. Huge congrats to @Alibaba_Qwen on this achievement!
https://x.com/arena/status/2059297720079393107
Announcing Surya OCR 2: – 650M params – 83.3% olmocr bench score (top under 3B) – 87% on internal 91-lang benchmark – 5 pages/s on RTX 5090 – Runs on CPU, GPU, MPS
https://x.com/VikParuchuri/status/2059675773712167423
Help us produce the most useful work on AI by taking our 5-minute survey:
https://t.co/W2tLu3e4WW (You can sign up at the end to join our compensated user research panel.)
https://x.com/EpochAIResearch/status/2059336781208924566
Introducing BenchBench – by Rohit Krishnan
https://www.strangeloopcanon.com/p/introducing-benchbench
MaxSim v2 is out! Now finally with backprop so you can do training with this one. Bench compared to naive torch: – 10.33× faster on H200 – 11.94× faster on A100 And a bunch of other improvements like better bounds check, testing and better docs + examples of the APIs.
https://x.com/ErikKaum/status/2059659837219156453
MiniMax just teased their Sparse Attention architecture for M3. The benchmarks show 9.7x prefilling speedup and 15.6x decoding speedup at 1M tokens vs M2. MiniMax deliberately went back to full attention for M2 because efficient attention wasn’t production-ready. Their pretrain
https://x.com/kimmonismus/status/2059302121489486335
that’s an interesting eval code summary dishonesty at an all time low
https://x.com/scaling01/status/2060042892903678414
The experiments conducted in this post illustrate how early we are as an industry on eval tooling. Some takeaways and related thoughts: 1. Naively applying automation (which many current frameworks do) is likely to fail. 2. It’s easy to get fooled that automation (esp
https://x.com/HamelHusain/status/2057875320011882923
This is the first code bench that actually aligns with how it feels to use these models coding.
https://x.com/theo/status/2059352130289651925
I don’t think anyone has a good intuitive sense about what this means, and that failure of imagination is a generally bad thing for planning, investment, and policy. I also don’t have an easy solution (funnily enough, AIs are cliche at imagining the AI future, so no help there)
https://x.com/emollick/status/2057517583914389847
Most models are only evaluated on a fraction of the benchmarks out there. ArtifactLinker, our new system, predicts which ones would set a new state-of-the-art on benchmarks hosted on @HuggingFace, then runs the evaluation to verify. 🧵
https://x.com/allen_ai/status/2057838486204326078
Sonic 3.5 is now the #1 text to speech model on the @ArtificialAnlys leaderboard! You no longer have to trade off quality and latency – Sonic 3.5 also has the fastest time to first audio at 82ms end to end. See full benchmark results 👇
https://x.com/cartesia/status/2057880195403800633
Announcing ESMFold2, our new state-of-the-art structure prediction model capable of predicting structure from single sequences or MSAs. ESMFold2 improves on benchmarks of protein-protein interaction and is particularly strong on predictions of antibody-antigen complexes.
https://x.com/proteinrosh/status/2059633089702240598
MiniMax teases upcoming M3 model with new sparse attention mechanism and 15.6X long-context response speed boost | VentureBeat
https://venturebeat.com/technology/minimax-teases-upcoming-m3-model-with-new-sparse-attention-mechanism-and-15-6x-response-speed-boost
Huawei’s “Tao / τ Law”: Tech Paper, White Paper, or Strategic Manifesto? 🧠🚀 🌟Insights from Zhihu contributor 无我梦中 Huawei’s new paper, “A Time Scaling Theory for Multi-Layer Electronic Systems” by Tingbo He, is better read as a semi-technical white paper + strategic
https://x.com/ZhihuFrontier/status/2059118295580852374
Gemini 3.5 Flash has made huge progress from 3.1 Pro on GDPval, Flash is competing at the frontier, post training going strong 🙂
https://x.com/OfficialLoganK/status/2057682092583227881
Gemini 3.5 Flash is on the Pareto frontier of cost per intelligence on Vending Bench (a measure of a models ability to run a simulated store)!
https://x.com/OfficialLoganK/status/2058210519291777087
Gemini 3.5 Flash Looks Good For How Fast It Is | Don’t Worry About the Vase
https://thezvi.wordpress.com/2026/05/22/gemini-3-5-flash-looks-good-for-how-fast-it-is/
Gemini 3.5 Flash outperforms 3.1 Pro on many vision use cases (like the below Roboflow eval) while being ~6x faster on average 🤯 Gemini multimodal understanding for the win.
https://x.com/OfficialLoganK/status/2057888362011463988
Gemini 3.5 Flash ranks #1 on the APEX-Agents-AA benchmark, outperforming much larger models a whole size above it.
https://x.com/OfficialLoganK/status/2057460544643404125
some more recent ALE-Bench results: Grok-4.3 is pretty terrible, basically worse than all the frontier chinese models like Kimi-K2.6, DeepSeek-V4, GLM-5.1 and even Grok-4.2 lol Gemini 3.5 Flash only gets good with multiple iterations but gets mogged by Kimi-K2.6 and is also
https://x.com/scaling01/status/2057937081070944709
Gemini 3.5 flash release is underwhelming for browser agents. Slight improvement in performance over Gemini 3.1 pro, at a small increase in total cost
https://x.com/Alezander907/status/2057686331380359566
The HF science team just made async RL weight sync ~100x cheaper on bandwidth, and you don’t need a shared cluster anymore. The problem: every RL step, the trainer typically has to sync fresh weights to the inference engine. for a 7B in bf16 that’s ~14GB. for a frontier 1T fp8
https://x.com/ClementDelangue/status/2059989047947260203
Our team at @AIatMeta is excited to announce ATLAS: one of the largest automated formalization efforts to date. ATLAS contains Lean 4 formalizations of both statements and proofs from 25+ mathematics textbooks, spanning dozens of domains, for a total of 500k lines of code. We
https://x.com/arnal_charles/status/2060009395107377282
📢 New @heyjasper release ! 📢 MONET 🌸 : An Apache2.0 deduped and recaptioned dataset of 105M samples unlocking reproducible text-to-image research. Nano T2I 🖌️ : A codebase to train your own T2I model 🤗 @huggingface:
https://t.co/RrEDnjSzVg 💻:
https://t.co/gSgVcLXGO9 Very
https://x.com/CChadebec/status/2059983277306351674
Its very limiting that a big set of very hard problems that we have just lying around are Erdos problems. Don’t get me wrong, they are quite cool, but we really need hard problems repositories for many fields, including areas that have less specified answers & require judges.
https://x.com/emollick/status/2059012803009151444
Modern LLMs can do multiplication of 100-digit numbers without tools. So much for “”embers of autoregression””. Just scale the COT bro
https://x.com/teortaxesTex/status/2057826903721951273
[2605.22391] Epicure: Navigating the Emergent Geometry of Food Ingredient Embeddings
https://arxiv.org/abs/2605.22391
🚀 Better inference efficiency, lower costs, broader access. MiMo-V2.5 Series API pricing is now permanently reduced — by up to 99% compared to previous pricing. ✨ Unified pricing across all context lengths. MiMo Token Plans have also been upgraded: • 5-8× more usable tokens
https://x.com/XiaomiMiMo/status/2059314052892099070?s=20
🚨New Optimizer Paper AMUSE: Anytime MUon with Stable gradient Evaluation AMUSE combines Muon with Schedule-Free-style gradient evaluation for stable anytime training without LR decay. • Stronger 124M / 720M / 1B pretraining • Strong ImageNet / ViT fine-tuning performance.
https://x.com/jueunkim_0525/status/2059127584601055426
2026 AI Engineering Survey
https://notion.qualtrics.com/jfe/form/SV_bP07tSVMXH7ePCS
AI 101: From Tokens to Answers: What Actually Happens During LLM Inference
https://x.com/TheTuringPost/status/2057589545500049671
ai infra is going VERTICAL
https://x.com/swyx/status/2059463182297747527
An interesting method for multi-reward RL DVAO trains an LLM with several rewards at once, automatically deciding which rewards deserve more attention at each moment of training. Everything depends on the strength of the learning signal ↓ For example, if you train a model for
https://x.com/TheTuringPost/status/2059439959644467397
Behind the MiMo API Price Reduction: The deepest price cut, up to 99%, is for Input (Cache Hit). The core reason is our inference framework now supports hierarchical KV cache optimization for SWA. Production inference engine tests show this optimization increases cached token
https://x.com/_LuoFuli/status/2059618247553745204
Clouded Judgement 5.22.26 – The Neocloud Boom
https://cloudedjudgement.substack.com/p/clouded-judgement-52226-the-neocloud
For over a decade, we’ve accepted that end-to-end backprop is the only way to train deep networks. But holding the entire network in memory all at once is why AI training is hitting a resource wall. We found a new way to break the network into blocks and train them
https://x.com/hardmaru/status/2059648995132367277
Generative Recursive reAsoning Models (GRAM) – a reasoning model that goes both deeper and wider This framework, introduced by @KAIST_AI, @Mila_Quebec and others, creates multiple possible reasoning paths in parallel, exploring different hypotheses and solution strategies at the
https://x.com/TheTuringPost/status/2057642646345220576
Great work on Vector Policy Optimization (VPO). The standard scalar-reward view of post-training is inherently lossy: compressing a trajectory into one number discards a lot of useful structure, such as which sub-goals were met, where the reasoning failed, and what tradeoffs
https://x.com/FeiziSoheil/status/2057889865362993561
I just pushed support for training Z-Image L2P 1k with AI Toolkit. This is a pixel space variant of Z-Image that has a cool unet on the end to recover that lost high frequency that a lot of pixel space models miss. Compute cost is about the same as normal Z-Image.
https://x.com/ostrisai/status/2057931161889095928
I made a mistake. I underestimated Julius again. T3 Code’s remote feature is so far ahead of the remote control options offered by, well, everything else. 2 clicks to get a URL. Now I’m running a bunch of worktrees on my Mac Mini. Tailscale support built in too.
https://x.com/theo/status/2057960907997876412
I’m excited to share new research work from the Snowflake AI Research Team focused on advancing enterprise AI systems. Arctic-Text2SQL-R2 is a reasoning model designed for enterprise SQL generation. Trained on Snowflake-native data and optimized for real-world enterprise SQL
https://x.com/dwarak/status/2059686825086902398#m
imagine thinking you invented an LLM talking to another LLM and it was in December 2025
https://x.com/jxmnop/status/2060109869399916770
Introducing a minimal training harness built on prime-rl and verifiers, so you can now train your own RLMs without sandboxes! All available in the `training/` folder in the RLM GitHub repo! We train RLM-Qwen3-30B-A3B-v0.1, using RL on a separate split of environments
https://x.com/a1zhang/status/2059633834094678173
Introducing Apex: A Fast, Specialized Model for React Native
https://www.callstack.com/blog/introducing-apex-a-fast-specialized-model-for-react-native
Introducing the Cursor Developer Habits Report. We’re sharing some of our findings on how software development is changing. It’s based on the most comprehensive dataset on AI coding in the world, across all model families.
https://x.com/cursor_ai/status/2060025063899058458
Language Models Need Sleep “”Transformer-based large language models are increasingly used for long-horizon tasks; however, their attention mechanism scales poorly with context length. To handle this, we study a sleep-like consolidation mechanism in which a model periodically
https://x.com/iScienceLuvr/status/2059221770075562113
MoE (8): Enforcing Sequence-Level Balance
https://t.co/2H2wpwzQnS This article explores how to achieve sequence-level load balancing without incurring any loss penalty. Starting from the original Quantile Balancing (QB), we gradually derive a new method called Moving Quantile
https://x.com/Jianlin_S/status/2057719868917793221
new minimax sparse attention compared to deepseek v3.2 (DSA) and v4 (CSA) main changes: – based on GQA not MLA – block level selection like in CSA but attention is done on the real KV, not in the compressed dimension
https://x.com/eliebakouch/status/2059321928205156568
New paper: We present a “”Unified Neural Scaling Law”” functional form that accurately models & extrapolates the multivariate scaling behaviors of artificial neural networks as the variables listed in this attached video are varied. (1/N)
https://x.com/ethanCaballero/status/2059686905105563907
Papers with Code
https://paperswithcode.co/methods/on-policy-distillation
RL doesn’t work on Slurm. Modern RL pipelines are multiple services with different hardware, stable networking, and independent recovery needs. Slurm allocates nodes. The rest becomes glue code. We saw this across OpenRLHF, veRL, NeMo RL, and TRL. Here’s how SkyPilot Job
https://x.com/skypilot_org/status/2057854003648598312
RL has almost always meant trying to maximize a scalar reward. Very expressive in theory, but do you have only ONE scalar reward? Preferences & tradeoffs are complex & high-dimensional! Vector Policy Optimization (VPO) trains LLMs to anticipate diverse environments and goals!
https://x.com/lateinteraction/status/2057854814395019623
This is what we have been working on for the last 6 months or so at the AI Snowflake Research: Zero Redundancy Rollouts (ZoRRo):
https://t.co/OqiEPscuRL If you do RL and you want it to be much faster make sure to have a look.
https://x.com/StasBekman/status/2059718503318655314
We’ve created the world’s fastest PDF parser ⚡️ And it’s more accurate than any other open-source, model-free PDF parser out there (pymupdf, pypdf, markitdown, pdftotext, opendataloader, pymupdf4llm) Introducing LiteParse v2 – we rewrote the entire library into Rust and
https://x.com/jerryjliu0/status/2059710330016817501
What is the key bottleneck to scaling looped transformers (LT)? A major challenge is their speed: the looped operation is coupled w/ full quadratic attention. More loop, more powerful, but much slower. Introducing LT2: linear-time looped transformers that loop over linear
https://x.com/ChunyuanDeng/status/2057826955236462715
Why does deep learning generalize? What does weight decay really do? Can algorithmic information theory address these questions? In my latest preprint, I give a proof that the minimum neural weight norm matches the minimum program length (aka Kolmogorov Complexity), up to a
https://x.com/Tiberiu_Musat_/status/2059562156102746148
Why KV cache is one of the main reasons LLMs are fast? KV cache is what connects attention mechanism with generation stage of autoregressive models. These models generate text token by token, but each new token still attends to all previous ones. → To optimize decode phase,
https://x.com/TheTuringPost/status/2058744631194791974
Your RL post-training may be sabotaging your LLM’s test-time scaling! Conventional RL pretends that you can collapse all reward signals *upfront* into a single *scalar reward*. We introduce Vector Policy Optimization (VPO), which natively maximizes *vector-valued* rewards,
https://x.com/RyanBoldi/status/2057847412819906658
Cursor · The Cursor Developer Habits Report
https://cursor.com/insights
DeepSWE
https://deepswe.datacurve.ai/blog
I have been very impressed by @SemiAnalysis_ . I think of myself as a wide ranging systems engineer, looking for value at every level from the chip specs to the user interface, but SA exposes me to additional levels of “”the system””, both above (datacenters) and below
https://x.com/ID_AA_Carmack/status/2059382254191652896
Today, @MichaelElabd, @QuantumArjun, and I are excited to announce Trajectory. We are a research lab and product company building the platform for Continual Learning. Our platform unlocks the signal already sitting in product usage, so companies can continuously post-train
https://x.com/rronak_/status/2059644771262730624
PIXLRelight: Controllable Relighting via Intrinsic Conditioning”” TL;DR: physically grounded feed-forward relighting framework combining PBR conditioning and transformer rendering for controllable single-image illumination editing
https://x.com/Almorgand/status/2059759763546620149
Most companies talk about vector search. Few share what it actually takes to scale to 100M+ embeddings in production. Başak Eskili from @bookingcom joined the Weaviate Podcast to break down their AI journey, and it’s packed with insights about what building production systems
https://x.com/weaviate_io/status/2059227285639581729
the role of tactile sensing and force/torque feedback as a key enabler for developing reliable, touch-enabled Physical AI hardware?
https://x.com/IlirAliu_/status/2059104814919770206





Leave a Reply