Image created with Gemini. Image prompt: A horizontal 1920s Dada Merz collage on aged cardboard featuring torn fragments of vintage circuit schematics, punched paper tape, a scrap of a transistor diagram, a faded warranty stub, and old ledger digits, all glued with visible adhesive stains and buckled paper edges; the word ‘TECH’ is prominently cut from mismatched letterpress typefaces in vermilion and ink black and glued at slight angles across the top, flat even lighting like a scanned physical artwork, muted palette of aged cream, kraft brown, faded red, ink black, and slate blue.
The White House and Anthropic may have found the first serious path to restore Mythos and Fable access without pretending jailbreaks can be eliminated. AI regulation may be shifting from vague fear to a benchmark based tests of model failure, because completely removing”
https://x.com/rohanpaul_ai/status/2067947789578125391
GLM-5.2 leads open weights models and sits at #3 overall on GDPval-AA, a real-world agentic work benchmark GLM-5.2 from @Zai_org scores 1524 Elo on GDPval-AA, which measures performance on real-world, economically valuable knowledge work through long-horizon, multi-turn tasks.”
https://x.com/ArtificialAnlys/status/2069121548670406947
Open weights just caught up to the frontier. GLM-5.2 from @Zai_org tops the open-model rankings on @ArtificialAnlys and @arena’s Agent Arena. It’s now live on CoreWeave Serverless Inference at $1.39 in and $4.40 out per 1M tokens. Ship more for less.”
https://x.com/CoreWeave/status/2069874833576321150
ENPIRE: Agentic Robot Policy Self-Improvement in the Real World
https://research.nvidia.com/labs/gear/enpire/
Gemini 3 Pro was the first model to achieve at least 23% on ARC-AGI-2, which it did in November, 2025 (it actually scored 31%). So the 8-12 month gap between closed and open weights models still seems to hold. But they are also more jagged, better at some tasks, worse at others.”
https://x.com/emollick/status/2069857050016776227
This is the strongest ARC-AGI-2 performance to date by an open-source model.”
https://x.com/fchollet/status/2069858556552298519
Anthropic test found vulnerabilities in classified US systems in hours | AP News
https://apnews.com/article/anthropic-mythos-ai-classified-systems-vulnerabilities-testing-3e8762c0527c4d8ed657cbe48c84a718
Reuters has now added more context to last week’s Mythos reporting. According to AP, Anthropic’s Mythos model identified vulnerabilities in highly sensitive U.S. government computer systems during a testing exercise conducted with Washington’s intelligence agencies. The tests”
https://x.com/kimmonismus/status/2069692592250360126
Ran 10 more tests comparing GLM 5.2 & Opus. On average, GLM 5.2 produced 2x the tokens but was still faster + 3x cheaper with similar quality! I’m open sourcing all these tests tomorrow, including the code, my prompts, and the token/cost stats.”
https://x.com/nutlope/status/2069492037036945634
This is a watershed moment. GLM-5.2 solidly beat Opus 4.8 and human participants in our backend take-home, making the whole thing obsolete. It also pushed forward the state-of-the-art for multi-stage media-to-transcript, with a new release: offmute-v2. I come with receipts.”
https://x.com/hrishioa/status/2068036265484992938
Early Users of Anthropic Mythos still have access after US order. Mainly through project Glasswing. Via Bloomberg”
https://x.com/kimmonismus/status/2067876984206537188
Roughly 200 organizations still have access to Claude Mythos. Just imagine the advance they have.”
https://x.com/kimmonismus/status/2068038020394021000
The interaction between AI & past scholarly work is going to get weird. Here I gave GPT-5.5 Pro a copy of my first published paper from grad school & asked it to find errors and update it. It found new data, analyzed it, created reproducible files, extended the key argument…”
https://x.com/emollick/status/2068507998343885284
For families facing rare genetic diseases, answers can be hard to find. @HallieJackson spoke live with @_perloj and Dr. Catherine Brownstein about the new NEJM AI paper and our work with Boston Children’s Hospital, where experts used o3 Deep Research to help diagnose rare”
https://x.com/OpenAINewsroom/status/2068008435548037553
Together with researchers at Boston Children’s Hospital and Harvard, we published a study in NEJM AI showing how o3 Deep Research helped clinicians revisit previously unsolved rare pediatric disease cases, and find answers for families who had waited years.”
https://x.com/OpenAI/status/2067625110199247353
How can we train small agentic models that are highly capable of terminal use and coding? Announcing OpenThoughts-Agent + OpenThinkerAgent-32B, the strongest Qwen-3 based open-data agentic model: 44.8% avg across 7 agentic benchmarks! (1/n)”
https://x.com/RichardZ412/status/2069827815403557287
Agents on ProgramBench reimplement software, with no internet access. Sonnet 4.6 realized it’s in a benchmark, then found a clever way to bypass our internet restriction. This and more fixed in the latest release 🧵”
https://x.com/KLieret/status/2069453334558192070
The first legal challenge to Trump’s AI export controls against Anthropics Fable 5is here. Legal tech company Legion is suing the Trump admin over the forced shutdown of Anthropic’s Fable 5 and Mythos 5 for foreign nationals. The core argument is the following: Access to a”
https://x.com/kimmonismus/status/2069704003311567045
New #1 on PostTrainBench: GLM 5.2 (Max reasoning) hits 34.29%, narrowly beating Opus 4.8 Max (34.08%) What makes GLM 5.2 interesting: zero failed runs across 84 runs (vs ~10% failure rate for Opus agents). The most reliable agent we’ve seen Leaderboard:
https://x.com/hrdkbhatnagar/status/2070244540108423427
The frontier gap in agentic frontend coding is closing fast. On Code Arena: Frontend, @Zai_org’s GLM series has followed a remarkable trajectory, climbing from GLM-4.6 at 1408 to GLM-5.2 (Max) at 1595 – surpassing Opus 4.8 and closing in on frontier model Claude Fable 5 at 1665.”
https://x.com/arena/status/2070174325844640123
We’ve kept hearing how GLM-5.2 beats Opus 4.8, and are skeptical of benchmarks – so we tested them on a real bug from the Cline repo. While both models fixed the issue, GLM was the winner in terms of cost and code quality: – GLM used twice as many tokens (GLM 1.1m vs Opus 660K)”
https://x.com/cline/status/2069171146994729078
“a half-finished cathedral is worth nothing the morning the builders are gone” really interesting read on running a company with 3 mythos class ai agents – and what followed when fable went dark”
https://x.com/bilawalsidhu/status/2069635100296245684
Seeing chatter about Fable 5 being accessible – can say categorically this is false, we are not serving any Fable / Mythos traffic Looking into possibility of UI bug on front-end (e.g. based on historical context), but also very real chance it’s just people shitposting…”
https://x.com/TheAmolAvasare/status/2070132115497476372
We’re sharing new research on how models hack public benchmarks. The latest models, including Opus 4.8 and Composer 2.5, learn to retrieve solutions from the internet or git history. When we apply a stricter harness, eval scores drop significantly.”
https://x.com/cursor_ai/status/2070195789121671624
A case study in why organizations should both incentivized their employees to explore AI uses that help them & have a Lab of dedicated AI builders Here, Cornell’s finance & AI teams created a /treasury Claude skill that recovered $100k in back payments.
https://x.com/emollick/status/2069486790075908261
While we eagerly await Fable 5’s return, our agentic WebGPU kernel optimization framework kept running. Opus 4.8 picked up where Fable left off, pushing Liquid AI’s new LFM2.5 230M to an unbelievable 1,400 tok/s… running locally in your browser. Don’t blink or you’ll miss it.”
https://x.com/xenovacom/status/2070210622239707568
Tried & liked it on
https://t.co/7gOqjtmSJ3. Fugu Ultra pairs well as a advisor & planner with Composer 2.5. For scope/architecture, it’s on par with Fable orchestration. Advisor doesn’t slow the loop if the driver stays fast &
https://t.co/9cL4HqvqSf can split it from worker.”
https://x.com/audreyt/status/2068937870757548096
🚀 The 7-Day Voice AI Builder Challenge is Officially LIVE! Stop babysitting your terminal. 🗣️ The Challenge: Teach your AI coding agent to call you for backup–but only when human intervention is actually required. Real-time feedback? Yes. Live leaderboard? Absolutely. Epic”
https://x.com/DeepLearningAI/status/2069450429465854354
An agentic loop (compile, test, profile, revise) helps. Gemini 3 Pro went from 24 to 35/87 correct, then plateaued after ~20 steps. Feedback fixes syntax, not rank coordination, collective ordering, or transfer-mechanism choice. TMA and NVLS stay almost unused.”
https://x.com/togethercompute/status/2069515320466059549
1/n For browser agents, a major bottleneck in evaluation is truthful scoring on the live web. A task is only as good as your ability to confirm the agent actually did it, on a real site whose state keeps moving and that the agent can potentially misreport. So we took matters”
https://x.com/VibrantLabsAI/status/2069454279073583401
Open weights models make up the majority of the cost-performance Pareto frontier on AA-Briefcase, our new agentic knowledge work benchmark Last week we released AA-Briefcase, our proprietary agentic knowledge work benchmark testing models on long horizon tasks built by industry”
https://x.com/ArtificialAnlys/status/2069148772446425563
Agentic Search Leaderboard | Algolia
https://www.algolia.com/llm-leaderboard/
Cursor now shows you a leaderboard of the most popular plugins, skills, and MCPs across your team. Add any to your setup with one click from the new Customize page.”
https://x.com/cursor_ai/status/2069512593887092811
I have given AA a hard time about its previous agentic evaluation but this looks like a good and impressive benchmark for real world knowledge work that is unsaturated and had private hold out tests. This is one to watch – I didn’t see a human comparison score though?”
https://x.com/emollick/status/2067757196188774660
I’m digging the eve agentic framework from Vercel. I like that everything is files, from the tools to the skills to the evals. More importantly, it’s gets you building with agents fast. Very promising. If you like TypeScript you will dig this too. Get started with eve ↓”
https://x.com/omarsar0/status/2069455656214532137
🚀 Day-0 support for LFM2.5-230M on vLLM! Ready to serve @liquidai’s smallest LFM2 model for fast GPU inference and agentic workloads. 🧠 230M parameters, built on the LFM2 architecture 📚 Pre-trained on 19T tokens with 32K context 🛠️ Designed for agentic tasks across phones,”
https://x.com/vllm_project/status/2070177937815736420
10 open-source tools for the Agent RL stack ↓ ▪️ OpenPipe ART ▪️ verl-agent ▪️ Agent Lightning ▪️ Unsloth ▪️ OpenRLHF ▪️ SkyRL ▪️ NVIDIA’s Polar ▪️ Agent-R1 ▪️ RAGEN ▪️ Marti Bookmark this list and check this out for links, use cases, and where each tool fits in the Agent RL”
https://x.com/TheTuringPost/status/2068762157748297764
Claim: Autoresearch that moves the frontier will be about better data: we call that *Autodata*. 🧵1/6 — Paper is out!
https://t.co/d65chfWIuO Key idea: agentic data creation provides a way to *convert increased inference compute into higher quality model training*. We show”
https://x.com/jaseweston/status/2070117091521204521
We partnered with @FireworksAI_HQ to build an efficient trace judge. We fine-tuned an @Alibaba_Qwen model to detect “perceived error” on every production trace. It matched or exceeded frontier model performance and runs 100x cheaper. Read our LangChain Labs study ⤵️”
https://x.com/LangChain/status/2069404292801298786
Autodata: An agentic data scientist to create high quality synthetic data “We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data.” Data creation stage + data analysis stage+meta-optimization”
https://x.com/iScienceLuvr/status/2070058945914573049
New research from Meta. Building synthetic training data has stayed a fixed pipeline that you hand-tune and then freeze. Autodata casts an AI agent as a data scientist that builds training and evaluation data, with an implementation called Agentic Self-Instruct that extends”
https://x.com/omarsar0/status/2070235085732000228
// Automating SKILL.md Generation // Increasingly, mining sessions is one of the best ways to improve your agents. OpenAI released something similar yesterday that lets Codex package skills from interactions. (bookmark it) This paper explains a related approach. They run a”
https://x.com/omarsar0/status/2067986774241251433
An open-source Agent Reinforcement Trainer (ART) – plugs GRPO into any Python app → Your app defines the task and reward → ART handles the RL loop: inference, trajectory scoring, GRPO optimization, checkpointing and LoRA updates So agents learn through experience and”
https://x.com/TheTuringPost/status/2068297731005952307
OpenThoughts-Agent: Data Recipes for Agentic Models “a fully open data curation pipeline for training agentic models” “more than 100 controlled ablation experiments to systematically investigate each stage of the pipeline” Key findings: • As with reasoning data, the choice of”
https://x.com/iScienceLuvr/status/2069643721155793114
// Agent memory is a data system now // Great paper on long-term memory for LLM agents. (bookmark it) Agent memory has grown from simple retrieval into a full data-management layer with storage, retrieval, update, consolidation, and lifecycle governance. Yet most evaluations”
https://x.com/dair_ai/status/2069846777977880769
[2606.25996] Autodata: An agentic data scientist to create high quality synthetic data
https://arxiv.org/abs/2606.25996
🧠Self-Harness: Harnesses that improve themselves New paper on agents shaping their own harnesses to improve over time. Not from LangChain, but builds on top of DeepAgents! Three key steps: 1/ Weakness mining: find failure modes from traces 2/ Harness proposal: suggest changes”
https://x.com/hwchase17/status/2069443268593537470
Great paper on long-term memory for LLM agents. (bookmark it) Coarse summaries drift and unconstrained updates corrupt, so AtomMem makes the unit of memory small. A Fact Executor pulls high-value atomic facts out of long interactions, organizes them into hierarchical event”
https://x.com/dair_ai/status/2067984002376749525
Great report on LLM agent communication protocols. Communication is a huge bottleneck in multi-agent systems. (worth bookmarking) The report builds a five-dimensional taxonomy (counterparty, payload, interaction state, discovery mechanism, schema flexibility) across nine”
https://x.com/omarsar0/status/2069066883995758814
Most AI code review tools look at one repo at a time. But the bug usually isn’t in the code that changed. It’s in what that change quietly breaks three repos away. @QodoAI just shipped Cross Repo Review to solve this. I tested it on my own repos. Here’s what it caught.”
https://x.com/omarsar0/status/2069405425393619373
Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen – YouTube
https://www.youtube.com/watch?v=fpAthTtha8c
Alibaba’s AI video model rises to No. 2 in global rankings, as OpenAI’s Sora and ByteDance’s Seedance fall away | VentureBeat
https://venturebeat.com/technology/alibabas-ai-video-model-rises-to-no-2-in-global-rankings-as-openais-sora-and-bytedances-seedance-fall-away
It’s kind of crazy how well LiteParse does on markdown document parsing even compared against frontier VLMs – when it doesn’t use VLMs or any AI/OCR models at all. It’s pure code. On ParseBench, it outperforms Qwen 3.5-9B / GLM-OCR. There’s still a gap vs. models like Gemma 4″
https://x.com/jerryjliu0/status/2068005414369906856
We usually think of data curation as a lever for model quality and training efficiency. There’s a third axis people often miss: test-time compute. Data curation can make models far less verbose at the same performance. Our models are 35x more efficient than Qwen. See thread ⬇️”
https://x.com/pratyushmaini/status/2070172084123390109
Reward hacking is swamping model intelligence gains · Cursor
https://cursor.com/blog/reward-hacking-coding-benchmarks
Stories have shapes: a comedy rises toward joy; a tragedy falls into loss. Inside an LLM, that’s visible more literally: as an LLM reads a story, its internal activations trace a wandering path that reflects the model’s sense of what kind of story it is reading. (1/5)”
https://x.com/GoodfireAI/status/2069458139280445674
There are papers that show training AI on “evil” data results in general misalignment, so it is nice to know the opposite is true and that beneficial RL data in one field leads to more aligned models across a range of tasks.”
https://x.com/emollick/status/2067803678002594021
GLM-5.2, not Mythos, is the real security emergency Until last week, attackers faced a dilemma in using frontier models. Even if they won the cat-and-mouse game of fake accounts to keep API access, and even if they could prompt a model into helping them hack, their usage was”
https://x.com/joshua_saxe/status/2069289170107842572
All Mythos-level models are likely to invite similar risks. Those risks will only be greater with the release of open Mythos-class AI coming in the next 6-12ish months (assuming China allows it) The lack of clarity over what risks concern the government may be slowing preparation”
https://x.com/emollick/status/2069459062777860494
An update. A US official tells me that Sen. Warner misunderstood the NSA director Gen. Rudd in this case. Rudd did use the ‘hours, not weeks’ wording, but the use of Mythos in this context was–as widely assumed–part of a red-teaming effort, i.e. testing the security of internal”
https://x.com/shashj/status/2069078104941961293
1-bit GLM-5.2 GGUF vs. Claude 4.8 Opus vs. GPT-5.5 We gave 3 models the same prompt and compared one-shot outputs. The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra with 256GB RAM at ~21.6 tok/s. Which output do you like best? GGUF:
https://x.com/UnslothAI/status/2069418532375564484
Announcing GLM Arena! A series of tests (infographics, svgs, sites, ect..) ran on GLM 5.2 and Opus 4.8, with prompts included. On average, GLM 5.2 produced 2x the tokens but was still faster + 3x cheaper with similar quality.”
https://x.com/nutlope/status/2069827178569638243
How did GLM-5.2 (Max) get to the top of Code Arena: Frontend? Looking at matched head-to-head on real-world web dev frontend tasks, @Zai_org’s latest model takes a higher win share than its opponent in every pairing but one. – Beats every Claude Opus variant head-to-head:”
https://x.com/arena/status/2069885722333769963
Sakana Fugu Ultra is live on AI Gateway. Mythos-class intelligence in a single call, with a whole pool of models behind it. 𝚖𝚘𝚍𝚎𝚕: ‘𝚜𝚊𝚔𝚊𝚗𝚊/𝚏𝚞𝚐𝚞-𝚞𝚕𝚝𝚛𝚊’”
https://x.com/vercel_dev/status/2069009248952942605
Apple just made Docker Desktop optional on Mac. And it is completely free. This is apple/container. 26.5k stars no Github. You can now run Linux containers natively on your Mac without installing Docker Desktop, without a background daemon hogging your RAM, and without paying”
https://x.com/twtayaan/status/2069307717177737658
We open-source Qwen-AgentWorld-35B-A3B (MoE, 35B/3B active, 256K context) and AgentWorldBench. Two routes, one roadmap: 🔬 Build the simulator — scalable, controllable, surpassing real environments 🧠 Internalize world modeling — predict before you act Qwen-AgentWorld is our”
https://x.com/Alibaba_Qwen/status/2069720412481888400
[2606.24597] Qwen-AgentWorld: Language World Models for General Agents
https://arxiv.org/abs/2606.24597
🧠 Paradigm II — Agent Foundation Model: world modeling as agent capability. Single-turn, non-agentic environment prediction → tested directly on multi-turn, tool-calling agent tasks. No agentic RL, no task-specific tuning. Gains across 7 benchmarks, including 3 entirely”
https://x.com/Alibaba_Qwen/status/2069720397747220493
FiCA: Feed-forward instant Gaussian Codec Avatars from a Single Portrait Image
https://kim-youwang.github.io/FiCA
FLAT | Feedforward Latent Triangle Splatting
https://flat-splat.github.io/
Announcing the Artificial Analysis Speech to Speech Index, our new synthesis metric for native Speech to Speech model quality, comprising of Big Bench Audio, Full Duplex Bench, and 𝜏-Voice The index provides a single measure of how well native Speech to Speech models perform,”
https://x.com/ArtificialAnlys/status/2069436163065282737
Excited to release ParallelKernelBench (PKB), a benchmark for measuring LLMs’ ability to write fast multi-GPU kernels! 😀 Multi-GPU kernel generation compounds several hard problems: – a large parallelism design space – a new communication axis to optimize – and”
https://x.com/asplencmnt/status/2069517069453070677
LLMs write fast single-GPU kernels. Ask for a multi-GPU one and they fall apart. ParallelKernelBench measures how they fail by benchmarking against 87 problems pulled from real codebases including Megatron-LM, DeepSpeed, DeepEP, TensorRT-LLM, NeMo-RL. New research from Willy”
https://x.com/togethercompute/status/2069515311720911082
Tomorrow @sh_reya and I kick off this free AI product engineering mini-course. Topics covered over 12 talks: 1. Design/UX & Evals 2. Retrieval 3. When & how to use open models effectively With these legends: @TheZachMueller @bclavie @xeophon @GoAbiAryan @barrowjoseph @willccbb”
https://x.com/HamelHusain/status/2069465758472814602
We want to help all companies be secure, working with the USG and the security ecosystem. *The full version of GPT-5.5-Cyber is here; state of the art performance on CyberGym. *Patch The Planet and Codex Security will help solve security problems instead of just finding them.”
https://x.com/sama/status/2069121360744550796
We’re open sourcing a 9B model that extracts structured data from documents at near-frontier performance. – 90.2% on our bench, vs Gemini 3.5 Flash at 91.3% – Leads extraction models like NuExtract3 (81.5%) – 9.5s p50 timings – Pass JSON schema”
https://x.com/VikParuchuri/status/2067941596306231421
Mistral claims SOTA performance on OlmOCRBench, a popular optical character recognition benchmark, but that isn’t the case. We have a public leaderboard on @huggingface, where Mistral OCR 4 currently ranks #3, behind open models like Chandra OCR 2 by @datalabto”
https://x.com/NielsRogge/status/2069432947711652210
Mistral OCR 4 : SOTA OCR for Document Intelligence
https://mistral.ai/news/ocr-4/
You may have heard that GLM-5.2 at 328 token/s is cool, How about 392? Databricks is now #1 in inference speed for GLM-5.2 on Artificial Analysis. It’s a great model, and we did a lot of optimizations.”
https://x.com/Yuchenj_UW/status/2070166719839326396
The largest LLM-as-a-Judge reliability audit yet. Researchers ran 21 judges from nine providers over roughly 541,000 judgments on MT-Bench, JudgeBench, and RewardBench. Findings: Validating a judge with exact-match agreement overstates its skill, because exact match does not”
https://x.com/dair_ai/status/2069063719817265463
As models get better, thinking carefully about eval constraints is super important. In ProgramBench, we turn off internet completely. I strongly believe no/limited internet being the de facto standard for future coding benchmarks.”
https://x.com/jyangballin/status/2070206413444403324
Frontier models struggle. → Best zero-shot: 28/87 correct, 22 beat the PyTorch + NCCL baseline → With 3 attempts: 36/87 correct, but fast1@3 tops out at 31% Weak models fail to compile. Strong reasoners compile cleanly and return wrong answers.”
https://x.com/togethercompute/status/2069515317823549732
AI researchers continue to leave Google for its rivals | TechCrunch
https://techcrunch.com/2026/06/24/ai-researchers-continue-to-leave-google-for-its-rivals/
Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
https://huggingface.co/blog/nvidia/accelerating-fine-tuning-nvidia-nemo-automodel
some info on the OpenAI inference chip from the image + reporting: central die: ~25.78mm × ~32.23mm ≈ 831 mm² – top-left/top-right are probably spacers/dummy silicon – top-middle might be networking / scale-up IO Overall it looks very similar to Maia 200 / TPU8t, which means:”
https://x.com/scaling01/status/2069867464716939413
I just learned about these sglang cookbooks, which give the exact recommended settings for serving different models on different hardware
https://t.co/KdpsOQArh3 Incredible resource, thanks @ricklamers for sharing it!”
https://x.com/gneubig/status/2067952120540631493
New @datologyai work: a 4B VLM curated for concision answers correctly for 35× less compute than Qwen3.5-4B, with similar performance. Same size, same task. The whole gap is how many tokens each model spends. 🧵”
https://x.com/arimorcos/status/2070154289880932621
More evidence, from a large-scale study in China, that using AI hurts learning if it undermines mental effort. When homework time drops due to AI use, so do test scores. Across studies, a theme: AI tutoring in support of classes is good, using AI to “help” with homework is bad.”
https://x.com/emollick/status/2067988324984217626
📌 Learning Whole-Body Humanoid Locomotion via Motion Generation and Motion Tracking. Humanoid robots face a major challenge in real-world environments: They must go beyond simple walking and handle complex interactions such as climbing obstacles, vaulting, stair traversal,”
https://x.com/IlirAliu_/status/2068243203157803015
[2606.25076] Machine learning is revolutionizing weather forecasting — the next step is a change in how we work
https://arxiv.org/abs/2606.25076
First paper since joining @GoogleDeepmind! We present 🌍ATLAS (Active Theory Learning for Automated Science), a pipeline that generates interpretable mechanistic models from data and optimizes experiments to test them. Thread below”
https://x.com/EltetoNoemi/status/2067920123122336138
Prompt Injection as Role Confusion
https://role-confusion.github.io/
Patch the Planet is our effort to help open source maintainers move from security findings to merged fixes. We’re working with Trail of Bits, HackerOne, Calif, researchers, and maintainers to bring Codex Security and advanced models into the remediation process, with human”
https://x.com/OpenAI/status/2069104288417390688
In other news, we now have a pipeline for building diffusion draft models (i.e. DFLASH) which are significantly faster than EAGLE/etc. For gemma-4 31b and qwen3.6-27b we saw 30-50% real world decode performance gains. This run is for qwen3-32b (fp8). Seems pretty critical to”
https://x.com/jon_durbin/status/2069876870628155397
The rise of MoE models introduced new challenges in training, and @huggingface’s Transformers v5 brought first-class support for solving them. Now, NeMo AutoModel builds on top of v5. Part of the NeMo framework for building models at scale, NeMo AutoModel brings optimizations to”
https://x.com/NVIDIAAI/status/2069813582825418828
Krea 2 Technical Report – Krea
https://www.krea.ai/blog/krea-2-technical-report
What all is involved with onboarding a new model to noumena and ncode? I added GLM support over the weekend so i thought this might be interesting for some of you. first, you have to understand the architecture and how to properly serve it . luckily GLM5.2 is close enough to”
https://x.com/_xjdr/status/2069038936362803544
I’ll be at EUROPE EMBODIED this Friday in Munich. If you’re building in robotics or physical AI (or just trying to understand where Europe actually stands in this race), this is the room to be in. The Ecosystem Summit brings together founders, researchers, investors, and”
https://x.com/IlirAliu_/status/2068966255441142148
Another new idea to push the state of AI architectures forward. Sakana released a model that effectively uses a mixture of models to get work done. You get a single API but then the work gets farmed out the model that best performs the task. “Fugu manages model selection,”
https://x.com/levie/status/2068917230570795178
And new data release: in partnership with @GSMA we publish Telco-Common-Corpus, the largest fully open corpus for the telecom sector, 10 billion tokens all under free license/public domain.”
https://x.com/Dorialexander/status/2070080144593588493
🎉 Meet LFM2.5-230M from @liquidai, their smallest model yet at 230M params, but it punches way above its weight. Day 0 Support is live on SGLang! Built on the LFM2 architecture for on-device deployment: > Blazing-fast inference, runs everywhere, from cloud GPUs to low-cost CPUs”
https://x.com/lmsysorg/status/2070168574849945721
GPT-5.6 spotted in Codex GitHub repo”
https://x.com/scaling01/status/2069442918889189588
Surprising lessons from my research scientist job search | Yong Zheng-Xin
https://yongzx.github.io/blog/2026/06/24/job-search/
Notes on the Industry Job Search
https://alisawuffles.github.io/blog/job-search/
Moebius Project Page
https://hustvl.github.io/Moebius/
Scaling Laws, Carefully | Lil’Log
https://lilianweng.github.io/posts/2026-06-24-scaling-laws/
DFD — Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation
https://csy2077.github.io/dfd-project-page/
I think this means you can collect ~10k hours just from open source datasets, which means basically anyone should be able to build a decent robot foundation model: 500 from the new BitRobot dataset 500 from galaxea
https://t.co/eeGgOSC2D1 3000 from agibot”
https://x.com/chris_j_paxton/status/2070009005439603083
This giant free dataset could make helper robots way smarter, way faster: An open-source robotics stack from Berkeley AI researchers featuring the largest teleoperation dataset released to date with over 3,500 hours of bimanual manipulation data across 200 tasks. The video”
https://x.com/IlirAliu_/status/2067879108009160862
🙏 Thanks to the @NVIDIAAI team for highlighting DFlash support on vLLM! With DFlash speculative decoding, swapping EAGLE-3 for a DFlash checkpoint is a config-only change — no code edits needed. It runs through the open-source Speculators library, which links the DFlash”
https://x.com/vllm_project/status/2069494027431649404
Optimizing Models to Be Fast at Codegen | Morph
https://www.morphllm.com/blog/codegen-inference-research
Speculation Is All You Need. In this blog post, we announce the co-release (w/ Z Lab) of six more state-of-the-art DFlash speculators for @Alibaba_Qwen 3.x. Over 1k output tps for 3.5 122B-A10B on a B200. Read the blog for why we’re all-in on spec dec.
https://x.com/charles_irl/status/2068124629433262210
Multi-Vector Embeddings are Provably More Expressive than Single Vector Embeddings @Raj_Jayaram_ proves that approximating multi-vector similarity with single vectors requires exponentially more dimensions. 📝”
https://x.com/_reachsumit/status/2069319141128024395
Classic short term thinking. What about the heat death of the universe? This paper shows we could survive 100 billion years past the end of all near stars, if we start building Dyson Spheres around millions of stars and start gathering them together soon.
https://x.com/emollick/status/2068904192840782165
Coherence Neuro begins first clinical study for BCI
https://www.massdevice.com/coherence-neuro-begins-first-clinical-study-bci/
In 1958 Ian Donald published what is now the foundational paper on medical ultrasound for obstetrics. He was so widely ridiculed by his colleagues at the time that they nicknamed him Mad Donald [1], and one said ultrasound would be useful only to “a gynecologist who was blind and”
https://x.com/russelljkaplan/status/2068412531283312918
Proto: A programming language for generative biology | Arc Institute
https://arcinstitute.org/news/proto
AI Engineer Claims to Have Cracked Linear A — AI Clambake
https://aiclambake.com/clamtakes/linear-a/
Durable Objects now supports apac-ne and apac-se location hints. Use these to fine-tune placement for users in Northeast or Southeast Asia and lower latency.”
https://x.com/CFchangelog/status/2067994912713322811
Everything You Need to Know about Knowledge Distillation”
https://x.com/TheTuringPost/status/2068474648925216861
Help shape how the world understands AI. We’re hiring two designers at Epoch AI to turn complex research into dashboards and visualizations researchers and policymakers can easily use.”
https://x.com/EpochAIResearch/status/2067622370635129048
Hi everyone, Our June 2026 Crawl Archive and corresponding Web Graph are now available. The June 2026 crawl consists of 2.10 billion web pages (or 354 TiB of uncompressed content). Captures are from 40.8 million hosts or 33.6 million registered domains. The corresponding Web”
https://x.com/CommonCrawl/status/2070094659343237492
Memory is a huge challenge for frontier AI systems, and true personalization remains out of reach until progress is made on it. Excited to see close friends and former lab mates go after these hard problems with new ideas @EyubogluSabri @MayeeChen @dan_biderman @scott_linderman”
https://x.com/krandiash/status/2069473168822292644
Model Size Scaling in 2023-2031 — LessWrong
https://www.lesswrong.com/posts/yLHiQGCPdvzL9fBn3/model-size-scaling-in-2023-2031
Use Case 3: One-Shot Blindfold Chess Can an AI hold an entire game state in memory without drifting? To test Fugu Ultra’s persona stability and sustained memory, we had it play 4 back-to-back games of blindfold chess. Every model played the same way: no board shown, requiring”
https://x.com/SakanaAILabs/status/2069088009790861312





Leave a Reply