Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: Using the provided reference image, preserve every detail — the marigold orange backdrop, the seated young woman in her purple-and-white windbreaker with closed-eyes contented smile, the tattooed man in the red beanie and layered red vest leaning in mid-serenade, the cinematic studio lighting and shallow depth of field — but replace ONLY the black handheld microphone with a vertically-held green printed circuit board dense with chips, capacitors, and gold traces, gripped exactly like a mic at the same scale and position to his mouth, photographed with seamless realism and matching warm light. After generating the image, overlay the text “Tech” in the upper-left corner of the frame in large, bold, all-caps ITC Avant Garde Gothic Pro Medium (or a near-identical geometric sans-serif if unavailable), pure white (#FFFFFF), with no date, subtitle, drop shadow, or outline. The text should be substantial in scale — taking up a meaningful portion of the upper-left area — with comfortable margin from the top and left edges, set against the negative space of the orange backdrop so it does not overlap or obscure the singer, the seated woman, or the replaced object.
give your agents a browser. Browser Run (fka Browser Rendering) really sprinted for Agents Week 🏃♀️ quick look at what shipped 1) Browser Rendering –> Browser Run (renamed!) 2) Live View – realtime view of browser sessions 3) Human in the Loop – intervene when your agent needs
https://x.com/kathyyliao/status/2044479579382026484
So the concern over Mythos and cybersecurity seems warranted.
https://x.com/emollick/status/2043810051979157680
This was… an interesting one. Reminder that we run independent evals on our cyber ranges that labs don’t have access to. Exploitation capabilities are getting seriously good. Mythos is the first model to complete our full 32-step corporate network attack sim E2E.
https://x.com/ekinomicss/status/2043688793085992970
Anthropic launched Claude Opus 4.7 today, the new #1 in our GDPval-AA benchmark for performance on agentic real-world work tasks Opus 4.7 scored 1753 on GDPval-AA at launch with its ‘max’ effort setting, surpassing GPT-5.4 xhigh. This is a significant upgrade, placing Opus back
https://x.com/ArtificialAnlys/status/2044856740970402115
Anthropic says Opus 4.7 hits 80.6% on Document Reasoning — up from 57.1%. But “”reasoning about documents”” ≠ “”parsing documents for agents.”” We ran it on ParseBench. → Charts: 13.5% → 55.8% (+42.3) — huge → Formatting: 64.2% → 69.4% (+5.2) → Content: 89.7% → 90.3%
https://x.com/llama_index/status/2044886527352647859
Anthropic’s Opus 4.7 just seized the #1 spot on the Vals Index with a score of 71.4%, a massive jump from the previous best (67.7%). It also ranks #1 on Vibe Code Bench, Vals Multimodal, Finance Agent, Mortgage Tax, SAGE, SWE-Bench, and Terminal Bench 2.
https://x.com/ValsAI/status/2044792518953533777
big jump in coding capabilities by Claude 4.7 Opus SWE-Bench Pro 64.3% SWE-Bench Verified 87.6% TerminalBench 69.4% but interestingly, I think they kept CyberGym scores artificially low
https://x.com/scaling01/status/2044784563201708379
Claude 4.7 Opus has an Elo of 1753 on GDPVal-AA
https://x.com/scaling01/status/2044784781368365233
Claude Opus 4.7 is out! Benchmark scores look pretty strong, but clearly much worse than Mythos. It’s a nerfed Mythos, they deliberately reduced cyber capabilities during training.
https://x.com/Yuchenj_UW/status/2044787564440334350
Document Arena update: four new models are reshaping the top ranks – including two open models! – #1 Claude Opus 4.6 Thinking is new, keeping @AnthropicAI in the top 3 – #8 Kimi-K2.5 Thinking by @Kimi_Moonshot now the best open model (Modified MIT) – #10 Gemma-4-31b by
https://x.com/arena/status/2044437193205395458
Document reasoning increased by A LOT for Opus 4.7
https://x.com/scaling01/status/2044784878965703100
Introducing Claude Opus 4.7 \ Anthropic
https://www.anthropic.com/news/claude-opus-4-7
New Anthropic Fellows research: developing an Automated Alignment Researcher. We ran an experiment to learn whether Claude Opus 4.6 could accelerate research on a key alignment problem: using a weak AI model to supervise the training of a stronger one.
https://x.com/AnthropicAI/status/2044138481790648323
Nontheless Opus 4.7 scores much higher on Firefox shell exploitation
https://x.com/scaling01/status/2044788243435069764
OpenAI just dropped a major Codex update, one hour after Anthropic’s Opus 4.7. Whats new: background computer use on macOS (Codex clicks and types on your Mac while you keep working), in-app browser, image generation via gpt-image-1.5, persistent memory, long-running
https://x.com/kimmonismus/status/2044832303075995994
Opus 4.7 first-hour impressions Ran the canvas tree growth test twice. 4.6: nailed the animation both times 4.7: static tree, no growth animation — twice 4.7’s thinking is noticeably shorter and faster though (trimmed some 4.6 thinking in the clip for pacing). Not the upgrade
https://x.com/stevibe/status/2044800069661254064
Opus 4.7 scores 92% on ARC-AGI-1 and 75.83% on ARC-AGI-2
https://x.com/scaling01/status/2044791039605506344
The new Opus 4.7 model places #1 on our Vibe Code Benchmark, at 71%. When we first released the benchmark 4.5 months ago, no model scored above 25%. This benchmark tests a model’s ability to create a fully functional web application from the ground up.
https://x.com/ValsAI/status/2044791415524471099
We comprehensively benchmarked Opus 4.7 on document understanding. We evaluated it through ParseBench – our comprehensive OCR benchmark for enterprise documents where we evaluate tables, text, charts, and visual grounding. The results 🧑🔬: – Opus 4.7 is a general improvement
https://x.com/jerryjliu0/status/2044902620746363016
What are the largest software engineering tasks AI can perform? In our new benchmark, MirrorCode, Claude Opus 4.6 reimplemented a 16,000-line bioinformatics toolkit — a task we believe would take a human engineer weeks. Co-developed with @METR_Evals. Details in thread.
https://x.com/EpochAIResearch/status/2042624189421752346
What you need to know about Opus 4.7 * Takes instructions literally * Better vision means improved computer use and producing slides and other visual artifacts * Optimized for large-scale real-world analysis * Better at using file system-based memory
https://x.com/omarsar0/status/2044797480471044536
Wow I can already say after just 5 hours using @AnthropicAI Opus 4.7 that this is the first model that “”gets”” what I’m doing when I’m working. It feels aligned with me in a way no previous model did. (4.6 actively worked against me. I hated it. So this is *very* exciting!)
https://x.com/jeremyphoward/status/2044942799511191559
Anthropic co-founder confirms the company briefed the Trump administration on Mythos | TechCrunch
Anthropic co-founder confirms the company briefed the Trump administration on Mythos
First model from Anthropic, which openly acknowledges it isn’t the best model they have
https://x.com/nrehiew_/status/2044791293080121553
Internal Anthropic survey on Claude Mythos Preview 12/18 people thought that Mythos can manage day long ambiguous tasks 8/18 thought that it can execute week long tasks
https://x.com/scaling01/status/2044787521691742338
Nearly 1/3 of surveyed people in Anthropic now think entry-level engineers and researchers are likely replaced by Mythos within 3 months
https://x.com/arankomatsuzaki/status/2044808883928186936
Read OpenAI’s latest internal memo about beating the competition — including Anthropic | The Verge
https://www.theverge.com/ai-artificial-intelligence/911118/openai-memo-cro-ai-competition-anthropic
Five hyperscalers now own over two-thirds of global AI compute
https://epochai.substack.com/p/five-hyperscalers-now-own-over-two
Today, we’re introducing Skills in @GoogleChrome, a new way to build one-click workflows for your most frequently used AI prompts — like asking for ingredient substitutions to make a recipe vegan, generating side-by-side shopping comparisons across multiple tabs, or scanning long
https://x.com/Google/status/2044106378655215625
Turn your best AI prompts into one-click tools in Chrome
https://blog.google/products-and-platforms/products/chrome/skills-in-chrome/
Gemini Robotics ER 1.6: Enhanced Embodied Reasoning — Google DeepMind
https://deepmind.google/blog/gemini-robotics-er-1-6/
Instead of writing complex code, the team interacted with Spot using plain English. We built a bridge between Gemini Robotics ER and Spot’s system, giving the AI a basic set of tools to move freely, take photos, and grab things – enabling it to carry out more complex tasks.
https://x.com/GoogleDeepMind/status/2044763631858909269
Introducing Gemini Robotics ER 1.6, our new SOTA robotics model 🤖 which excels at visual and spacial reasoning, now available via the Gemini API!
https://x.com/OfficialLoganK/status/2044080025474126065
Robotics is making progress! 🤖 We just released @GoogleDeepMind Gemini Robotics-ER 1.6 for enhanced embodied reasoning. – Unlocks instrument reading capabilities for complex gauges and sight glasses. – Achieves 93% success on instrument reading tasks using agentic vision. –
https://x.com/_philschmid/status/2044071114578509971
We teamed up with @BostonDynamics to power their robot Spot with Gemini Robotics embodied reasoning models. This means it can better understand its surroundings, identify objects and follow simple commands – like tidying up a room.
https://x.com/GoogleDeepMind/status/2044763625680765408
We’re rolling out an upgrade designed to help robots reason about the physical world. 🤖 Gemini Robotics-ER 1.6 has significantly better visual and spatial understanding in order to plan and complete more useful tasks. Here’s why this is important 🧵
https://x.com/GoogleDeepMind/status/2044069878781390929
Sub-32B open weights models now offer GPT-5 level intelligence with Qwen3.5 27B (Reasoning) matching GPT-5 (medium) at 42 and Gemma 4 31B (Reasoning) matching GPT-5 (low) at 39 on the Artificial Analysis Intelligence Index @Alibaba_Qwen’s Qwen3.5 and @GoogleDeepMind’s Gemma 4
https://x.com/ArtificialAnlys/status/2043929874537296026
Banger paper from NVIDIA. Agentic reasoning needs models that are not just capable, but efficient at long-context inference. The agent model layer is moving toward open, long-context, high-throughput architectures. This paper introduces Nemotron 3 Super, an open 120B parameter
https://x.com/dair_ai/status/2044452957023047943
NVIDIA Launches Ising, the World’s First Open AI Models to Accelerate the Path to Useful Quantum Computers | NVIDIA Newsroom
https://nvidianews.nvidia.com/news/nvidia-launches-ising-the-worlds-first-open-ai-models-to-accelerate-the-path-to-useful-quantum-computers
We’ve been developing a multi-agent system that builds and maintains complex software autonomously. Recently, we partnered with NVIDIA to apply it to optimizing CUDA kernels. In 3 weeks, it delivered a 38% geomean speedup across 235 problems.
https://x.com/cursor_ai/status/2044136953239740909
Today, we released Lyra 2.0, a framework for generating persistent, explorable 3D worlds at scale, from NVIDIA Research. Generating large-scale, complex environments is difficult for AI models. Current models often “forget” what spaces look like and lose track of movement over
https://x.com/NVIDIAAIDev/status/2044445645109436672
⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agentic coding on par with models 10x its active size 📷 Strong multimodal perception and reasoning ability 🧠 Multimodal thinking + non-thinking modes
https://x.com/Alibaba_Qwen/status/2044768734234243427
LM Performance:Qwen3.6-35B-A3B outperforms the dense 27B-param Qwen3.5-27B on several key coding benchmarks and dramatically surpasses its direct predecessor Qwen3.5-35B-A3B, especially on agentic coding and reasoning tasks.
https://x.com/Alibaba_Qwen/status/2044768738294268199
VLM Performance:Qwen3.6 is natively multimodal, and Qwen3.6-35B-A3B showcases perception and multimodal reasoning capabilities that far exceed what its size would suggest, with only around 3 billion activated parameters. Across most vision-language benchmarks, its performance
https://x.com/Alibaba_Qwen/status/2044768742761189762
Alibaba released Qwen3.6-35B-A3B today. Big jump compared to Qwen 3.5-35B model. It’s a sparse MoE, 35B total params, only 3B active. Natively multimodal, thinking and non-thinking modes. Hardfacts: SWE-bench Verified: 73.4, near dense Qwen3.5-27B (75.0), way ahead of
https://x.com/kimmonismus/status/2044780695361290347
All is not lost. Duckerton is still possible. Here is Seedance 2.0 with the same prompt.
https://x.com/emollick/status/2042455596834660479
Agent evals are drifting away from production reality. Most benchmarks use clean tasks, well-specified requirements, deterministic metrics, and retrospective curation. Production work is messier, with implicit constraints, fragmented multimodal inputs, undeclared domain
https://x.com/dair_ai/status/2044773323914322393
Doing my “”large codebase modernization”” bench. Cooked for 32 minutes. Looking reasonable so far but it missed the changes to the Link component in Next.js (almost everything has missed this to be fair)
https://x.com/theo/status/2044907295205961806
Introducing FrontierSWE, an ultra-long horizon coding benchmark. We test agents on some of the hardest technical tasks like optimizing a video rendering library or training a model to predict the quantum properties of molecules. Despite having 20 hours, they rarely succeed
https://x.com/MatternJustus/status/2044876224896565679
Scaling to ultra-long horizon agents requires novel benchmarks and RL environments. FrontierSWE by @ProximalHQ is exactly that: 11h average runtime, open-ended tasks like end-to-end model optimization, and frontier agents fail almost all of them. We co-designed granite_inf,
https://x.com/vincentweisser/status/2044923733048222197
Turns out we can get SOTA on agentic benchmarks with a simple test-time method! Excited to introduce LLM-as-a-Verifier. Test-time scaling is effective, but picking the “”winner”” among many candidates is the bottleneck. We introduce a way to extract a cleaner signal from the
https://x.com/Azaliamirh/status/2043813128690192893
we just shipped Kernels, it’s a new repo at @huggingface 💚 it allows for packaging and distribution of optimized kernels 🔥 vibe-optimize Kernels, benchmark gains and share them on Hub 🫵
https://x.com/mervenoyann/status/2044080953648128073
We partnered with @ProximalHQ to run five frontier coding agents on a hard task: rebuild the full Wan 2.1 text-to-video pipeline on MAX (no PyTorch, no diffusers) in 20 hours as part of their new Frontier-SWE benchmark. Two nearly pulled it off. Every model understood the
https://x.com/Modular/status/2044879525881024968
Current frontier models are increasingly saturating common AI benchmarks. Are they still useful? We think benchmarks remain important, but they can both over- and understate AI capabilities. To better survey this space, the field is turning to a new paradigm: open-world evals.
https://x.com/steverab/status/2044852672562426216
I’m pleased to share that our search team has open sourced an embedding model called Harrier that is currently ranking #1 on the multilingual MTEB-v2 benchmark leaderboard. Harrier delivers SOTA performance on retrieval quality, semantic matching, and contextual analysis across
https://x.com/JordiRib1/status/2041550352739164404
Inside VAKRA: Reasoning, Tool Use, and Failure Modes of Agents
https://huggingface.co/blog/ibm-research/vakra-benchmark-analysis
Our latest Live model is # 1 on Tau Voice Bench! Excited to see this new frontier of voice models cross the chasm of usability in production.
https://x.com/OfficialLoganK/status/2042672082425712935
significant improvement on coding and agentic benchmarks. better at computer vision and a new xhigh mode
https://x.com/dejavucoder/status/2044786310746186094
We’re open sourcing the first document OCR benchmark for the agentic era, ParseBench. Document parsing is the foundation of every AI agent that works with real-world files. ParseBench is a benchmark that measures parsing quality specifically for agent knowledge work: ✅ It
https://x.com/jerryjliu0/status/2043721536922955918
[2604.08407] Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain
https://arxiv.org/abs/2604.08407
// Artifacts as Memory Beyond the Agent Boundary // An agent doesn’t always need a bigger memory buffer. Sometimes the environment itself remembers on the agent’s behalf. New research formalizes this intuition mathematically for the first time. The work introduces a formal
https://x.com/dair_ai/status/2044066936045351317
// Multi-User LLM Agents // Every agent framework assumes one user giving instructions. But deploy an agent into a team workflow, and suddenly it has multiple bosses with conflicting goals, private information, and different authority levels. This work formalizes multi-user
https://x.com/omarsar0/status/2044067923787165799
🚀 deepagents 0.5 release 👉 Async subagents – kick off background tasks on any Agent Protocol backed server while you continue to interact with the main agent. Start multiple background tasks in parallel, keep the conversation going, and collect results as they come in. Tasks
https://x.com/LangChain/status/2044086454230626733
3 months ago I started building a coding agent that runs in the cloud. It’s since written every line of code I’ve shipped, including itself. Today, I’m open sourcing it. Introducing Open Agents.
https://x.com/nicoalbanese10/status/2043745569278251112
Agent Lee is an in-dashboard agent that shifts Cloudflare’s interface from manual tab-switching to a single prompt. Using sandboxed TypeScript, it helps you troubleshoot and manage your stack as a grounded technical collaborator.
https://x.com/Cloudflare/status/2044406215208316985
As AI agents accelerate coding, what is the future of software engineering? Some trends are clear, such as the Product Management Bottleneck, referring to the idea that we are more constrained by deciding what to build rather than the actual building. But many implications, like
https://x.com/AndrewYNg/status/2043742105852621052
copilot –remote Take your coding agent session with you anywhere!
https://x.com/pierceboggan/status/2043717775265562701
hermes-lcm v0.2.0 is out! Lossless context management for Hermes Agent — every message persisted, hierarchical DAG summaries, agent tools to drill back into anything that was compacted. No more lossy flat summaries. What’s new since launch: – 6 agent tools (grep, describe,
https://x.com/SteveSchoettler/status/2043870709613768820
Humwork A2P marketplace connects AI agents with experts
https://www.testingcatalog.com/humwork-a2p-marketplace-connects-ai-agents-with-experts/
I am more and more convinced that this is the future of software development UI. @cursor_ai is the closest in my opinion A list of work you’re working on parallel, the agent in the middle, and most importantly, the thing you’re building on the right. Because you want to see what
https://x.com/kieranklaassen/status/2044108436087157220
I’m noticing some really big shifts in how AI models starts to handle memory. @ECNUER and others introduced Memory Intelligence Agent (MIA) that highlights the importance of storing the whole problem-solving journey – how to perform tasks. It turns memory into something closer
https://x.com/TheTuringPost/status/2042386614568325404
les fucking go… for agents to push kernels to the hub, do: > pip install kernels > kernels skills add > <start agent> > “”write an RMSNorm kernel for h100 and push to Hugging Face Hub”” bam, you are a kernel author!
https://x.com/ben_burtenshaw/status/2044114277745807684
Long-horizon AI research agents are mostly a state-management problem. It is not enough for an agent to reason well in the next turn. ML research requires task setup, implementation, experiments, debugging, and evidence tracking over hours or days. This new paper introduces
https://x.com/omarsar0/status/2044436099121209546
Most AI assistants wait for you to ask. But a truly useful agent should notice you need help before you say anything. New research takes a serious shot at building proactive agents that work in real time. The work introduces PASK with three components: IntentFlow for streaming
https://x.com/dair_ai/status/2044145437456904438
Redesigning the Service Role for the AI Agent Era
https://www.asapp.com/webinars/redesigning-the-service-role-for-the-ai-agent-era
Speeding up GPU kernels by 38% with a multi-agent system · Cursor
https://cursor.com/blog/multi-agent-kernels
The 80/20 of multi-agent teams for non-technical people: Stop making one AI agent do everything. Build a team of 4: 1. Orchestrator: plans the work, routes tasks, synthesizes results 2. Researcher: gathers sources, verifies claims, flags uncertainty 3. Writer: turns raw
https://x.com/coreyganim/status/2043627229205193211
The crazy part? This was done (nearly) fully autonomously! Only 8 prompts from the human in the loop. Just a Hermes agent, a skill, and a dream. 🐉 I told my AI agent “”use obliteratus to find the best way to get the guardrails off Gemma 4 E4B”” It loaded the OBLITERATUS skill
https://x.com/elder_plinius/status/2044462515443372276
Love seeing this open-sourced. Had a great chat with @nicoalbanese10 some weeks ago where he hinted to something like this. Great reference architecture for cloud coding agents. Open Agents gives you the full stack: UI, auth, workflows, sandbox. #DeepAgent from @LangChain takes
https://x.com/bromann/status/2043886229650067729
Marcus Hutchins, the guy famous for stopping the WannaCry Ransomware, probably has the best take on Mythos doing vulnerability research
https://x.com/ananayarora/status/2043381424594837789
The Mythos Threshold – Joe Reis
https://joereis.substack.com/p/the-mythos-threshold
What I learned this week – Pretraining parallelisms, Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RL, On why pretraining runs fails
https://www.dwarkesh.com/p/what-i-learned-april-15
2 prompts deep into Opus 4.7 and benchmarks don’t do it justice. Way better behavior and instruction following. Pretty massive improvement in actual usage.
https://x.com/mweinbach/status/2044801022439137566
3. Tell the model how to verify its changes. Put your testing workflow in your claude.md, or add a /verify-app skill. Opus 4.7 is better at verifying it’s work, and it’s helpful to share any local dev tips that are hard to discover.
https://x.com/_catwu/status/2044808538351100377
after ~10 million tokens Mythos is much more efficient than other models it reaches the same performance as Opus with ~40% the tokens
https://x.com/scaling01/status/2043700788245963167
Claude Opus 4.7 is now available as an Agent Preview inside of Devin! Anthropic has clearly optimized Claude Opus 4.7 for long-horizon autonomy, unlocking a class of deep investigation work we couldn’t reliably run before. Claude Opus 4.7 model costs within Devin will be
https://x.com/cognition/status/2044844661076902082
Claude Opus 4.7 is now available in Cursor. We’ve found it to be impressively autonomous and more creative in its reasoning. We’re launching it with 50% off for a limited time. Enjoy!
https://x.com/cursor_ai/status/2044785960899236341
Claude Opus 4.7 is out! Handles ambiguous, multi-step work even better than 4.6. Cursor’s internal bench cleared 70%, up from 58% on 4.6. Notion saw a 14% lift on their evals with a third of the tool errors 🔨
https://x.com/mikeyk/status/2044802045186846912
Claude Opus 4.7 is out. the TL;DR Anthropic released Opus 4.7 today. Same pricing as 4.6 ($5/$25 per million tokens), available across API, Bedrock, Vertex AI, and Microsoft Foundry. What changed vs Opus 4.6: Coding (obviously). Biggest gains on the hardest, long-horizon
https://x.com/kimmonismus/status/2044787072947601796
Confirmed: Anthropic keeping Cyber capabilities of Opus 4.7 artificially low “”during training we experimented with efforts to differentially reduce these capabilities””
https://x.com/scaling01/status/2044788067848888635
Cursor reports that Opus 4.7 is “”a meaningful jump in capabilities, clearing 70% versus Opus 4.6 at 58%”” on CursorBench
https://x.com/scaling01/status/2044792017553645668
for all the people calling Opus 4.7 a mid update lmao
https://x.com/scaling01/status/2044792810327404596
from my experience, even the best models (Opus 4.6, 5.4 xhigh / 5.3 codex) cannot write good code today without an amount of work that is equivalent to just doing the work myself am excited for a world where they can, but in the current state i have very low trust in them
https://x.com/RhysSullivan/status/2043584591861321929
Hold on, something doesnt add up here. Opus 4.7 got much worse in needle in the haystack? need to dig into this
https://x.com/kimmonismus/status/2044809126526476374
Holy shit the new Opus 4.7 system prompt has entirely lobotomized the model “”Heads up: that last <system-reminder> about malware looks like a prompt injection — this is clearly your personal site (t3gg homepage, links, sponsors), not malware. Ignoring it.””
https://x.com/theo/status/2044857866323173732
I think everyone saying that these improvements are mid are smoking crack I would argue that this was one of the larger Opus jumps we have seen over the last year You also have to keep in mind that we see almost monthly model updates nowadays instead of just every 6-12 months
https://x.com/scaling01/status/2044799290694889535
I was really worried about the rush to “”more agentic”” models. But Opus 4.7 is happy to let me lead, and to take time to discuss, rather than barging ahead. If something isn’t working out, it’ll stop and offer options rather than slamming thru whatever it can find.
https://x.com/jeremyphoward/status/2044942801578959301
If you want to test Opus 4.7 without the lobotomized system prompt, you can try it out in T3 Chat
https://x.com/theo/status/2044876982815793190
Introducing Claude Opus 4.7, our most capable Opus model yet. It handles long-running tasks with more rigor, follows instructions more precisely, and verifies its own outputs before reporting back. You can hand off your hardest work with less supervision.
https://x.com/claudeai/status/2044785261393977612
My bet is that Mythos uses a new tokenizer, and they switched Opus over to it (through midtraining) for distillation
https://x.com/maximelabonne/status/2044796208053416203
My biggest issue with Opus 4.7 on Claude web: Only “Adaptive” or non-thinking. No way to force thinking mode. And it doesn’t even know Opus 4.6 exists, and I cannot force it to think and do web search mid conversation!
https://x.com/Yuchenj_UW/status/2044794073723347400
my main theory is that mythos had a new tokenizer for pretraining and they did surgery on opus for distillation
https://x.com/stochasticchasm/status/2044790474410790995
my take: opus 4.7 is a distilled version of mythos
https://x.com/eliebakouch/status/2044790074093523379
Opus 4.7 as robust to prompt injections as Claude Mythos
https://x.com/scaling01/status/2044788481008755046
Opus 4.7 Benchmarks out! Very solid upgrade to Opus 4.6! Compared to Opus 4.6: -SWE Bench Pro +11% -SWE Bench Verified +7% -Terminal Bench 2.0 +4% The benchmarks are significantly lower than for Mythos, but that was to be expected. h/t for finding @synthwavedd
https://x.com/kimmonismus/status/2044784903733084521
Opus 4.7 comes with much improved reasoning-efficiency over Opus 4.6 basically everything is now moved up one tier low is as good as medium medium as good as high high as good as max
https://x.com/scaling01/status/2044785467942453698
Opus 4.7 deleting all long-context gains from Opus 4.6 lol
https://x.com/scaling01/status/2044791314898723179
Opus 4.7 has a new tokenizer. This means it’s also a new base model. Glory days of pretraining still very much going.
https://x.com/natolambert/status/2044788470179332533
opus 4.7 is here on claude platform / app
https://x.com/dejavucoder/status/2044784097378316327
Opus 4.7 is live in Claude Code today! The model performs best if you treat it like an engineer you’re delegating to, not a pair programmer you’re guiding line by line. Here are three workflow shifts we recommend for this model 🧵
https://x.com/_catwu/status/2044808533905178822
Opus 4.7 is now available in @MagicPathAI. From our early testing, the model is really strong at long tasks when design requires lots of changes, image-to-code, and overall produces cleaner, more reusable React components.
https://x.com/skirano/status/2044804877696516442
Opus 4.7 is WORSE than 4.6 on Long Context?
https://x.com/nrehiew_/status/2044795171213291614
Opus 4.7 much less likely to sudo rm -rf (taking destructive actions in production envs)
https://x.com/scaling01/status/2044789371837001779
Opus 4.7 uses a different tokenizer from Opus 4.6. So either: – Anthropic has a way to change tokenizer between finetunes – It is just new special tokens which implies they uses special tokens liberally within messages and not just as part of the chat template
https://x.com/nrehiew_/status/2044792314825228690
Opus 4.7 uses more thinking tokens, so we’ve increased rate limits for all subscribers to make up for it. Enjoy!
https://x.com/bcherny/status/2044839936235553167
Opus is going to be a bioweapon risk at this pace
https://x.com/scaling01/status/2044785139905913077
Some of my favorite things in Opus 4.7: – Very good at async work and following instructions – Effort levels are far more predictable for token control (+ new xhigh level) – No more downscaling of high-res images – Noticeably more taste in UIs, slides, docs
https://x.com/alexalbert__/status/2044788914813292583
Unfortunately they didn’t include a chart for GraphWalks scores: Opus 4.6 – 38.7% Opus 4.7 – 58.6% This would make clearer that long-context didn’t suffer as much as MRCR suggests.
https://x.com/scaling01/status/2044823423013020088
wait why is there an INSANE gap on long context benchmarks between opus 4.6 and 4.7??? this is crazy
https://x.com/eliebakouch/status/2044798168211100096
We’ve set the default effort level for Opus 4.7 to xhigh in Claude Code. You can use /effort to adjust this. Excited for you to try Claude Code with Opus 4.7 and let us know your feedback!
https://x.com/_catwu/status/2044808539663978970
Shocking result on my pelican benchmark this morning, I got a better pelican from a 21GB local Qwen3.6-35B-A3B running on my laptop than I did from the new Opus 4.7! Qwen on the left, Opus on the right
https://x.com/simonw/status/2044830134885306701
@stochasticchasm yeah they tend to forget that releases are now monthly and now bi-anually
https://x.com/scaling01/status/2044795960224592329
Anthropic Changes Pricing to Bill Firms Based on AI Use as Demand Jumps — The Information
https://www.theinformation.com/articles/anthropic-changes-pricing-bill-firms-based-ai-use-amid-compute-crunch
Anthropic introduced xhigh reasoning effort
https://x.com/scaling01/status/2044785557058814059
Anthropic loses Claude Code trust in black-box fight
https://www.implicator.ai/claude-probably-wasnt-secretly-nerfed-anthropic-made-the-black-box-too-dark/
Anthropic tests Claude Code upgrade to rival Codex Superapp
https://www.testingcatalog.com/anthropic-tests-claude-code-upgrade-to-rival-codex-superapp/
anthropic? you mean the greedy token guzzler company?
https://x.com/dejavucoder/status/2044798065530528061
every engineer at anthropic has been using mythos for ~1.5 months. meanwhile, their uptime is horrendous, claude code still has rendering bugs, etc. one could conclude that it won’t be the end of software engineering.
https://x.com/benhylak/status/2042051048261722467
GitHub reports similar improvements
https://x.com/scaling01/status/2044792459125834029
OpenAI has released a plugin that lets you call Codex directly within Anthropic’s Claude Code environment It turns Claude Code into a multi-agent setup with Codex as a specialized coding assistant This gives you: – High-quality code reviews – Delegation of real tasks
https://x.com/TheTuringPost/status/2044561927905677558
So we now have a pretty good picture of the state of the frontier AI model makers. US closed source models continue to lead. Google, OpenAI, and Anthropic stand well ahead of the pack, and may have signs of recursive self-improvement. xAI has fallen from frontier status for now
https://x.com/emollick/status/2042088011748290750
The pace at which Anthropic is shipping Opus variants is a very new thing in the industry.
https://x.com/_arohan_/status/2044791678180167804
The pace at which useful things are shipping also seems to be accelerating. Model releases are coming faster, of course, but so are significant application and enterprise products (especially from Anthropic). Almost certainly faster than the market can track or absorb information
https://x.com/emollick/status/2042434850003534077
we were literally stuck at 80% SWE-Bench Verified for months and just jumped to almost 90% and you guys call it mid …
https://x.com/scaling01/status/2044790717722034511
Yeah folks, it’s gonna be harder in the future to ensure OpenClaw still works with Anthropic models.
https://x.com/steipete/status/2042615534567457102
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts – Apple Machine Learning Research
https://machinelearning.apple.com/research/cram-less
Masked Depth Modeling for Spatial Perception”” TL;DR: treats missing depth as a learning signal to reconstruct accurate 3D geometry from noisy RGB-D inputs, enabling robust perception in real-world conditions
https://x.com/Almorgand/status/2042639933194575985
(14) ARC-AGI-3 – YouTube
@NBCNews on our recent AI usage survey:
https://x.com/EpochAIResearch/status/2044208011024142594
Buckle up everyone, your API costs are going up, not down.
https://x.com/madiator/status/2044801082359210215
New Eval mode: Battles in Direct We sample two random anonymous models during Direct chats – enabling pairwise comparison beyond turn 1. Why this matters: • Evaluates under longer context + multi-turn dependency • Captures failure modes: drift, consistency, recovery • Closer
https://x.com/arena/status/2044096836114493609
New models, new prompts. Perhaps the most valuable reason to be using GEPA. If you’ve got GEPA set up, migrating prompts takes a couple clicks. If you don’t? Get ready for some tedious prompt engineering over the next week or two.
https://x.com/dbreunig/status/2044794013375770915
Today we’re releasing SWE-check, a specialized bug detection model we RL-trained with @appliedcompute that matches frontier performance on internal in-distribution evals and makes meaningful progress on out-of-distribution evals, all while running 10x faster.
https://x.com/cognition/status/2044174496312242544
We are excited to host @ProximalHQ’s FrontierSWE on the Environments Hub as a launch partner. As an ultra-long horizon coding evaluation, even today’s frontier models struggle to solve the tasks after running for hours.
https://x.com/PrimeIntellect/status/2044878952020554083
AI is changing our jobs: among people who use AI regularly at work, 27% say AI has replaced some of their tasks; 21% say it has enabled new tasks. This and other workplace usage findings from our new Epoch AI/Ipsos survey on AI usage in 🧵
https://x.com/EpochAIResearch/status/2042302337059078605
We estimate that Gemini 3.1 Pro with thinking level `high` has a 50%-time-horizon of around 6.4 hrs (95% CI of 4 hrs to 12 hrs) on our suite of software tasks.
https://x.com/METR_Evals/status/2044463380057194868
Memory Caching: RNNs with Growing Memory”” Google’s new paper proposes a simple way to give recurrent models a memory that grows with sequence length. So instead of forcing an RNN to compress the full past into 1 fixed hidden state, it caches memory checkpoints across
https://x.com/askalphaxiv/status/2043782770657219010
NEW Research from Google. Integration test failures are painful because the signal is buried in messy logs. Massive output, heterogeneous systems, low signal-to-noise ratio, and unclear root causes. This paper introduces Auto-Diagnose, an LLM-based tool deployed inside Google’s
https://x.com/omarsar0/status/2044769798845079665
The PR you would have opened yourself
https://huggingface.co/blog/transformers-to-mlx
It was a pleasure to sit down with @FidlerSanja, VP of AI Research at NVIDIA, leading company’s Spatial Intelligence Lab, who is actively building the next major frontier of AI – physical AI. During GTC, where her lab introduced AlpaDream, we discussed: • If Transformers are
https://x.com/TheTuringPost/status/2042512295742656776
Rethinking AI TCO: Why Cost per Token Is the Only Metric That Matters
https://blogs.nvidia.com/blog/lowest-token-cost-ai-factories/
Figure and Hark just took an entire data center of NVIDIA B200s – every rack in the building Figure will be using these to predict physics and Hark will train next generation multi-modal models
https://x.com/adcock_brett/status/2042675641037000868
What are world models actually? @FidlerSanja, VP of AI Research at NVIDIA, leading company’s Spatial Intelligence Lab, explains in our interview If you want to learn about the major next frontier in AI, watch the full conversation:
https://x.com/TheTuringPost/status/2043962055531868554
new open-source Bonsai models are out 🔥 > ternary weights in 8B (1.75 GB), 4B (0.86 GB), and 1.7B (0.37 GB) > comes in MLX, ONNX weights and WebGPU browser demo 😍 > a2.0 licensed 👏
https://x.com/mervenoyann/status/2044841709075411047
🎉 Congrats @Alibaba_Qwen on the first open-weight Qwen3.6! Stronger agentic coding and a new thinking preservation option to retain reasoning context across turns. Same architecture as Qwen3.5, so serving teams can upgrade in place. Day-0 support in vLLM v0.19+. Thinking, tool
https://x.com/vllm_project/status/2044787721538060784
Introducing Nucleus-Image: the first sparse Mixture-of-Experts diffusion model 17B parameters. Only 2B active. 10x more parameter-efficient than leading diffusion models. Toe-to-toe with GPT Image 1, Imagen 4, and Qwen-Image: from pure pre-training alone. No DPO. No RL. No
https://x.com/withnucleusai/status/2044412335473713284
Qwen/Qwen3-Coder-Next · Hugging Face
https://huggingface.co/Qwen/Qwen3-Coder-Next
We built FrogsGame as a new task for evaluating AI’s posttraining skills! It’s a tool-using RL environment built around a blind-start interaction loop. Frontier agents get a container with the Qwen3-8B tokenizer, board-generating scaffolding, and @tinkerapi for remote training
https://x.com/karinanguyen/status/2044885375085339023
2-bit Qwen3.6-35B-A3B did a complete repo bug hunt with evidence, repro, fixes, tests and a PR writeup. 🔥 Run it locally in Unsloth Studio with just 13GB RAM. 2-bit Qwen3.6 GGUF made 30+ tool calls, searched 20 sites and executed Python code. GitHub:
https://x.com/UnslothAI/status/2044858346948464743
Qwen3.6-35B-A3B can now be run locally!💜 The model is the strongest mid-sized LLM on nearly all benchmarks. Run on 23GB RAM via Unsloth Dynamic GGUFs. GGUFs to run:
https://t.co/VlyW8UwDjw Guide:
https://x.com/UnslothAI/status/2044786492451778988
[2604.09168] ELT: Elastic Looped Transformers for Visual Generation
https://arxiv.org/abs/2604.09168
[PATCH] Add git-request-pull-script, a short script that generates a summary of pending changes – Ryan Anderson
https://lore.kernel.org/git/20050726073036.GJ6098@mythryan2.michonline.com/
@kimmonismus We kept MRCR in the system card for scientific honesty, but we’ve actually been phasing it out slowly. Two reasons: (1) it’s built around stacking distractors to trick the model, which isn’t how people actually use long context, and (2) we care more about applied long-context
https://x.com/bcherny/status/2044826315849888207
📢 Super excited to announce Parcae! We’ve been thinking about scaling laws and the “”right”” way to get more FLOPs. Turns out layer looping – with the right parameterization – gives you a new axis to scale! Parcae matches Transformers 2x their size (w/ the same data), and
https://x.com/realDanFu/status/2044459930149941304
🔐 One deployment, isolated data per user. Add custom auth so every user gets their own scoped threads, runs, and conversation history — with per-user data isolation and role-based access using any auth provider. Full walkthrough:
https://t.co/zDsViKLL2I Docs:
https://x.com/LangChain/status/2044098386270310783
🤯Topology & UV — the NO.1 headaches in #3D GenAI. 🔥We just move closer to BOTH at the same time. Introducing SATO: Strips as Tokens, a new autoregressive model for topology & UV, has been conditionally accepted to #SIGGRAPH 2026. Will available at #Hyper3D More details👇
https://x.com/DeemosTech/status/2044067290908635418
A must-read survey on On-Policy Distillation (OPD) for LLMs Shows how distillation is moving from static teacher data to interactive learning, where models learn from their own mistakes. Covers: – Why off-policy distillation fails (the problem of exposure bias) – OPD approach
https://x.com/TheTuringPost/status/2042634001404629322
A New Era of Databases: Lakebase | Databricks Blog
https://www.databricks.com/blog/what-is-a-lakebase?scid=701Vp000004h4j2IAA&dclid=CJOd2qO_-pMDFTIQiAkdlXMFmw&gad_source=7&gad_campaignid=23676083666
A new form of machine – Neural Computers (NCs)! @AIatMeta and KAUST proposed the model that becomes the computer itself, unifying computation, memory, and I/O into a one learned system. ▪️ NCs “”run”” interfaces directly from data: • In video gen setup, the model predicts
https://x.com/TheTuringPost/status/2043481205631504450
A new programming model for durable execution – Vercel
https://vercel.com/blog/a-new-programming-model-for-durable-execution
AI Horseless Carriages | koomen.dev
https://koomen.dev/essays/horseless-carriages/
Another great AiE in the books, this time in 🇪🇺 Europe for Arize AI and @arizephoenix So great to see all the homies again and learn a few things myself in one of my favorite cities Great themes this year: – maturing harness engineering patterns – context engineering /
https://x.com/dat_attacked/status/2043647001749836253
At @github @moraes_c_ and team are seriously fixing OSS pain points 🔥🔥🔥 1. Disable pull requests in repositories 2. Limit pull request creation to collaborators 3. Repository member role labels in the pull request list 4. A dedicated way to flag low quality comments 5.
https://x.com/SamMorrowDrums/status/2044375099738825103
Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D”” TL;DR: lifts open-vocabulary 2D detections into consistent 3D boxes using a transformer and multi-view geometric fusion
https://x.com/Almorgand/status/2042255903001391138
Can we learn a generalizable system prompt only using two examples? How do we select such a small subset for prompt optimization? P1 proposes a simple approach which selects 2 examples from AIME 24, applies RL-based optimization (policy gradient) to learn a single system
https://x.com/WenSun1/status/2043755261954011484
CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation”” TL;DR: generates camera-controlled cinematic videos from scene images by injecting implicit 3D scene priors into diffusion models, achieving strong consistency under large viewpoint changes
https://x.com/Almorgand/status/2044095955398492580
Defeating Nondeterminism in LLM Inference – Thinking Machines Lab
https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/
Does more common (frequent) text help LLMs perform better? A new paper “”Adam’s Law”” explains: if 2 sentences mean the same thing, LLMs tend to “”favor”” the more common wording because they’ve likely seen it more during training. Researchers proposed 3 things: ▪️ Textual
https://x.com/TheTuringPost/status/2042934417208025268
Don’t remember when I was this excited about a clicky thing clicking things. The sky team cooked with this one, it feels so polished (wish it was 2x faster though 😭) Excerpt from today’s @thursdai_pod ep dropping soon
https://x.com/altryne/status/2044898285299929181
Exciting new finding: LLMs struggle to *discover* a latent planning strategy for a task that is trivial when taught step-by-step. Scaling helps surprisingly little: from an 8-layer model to GPT-5.4 buys only 4 extra steps. We argue this is good news, for CoT monitoring ⤵️
https://x.com/LauraRuis/status/2043715536186384775
Generate 3D-printable cubes/cuboids with ArUco or AprilTag fiducial markers on all 6 faces, then detect their 6-DOF pose from a camera. Aprilcube is a two-part pipeline: Generator -> Creates a multi-color 3MF file with markers on every face, ready for dual-color 3D printing
https://x.com/IlirAliu_/status/2042664687771205966
Great blogpost from @pcuenq on making a new skill + test harness to automate porting new models from Transformers to mlx-lm
https://x.com/awnihannun/status/2044804263621312993
HappyHorse is still in the stable. 🐴 No official website yet — anything you’ve seen out there isn’t us. We’re part of Alibaba’s ATH AI Innovation Unit, and when we’re ready, you’ll know. Stay tuned. #HappyHorse #Alibaba #ATH
https://x.com/HappyHorseATH/status/2042454844842336461?ref_src=twsrc%5Etfw%7Ctwcamp%5Etweetembed%7Ctwterm%5E2042454844842336461%7Ctwgr%5Ecaae56671b2e17af67dbccc2955ce64dbf5d08d0%7Ctwcon%5Es1_&ref_url=https%3A%2F%2Fwww.bloomberg.com%2Fnews%2Farticles%2F2026-04-10%2Fstealth-alibaba-video-ai-model-tops-global-ranking-on-debut
Harness Engineering Derived from what Models can’t do alone: It feels like a good time to step back and reshare some basic mental models for why harnesses exist in the first place – working backwards from The Model Why the Harness Exists: Harnesses exist to augment and shape
https://x.com/Vtrivedy10/status/2043870915059236966
Hi, we are releasing ColGrep 1.2.0 ColGrep now incorporate BM25 trigrams to further enhance our multi-vector models using hybrid search. Now, ColGrep print relative paths by default (fewer tokens per result) Exact same features as GREP Improved CUDA usage and installation
https://x.com/raphaelsrty/status/2043676936442875954
How Missions Work | Factory.ai
https://factory.ai/news/missions-architecture
i love that @theo has this disclaimer in the lawn repo i keep telling you we need to reinvent everything: PRs, github, code collaboration, etc. burn everything down man it’s not 2022 anymore
https://x.com/thekitze/status/2030222687084359871?s=46
I moved today, tried to setup internet, provider needed to send a tech, earliest appointment was Wednesday. Got a @Starlink Standard 4 delivered and online in less than 1 hour, truly sci-fi, ty to SpaceX team and Elon.
https://x.com/OfficialLoganK/status/2043133990568366435
I-DLM: Introspective Diffusion Language Models
https://introspective-diffusion.github.io/
Instead of the gold standard, we can imagine an inference standard of exchange, the FLOP. (As opposed to tokens, this accounts for AI ability) With some AI help, I figure $1 buys roughly 10^17 managed-LLM inference FLOPs. So that $4 coffee would cost half an exaFLOP, choom.
https://x.com/emollick/status/2044501483757179229
Introducing DDTree: accelerates speculative decoding by drafting a tree with one block diffusion pass, then verifying multiple likely continuations together. Paper:
https://t.co/cgYBw70O5i Project page:
https://t.co/ygFukxrZLB Code:
https://x.com/liranringel/status/2043813397972607477
Introduction to recursive-mode – recursive-mode
https://recursive-mode.dev/introduction
just went live on european TBPN! exclusive preview of the @aiDotEngineer Europe Build Day today
https://x.com/swyx/status/2041504079008919915
Kiro CLI 2.0: a new look and feel, headless CI/CD pipelines, and Windows support – Kiro
https://kiro.dev/blog/cli-2-0/
Messy Middle of Installation -> GPT Meets GPT
https://x.com/TheTuringPost/status/2043842151876907325
Neat experiment finds AI fact checks are rated as more helpful & less ideological than human ones “”LLM-generated Community Notes can achieve broader cross-ideological acceptance than human-written notes, receiving more positive ratings from raters across the political spectrum””
https://x.com/emollick/status/2042797671094665254
Neural Computers explained It looks like a computer running code, but it’s just a model rendering the simulation. No real computation What is actually happening here – and is this the future of computing or just a very convincing illusion?
https://x.com/TheTuringPost/status/2044170497944908054
Oh no, it “”modernized”” by using a Next.js version from over a year ago 🙃
https://x.com/theo/status/2044907768487047567
Oh yeah, there’s pull requests now – The GitHub Blog
One of the coolest upgrades of attention ↓ New Interleaved Head Attention (IHA) lets multiple heads mix together before applying attention. What changes? Regular attention gives many isolated views – IHA gives many different views that mix and collaborate. IHA creates
https://x.com/TheTuringPost/status/2043810950864920890
Parcae: Doing more with fewer parameters using stable looped models
https://www.together.ai/blog/parcae
PrismML — Introducing Ternary Bonsai: Top Intelligence at 1.58 Bits
https://prismml.com/news/ternary-bonsai
Research we co-authored on subliminal learning–how LLMs can pass on traits like preferences or misalignment through hidden signals in data–was published today in @Nature. Read the paper:
https://x.com/AnthropicAI/status/2044493337835802948
SciPredict: Can LLMs Predict the Outcomes of Research Experiments in Natural Sciences? “”SciPredict addresses two critical questions: (a) can LLMs predict the outcome of scientific experiments with sufficient accuracy? and (b) can such predictions be reliably used in the
https://x.com/iScienceLuvr/status/2043977751506428323
Self-reported side effects of semaglutide and tirzepatide in online communities | Nature Health
https://www.nature.com/articles/s44360-026-00108-y
Speculative decoding is one of the highest-leverage inference optimizations out there. Here are tactical steps on how to train EAGLE-3 heads to work in production.
https://x.com/baseten/status/2043762663235432855
State of AI 2025: 100T Token LLM Usage Study | OpenRouter
https://openrouter.ai/state-of-ai
Surprisingly, the transport layer! @cmpatino_ found that vLLM sends logprobs over the wire in JSON format which is sloooow. Switching to binary NumPy arrays gave a 1.4x speed-up out of the box. Neat 🙂
https://x.com/_lewtun/status/2043690765227102335
The growing KV-cache of attention is the key component for the long-context understanding of LLMs, but what holds back long-term memory modules (e.g., Titans)? What if we could have the compression power of Titans but with a growing memory similar to Transformers? Memory
https://x.com/behrouz_ali/status/2043743704335192095
Version numbers are not a very useful way to understand model ability gains at this stage. Unfortunately that means that if you aren’t following closely, you would expect that 5.4 is a small gain over 5, or 4.6 a small gain over 4. That just isn’t the case, though.
https://x.com/emollick/status/2044134715867750568
Vouch | Hacker News
https://news.ycombinator.com/item?id=46930961
We partnered with University of Chicago economist @SuproteemSarkar to study how more capable models have changed the way people use Cursor. Across 500 teams, we find that developers are tackling more ambitious work with AI, with a 68% increase in high-complexity tasks this year.
https://x.com/cursor_ai/status/2044841478913130930
We shipped a new repo type called “”kernel”” on the Hub. We want to democratize the whole ping-pong around packaging, distributing, and using custom kernels. This repo type is only available to a few community partners, @sgl_project being the first! Hop in 🧵for more details.
https://x.com/RisingSayak/status/2043984021521346575
We’ve been thinking a lot about scaling laws, wondering if there is a more effective way to scale FLOPs without increasing parameters. Turns out the answer is YES – by looping blocks of layers during training. We find that predictable scaling laws exist for layer looping,
https://x.com/hayden_prairie/status/2044453231913537927
What compression looks like on @vllm_project. Same Gemma 4 31B. Red Hat AI’s quantized version runs at nearly 2x tokens/sec, half the memory, 99%+ accuracy retained. Open source. Quantized with LLM Compressor. Links in comments. 🙏 @_soyr_ for the 2-minute demo.
https://x.com/RedHat_AI/status/2043709783102906489
What if you could get 1.3B Transformer quality from a 770M model? That’s not a compression result. It’s a different architecture. Parcae, from @realDanFu (Together AI’s VP of Kernels) and his lab at UCSD, passes activations through the same layers multiple times — stably, for
https://x.com/togethercompute/status/2044454051543453745
Why you should work much harder RIGHT NOW – Marginal REVOLUTION
https://marginalrevolution.com/marginalrevolution/2026/03/why-you-should-work-much-harder-right-now.html
Wish there was information about where this data came from, but this is a very significant change. Since AI use comes from experience, the persistent gender gap in AI use across every study of AI was something that a lot of scholars were concerned about.
https://x.com/emollick/status/2044486831883137460
Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction”” TL;DR: scalable test-time training with global memory modules enables accurate kilometer-scale 3D reconstruction from long RGB video sequences
https://x.com/Almorgand/status/2044468554754412564
Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment”” TL;DR: improves 3D reconstruction by aligning features across views using self-distilled geometry-aware representations
https://x.com/Almorgand/status/2042631239601930681
Spark 2.0 is here! 🚀 We’re redefining what’s possible on the web with a streamable LoD system for 3D Gaussian Splatting. Built on Three.js, you can now stream massive 100M+ splat worlds to any device from mobile to VR using WebGL2. All open-source. Dive into the tech 👇
https://x.com/sparkjsdev/status/2044090505982816449
Static 3D generation isn’t enough. We need assets ready for animation. Our new #SIGGRAPH work, AniGen, takes a single image and generates the 3D shape, skeleton, and skinning weights all at once. Code is fully open-sourced! Kudos to @KyrieIr31012755 and @VastAIResearch 🧵(1/4)
https://x.com/yanpei_cao/status/2044094818872377720
The future of sports is immersive. We can already reproduce entire games in 3d, track every player down to their pose and heartbeat – but barely any of that makes it to your living room. Here’s everything you need to known about 3d sports tech.
https://x.com/bilawalsidhu/status/2043085376349442077
Introducing Kernels on the Hugging Face Hub ✨ What if shipping a GPU kernel was as easy as pushing a model? – Pre-compiled for your exact GPU, PyTorch & OS – Multiple kernel versions coexist in one process – torch.compile compatible – 1.7x-2.5x speedups over PyTorch baselines
https://x.com/ClementDelangue/status/2044053580504584349
Before he wrote AI 2027, he predicted the world in 2026. How did he do?
https://asteriskmag.substack.com/p/before-he-wrote-ai-2027-he-predicted
Our paper landed in Nature Health today! Healthcare is one of the most high-stakes, high-potential applications of AI. So we set out to understand how people actually use it in our AI products today.
https://x.com/mustafasuleyman/status/2044817893460996487
🚀 Great to see vLLM powering OCR at this scale — Chandra-OCR-2 (5B) serving ~60 papers/hour per L40S across 16 parallel jobs. The full pipeline breakdown is a great read 👇 🔗
https://x.com/vllm_project/status/2043964594679636260
We’re open-sourcing webAI-ColVec1. #1 on ViDoRe V3. Two of the top three spots. For a long time, multimodal RAG looked like a scaling problem. It’s not. Frontier-level retrieval. No OCR. No preprocessing. Built for real documents.
https://x.com/thewebAI/status/2044435998508240926
Securing non-human identities: automated revocation, OAuth, and scoped permissions
https://blog.cloudflare.com/improved-developer-security/
[2604.13036] Lyra 2.0: Explorable Generative 3D Worlds
https://arxiv.org/abs/2604.13036





Leave a Reply