Image created with OpenAI GPT-Image-1. Image prompt: 1966 Kodachrome photo-look, thin white frame, forest-green title band in upper left with stacked yellow/white serif text reading “TECH”, band waving goodbye from behind the fence scene featuring a neon sign “TECH” reflected in a small puddle; gentle film grain, overcast daylight

How @Uber used LangGraph to build AI developer agents that generate thousands of daily code fixes and saved 21,000+ hours — serving an organization of 5,000 developers working with hundreds of millions of lines of code. Watch their full session here: https://x.com/LangChainAI/status/1932493346498543898

Apple doesn’t report benchmarks for their AIs, reporting on an ill-documented head-to-head evaluation But even by their standards, Apple’s latest on device models are mostly worse than the open Gemma 3-4B from Google or Qwen 3-4B And their server LLM is similar to Llama 4 Scout https://x.com/emollick/status/1932420903515590997

Qualcomm strengthens AI portfolio with $2.4 billion Alphawave deal | Reuters https://www.reuters.com/world/uk/qualcomm-acquire-uks-alphawave-24-billion-2025-06-09/

Our vision is for AI that uses world models to adapt in new and dynamic environments and efficiently learn new skills. We’re sharing V-JEPA 2, a new world model with state-of-the-art performance in visual understanding and prediction. V-JEPA 2 is a 1.2 billion-parameter model, https://x.com/AIatMeta/status/1932808881627148450

o3 (left) got an idle question wrong, o3-pro nailed it. Good first impression 🙂 (Q: If a full-size crossbow shoots 160m, estimate what a half-size replica would shoot…) https://x.com/johnowhitaker/status/1932821323979632783

RT @WesRothMoney: o3 pro one-shotted the Tower of Hanoi 10 disk problem (one of the more contested problems in Apple’s “”The Illusion of Th…”” / X https://x.com/code_star/status/1932679839682867296

HOLY SHIT IT’S FUCKING REAL LET THE PRICE WARS BEGIN OpenAI updated their pricing page. o3 is now cheaper than GPT-4o, but more importantly, cheaper than Sonnet 4 and Gemini 2.5 Pro I would cry if I were Anthropic and Google! https://x.com/scaling01/status/1932488441100468438

OpenAI just killed Claude 4 and Gemini 2.5 Pro if that 80% price drop is true (docs still show old pricing) It would also mean o3 would be cheaper than GPT-4o ? https://x.com/scaling01/status/1932437241592152161

After the o3 price reduction, we retested the o3-2025-04-16 model on ARC-AGI to determine whether its performance had changed. We compared the retest results with the original results and observed no difference in performance.”” / X https://x.com/arcprize/status/1932836756791177316

ARC-AGI-1 results for o3-pro and o3-high are in o3-pro (high) does not beat o3-high despite being slightly above 8 times more expensive https://x.com/scaling01/status/1932539254703321399

ARC-AGI-2 don’t look good for o3-pro (high) o3-pro (high) does not beat o3-high despite being 9 times more expensive https://x.com/scaling01/status/1932539573432684779

Been playing with o3-pro for a bit. It is quite smart. One problem it solved where every other model has failed is making word ladder from SPACE to EARTH. (Probably not contamination: the answer is different than the only online answer, which is for EARTH to SPACE in any case) https://x.com/emollick/status/1932533635984355792

o3-pro is much stronger than o3:”” / X https://x.com/gdb/status/1932561536268329463

o3-pro is rolling out now for all chatgpt pro users and in the api. it is really smart! i didnt believe the win rates relative to o3 the first time i saw them.”” / X https://x.com/sama/status/1932532561080975797

RT @GregKamradt: After the o3 price drop it made sense to test it on SnakeBench wow – it’s the new #1 model (out of 71 tested) It made th…”” / X https://x.com/imjaredz/status/1932898036466004317

The 80% price drop of o3 came with no performance trade-offs”” / X https://x.com/emollick/status/1932846451681337674

i like this take: “”The plan o3 gave us was plausible, reasonable; but the plan o3 Pro gave us was specific and rooted enough that it actually changed how we are thinking about our future.”””” / X https://x.com/sama/status/1932533208366608568

o3 was considerably less verbose in responses in our Artificial Analysis Intelligence Index eval set than Gemini 2.5 Pro & DeepSeek R1 but more than Claude 4 Opus https://x.com/ArtificialAnlys/status/1932489580592435301

In expert evaluations, reviewers consistently prefer OpenAI o3-pro over o3, highlighting its improved performance in key domains—including science, education, programming, data analysis, and writing. Reviewers also rated o3-pro consistently higher for clarity, comprehensiveness, https://x.com/OpenAI/status/1932530411651150013

New paper shows a familiar result on LLMs & medicine: Doctors given clinical vignettes produce significantly more accurate diagnoses when using a custom GPT built with the (obsolete) GPT-4 than doctors with Google/Pubmed but not AI. Yet AI alone is as accurate as doctors + AI. https://x.com/emollick/status/1931907652118069510

This graph (shows a steep decline in organic search traffic) https://x.com/fdaudens/status/1932501681628905788

What an incredible trajectory of performance improvements for the reasoning models since the original o1-preview ! 60%+ winrates are comparatively huge, and few model upgrades achieved this historically”” / X https://x.com/BorisMPower/status/1932556016455201145

What “”Working”” Means in the Era of AI Apps | Andreessen Horowitz https://a16z.com/revenue-benchmarks-ai-apps/

ScreenSuite – The most comprehensive evaluation suite for GUI Agents! https://huggingface.co/blog/screensuite

At WWDC 2025, Apple showed off only a handful of AI upgrades, including: —New Live translation for FaceTime, Messages, and calls —Visual intelligence via screenshots —AI-powered intelligent actions in Shortcuts —AI “”Workout Buddy”” on Apple Watch https://x.com/rowancheung/status/1932341247810678845

Vision Transformers have high computational costs. Existing token reduction methods like pruning and merging are exclusive, causing significant information loss and needing post-training to recover performance. This paper presents Token Transforming, a unified many-to-many https://x.com/rohanpaul_ai/status/1932718446648918269

reasoning LLMs are bad at puzzles if they are too hard”” is a less catchy title than The Illusion of Reasoning huh”” / X https://x.com/andersonbcdefg/status/1931821352463577482

Evals now supports tool use. 🛠️ You can now use tools and Structured Outputs when completing eval runs, and evaluate tool calls based on the arguments passed and responses returned. This supports tools that are OpenAI-hosted, MCP, and non-hosted. Read more in our guides below. https://x.com/OpenAIDevs/status/1932169029147557924

How do you evaluate Voice Agents? This is the talk for you with @kwindla He provides code/Github repo & does some fun demos in this talk (links in yt description) https://x.com/HamelHusain/status/1932204210994704625

📊Benchmarking Multi-Agent Architectures As more systems become multi-agent, this begs the question: how do you best orchestrate across multiple agents? We did some initial benchmarking, including some improvements to our supervisor approach Blog: https://x.com/LangChainAI/status/1932825652312600810

Glass with Deep Reasoning achieves new state-of-the-art performance on common clinical benchmarks. ✅ 97% on USMLE Steps 1–3 ✅ 98% on JAMA Clinical Challenge cases ✅ 90% on NEJM Clinicopathologic Case Conferences Available to clinicians at https://x.com/GlassHealthHQ/status/1933291603906736328

RT @HuggingPapers: LLMs often guess in math proofs! 😱 A new study, IneqMath, reveals LLMs struggle to construct rigorous proofs, even when…”” / X https://x.com/_akhaliq/status/1932894338616574091

Super impressive to see the new Gemini 2.5 Pro (06-05) climbing the public leaderboards! 🚀  > Best model at 192k tokens on Live Fiction > Number #1 on SimpleBench with 62.4% > Strongest over Document Processing model in IDP > Best cost-performance on Aider Kudos to everyone https://x.com/_philschmid/status/1932723220379049999

The new Gemini 2.5 Pro is SOTA at long context, especially capable on higher number of items being retrieved (needles) as shown below! https://x.com/OfficialLoganK/status/1931078494337073409

RT @xennygrimmato_: Gemini 2.5 Pro solved *all* JEE Advanced 2025 problems from the Mathematics Section (both Paper 1 and Paper 2)! * Goo…”” / X https://x.com/dilipkay/status/1932754214469402630

First time we experience a “perfect storm” of this magnitude. Half of the internet was down. On the bright side, @Replit is coming back! Thanks for powering through with us. https://x.com/pirroh/status/1933269623979585695

This paper proposes Agents Co-Evolution (ACE), a framework where LLMs guide Reinforcement Learning training offline, enabling superior performance in complex decision tasks. Methods 🔧: → ACE introduces a dual-role trajectory refinement mechanism. → LLMs act as a Policy https://x.com/rohanpaul_ai/status/1931249013761999118

Interesting argument: solving the Tower of Hanoi requires thousands of moves. Reasoning models are trained so that they will not “”think”” long enough to solve the puzzle at high complexity, not because LLMs can’t do it, but because deployed models are built to not think that long.”” / X https://x.com/emollick/status/1932096155049169155

Paper page – GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents https://huggingface.co/papers/2506.03143

The Utility of Interpretability — Emmanuel Amiesen – YouTube https://www.youtube.com/watch?v=9YQW2mH9FyA

When people talk about AGI and super intelligence and all of that, our conviction is that it will come from the community and the whole field collaborating on this topic.”” this convo btw @operationdanish and @ClementDelangue is packed with sharp insights https://x.com/fdaudens/status/1932136783443001432

This is a fantastic application of applied interpretability! When using llms to review resumes, prior debiasing techniques break in more realistic settings. But simply finding and removing gender or race directions remains effective, beating existing than baselines!”” / X https://x.com/NeelNanda5/status/1933645976889422110

We are excited to announce Trinity, an autoformalization system for verified superintelligence that we have developed at @morph_labs. We have used it to automatically formalize in Lean a classical result of de Bruijn that the abc conjecture is true almost always. https://x.com/morph_labs/status/1933181394588483868?s=46

I replicated the Anthropic alignment faking experiment on other models, and they didn’t fake alignment – LessWrong 2.0 viewer https://www.greaterwrong.com/posts/pCMmLiBcHbKohQgwA/i-replicated-the-anthropic-alignment-faking-experiment-on

just learned about “”model diffing”” from Anthropic. buried in an october blogpost; feels really novel. training a ‘crosscoder’ between two models of the same family produces interpretable diffs. here post-training clearly adds refusals, QA, math, etc. pretty amazing stuff https://x.com/jxmnop/status/1933571979975487996

RT @jiaxinwen22: New Anthropic research: We elicit capabilities from pretrained models using no external supervision, often competitive or…”” / X https://x.com/akbirkhan/status/1933323897526759553

Everyone talks about these two papers, but what do many people miss about them? 1. The Illusion of Thinking by Apple 2. How much do language models memorize? by FAIR, Google DeepMind, Cornell University, and NVIDIA They describe the same underlying breakdown – a model’s coping https://x.com/TheTuringPost/status/1932912650444550470

Every call I have had this week has had someone ask a question about the Apple paper. I think its worth reflecting on why any time a “”AI must fail”” paper comes out (also: model collapse), it gets a lot of buzz & why the many “”AI does this well”” papers don’t. Discomfort with AI?”” / X https://x.com/emollick/status/1932469064363814947

Good analysis (rebuttal) of Apple’s “”illusion of thinking”” Hanoi towers experiment.”” / X https://x.com/giffmana/status/1931801836052189191

I don’t even agree with the Apple paper but this is an extremely midwit take https://x.com/iScienceLuvr/status/1931877956257005904

I think the Apple paper on the limits of reasoning models in particular tests is useful & important, but the “LLMs are hitting a wall” narrative on X around it feels premature at best. Reminds me of the buzz over model collapse – limitations that were overcome quickly in practice”” / X https://x.com/emollick/status/1931449878653403569

The “”reasoning doesn’t exist”” Apple paper drives me crazy. Take logic puzzle like Tower of Hanoi w/ 10s to 1000000s of moves to solve correctly. Check first step where an LLM makes mistake. Long problems aren’t solved. Fewer thought tokens/early mistakes on longer problems. 1/11 https://x.com/Afinetheorem/status/1931853801293484358

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity – Apple Machine Learning Research https://machinelearning.apple.com/research/illusion-of-thinking

the-illusion-of-thinking.pdf https://ml-site.cdn-apple.com/papers/the-illusion-of-thinking.pdf

CVPR 2025: Exclusive Talk with Paul E. (CEO / CTO @ 3LC.AI) – YouTube https://www.youtube.com/watch?v=msbPP5td0cA&list=PLaU7MWI8yG9Uy8P_3K5R4_H6HZ5AsYsxD&index=13&t=1s

CVPR 2025: Exclusive Talk with Tom Bishop (Chief Technology Officer at GLASS Imaging) – YouTube https://www.youtube.com/watch?v=97flahWCRGA&list=PLaU7MWI8yG9Uy8P_3K5R4_H6HZ5AsYsxD&index=15&t=1s

CVPR 2025: Exclusive Talk with Yotam Azriel (CEO & CTO & Co-Founder at TensorLeap) – YouTube https://www.youtube.com/watch?v=x0csNvsdrvA&list=PLaU7MWI8yG9Uy8P_3K5R4_H6HZ5AsYsxD&index=12&t=107s

CVPR 2025: Graph Neural Network Combining Event Stream and Periodic Aggregation…… – YouTube https://www.youtube.com/watch?v=8_AkywWE9GE&list=PLaU7MWI8yG9Uy8P_3K5R4_H6HZ5AsYsxD&index=8&t=3s

CVPR 2025: Motion Prompting: Controlling Video Generation with Motion Trajectories – YouTube https://www.youtube.com/watch?v=LpnOZv4ziGA&list=PLaU7MWI8yG9Uy8P_3K5R4_H6HZ5AsYsxD&index=22

CVPR 2025: SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding – YouTube https://www.youtube.com/watch?v=uqjDXUgjR3c&list=PLaU7MWI8yG9Uy8P_3K5R4_H6HZ5AsYsxD&index=17&t=8s

CVPR 2025: The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour… – YouTube https://www.youtube.com/watch?v=HGMjyhWGx-4&list=PLaU7MWI8yG9Uy8P_3K5R4_H6HZ5AsYsxD&index=16&t=43s

🎙️ After serving millions of users through our text-to-speech platform, one need kept coming up: fine-grained AI speech editing – the ability to modify existing speech. Today, we’re open-sourcing PlayDiffusion, a diffusion-based inpainting model built for that exact purpose. https://x.com/PlayAIOfficial/status/1929558863319330822

eval work and staring at data are both incredibly important and incredibly boring”” / X https://x.com/finbarrtimbers/status/1933278968859468161

Get 2x faster for reward model serving and sequence classification inference through @UnslothAI! Nice benchmarks Kyle!”” / X https://x.com/danielhanchen/status/1932965003621204391

Interesting findings on potential limitations of reasoning models.”” / X https://x.com/emollick/status/1930720378361672130

my timeline is flooded with “”haiku is cracked”” tweets so i got out my spreadsheet again to see if this is legit dear reader, haiku officially cracked the cost-elo frontier https://x.com/swyx/status/1772799201023557697

We Made Top AI Models Compete in a Game of Diplomacy. Here’s Who Won. https://every.to/diplomacy

Chris is joining us for our next (and last!) live evals course next month! https://x.com/HamelHusain/status/1932657675294421061

Big question in AI is whether new entrants can still hope to reach the state-of-the-art, or whether learning curve plus compute needs are high enough that this is impossible. xAI did it with massive compute & hiring investment. But otherwise is the list of competitors fixed?”” / X https://x.com/emollick/status/1932945672300314659

AI + Continously updated personal health data fed to it can change healthcare. Correlations that aren’t obvious to humans jump out fast. AI isn’t YET replacing your doctor. It’s augmenting your ability to ask better questions. Most people don’t track enough data to even ask https://x.com/rohanpaul_ai/status/1931298548831731769

Corporate AI adoption may be leveling off, according to Ramp data | TechCrunch https://techcrunch.com/2025/06/09/corporate-ai-adoption-may-be-leveling-off-according-to-ramp-data/

Announcing HMAR – Efficient Hierarchical Masked Auto-Regressive Image Generation, led by @KumbongHermann! HMAR is hardware-efficient, reformulates autoregressive image generation in a way that can take advantage of tensor cores. Hermann is presenting it at CVPR this week!”” / X https://x.com/realDanFu/status/1932163091174981978

AlphaWrite: Inference time compute Scaling for Writing | Toby Simonds https://tobysimonds.com/research/2025/06/06/AlphaWrite.html

If I want to compare the generation costs, latency and other attributes for all the video models out there — is there a good resource you would recommend? Interested in everything from the Chinese open weights models to Veo 3 etc.”” / X https://x.com/bilawalsidhu/status/1932799827324064026

Announcing Magistral — our @MistralAI first reasoning model — excelling in domain-specific, transparent, and multilingual reasoning. https://x.com/sophiamyang/status/1932451856447586312

Announcing Magistral, our first reasoning model designed to excel in domain-specific, transparent, and multilingual reasoning. https://x.com/MistralAI/status/1932441507262259564

Magistral | Mistral AI https://mistral.ai/news/magistral

Mistral just released their reasoning models: Magistral-Small and Magistral-Medium Magistral Small is open-source and based on the 24B Mistral-Small 3.1, and can run on a single RTX 4090. Unfortunately it gets crushed by Qwen3-32B and Qwen3-30B-A3B Download Link: https://x.com/scaling01/status/1932445360380612712

Mistral really cooked – 24B, Based on Mistral Small 3.1, Multilingual, 128K context (40k effective), Apache 2.0 licensed! 🔥 Works on MLX, llama.cpp, transformers, vllm and more ⚡ https://x.com/reach_vb/status/1932449015657836730

The Mistral team at it again with Magistral! GRPO with edits: 1. Removed KL Divergence 2. Normalize by total length (Dr. GRPO style) 3. Minibatch normalization for advantages 4. Relaxing trust region Paper: https://x.com/danielhanchen/status/1932451325398413518

Mistrals paper is the best practical paper on doing reasoning rl since deepseek r1 paper fwiw will do a writeup later to go through it 🤓 if i get the time”” / X https://x.com/Teknium1/status/1932580993132790232

DesignBench provides a benchmark for multimodal LLMs evaluating front-end engineering across popular frameworks and tasks like generation, edit, and repair. Methods 🔧: → DesignBench contains 900 real-world webpage samples for HTML/CSS, React, Vue, and Angular frameworks. → https://x.com/rohanpaul_ai/status/1932279554954940445

[2411.12915] VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge https://arxiv.org/abs/2411.12915

CVPR 2025: VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge – YouTube https://www.youtube.com/watch?v=_Z2KMfDXkwY&list=PLaU7MWI8yG9Uy8P_3K5R4_H6HZ5AsYsxD&index=5&t=4s

Graph Neural Network Combining Event Stream and Periodic Aggregation for Low-Latency Event-based Vision https://openaccess.thecvf.com/content/CVPR2025/papers/Dampfhoffer_Graph_Neural_Network_Combining_Event_Stream_and_Periodic_Aggregation_for_CVPR_2025_paper.pdf

RT @mattshumer_: I really can’t believe I’m saying this, but for non-code tasks, o3 Pro feels MILES ahead of Claude Opus 4″” / X https://x.com/imjaredz/status/1932657322204987718

my minions returned and they observed no performance difference between o3 versions”” / X https://x.com/scaling01/status/1932839048273670563

OpenAI slashed o3’s price by 80% and dropped a new o3-pro that: —Thinks longer to boost performance —Beats rivals on PhD-level math & science —Can web search and data analysis (but no image generation & canvas) —Is available to ChatGPT Pro and Team users https://x.com/rowancheung/status/1932694638785122610

pretty embarrassing for OpenAI, at this point one could expect o3 to crush this kind of trickery with contempt. muh ARC-AGI! …I’ve started to fear that as models and datasets get bigger, we’ll be seeing persistent sloppiness because they fail to develop crisp general features. https://x.com/teortaxesTex/status/1933371065863909638

RT @flavioAd: I’ve been secretly testing o3-pro for a while now 👀 Extremely cheaper, faster, and way more precise than o1-pro (and coding…”” / X https://x.com/OpenAIDevs/status/1932538094168801492

RT @LechMazur: o3-pro sets a new record on the Extended NYT Connections, surpassing o1-pro! 82.5 → 87.3. This benchmark evaluates LLMs us…”” / X https://x.com/SebastienBubeck/status/1932656485341032719

We ran this eval yesterday before price drop 😆🫠 @OpenAI”” / X https://x.com/StringChaos/status/1932642264163180844

BREAKING: Apple just proved AI “”reasoning”” models like Claude, DeepSeek-R1, and o3-mini don’t actually reason at all. They just memorize patterns really well. Here’s what Apple discovered: (hint: we’re not as close to AGI as the hype suggests) https://x.com/RubenHssd/status/1931389580105925115

Wow, end-to-end omni model from Ant Group: Ming Lite Omni – can hear, speak, and generate images – competitive to GPT4o 🔥 Some notes on the paper and the release: > GUI tasks: +9% accuracy over Qwen2.5VL-7B on AITZ(EM) > Audio understanding: 6/13 SOTA results on public https://x.com/reach_vb/status/1933458455794229317

Most AI labs talk about merely “augmenting” humans at work. They say this because AI currently falls short, not out of some deep conviction. I’m not here to bullshit anyone. Mechanize’s explicit goal is to automate all work as quickly as possible. https://x.com/tamaybes/status/1932841955542904919

[2505.17548] H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips https://arxiv.org/abs/2505.17548

[2505.23075] Second Opinion Matters: Towards Adaptive Clinical AI via the Consensus of Expert Model Ensemble https://arxiv.org/abs/2505.23075

[2506.05231v1] Progressive Tempering Sampler with Diffusion https://arxiv.org/abs/2506.05231v1

[2506.08007] Reinforcement Pre-Training https://arxiv.org/abs/2506.08007

❓Why do some AI products explode in adoption while others struggle? It’s not just related to model capabilities – there’s also a lot of UX work that can be put into the product to make it be more likely to be adopted My friend @assaf_elovic had some great insights, so we wrote https://x.com/hwchase17/status/1933528162799136988

🚨We have a new prompting report: Prompting a model with Chain of Thought is a common prompt engineering technique, but we find simple Chain-of-Thought prompts don’t help recent frontier LLMs, including reasoning & non-reasoning models, perform any better (but do increase costs) https://x.com/emollick/status/1931716329922249206

30x daily maintenance job speedup from one simple hack: 41h → 82m 🧹 Bottleneck was paging through S3/GCS prefixes (1:1 w/namespaces). Added a read-through cache of every Nth prefix (N=page_size/2) so now we can list with arbitrary concurrency. https://x.com/turbopuffer/status/1932916345571848610

And to close out a trio of diffusion papers… Super excited to announce Grafting – a method for distilling pretrained diffusion transformers into *new architectures*, led by @keshigeyan! Swap attention for new primitives for 2% pretraining cost, exciting for modeling research!”” / X https://x.com/realDanFu/status/1932494049061318821

As we’ve hit peak real data, synthetic data can open new possibilities, such as: – Filling data gaps – Protecting privacy – Cutting costs – Reducing models bias But bad synthetic data can lead to model collapse – a loop of self-reinforcing errors. To get quality, we still need https://x.com/TheTuringPost/status/1933297536300929411

exciting to see that hybrid models maintain reasoning performance with few attention layers. benefits of linear architectures are prominent for long reasoning traces, when efficiency is bottlenecked by decoding – seems like a free win if reasoning ability is preserved as well!”” / X https://x.com/_albertgu/status/1932844922241233019

Heard that some eng teams at big co’s are now are now testing their API designs against LLMs before release. They run evals to see which API structure is easiest for the model to work with and redesign it if the model struggles to understand the format. I expect this to scale to”” / X https://x.com/alexalbert__/status/1933177502777913596

I dont think Ive seen anyone discuss this, but forgive me if Im not the first one to figure this out: did you know you can literally merge two transformer of width h with same arch into one transformer of width 2h, such that f_merge = (f1 + f2)/2 ? Just concat everything except”” / X https://x.com/cloneofsimo/status/1931566076116324392

I spent 20$ on poking holes in their shitty paper lol I think it’s time to stop this little ADHD project”” / X https://x.com/scaling01/status/1931818332321149416

I was wondering why LLMs can do so many steps on Hanoi but not in the other games… In the paper they use the minimum/optimal path length as a proxy for problem difficulty or compositional depth, like 2^N – 1 for Towers of Hanoi, (N+1)^2 – 1 for Checker Jumping and some linear https://x.com/scaling01/status/1931854370716426246

I wish I was as eloquent at story-telling as Joe. Some quotes here but please read the full original below. > I think for many people, myself included, you need to see this play out in some form to understand that prompts are ephemeral artifacts. Prompts are not source code.”” / X https://x.com/lateinteraction/status/1932551576100667416

If you want to remain competitive, and ensure that your model improvements continue in the near and long term you MUST be investing in data curation. Very exciting to see these latest results from @datologyai, which makes building better datasets suck far less.”” / X https://x.com/sarahcat21/status/1932447722659082577

it really is incredible what kinds of things become possible when RL on LLMs works. clearly we’re just getting started”” / X https://x.com/jxmnop/status/1933359925415325980

LLMs rely on text-only reasoning, so numeric steps drift and answers get long. Code-Integrated Reasoning (CIR) lets the model insert Python, execute it, keep the result, and learn this loop with a steadier reinforcement-learning recipe. Accuracy rises and responses shrink. So https://x.com/rohanpaul_ai/status/1931307165027012732

LLMs struggle to optimize intermediate reasoning steps using only full trajectory rewards. This new method estimates mathematical expectations of rewards at different steps through tree sampling, providing fine-grained guidance without a separate reward model. TREERPO https://x.com/rohanpaul_ai/status/1932064639078060169

more evidence that kv caches have a lot of room for compression.”” / X https://x.com/gallabytes/status/1932475390293156275

One thing we discovered in our new paper is that this is now a very bad prompt Asking modern models to “answer directly” results in less accurate outcomes than not including that instruction. Models do best when they can “think,” and telling them to be direct can hurt reasoning.”” / X https://x.com/emollick/status/1931849764183703864

Oxford Languages Whitepaper | The Strategic Value of Lexical Data in AI Development – https://languages.oup.com/the-strategic-value-of-lexical-data-in-ai-development/

Recent Frontier Models Are Reward Hacking – METR https://metr.org/blog/2025-06-05-recent-reward-hacking/

RT @gm8xx8: Cartridges: Storing long contexts in tiny caches with self-study – train-once, reusable memory via SELF-STUDY – 38.6× less mem…”” / X https://x.com/simran_s_arora/status/1932304114857509305

RT @jiaxinwen22: want to clarify some common misunderstandings – this paper is about elicitation, not self-improvement. – we’re not adding…”” / X https://x.com/jeremyphoward/status/1933364618371739948

RT @Michael_D_Moor: Excited to announce MIRIAD — a large-scale dataset of 5,821,948 medical question-answer pairs, each rephrased from pass…”” / X https://x.com/lateinteraction/status/1932159654068674880

RT @qx_dong: ⏰ We introduce Reinforcement Pre-Training (RPT🍒) — reframing next-token prediction as a reasoning task using RLVR ✅ Gen…”” / X https://x.com/kylebrussell/status/1932454923322450186

RT @vineettiruvadi: Took just 20 minutes to make a synthetic clinical note generator with @DSPyOSS. It’s far from perfect but instead of j…”” / X https://x.com/lateinteraction/status/1932810369262887015

RT @yandexcom: Yandex released Yambda — a massive public dataset for recommender systems. It includes nearly 5B anonymized user–track inter…”” / X https://x.com/_akhaliq/status/1932872791768117483

stop building parser pipelines 👋🏻 there’s a new document parser that is small, fast, Apache 2.0 licensed and is better than all the other ones! 😱 MonkeyOCR is a 3B model that can parse everything (charts, formules, tables etc) in a document 🤠 https://x.com/mervenoyann/status/1932830954722054345

The Common Pile v0.1 | EleutherAI Blog https://blog.eleuther.ai/common-pile/

The origin story of “AI as Normal Technology”, and lessons learned Many people have asked how the “AI as Normal Technology” paper came to be. This paper has been an (ongoing) journey for me and @sayashk in developing not just the substance of our arguments but also learning how”” / X https://x.com/random_walker/status/1932410929112572384

This paper proposes Table-R1 models using two post-training strategies, distillation and reinforcement learning, enabling inference-time scaling to improve table reasoning performance and generalization, even with smaller models. Methods 🔧: → The paper creates a large dataset https://x.com/rohanpaul_ai/status/1931218059412250961

We’re excited to introduce Text-to-LoRA: a Hypernetwork that generates task-specific LLM adapters (LoRAs) based on a text description of the task. Catch our presentation at #ICML2025! Paper: https://x.com/SakanaAILabs/status/1932972420522230214

What if an LLM could update its own weights? Meet SEAL🦭: a framework where LLMs generate their own training data (self-edits) to update their weights in response to new inputs. Self-editing is learned via RL, using the updated model’s downstream performance as reward. https://x.com/jyo_pari/status/1933350025284702697

Dwarkesh Patel on Continual Learning – by Zvi Mowshowitz https://thezvi.substack.com/p/dwarkesh-patel-on-continual-learning

Hackaprompt 2.0 https://www.hackaprompt.com/track/pliny

[2506.07330] JavelinGuard: Low-Cost Transformer Architectures for LLM Security https://www.arxiv.org/abs/2506.07330

Think Like a Person Before Responding Online hate speech needs effective automated counter-narratives (CNs), but tone, accessibility, and ethics pose concerns. This paper evaluates LLMs’ counter-narrative generation using persona prompts, assessing four dimensions. Methods 🔧: https://x.com/rohanpaul_ai/status/1931233410690818528

Our previous projects at MedARC on fMRI mind reading focused on reconstructing what a person is simultaneously seeing. But that is not true mind reading. True mind reading would reconstruct directly from imagination! Researchers at UMN (+my cofounder @humanscotti) have created https://x.com/iScienceLuvr/status/1932945933521817988

RL LLM training for medicine is very underexplored right now and there is still significant room to improve clinical reasoning abilities It is difficult though because how do you turn medicine into verifiable problems? This is what I spend a lot of time thinking about now”” / X https://x.com/iScienceLuvr/status/1931694421239902474

[2412.02700] Motion Prompting: Controlling Video Generation with Motion Trajectories https://arxiv.org/abs/2412.02700

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading