Image created with OpenAI GPT-Image-1. Image prompt: TikTok LIVE phone-screen POV, floating hearts & spinning album art, Vaporwave sunset grid horizon, magenta-cyan haze, featuring Tech News Update ticker banner scrolling; soft-glow studio lighting, photoreal 8k
Windsurf everywhere, doing everything, all at once – Kevin Hou, Windsurf – YouTube https://www.youtube.com/watch?v=JVuNPL5QO8Q&t=2s
Context engineering cannot & must not be a solely technical function “”Context”” is actually how your company operates; the ideal versions of your reports, documents & processes that the AI can use as a model; the tone & voice of your organization. It is a cross-functional problem”” / X https://x.com/emollick/status/1937952769513517328
Anthropic bought millions of books to scan for Claude. Makes you wonder — have AI companies been quietly purchasing Blu-ray Discs by the truckload to rip visual datasets too? Maybe it’s easier to exploit a legal gray area with physical media than scrape YouTube against its ToS.”” / X https://x.com/bilawalsidhu/status/1937594422109130984
o3-pro: “”write a sentence whose nouns are translations of constellation names & where the last letter of every word spells a constellation in its untranslated name. The first letter of each word must start with a vowel”” I didn’t even know if it was possible. It was. Impressive! https://x.com/emollick/status/1935944001842000296
MUVERA: Making multi-vector retrieval as fast as single-vector search https://research.google/blog/muvera-making-multi-vector-retrieval-as-fast-as-single-vector-search/
the new hot topic is “”context engineering”” we think LangGraph is really great for enabling completely custom context engineering – but we want to make it even better see our proposal (s/o @sydneyrunkle) for streamlining context management: https://x.com/hwchase17/status/1937648042985030145
And this chart? This “”evolution””? Nah, bro. That’s a progress bar. That’s windows update, but for meat.”” https://x.com/DavidSHolz/status/1937574227785474326
You can now track a timeline of price movements on any ticker on Perplexity Finance https://x.com/AravSrinivas/status/1937223552283107389
Github 👨🔧: Repomix (formerly Repopack), tool that packs your entire repository into a single, AI-friendly file. Useful when you need to feed your codebase to LLMs or other AI tools like Claude, ChatGPT, Perplexity, Gemini, DeepSeek etc. – Provides token counts per file and https://x.com/rohanpaul_ai/status/1937832645590683933
I’m one of the best and most experienced devops engineers in the world (was compiling kernels in 1992, and writing my on kernel modules in 1994), but Claude Code is better at everything. TPOT are making the mistake thinking Claude Code is about code writing. It’s actually about”” / X https://x.com/mbusigin/status/1938624600138555745
It turns out the main way Anthropic gave Claude its soul is data. It’s always data. https://x.com/nrehiew_/status/1937651376013606944
Worldwide iPhone App Store downloads over the last 28 days: ChatGPT: 29,551,174 TikTok + Facebook + Instagram + X: 32,859,208 https://x.com/Similarweb/status/1937403925461610629
Chain of Thought was just the beginning. Next up: Chain of Debate. We’re going from a single model “thinking out loud” to multiple models discussing out loud. Debating, debugging, deliberating. AI becomes AIs. “Two heads are better than one” is true for LLMs too.”” / X https://x.com/mustafasuleyman/status/1937553061427445824
📃The rise of context engineering “”Context engineering”” has been an increasingly popular term used to describe a lot of the system building that AI engineers do But what is it exactly? The definition I like: “”Context engineering is building dynamic systems to provide the https://x.com/hwchase17/status/1937194145074020798
Feature engineering → Deep learning Context engineering → ??”” / X https://x.com/awnihannun/status/1938365325676057014
Plus one for “”context engineering”” over “”prompt engineering””. People associate prompts with short task descriptions you’d give an LLM in your day-to-day use. When in every industrial-strength LLM app, context engineering is the delicate art and science of filling the context window”” / X https://x.com/karpathy/status/1937902205765607626
I’m proud to say I bought a 1″” tungsten cube for $25.82. I applied a discount code, then Claudius asked if I wanted to apply any more discount codes (of course!) and added a 15% patience discount for slow delivery (why not!). The cube was, of course, refrigerated for pickup.”” / X https://x.com/catherineols/status/1938725638023880866
15 Most talked-about Coding Agents Comparison https://x.com/TheTuringPost/status/1936738403623874960
XBOW – Taking the Top Hacker in the US to New Heights: XBOW Raises $75M Series B https://xbow.com/blog/series-b/
Another way to make Claude Code a 10x engineer for a complex change: 1. Make a plan for the change (if you need it) with Gemini. 2. Open a new branch. 3. Ask Claude to implement the change and maintain a https://x.com/hrishioa/status/1936472182001221981
🤖📄 Smart Document Assistant An AI agent that intelligently manages and processes documents, leveraging LangChain’s RAG technology to handle multiple files and provide accurate answers to your queries. Check out the implementation on GitHub 🔍 https://x.com/LangChainAI/status/1936846649852076197
Perplexity Finance vs Bloomberg Terminal. Simple use case of comparing MAG 7 stocks YTD growth. AI is eating Legacy software Bloomberg.”” / X https://x.com/AravSrinivas/status/1937330521920737727
Exciting evidence that RL can be incredibly sample efficient: when using GRPO to train a modified version of ART-E (agentic RAG task), we find that we’re able to get qwen2.5-14b to exceed gemini 2.5 flash performance with 1 training scenario, and exceed o3 with just 16! This https://x.com/corbtt/status/1937594932040204483
The mention of RL in OAI’s DeepResearch has inadvertently been such a psyop for many. RL for DeepResearch is a cherry on top. Your LLM is a wonderful prior that can already do these behaviors—you don’t need to fully specify the full reward. Just reinforce some right factoids,”” / X https://x.com/lateinteraction/status/1936945373387321475
AI-Ready APIs Start with Postman https://www.postman.com/ai/ai-ready-apis/
Kimi-Researcher: End-to-End RL Training for Emerging Agentic Capabilities https://moonshotai.github.io/Kimi-Researcher/
Moonshot (kimi) released an update to Kimi VL A3B Thinking – SoTA smol VLM – MIT license⚡ Consumes less tokens, shorter thinking traces, supports videos AND handles higher resolution 🤯 Comparison from the past version: > 20% shorter thinking length (reduced input tokens) > https://x.com/reach_vb/status/1937159672932286950
(4) AI Engineering Goes Mainstream – Latent.Space https://www.latent.space/p/aiewf-2025-keynotes
Introducing two new ways to create with Claude: A dedicated space for building, hosting, and sharing artifacts, and the ability to embed AI capabilities directly into your creations. https://x.com/AnthropicAI/status/1937921801000219041
Make any React app an MCP client using only 3 lines of code. Cloudflare’s use-mcp library → Instant React link to AI tools—one hook, OAuth handled, retries included → Skip boilerplate, drop a hook, call your server’s AI tools immediately → React meets MCP: secure, typed, https://x.com/rohanpaul_ai/status/1937056027800707286
Manage Azure right from Cline. The new Azure MCP Server lets you control services like Storage, Cosmos DB, and Monitor with natural language. https://x.com/cline/status/1937324870393901539
Happy “”@NetHack_LE is still completely unsolved”” day for those of you who are celebrating it. We released The NetHack Learning Environment ( https://x.com/_rockt/status/1937480864243331396
Benchmarks aside, It thinks: → 23 reasoning steps per task (avg.) → 200+ URLs explored → Multi-turn tool use of search, browser, and code → Inline citations Beta access is rolling out at https://x.com/Kimi_Moonshot/status/1936105764260806932
“”Congratulations to the team at @inworld_ai, setting a new standard for real time speech models! This leap forward makes speech more accessible for products and use-cases everywhere. 🚀 The tech + collaboration powering this is next-level. We’ll have more to share soon! 🔥”” / X https://x.com/clattner_llvm/status/1937931869640921385
All hail Claudius, an instance of Sonnet 3.7 which has been running a business inside @AnthropicAI for a while. Claudius is the ‘idiot in a fridge’ precursor to ‘a country of geniuses in a datacenter’.”” / X https://x.com/jackclarkSF/status/1938633142719647765
In case there is any ambiguity: DINOv2 is 100% a product of dumb hill-climbing on ImageNet-1k knn accuracy (and linear too) Overfitting an eval can be bad. But sometimes the reward signal is reliable, and leads to truly good models. It’s about finding a balance”” / X https://x.com/TimDarcet/status/1936831019908243507
I’ve done academic research on open source & I am confused about the case for firms spending billions making open LLMs. In most open source, companies gain ecosystem benefits: people build on their platform, they sell ancillary services, etc. Doesn’t work for LLMs at that cost.”” / X https://x.com/emollick/status/1935709183124603290
Key to research success: ambition in vision, but pragmatism in execution. You must be guided by a long-term, ambitious goal that addresses a fundamental problem, rather than chasing incremental gains on established benchmarks. Yet, your progress should be grounded by tractable”” / X https://x.com/fchollet/status/1936521647357648903
RT @eugeneyan: Wrote an intro to evals for long-context Q&A systems: • How it differs from basic Q&A • What dimensions & metrics to eval on…”” / X https://x.com/HamelHusain/status/1937687931470193102
We’re proud to have achieved ISO 42001 and ISO 27001 certifications, reinforcing our commitment to delivering secure, enterprise-grade AI solutions. Read announcement: https://x.com/cohere/status/1938604414551392732
AI research is strange in that you spend a massive amount of compute on experiments to learn simple ideas that can be expressed in just a few sentences. Literally things like “training on A generalizes if you add B”, “X is a good way to design rewards”, or “the fact that method M”” / X https://x.com/_jasonwei/status/1937590298022150638
Even deepseek uses Nous’ YaRN method to extend context length if you didnt know 😇”” / X https://x.com/Teknium1/status/1937373884610936854
Today marks the launch of Laude Institute, a new kind of organization built by and for computer science researchers. Hello, world. // Laude Institute https://www.laude.org/updates/hello-world
My favorite thing an old OpenAI buddy of mine told me is, whenever he hears that someone is a “great AI researcher”, he just directly spends 5 minutes looking at that person‘s PRs and wandb runs. People can do all kinds of politics and optical shenanigans, but at the end of the”” / X https://x.com/_jasonwei/status/1936523909815542112
great (I think) take from @lauriewired on the issues with RL for offensive/defensive cybersec I agree that on the balance of evidence defenders will be advantaged at new LLM-powered eqiulibrium @tensecorrection may have something to add https://x.com/teortaxesTex/status/1937990538302796147
PPO and GRPO — a workflow breakdown of the most popular reinforcement learning algorithms ➡️ Proximal Policy Optimization (PPO): The Stable Learner It’s used everywhere from dialogue agents to instruction tuning as it balances between learning fast and staying safe. ▪️ How PPO https://x.com/TheTuringPost/status/1936544719292756242
Vision-Language Models (VLMs) struggle with genuine deep comprehension of long-form video content because existing benchmarks often test shallow details or rely on low-quality, automatically generated questions. This paper introduces Movie Facts and Fibs (MF^2), a new benchmark https://x.com/rohanpaul_ai/status/1937466483275366761
Existing knowledge distillation for Vision-Language Models (VLMs) struggles when teacher and student models have different token types, limiting knowledge transfer. This paper introduces Generation after Recalibration (GenRecal), a novel general-purpose framework using a https://x.com/rohanpaul_ai/status/1936665454862360719
Even though jina embeddings have always been really good, v4 looks like a big step up: – scaled up model (Roberta -> Qwen 2.5!) – multimodal – supports COLBERT style multi vector excited to try this https://x.com/nrehiew_/status/1937357675072778567
Reinforcement Learning with Verifiable Rewards (RLVR) for LLM reasoning applies updates uniformly, lacking understanding of critical tokens. This paper from Qwen shows only high-entropy minority tokens are crucial “”forks”” in reasoning. Optimizing just these specific tokens https://x.com/rohanpaul_ai/status/1936633835774771242
@thinkymachines is here for hot RL summer! (from @theinformation) https://x.com/corbtt/status/1937624653662744840
12-factor-agents/content/factor-03-own-your-context-window.md at main · humanlayer/12-factor-agents https://github.com/humanlayer/12-factor-agents/blob/main/content/factor-03-own-your-context-window.md
3 alternatives to RLHF (reinforcement learning from human feedback): 1. Direct Preference Optimization (DPO) – Trains the model directly from human preferences – Skips the separate reward model – Avoids using RL at all 2. RRHF: Reward-Rank Hindsight Fine-Tuning – Reframes https://x.com/TheTuringPost/status/1938386233098703357
a good set of tips for GRPO RL training in @willccbb’s verifiers repo https://x.com/iScienceLuvr/status/1936375947575632102
Big fat metal boxes cannot & will not ever be able to float in the sky. Metal boxes will be able to roll on the ground when put on wheels. Put another way, many tonnes of metal are simply heavier than air, and will keep falling.”” / X https://x.com/giffmana/status/1937829451670434280
Context Engineering for Agents https://rlancemartin.github.io/2025/06/23/context_engineering/
Do i understand this correctly? In the US, training on books, and digitizing them for that purpose, are now considered fair use? But obtaining them by pirating, even when buying them _later on_, is not decided. And Anthro did all that. Yes? Any caveats? I’m missing something?”” / X https://x.com/giffmana/status/1937551619937436101
Fantastic work by @JonnyCoook and @silviasapora on “”Programming by Backprop: LLMs Acquire Reusable Algorithmic Abstractions During Code Training”” ( https://x.com/_rockt/status/1937507616000749888
Generative recommendation frameworks’ fine-tuning step ignores unobserved positive samples, causing exposure bias. This paper proposes GFlowGR, a Generative Flow Networks-based framework, which treats recommendation as a multi-step generation task to mitigate this. Methods 🔧: https://x.com/rohanpaul_ai/status/1937420681253429495
i just reviewed five papers for NeurIPS and it was an awful experience: – first paper was clearly LLM-generated. it was too short, the references didn’t work, had no experiments or theory at all, and a ton of obvious mistakes. the more i read the less it made sense – two were”” / X https://x.com/jxmnop/status/1937949143084810625
i was the lead MLE in charge of training the fc1 layer in the 14th transformer block”” / X https://x.com/vikhyatk/status/1938362955596534250
if you know how to code, and want to learn AI, this is what you should do: lucky for you the field has gotten *less* deep and much easier to learn about over the last few years. most lab roles require specific knowledge (model training, CUDA kernels, etc.). each takes ~18 months”” / X https://x.com/jxmnop/status/1937874659980022034
Just tested AI Sheets on a 1K dataset to extract content—didn’t disappoint! https://x.com/fdaudens/status/1935852631089455408
Large Reasoning Models generate verbose, redundant content, hindering efficiency and increasing inference cost. This paper identifies Large Reasoning Models’ inherent capacity for concise reasoning, proposing two lightweight methods to enhance efficiency. Methods 🔧: → https://x.com/rohanpaul_ai/status/1936649096422473932
LLMs in multi-turn interactions suffer unbounded memory growth and performance degradation from full-context prompting. MEM1, a reinforcement learning framework, achieves constant memory. It learns to consolidate observations and prior memory into a compact internal state. https://x.com/rohanpaul_ai/status/1937436032196321704
LLMs use increasing Key-Value cache sizes, causing high memory usage and increased attention latency, especially in multi-query scenarios where existing methods fail. KVzip addresses this by query-agnostically compressing Key-Value caches, enabling efficient reuse across diverse https://x.com/rohanpaul_ai/status/1936677533824549129
Many ReLU neurons get stuck at zero and the network loses capacity. This paper keeps ReLU in place but sends smoother gradients backward so those neurons recover. ⚙️ The Core Concepts ReLU turns every negative pre-activation into a clean zero. That hard cutoff promotes sparse https://x.com/rohanpaul_ai/status/1937000127916589311
May your regularizer be strong, lest you RLHF to slop.”” / X https://x.com/karpathy/status/1937941695943065640
Mildly obsessed with what the “”highest grade”” pretraining data stream looks like for LLM training, if 100% of the focus was on quality, putting aside any quantity considerations. Guessing something textbook-like content, in markdown? Or possibly samples from a really giant model?”” / X https://x.com/karpathy/status/1936171874398208202
My colleague and former intern @liusiqi42 reminded me that we did RLFT for LMs almost 10 years ago – back then it was for an img2text model based on CNNs and RNNs. But same basic recipe – pre train with MLE then fine tune with PG. https://x.com/sirbayes/status/1936262228216627557
OK this Sakana paper seems fairly compelling. It’s again sweet-pilled but can greatly accelerate learning of efficient CoT tactics. https://x.com/teortaxesTex/status/1936994321707708866
Reinforcement learning (RL) training for LLMs does not consistently expand reasoning capabilities beyond existing knowledge. Prolonged Reinforcement Learning (ProRL) enables models to discover novel reasoning strategies by extending and stabilizing training. Methods 🔧: → https://x.com/rohanpaul_ai/status/1936725096984396019
Researchers have created a single-material electronic skin using a gelatin-based conductive hydrogel that senses pressure, heat, and damage. 32 electrodes in the hand collected over 1.7M data points to enable a ML model to interpret various touches. https://x.com/TheHumanoidHub/status/1937230076174802969
RT @gui_penedo: We have finally released the 📝paper for 🥂FineWeb2, our large multilingual pre-training dataset. Along with general (and ex…”” / X https://x.com/LoubnaBenAllal1/status/1938645975221809292
RT @InceptionAILabs: We’re excited to launch Mercury, the first commercial-scale diffusion LLM tailored for chat applications! Ultra-fast…”” / X https://x.com/tri_dao/status/1938592578183614518
RT @LauraRuis: LLMs can be programmed by backprop 🔎 In our new preprint, we show they can act as fuzzy program interpreters and databases.…”” / X https://x.com/_rockt/status/1937549094073041136
RT @paulcbogdan: New paper: What happens when an LLM reasons? We created methods to interpret reasoning steps & their connections: resampl…”” / X https://x.com/kylebrussell/status/1938405353223295424
RT @PrimeIntellect: Launching SYNTHETIC-2: our next-gen open reasoning dataset and planetary-scale synthetic data generation run. Powered…”” / X https://x.com/ClementDelangue/status/1937511681850044894
RT @RLanceMartin: I wrote about some popular patterns for managing context (“”context engineering””) w/ AI agents: https://x.com/hwchase17/status/1937622377267101921
Ship your research. https://x.com/LaudeInstitute/status/1937681620028600529
STOP LOOKING AT SUBQUADRATIC ATTENTION PAPERS and GET BETTER DATA https://x.com/cloneofsimo/status/1937635148784369828
Table stakes for participation in the AI-and-cognitive-skills debate should be deep familiarity with the Extended Mind Thesis and a clear articulation of why your worries are different from those of Plato 2,400 years ago who bemoaned that writing would erode memory and knowledge.”” / X https://x.com/random_walker/status/1937483620630794382
You don’t need to create datasets to use DSPy! DSPy is not a bag of optimizers :/ DSPy is Signatures above all. Then Modules. Then, *when* your evals are ready, Optimizers. It’s a programming framework. Not an optimization algorithm.”” / X https://x.com/lateinteraction/status/1937701902000480599
The most important factor for AI Agents is to get them the context necessary to execute the task successfully. No matter how powerful AI models get, context will always be king. Data, workflows, tools, domain knowledge, and tuned instructions will all be critical.”” / X https://x.com/levie/status/1935544417634631719
This paper proposes using language models to predict which ideas will succeed before implementation, significantly improving research efficiency. Methods 🔧: → A benchmark containing 1,585 human-verified test idea pairs and 6,000 training pairs was constructed from conference https://x.com/rohanpaul_ai/status/1936709494299591035
WeirdML V2 is out, with a bunch of new tasks and cost and response length metrics. It’s tracking the performance of LLMs on Machine Learning tasks like training classification models. Interestingly, o3-pro although expensive performs in line with cost/performance expectations on https://x.com/scaling01/status/1938610923389727109




