Image created with OpenAI GPT-Image-1. Image prompt: Cheesy late-night infomercial freeze-frame—split-screen bar-chart growth animation featuring binder “BENCHMARKS BOOSTER-KIT™”; teal headlines, 35 mm grain, high-resolution
If you are beginning a software engineering career, you MUST be at the the top 1%, else it will be very difficult. OpenAI’s SWE AI coding agent Codex merged 352K+ Pull Request with 85.5% success rate. And this number is just for previous 35 days. This repo tracks the opened https://x.com/rohanpaul_ai/status/1936041618554954142
II-Medical-8B-1706 is our latest state of the art open medical model 💡 Outperforms the latest @Google MedGemma 27b model with 70% less parameters 🤏 Quantised GGUF weights, works on <8 Gb RAM 🚀 One more step to the universal health knowledge access that everyone deserves ⚕️ https://x.com/ii_posts/status/1934959488710094990
Intelligent Internet introduced II-Medical-8B-1706, an updated version of its open medical model Capable of running on <8GB RAM, the AI outperformed Google’s MedGemma 27B across benchmarks despite 70% fewer parameters https://x.com/rowancheung/status/1935247303524114645
II-Medical-8B-1706, an 8B param health model reached GPT-4o/4.1/4.5 benchmarks, surpassing physicians. Intelligence is gradually ceasing to be human-exclusive. https://x.com/rohanpaul_ai/status/1935309527832056065
in the last 35 days, @OpenAI codex has merged 345,000 PRs on github. 345,000. AI is eating software engineering https://x.com/AnjneyMidha/status/1935865723328590229
Who did it best? Simple svg prompt, one-shot : r/singularity https://www.reddit.com/r/singularity/comments/1l486ji/who_did_it_best_simple_svg_prompt_oneshot/
The progress of Gemini over the last year + https://x.com/OfficialLoganK/status/1935136191927501235
RT @rohanpaul_ai: This is really BAD news of LLM’s coding skill. ☹️ The best Frontier LLM models achieve 0% on hard real-life Programming…”” / X https://x.com/sainingxie/status/1934994111536251361
A useful piece on criticizing AI: “its all PR” and “they are just parrots” are increasingly dead ends in a world where AI clearly can do effectively novel & important tasks. AI calls for robust criticism, but that criticism needs to be more grounded in the current state of LLMs.”” / X https://x.com/emollick/status/1935020901172387979
When people are delighted… it when humanity loses to AI…..There’s a handful of personal use cases that make me really bullish on hyper personalized utility. One of them is using ChatGPT as a personal running coach. Fed it all my run stats going back years, and said hey I have a race on this date, I wanna run x pace and keep my HR”” / X https://x.com/raizamrtn/status/1935781113513091107
[2506.11763] DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents https://arxiv.org/abs/2506.11763
Introducing ALE-Bench, ALE-Agent! Towards Automating Long-Horizon Algorithm Engineering for Hard Optimization Problems Blog: https://x.com/SakanaAILabs/status/1934767254715117812
One challenge no AI model has been able to do well: “”create a coherent, thematic puzzle for a D&D game. The puzzle should be challenging, but solvable”” The current big models are much more on theme than older ones, but still are either too easy or hard (And love similar puzzles) https://x.com/emollick/status/1934854293649006837
Sakana AI developed a new coding agent, ALE-Agent, trained to solve NP-hard optimization problems. Our agent participated in a live coding competition, the challenging AtCoder Heuristic Contest, and ranked #21 out of 1,000 human participants! Learn more: https://x.com/hardmaru/status/1934767617895747862
Switching LLMs wastes tokens. That’s why I built an n8n AI Agent that picks the best LLM for best performance & cost. • n8n JSON template • Prompts included • Setup guide It’s 100% FREE! Just: • Follow • Like • Reply “”ROUTER”” I’ll DM you. https://x.com/TheVeller/status/1922023956980101205
Why did the new coding LM Kimi-Dev-72B have a 43% accuracy drop when used in a different harness? The reason lies in the difference between agentic and agentless approaches to doing bug-fixing on repos, as explained in the thread!”” / X https://x.com/gneubig/status/1935028296565309807
Error checking is a great application of generative AI capabilities and there are low-hanging fruits in just about every domain: – Software: automatic detection of security vulnerabilities – Writing: identifying logical gaps, unclear structure, and weak arguments (can be seen as”” / X https://x.com/random_walker/status/1935311882857947507
Tracing + Evals w/o LangChain/Graph How to get the benefits of LangSmith (evals + tracing) + Studio (testing) w/o using LangChain or LangGraph? Here, we walk through the from scratch, using a non-LangChain/Graph agent as an example! 📽️: https://x.com/LangChainAI/status/1935706402896707657
RT @googleaidevs: 🩺 Get started with MedGemma, a collection of Gemma 3 variants built for medical text and image comprehension. Choose betw…”” / X https://x.com/osanseviero/status/1936096973691539652
Investment analyst Mary Meeker has released “Trends — Artificial Intelligence (May ’25),” her first tech market survey since 2019. 👉 The 340-page, data-rich report argues that AI’s breakneck adoption and escalating capital spending are fueling both record opportunities and https://x.com/DeepLearningAI/status/1934823324183396517
RT @lmarena_ai: 🚨Breaking: New DeepSeek-r1 (0528) just tied for #1 in WebDev Arena, matching Claude Opus 4! More highlights: 💠 #6 Overall…”” / X https://x.com/ClementDelangue/status/1934714392693588415
New test can help driverless cars make ‘moral’ decisions https://techxplore.com/news/2025-06-driverless-cars-moral-decisions.html
📚 We just enhanced YourBench with support for cross-document question generation! Now you can create evaluation datasets where questions span across multiple documents, not just one Just add a cross_document section to your config and set `enable: true` • you’ll get a https://x.com/ailozovskaya/status/1934962439247851889
60.4% on SWE-bench Verified in a 72B package? https://x.com/scaling01/status/1934746243286319435
98.5th percentile for 10 cents is now considered “”BAD NEWS for LLMs”” https://x.com/scaling01/status/1935058896806457427
Almost all randomized controlled trials on the impacts of AI on innovation, productivity, & job performance pre-dates reasoner models. The ones we do have (a couple of tests of o1-preview in law & medicine) suggest they may lead to a large jump in many fields, but we don’t know”” / X https://x.com/emollick/status/1933282450949689851
stop using VLMs blindly ✋🏻 compare different VLM outputs on a huge variety of inputs (from reasoning to OCR!) 🔥 > has support for multiple VLMs: Gemma 3, Qwen2.5VL, Llama4 > recommend us new models or inputs, we’ll add 🫡 https://x.com/mervenoyann/status/1935708014645784713
o3-pro does by far the best so far at my benchmark (scroll quote tweet thread for others): “”create a visually interesting shader that can run in twigl app make it like the ocean in a storm”” It did take 21 minutes for o3-pro to think (and another 19 to fix a small shader error) https://x.com/emollick/status/1932995067091800066
For years, we’ve been saying that bigger isn’t always better for AI and that smaller specialized models are usually faster, cheaper and more accurate for your specific constraints. So super happy to release the long-overdue capability of finding the best model based on size on https://x.com/ClementDelangue/status/1934672721066991908
Here’s a downloadable preview of the first chapter of our book on AI Evals written by @sh_reya and I, with a full table of contents. We are currently using this in our course and plan on eventually expanding it into a book. Feedback on TOC welcome! https://x.com/HamelHusain/status/1933912566910378384
It seems like benchmarking papers are increasingly being discussed as if they were proofs of AI limitations in the long run. Benchmarks are not useful if they are already saturated (and the fact that AI is at the 98.5% of human coders on a task set seems to be important as well)”” / X https://x.com/emollick/status/1935107962835444080
Part 2 of this mystery. Spotted on reddit. In my test not 100% reproducible but still quite reproducible. 🤔 https://x.com/karpathy/status/1935404600653492484
Ten years ago today, @OriolVinyalsML and I published this paper on arxiv. Back then, we didn’t know how to evaluate chatbots so we chatted with model and showed the samples in the paper. Glad that 10 years later, chatbots are still cool – and vibe-checking is going strong :-)”” / X https://x.com/quocleix/status/1936170043332825164
Excited to release AbstentionBench — our paper and benchmark on evaluating LLMs’ *abstention*: the skill of knowing when NOT to answer! Key finding: reasoning LLMs struggle with unanswerable questions and hallucinate! Details and links to paper & open source code below!
https://x.com/polkirichenko/status/1934730967446638644
There’s a tech report for gemini 2.5 now. If i don’t see this citation on the first page of any LLM paper i’m not gonna read it. 😃”” / X https://x.com/agihippo/status/1935015620250305018
With just 5 nodes in n8n + Apify, I’ve automated Instagram market research. Here’s exactly how I built it: https://x.com/samruddhi_mokal/status/1924379123096420366
o3-pro is rolling out now for all chatgpt pro users and in the api. it is really smart! i didnt believe the win rates relative to o3 the first time i saw them.”” / X https://x.com/sama/status/1932532561080975797
Redditor says ChatGPT saved his wife’s life by correcting a doctor’s fatal misdiagnosis. Comments are filled with people sharing their own stories. I don’t understand the AI haters at all. This technology saves lives. https://x.com/deedydas/status/1933370776264323164
Sam says Zuck🦎 is luring OpenAI researchers with $100M signing bonuses and $100M+ yearly salaries : r/ChatGPT https://www.reddit.com/r/ChatGPT/comments/1leciub/sam_says_zuck_is_luring_openai_researchers_with/
Back-of-the envelope it seems like each ChatGPT 4o query costs less than a cent, given the .34 Wh energy use per average prompt & a billion prompts a day & public GPU costs/per hour (training costs of $100M+ are basically meaningless per query). Pretty profitable at $20/month.”” / X https://x.com/emollick/status/1933020534498865574
OpenAI Is Phasing Out Its Work With Scale AI After Meta Deal – Bloomberg https://www.bloomberg.com/news/articles/2025-06-18/openai-is-phasing-out-its-work-with-scale-ai-after-meta-deal?embedded-checkout=true
Moonshot AI launched Kimi-Dev-72B, a new open-source coding model for software engineering tasks It achieves SOTA results on SWE-bench Verified software tasks, surpassing open-source rivals like DeepSeek R1, V3, and Devstral https://x.com/rowancheung/status/1934881573490331768
Thrilled to introduce Kimi-Dev-72B, our new open-source coding LLM for software engineering tasks. Kimi-Dev-72B achieves 60.4% resolve rate on SWE-bench Verified, setting a new SoTA result among open-source models. (1/5) https://x.com/yang_zonghan/status/1934652763985838585
UPDATE: This video was from last Saturday – robot speed was 4.05 seconds/package Yesterday, I saw it running at 3.54 seconds/package That’s a 13% speed-up in just 6 days 🤯 https://x.com/adcock_brett/status/1933970257028530514




