Image created with gemini-2.5-flash-image with claude-sonnet-4-5-20250929. Image prompt: A grand English courtroom interior with stone walls and Gothic windows, featuring an ornate brass balance scale as the centerpiece where one side holds leather-bound legal tomes and the other side displays floating holographic AI performance metrics and test scores, warm judicial lighting, painted in the style of a classical oil painting with rich browns and golds.
Reasoning models (apparently without tool use) scored #1 (OpenAI) & tied for #2 (Google) in the International Collegiate Programming Contest Its been one year since reasoners were first announced, it is genuinely surprising how good they have gotten at hard problems, so quickly https://x.com/emollick/status/1968402884627697950
Last week, our reasoning models took part in the 2025 International Collegiate Programming Contest (ICPC), the world’s premier university-level programming competition. Our system solved all 12 out of 12 problems, a performance that would have placed first in the world (the best”” / X https://x.com/merettm/status/1968363783820353587
Our general-purpose reasoning models solved all 12 problems at the 2025 International Collegiate Programming Contest (ICPC) World Finals, the world’s top university programming competition which was enough for a 1st-place human ranking.”” / X https://x.com/OpenAI/status/1968368133024231902
1/n I’m really excited to share that our @OpenAI reasoning system got a perfect score of 12/12 during the 2025 ICPC World Finals, the premier collegiate programming competition where top university teams from around the world solve complex algorithmic problems. This would have https://x.com/MostafaRohani/status/1968360976379703569
Demis Hassabis: calling today’s chatbots “PhD intelligences” is nonsense. They can dazzle at a PhD level one moment and fail high school math the next. True AGI won’t make trivial mistakes. It will reason, adapt, and learn continuously. We’re still 5–10 years away. https://x.com/vitrupo/status/1966752552025792739
Google Gemini is the top free iPhone app https://9to5google.com/2025/09/13/gemini-top-free-apple-app-store/
Made it to no.1 in the App Store. Congrats to the @GeminiApp team for all their hard work, and this is just the start, so much more to come!”” / X https://x.com/demishassabis/status/1966931091346125026
(1/3) Thrilled to announce a new Gemini breakthrough! Building on our success at IMO this year, an advanced version of Gemini Deep Think achieved gold-medal level performance at the ICPC 2025 World Finals – one of the world’s leading competitive programming competitions.”” / X https://x.com/quocleix/status/1968361041487904855
(2/3) Our model solved 10 out of 12 problems to achieve gold medal level. We were able to achieve this through breakthroughs in parallel thoughts, multi-step reasoning, and novel reinforcement learning techniques. You can find Gemini’s solutions here: https://x.com/quocleix/status/1968361222849642929
An advanced version of Gemini 2.5 Deep Think has achieved gold-medal level performance at the ICPC 2025 – one of the world’s most prestigious programming contests. 🏅 Building on the model’s success in math at the IMO, this marks another historic milestone for advanced AI. 🧵 https://x.com/GoogleDeepMind/status/1968361776321323420
Incredible milestone: an advanced version of Gemini 2.5 Deep Think achieved gold-medal performance at the ICPC World Finals, a top global programming competition, solving an impressive 10/12 problems. Such a profound leap in abstract problem-solving – congrats to @googledeepmind!”” / X https://x.com/sundarpichai/status/1968365605851218328
AI has officially beaten me at the ICPC World Finals. It reminds me of a rare ICPC skill: being able to quickly read a teammate’s code and spot bugs. This skill takes years to train, and explains why AI often makes coding slower (see arXiv:2507.09089). No matter how strong AI”” / X https://x.com/ZeyuanAllenZhu/status/1968568919482089764
amazing to get all 12 problems correct!”” / X https://x.com/sama/status/1968474300026859561
ICPC is a very hard and meaningful challenge:”” / X https://x.com/gdb/status/1968415631906324792
perfect score on the 2025 ICPC programming competition from our latest reasoning system:”” / X https://x.com/gdb/status/1968404060001968429
🐻Qwen3-Next just dropped on Together AI 80B parameters, 3B activated. Two models: ⚡Thinking: Outperforms Gemini-2.5-Flash-Thinking on reasoning benchmarks 🧠Instruct: Matches 235B model performance on key tasks Available now via our API 🚀 https://x.com/togethercompute/status/1966932629078634543
Qwen3 Next 80B A3B Thinking outperforms higher-cost and closed models like Gemini 2.5 Flash Thinking on benchmarks, nearing Qwen’s flagship model quality at a fraction the size. We have it ready to deploy in our model library, running on @nvidia and the Baseten Inference Stack. https://x.com/basetenco/status/1967688601640288288
📢 @Alibaba_Qwen new open-source model Qwen3-Next-80B-A3B is making waves. With a hybrid architecture & strong long-context reasoning, it’s sparking intense debate in the Zhihu community🔥 🔧 Zhihu contributor toyama nao with evalution: TLDR: A new “”gatekeeper”” for open-source https://x.com/ZhihuFrontier/status/1966415278922989813
🚨 Top 10 Open Model Leaderboard Update New open models have entered the Text Arena, and the top 10 rankings by provider have shifted for September! 🔹Qwen-3-235b-a22b-instruct from @Alibaba_Qwen holds the crown at #1 🏆 🔹Longcat-flash-chat from @Meituan_LongCat makes a strong https://x.com/arena/status/1968705194868535749
Alibaba has released Qwen3 Next 80B: an open weights hybrid reasoning model that achieves DeepSeek V3.1-level intelligence with only 3B active parameters Key takeaways: 💡 Novel architecture: First model to introduce @Alibaba_Qwen’s ‘Qwen3-Next’ foundation models, with several https://x.com/ArtificialAnlys/status/1966523300781428788
The new open-source Qwen3-Next Instruct and Thinking models put state-of-the-art long-context reasoning into the hands of everyone. We collaborated with #opensource frameworks from SGLang (@lmsysorg) and @vllm_project to enable communities to deploy Qwen3-Next across the https://x.com/NVIDIAAIDev/status/1967575419638468667
(1/n) Scheming has been a key concern in AI safety for 20+ years. It’s when an AI acts aligned while hiding true goals. New OpenAI + Apollo research found scheming in every tested frontier model, though no harmful scheming has been seen in production traffic.”” / X https://x.com/woj_zaremba/status/1968360708808278470
Today we’re releasing research with @apolloaievals. In controlled tests, we found behaviors consistent with scheming in frontier models—and tested a way to reduce it. While we believe these behaviors aren’t causing serious harm today, this is a future risk we’re preparing”” / X https://x.com/OpenAI/status/1968361701784568200
This is significant progress, but we have more work to do. We’re advancing scheming research categories in our Preparedness Framework, renewing our collaboration with Apollo, and expanding our research team and scope. And because solving scheming will go beyond any single lab,”” / X https://x.com/OpenAI/status/1968361716770816398
This OpenAI update on anti-scheming is exceptionally good for an AIco, clearing an (extremely low) bar of “”Exhibiting some idea of some problems that might arise in scaling the work to ASI”” and “”Not immediately claiming to have fixed everything already.”” https://x.com/ESYudkowsky/status/1968388335354921351
IntrEx: The first dataset for engagement modeling in educational dialogues It provides sequence-level annotations for interestingness & expected interestingness in teacher-student chats, collected from over 100 second-language learners. https://x.com/HuggingPapers/status/1967562091570827588
Alibaba’s WebSailor-V2: SOTA open-source web agents arrive A groundbreaking framework, powered by synthetic data and a dual-environment RL pipeline, achieves state-of-the-art results on BrowseComp & HLE. It outperforms existing open-source models and closes the gap to https://x.com/HuggingPapers/status/1968346179894235444
1/7 We’re launching Tongyi DeepResearch, the first fully open-source Web Agent to achieve performance on par with OpenAI’s Deep Research with only 30B (Activated 3B) parameters! Tongyi DeepResearch agent demonstrates state-of-the-art results, scoring 32.9 on Humanity’s Last Exam, https://x.com/Ali_TongyiLab/status/1967988004179546451
🚀 Kimi K2 Official Turbo API — 50% OFF for 30 days Code faster, ship sooner. Try it now: https://x.com/Kimi_Moonshot/status/1967829577037910427
Our engineer wrote about the thinking and technical story behind Checkpoint Engine. 👉 https://x.com/Kimi_Moonshot/status/1967923416008462785
Excited to release a preview of Moondream 3. A 9B param, 2B active MoE vision language model that makes no compromises; offering state-of-the-art visual reasoning while still retaining an efficient and deployment-friendly form factor. https://x.com/vikhyatk/status/1968800178640429496
when chatgpt said moondream wasn’t a frontier model, i took it personally”” / X https://x.com/vikhyatk/status/1968811248381784167
OpenAI claims hallucinations persist because evaluations reward guessing and that GPT-5 is better calibrated. Do results from HAL support this conclusion? On AssistantBench, a general web search benchmark, GPT-5 has higher precision and lower guess rates than o3! https://x.com/PKirgis/status/1966547382033936577
OpenAI has finally fixed their SWEBench errors and we can now finally apples to apples compare their scores over the entire 500 sample set (the fact that it took this long says alot about how much they care about SWEBench internally and maybe there’s a lesson here) https://x.com/nrehiew_/status/1967781400528245221
OpenAI just revealed that they have an internal unreleased SWE-bench-style benchmark for large ‘refactoring’ PRs, like the one mentioned here that edits 3.5k lines across 232 files. Their new model gets 51% accuracy on this benchmark. Who wants to make a public version of this? https://x.com/OfirPress/status/1967652031704994131
OpenAI’s Models Are Getting Too Smart For Their Human Teachers — The Information https://www.theinformation.com/articles/openais-models-getting-smart-human-teachers
GPT-5 is the best model for code quality out there 2 years ago, we created the world’s hardest software design quiz. Only 5 questions, multiple choice. Yet only about 3% of software engineers get them. The average score is somewhere between 2 and 3. Supposedly brilliant models https://x.com/jimmykoppel/status/1968683689421701413
The jaggedness of AI remains even as models have rapidly come to exceed human abilities in many of the hardest timed math & science contests. Yet there is much less progress on good puns. True AGI would be figure out our limits in more than calculus (sorry, but also seriously). https://x.com/emollick/status/1968447706969329718
Hey Claude, ChatGPT, Gemini: “”I am time traveling back to the 75 BC Rome for one day. I can’t bring anything back. What is the one thing I could learn that would most advance today’s knowledge and what is one thing I could do there that would make me richest today”” Pretty good https://x.com/emollick/status/1967009330789589077
Evals now support native audio inputs and audio graders. Evaluate model audio responses, with no text transcription needed. Get started in the Cookbook guide: https://t.co/V8qD5XFNqt https://t.co/tZuaCYccnQ” / X
https://x.com/OpenAIDevs/status/1965923707085533368
Clementine just dropped an incredibly useful guide to evaluations in 2025 ✨ The key insight: we’re transitioning from testing knowledge retention to measuring practical problem-solving ability. Her framework spans core capabilities, integrated assistant tasks, adaptive”” / X https://x.com/joelniklaus/status/1968596729852231813
New SOTA on ARC-AGI – V1: 79.6%, $8.42/task – V2: 29.4%, $30.40/task Custom submissions by @jerber888 and @_eric_pang_ are now the best known solutions to ARC-AGI Both: * Are open source * Use Grok 4 * Implement program-synthesis outer loops with test-time adaptation https://x.com/arcprize/status/1967998885701538060
Fireworks passed ASIC speed! First time, GPU based inference crossed an ASIC provider. Benchmark credit to AA Model: GPT-OSS-120B Speed: 540 TPS Legend: Purple – Fireworks on B200; Orange – Groq https://x.com/lqiao/status/1967641702484807695
Last week we found an issue with SWE-Bench, allowing agents to cheat by looking at future commits. Instead of celebrating the SWE-Bench Devs for quickly fixing the issue and being transparent, the HN crowd is dunking on them and drawing wildly inaccurate conclusions about”” / X https://x.com/TacoCohen/status/1966421688846778561
ARC just published new #1 and #2 reproducible SOTA scores on our public leaderboard from @jerber888 and @_eric_pang_. And their code is now open source! My analysis below — includes suggestions for application layer AI and future research directions. New SOTA: – v1: 79.6%,”” / X https://x.com/mikeknoop/status/1967999305983381630
Frontier Models Struggle The results are revealing: even the most advanced LLMs achieve a task success rate below 60%. Performance degrades substantially as task difficulty increases, with the top model, GPT-5, scoring only 39.02% on hard tasks. https://x.com/omarsar0/status/1966525793586360384
GenExam: The first multidisciplinary text-to-image exam is now on Hugging Face This new benchmark challenges T2I models with 1,000 rigorous, exam-style prompts across 10 subjects. It comes with ground-truth images and detailed scoring for semantic correctness and visual https://x.com/HuggingPapers/status/1968527551703433595
Qwen3 Next 80B used ~100M tokens with reasoning and ~25M without reasoning to run the Artificial Analysis Intelligence Index, slightly less verbose than Qwen3 235B 2507 with reasoning, and similar to it without reasoning https://x.com/ArtificialAnlys/status/1966523306338893979
Check out Seedream 4 and Seedream 4 High Res in the Arena in Battle, Side by Side and Direct modes here: https://x.com/arena/status/1966673632069132770
There was something deeply satisfying about ImageNet. It had a well curated training set. A clearly defined testing protocol. A competition that rallied the best researchers. And a leaderboard that spawned ResNets and ViTs, and ultimately changed the field for good. Then NLP”” / X https://x.com/DrJimFan/status/1966877464598094334
With over 4.5k dedicated votes in the Text-to-Image modality, Seedream 4 ranks at #5. 🥇Gemini 2.5 Flash Image is tied with Image 4.0 Ultra Generate for #1. 🥉GPT-Image-1 and Image 4.0 Generate Preview rank tied for #3. Check out the leaderboard details for Image Edit and https://x.com/arena/status/1966562486897029274
🚨 Leaderboard Update: With over 43k votes collected, the community has spoken! 🥈 Seedream 4 by ByteDance has landed at #2 on the Image Edit Leaderboard 🔸 It is also ranked #5 for Text-to-Image Real prompts and votes at scale illustrate sharper confidence intervals and more https://x.com/arena/status/1966562484506230922
🚨New Model update before the weekend 📣 By popular demand, we’ve added a “”High Res”” version of Seedream 4 that supports an output at 4096×4096 dimensions. We’ll see how this version of Seedream 4 stacks up vs. all the other top Image generation models soon. https://x.com/arena/status/1966673628327801255
Prompt for Nano Banana / Seedream: “Imagine what an entity sees that exists outside of time in a higher dimension that can concurrently visualize everything that has ever happened or will ever happen when looking at [insert point of interest]. Now generate that image projected https://x.com/bilawalsidhu/status/1966191138530013661
Our lightweight open-source eval library “”lighteval”” now ships with 7,000+ (!!) benchmarks baked in. Running it locally is literally a one-liner: >> lighteval vllm “”model_name=gpt2″” “”leaderboard|truthfulqa:mc|0″” (there is also a Python API for in/post-training evals ofc)”” / X https://x.com/Thom_Wolf/status/1967926861889163304
HunyuanImage 2.1 is the new leading open weights text to image model from @TencentHunyuan , surpassing HiDream-I1-Dev and Qwen-Image in the Artificial Analysis Image Arena! HunyuanImage 2.1 is the latest release from Tencent – a 17B DiT text-to-image model natively supporting https://x.com/ArtificialAnlys/status/1967800071115903358
First test of MLX batch generation PR on Mac Studio M3 Ultra 512GB with Qwen3-1.7B (4K ctx, 64 tokens) 🔥 Batch generation = WOW bf16 vs 4bit (avg of 3 runs) Batch of 1 → 127 vs 237 t/s 5 → 365 vs 515 t/s 10 → 556 vs 625 t/s 15 → 672 vs 617 t/s MLX vllm not a dream anymore! https://x.com/ivanfioravanti/status/1966903782400545196
LM Studio now supports Qwen3-Next with MLX on Mac! 🧵 https://x.com/lmstudio/status/1967985102845366280
Woah, 66 tok/s on a Macbook M4 Max 64GB with qwen3-next-80b-a3b-instruct-mlx@4bit, which uses about 41GB. Amazing job to the folks working on MLX, aware of at least these guys: @ivanfioravanti @ActuallyIsaak @awnihannun https://x.com/rwojo/status/1967767157250592899
Check out the actual speed (not yet the final version) of Qwen3-Next-80B-A3B-Instruct on Apple MLX! 🔥 4-bit: 67 TPS 8-bit: 58 TPS bf16: 48 TPS Movie normal speed, only waiting times removed. @awnihannun and @ActuallyIsaak did it and I bet there is still room for improvement 💪 https://x.com/ivanfioravanti/status/1966866942461177925
@Alibaba_Qwen Massive efficiency gains for long contexts. 262K context native, extensible to 1M+ tokens. Perfect for: ⚡ Repository-scale code analysis 🧠 Complex reasoning tasks 📄 Long document processing Both models available now → Instruct: https://x.com/togethercompute/status/1966933240683319556
This thread misses the point. Evals are not a scam. They’re misunderstood. Here’s how I think about evals ➡️ Logging is not evals. That’s like saying completing an exam is the same as getting graded on the exam. Imagine you take an exam. You start by quizzing yourself. You run”” / X https://x.com/rebeccatqian/status/1967758557174174027
Why Agents Fail The paper provides a fine-grained failure analysis, identifying seven common error types: ignoring requirements, overconfident self-solving, unproductive thinking, wrong tool selection, syntactic errors, semantic errors, and output parsing errors. Paper: https://x.com/omarsar0/status/1966525809302417436




