Image created with gemini-2.5-flash-image with claude-sonnet-4-5. Image prompt: Cinematic wide shot of an ornate art nouveau rocket launching from an Emerald City spire at dusk, exhaust trail and rocket components highlighted with glowing neon green and magenta object segmentation outlines, moody dramatic lighting with witch hat silhouettes embossed on the hull, Wicked-inspired color palette of emerald and rose gold
Gemini 3 Pro Preview has comparable speeds to Gemini 2.5 Pro, with 128 output tokens per second. This places it ahead of other frontier models including GPT-5.1 (high), Kimi K2 Thinking and Grok 4 https://x.com/ArtificialAnlys/status/1990813128226189811
The Artificial Analysis leaderboard shows Gemini 3 at 73%, GPT-5.1 at 70%, and Kimi at 67% – minor differences. On our leaderboard, Gemini is 47%, GPT-5.1 is 38%, and Kimi is 27% – Gemini 3 is substantially more capable on hard benchmarks. https://x.com/hendrycks/status/1991188104804208736
We estimate that Kimi K2 Thinking has a 50%-time-horizon of around 54 minutes (95% confidence interval of 25 to 100 minutes) on our agentic SWE tasks. Note that we conducted this evaluation through a third-party inference provider, which reduces our confidence in this estimate. https://x.com/METR_Evals/status/1991658241932292537
Kimi K2 Thinking is impressive. So I built a multi-agent deep researcher, Kimi Deep Researcher. It generates long research reports on any topic, powered by subagents (web searcher, analyzer, and synthesizer). It can do 100s of tool calls per session. Repo soon! https://x.com/omarsar0/status/1988974710592516454
🤗 Kimi-k2-Thinking has reached top performance on the latest IMO-level reasoning benchmark, AMO-Bench from Meituan Longcat!”” / X https://x.com/Kimi_Moonshot/status/1991139250566545886
Kimi-K2 Thinking gets the same score on METR as Claude 3.7 Sonnet as I was saying, open-source is 9 months behind frontier labs on agentic, long-context reasoning tasks it’s still an improvement and open-source models seem to be on their own exponential, but I heavily suspect https://x.com/scaling01/status/1991665386513748172
Open-source research agents have been lagging behind proprietary systems like OpenAI’s Deep Research. The gap has been frustrating for developers who want powerful, deep research agents without vendor lock-in. I’ve been building my own called Kimi Deep Researcher. Similarly, https://x.com/omarsar0/status/1990794651608219727
Perplexity Pro and Max subscribers now have access to Kimi-K2 Thinking and Gemini 3 Pro. https://x.com/perplexity_ai/status/1991614227950498236
🚨Leaderboard Update New model provider in the Arena: @DeepCogito has released Cogito v2.1 (MIT licensed) 🔹Top 10 Open Source Model for WebDev, rank #10 🔹Tie ranks #18 overall for WebDev This puts Cogito v2.1 on par with community favorites like Qwen 3 Coder Plus & Kimi K2 https://x.com/arena/status/1991211903331496351
As promised, Kimi K2 and Gemini 3 Pro are available for all Perplexity Pro and Max users. Grok 4.1 will be available soon.”” / X https://x.com/AravSrinivas/status/1991619527638151665





Leave a Reply