“we just used @Replit agent to build a browser agent 🤯 you might not want to give @OpenAI your $200 just yet… this cost $6.14 to build. here it is browsing @ProductHunt and making a list from a prompt… https://x.com/MakerThrive/status/1888998516250304900
OpenAI’s Operator agent helped me move, but I had to help it, too | TechCrunch https://techcrunch.com/2025/02/04/openais-operator-agent-helped-me-move-but-i-had-to-help-it-too/
“OpenAI introduced “deep research,” an AI-powered agent that searches the web and generates detailed research reports. Currently available exclusively to ChatGPT Pro users, the deep research tool uses OpenAI’s o3 model to process information, ask clarifying questions, and” / X https://x.com/DeepLearningAI/status/1890476096409194976
“”high-taste testers” feeling the agi https://x.com/bilawalsidhu/status/1891563147724530137
“trying GPT-4.5 has been much more of a “feel the AGI” moment among high-taste testers than i expected!” / X https://x.com/sama/status/1891533802779910471
“The new ChatGPT-4o upgrades, whatever they were, does make it seem to pull off some funny writing for the first time. Claude still feels more light and charming, but the 4o update is very noticeable. “the most disturbing yet absurd corporate memo you can come up with” “MORE” https://x.com/emollick/status/1891586403126919413
“GPQA: 448 multiple choice questions in 16 subdomains SuperGPQA: 26,529 mutiple choice questions across 285 graduate disciplines 😲 DeepSeek-R1 outperforms o1, o2-mini, Claude 3.5 Sonnet, etc. on this benchmark 🤔 https://x.com/iScienceLuvr/status/1892879645223375319
“LLMs are still incredibly bad at long context, severe drop in response quality from the best of the best models (o1, Claude, grok, DeepSeek), doesn’t really matter what model – it will choke” / X https://x.com/abacaj/status/1893024046469493212
“I can’t believe X users are so stupid. Not voting for o3-mini is insane. You can literally already distill 4o, Claude 3.5, Deepseekv3, etc into sizes that will run on phones.” / X https://x.com/dylan522p/status/1891682135255154775
“After the Grok-3 launch you have to consider xAI as a real competitor for SOTA models. Everything else is just cope. However, internally OpenAI, Anthropic and Google are likely ahead. But honestly not so sure about Google anymore, they need to drop a banger (Pro/Ultra with” / X https://x.com/scaling01/status/1891846484791820502
“i don’t recall ever caring about the algorithm powering google search i think we’re quickly arriving at the point where we also won’t care what model (gemini, claude, openai,…) powers our ai systems the better product or dev experience will win at the end of the day” / X https://x.com/omarsar0/status/1891570913327374496
“chatgpt 4o update 1 shot same prompt: I’d like to make a JS simulation of a sphere made up of ASCII numbers, rotating. The closest numbers should be pure white, and the farthest ones should fade to gray, on a black background. https://x.com/_akhaliq/status/1891249188366701011
OpenAI Rejects Elon Musk’s $97.4 Billion Bid for Control of the Company – The New York Times https://www.nytimes.com/2025/02/14/technology/openai-elon-musk.html
“400 million weekly active users on ChatGPT: https://x.com/gdb/status/1892749291233693849
“Today we’re launching SWE-Lancer—a new, more realistic benchmark to evaluate the coding performance of AI models. SWE-Lancer includes over 1,400 freelance software engineering tasks from Upwork, valued at $1 million USD total in real-world payouts. https://x.com/OpenAI/status/1891911123517018521
OpenAI tops 400 million users despite DeepSeek’s emergence https://www.cnbc.com/2025/02/20/openai-tops-400-million-users-despite-deepseeks-emergence.html
“AI NEWS: OpenAI’s former CTO Mira Murati just announced her secretive new venture to take on top AI labs in the world Plus, more news from OpenAI, Perplexity, and AI Pin company Humane, Meta, Stanford, Fiverr, and xAI. Here’s what you need to know:” / X https://x.com/rowancheung/status/1892132024124572024
“OpenAI announces SWE-Lancer Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? https://x.com/_akhaliq/status/1891721712296747126
Operator is now rolling out to Pro users in Australia, Brazil, Canada, India, Japan, Singapore, South Korea, the UK, and most places ChatGPT is available. Still working on making Operator available in the EU, Switzerland, Norway, Liechtenstein & Iceland—we’ll keep you updated! / X https://x.com/OpenAI/status/1892832374997631250
“Great example of the jagged frontier- even smart AIs have weaknesses at basic tasks. Reading clocks is hard because AI vision systems are crude & calendar facts require vision & math The best model, Gemini, gets only 22% of clocks rights, while o1 gets 80% of calendar questions. https://x.com/emollick/status/1890460620262134176
“Trying Deep Research with Grok 3. It shows promise, but not there yet. Generally accurate (I spotted a minor hallucination), but not as comprehensive as Google Deep Research nor close to as insightful as OpenAI’s Deep Research in the actual analysis of information. Early days. https://x.com/emollick/status/1892010991250018357
“In the Coding category, Grok-3 surpassed top reasoning models like o1 and Gemini-thinking. https://x.com/lmarena_ai/status/1891706272711381237
“LMAO, so Google just dropped infinite memory for Gemini before OpenAI did for ChatGPT. It can now recall past conversations , you can refer to something discussed a week ago. How long has OpenAI been working on this? 😆 Note: To ask Gemini to reference past chats, you need Gemini https://x.com/ai_for_success/status/1890377941579891003
“Here are the benchmark numbers: Grok 3 significantly outperforms other models in its category such as Gemini 2 Pro and GPT-4o. Even Grok-3 mini shows to be competitive. https://x.com/omarsar0/status/1891706611023938046
“TL;DR grok3 is fine and passes the vibe check of frontier level quality but its not better than R1 or o1-pro for me for most things i do. overall much better than i had expected, i put it in the gemini category but for me its still pretty far below the usefulness of R1 and the” / X https://x.com/_xjdr/status/1891911178147987513?s=46
“chatgpt 4o update this took a couple tries same prompt: Write a p5.js script that simulates 100 colorful balls bouncing inside a sphere. Each ball should leave behind a fading trail showing its recent path. The container sphere should rotate slowly. Make sure to implement https://x.com/_akhaliq/status/1891296414174490897
Flavio Adamo on X: “🚨 o3-mini crushed DeepSeek R1 🚨 “write a Python program that shows a ball bouncing inside a spinning hexagon. The ball should be affected by gravity and friction, and it must bounce off the rotating walls realistically” https://t.co/xEvPDzzbVk” / X
https://x.com/flavioAd/status/1885449107436679394
“Actually I quite like the new ChatGPT 4o personality, whatever they did. – it’s a lot more chill / conversational, feels a bit more like talking to a friend and a lot less like to your HR partner – now has a pinch of sassy, may defend itself e.g. when accused of lying – a lot of” / X https://x.com/karpathy/status/1891213379018400150
“i have no idea what we did but the new chatgpt is like cool now” / X https://x.com/mckbrando/status/1891280957568610454
“Dialing in that “found footage” vibe with OpenAI Sora. Adding a gentle VHS filter hides some of the artifacts while creating that retro aesthetic. https://x.com/bilawalsidhu/status/1891311402352099551
“OpenAI Deep Research is not built for prediction (people always ask about stock & demand forecasts) it is built for analysis: make an argument about a point of view and the evidence to support it. This is what lawyers, accountants, academics, analysts, and entrepreneurs do a lot.” / X https://x.com/emollick/status/1890063434022347257
“A new version of @OpenAI’s ChatGPT-4o is now live on Arena leaderboard! Currently tied for #1 in categories: 💠Overall 💠Creative Writing 💠Coding 💠Instruction Following 💠Longer Query 💠Multi-Turn This is a jump from #5 since the November update. Math continues to be an area https://x.com/lmarena_ai/status/1890477460380348916
“If reasoning models had not been invented at the end of 2023 / early 2024, the AI hype would have died this year after the GPT-5 release. What they call GPT-4.5, really is GPT-5. But it would have been a disappointing one without the o3 reasoning on top. The promise of scaling” / X https://x.com/scaling01/status/1892733059000148137
“Good point! Let’s do an apples-to-apples comparison. Here’s OpenAI’s Deep Research native: https://x.com/hrishioa/status/1891511609098387487
“It retrospect it is surprising that OpenAI released o1-preview. As soon as they showed off reasoning, everyone copied it immediately. They could have waited until full o1 (or o3), but I guess the advantage of continuing to be the leader outweighed the actual edge from the model.” / X https://x.com/emollick/status/1891929362074653049
“one slide that people are sleeping on the last 6 months of o1/o3/Agents ships has doubled @ChatGPTapp users and the path is now pretty darn clear to reach 1B weekly -active- users by end of 2025 (!!!) “”Products don’t really get that interesting to turn into businesses until https://x.com/swyx/status/1892982602199429455
“The new GPT4o is palpably better — a little heavy w/ emojis and bold text but damn does it feel a lot more intelligent and insightful than the LinkedIn MBA that was the previous version. Really good stuff OpenAI.” / X https://x.com/bilawalsidhu/status/1890942354644738289
“This was fun: “o1, build a simulator of a D&D guild hall. Persistent characters come in, get quests, interact with each other, leave & return, make it procedurally generated” I kept asking it to add other ideas (relationships, etc) 8 times, got no errors. Desire-based coding! https://x.com/emollick/status/1890215778529767795
“GPT-4o is now definitely different. Beyond that, hard to say. It seems “smarter” and is much more personable and less likely to refuse requests… but my Innovation GPT (with over 10k uses) now no longer works with GPT-4o. Now I have to check all of my deployed GPTs, I guess. https://x.com/emollick/status/1890856663738925545
OpenAI looking at 16 states for data center campuses tied to Stargate https://www.cnbc.com/2025/02/06/openai-looking-at-16-states-for-data-center-campuses-tied-to-stargate.html
“If the light blue part is best of N scores, this means that Grok 3 reasoning is inherently an ~o1 level model. This means the capabilities gap between OpenAI and xAI is ~9 months. Also what is the difference between “think” and “big brain” https://x.com/nrehiew_/status/1891710589115715847
“all openai users are high-taste testers 🥰🫵🫶💛” / X https://x.com/aidan_mclau/status/1892991117924184118
“Here’s the final results – built in 15 mnts: Rundown: – I used @CodeGuidedev Starter kit with @boltdotnew – generated PRD + app flow with CodeGuide – uploaded docs to Bolt – used 21st .dev for Ready-made UI components – OpenAI API (for report gen.) – Jina ai (for scraping) https://x.com/cj_zZZz/status/1888627087650795614
“Revisiting the Test-Time Scaling of o1-like Models Do they Truly Possess Test-Time Scaling Capabilities? https://x.com/_akhaliq/status/1892071152215785526
“Openator (open source version of @OpenAI operator) is traveling on the WebVoyager bench🤘 https://x.com/kevinpiac/status/1890133149050663177
“The Guardian Media Group inks a deal with OpenAI https://x.com/fdaudens/status/1890502321047568705
Reasoning best practices – OpenAI API https://platform.openai.com/docs/guides/reasoning-best-practices
“The new GPT-4o be like… “I’m just a chill guy” https://x.com/bilawalsidhu/status/1891261534245925346
“Perplexity just announced Deep Research (PDR)! I’m now testing and comparing it with OpenAI’s Deep Research (ODR). I still think the o3 variant powering ODR is a massive advantage. 20.5% (PDR) vs. 26.6% (ODR) on Humanity’s Last Exam. https://x.com/omarsar0/status/1890525249977872640
“Perplexity Deep Research is quite close to OpenAI o3 on the Humanity Last Exam Benchmark despite being an order of magnitude faster and cheaper. This is possible because DeepSeek is open source and cheap and fast. https://x.com/AravSrinivas/status/1890486069361025040
“SemiAnalysis is hosting Blackwell & low level GPU Hackathon 🚀 Hacking, prizes, and insights from top industry leaders like @cHHillee, @tri_dao, @marksaroufim, Phil Tillet from TogetherAI, GPUMode, OpenAI, Coreweave, Lambda, etc Limited spots, apply here! https://x.com/dylan522p/status/1893026079931277636
“Changing conspiracy theory beliefs is very hard, but a replicated finding shows a short chat with GPT-4 changes people’s belief in conspiracy theories for the long term. Why? Not tricks, it’s that AI provides relevant facts and evidence tailored to each person’s specific beliefs https://x.com/emollick/status/1892202893165425071
“In April 2024, a few words like “delve” had a viral moment, when researchers pointed out that GPT-4 used it often, and the spread of the word showed AI use was common in scientific papers. Did that result in a decrease in AI use? No, but people started avoiding the word “delve!”” / X https://x.com/emollick/status/1892055335197684065
“This is my o1-pro workflow guide for building apps from a starting template. It’s honestly *crazy* good. 4hr+ detailed 6-prompt workflow to maximize your ability to build with AI. Great for beginners & experienced devs. Full tutorial here tomorrow at 12pm PT. https://x.com/mckaywrigley/status/1890457528087044401
“Based on the announcement (& not using the model, yet): 1) X has caught up with the frontier of released models VERY quickly, if they continue to scale this fast, they are a major player 2) Grok 3 is closely following the OpenAI playbook 3) Not sure who will use API at this point” / X https://x.com/emollick/status/1891714787022639373
“Less is More for Reasoning (LIMO): a 32B model fine-tuned with 817 examples can beat o1-preview on math reasoning! 🤯 Do we really need o1’s huge RL procedure to see reasoning emerge? It seems not. Researchers from Shanghai Jiaotong University just demonstrated that carefully https://x.com/AymericRoucher/status/1891822202812760206
“xAI arrives at the frontier: Grok 3 is poised to be the world’s new leading model, likely only surpassed by OpenAI’s unreleased o3 model Key takeaways: ➤ Grok 3 is now the leading non-reasoning model, pushing pre-training to new limits ➤ Grok 3 Reasoning likely beats o3-mini https://x.com/ArtificialAnlys/status/1891853619907133702
“Come test out OpenAI’s o3-mini-high for yourself in the Arena at: https://x.com/lmarena_ai/status/1892979592727597259
“In my testing it was at least as good in thinking mode then o3-full deep research was, despite that not being listed here – Interesting to note that grok-3mini seems generally better than full, my guess is that this means they didnt distill full into mini like I assume OpenAI https://x.com/Teknium1/status/1891715974992408738
OpenAI now reveals more of its o3-mini model’s thought process | TechCrunch https://techcrunch.com/2025/02/06/openai-now-reveals-more-of-its-o3-mini-models-thought-process/
“@OpenAI’s o3-mini-high is now available in the Arena! We noticed general improvements over o3-mini, and some notable highlights: ⬆️ Significant jump over o3-mini (1332 vs 1306) 🏅 Tie ranked #1 in coding, math & hard prompts https://x.com/lmarena_ai/status/1892979590018277669
“Everyone vote for o3-mini type model to be open-sourced please 🥺🥺🥺 We can distill or quantize a phone sized model dw the open-source community will work its magic!!” / X https://x.com/iScienceLuvr/status/1891669417332805739
“joke aside it would be great to have o3-mini so it would be community distilling it instead” / X https://x.com/mervenoyann/status/1891772390301941796
“go vote for o3-mini. if you voted for the phone model, please explain yourself in the comments.” / X https://x.com/gallabytes/status/1891674566931497410
“.@giffmana internal propaganda is working 📿 (please vote for o3-mini)” / X https://x.com/eliebakouch/status/1891675065021853805
“Introducing o3-mini Deal Finder 🛒 Have AI search the web and find the best place to buy a product. Simply enter the product name, and it will search the web providing a review summary and the best deal. Powered by @firecrawl_dev and OpenAI’s o3-mini. https://x.com/ericciarla/status/1892257110047805660
“ChatGPT now has 400M weekly active users. Extra motivation to ship more great things for you all. What’s one thing you’d love to see in ChatGPT? Big or small. https://x.com/kevinweil/status/1892757921118703970
“I like this, but found o1 Pro (which wasn’t tested) did surprisingly well on the sample puzzles (and I do wonder how much the limitations of vision play in). https://x.com/emollick/status/1890199820956315971




