“”high-taste testers” feeling the agi https://x.com/bilawalsidhu/status/1891563147724530137

“AlphaMaze: Teaching a 1.5B LLM to think visually and solve ARC-AGI like puzzles! 🤯 Powered by DeepSeek R1 1.5B + GRPO All with Apache licensed checkpoints and dataset 🤗 https://x.com/reach_vb/status/1892999150255440012

“Current LLM interpretability methods are limited. This paper explores rewriting parts of an LLM using natural language for better understanding. Proposes a method to partially rewrite a Transformer layer with natural language. It uses sparse representations and LLMs to explain https://x.com/rohanpaul_ai/status/1891410736330502288

“So no wall so far (not that we were expecting an overall wall given the rise of test-time compute), and the moats appear to be the usual: speed of execution, good partnerships and ecosystems, CapEx (I suspect the labs also believe speed to AGI and flywheels from AI development)” / X https://x.com/emollick/status/1891717323112747262

“What Dance Would You Like to Perform with Unitree G1? With the upgraded algorithm, G1 can learn any dance. Leave a comment to tell us what dance you’d like to see!😘 #Unitree #AGI #EmbodiedAI #SpringFestivalGalaRobot #AI #Humanoid #Bipedal #WorldModel #Dance https://x.com/UnitreeRobotics/status/1890377207169548775

“trying GPT-4.5 has been much more of a “feel the AGI” moment among high-taste testers than i expected!” / X https://x.com/sama/status/1891533802779910471

“🚀 Day 0: Warming up for #OpenSourceWeek! We’re a tiny team @deepseek_ai exploring AGI. Starting next week, we’ll be open-sourcing 5 repos, sharing our small but sincere progress with full transparency. These humble building blocks in our online service have been documented,” / X https://x.com/deepseek_ai/status/1892786555494019098

“We are at the early days of Reasoners, and will see a lot of ability growth as people figure out what sorts of System 2 simulation is best & how to build good reasoning chains. I suspect computer scientists would learn a lot from cognitive science, social psychology, & education” / X https://x.com/emollick/status/1890863268446498992

AI cracks superbug problem in two days that took scientists years https://www.bbc.com/news/articles/clyz6e9edy3o

“Great example of the jagged frontier- even smart AIs have weaknesses at basic tasks. Reading clocks is hard because AI vision systems are crude & calendar facts require vision & math The best model, Gemini, gets only 22% of clocks rights, while o1 gets 80% of calendar questions. https://x.com/emollick/status/1890460620262134176

“Deep Research on Perplexity scores 21.1% on Humanity’s Last Exam, outperforming Gemini Thinking, o3-mini, o1, DeepSeek-R1, and other top models. We also have optimized Deep Research for speed. https://x.com/perplexity_ai/status/1890452359773405675

Flavio Adamo on X: “🚨 o3-mini crushed DeepSeek R1 🚨 “write a Python program that shows a ball bouncing inside a spinning hexagon. The ball should be affected by gravity and friction, and it must bounce off the rotating walls realistically” https://t.co/xEvPDzzbVk” / X
https://x.com/flavioAd/status/1885449107436679394

“It’s kind of funny that this entire RL wave is centered on reasoning. There are very few problems in the world where you get sparser rewards than when you’re just… thinking.” / X https://x.com/lateinteraction/status/1892685182374691304

“@nearcyan well there’s a difference between prompting with chain-of-thought (which people were doing mostly with base models i think) and RL on chain-of-thought which is what is becoming more popular recently” / X https://x.com/iScienceLuvr/status/1892863034865177038

“> every line shared becomes collective momentum that accelerates the journey. > No ivory towers – just pure garage-energy and community-driven innovation. You can just tell, R1 “hands” wrote this. As I’ve been saying: a uniquely cyborgist lab. https://x.com/teortaxesTex/status/1892793229092827619

“TLDR: “Artificial General Intelligence” is an ill-defined concept and we probably shouldn’t waste time talking about it, or speculating when it might arrive instead, let’s directly measure AI productivity per unit of human input https://x.com/jxmnop/status/1893002519409721565

“why did it take 2-4 years after “let’s think step by step” was found for chain of thought models to become common? (context in 2nd tweet)” / X https://x.com/nearcyan/status/1892861840033501603

“3 things that I’ve been using to help Deep Research along 1. Chat before you run the search: The more context you can add by chatting back and forth, the better. Humans are still really bad at putting all the context down first-try into a chatbox. Ask your coworkers.” / X https://x.com/hrishioa/status/1891350088527835137

One response to “AGI: AI News Week Ending 02/21/2025”

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading