About This Week’s Covers
This week’s cover reflects how AI vision empowers robots to perceive and interact with everyday environments. As language models move from being “chatbots” to include vision and object recognition, it will force the general public away (hopefully) from thinking of AI as simply a writing tool. The prompt was generated by Claude 3.5 Sonnett and the image is from MidJourney.
The rest of the covers were created using Claude 3.5 and the Ideogram API, with the theme ‘an embodied robot Christmas + category name.’ The results were very creative. The AGI Star of Bethlehem over the city is pretty nuts. Also robot carolers for audio, programming a complex light strand for Perplexity (taking the name literally), and autonomous sleighs!

This Week’s Executive Summaries
OpenAI’s o3 Is Freaking Everyone Out – Joy/Doom Is Flooding Industry Chatter
First, it’s important to explain three tests you’re going to see in this summary, ARC-AGI, Frontier Math, and SWE-Bench:
ARC-AGI and EpochAI’s Frontier Math are tools designed to measure how advanced AI systems are becoming. ARC-AGI tests an AI’s ability to think like humans in creative and problem-solving tasks, helping researchers understand if it’s approaching “general intelligence.” Frontier Math, on the other hand, evaluates how well AI handles complex math and logic problems. SWE-Bench is a benchmark designed to evaluate how well AI systems can handle real-world software engineering tasks. Think of it as a way to see if AI can not just write code, but also debug, optimize, and handle complex programming projects like a professional developer. All three tests are like report cards for AI, showing how close it is to mastering human-like thinking.
Put on your seatbelt.
OpenAI’s o3 System Sets Record on ARC-AGI Benchmark
OpenAI’s o3 model has achieved a groundbreaking 75.7% score on the ARC-AGI-1 Semi-Private Evaluation set, working within the $10,000 compute limit of the public leaderboard. A more resource-intensive configuration (172x compute) pushed the score to an impressive 87.5%. This leap showcases unprecedented task adaptation capabilities, far surpassing previous GPT-family performance. For comparison, progress from GPT-3 to GPT-4o took four years to achieve just 5% on the same benchmark.
Arcprize | adcock_brett | emollick
OpenAI’s o3 Breaks New Ground in Advanced Math Benchmark
OpenAI’s o3 achieved a remarkable 25% success rate on the newly released FrontierMath benchmark, far surpassing GPT-4’s sub-2% performance. Designed by over 60 expert mathematicians, FrontierMath features problems requiring deep theoretical understanding and hours of expert effort to solve. This result underscores a significant breakthrough for o3, as it tackles challenges across advanced fields like number theory, algebraic topology, and combinatorics. The achievement highlights the model’s ability to move beyond pattern recognition and engage with complex, abstract reasoning.
Rohanpaul_ai | rowancheung |
SWE-Bench and Codeforce
“OpenAI o3 does 71.7% on SWE-Bench verified!! And 2727 on Codeforces! Unbelievable.”
deedydas
“OpenAI o3 is 2727 on Codeforces which is equivalent to the #175 best human competitive coder on the planet. This is an absolutely superhuman result for AI and technology at large.”
https://twitter.com/deedydas/status/1870175212328608232
“And that’s why OpenAI CFO Sarah Friar was saying a few days back about a possible $2000/month subscription. Thousands of Companies will still save significant payroll vs hiring a Graduate Engineer.”
https://twitter.com/rohanpaul_ai/status/1870760234228347017
Two charts that will get your attention…


OpenAI trained o1 and o3 to ‘think’ about its safety policy | TechCrunch
Whenever you see “safety” alongside OpenAI, you know it’s important, but it usually implies a ham-fisted Nerfed user experience. OpenAI’s new o1 and o3 models were designed to reference the company’s safety policies in real time, using a method called “deliberative alignment.” This approach allows the AI to break down prompts, consult safety guidelines, and refuse unsafe requests—like advice on illegal activities—while still answering appropriate ones (in theory). Instead of relying on human-written data, OpenAI trained the new models with AI-generated synthetic examples, making the process faster and scalable. As models like o3 become smarter, the risks of misalignment grow, with the potential for serious harm if they act irresponsibly.
techcrunch
AI Reshaping Job Values: Brain Work vs. Hands-On Skills
Years ago, Sam Altman predicted a major shift in job value as AI grows more capable. Roles involving computer-based tasks like coding, design, and writing will become less valuable. Meanwhile, hands-on jobs like plumbing, surgery, and logistics, which require physical presence and skills, will gain value because AI struggles with physical tasks. The flip challenges the traditional view that “brain power” jobs are more prestigious. Instead, there’s growing recognition of the importance of practical, tangible skills. Beyond economics, this affects how people will find meaning in work, with more emphasis on creating and fixing in the physical world. That said, I’m not sure Altman saw how quickly embodied robots would come on the heels of AI innovation.
Rohanpaul_ai
Sriram Krishnan Named Senior Policy Advisor for AI in Trump Administration
President-elect Donald Trump has appointed Sriram Krishnan, former general partner at Andreessen Horowitz, as senior policy advisor for AI in the White House Office of Science and Technology Policy. Krishnan will lead efforts to shape AI policy across government, collaborating with the president’s science and technology council as well as ex-PayPal COO David Sacks, Trump’s crypto and AI “czar.” A seasoned tech executive with experience at Microsoft, Facebook, and Twitter, Krishnan gained prominence as a podcast host and ally of Elon Musk. When it comes to concerns about AI’s impact on the internet, he advocated for tech-based solutions rather than publisher alliances and regulations. It’s easy to see why the tech folks want him.
techcrunch
Claude Can Handle Huge Excel Files
“Claude can now analyze large Excel files (up to 30MB) that would typically exceed its context window”
Alexalbert__
AI Visuals and Charts: Week Ending 12/27/2024
Microsoft v. The World
Microsoft acquires twice as many Nvidia AI chips as tech rivals
https://archive.md/vfMmV
AI Video Examples To Be Sure You Know How Good It’s Getting
“Google Veo prompt: “British chef slices and reveals a Beef Wellington”
The Heist – YouTube
“veo 2 is very good with photorealistic animals –pretty good consistency across shots; but not perfect. after the cut notice this: the headphones & glasses and drink stay the same — but the cocktail umbrella changes color from red to green.
“Pika 2.0 is pretty lit
“Slow down! Playing with the latest Kling 1.6
https://x.com/jamesyeung18/status/1869738357586358356
Top 37 Links of The Week – Organized by Category
Agents and Copilots: AI News Week Ending 12/27/2024
“I’ve replaced 100% of my marketing & sales dept with AI in 2024. It’s literally just me + AI Agents, AI wrappers & AI workflows now. My goal for 2025 is to replace 90% of support, operations, and the rest. I’m not alone 🧵 :
“Announcing Our $3.8M Seed Round to Make Voice AI Agents Reliable 🎙 If 2024 was the year of demos, 2025 is the year of reliable voice agents. We’re excited to share that @HammingAI has raised $3.8M in seed funding led by @MischiefVC, with participation from @ycombinator,
“We’ve raised $34m to accelerate our goal of equipping every accountant with a team of AI agents. Led by Keith Rabois at Khosla Ventures with Nat Friedman & Daniel Gross and continued participation from Better Tomorrow Ventures, BoxGroup, and Avid Ventures. Excited to also
Anthropic: AI News Week Ending 12/27/2024
Building effective agents \ Anthropic
Apple: AI News Week Ending 12/27/2024
“Big news from Apple! Their researchers have unveiled the open-source ReDrafter method, supercharging Nvidia GPUs for a whopping 2.7x faster token generation. This could revolutionize model creation for Apple Intelligence. Get the scoop here:
Apple reportedly developing AI server chip with Broadcom | TechCrunch
Augmented and Virtual Reality (AR/VR): AI News Week Ending 12/27/2024
New physics sim trains robots 430,000 times faster than reality – Ars Technica
Robot Training
“We’re excited to announce the official release of our Genesis Simulator!”
Genesis — Genesis 0.2.0 documentation
Autonomous Vehicles: AI News Week Ending 12/27/2024
Driv3R: Learning Dense 4D Reconstruction for Autonomous Driving
Chips, Hardware, and Infrastructure: AI News Week Ending 12/27/2024
Apple reportedly developing AI server chip with Broadcom | TechCrunch
Consumer Products: AI News Week Ending 12/27/2024
“LangChain State of AI 2024 What LLMs are the most widely used today? What metrics are commonly used for evals? Are developers finding success in building agents? Our State of AI 2024 report shows where the AI ecosystem is headed, based on data from LangSmith. Key 5
“By 2030, all phones will go extinct. While everybody is watching AI, Google & Samsung quietly reinvented reality — making screens obsolete. What you need to know about Google’s AI glasses and how they’ll change the way we see, think, and live:
Ethics/Legal/Security: AI News Week Ending 12/27/2024
“Nice paper from @GoogleDeepMind When models share work, they accidentally share your secrets too. MoE models can leak user prompts through expert routing vulnerabilities in batched processing. Expert-Choice-Routing and Token dropping in MoE creates a backdoor to steal user
“This is real – the OpenAI website already has references to “O3 Mini Safety Testing Call” form
“Along with the o3 announcement OpenAI also dropped a cute little paper: “Deliberative Alignment: Reasoning Enables Safer Language Models” The paper introduces a new training approach for LLMs called “Deliberative Alignment.” This method teaches the model safety specifications
“if you are a safety researcher, please consider applying to help test o3-mini and o3. excited to get these out for general availability soon. extremely proud of all of openai for the work and ingenuity that went into creating these models; they are great.” / X
Google: AI News Week Ending 12/27/2024
Google’s new Trillium AI chip delivers 4x speed and powers Gemini 2.0 | VentureBeat
Google says new quantum chip is faster than most powerful supercomputer
Microsoft: AI News Week Ending 12/27/2024
“Something special is coming this holiday break 🎁 As of today, Vision is rolling out to our U.S. Copilot Pro subscribers on Windows—so if that includes you, keep your eyes peeled in Edge! And if you still have some Christmas sweater shopping to do…
Microsoft’s growing AI health ambitions | Semafor
OpenAI: AI News Week Ending 12/27/2024
“ChatGPT’s new Projects feature can organize your AI clutter | TechRadar”
“Breaking news! Alec Radford departs OpenAI! As one of their star researchers, he was first author on GPT, GPT-2, CLIP, and Whisper papers.
“OpenAI’s o3 is absolutely incredible, and I am so excited about it, but its nowhere near to AGI. “early data points suggest that the upcoming ARC-AGI-2 benchmark will still pose a significant challenge to o3, potentially reducing its score to under 30% even at high compute
“Independent evaluations of OpenAI’s o3 suggest that it passed benchmarks that were previously considered far out of reach for AI including achieving a score on ARC-AGI that was associated with actually achieving AGI (though the creators of the benchmark don’t think it o3 is AGI)” / X
“🧭 As o3 tackles ARC-AGI’s Semi-Private Evaluation, here’s what this benchmark is about: Created by @fchollet, it tests AI’s ability to solve puzzles it’s never seen before—using only basic human priors like geometry & counting. It strips away language & cultural knowledge.
“imo the improvements on FrontierMath are even more impressive than ARG-AGI. Jump from 2% to 25% Terence Tao said the dataset should “resist AIs for several years at least” and “These are extremely challenging. I think that in the near term basically the only way to solve them,” / X
“Today OpenAI announced o3, its next-gen reasoning model. We’ve worked with OpenAI to test it on ARC-AGI, and we believe it represents a significant breakthrough in getting AI to adapt to novel tasks. It scores 75.7% on the semi-private eval in low-compute mode (for $20 per task
“The thing about o3 is that my fellow academics were literally just starting to test o1 on hard problems in their field. It showed a lot or promise. Lots of interesting discussions about how to validate and incorporate it into academic work. Then o3 is announced & it is obsolete” / X
“Evals on 03 Model from OpenAI. More than 20% better than o1 models. On real-world software tasks.
ChatGPT now understands real-time video, seven months after OpenAI first demoed it | TechCrunch
“Big OpenAI personnel news w/ @erinkwoo: Alec Radford, the lead author of OpenAI’s original GPT paper, is leaving to pursue independent research.
o3: Stratospheric reasoning – Dr Alan D. Thompson – LifeArchitect.ai
(2) o3: The grand finale of AI in 2024 – by Nathan Lambert
Podcasts/YouTube/Op-Eds: AI News Week Ending 12/27/2024
o3 – wow – YouTube
Nobel Minds 2024 – YouTube
Twitter/X/Grok: AI News Week Ending 12/27/2024
“Announcing our Series C of $6B to accelerate our progress Investors participating include a16z, Blackrock, Fidelity, Kingdom Holdings, Lightspeed, MGX, Morgan Stanley, OIA, QIA, Sequoia Capital, Valor Equity Partners, Vy Capital, Nvidia, AMD, among others” / X





Leave a Reply