About This Week’s Covers
This week’s cover again represents the future of embodied AI, where robots learn physical skills through simulation and computer vision. Claude Sonnet 3.5 wrote the prompt and suggested a triptych format, which I’d never tried before. The first panel shows the physics simulation, the middle panel reveals the computer vision systems, and the final panel shows the robots gracefully skating on a winter lake, having successfully transferred their learning to the physical world.
I gave the prompt to MidJourney, FLUX, and Ideogram, and as one would expect, Ideogram was the best at handling the complex composition. Claude wrote the prompt “Triptych of robots learning to ice skate. Panel 1: Simple wireframe simulation with basic physics visualization. Panel 2: Computer vision view with segmentation colors and depth maps showing the rink, skaters, and motion paths. Panel 3: Photorealistic scene of robots gracefully skating on a frozen lake at sunset, their movements fluid and natural. Unified cool color palette with ice blues and winter purples. Each panel flows into the next showing the learning progression.” The font is Domyouji, the closest Adobe font to the NVIDIA font, Handel Gothic.
The rest of the covers were created using Claude 3.5 and the Ideogram API, with the theme “winter wonderland + category name.” We’re at the point where AI often does better with more room to be creative. But it’s important to set it up. My prompt to Claude was “I’d like to make the theme winter wonderland. The idea of a winter wonderland can be a variety of settings, places and concepts, based on what you think best captures the spirit of each specific category within that theme. Does that sound good?” Then I give it the category list and the rubric for the txt file that powers the Ideogram API via a Python script.
Note the audio wavelength on the frosted window glass, the conference table made of ice, and the magical printing press making flying books of snow, below:


This Week’s Executive Summaries
OpenAI’s New Model Starts to Surpass Physicians in Diagnostic Reasoning
OpenAI’s new AI model, o1-preview, achieves superhuman performance in medical diagnostic tasks, according to a recent paper. In a test using challenging cases from the New England Journal of Medicine’s Clinical Pathological Conferences, o1-preview successfully diagnosed 80% of the hardest cases compared to 30% by doctors. The model excels at differential diagnosis, clinical reasoning, and treatment planning, vastly outperforming both previous AI versions and human physicians. Unlike multiple-choice tests, these cases required complex reasoning based on detailed patient vignettes. For example, when evaluating phosphate-wasting conditions, o1-preview proposed a comprehensive testing plan to systematically rule out potential causes. In another case of unexplained hyperammonemia, it recommended a prioritized and thorough sequence of tests, emphasizing common causes first. To paraphrase VC Deedy Das, it’s a bad idea to not also consult an AI model, just to see what it says.
I give AI all of my lab tests and it does a great job explaining each one to me in layman’s terms. The sheer amount of time and patience for my general practitioner to explain it would be over an hour. And to be honest, I don’t think a local family doctor is going to know the intricacies of the minor tests as well as AI. It’s simply not a fair challenge. As AI starts to hallucinate less than people (who let’s face it also hallucinate and space out)… it’s going to win more and more often.
Emollick | deedydas | arxiv
Meta Releases Apollo: Advancing Video Understanding in Large Multimodal Models
It’s definitely time to stop thinking of AI as chatbots. They process and understand video and audio now as much as text. Meta introduced Apollo, a family of video-focused Large Multimodal Models (LMMs) designed to tackle the challenges of video understanding. For the nerds in the room, Apollo leverages insights from a comprehensive study on the computational demands and design principles of video-LMMs, including a concept called “Scaling Consistency,” which ensures design choices from smaller models scale effectively to larger ones. Notably, Apollo-3B surpasses many 7B models with a score of 55.1 on LongVideoBench, while Apollo-7B leads its category with a 70.9 on MLVU and 63.3 on Video-MME.
_akhaliq | huggingface
Simulations Are Rapidly Transforming Robotics – It’s the ONE TREND to Watch
Advanced AI-driven simulations are reshaping robotics by enabling rapid, scalable training (aka 430,000x faster than real life). In simulations, robots can practice millions of skills across billions of scenarios, using computing power to exponentially accelerate learning compared to linear human-led training (the Matrix on crack). Three key trends are emerging: massive parallelization on GPUs for high-speed simulation, automated generative pipelines for building realistic 3D environments, and neural network-based simulators for dynamic and adaptive learning. Projects like Genesis embody these advancements, combining open-source accessibility with cutting-edge tools for sim-to-real training.
Read this entire Tweet from Dr. Jim Fan at NVIDEA:
“If an AI can control 1,000 robots to perform 1 million skills in 1 billion different simulations…”
https://x.com/DrJimFan/status/1869795912597549137
Genesis: Simulating Real-World Physics For Robot Training
Genesis helps simulate how objects move, interact, and behave in the real world. It’s built for tasks like teaching robots new skills, studying how people or animals move, and creating lifelike animations. The platform combines three tools in one: a physics simulator to model physical behavior, a rendering system to make realistic visuals, and a data generator that creates videos, motions, and other resources based on simple instructions. What makes Genesis stand out is its speed—it can simulate robotic actions 430,000 times faster than real life.
Genesis
Watch this Genesis demo video- it’s insane: https://x.com/zhou_xian_/status/1869511650782658846
AI Training Data Concentrates Power in Big Tech’s Hands
A new study by the Data Provenance Initiative reveals that the sources of data used to train AI models are increasingly concentrated, favoring large technology companies. Over 70% of video and image training data comes from YouTube, owned by Google, and 90% of analyzed datasets originate in North America and Europe. Exclusive data-sharing deals between major tech players and platforms like Reddit further exacerbate unequal access, sidelining smaller companies and researchers. This trend raises concerns about monopolization and bias, as much of the data reflects the priorities of profit-driven corporations rather than organic human experience. Researchers also highlighted challenges in data transparency, as many datasets lack clear provenance or carry restrictive licenses that complicate their use. The shift toward indiscriminate web scraping since 2018 has widened the gap between curated and unstructured data, with implications for the fairness and inclusivity of AI systems. Worth a read, at the links below.
Technologyreview | fdaudens
AI Academic and Research Fraud Is Potentially a Huge Problem
“Researchers used AI to generate 288 complete academic finance papers predicting stock returns, complete with plausible theoretical frameworks & citations. Each paper looks legit. They did this to show how easy it now is to mass produce ‘credible’ research. Academia isn’t ready.”
emollick
OpenAI Takes The Lead In Coding – Dethrones Claude?
“The new o1 model from 12/17 is #1 on Livebench AI and scores 91.58 in Reasoning. OpenAI also beats Sonnet in Coding.” Personally, I’ll need to see it to believe it.
bindureddy
Google Releases Veo 2, Claiming It Outperforms OpenAI’s Sora in AI Video (I agree)
Google has introduced Veo 2, the latest version of its AI video generation model, touting improvements in consistency, realism, physics, and human expression. Veo 2 allows users to specify genres, cinematic effects, and lens styles, to create 4K resolution videos. Early internal tests show Veo 2 outperforms OpenAI’s Sora in user preference and prompt adherence, despite occasional artifacts like extra fingers (nobody’s perfect – but see the examples, below!) The world of AI video generation faces serious competition. OpenAI’s Sora, RunwayML’s Gen-3 Alpha Turbo, Kling, and Pika are all crushing it with new features (albeit in different ways).
Deepmind | rowancheung | venturebeat
“A pair of hands skillfully slicing a ripe tomato on a wooden cutting board” #veo
https://twitter.com/agrimgupta92/status/1868745017571131582
“Veo 2 prompt: “Eating soup like they do in Europe, the old fashioned way”
https://twitter.com/bilawalsidhu/status/1869182358454505959
Databricks: The $62 Billion Company You’ve Probably Never Heard Of…
Databricks Inc. is raising $10 billion in funding, bringing its valuation to $62 billion and reinforcing its status as one of the most valuable private tech companies globally. Databricks is known for its cloud hosting and software, and has seen rapid growth, with revenue expected to surpass $3 billion annually by January 2025. Its Databricks SQL product, which competes with Snowflake, achieved a $600 million revenue run rate, growing over 150% annually. While speculation about an IPO continues, the company’s CEO, Ali Ghodsi, suggests this funding offers flexibility, enabling the firm to remain private for now while rewarding employees with liquidity.
yahoo
ChatGPT search is available to all Free users (I’m not a fan)
“Search the web in a faster, better way—available globally on http://chatgpt.com and our mobile and desktop apps for all logged-in users.” I tried it a few times, and the bottom line is Google’s search has so many strong widgets that while I hate the search results on Google, I do love the weather, maps, and little integrated helpers (like a stopwatch)… too much to quit for a new default. BUT… the more I use regular GPT and Claude and the more they get what I need… the less I go to Google. It’s like a scale tipping rather than a switch.
OpenAI | youtube
Follow-up to Last Week: Meta Revolutionizes AI with Token-Free Transformers (nerdy but important)
Meta has unveiled a groundbreaking approach to training transformers, eliminating the need for tokens and operating directly on raw byte sequences. Their method uses Dynamic Patching, where a local model organizes bytes into variable-sized patches—large for simple text, small for complex text—based on entropy. These patches are then processed by a transformer, bypassing traditional tokenization. The approach ends with a small decoder that reconstructs byte sequences from the patches. This innovation reduces tokenization errors, enhances model robustness, and provides a scalable path forward for large language models.
LiorOnAI
AI Visuals and Charts: Week Ending 12/20/2024
“Tesla demonstrated Optimus walking across mulched, uneven ground, blind without vision Looks like they used neural nets processing other sensors to maintain balance
Top 49 Links of The Week – Organized by Category
Agents and Copilots
Devin
“Cognition Labs officially launched Devin, its AI agent coding assistant Devin integrates directly with development workflows through Slack, GitHub, and IDE extensions (beta) It’s starting at $500/month for unlimited team access
Gemini Code Assist adds tools to aid developer workflows | Google Cloud Blog
Augmented and Virtual Reality (AR/VR)
“Check out this Stereo4D paper from Google DeepMind. It’s a pretty clever approach to a persistent problem in computer vision — getting good training data for how things move in 3D. The key insight is using VR180 videos — those stereo fisheye videos we launched back in 2017 for
Robot Training
Generative Video WorldSim, Diffusion, Vision, Reinforcement Learning and Robotics — ICML 2024 Part 1
World Models
“Ed Catmull (Pixar founder) just joined Odyssey’s board as they reveal their first 3D breakthrough: Instantly turning any 2D image into an editable 3D world. World Labs isn’t alone anymore. Here’s the TL;DR without the hype 🧵
Ethics/Legal/Security
“Lockheed forms subsidiary to help defense companies adopt AI | Reuters”
AI data centers could be built on federal lands under Biden draft order – The Washington Post
SoftBank CEO and Trump announce $100 billion U.S. investment
“Breaking news from Chatbot Arena⚡🤔 @GoogleDeepMind’s Gemini-2.0-Flash-Thinking debuts as #1 across ALL categories! The leap from Gemini-2.0-Flash: – Overall: #3 → #1 – Overall (Style Control): #4 → #1 – Math: #2 → #1 – Creative Writing: #2 → #1 – Hard Prompts: #1 → #1
Imagery
“Topaz really cooked with their new upscaling model called “redefine” — basically every CSI “enhance” meme you’ve seen IRL. Settings: – 4x Upscale – Creativity: 2 – Texture: 3 – No prompt It’s basically the Topaz take on the magnific style of “creative upscaling” where you use”
Grok Image Generation Release
“IT’S FINALLY HERE! 🔥 Magnific’s Super Real 🔥 The most amazing state-of-the-art generator for REALISTIC images specially designed for professionals (architecture, interior design, films, photography, etc). You have never seen a level of realism like this ✨ Info & prompts 👇
MidJourney
“Midjourney launched ‘Patchwork,’ a collaborative multiplayer world-building tool It’s essentially a storyboarding tool that will be expanded to populate worlds eventually
“Today we’re releasing “Moodboards” which let you personalize our models using collections of images. We’ve also added support for multiple personalization profiles and made image ranking 5x faster. It’s finally time to use custom Midjourney models for all your projects <3
Locally Run Models
“We released Gemini 2.0 Flash Thinking today! ⚡️🤔 It’s a small step towards improved reasoning via inference-time compute, built on top of our small and mighty 2.0 Flash!” / X
“Just 10 days after o1’s public debut, we’re thrilled to unveil the open-source version of the groundbreaking technique behind its success: scaling test-time compute 🧠💡 By giving models more “time to think,” LLaMA 1B outperforms LLaMA 8B in math—beating a model 8x its size.
Surrey announces world’s first AI model for near-instant image creation on consumer-grade hardware | University of Surrey
Multimodality
“We launched a new, interactive video analyzer in @googleaistudio that you can use to test Gemini 2.0 for video understanding! 🎉 It parses the model’s timecode results to make it way easier to explore results. Try it here:
“Check out this Stereo4D paper from Google DeepMind. It’s a pretty clever approach to a persistent problem in computer vision — getting good training data for how things move in 3D. The key insight is using VR180 videos — those stereo fisheye videos we launched back in 2017 for
Segmentation
“Excited to share StableAnimator, an open-source human image animation model that excels in identity preservation! It transforms reference images into stunning, pose-driven animations with remarkable fidelity.
OpenAI
Hundreds of OpenAI’s current and ex-employees are about to get a huge payday by cashing out up to $10 million each in a private stock sale | Fortune
“LiveBench Coding results for GPT-4o, o1-preview and o1 Sonnet 3.5 v2 is at 0.67 o1 looks like it’s at 0.76 !!
What David Sacks as AI czar (with Elon Musk as wingman) could mean for OpenAI | Fortune
“o1-preview is far superior to doctors on reasoning tasks and it’s not even close, according to OpenAI’s latest paper. AI does ~80% vs ~30% on the 143 hard NEJM CPC diagnoses. It’s dangerous now to trust your doctor and NOT consult an AI model. Here are some actual tasks: 1/5
Open Source
Meta/Llama
Meta launches Llama 3.3, shrinking powerful 405B open model | VentureBeat
“As we wrap up 2024, we’re sharing an update on our progress with Llama and the impact it’s having around the world. Read the full update here ➡️
Perplexity
AI Startup Perplexity Closes Funding Round at $9 Billion Value
Publishing
“Reddit tests a conversational AI search tool | TechCrunch”
Partnering with Creative Artists Agency on responsible AI tools for talent – YouTube Blog
YouTube is letting creators opt in to allowing third-party AI training – The Verge
(RAG) Retrieval-Augmented Generation
“What is the best LLM for RAG and grounded long-form inputs? @GoogleDeepMind DeepMind FACTS is a new benchmark for how well LLMs can generate factually accurate, long-form responses while remaining faithful to provided source documents. 👀 TL;DR; 📊 1,719 examples (860 public,
“Notion, the much-loved workspace platform, is redefining productivity with Notion AI and Cohere Rerank. Every search and Notion AI interaction now runs through Rerank for fast and relevant results. Rerank delivers: – Enhanced search accuracy – Unmatched cost efficiency – A
Robotics and Embodiment
“New paper presents Exbody2, a whole-body tracking framework enabling humanoid robots to mimic dynamic, human-like motions (e.g., running, dancing) with stability. Trained via RL in simulation and transferred to real robots, it outperforms prior methods.
Genesis
Generative Video WorldSim, Diffusion, Vision, Reinforcement Learning and Robotics — ICML 2024 Part 1
Figure
“<24 hours ago we unboxed our first commercial humanoid robot shipment And today F.02 robots are successfully performing the end-to-end use case at our client site The transfer rate we saw from HQ to commercial site was incredible” / X
“Update: All robots are in good condition and walking around This is our 2nd generation humanoid robot, Figure-02, designed to be a universal interface to the physical world We built a human world and the interface will have a human form” / X
“It’s official: F.02 humanoid robots have arrived at our commercial customer The robots are connecting to the network and performing pre-checks this morning Our time from filing the C-Corp to shipping commercially was 31 months
“Exciting news – today, Figure officially became a revenue-generating company! This week, we delivered F.02 humanoid robots to our commercial client, & they’re currently hard at work It marks 31 months from filing our C-Corp to getting to humanoid robot revenue
Science and Medicine
“Everything you love about generative models — now powered by real physics! Announcing the Genesis project — after a 24-month large-scale research collaboration involving over 20 research labs — a generative physics engine able to generate 4D dynamical worlds powered by a physics
Study claims AI could boost detection of breast cancer by 21% | TechCrunch
Video News
Google (Veo)
“BREAKING: Google just dropped Veo 2 and Imagen 3 — their next gen video and image generation models. Turns out Google’s been closing the gap quietly — not just on LLMs, but on visual creation too. Here’s everything you need to know w/o the hype 🧵
“Today, we’re announcing Veo 2: our state-of-the-art video generation model which produces realistic, high-quality clips from text or image prompts. 🎥 We’re also releasing an improved version of our text-to-image model, Imagen 3 – available to use in ImageFX through
Pika
“Pika Labs launched Pika 2.0 It’s the startup’s new image-to-video model that lets you combine characters, objects, clothes, and locations in an AI-generated video
“Our holiday gift to you: Pika 2.0 is here. Not just for pros. For actual people. (Even Europeans!) Now available at
Runway
Runway Talent Networkhttps://talent.runwayml.com/





Leave a Reply