About This Week’s Covers

For the second week, I used the Ideogram API and a text file to auto-generate 33 cover images in various styles, downloading them directly to my local drive. This saves me an hour per week and helps me catch up, as I’m six weeks behind. The downside is this there are no retakes or multiple images to choose from. It’s also less creative, but I plan to improve the process once I’m caught up.

This week, I ran the script to create covers inspired by Jean-Michel Basquiat. The results weren’t great, but a few examples are below. For the main cover, I gave MidJourney, Ideogram, and FLUX the same prompt: “a nuclear power plant scrawled in the style of Jean-Michel Basquiat. The cooling tower is wearing one of Basquiat’s iconic yellow crowns.” While MidJourney and Ideagram performed well, FLUX won the test. Examples of the competition images are also below.

Ideagram’s cover entry was strong and had a nice style including a bit of commentary:

MidJourney’s cover attempts had a blend of photorealism and graffiti:

Examples of the few standout API-generated category covers:

This Week’s Executive Summaries

Google Turns to Nuclear Reactors to Power AI Datacenters
Who would have guessed this would be a thing two years ago? Google struck a deal with California-based Kairos Power to acquire six small nuclear reactors (SMRs – Small Modular Reactor) to meet the growing energy demands of AI and cloud storage. The first reactor is slated for completion by 2030, with the rest by 2035, providing Google with 500 megawatts of round-the-clock power. This comes as other tech giants, including Microsoft and Amazon, are also turning to nuclear energy to power their operations.   
Wsj | emollick | theguardian

AI Helps U.S. Treasury Recover $4 Billion in Fraud in 2024 
The Treasury Department’s use of machine learning AI has significantly helped combat financial fraud, recovering $1 billion in check fraud and over $4 billion in total fraud during fiscal 2024—a six-fold increase from the previous year. By analyzing data sets and detecting hidden patterns in milliseconds, AI tools have transformed the way the government safeguards taxpayer dollars. This comes as online payment fraud is projected to exceed $362 billion by 2028.  The Treasury is expanding its AI toolset this year, exploring fraud-detection methods used by banks and partnering with states to fight unemployment insurance fraud.
cnn

Meta Advances Point Tracking with CoTracker3 – One of my favorite topics
Meta’s CoTracker3 is a new AI tool designed to track specific points in videos, like following a moving object or feature, even if it’s temporarily blocked from view or goes off-screen. This is similar to segmentation (outlining an object) and depthing (color shading objects based on their distance from the camera).  Tracking objects is going to be a big part of driverless cars and embodied robots, so when you watch these videos, imagine you are a friendly version of The Terminator or Predator.  CoTracker3 improves on older tools by using real-world videos to train, thanks to a clever technique called “pseudo-labeling,” where other AI systems generate training data automatically. Unlike traditional methods that rely on synthetic video (which doesn’t always match reality), CoTracker3 bridges the synthetic and real-lift, delivering more accurate and reliable results with 1,000 times less training data than earlier models. It’s worth seeing in action at Hugging Face.
Cotracker3 | AIatMeta 

TANGO: Creating Realistic Gesture Videos That Match Audio
Another one of my favorite topics and tangentially related to the point tracking link above.

TANGO is a new tool that creates videos where a person’s gestures perfectly match speech from any audio clip. Using just a short movie of someone talking and a new audio recording, TANGO generates high-quality, natural-looking videos of the person speaking the audio and gesturing. It improves on older methods by fixing common problems like gestures not syncing well with speech or awkward transitions between frames.  This is one of my favorite types of AI tech (similar to LivePortrait or Viggle), because it combines a blend of segmentation, latent consistency, depthing, and diffusion in ways that are hard to pick apart.  Again, think of an embodied robot or vision model learning emotions and context from audio and gestures. 
pantomatrix

Apple Study: AI Models Struggle with Math for Real Reasoning
A new study from Apple shows that AI models, like chatbots, don’t truly “understand” math—they often rely on recognizing patterns instead of reasoning like humans do. While these models do well on simple math tests, like grade-school problems, their abilities fall apart when tested in more realistic ways.  The study introduced a new test called GSM-Symbolic to see how AI handles changes in math problems, like using different numbers, adding tricky but irrelevant details, or making problems harder. Results showed the AI often gets confused: its accuracy drops sharply when questions get tougher or when unnecessary info is added. For example, one model’s accuracy fell by 65% just because irrelevant details were included. Compared to the older test (GSM8K), the AI struggled more with GSM-Symbolic. It performed worse when numbers changed compared to names, and its scores dropped as problems became more complex. This set-back of sorts suggests that AI isn’t really “thinking” about math—it’s just guessing based on patterns.
rohanpaul_ai | wired

OpenAI Releases Swarm: A Tool to Test Multiple AI Agents Working Together
OpenAI’s Swarm is an unsupported experimental tool for developers to tinker with how multiple AI agents can work together. It’s not meant for production use—more like a learning tool to try out ideas. The framework is based on two concepts: “Agents,” which handle tasks with specific tools and instructions, and “handoffs,” where one agent passes the task to another if needed. This makes it easy to build systems where AI handles different jobs, like triaging customer service requests or acting as a personal shopper. Swarm is lightweight, runs on your computer (not the cloud), and doesn’t save info between tasks.
AtomSilverman

Adobe’s Firefly AI Editor Now Includes Videos 
Adobe has launched a new AI tool, called the Firefly Video Model, that can help people create and edit videos with just a few clicks or even by typing simple instructions. Announced at Adobe MAX 2024, this tool is designed to be safe for businesses to use and is in a testing phase for the time being.  Firefly Video lets users extend video clips, smooth out transitions, or turn written descriptions into video scenes. For businesses, Adobe added features like dubbing voices into other languages and editing large numbers of images quickly. The tools are already in use by major companies like Pepsi and Mattel.  Firefly tools are built using licensed material and they include a special “label” that shows the content was made with AI. The tool is free to try while it’s in testing.
Adobe | theverge 

AI Visuals and Charts: Week Ending 10/18/2024

TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio-Motion Embedding and Diffusion Interpolation 

“New AI research from Meta – CoTracker3 Simpler and Better Point Tracking by Pseudo-Labelling Real Videos. More details ➡️ 

CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos

Top 33 Links of The Week – Organized by Category 

Agents and Copilots

“2/ A single person ran 30k agents on @ottogrid_ai within 48 hours! @HomamMalk mentioned that they are now aggregating ~250k data points on the platform per day. This is amazing @SullyOmarr 

“📃Open Canvas Open Canvas is an open source web application for collaborating with agents to better write documents. It is inspired by OpenAI’s “Canvas”, but with a few key differences: 📂Open Source 🧠Built in memory ✒️Start from existing documents 📂Open Source: All the 

AGI (Artificial General Intelligence)

LLMs don’t do formal reasoning – and that is a HUGE problem

Sabotage-Evaluations-for-Frontier-Models.pdf

“We’ve published a significant update to our Responsible Scaling Policy, which matches safety and security measures to an AI model’s capabilities. Read more here: 

Anthropic just made it harder for AI to go rogue with its updated safety policy | VentureBeat

Anthropic’s Responsible Scaling Policy, October 15, 2024

Sabotage evaluations for frontier models \ Anthropic

Audio

NotebookLM update: Audio Overview controls and team collaborations

“New NotebookLM updates, rolling out today: 🎧Pass a note to the hosts – you can now click on ‘Customize’ in Audio Overviews to give additional instructions, such as focusing on a specific topic, source, or even adjusting the audience it’s optimized for. 🎧You can now minimize” / X

“BREAKING 🚨: NotebookLM is launching an early access program for Businesses as well as introducing a customisation feature for Audio Overviews and background listening support. More details below 👇 

Google’s NotebookLM Now Lets You Customize Its AI Podcasts | WIRED

Ethics/Legal/Security

Sam Altman’s Worldcoin becomes World and shows new iris-scanning Orb to prove your humanity | TechCrunch

‘Blade Runner 2049’ Producer Sues Elon Musk’s Tesla Over AI Images

“You really, really should not trust audio clips anymore Even a couple months ago, it used to take a commercial service to clone a voice. No more. Here is me creating a voice clone of myself using just a 10 second reference clip on my home computer This is all real time, no cuts 

“A German court dismissed a copyright lawsuit against LAION, a nonprofit that compiles large-scale image datasets for training models like Midjourney and Stable Diffusion. Learn more in #TheBatch: 

Google

Google supercharges Shopping tab with AI and personalized recommendation feed | TechCrunch

Google Shopping is getting a ‘for you’ feed of products – The Verge

The new Google Shopping is rebuilt with AI

International

China’s Alibaba claims AI translation tool beats Google, ChatGPT

“We are proud to present the latest model ⚡️Yi-Lightning ⚡️ now #6 in the world, higher than the original GPT-4o released 5 months ago. Also humbled that @01AI_Yi is ranked #3 LLM player on @lmarena_ai Chatbot Arena — open after OpenAI, Google and tie with xAI to serve other” / X

Meta

“📢 Meta Spirit: The first open source multimodal language model from @AIatMeta 👏 Meta Spirit LM offers: – Seamless integration of speech and text in a multimodal language model – Word-level interleaving of speech and text datasets – Cross-modality generation capabilities – Two”  https://twitter.com/rohanpaul_ai/status/1848000249920364838

OpenAI

Prominent Microsoft AI Researcher to Join OpenAI — The Information

Former Palantir CISO Dane Stuckey joins OpenAI to lead security | TechCrunch

Perplexity

“Introducing Internal Knowledge Search (our most-requested Enterprise feature)! For the first time, you can search through both your organization’s files and the web simultaneously, with one product. 

“Perplexity Finance: real time stock prices, deep dives into a company’s financials, comparing multiple companies, studying 13f’s of hedge funds, etc. The UI is just delightful! 

Publishing

“Will AI Agents Kill Advertisements? Yes, Yes, No, and Yes” / X

Filmmakers Are Worried About AI. Big Tech Wants Them to See ‘What’s Possible’ | WIRED

Robotics and Embodiment

“Our robots are helping us get to the next level, powering the next gen fulfillment facility at Amazon. How? By using #AI to ensure the right products are in fulfillment centers near customers, maximizing efficiency for our employees, lifting heavy objects for us, and so much 

Science and Medicine

Generative AI can ease administrative burden in healthcare

AI Medical Imagery Model Offers Fast, Cost-Efficient Expert Analysis | NVIDIA Technical Blog

GE HealthCare announces CareIntellect for Oncology, harnessing AI to give clinicians an easy way to see the patient journey in a single view

Twitter/X/Grok

Elon Musk’s X is changing its privacy policy to allow third parties to train AI on your posts | TechCrunchhttps://techcrunch.com/2024/10/17/elon-musks-x-is-changing-its-privacy-policy-to-allow-third-parties-to-train-ai-on-your-posts/

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading