This Week’s Covers

This week’s cover celebrates Figure Robotics’ new humanoid robot and OpenAI’s rumored release of a model capable of advanced reasoning, codenamed “Strawberry.” The main image is a PR photo from Figure, which depicts their robot standing in front of a large, blank movie screen. The composite was created using Adobe Photoshop’s “remove background” feature to form layers, and strawberries were added using Photoshop’s “generative fill” function. The font used is Neue Haas Grotesk, Figure’s core brand font.

The category images this week are robot-themed and generated using Flux. It took me about twelve covers to find my rhythm. I prefer MidJourney’s interface because it keeps an archive of all images and prompts, while Flux doesn’t save past images once a new one is generated. MidJourney still holds its own, and I remain partial to it; however, when Flux nails it, the results are unparalleled. Just amazing.

Each prompt I used was unique, but the most successful format was something like: “[Describe main setting] [a futuristic robot, sleek smooth humanoid design. Smooth, glossy black faceplate with no visible facial features, high-tech, minimalist appearance. The robot’s body is matte black or dark gray, with articulated joints and mechanical parts that resemble those of a human, including fingers] [Describe how to display the category name with “CATEGORY NAME”].   Here are a few of my favorites from this week. Remember, these are all made with Flux, not MidJourney:

Executive Summary: Week Ending 08/09/2024

Figure Robotics Unveils Advanced AI Robot, Figure 02
Figure Robotics introduces Figure 02, a cutting-edge AI-powered robot, capable of conversation and common-sense visual reasoning with the world around it.  The robot’s key improvements over the 01 include enhanced AI-driven vision with six RGB cameras, 16-degree freedom hands with human-equivalent strength, and speech-to-speech capabilities for seamless conversations. With onboard Visual Language Models (VLM) for fast common-sense reasoning and a powerful 2.25 KWh battery providing 50% more energy, Figure 02 also boasts triple the processing power of its predecessor.  
Twitter https://www.figure.ai/ 

OpenAI Introduces Structured Outputs for API Reliability
OpenAI has unveiled a new feature called Structured Outputs, which ensures model-generated outputs adhere strictly to developer-supplied JSON Schemas. This might sound nerdy, but it’s incredible. Structured Outputs allow developers to build more reliable applications, such as data-fetching assistants and multi-step workflows. OpenAI’s latest model, gpt-4o-2024-08-06, achieves perfect scores in following complex JSON schemas, marking a major improvement in tool call reliability.
LangChainAI | rohanpaul_ai | https://openai.com/index/introducing-structured-outputs-in-the-api/ 

GPT-4 Effectively Simulates People, Accurately Replicating Social Science Experiments
A recent study shows that GPT-4 can simulate human responses well enough to replicate social science experiments with high accuracy. By generating survey responses from “AI people” based on randomly assigned demographic characteristics, researchers tested the predictive power of large language models (LLMs). The study, which spanned 70 different experiments, found a strong correlation (r = .85) between simulated responses and real-world observed effects. This suggests that LLMs could be valuable tools in predicting outcomes of social science research, potentially transforming the way experiments are conducted.
Emollick | RobbWiller

OpenAI Leadership Shake-up Raises Doubts About AGI Progress
OpenAI has seen a series of key departures, fueling skepticism about the company’s progress toward AGI (Artificial General Intelligence). Co-founder Greg Brockman is taking a sabbatical through the end of 2024, and John Schulman left for rival Anthropic to deepen his focus on AI alignment and ‘engage in more technical work’.  OpenAI co-founder and VP of Consumer Product Peter Deng also resigned.  Brockman reassured that he is returning and OpenAI’s mission remains unfinished, but some in the AI community question why prominent figures are leaving if a breakthrough is imminent.  Critics suggest the hype around AGI may be overblown, with recent advancements plateauing around GPT-4-level capabilities.   My optimistic theory is that OpenAI is in a tactical mode, training and building, giving Brockman a window of opportunity to relax before the next phase of innovation kicks in.  The other departures are tough to explain. 
arstechnica | cnbc | techcrunch

Elon Musk Files New Lawsuit Against OpenAI, Citing ‘Deceit of Shakespearean Proportions’
Elon Musk has reignited his legal battle against OpenAI, alleging that his former partners, including CEO Sam Altman, manipulated him into co-founding the company. Musk claims the shift from OpenAI’s non-profit origins to a for-profit enterprise breached their agreement to prioritize AI development for humanity’s benefit. Filed in a California court, the lawsuit accuses OpenAI of racketeering and wire fraud, seeking compensation for damages. OpenAI denies the allegations, referencing earlier communications that suggest Musk supported the company’s evolution.  
theguardian | arstechnica

Nvidia is Scraping ‘A Human Lifetime’ of Videos Per Day to Train AI
Leaked documents reveal Nvidia’s large-scale data scraping practices, where the company collects about 80 years’ worth of videos daily to train its AI models. This includes using millions of videos from YouTube, Netflix, and other sources to develop tools for its Omniverse, autonomous vehicles, and digital avatars. Nvidia claims it adheres to copyright laws, though internal discussions show creative methods to bypass restrictions, such as using Google Cloud to access the YouTube-8M dataset. The documents highlight concerns about copyright and personal data usage, raising broader ethical questions about AI training transparency.
404media | pcgamer

OpenAI Delays Voice Features Due to Concerns of Emotional Dependence and ‘Novel Risks’
OpenAI’s recent safety review highlights concerns about its new GPT-4 voice mode, which can mimic human conversation with lifelike accuracy. The company is wary that users may become emotionally reliant on the AI, leading to diminished human interaction, especially as the tool can recognize emotional tones and respond accordingly. OpenAI also raised concerns about potential security risks, like GPT-4 convincingly imitating people or secretly personalizing responses based on individual voices.  
cnn | emollick

China Circumvents U.S. A.I. Chip Bans Through Smuggling and Front Companies
Despite U.S. efforts to restrict China’s access to advanced Nvidia A.I. microchips, a thriving underground market is helping China bypass the bans. Vendors in Shenzhen openly offer U.S.-made chips, essential for powering advanced technologies with potential military applications. American companies, including Nvidia, Intel, and Microsoft, are under scrutiny, as Chinese firms use front companies and fraudulent shipping practices to acquire restricted chips. The U.S. has implemented export controls to slow China’s military advancements, but businesses are exploiting loopholes, raising concerns about enforcement effectiveness and the potential for Chinese technological progress despite the restrictions.
nytimes

Grounded SAM: A Powerful Tool for Visual Recognition and Editing
Seeing and tracking objects in videos is getting more accurate and sophisticated.  Now, AI can not only find, select, and track objects in videos, it can respond to plain language instructions without training on what’s in the video.  Grounded SAM combines two advanced tools to help computers recognize and select objects in videos based on simple text descriptions. By linking different models, it can handle a variety of tasks, like automatically labeling parts of an image, editing images, and even analyzing 3D human movements. It’s especially good at recognizing objects without needing specific training, performing well on benchmarks that test its ability to identify things in real-world settings.  This is going to be the key to automated vehicles, robot embodiment, surveillance and applying data to video (i.e. sports, traffic, animals, plants, etc).
Must see video: Tianhe_Ren

Tesla’s New Vision System Could Make Cars and Robots Smarter
Tesla has a new invention that could change how robots understand the world around them. Instead of using extra sensors like radar or LiDAR, Tesla’s system uses only cameras to help robots “see” their surroundings. It works by breaking down the environment into small 3D blocks (called voxels) and figuring out if each block has an object in it, what kind of object it is (like a car or a person), and if it’s moving. This helps robots make quick decisions to avoid obstacles, move around safely, and even pick up objects. Tesla’s system could make robots like Optimus much smarter and better at doing tasks in places where people live and work.
Video examples: seti_park

Tesla Dojo: Musk’s Vision for an AI Supercomputer and Autonomous Vehicles
Elon Musk’s Dojo is a custom-built supercomputer that plays a crucial role in Tesla’s push toward fully autonomous driving. Dojo is designed to train Tesla’s “Full Self-Driving” (FSD) neural networks using video data from millions of Tesla cars. Remember how much I talk about segmentation and object tracking, and Jim Fan’s simulation lab? – this is Elon’s version. This massive data collection supports Tesla’s vision of achieving self-driving without relying on external sensors like lidar, instead using only cameras (see above story as well). The Dojo project also involves Tesla’s development of its own D1 chips, which aim to improve the efficiency and power of AI training while reducing reliance on Nvidia’s GPUs. As Tesla continues to advance Dojo, it could transform the company into a major AI player, not just in vehicles but potentially in broader AI applications. While risks remain, analysts predict Dojo could significantly increase Tesla’s market value by enabling new technologies like robotaxis and AI-powered services.
techcrunch

AI Visuals and Charts 

Incredible augmented reality demo of an empty glass (real?) filling with fluid (virtual)

“AI is going to seriously have us questioning reality, especially ones delivered on a screen”

Image to video creation using FLUX + RunwayML – a clip of a woman speaking at a TED conference

Nicolas Neubert on X: “Let’s bring in some motion, shall we?” https://twitter.com/iamneubert/status/1821970292014768420  

Top 64 Links of The Week by Category

Apple 

“Apple is once again rumored to charge $10-20/month for its advanced Apple Intelligence features (via CNBC). Would you pay for Apple Intelligence?

Artificial General Intelligence (AGI)

“Future AI won’t be tricked or manipulated by simple tactics. They might even perceive it as disingenuous and manipulative. So it’s important to just be a good person. Future AIs are watching.” / X

Audio

Meta is reportedly offering millions to use Hollywood voices in AI projects

Meta courts celebs like Awkwafina to voice AI assistants ahead of Meta Connect – The Verge

HeyGen

“Olympic shooter Yusuf’s interview translated into English by @HeyGen_Official. Pretty good! Can’t wait till this tech is more like a filter I can turn on for any video.

Chips, Hardware, and Infrastructure

OpenAI Makes a $60 Million Hardware Startup Bet — The Information

Dell Makes Cuts to Boost AI Pivot, Reportedly Laying Off 12,500 Employees | PCMag – https://www.pcmag.com/news/dell-makes-cuts-to-boost-ai-pivot-reportedly-laying-off-12500-employees

Intel

Intel CEO Pat Gelsinger’s Dream Job Takes a Nightmarish Turn – WSJ 

Intel reportedly gave up a chance to buy a stake in OpenAI in 2017 | Tom’s Hardware 

“hello I’m new to the stock market is it good when the intel ceo starts praying” / X

Intel CEO Pat Gelsinger on X: ““Let your eyes look straight ahead; fix your gaze directly before you. Give careful thought to the paths for your feet and be steadfast in all your ways” Proverbs 4:25-26” / X – https://x.com/PGelsinger/status/1820129317122080977 

NVIDIA

NVIDIA AI Developer on X: “Meet James, an interactive digital human knowledgeable about our products. Talk to James ➡️ https://t.co/YXoT6fRD8t 💬 James uses a collection of NVIDIA and @elevenlabsio digital human technologies to provide natural and immersive responses. https://t.co/53IfOg0w8O” / X – https://twitter.com/NVIDIAAIDev/status/1821637513728950441

Nvidia reportedly delays its next AI chip due to a design flaw – The Verge

Consumer Products

Wendy’s bringing Spanish AI ordering to drive-thrus in Florida | WFLA

Taco Bell’s drive-thru AI might take your next order – The Verge

Education 

Colleges offer AI degrees — could they give job seekers an edge? – https://www.nbcnews.com/tech/tech-news/ai-degree-major-college-university-schools-rcna163462

Ethics/Legal/Security 

Five US states push Musk to fix AI chatbot over election misinformation

“Could AI agents develop a language incomprehensible to humans?💬 Our CEO @ronbodkin discusses decentralizing AI for transparency at @BerkeleyRDI’s panel during @StanfordCrypto’s Blockchain Week. 🔍 With:

YouTube testing a feature like Twitter Community Notes

Argentina will use AI to ‘predict future crimes’ but experts worry for citizens’ rights | Argentina | The Guardian

Inside Turing, the company gathering ‘human data’ for every major AI company | Semafor

near on X: “heavenbanning, the hypothetical practice of banishing a user from a platform by causing everyone that they speak with to be replaced by AI models that constantly agree and praise them, but only from their own perspective, is entirely feasible with the current state of AI/LLMs https://t.co/PlPYROHXnn” / X – https://twitter.com/nearcyan/status/1532076277947330561 

“Fei-Fei agrees with the overwhelming majority of AI scientists: SB1047 won’t solve anything and will harm AI R&D in academia, little tech, and the open source community.”  

“AI regulators in the West should read up on the history of the printing press in the Ottoman Empire, wherein its adoption was inhibited for centuries to 1. protect the scribe industry (!) and 2. prevent “misinformation” All of which ultimately led to the region’s downfall 

Nvidia Scrambles for a Response to Antitrust Scrutiny – The New York Times

OpenAI has built a text watermarking method to detect ChatGPT-written content — company has mulled its release over the past year | Tom’s Hardware

Exclusive: Experts Pen Support for California AI Safety Bill | TIME

Google 

Google Unveils New Gemini-AI Powered TV Streamer

Google brings Gemini-powered search history and Lens to Chrome desktop | TechCrunch

International 

“Chinese open weights model that easily surpasses all previous models, both closed and open, at MATH:” / X

Meta

Zuckerberg touts Meta’s latest video vision AI with Nvidia CEO Jensen Huang | TechCrunch

“The methods from this paper were able to reliably jailbreak the most difficult target models with prompts that appear similar to human-written prompts. Achieves attack success rates > 93% for Llama-2-7B, Llama-3-8B, and Vicuna-7B, while maintaining model-measured perplexity <

Meta is reportedly offering millions to use Hollywood voices in AI projects

Meta courts celebs like Awkwafina to voice AI assistants ahead of Meta Connect – The Verge

Microsoft

Palantir and Microsoft Partner to Deliver Enhanced Analytics and AI Services to Classified Networks for Critical National Security Operations – Stories

Microsoft scales back AI partnership with Emirati firm amid concerns over China ties – POLITICO

Every Microsoft employee is now being judged on their security work – The Verge

Multimodality

AI can see what’s on your screen by reading HDMI electromagnetic radiation | TechSpot – https://www.techspot.com/news/104015-ai-can-see-what-screen-reading-hdmi-electromagnetic.html

“Want a robot to assist you in the kitchen **without any instructions** simply by watching you?🤖🏠 🚀 Presenting our recent paper on action anticipation from short video context for human-robot collaboration, accepted at Robotics and Automation Letters (RA-L).

“Computer vision + Journalism + #Olympics = 😍 – The @nytimes used computer vision to detect the positions of the athletes on photos taken every 100ms – Speeds were then computed by combining their positions and the timestamp of each photograph – Manual verification and

OpenAI

OpenAI invests in a webcam company turned AI startup. – The Verge

“OpenAI is leading a $60 million investment in Opal, which sells $300 professional webcams and plans to develop other types of devices powered by OpenAI’s AI models. w/ @steph_palazzolo

Open Source

“Mistral Large 2 (2407) is now on @lmsysorg. It performs extremely well in the Coding, Hard Prompts, Math, and Longer Query categories, where it outperforms GPT4-Turbo and Claude 3 Opus. It is also doing very well in Instruction Following where it ranks above Llama 3.1 405B.

Qwen

“CONGRATS to @Alibaba_Qwen team on Qwen2-Math-72B outperforming GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro, Llama-3.1-405B on a series of math benchmarks 👏👏👏

Introducing Qwen2-Math | Qwen

Podcasts/YouTube/Op-Eds

My Vision: How AI Changes the Internet in 6 Phases

Publishing

“A powerful application of GPT in journalism is to turn @AP wire stories into the @axios “smart brevity” format. Here’s an example of a story about Amazon’s recent launch of RxPass, a new drug prescription service. 

Reddit considers search ads, paywalled content for the future | Ars Technica

Reddit to test AI-powered search result pages | TechCrunch

Robotics and Embodiment

New Scientist on X: “Watch a robot peel vegetables like people do – by holding a squash, radish or pumpkin in one hand, while it is peeled by the other. https://t.co/NECplBer1c” / X – https://twitter.com/newscientist/status/1820712997246939535 

Science and Medicine

Fully-automatic robot dentist performs world’s first human procedure

Neuralink implanted second trial patient with brain chip, Musk says | Reuters

AI predicts 3D structures of receptors for drug development

‘Game changer’ AI detects hidden heart attack risk, say scientists

Video

HeyGen 

“Olympic shooter Yusuf’s interview translated into English by @HeyGen_Official. Pretty good! Can’t wait till this tech is more like a filter I can turn on for any video.

The Rest of Video News

Model a full head from a single portrait image

“🚨Head360: Learning a Parametric 3D Full-Head for Free-View Synthesis in 360° 🌟𝐏𝐫𝐨𝐣:

“Mind blowing video generation model from @thukeg I’ve seen so many AI generated videos and while the examples below might be a bit cherry-picked (like any other demos), I’m still feel totally ASTONISHED by the quality that it deliver. Great job team!

Twitter/X/Grok

Elon Musk continues to siphon Tesla talent to train xAI’s Grok | Electrek

Neuralink installs its brain implant into a second human patient https://www.teslarati.com/neuralink-brain-implant-second-human-patient/

Credits/Sources


Medium shot.  A modern computer chip factory.  In the background, a large metal box is stampted with the text “Ethan B. Holland” in large letters.   In the foreground, an orange robotic arm is extended toward the camera, holding a bouquet of fresh flowers.  A delicate tag on the flowers reads “Credits” in a clear font. – FLUX.1 [dev]

Most of these weekly links come from just a few prolific oversharing sources. Please follow them, as they work hard to find the news each week and they make it a lot easier for me to compile.

For previous issues, please visit the archives!

MidJourney Prompt: a cozy home library at a beach house, a gorgeous summer afternoon. a humanoid robot reads a book. –chaos 20 –ar 4:3 –style raw –v 6.1 Upscaled with Magnific.ai. GPT suggested Lato for the August/Summer font because it is “a friendly and warm sans-serif font with soft curves, making it feel fresh and approachable.”

Thanks for reading!

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading