About This Week’s Cover

The two big stories this week were Sam Altman’s essay “The Intelligence Age” that predicts we’ll see superhuman intelligence within 10,000 days (cue Donnie Darko) and the resignation of OpenAI CTO Mira Murati (along with the chief of research, Bog McGrew and VP of Post Training Barrett Zoph).  The cover combines the two stories in an abandoned office.  Image created with Ideogram with the prompt: Interior shot of an empty office in Silicon Valley. The office is open and spartan with soaring glass windows and light wooden accents. It is abandoned and dusty. In the foreground a binder lies on a desk with a label that reads “The Intelligence Age”. Next to it is an envelope labeled “Mira Resignation”.  Edited with Photoshop.  

This Week’s Executive Summaries

Sam Altman writes a lofty essay: “The Intelligence Age” In The Intelligence Age, OpenAI CEO Sam Altman predicts a not-so-distant future where AI transforms daily life and brings a new era of prosperity.  Altman is amazed by how well today’s AI—especially deep learning—actually works, calling it a turning point in history. He thinks we could have super-smart AI systems, even “superintelligence,” within just a few thousand days, or around 10 years.  Altman believes AI will soon let us do things that seem like magic today—think personal AI tutors, virtual medical assistants, and tools to solve huge problems like climate change. Just like electricity and the internet changed the world, Altman sees AI as the next big leap. He warns that AI’s benefits should be available to everyone, not just the wealthy, or it could cause new divisions and wars.  The possibilities, he predicts, are mind-blowing: from personalized learning for all to breakthroughs in health and science that make today’s progress look small. It’s a barn-burner and worth a skim.
Samaltman

Two reflections on OpenAI’s recent models:
“OpenAI staff who worked on the new o1 model say that the AI is “spiritual” and “oddly human” in how it can reason, reflect and question itself 
tsarnick

“Comparing the creativity of a representative human sample to GPT-4 finds “the creative ideas produced by AI chatbots are rated more creative than those created by humans… Augmenting humans with AI improves human creativity, albeit not as much as ideas created by ChatGPT alone.” 
Emollick

OpenAI’s CTO Mira Murati Steps Down – Organizational Shakeup
OpenAI CTO Mira Murati has announced her departure from the company, alongside Chief Research Officer Bob McGrew and VP of Post Training Barret Zoph, marking a major leadership transition ahead of the company’s annual Dev Day event. In a post on X, Murati said she is leaving to pursue new personal and professional explorations. CEO Sam Altman acknowledged the abrupt nature of the changes, citing the intense pace and unique demands of OpenAI’s mission. Over her six-year tenure, Murati played  a pivotal role in the release of ChatGPT and DALL-E, while also establishing foundational safety and usability standards. The reshuffle aligns with OpenAI’s ongoing evolution, as the company recently secured new funding and is reportedly moving toward a for-profit model.
Theverge | sama | miramurati | techcrunch

OpenAI Releases ChatGPT Advanced Voice

Rather than an executive summary, here are examples of it in action:

“amazing. chatgpt advanced voice perfectly handles a sentence mixing three languages at once something truly different about how this works 
https://twitter.com/localghost/status/1838981239493308534

“Advanced Voice is rolling out to all Plus and Team users in the ChatGPT app over the course of the week. While you’ve been patiently waiting, we’ve added Custom Instructions, Memory, five new voices, and improved accents. It can also say “Sorry I’m late” in over 50 languages. 
https://twitter.com/OpenAI/status/1838642444365369814

“Advanced Voice rolling out broadly, enabling fluid voice conversation with ChatGPT. Makes you realize how unnatural typing things into a computer really is:” / X
https://twitter.com/gdb/status/1838662392970150023

“ChatGPT Advanced Voice Mode counting as fast as it can to 10, then to 50 (this blew my mind – it stopped to catch its breath like a human would) 
https://twitter.com/CrisGiardina/status/1818627205217272098

“Advanced Voice in ChatGPT tunes my guitar. 
https://twitter.com/skirano/status/1838722728443904120

 “I spent 24 hours with @ChatGPTapp’s all-new Advanced Voice Mode feature. It is absolutely a peek at the future. Here’s my full review: 
https://twitter.com/danshipper/status/1819735498870440139

“This is madness 🤯 OpenAI Advanced Voice Mode is shockingly unbelievable. Instruction: Could you tell me a story about a bear and a fox? Use your voice but make the character voices super dramatic and use a slithering sound for a snake. 
https://twitter.com/slow_developer/status/1838816505640739052

“Because at heart I’m a child – I had Advanced Voice perform a monologue from Hamlet in his best thespian tone and then repeat it as if it were in a middle school play. It can also translate and perform in modern English.
https://twitter.com/RyanMorrisonJer/status/1838971228205298102

A bit edgy. But advanced voice has a perfect rural Australian accent. 
https://twitter.com/AtomSilverman/status/1838983919116435522

You can get advanced voice to rap and beatbox
https://twitter.com/AtomSilverman/status/1838983922861957378

Meta creates augmented reality glasses prototype
“We just unveiled Orion, our full AR glasses prototype that we’ve been working on for nearly a decade. When we started on this journey, our teams predicted that we had a 10% chance (at best) of success. This was our project to see if our dream AR glasses—wide FOV display, less than 100 grams, wireless—were actually possible to build. Not only do they work, we’ll be using them internally as a time machine to help build the core experiences and interaction paradigms needed for the consumer AR glasses we plan to launch in the coming years.”
Ylecun | NathieVR

Alibaba and Nvidia Team Up to Boost Autonomous Driving in China
Alibaba Cloud and Nvidia are partnering to bring smarter autonomous driving technology to Chinese carmakers. Announced at Alibaba’s recent Apsara Conference, the collaboration uses Alibaba’s AI models on Nvidia’s Drive platform to power advanced in-car features, like voice-controlled assistants that can hold conversations, make recommendations, and even help with navigation. Chinese electric vehicle brands like Li Auto and Geely’s Zeekr plan to use this technology to enhance the driving experience for their vehicles. This partnership also explores leveraging Nvidia to make computing faster and more energy-efficient, helping businesses move more of their AI operations to the cloud. Although U.S. restrictions limit Nvidia’s sales of its highest-end chips in China, the company continues to support the growing demand for AI-powered services in the country.
Scmp

It was a slow week overall as far as huge announcements go.   Agent builder Replit claims improved accuracy and there were several open source announcements (see top links below).  Agents in general are the hot topic to follow, and if you’re a futurist, I highly recommend keeping track of embodied robots.  This guy in the visuals and charts section (immediately below), is an incredibly creative person.  If you’re looking for the mix of how AI can augment creativity rather than compete, he’s a key creator to follow.

AI Visuals and Charts: Week Ending 09/27/2024

“Now the BTS. I recorded myself lip-syncing the song to transfer my facial movements onto images and videos. The parts where there’s other movements, I used Runway and Lumalabs to get the images to move and used liveportrait after to get the mouth movements. #kanye #ye #comfyui 
https://twitter.com/8bit_e/status/1839732374336090625

Top 30 Links of The Week – Organized by Category 

Agents and Copilots: AI News Week Ending 09/27/2024

“In the two weeks post launch, our AI team has reduced adverse experiences in the Replit Agent by over 80%. If you’ve been holding out to give it a shot, now is a great time. 

“Using replit agent feels like talking to a really eager APM. Also now I have a clean GUI for YT-DLP to download YouTube videos, and never need to deal with a sketchy app or website again. 

“the “end state” of most ai apps probably wont be a platform its a gonna be an ai agent talking with you on slack doing the work of 10 employees at once and then when its done, it sends you the finished project” / X

“We are currently hiring for an open research scientist position working on my team in multi-agent artificial general intelligence: 

“O1 preview scores 98% on a planning benchmark, open up new category: the Large Reasoning Model (LRM) The king is dead long live the king” / X

Rabbit’s web-based ‘large action model’ agent arrives on r1 on October 1 | TechCrunch

Anthropic

“Rumor has it, Anthropic is dropping a new model tomorrow.” / X

“Anthropic will likely drop a new model tomorrow….. It should be super exciting.” / X

Ethics/Legal/Security

“SB-1047 would definitely have a chilling effect on open source AI and the entire AI ecosystem. Very much hoping that Governor @GavinNewsom will veto it.” / X

“Thank you Governor @GavinNewsom for vetoing SB-1047. The open source AI community as grateful for your sensible decision. 

An AI can beat CAPTCHA tests 100 per cent of the time | New Scientist

Google

Updated production-ready Gemini models, reduced 1.5 Pro pricing, increased rate limits, and more – Google Developers Blog

Meta

“With Llama 3.2 we released our first-ever lightweight Llama models: 1B & 3B. These models empower developers to build personalized, on-device agentic applications with capabilities like summarization, tool use and RAG where data never leaves the device. 

“Zuck’s AI strategy in a nutshell: – Free frontier model for devs – Undercut closed-source rivals – Crowdsource best use cases for multimodal AI – Skip cloud wars; instead monetize biz/creator agents – Harvest fresh data/content for Meta ecosystem – Profit Zuck’s XR strategy in 

Worst idea in human history

“4. Meta is testing ‘Imagined for you’ AI-generated content that will show up on users Facebook and Instagram Feeds. You can “tap a post to take the content in a new direction” or “swipe to see more content imagined for you in real-time” AI-generated content, tailored to each 

Multimodality 

“Meet Molmo: a family of open, state-of-the-art multimodal AI models. Our best model outperforms proprietary systems, using 1000x less data. Molmo doesn’t just understand multimodal data—it acts on it, enabling rich interactions in both the physical and virtual worlds. Try it 

“Molmo is a new fancy multimodal model, but people are missing the data sections! The data pipelines and sections are a huge gem 💎 Stage 1 – create a dense captioning dataset of 712k images/1.3M captions No VLM was used to generate data 1. Obtained images from 50 high-level 

Warner Bros. Discovery, Google Tap AI For Captioning

OpenAI 

OpenAI says the latest ChatGPT can ‘think’ – and I have thoughts | Technology | The Guardian

OpenAI just unleashed an alien of extraordinary ability

Apple not investing in OpenAI after all, new report says – 9to5Mac

Turning OpenAI Into a Real Business Is Tearing It Apart – WSJ

Open Source 

“New Open source SoTA Multimodal (Vision) Language model dropped – Molmo Outperformed Claude 3.5 Sonnet, GPT4V, Gemini 1.5 Pro – using 1000x less data. 🤯 🗣️ Novel dataset (PixMo) with detailed human-spoken image captions 🧠 Architecture: Vision encoder + LLM 🔓 Open weights, 

“Meet Molmo: a family of open, state-of-the-art multimodal AI models. Our best model outperforms proprietary systems, using 1000x less data. Molmo doesn’t just understand multimodal data—it acts on it, enabling rich interactions in both the physical and virtual worlds. Try it 

“Open Dataset release by @OpenAI! 👀 OpenAI just released a Multilingual Massive Multitask Language Understanding (MMMLU) dataset on @huggingface! 🌍 MMLU test set available in 14 languages, including Arabic, German, Spanish, French,…. 🧠 Covers 57 categories from elementary to 

“The really MASSIVE week continues with 🚀 Llama 3.2 Release – Introduces 1B and 3B text models for edge devices, 11B and 90B vision models – All models support 128K token context – 1B/3B outperform Gemma 2 2.6B and Phi 3.5-mini on key tasks – 11B/90B vision models competitive 

“We have GPT-4 for coding at home! I looked up @OpenAI GPT-4 0613 results for various benchmarks and compared them with @Alibaba_Qwen 2.5 7B coder. 👀 > 15 months after the release of GPT-0613, we have an open LLM under Apache 2.0, which performs just as well. 🤯 > GPT-4 pricing 

“GPT-4 for coding at home! Qwen 2.5 Coder 7B outperforms other @OpenAI GPT-4 0613 and open LLMs < 33B, including @BigCodeProject StartCoder, @MistralAI Codestral, or Deepseek, and is released under Apache 2.0. 🤯 Details: 🚀 Three model sizes: 1.5B, 7B, and 32B (coming soon) up 

Podcasts/YouTube/Op-Eds

“5 Predictions on AI Agents (summary slide from my speeches) 1) AI agents bring the internet to humans: 1) AI agents centralize information, handling internet tasks and purchases. 2) Content delivered in a multi-modal format humans choose. 2) Media and revenue models shift 

Robotics and Embodiment 

Boston Dynamics’ Spot can now autonomously unlock doors | TechCrunch

Boston Dynamics partners with Assa Abloy to let the dogs in – The Verge

Video

Alibaba launches over 100 new AI models, releases text-to-video generation

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading