This week’s cover is a play on The Wolf Of Wall Street, inspired by Sam Altman’s return as CEO of OpenAI and created with Dall-E 3.
Executive Summary
This week had a lot of big news (beyond Sam Altman returning as CEO of OpenAI). Each one of these topics stands alone and is worth knowing. See examples in the top stories section.
- ChatGPT Voice: Voice interface to talk to ChatGPT has rolled out for all free users
- Audio to Instrument: Sing, tap, or kazoo your way into any instrument
- Stable Video Diffusion: Generative video based on images
- Eleven Labs Audio Mapping: Swap audio clones on top of audio samples
- Amazon has deployed 750,000 robots: 400,00 in the last two years
- AI Coming For Alexa: Amazon.com on Friday announced it is trimming jobs at its Alexa voice assistant unit, citing shifting business priorities and a greater focus on generative artificial intelligence.
Top Stories
These are the links to click if you only pick a few. Even if they look boring, click them! I did the work, so you don’t have to worry. All are 10/10 would recommend.
- AI Voice to Instrument: You have to see it to believe it.
- Homemade Commercial Quality Video With AI: Absolutely bonkers
- My favorite AI project this week comes from the incredibly talented Martin Haerlin and Hauke Hilberg. It’s still very early days for text-to-video and video-to-video… but this 50 sec clip really brings to life the potential and possibilities for generative AI in marketing.
- https://www.linkedin.com/posts/harbech_my-favourite-ai-project-this-week-comes-from-activity-7136298249912475648-ZGcf
- https://twitter.com/Martin_Haerlin/status/1724066090882392219
- Behind the scenes: https://www.instagram.com/p/C0cc8Tfs9gF/
- Stable Video Diffusion: Animate still frames. Mind-blowing.
- Turing Memes into Videos with Stable Video Diffusion
- Animating Album Covers with Stable Video Diffusion
- Here’s our latest LABs experiment where we ran several album covers through the newly released SVD model to bring them to life. Stable Video Diffusion (SVD) is a foundational model for generating videos from text prompts or a single image.
- https://www.linkedin.com/posts/t-da_ai-svd-stablediffusion-activity-7135040448061972480-reRC
- De-Aging Harrison Ford with Stable Diffusion + ControlNet + EbSynth + Fusion
The Rest: AI News of The Week
Don’t let the volume overwhelm you. Have fun and skim it. The links are organized by topic, sorted from ‘coolest’ to ‘least cool’, and each topic is clearly defined with a headline. I do the work so you don’t have to! The links descriptions are often pulled directly from tweets or articles, so it’s not always my voice. Pause when you see something that interests you. Reach out to me any time. I enjoy sharing and discussing these items!
OpenAI: Weekend of Chaos
Employee Loyalty
Breaking: 505 of 700 employees at OpenAI tell the board to resign.
more senior departures from openai tonight:
– GPT-4 lead / director of research, jakub pachocki
– head of AI risk, aleksander madry
– open-source baselines researcher
– sam altman
– greg brockman
this is just day 1
There are now over 700 signers and counting (out of ~750 employees) to our letter calling for unity & board ouster. The 550 the media reported was just the signatures we could get between 1:30am and 5am PT last night. This was everyone staying up and calling each other in support
OpenAI Staff Threatens Exodus, Jeopardizing Company’s Future
A board member who was part of Sam Altman’s ouster as chief executive joined a majority of the company’s staff in calling for the decision’s reversal.
Microsoft hires former OpenAI CEO Sam Altman
Greg Brockman, OpenAI co-founder, is also joining Microsoft to lead a new advanced AI research team.
We are going to build something new & it will be incredible. Initial leadership (more soon):
@merettm @sidorszymon @aleks_madry @sama @gdb
The mission continues.
I’m super excited to have you join as CEO of this new group, Sam, setting a new pace for innovation. We’ve learned a lot over the years about how to give founders and innovators space to build independent identities and cultures within Microsoft, including GitHub, Mojang Studios, and LinkedIn, and I’m looking forward to having you do the same.
New CEO of OpenAI Threatens to Quit
Today I got a call inviting me to consider a once-in-a-lifetime opportunity: to become the interim CEO of OpenAI.
The new CEO of OpenAI, Emmett Shear, will quit if the OpenAI board can’t provide evidence of why they fired Sam Altman.
Ilya Sutskever
I deeply regret my participation in the board’s actions. I never intended to harm OpenAI. I love everything we’ve built together and I will do everything I can to reunite the company.
Open AI Statement to Rehire Sam
We have reached an agreement in principle for Sam Altman to return to OpenAI as CEO with a new initial board of Bret Taylor (Chair), Larry Summers, and Adam D’Angelo.
We are collaborating to figure out the details. Thank you so much for your patience through this.
Sam Altman Returns
i love openai, and everything i’ve done over the past few days has been in service of keeping this team and its mission together. when i decided to join msft on sun evening, it was clear that was the best path for me and the team. with the new board and w satya’s support, i’m looking forward to returning to openai, and building on our strong partnership with msft.
Sam Altman and Greg Brockman Statements
we have more unity and commitment and focus than ever before.
we are all going to work together some way or other, and i’m so excited.
one team, one mission.
we are so back
the openai leadership team, particularly mira brad and jason but really all of them, have been doing an incredible job through this that will be in the history books.
incredibly proud of them.
Microsoft CEO Satya Nandella
We are encouraged by the changes to the OpenAI board. We believe this is a first essential step on a path to more stable, well-informed, and effective governance. Sam, Greg, and I have talked and agreed they have a key role to play along with the OAI leadership team in ensuring OAI continues to thrive and build on its mission. We look forward to building on our strong partnership and delivering the value of this next generation of AI to our customers and partners.
OpenAI: The Board
A Timeline of the OpenAI Board
https://loeber.substack.com/p/a-timeline-of-the-openai-board
Larry Summers – New OpenAI Board Member
More and more, I think ChatGPT is coming for the cognitive class. It’s going to replace what doctors do — hearing symptoms and making diagnoses — before it changes what nurses do — helping patients get up and handle themselves in the hospital.
Sam Altman’s back. Here’s who’s on the new OpenAI board and who’s out
OpenAI: News Coverage
A timeline of Sam Altman’s firing from OpenAI — and the fallout
https://techcrunch.com/2023/11/29/a-timeline-of-sam-altmans-firing-from-openai-and-the-fallout/
How a Fervent Belief Split Silicon Valley—and Fueled the Blowup at OpenAI (Paywall)
Sam Altman’s firing showed the influence of effective altruism and its view that AI development must slow down; his return marked its limits
https://www.wsj.com/tech/ai/openai-blowup-effective-altruism-disaster-f46a55e8
Read Microsoft’s internal memos about the chaos at OpenAI
https://www.theverge.com/2023/11/22/23972572/microsoft-internal-memo-kevin-scott-openai
Before Altman’s Ouster, OpenAI’s Board Was Divided and Feuding
Sam Altman confronted a member over a research paper that discussed the company, while directors disagreed for months about who should fill board vacancies.
The Sam Altman drama points to a deeper split in the tech world
On one side are the “doomers”, who believe that, left unchecked, ai poses an existential risk to humanity and hence advocate stricter regulations. Opposing them are “boomers”, who play down fears of an ai apocalypse and stress its potential to turbocharge progress.
https://finance.yahoo.com/news/sam-altman-drama-points-deeper-182335648.html
OpenAI’s board of directors approached Dario Amodei, the co-founder and CEO of rival large-language model developer Anthropic, about a potential merger of the two companies
https://www.theinformation.com/articles/openai-approached-anthropic-about-merger
Q* Super Intelligence: A Leading Theory Around Sam Altman (truth + hyperbole)
It’s pronounced Q-Star.
OpenAI researchers warned board of AI breakthrough ahead of CEO ouster, sources say
Altman Sought Billions For Chip Venture Before OpenAI Ouster
Altman was fundraising in the Middle East for new chip venture
The project, code-named Tigris, is intended to rival Nvidia
The news that OpenAI’s breakthrough involves something called Q* (Q star) suggests it’s related. Q-learning is a class of reinforcement learning and not new, however there’s been recent progress in combining Q-learning with transformers and LLMs. Tesla uses deep Q-learning for self-driving, for example. There’s even speculation that Google’s long-awaited Gemini model employs a version of it. Q-learning is a “model free” approach to RL as it can work even if the environment is complex and randomly changing, rather than requiring a set of well-defined rules like Chess. Q-learning is popular for single-agent games as, by default, it models other agents as simply features in its environment to navigate around, rather than as distinct agents with their own internal states. (Note this is also the basic definition of sociopathy.)
Finding Q* is equivalent to having the best possible Markov decision process. In other words, no matter what life throws your way, you always find a way to win.
What is OpenAI Project Q*? AGI Superintelligence Explained
With OpenAI’s leadership in chaos, here’s everything you need to know about the companies clandestine AGI project.
https://tech.co/news/what-is-openai-project-q-star-agi-superintelligence
Building Intelligent Q* Agents with Microsoft’s AutoGen: A Comprehensive Guide
Q* Did OpenAI Achieve AGI? OpenAI Researchers Warn Board of Q-Star | Caused Sam Altman to be Fired?
A day before Sam was fired: “Is this a tool we’ve built or a creature we have built?”
Q* Hypothesis | Is this a hybrid of GPT and AlphaGO? AI self-play and synthetic data.
OpenAI’s Q* is the biggest thing since Word2Vec
https://www.youtube.com/watch?v=3d0kk88IE8c
Audio
Eleven Labs
Unleash the power of our cutting-edge technology to generate realistic, captivating speech in a wide range of languages.
https://elevenlabs.io/speech-synthesis
ChatGPT Voice rolled out for all free users. Give it a try — totally changes the ChatGPT experience:
“ChatGPT with voice” opens up to everyone on iOS and Android
All Android and iOS users can soon tap a headphone icon and start chatting.
https://arstechnica.com/gadgets/2023/11/chatgpt-with-voice-opens-up-to-everyone-on-ios-and-android/
You can now transcribe 2.5 hours of audio in 98 seconds, locally. A new implementation called insanely-fast-whisper is blowing up on Github. It works on Mac or Nvidia GPUs and uses the Whisper + Pyannote library to speed up transcriptions and speaker segmentations.
https://www.linkedin.com/feed/update/urn:li:activity:7136072189933408256/
Native whisper.cpp server with OAI-like API is now available
This is a very convenient way to run an efficient local transcription service locally on any kind of hardware (CPU, GPU (CUDA or Metal) or ANE)
https://github.com/ggerganov/whisper.cpp/pull/1380
AI Voice to Instrument: Coolest thing of the week
Video
Incredible Video – Creatively Combining Multiple Tools
My favorite AI project this week comes from the incredibly talented Martin Haerlin and Hauke Hilberg. It’s still very early days for text-to-video and video-to-video… but this 50 sec clip really brings to life the potential and possibilities for generative AI in marketing.
Making of: https://www.instagram.com/p/C0cc8Tfs9gF/
Stable Video Diffusion
De-Aging Harrison Ford with Stable Diffusion + ControlNet + EbSynth + Fusion
https://www.linkedin.com/feed/update/urn:li:activity:7131657496791760896/
Turing Memes into Videos with Stable Video Diffusion
https://www.linkedin.com/feed/update/urn:li:activity:7134128925986762752/
Stability AI just released an open source image-to-video AI. The visual quality is a game-changer, this looks already better than RunwayML. Here’s everything you need to know (with stunning examples):
Introducing Stable Video Diffusion
Today, we are releasing Stable Video Diffusion, our first foundation model for generative video based on the image model Stable Diffusion.
Now available in research preview, this state-of-the-art generative AI video model represents a significant step in our journey toward creating models for everyone of every type.
https://stability.ai/news/stable-video-diffusion-open-ai-video-model
Animating Album Covers with Stable Video Diffusion
Here’s our latest LABs experiment where we ran several album covers through the newly released SVD model to bring them to life. Stable Video Diffusion (SVD) is a foundational model for generating videos from text prompts or a single image.
https://www.linkedin.com/posts/t-da_ai-svd-stablediffusion-activity-7135040448061972480-reRC
Stable Video Diffusion is very cool! We can feed the generated 3 multi-view frames (no poses needed) into our PF-LRM model and get 3D NeRF in ~1.3 seconds. Cool result!
This is image-to-video with Stable Video Diffusion (SVD). The temporal coherence is impressive!
FAL Stable Video Diffusion
Mid-Journey Motion Brush
AI Video Keeps Getting Better! Motion Brush is Here!
I’ve got a full video on Gen-2’s Motion Brush coming up later today, but I loved this example I generated for it, so here’s a little preview!
Apple
Apple is currently using LLM to completely revamp Siri into the ultimate virtual assistant and is preparing to develop it into Apple’s most powerful killer AI app.
This integrated development effort is actively underway, and the first product is expected to be unveiled at WWDC 2024, with plans for it to be standard on the iPhone 16 models and onwards.
Apple’s AI-powered Siri assistant could land as soon as WWDC 2024
Fresh rumors says it could be standard on the iPhone 16
https://www.techradar.com/phones/apple-plans-to-reinvent-siri-with-on-device-ai-for-the-iphone-16
Augmented Reality: Frontline Work
AI + Mixed Reality for the Frontline Workers | Copilot in Dynamics 365 Guides
AI Agents
AI Is Navigating the Web Even Without Formally Structured Content
The beginning of the end for Siri (and maybe for a lot of apps).
GPT-4V was able to navigate an iPhone screen using visuals alone, and had a 91% success rate when deciding what actions when given commands like “shop for a milk frother that costs between $50 and $100.” It was actually able to take the right action 75% of the time.
This is the AI-as-agent future that people are talking about. It does seem to be coming.
GPT-4 Turbo with Google Web Browsing (Assistants API)
Autonomous AI Video Analysis 2.0 | GPT-4V Turbo x Whisper
ICYMI Bard can now help with understanding YouTube videos
Very real application going into this weekend:
https://bard.google.com/share/ef6d680d9d95
If you want to get a glimpse of the future, even dimly, do this with your phone: have a shortcut to launch ChatGPT in voice mode. Pick a voice you like, first. It is still clunky, but you can absolutely see how it replaces the Siris of the world & also how AI agents might work.
screenshot-to-code:
upload a screenshot of any website, watch as AI progressively builds the html, iteratively improving the generated code by comparing it against the screenshot repeatedly.
Inflection 2
Thrilled to announce that Inflection-2 is now the 2nd best LLM in the world! It will be powering http://Pi.ai very soon. And available to select API partners in time. Tech report linked…Come run with us!
https://inflection.ai/inflection-2
Inflection AI introduces Inflection-2: The Next Step Up blog: https://inflection.ai/inflection-2
The new model, Inflection-2, is substantially more capable than Inflection-1, demonstrating much improved factual knowledge, better stylistic control, and dramatically improved reasoning.
Inflection-2 was trained on 5,000 NVIDIA H100 GPUs in fp8 mixed precision for ~10²⁵ FLOPs. This puts it into the same training compute class as Google’s flagship PaLM 2 Large model, which Inflection-2 outperforms on the majority of the standard AI performance benchmarks, including the well known MMLU, TriviaQA, HellaSwag & GSM8k.
Anthropic Claude 2.1
Anthropic: Introducing Claude 2.1
Our latest model, Claude 2.1, is now available over API in our Console and is powering our claude.ai chat experience. Claude 2.1 delivers advancements in key capabilities for enterprises—including an industry-leading 200K token context window, significant reductions in rates of model hallucination, system prompts and our new beta feature: tool use. We are also updating our pricing to improve cost efficiency for our customers across models.
https://www.anthropic.com/index/claude-2-1
Claude 2.1 (200K Tokens) – Pressure Testing Long Context Recall
We all love increasing context lengths – but what’s performance like?
Anthropic reached out with early access to Claude 2.1 so I repeated the “needle in a haystack” analysis I did on GPT-4
Claude 2.1 API pricing from AnthropicAI has been reduced to make it cheaper per-token than OpenAI GPT 4 Turbo.
Google Delays Launch of GPT-4 Rival Gemini
Google quietly open sourced a 1.6 trillion parameter MOE model
I have yet to use Bard extensions.
Amazon
Robots and Embodiment
Amazon just revealed they now have more than 750,000 robots deployed. 🤯
2013: 1,000
2014: 15,000
2017: 100,000
2019: 200,000
2021: 350,000
2022: 520,000
2023: 750,000
400,000 additional robotic units in roughly two years.
https://www.linkedin.com/feed/update/urn:li:activity:7121124632673300480/
Amazon LLM
Amazon reportedly training AI with twice as many parameters as gpt-4
https://futurism.com/the-byte/amazon-training-ai-twice-parameters-gpt-4
AI Taking Alexa Jobs
Amazon.com to cut ‘several hundred’ Alexa jobs. Amazon.com on Friday announced it is trimming jobs at its Alexa voice assistant unit, citing shifting business priorities and a greater focus on generative artificial intelligence.
https://www.reuters.com/technology/amazoncom-cut-several-hundred-alexa-jobs-2023-11-17/
Learning/Education
Andrej Karpathy Intro to LLMs
New YouTube video: 1hr general-audience introduction to Large Language Models
Based on a 30 min talk I gave recently; It tries to be a non-technical intro, covers mental models for LLM inference, training, finetuning, the emerging LLM OS and LLM Security.
IBM launches free Artificial Intelligence training courses
https://www.rte.ie/news/2023/1122/1417785-ibm-ai-training/
Amazon aims to provide free AI skills training to 2 million people by 2025 with its new ‘AI Ready’ commitment
https://www.aboutamazon.com/news/aws/aws-free-ai-skills-training-courses
Carnegie Learning Announces LiveHint AI™
https://finance.yahoo.com/news/carnegie-learning-announces-livehint-ai-170200091.html
An increasing number of papers are finding that AI is a reasonably good “grader” – often matching human evaluators. We don’t know all the biases these systems have, though, and this paper finds an interesting one: (older) AIs prefer their own writing most. https://arxiv.org/abs/2311.09766
Ethics/Legal
It is not an exaggeration to say that many (most?) people are never going to write their own first drafts ever again. Writing and our relationship to it is about to change.
These tools are all ones I have access to now, not some sort of mockup or future release.
Apollo Research
We think some of the greatest risks come from advanced AI systems that can evade standard safety evaluations by exhibiting strategic deception. Our goal is to understand AI systems well enough to prevent the development and deployment of deceptive AIs. We intend to develop a holistic and far-ranging model evaluation suite that includes behavioral tests, fine-tuning, and interpretability approaches to detect deception and potentially other misaligned behavior.
https://www.apolloresearch.ai/
Researchers develop AI-powered model to predict stock market trends
https://techxplore.com/news/2023-11-ai-powered-stock-trends.html
Bill Gates teases the possibility of a 3-day work week where ‘machines can make all the food and stuff’
https://fortune.com/2023/11/23/bill-gates-microsoft-3-day-work-week-machines-make-food/
Microsoft
Microsoft rebrands its AI-powered Bing Chat as Copilot
The company has also announced more Copilot AI features for its 365 apps.
https://www.engadget.com/microsoft-rebrands-its-ai-powered-bing-chat-as-copilot-160027250.html
One really interesting choice Microsoft made with Copilot was to include “coaching” as an option. I am not sure how many people will use it over just having the AI write content, but the advice is pretty solid, and shows an educational/mentoring use of AI to improve performance.
Health
BIG new study: A new AI cancer detection model trained to spot pancreatic malignancies—the most deadly solid cancer— outperformed expert radiologists
NeuralLink PRIME Study
Check out our latest video to learn more about our PRIME Study!
Machine Learning Methods May Improve Brain Tumor Characterization
Researchers combining mass spectrometry with machine learning to assess meningioma tumors found that the tools were 87% accurate in classifying tumor grade.
https://healthitanalytics.com/news/machine-learning-methods-may-improve-brain-tumor-characterization
AI in Newsrooms
Building AI-Powered Projects for Local Newsrooms
https://generative-ai-newsroom.com/building-ai-powered-projects-for-local-newsrooms-f00d105d05ff
Hardware/Chips
Nvidia’s revenue triples as AI chip boom continues
https://www.cnbc.com/2023/11/21/nvidia-nvda-q3-earnings-report-2024.html
Humane Pin
Video demonstrating the AI Pin was recently posted on the Humane discord by Imran Chaudhri.
3D Rendering
Gaussian Splatting is incredible. You can now capture anything in 3D from your phone with Luma, import it into Spline, crop, adjust, and embed on your site.
Applying textures generated by Stable Diffusion in real time, locally, in a 3D environment
Finally had a chance to play with LCM-Lora and real time generative AI and it lives up to the hype.
Text-to-3D is becoming amazing. A team just designed an entire miniature world using nothing but Luma AI’s text-to-3D tool and Blender’s magic.
Development/Technical Stories
Towards Accurate AI: Complementary Methods for RAG systems
https://ai.plainenglish.io/towards-accurate-ai-complementary-methods-for-rag-systems-584554106d44
Blending Fine-Tuning and RAG for Collaborative Filtering with LLMs
Deci just released a game changer in Pose Estimation: YOLO-NAS Pose.
Their new open-source model combines object detection and keypoint recognition and already beats YOLOv8 variants in accuracy and performance benchmarks.
This paper introduces a novel prompting technique called XOT that unlocks the potential of large language models (LLMs) like GPT-3 and GPT-4 for complex problem-solving
Put very simply, S2A is about prompting a model to rephrase a query as a question based on only the important information in its context.
https://www.linkedin.com/feed/update/urn:li:activity:7133050136833736704/
I’ve discovered how to use ChatGPT’s Code Interpreter to create functional, small-scale Neural Networks, a truly amazing feat in the area of GPT data analysis
https://www.linkedin.com/feed/update/urn:li:activity:7134572619726557184/
Increasing signs that synthetic, AI-created data is beneficial to training LLMs under some circumstances, first with Code Llama, now with Orca 2
(The animal names have to stop, folks. The only thing worse than a misaligned AGI is a misaligned AGI named Fluffy Fast WaLLaby-3)
Leveraging LLMs, we train LEO with real and synthetic 3D data across a diverse spectrum of tasks. It’s thrilling to see LEO surpass current state-of-the-art SOTA methods in most benchmarked tasks, all under a single, unified model.
Introducing the new Amazon EC2 DL2q instance – cost-efficient, high-performance AI inference now available!
Eye On AI: Bain Capital Ventures Launches BCV Labs In Search Of New AI Deals
https://news.crunchbase.com/ai/bain-capital-launches-bcv-labs-startup-venture
AGI’s Impact on Tech, SaaS Valuations
https://nextword.substack.com/p/agis-impact-on-tech-saas-valuations
Stable Video Diffusion Image-to-Video Model Card
https://huggingface.co/stabilityai/stable-video-diffusion-img2vid
ChatGPT is getting much better at word games and acrostics (see the first letter of each sentence, which it came up with itself), this used to be a weak spot for GPT-4.
Model Context Windows Chart
OpenAI has put ChatGPT Plus sign-ups on pause
After announcing premium-tier users can build their own chatbots, CEO Sam Altman says its Plus subscription has exceeded capacity
https://qz.com/openai-has-put-chatgpt-plus-sign-ups-on-pause-1851025002
Explaining the SDXL latent space
https://huggingface.co/blog/TimothyAlexisVass/explaining-the-sdxl-latent-space
A useful reading for anyone interested in AI risks is the official GPT-4 system card from back in March. It can help you understand what an advanced LLM without guardrails can do, since they give “before” and “after” examples. Appendix D has details.
As suspected, OAI invented a way to overcome training data limitations with synthetic data. When trained with enough examples, models begin to generalize nicely! Great news for open source and decentralized AI – we are no longer beholden to the data rich companies
People are just starting to realize the power of Local LLMs. Especially with new Apple chips. It’s a game-changer. Let me show you:
2024 will be a breakout year for knowledge graphs and ontological engineering in general.
Credits/Sources
Most of these links come from just a few incredible sources. Please follow them:
- Robert Scoble – https://x.com/Scobleizer
- Ethan Mollick – https://www.linkedin.com/in/emollick/
- David Armano – https://www.linkedin.com/in/darmano/
- Alan Thompson – https://lifearchitect.ai/
- Theoretically Media – https://www.youtube.com/@TheoreticallyMedia
- The Rundown – https://www.therundown.ai/
- Borriss – https://twitter.com/_Borriss_
- Bilawal Sidhu – https://twitter.com/bilawalsidhu/
- TLDR – https://tldr.tech/ai





Leave a Reply