This week’s cover image reflects the PR drama surrounding Google manually modifying the Gemini product demo to improve performance. Generated by Dalle-3, upscaled with Magnific AI, and layered in Photoshop.
Executive Summary
- Google Gemini: The top story is the release of Google Gemini, which promises to compete with, if not eclipse, OpenAI in 2024.
- Google AlphaCode2: Google also released AlphaCode2, a computer-coding AI that can beat 85% of human competitive programmers.
- Video: AI generated video continues to improve very rapidly, and this week there are several amazing demo videos worth watching.
- Image Upscaling: Image upscaling boosts AI imagery to photographic realism.
- Meta Imagine: Released their own text to image site, called Imagine, to compete with MidJourney and Dalle
- Slow Down Ahead?: A few leading minds are starting to wonder if the next phase of AI will be more difficult to reach than previously thought. This implies that OpenAI may lose its wide advantage as competitors plateau.
- Elon’s AI: X’s AI, called Grok, is now available to all paying users and integrates with X.
- Apple: Apple is doubling down on AI, and many think iPhones will soon be able to run AI locally on the device.
Top 8 Stories
These are the links to click if you only pick a few. Even if they look boring, click them! I did the work, so you don’t have to worry. All are 10/10 would recommend.
- Upscaling: Grand Theft Auto Example Is Incredible
- Animating Still Photos
- Hot Swapping Video Elements on the Fly with Pika
- Turning Harry Potter into Anime in Real Time Using Latent Consistency
- Meta’s AI image generator is available as a standalone website
- Google unveils AlphaCode 2, outcodes 85% of human competition
- Pilotless FedEx, Reliable Robotics Plane Completes Flight
- Google Gemini Demo
The Rest: AI News of The Week
Don’t let the volume overwhelm you. Have fun and skim it. The links are organized by topic, sorted from ‘coolest’ to ‘least cool’, and each topic is clearly defined with a headline. I do the work so you don’t have to! The links descriptions are often pulled directly from tweets or articles, so it’s not always my voice. Pause when you see something that interests you. Reach out to me any time. I enjoy sharing and discussing these items!
AVideo
Image to Video – People – (Magic Animate, Animate Anyone, HumanAIGC)
HumanAIGC – Strong Example of Image -> Video
Magic Animate Examples: Less than 48 hours since MagicAnimate’s public launch
- Dancing statue- https://twitter.com/HirokaKoizumi/status/1731888505192685865
- MidJourney TikTok star – https://twitter.com/blizaine/status/1731863651013845290
- Another example – https://twitter.com/thibaudz/status/1731995356546630085
- Leaping woman from pose – https://twitter.com/toyxyz3/status/1732092223515394424
MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model
https://huggingface.co/spaces/zcxu-eric/magicanimate
‘Animate Anyone’ heralds the approach of full-motion deepfakes
https://techcrunch.com/2023/12/04/animate-anyone-heralds-the-approach-of-full-motion-deepfakes/
Alibaba’s ‘Animate Anyone’ Is Trained on Scraped Videos of Famous TikTokers
https://www.404media.co/alibaba-animate-anyone-ai-generated-tiktok/
Text to Video: Pika is the star of December
Mastering Pika 1.0 – Tutorial & Look at the New AI Video Generator
Must see: hot swapping environments and people in a video
Pika Examples
- Gannet on the beach: https://twitter.com/TomLikesRobots/status/1732360554239381969
- Rabbit in the jungle: https://twitter.com/MatanCohenGrumi/status/1730064185714004293
- Pika in-painting: https://twitter.com/mrjonfinger/status/1732198293353152689
- People: https://twitter.com/MatthieuGB/status/1732354045359095951
- Kid in a forest watching a UFO: https://twitter.com/DaveJWVillalva/status/1732159746382409907
- Flower bud: https://twitter.com/chaseleantj/status/1732333917288685641
Runway
Runway partners with Getty Images to build enterprise ready AI tools
https://runwayml.com/blog/runway-partners-with-getty-images/
Latent Consistency
Turning Harry Potter into Anime in Real Time
ComfyUI LCM+Stable fast real time test
Video Creation Workflows/How-To
“As promised, here was my process with this Attack on Titan animation (with motion capture)”
“I created a synthetic character using MidJourney, then used Face-fusion to wrap its face on my face as I mimed the lines to a prerecorded Elevenlabs track. I then placed it back on its Runway animated body using After Effects.“
4k result: https://www.youtube.com/watch?v=NRVh6Cjd-Vo
Other Video News
Relightable Gaussian Codec Avatars from Meta
https://shunsukesaito.github.io/rgca/
Unreal Engine 5 Powered Coffee Foam Generator Looks Amazing
https://80.lv/articles/this-ue5-powered-coffee-foam-generator-looks-amazing/
AI Video News – Weekly show
AI Images
Upscaling
Upscaling PlayStation1 Lara Croft to photorealism (revisited)
https://twitter.com/javilopen/status/1730987030971142519/photo/1
https://twitter.com/javilopen/status/1730987030971142519/photo/2
The Best AI Upscaler Makes GTA 6 Look AMAZING
Upscaling Grand Theft Auto
Upscaling emojis into photographs
Upscaling MidJourney
The quality of these AI generated nature scenes is incredible.
“Half-Life: Opposing Force gets a visual upgrade with generative AI. Creative upscaling tools like Magnific and Krea are a lot of fun! Pretty much an “enhance 3D render” button. How soon until this tech is running in realtime on your GPU? It’d be like NVIDIA DLSS on steroids. The key to continuity in AI images is getting closer!
Meta’s Releases Text to Image Engine to Compete with Dalle and MidJourney
Meta’s AI image generator is available as a standalone website
Update on Meta AI’s 20 new features
Meta will let you ‘reimagine’ your friends’ AI-generated images
https://www.theverge.com/2023/12/6/23990896/meta-ai-reimagine-images-chatbot
Consistent Characters
This is a demo of Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models.
https://huggingface.co/spaces/baulab/ConceptSliders
Kandisky 3.0
“We present Kandinsky 3.0, a large-scale text-to-image generation model based on latent diffusion, continuing the series of text-to-image Kandinsky models and reflecting our progress to achieve higher quality and realism of image generation.”
https://ai-forever.github.io/Kandinsky-3/
https://fusionbrain.ai/editor/
Google/AlphaCode2
Google’s AlphaCode 2 beats 85% of competitive programers and solves 1.7x more problems
https://storage.googleapis.com/deepmind-media/AlphaCode2/AlphaCode2_Tech_Report.pdf
AlphaCode 2 is the hidden champion of Google’s Gemini project
https://the-decoder.com/alphacode-2-is-the-hidden-champion-of-googles-gemini-project
Google unveils AlphaCode 2, powered by Gemini
In a subset of programming competitions hosted on Codeforces, a platform for programming contests, AlphaCode 2 — coding in languages spanning Python, Java, C++ and Go — performed better than an estimated 85% of competitors on average, according to Google.
https://techcrunch.com/2023/12/06/deepmind-unveils-alphacode-2-powered-by-gemini/
Google/Gemini
Side Story: Benchmark Scrutiny Dominates Discussion
Google’s Gemini Looks Remarkable, But It’s Still Behind OpenAI
“The tech giant’s latest AI model is only marginally better than the one from OpenAI that’s been out for eight months.”
How Google’s Gemini video really worked
Google’s best Gemini demo was faked
https://techcrunch.com/2023/12/07/googles-best-gemini-demo-was-faked/
Main Story: Gemini Launch and Video Walk-Throughs
Gemini overview page
https://deepmind.google/technologies/gemini/
Introducing Gemini: our largest and most capable AI model
https://blog.google/technology/ai/google-gemini-ai/
Google launches Gemini—a powerful AI model it says can surpass GPT-4
Google claims Gemini beats GPT-4 in “30 of the 32 widely used academic benchmarks.”
Gemini:AFamilyofHighlyCapable MultimodalModels (PDF overview)
https://storage.googleapis.com/deepmind-media/gemini/gemini_1_report.pdf
“Actually the Blue Duck Google demo is a very big deal. It shows how AI agents will work. Instead of your physical desk, they’ll see your desktop. Be an expert at all your apps, walk you through your work and learning, click the buttons and much more. Working on a computer will be a completely different experience in just 1-2 years. It will much more than 10x your productivity. I think it has the potential to be a great era for everyone, who can come up with jobs for themselves.”
How it’s Made: Interacting with Gemini through multimodal prompting
https://developers.googleblog.com/2023/12/how-its-made-gemini-multimodal-prompting.html
“Less noticed in today’s Google Bard news today was the fact that it was trained and is running on a new version of the company’s homegrown TPU chips for AI. Also there’s a new AI Hypercomputer.”
Enabling next-generation AI workloads: Announcing TPU v5p and AI Hypercomputer
Gemini Nano runs on a phone and without the internet, beginning an era of on-device LLMs, one that fits in your pocket
Google has quietly pushed back the launch of next-gen AI model Gemini until next year, report says (there are multiple “Gemini”s and the strongest one is delayed).
https://www.businessinsider.com/google-delays-launch-of-gemini-ai-to-early-2024-report-2023-12
Google Gemini: Explainer Videos by Topic
Full Playlist
- Intro to Gemini: Gemini: Google’s newest and most capable AI model: https://youtu.be/jV1vkHv4zq8?si=jPXZIfw-asa6unRr
- Hands-on with Gemini: Interacting with multimodal AI:
- https://youtu.be/UIZAiXYceBI?si=643AzNF1YbZbGevQ
- Gemini: Unlocking insights in scientific literature: https://youtu.be/sPiOP_CB54A?si=qHyvZ1JpdZk3JO_w
- Gemini: Processing and understanding raw audio: https://youtu.be/D64QD7Swr3s?si=oQZuWbmFYkWvy_1J
- Gemini: Excelling at competitive programming:
- https://youtu.be/LvGmVmHv69s?si=7P_a393yYDQ3SOt0
- Gemini: Reasoning about user intent to generate bespoke experiences:
- https://youtu.be/v5tRc_5-8G4?si=kTJkc_E5c9lQTRwe
- Gemini: Explaining reasoning in math and physics:
- https://youtu.be/K4pX1VAxaAI?si=4-T3C2nphVEZe3L1
- Testing Gemini: Finding connections:
- https://youtu.be/Rn30RMhEBTs?si=J0YofZ1B3amIWG2_
- Testing Gemini: Guess the movie:
- https://youtu.be/aRyuMNwn02w?si=nKImdTX1bvcDQyFa
- Testing Gemini: Emoji Kitchen:
- https://youtu.be/ki8kRJPXCW0?si=PayQPGlhSm71nDFp
- Testing Gemini: Fit check: https://youtu.be/HP2pNdCRT5M?si=46-FVaAAdqNrv0bv
- Testing Gemini: Turning images into code:
- https://youtu.be/NHLnjWTEZps?si=QI1E9Gl3gAnKsY5N
- Testing Gemini: Understanding environments:
- https://youtu.be/JPwU1FNhMOA?si=CiJKdn0Br4SH8UBm
Multimodality/Vision
People are underestimating what GPT-4V can already do. Using a sign with half-obscured text, it guessed the location of a trip to Hershey amusement park, tracked who was in which picture, figured out the context, and made inferences about the sequence of events and activities.
LLM writes humorous captions for any photo
https://zhongshsh.github.io/CLoT/
Robotics/Embodiment
Pilotless FedEx, Reliable Robotics Plane Completes Flight
https://www.ttnews.com/articles/pilotless-fedex-plane
Josh Bongard: The roboticist who wants to bring AI into contact with the real world
This cyborg cockroach could be the future of earthquake search and rescue
https://www.nature.com/articles/d41586-023-03801-0
Apple
Apple released an ML framework for Apple Silicon, finally. MLX is an efficient machine learning framework specifically designed for Apple silicon (i.e. your laptop!) This may be Apple’s biggest move on open-source AI so far: MLX, a PyTorch-style NN framework optimized for Apple Silicon, e.g. laptops with M-series chips.
Science/Education
‘Google for wildlife sounds’: Australian conservation research gets an AI boost
Researchers can upload their recordings, and match them to bird calls from around the country.
SchoolXpress turns your books, handwritten notes, classwork and documents into interactive learning content that simplifies concepts and delivers on your learning objectives. https://www.schoolxpress.ai
Sperm whales have equivalents to human vowels.
We uncovered spectral properties in whales’ clicks that are recurrent across whales, independent of traditional types, and compositional.
We got clues to look into spectral properties from our AI interpretability technique CDEV.
Absci Announces Collaboration with AstraZeneca to Advance AI-Driven Oncology Candidate
https://finance.yahoo.com/news/absci-announces-collaboration-astrazeneca-advance-123000837.html
OpenAI
New report illuminates why OpenAI board said Altman “was not consistently candid”
https://arstechnica.com/ai/2023/12/openai-board-reportedly-felt-manipulated-by-ceo-altman
The Inside Story of Microsoft’s Partnership with OpenAI
https://www.newyorker.com/magazine/2023/12/11/the-inside-story-of-microsofts-partnership-with-openai
Elon Musk told OpenAI to move faster right before he left the company in 2018: NYT
https://www.businessinsider.com/elon-musk-told-openai-to-move-faster-before-he-left-2023-12
The OpenAI Board Member Who Clashed With Sam Altman Shares Her Side
https://www.wsj.com/tech/ai/helen-toner-openai-board-2e4031ef
OpenAI’s GPT store delayed to next year
https://www.theverge.com/2023/12/1/23984497/openai-gpt-store-delayed-ai-gpt
OpenAI Agreed to Buy $51 Million of AI Chips From a Startup Backed by CEO Sam Altman
https://www.wired.com/story/openai-buy-ai-chips-startup-sam-altman/
OpenAI employees really, really did not want to go work for Microsoft
https://www.businessinsider.com/openai-employees-did-not-want-to-work-for-microsoft-2023-12
Amazon
Amazon’s Q generative AI chatbot allegedly leaks location of AWS data centers – report
Business/Enterprise
McKinsey Sees AI Adding Up to $340 Billion to Wall Street Profit
McDonald’s taps Google for ‘Ask Pickles’ AI chatbot to help fix ice cream machines
Called “Ask Pickles,” the bot will be trained on everything from manuals to data generated by equipment at restaurants. It will give workers guidance on the spot, potentially boosting productivity in an industry where every second counts.
https://finance.yahoo.com/news/mcdonalds-taps-google-ask-pickles-183450432.html
https://www.theverge.com/2023/12/6/23990900/mcdonalds-google-ai-cloud-generative
Solve Intelligence helps attorneys draft patents for IP analysis and generation
https://www.solveintelligence.com/
A Bearish POV Emerges
The step from GPT4 to GPT5 might be so difficult that there is an AI bubble.
“Calling it now: The $86B OpenAI tender will someday be seen as the WeWork moment of AI.
GPT-5 will either be significantly delayed or not meet expectations.
Companies will struggle to put GPT-4 and 5 into production (see below)
Competition will increase, margins will be thin
The profits won’t justify the valuation, esp after MSFT’s hefty cut is taken out.”
Google Gemini seems to have by many measures matched (or slightly exceeded) GPT-4, but not to have blown it away.
From a commercial standpoint GPT-4 is no longer unique. That’s a huge problem for OpenAI, especially post drama, when many customers are now seeking a backup plan.
From a technical standpoint, the key question is: are LLMs close to a plateau?
Note that Gates and Altman have both been dropping hints, and GPT-5 isn’t here after a year despite immense commercial desire. The fact that Google, with all its resources, did NOT blow away GPT-4 could be telling.
Twitter/X/Grok
Grok officially launches to all (thread)
Elon Musk’s AI firm xAI files to raise up to $1 billion in equity offering
https://finance.yahoo.com/news/elon-musks-xai-files-raise-193558630.html
Legal/Ethics
A game to try to get an LLM to give away sensitive information
Evaluating and Mitigating Discrimination in Language Model Decisions
https://www.anthropic.com/index/evaluating-and-mitigating-discrimination-in-language-model-decisions
IBM and Meta Launch the AI Alliance in collaboration with over 50 Founding Members and Collaborators globally including AMD, Anyscale, CERN, Cerebras, Cleveland Clinic, Cornell University, Dartmouth, Dell Technologies, EPFL, ETH, Hugging Face, Imperial College London, Intel, INSAIT, Linux Foundation, MLCommons, MOC Alliance operated by Boston University and Harvard University, NASA, NSF, Oracle, Partnership on AI, Red Hat, Roadzen, ServiceNow, Sony Group, Stability AI, University of California Berkeley, University of Illinois, University of Notre Dame, The University of Tokyo, Yale University and others
https://ai.meta.com/blog/ai-alliance/
Uber Eats is using AI for pictures of food. It doesn’t know that “pie” means pizza, and it invented a brand of ranch dressing called “Lelnach”
Announcing Purple Llama: Towards open trust and safety in the new world of generative AI
https://ai.meta.com/blog/purple-llama-open-trust-safety-generative-ai/
G42’s Ties To China Run Deep
More details on the Emirati technology firm at the nexus of China and the U.S.’s technological and geopolitical rivalry.
https://www.thewirechina.com/2023/12/03/g42s-ties-to-china-run-deep-g42-peng-xiao/
Microsoft
Microsoft’s Copilot is getting OpenAI’s latest models and a new code interpreter
Audio
Google’s new AI experiment composes abstract musical clips inspired by instruments
https://artsandculture.google.com/experiment/8QFo2oQr2uT3pg
SeamlessExpressive, a new AI translation model by research teams at Meta, enables high-quality speech translation that maintains the speaker’s vocal style, tone and unique expressions in translated outputs.
Try the demo with your own voice
https://seamless.metademolab.com/expressive
Technical/Dev/IT
Liquid AI, a new MIT spinoff, wants to build an entirely new type of AI
“Liquid neural networks consist of “neurons” governed by equations that predict each individual neuron’s behavior over time, like most other modern model architectures. The “liquid” bit in the term “liquid neural networks” refers to the architecture’s flexibility; inspired by the “brains” of roundworms, not only are liquid neural networks much smaller than traditional AI models, but they require far less compute power to run”
“GPT-3 contains about 175 billion parameters and ~50,000 neurons — “parameters” being the parts of the model learned from training data that essentially define the skill of the model on a problem (in GPT-3’s case generating text). By contrast, a liquid neural network trained for a task like navigating a drone through an outdoor environment can contain as few as 20,000 parameters and fewer than 20 neurons.”
An Opinionated Guide to Which AI to Use: ChatGPT Anniversary Edition
https://www.oneusefulthing.org/p/an-opinionated-guide-to-which-ai
Quantifying ChatGPT’s gender bias
https://www.aisnakeoil.com/p/quantifying-chatgpts-gender-bias
“Quadratic attention has been indispensable for information-dense modalities such as language… until now. Announcing Mamba: a new SSM arch. that has linear-time scaling, ultra long context, and most importantly–outperforms Transformers everywhere we’ve tried.”
“VideoRF: Rendering Dynamic Radiance Fields as 2D Feature Video Streams”
A few months ago, Meta released Segment Anything Model (SAM). It’s already in photography apps, medicine, and video-generation. Now we’ve invented SAM’s little brother, EfficientSAM. It is small but mighty! With 20x fewer parameters and 20x faster runtime, EfficientSAM is within 2 points (44.4 AP vs 46.5 AP) of the original SAM model.
How’d we do it? I have 2 words for you: Masked Autoencoders. Check out the following for more details!
Excited to share ReconFusion! 3D reconstruction of real-world scenes from only a few photos, powered by diffusion priors
LLMs fine-tuned with RLHF are known to be poorly calibrated. We found that they can actually be quite good at *verbalizing* their confidence.
HybridNeRF: Efficient Neural Rendering via Adaptive Volumetric Surfaces
LooseControl: Lifting ControlNet for Generalized Depth Conditioning
Mistral AI just dropped a mixtral-8x7b-32kseqle model as a torrent link! 32K context, 8x mixture of experts.
How Aws Can Undercut Nvidia With Homegrown Ai Compute Engines
Ai Model Visualization Tool
OneLLM: One Framework to Align All Modalities with Language
Long Context Prompting For Claude 2.1
Claude 2.1’s performance when retrieving an individual sentence across its full 200K token context window. This experiment uses a prompt technique to guide Claude in recalling the most relevant sentence.
https://www.anthropic.com/index/claude-2-1-prompting
Introducing Stable LM Zephyr 3B: A New Addition to Stable LM, Bringing Powerful LLM Assistants to Edge Devices
Stable LM Zephyr 3B is a 3 billion parameter Large Language Model (LLM), 60% smaller than 7B models, allowing accurate, and responsive output on a variety of devices without requiring high-end hardware.
https://stability.ai/news/stablelm-zephyr-3b-stability-llm
Inside Apple’s chip lab, home to the most ‘profound change’ at the company in decades
OpenAI Rival Mistral Nears $2 Billion Valuation With Andreessen Horowitz Backing
https://archive.is/4F3dT#selection-4559.0-4566.0
GPT4
Bing is a little confusing, it only uses GPT-4 in Creative (Purple) or Precise (Green) mode. Balanced (Blue) does not use GPT-4. If you are doing web searches, Precise mode will give you less hallucination. If you are doing anything else, use Creative Mode.
Self-operating-computer + AI to AI Simulation Test Project feat Q*
“Today we check out the open source project the Self-operating-computer and a system where 2 AIs simulate a conversation discussing a uploaded image.”
5 LLM Security Threats- The Future of Hacking?
“Today we check what could be the future of hacking and LLM attacks with Jailbreaks and Prompt Injections on LLMs and Multimodal Models”
Sam Altman’s Brain Chips | Rain Neuromorphic Chips | UAE Funds and US National Security and Q*
Credits/Sources
Most of these links come from just a few incredible sources. Please follow them:
- Robert Scoble: https://x.com/Scobleizer
- Ethan Mollick: https://www.linkedin.com/in/emollick/
- David Armano: https://www.linkedin.com/in/darmano/
- Alan Thompson: https://lifearchitect.ai/
- Theoretically Media: https://www.youtube.com/@TheoreticallyMedia
- The Rundown: https://www.therundown.ai/
- Borriss: https://twitter.com/_Borriss_
- Bilawal Sidhu: https://twitter.com/bilawalsidhu/
- TLDR: https://tldr.tech/ai
- Jeremiah Owyang: https://twitter.com/jowyang
- Wes Roth: https://www.youtube.com/@WesRoth





Leave a Reply