This week’s cover image reflects the PR drama surrounding Google manually modifying the Gemini product demo to improve performance. Generated by Dalle-3, upscaled with Magnific AI, and layered in Photoshop.

Executive Summary

  • Google Gemini: The top story is the release of Google Gemini, which promises to compete with, if not eclipse, OpenAI in 2024.  
  • Google AlphaCode2: Google also released AlphaCode2, a computer-coding AI that can beat 85% of human competitive programmers.
  • Video: AI generated video continues to improve very rapidly, and this week there are several amazing demo videos worth watching.  
  • Image Upscaling: Image upscaling boosts AI imagery to photographic realism.  
  • Meta Imagine: Released their own text to image site, called Imagine, to compete with MidJourney and Dalle
  • Slow Down Ahead?: A few leading minds are starting to wonder if the next phase of AI will be more difficult to reach than previously thought.  This implies that OpenAI may lose its wide advantage as competitors plateau.  
  • Elon’s AI: X’s AI, called Grok, is now available to all paying users and integrates with X.  
  • Apple: Apple is doubling down on AI, and many think iPhones will soon be able to run AI locally on the device.

Top 8 Stories

These are the links to click if you only pick a few.  Even if they look boring, click them!  I did the work, so you don’t have to worry.  All are 10/10 would recommend.

The Rest: AI News of The Week

Don’t let the volume overwhelm you.  Have fun and skim it. The links are organized by topic, sorted from ‘coolest’ to ‘least cool’, and each topic is clearly defined with a headline.  I do the work so you don’t have to!   The links descriptions are often pulled directly from tweets or articles, so it’s not always my voice.  Pause when you see something that interests you.  Reach out to me any time.  I enjoy sharing and discussing these items!

AVideo 

Image to Video – People –  (Magic Animate, Animate Anyone, HumanAIGC)

HumanAIGC – Strong Example of Image -> Video

Magic Animate Examples: Less than 48 hours since MagicAnimate’s public launch

MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model

https://huggingface.co/spaces/zcxu-eric/magicanimate

‘Animate Anyone’ heralds the approach of full-motion deepfakes

https://techcrunch.com/2023/12/04/animate-anyone-heralds-the-approach-of-full-motion-deepfakes/

Alibaba’s ‘Animate Anyone’ Is Trained on Scraped Videos of Famous TikTokers

https://www.404media.co/alibaba-animate-anyone-ai-generated-tiktok/

Text to Video: Pika is the star of December

Mastering Pika 1.0 – Tutorial & Look at the New AI Video Generator

https://www.youtube.com/watch?v=cte-UPyrhIs

Must see: hot swapping environments and people in a video

Pika Examples

Runway

Runway partners with Getty Images to build enterprise ready AI tools

https://runwayml.com/blog/runway-partners-with-getty-images/

Latent Consistency 

Turning Harry Potter into Anime in Real Time

ComfyUI LCM+Stable fast real time test

Video Creation Workflows/How-To

“As promised, here was my process with this Attack on Titan animation (with motion capture)”

“I created a synthetic character using MidJourney, then used Face-fusion to wrap its face on my face as I mimed the lines to a prerecorded Elevenlabs track.  I then placed it back on its Runway animated body using After Effects.“

4k result: https://www.youtube.com/watch?v=NRVh6Cjd-Vo 

Other Video News

Relightable Gaussian Codec Avatars from Meta

https://shunsukesaito.github.io/rgca/

Unreal Engine 5 Powered Coffee Foam Generator Looks Amazing

https://80.lv/articles/this-ue5-powered-coffee-foam-generator-looks-amazing/

AI Video News – Weekly show  

AI Images

Upscaling

Upscaling PlayStation1 Lara Croft to photorealism (revisited)

https://twitter.com/javilopen/status/1730987030971142519/photo/1

https://twitter.com/javilopen/status/1730987030971142519/photo/2

The Best AI Upscaler Makes GTA 6 Look AMAZING

https://youtu.be/VCP1R5-Zywc?si=bPDAQ6Eckjdwszl-

Upscaling Grand Theft Auto 

Upscaling emojis into photographs

https://www.linkedin.com/posts/eric-vyacheslav-156273169_this-is-incredible-ai-can-now-transform-activity-7140706133148688384-lcGQ

Upscaling MidJourney

The quality of these AI generated nature scenes is incredible.

“Half-Life: Opposing Force gets a visual upgrade with generative AI.  Creative upscaling tools like Magnific and Krea are a lot of fun! Pretty much an “enhance 3D render” button.  How soon until this tech is running in realtime on your GPU? It’d be like NVIDIA DLSS on steroids. The key to continuity in AI images is getting closer! 

Meta’s Releases Text to Image Engine to Compete with Dalle and MidJourney

Meta’s AI image generator is available as a standalone website

https://imagine.meta.com/

https://www.engadget.com/metas-ai-image-generator-is-available-as-a-standalone-website-185953058.html

Update on Meta AI’s 20 new features

Meta will let you ‘reimagine’ your friends’ AI-generated images

https://www.theverge.com/2023/12/6/23990896/meta-ai-reimagine-images-chatbot

Consistent Characters

This is a demo of Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models.

https://huggingface.co/spaces/baulab/ConceptSliders

https://sliders.baulab.info/

Kandisky 3.0 

“We present Kandinsky 3.0, a large-scale text-to-image generation model based on latent diffusion, continuing the series of text-to-image Kandinsky models and reflecting our progress to achieve higher quality and realism of image generation.”

https://ai-forever.github.io/Kandinsky-3/

https://fusionbrain.ai/editor/

Google/AlphaCode2

Google’s AlphaCode 2 beats 85% of competitive programers and solves 1.7x more problems

https://storage.googleapis.com/deepmind-media/AlphaCode2/AlphaCode2_Tech_Report.pdf

AlphaCode 2 is the hidden champion of Google’s Gemini project 

https://the-decoder.com/alphacode-2-is-the-hidden-champion-of-googles-gemini-project

Google unveils AlphaCode 2, powered by Gemini

In a subset of programming competitions hosted on Codeforces, a platform for programming contests, AlphaCode 2 — coding in languages spanning Python, Java, C++ and Go — performed better than an estimated 85% of competitors on average, according to Google.

https://techcrunch.com/2023/12/06/deepmind-unveils-alphacode-2-powered-by-gemini/

Google/Gemini

Side Story: Benchmark Scrutiny Dominates Discussion

Google’s Gemini Looks Remarkable, But It’s Still Behind OpenAI

“The tech giant’s latest AI model is only marginally better than the one from OpenAI that’s been out for eight months.”

https://www.bloomberg.com/opinion/articles/2023-12-07/google-s-gemini-ai-model-looks-remarkable-but-it-s-still-behind-openai-s-gpt-4

How Google’s Gemini video really worked

Google’s best Gemini demo was faked

https://techcrunch.com/2023/12/07/googles-best-gemini-demo-was-faked/

Main Story: Gemini Launch and Video Walk-Throughs

Gemini overview page

https://deepmind.google/technologies/gemini/

Introducing Gemini: our largest and most capable AI model

https://blog.google/technology/ai/google-gemini-ai/

Google launches Gemini—a powerful AI model it says can surpass GPT-4

Google claims Gemini beats GPT-4 in “30 of the 32 widely used academic benchmarks.”

https://arstechnica.com/information-technology/2023/12/google-launches-gemini-a-powerful-ai-model-it-says-can-surpass-gpt-4

Gemini:AFamilyofHighlyCapable MultimodalModels (PDF overview)

https://storage.googleapis.com/deepmind-media/gemini/gemini_1_report.pdf

“Actually the Blue Duck Google demo is a very big deal. It shows how AI agents will work. Instead of your physical desk, they’ll see your desktop. Be an expert at all your apps, walk you through your work and learning, click the buttons and much more. Working on a computer will be a completely different experience in just 1-2 years. It will much more than 10x your productivity. I think it has the potential to be a great era for everyone, who can come up with jobs for themselves.”

How it’s Made: Interacting with Gemini through multimodal prompting

https://developers.googleblog.com/2023/12/how-its-made-gemini-multimodal-prompting.html

“Less noticed in today’s Google Bard news today was the fact that it was trained and is running on a new version of the company’s homegrown TPU chips for AI. Also there’s a new AI Hypercomputer.”

Enabling next-generation AI workloads: Announcing TPU v5p and AI Hypercomputer

https://cloud.google.com/blog/products/ai-machine-learning/introducing-cloud-tpu-v5p-and-ai-hypercomputer

Gemini Nano runs on a phone and without the internet, beginning an era of on-device LLMs, one that fits in your pocket

Google has quietly pushed back the launch of next-gen AI model Gemini until next year, report says (there are multiple “Gemini”s and the strongest one is delayed).

https://www.businessinsider.com/google-delays-launch-of-gemini-ai-to-early-2024-report-2023-12

Google Gemini: Explainer Videos by Topic

Full Playlist

Multimodality/Vision

People are underestimating what GPT-4V can already do.  Using a sign with half-obscured text, it guessed the location of a trip to Hershey amusement park, tracked who was in which picture, figured out the context, and made inferences about the sequence of events and activities.

LLM writes humorous captions for any photo

https://zhongshsh.github.io/CLoT/

Robotics/Embodiment

Pilotless FedEx, Reliable Robotics Plane Completes Flight

https://www.ttnews.com/articles/pilotless-fedex-plane

Josh Bongard: The roboticist who wants to bring AI into contact with the real world

https://www.newscientist.com/article/2406229-the-roboticist-who-wants-to-bring-ai-into-contact-with-the-real-world/

This cyborg cockroach could be the future of earthquake search and rescue

https://www.nature.com/articles/d41586-023-03801-0

Apple

Apple released an ML framework for Apple Silicon, finally. MLX is an efficient machine learning framework specifically designed for Apple silicon (i.e. your laptop!)  This may be Apple’s biggest move on open-source AI so far: MLX, a PyTorch-style NN framework optimized for Apple Silicon, e.g. laptops with M-series chips.

Science/Education

‘Google for wildlife sounds’: Australian conservation research gets an AI boost

Researchers can upload their recordings, and match them to bird calls from around the country.

https://www.smh.com.au/technology/google-for-wildlife-sounds-huge-boost-for-conservation-research-20231127-p5en31.html

SchoolXpress turns your books, handwritten notes, classwork and documents into interactive learning content that simplifies concepts and delivers on your learning objectives. https://www.schoolxpress.ai 

Sperm whales have equivalents to human vowels.

We uncovered spectral properties in whales’ clicks that are recurrent across whales, independent of traditional types, and compositional.

We got clues to look into spectral properties from our AI interpretability technique CDEV.

Absci Announces Collaboration with AstraZeneca to Advance AI-Driven Oncology Candidate

https://finance.yahoo.com/news/absci-announces-collaboration-astrazeneca-advance-123000837.html

OpenAI

New report illuminates why OpenAI board said Altman “was not consistently candid”

https://arstechnica.com/ai/2023/12/openai-board-reportedly-felt-manipulated-by-ceo-altman

The Inside Story of Microsoft’s Partnership with OpenAI

https://www.newyorker.com/magazine/2023/12/11/the-inside-story-of-microsofts-partnership-with-openai

Elon Musk told OpenAI to move faster right before he left the company in 2018: NYT

https://www.businessinsider.com/elon-musk-told-openai-to-move-faster-before-he-left-2023-12

The OpenAI Board Member Who Clashed With Sam Altman Shares Her Side

https://www.wsj.com/tech/ai/helen-toner-openai-board-2e4031ef

OpenAI’s GPT store delayed to next year

https://www.theverge.com/2023/12/1/23984497/openai-gpt-store-delayed-ai-gpt

OpenAI Agreed to Buy $51 Million of AI Chips From a Startup Backed by CEO Sam Altman

https://www.wired.com/story/openai-buy-ai-chips-startup-sam-altman/

OpenAI employees really, really did not want to go work for Microsoft

https://www.businessinsider.com/openai-employees-did-not-want-to-work-for-microsoft-2023-12

Amazon

Amazon’s Q generative AI chatbot allegedly leaks location of AWS data centers – report

https://www.datacenterdynamics.com/en/news/amazons-q-generative-ai-chatbot-leaks-location-of-aws-data-centers

Business/Enterprise

McKinsey Sees AI Adding Up to $340 Billion to Wall Street Profit

https://www.bloomberg.com/news/articles/2023-12-05/ai-could-add-340-billion-to-wall-street-profits-mckinsey-says

McDonald’s taps Google for ‘Ask Pickles’ AI chatbot to help fix ice cream machines

Called “Ask Pickles,” the bot will be trained on everything from manuals to data generated by equipment at restaurants. It will give workers guidance on the spot, potentially boosting productivity in an industry where every second counts.

https://finance.yahoo.com/news/mcdonalds-taps-google-ask-pickles-183450432.html

https://www.bloomberg.com/news/articles/2023-12-06/mcdonald-s-mcd-getting-ai-chatbot-from-google-googl-for-restaurant-crew

https://www.theverge.com/2023/12/6/23990900/mcdonalds-google-ai-cloud-generative

Solve Intelligence helps attorneys draft patents for IP analysis and generation

https://techcrunch.com/2023/11/28/solve-intelligences-ai-solution-helps-attorneys-draft-patents-for-ip-analysis-and-generation/

https://www.solveintelligence.com/

A Bearish POV Emerges

The step from GPT4 to GPT5 might be so difficult that there is an AI bubble.

“Calling it now:  The $86B OpenAI tender will someday be seen as the WeWork moment of AI. 

GPT-5 will either be significantly delayed or not meet expectations.

Companies will struggle to put GPT-4 and 5 into production (see below)

Competition will increase, margins will be thin

The profits won’t justify the valuation, esp after MSFT’s hefty cut is taken out.”

Google Gemini seems to have by many measures matched (or slightly exceeded) GPT-4, but not to have blown it away.

From a commercial standpoint GPT-4 is no longer unique. That’s a huge problem for OpenAI, especially post drama, when many customers are now seeking a backup plan.

From a technical standpoint, the key question is: are LLMs close to a plateau? 

Note that Gates and Altman have both been dropping hints, and GPT-5 isn’t here after a year despite immense commercial desire. The fact that Google, with all its resources, did NOT blow away GPT-4 could be telling.

Twitter/X/Grok

Grok officially launches to all (thread)

Elon Musk’s AI firm xAI files to raise up to $1 billion in equity offering

https://finance.yahoo.com/news/elon-musks-xai-files-raise-193558630.html

Legal/Ethics

A game to try to get an LLM to give away sensitive information

https://gandalf.lakera.ai/

Evaluating and Mitigating Discrimination in Language Model Decisions

https://www.anthropic.com/index/evaluating-and-mitigating-discrimination-in-language-model-decisions

IBM and Meta Launch the AI Alliance in collaboration with over 50 Founding Members and Collaborators globally including AMD, Anyscale, CERN, Cerebras, Cleveland Clinic, Cornell University, Dartmouth, Dell Technologies, EPFL, ETH, Hugging Face, Imperial College London, Intel, INSAIT, Linux Foundation, MLCommons, MOC Alliance operated by Boston University and Harvard University, NASA, NSF, Oracle, Partnership on AI, Red Hat, Roadzen, ServiceNow, Sony Group, Stability AI, University of California Berkeley, University of Illinois, University of Notre Dame, The University of Tokyo, Yale University and others

https://ai.meta.com/blog/ai-alliance/

Uber Eats is using AI for pictures of food. It doesn’t know that “pie” means pizza, and it invented a brand of ranch dressing called “Lelnach”

Announcing Purple Llama: Towards open trust and safety in the new world of generative AI

https://ai.meta.com/blog/purple-llama-open-trust-safety-generative-ai/

G42’s Ties To China Run Deep

More details on the Emirati technology firm at the nexus of China and the U.S.’s technological and geopolitical rivalry.

https://www.thewirechina.com/2023/12/03/g42s-ties-to-china-run-deep-g42-peng-xiao/

Microsoft

Microsoft’s Copilot is getting OpenAI’s latest models and a new code interpreter

https://www.theverge.com/2023/12/5/23989052/microsoft-copilot-gpt-4-turbo-openai-models-code-interpreter-feature

Audio

Google’s new AI experiment composes abstract musical clips inspired by instruments

https://artsandculture.google.com/experiment/8QFo2oQr2uT3pg

https://www.engadget.com/googles-new-ai-experiment-composes-abstract-musical-clips-inspired-by-instruments-203732054.html

SeamlessExpressive, a new AI translation model by research teams at Meta, enables high-quality speech translation that maintains the speaker’s vocal style, tone and unique expressions in translated outputs.

Try the demo with your own voice 

https://seamless.metademolab.com/expressive

Technical/Dev/IT

Liquid AI, a new MIT spinoff, wants to build an entirely new type of AI

“Liquid neural networks consist of “neurons” governed by equations that predict each individual neuron’s behavior over time, like most other modern model architectures. The “liquid” bit in the term “liquid neural networks” refers to the architecture’s flexibility; inspired by the “brains” of roundworms, not only are liquid neural networks much smaller than traditional AI models, but they require far less compute power to run”

“GPT-3 contains about 175 billion parameters and ~50,000 neurons — “parameters” being the parts of the model learned from training data that essentially define the skill of the model on a problem (in GPT-3’s case generating text). By contrast, a liquid neural network trained for a task like navigating a drone through an outdoor environment can contain as few as 20,000 parameters and fewer than 20 neurons.”

https://techcrunch.com/2023/12/06/liquid-ai-a-new-mit-spinoff-wants-to-build-an-entirely-new-type-of-ai/

An Opinionated Guide to Which AI to Use: ChatGPT Anniversary Edition

https://www.oneusefulthing.org/p/an-opinionated-guide-to-which-ai

Quantifying ChatGPT’s gender bias

https://www.aisnakeoil.com/p/quantifying-chatgpts-gender-bias

“Quadratic attention has been indispensable for information-dense modalities such as language… until now.  Announcing Mamba: a new SSM arch. that has linear-time scaling, ultra long context, and most importantly–outperforms Transformers everywhere we’ve tried.”

“VideoRF: Rendering Dynamic Radiance Fields as 2D Feature Video Streams”

A few months ago, Meta released Segment Anything Model (SAM). It’s already in photography apps, medicine, and video-generation. Now we’ve invented SAM’s little brother, EfficientSAM. It is small but mighty!  With 20x fewer parameters and 20x faster runtime, EfficientSAM is within 2 points (44.4 AP vs 46.5 AP) of the original SAM model. 

How’d we do it? I have 2 words for you: Masked Autoencoders. Check out the following for more details!

Excited to share ReconFusion! 3D reconstruction of real-world scenes from only a few photos, powered by diffusion priors

LLMs fine-tuned with RLHF are known to be poorly calibrated.  We found that they can actually be quite good at *verbalizing* their confidence.

HybridNeRF: Efficient Neural Rendering via Adaptive Volumetric Surfaces

LooseControl: Lifting ControlNet for Generalized Depth Conditioning 

Mistral AI just dropped a mixtral-8x7b-32kseqle model as a torrent link!  32K context, 8x mixture of experts.

How Aws Can Undercut Nvidia With Homegrown Ai Compute Engines

https://www.nextplatform.com/2023/12/04/how-aws-can-undercut-nvidia-with-homegrown-ai-compute-engines/

Ai Model Visualization Tool

https://bbycroft.net/llm

OneLLM: One Framework to Align All Modalities with Language

https://onellm.csuhan.com/

Long Context Prompting For Claude 2.1

Claude 2.1’s performance when retrieving an individual sentence across its full 200K token context window. This experiment uses a prompt technique to guide Claude in recalling the most relevant sentence.

https://www.anthropic.com/index/claude-2-1-prompting

Introducing Stable LM Zephyr 3B: A New Addition to Stable LM, Bringing Powerful LLM Assistants to Edge Devices

Stable LM Zephyr 3B is a 3 billion parameter Large Language Model (LLM), 60% smaller than 7B models, allowing accurate, and responsive output on a variety of devices without requiring high-end hardware. 

https://stability.ai/news/stablelm-zephyr-3b-stability-llm

Inside Apple’s chip lab, home to the most ‘profound change’ at the company in decades

https://www.cnbc.com/2023/12/01/how-apple-makes-its-own-chips-for-iphone-and-mac-edging-out-intel.html

OpenAI Rival Mistral Nears $2 Billion Valuation With Andreessen Horowitz Backing

https://archive.is/4F3dT#selection-4559.0-4566.0

GPT4  

Bing is a little confusing, it only uses GPT-4 in Creative (Purple) or Precise (Green) mode. Balanced (Blue) does not use GPT-4.  If you are doing web searches, Precise mode will give you less hallucination. If you are doing anything else, use Creative Mode.

Self-operating-computer + AI to AI Simulation Test Project feat Q*

“Today we check out the open source project the Self-operating-computer and a system where 2 AIs simulate a conversation discussing a uploaded image.”

5 LLM Security Threats- The Future of Hacking?

“Today we check what could be the future of hacking and LLM attacks with Jailbreaks and Prompt Injections on LLMs and Multimodal Models”

Sam Altman’s Brain Chips | Rain Neuromorphic Chips | UAE Funds and US National Security and Q*

Credits/Sources

Most of these links come from just a few incredible sources.  Please follow them:

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading