I decided to try a theme with this week’s cover imagery to see how creative MidJourney could be with simple prompts. Each category cover image is a name tag + art style. It was pretty neat to see the variances. The goal is not perfection. By posting the mistakes, we’ll get to see how imagery improves over time. Here is the prompt for the cover:

a psychedelic art name tag that reads “Googles” –ar 5:3 –style raw

Gemini

Gemini Flash – Google DeepMind 

“Today, we’re excited to introduce a new Gemini model: 1.5 Flash. ⚡ It’s a lighter weight model compared to 1.5 Pro and optimized for tasks where low latency and cost matter – like chat applications, extracting data from long documents and more. #GoogleIO (plus a huge context window) 

Flash has a one-million-token context window by default, which means you can process one hour of video, 11 hours of audio, codebases with more than 30,000 lines of code, or over 700,000 words.

Gemini 1.5 Pro updates, 1.5 Flash debut and 2 new Gemma models

“Introducing Gemini 1.5 Flash ⚡ It’s a lighter-weight model, optimized for tasks where low latency and cost matter most. Starting today, developers can use it with up to 1 million tokens in Google AI Studio and Vertex AI. #GoogleIO 

Google is building its Gemini Nano AI model into Chrome on the desktop | TechCrunch

At the Google I/O 2024 developer conference on Tuesday, Google announced that it is building Gemini Nano, the smallest of its AI models, directly into the Chrome desktop client, starting with Chrome 126.

“Today we have published our updated Gemini 1.5 Model Technical Report. As @JeffDean highlights, we have made significant progress in Gemini 1.5 Pro across all key benchmarks; TL;DR: 1.5 Pro > 1.0 Ultra, 1.5 Flash (our fastest model) ~= 1.0 Ultra. https://twitter.com/OriolVinyalsML/status/1791521517211107515

“One other thing in the updated Gemini 1.5 Pro report: we show how a research model that is a mathematics-specialized version of 1.5 Pro achieves a record score of 91.1% on the MATH benchmark (the SOTA just 3 years ago, in May, 2021 was 6.9%!).”

“Introducing LearnLM: our new family of models based on Gemini and fine-tuned for learning. LearnLM applies educational research to make our products — like Search, Gemini and YouTube — more personal, active and engaging for learners. #GoogleIO 

“Education is one of the areas in which LLMs can do the most immediate good, even with their limitations, so I was excited to see that Google is fine tuning a tutor LLM. Also, the comparison they used was Gemini 1.0 running a variation of our tutor prompt! The prompt alone did ok 

DeepMind

Google DeepMind CEO on Drug Discovery, Hype, Isomorphic – YouTube

Other Google News

Google I/O 2024: Here’s everything Google just announced | TechCrunch

Google I/O 2024: News and announcements

Google Gemini updates: Flash 1.5, Gemma 2 and Project Astra

Google I/O 2024: New generative AI experiences in Search

Google is overhauling its search results page with AI overviews and Gemini organization – The Verge

“Coming soon, we’ll bring new multi-step reasoning capabilities to Google Search. It breaks your bigger question down into parts and figures out which problems to solve and in what order, so research that might’ve taken you minutes or even hours can be done in seconds. #GoogleIO 

“We’re sharing Project Astra: our new project focused on building a future AI assistant that can be truly helpful in everyday life.  Watch it in action, with two parts – each was captured in a single take, in real time. ↓ #GoogleIO 

“Introducing Veo: our most capable generative video model. 🎥 It can create high-quality, 1080p clips that can go beyond 60 seconds. From photorealism to surrealism and animation, it can tackle a range of cinematic styles. 🧵 #GoogleIO 

Veo – Google DeepMind

“Google DeepMind just launched Veo, it’s Sora competitor. It generates high-quality videos in 1080P from text, image, and video prompts. 

“Veo could help make high-quality video production accessible to everyone. ✨ It can understand many kinds of effects and even captures the nuance and tone of a prompt – offering an unprecedented level of creative control. → 

“Last year we introduced SynthID, which adds imperceptible watermarks to AI-generated images and audio, making them easier to distinguish. Today we’re expanding SynthID to text and video outputs, including our new Veo model. #GoogleIO 

“We’re introducing Imagen 3: our highest quality text-to-image generation model yet. 🎨 It produces visuals with incredible detail, realistic lighting and fewer distracting artifacts. From quick sketches to very high-res imagery, here’s a look at what it can create. 👀 #GoogleIO 

“Today we’re introducing Imagen 3, @GoogleDeepMind’s most capable image generation model yet. It understands prompts the way people write, creates more photorealistic images and is our best model for rendering text. #GoogleIO 

“Google announced its new text-to-music tool. it’s really good, wow. You can even mix different tracks/prompts in real time. 

“Together with @YouTube, we’ve been building Music AI Sandbox, a suite of AI tools to transform how music can be created. 🎵 To help us design and test them, we’ve been working closely with musicians, songwriters and producers. ↓ #GoogleIO 

Music AI Demos | Experiments with Music AI Sandbox – YouTube – https://www.youtube.com/playlist?list=PLqYmG7hTraZA7o7KkLWoVscoELWRGu3Xg  

I/O 2024: New ways to experience Google AI on Android

“Whether you need a yoga bestie or calculus tutor, in the coming months you’ll be able to customize Gemini, saving time when you have specific ways you interact with Gemini again and again. We’re calling these Gems. #GoogleIO 

“Get a sneak peek of Gemma 2, our next generation of models that will include a 27B parameter instance launching in a few weeks. Built on new architecture, Gemma 27B outperforms models twice its size and can run on a single TPU host in Vertex AI. #GoogleIO 

Google is bringing Project Starline’s ‘magic window’ experience to real video calls – The Verge

“Trillium is our latest generation of TPUs and delivers a 4.7x improvement in compute performance per chip over the previous generation, TPU v5e. #GoogleIO”

Introducing Trillium, sixth-generation TPUs | Google Cloud Blog

“This summer, we’re expanding Gemini’s multimodal capabilities — including the ability to have an in-depth two-way conversation using your voice. This new experience is called Live. #GoogleIO 

“Sir Demis Hassabis just showed a super low latency demo of Google’s multimodal AI assistant on your phone AND augmented reality glasses. Clearly they’ve been cooking this for a while. The race is on! 

“Google is doing something very interesting by building specialized versions of its frontier models for math, healthcare, and education (so far). The benchmarks on all of these are pretty impressive, and it seems to be beyond what can be done with traditional fine tuning alone.”

“Does AlphaZero count as training on synthetic data? There’s no human grandmaster data at all. AlphaZero expands its strategies & wisdom indefinitely with self-driven exploration and compute. The input is just a simple Go/Chess simulator that implements the game rules.”

Heads up! You’ve scrolled to the end of this category. There may have been just one or two links (above), so go back up and double check to be sure you didn’t quickly scroll down past it.

Be Sure To Read This Week’s Main Post:

This week’s executive overview and top links are here:

AI News #33: Week Ending 05/17/2024 with Executive Summary and Top 58 Links

The post you just read is an deep dive extension of my weekly newsletter, This Week In AI, an executive summary of the top things to know in AI. Each week, I create an accessible overview for laypeople to feel confident they are conversant with the week’s AI developments. I include a curated list of must-click links of the week, to offer everyone a hands-on opportunity to explore the most intriguing updates in artificial intelligence across various categories, including robotics, imagery, video, AR/VR, science, ethics, and more. Beyond the overview, I post these topic-based deeper dives (below). If you haven’t read this week’s overview, I recommend starting there.

Credits/Sources

Most of these weekly links come from just a few prolific oversharing sources. Please follow them, as they work hard to find the news each week and they make it a lot easier for me to compile.

For previous issues, please visit the archives!

Thanks for reading!

2 responses to “Google AI News: Week Ending 05/17/2024”

  1. […] Google AI News of the Week: Individual company products will often be placed in the categories they match (image, audio, agents, robots, etc). Occasionally, I’ll dedicate space to a company’s news if it’s broadThis week’s latest Google AI news: https://ethanbholland.com/2024/05/17/google-ai-news-week-ending-05-17-2024/ […]

  2. […] a psychedelic art name tag that reads “Google” –ar 5:3 –style raw […]

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading