This week’s cover reflects the launch of Mistral 8x, an open source LLM that beats GPT 3.5 and has no guardrails. Made with Dalle-3, Magific.ai, Text Studio (https://www.textstudio.com/), and Photoshop.

Executive Summary

  • Open-Source AI: The top story is the release of Mistral, an AI model that is free and anyone can download or modify.  It’s better than GPT 3.5 and has no safety guardrails.  This means there will be 1000s of custom trained AI’s out in the wild… being used for every good and bad use case you can imagine.  No internet connection needed.
  • Small AI’s: Powerful AI’s can now fit on your phone and run without the internet.  This week alone covers eight small models.
  • AI Newscast Goes Viral: A full-length newscast with AI anchors and AI stories goes viral.  Personally, I see bite-sized dynamic news as stronger than one-size-fits-all linear shows, but the proof of concept grabbed attention and is important to note.
  • What’s Up At OpenAI: When directly asked “What happened”, Sam Altman states that as OpenAI approaches superintelligence, “people get more intense.”   In Elon Musk’s infamous interview (when he cursed at advertisers), Elon also said he thinks OpenAI discovered something big and Ilya had a reason to fire Sam Altman. 
  • Google Solves Unsolved Math Problem: DeepMind used a large language model to discover solutions for the cap set problem 
  • AI Price Wars: Google is pricing out OpenAI by giving 60 API queries on Gemini Pro per minute for free.  Behind the scenes, there has generally been an approximate hundredfold cost reduction in one year for AI companies running models.
  • Zoom AI Companion: Zoom introduces meeting summaries and other productivity tools within conference calls.
  • Midjourney Alpha: The most photorealistic image generator launches its standalone product, independent of the Discord interface
  • Anthropic Google Integration: Anthropic released a prompt engineering API tool that makes Claude accessible via spreadsheets.  Users can test and refine prompts within their day-to-day workflows and collaborate with teammates. 
  • Audio Restoration: Resemble Enhance transforms old noisy audio into crystal clear speech

Top Stories

These are the must-click links if you only have time for a few.  Even if they look boring, click them!  I did the work, so you don’t have to worry.  All are 10/10 would recommend.

The Rest: AI News of The Week

Don’t let the volume overwhelm you.  Have fun and skim it. The links are organized by topic, sorted from coolest to least cool, and each topic is clearly defined with a headline.  The links descriptions are often pulled directly from tweets or articles, so it’s not always my voice.  Pause when you see something that interests you.  Reach out to me any time.  I enjoy sharing and discussing these items.

Mistral – Open Source

“For those who don’t follow AI closely:

1) An open source model (free, anyone can download or modify) beats GPT-3.5

2) It has no safety guardrails

There are good things about this release, but also regulators, IT security experts, etc. should note the genie is out of the bottle.”

“Today, the team is proud to release Mixtral 8x7B, a high-quality sparse mixture of experts model (SMoE) with open weights. Licensed under Apache 2.0. Mixtral outperforms Llama 2 70B on most benchmarks with 6x faster inference. It is the strongest open-weight model with a permissive license and the best model overall regarding cost/performance trade-offs. In particular, it matches or outperforms GPT3.5 on most standard benchmarks.”

https://mistral.ai/news/mixtral-of-experts/

“Very excited to release our second model, Mixtral 8x7B, an open weight mixture of experts model.  Mixtral matches or outperforms Llama 2 70B and GPT3.5 on most benchmarks, and has the inference speed of a 12B dense model. It supports a context length of 32k tokens. (1/n)”

“Mistral AI, a Paris-based OpenAI rival, closed its $415 million funding round”

https://techcrunch.com/2023/12/11/mistral-ai-a-paris-based-openai-rival-closed-its-415-million-funding-round/

“Mistral AI brings the strongest open generative models to the developers, along with efficient ways to deploy and customize them for production.  We’re opening a beta access to our first platform services today. We start simple: la plateforme serves three chat endpoints for generating text following textual instructions and an embedding endpoint. Each endpoint has a different performance/price tradeoff.”

https://mistral.ai/news/la-plateforme/

“Mixtral 8x7B is insanely good — here’s my walkthrough on deploying it and using it as an open AI agent”

“How Mixstral Works”

Mistral AI

“French startup Mistral AI just released Mixtral, an open-source 45B parameter AI model.  Mixtral matches or outperforms LLaMA 2 and GPT-3.5 on most benchmarks while running 6x faster.”

Mixtral of experts

A high quality Sparse Mixture-of-Experts.

https://mistral.ai/news/mixtral-of-experts/

“Only about a year after the launch of ChatGPT-3.5, I now have a GPT-3.5 class AI running on my home computer that is open source, free, reasonably fast, & doesn’t require an Internet connection (Mixtral 8x)”

Small/Locally-Run AI

Samsung

Samsung unveils its generative AI model Samsung Gauss

Samsung Gauss consists of language, code, and image models and will be applied to the company’s various products in the future.

https://www.zdnet.com/article/samsung-unveils-its-generative-ai-model-samsung-gauss/

Microsoft

Phi-2: The surprising power of small language models

Outperforms the 7B parameter Mistral and Llama 2 models, and even surpasses Google’s 3B parameter Gemini Nano 2 on some tests.

Microsoft Phi-2 is a Transformer with 2.7 billion parameters.

https://huggingface.co/microsoft/phi-2

Phi-2: The surprising power of small language models

You can run `phi-2` on your own device (see my code example for running accelerated on Metal backend). It’s available under Microsoft’s Research License.

Mistral 7B

After playing with this new mistral model for the last 24 hours, pretty sure if you go through fine tuning and rlhf you would get a > gpt-3.5 local model

I have now run one of the more powerful, open source LLMs (Mistral 7B) directly on my iPhone. No internet needed.

Obsidian

Worlds smallest multi-modal LLM. First multi-modal model in size 3B

https://huggingface.co/NousResearch/Obsidian-3B-V0.5

EdgeSAM

Introducing EdgeSAM, the first SAM variant that can run at over 30 FPS on an iPhone 14 with minimal compromise in performance.

https://mmlab-ntu.github.io/project/edgesam/

Segmind Vega

Introducing Segmind Vega: Compact model for Real-time Text-To-Image

Segmind has introduced two new open-source text-to-image models, Segmind-VegaRT (Real Time) and Segmind-Vega, the fastest and smallest, open source models for image generation at the highest resolution.

https://blog.segmind.com/segmind-vega/

Upstage/SOLAR-10.7B

Introducing Upstage/SOLAR-10.7B.   It’s compact, yet powerful, delivering the best performance among models under 30B. The secret? We use vertical depth upscaling instead of horizontal upscaling methods like MOE.

gigaGPT

Introducing gigaGPT: GPT-3 sized models in 565 lines of code

GigaGPT is Cerebras’ implementation of Andrei Karpathy’s nanoGPT – the simplest and most compact code base to train and fine-tune GPT models.https://www.cerebras.net/blog/introducing-gigagpt-gpt-3-sized-models-in-565-lines-of-code

Science/Medicine

Google Solves Unsolved Math Problem

“A highly important open question about AIs has been whether they can actually discover new things.  Today, a new Nature paper has shown that LLMs can.  A model ‘discovered new solutions for the cap set problem, a longstanding open problem in mathematics’”

Google DeepMind used a large language model to solve an unsolvable math problem

DeepMind AI outdoes human mathematicians on unsolved problem

Large language model improves on efforts to solve combinatorics problems inspired by the card game Set.

https://www.nature.com/articles/d41586-023-04043-w

Google DeepMind used a large language model to solve an unsolved math problem

They had to throw away most of what it produced but there was gold among the garbage.

https://www.technologyreview.com/2023/12/14/1085318/google-deepmind-large-language-model-solve-unsolvable-math-problem-cap-set/

FunSearch: Making new discoveries in mathematical sciences using Large Language Models

https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models/

Google Medical Models

Today Google announced MedLM, a suite of new health-care-specific AI models designed to help clinicians and researchers carry out complex studies, summarize doctor-patient interactions and more.

Google is rolling out new AI models for health care. Here’s how doctors are using them

https://www.cnbc.com/2023/12/13/how-doctors-are-using-googles-new-ai-models-for-health-care.html

General

This is a collection of awesome articles about CLIP in medical imaging 

https://github.com/zhaozh10/awesome-clip-in-medical-imaging

Ultra-thin flat cameras are possible with nanophotonic optics! Excited to share recent work that shrinks the entire optical stack down to a 700-nanometer thick layer of optics on the sensor cover glass.

Neural Stress Field is a reduced-order modeling framework for solid, fluid, and fracture mechanics. After training, it enables 10X speedup for simulations of tearing and other damage phenomena to everyday materials such as bread!
https://twitter.com/peterchencyc/status/1734735049822806302

Ethics/Legal

AI Romance Bot

Excited to announce v(1.0) of Digi, the future of AI Romantic Companionship, for IOS and Android.  A quick thread on features, and where we go from here.  (Designed by lead animator at Pixar with high-speed audio developed in-house.)

https://digi.ai/

Sports Illustrated’s Fake AI Author Bios and Articles

Sports Illustrated publisher fires CEO in latest round of exec terminations after AI scandal

https://www.nbcnews.com/news/us-news/sports-illustrated-publisher-fires-ceo-latest-exec-terminations-ai-sca-rcna129236

Sports Illustrated Published Articles by Fake, AI-Generated Writers

We asked them about it — and they deleted everything.

https://futurism.com/sports-illustrated-ai-generated-writers

AI Drew (fake SI author)

“Drew likes to say that he grew up in the wild, which is partially true. He grew up in a farmhouse, surrounded by woods, fields, and a creek. Drew has spent much of his life outdoors, and is excited to guide you through his never-ending list of the best products to keep you from falling to the perils of nature. Nowadays, there is rarely a weekend goes by where Drew isn’t out camping, hiking, or just back on his parents’ farm.”

https://web.archive.org/web/20221205082417/https://www.si.com/review/author/drewortiz/ 

Supervising Superintelligence

A core challenge for aligning future superhuman AI systems (superalignment) is that humans will need to supervise AI systems much smarter than them. We study a simple analogy: can small models supervise large models?

https://openai.com/research/weak-to-strong-generalization

In the future, humans will need to supervise AI systems much smarter than them.

We study an analogy: small models supervising large models.

Our new research white paper identifies seven practices for keeping increasingly agentic AI systems safe and accountable as they become more common and more capable.

OpenAI Practices for Governing Agentic AI Systems

Our new research white paper identifies seven practices for keeping increasingly agentic AI systems safe and accountable as they become more common and more capable.

We are providing research grants for work on a range of open questions.

https://openai.com/research/practices-for-governing-agentic-ai-systems

MIT group releases white papers on governance of AI

The series aims to help policymakers create better oversight of AI in society.

https://news.mit.edu/2023/mit-group-releases-white-papers-governance-ai-1211

Laws and Policies

Europe reaches a deal on the world’s first comprehensive AI rules

https://apnews.com/article/ai-act-europe-regulation-59466a4d8fd3597b04542ef25831322c

Microsoft and Labor Unions Form ‘Historic’ Alliance on AI

https://finance.yahoo.com/news/microsoft-labor-unions-form-historic-142333100.html

Poisoning AI Model Training Data

The adversarial data arms race for large AI models has officially begun. I knew poisoning LLM’s training data was possible but apparently it’s live in the wild.

Introducing, VonGoom: A  method for data poisoning large language models to introduce bias, requiring as few as 100 poisoned examples within training data.

Even a single article has the power to bias a model to hold any opinion an attacker wants…

https://delcomplex.com/vonGoom

Prompt Attacks

Weak cybersec + ease of adversarial attacks remain important reasons for the somewhat slow-ish corp. adoption of LMs at scale: “Through prompt injection, an adversary can not only extract the customized system prompts but also access the uploaded files.”

https://ar5iv.labs.arxiv.org/html/2311.11538

Google’s Self-Improving Agent

Google Develops Self-Improving LLM Agent

ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

https://huggingface.co/papers/2312.10003

Turing Test Drama

How 3 Turing Awardees Republished Key Methods and Ideas Whose Creators They Failed to Credit

https://people.idsia.ch/~juergen/ai-priority-disputes.html

Robo Calls 

Meet Ashley, the world’s first AI-powered political campaign caller

“This is going to scale fast,” said 30-year-old Ilya Mouzykantskii, the London-based CEO of Civox, the company behind Ashley. “We intend to be making tens of thousands of calls a day by the end of the year and into the six digits pretty soon. This is coming for the 2024 election and it’s coming in a very big way. … The future is now.”

https://www.reuters.com/technology/meet-ashley-worlds-first-ai-powered-political-campaign-caller-2023-12-12/

Phishing Attacks

Large Language Models Can Be Used To Effectively Scale Spear Phishing Campaignshttps://arxiv.org/ftp/arxiv/papers/2305/2305.06972.pdf

OpenAI Drama: Continued

Sam Altman on OpenAI, Future Risks and Rewards, and Artificial General Intelligence

https://time.com/6344160/a-year-in-time-ceo-interview-sam-altman/

OpenAI leaders warned of abusive behavior before Sam Altman’s ouster

https://www.msn.com/en-us/money/companies/warning-from-openai-leaders-helped-trigger-sam-altman-s-ouster/ar-AA1ldAfV

The real research behind the wild rumors about OpenAI’s Q* project

https://www.understandingai.org/p/how-to-think-about-the-openai-q-rumors

After moving to oust Sam Altman, OpenAI cofounder and chief scientist Ilya Sutskever is in a sort of limbo, and nobody seems to know what will happen next.

https://futurism.com/the-byte/ilya-sutskever-openai-limbo?utm_source=tldrnewsletter

In a now infamous interview for his remarks to advertisers, Elon also said he thinks OpenAI discovered something big and Ilya had a reason to fire Sam Altman.  Skip to the 48 minute mark.

https://www.rev.com/blog/transcripts/dealbook-summit-2023-elon-musk-interview-transcript

Sam Altman directly says that as OpenAI approaches superintelligence, people get more intense… in response to “What happened”?  Surprised me. https://www.youtube.com/watch?v=e1cf58VWzt8

Embodiment/Robotics

Autonomous drone flies through forest canopy to avoid collisions and maps dense forest

Mountain View company touts aviation history with flight of crewless cargo plane

Uncrewed flight creates headway at South Bay airport

https://www.paloaltoonline.com/news/2023/12/12/mountain-view-company-touts-aviation-history-with-flight-of-crewless-cargo-plane

Tesla’s Humanoid Robot Is Now 30% Faster, 22lbs Lighter. 

https://www.zerohedge.com/technology/teslas-humanoid-robot-now-30-faster-22lbs-lighter

Here are 17 real demos that look straight out of a sci-fi movie:

Tesla Optimus Gen 2 Demo

“Remarkably this approach enables Alter3 to adopt various poses, such as a ‘selfie’ stance or ‘pretending to be a ghost,’ and generate sequences of actions over time without explicit programming for each body part.”

Thrilled to announce Octo, an open-source robot foundation model! Octo is a sota generalist robot policy based on transformer+diffusion. Most importantly, you can finetune Octo *today* with flexible observation and action spaces on your robot setup!

Remotely Operated Trucks

https://einride.tech/autonomous/vehicles

Enterprise/Business

Zoom AI Companion

Zoom AI Companion empowers you to increase productivity, improve team effectiveness and enhance your skills. Using Zoom’s unique federated approach to AI, you can expect high-quality results when drafting emails and chat messages, summarizing meetings and chat threads, brainstorming creatively and much more – all in the simple, easy-to-use Zoom experience you know and love.

https://www.zoom.com/en/ai-assistant/

Podcast Tool

Toasty Transcribes Podcasts and Learns from Them

Human Resources

Review job applications 10x faster

Dover’s AI sorting and chat-powered assistant helps you focus on the best resumes, and quickly get back to everyone else.

https://www.dover.com/ai-sorting

Start-ups

Why a Vertical Approach is Key to Building Enduring AI Applications

At Greylock, we believe this convergence of large language models and the embrace of differentiated models creates ideal conditions for founders who want to solve long-standing problems.

https://greylock.com/greymatter/vertical-ai/

Journalism/Reporting

Virtual News 

Channel1 AI News Demo of Present Capabilities

“Our generated anchors deliver stories that are informative, heartfelt and entertaining.”

I don’t understand why they made a full length show, when bite-sized personalized clips would be what I’d expect/want…. But it went viral.  The future will be agent driven on-demand.  One-sized fits all is going away (everything will be “for you page” style or on request, IMO)

https://www.channel1.ai/

Deals

OpenAI inks deal with Axel Springer on licensing news for model training

https://techcrunch.com/2023/12/13/openai-inks-deal-with-axel-springer-on-licensing-news-for-model-training/

Partnership with Axel Springer to deepen beneficial use of AI in journalism

https://openai.com/blog/axel-springer-partnership

Axel Springer and OpenAI partner to deepen beneficial use of AI in journalism

https://www.axelspringer.com/en/ax-press-release/axel-springer-and-openai-partner-to-deepen-beneficial-use-of-ai-in-journalism

ChatGPT to summarize Politico and Business Insider articles in ‘first of its kind’ deal

https://www.theguardian.com/technology/2023/dec/13/openai-axel-springer-chatgpt-story-writing-business-insider-politico

Video

Pika Keeps Getting Better

AI can now Outfit AND Animate Anyone from a single  image 

This Outfit Anyone + Animate Anyone is integrated together to take an image of a model, dress them with any style, then animate them

Demo Attack on Titan Style Video with Workflow Explanation

https://twitter.com/DeepMotionInc/status/1732895915072266407  

Stanford Releases WALT

We introduce W.A.L.T, a diffusion model for photorealistic video generation. Our model is a transformer trained on image and video generation in a shared latent space.

https://walt-video-diffusion.github.io/

Tik Tok/Alibaba

Chinese Tech Giant Alibaba Unveils New AI Video Tool

Alibaba says its I2VGen-XL model can handle “visualization, sampling, training, inference, join training using images and videos, acceleration, and more.

https://decrypt.co/210018/alibaba-ai-text-to-video-generative-cloud

Chat With YouTube

https://chatwithyoutube.pro/chat/

Imagery

Midjourney Alpha is officially here!!

If you’ve generated 10k or more images you should have access.

Adobe Generative Fill Demo (adding hair)

Krea AI – announcing Overlay View

Instagram introduces GenAI powered background editing tool

https://techcrunch.com/2023/12/14/instagram-introduces-gen-ai-powered-background-editing-tool/

“Here is my now traditional ‘Photo of otter on a plane using wifi’ test (I take one of the first four images) in Meta’s Imagine (Free), Bing’s DALL-E3 (Free), ChatGPT ($) and Midjourney ($)”

Google debuts Imagen 2 with text and logo generation

https://techcrunch.com/2023/12/13/google-debuts-imagen-2-with-text-and-logo-generation/

https://imagen.research.google/

https://deepmind.google/technologies/imagen-2/

Fooocus is a rethinking of Stable Diffusion and Midjourney’s designs:

Learned from Stable Diffusion, the software is offline, open source, and free.

Learned from Midjourney, the manual tweaking is not needed, and users only need to focus on the prompts and images.

https://github.com/lllyasviel/Fooocus

“Now add a walrus: Prompt engineering in DALL-E 3” (super fun read for power users)

https://simonwillison.net/2023/Oct/26/add-a-walrus/

Vision

ChatGPT Vision for digitizing journal entries

Is GPT-Vision still the best?

Price Wars

Google is pricing out OpenAI.  60 API queries on Gemini Pro per minute for free

When ChatGPT launched, Sam Altman estimated that it cost “several cents” per chat to run. Assuming roughly 1K tokens per chat, there’s been approximately a hundredfold cost reduction in one year, for roughly the same capability level.

Why Stability AI is launching a subscription fee

Stability AI will charge commercial customers for the use of its most advanced models, pivoting away from being fully open source

https://sifted.eu/articles/stability-business-model

Anthropic

We’ve released a new prompt engineering tool that makes Claude accessible via spreadsheets.

With Claude for Google Sheets, API users can test and refine prompts within their day-to-day workflows and easily collaborate with teammates.

https://docs.anthropic.com/claude/docs/using-claude-for-sheets

Audio

Remarkable restoration and clarity

Today, we introduce Resemble Enhance — our latest AI-powered model! Enhance is an open-source speech enhancement model that transforms noisy audio into noteworthy speech!

https://huggingface.co/spaces/ResembleAI/resemble-enhance

Meta Audioboc

Starting today you can try our new foundation research model for audio generation. The demo includes Zero shot TTS, Text to sound effects, Infilling and more!

https://audiobox.metademolab.com/

I just tried Google’s new AI music generator MusicFX — and it’s the best one yet

https://www.tomsguide.com/features/i-just-tried-googles-new-ai-music-generator-musicfx-and-its-the-best-one-yet

Google

Gemini Demo Backlash

A Remake of the Google Gemini Fake Demo, Except Using GPT-4 and It’s Real

Gemini Reviews

Interesting that Gemini AI Studio has this UI for adjusting safety settings of content; I don’t think I’ve ever seen a product take this direction before

First Impressions with Google’s Gemini

The Roboflow team has analyzed Gemini across a range of standard prompts that we have used to evaluate other LMMs, including GPT-4 with Vision, LLaVA, and CogVLM. Our goal is to  better understand what Gemini can and cannot do well at the time of writing this piece.

https://blog.roboflow.com/first-impressions-with-google-gemini

“It’s time for developers and enterprises to build with Gemini Pro

Learn more about how to integrate Gemini Pro into your app or business at ai.google.dev.”

https://blog.google/technology/ai/gemini-api-developers-cloud/

Google Notebook

Notebook released to public

https://notebooklm.google.com/

https://blog.google/technology/ai/notebooklm-google-ai/

Life Tracking

Project Ellmann, named after biographer and literary critic Richard David Ellmann, will take users’ search results and photos to make a chatbot able to answer “previously impossible questions.”

https://www.theverge.com/2023/12/8/23993641/google-is-reportedly-working-on-a-project-that-lets-ai-models-tell-someones-life-story

Google’s Project Ellmann Shows Your Life in Photos

https://gizmodo.com/google-gemini-ellmann-photos-privacy-1851089179

Google weighs Gemini AI project to tell people their life story using phone data, photos

https://www.cnbc.com/2023/12/08/google-weighing-project-ellmann-uses-gemini-ai-to-tell-life-stories.html

Report: Google working on ‘Pixie’ AI assistant for Pixel 9, discussed glasses with object recognition

https://9to5google.com/2023/12/14/pixel-9-pixie-ai-assistant/

Humane Pin PR Fails

Sadly, Humane is an example of how “NOT” to do marketing 

A horrible product marketing appearance

Augmented Reality/VirtualReality

Text-to-AR – Get your ideas into functioning AR filters in real-time. 

Real Time at 30 FPS – FaceSwapping Jamie Fox

Try the demo yourself here: https://fal.ai/camera  

It’s hard to convey how playful using Hand Tracking to animate characters in Mixed Reality is … but reaching that point where the technology disappears in an experience always feels magical

What if you could have your own AI 3D AR character take you to posters of interest to you and explain it, translated “in language of your own latent space”, even when the topic is left field and the presenter isn’t there? 

Meta Teases Render Of Advanced ‘Mirror Lake’ Headset It Says Is “Practical To Build Now”

https://www.uploadvr.com/meta-mirror-lake-advanced-prototype-render/

SMERF: Streamable Memory Efficient Radiance Fields for Real-Time Large-Scene Exploration

Introducing SMERF: a streamable, memory-efficient method for real-time exploration of large, multi-room scenes on everyday devices. Our method brings the realism of Zip-NeRF to your phone or laptop!

https://smerf-3d.github.io/

The First Ever TikTok Effect House Event

https://effecthouse.tiktok.com/open-house/

https://campaignme.com/tiktok-launches-virtual-ar-showcase-event/

Nuvo: Neural UV Mapping for Unruly 3D Representations

iOS 17.2 arrives with new Journal app and spatial video capture support

https://www.theverge.com/2023/12/11/23990347/ios-17-2-released-iphone-update-journal-app-spatial-video

Introducing Stable Zero123: Quality 3D Object Generation from Single Images

https://stability.ai/news/stable-zero123-3d-generation

Introducing SMERF: a streamable, memory-efficient method for real-time exploration of large, multi-room scenes on everyday devices. Our method brings the realism of Zip-NeRF to your phone or laptop!

AI-Powered Microdisplay Adapts to Users’ Eyesight “NeuralDisplay” could make AR less squinty, blurry, and nausea-inspiring

https://spectrum.ieee.org/augmented-reality-display-adapts

Chat-3D v2: Bridging 3D Scene and Large Language Models with Object Identifiers

https://arxiv.org/abs/2312.08168v1

OpenAI: Other News

The OpenAI Startup Fund was founded on two core beliefs. First, new and powerful AI systems will give rise to a new wave of transformative startups.  Today, we’re opening applications for Converge 2: the second cohort of our six-week program for exceptional engineers, designers, researchers, and product builders using AI to reimagine the world.

https://openai.fund/news/converge-2

“we have re-enabled chatgpt plus subscriptions!”

“GPT 4.5 was trending all day today on X, with a reported leak from Reddit showing new pricing with enhanced multimodal capabilities like video and 3D.  Sam Altman later denied the rumors, saying the GPT 4.5 leak was not legit.

OpenAI Launches Dev Account on X

India’s Aligned AI

We are extremely happy to launch India’s first AI computing stack that is unique to our context, connecting our future to our roots. AI will certainly transform everything, making India the most productive, efficient and empowered economy in the world.

Technical/Deep Dives

Steering at the Frontier: Extending the Power of Prompting

Today, we’re sharing information on Medprompt and other approaches to steering frontier models in promptbase(opens in new tab), a collection of resources on GitHub. Our goal is to provide information and tools to engineers and customers to evoke the best performance from foundation models. We’ll start by including scripts that enable replication of our results using the prompting strategies that we present here. We’ll be adding more sophisticated general-purpose tools and information over the coming weeks. 

Rethinking Compression: Reduced Order Modelling of Latent Features in Large Language Models

https://arxiv.org/abs/2312.07046v1

To reinvent how enterprises work, we are developing full-stack AI products that quickly learn to increase productivity by automating time-consuming and monotonous workflows. For instance, our technology will make data analysts 10X faster and give business users the tools to become independent data driven decision makers themselves. It will also identify the biggest risks and suggest improvements in organizations’ supply chains.

https://essential.ai/

As ChatGPT gets “lazy,” people test “winter break hypothesis” as the cause.  (Note – since then it appears to be a combination of a bit of winter break but mostly human guidance from OpenAI to take the foot off the gas to save compute)

https://arstechnica.com/information-technology/2023/12/is-chatgpt-becoming-lazier-because-its-december-people-run-tests-to-find-out/

Answer.AI is a new kind of AI R&D lab which creates practical end-user products based on foundational research breakthroughs.

https://www.answer.ai/posts/2023-12-12-launch.html

A lot of non-technical people don’t realize how amazing huggingface  is for anyone even mildly curious about AI.  It has over 430K AI models hosted on it, including model demos hosted by OpenAI, Google etc.  Quick intro guide for non-technical folks on how to get started:

We introduce X-InstructBLIP, a simple and effective scalable cross-modal framework to empower LLMs to handle tasks across modalities such as text, image, video, sound, and 3D.

For people who are not in the weeds of ML this might be entertaining. The best paper award for this years NeurIPS went to this paper (which I love) that showed claims of emergence in LLMs are nonsensical.

A thought provoking slide deck by Benedict Evans (self-walk through 87 slides)

https://pitch.com/v/2jmm9p/538fe91f-6878-4203-8b49-c66aae06e9a1

Quantization is a technique used to reduce the size and increase the efficiency of deep learning models.  Here is a list of 23 LLM quantization techniques you need to know about:

Microsoft “Guidance” is a tool that allows for precise templating of LLM outputs for better integrations as part of a more complex whole by manipulation of token probabilities.

https://github.com/guidance-ai/guidance

By popular demand, log probabilities are now in the Chat Completions API!

With logprobs for top output tokens, you can build features such as better autocomplete, classification, keyword suggestion, or assess the model’s confidence in its predictions.

https://platform.openai.com/docs/api-reference/chat/create#chat-create-logprobs

Towards a Generalized Multimodal Foundation Model

Along the track of X-Decoder, SEEM, and FIND, our goal is to build a foundation model that is extendable to a wide range of vision and vision-language tasks without any overhead!

https://x-decoder-vl.github.io/

Anyscale Endpoints: JSON Mode and Function calling Features

JSON mode ensures the outputs from our Large Language Models (LLMs) are not only valid JSON, but also tailored to your specific schema requirements.

https://www.anyscale.com/blog/anyscale-endpoints-json-mode-and-function-calling-features

The race between Intel, Samsung, and TSMC to ship the first 2 nm chip

https://arstechnica.com/gadgets/2023/12/the-race-between-intel-samsung-and-tsmc-to-ship-the-first-2nm-chip/

Intel unveils new AI chip to compete with Nvidia and AMD

https://www.cnbc.com/2023/12/14/intel-unveils-gaudi3-ai-chip-to-compete-with-nvidia-and-amd.html

explained: latent consistency models  (technical)

https://naklecha.notion.site/explained-latent-consistency-models-13a9290c0fd3427d8d1a1e0bed97bde2

Mamba-Chat is the first chat language model based on a state-space model architecture, not a transformer. The model is based on Albert Gu’s and Tri Dao’s work Mamba: Linear-Time Sequence Modeling with Selective State Spaces (paper) as well as their model implementation. This repository provides training / fine-tuning code for the model based on some modifications of the Huggingface Trainer class.

https://github.com/havenhq/mamba-chat

Chat With All LLMs At Once

Large Language Models (LLMs) based AI bots are amazing. However, their behavior can be random and different bots excel at different tasks. If you want the best experience, don’t try them one by one. ChatALL (Chinese name: 齐叨) can send prompt to several AI bots concurrently, help you to discover the best results.

https://github.com/sunner/ChatALL

Paving the way to efficient architectures: StripedHyena-7B, open source models offering a glimpse into a world beyond Transformers. One of the focus areas at Together Research is new architectures for long context, improved training, and inference performance over the Transformer architecture. Spinning out of a research program from our team and academic collaborators, with roots in signal processing-inspired sequence models, we are excited to introduce the StripedHyena models.

https://www.together.ai/blog/stripedhyena-7b

Better, Cheaper, Faster LLM Alignment with KTO

Today, we’re releasing a method called Kahneman-Tversky Optimization (KTO) that makes it easier and cheaper than ever before to align LLMs on your data without compromising performance.

PyTorch 2 Internals

Cool walk-through of latest innovations

https://www.slideshare.net/perone/pytorch-2-internals

OpenAssistants is a collection of open source libraries aimed at developing robust AI assistants rather than autonomous agents. By focusing on specific tasks and incorporating human oversight, OpenAssistants strives to minimize error rates typically found in agentic systems.

https://github.com/definitive-io/openassistants

Education: Roadmap To Learn Generative AI In 2024 

https://github.com/krishnaik06/Roadmap-To-Learn-Generative-AI-In-2024

Human brain-like supercomputer with 228 trillion links coming in 2024

https://interestingengineering.com/innovation/human-brain-supercomputer-coming-in-2024

Multimodal model support

Ollama now supports multimodal models that can describe what they see. To use a multimodal model with ollama run, include the full path of a png or jpeg image in the prompt:

https://github.com/jmorganca/ollama/releases/tag/v0.1.15

Introducing Unsloth: 30x faster LLM training

https://unsloth.ai/introducing

KwaiAgents is a series of Agent-related works open-sourced by the KwaiKEG from Kuaishou Technology. 

https://github.com/kwaikeg/kwaiagents

Podcast: Two Titans on the Future of AI (with Reid Hoffman & Vinod Khosla)

https://www.newcomer.co/p/two-titans-on-the-future-of-ai-with

Learning Naturally Aggregated Appearance for Efficient 3D Editing

Neural radiance fields, which represent a 3D scene as a color field and a density field, have demonstrated great progress in novel view synthesis yet are unfavorable for editing due to the implicitness. In view of such a deficiency, we propose to replace the color field with an explicit 2D appearance aggregation, also called canonical image, with which users can easily customize their 3D editing via 2D image processing.

https://felixcheng97.github.io/AGAP/

CorresNeRF: Image Correspondence Priors for Neural Radiance Fields

Neural Radiance Fields (NeRFs) have achieved impressive results in novel view synthesis and surface reconstruction tasks. However, their performance suffers under challenging scenarios with sparse input views. We present CorresNeRF, a novel method that leverages image correspondence priors computed by off-the-shelf methods to supervise NeRF training. We design adaptive processes for augmentation and filtering to generate dense and high-quality correspondences. The correspondences are then used to regularize NeRF training via the correspondence pixel reprojection and depth loss terms.

https://yxlao.github.io/corres-nerf/

Promising or Elusive? Unsupervised Object Segmentation from Real-world Single Images

https://vlar-group.github.io/UnsupObjSeg.html

Credits/Sources

Most of these links come from just a few incredible sources.  Please follow them:

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading