Welcome to State of AI Report 2024

“🚨New🚨 We are asking a fundamental question: how far can we push in-context learning for instruction following and how does it compare to fine-tuning? TL;DR: you should, of course, fine-tune, but the scaling laws are similar, at least in the small-sample regime: Key findings 

Addition is All You Need for Energy-Efficient Language Models: Reduce energy costs by 95% using integer adders instead of floating-point multipliers. : r/LocalLLaMA

“When training a visual encoder with self-supervised learning, we know for a fact that using a decoder with a reconstruction loss doesn’t work nearly as well as using a joint embedding architecture with feature prediction loss and a collapse prevention mechanism. This paper from” / X

[2410.02707] LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations

[2410.07073] Pixtral 12B

[2410.04671v1] CAR: Controllable Autoregressive Modeling for Visual Generation

“HyperCloning enables efficient transfer of knowledge from smaller pre-trained models to accelerate training of larger language models. **Original Problem** 🔍: Pre-training large language models from scratch is extremely slow and costly. Small models are cheaper to train but 

“Fantastic New Paper – PrefixQuant is the first to enable efficient per-tensor static quantization to outperform expensive per-token dynamic quantization. **Results** 📊: • W4A4KV4 Llama-3-8B: 7.43 WikiText2 perplexity, 71.08% average accuracy on 5 tasks • Outperforms QuaRot 

Production AI Engineering starts with Evals — with Ankur Goyal of Braintrust

Entropixplained

“software, before modern AI, was unique in that it was possible to create massively profitable/successful businesses with very little capital with LLMs, this is less true; we should expect a radical impact on the industry, as it looks a lot more similar to traditional industry” / X

“What is a Mixture of Experts (MoE), and why are they successful? @MaartenGr just published a new visual guide on the Mixture of Experts (MoE) to explain the two main components of MoE: Experts and the Router. đź‘€ TL;DR: đź§  MoE consists of multiple “expert” neural networks and a 

“After the results we published in August, there was tremendous interest in how the latest releases from @OpenAI and @GoogleDeepMind perform against our benchmarks. We evaluate: * OpenAI o1-preview and o1-mini * Google Gemini 1.5 Pro, Gemini 1.5 Flash (May release) (2/5)” / X

“I actively avoid any open-source projects with poor or non-existent test coverage. Untested code is non-working code. I don’t trust developers who don’t know how to automatically test their code.” / X

““Overall, we found no evidence of formal reasoning in language models (…). Their behavior is better explained by sophisticated pattern matching”” / X

The Era of 1-bit LLMs • Buttondown

Zoomtopia 2024: Unveiling AI-first work platform innovations | Zoom – Zoom

Differential Transformer | Hacker News

“The Top ML Papers of the Week (Oct 7 – Oct 13): – ToolGen – Astute RAG – MLE-Bench – Differential Transformer – Addition Is All You Need – Long-Context LLMs Meet RAG Read on for more:” / X

[2410.04733v1] PredFormer: Transformers Are Effective Spatial-Temporal Predictive Learners

[2410.03617] What Matters for Model Merging at Scale?

[2410.05270v1] Fine-Tuning CLIP’s Last Visual Projector: A Few-Shot Cornucopia

“This Nov-2023 Paper gives some good ideas on Prompt Caching and Low-Latency Inference using Prompt Markup Language (PML). Prompt Cache significantly reduces latency in time-to-first-token, ranging from 8X for GPU-based inference to 60X for CPU-based inference, all while 

2410.07095v2.pdf

chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/https://arxiv.org/pdf/2410.07095

[2410.05258] Differential Transformer

[2410.01784v1] OmniGenBench: Automating Large-scale in-silico Benchmarking for Genomic Foundation Models

JEDi

“FEATURE RELEASE🚨: You can now build your own Gumloop nodes with AI 300+ people have asked for this and it’s finally out! Check out how it works here: 

“Yeah, I don’t believe Large Language Models are smart. Two things going on: 1. They are fantastic memorization and interpolation machines that have memorized the entire Internet and can mix and match different pieces to answer questions. 2. We don’t fully understand how they” / X

The real data wall is billions of years of evolution

Three Subtle Examples of Data Leakage — LessWrong

Update on Reflection-70Bhttps://glaive.ai/blog/post/reflection-postmortem

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading