Welcome to State of AI Report 2024
“🚨New🚨 We are asking a fundamental question: how far can we push in-context learning for instruction following and how does it compare to fine-tuning? TL;DR: you should, of course, fine-tune, but the scaling laws are similar, at least in the small-sample regime: Key findings
Addition is All You Need for Energy-Efficient Language Models: Reduce energy costs by 95% using integer adders instead of floating-point multipliers. : r/LocalLLaMA
“When training a visual encoder with self-supervised learning, we know for a fact that using a decoder with a reconstruction loss doesn’t work nearly as well as using a joint embedding architecture with feature prediction loss and a collapse prevention mechanism. This paper from” / X
[2410.02707] LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
[2410.07073] Pixtral 12B
[2410.04671v1] CAR: Controllable Autoregressive Modeling for Visual Generation
“HyperCloning enables efficient transfer of knowledge from smaller pre-trained models to accelerate training of larger language models. **Original Problem** 🔍: Pre-training large language models from scratch is extremely slow and costly. Small models are cheaper to train but
“Fantastic New Paper – PrefixQuant is the first to enable efficient per-tensor static quantization to outperform expensive per-token dynamic quantization. **Results** 📊: • W4A4KV4 Llama-3-8B: 7.43 WikiText2 perplexity, 71.08% average accuracy on 5 tasks • Outperforms QuaRot
Production AI Engineering starts with Evals — with Ankur Goyal of Braintrust
Entropixplained
“software, before modern AI, was unique in that it was possible to create massively profitable/successful businesses with very little capital with LLMs, this is less true; we should expect a radical impact on the industry, as it looks a lot more similar to traditional industry” / X
“What is a Mixture of Experts (MoE), and why are they successful? @MaartenGr just published a new visual guide on the Mixture of Experts (MoE) to explain the two main components of MoE: Experts and the Router. đź‘€ TL;DR: đź§ MoE consists of multiple “expert” neural networks and a
“After the results we published in August, there was tremendous interest in how the latest releases from @OpenAI and @GoogleDeepMind perform against our benchmarks. We evaluate: * OpenAI o1-preview and o1-mini * Google Gemini 1.5 Pro, Gemini 1.5 Flash (May release) (2/5)” / X
“I actively avoid any open-source projects with poor or non-existent test coverage. Untested code is non-working code. I don’t trust developers who don’t know how to automatically test their code.” / X
““Overall, we found no evidence of formal reasoning in language models (…). Their behavior is better explained by sophisticated pattern matching”” / X
The Era of 1-bit LLMs • Buttondown
Zoomtopia 2024: Unveiling AI-first work platform innovations | Zoom – Zoom
Differential Transformer | Hacker News
“The Top ML Papers of the Week (Oct 7 – Oct 13): – ToolGen – Astute RAG – MLE-Bench – Differential Transformer – Addition Is All You Need – Long-Context LLMs Meet RAG Read on for more:” / X
[2410.04733v1] PredFormer: Transformers Are Effective Spatial-Temporal Predictive Learners
[2410.03617] What Matters for Model Merging at Scale?
[2410.05270v1] Fine-Tuning CLIP’s Last Visual Projector: A Few-Shot Cornucopia
“This Nov-2023 Paper gives some good ideas on Prompt Caching and Low-Latency Inference using Prompt Markup Language (PML). Prompt Cache significantly reduces latency in time-to-first-token, ranging from 8X for GPU-based inference to 60X for CPU-based inference, all while
2410.07095v2.pdf
chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/https://arxiv.org/pdf/2410.07095
[2410.05258] Differential Transformer
[2410.01784v1] OmniGenBench: Automating Large-scale in-silico Benchmarking for Genomic Foundation Models
JEDi
“FEATURE RELEASE🚨: You can now build your own Gumloop nodes with AI 300+ people have asked for this and it’s finally out! Check out how it works here:
“Yeah, I don’t believe Large Language Models are smart. Two things going on: 1. They are fantastic memorization and interpolation machines that have memorized the entire Internet and can mix and match different pieces to answer questions. 2. We don’t fully understand how they” / X
The real data wall is billions of years of evolution
Three Subtle Examples of Data Leakage — LessWrong
Update on Reflection-70Bhttps://glaive.ai/blog/post/reflection-postmortem





Leave a Reply