New LLM optimization technique slashes memory costs up to 75% | VentureBeat

New LLM optimization technique slashes memory costs up to 75%

“I’m near certain a straight line from the max LR at 2k to 0 LR at the end would have worked better. Linear Decay is the One True Schedule” / X
https://x.com/aaron_defazio/status/1872481458184745374

[2408.09674v1] Implicit Grid Convolution for Multi-Scale Image Super-Resolution
https://arxiv.org/abs/2408.09674v1

“New from Meta FAIR — Byte Latent Transformer: Patches Scale Better Than Tokens introduces BLT, which for the first time, matches tokenization-based LLM performance at scale with significant improvements in inference efficiency & robustness. Paper ➡️
https://x.com/AIatMeta/status/1872735578846052817

“Newly published research for generative retrieval for recommendations from teams at Meta. – Preference Discerning with LLM-Enhanced Generative Retrieval ➡️
https://x.com/AIatMeta/status/1873798477928714460

“This paper examines how open-source LLMs are challenging closed-source models through innovative techniques like Low-Rank Adaptation and Flash Attention. 🛠️ Topics discussed in this Paper: → Open-source models like LLaMA and BLOOM leverage community-driven development and
https://x.com/rohanpaul_ai/status/1873959124830236791

“We played around with this a bit while I was at FAIR. Essentially we saw router collapse (router always picked single expert in TC) in the earlier layers if you didn’t very carefully balance experts/upcycle or use EC. We saw a similar phenomena in the very last layers as well.” / X
https://x.com/ArmenAgha/status/1872426813865201700

“Is there a clear analysis of how gradient descent manages to train a “a top-k routing MoE” ? It is simply that if the k-th looks good, it gets moved to a better rank, and by some magic it may move something beyond k into the top k elsewhere?” / X
https://x.com/francoisfleuret/status/1872370360307568964

“Let’s start the new year with a buffet of papers 🎉. I rarely recommend reading list but this one covers it well for AI practitioners. All you can eat!” / X
https://x.com/DrJimFan/status/1874490807652356377

⭐️ Fast LLM Inference From Scratch
https://andrewkchan.dev/posts/yalm.html

Chain of Continuous Thoughts | Ben Congdon
https://benjamincongdon.me/blog/2024/12/14/Chain-of-Continuous-Thoughts/

Computing inside an AI | Will Whitney
https://willwhitney.com/computing-inside-ai.html

“LLMs think better with connected evidence chains Chain of Evidence (CoE) helps LLMs make better decisions by ensuring knowledge pieces are both relevant and mutually supportive, like evidence in a criminal case. —– 🔍 Original Problem: → LLMs struggle with outdated
https://x.com/rohanpaul_ai/status/1873806409307218388

rnoti-p6.pdf

Click to access rnoti-p6.pdf

Cerebras Demonstrates Trillion Parameter Model Training on a Single CS-3 System – Cerebras
https://cerebras.ai/press-release/cerebras-demonstrates-trillion-parameter-model-training-on-a-single-cs-3-system

[2411.02306v1] Targeted Manipulation and Deception Emerge when Optimizing LLMs for User Feedback
https://arxiv.org/abs/2411.02306v1

“Teaching LLMs the art of knowing their knowledge boundaries LLMs struggle with reliability when they lack knowledge, leading to hallucinations. This paper introduces Contrastive Decoding with Abstention (CDA), enabling models to either generate accurate responses or abstain when
https://x.com/rohanpaul_ai/status/1873805807537840408

“@dh7net Evals are *definitely* not, however, “useless”. Some evals are actually very useful. You do, however, need to study the eval datasets closely, and the cost used to score against them, to understand their limitations.” / X
https://x.com/jeremyphoward/status/1872411782884802803

The AI revolution is running out of data. What can researchers do?
https://www.nature.com/articles/d41586-024-03990-2

[2411.06790v1] Large-scale moral machine experiment on large language models
https://arxiv.org/abs/2411.06790v1

2412.00396v1
https://arxiv.org/pdf/2412.00396v1

[2411.13145v1] Globally Correlation-Aware Hard Negative Generation
https://arxiv.org/abs/2411.13145v1

[2411.10083v1] Xmodel-1.5: An 1B-scale Multilingual LLM
https://arxiv.org/abs/2411.10083v1

[2410.05265v1] PrefixQuant: Static Quantization Beats Dynamic through Prefixed Outliers in LLMs
https://arxiv.org/abs/2410.05265v1

“The paper behind ModernBERT Makes encoder models fast and powerful again with smart architecture choices. ModernBERT introduces efficient encoder-only transformers with modern optimizations, achieving state-of-the-art performance while maintaining fast inference and low memory
https://x.com/rohanpaul_ai/status/1872712284302393532

[2408.03219] Learning to Learn without Forgetting using Attention
https://arxiv.org/abs/2408.03219

“recently started using ghstack to manage my pull requests. at my last job I used to just commit on my main branch and open a PR for each commit — this makes it easy to do the same with GitHub
https://x.com/vikhyatk/status/1872394404398588225

huggingface/smolagents: 🤗 smolagents: a barebones library for agents. Agents write python code to call tools and orchestrate other agents.
https://github.com/huggingface/smolagents

[2404.01332] Explaining Large Language Models Decisions Using Shapley Values
https://arxiv.org/abs/2404.01332

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading