GitHub – maclong01/DeBiFormer: [ACCV 2024 ] Official code for “DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention”

“Some cool tricks in this model! – “We apply a LoRA projector to each shared MLP block…” (nice hack for weight sharing without using exact duplicates) – Only 3T tokens – Annealing fast on super high-quality data” / X

Introduction | W&B Weave

Weave is a lightweight toolkit for tracking and evaluating LLM applications, built by Weights & Biases.

“New work on linearizing LLMs! Like subquadratic capabilities? Like modern 7B+ LLMs? But don’t have the budget to pre-train billions of parameters on trillions of tokens to get subquadratic, 7B+ LLMs? Then check out LoLCATs, our new work led by @mzhangio that *converts existing* 

“✨ CrewAI 0.74.2 is out! ✨ 📈 100x Faster to install dependencies! HUGE! 🚣 UV migration 🛠️ Adapt Tools CLI to UV 🧠 New Memory Base 📃 Update Docs and Bug fixes RT please? 🙏 More to come!” / X

“Just put together a short Jupyter notebook with tips and tricks for reducing memory usage when loading larger and larger models (like LLMs) in PyTorch: 

“LLMs demonstrate ability to process multiple in-context learning tasks in a single inference pass. Reveals LLMs’ capacity for task superposition **Original Problem** 🔍: LLMs demonstrate remarkable in-context learning capabilities, but their ability to perform multiple 

“A common misconception about Transformers is to believe that they’re a sequence-processing architecture. They’re not. They’re a *set-processing* architecture. Transformers are 100% order-agnostic (which was the big innovation compared to RNNs, back in late 2016 — you compute” / X

“EPIC data lab @UCBEPIC advance at UC Berkeley kicks off with JD @jdzamfi talking about the role of exposing problem and solution alternatives in LLM-assisted interface design. Moral: helpful but easy to overwhelm!” / X

“@salomartin @AnimaAnandkumar No, this is intuition-guided reasoning where the reasoning is provided by the Lean theorem prover (a symbolic, discrete search system of great sophistication) and the intuition is provided by a LLM. At a glance it seems like very good work, and it follows closely the template of” / X

“the best and worst aspect of modern deep learning is the culture of empiricism it’s not required to have a theoretical justification for proposed changes as long as they “work” which leads to the incentive to have weak baselines (happened allllll the time in RL research)” / X

“been seeing a lot of cool use cases for @ottogrid_ai. One that keeps coming is Contact Information Lookup Instead of spending all day researching contact information for each person manually, you can upload a hundred or even thousands of names and get all that information in 

[2410.11181v1] DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection

LoLCATs Blog Part 2: How to Linearize LLMs for Me and You · Hazy Research

“Released diagen yesterday, but how does it work? 1. Generate @terrastruct d2 diagrams with the model of your choice. Sonnet seems best, o1 seems needlessly expensive, gemini-flash is insane if you do a few rounds of visual reflection. What’s visual reflection? 👇” / X

“The gradient accumulation fix is now in the main branch of transformers! Thank you to the entire @huggingface team, especially @TheZachMueller and @art_zucker for collabing with us to fix it! 🤗🦥” / X

“Instead of finetuning your LLMs, try dynamic few-shot prompting instead 💡 With dynamic few-shot prompting, instead of injecting a fixed set of examples into the prompt, you retrieve a dynamic set of examples based on the query – so you find relevant examples that are relevant 

“nanoGPT speedrun: Nice work from @kellerjordan0 adapting the nanoGPT/llmc PyTorch training code into a benchmark training a 124M Transformer to a fixed validation loss target. Current SOTA is 3.8X more token-efficient training (2.7B vs. 10B tokens)” / X

Introducing the prompt() Function: Use the Power of LLMs with SQL! – MotherDuck Blog

INTELLECT–1: Launching the First Decentralized Training of a 10B Parameter Model

“Automated Zero-cost proxy design enhances efficiency in evaluating language model architectures. **Original Problem** 🔍: Existing Zero-cost (ZC) proxies for Neural Architecture Search heavily rely on expert knowledge and trial-and-error, limiting their effectiveness for 

“Really happy to see that the LLM Engineer’s Handbook is #1 New Release in Neural Networks We put an insane amount of work into this book. I hope it’ll help a new generation of LLM engineers and tinkerers to build production-level AI systems. 

“Model / weight merging is the most simple, practical, and efficient way to combine the skills of multiple LLMs. The best example of merging being used for this purpose is Prometheus-2… TL;DR: Model merging works surprisingly well for combining the skills of separate LLMs; e.g., 

“The word frequencies in all languages follow a power law (Zipfian) distribution. This means that, during pretraining, the LLM sees a small number of words a lot, and lots of different words vary rarely 2/n 

“Is the Transformers Architecture being challenged? Earlier this week, @ZyphraAI released a new state-of-the-art 7B LLM outperforming @AIatMeta Llama 3.1, @GoogleDeepMind Gemma 2, and @MistralAI using hybrid SSM-attention architecture. 👀 TL;DR: 🧠 Hybrid Architecture using 6x 

“Let’s fly through the stable diffusion embedding space together, each image takes ~30ms to generate: (Many thanks to @Kveykva for doing most of this frontend work) 

“wtf, @mikeshou1’s 1B omnimodel seems to beat @ArmenAgha’s 34B Chameleon in VQA AND is competitive with SDXL etc image gen models AND can generate keyframes for video gen (eg Sora) all in the same model, on 35m labeled images + RefinedWeb chat wat is this sorcery if 

[2410.10133v1] TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control

Welcome to State of AI Report 2024

[2410.11081] Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models

“Super excited to share our work on LOTUS, a query engine for reasoning over large corpuses of data with LLMs! Joint work w/ the amazing @sid_jha1, @matei_zaharia & @guestrin Read the paper: 

“In-Context Reinforcement Learning (ICRL) unlocks new learning paradigms for LLMs, enabling adaptation through reward signals alone, without parameter updates. This paper’s algorithm increases test-time compute, as well as a compute-bound approximation. **Original Problem** 🤔: 

“torchtitan repo is kinda a miracle? all the parallelism you could need and i dont think it looks like theres any changes you need to make to the actual model itself.” / X

[2406.20094v1] Scaling Synthetic Data Creation with 1,000,000,000 Personas

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading