A dusty book on top of a computer with the word “tech papers” on the cover.

“It’s a bit sad and confusing that LLMs (“Large Language Models”) have little to do with language; It’s just historical. They are highly general purpose technology for statistical modeling of token streams. A better name would be Autoregressive Transformers or something. They” / X

“If you’ve followed my recent posts on model merging, I just published a long-form survey on this topic. It covers 50+ papers from the 1990s until now, including everything from basic concepts to the recent application of model merging to LLM alignment. See image for details! 

[2409.11355] Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think

“What if self-attention isn’t the end-all be-all? 🤔 Paper Podcast – “Masked Mixers For Language Generation And Retrieval” **Key Insights from this Paper** 💡: • Transformers exhibit poor input representation accuracy in deeper layers • Masked mixers with convolutions retain 

“I’m doing a podcast with the Cursor team. If you have questions / feature requests to discuss (including super-technical topics) let me know! For those not familiar, Cursor is a code editor based on VSCode that adds a lot of powerful features for AI-assisted coding. I’ve been” / X

A future full of opportunities, Made On YouTube – YouTube Blog

“📚 Paper Podcast – “Chain of Thought Empowers Transformers to Solve Inherently Serial Problems” CoT enables transformers to expand their problem-solving capabilities beyond parallel-only limitations. Augmenting transformer expressiveness, particularly for inherently sequential 

“Can we use incorrect results to improve the Chain of Thoughts and LLMs? Yes, V-STaR utilizes preference pairs generated during the self-reflection to train a Verifier with DPO to judge the correctness of model-generated solutions during inference. 👀 Implementation 1️⃣ Select a 

“Caching will make your LLM application cheaper and faster to run. But caching is hard. As the famous saying goes, “There are 2 hard problems in computer science: cache invalidation, naming things, and off-by-1 errors.” Here is how caching works at a very high level: 1. A new 

“This one log is more valuable than 10 papers on MCTS for mathematical reasoning, and completely different from my speculation meditate on it as long as it takes 

“Confirmation of one of my points of analysis for the future of data annotation companies: even the best still end up with wild amounts of language model outputs. Prevention/detection is near impossible. H/t to @jefrankle for sharing 

“Today, we release several Moshi artifacts: a long technical report with all the details behind our model, weights for Moshi and its Mimi codec, along with streaming inference code in Pytorch, Rust and MLX. More details below 🧵 ⬇️ Paper: 

“🧠 AI models: baking without recipes? 🍰 This is probably the clearest explanation you can get on model training and how inputs affect outputs: @mmitchell_ai used a clever analogy during her Senate testimony to explain how AI models are created. 🥚 Baking: more eggs = puffier 

[2409.12917] Training Language Models to Self-Correct via Reinforcement Learning

[2409.10038v1] On the Diagram of Thought

[2409.12640] Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries

“Epic Arnie-approved weekend at our first ever @weights_biases SF hackathon Winners demo recordings to come including: 🥇 Knowledge Graph evals (featuring prolog 👀) 🥈OptoPrompt with a 👌prompt optimizer 🥉Lets Get Creative with glimpses into measuring creativity of LLMs 

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading