An 80s print out from a old printer with dot matric text reading “Tech Papers, Training, and Development”
“@Yuchenj_UW @NousResearch DiLoCo is a variant of federated averaging, where it requires nodes to still perform a full-bandwidth synchronization every N steps. DisTrO is a drop-in to any single optimizer & low-bandwidth synchronization every step, and could even be used in tandem with DiLoCo!” / X
“Scaling Laws, Economics and the AI “Game of Emperors.”” / X
Box AI Developer Zone – Box Developer Documentation
Data Exfiltration from Slack AI via indirect prompt injection
“Quantize Llama-3.1-8B with AutoRound library AutoRound lib is a very useful work from from @intel Neural Compressor team. ✨ `pip install auto-round` Advanced Quantization Algorithm for LLMs. This is official implementation of “Optimize Weight Rounding via Signed Gradient
Releasing Re-LAION 5B: transparent iteration on LAION-5B with additional safety fixes | LAION
“Struggling to keep up with AI’s breakneck pace? I sift through tons of AI sources every day. Here’s what catched my attention today. 🧶” / X
(2) Inference is FREE and INSTANT – Source Code by Fume
“Last Week in AI was on fire! 🔥 From serverless LLMs to long-form content generation, here are the highlights: 👉
GitHub – NousResearch/Hermes-Function-Calling
“Model merging is a popular research topic with applications to LLM alignment and specialization. But, did you know this technique has been studied since the 90s? Here’s a brief timeline… (Stage 0) Original work on model merging dates back to the 90s [1], where authors showed
“From LoRA to iLoRA 😆 – Nice proposal in this paper. iLoRA personalizes LLM recommendations by integrating LoRA with Mixture of Experts for improved accuracy. Instance-wise LoRA tailors recommendations to individual users, enhancing sequential recommendation performance.
“It’s 2024, and many teams out there are still stuck in basic mode. Few teams can build Machine Learning systems at scale. Training a model is the easy part, but that’s maybe 2% of the work. Building a system that works at scale requires engineering practices that have been” / X
GameNGen
Zyphra
AI companies are pivoting from creating gods to building products. Good.
Announcing Together Inference Engine 2.0 with new Turbo and Lite endpoints
Falling LLM Token Prices and What They Mean for AI Companies
“Does style matter over substance in Arena? Can models “game” human preference through lengthy and well-formatted responses? Today, we’re launching style control in our regression model for Chatbot Arena — our first step in separating the impact of style from substance in
Why should anyone boot you up? | Onur Solmaz blog
“Inference: 20 tokens per second per user is all you need. The interesting thing about online inference is that unlike normal webserving it doesn’t have to be as fast as possible, since it doesn’t have to return the full generated response at once. Moreover depending on the
“AI in the IDE is great. But also pay attention to AI in the Command Line One perspective I find helpful in making sense of the tech shift into AI, is to sometimes think of it as your existing software tools becoming better and smarter (as opposed to a distinct technology). For” / X
Introducing Pharia-1-LLM: transparent and compliant
How Fireworks evaluates quantization precisely and interpretablyhttps://fireworks.ai/blog/fireworks-quantization





Leave a Reply