“The current AI wave is a lot of fun, but don’t forget about the fundamentals: Learn how neural networks work, loss functions, optimization techniques, activation functions, and the art of training and evaluating these networks. That’s how you prepare for the next wave.” / X

“Would you look at that, sigmoid loss with bias is a huge win! Sounds familiar? 😁” / X

OmniParser (w/ Transformers.js)

Recraft introduces a revolutionary AI model that thinks in design language

“Reminder for those who think they can create anything they want using AI without actually knowing how to build software. 

“Python is now #1 like it should be. 

“@random_walker 1% productivity boost is pretty big though? Especially if only 3% used ai??” / X

We’re forking Flutter. This is why.

“tinygrad is running the funniest goodhart around right now they’re obsessed with talking about how their library uses fewer lines of code than pytorch, so their codebase is growing horizontally instead of vertically some parts are borderline unreadable to humans 

“The Top ML Papers of the Week (Oct 28 – Nov 3): – MrT5 – SimpleQA – Multimodal RAG – o1 Replication Journey – Geometry of Concepts in LLMs – Relaxed Recursive Transformers Read on for more:” / X

Easier, Better, Faster, Cuter · Hazy Research

Enchant whitepaper: Breaking the data wall between lab and clinic

[2211.17192] Fast Inference from Transformers via Speculative Decoding

[2410.23856v1] Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?

“Octoverse 2024 report is out for GitHub data. 🌍 Developer Growth & Distribution – 5.2B total contributions across 518M projects in 2024 – 98% YoY growth in AI projects – Python overtook JavaScript as most used language – US maintains largest developer population – India 

“We wrote a paper: estimating body and hand motion from a pair of glasses! 👓 website: 

GitHub Next | GitHub Spark

[2410.16256v1] CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution

“Very interesting revelations in this paper. 💡 Mixture-of-Experts (MoE) trade reasoning power for memory efficiency in LLM architectures More experts don’t make LLMs smarter, just better at memorizing 🤔 Original Problem: Mixture-of-Experts (MoE) architecture lets LLMs scale 

Fine-tuning LLMs to 1.58bit: extreme quantization made easy

[2410.20750v1] ODRL: A Benchmark for Off-Dynamics Reinforcement Learning

“Future ML specialization: Inference or Training? Very soon training LLMs will become a domain of a few companies and there will be very little need in experts in LLM training. Especially when LLMs will be at the level of CV cats-vs-dogs quality. Inference expertise on the other” / X

“🏎️ Want faster, better quantized LLMs? Check out QTIP, led by @tsengalb99! It achieves a SOTA combination of quality and inference speed using incoherence processing & trellis coded quantization—outperforming methods like QuIP#! 👉 

“Excited to open-source a new hallucinations eval called SimpleQA! For a while it felt like there was no great benchmark for factuality, and so we created an eval that was simple, reliable, and easy-to-use for researchers. Main features of SimpleQA: 1. Very simple setup: there 

Using Reinforcement Learning and $4.80 of GPU Time to Find the Best HN Post Ever (RLHF Part 1) – OpenPipe

New from Universe 2024: Get the latest previews and releases – The GitHub Bloghttps://github.blog/news-insights/product-news/universe-2024-previews-releases

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading