Times Square LED sign ticker reads “What’s next in DSPy 2.5? And DSPy 3.0?” A busy street. Sunny day. ideogram.ai

“The central purpose of the AI Engineer is turning existing foundation model capabilities into useful products. The reason this is a useful statement: – Puts useful constraints on safety, agi, regulation yapping outside of zone of influence – “Useful products” is the one thing 

“🆕 @latentspacepod: Is finetuning GPT4o worth it? w/ @AlistairPullen of @cosine_sh Betteridge’s law says no: with 59 different flavors of RAG, and >2million token context + prompt caching, it’s reasonable to believe that “in context learning is all you need”. But Genie is the 

“More people should play with base models. They have a distinct feeling of simulation and “intelligence grown from the world” rather than injection-molded from obedience that instruction-tuned models have. They are state-of-the-art world simulators, not question-answerers. 

“Powerful Paper for LLM significantly speeding up LLM Inference. MARLIN achieves near-optimal 3.87x speedup up to batch size 32 on NVIDIA A10 🤯 Key Insights 💡: • LLM inference remains memory-bound even at larger batch sizes • Careful pipelining and partitioning can maintain 

Coinbase CEO Brian Armstrong: AI ‘should have crypto wallets’

“I wrote a blogpost “On the speed of ViTs and CNNs”. Addresses the following concerns I often hear: – worry about ViTs speed at high resolution. – how high resolution do I need? – is it super important to keep the aspect ratio? I think @ylecun might like it too! Link below 

“🧵What’s next in DSPy 2.5? And DSPy 3.0? I’m excited to share an early sketch of the DSPy Roadmap, a document we’ll expand and maintain as more DSPy releases ramp up. The goal is to communicate our objectives, milestones, & efforts and to solicit input—and help!—from everyone. 

“The transformer-land and diffusion-land have been separate for too long. There were many attempts to unify before, but they lose simplicity and elegance. Time for a transfusion🩸to revitalize the merge!” / X

[2408.07541] DifuzCam: Replacing Camera Lens with a Mask and a Diffusion Model

“You’ve gotta learn how to use Docker. I don’t see how you could get into building and deploying software without a basic knowledge of using Docker images and containers. I learned this the hard way. Little did I know this was going to become one of the main differentiators in 

“Think you understand classifier-free diffusion guidance? Think again! These two papers beg to differ😁 

“I worked on Metamate last year with Aparna Ramani and Zach Rait, secured funding, brought together the initial team (@Vjeux @bolinfest, etc.) built out the ML side of things with @shahin_sefati . It’s been really cool to build a vertical GenAI product for Meta-internal” / X

Accurate computation of quantum excited states with neural networks | Science

“Zero-shot DUP prompting achieves SOTA results on math reasoning tasks across various LLMs. ✨ Problem 🔍: LLMs struggle with complex math word problems due to semantic misunderstanding, calculation errors, and step-missing errors. Prior studies focused on addressing calculation 

VMP: Versatile Motion Priors for Robustly Tracking Motion on Physical Characters

(2) The AI OS (Sept 2023 Recap) – by swyx & Alessio

“Introducing The AI Scientist: The world’s first AI system for automating scientific research and open-ended discovery! 

BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning

“Replace LoRA with Flexora (flexible low-rank adaptation) 🔥 Flexora’s flexible approach to LoRA fine-tuning yields superior results and reduces training parameters by up to 50% 🤯 Introduces adaptive layer selection for LoRA Key Insights 💡: • Selective layer fine-tuning can 

[2408.10914] To Code, or Not To Code? Exploring Impact of Code in Pre-training

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading