Times Square LED sign ticker reads “What’s next in DSPy 2.5? And DSPy 3.0?” A busy street. Sunny day. ideogram.ai
“The central purpose of the AI Engineer is turning existing foundation model capabilities into useful products. The reason this is a useful statement: – Puts useful constraints on safety, agi, regulation yapping outside of zone of influence – “Useful products” is the one thing
“🆕 @latentspacepod: Is finetuning GPT4o worth it? w/ @AlistairPullen of @cosine_sh Betteridge’s law says no: with 59 different flavors of RAG, and >2million token context + prompt caching, it’s reasonable to believe that “in context learning is all you need”. But Genie is the
“More people should play with base models. They have a distinct feeling of simulation and “intelligence grown from the world” rather than injection-molded from obedience that instruction-tuned models have. They are state-of-the-art world simulators, not question-answerers.
“Powerful Paper for LLM significantly speeding up LLM Inference. MARLIN achieves near-optimal 3.87x speedup up to batch size 32 on NVIDIA A10 🤯 Key Insights 💡: • LLM inference remains memory-bound even at larger batch sizes • Careful pipelining and partitioning can maintain
Coinbase CEO Brian Armstrong: AI ‘should have crypto wallets’
“I wrote a blogpost “On the speed of ViTs and CNNs”. Addresses the following concerns I often hear: – worry about ViTs speed at high resolution. – how high resolution do I need? – is it super important to keep the aspect ratio? I think @ylecun might like it too! Link below
“🧵What’s next in DSPy 2.5? And DSPy 3.0? I’m excited to share an early sketch of the DSPy Roadmap, a document we’ll expand and maintain as more DSPy releases ramp up. The goal is to communicate our objectives, milestones, & efforts and to solicit input—and help!—from everyone.
“The transformer-land and diffusion-land have been separate for too long. There were many attempts to unify before, but they lose simplicity and elegance. Time for a transfusion🩸to revitalize the merge!” / X
[2408.07541] DifuzCam: Replacing Camera Lens with a Mask and a Diffusion Model
“You’ve gotta learn how to use Docker. I don’t see how you could get into building and deploying software without a basic knowledge of using Docker images and containers. I learned this the hard way. Little did I know this was going to become one of the main differentiators in
“Think you understand classifier-free diffusion guidance? Think again! These two papers beg to differ😁
“I worked on Metamate last year with Aparna Ramani and Zach Rait, secured funding, brought together the initial team (@Vjeux @bolinfest, etc.) built out the ML side of things with @shahin_sefati . It’s been really cool to build a vertical GenAI product for Meta-internal” / X
Accurate computation of quantum excited states with neural networks | Science
“Zero-shot DUP prompting achieves SOTA results on math reasoning tasks across various LLMs. ✨ Problem 🔍: LLMs struggle with complex math word problems due to semantic misunderstanding, calculation errors, and step-missing errors. Prior studies focused on addressing calculation
VMP: Versatile Motion Priors for Robustly Tracking Motion on Physical Characters
(2) The AI OS (Sept 2023 Recap) – by swyx & Alessio
“Introducing The AI Scientist: The world’s first AI system for automating scientific research and open-ended discovery!
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
“Replace LoRA with Flexora (flexible low-rank adaptation) 🔥 Flexora’s flexible approach to LoRA fine-tuning yields superior results and reduces training parameters by up to 50% 🤯 Introduces adaptive layer selection for LoRA Key Insights 💡: • Selective layer fine-tuning can
[2408.10914] To Code, or Not To Code? Exploring Impact of Code in Pre-training





Leave a Reply