“Here’s a breakdown of most of the top model scores on aider’s code editing benchmark: 84% Claude 3.5 Sonnet 10/22 80% o1-preview 77% Claude 3.5 Sonnet 06/20 72% DeepSeek V2.5 72% GPT-4o 08/06 71% o1-mini 68% Claude 3 Opus” / X
“We just launched our new database of Machine Learning Hardware! 🚨 This database covers key data on hardware used to train AI models during the deep learning era. So far, we have collected data on over 100 accelerators (GPUs and TPUs). Read on for some key insights! đź§µ
How I Studied LLMs in Two Weeks: A Comprehensive Roadmap | Towards Data Science
“Learn how to leverage LLMs for efficient report generation! 🚀📊 Report generation is a major use-case for our users, so we built a demo using Arxiv data. The key principles to get here are: ➡️ Techniques for extracting key information from complex documents, such as Pydantic
“interesting: truncating the lowest RoPE frequencies helps with length extrapolation, because for small sequence lengths the lower frequency channels change so infrequently that the model uses them as semantic channels, which hurts performance when you increase context length
“Reinforcement learning (RL) is commonly used to finetune LLMs based on human feedback, but did you know that we can also use RL to automatically improve our prompts? Prompt engineering is the act of discovering (via trial and error) prompts that perform well. Small changes or
“#AIPythonforBeginners has equipped learners with practical Python skills. From automating tasks to integrating LLMs into the process, learners all over the world are applying these skills in impactful ways. Ready to get started? Sign up today:
“Software development’s evolution into an “engineering” discipline has caused major problems for people and companies: 1. In software there’s little gap between designing and building In civil engineering, there is a big difference in knowing how to design a bridge and knowing
Sparse Crosscoders for Cross-Layer Features and Model Diffing
“What Matters In Transformers?” is an interesting paper (https://arxiv.org/abs/2406.15786) that finds you can actually remove half of the attention layers in LLMs like Llama without noticeably reducing modeling performance.
[2410.17001v1] Joint Point Cloud Upsampling and Cleaning with Octree-based CNNs
“1/ We are excited to share a milestone in our journey toward developing a physical AI foundation model. In a recent paper by the Archetype AI team, “A Phenomenological AI Foundation Model for Physical Signals,” we demonstrate how an AI foundation model can effectively encode and
GitHub – vectara/hallucination-leaderboard: Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents
“MASSIVE claim in this Paper for LLM Training 🤯 This new Linear-complexity Multiplication (L-Mul) algorithm can reduce energy costs by 95% for element-wise tensor multiplications and 80% for dot products in large language models, while maintaining or even improving precision
“🧨 diffusers 🤝 bitsandbytes ⚡️ We’re shipping native quantization support in diffusers, starting with bitsandbytes 🤗 Follow along this đź§µ to know more (inference & training) 1/n
[2410.13928] Automatically Interpreting Millions of Features in Large Language Models
[2410.11081] Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models





Leave a Reply