A floral poster in the style of Kehinde Wiley with a reseach paper theme and the word “Tech” in bold text
“@francoisfleuret Modern transformers are well-behaved according to well-behaved scaling laws (which are a function of (num. of tokens, num. of parameters). This allows one to find the hyperparameters at a smaller scale, and then keep scaling up parameters and data according to some power law.” / X
[2409.20081v1] ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
“Are we ready for a world where our data is exposed at a glance? @CaineArdayfio and I offer an answer to protect yourself here:
“🎊 We’re excited to announce our new export metrics integration, allowing you to easily export model inference metrics to your favorite observability platforms like @grafana Cloud! After working with our customers, including @DescriptApp, @rimelabs, and @Pixis_AI, we
Cross Capabilities of LLMs
Entropy-based sampling
“We’ve introduced Test Mode: Now you can update and iterate on your tables for free.
Table Extraction using LLMs: Unlocking Structured Data from Documents
What comes after? – by Rohit Krishnan – Strange Loop Canon
Why bigger is not always better in AI | MIT Technology Review
“now that 1B models are getting good, the potential for doing crazy inference time search is massive that’s what I’m most excited for in LLMs right now (and conditional compute more broadly)” / X
“wild that in cloudflare sqlite you can: – write synchronous queries but get the performance of async queries – write “dumb” O(n+1) queries but get the performance of O(1) query – rollback state to any point in last 30 days = cheap always on disaster recovery and it just fell
“The AI gold rush is here, and freelancers are poised to win big. Here’s why: 1. AI specialists are the new “lucky ones” – high demand, short supply 2. Complexity of AI systems = higher value for expertise 3. Every industry needs AI integration, but lacks in-house talent 4. AI” / X
“An exhaustive work in this 103 page long Synthetic Data Generation paper. “Comprehensive Exploration of Synthetic Data Generation: A Survey” 👨🔧 Surveys 417 Synthetic Data Generation (SDG) models over the last decade. 📌 Covers 20 distinct model types, further categorized into
“We’re hiring at @e2b_dev: ✶ Product/designer engineer ✶ Distributed systems engineer We work in-person in SF. Send me your CV and GitHub!
[2409.19977v1] Knowledge Graph Embedding by Normalizing Flows
“A super interesting Paper getting new values from good old RNN with a huge Computational Efficiency win 🥇 Finds that by removing their hidden state dependencies from their input, forget, and update gates, LSTMs and GRUs no longer need to backpropagate through time (BPTT) and
[2409.18786v1] A Survey on the Honesty of Large Language Models
[2410.01556v1] Integrative Decoding: Improve Factuality via Implicit Self-consistency
MaskLLM
Distributed Training of Deep Learning models – Part ~ 1
[2410.02525] Contextual Document Embeddings
[2409.20370] The Perfect Blend: Redefining RLHF with Mixture of Judges
[2410.01600] ENTP: Encoder-only Next Token Prediction
“This is how you turn on float8 training or inference on any Keras model: `.quantize(policy)` also works for any other supported form of quantization, e.g. “int8” (mind you, that one is inference-only)
“I ❤️ @AEStudioLA’s work on bio-inspired & other neglected approaches to AI safety “We add a loss function that says, do your task but also predict yourself. We don’t lose performance, but end up with simpler internals” Seems important! 👏@juddrosenblatt & @mikevaiana👏 🔗👇
“@francoisfleuret There’s three parts. 1. Fitting as large of a network and as large of a batch-size as possible onto the 10k/100k/1m H100s — parallelizing and using memory-saving tricks. 2. Communicating state between these GPUs as quickly as possible 3. Recovering from failures (hardware,” / X
Introducing My Reasoning Model: No Tags, Just Logic : r/LocalLLaMA
MIT spinoff Liquid debuts small, efficient non-transformer AI models | VentureBeat
“A nice 51 mint Blog on Transformers Inference Optimization Toolset Techniques span algorithmic improvements to low-level hardware optimizations. Effective implementation requires understanding both algorithms and hardware. As LLMs grow, new optimization methods continue emerging
“Thanks to @willknight, I’m discovering the Sundai Club, a group of Harvard and MIT students, developers and product managers who code for good. They built a tool to help reporters covering AI identify potentially interesting papers posted to the Arxiv. Check out the article:
“20. Programming Skills and Future Jobs (2h17m) A longer segment here because its insightful about the future of programming. Will there still be jobs for programmers in the future? Is programming becoming more fun? Less boilerplate and more creativity. Humans in the driving
“One reason the AI risk debate is so polarized is bad epistemology. Skeptics often shy away from cost-benefit reasoning under uncertainty. Conversely, many doomers are far too Bayesian. My new post (link in replies) argues that Bayesianism is very incomplete as an epistemology.
“LLM Pricing Space Updated! 🔥 LLM pricing keeps dropping! Included the latest updates from @OpenAI, @GoogleDeepMind, @MistralAI, @cohere, @Cloudflare, and more. 🤑 Highlights: > OpenAI: 01 preview & 01 mini added. > Google Deepmin: Gemini Pro 2x prices decreased, starting today. https://x.com/_philschmid/status/1841488046752997548





Leave a Reply