“Research on LLMs is moving quickly, and even models / techniques that have been state-of-the-art for a long time (e.g., GPT-4 and Mixtral) are being quickly dethroned. Here’s a list of my top ten AI developments (each with a brief summary) over the last few months… [1] DBRX is…  https://twitter.com/cwolferesearch/status/1774817920704512013

“A new method was able to delete 40% of LLM layers with no drop in accuracy. This makes them mich cheaper and faster. The method combines pruning, quantization and PEFT. They tested this across various open source models. Each family of models had a maximum amount of layers…  https://twitter.com/AlphaSignalAI/status/1774858806817906971 

“What are the LLMs with the most output tokens these days? GPT-4 and Claude 3 are both 4096. Gemini Pro 1.5 is 8192 This really matters for structured data extraction: even with 1m of input tokens you can’t scrape a big webpage into a CSV file if you run out of output tokens” / X – https://twitter.com/simonw/status/1774168227180147011 

“Google presents Mixture-of-Depths Dynamically allocating compute in transformer-based language models Transformer-based language models spread FLOPs uniformly across input sequences. In this work we demonstrate that transformers can instead learn to dynamically allocate  https://twitter.com/_akhaliq/status/1775740222120087847

The cost of AI reasoning over time. – by Karina Nguyen – https://semaphore.substack.com/p/the-cost-of-reasoning-in-raw-intelligence 

“Put the information you want the AI to comment into the prompt FIRST and only then ask the question. The performance differences you get from non-obvious differences in prompting approaches are quite large, but it is really hard to know what works well in advance.” / X – https://twitter.com/emollick/status/1774255147960430931 

“Mojo 🔥, the programming language that turns Python into a beast, went open-source. This is a huge step and great news for the Python and AI communities! With Mojo 🔥 you can write Python code or scale all the way down to metal code. It’s fast!  https://twitter.com/svpino/status/1774406305148805525 

– Your AI Product Needs Evals – https://hamel.dev/blog/posts/evals/ 

“🚀Building a Perplexity Style LLM Answer Engine: Frontend to Backend Tutorial This repo has been taking off 📈 over the past week – and for good reason Great introduction to building an answer engine from scratch! Video:  https://twitter.com/LangChainAI/status/1774502671669501973 

“I recorded a new tutorial: How to evaluate a RAG application. It’s a 50-minute YouTube video. I built everything from scratch. My goal with these videos is for you to learn, not memorize. I hope you find this approach helpful. Link in next tweet so X doesn’t bury this post.” / X – https://twitter.com/svpino/status/1774496892095009223 

“Huawei presents DiJiang: Efficient Large Language Models through Compact Kernelization Achieves comparable performance with LLaMA2-7B on various benchmark while requiring only about 1/50 pretraining cost  https://twitter.com/arankomatsuzaki/status/1774631022325350727

“📄Jamba whitepaper is out! The whitepaper details our in-depth ablations on this novel hybrid SSM-Transformer architecture, and how we chose to interleave Mamba, Transformer and MoE.  https://twitter.com/AI21Labs/status/1774824070053331093 

“FastEmbed now allows you to generate efficient and interpretable sparse vector embeddings using the SPLADE++ model. 🚀 Get started with @NirantK’s guide:  https://twitter.com/qdrant_engine/status/1774723490567860634 

“LangChain Financial Assistant 🦜 Our repo hit 100 stars this weekend. To celebrate, I added three Buffett-inspired tools today. • calculate owner earnings • calculate return on equity • calculate return on invested capital The neat thing about our agent is that you can…  https://twitter.com/virattt/status/1774909569723932850 

“instructor 1.0.0 is here  https://twitter.com/jxnlco/status/1774813661900558442 

“RAGFlow, the deep document understanding based RAG engine is open sourced now ✨ The library looks really good and it tries to understand semantic document structure and layout. RAGFlow could also work with LLMs deployed on premise  https://twitter.com/rohanpaul_ai/status/1774890566179774733 

“New LlamaIndex Webinar 🚨 – come learn how to do retrieval-augmented fine-tuning (RAFT)! Doing RAG is like taking an open-book exam without studying. It’s marginally better than a closed-book exam where the LLM has to memorize information beforehand (through fine-tuning). But…  https://twitter.com/llama_index/status/1774814982322172077 

“RAG is effective if the information retrieved from a database as a result of a query is relevant to the query and its application. “Advanced Retrieval for AI with @trychroma” teaches techniques to improve the relevancy of retrieved results. Join today:  https://twitter.com/DeepLearningAI/status/1774919695373603040 

“This is an excellent tutorial by @mesudarshan showing you how to build advanced PDF RAG with LlamaParse and purely local models for embedding, LLMs, and reranking (@GroqInc and FastEmbed by @qdrant_engine, flag-embedding-reranker) Having a good extraction step is super important…  https://twitter.com/llama_index/status/1774832426000515100 

Full Steam Ahead: The 2024 MAD (Machine Learning, AI & Data) Landscape – Matt Turck – https://mattturck.com/mad2024/ 

[2404.02082v1] WcDT: World-centric Diffusion Transformer for Traffic Scene Generation – https://arxiv.org/abs/2404.02082v1 

Paper page – Octopus v2: On-device language model for super agent – https://huggingface.co/papers/2404.01744 

Replit — Building LLMs for Code Repair – https://blog.replit.com/code-repair 

[2403.19928v2] DiJiang: Efficient Large Language Models through Compact Kernelization – https://arxiv.org/abs/2403.19928v2 

[2403.20101] RealKIE: Five Novel Datasets for Enterprise Key Information Extraction – https://arxiv.org/abs/2403.20101 

“@figma @skirano @everartai We just announced Code Repair, the world’s first low-latency program repair AI agent. Informed by Replit’s unique data on developer intuition, and grounded in real-world use cases to automatically fix your code in the background.  https://twitter.com/Replit/status/1775274756239176051?s=20 

JetMoE – https://research.myshell.ai/jetmoe 

[2404.02258] Mixture-of-Depths: Dynamically allocating compute in transformer-based language models – https://arxiv.org/abs/2404.02258 

How to win at Vertical AI – by Sangeet Paul Choudary – https://platforms.substack.com/p/how-to-win-at-vertical-ai 

With Brave Leo on iOS today, the browser AI assistant is now available on all platforms | Brave – https://brave.com/blog/leo-ios/ 

Bringing serverless GPU inference to Hugging Face users – https://huggingface.co/blog/cloudflare-workers-ai

Heads up! You’ve scrolled to the end of this category. There may have been just one or two links (above), so go back up and double check to be sure you didn’t quickly scroll down past it.

Be Sure To Read This Week’s Main Post:

This week’s executive overview and top links are here:

AI News #27: Week Ending 04/05/2024 with Executive Summary and Top 48 Links

The post you just read is an deep dive extension of my weekly newsletter, This Week In AI, an executive summary of the top things to know in AI. Each week, I create an accessible overview for laypeople to feel confident they are conversant with the week’s AI developments. I include a curated list of must-click links of the week, to offer everyone a hands-on opportunity to explore the most intriguing updates in artificial intelligence across various categories, including robotics, imagery, video, AR/VR, science, ethics, and more. Beyond the overview, I post these topic-based deeper dives (below). If you haven’t read this week’s overview, I recommend starting there.

Credits/Sources

Most of these weekly links come from just a few prolific oversharing sources. Please follow them, as they work hard to find the news each week and they make it a lot easier for me to compile.

For previous issues, please visit the archives!

Thanks for reading!

One response to “Tech and Development News: Week Ending 04/05/2024”

  1. […] X/Twitter/Grok: Grok is one of several AI’s developed by X, and it’s a bit blended in with Telsa and other Elon Musk technology. Not every week will have a Grok section, but like Meta, Google, Apple, and OpenAI, X will be in the news enough to have its own section.This week’s latest X news: https://ethanbholland.com/2024/04/06/tech-and-development-news-week-ending-04-05-2024/ […]

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading