An office in the Flintstone’s prehistoric cartoon town of Bedrock with name “Amazon” on it. –ar 5:3 –style raw
Cohere
“Now you can deploy our Command R model series for your business use cases on Amazon Bedrock!
New Cohere Toolkit Accelerates Generative AI Application Development
“Extract data from financial queries📊 Challenge: given user query, extract ticker, year, etc from query. Then use ticker, etc. as filters for DB query. I used @cohere cmd r+ for extraction. It was wicked fast. How it works: • user submits query • cmd r+ extracts metadata
Meta/Llama
“I found a new tool calling champion Llama3 70b on @GroqInc Challenge: given user query, extract financial quarters and years. Example: “How did revenue change between Q4 2023 and year before that?” The 70b model: • passed the task • was very fast • had best pricing I
“AI Agents creating youtube content?! And a sneak peek of crewAI+?! @tonykipkemboi just dropped amazing content on @crewAIInc using @ollama and @GroqInc
“We’ve been blessed again with new LLaVA-like models based on LLaMA 3 & Phi-3 🤩 Also passes the baklava benchmark 🤝✅
“llama-3 70B takes 3rd place, replacing haiku. full results on right
“llama-3 models did very poorly on this benchmark, simply because their context length is *limited to 8k*. But… with zero-training (actually just a simple 2 line config) you can get 32k context out of llama-3 models with *exceptional* quality. llama-3 8B surpasses many models
“”Extending Llama-3’s Context Ten-Fold Overnight” – with only 3.5K synthetic training samples generated by GPT-4 📌 During training, the question-answer pairs for the same context are organized into a multi-turn conversation. The model is fine-tuned to correctly answer the
Result: Llama 3 EXL2 quant quality compared to GGUF and Llama 2 : r/LocalLLaMA
“Groq-powered inference for Llama 3 is now available on Poe! You can use Llama-3-70b-Groq and experience the state-of-the-art open source model with near-instant streaming. (1/2)
Nvidia has published a competitive llama3-70b QA/RAG fine tune : r/LocalLLaMA
Meta’s Llama 3 400b: Multi-modal , longer context, potentially multiple models : r/LocalLLaMA
Llama3_8B 256K Context : EXL2 quants : r/LocalLLaMA
Llamafile’s progress, four months in – Mozilla Hacks – the Web developer blog
“Researchers at @ICepfl & @YaleMed teamed up to build Meditron, an LLM suite for low-resource medical settings. With Llama 3, their new model outperforms most open models in its parameter class on benchmarks like MedQA & MedMCQA. More details ➡️
Anyone tried new dolphin-2.9-llama3-8b-256k? : r/LocalLLaMA
“Back-of-the-envolope for cost of Llama 3 4-bit 70B on an M2 Ultra with MLX: $0.2 / million tokens. Here’s my calculation, double check it: -4-bit 70B model generates at 15 toks/sec on an M2 Ultra – Consumes 60W power – $0.18 per Kilowatt-hour average for US 1000000 tokens * (1” / X
“FIRST EVER Long Context (128K) Llama-3 70B – Llama-3-Giraffe-70B From Abacus AI The biggest issue with the Llama-3 family is that they are fairly small context models. So, even though they are as performant as GPT-4 on critical benchmarks, they can’t be used in the real world
“LLaMA-70b inferencing using only a single GPU and achieving 1.69x-2.65x higher normalized inference throughput than the FP16 baseline. with Six-bit quantization (FP6) 🔥 Deepspeed has just recently released this Paper and also integrated the FP6 quantization – “FP6-LLM:
“Llama 3 degrades more than Llama 2 when quantized. Probably because Llama 3, trained on a record 15T tokens, captures extremely nuanced data relationships, utilizing even the minutest decimals in BF16 precision fully. Making it more sensitive to quantization degradation.
“Meta’s groundbreaking paper – “Better & Faster Large Language Models via Multi-token Prediction” ✨ Models trained with 4-token prediction are up to 3 times faster at inference, even with large batch sizes. 🔥 📌 Large language models such as GPT and Llama are trained with a
https://x.com/rohanpaul_ai/status/1785666587879444645″Nvidia has published a competitive llama3-70b question answering and retrieval-augumented generation (RAG) fine tune – ChatQA-1.5 Benches look good, but missing that “llama 3” in the name as per the requirements of meta 😃
“Performing Complex Financial Calculations with Agentic RAG 🧮📈 Let’s build a financial assistant that can calculate percentage evolution, compound annual growth rate (CAGR), and P/E ratios over unstructured financial reports, without any human data transformation! This post by
“⚡️Multi-Agent RAG YouTube Workshop⚡️ If there’s anything better than agentic RAG, it’s multi-agent RAG! In this event, we’ll explore the big idea behind “multi-agent” applications. These types of workflows combine multiple independent agents, which can be structured to work” / X
“A Reference Architecture for Advanced RAG with @llama_index and Bedrock 📖 If you’re looking to build advanced RAG in the AWS ecosystem, our code repo gives you the perfect reference material for both advanced parsing and agentic reasoning, while using a full suite of AWS
“Performing Complex Financial Calculations with Agentic RAG 🧮📈 Let’s build a financial assistant that can calculate percentage evolution, compound annual growth rate (CAGR), and P/E ratios over unstructured financial reports, without any human data transformation! This post by
“LlamaParse 🤝 Bedrock Knowledge Base If you’re an AI developer in the AWS/Bedrock ecosystem, you can still take advantage of the advanced parsing capabilities that LlamaParse provides to build advanced RAG over complex PDFs, and we have a full tutorial on doing so ✅
“A Reference Architecture for Advanced RAG with @llama_index and Bedrock 📖 If you’re looking to build advanced RAG in the AWS ecosystem, our code repo gives you the perfect reference material for both advanced parsing and agentic reasoning, while using a full suite of AWS
Self-Learning Llama-3 Voice Agent with Function Calling and Automatic RAG : r/LocalLLaMA
“Introducing Panza, a personalized LLM email assistant, running entirely on-device! [1/6] * Panza adapts LLaMA-3-8B to match your unique writing style; * Can be fine-tuned and executed on a single GPU (free Colab version available!). Give it a try:
“Introducing LlamaIndex.TS version 0.3! Featuring: ⭐️ Agent support including ReAct, Anthropic and OpenAI agents, as well as a generic AgentRunner class ⭐️ Standardized Web Streams compatible with React 19, Deno, and Node 22 ⭐️ More comprehensive type system ⭐️ Enhanced support
“AI Agents creating youtube content?! And a sneak peek of crewAI+?! @tonykipkemboi just dropped amazing content on @crewAIInc using @ollama and @GroqInc
Other Open Source News
Home Assistant’s next era begins now – The Verge
“Introduce OpenVoice V2 – a Text-to-Speech model that can clone any voice and speak in any language. Developed by MyShell and @MIT_CSAIL researchers. 🌐 Imagine your voice going global in multiple languages. 🔊 OpenVoice V2 breaks the language barrier and redefines voice https://x.com/myshell_ai/status/1783161876052066793

Heads up! You’ve scrolled to the end of this category. There may have been just one or two links (above), so go back up and double check to be sure you didn’t quickly scroll down past it.
Be Sure To Read This Week’s Main Post:
This week’s executive overview and top links are here:
AI News #31: Week Ending 05/03/2024 with Executive Summary and Top 95 Links
The post you just read is an deep dive extension of my weekly newsletter, This Week In AI, an executive summary of the top things to know in AI. Each week, I create an accessible overview for laypeople to feel confident they are conversant with the week’s AI developments. I include a curated list of must-click links of the week, to offer everyone a hands-on opportunity to explore the most intriguing updates in artificial intelligence across various categories, including robotics, imagery, video, AR/VR, science, ethics, and more. Beyond the overview, I post these topic-based deeper dives (below). If you haven’t read this week’s overview, I recommend starting there.
- Agents/Copilots
- Amazon
- Apple
- Artificial General Intelligence (AGI)
- Augmented and Virtual Reality (AR/VR)
- Autonomous Vehicles
- AI Audio
- Business and Enterprise AI
- Chips and Hardware
- Consumer Products
- Education
- Ethics/Legal Security
- Images/Photos
- International AI News
- Locally Run AI Models
- Mobile
- Meta
- Microsoft
- OpenAI
- Open Source
- Podcasts/YouTube
- Publishing and News
- Retrieval-Augmented Generation (RAG) News
- Robots and Embodiment
- Science and Medicine
- Video
- Vision/Multimodality
- X/Twitter/Grok
- Tech and Development
Credits/Sources

Most of these weekly links come from just a few prolific oversharing sources. Please follow them, as they work hard to find the news each week and they make it a lot easier for me to compile.
- Robert Scoble: https://x.com/Scobleizer
- Ethan Mollick: https://www.linkedin.com/in/emollick/
- Alan Thompson: https://lifearchitect.ai/
- Theoretically Media: https://www.youtube.com/@TheoreticallyMedia
- The Rundown: https://www.therundown.ai/
- Bilawal Sidhu: https://twitter.com/bilawalsidhu/
- TLDR: https://tldr.tech/ai
- Jeremiah Owyang: https://twitter.com/jowyang
- Nick St. Pierre: https://twitter.com/nickfloats
- Dr. Jim Fan: https://twitter.com/DrJimFan
- All About AI: https://www.youtube.com/@AllAboutAI
- Marshall Kirkpatrick: https://aitimetoimpact.com/
- AI News (Smol Talk): https://buttondown.email/ainews/archive/
For previous issues, please visit the archives!

Thanks for reading!





Leave a Reply