the floodgates are open at a huge damn, water rushes out spectacularly. The damn is inscribed with the word “Meta” in large letters etched across its concrete wall. –ar 16:9 –style raw

Hugging Face

“we just shipped HuggingChat on iOS 💬 The app is super polished and gives you access to the community’s best open AI models, on the go. Give it a try! link to Appstore below ⤵️  https://twitter.com/julien_c/status/1780929979234017651

Meta

“The upcoming Llama-3-400B+ will mark the watershed moment that the community gains open-weight access to a GPT-4-class model. It will change the calculus for many research efforts and grassroot startups. I pulled the numbers on Claude 3 Opus, GPT-4-2024-04-09, and Gemini.  https://twitter.com/DrJimFan/status/1781006672452038756

“Because anyone can work with them, open models are likely to improve very quickly, creating a lot of capabilities focused on factors ranging from speed to costs. Here is the new Llama 3 70B being served by Groq (with a q) at 224 tokens/second. This is real-time of me using it.  https://twitter.com/emollick/status/1781371768835600849

“Congrats to @AIatMeta on Llama 3 release!! 🎉  https://twitter.com/karpathy/status/1781028605709234613

“Introducing Meta Llama 3: the most capable openly available LLM to date. Today we’re releasing 8B & 70B models that deliver on new capabilities such as improved reasoning and set a new state-of-the-art for models of their sizes. Today’s release includes the first two Llama 3…  https://twitter.com/AIatMeta/status/1780997403979735440

“I think Meta and Llama-3 is the final nail in the coffin to several misconceptions I’ve been fighting against for the last year. Llama-3 Chat was trained on over 10M Instruction/Chat samples, and is one of the only finetunes that shows significant improvements to MMLU.…  https://twitter.com/Teknium1/status/1781345814633390579 

“Meta is playing the long game and will go down in history as a seminal company! Most AI innovation in the OSD ecosystem will happen on the Llama architecture! The probability that the next breakthrough happens because of these Llama models is very high! Secretive closed…” / X – https://twitter.com/bindureddy/status/1781152808072626460 

“HISTORIC MOMENT! Llama-3 70B numbers are INSANE! At 82.0 MMLU, it’s FAR AND AWAY the best OSS Model. GSM-8K, Math, and Human Eval are MIND BLOWING as well. The OSS community is definitely going to beat GPT-4 in a matter of weeks!! Xmas came very very early  https://twitter.com/bindureddy/status/1780993893645132228 

“Meta released their open source AI, Llama 3, today. As a key leader in LLMs, their models are often the most advanced open source ones out there. Based on benchmarks, the current model is not quite GPT-4 class, but their larger one (still training) will reach GPT-4 level.  https://twitter.com/emollick/status/1780994637085221226 

“”Llama 3 is a very capable looking model release from Meta. Sticking to fundamentals, spending a lot of quality time on solid systems and data work, exploring the limits of long-training models. Also very excited for the 400B model, which could be the first GPT-4 grade open…” / X – https://twitter.com/_Borriss_/status/1781033686852391125 

Llama 3 is not very censored · Ollama Blog – https://ollama.com/blog/llama-3-is-not-very-censored

Meta Llama 3 – https://llama.meta.com/llama3/

“Some technical details: – standard decoder-only transformer – Llama 2’s vocab is 128K tokens – trained on sequences of 8k tokens – applies grouped query attention (GQA) – pretrained on over 15T tokens – post-training includes a combination of SFT, rejection sampling, PPO, and…” / X – https://twitter.com/omarsar0/status/1780992539891249466 

Introducing Meta Llama 3: The most capable openly available LLM to date – https://ai.meta.com/blog/meta-llama-3/ 

“Llama 3 is officially the fastest model from release to #1 trending on Hugging Face – in just a few hours. 30,000 new models have been released based on llama 1 & 2 so I can’t wait to see the impact that the third and most powerful version will have on the ecosystem! 🚀🚀🚀  https://twitter.com/ClementDelangue/status/1781068939641999388 

Mistral

“We just released Mixtral-8x22B-v0.1 and Mixtral-8x22B-Instruct-v0.1: – Free to use under Apache 2.0 license – Outperforms all open models – Native function calling – Masters English, French, Italian, German and Spanish. – Seq_len = 64K  https://twitter.com/dchaplot/status/1780598435823198366

“The official Apache 2 Mixtral 8x22B Instruct model is out! 🔥 🌍Multilingual (en/fr/it/de/es) 🧠Math and code capabilities ✏️Native function calling ⚡️39B active params 🤯64k context window Model:  https://twitter.com/osanseviero/status/1780595541711454602 

Cheaper, Better, Faster, Stronger | Mistral AI | Frontier AI in your hands – https://mistral.ai/news/mixtral-8x22b/ 

Mistral, an OpenAI Rival in Europe, in Talks to Raise Capital at a $5 Billion Valuation — The Information – https://www.theinformation.com/articles/mistral-an-openai-rival-in-europe-in-talks-to-raise-capital-at-a-5-billion-valuation

France’s Mistral AI seeks funding at $5 bln valuation, The Information reports | Reuters – https://www.reuters.com/technology/frances-mistral-ai-seeks-funding-5-bln-valuation-information-reports-2024-04-17/ 

“Mixtral 8x22B Instruct is out. It significantly outperforms existing open models, and only uses 39B active parameters (making it significantly faster than 70B models during inference). 1/n  https://twitter.com/GuillaumeLample/status/1780602023203029351 

“I’m pretty confident that strong fine-tunes of Mixtral-8x22B will significantly close this gap even further E.g. our Zephyr recipe involved just 7k examples and expanding this to a diverse corpus like OpenHermes with ~1M examples will likely produce something that’s competitive…” / X – https://twitter.com/_lewtun/status/1779804085677404583

“Mixtral 8x22B Has The Best Cost-To-Performance Ratio. > great base model > apache 2 license > very good MMLU combined with the MOE architecture gives it one of the best performance ratios > fine-tunes on this base model will get us past GPT-4 performance Expecting Llama-3 to…  https://twitter.com/bindureddy/status/1780609164223627291

“Awesome performance from today’s Mistral’s release of Mixtral 8x22B Instruct. math performance, with a score of 90.8% on GSM8K maj@8 and a Math maj@4 score of 44.6%. The most efficient performance/cost ratio on MMLU.  https://twitter.com/rohanpaul_ai/status/1780605940842021327 

“There are so many new open-source AI models that there is not enough space on this graph anymore!! And the best part is that Llama-3 isn’t out yet!  https://twitter.com/bindureddy/status/1780797091465527736 

“The latest model from @MistralAI, the hugely powerful 8x22b, defines the state of the art in open models and is now available! As usual, we have day 0 support: check out @ravithejads’s Mistral cookbook showing ➡️ RAG ➡️ Query routing ➡️ Tool use  https://twitter.com/llama_index/status/1780646484712788085 

“More Mistral models and more details! 🚀 Mistral AI just uploaded the 8x22B model and a brand new instruct 8x22B with function calling support on Hugging Face, along with blog post and evaluation. 🔥 Confirmed and new Insights of 8x22B: 🌎 Fluent in 5 languages: English, French,…  https://twitter.com/_philschmid/status/1780598146470379880 

Reka

“Along with Core, we have published a technical report detailing the training, architecture, data, and evaluation for the Reka models.  https://twitter.com/RekaAILabs/status/1779894626083864873

“Meet Reka Core, our best and most capable multimodal language model yet. 🔮 It’s been a busy few months training this model and we are glad to finally ship it! 💪 Core has a lot of capabilities, and one of them is understanding video — let’s see what Core thinks of the 3 body…  https://twitter.com/RekaAILabs/status/1779894622334189592 

“We evaluate Core on standard benchmarks for both text and multimodal, along with a blind third-party human evaluation.  https://twitter.com/RekaAILabs/status/1779894623848304777 

Wizard

“🔥Today we are announcing WizardLM-2, our next generation state-of-the-art LLM. New family includes three cutting-edge models: WizardLM-2 8x22B, 70B, and 7B – demonstrates highly competitive performance compared to leading proprietary LLMs. 📙Release Blog:…  https://twitter.com/WizardLM_AI/status/1779899325868589372

“🧙‍♀️ WizardLM-2 8x22B is our most advanced model, and just slightly falling behind GPT-4-1106-preview. 🧙 WizardLM-2 70B reaches top-tier capabilities in the same size. 🧙‍♀️ WizardLM-2 7B even achieves comparable performance with existing 10x larger opensource leading models. The…  https://twitter.com/WizardLM_AI/status/1779899329844760771 

Other Open Source News

Open source groups fear xz repeat as suspicious messages fly • The Register – https://www.theregister.com/2024/04/16/xz_style_attacks_continue/

“🚀 Introducing Pile-T5! 🔗 We (EleutherAI) are thrilled to open-source our latest T5 model trained on 2T tokens from the Pile using the Llama tokenizer. ✨ Featuring intermediate checkpoints and a significant boost in benchmark performance. Work done by @lintangsutawika, me…  https://twitter.com/arankomatsuzaki/status/1779891910871490856 

Introducing Idefics2: A Powerful 8B Vision-Language Model for the community – https://huggingface.co/blog/idefics2 

“Introducing Idefics2, the strongest Vision-Language-Model (VLM) < 10B! 🚀 Idefics2 comes with significantly enhanced capabilities in OCR, document understanding, and visual reasoning. 💬📄🖼️ TL;DR; 📚 8B base and instruction variant 🖼️ Image + text inputs ⇒ Text output 📷…  https://twitter.com/_philschmid/status/1779922877589889400 

“⚔️ Closed-source vs. Open-weight LLMs The gap between the best-performing closed-source and open-source models is also narrowing in terms of Arena ELO ( https://twitter.com/maximelabonne/status/1779801605702836454 

“#1 This is for open-domain generalist chat, which is the task where prioprietary has the biggest structural advantage (more costly/bigger datasets) #2 This is only taking into account user preferences (open-source is faster/cheaper/safer). #3 IMO the gap between most of the…” / X – https://twitter.com/ClementDelangue/status/1779805019841142791

“@SnowflakeDB is open sourcing the best embedding models in the world! 🚀🚀 They are now available open source in @huggingface We are releasing it under the Apache 2 license so that it is easy for the OSS community to experiment with them freely.🎁🎁 These impressive models…  https://twitter.com/RamaswmySridhar/status/1780225794402627946

Heads up! You’ve scrolled to the end of this category. There may have been just one or two links (above), so go back up and double check to be sure you didn’t quickly scroll down past it.

Be Sure To Read This Week’s Main Post:

This week’s executive overview and top links are here:

AI News #28: Week Ending 04/12/2024 with Executive Summary and Top 48 Links

The post you just read is an deep dive extension of my weekly newsletter, This Week In AI, an executive summary of the top things to know in AI. Each week, I create an accessible overview for laypeople to feel confident they are conversant with the week’s AI developments. I include a curated list of must-click links of the week, to offer everyone a hands-on opportunity to explore the most intriguing updates in artificial intelligence across various categories, including robotics, imagery, video, AR/VR, science, ethics, and more. Beyond the overview, I post these topic-based deeper dives (below). If you haven’t read this week’s overview, I recommend starting there.

Credits/Sources

Most of these weekly links come from just a few prolific oversharing sources. Please follow them, as they work hard to find the news each week and they make it a lot easier for me to compile.

For previous issues, please visit the archives!

Thanks for reading!

One response to “Open Source AI News: Week Ending 04/19/2024”

  1. […] Open Source Models: An open source AI model refers to a class of artificial intelligence models with public source code. They can be inspected, copied, installed, and customized on private computers. In contrast, a closed source model is proprietary and owned by a company that you pay to use (like PowerPoint or Photoshop). One of the most famous open source language models is a French model called Mistral. Its code is completely publicly available, and anyone can download it and customize it. On one hand, open source is a transparent and powerful way to democratize AI, but on the other hand, open source models circumvent the guard rails and copyright protections that private companies implement. Open source models are the wild west of artificial intelligence, but also the potential saving grace (depending on who you ask). It’s a bit like gun control debates but for computing power.This weeks’s latest open source news: https://ethanbholland.com/2024/04/19/open-source-ai-news-week-ending-04-19-2024/ […]

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading