A huge modern humanoid robot with “Reflection” written on it looks into a mirror. Only a small robot mouse is reflected back in the mirror’s reflection.

Meta/Llama

“Direct modeling of JPEG and AVC bytes yields high-quality image and video generation using regular LLMs with Llama Architecture. So here basically, the file encoding approach (Direct modeling of JPEG and AVC bytes) outperforms vector quantization in image generation, especially 

“Llama 3 405B crosses 100 TPS barrier on Together APIs with a new inference engine release we rolled out this weekend powered by our new TKC! The continuous benchmark by @ArtificialAnlys shows 106.9 TPS. For context, we serve this model on @Nvidia H100 GPUs. API @ 

“Carry an intelligence, of sort, with you anywhere. (Also why it is impossible to restrict the use of GPT-4 class AIs. Anyone can download one, though Llama 405B is not really home PC compatible) 

“@emollick I like telling those people that I have a USB stick with Llama 3.1 70B – a very capable model – that runs directly on my laptop That’s not going away You can grab a copy from here, packaged with the software needed to run it on any modern OS 

Qwen

“Alibaba just unveiled Qwen2-VL, a new vision-language AI model that beat GPT-4o in some benchmarks It excels at problem-solving, math, and doc analysis The GPT-4o model they compared it to was the old model before updates, but still impressive 

Alibaba’s Qwen2-VL AI can analyze videos more than 20 min long | VentureBeat

Reflection

“Reflection 70b just dropped and is beating every other model, including GPT4o and Claude 3.5. How did this happen? Here’s my conversation with Matt Shumer (@mattshumer_) and Sahil Chaudhary (@csahil28), the authors of Reflection 70b. 

Even 4bit quants of Reflection 70b are amazing : r/LocalLLaMA

“Strong claims put by Reflection-70B creators https://t.co/473SKreMRI argue that techniques they used for fine-tuning the model (“reflection tuning”) make their model “world’s top open source model”. On common benchmarks, the scores seem to indicate model’s superiority. 2/n” / X

“@mattshumer_ @GlaiveAI Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!” / X

“I’m excited to announce Reflection 70B, the world’s top open-source model. Trained using Reflection-Tuning, a technique developed to enable LLMs to fix their own mistakes. 405B coming next week – we expect it to be the best model in the world. Built w/ @GlaiveAI. Read on ⬇️: 

HyperWrite debuts Reflection 70B, most powerful open source LLM | VentureBeat

“Major grifter energy. – 99.2% performance on GSM8k even though GSM8k has more than >1% error rate. -Doesn’t disclose that he’s an investor in Glaive. – Gives misleading details about RLHF (the model was finetuned from llama3.1-70b-Instruct).” / X

“Reflection 70B scored 42% on the aider code editing benchmark, well below Llama3 70B at 49%. I modified aider to ignore the <thinking/reflection> tags. This model won’t work properly with the released aider. 

Reflection 70B: Hype? : r/LocalLLaMA

“After verifying the required setup (with system prompt, no prefilling), I can safely say Reflection does not do well on BigCodeBench-Hard, at least. Complete: 20.3 (vs 28.4 from Llama3.1-70B) Instruct: 14.9 (vs 23.6 from Llama3.1-70B) The CoT/thinking/reflection process” / X

“I gave this classic riddle to Reflection Llama by @mattshumer_ and I’m so insanely impressed. Here’s what I got: Riddle used as input: A man walks into a shop and steals a $100 bill from the cash register. He then uses that $100 bill to buy $70 worth of goods from the same” / X

“@mattshumer_ @GlaiveAI Best 70B model I have seen by far, and on par with Claude Sonnet 3.5 for many queries I could try. Now waiting for the gguf! 

Snowflake

“There’s a lot of buzz about “founder mode” and how to scale an organization effectively. While I’m not the founder of Snowflake, I’ve been a founder before, and now my focus is on scaling a rapidly growing business. Here’s what “founder mode” means to me: When I became CEO of” / X

Other Open Source News

“OLMoE: Open Mixture-of-Experts Language Models abs: https://t.co/68jxJS6GLc model: https://t.co/h4tLvJRZqv “OLMOE-1B-7B has 7 billion (B) parameters but uses only 1B per input token. We pretrain it on 5 trillion tokens and further adapt it to create OLMOE-1B-7B-INSTRUCT. Our 

“Buildel: Opensource AI Automation Platform. Designed to empower users to create versatile and dynamic workflows tailored to their specific needs. 🔀 Multiple Providers – We support multiple providers for the same type of block. Use OpenAI, Google, Mistral and many more. 💻 

“Looks like a SOTA open source text-to-music model (a rectified flow dit) is out. Paper here: 

Show HN: AnythingLLM – Open-Source, All-in-One Desktop AI Assistant | Hacker News

Debate over “open source AI” term brings new push to formalize definition – Ars Technica

Show HN: Laminar – Open-Source DataDog + PostHog for LLM Apps, Built in Rust | Hacker News

“DeepSeek 2.5 is out! A powerful MOE with 238B params with 160 experts and 16B active params 👀Chat and code capabilities 🔥Function calling 💻JSON output and FIM completion 📏128k context length They also dropped DeepSeek-Coder-V2 Check it out! 

“🚀 RWKV.cpp AI system – is being deployed to 0.5 billion installs globally Making it one of the world’s most deployed, truly open-source (apache2) AI solutions out there As it ships with every Windows 11 system today!! Our group’s code install count, went from ~100k -> 0.5B 

“A sanity check on the best published tricks in MoE training, with likely the best dataset to date. DeepSeek-MoE scores 1 of 2: granularity yes, shared experts no. I think these results are very scale dependent. Also, switching to expert-wise bias for balancing might make it 2/2 

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading