Image created with Flux Pro v1.1 Ultra. Image prompt: Assembly instruction diagram for a video camera gimbal stabilizer, three-axis assembly view, modern creator style, red and black color scheme, white studio background, “BYTEDANCE” in bold modern font, motor positions and balance points clearly marked

ByteDance’s Bagel 14B MOE (7B active) Multimodal with image generation (open source, apache license) is just an incredible modle. A unified multimodal model rivalling GPT-4o and Gemini 2.0, with 7B active params (14B total), 40K context, 88% GenEval and 85% understanding, https://x.com/rohanpaul_ai/status/1927705853580509607

ByteDance Seed introduces: Emerging Properties in Unified Multimodal Pretraining “”In this work, we introduce BAGEL, an open-source foundational model that natively supports multimodal understanding and generation. BAGEL is a unified, decoder-only model pretrained on trillions https://x.com/iScienceLuvr/status/1925162040534208758

A new recipe for training multimodal models 👉 Mixed together various data types: text next to images, video frames after captions, then webpages, etc. This way the model learns to connect what it reads with what it sees. ByteDance proposed and implemented this idea in their https://x.com/TheTuringPost/status/1927123359969468420

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading