Image created with Flux Pro v1.1 Ultra. Image prompt: Assembly instruction diagram for a video camera gimbal stabilizer, three-axis assembly view, modern creator style, red and black color scheme, white studio background, “BYTEDANCE” in bold modern font, motor positions and balance points clearly marked
ByteDance’s Bagel 14B MOE (7B active) Multimodal with image generation (open source, apache license) is just an incredible modle. A unified multimodal model rivalling GPT-4o and Gemini 2.0, with 7B active params (14B total), 40K context, 88% GenEval and 85% understanding, https://x.com/rohanpaul_ai/status/1927705853580509607
ByteDance Seed introduces: Emerging Properties in Unified Multimodal Pretraining “”In this work, we introduce BAGEL, an open-source foundational model that natively supports multimodal understanding and generation. BAGEL is a unified, decoder-only model pretrained on trillions https://x.com/iScienceLuvr/status/1925162040534208758
A new recipe for training multimodal models 👉 Mixed together various data types: text next to images, video frames after captions, then webpages, etc. This way the model learns to connect what it reads with what it sees. ByteDance proposed and implemented this idea in their https://x.com/TheTuringPost/status/1927123359969468420




