11 citations · 12 across the 3 of their papers we have counts for
3 papers
cs.CV2025
From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs
Mingxiao Li, Fang Qu, Zhanpeng Chen +5
While MLLMs perform well on perceptual tasks, they lack precise multimodal alignment, limiting performance. To address this challenge, we propose Vision Dynamic Embedding-Guided Pr…
cs.LG2024★ 1 cited
Revisiting MoE and Dense Speed-Accuracy Comparisons for LLM Training
Xianzhi Du, Tom Gunter, Xiang Kong +5
Mixture-of-Experts (MoE) enjoys performance gain by increasing model capacity while keeping computation cost constant. When comparing MoE to dense models, prior work typically adop…
cs.CV2024★ 11 cited
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier +29
In this work, we discuss building performant Multimodal Large Language Models (MLLMs). In particular, we study the importance of various architecture components and data choices. T…