1 paper · 1 filter
Yilong Li, Shuai Zhang, Yijing Zeng +5
Large Multimodal Models (LMMs) are inherently modular, comprising vision and audio encoders, a projector, and a language backbone. Yet existing systems execute them monolithically,…