1 paper
Yin Xie, Kaicheng Yang, Peirou Liang +7
Large Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to…