1 paper
Jiaxin Fan, Wenpo Song
Large Multimodal Models (LMMs) have achieved strong performance in vision-language understanding, yet many existing approaches rely on large-scale architectures and coarse supervis…