1 paper
Mingxiao Li, Fang Qu, Zhanpeng Chen +5
While MLLMs perform well on perceptual tasks, they lack precise multimodal alignment, limiting performance. To address this challenge, we propose Vision Dynamic Embedding-Guided Pr…