1 paper · 1 filter
Jianfei Zhao, Feng Zhang, Xin Sun +3
Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to text tokens. Visual tokens, th…