1 paper
Zicheng Kong, Dehua Ma, Zhenbo Xu +9
Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly…