1 paper
Jiayi Zhou, Jiaming Ji, Boyuan Chen +6
Training multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low acc…