1 paper · 1 filter
Deqing Fu, Tong Xiao, Rui Wang +5
Although reward models have been successful in improving multimodal large language models, the reward models themselves remain brutal and contain minimal information. Notably, exis…