1 paper
Xiangxiang Zhang, Jingxuan Wei, Donghong Zhong +31
Existing Vision-Language Models often struggle with complex, multi-question reasoning tasks where partial correctness is crucial for effective learning. Traditional reward mechanis…