1 paper
Fatemeh Pesaran Zadeh, Yoojin Oh, Gunhee Kim
Aligning large VLMs with human preferences is a challenging task, as methods like RLHF and DPO often overfit to textual information or exacerbate hallucinations. Although augmentin…