1 paper · 1 filter
Fatemeh Pesaran Zadeh, Yoojin Oh, Gunhee Kim
Aligning large VLMs with human preferences is a challenging task, as methods like RLHF and DPO often overfit to textual information or exacerbate hallucinations. Although augmentin…