2 papers
cs.AI2025
When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
Yifan Xu, Xichen Ye, Yifan Chen +1
Quality of datasets plays an important role in large language model (LLM) alignment. In collecting human feedback, however, preference flipping is ubiquitous and causes corruption…
cs.CV2025
Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
Chengzhi Yu, Yifan Xu, Yifan Chen +1
Recently, large vision-language models (LVLMs) have risen to be a promising approach for multimodal tasks. However, principled hallucination mitigation remains a critical challenge…