3 papers
cs.CV2026
Claim-Level Rubric Rewards for Video Caption Reinforcement Learning
Mingqi Gao, Hongyuan Dong, Yifei Chen +6
In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense vi…
cs.CV2024
Benchmarking and Improving Detail Image Caption
Hongyuan Dong, Jiawen Li, Bohong Wu +3
Image captioning has long been regarded as a fundamental task in visual understanding. Recently, however, few large vision-language model (LVLM) research discusses model's image ca…
cs.CV2024
Unveiling the Tapestry of Consistency in Large Vision-Language Models
Yuan Zhang, Fei Xiao, Tao Huang +7
Large vision-language models (LVLMs) have recently achieved rapid progress, exhibiting great perception and reasoning abilities concerning visual information. However, when faced w…