3 papers
cs.CV2026
Claim-Level Rubric Rewards for Video Caption Reinforcement Learning
Mingqi Gao, Hongyuan Dong, Yifei Chen +6
In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense vi…
cs.CV2024
Unveiling the Tapestry of Consistency in Large Vision-Language Models
Yuan Zhang, Fei Xiao, Tao Huang +7
Large vision-language models (LVLMs) have recently achieved rapid progress, exhibiting great perception and reasoning abilities concerning visual information. However, when faced w…
cs.CV2024
Benchmarking and Improving Detail Image Caption
Hongyuan Dong, Jiawen Li, Bohong Wu +3
Image captioning has long been regarded as a fundamental task in visual understanding. Recently, however, few large vision-language model (LVLM) research discusses model's image ca…