4 papers · 1 filter
Mitigating Multimodal Hallucination via Phase-wise Self-reward
Yu Zhang, Chuyang Sun, Kehai Chen +3
Large Vision-Language Models (LVLMs) still struggle with vision hallucination, where generated responses are inconsistent with the visual input. Existing methods either rely on lar…
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
Zhiyu Yin, Zhipeng Liu, Kehai Chen +5
While current video generation focuses on text or image conditions, practical applications like video editing and vlogging often need to seamlessly connect separate clips. In our w…
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
Yingjie Zhu, Xuefeng Bai, Kehai Chen +6
Large Vision-Language Models (LVLMs) have achieved remarkable success across a wide range of multimodal tasks, yet their robustness to spatial variations remains insufficiently und…
A Survey: Spatiotemporal Consistency in Video Generation
Zhiyu Yin, Kehai Chen, Xuefeng Bai +7
Video generation aims to produce temporally coherent sequences of visual frames, representing a pivotal advancement in Artificial Intelligence Generated Content (AIGC). Compared to…