3 papers
cs.CL2026
GAVEL: Grounded Caption Error Verification and Localization
Zixian Gao, Atsushi Hashimoto, Kuniaki Saito
Vision-language models (VLMs) often produce hallucinated or inconsistent outputs, where text and images are not properly aligned. Addressing this issue requires not only detecting…
cs.CV2026
Action Images: End-to-End Policy Learning via Multiview Video Generation
Haoyu Zhen, Zixian Gao, Qiao Sun +7
World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model the future states. However, exist…
cs.CV2025
Mitigating Object Hallucination via Robust Local Perception Search
Zixian Gao, Chao Yang, Zhanhui Zhou +2
Recent advancements in Multimodal Large Language Models (MLLMs) have enabled them to effectively integrate vision and language, addressing a variety of downstream tasks. However, d…