2 papers
cs.CV2025
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
Zongxia Li, Xiyang Wu, Guangyao Shi +6
Vision-Language Models (VLMs) have achieved strong results in video understanding, yet a key question remains: do they truly comprehend visual content or only learn shallow correla…
cs.CV2024
HawkI: Homography & Mutual Information Guidance for 3D-free Single Image to Aerial View
Divya Kothandaraman, Tianyi Zhou, Ming Lin +1
We present HawkI, for synthesizing aerial-view images from text and an exemplar image, without any additional multi-view or 3D information for finetuning or at inference. HawkI use…