12 papers
A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks
Ishani Mondal, Aparna Garimella, Ananya Sai +2
Automatically generated videos from scientific papers are increasingly used for education and research dissemination. However, existing evaluation metrics mainly measure visual qua…
Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures
Ishani Mondal, Javad Baghirov, Jordan Boyd-Graber
Scientific figures compress complex pipelines into a single canvas, yet understanding them requires paper-grounded, step-by-step narration aligned with visual highlights a capabili…
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
Feng Gu, Zongxia Li, Carlos Rafael Colon +3
Event annotation is important for identifying market changes, monitoring breaking news, and understanding sociological trends. Although expert annotators set the gold standards, hu…
Self-Rewarding Vision-Language Model via Reasoning Decomposition
Zongxia Li, Wenhao Yu, Chengsong Huang +8
Vision-Language Models (VLMs) often suffer from visual hallucinations: generating things that are not consistent with visual inputs and language shortcuts, where they skip the visu…
CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding
Ishani Mondal, Yiwen Song, Mihir Parmar +4
Long-form visual storytelling requires maintaining continuity across shots, including consistent characters, stable environments, and smooth scene transitions. While existing gener…
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
Zongxia Li, Xiyang Wu, Guangyao Shi +6
Vision-Language Models (VLMs) have achieved strong results in video understanding, yet a key question remains: do they truly comprehend visual content or only learn shallow correla…