3 papers
cs.CV2026
Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia
Burak Satar, Zhixin Ma, Cheng Yu-Tong +3
Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this int…
cs.CV2025
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
Burak Satar, Zhixin Ma, Patrick A. Irawan +4
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultur…
cs.CV2024
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
Muhammet Furkan Ilaslan, Ali Koksal, Kevin Qinhong Lin +3
Large Language Model (LLM)-based agents have shown promise in procedural tasks, but the potential of multimodal instructions augmented by texts and videos to assist users remains u…