Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia
Burak Satar, Zhixin Ma, Cheng Yu-Tong +3
Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this int…
cs.CV2025
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
Burak Satar, Zhixin Ma, Patrick A. Irawan +4
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultur…
cs.CV2024
Improving Interpretable Embeddings for Ad-hoc Video Search with Generative Captions and Multi-word Concept Bank
Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan
Aligning a user query and video clips in cross-modal latent space and that with semantic concepts are two mainstream approaches for ad-hoc video search (AVS). However, the effectiv…