6 papers
Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia
Burak Satar, Zhixin Ma, Cheng Yu-Tong +3
Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this int…
Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use
Zhixin Ma, Yutong Zhou, Yongqi Li +2
Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the "brain" of embodied AI, instructing robots to interact with the physical worl…
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
Burak Satar, Zhixin Ma, Patrick A. Irawan +4
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultur…
Multimodal LLM-based Query Paraphrasing for Video Search
Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan +3
Text-to-video retrieval answers user queries through searches based on concepts and embeddings. However, due to limitations in the size of the concept bank and the amount of traini…
Robust Relevance Feedback for Interactive Known-Item Video Search
Zhixin Ma, Chong-Wah Ngo
Known-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to inf…
PolySmart and VIREO @ TRECVid 2024 Ad-hoc Video Search
Jiaxin Wu, Chong-Wah Ngo, Xiao-Yong Wei +1
This year, we explore generation-augmented retrieval for the TRECVid AVS task. Specifically, the understanding of textual query is enhanced by three generations, including Text2Tex…