activity
20242026
collaborators

6 papers

cs.CV2026

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

Burak Satar, Zhixin Ma, Cheng Yu-Tong +3

Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this int…

cs.CL2026

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

Zhixin Ma, Yutong Zhou, Yongqi Li +2

Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the "brain" of embodied AI, instructing robots to interact with the physical worl…

cs.CV2025

Seeing Culture: A Benchmark for Visual Reasoning and Grounding

Burak Satar, Zhixin Ma, Patrick A. Irawan +4

Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultur…

cs.MM2025

Multimodal LLM-based Query Paraphrasing for Video Search

Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan +3

Text-to-video retrieval answers user queries through searches based on concepts and embeddings. However, due to limitations in the size of the concept bank and the amount of traini…

cs.IR2025

Robust Relevance Feedback for Interactive Known-Item Video Search

Zhixin Ma, Chong-Wah Ngo

Known-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to inf…

cs.IR2024

PolySmart and VIREO @ TRECVid 2024 Ad-hoc Video Search

Jiaxin Wu, Chong-Wah Ngo, Xiao-Yong Wei +1

This year, we explore generation-augmented retrieval for the TRECVid AVS task. Specifically, the understanding of textual query is enhanced by three generations, including Text2Tex…