2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2025
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
Burak Satar, Zhixin Ma, Patrick A. Irawan +4
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultur…
cs.IR2025★ 2 cited
Robust Relevance Feedback for Interactive Known-Item Video Search
Zhixin Ma, Chong-Wah Ngo
Known-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to inf…
cs.IR2024
PolySmart and VIREO @ TRECVid 2024 Ad-hoc Video Search
Jiaxin Wu, Chong-Wah Ngo, Xiao-Yong Wei +1
This year, we explore generation-augmented retrieval for the TRECVid AVS task. Specifically, the understanding of textual query is enhanced by three generations, including Text2Tex…