2 papers
cs.CV2026
Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs
Ziling Huang, Shin'ichi Satoh
Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capt…
cs.CV2025
ReSeDis: A Dataset for Referring-based Object Search across Large-Scale Image Collections
Ziling Huang, Yidan Zhang, Shin'ichi Satoh
Large-scale visual search engines are expected to solve a dual problem at once: (i) locate every image that truly contains the object described by a sentence and (ii) identify the…