Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs
Ziling Huang, Shin'ichi Satoh
Multimodal Large Language Models (MLLMs) have made strong progress in video understanding, yet long videos remain difficult: the visual token budget grows with video length, so tem…
cs.CV2026
Temporal Tree of Thought: Reasoning-Guided Visual Cue Search for Long-Video Understanding
Ziling Huang, Shin'ichi Satoh
Long-video understanding remains challenging for Multimodal Large Language Models (MLLMs) due to limited context length. Uniform sampling may miss crucial moments, while agent-base…
cs.CV2025
ReSeDis: A Dataset for Referring-based Object Search across Large-Scale Image Collections
Ziling Huang, Yidan Zhang, Shin'ichi Satoh
Large-scale visual search engines are expected to solve a dual problem at once: (i) locate every image that truly contains the object described by a sentence and (ii) identify the…