3 papers
cs.CV2026
DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation
Ruizhe Zeng, Siyu Cao, Lu Zhang +1
Reasoning segmentation aims to predict pixel-wise masks for targets given complex language queries. Existing approaches leverage Multimodal Large Language Models (MLLMs) for vision…
cs.CV2026
Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval
Yiming Ding, Siyu Cao, Luyuan Jiao +4
Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each…
cs.CV2026
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
Siyu Cao, Lu Zhang, Ruizhe Zeng +1
Sports videos are a challenging domain for multimodal understanding because they involve complex and dynamic human activities. Despite rapid progress in Multimodal Large Language M…