From the 1 of 4 linked papers with an AI index.
4 papers
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
Siyu Cao, Lu Zhang, Ruizhe Zeng +1
The paper introduces SportsTime, a large benchmark of long-form sports videos with detailed temporal evidence annotations, and proposes the Chain-of-Time Reasoning (CoTR) framework…
DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation
Ruizhe Zeng, Siyu Cao, Lu Zhang +1
Reasoning segmentation aims to predict pixel-wise masks for targets given complex language queries. Existing approaches leverage Multimodal Large Language Models (MLLMs) for vision…
HunyuanImage 3.0 Technical Report
Tencent Hunyuan Foundation Model Team
We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module pub…
Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval
Yiming Ding, Siyu Cao, Luyuan Jiao +4
Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each…