benchmark dataset 1evidence grounding 1multimodal large language models 1sports video analysis 1temporal compositional reasoning 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
Siyu Cao, Lu Zhang, Ruizhe Zeng +1
The paper introduces SportsTime, a large benchmark of long-form sports videos with detailed temporal evidence annotations, and proposes the Chain-of-Time Reasoning (CoTR) framework…
cs.CV2026
DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation
Ruizhe Zeng, Siyu Cao, Lu Zhang +1
Reasoning segmentation aims to predict pixel-wise masks for targets given complex language queries. Existing approaches leverage Multimodal Large Language Models (MLLMs) for vision…
cs.GR2024
LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
Haozhou Pang, Tianwei Ding, Lanshan He +3
In this work, we present LLM Gesticulator, an LLM-based audio-driven co-speech gesture generation framework that synthesizes full-body animations that are rhythmically aligned with…