3 papers
cs.CV2026
Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning
Mingwen Zhang, Jisheng Dang, Minqiang Yang +3
Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. In many current training setups,…
cs.AI2026
PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos
Siyao Yan, Bo Han, Jisheng Dang +7
Video multimodal large language models support language guided video segmentation, but they often show spatio temporal inconsistencies, e.g., jitter, drift, and identity switches.…
cs.AI2025
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
Jisheng Dang, Yizhou Zhang, Hao Ye +6
Fine-grained video captioning aims to generate detailed, temporally coherent descriptions of video content. However, existing methods struggle to capture subtle video dynamics and…