Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos
Siyao Yan, Bo Han, Jisheng Dang +7
Video multimodal large language models support language guided video segmentation, but they often show spatio temporal inconsistencies, e.g., jitter, drift, and identity switches.…
cs.AI2026
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
Jisheng Dang, Yizhou Zhang, Hao Ye +6
Fine-grained video captioning aims to generate detailed, temporally coherent descriptions of video content. However, existing methods struggle to capture subtle video dynamics and…