collaborators

6 papers

cs.CV2026

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

Yifan Xu, Zihao Wang, Zhixiao Wang +6

Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occ…

cs.CV2026

LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization

Zhenpeng Huang, Jiaqi Li, Zihan Jia +6

We present LongVPO, a novel two-stage Direct Preference Optimization framework that enables short-context vision-language models to robustly understand ultra-long videos without an…

cs.CV2025

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Xinhao Li, Ziang Yan, Desen Meng +7

Reinforcement Learning (RL) benefits Large Language Models (LLMs) for complex reasoning. Inspired by this, we explore integrating spatio-temporal specific rewards into Multimodal L…

cs.CV2025

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

Jun Zhang, Desen Meng, Zhengming Zhang +3

Despite the remarkable performance of multimodal large language models (MLLMs) across diverse tasks, the substantial training and inference costs impede their advancement. In this…

cs.CV2025

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Desen Meng, Rui Huang, Zhilin Dai +8

While recent advances in reinforcement learning have significantly enhanced reasoning capabilities in large language models (LLMs), these techniques remain underexplored in multi-m…

cs.CV2025

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

Yifan Xu, Xinhao Li, Yichun Yang +3

Video understanding, including video captioning and retrieval, is still a great challenge for video-language models (VLMs). The existing video retrieval and caption benchmarks only…