3 papers
cs.CV2025
CVBench: Benchmarking Cross-Video Synergies for Complex Multimodal Reasoning
Nannan Zhu, Yonghao Dong, Teng Wang +9
While multimodal large language models (MLLMs) exhibit strong performance on single-video tasks (e.g., video question answering), their capability for spatiotemporal pattern reason…
cs.CV2025
Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
Dang Jisheng, Wu Xudong, Wang Bimei +7
Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dyna…
cs.CV2025
Reinforcing Video Reasoning with Focused Thinking
Jisheng Dang, Jingze Wu, Teng Wang +6
Recent advancements in reinforcement learning, particularly through Group Relative Policy Optimization (GRPO), have significantly improved multimodal large language models for comp…