6 papers
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
Jiayi Luo, Qiyan Liu, Tengyang Wang +8
Autoregressive (AR) video generation has emerged as a promising paradigm for long-horizon video synthesis, where each frame is generated conditioned on previously generated tokens.…
AutoTraces: Autoregressive Trajectory Forecasting via Multimodal Large Language Models
Teng Wang, Yanting Lu, Ruize Wang
We present AutoTraces, an autoregressive vision-language-trajectory model for robot trajectory forecasting in humam-populated environments, which harnesses the inherent reasoning c…
AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videos
Teng Yan, Yihan Liu, Jiongxu Chen +3
Long-term language-guided referring in fixed-view videos is challenging: the referent may be occluded or leave the scene for long intervals and later re-enter, while framewise refe…
CVBench: Benchmarking Cross-Video Synergies for Complex Multimodal Reasoning
Nannan Zhu, Yonghao Dong, Teng Wang +9
While multimodal large language models (MLLMs) exhibit strong performance on single-video tasks (e.g., video question answering), their capability for spatiotemporal pattern reason…
Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
Dang Jisheng, Wu Xudong, Wang Bimei +7
Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dyna…
Reinforcing Video Reasoning with Focused Thinking
Jisheng Dang, Jingze Wu, Teng Wang +6
Recent advancements in reinforcement learning, particularly through Group Relative Policy Optimization (GRPO), have significantly improved multimodal large language models for comp…