5 papers
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
Xiaodong Wang, Langling Huang, Zhirong Wu +4
The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interacti…
COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
Canming Xia, Peixi Peng, Guang Tan +4
Visual reinforcement learning (RL) suffers from poor sample efficiency due to high-dimensional observations in complex tasks. While existing works have shown that vision-language m…
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
Xiaodong Wang, Jinfa Huang, Li Yuan +1
Most Video Large Language Models (Video-LLMs) adopt preference alignment techniques, e.g., DPO~\citep{rafailov2024dpo}, to optimize the reward margin between a winning response ($y…
LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model
Xiaodong Wang, Zhirong Wu, Peixi Peng
Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error ac…
ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos
Xiaodong Wang, Peixi Peng
Real-world driving requires people to observe the current environment, anticipate the future, and make appropriate driving decisions. This requirement is aligned well with the capa…