collaborators

5 papers

cs.RO2026

Vid2WAM: Distilling Video Diffusion Priors into World Action Models

Chenhao Qiu, Ruixiang Wang, Runyi Zhao +7

World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by…

cs.RO2026

-WM: A Unified Video-Action World Model for Robotic Manipulation

Pengfei Zhou, Shengcong Chen, Di Chen +17

Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present -World…

cs.CV2026

VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority

Chenhao Qiu, Yechao Zhang, Xin Luo +2

Long video question answering requires locating sparse, time-scattered visual evidence within highly redundant content. Although current MLLMs perform well on short videos, long vi…

cs.RO2026

OFlow: Injecting Object-Aware Temporal Flow Matching for Robust Robotic Manipulation

Kuanning Wang, Ke Fan, Chenhao Qiu +5

Robust robotic manipulation requires not only predicting how the scene evolves over time, but also recognizing task-relevant objects in complex scenes. However, existing VLA models…

cs.CV2025

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

Peng Xu, Shengwu Xiong, Jiajun Zhang +125

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…