collaborators

19 papers

cs.CV2026

PhysFlow: Physics-Aware Optical Flow for Motion Controllable Video Generation

Cong Wang, Hanxin Zhu, Yonglin Tian +5

Video generation models have recently attracted substantial attention for their ability to generate visually compelling videos, yet ensuring physically consistent and plausible dyn…

cs.AI2026

IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

Rongze Tang, Jianjie Fang, Zhaolu Wang +8

World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing appr…

cs.CV2026

Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis

Liang Xu, Chengqun Yang, Zili Lin +6

The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approache…

cs.RO2026

EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation

Jiayi Luo, Hanxin Zhu, Chen Gao +5

Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable p…

cs.CV2026

CoT-Edit: Let CoT Guide Instruction Video Editing

Sen Liang, Fengbin Guan, Youliang Zhang +2

Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relationships and physical constrain…

cs.CV2026

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing

Sen Liang, Cong Wang, Zhentao Yu +8

Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real-world scenarios. To bridge…