2 citations · 3 across the 12 of their papers we have counts for
9 papers · 1 filter
Controllable Video Object Insertion via Multi-View Priors
Qi Xia, Peishan Cong, Yichen Yao +3
Video object insertion places a user-specified object in an existing dynamic scene. Existing methods typically condition generation on text or a single reference image. Consequentl…
Driving Like Yourself: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving
Xiaoru Dong, Ruiqin Li, Xiao Han +7
Human driving behavior is inherently diverse, yet most end-to-end autonomous driving (E2E-AD) systems learn a single average driving style, neglecting individual differences. Achie…
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
Yu Li, Yuenan Hou, Yingmei Wei +4
Multi-modal 3D understanding is a fundamental task in computer vision. Previous multi-modal fusion methods typically employ a single, dense fusion network, struggling to handle the…
STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation
Jiamin Wang, Yichen Yao, Xiang Feng +5
The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approac…
HUMOF: Human Motion Forecasting in Interactive Social Scenes
Caiyi Sun, Yujing Sun, Xiao Han +5
Complex scenes present significant challenges for predicting human behaviour due to the abundance of interaction information, such as human-human and humanenvironment interactions.…
FreeCap: Hybrid Calibration-Free Motion Capture in Open Environments
Aoru Xue, Yiming Ren, Zining Song +3
We propose a novel hybrid calibration-free method FreeCap to accurately capture global multi-person motions in open environments. Our system combines a single LiDAR with expandable…