3 papers
cs.CV2026
Controllable Video Object Insertion via Multi-View Priors
Qi Xia, Xia Qi, Peishan Cong +4
Video object insertion places a user-specified object in an existing dynamic scene. Existing methods typically condition generation on text or a single reference image. Consequentl…
cs.CV2025
STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation
Jiamin Wang, Yichen Yao, Xiang Feng +5
The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approac…
cs.RO2024
RealDex: Towards Human-like Grasping for Robotic Dexterous Hand
Yumeng Liu, Yaxun Yang, Youzhuo Wang +9
In this paper, we introduce RealDex, a pioneering dataset capturing authentic dexterous hand grasping motions infused with human behavioral patterns, enriched by multi-view and mul…