3 papers
cs.CV2026
EC: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control
Qiao Gu, Lingni Ma, Adam W Harley +3
Controllable and physically grounded egocentric video generation is essential for embodied agents to reason about how their own and others' actions manifest and change the world. C…
cs.RO2025
SAFE: Multitask Failure Detection for Vision-Language-Action Models
Qiao Gu, Yuanliang Ju, Shengxiang Sun +4
While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on no…
cs.RO2025
SICNav-Diffusion: Safe and Interactive Crowd Navigation with Diffusion Trajectory Predictions
Sepehr Samavi, Anthony Lem, Fumiaki Sato +5
To navigate crowds without collisions, robots must interact with humans by forecasting their future motion and reacting accordingly. While learning-based prediction models have sho…