8 papers
EC: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control
Qiao Gu, Lingni Ma, Adam W Harley +3
Controllable and physically grounded egocentric video generation is essential for embodied agents to reason about how their own and others' actions manifest and change the world. C…
SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online Transfer
Nathan Samuel de Lara, Florian Shkurti
Modern offline Reinforcement Learning (RL) methods find performant actor-critics, however, fine-tuning these actor-critics online with value-based RL algorithms typically causes im…
The Impact of Class Uncertainty Propagation in Perception-Based Motion Planning
Jibran Iqbal Shah, Andrei Ivanovic, Kelly Zhu +4
Autonomous vehicles (AVs) are being increasingly deployed in urban environments. In order to operate safely and reliably, AVs need to account for the inherent uncertainty associate…
What Do You Need for Compositional Generalization in Diffusion Planning?
Quentin Clark, Florian Shkurti
In policy learning, stitching and compositional generalization refer to the extent to which the policy is able to piece together sub-trajectories of data it is trained on to genera…
Scalable Policy Evaluation with Video World Models
Wei-Cheng Tseng, Jinwei Gu, Qinsheng Zhang +4
Training generalist policies for robotic manipulation has shown great promise, as they enable language-conditioned, multi-task behaviors across diverse scenarios. However, evaluati…
SAFE: Multitask Failure Detection for Vision-Language-Action Models
Qiao Gu, Yuanliang Ju, Shengxiang Sun +4
While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on no…