9 papers
USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning
Yuchen Xie, Xinyu Zhou, Kuangji Zuo +4
Embodied Visual Tracking (EVT) requires an agent to continuously follow a specified target while actively moving through dynamic environments. However, prevailing EVT paradigms pre…
ADAPT: Analytical Disturbance-Aware Policy Training for Humanoid Locomotion
Bofan Lyu, Jindou Jia, Kuangji Zuo +7
Humanoids deployed in human-centered environments must handle force-interactive tasks, where external contacts introduce unexpected disturbances that disrupt locomotion accuracy an…
EM-Fall: Embodied mmWave Sensing for Day-and-Night Fall Detection on Humanoid Robots
Yanshuo Lu, Yuxuan Hu, Shenghai Yuan +5
Falls are one of the leading causes of injury and hospitalization among elderly individuals, making reliable fall awareness an essential capability for safety monitoring in residen…
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
Kuangji Zuo, Gen Li, Bofan Lyu +9
Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insuff…
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
Geng Li, Guohao Chen, Ting Chen +6
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning…
FLASH: Efficient Visuomotor Policy via Sparse Sampling
Jiaqi Bai, Jindou Jia, Yuxuan Hu +5
Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference…