collaborators

9 papers

cs.CV2026

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning

Yuchen Xie, Xinyu Zhou, Kuangji Zuo +4

Embodied Visual Tracking (EVT) requires an agent to continuously follow a specified target while actively moving through dynamic environments. However, prevailing EVT paradigms pre…

cs.RO2026

ADAPT: Analytical Disturbance-Aware Policy Training for Humanoid Locomotion

Bofan Lyu, Jindou Jia, Kuangji Zuo +7

Humanoids deployed in human-centered environments must handle force-interactive tasks, where external contacts introduce unexpected disturbances that disrupt locomotion accuracy an…

cs.RO2026

EM-Fall: Embodied mmWave Sensing for Day-and-Night Fall Detection on Humanoid Robots

Yanshuo Lu, Yuxuan Hu, Shenghai Yuan +5

Falls are one of the leading causes of injury and hospitalization among elderly individuals, making reliable fall awareness an essential capability for safety monitoring in residen…

cs.RO2026

Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation

Kuangji Zuo, Gen Li, Bofan Lyu +9

Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insuff…

cs.CV2026

OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning

Geng Li, Guohao Chen, Ting Chen +6

Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning…

cs.RO2026

FLASH: Efficient Visuomotor Policy via Sparse Sampling

Jiaqi Bai, Jindou Jia, Yuxuan Hu +5

Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference…