3 papers
cs.CV2026
Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation
Yifan Li, Xinyu Zhou, Yunhao Ge +1
Robotic manipulation is often specified through language instructions or task identifiers, yet cluttered environments with similar objects are better handled by spatially indicatin…
cs.RO2025
Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation
Teqiang Zou, Hongliang Zeng, Yuxuan Nong +6
Most Vision-Language-Action (VLA) systems integrate a Vision-Language Model (VLM) for semantic reasoning with an action expert generating continuous action signals, yet both typica…
cs.CV2025
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
Yifan Li, Xin Li, Tianqin Li +3
Vision foundation models (VFMs) have demonstrated remarkable performance across a wide range of downstream tasks. While several VFM adapters have shown promising results by leverag…