10 papers
APEX: Adaptive Policy Execution for Precise Manipulation
Mengfei Zhao, Chenxi Jiang, Tuo An +2
Modern imitation learning methods, including visuomotor and Vision-Language-Action (VLA) policies, typically output high-level action references that are executed by low-level cont…
MARS Policy: Multimodality Only When It Matters
Jindou Jia, Tuo An, Yuxuan Hu +7
Imitation learning has become a cornerstone for solving complex robotic manipulation tasks. In particular, multimodality, which enables robots to capture diverse yet valid behavior…
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
Geng Li, Guohao Chen, Ting Chen +6
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning…
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
Jingliang Li, Jindou Jia, Tuo An +7
When told to "cut the cake," a robot must choose the knife over nearby scissors, despite both objects affording the same cutting function. In real-world scenes, multiple objects ma…
Feedback World Model Enables Precise Guidance of Diffusion Policy
Tuo An, Jindou Jia, Gen Li +8
World models aim to improve robotic decision making by predicting the consequences of actions. However, in practice, their predictions often become unreliable once the robot encoun…
FLASH: Efficient Visuomotor Policy via Sparse Sampling
Jiaqi Bai, Jindou Jia, Yuxuan Hu +5
Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference…