8 papers
FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning
Chuhao Zhou, Liquan Wang, Shuxin Cao +5
While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same function to novel ones -- a gap we formalize a…
Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning
Yanzhe Tang, Xinyu Shao, Yuxuan Hu +6
While end-to-end Vision-Language-Action (VLA) models show promise in robotic manipulation, their monolithic paradigm inherently couples semantic reasoning and spatial control. This…
Rethinking Implicit Spatial Representation in Visuomotor Policy Learning
Xiangyu Chen, Yuxuan Hu, Chuhao Zhou +1
Generative model-based imitation learning has become a widely adopted paradigm for robotic manipulation, where policy performance depends critically on the conditioned visual repre…
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
Yikai Tang, Haoran Geng, Jindou Jia +5
Imitation learning has emerged as a crucial approach for acquiring visuomotor skills from demonstrations, where designing effective observation encoders is essential for policy gen…
MARS Policy: Multimodality Only When It Matters
Jindou Jia, Tuo An, Yuxuan Hu +7
Imitation learning has become a cornerstone for solving complex robotic manipulation tasks. In particular, multimodality, which enables robots to capture diverse yet valid behavior…
FLASH: Efficient Visuomotor Policy via Sparse Sampling
Jiaqi Bai, Jindou Jia, Yuxuan Hu +5
Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference…