7 papers
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
Zihao Zheng, Sicheng Tian, Zhihao Mao +8
Vision-Language-Action (VLA) models have emerged as the mainstream of embodied intelligence. Recent VLA models have expanded their input modalities from 2D-only to 2D+3D paradigms,…
Adaptor: Advancing Assistive Teleoperation with Few-Shot Learning and Cross-Operator Generalization
Yu Liu, Yihang Yin, Tianlv Huang +8
Assistive teleoperation enhances efficiency via shared control, yet inter-operator variability, stemming from diverse habits and expertise, induces highly heterogeneous trajectory…
EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities
Travis Davies, Yiqi Huang, Alexi Gladstone +5
Implicit policies parameterized by generative models, such as Diffusion Policy, have become the standard for policy learning and Vision-Language-Action (VLA) models in robotics. Ho…
Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
Travis Davies, Yiqi Huang, Yunxin Liu +3
Scaling Transformer policies and diffusion models has advanced robotic manipulation, yet combining these techniques in lightweight, cross-embodiment learning settings remains chall…
Spatial RoboGrasp: Generalized Robotic Grasping Control Policy
Yiqi Huang, Travis Davies, Jiahuan Yan +3
Achieving generalizable and precise robotic manipulation across diverse environments remains a critical challenge, largely due to limitations in spatial perception. While prior imi…
CoinRobot: Generalized End-to-end Robotic Learning for Physical Intelligence
Yu Zhao, Huxian Liu, Xiang Chen +3
Physical intelligence holds immense promise for advancing embodied intelligence, enabling robots to acquire complex behaviors from demonstrations. However, achieving generalization…