14 papers
StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation
Xiao Liu, Yuguang Yang, Xi Wang +6
Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the fut…
VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation
Shuai Tian, Yupeng Zheng, Yuhang Zheng +7
Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observat…
StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation
Guangda Liu, Yiquan Wang, Chengwei Li +6
Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is widely used in knowledge distillation, mod…
TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation
Yujie Zang, Yuhang Zheng, Xian Nie +7
Contact-rich manipulation requires robots to continuously perceive and regulate evolving physical interactions under dynamic contact transitions or complex surface geometries. Rece…
Learning High-Frequency Continuous Action Chunks in Latent Space
Kunyun Wang, Yuhang Zheng, Yupeng Zheng +2
Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal consistency at moderate action…
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
Yupeng Zheng, Xiang Li, Songen Gu +12
Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency and a lack of high-level know…