activity
20242026
collaborators

14 papers

cs.RO2026

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

Xiao Liu, Yuguang Yang, Xi Wang +6

Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the fut…

cs.RO2026

VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation

Shuai Tian, Yupeng Zheng, Yuhang Zheng +7

Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observat…

cs.LG2026

StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation

Guangda Liu, Yiquan Wang, Chengwei Li +6

Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is widely used in knowledge distillation, mod…

cs.RO2026

TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation

Yujie Zang, Yuhang Zheng, Xian Nie +7

Contact-rich manipulation requires robots to continuously perceive and regulate evolving physical interactions under dynamic contact transitions or complex surface geometries. Rece…

cs.RO2026

Learning High-Frequency Continuous Action Chunks in Latent Space

Kunyun Wang, Yuhang Zheng, Yupeng Zheng +2

Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal consistency at moderate action…

cs.RO2026

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance

Yupeng Zheng, Xiang Li, Songen Gu +12

Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency and a lack of high-level know…