From the 1 of 13 linked papers with an AI index.
13 papers
WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos
Jiahao Liu, Zhongpu Xia, Shuai Tian +13
WALA is a framework that learns executable latent actions for robot manipulation by pretraining on both action‑labeled demonstrations and unlabeled videos, predicting future change…
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
Kailin Lyu, Di Wu, Pengwei Zhang +12
Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense re…
VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation
Shuai Tian, Yupeng Zheng, Yuhang Zheng +7
Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observat…
TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation
Yujie Zang, Yuhang Zheng, Xian Nie +7
Contact-rich manipulation requires robots to continuously perceive and regulate evolving physical interactions under dynamic contact transitions or complex surface geometries. Rece…
Dynamic Resilient Spatio-Semantic Memory with Hybrid Localization for Mobile Manipulation
Zhijie Yan, Shufei Li, Ze Zhang +3
Reliable mobile manipulation in dynamic indoor environments requires a scene representation that remains geometrically consistent, semantically queryable, and computationally bound…
Learning High-Frequency Continuous Action Chunks in Latent Space
Kunyun Wang, Yuhang Zheng, Yupeng Zheng +2
Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal consistency at moderate action…