4 papers
Data Pyramid for Embodied Manipulation: A Survey
Yifan Ye, Yankai Fu, Yaoxu Lv +26
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations w…
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
Kailin Lyu, Di Wu, Pengwei Zhang +12
Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense re…
UniTacVLA: Unified Tactile Understanding and Prediction in Vision Language Action Models
Xidong Zhang, Yichi Zhang, Jiaxin Shi +5
Vision-language-action (VLA) models have achieved strong performance in many robotic manipulation tasks, yet remain limited in contact-rich dexterous manipulation. To overcome this…
Touch-R1: Reinforcing Touch Reasoning in MLLMs
Yingxin Lai, Yafei Zhou, Fucai Zhu +2
While rule-based reinforcement learning has recently catalyzed explicit reasoning in multimodal models, tactile reasoning remains largely underexplored. Existing tactile-language m…