activity
20212026
most citedPlanAgent: A Multi-modal Large Language Agent for Closed-loop Vehicle Motion Planning

5 citations · 6 across the 19 of their papers we have counts for

collaborators

20 papers

cs.RO2026

GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

Yupeng Zheng, Xiang Li, Songen Gu +11

Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction obje…

cs.RO2026

WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos

Jiahao Liu, Zhongpu Xia, Shuai Tian +13

Generalizable robot policies typically rely on action-labeled robot demonstrations, which are expensive to collect and difficult to scale. In contrast, large-scale human and robot…

cs.RO2026

VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation

Shuai Tian, Yupeng Zheng, Yuhang Zheng +7

Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observat…

cs.RO2026

TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation

Yujie Zang, Yuhang Zheng, Xian Nie +7

Contact-rich manipulation requires robots to continuously perceive and regulate evolving physical interactions under dynamic contact transitions or complex surface geometries. Rece…

cs.AI2026

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

Kailin Lyu, Di Wu, Pengwei Zhang +12

Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense re…

cs.RO2026

Dynamic Resilient Spatio-Semantic Memory with Hybrid Localization for Mobile Manipulation

Zhijie Yan, Shufei Li, Ze Zhang +3

Reliable mobile manipulation in dynamic indoor environments requires a scene representation that remains geometrically consistent, semantically queryable, and computationally bound…