activity
20242026
collaborators
Showing cs.ROShow all

5 papers · 1 filter

cs.RO2026

ActionMap: Robot Policy Learning via Voxel Action Heatmap

Pei Yang, Hai Ci, Yanzhe Chen +3

Vision-language-action (VLA) models have advanced rapidly across backbones, training recipes, and data scale, yet the action decoder, which converts the backbone's hidden state int…

cs.RO2026

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation

Yanzhe Chen, Kevin Yuchen Ma, Qi Lv +4

While Vision-Language-Action (VLA) models offer broad general capabilities, deploying them on specific hardware requires real-world adaptation to bridge the embodiment gap. Since r…

cs.RO2025

Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Haoming Song, Delin Qu, Yuanqi Yao +9

Humans practice slow thinking before performing actual actions when handling complex tasks in the physical world. This thinking paradigm, recently, has achieved remarkable advancem…

cs.RO2025

Few-Shot Vision-Language Action-Incremental Policy Learning

Mingchen Song, Xiang Deng, Guoqiang Zhong +5

Recently, Transformer-based robotic manipulation methods utilize multi-view spatial representations and language instructions to learn robot motion trajectories by leveraging numer…

cs.RO2024

EPD: Long-term Memory Extraction, Context-awared Planning and Multi-iteration Decision @ EgoPlan Challenge ICML 2024

Letian Shi, Qi Lv, Xiang Deng +1

In this technical report, we present our solution for the EgoPlan Challenge in ICML 2024. To address the real-world egocentric task planning problem, we introduce a novel planning…