3 papers
cs.RO2025
FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
Guangyan Chen, Meiling Wang, Te Cui +8
Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in foundation models, particularly Vis…
cs.RO2025
STEP Planner: Constructing cross-hierarchical subgoal tree as an embodied long-horizon task planner
Tianxing Zhou, Zhirui Wang, Haojia Ao +5
The ability to perform reliable long-horizon task planning is crucial for deploying robots in real-world environments. However, directly employing Large Language Models (LLMs) as a…
cs.RO2024
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
Guanyan Chen, Meiling Wang, Te Cui +9
Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have…