6 papers
HOST:Robots Acquire Manipulation Skills in Seconds from a Single Human Video
Guangyan Chen, Meiling Wang, Te Cui +9
The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-…
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
Guangyan Chen, Meiling Wang, Qi Shao +10
Developing robust and general-purpose manipulation policies represents a fundamental objective in robotics research. While Vision-Language-Action (VLA) models have demonstrated pro…
GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
Ye Chen, Zichen Zhou, Jianyu Dou +3
In recent years, visual representation learning has gained widespread attention in robotic imitation learning. However, in complex Out-of-Distribution(OOD) settings characterized b…
FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
Guangyan Chen, Meiling Wang, Te Cui +8
Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in foundation models, particularly Vis…
Human Demonstrations are Generalizable Knowledge for Robots
Te Cui, Tianxing Zhou, Zicai Peng +6
Learning from human demonstrations is an emerging trend for designing intelligent robotic systems. However, previous methods typically regard videos as instructions, simply dividin…
TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models
Meng Yu, Te Cui, Qitong Chu +3
Reliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rel…