5 papers
LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
Youngjin Hong, Houjian Yu, Mingen Li +1
Learning generalizable policies for robotic manipulation increasingly relies on large-scale models that map language instructions to actions (L2A). However, this one-way paradigm o…
Hierarchical DLO Routing with Reinforcement Learning and In-Context Vision-language Models
Mingen Li, Houjian Yu, Yixuan Huang +3
Long-horizon routing tasks of deformable linear objects (DLOs), such as cables and ropes, are common in industrial assembly lines and everyday life. These tasks are particularly ch…
Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
Houjian Yu, Zheming Zhou, Min Sun +5
Enabling robots to grasp objects specified through natural language is essential for effective human-robot interaction, yet it remains a significant challenge. Existing approaches…
A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping
Houjian Yu, Mingen Li, Alireza Rezazadeh +2
The language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasp…
Attribute-Based Robotic Grasping with Data-Efficient Adaptation
Yang Yang, Houjian Yu, Xibai Lou +2
Robotic grasping is one of the most fundamental robotic manipulation tasks and has been the subject of extensive research. However, swiftly teaching a robot to grasp a novel target…