1 paper · 1 filter
Zirui Song, Guangxian Ouyang, Mingzhe Li +10
Large Vision-Language Models (LVLMs) have recently advanced robotic manipulation by leveraging vision for scene perception and language for instruction following. However, existing…