1 paper
Junzhou Chen, Jindong Wang, Gang Zhou
Vision-language models and vision-language action models endow the robot with unprecedented capabilities. However, the input of video and high-resolution images yields a massive nu…