1 paper
Yichang Jian, Boyuan Xiao, Zhenyuan Huang +2
Planning from raw visual input remains a significant challenge for current Vision-Language Models (VLMs), when the complexity of input is beyond their one-step perception capabilit…