1 paper
Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste +2
Recent work shows that pre-trained, instruction-tuned vision-language models (VLMs) perform well at mapping from instructions and observations to high-level plans, but struggle to…