1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste +2
Recent work shows that pre-trained, instruction-tuned vision-language models (VLMs) perform well at mapping from instructions and observations to high-level plans, but struggle to…