1 citations · 1 across the 6 of their papers we have counts for
1 paper · 2 filters
David Venuto, Sami Nur Islam, Martin Klissarov +3
Pre-trained Vision-Language Models (VLMs) are able to understand visual concepts, describe and decompose complex tasks into sub-tasks, and provide feedback on task completion. In t…