1 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Haobin Jiang, Zongqing Lu
Generalization is a pivotal challenge for agents following natural language instructions. To approach this goal, we leverage a vision-language model (VLM) for visual grounding and…