1 paper · 1 filter
Angelos Mavrogiannis, Dehao Yuan, Yiannis Aloimonos
There has been a lot of interest in grounding natural language to physical entities through visual context. While Vision Language Models (VLMs) can ground linguistic instructions t…