1 paper · 1 filter
Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh +6
Interactive and embodied tasks pose at least two fundamental challenges to existing Vision & Language (VL) models, including 1) grounding language in trajectories of actions and ob…