14 citations · 37 across the 11 of their papers we have counts for
1 paper · 1 filter
Liang Chen, Hongcheng Gao, Tianyu Liu +5
Vision-Language Models (VLMs) excel in many direct multimodal tasks but struggle to translate this prowess into effective decision-making within interactive, visually rich environm…