1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Wenxuan Zhang, Yuhui Wang, Donggang Jia +5
Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end t…