1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Wentao Tan, Bowen Wang, Heng Zhi +15
Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs ge…