5 citations · 10 across the 8 of their papers we have counts for
1 paper · 1 filter
Kai Zhang, Jianwei Yang, Jeevana Priya Inala +4
Despite the promising results of large multimodal models (LMMs) in complex vision-language tasks that require knowledge, reasoning, and perception abilities together, we surprising…