1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Kai Zhang, Jianwei Yang, Jeevana Priya Inala +4
Despite the promising results of large multimodal models (LMMs) in complex vision-language tasks that require knowledge, reasoning, and perception abilities together, we surprising…