1 paper · 1 filter
Kai Zhang, Jianwei Yang, Jeevana Priya Inala +4
Despite the promising results of large multimodal models (LMMs) in complex vision-language tasks that require knowledge, reasoning, and perception abilities together, we surprising…