1 paper
Pengcheng Zheng, Chaoning Zhang, Jiarong Mo +8
Large multimodal models (LMMs) have achieved impressive performance on various vision-language tasks, but their substantial computational and memory costs hinder their practical de…