1 paper
Gursimran Singh, Xinglu Wang, Yifan Hu +9
Large Multimodal Models (LMMs) extend Large Language Models (LLMs) by handling diverse inputs such as images, audio, and video, but at the cost of adding a multimodal encoding stag…