1 paper
Vatsal Agarwal, Matthew Gwilliam, Gefen Kohavi +3
Recent advances in multimodal large language models (MLLMs) have enabled image-based question-answering capabilities. However, a key limitation is the use of CLIP as the visual enc…