1 paper
Matthieu Futeral, Armel Zebaze, Pedro Ortiz Suarez +5
Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. (2022) showed that…