1 paper
Sara Ghazanfari, Alexandre Araujo, Prashanth Krishnamurthy +2
Multi-modal Large Language Models (MLLMs) have recently exhibited impressive general-purpose capabilities by leveraging vision foundation models to encode the core concepts of imag…