1 paper
Mirco Bonomo, Simone Bianco
Multimodal Large Language Models (MLLMs) have achieved notable performance in computer vision tasks that require reasoning across visual and textual modalities, yet their capabilit…