1 paper
Hyeonjun An, Sihyun Kim, Chaerim Lim +9
Multimodal Large Language Models (MLLMs) have achieved remarkable advances by integrating text, image, and audio understanding within a unified architecture. However, existing dist…