1 paper · 2 filters
Seunghee Kim, Ingyu Bang, Seokgyu Jang +5
Multimodal Large Language Models (MLLMs) have increasingly supported omni-modal processing across text, vision, and speech. However, existing evaluation frameworks for such models…