1 paper · 1 filter
Jihoo Jung, Chaeyoung Jung, Ji-Hoon Kim +1
Audio-visual large language models (AVLLMs) have recently emerged as a powerful architecture capable of jointly reasoning over audio, visual, and textual modalities. In AVLLMs, the…