4 papers
OmniKVQuant: KV Cache Quantization for Omni-LLMs
Suho Yoo, Hyunjong Ok, Jongmin Choi +2
As Omni-modal large language models (Omni-LLMs) take in audio, video and text together, their KV cache memory cost grows. KV cache quantization is the de facto approach in text-onl…
Tracing Audio Grounding and Answer Selection in Audio LLMs
Hyebin Cho, Suho Yoo, Jihoo Jung +1
Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the answer by reasoning from textual cues or linguistic priors rather than…
Probing Cross-modal Information Hubs in Audio-Visual LLMs
Jihoo Jung, Chaeyoung Jung, Ji-Hoon Kim +1
Audio-visual large language models (AVLLMs) have recently emerged as a powerful architecture capable of jointly reasoning over audio, visual, and textual modalities. In AVLLMs, the…
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
Jihoo Jung, Ji-Hoon Kim, Doyeop Kwak +3
We introduce UNMIXX, a novel framework for multiple singing voices separation (MSVS). While related to speech separation, MSVS faces unique challenges: data scarcity and the highly…