9 papers
Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs
Chaeyoung Jung, Kyeongha Rho, Joon Son Chung
Omnimodal Large Language Models (Omni-LLMs) incur substantial computational overhead due to the large number of multimodal input tokens they process, making token reduction essenti…
Probing Cross-modal Information Hubs in Audio-Visual LLMs
Jihoo Jung, Chaeyoung Jung, Ji-Hoon Kim +1
Audio-visual large language models (AVLLMs) have recently emerged as a powerful architecture capable of jointly reasoning over audio, visual, and textual modalities. In AVLLMs, the…
FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference
Chaeyoung Jung, Youngjoon Jang, Seungwoo Lee +1
In this work, we present FastAV, the first token pruning framework tailored for audio-visual large language models (AV-LLMs). While token pruning has been actively explored in stan…
Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models
Chaeyoung Jung, Youngjoon Jang, Jongmin Choi +1
The goal of this work is to enhance balanced multimodal understanding in audio-visual large language models (AV-LLMs) by addressing modality bias without additional training. In cu…
AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
Chaeyoung Jung, Youngjoon Jang, Joon Son Chung
Hallucination remains a major challenge in multimodal large language models (MLLMs). To address this, various contrastive decoding (CD) methods have been proposed that contrasts or…
InfiniteAudio: Infinite-Length Audio Generation with Consistency
Chaeyoung Jung, Hojoon Ki, Ji-Hoon Kim +2
This paper presents InfiniteAudio, a simple yet effective strategy for generating infinite-length audio using diffusion-based text-to-audio methods. Current approaches face memory…