8 papers
Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs
Han Yin, Jung-Woo Choi
Recently, Large Audio Language Models (LALMs) have progressed rapidly, demonstrating their strong efficacy in universal audio understanding through cross-modal integration. To eval…
Neural acoustic multipole splatting for room impulse response synthesis
Geonwoo Baek, Jung-Woo Choi
Room Impulse Response (RIR) prediction at arbitrary receiver positions is essential for practical applications such as spatial audio rendering. We propose Neural Acoustic Multipole…
DISPATCH: Distilling Selective Patches for Speech Enhancement
Dohwan Kim, Jung-Woo Choi
In speech enhancement, knowledge distillation (KD) compresses models by transferring a high-capacity teacher's knowledge to a compact student. However, conventional KD methods trai…
DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis
Dongheon Lee, Younghoo Kwon, Jung-Woo Choi
We propose DeepASA, a multi-purpose model for auditory scene analysis that performs multi-input multi-output (MIMO) source separation, dereverberation, sound event detection (SED),…
Sound Separation and Classification with Object and Semantic Guidance
Younghoo Kwon, Jung-Woo Choi
The spatial semantic segmentation task focuses on separating and classifying sound objects from multichannel signals. To achieve two different goals, conventional methods fine-tune…
Self-Guided Target Sound Extraction and Classification Through Universal Sound Separation Model and Multiple Clues
Younghoo Kwon, Dongheon Lee, Dohwan Kim +1
This paper introduces a multi-stage self-directed framework designed to address the spatial semantic segmentation of sound scene (S5) task in the DCASE 2025 Task 4 challenge. This…