activity
20242026
collaborators

8 papers

cs.SD2026

Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs

Han Yin, Jung-Woo Choi

Recently, Large Audio Language Models (LALMs) have progressed rapidly, demonstrating their strong efficacy in universal audio understanding through cross-modal integration. To eval…

eess.AS2026

Neural acoustic multipole splatting for room impulse response synthesis

Geonwoo Baek, Jung-Woo Choi

Room Impulse Response (RIR) prediction at arbitrary receiver positions is essential for practical applications such as spatial audio rendering. We propose Neural Acoustic Multipole…

cs.SD2026

DISPATCH: Distilling Selective Patches for Speech Enhancement

Dohwan Kim, Jung-Woo Choi

In speech enhancement, knowledge distillation (KD) compresses models by transferring a high-capacity teacher's knowledge to a compact student. However, conventional KD methods trai…

eess.AS2025

DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis

Dongheon Lee, Younghoo Kwon, Jung-Woo Choi

We propose DeepASA, a multi-purpose model for auditory scene analysis that performs multi-input multi-output (MIMO) source separation, dereverberation, sound event detection (SED),…

eess.AS2025

Sound Separation and Classification with Object and Semantic Guidance

Younghoo Kwon, Jung-Woo Choi

The spatial semantic segmentation task focuses on separating and classifying sound objects from multichannel signals. To achieve two different goals, conventional methods fine-tune…

eess.AS2025

Self-Guided Target Sound Extraction and Classification Through Universal Sound Separation Model and Multiple Clues

Younghoo Kwon, Dongheon Lee, Dohwan Kim +1

This paper introduces a multi-stage self-directed framework designed to address the spatial semantic segmentation of sound scene (S5) task in the DCASE 2025 Task 4 challenge. This…