3 papers
cs.SD2026
A Sensitivity Analysis of Multi-Event Audio Grounding in Audio LLMs
Taehan Lee, Jaehan Jung, Hyukjun Lee
Audio LLMs have shown a strong ability to understand audio samples, yet their reliability in complex acoustic scenes remains under-explored. Unlike prior work limited to small scal…
cs.SD2026
SAM: A Mamba-2 State-Space Audio-Language Model
Taehan Lee, Jaehan Jung, Hyukjun Lee
We present SAM, a State-space Audio-language Model that integrates an audio encoder with a Mamba-2 backbone. SAM-2.7B achieves 21.1 mAP on AudioSet and 17.6 SPICE on AudioCaps, mat…
cs.SD2025
Token Pruning in Audio Transformers: Optimizing Performance and Decoding Patch Importance
Taehan Lee, Hyukjun Lee
Vision Transformers (ViTs) have achieved state-of-the-art performance across various computer vision tasks, but their high computational cost remains a challenge. Token pruning has…