2 papers
cs.SD2026
VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding
Jashin Ye, Dongxiao Wang, Yixuan Ye +10
While large audio language models (LALMs) have achieved remarkable progress in audio processing at the second- or minute-level scale, understanding hour-level audio remains a funda…
cs.SD2024
DENSE: Dynamic Embedding Causal Target Speech Extraction
Yiwen Wang, Zeyu Yuan, Xihong Wu
Target speech extraction (TSE) focuses on extracting the speech of a specific target speaker from a mixture of signals. Existing TSE models typically utilize static embeddings as c…