3 papers
cs.MM2026
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models
Yuwen Wang, Tian-Hao Zhang, Minghao Cai +7
Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rath…
cs.SD2026
SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios
Ziyang Jiang, Yu Chen, Zexu Pan +5
Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-b…
cs.SD2026
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
Ziyang Jiang, Jiahe Lei, Xueyan Chen +4
Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has…