Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
Xiaomi-CocktailASR-1 Technical Report
Yiru Zhang, Hang Su, Lichun Fan +10
Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party prob…
cs.SD2025
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
Wenhuan Lu, Xinyue Song, Wenjun Ke +3
Audio grounding, or speech-driven open-set object detection, aims to localize and identify objects directly from speech, enabling generalization beyond predefined categories. This…
cs.SD2024
Robust Channel Learning for Large-Scale Radio Speaker Verification
Wenhao Yang, Jianguo Wei, Wenhuan Lu +2
Recent research in speaker verification has increasingly focused on achieving robust and reliable recognition under challenging channel conditions and noisy environments. Identifyi…