Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
Xiaomi-CocktailASR-1 Technical Report
Yiru Zhang, Hang Su, Lichun Fan +10
Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party prob…
cs.SD2025
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
Yiru Zhang, Hang Su, Lichun Fan +2
Target Speaker Automatic Speech Recognition (TS-ASR) aims to transcribe the speech of a specified target speaker from multi-speaker mixtures in cocktail party scenarios. Recent adv…