3 papers
eess.AS2026
Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis
Mingyue Huo, Yuheng Zhang, Hao Zhang
Speech enhancement (SE) can substantially improve perceptual quality, yet enhanced speech does not necessarily improve automatic speech recognition (ASR). Existing remedies, such a…
eess.AS2026
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
Mingyue Huo, Yiwen Shao, Yuheng Zhang
We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs:…
eess.AS2025
Identifying and Calibrating Overconfidence in Noisy Speech Recognition
Mingyue Huo, Yuheng Zhang, Yan Tang
Modern end-to-end automatic speech recognition (ASR) models like Whisper not only suffer from reduced recognition accuracy in noise, but also exhibit overconfidence - assigning hig…