3 papers
eess.AS2025
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
Mingyue Huo, Wei-Cheng Tseng, Yiwen Shao +2
Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward…
eess.AS2025
Identifying and Calibrating Overconfidence in Noisy Speech Recognition
Mingyue Huo, Yuheng Zhang, Yan Tang
Modern end-to-end automatic speech recognition (ASR) models like Whisper not only suffer from reduced recognition accuracy in noise, but also exhibit overconfidence - assigning hig…
eess.AS2025
Beyond Speaker Identity: Text Guided Target Speech Extraction
Mingyue Huo, Abhinav Jain, Cong Phuoc Huynh +4
Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker's identity like enrollment audio, face images, or videos, which may not always be available.…