From the 1 of 5 linked papers with an AI index.
5 papers
Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis
Mingyue Huo, Yuheng Zhang, Hao Zhang
The paper introduces an inference‑time polar projection method to diagnose how components of STFT‑domain speech enhancement affect automatic speech recognition, revealing that magn…
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
Mingyue Huo, Yiwen Shao, Yuheng Zhang
We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs:…
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo +3
Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored. Crucially, there is no consensus on whether audio-langu…
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
Mingyue Huo, Wei-Cheng Tseng, Yiwen Shao +2
Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward…
Beyond Speaker Identity: Text Guided Target Speech Extraction
Mingyue Huo, Abhinav Jain, Cong Phuoc Huynh +4
Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker's identity like enrollment audio, face images, or videos, which may not always be available.…