From the 1 of 8 linked papers with an AI index.
5 papers · 1 filter
Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis
Mingyue Huo, Yuheng Zhang, Hao Zhang
The paper introduces an inference‑time polar projection method to diagnose how components of STFT‑domain speech enhancement affect automatic speech recognition, revealing that magn…
TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion
Zihan Zhang, Shreeram Suresh Chandra, Zongyang Du +5
Traditional emotional voice conversion (EVC) conditions generation on explicit target emotions like labels or references, defining the target affective state but omitting the direc…
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo +3
Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored. Crucially, there is no consensus on whether audio-langu…
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM
Jiatong Shi, Chunlei Zhang, Jinchuan Tian +4
Recent advances in speech language models (LLMs) have extended textual LLMs to the speech domain, but balancing speech understanding and generation remains challenging, especially…
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
Mingyue Huo, Wei-Cheng Tseng, Yiwen Shao +2
Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward…