25 papers
Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
Ye-Xin Lu, Xin Wang, Yang Ai +3
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or enta…
LatentFlowSR: High-Fidelity Audio Super-Resolution via Noise-Robust Latent Flow Matching
Fei Liu, Yang Ai, Hui-Peng Du +2
Audio super-resolution aims to recover missing high-frequency details from bandwidth-limited low-resolution audio, thereby improving the naturalness and perceptual quality of the r…
Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes
Xiao-Hang Jiang, Han-Jie Guo, Ying-Si Liang +4
Ambient clinical scribes increasingly combine Automatic Speech Recognition with Large Language Models to automate documentation. However, traditional metrics like Word Error Rate m…
VoCodec: A Low-bitrate Streamable Neural Speech Codec with Voicing-driven Quantization
Xiao-Hang Jiang, Yang Ai, Rui-Chen Zheng +3
Neural speech codecs are key to speech transmission and storage, but most use uniform quantization across frames, allocating the same bitrate regardless of content and wasting bits…
An Ultra-Low-Bitrate Neural Speech Codec with Plain-to-Pseudo Synergistic Vector Quantization
Xiao-Hang Jiang, Yang Ai, Fei Liu +4
Most neural speech codecs use residual vector quantization (RVQ), in which later VQs contribute less but consume the same bitrate, leading to inefficiency. We propose P2PSynCodec,…
UniVocal: Unified Speech-Singing Code-Switching Synthesis
Yufei Shi, Qian Chen, Wen Wang +3
We propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis - a task where transitions ar…