activity
20242026
collaborators

25 papers

eess.AS2026

Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling

Ye-Xin Lu, Xin Wang, Yang Ai +3

Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or enta…

cs.SD2026

LatentFlowSR: High-Fidelity Audio Super-Resolution via Noise-Robust Latent Flow Matching

Fei Liu, Yang Ai, Hui-Peng Du +2

Audio super-resolution aims to recover missing high-frequency details from bandwidth-limited low-resolution audio, thereby improving the naturalness and perceptual quality of the r…

cs.SD2026

Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes

Xiao-Hang Jiang, Han-Jie Guo, Ying-Si Liang +4

Ambient clinical scribes increasingly combine Automatic Speech Recognition with Large Language Models to automate documentation. However, traditional metrics like Word Error Rate m…

eess.AS2026

VoCodec: A Low-bitrate Streamable Neural Speech Codec with Voicing-driven Quantization

Xiao-Hang Jiang, Yang Ai, Rui-Chen Zheng +3

Neural speech codecs are key to speech transmission and storage, but most use uniform quantization across frames, allocating the same bitrate regardless of content and wasting bits…

eess.AS2026

An Ultra-Low-Bitrate Neural Speech Codec with Plain-to-Pseudo Synergistic Vector Quantization

Xiao-Hang Jiang, Yang Ai, Fei Liu +4

Most neural speech codecs use residual vector quantization (RVQ), in which later VQs contribute less but consume the same bitrate, leading to inefficiency. We propose P2PSynCodec,…

cs.SD2026

UniVocal: Unified Speech-Singing Code-Switching Synthesis

Yufei Shi, Qian Chen, Wen Wang +3

We propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis - a task where transitions ar…