65 citations · 77 across the 26 of their papers we have counts for
17 papers · 1 filter
CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction
Hui-Peng Du, Yang Ai, Xiao-Hang Jiang +2
This paper targets a new scenario that integrates speech separation with speech compression, aiming to disentangle multiple speakers while producing discrete representations for ef…
Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation Simulation
Rui-Chen Zheng, Yang Ai, Hui-Peng Du +1
Noise robustness remains a critical challenge for deploying neural speech codecs in real-world acoustic scenarios where background noise is often inevitable. A key observation we m…
A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation
En-Wei Zhang, Hui-Peng Du, Xiao-Hang Jiang +2
While many current neural speech codecs achieve impressive reconstructed speech quality, they often neglect latency and complexity considerations, limiting their practical deployme…
A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
Hui-Peng Du, Yang Ai, Zhen-Hua Ling
The majority of mainstream neural vocoders primarily focus on speech quality and generation speed, while overlooking latency, which is a critical factor in real-time applications.…
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
Ye-Xin Lu, Yu Gu, Kun Wei +3
This paper presents DAIEN-TTS, a zero-shot text-to-speech (TTS) framework that enables ENvironment-aware synthesis through Disentangled Audio Infilling. By leveraging separate spea…
Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
Rui-Chen Zheng, Wenrui Liu, Hui-Peng Du +6
Existing speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This u…