11 papers
VoCodec: A Low-bitrate Streamable Neural Speech Codec with Voicing-driven Quantization
Xiao-Hang Jiang, Yang Ai, Rui-Chen Zheng +3
Neural speech codecs are key to speech transmission and storage, but most use uniform quantization across frames, allocating the same bitrate regardless of content and wasting bits…
An Ultra-Low-Bitrate Neural Speech Codec with Plain-to-Pseudo Synergistic Vector Quantization
Xiao-Hang Jiang, Yang Ai, Fei Liu +4
Most neural speech codecs use residual vector quantization (RVQ), in which later VQs contribute less but consume the same bitrate, leading to inefficiency. We propose P2PSynCodec,…
CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction
Hui-Peng Du, Yang Ai, Xiao-Hang Jiang +2
This paper targets a new scenario that integrates speech separation with speech compression, aiming to disentangle multiple speakers while producing discrete representations for ef…
Universal Discrete-Domain Speech Enhancement
Fei Liu, Yang Ai, Ye-Xin Lu +3
In real-world scenarios, speech signals are inevitably corrupted by various types of interference, making speech enhancement (SE) a critical task for robust speech processing. Howe…
Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
Rui-Chen Zheng, Wenrui Liu, Hui-Peng Du +6
Existing speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This u…
Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation Simulation
Rui-Chen Zheng, Yang Ai, Hui-Peng Du +1
Noise robustness remains a critical challenge for deploying neural speech codecs in real-world acoustic scenarios where background noise is often inevitable. A key observation we m…