activity
20182026
most citedWaveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extension

65 citations · 77 across the 26 of their papers we have counts for

collaborators
Showing eess.ASShow all

17 papers · 1 filter

eess.AS2026

CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction

Hui-Peng Du, Yang Ai, Xiao-Hang Jiang +2

This paper targets a new scenario that integrates speech separation with speech compression, aiming to disentangle multiple speakers while producing discrete representations for ef…

eess.AS2025

Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation Simulation

Rui-Chen Zheng, Yang Ai, Hui-Peng Du +1

Noise robustness remains a critical challenge for deploying neural speech codecs in real-world acoustic scenarios where background noise is often inevitable. A key observation we m…

eess.AS2025

A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation

En-Wei Zhang, Hui-Peng Du, Xiao-Hang Jiang +2

While many current neural speech codecs achieve impressive reconstructed speech quality, they often neglect latency and complexity considerations, limiting their practical deployme…

eess.AS2025

A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction

Hui-Peng Du, Yang Ai, Zhen-Hua Ling

The majority of mainstream neural vocoders primarily focus on speech quality and generation speed, while overlooking latency, which is a critical factor in real-time applications.…

eess.AS2025

DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis

Ye-Xin Lu, Yu Gu, Kun Wei +3

This paper presents DAIEN-TTS, a zero-shot text-to-speech (TTS) framework that enables ENvironment-aware synthesis through Disentangled Audio Infilling. By leveraging separate spea…

eess.AS2025

Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding

Rui-Chen Zheng, Wenrui Liu, Hui-Peng Du +6

Existing speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This u…