activity
20182026
most citedWaveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extension

65 citations · 77 across the 26 of their papers we have counts for

collaborators
Showing cs.SDShow all

14 papers · 1 filter

cs.SD2025

Universal Discrete-Domain Speech Enhancement

Fei Liu, Yang Ai, Ye-Xin Lu +3

In real-world scenarios, speech signals are inevitably corrupted by various types of interference, making speech enhancement (SE) a critical task for robust speech processing. Howe…

cs.SD2025

Neural Speech Separation with Parallel Amplitude and Phase Spectrum Estimation

Fei Liu, Yang Ai, Zhen-Hua Ling

This paper proposes APSS, a novel neural speech separation model with parallel amplitude and phase spectrum estimation. Unlike most existing speech separation methods, the APSS dis…

cs.SD2025

Universal Preference-Score-based Pairwise Speech Quality Assessment

Yu-Fei Shi, Yang Ai, Zhen-Hua Ling

To compare the performance of two speech generation systems, one of the most effective approaches is estimating the preference score between their generated speech. This paper prop…

cs.SD2025

A Streamable Neural Audio Codec with Residual Scalar-Vector Quantization for Real-Time Communication

Xiao-Hang Jiang, Yang Ai, Rui-Chen Zheng +1

This paper proposes StreamCodec, a streamable neural audio codec designed for real-time communication. StreamCodec adopts a fully causal, symmetric encoder-decoder structure and op…

cs.SD2024

Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion

Yu-Fei Shi, Yang Ai, Ye-Xin Lu +2

We participated in track 2 of the VoiceMOS Challenge 2024, which aimed to predict the mean opinion score (MOS) of singing samples. Our submission secured the first place among all…

cs.SD2024

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram

Xiao-Hang Jiang, Hui-Peng Du, Yang Ai +2

This paper proposes ESTVocoder, a novel excitation-spectral-transformed neural vocoder within the framework of source-filter theory. The ESTVocoder transforms the amplitude and pha…