activity
20232026
collaborators
Showing cs.SDShow all

9 papers · 1 filter

cs.SD2026

LatentFlowSR: High-Fidelity Audio Super-Resolution via Noise-Robust Latent Flow Matching

Fei Liu, Yang Ai, Hui-Peng Du +2

Audio super-resolution aims to recover missing high-frequency details from bandwidth-limited low-resolution audio, thereby improving the naturalness and perceptual quality of the r…

cs.SD2025

A Stage-Wise Learning Strategy with Fixed Anchors for Robust Speaker Verification

Bin Gu, Lipeng Dai, Huipeng Du +2

Learning robust speaker representations under noisy conditions presents significant challenges, which requires careful handling of both discriminative and noise-invariant propertie…

cs.SD2025

Universal Discrete-Domain Speech Enhancement

Fei Liu, Yang Ai, Ye-Xin Lu +3

In real-world scenarios, speech signals are inevitably corrupted by various types of interference, making speech enhancement (SE) a critical task for robust speech processing. Howe…

cs.SD2024

Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion

Yu-Fei Shi, Yang Ai, Ye-Xin Lu +2

We participated in track 2 of the VoiceMOS Challenge 2024, which aimed to predict the mean opinion score (MOS) of singing samples. Our submission secured the first place among all…

cs.SD2024

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram

Xiao-Hang Jiang, Hui-Peng Du, Yang Ai +2

This paper proposes ESTVocoder, a novel excitation-spectral-transformed neural vocoder within the framework of source-filter theory. The ESTVocoder transforms the amplitude and pha…

cs.SD2024

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features

Yu-Fei Shi, Yang Ai, Ye-Xin Lu +2

Assessing the naturalness of speech using mean opinion score (MOS) prediction models has positive implications for the automatic evaluation of speech synthesis systems. Early MOS p…