4 papers
MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation
Di Zhu, Zixuan Li
Distributional metrics such as Fréchet Audio Distance cannot score individual music clips and correlate poorly with human judgments, while the only per-sample learned metric achie…
Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
Zixuan Li, Xueliang Zhang, Lei Miao +3
Audio-Visual Target Speaker Extraction (AVTSE) aims to isolate a target speaker's voice in a multi-speaker environment with visual cues as auxiliary. Most of the existing AVTSE met…
Loud-loss: A Perceptually Motivated Loss Function for Speech Enhancement Based on Equal-Loudness Contours
Zixuan Li, Xueliang Zhang, Changjiang Zhao +5
The mean squared error (MSE) is a ubiquitous loss function for speech enhancement, but its problem is that the error cannot reflect the auditory perception quality. This is because…
Robust Target Speaker Direction of Arrival Estimation
Zixuan Li, Shulin He, Xueliang Zhang
In multi-speaker environments the direction of arrival (DOA) of a target speaker is key for improving speech clarity and extracting target speaker's voice. However, traditional DOA…