4 papers
SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment
Fengyuan Cao, Xinyu Liang, Fredrik Cumlin +4
Designing a speech quality assessment (SQA) system for estimating mean-opinion-score (MOS) of multi-rate speech with varying sampling frequency (16-48 kHz) is a challenging task. T…
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
Xinyu Liang, Fredrik Cumlin, Victor Ungureanu +3
Self-supervised learning (SSL) models like Wav2Vec2, HuBERT, and WavLM have been widely used in speech processing. These transformer-based models consist of multiple layers, each c…
Multivariate Probabilistic Assessment of Speech Quality
Fredrik Cumlin, Xinyu Liang, Victor Ungureanu +3
The mean opinion score (MOS) is a standard metric for assessing speech quality, but its singular focus fails to identify specific distortions when low scores are observed. The NISQ…
Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models
Fredrik Cumlin, Xinyu Liang, Victor Ungureanu +3
In this article, we provide an experimental observation: Deep neural network (DNN) based speech quality assessment (SQA) models have inherent latent representations where many type…