3 papers
eess.AS2025
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
Xinyu Liang, Fredrik Cumlin, Victor Ungureanu +3
Self-supervised learning (SSL) models like Wav2Vec2, HuBERT, and WavLM have been widely used in speech processing. These transformer-based models consist of multiple layers, each c…
eess.AS2025
Multivariate Probabilistic Assessment of Speech Quality
Fredrik Cumlin, Xinyu Liang, Victor Ungureanu +3
The mean opinion score (MOS) is a standard metric for assessing speech quality, but its singular focus fails to identify specific distortions when low scores are observed. The NISQ…
eess.AS2025
Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models
Fredrik Cumlin, Xinyu Liang, Victor Ungureanu +3
In this article, we provide an experimental observation: Deep neural network (DNN) based speech quality assessment (SQA) models have inherent latent representations where many type…