4 papers
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
Xinyu Liang, Fredrik Cumlin, Victor Ungureanu +3
Self-supervised learning (SSL) models like Wav2Vec2, HuBERT, and WavLM have been widely used in speech processing. These transformer-based models consist of multiple layers, each c…
Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
Fredrik Cumlin, Xinyu Liang, Anubhab Ghosh +1
Non-intrusive speech quality assessment (SQA) systems suffer from limited training data and costly human annotations, hindering their generalization to real-time conferencing calls…
Multivariate Probabilistic Assessment of Speech Quality
Fredrik Cumlin, Xinyu Liang, Victor Ungureanu +3
The mean opinion score (MOS) is a standard metric for assessing speech quality, but its singular focus fails to identify specific distortions when low scores are observed. The NISQ…
Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models
Fredrik Cumlin, Xinyu Liang, Victor Ungureanu +3
In this article, we provide an experimental observation: Deep neural network (DNN) based speech quality assessment (SQA) models have inherent latent representations where many type…