collaborators

5 papers

eess.AS2026

DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning

Xinyu Liang, Fredrik Cumlin, Victor Ungureanu +3

We introduce DNSMOS-C, a compact end-to-end speech quality assessment model that extends the DNSMOS Pro framework by integrating a MOS-guided triplet-based contrastive loss. Applie…

eess.AS2026

SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment

Fengyuan Cao, Xinyu Liang, Fredrik Cumlin +4

Designing a speech quality assessment (SQA) system for estimating mean-opinion-score (MOS) of multi-rate speech with varying sampling frequency (16-48 kHz) is a challenging task. T…

eess.AS2025

Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech

Xinyu Liang, Fredrik Cumlin, Victor Ungureanu +3

Self-supervised learning (SSL) models like Wav2Vec2, HuBERT, and WavLM have been widely used in speech processing. These transformer-based models consist of multiple layers, each c…

eess.AS2025

Multivariate Probabilistic Assessment of Speech Quality

Fredrik Cumlin, Xinyu Liang, Victor Ungureanu +3

The mean opinion score (MOS) is a standard metric for assessing speech quality, but its singular focus fails to identify specific distortions when low scores are observed. The NISQ…

eess.AS2025

Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models

Fredrik Cumlin, Xinyu Liang, Victor Ungureanu +3

In this article, we provide an experimental observation: Deep neural network (DNN) based speech quality assessment (SQA) models have inherent latent representations where many type…