9 papers
DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning
Xinyu Liang, Fredrik Cumlin, Victor Ungureanu +3
We introduce DNSMOS-C, a compact end-to-end speech quality assessment model that extends the DNSMOS Pro framework by integrating a MOS-guided triplet-based contrastive loss. Applie…
SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment
Fengyuan Cao, Xinyu Liang, Fredrik Cumlin +4
Designing a speech quality assessment (SQA) system for estimating mean-opinion-score (MOS) of multi-rate speech with varying sampling frequency (16-48 kHz) is a challenging task. T…
DNS: Data-driven Nonlinear Smoother for Complex Model-free Process
Fredrik Cumlin, Anubhab Ghosh, Saikat Chatterjee
We propose data-driven nonlinear smoother (DNS) to estimate a hidden state sequence of a complex dynamical process from a noisy, linear measurement sequence. The dynamical process…
Rho-Perfect: Correlation Ceiling For Subjective Evaluation Datasets
Fredrik Cumlin
Subjective ratings contain inherent noise that limits the model-human correlation, but this reliability issue is rarely quantified. In this paper, we present -Perfect, a practi…
VSE: Variational state estimation of complex model-free process
Gustav Norén, Anubhab Ghosh, Fredrik Cumlin +1
We design a variational state estimation (VSE) method that provides a closed-form Gaussian posterior of an underlying complex dynamical process from (noisy) nonlinear measurements.…
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
Xinyu Liang, Fredrik Cumlin, Victor Ungureanu +3
Self-supervised learning (SSL) models like Wav2Vec2, HuBERT, and WavLM have been widely used in speech processing. These transformer-based models consist of multiple layers, each c…