13 papers
DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning
Xinyu Liang, Fredrik Cumlin, Victor Ungureanu +3
We introduce DNSMOS-C, a compact end-to-end speech quality assessment model that extends the DNSMOS Pro framework by integrating a MOS-guided triplet-based contrastive loss. Applie…
Generative Modeling for Physiological Signals
Xinqi Bao, Ernest Kamavuako, Saikat Chatterjee
Physiological signals support clinical diagnosis, health monitoring, rehabilitation, wearable sensing, and human--machine interaction. However, their applications are often constra…
Diffusion-Based Heart Sound Generation: Evaluation with Physiological Signal Metrics, Classifiers, and Expert Listening
Xinqi Bao, Jia Bi, Xin Chen +2
Publicly available phonocardiogram (PCG) datasets remain limited in size and pathological diversity, constraining both auscultation training and the generalisation of automated hea…
Semi-Supervised Model-Free Bayesian State Estimation from Compressed Measurements
Anubhab Ghosh, Yonina C. Eldar, Saikat Chatterjee
We consider data-driven Bayesian state estimation from compressed measurements (BSCM) of a model-free process. The dimension of the temporal measurement vector is lower than that o…
pDANSE: Particle-based Data-driven Nonlinear State Estimation from Nonlinear Measurements
Anubhab Ghosh, Yonina C. Eldar, Saikat Chatterjee
We consider the problem of designing a data-driven nonlinear state estimation (DANSE) method that uses (noisy) nonlinear measurements of a process whose underlying state transition…
SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment
Fengyuan Cao, Xinyu Liang, Fredrik Cumlin +4
Designing a speech quality assessment (SQA) system for estimating mean-opinion-score (MOS) of multi-rate speech with varying sampling frequency (16-48 kHz) is a challenging task. T…