4 papers
A Mixed-Behavior Vote Model for Multimedia Subjective Quality Votes, Means, and Variances
Jaden Pieper, Stephen D. Voran
The relationship between subjective test vote variance and vote mean (or MOS) is well-studied, and the mathematically admissible vote variance region has been previously defined. W…
Unseen but not Unknown: Using Dataset Concealment to Robustly Evaluate Speech Quality Estimation Models
Jaden Pieper, Stephen D. Voran
We introduce Dataset Concealment (DSC), a rigorous new procedure for evaluating and interpreting objective speech quality estimation models. DSC quantifies and decomposes the perfo…
Why some audio signal short-time Fourier transform coefficients have nonuniform phase distributions
Stephen D. Voran
The short-time Fourier transform (STFT) represents a window of audio samples as a set of complex coefficients. These are advantageously viewed as magnitudes and phases and the over…
AlignNet: Learning dataset score alignment functions to enable better training of speech quality estimators
Jaden Pieper, Stephen D. Voran
We develop two complementary advances for training no-reference (NR) speech quality estimators with independent datasets. Multi-dataset finetuning (MDF) pretrains an NR estimator o…