AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech
arXiv:1611.09207
Abstract
Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep recurrent neural network whose inputs consist solely of a raw waveform. Our best models provide utterance-level estimates of MOS only moderately inferior to sampled human ratings, as shown by Pearson and Spearman correlations. When multiple utterances are scored and averaged, a scenario common in synthesizer quality assessment, AutoMOS achieves correlations approaching those of human raters. The AutoMOS model has a number of applications, such as the ability to explore the parameter space of a speech synthesizer without requiring a human-in-the-loop.
4 pages, 2 figures, 2 tables, NIPS 2016 End-to-end Learning for Speech and Audio Processing Workshop
Cited by in corpus (10)
- Deep Learning Based Assessment of Synthetic Speech Naturalness
- NORESQA: A Framework for Speech Quality Assessment using Non-Matching References
- SQuId: Measuring Speech Naturalness in Many Languages
- RAMP: Retrieval-Augmented MOS Prediction via Confidence-based Dynamic Weighting
- Investigating Content-Aware Neural Text-To-Speech MOS Prediction Using Prosodic and Linguistic Features
- MBNet: MOS Prediction for Synthesized Speech with Mean-Bias Network
- CCATMos: Convolutional Context-aware Transformer Network for Non-intrusive Speech Quality Assessment
- Predicting pairwise preferences between TTS audio stimuli using parallel ratings data and anti-symmetric twin neural networks
- A Pyramid Recurrent Network for Predicting Crowdsourced Speech-Quality Ratings of Real-World Signals
- SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction