2 citations · 4 across the 3 of their papers we have counts for
14 papers
Using Rater and System Metadata to Explain Variance in the VoiceMOS Challenge 2022 Dataset
Michael Chinen, Jan Skoglund, Chandan K A Reddy +2
Non-reference speech quality models are important for a growing number of applications. The VoiceMOS 2022 challenge provided a dataset of synthetic voice conversion and text-to-spe…
SoundStream: An End-to-End Neural Audio Codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran +2
We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs. SoundStrea…
Handling Background Noise in Neural Speech Generation
Tom Denton, Alejandro Luebs, Felicia S. C. Lim +4
Recent advances in neural-network based generative modeling of speech has shown great potential for speech coding. However, the performance of such models drops when the input is n…
WARP-Q: Quality Prediction For Generative Neural Speech Codecs
Wissam A. Jassim, Jan Skoglund, Michael Chinen +1
Good speech quality has been achieved using waveform matching and parametric reconstruction coders. Recently developed very low bit rate generative codecs can reconstruct high qual…
Generative Speech Coding with Predictive Variance Regularization
W. Bastiaan Kleijn, Andrew Storus, Michael Chinen +5
The recent emergence of machine-learning based generative models for speech suggests a significant reduction in bit rate for speech codecs is possible. However, the performance of…
ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric
Michael Chinen, Felicia S. C. Lim, Jan Skoglund +3
Estimation of perceptual quality in audio and speech is possible using a variety of methods. The combined v3 release of ViSQOL and ViSQOLAudio (for speech and audio, respectively,)…