57 citations · 71 across the 2 of their papers we have counts for
2 papers
cs.CL2016★ 57 cited
AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech
Brian Patton, Yannis Agiomyrgiannakis, Michael Terry +3
Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opin…
cs.SD2016★ 14 cited
CNN Architectures for Large-Scale Audio Classification
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis +10
Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of…