14 citations · 15 across the 4 of their papers we have counts for
4 papers
Learning Music Representations with wav2vec 2.0
Alessandro Ragano, Emmanouil Benetos, Andrew Hines
Learning music representations that are general-purpose offers the flexibility to finetune several downstream tasks using smaller datasets. The wav2vec 2.0 speech representation mo…
Using Rater and System Metadata to Explain Variance in the VoiceMOS Challenge 2022 Dataset
Michael Chinen, Jan Skoglund, Chandan K A Reddy +2
Non-reference speech quality models are important for a growing number of applications. The VoiceMOS 2022 challenge provided a dataset of synthetic voice conversion and text-to-spe…
Exploring the influence of fine-tuning data on wav2vec 2.0 model for blind speech quality prediction
Helard Becerra, Alessandro Ragano, Andrew Hines
Recent studies have shown how self-supervised models can produce accurate speech quality predictions. Speech representations generated by the pre-trained wav2vec 2.0 model allows c…
More for Less: Non-Intrusive Speech Quality Assessment with Limited Annotations
Alessandro Ragano, Emmanouil Benetos, Andrew Hines
Non-intrusive speech quality assessment is a crucial operation in multimedia applications. The scarcity of annotated data and the lack of a reference signal represent some of the m…