9 papers
MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres
Wenhao Feng, Yuxun Tang, Jiatong Shi +1
Singing voice synthesis (SVS) has progressed rapidly, yet its ability to generalize across diverse musical genres remains underexplored. Existing benchmarks are heavily biased towa…
SingMOS-Pro: An Comprehensive Benchmark for Singing Quality Assessment
Yuxun Tang, Lan Liu, Wenhao Feng +5
Singing voice generation progresses rapidly, yet evaluating singing quality remains a critical challenge. Human subjective assessment, typically in the form of listening tests, is…
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
Jiatong Shi, Jinchuan Tian, Yihan Wu +17
Neural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance do…
Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
Yuning Wu, Jiatong Shi, Yifeng Yu +7
This research presents Muskits-ESPnet, a versatile toolkit that introduces new paradigms to Singing Voice Synthesis (SVS) through the application of pretrained audio models in both…
SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
Yuxun Tang, Jiatong Shi, Yuning Wu +1
In speech generation tasks, human subjective ratings, usually referred to as the opinion score, are considered the "gold standard" for speech quality evaluation, with the mean opin…
SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models
Yuxun Tang, Yuning Wu, Jiatong Shi +1
Discrete representation has shown advantages in speech generation tasks, wherein discrete tokens are derived by discretizing hidden features from self-supervised learning (SSL) pre…