1 citations · 2 across the 3 of their papers we have counts for
Showing cs.SDShow all
3 papers · 1 filter
cs.SD2025
PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
Yujia Xiao, Liumeng Xue, Lei He +8
Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into…
cs.SD2023
StyleSpeech: Self-supervised Style Enhancing with VQ-VAE-based Pre-training for Expressive Audiobook Speech Synthesis
Xueyuan Chen, Xi Wang, Shaofei Zhang +4
The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these is…
cs.SD2023★ 1 cited
Large-Scale Automatic Audiobook Creation
Brendan Walsh, Mark Hamilton, Greg Newby +8
An audiobook can dramatically improve a work of literature's accessibility and improve reader engagement. However, audiobooks can take hundreds of hours of human effort to create,…