4 papers
VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
Kyungguen Byun, Jason Filos, Erik Visser +1
Noise suppression (NS) algorithms are effective in improving speech quality in many cases. However, aggressive noise suppression can damage the target speech, reducing both speech…
Highly Controllable Diffusion-based Any-to-Any Voice Conversion Model with Frame-level Prosody Feature
Kyungguen Byun, Sunkuk Moon, Erik Visser
We propose a highly controllable voice manipulation system that can perform any-to-any voice conversion (VC) and prosody modulation simultaneously. State-of-the-art VC systems can…
Parameter Efficient Audio Captioning With Faithful Guidance Using Audio-text Shared Latent Representation
Arvind Krishna Sridhar, Yinyi Guo, Erik Visser +1
There has been significant research on developing pretrained transformer architectures for multimodal-to-text generation tasks. Albeit performance improvements, such models are fre…
Detecting False Alarms and Misses in Audio Captions
Rehana Mahfuz, Yinyi Guo, Arvind Krishna Sridhar +1
Metrics to evaluate audio captions simply provide a score without much explanation regarding what may be wrong in case the score is low. Manual human intervention is needed to find…