paper

Controllable and Interpretable Singing Voice Decomposition via Assem-VC

arXiv:2110.12676

Abstract

We propose a singing decomposition system that encodes time-aligned linguistic content, pitch, and source speaker identity via Assem-VC. With decomposed speaker-independent information and the target speaker's embedding, we could synthesize the singing voice of the target speaker. In conclusion, we made a perfectly synced duet with the user's singing voice and the target singer's converted singing voice.

Accepted to NeurIPS Workshop on ML for Creativity and Design 2021 (Oral)

References in corpus (1)