activity
20212026
most citedThe Multi-speaker Multi-style Voice Cloning Challenge 2021

3 citations · 4 across the 6 of their papers we have counts for

collaborators

7 papers

cs.SD2026

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Ziyang Ma, Zhikang Niu, Wenming Tu +30

We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To sup…

eess.AS2025

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Hanzhao Li, Yuke Li, Xinsheng Wang +4

Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user ne…

eess.AS2023

MSM-VC: High-fidelity Source Style Transfer for Non-Parallel Voice Conversion by Multi-scale Style Modeling

Zhichao Wang, Xinsheng Wang, Qicong Xie +4

In addition to conveying the linguistic content from source speech to converted speech, maintaining the speaking style of source speech also plays an important role in the voice co…

cs.SD2022

UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice Synthesis

Yi Lei, Shan Yang, Xinsheng Wang +4

Text-to-speech (TTS) and singing voice synthesis (SVS) aim at generating high-quality speaking and singing voice according to textual input and music scores, respectively. Unifying…

eess.AS20221 cited

Expressive-VC: Highly Expressive Voice Conversion with Attention Fusion of Bottleneck and Perturbation Features

Ziqian Ning, Qicong Xie, Pengcheng Zhu +5

Voice conversion for highly expressive speech is challenging. Current approaches struggle with the balancing between speaker similarity, intelligibility and expressiveness. To addr…

cs.CV2021

AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person

Xinsheng Wang, Qicong Xie, Jihua Zhu +2

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. I…