activity
20192022
most citedVISinger 2: High-Fidelity End-to-End Singing Voice Synthesis Enhanced by Digital Signal Processing Synthesizer

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.SD20221 cited

VISinger 2: High-Fidelity End-to-End Singing Voice Synthesis Enhanced by Digital Signal Processing Synthesizer

Yongmao Zhang, Heyang Xue, Hanzhao Li +4

End-to-end singing voice synthesis (SVS) model VISinger can achieve better performance than the typical two-stage model with fewer parameters. However, VISinger has several problem…

eess.AS2022

Audio Deep Fake Detection System with Neural Stitching for ADD 2022

Rui Yan, Cheng Wen, Shuran Zhou +3

This paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge\cite{Yi2022ADD}. The very same system was used for both two ro…

eess.AS2022

Time Domain Adversarial Voice Conversion for ADD 2022

Cheng Wen, Tingwei Guo, Xingjun Tan +5

In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) sy…

cs.SD2022

Audio-Visual Wake Word Spotting System For MISP Challenge 2021

Yanguang Xu, Jianwei Sun, Yang Han +7

This paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage…

eess.AS2020

DiDiSpeech: A Large Scale Mandarin Speech Corpus

Tingwei Guo, Cheng Wen, Dongwei Jiang +8

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the…

cs.CL2019

DELTA: A DEep learning based Language Technology plAtform

Kun Han, Junwen Chen, Hui Zhang +20

In this paper we present DELTA, a deep learning based language technology platform. DELTA is an end-to-end platform designed to solve industry level natural language and speech pro…