14 citations · 14 across the 10 of their papers we have counts for
10 papers
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
Yongyi Zang, Jiatong Shi, You Zhang +8
Recent singing voice synthesis and conversion advancements necessitate robust singing voice deepfake detection (SVDD) models. Current SVDD datasets face challenges due to limited c…
GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
Zehua Kcriss Li, Meiying Melissa Chen, Yi Zhong +2
Expressive speech synthesis aims to generate speech that captures a wide range of para-linguistic features, including emotion and articulation, though current research primarily em…
Toward Fully Self-Supervised Multi-Pitch Estimation
Frank Cwitkowitz, Zhiyao Duan
Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supe…
EDMSound: Spectrogram Based Diffusion Models for Efficient and High-Quality Audio Synthesis
Ge Zhu, Yutong Wen, Marc-André Carbonneau +1
Audio diffusion models can synthesize a wide variety of sounds. Existing models often operate on the latent domain with cascaded phase recovery modules to reconstruct waveform. Thi…
Mitigating Cross-Database Differences for Learning Unified HRTF Representation
Yutong Wen, You Zhang, Zhiyao Duan
Individualized head-related transfer functions (HRTFs) are crucial for accurate sound positioning in virtual auditory displays. As the acoustic measurement of HRTFs is resource-int…
SingNet: A Real-time Singing Voice Beat and Downbeat Tracking System
Mojtaba Heydari, Ju-Chiang Wang, Zhiyao Duan
Singing voice beat and downbeat tracking posses several applications in automatic music production, analysis and manipulation. Among them, some require real-time processing, such a…