175 citations · 473 across the 30 of their papers we have counts for
14 papers · 1 filter
Self Voice Conversion as an Attack against Neural Audio Watermarking
Yigitcan Özer, Wanying Ge, Zhe Zhang +2
Audio watermarking embeds auxiliary information into speech while maintaining speaker identity, linguistic content, and perceptual quality. Although recent advances in neural and d…
Zero-Day Audio DeepFake Detection via Retrieval Augmentation and Profile Matching
Xuechen Liu, Xin Wang, Junichi Yamagishi
Modern audio deepfake detectors built on foundation models and large training datasets achieve promising detection performance. However, they struggle with zero-day attacks, where…
LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech
Xuechen Liu, Wanying Ge, Xin Wang +1
This study introduces LENS-DF, a novel and comprehensive recipe for training and evaluating audio deepfake detection and temporal localization under complicated and realistic audio…
MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling
Jingjing Tang, Xin Wang, Zhe Zhang +3
Generating expressive audio performances from music scores requires models to capture both instrument acoustics and human interpretation. Traditional music performance synthesis pi…
A Comparative Study on Proactive and Passive Detection of Deepfake Speech
Chia-Hua Wu, Wanying Ge, Xin Wang +3
Solutions for defending against deepfake speech fall into two categories: proactive watermarking models and passive conventional deepfake detectors. While both address common threa…
Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
Jingjing Tang, Erica Cooper, Xin Wang +2
This paper presents an integrated system that transforms symbolic music scores into expressive piano performance audio. By combining a Transformer-based Expressive Performance Rend…