3 citations · 3 across the 2 of their papers we have counts for
7 papers
Towards Perception-Informed Latent HRTF Representations
You Zhang, Andrew Francl, Ruohan Gao +3
Personalized head-related transfer functions (HRTFs) are essential for ensuring a realistic auditory experience over headphones, because they take into account individual anatomica…
Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation
Frank Cwitkowitz, Zhiyao Duan
Multi-Pitch Estimation (MPE) continues to be a sought after capability of Music Information Retrieval (MIR) systems, and is critical for many applications and downstream tasks invo…
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
You Zhang, Baotong Tian, Lin Zhang +1
Neural speech editing enables seamless partial edits to speech utterances, allowing modifications to selected content while preserving the rest of the audio unchanged. This useful…
HARP 2.0: Expanding Hosted, Asynchronous, Remote Processing for Deep Learning in the DAW
Christodoulos Benetatos, Frank Cwitkowitz, Nathan Pruyne +4
HARP 2.0 brings deep learning models to digital audio workstation (DAW) software through hosted, asynchronous, remote processing, allowing users to route audio from a plug-in inter…
Audio Visual Segmentation Through Text Embeddings
Kyungbok Lee, You Zhang, Zhiyao Duan
The goal of Audio-Visual Segmentation (AVS) is to localize and segment the sounding source objects from video frames. Research on AVS suffers from data scarcity due to the high cos…
SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
You Zhang, Yongyi Zang, Jiatong Shi +3
With the advancements in singing voice generation and the growing presence of AI singers on media platforms, the inaugural Singing Voice Deepfake Detection (SVDD) Challenge aims to…