6 papers
XAttnMark: Learning Robust Audio Watermarking with Cross-Attention
Yixin Liu, Lie Lu, Jihui Jin +2
The rapid proliferation of generative audio synthesis and editing technologies has raised serious concerns about copyright infringement, data provenance, and the spread of misinfor…
Are Deep Speech Denoising Models Robust to Adversarial Noise?
Will Schwarzer, Neel Chaudhari, Philip S. Thomas +2
Deep noise suppression (DNS) models enjoy widespread use throughout a variety of high-stakes speech applications. However, we show that four recent DNS models can each be reduced t…
Decomposing multimodal embedding spaces with group-sparse autoencoders
Chiraag Kaushik, Davis Barch, Andrea Fanelli
The Linear Representation Hypothesis asserts that the embeddings learned by neural networks can be understood as linear combinations of features corresponding to high-level concept…
Audio-Visual Speech Separation via Bottleneck Iterative Network
Sidong Zhang, Shiv Shankar, Trang Nguyen +2
Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to ob…
Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval
Shanti Stewart, Gouthaman KV, Lie Lu +1
Content creators often use music to enhance their videos, from soundtracks in movies to background music in video blogs and social media content. However, identifying the best musi…
Accent Conversion with Articulatory Representations
Yashish M. Siriwardena, Nathan Swedlow, Audrey Howard +4
Conversion of non-native accented speech to native (American) English has a wide range of applications such as improving intelligibility of non-native speech. Previous work on this…