Showing cs.SDShow all
2 papers · 1 filter
cs.SD2025
The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
Patrick O'Reilly, Julia Barnett, Hugo Flores GarcÃa +4
Musicians and nonmusicians alike use rhythmic sound gestures, such as tapping and beatboxing, to express drum patterns. While these gestures effectively communicate musical ideas,…
cs.SD2025
Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech
Patrick O'Reilly, Zeyu Jin, Jiaqi Su +1
In the audio modality, state-of-the-art watermarking methods leverage deep neural networks to allow the embedding of human-imperceptible signatures in generated audio. The ideal is…