3 papers
cs.SD2026
Woosh: A Sound Effects Foundation Model
Gaëtan Hadjeres, Marc Ferras, Khaled Koutini +7
The audio research community depends on open generative models as foundational tools for building novel approaches and establishing baselines. In this report, we present Woosh, Son…
cs.CL2024
Exploring neural oscillations during speech perception via surrogate gradient spiking neural networks
Alexandre Bittar, Philip N. Garner
Understanding cognitive processes in the brain demands sophisticated models capable of replicating neural dynamics at large scales. We present a physiologically inspired speech rec…
cs.SD2023
Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
Alexandre Bittar, Paul Dixon, Mohammad Samragh +2
Using a vision-inspired keyword spotting framework, we propose an architecture with input-dependent dynamic depth capable of processing streaming audio. Specifically, we extend a c…