42 citations · 42 across the 9 of their papers we have counts for
4 papers · 1 filter
Multi-view Audio and Music Classification
Huy Phan, Huy Le Nguyen, Oliver Y. Chén +4
We propose in this work a multi-view learning approach for audio and music classification. Considering four typical low-level representations (i.e. different views) commonly used f…
Self-Attention Generative Adversarial Network for Speech Enhancement
Huy Phan, Huy Le Nguyen, Oliver Y. Chén +4
Existing generative adversarial networks (GANs) for speech enhancement solely rely on the convolution operation, which may obscure temporal dependencies across the sequence input.…
Spatio-Temporal Attention Pooling for Audio Scene Classification
Huy Phan, Oliver Y. Chén, Lam Pham +4
Acoustic scenes are rich and redundant in their content. In this work, we present a spatio-temporal attention pooling layer coupled with a convolutional recurrent neural network to…
Beyond Equal-Length Snippets: How Long is Sufficient to Recognize an Audio Scene?
Huy Phan, Oliver Y. Chén, Philipp Koch +4
Due to the variability in characteristics of audio scenes, some scenes can naturally be recognized earlier than others. In this work, rather than using equal-length snippets for al…