94 citations · 119 across the 8 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2021★ 1 cited
Raw Waveform Encoder with Multi-Scale Globally Attentive Locally Recurrent Networks for End-to-End Speech Recognition
Max W. Y. Lam, Jun Wang, Chao Weng +2
End-to-end speech recognition generally uses hand-engineered acoustic features as input and excludes the feature extraction module from its joint optimization. To extract learnable…
cs.SD2018
Polyphonic audio tagging with sequentially labelled data using CRNN with learnable gated linear units
Yuanbo Hou, Qiuqiang Kong, Jun Wang +1
Audio tagging aims to detect the types of sound events occurring in an audio recording. To tag the polyphonic audio recordings, we propose to use Connectionist Temporal Classificat…