4 citations · 6 across the 4 of their papers we have counts for
5 papers
Train your classifier first: Cascade Neural Networks Training from upper layers to lower layers
Shucong Zhang, Cong-Thanh Do, Rama Doddipatla +3
Although the lower layers of a deep neural network learn features which are transferable across datasets, these layers are not transferable within the same dataset. That is, in gen…
On the Usefulness of Self-Attention for Automatic Speech Recognition with Transformers
Shucong Zhang, Erfan Loweimi, Peter Bell +1
Self-attention models such as Transformers, which can capture temporal relationships without being limited by the distance between events, have given competitive speech recognition…
Stochastic Attention Head Removal: A simple and effective method for improving Transformer Based ASR Models
Shucong Zhang, Erfan Loweimi, Peter Bell +1
Recently, Transformer based models have shown competitive automatic speech recognition (ASR) performance. One key factor in the success of these models is the multi-head attention…
When Can Self-Attention Be Replaced by Feed Forward Layers?
Shucong Zhang, Erfan Loweimi, Peter Bell +1
Recently, self-attention models such as Transformers have given competitive results compared to recurrent neural network systems in speech recognition. The key factor for the outst…
Acoustic Model Adaptation from Raw Waveforms with SincNet
Joachim Fainberg, Ondřej Klejch, Erfan Loweimi +2
Raw waveform acoustic modelling has recently gained interest due to neural networks' ability to learn feature extraction, and the potential for finding better representations for a…