60 citations · 174 across the 16 of their papers we have counts for
8 papers · 1 filter
TaL: a synchronised multi-speaker corpus of ultrasound tongue imaging, audio, and lip videos
Manuel Sam Ribeiro, Jennifer Sanger, Jing-Xuan Zhang +4
We present the Tongue and Lips corpus (TaL), a multi-speaker corpus of audio, ultrasound tongue imaging, and lip videos. TaL consists of two parts: TaL1 is a set of six recording s…
On the Usefulness of Self-Attention for Automatic Speech Recognition with Transformers
Shucong Zhang, Erfan Loweimi, Peter Bell +1
Self-attention models such as Transformers, which can capture temporal relationships without being limited by the distance between events, have given competitive speech recognition…
Stochastic Attention Head Removal: A simple and effective method for improving Transformer Based ASR Models
Shucong Zhang, Erfan Loweimi, Peter Bell +1
Recently, Transformer based models have shown competitive automatic speech recognition (ASR) performance. One key factor in the success of these models is the multi-head attention…
Leveraging speaker attribute information using multi task learning for speaker verification and diarization
Chau Luu, Peter Bell, Steve Renals
Deep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between…
Word Error Rate Estimation Without ASR Output: e-WER2
Ahmed Ali, Steve Renals
Measuring the performance of automatic speech recognition (ASR) systems requires manually transcribed data in order to compute the word error rate (WER), which is often time-consum…
When Can Self-Attention Be Replaced by Feed Forward Layers?
Shucong Zhang, Erfan Loweimi, Peter Bell +1
Recently, self-attention models such as Transformers have given competitive results compared to recurrent neural network systems in speech recognition. The key factor for the outst…