activity
20002021
most citedDynamic Evaluation of Neural Sequence Models

60 citations · 174 across the 16 of their papers we have counts for

collaborators
Showing 2020Show all

8 papers · 1 filter

eess.AS2020

TaL: a synchronised multi-speaker corpus of ultrasound tongue imaging, audio, and lip videos

Manuel Sam Ribeiro, Jennifer Sanger, Jing-Xuan Zhang +4

We present the Tongue and Lips corpus (TaL), a multi-speaker corpus of audio, ultrasound tongue imaging, and lip videos. TaL consists of two parts: TaL1 is a set of six recording s…

cs.CL2020★ 4 cited

On the Usefulness of Self-Attention for Automatic Speech Recognition with Transformers

Shucong Zhang, Erfan Loweimi, Peter Bell +1

Self-attention models such as Transformers, which can capture temporal relationships without being limited by the distance between events, have given competitive speech recognition…

cs.CL2020

Stochastic Attention Head Removal: A simple and effective method for improving Transformer Based ASR Models

Shucong Zhang, Erfan Loweimi, Peter Bell +1

Recently, Transformer based models have shown competitive automatic speech recognition (ASR) performance. One key factor in the success of these models is the multi-head attention…

cs.SD2020

Leveraging speaker attribute information using multi task learning for speaker verification and diarization

Chau Luu, Peter Bell, Steve Renals

Deep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between…

eess.AS2020

Word Error Rate Estimation Without ASR Output: e-WER2

Ahmed Ali, Steve Renals

Measuring the performance of automatic speech recognition (ASR) systems requires manually transcribed data in order to compute the word error rate (WER), which is often time-consum…

eess.AS2020★ 2 cited

When Can Self-Attention Be Replaced by Feed Forward Layers?

Shucong Zhang, Erfan Loweimi, Peter Bell +1

Recently, self-attention models such as Transformers have given competitive results compared to recurrent neural network systems in speech recognition. The key factor for the outst…