6 citations · 13 across the 8 of their papers we have counts for
14 papers · 1 filter
It's not what you said, it's how you said it: discriminative perception of speech as a multichannel communication system
Sarenne Wallbridge, Peter Bell, Catherine Lai
People convey information extremely effectively through spoken interaction using multiple channels of information transmission: the lexical channel of what is said, and the non-lex…
Segmenting Subtitles for Correcting ASR Segmentation Errors
David Wan, Chris Kedzie, Faisal Ladhak +6
Typical ASR systems segment the input audio into utterances using purely acoustic information, which may not resemble the sentence-like units that are expected by conventional mach…
On the Usefulness of Self-Attention for Automatic Speech Recognition with Transformers
Shucong Zhang, Erfan Loweimi, Peter Bell +1
Self-attention models such as Transformers, which can capture temporal relationships without being limited by the distance between events, have given competitive speech recognition…
Stochastic Attention Head Removal: A simple and effective method for improving Transformer Based ASR Models
Shucong Zhang, Erfan Loweimi, Peter Bell +1
Recently, Transformer based models have shown competitive automatic speech recognition (ASR) performance. One key factor in the success of these models is the multi-head attention…
Subtitles to Segmentation: Improving Low-Resource Speech-to-Text Translation Pipelines
David Wan, Zhengping Jiang, Chris Kedzie +3
In this work, we focus on improving ASR output segmentation in the context of low-resource language speech-to-text translation. ASR output segmentation is crucial, as ASR systems s…
Multi-scale Octave Convolutions for Robust Speech Recognition
Joanna Rownicka, Peter Bell, Steve Renals
We propose a multi-scale octave convolution layer to learn robust speech representations efficiently. Octave convolutions were introduced by Chen et al [1] in the computer vision f…