29 citations · 150 across the 31 of their papers we have counts for
11 papers · 1 filter
Does Joint Training Really Help Cascaded Speech Translation?
Viet Anh Khoa Tran, David Thulke, Yingbo Gao +2
Currently, in speech translation, the straightforward approach - cascading a recognition system with a translation system - delivers state-of-the-art results. However, fundamental…
On Architectures and Training for Raw Waveform Feature Extraction in ASR
Peter Vieting, Christoph Lüscher, Wilfried Michel +2
With the success of neural network based modeling in automatic speech recognition (ASR), many studies investigated acoustic modeling and learning of feature extractors directly bas…
Early Stage LM Integration Using Local and Global Log-Linear Combination
Wilfried Michel, Ralf Schlüter, Hermann Ney
Sequence-to-sequence models with an implicit alignment mechanism (e.g. attention) are closing the performance gap towards traditional hybrid hidden Markov models (HMM) for the task…
Investigation of Large-Margin Softmax in Neural Language Modeling
Jingjing Huo, Yingbo Gao, Weiyue Wang +2
To encourage intra-class compactness and inter-class separability among trainable feature vectors, large-margin softmax methods are developed and widely applied in the face recogni…
A systematic comparison of grapheme-based vs. phoneme-based label units for encoder-decoder-attention models
Mohammad Zeineldeen, Albert Zeyer, Wei Zhou +3
Following the rationale of end-to-end modeling, CTC, RNN-T or encoder-decoder-attention models for automatic speech recognition (ASR) use graphemes or grapheme-based subword units…
A New Training Pipeline for an Improved Neural Transducer
Albert Zeyer, André Merboldt, Ralf Schlüter +1
The RNN transducer is a promising end-to-end model candidate. We compare the original training criterion with the full marginalization over all alignments, to the commonly used max…