29 citations · 145 across the 24 of their papers we have counts for
63 papers
Does Joint Training Really Help Cascaded Speech Translation?
Viet Anh Khoa Tran, David Thulke, Yingbo Gao +2
Currently, in speech translation, the straightforward approach - cascading a recognition system with a translation system - delivers state-of-the-art results. However, fundamental…
Mask More and Mask Later: Efficient Pre-training of Masked Language Models by Disentangling the [MASK] Token
Baohao Liao, David Thulke, Sanjika Hewavitharana +2
The pre-training of masked language models (MLMs) consumes massive computation to achieve good results on downstream NLP tasks, resulting in a large carbon footprint. In the vanill…
Controllable Factuality in Document-Grounded Dialog Systems Using a Noisy Channel Model
Nico Daheim, David Thulke, Christian Dugast +1
In this work, we present a model for document-grounded response generation in dialog that is decomposed into two components according to Bayes theorem. One component is a tradition…
Monotonic segmental attention for automatic speech recognition
Albert Zeyer, Robin Schmitt, Wei Zhou +2
We introduce a novel segmental-attention model for automatic speech recognition. We restrict the decoder attention to segments to avoid quadratic runtime of global attention, bette…
Is Encoder-Decoder Redundant for Neural Machine Translation?
Yingbo Gao, Christian Herold, Zijian Yang +1
Encoder-decoder architecture is widely adopted for sequence-to-sequence modeling tasks. For machine translation, despite the evolution from long short-term memory networks to Trans…
Revisiting Checkpoint Averaging for Neural Machine Translation
Yingbo Gao, Christian Herold, Zijian Yang +1
Checkpoint averaging is a simple and effective method to boost the performance of converged neural machine translation models. The calculation is cheap to perform and the fact that…