29 citations · 32 across the 6 of their papers we have counts for
6 papers · 1 filter
Mask More and Mask Later: Efficient Pre-training of Masked Language Models by Disentangling the [MASK] Token
Baohao Liao, David Thulke, Sanjika Hewavitharana +2
The pre-training of masked language models (MLMs) consumes massive computation to achieve good results on downstream NLP tasks, resulting in a large carbon footprint. In the vanill…
Controllable Factuality in Document-Grounded Dialog Systems Using a Noisy Channel Model
Nico Daheim, David Thulke, Christian Dugast +1
In this work, we present a model for document-grounded response generation in dialog that is decomposed into two components according to Bayes theorem. One component is a tradition…
Investigation on Data Adaptation Techniques for Neural Named Entity Recognition
Evgeniia Tokarchuk, David Thulke, Weiyue Wang +2
Data processing is an important step in various natural language processing tasks. As the commonly used datasets in named entity recognition contain only a limited number of sample…
Cascaded Span Extraction and Response Generation for Document-Grounded Dialog
Nico Daheim, David Thulke, Christian Dugast +1
This paper summarizes our entries to both subtasks of the first DialDoc shared task which focuses on the agent response prediction task in goal-oriented document-grounded dialogs.…
On Sampling-Based Training Criteria for Neural Language Modeling
Yingbo Gao, David Thulke, Alexander Gerstenberger +3
As the vocabulary size of modern word-based language models becomes ever larger, many sampling-based training criteria are proposed and investigated. The essence of these sampling…
Efficient Retrieval Augmented Generation from Unstructured Knowledge for Task-Oriented Dialog
David Thulke, Nico Daheim, Christian Dugast +1
This paper summarizes our work on the first track of the ninth Dialog System Technology Challenge (DSTC 9), "Beyond Domain APIs: Task-oriented Conversational Modeling with Unstruct…