53 citations · 82 across the 4 of their papers we have counts for
4 papers
Improving Low Compute Language Modeling with In-Domain Embedding Initialisation
Charles Welch, Rada Mihalcea, Jonathan K. Kummerfeld
Many NLP applications, such as biomedical data and technical support, have 10-100 million tokens of in-domain data and limited computational resources for learning from it. How sho…
The Eighth Dialog System Technology Challenge
Seokhwan Kim, Michel Galley, Chulaka Gunasekara +18
This paper introduces the Eighth Dialog System Technology Challenge. In line with recent challenges, the eighth edition focuses on applying end-to-end dialog technologies in a prag…
Outlier Detection for Improved Data Quality and Diversity in Dialog Systems
Stefan Larson, Anish Mahendran, Andrew Lee +6
In a corpus of data, outliers are either errors: mistakes in the data that are counterproductive, or are unique: informative samples that improve model robustness. Identifying outl…
Dialog System Technology Challenge 7
Koichiro Yoshino, Chiori Hori, Julien Perez +14
This paper introduces the Seventh Dialog System Technology Challenges (DSTC), which use shared datasets to explore the problem of building dialog systems. Recently, end-to-end dial…