33 citations · 41 across the 6 of their papers we have counts for
10 papers · 1 filter
MorphPiece : A Linguistic Tokenizer for Large Language Models
Haris Jabbar
Tokenization is a critical part of modern NLP pipelines. However, contemporary tokenizers for Large Language Models are based on statistical analysis of text corpora, without much…
Uncertainty-Aware Natural Language Inference with Stochastic Weight Averaging
Aarne Talman, Hande Celikkanat, Sami Virpioja +2
This paper introduces Bayesian uncertainty modeling using Stochastic Weight Averaging-Gaussian (SWAG) in Natural Language Understanding (NLU) tasks. We apply the approach to standa…
FinChat: Corpus and evaluation setup for Finnish chat conversations on everyday topics
Katri Leino, Juho Leinonen, Mittul Singh +2
Creating open-domain chatbots requires large amounts of conversational data and related benchmark tasks to evaluate them. Standardized evaluation tasks are crucial for creating aut…
Effects of Language Relatedness for Cross-lingual Transfer Learning in Character-Based Language Models
Mittul Singh, Peter Smit, Sami Virpioja +1
Character-based Neural Network Language Models (NNLM) have the advantage of smaller vocabulary and thus faster training times in comparison to NNLMs based on multi-character units.…
Subword RNNLM Approximations for Out-Of-Vocabulary Keyword Search
Mittul Singh, Sami Virpioja, Peter Smit +1
In spoken Keyword Search, the query may contain out-of-vocabulary (OOV) words not observed when training the speech recognition system. Using subword language models (LMs) in the f…
Transfer learning and subword sampling for asymmetric-resource one-to-many neural translation
Stig-Arne Grönroos, Sami Virpioja, Mikko Kurimo
There are several approaches for improving neural machine translation for low-resource languages: Monolingual data can be exploited via pretraining or data augmentation; Parallel c…