activity
20172023
most citedAutomatic Speech Recognition with Very Large Conversational Finnish and Estonian Vocabularies

33 citations · 41 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2023

MorphPiece : A Linguistic Tokenizer for Large Language Models

Haris Jabbar

Tokenization is a critical part of modern NLP pipelines. However, contemporary tokenizers for Large Language Models are based on statistical analysis of text corpora, without much…

cs.CL20233 cited

Uncertainty-Aware Natural Language Inference with Stochastic Weight Averaging

Aarne Talman, Hande Celikkanat, Sami Virpioja +2

This paper introduces Bayesian uncertainty modeling using Stochastic Weight Averaging-Gaussian (SWAG) in Natural Language Understanding (NLU) tasks. We apply the approach to standa…

cs.CL2020

FinChat: Corpus and evaluation setup for Finnish chat conversations on everyday topics

Katri Leino, Juho Leinonen, Mittul Singh +2

Creating open-domain chatbots requires large amounts of conversational data and related benchmark tasks to evaluate them. Standardized evaluation tasks are crucial for creating aut…

cs.CL2020

Effects of Language Relatedness for Cross-lingual Transfer Learning in Character-Based Language Models

Mittul Singh, Peter Smit, Sami Virpioja +1

Character-based Neural Network Language Models (NNLM) have the advantage of smaller vocabulary and thus faster training times in comparison to NNLMs based on multi-character units.…

cs.CL20205 cited

Subword RNNLM Approximations for Out-Of-Vocabulary Keyword Search

Mittul Singh, Sami Virpioja, Peter Smit +1

In spoken Keyword Search, the query may contain out-of-vocabulary (OOV) words not observed when training the speech recognition system. Using subword language models (LMs) in the f…

cs.CL2020

Transfer learning and subword sampling for asymmetric-resource one-to-many neural translation

Stig-Arne Grönroos, Sami Virpioja, Mikko Kurimo

There are several approaches for improving neural machine translation for low-resource languages: Monolingual data can be exploited via pretraining or data augmentation; Parallel c…