activity
20172022
most citedAutomatic Speech Recognition with Very Large Conversational Finnish and Estonian Vocabularies

33 citations · 58 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

13 papers · 1 filter

cs.CL20222 cited

When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and its Intensity

Khalid Alnajjar, Mika Hämäläinen, Jörg Tiedemann +2

Prerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatical…

cs.CL20222 cited

Finnish Parliament ASR corpus - Analysis, benchmarks and statistics

Anja Virkkunen, Aku Rouhe, Nhan Phan +1

Public sources like parliament meeting recordings and transcripts provide ever-growing material for the training and evaluation of automatic speech recognition (ASR) systems. In th…

cs.CL2022

Lahjoita puhetta -- a large-scale corpus of spoken Finnish with some benchmarks

Anssi Moisio, Dejan Porjazovski, Aku Rouhe +5

The Donate Speech campaign has so far succeeded in gathering approximately 3600 hours of ordinary, colloquial Finnish speech into the Lahjoita puhetta (Donate Speech) corpus. The c…

cs.CL2020

FinChat: Corpus and evaluation setup for Finnish chat conversations on everyday topics

Katri Leino, Juho Leinonen, Mittul Singh +2

Creating open-domain chatbots requires large amounts of conversational data and related benchmark tasks to evaluate them. Standardized evaluation tasks are crucial for creating aut…

cs.CL2020

Effects of Language Relatedness for Cross-lingual Transfer Learning in Character-Based Language Models

Mittul Singh, Peter Smit, Sami Virpioja +1

Character-based Neural Network Language Models (NNLM) have the advantage of smaller vocabulary and thus faster training times in comparison to NNLMs based on multi-character units.…

cs.CL20205 cited

Subword RNNLM Approximations for Out-Of-Vocabulary Keyword Search

Mittul Singh, Sami Virpioja, Peter Smit +1

In spoken Keyword Search, the query may contain out-of-vocabulary (OOV) words not observed when training the speech recognition system. Using subword language models (LMs) in the f…