activity
20172022
most citedData augmentation using prosody and false starts to recognize non-native children's speech

15 citations · 21 across the 6 of their papers we have counts for

collaborators

9 papers

eess.AS20221 cited

End-to-end Ensemble-based Feature Selection for Paralinguistics Tasks

Tamás Grósz, Mittul Singh, Sudarsana Reddy Kadiri +2

The events of recent years have highlighted the importance of telemedicine solutions which could potentially allow remote treatment and diagnosis. Relatedly, Computational Paraling…

eess.AS202015 cited

Data augmentation using prosody and false starts to recognize non-native children's speech

Hemant Kathania, Mittul Singh, Tamás Grósz +1

This paper describes AaltoASR's speech recognition system for the INTERSPEECH 2020 shared task on Automatic Speech Recognition (ASR) for non-native children's speech. The task is t…

cs.CL2020

FinChat: Corpus and evaluation setup for Finnish chat conversations on everyday topics

Katri Leino, Juho Leinonen, Mittul Singh +2

Creating open-domain chatbots requires large amounts of conversational data and related benchmark tasks to evaluate them. Standardized evaluation tasks are crucial for creating aut…

eess.AS2020

Aalto's End-to-End DNN systems for the INTERSPEECH 2020 Computational Paralinguistics Challenge

Tamás Grósz, Mittul Singh, Sudarsana Reddy Kadiri +2

End-to-end neural network models (E2E) have shown significant performance benefits on different INTERSPEECH ComParE tasks. Prior work has applied either a single instance of an E2E…

cs.CL2020

Effects of Language Relatedness for Cross-lingual Transfer Learning in Character-Based Language Models

Mittul Singh, Peter Smit, Sami Virpioja +1

Character-based Neural Network Language Models (NNLM) have the advantage of smaller vocabulary and thus faster training times in comparison to NNLMs based on multi-character units.…

cs.CL20205 cited

Subword RNNLM Approximations for Out-Of-Vocabulary Keyword Search

Mittul Singh, Sami Virpioja, Peter Smit +1

In spoken Keyword Search, the query may contain out-of-vocabulary (OOV) words not observed when training the speech recognition system. Using subword language models (LMs) in the f…