activity
20172023
most citedPretraining Approaches for Spoken Language Recognition: TalTech Submission to the OLR 2021 Challenge

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL2023

Dialect Adaptation and Data Augmentation for Low-Resource ASR: TalTech Systems for the MADASR 2023 Challenge

Tanel Alumäe, Jiaming Kong, Daniil Robnikov

This paper describes Tallinn University of Technology (TalTech) systems developed for the ASRU MADASR 2023 Challenge. The challenge focuses on automatic speech recognition of diale…

eess.AS2022

Collar-aware Training for Streaming Speaker Change Detection in Broadcast Speech

Joonas Kalda, Tanel Alumäe

In this paper, we present a novel training method for speaker change detection models. Speaker change detection is often viewed as a binary sequence labelling problem. The main cha…

eess.AS20221 cited

Pretraining Approaches for Spoken Language Recognition: TalTech Submission to the OLR 2021 Challenge

Tanel Alumäe, Kunnar Kukk

This paper investigates different pretraining approaches to spoken language identification. The paper is based on our submission to the Oriental Language Recognition 2021 Challenge…

eess.AS2020

VoxLingua107: a Dataset for Spoken Language Recognition

Jörgen Valk, Tanel Alumäe

This paper investigates the use of automatically collected web audio data for the task of spoken language recognition. We generate semi-random search phrases from language-specific…

cs.SD2018

Weakly Supervised Training of Speaker Identification Models

Martin Karu, Tanel Alumäe

We propose an approach for training speaker identification models in a weakly supervised manner. We concentrate on the setting where the training data consists of a set of audio re…

cs.CL2017

Low-Resource Neural Headline Generation

Ottokar Tilk, Tanel Alumäe

Recent neural headline generation models have shown great results, but are generally trained on very large datasets. We focus our efforts on improving headline quality on smaller d…