activity
20182024
most citedKazakhTTS: An Open-Source Kazakh Text-to-Speech Synthesis Dataset

30 citations · 38 across the 13 of their papers we have counts for

collaborators

19 papers

eess.AS2024

Dual-Pipeline with Low-Rank Adaptation for New Language Integration in Multilingual ASR

Yerbolat Khassanov, Zhipeng Chen, Tianfeng Chen +5

This paper addresses challenges in integrating new languages into a pre-trained multilingual automatic speech recognition (mASR) system, particularly in scenarios where training da…

eess.AS2023

Multilingual Text-to-Speech Synthesis for Turkic Languages Using Transliteration

Rustem Yeshpanov, Saida Mussakhojayeva, Yerbolat Khassanov

This work aims to build a multilingual text-to-speech (TTS) synthesis system for ten lower-resourced Turkic languages: Azerbaijani, Bashkir, Kazakh, Kyrgyz, Sakha, Tatar, Turkish,…

eess.AS2022

Random Utterance Concatenation Based Data Augmentation for Improving Short-video Speech Recognition

Yist Y. Lin, Tao Han, Haihua Xu +6

One of limitations in end-to-end automatic speech recognition (ASR) framework is its performance would be compromised if train-test utterance lengths are mismatched. In this paper,…

eess.AS2022

KazakhTTS2: Extending the Open-Source Kazakh TTS Corpus With More Data, Speakers, and Topics

Saida Mussakhojayeva, Yerbolat Khassanov, Huseyin Atakan Varol

We present an expanded version of our previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In the new KazakhTTS2 corpus, the overall size has increased from 93 h…

cs.CL2021

KazNERD: Kazakh Named Entity Recognition Dataset

Rustem Yeshpanov, Yerbolat Khassanov, Huseyin Atakan Varol

We present the development of a dataset for Kazakh named entity recognition. The dataset was built as there is a clear need for publicly available annotated corpora in Kazakh, as w…

cs.CV2021

A Study of Multimodal Person Verification Using Audio-Visual-Thermal Data

Madina Abdrakhmanova, Saniya Abushakimova, Yerbolat Khassanov +1

In this paper, we study an approach to multimodal person verification using audio, visual, and thermal modalities. The combination of audio and visual modalities has already been s…