5 papers
100,000+ Movie Reviews from Kazakhstan: Russian, Kazakh, and Code-Switched Texts
Rustem Yeshpanov
We present a new publicly available corpus of 100,502 movie reviews from Kazakhstan collected from kino.kz, spanning 2001-2025 and covering 4,943 unique titles. The dataset is mult…
Using Songs to Improve Kazakh Automatic Speech Recognition
Rustem Yeshpanov
Developing automatic speech recognition (ASR) systems for low-resource languages is hindered by the scarcity of transcribed corpora. This proof-of-concept study explores songs as a…
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
Adal Abilbekov, Saida Mussakhojayeva, Rustem Yeshpanov +1
This study focuses on the creation of the KazEmoTTS dataset, designed for emotional Kazakh text-to-speech (TTS) applications. KazEmoTTS is a collection of 54,760 audio-text pairs,…
KazParC: Kazakh Parallel Corpus for Machine Translation
Rustem Yeshpanov, Alina Polonskaya, Huseyin Atakan Varol
We introduce KazParC, a parallel corpus designed for machine translation across Kazakh, English, Russian, and Turkish. The first and largest publicly available corpus of its kind,…
KazSAnDRA: Kazakh Sentiment Analysis Dataset of Reviews and Attitudes
Rustem Yeshpanov, Huseyin Atakan Varol
This paper presents KazSAnDRA, a dataset developed for Kazakh sentiment analysis that is the first and largest publicly available dataset of its kind. KazSAnDRA comprises an extens…