speech processing

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants

arXiv:2607.14310 · doi:10.57967/hf/9194

summary

The paper presents Dialogs, a high-quality Russian conversational speech corpus recorded in a studio, featuring expressive prosody, turn‑taking, and emotion/style annotations, and demonstrates its usefulness for training expressive dialogue TTS models.

Abstract

We introduce Dialogs, a studio-quality Russian conversational speech corpus for dialog assistants. The dataset contains 20.6 hours of face-to-face acted dialogs recorded in a professional studio (44.1 kHz stereo) and segmented into 11,796 utterances across 3 speakers. Unlike read-speech resources, Dialogs captures turn-taking rhythm and expressive prosody, and provides per-utterance style/emotion labels spanning 12 categories. We validate corpus quality with crowd MOS tests, showing comparable audio quality and intelligibility to strong Russian studio baselines while achieving higher ratings for expressiveness and conversational naturalness. Finally, we train a VITS2 model as a proof of concept, demonstrating that Dialogs supports training expressive, dialog-like TTS despite limited per-speaker data.

4 pages, 1 figure, 5 tables. Interspeech 2026

Topics & keywords

#speech corpus#russian language#dialogue systems#expressive tts#prosodystudio-quality recordings44.1 kHz stereostyle/emotion labelsVITS2 modelturn-taking rhythm
Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants · wovepaper