3 papers
cs.LG2025
VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks
Daria Diatlova, Nikita Balagansky, Alexander Varlamov +1
Conventional methods for aggregating layers in fine-tuned self-supervised speech models, such as using the final layer or weighted sum, suffer from information bottlenecks and stat…
cs.LG2025
NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
Maksim Borisov, Egor Spirin, Daria Diatlova
Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we in…
cs.LG2024
Adapting WavLM for Speech Emotion Recognition
Daria Diatlova, Anton Udalov, Vitalii Shutov +1
Recently, the usage of speech self-supervised models (SSL) for downstream tasks has been drawing a lot of attention. While large pre-trained models commonly outperform smaller mode…