4 papers
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
Nikolaos Ellinas, Alexandra Vioni, Panos Kakoulidis +8
This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-base…
Improved Text Emotion Prediction Using Combined Valence and Arousal Ordinal Classification
Michael Mitsios, Georgios Vamvoukakis, Georgia Maniati +13
Emotion detection in textual data has received growing interest in recent years, as it is pivotal for developing empathetic human-computer interaction systems. This paper introduce…
Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation
Nikolaos Ellinas, Georgios Vamvoukakis, Konstantinos Markopoulos +7
This paper presents a method for end-to-end cross-lingual text-to-speech (TTS) which aims to preserve the target language's pronunciation regardless of the original speaker's langu…
Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations
Panos Kakoulidis, Nikolaos Ellinas, Georgios Vamvoukakis +8
In this paper, we propose a singing voice synthesis model, Karaoker-SSL, that is trained only on text and speech data as a typical multi-speaker acoustic model. It is a low-resourc…