activity
20222025
most citedControllable speech synthesis by learning discrete phoneme-level prosodic representations

6 citations · 8 across the 7 of their papers we have counts for

collaborators

7 papers

cs.SD2025

Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders

Nikolaos Ellinas, Alexandra Vioni, Panos Kakoulidis +8

This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-base…

cs.SD2024

Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling

Sotirios Karapiperis, Nikolaos Ellinas, Alexandra Vioni +4

Most of the prevalent approaches in speech prosody modeling rely on learning global style representations in a continuous latent space which encode and transfer the attributes of r…

cs.LG20242 cited

Improved Text Emotion Prediction Using Combined Valence and Arousal Ordinal Classification

Michael Mitsios, Georgios Vamvoukakis, Georgia Maniati +13

Emotion detection in textual data has received growing interest in recent years, as it is pivotal for developing empathetic human-computer interaction systems. This paper introduce…

cs.SD2024

Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations

Panos Kakoulidis, Nikolaos Ellinas, Georgios Vamvoukakis +8

In this paper, we propose a singing voice synthesis model, Karaoker-SSL, that is trained only on text and speech data as a typical multi-speaker acoustic model. It is a low-resourc…

cs.SD20226 cited

Controllable speech synthesis by learning discrete phoneme-level prosodic representations

Nikolaos Ellinas, Myrsini Christidou, Alexandra Vioni +4

In this paper, we present a novel method for phoneme-level prosody control of F0 and duration using intuitive discrete labels. We propose an unsupervised prosodic clustering proces…

cs.SD2022

Predicting phoneme-level prosody latents using AR and flow-based Prior Networks for expressive speech synthesis

Konstantinos Klapsas, Karolos Nikitaras, Nikolaos Ellinas +5

A large part of the expressive speech synthesis literature focuses on learning prosodic representations of the speech signal which are then modeled by a prior distribution during i…