activity
20212025
most citedHigh Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent Latency

25 citations · 53 across the 11 of their papers we have counts for

collaborators

11 papers

cs.SD2025

Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders

Nikolaos Ellinas, Alexandra Vioni, Panos Kakoulidis +8

This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-base…

cs.LG2024★ 2 cited

Improved Text Emotion Prediction Using Combined Valence and Arousal Ordinal Classification

Michael Mitsios, Georgios Vamvoukakis, Georgia Maniati +13

Emotion detection in textual data has received growing interest in recent years, as it is pivotal for developing empathetic human-computer interaction systems. This paper introduce…

cs.SD2022

Generating Multilingual Gender-Ambiguous Text-to-Speech Voices

Konstantinos Markopoulos, Georgia Maniati, Georgios Vamvoukakis +9

The gender of any voice user interface is a key element of its perceived identity. Recently, there has been increasing interest in interfaces where the gender is ambiguous rather t…

cs.SD2022★ 1 cited

Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation

Nikolaos Ellinas, Georgios Vamvoukakis, Konstantinos Markopoulos +7

This paper presents a method for end-to-end cross-lingual text-to-speech (TTS) which aims to preserve the target language's pronunciation regardless of the original speaker's langu…

cs.SD2022★ 5 cited

Fine-grained Noise Control for Multispeaker Speech Synthesis

Karolos Nikitaras, Georgios Vamvoukakis, Nikolaos Ellinas +7

A text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model th…

eess.AS2022★ 3 cited

Karaoker: Alignment-free singing voice synthesis with speech training data

Panos Kakoulidis, Nikolaos Ellinas, Georgios Vamvoukakis +5

Existing singing voice synthesis models (SVS) are usually trained on singing data and depend on either error-prone time-alignment and duration features or explicit music score info…