6 citations · 6 across the 4 of their papers we have counts for
5 papers
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
Nikolaos Ellinas, Alexandra Vioni, Panos Kakoulidis +8
This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-base…
MambaRate: Speech Quality Assessment Across Different Sampling Rates
Panos Kakoulidis, Iakovi Alexiou, Junkwang Oh +4
We propose MambaRate, which predicts Mean Opinion Scores (MOS) with limited bias regarding the sampling rate of the waveform under evaluation. It is designed for Track 3 of the Aud…
Controllable speech synthesis by learning discrete phoneme-level prosodic representations
Nikolaos Ellinas, Myrsini Christidou, Alexandra Vioni +4
In this paper, we present a novel method for phoneme-level prosody control of F0 and duration using intuitive discrete labels. We propose an unsupervised prosodic clustering proces…
Predicting phoneme-level prosody latents using AR and flow-based Prior Networks for expressive speech synthesis
Konstantinos Klapsas, Karolos Nikitaras, Nikolaos Ellinas +5
A large part of the expressive speech synthesis literature focuses on learning prosodic representations of the speech signal which are then modeled by a prior distribution during i…
Learning utterance-level representations through token-level acoustic latents prediction for Expressive Speech Synthesis
Karolos Nikitaras, Konstantinos Klapsas, Nikolaos Ellinas +6
This paper proposes an Expressive Speech Synthesis model that utilizes token-level latent prosodic variables in order to capture and control utterance-level attributes, such as cha…