activity
20222025
collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2025

Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders

Nikolaos Ellinas, Alexandra Vioni, Panos Kakoulidis +8

This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-base…

cs.SD2025

MambaRate: Speech Quality Assessment Across Different Sampling Rates

Panos Kakoulidis, Iakovi Alexiou, Junkwang Oh +4

We propose MambaRate, which predicts Mean Opinion Scores (MOS) with limited bias regarding the sampling rate of the waveform under evaluation. It is designed for Track 3 of the Aud…

cs.SD2024

Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling

Sotirios Karapiperis, Nikolaos Ellinas, Alexandra Vioni +4

Most of the prevalent approaches in speech prosody modeling rely on learning global style representations in a continuous latent space which encode and transfer the attributes of r…

cs.SD2024

Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations

Panos Kakoulidis, Nikolaos Ellinas, Georgios Vamvoukakis +8

In this paper, we propose a singing voice synthesis model, Karaoker-SSL, that is trained only on text and speech data as a typical multi-speaker acoustic model. It is a low-resourc…

cs.SD2022

Predicting phoneme-level prosody latents using AR and flow-based Prior Networks for expressive speech synthesis

Konstantinos Klapsas, Karolos Nikitaras, Nikolaos Ellinas +5

A large part of the expressive speech synthesis literature focuses on learning prosodic representations of the speech signal which are then modeled by a prior distribution during i…

cs.SD2022

Learning utterance-level representations through token-level acoustic latents prediction for Expressive Speech Synthesis

Karolos Nikitaras, Konstantinos Klapsas, Nikolaos Ellinas +6

This paper proposes an Expressive Speech Synthesis model that utilizes token-level latent prosodic variables in order to capture and control utterance-level attributes, such as cha…