most citedHigh Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent Latency

25 citations · 53 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD2021★ 6 cited

Prosodic Clustering for Phoneme-level Prosody Control in End-to-End Speech Synthesis

Alexandra Vioni, Myrsini Christidou, Nikolaos Ellinas +7

This paper presents a method for controlling the prosody at the phoneme level in an autoregressive attention-based text-to-speech system. Instead of learning latent prosodic featur…

cs.SD2021★ 5 cited

Word-Level Style Control for Expressive, Non-attentive Speech Synthesis

Konstantinos Klapsas, Nikolaos Ellinas, June Sig Sung +2

This paper presents an expressive speech synthesis architecture for modeling and controlling the speaking style at a word level. It attempts to learn word-level stylistic and proso…

cs.SD2021★ 2 cited

Improved Prosodic Clustering for Multispeaker and Speaker-independent Phoneme-level Prosody Control

Myrsini Christidou, Alexandra Vioni, Nikolaos Ellinas +7

This paper presents a method for phoneme-level prosody control of F0 and duration on a multispeaker text-to-speech setup, which is based on prosodic clustering. An autoregressive a…

cs.SD2021★ 2 cited

Rapping-Singing Voice Synthesis based on Phoneme-level Prosody Control

Konstantinos Markopoulos, Nikolaos Ellinas, Alexandra Vioni +8

In this paper, a text-to-rapping/singing system is introduced, which can be adapted to any speaker's voice. It utilizes a Tacotron-based multispeaker acoustic model trained on read…

cs.SD2021★ 13 cited

Cross-lingual Low Resource Speaker Adaptation Using Phonological Features

Georgia Maniati, Nikolaos Ellinas, Konstantinos Markopoulos +5

The idea of using phonological features instead of phonemes as input to sequence-to-sequence TTS has been recently proposed for zero-shot multilingual speech synthesis. This approa…

cs.SD2021★ 25 cited

High Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent Latency

Nikolaos Ellinas, Georgios Vamvoukakis, Konstantinos Markopoulos +7

This paper presents an end-to-end text-to-speech system with low latency on a CPU, suitable for real-time applications. The system is composed of an autoregressive attention-based…