activity
20202022
most citedSpeaking Speed Control of End-to-End Speech Synthesis using Sentence-Level Conditioning

1 citations · 2 across the 6 of their papers we have counts for

collaborators

7 papers

eess.AS2022

An Empirical Study on L2 Accents of Cross-lingual Text-to-Speech Systems via Vowel Space

Jihwan Lee, Jae-Sung Bae, Seongkyu Mun +4

With the recent developments in cross-lingual Text-to-Speech (TTS) systems, L2 (second-language, or foreign) accent problems arise. Moreover, running a subjective evaluation for su…

eess.AS20211 cited

GANSpeech: Adversarial Training for High-Fidelity Multi-Speaker Speech Synthesis

Jinhyeok Yang, Jae-Sung Bae, Taejun Bak +2

Recent advances in neural multi-speaker text-to-speech (TTS) models have enabled the generation of reasonably good speech quality with a single model and made it possible to synthe…

eess.AS2021

Hierarchical Context-Aware Transformers for Non-Autoregressive Text to Speech

Jae-Sung Bae, Tae-Jun Bak, Young-Sun Joo +1

In this paper, we propose methods for improving the modeling performance of a Transformer-based non-autoregressive text-to-speech (TNA-TTS) model. Although the text encoder and aud…

eess.AS2021

FastPitchFormant: Source-filter based Decomposed Modeling for Speech Synthesis

Taejun Bak, Jae-Sung Bae, Hanbin Bae +2

Methods for modeling and controlling prosody with acoustic features have been proposed for neural text-to-speech (TTS) models. Prosodic speech can be generated by conditioning acou…

eess.AS2021

A Neural Text-to-Speech Model Utilizing Broadcast Data Mixed with Background Music

Hanbin Bae, Jae-Sung Bae, Young-Sun Joo +2

Recently, it has become easier to obtain speech data from various media such as the internet or YouTube, but directly utilizing them to train a neural text-to-speech (TTS) model is…

eess.AS20201 cited

Speaking Speed Control of End-to-End Speech Synthesis using Sentence-Level Conditioning

Jae-Sung Bae, Hanbin Bae, Young-Sun Joo +3

This paper proposes a controllable end-to-end text-to-speech (TTS) system to control the speaking speed (speed-controllable TTS; SCTTS) of synthesized speech with sentence-level sp…