collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2024

DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech

Jan Melechovsky, Ambuj Mehrish, Berrak Sisman +1

Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by…

eess.AS2024

Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training

Jan Melechovsky, Ambuj Mehrish, Berrak Sisman +1

With rapid globalization, the need to build inclusive and representative speech technology cannot be overstated. Accent is an important aspect of speech that needs to be taken into…

eess.AS2024

Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder

Jan Melechovsky, Ambuj Mehrish, Berrak Sisman +1

Accent plays a significant role in speech communication, influencing one's capability to understand as well as conveying a person's identity. This paper introduces a novel and effi…

eess.AS2024

MidiCaps: A large-scale MIDI dataset with text captions

Jan Melechovsky, Abhinaba Roy, Dorien Herremans

Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist due to the lack of a captioned MIDI dataset. This w…

eess.AS2024

Mustango: Toward Controllable Text-to-Music Generation

Jan Melechovsky, Zixun Guo, Deepanway Ghosal +3

The quality of the text-to-music models has reached new heights due to recent advancements in diffusion models. The controllability of various musical aspects, however, has barely…