3 papers
eess.AS2026
Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
Lucas H. Ueda, João G. T. Lima, Paula D. P. Costa
Speech Emotion Recognition (SER) in real-world scenarios remains challenging due to severe class imbalance and the prevalence of spontaneous, natural speech. While recent approache…
eess.AS2026
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
Lucas H. Ueda, João G. T. Lima, Pedro R. Corrêa +3
This paper presents SelfTTS, a text-to-speech (TTS) model designed for cross-speaker style transfer that eliminates the need for external pre-trained speaker or emotion encoders. T…
eess.AS2024
Improving Data Augmentation-based Cross-Speaker Style Transfer for TTS with Singing Voice, Style Filtering, and F0 Matching
Leonardo B. de M. M. Marques, Lucas H. Ueda, Mário U. Neto +4
The goal of cross-speaker style transfer in TTS is to transfer a speech style from a source speaker with expressive data to a target speaker with only neutral data. In this context…