5 papers · 1 filter
EmoSURA: Towards Accurate Evaluation of Detailed and Long-Context Emotional Speech Captions
Xin Jing, Andreas Triantafyllopoulos, Jiadong Wang +3
Recent advancements in speech captioning models have enabled the generation of rich, fine-grained captions for emotional speech. However, the evaluation of such captions remains a…
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
Xin Jing, Jiadong Wang, Andreas Triantafyllopoulos +4
The ambiguity of human emotions poses several challenges for machine learning models, as they often overlap and lack clear delineating boundaries. Contrastive language-audio pretra…
Audio-based Kinship Verification Using Age Domain Conversion
Qiyang Sun, Alican Akman, Xin Jing +2
Audio-based kinship verification (AKV) is important in many domains, such as home security monitoring, forensic identification, and social network analysis. A key challenge in the…
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
Xin Jing, Kun Zhou, Andreas Triantafyllopoulos +1
While current emotional text-to-speech (TTS) systems can generate highly intelligible emotional speech, achieving fine control over emotion rendering of the output speech still rem…
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
Xin Jing, Andreas Triantafyllopoulos, Björn Schuller
Contrastive language-audio pretraining (CLAP) has recently emerged as a method for making audio analysis more generalisable. Specifically, CLAP-style models are able to `answer' a…