5 papers
Voice Timbre Attribute Detection with Compact and Interpretable Training-Free Acoustic Parameters
Aemon Yat Fei Chiu, Yujia Xiao, Qiuqiang Kong +1
Voice timbre attribute detection (vTAD) is the task of determining the relative intensity of timbre attributes between speech utterances. Voice timbre is a crucial yet inherently c…
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
Aemon Yat Fei Chiu, Kei Ching Fung, Roger Tsz Yeung Li +2
Enhancing explainability in speech self-supervised learning (SSL) is important for developing reliable SSL-based speech processing systems. This study probes how speech SSL models…
CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025
Aemon Yat Fei Chiu, Jingyu Li, Yusheng Tian +2
This paper presents the Voice Timbre Attribute Detection (vTAD) systems developed by the Digital Signal Processing & Speech Technology Laboratory (DSP&STL) of the Department of Ele…
PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
Yujia Xiao, Liumeng Xue, Lei He +8
Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into…
An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems
Jingyu Li, Aemon Yat Fei Chiu, Tan Lee
Language mismatch is among the most common and challenging domain mismatches in deploying speaker verification (SV) systems. Adversarial reprogramming has shown promising results i…