9 papers
Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning
Zezhong Jin, Xiaoyu Wang, Zhe Li +4
Residual connections are fundamental to deep speaker recogni- tion models, such as ECAPA-TDNN and ResNet. However, standard identity mapping limits information flow to a sin- gle p…
EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation
Zilong Huang, Kong Aik Lee, Junjie Li +2
Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC ofte…
EII-SCL: Harnessing Emotional Inertia for Multimodal Emotion Recognition in Conversation
Zilong Huang, Kong Aik Lee, Chong-Xin Gan +3
Multimodal emotion recognition in conversation (MERC) achieves accurate predictions by integrating multimodal and contextual information in dialogues. While current MERC approaches…
Text-Independent Speaker Verification Using Discrete Audio Tokens
Zheng Liang, Junjie Li, Kong Aik Lee
Neural audio codecs (NACs) enable efficient audio compression and have achieved success in downstream tasks such as speech synthesis. However, their discrete representations consis…
Towards Robust Uncertainty-Aware Speaker Modeling
Junjie Li, Yang Xiao, Kong Aik Lee
Speaker embeddings aggregate frame-level acoustic features into compact representations for speaker recognition. Recent uncertainty-aware speaker modeling approaches further charac…
UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition
Chong-Xin Gan, Peter Bell, Man-Wai Mak +4
The joint training of speech enhancement and speaker embedding networks for speaker recognition is widely adopted under noisy acoustic environments. While effective, this paradigm…