5 papers
Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning
Zezhong Jin, Xiaoyu Wang, Zhe Li +4
Residual connections are fundamental to deep speaker recogni- tion models, such as ECAPA-TDNN and ResNet. However, standard identity mapping limits information flow to a sin- gle p…
EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation
Zilong Huang, Kong Aik Lee, Junjie Li +2
Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC ofte…
UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition
Chong-Xin Gan, Peter Bell, Man-Wai Mak +4
The joint training of speech enhancement and speaker embedding networks for speaker recognition is widely adopted under noisy acoustic environments. While effective, this paradigm…
Spectral-Aware Low-Rank Adaptation for Speaker Verification
Zhe Li, Man-wai Mak, Mert Pilanci +2
Previous research has shown that the principal singular vectors of a pre-trained model's weight matrices capture critical knowledge. In contrast, those associated with small singul…
Subtractive Training for Music Stem Insertion using Latent Diffusion Models
Ivan Villa-Renteria, Mason L. Wang, Zachary Shah +4
We present Subtractive Training, a simple and novel method for synthesizing individual musical instrument stems given other instruments as context. This method pairs a dataset of c…