collaborators

6 papers

cs.SD2026

Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning

Zezhong Jin, Xiaoyu Wang, Zhe Li +4

Residual connections are fundamental to deep speaker recogni- tion models, such as ECAPA-TDNN and ResNet. However, standard identity mapping limits information flow to a sin- gle p…

cs.MM2026

EII-SCL: Harnessing Emotional Inertia for Multimodal Emotion Recognition in Conversation

Zilong Huang, Kong Aik Lee, Chong-Xin Gan +3

Multimodal emotion recognition in conversation (MERC) achieves accurate predictions by integrating multimodal and contextual information in dialogues. While current MERC approaches…

cs.SD2026

LISE : Listenable Interpretable Speaker Embeddings

Xiaoliang Wu, Chongxin Gan, Ke Liu +2

Deep neural network-based automatic speaker verification (ASV) systems achieve impressive performance but their embedding representations remain opaque, lacking a structured and pe…

cs.SD2026

Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training

Yanru Wu, Jianning Wang, Chongxin Gan +1

Training general-purpose Audio Large Language Models (ALLMs) across diverse datasets is essential for holistic audio understanding, yet it faces significant challenges due to datas…

eess.AS2026

UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition

Chong-Xin Gan, Peter Bell, Man-Wai Mak +4

The joint training of speech enhancement and speaker embedding networks for speaker recognition is widely adopted under noisy acoustic environments. While effective, this paradigm…

cs.CL2025

TrInk: Ink Generation with Transformer Network

Zezhong Jin, Shubhang Desai, Xu Chen +8

In this paper, we propose TrInk, a Transformer-based model for ink generation, which effectively captures global dependencies. To better facilitate the alignment between the input…