9 papers
EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation
Zilong Huang, Kong Aik Lee, Junjie Li +2
Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC ofte…
Text-Independent Speaker Verification Using Discrete Audio Tokens
Zheng Liang, Junjie Li, Kong Aik Lee
Neural audio codecs (NACs) enable efficient audio compression and have achieved success in downstream tasks such as speech synthesis. However, their discrete representations consis…
Towards Robust Uncertainty-Aware Speaker Modeling
Junjie Li, Yang Xiao, Kong Aik Lee
Speaker embeddings aggregate frame-level acoustic features into compact representations for speaker recognition. Recent uncertainty-aware speaker modeling approaches further charac…
SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification
Junyi Peng, OldÅich Plchot, Xiao Song +9
Modern speaker verification (SV) systems rely on speaker embeddings that are effective but difficult to interpret or query in natural language. Most existing speech-text corpora ta…
QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection
Duc-Tuan Truong, Tianchi Liu, Ruijie Tao +3
Recent work shows that one-class learning can detect unseen deepfake attacks by modeling a compact distribution of bona fide speech around a single centroid. However, the single-ce…
TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs
Jing Peng, Chenghao Wang, Yi Yang +5
Speech LLM post-training increasingly relies on efficient cross-modal alignment and robust low-resource adaptation, yet collecting large-scale audio-text pairs remains costly. Text…