collaborators

9 papers

cs.MM2026

EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation

Zilong Huang, Kong Aik Lee, Junjie Li +2

Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC ofte…

eess.AS2026

Text-Independent Speaker Verification Using Discrete Audio Tokens

Zheng Liang, Junjie Li, Kong Aik Lee

Neural audio codecs (NACs) enable efficient audio compression and have achieved success in downstream tasks such as speech synthesis. However, their discrete representations consis…

cs.SD2026

Towards Robust Uncertainty-Aware Speaker Modeling

Junjie Li, Yang Xiao, Kong Aik Lee

Speaker embeddings aggregate frame-level acoustic features into compact representations for speaker recognition. Recent uncertainty-aware speaker modeling approaches further charac…

eess.AS2026

SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification

Junyi Peng, Oldřich Plchot, Xiao Song +9

Modern speaker verification (SV) systems rely on speaker embeddings that are effective but difficult to interpret or query in natural language. Most existing speech-text corpora ta…

cs.SD2026

QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection

Duc-Tuan Truong, Tianchi Liu, Ruijie Tao +3

Recent work shows that one-class learning can detect unseen deepfake attacks by modeling a compact distribution of bona fide speech around a single centroid. However, the single-ce…

eess.AS2026

TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs

Jing Peng, Chenghao Wang, Yi Yang +5

Speech LLM post-training increasingly relies on efficient cross-modal alignment and robust low-resource adaptation, yet collecting large-scale audio-text pairs remains costly. Text…