3 papers
cs.SD2025
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
Shangkun Huang, Yuxuan Du, Jingwen Yang +5
This paper presents the system developed to address the MISP 2025 Challenge. For the diarization system, we proposed a hybrid approach combining a WavLM end-to-end segmentation met…
cs.SD2025
Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection
Shangkun Huang, Jing Deng, Jintao Kang +1
The performance bottleneck of Automatic Speech Recognition (ASR) in stuttering speech scenarios has limited its applicability in domains such as speech rehabilitation. This paper p…
eess.AS2025
SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker Recognition
Rongjin Li, Weibin Zhang, Dongpeng Chen +2
In conventional deep speaker embedding frameworks, the pooling layer aggregates all frame-level features over time and computes their mean and standard deviation statistics as inpu…