3 papers
cs.SD2024
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
Xiaohui Zhang, Wenjie Fu, Mangui Liang
Speech Emotion Recognition (SER) is still a complex task for computers with average recall rates usually about 70% on the most realistic datasets. Most SER systems use hand-crafted…
cs.SD2024
Soft-Weighted CrossEntropy Loss for Continous Alzheimer's Disease Detection
Xiaohui Zhang, Wenjie Fu, Mangui Liang
Alzheimer's disease is a common cognitive disorder in the elderly. Early and accurate diagnosis of Alzheimer's disease (AD) has a major impact on the progress of research on dement…
cs.SD2023
TST: Time-Sparse Transducer for Automatic Speech Recognition
Xiaohui Zhang, Mangui Liang, Zhengkun Tian +2
End-to-end model, especially Recurrent Neural Network Transducer (RNN-T), has achieved great success in speech recognition. However, transducer requires a great memory footprint an…