activity
20192022
most citedIntegrating the Data Augmentation Scheme with Various Classifiers for Acoustic Scene Modeling

67 citations · 104 across the 19 of their papers we have counts for

collaborators

20 papers

eess.AS20225 cited

Deepfake Detection System for the ADD Challenge Track 3.2 Based on Score Fusion

Yuxiang Zhang, Jingze Lu, Xingming Wang +5

This paper describes the deepfake audio detection system submitted to the Audio Deep Synthesis Detection (ADD) Challenge Track 3.2 and gives an analysis of score fusion. The propos…

cs.CL2022

Summary on the ISCSLP 2022 Chinese-English Code-Switching ASR Challenge

Shuhao Deng, Chengfei Li, Jinfeng Bai +6

Code-switching automatic speech recognition becomes one of the most challenging and the most valuable scenarios of automatic speech recognition, due to the code-switching phenomeno…

cs.CV20221 cited

Audio-Visual Scene Classification Using A Transfer Learning Based Joint Optimization Strategy

Chengxin Chen, Meng Wang, Pengyuan Zhang

Recently, audio-visual scene classification (AVSC) has attracted increasing attention from multidisciplinary communities. Previous studies tended to adopt a pipeline training strat…

cs.SD2022

Back-ends Selection for Deep Speaker Embeddings

Zhuo Li, Runqiu Xiao, Zihan Zhang +3

Probabilistic Linear Discriminant Analysis (PLDA) was the dominant and necessary back-end for early speaker recognition approaches, like i-vector and x-vector. However, with the de…

cs.SD20222 cited

CTA-RNN: Channel and Temporal-wise Attention RNN Leveraging Pre-trained ASR Embeddings for Speech Emotion Recognition

Chengxin Chen, Pengyuan Zhang

Previous research has looked into ways to improve speech emotion recognition (SER) by utilizing both acoustic and linguistic cues of speech. However, the potential association betw…

cs.CL2022

Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset

Zehui Yang, Yifan Chen, Lei Luo +9

This paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC. The MagicData-RAMC corpus contains 180 hours of conversatio…