activity
20202022
most citedExploiting Pre-Trained ASR Models for Alzheimer's Disease Recognition Through Spontaneous Speech

10 citations · 14 across the 6 of their papers we have counts for

collaborators

6 papers

eess.AS20221 cited

Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification

Jingyu Li, Yusheng Tian, Tan Lee

Deep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism…

eess.AS20221 cited

An Investigation on Applying Acoustic Feature Conversion to ASR of Adult and Child Speech

Wei Liu, Jingyu Li, Tan Lee

The performance of child speech recognition is generally less satisfactory compared to adult speech due to limited amount of training data. Significant performance degradation is e…

eess.AS202110 cited

Exploiting Pre-Trained ASR Models for Alzheimer's Disease Recognition Through Spontaneous Speech

Ying Qin, Wei Liu, Zhiyuan Peng +4

Alzheimer's disease (AD) is a progressive neurodegenerative disease and recently attracts extensive attention worldwide. Speech technology is considered a promising solution for th…

eess.AS2021

Improving Text-Independent Speaker Verification with Auxiliary Speakers Using Graph

Jingyu Li, Si-Ioi Ng, Tan Lee

The paper presents a novel approach to refining similarity scores between input utterances for robust speaker verification. Given the embeddings from a pair of input utterances, a…

eess.AS2021

Detection of Consonant Errors in Disordered Speech Based on Consonant-vowel Segment Embedding

Si-Ioi Ng, Cymie Wing-Yee Ng, Jingyu Li +1

Speech sound disorder (SSD) refers to a type of developmental disorder in young children who encounter persistent difficulties in producing certain speech sounds at the expected ag…

eess.AS20202 cited

Text-Independent Speaker Verification with Dual Attention Network

Jingyu Li, Tan Lee

This paper presents a novel design of attention model for text-independent speaker verification. The model takes a pair of input utterances and generates an utterance-level embeddi…