Multi-Modal Detection of Alzheimer's Disease from Speech and Text
arXiv:2012.00096
Abstract
Reliable detection of the prodromal stages of Alzheimer's disease (AD) remains difficult even today because, unlike other neurocognitive impairments, there is no definitive diagnosis of AD in vivo. In this context, existing research has shown that patients often develop language impairment even in mild AD conditions. We propose a multimodal deep learning method that utilizes speech and the corresponding transcript simultaneously to detect AD. For audio signals, the proposed audio-based network, a convolutional neural network (CNN) based model, predicts the diagnosis for multiple speech segments, which are combined for the final prediction. Similarly, we use contextual embedding extracted from BERT concatenated with a CNN-generated embedding for classifying the transcript. The individual predictions of the two models are then combined to make the final classification. We also perform experiments to analyze the model performance when Automated Speech Recognition (ASR) system generated transcripts are used instead of manual transcription in the text-based model. The proposed method achieves 85.3% 10-fold cross-validation accuracy when trained and evaluated on the Dementiabank Pitt corpus.
9 pages, 3 figures, Accepted in BIOKDD 2021
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- How transferable are features in deep neural networks?
- Learning Transferable Features with Deep Adaptation Networks
- YouTube-8M: A Large-Scale Video Classification Benchmark
Cited by in corpus (3)
- Detecting Dementia from Speech and Transcripts using Transformers
- Context-aware attention layers coupled with optimal transport domain adaptation and multimodal fusion methods for recognizing dementia from spontaneous speech
- A Multimodal Approach for Dementia Detection from Spontaneous Speech with Tensor Fusion Layer