Deep Insights into Cognitive Decline: A Survey of Leveraging Non-Intrusive Modalities with Deep Learning Techniques
arXiv:2410.18972 · doi:10.1016/j.asoc.2025.113787
Abstract
Cognitive decline is a natural part of aging. However, under some circumstances, this decline is more pronounced than expected, typically due to disorders such as Alzheimer's disease. Early detection of an anomalous decline is crucial, as it can facilitate timely professional intervention. While medical data can help, it often involves invasive procedures. An alternative approach is to employ non-intrusive techniques such as speech or handwriting analysis, which do not disturb daily activities. This survey reviews the most relevant non-intrusive methodologies that use deep learning techniques to automate the cognitive decline detection task, including audio, text, and visual processing. We discuss the key features and advantages of each modality and methodology, including state-of-the-art approaches like Transformer architecture and foundation models. In addition, we present studies that integrate different modalities to develop multimodal models. We also highlight the most significant datasets and the quantitative results from studies using these resources. From this review, several conclusions emerge. In most cases, text-based approaches consistently outperform other modalities. Furthermore, combining various approaches from individual modalities into a multimodal model consistently enhances performance across nearly all scenarios.
References in corpus (28)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Distributed Representations of Sentences and Documents
- LLaMA: Open and Efficient Foundation Language Models
- Longformer: The Long-Document Transformer
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
- A Survey of Large Language Models
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Deep Learning for Audio Signal Processing
- Voice Recognition Algorithms using Mel Frequency Cepstral Coefficient (MFCC) and Dynamic Time Warping (DTW) Techniques
- Publicly Available Clinical BERT Embeddings
- Spanish Pre-trained BERT Model and Evaluation Data
- MFCC-based Recurrent Neural Network for Automatic Clinical Depression Recognition and Assessment from Speech
- ActionCLIP: A New Paradigm for Video Action Recognition
- Deep Transfer Learning for Automatic Speech Recognition: Towards Better Generalization
- ConvBERT: Improving BERT with Span-based Dynamic Convolution
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
- Explainable Identification of Dementia from Transcripts using Transformer Networks
- InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining
- Detecting Dementia from Speech and Transcripts using Transformers
- Speech Emotion Recognition via Contrastive Loss under Siamese Networks
- MC-ViViT: Multi-branch Classifier-ViViT to detect Mild Cognitive Impairment in older adults using facial videos
- The Dem@Care Experiments and Datasets: a Technical Report
- CogniAlign: Word-Level Multimodal Speech Alignment with Gated Cross-Attention for Alzheimer's Detection
- How word semantics and phonology affect handwriting of Alzheimer's patients: a machine learning based analysis