Twins-PainViT: Towards a Modality-Agnostic Vision Transformer Framework for Multimodal Automatic Pain Assessment using Facial Videos and fNIRS
arXiv:2407.19809 · doi:10.1109/ACIIW63320.2024.00007
Abstract
Automatic pain assessment plays a critical role for advancing healthcare and optimizing pain management strategies. This study has been submitted to the First Multimodal Sensing Grand Challenge for Next-Gen Pain Assessment (AI4PAIN). The proposed multimodal framework utilizes facial videos and fNIRS and presents a modality-agnostic approach, alleviating the need for domain-specific models. Employing a dual ViT configuration and adopting waveform representations for the fNIRS, as well as for the extracted embeddings from the two modalities, demonstrate the efficacy of the proposed method, achieving an accuracy of 46.76% in the multilevel pain assessment task.
References in corpus (5)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks
- AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild
- A Full Transformer-based Framework for Automatic Pain Estimation using Videos
- Multi-task Neural Networks for Pain Intensity Estimation using Electrocardiogram and Demographic Factors