paper

Neurai-VN Benchmark: Standardized Machine Learning Models for Multimodal Digital Phenotyping in Mental Health Classification

arXiv:2607.25232

Abstract

Digital phenotyping (DP) using smartphones and wearable devices has emerged as a promising approach for assessing mental health, particularly depression and anxiety. However, progress remains difficult to evaluate because of heterogeneity across datasets and inconsistencies in preprocessing pipelines. In this work, we introduce a reproducible machine learning benchmark using the Neurai-VN dataset, a multimodal digital phenotyping dataset collected from 100 Vietnamese adults over two weeks. We define four binary classification tasks evaluated using standardized subject-wise cross-validation. Representative linear, tree-based, and neural baseline models are evaluated systematically across predefined feature-group configurations. Mean subject-level F1 scores across five cross-validation folds reached 0.71 for Healthy Control vs. Depression and Healthy Control vs. Clinical, while Healthy Control vs. Anxiety and Depression vs. Anxiety achieved 0.69 and 0.56, respectively. These baseline results provide reproducible baselines for future research on multimodal DP for mental health classification tasks. The code to reproduce the benchmark is available at https://github.com/neurai-vn/Neurai-VN-benchmark.