Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram
arXiv:2607.22980
Abstract
Brain-computer interfaces (BCIs) have long sought calibration-free operation, yet classifiers are typically benchmarked on discrimination alone. Discrimination is blind to calibration, a meaningful gap given that electroencephalogram (EEG) signals are nonstationary and point-estimate classifiers can become overconfident under distribution shift. We conducted a large-scale study contrasting Bayesian complete-pooling models against frequentist baselines for cross-subject, left-hand versus right-hand motor imagery EEG classification across 20 datasets. Each of six frequentist pipelines was paired with an analogous Bayesian pipeline sharing identical feature engineering, fit via Markov chain Monte Carlo posterior sampling. Our primary metric was the Brier score, decomposed into reliability and resolution, alongside the area under the receiver operating characteristic curve for discrimination and Shannon entropy for sharpness. For each metric we fit a random-effects meta-analysis with Knapp-Hartung adjustment, verified by leave-one-out influence analysis. Bayesian complete-pooling produced statistically significant improvements in reliability and increases in predictive uncertainty (lower sharpness), but 95% confidence intervals bounded both effects near zero, and neither held significance when datasets with flagged posterior sampling were excluded. Brier score, resolution, and discrimination showed no significant differences. Between-study heterogeneity was low across all metrics. Bayesian pipelines consumed roughly thirteen times more energy than their frequentist counterparts, a cost that remains modest relative to common household appliances. Bayesian complete-pooling alone offers limited benefit for cross-subject motor imagery classification, but its modest cost makes partial-pooling across subjects and sessions a feasible next step.