brain-computer interfaces

Does EEG Foundation Models Transfer to Speech? A Benchmark on Overt and Imagined Speech Decoding

arXiv:2607.27268

summary

The paper benchmarks two EEG foundation models on overt and imagined speech decoding tasks, comparing them to established convolutional baselines, and finds that large‑scale EEG pretraining does not consistently improve performance for speech production.

Abstract

EEG foundation models pretrained on thousands of hours have shown large gains over task-specific networks for motor imagery, seizure detection, sleep staging, and emotion recognition, but their transfer to speech decoding-arguably the most demanding non-invasive BCI application-remains untested. We present the first systematic benchmark of EEG foundation models against strong convolutional baselines for speech decoding, using two corpora: UGR-MINDVOICE (overt and covert Iberian Spanish) and BCI Competition 2020 Track 3 (imagined speech). We compare two foundation models (LaBraM, EEGMamba) against three established baselines (EEGNet, ShallowFBCSPNet, EEGConformer) under a unified preprocessing and fine-tuning protocol. Large-scale EEG pretraining yields no consistent advantage over a 16K-parameter CNN on speech tasks, indicating that current general-purpose EEG pretraining does not yet transfer to speech production and motivating speech-specific foundation models.

6 pages, 1 figure, 3 tables, submitted to IberSPEECH 2026

Topics & keywords

#eeg foundation models#speech decoding#overt speech#imagined speech#transfer learningEEGMambaLaBraMEEGNetfine-tuningconvolutional neural networkEEG pretraining