EEG Foundation Models Fail to Beat Classical Features in Clinical Benchmark
A new preprint on arXiv (2607.24519) evaluates five EEG foundation models against traditional feature-based methods using several clinical benchmarks, including the Korean CAUEEG dataset. The results indicate that traditional features outperform all foundation models in classifying normal, mild cognitive impairment, and dementia, achieving a macro-AUROC of 0.734 with a sensitivity of 0.736. In comparison, BIOT-bipolar16 received a score of 0.677, CBraMod 0.669, and REVE 0.568. The no-overlap subset upheld the same ranking (0.717 vs 0.565). Interestingly, all five encoders had a dataset identity decoding score of 1.000, both before and after applying in-fold PCA-50, showing a dependence on dataset membership rather than causal effects. The research highlights the importance of negative-control protocols for evaluating clinical EEG models.
Key facts
- Five EEG foundation models were evaluated on five tasks across four benchmark datasets plus Korean CAUEEG.
- Classical features achieved 0.734 macro-AUROC on matched CAUEEG classification, outperforming all foundation models.
- BIOT-bipolar16 scored 0.677, CBraMod 0.669, and REVE 0.568 on the same task.
- The annotated no-overlap held-out subset showed classical features at 0.717 vs REVE at 0.565.
- All five encoders decoded dataset identity at 1.000 before and after in-fold PCA-50.
- Label permutations collapsed to chance, but balanced subsamples remained at 1.000.
- The study establishes dataset membership, not a causal site, geography, or population effect.
- A matched fully randomly initialized encoder was included as a control.
Entities
Institutions
- arXiv
Locations
- Korea