ARTFEED — Contemporary Art Intelligence

EEG Foundation Models Fail to Beat Classical Features in Clinical Benchmark

other · 2026-08-17

A new preprint on arXiv (2607.24519) evaluates five EEG foundation models against traditional feature-based methods using several clinical benchmarks, including the Korean CAUEEG dataset. The results indicate that traditional features outperform all foundation models in classifying normal, mild cognitive impairment, and dementia, achieving a macro-AUROC of 0.734 with a sensitivity of 0.736. In comparison, BIOT-bipolar16 received a score of 0.677, CBraMod 0.669, and REVE 0.568. The no-overlap subset upheld the same ranking (0.717 vs 0.565). Interestingly, all five encoders had a dataset identity decoding score of 1.000, both before and after applying in-fold PCA-50, showing a dependence on dataset membership rather than causal effects. The research highlights the importance of negative-control protocols for evaluating clinical EEG models.

Key facts

  • Five EEG foundation models were evaluated on five tasks across four benchmark datasets plus Korean CAUEEG.
  • Classical features achieved 0.734 macro-AUROC on matched CAUEEG classification, outperforming all foundation models.
  • BIOT-bipolar16 scored 0.677, CBraMod 0.669, and REVE 0.568 on the same task.
  • The annotated no-overlap held-out subset showed classical features at 0.717 vs REVE at 0.565.
  • All five encoders decoded dataset identity at 1.000 before and after in-fold PCA-50.
  • Label permutations collapsed to chance, but balanced subsamples remained at 1.000.
  • The study establishes dataset membership, not a causal site, geography, or population effect.
  • A matched fully randomly initialized encoder was included as a control.

Entities

Institutions

  • arXiv

Locations

  • Korea

Sources