Everything that produced the scores below — cohort, split, electrodes, window, what each method was allowed to learn, and what guessing would score — as the released protocol file states it. Compare scores within this protocol only.
Balanced accuracy (primary) · Macro F1 (secondary)
Results
Every method run under this protocol, with the score the released results file holds. Compare down this table only: other protocols differ in cohort, electrodes, window or chance level.
Spectral ridge
56.8% (54.6%–58.9%)
LaBraM
64.6% (60.5%–68.6%)
CBraMod
62.3% (58.1%–66.5%)
EEGNet
67.6% (62.8%–72.5%)
0%25%50%75%100%
Balanced accuracy, every method under this protocol. Dot: the estimate; line: descriptive 95% interval; dashed line: chance level.
67.6%62.8%–72.5%Single seed — the highest of the seeds run (see Stability)
0.637Mean across held-out participants
19
36
117.7 s
Scoring time includes fitting and prediction, may include accelerator waiting, and excludes data preparation. It is configuration-specific, not a hardware benchmark.
Not run under this protocol:CSP+LDA, ShallowFBCSPNet, Deep4Net, Standard CCA, Temporal ridge. A method missing here was not run on this protocol — that is not a failure.
Read with care
Thirty nonoverlapping 2-second windows from each of two conditions per person. Published signals were already cleaned with ICA; task performance groups are not evaluated.
The protocol, step by step
5 participant-disjoint folds. All recordings from a person stay together. Each person contributes to the held-out predictions once.
First 60 s of each recording; 30 contiguous nonoverlapping 2 s windows; exclude A2-A1 ear-difference and ECG; EDF physical volts converted to microvolts; subtract each channel's window mean; no rejection, filtering, resampling, or learned preprocessing.
One fixed seed (20260919); no early stopping or test-based tuning. EEGNet trains for 20 epochs per fold. Frozen encoders use training-only standardized ridge heads (alpha 100).
The benchmark detects condition (rest versus serial subtraction), not the good/bad count-quality participant grouping.
Only the first documented 60 seconds is retained even though EDF containers are longer.
The source README reports prior ICA artifact removal, so these are not untouched acquisition signals.
Open Data Commons Attribution License v1.0 applies; retain PhysioNet attribution.
EEGNet three-seed mean 67.52%; sample SD 0.14 percentage points; range 67.36–67.64%. Main table retains the original fixed seed; this is not a confidence interval.
Stability
EEGNet three-seed mean 67.52%; sample SD 0.14 percentage points; range 67.36–67.64%. Main table retains the original fixed seed; this is not a confidence interval.
Seeds run for EEGNet: 67.64% · 67.55% · 67.36%; mean 67.52%
Same participants, folds, preprocessing and 20-epoch budget. Three seeds measure initialization variability, not population uncertainty. Main table retains its preselected seed; no best-seed selection.
Pretraining exposure
Unknown unless explicitly documented; no unseen-pretraining claim.
Notes on the methods
One fixed configuration. Foundation encoders remain frozen; small networks train from scratch. These scores do not establish optimal fine-tuned performance. No individual predictions or participant-level results are distributed.
One fixed configuration. Foundation encoders remain frozen; small networks train from scratch. These scores do not establish optimal fine-tuned performance. No individual predictions or participant-level results are distributed. EEGNet three-seed mean 67.52%; sample SD 0.14 percentage points; range 67.36–67.64%. Main table retains the original fixed seed; this is not a confidence interval.
Model terms
Spectral ridge: Trained from scratch / deterministic reference; no third-party pretrained weights. Braindecode BSD-3-Clause; MNE/scikit-learn BSD where used.
LaBraM: Code/repository: MIT · Checkpoint: committed in that repository; no separate weight terms
CBraMod: Code: MIT · Weights: Apache-2.0 (official model card)
EEGNet: Trained from scratch / deterministic reference; no third-party pretrained weights. Braindecode BSD-3-Clause; MNE/scikit-learn BSD where used.
Source and licence
EEGMAT
CreditIgor Zyma, Ivan Seleznov, Anton Popov, Mariia Chernykh, Oleksii Shpenkov · EEG During Mental Arithmetic Tasks 1.0.0, PhysioNet, doi:10.13026/C2JQ1P. Study: Zyma et al. (2019), doi:10.3390/data4010014. PhysioNet platform: Pollard et al. (2026), doi:10.1038/s44360-026-00096-z.
BCI Report does not redistribute any recording. These are aggregate measurements computed by BCI Report under the licence above; the data belong to the people credited.
Permitted scopePersonal noncommercial research; aggregate results only
Public-data register noteAttribution license and study-specific approval/consent are documented; eligibility assumes aggregate-only output and exclusion of subject-info fields.