Everything that produced the scores below — cohort, split, electrodes, window, what each method was allowed to learn, and what guessing would score — as the released protocol file states it. Compare scores within this protocol only.
Balanced accuracy (primary) · Macro F1 (secondary)
Results
Every method run under this protocol, with the score the released results file holds. Compare down this table only: other protocols differ in cohort, electrodes, window or chance level.
Spectral ridge
72.2% (70.7%–73.8%)
LaBraM
71.1% (69.0%–73.2%)
CBraMod
71.3% (68.7%–73.7%)
EEGNet
56.5% (54.9%–58.3%)
0%25%50%75%100%
Balanced accuracy, every method under this protocol. Dot: the estimate; line: descriptive 95% interval; dashed line: chance level.
Scoring time includes fitting and prediction, may include accelerator waiting, and excludes data preparation. It is configuration-specific, not a hardware benchmark.
Not run under this protocol:CSP+LDA, ShallowFBCSPNet, Deep4Net, Standard CCA, Temporal ridge. A method missing here was not run on this protocol — that is not a failure.
Read with care
Balanced quality-screened scalp subset, not whole-night deployment prevalence and not ear-EEG. Frozen encoders average fifteen 2-second representations; EEGNet receives 10 epochs.
The protocol, step by step
5 participant-disjoint folds. All recordings from a person stay together. Each person contributes to the held-out predictions once.
Retain six named scalp electrodes; exclude mastoids. Drop any epoch with source per-channel missing-value flag in those six electrodes, actual nonfinite values, any channel std<0.01uV, or peak-to-peak>1000uV. Use full30s at200Hz, no further filtering/reference/amplitude transformation. Select up to30 evenly spaced eligible epochs per participant/class across available nights.
One fixed seed (20260919); no early stopping or test-based tuning. EEGNet trains for 10 epochs per fold. Frozen encoders use training-only standardized ridge heads (alpha 100).
Lightweight balanced quality-screened subset; scores do not describe natural sleep-stage prevalence or the entire73780epoch release.
Two original corrupt sessions and boundary epochs were excluded by the uploader.
200Hz data may have finite replacements despite original missing-value flags; source flags are therefore enforced instead of relying on finite checks alone.
One prediction per30s epoch. Frozen encoders pool15nonoverlapping2s segments; no sequence context across epochs.
All nights from a participant remain together; pretraining overlap unknown.
Quality thresholds fixed before any scores; source artifacts may remain.
Stability
One fixed seed and training budget; multi-seed sensitivity pending.
Pretraining exposure
Unknown unless explicitly documented; no unseen-pretraining claim.
Notes on the methods
One fixed configuration. Foundation encoders remain frozen; small networks train from scratch. These scores do not establish optimal fine-tuned performance. No individual predictions or participant-level results are distributed.
Model terms
Spectral ridge: Trained from scratch / deterministic reference; no third-party pretrained weights. Braindecode BSD-3-Clause; MNE/scikit-learn BSD where used.
LaBraM: Code/repository: MIT · Checkpoint: committed in that repository; no separate weight terms
CBraMod: Code: MIT · Weights: Apache-2.0 (official model card)
EEGNet: Trained from scratch / deterministic reference; no third-party pretrained weights. Braindecode BSD-3-Clause; MNE/scikit-learn BSD where used.
Source and licence
EESM19 scalp subset
CreditKaare B. Mikkelsen et al. · Accurate whole-night sleep monitoring with dry-contact ear-EEG (2019), doi:10.1038/s41598-019-53115-3; OpenNeuro ds005185 v1.0.2. Processed mirror: Zachary1150/EESM19-Processed.
BCI Report does not redistribute any recording. These are aggregate measurements computed by BCI Report under the licence above; the data belong to the people credited.
Permitted scopePersonal noncommercial research; aggregate results only
Privacy Only cohort aggregates are published here: no recording, no participant identifier, no per-person score. According to the 2025 data descriptor, the informed consent form did not mention publication, and before release the GDPR office of Region Midt judged the data fully anonymised: consent covered the study, and the public release rests on that anonymisation judgement.The full review note is in the protocol JSON ↓
Public-data register noteClear upstream CC0 and study ethics evidence; use only the pinned scalp-signal derivative and publish aggregates, with provenance disclosed. Consent covered the study, not publication; the public release rests on a GDPR anonymization assessment (amended 2026-09-22).