Nights evaluated
81 archive records, minus 1 the publisher declares has no consensus scoring: DOD-H 25, DOD-O 55.
Sleep staging · two simple baselines · two separate cohorts
Five-stage sleep staging of 30-second epochs against the consensus hypnogram the publisher stores, in two Dreem cohorts recorded at different centres with different equipment. Each cohort is its own experiment, with its own models, folds and intervals; nothing on this page compares them. Two deliberately simple CPU baselines show what ordinary accuracy hides.
Measured on: Dreem Open Datasets (DOD-H and DOD-O)
Because accuracy counts epochs, and sleep stages are far from equal in size: a stager can miss a whole stage and lose little accuracy. Against the publisher's consensus in two Dreem cohorts, analysed as separate experiments — 25 healthy sleepers and 55 people with obstructive sleep apnoea — a simple spectral ridge scored 72.3% accuracy but 56.7% balanced accuracy in the first and 72.5% accuracy but 49.3% balanced accuracy in the second, and never predicted N1 in either. A prior that always answers N2, the most frequent stage, already reaches 48.0% and 49.2% accuracy, yet only 20.2% and 20.5% balanced accuracy: a floor, not a chance level.
Experiment 1 of 2 · DOD-H
Recorded at French Armed Forces Biomedical Research Institute (IRBA), Fatigue and Vigilance Unit, Brétigny-sur-Orge, France. 25 nights, one per person, with 24,662 eligible 30-second epochs. 5 folds within the cohort, assigned before any outcome was seen: each tests 5 nights with models fitted on the other 20, so every night is held out once.
| Baseline | Accuracy | Balanced accuracy | Macro F1 | Cohen’s kappa |
|---|---|---|---|---|
| Training priorpredicts N2 for every epoch · a floor, not a chance level | 48.0%95% interval 44.6%–51.4% | 20.2%95% interval 20.0%–20.6% | 13.0%95% interval 12.3%–13.8% | 0.00095% interval 0.000–0.000 |
| Spectral ridgestandardised one-hot ridge, alpha 1, on 25 relative spectral values | 72.3%95% interval 67.8%–76.3% | 56.7%95% interval 53.1%–60.0% | 53.3%95% interval 49.0%–57.3% | 0.58495% interval 0.526–0.636 |
Paired interval +32.9 to +39.8 pp, computed on the same nights for both baselines: it excludes zero. Paired differences in accuracy, macro F1 and kappa are not published; each baseline’s own value is in the table above.
| Stage | Recall | Precision | F1 |
|---|---|---|---|
| Wake3,037 epochs | 64.8%95% interval 54.9%–73.9% | 76.4%95% interval 70.1%–82.3% | 66.5%95% interval 58.5%–73.8% |
| N11,505 epochs | 0.0%95% interval 0.0%–0.0%never predicted, so every N1 epoch is missed | —not defined: N1 was never predicted on any of the 25 nights | 0.0%95% interval 0.0%–0.0% |
| N211,879 epochs | 85.4%95% interval 77.4%–91.8% | 77.3%95% interval 72.8%–81.8% | 78.9%95% interval 74.0%–83.1% |
| N33,514 epochs | 59.4%95% interval 46.8%–71.2%on 24 of 25 nights; 1 without N3 in the consensus | 69.7%95% interval 55.7%–81.9%on 24 of 25 nights; 1 where N3 was never predicted | 56.2%95% interval 43.6%–67.7% |
| REM4,727 epochs | 73.6%95% interval 63.8%–82.6% | 63.8%95% interval 55.4%–71.5% | 64.6%95% interval 56.1%–72.2% |
The ridge predicted N1 on none of the 25 nights. Its N1 recall and F1 are a measured zero; its N1 precision has no value, because precision divides by the epochs predicted as N1 and there are none.
Experiment 2 of 2 · DOD-O
Recorded at Stanford Sleep Medicine Center, United States (clinical trial NCT03657329). 55 nights, one per person, with 53,161 eligible 30-second epochs. 5 folds within the cohort, assigned before any outcome was seen: each tests 11 nights with models fitted on the other 44, so every night is held out once.
| Baseline | Accuracy | Balanced accuracy | Macro F1 | Cohen’s kappa |
|---|---|---|---|---|
| Training priorpredicts N2 for every epoch · a floor, not a chance level | 49.2%95% interval 46.3%–52.2% | 20.5%95% interval 20.1%–21.2% | 13.3%95% interval 12.8%–13.9% | 0.00095% interval 0.000–0.000 |
| Spectral ridgestandardised one-hot ridge, alpha 1, on 25 relative spectral values | 72.5%95% interval 69.7%–75.0% | 49.3%95% interval 47.3%–51.2% | 47.4%95% interval 44.9%–49.7% | 0.54295% interval 0.504–0.578 |
Paired interval +26.7 to +30.7 pp, computed on the same nights for both baselines: it excludes zero. Paired differences in accuracy, macro F1 and kappa are not published; each baseline’s own value is in the table above.
| Stage | Recall | Precision | F1 |
|---|---|---|---|
| Wake10,427 epochs | 80.4%95% interval 75.5%–85.1% | 76.5%95% interval 72.0%–80.7% | 75.6%95% interval 71.7%–79.3% |
| N12,860 epochs | 0.0%95% interval 0.0%–0.0%never predicted, so every N1 epoch is missed | —not defined: N1 was never predicted on any of the 55 nights | 0.0%95% interval 0.0%–0.0% |
| N226,271 epochs | 92.8%95% interval 90.5%–94.7% | 71.0%95% interval 67.9%–74.0% | 79.8%95% interval 77.3%–82.0% |
| N35,500 epochs | 22.1%95% interval 15.4%–29.2%on 52 of 55 nights; 3 without N3 in the consensus | 63.4%95% interval 52.4%–73.9%on 49 of 55 nights; 6 where N3 was never predicted | 27.9%95% interval 20.3%–35.8%on 53 of 55 nights; 2 with N3 neither in the consensus nor predicted |
| REM8,103 epochs | 49.5%95% interval 43.1%–55.9%on 53 of 55 nights; 2 without REM in the consensus | 68.2%95% interval 60.9%–74.9%on 54 of 55 nights; 1 where REM was never predicted | 52.5%95% interval 46.0%–58.7%on 54 of 55 nights; 1 with REM neither in the consensus nor predicted |
The ridge predicted N1 on none of the 55 nights. Its N1 recall and F1 are a measured zero; its N1 precision has no value, because precision divides by the epochs predicted as N1 and there are none.
Cohorts and design
Two cohorts of the Dreem Open Datasets, each split into its own folds. Every count below is from the reviewed export.
Nights evaluated
81 archive records, minus 1 the publisher declares has no consensus scoring: DOD-H 25, DOD-O 55.
Eligible 30-second epochs
77,901 in all, minus 3 unscored and 75 with a zero or invalid channel scale.
Models fitted
Two cohorts × 5 folds × two baselines; none failed.
The consensus hypnogram as the publisher stores it, in five stages — Wake, N1, N2, N3, REM — for 30-second epochs. The individual scorers’ votes were not reconstructed.
Five EEG derivations (C3-M2, F3-F4, F3-M2, F3-O1, F4-O2) at 250 Hz. Each epoch and channel is divided by its own largest absolute value; Welch spectra then give delta, theta, alpha, sigma and beta power as a log share of 0.5–30 Hz power: 25 values per epoch.
The training prior uses the training folds’ stage frequencies, the same for every epoch, so it predicts N2 everywhere. The spectral ridge is a standardised one-hot ridge regression with alpha 1 on the 25 values. Nothing was tuned, calibrated on held-out data, oversampled or class-balanced, and no run was retried on its outcome; stages keep their natural prevalence.
Five folds within each cohort, assigned by a hash of each record before any outcome was seen. Every night is tested once per baseline, by a model fitted on the other four folds of its own cohort.
Computed per night, then averaged so each night counts once, however long. Balanced accuracy is the mean recall over the stages a night contains; macro F1 the mean F1 over the stages it contains or the model predicts. Cohen’s kappa is zero by construction for a prior that gives every epoch the same answer.
Pointwise 95% whole-night bootstrap within each cohort: 10,000 draws (PCG64, seed 20261003), the same draws for both baselines and every metric. Conditional on the fixed cross-validation predictions — refitting and fold assignment are not included — and not adjusted for multiple comparisons.
Held · no score
No neural-network or foundation-model score exists for the Dreem cohorts. The stored signal metadata says millivolts, while the publisher’s own converter treats the same arrays as microvolts. Foundation models and other amplitude-sensitive encoders are held because the source’s physical units disagree: they are not run on a guessed unit. The two baselines above divide each epoch and channel by a positive gain before computing relative spectra, so they do not depend on that unit. The hold says nothing about how accurate those models would be.
Methods & limits
Two fixed classical baselines on CPU, in two separate experiments. They set a floor and show what accuracy hides; they are not a ranking, a clinical tool or a model of expert scoring.
DOD-H and DOD-O have separate models, folds and intervals, and were recorded at different centres with different equipment. Their figures are not a controlled comparison of health status, and nothing here measures transfer from one cohort to the other.
The target is the consensus the publisher stores, not a reconstruction of each scorer’s vote. Nothing here is a diagnosis, a clinical tool or evidence of equivalence with human experts.
The stored metadata says millivolts; the upstream converter treats the arrays as microvolts. Dividing each epoch and channel by a positive gain makes the relative spectra independent of that scale, but it does not establish that calibration, reference, clipping or filtering are equivalent.
The ridge’s scores are not calibrated probabilities: no calibration, confidence or abstention claim is made, and no probability score is published.
Balanced accuracy averages only the stages a night contains. The prior predicts N2 everywhere, so on a night that lacks one of the five stages its balanced accuracy is one quarter rather than one fifth; that is why it sits slightly above one fifth. Do not read it as a chance line.
Natural stage prevalence, no class balancing or tuning, five particular derivations and these folds. Do not compare these figures with papers that use other channels, cohorts, preprocessing or splits.
Antoine Guillot, Fabien Sauvet, Emmanuel H. During and Valentin Thorey · Dreem Open Datasets: Multi-Scored Sleep Datasets to Compare Human and Automated Sleep Staging, IEEE Transactions on Neural Systems and Rehabilitation Engineering (2020), doi:10.1109/TNSRE.2020.3011181; preprint arXiv:1911.03221. Data: Dreem Open Datasets, Zenodo, doi:10.5281/zenodo.15900394. Official repository: github.com/Dreem-Organization/dreem-learning-open, revision 8b332d6827f5ae6a22f4bb97b4deef4273238ec3.
Deposit licence: MIT, as declared on the Zenodo record. It is recorded as the deposit’s declaration, not as a statement about every future use; BCI Report redistributes no upstream code or data.
Paper ↗ · Preprint ↗ · Zenodo deposit ↗ · Official repository at the pinned revision ↗ · MIT
DOD-H · Ethics approval: Approved by the Committees of Protection of Persons (CPP); declared to the French National Agency for Medicines and Health Products Safety; carried out in compliance with the French Data Protection Act, International Conference on Harmonization (ICH) standards and the Declaration of Helsinki (1964, as revised in 2013). Informed consent: not stated. The paper's DOD-H description contains no informed-consent sentence.
DOD-O · Ethics approval: not stated. No ethics committee or institutional review board is named for DOD-O in the paper. Informed consent: All trial participants gave informed written consent before taking part.
Each cohort is published as it stands: one has an ethics approval and no consent sentence, the other written consent and no named committee. Source of these statements: arXiv:1911.03221v4 (27 April 2020), II. Materials and Methods, A. Datasets, read 2026-10-03.
Data source: large-source-update.json · schema bci-report-large-source-update-v1 · generated 2026-10-03.
BCI Report (2026). Why can a sleep stager be right most of the time and still miss whole stages? https://bci.report/topics/sleep-staging/
Figures from release large-source-update-20261003 (2026-10-03). Cite the upstream datasets as well: their credits are on this page.
Every release is archived on Zenodo: doi:10.5281/zenodo.23123296. BibTeX for the site and its releases → · CITATION.cff ↗