Sleep staging · two simple baselines · two separate cohorts

Why can a sleep stager be right most of the time and still miss whole stages?

Five-stage sleep staging of 30-second epochs against the consensus hypnogram the publisher stores, in two Dreem cohorts recorded at different centres with different equipment. Each cohort is its own experiment, with its own models, folds and intervals; nothing on this page compares them. Two deliberately simple CPU baselines show what ordinary accuracy hides.

Measured on: Dreem Open Datasets (DOD-H and DOD-O)

Short answer

Because accuracy counts epochs, and sleep stages are far from equal in size: a stager can miss a whole stage and lose little accuracy. Against the publisher's consensus in two Dreem cohorts, analysed as separate experiments — 25 healthy sleepers and 55 people with obstructive sleep apnoea — a simple spectral ridge scored 72.3% accuracy but 56.7% balanced accuracy in the first and 72.5% accuracy but 49.3% balanced accuracy in the second, and never predicted N1 in either. A prior that always answers N2, the most frequent stage, already reaches 48.0% and 49.2% accuracy, yet only 20.2% and 20.5% balanced accuracy: a floor, not a chance level.

Experiment 1 of 2 · DOD-H

DOD-H: healthy sleepers

Recorded at French Armed Forces Biomedical Research Institute (IRBA), Fatigue and Vigilance Unit, Brétigny-sur-Orge, France. 25 nights, one per person, with 24,662 eligible 30-second epochs. 5 folds within the cohort, assigned before any outcome was seen: each tests 5 nights with models fitted on the other 20, so every night is held out once.

Training prior · Accuracy
48.0% (44.6%–51.4%)
Training prior · Balanced accuracy
20.2% (20.0%–20.6%)
Spectral ridge · Accuracy
72.3% (67.8%–76.3%)
Spectral ridge · Balanced accuracy
56.7% (53.1%–60.0%)
Mean over 25 nights, each counting once. Accuracy and balanced accuracy for both baselines. Dot: the mean; line: whole-night bootstrap 95% interval. There is no dashed line: the training prior is a floor this baseline sets, not a chance level.
Both baselines on the same 25 nights: mean over nights, with whole-night bootstrap 95% intervals. Accuracy is printed beside balanced accuracy on purpose.
BaselineAccuracyBalanced accuracyMacro F1Cohen’s kappa
Training priorpredicts N2 for every epoch · a floor, not a chance level48.0%95% interval 44.6%–51.4%20.2%95% interval 20.0%–20.6%13.0%95% interval 12.3%–13.8%0.00095% interval 0.000–0.000
Spectral ridgestandardised one-hot ridge, alpha 1, on 25 relative spectral values72.3%95% interval 67.8%–76.3%56.7%95% interval 53.1%–60.0%53.3%95% interval 49.0%–57.3%0.58495% interval 0.526–0.636
+36.5 ppSpectral ridge minus training prior, balanced accuracy, the same 25 nights

Well above the floor the prior sets

Paired interval +32.9 to +39.8 pp, computed on the same nights for both baselines: it excludes zero. Paired differences in accuracy, macro F1 and kappa are not published; each baseline’s own value is in the table above.

Spectral ridge, stage by stage: mean over the nights on which each value is defined, with whole-night bootstrap 95% intervals. Epochs: that stage’s count in the consensus. A dash means not defined — never zero.
StageRecallPrecisionF1
Wake3,037 epochs64.8%95% interval 54.9%–73.9%76.4%95% interval 70.1%–82.3%66.5%95% interval 58.5%–73.8%
N11,505 epochs0.0%95% interval 0.0%–0.0%never predicted, so every N1 epoch is missed—not defined: N1 was never predicted on any of the 25 nights0.0%95% interval 0.0%–0.0%
N211,879 epochs85.4%95% interval 77.4%–91.8%77.3%95% interval 72.8%–81.8%78.9%95% interval 74.0%–83.1%
N33,514 epochs59.4%95% interval 46.8%–71.2%on 24 of 25 nights; 1 without N3 in the consensus69.7%95% interval 55.7%–81.9%on 24 of 25 nights; 1 where N3 was never predicted56.2%95% interval 43.6%–67.7%
REM4,727 epochs73.6%95% interval 63.8%–82.6%63.8%95% interval 55.4%–71.5%64.6%95% interval 56.1%–72.2%

The ridge predicted N1 on none of the 25 nights. Its N1 recall and F1 are a measured zero; its N1 precision has no value, because precision divides by the epochs predicted as N1 and there are none.

Experiment 2 of 2 · DOD-O

DOD-O: people with obstructive sleep apnoea

Recorded at Stanford Sleep Medicine Center, United States (clinical trial NCT03657329). 55 nights, one per person, with 53,161 eligible 30-second epochs. 5 folds within the cohort, assigned before any outcome was seen: each tests 11 nights with models fitted on the other 44, so every night is held out once.

Training prior · Accuracy
49.2% (46.3%–52.2%)
Training prior · Balanced accuracy
20.5% (20.1%–21.2%)
Spectral ridge · Accuracy
72.5% (69.7%–75.0%)
Spectral ridge · Balanced accuracy
49.3% (47.3%–51.2%)
Mean over 55 nights, each counting once. Accuracy and balanced accuracy for both baselines. Dot: the mean; line: whole-night bootstrap 95% interval. There is no dashed line: the training prior is a floor this baseline sets, not a chance level.
Both baselines on the same 55 nights: mean over nights, with whole-night bootstrap 95% intervals. Accuracy is printed beside balanced accuracy on purpose.
BaselineAccuracyBalanced accuracyMacro F1Cohen’s kappa
Training priorpredicts N2 for every epoch · a floor, not a chance level49.2%95% interval 46.3%–52.2%20.5%95% interval 20.1%–21.2%13.3%95% interval 12.8%–13.9%0.00095% interval 0.000–0.000
Spectral ridgestandardised one-hot ridge, alpha 1, on 25 relative spectral values72.5%95% interval 69.7%–75.0%49.3%95% interval 47.3%–51.2%47.4%95% interval 44.9%–49.7%0.54295% interval 0.504–0.578
+28.8 ppSpectral ridge minus training prior, balanced accuracy, the same 55 nights

Well above the floor the prior sets

Paired interval +26.7 to +30.7 pp, computed on the same nights for both baselines: it excludes zero. Paired differences in accuracy, macro F1 and kappa are not published; each baseline’s own value is in the table above.

Spectral ridge, stage by stage: mean over the nights on which each value is defined, with whole-night bootstrap 95% intervals. Epochs: that stage’s count in the consensus. A dash means not defined — never zero.
StageRecallPrecisionF1
Wake10,427 epochs80.4%95% interval 75.5%–85.1%76.5%95% interval 72.0%–80.7%75.6%95% interval 71.7%–79.3%
N12,860 epochs0.0%95% interval 0.0%–0.0%never predicted, so every N1 epoch is missed—not defined: N1 was never predicted on any of the 55 nights0.0%95% interval 0.0%–0.0%
N226,271 epochs92.8%95% interval 90.5%–94.7%71.0%95% interval 67.9%–74.0%79.8%95% interval 77.3%–82.0%
N35,500 epochs22.1%95% interval 15.4%–29.2%on 52 of 55 nights; 3 without N3 in the consensus63.4%95% interval 52.4%–73.9%on 49 of 55 nights; 6 where N3 was never predicted27.9%95% interval 20.3%–35.8%on 53 of 55 nights; 2 with N3 neither in the consensus nor predicted
REM8,103 epochs49.5%95% interval 43.1%–55.9%on 53 of 55 nights; 2 without REM in the consensus68.2%95% interval 60.9%–74.9%on 54 of 55 nights; 1 where REM was never predicted52.5%95% interval 46.0%–58.7%on 54 of 55 nights; 1 with REM neither in the consensus nor predicted

The ridge predicted N1 on none of the 55 nights. Its N1 recall and F1 are a measured zero; its N1 precision has no value, because precision divides by the epochs predicted as N1 and there are none.

Cohorts and design

What was evaluated, and how

Two cohorts of the Dreem Open Datasets, each split into its own folds. Every count below is from the reviewed export.

Nights evaluated

80

81 archive records, minus 1 the publisher declares has no consensus scoring: DOD-H 25, DOD-O 55.

Eligible 30-second epochs

77,823

77,901 in all, minus 3 unscored and 75 with a zero or invalid channel scale.

Models fitted

20

Two cohorts × 5 folds × two baselines; none failed.

Target

The consensus hypnogram as the publisher stores it, in five stages — Wake, N1, N2, N3, REM — for 30-second epochs. The individual scorers’ votes were not reconstructed.

Signal and features

Five EEG derivations (C3-M2, F3-F4, F3-M2, F3-O1, F4-O2) at 250 Hz. Each epoch and channel is divided by its own largest absolute value; Welch spectra then give delta, theta, alpha, sigma and beta power as a log share of 0.5–30 Hz power: 25 values per epoch.

Two baselines

The training prior uses the training folds’ stage frequencies, the same for every epoch, so it predicts N2 everywhere. The spectral ridge is a standardised one-hot ridge regression with alpha 1 on the 25 values. Nothing was tuned, calibrated on held-out data, oversampled or class-balanced, and no run was retried on its outcome; stages keep their natural prevalence.

Folds

Five folds within each cohort, assigned by a hash of each record before any outcome was seen. Every night is tested once per baseline, by a model fitted on the other four folds of its own cohort.

Metrics

Computed per night, then averaged so each night counts once, however long. Balanced accuracy is the mean recall over the stages a night contains; macro F1 the mean F1 over the stages it contains or the model predicts. Cohen’s kappa is zero by construction for a prior that gives every epoch the same answer.

Uncertainty

Pointwise 95% whole-night bootstrap within each cohort: 10,000 draws (PCG64, seed 20261003), the same draws for both baselines and every metric. Conditional on the fixed cross-validation predictions — refitting and fold assignment are not included — and not adjusted for multiple comparisons.

Held · no score

Neural-network and foundation models: held

No neural-network or foundation-model score exists for the Dreem cohorts. The stored signal metadata says millivolts, while the publisher’s own converter treats the same arrays as microvolts. Foundation models and other amplitude-sensitive encoders are held because the source’s physical units disagree: they are not run on a guessed unit. The two baselines above divide each epoch and channel by a positive gain before computing relative spectra, so they do not depend on that unit. The hold says nothing about how accurate those models would be.

Its row in the holds register →

Methods & limits

What these baselines can and cannot say

Two fixed classical baselines on CPU, in two separate experiments. They set a floor and show what accuracy hides; they are not a ranking, a clinical tool or a model of expert scoring.

Two separate experiments, not a comparison

DOD-H and DOD-O have separate models, folds and intervals, and were recorded at different centres with different equipment. Their figures are not a controlled comparison of health status, and nothing here measures transfer from one cohort to the other.

Not clinical, not a diagnosis

The target is the consensus the publisher stores, not a reconstruction of each scorer’s vote. Nothing here is a diagnosis, a clinical tool or evidence of equivalence with human experts.

Physical units unresolved

The stored metadata says millivolts; the upstream converter treats the arrays as microvolts. Dividing each epoch and channel by a positive gain makes the relative spectra independent of that scale, but it does not establish that calibration, reference, clipping or filtering are equivalent.

Uncalibrated probabilities

The ridge’s scores are not calibrated probabilities: no calibration, confidence or abstention claim is made, and no probability score is published.

A floor, not a chance level

Balanced accuracy averages only the stages a night contains. The prior predicts N2 everywhere, so on a night that lacks one of the five stages its balanced accuracy is one quarter rather than one fifth; that is why it sits slightly above one fifth. Do not read it as a chance line.

Not comparable with other papers

Natural stage prevalence, no class balancing or tuning, five particular derivations and these folds. Do not compare these figures with papers that use other channels, cohorts, preprocessing or splits.

Dreem Open Datasets (DOD-H and DOD-O)

Antoine Guillot, Fabien Sauvet, Emmanuel H. During and Valentin Thorey · Dreem Open Datasets: Multi-Scored Sleep Datasets to Compare Human and Automated Sleep Staging, IEEE Transactions on Neural Systems and Rehabilitation Engineering (2020), doi:10.1109/TNSRE.2020.3011181; preprint arXiv:1911.03221. Data: Dreem Open Datasets, Zenodo, doi:10.5281/zenodo.15900394. Official repository: github.com/Dreem-Organization/dreem-learning-open, revision 8b332d6827f5ae6a22f4bb97b4deef4273238ec3.

Deposit licence: MIT, as declared on the Zenodo record. It is recorded as the deposit’s declaration, not as a statement about every future use; BCI Report redistributes no upstream code or data.

Paper ↗ · Preprint ↗ · Zenodo deposit ↗ · Official repository at the pinned revision ↗ · MIT

Consent and ethics, cohort by cohort

DOD-H · Ethics approval: Approved by the Committees of Protection of Persons (CPP); declared to the French National Agency for Medicines and Health Products Safety; carried out in compliance with the French Data Protection Act, International Conference on Harmonization (ICH) standards and the Declaration of Helsinki (1964, as revised in 2013). Informed consent: not stated. The paper's DOD-H description contains no informed-consent sentence.

DOD-O · Ethics approval: not stated. No ethics committee or institutional review board is named for DOD-O in the paper. Informed consent: All trial participants gave informed written consent before taking part.

Each cohort is published as it stands: one has an ethics approval and no consent sentence, the other written consent and no named committee. Source of these statements: arXiv:1911.03221v4 (27 April 2020), II. Materials and Methods, A. Datasets, read 2026-10-03.

Data source: large-source-update.json · schema bci-report-large-source-update-v1 · generated 2026-10-03.

Cite this page

BCI Report (2026). Why can a sleep stager be right most of the time and still miss whole stages? https://bci.report/topics/sleep-staging/

Figures from release large-source-update-20261003 (2026-10-03). Cite the upstream datasets as well: their credits are on this page.

Every release is archived on Zenodo: doi:10.5281/zenodo.23123296. BibTeX for the site and its releases → · CITATION.cff ↗