# Why can a sleep stager be right most of the time and still miss whole stages?

Sleep staging · two simple baselines · two separate cohorts

Five-stage sleep staging of 30-second epochs against the consensus hypnogram the publisher stores, in two Dreem cohorts recorded at different centres with different equipment. Each cohort is its own experiment, with its own models, folds and intervals; nothing on this page compares them. Two deliberately simple CPU baselines show what ordinary accuracy hides.

- [Reviewed aggregate JSON ↓](https://bci.report/data/large-source-update.json)
- 

Measured on: [Dreem Open Datasets (DOD-H and DOD-O)](https://bci.report/datasets/dreem-dod/)

## Short answer

Because accuracy counts epochs, and sleep stages are far from equal in size: a stager can miss a whole stage and lose little accuracy. Against the publisher's consensus in two Dreem cohorts, analysed as separate experiments — **25** healthy sleepers and **55** people with obstructive sleep apnoea — a simple spectral ridge scored **72.3%** accuracy but **56.7%** balanced accuracy in the first and **72.5%** accuracy but **49.3%** balanced accuracy in the second, and never predicted N1 in either. A prior that always answers N2, the most frequent stage, already reaches **48.0%** and **49.2%** accuracy, yet only **20.2%** and **20.5%** balanced accuracy: a floor, not a chance level.

Experiment 1 of 2 · DOD-H

## DOD-H: healthy sleepers

Recorded at French Armed Forces Biomedical Research Institute (IRBA), Fatigue and Vigilance Unit, Brétigny-sur-Orge, France. 25 nights, one per person, with 24,662 eligible 30-second epochs. 5 folds within the cohort, assigned before any outcome was seen: each tests 5 nights with models fitted on the other 20, so every night is held out once.

**Mean over 25 nights, each counting once.** Accuracy and balanced accuracy for both baselines. Dot: the mean; line: whole-night bootstrap 95% interval. There is no dashed line: the training prior is a floor this baseline sets, not a chance level.

- Training prior · Accuracy: 48.0% (44.6%–51.4%)
- Training prior · Balanced accuracy: 20.2% (20.0%–20.6%)
- Spectral ridge · Accuracy: 72.3% (67.8%–76.3%)
- Spectral ridge · Balanced accuracy: 56.7% (53.1%–60.0%)

Both baselines on the same 25 nights: mean over nights, with whole-night bootstrap 95% intervals. Accuracy is printed beside balanced accuracy on purpose.

| Baseline | Accuracy | Balanced accuracy | Macro F1 | Cohen’s kappa |
| --- | --- | --- | --- | --- |
| Training prior — predicts N2 for every epoch · a floor, not a chance level | 48.0% (95% interval 44.6%–51.4%) | 20.2% (95% interval 20.0%–20.6%) | 13.0% (95% interval 12.3%–13.8%) | 0.000 (95% interval 0.000–0.000) |
| Spectral ridge — standardised one-hot ridge, alpha 1, on 25 relative spectral values | 72.3% (95% interval 67.8%–76.3%) | 56.7% (95% interval 53.1%–60.0%) | 53.3% (95% interval 49.0%–57.3%) | 0.584 (95% interval 0.526–0.636) |

**+36.5 pp** Spectral ridge minus training prior, balanced accuracy, the same 25 nights

### Well above the floor the prior sets

Paired interval +32.9 to +39.8 pp, computed on the same nights for both baselines: it excludes zero. Paired differences in accuracy, macro F1 and kappa are not published; each baseline’s own value is in the table above.

Spectral ridge, stage by stage: mean over the nights on which each value is defined, with whole-night bootstrap 95% intervals. Epochs: that stage’s count in the consensus. A dash means not defined — never zero.

| Stage | Recall | Precision | F1 |
| --- | --- | --- | --- |
| Wake — 3,037 epochs | 64.8% (95% interval 54.9%–73.9%) | 76.4% (95% interval 70.1%–82.3%) | 66.5% (95% interval 58.5%–73.8%) |
| N1 — 1,505 epochs | 0.0% (95% interval 0.0%–0.0%) (never predicted, so every N1 epoch is missed) | — (not defined: N1 was never predicted on any of the 25 nights) | 0.0% (95% interval 0.0%–0.0%) |
| N2 — 11,879 epochs | 85.4% (95% interval 77.4%–91.8%) | 77.3% (95% interval 72.8%–81.8%) | 78.9% (95% interval 74.0%–83.1%) |
| N3 — 3,514 epochs | 59.4% (95% interval 46.8%–71.2%) (on 24 of 25 nights; 1 without N3 in the consensus) | 69.7% (95% interval 55.7%–81.9%) (on 24 of 25 nights; 1 where N3 was never predicted) | 56.2% (95% interval 43.6%–67.7%) |
| REM — 4,727 epochs | 73.6% (95% interval 63.8%–82.6%) | 63.8% (95% interval 55.4%–71.5%) | 64.6% (95% interval 56.1%–72.2%) |

The ridge predicted N1 on none of the 25 nights. Its N1 recall and F1 are a measured zero; its N1 precision has no value, because precision divides by the epochs predicted as N1 and there are none.

Experiment 2 of 2 · DOD-O

## DOD-O: people with obstructive sleep apnoea

Recorded at Stanford Sleep Medicine Center, United States (clinical trial NCT03657329). 55 nights, one per person, with 53,161 eligible 30-second epochs. 5 folds within the cohort, assigned before any outcome was seen: each tests 11 nights with models fitted on the other 44, so every night is held out once.

**Mean over 55 nights, each counting once.** Accuracy and balanced accuracy for both baselines. Dot: the mean; line: whole-night bootstrap 95% interval. There is no dashed line: the training prior is a floor this baseline sets, not a chance level.

- Training prior · Accuracy: 49.2% (46.3%–52.2%)
- Training prior · Balanced accuracy: 20.5% (20.1%–21.2%)
- Spectral ridge · Accuracy: 72.5% (69.7%–75.0%)
- Spectral ridge · Balanced accuracy: 49.3% (47.3%–51.2%)

Both baselines on the same 55 nights: mean over nights, with whole-night bootstrap 95% intervals. Accuracy is printed beside balanced accuracy on purpose.

| Baseline | Accuracy | Balanced accuracy | Macro F1 | Cohen’s kappa |
| --- | --- | --- | --- | --- |
| Training prior — predicts N2 for every epoch · a floor, not a chance level | 49.2% (95% interval 46.3%–52.2%) | 20.5% (95% interval 20.1%–21.2%) | 13.3% (95% interval 12.8%–13.9%) | 0.000 (95% interval 0.000–0.000) |
| Spectral ridge — standardised one-hot ridge, alpha 1, on 25 relative spectral values | 72.5% (95% interval 69.7%–75.0%) | 49.3% (95% interval 47.3%–51.2%) | 47.4% (95% interval 44.9%–49.7%) | 0.542 (95% interval 0.504–0.578) |

**+28.8 pp** Spectral ridge minus training prior, balanced accuracy, the same 55 nights

### Well above the floor the prior sets

Paired interval +26.7 to +30.7 pp, computed on the same nights for both baselines: it excludes zero. Paired differences in accuracy, macro F1 and kappa are not published; each baseline’s own value is in the table above.

Spectral ridge, stage by stage: mean over the nights on which each value is defined, with whole-night bootstrap 95% intervals. Epochs: that stage’s count in the consensus. A dash means not defined — never zero.

| Stage | Recall | Precision | F1 |
| --- | --- | --- | --- |
| Wake — 10,427 epochs | 80.4% (95% interval 75.5%–85.1%) | 76.5% (95% interval 72.0%–80.7%) | 75.6% (95% interval 71.7%–79.3%) |
| N1 — 2,860 epochs | 0.0% (95% interval 0.0%–0.0%) (never predicted, so every N1 epoch is missed) | — (not defined: N1 was never predicted on any of the 55 nights) | 0.0% (95% interval 0.0%–0.0%) |
| N2 — 26,271 epochs | 92.8% (95% interval 90.5%–94.7%) | 71.0% (95% interval 67.9%–74.0%) | 79.8% (95% interval 77.3%–82.0%) |
| N3 — 5,500 epochs | 22.1% (95% interval 15.4%–29.2%) (on 52 of 55 nights; 3 without N3 in the consensus) | 63.4% (95% interval 52.4%–73.9%) (on 49 of 55 nights; 6 where N3 was never predicted) | 27.9% (95% interval 20.3%–35.8%) (on 53 of 55 nights; 2 with N3 neither in the consensus nor predicted) |
| REM — 8,103 epochs | 49.5% (95% interval 43.1%–55.9%) (on 53 of 55 nights; 2 without REM in the consensus) | 68.2% (95% interval 60.9%–74.9%) (on 54 of 55 nights; 1 where REM was never predicted) | 52.5% (95% interval 46.0%–58.7%) (on 54 of 55 nights; 1 with REM neither in the consensus nor predicted) |

The ridge predicted N1 on none of the 55 nights. Its N1 recall and F1 are a measured zero; its N1 precision has no value, because precision divides by the epochs predicted as N1 and there are none.

Cohorts and design

## What was evaluated, and how

Two cohorts of the Dreem Open Datasets, each split into its own folds. Every count below is from the reviewed export.

Nights evaluated

80

81 archive records, minus 1 the publisher declares has no consensus scoring: DOD-H 25, DOD-O 55.

Eligible 30-second epochs

77,823

77,901 in all, minus 3 unscored and 75 with a zero or invalid channel scale.

Models fitted

20

Two cohorts × 5 folds × two baselines; none failed.

### Target

The consensus hypnogram as the publisher stores it, in five stages — Wake, N1, N2, N3, REM — for 30-second epochs. The individual scorers’ votes were not reconstructed.

### Signal and features

Five EEG derivations (C3-M2, F3-F4, F3-M2, F3-O1, F4-O2) at 250 Hz. Each epoch and channel is divided by its own largest absolute value; Welch spectra then give delta, theta, alpha, sigma and beta power as a log share of 0.5–30 Hz power: 25 values per epoch.

### Two baselines

The training prior uses the training folds’ stage frequencies, the same for every epoch, so it predicts N2 everywhere. The spectral ridge is a standardised one-hot ridge regression with alpha 1 on the 25 values. Nothing was tuned, calibrated on held-out data, oversampled or class-balanced, and no run was retried on its outcome; stages keep their natural prevalence.

### Folds

Five folds within each cohort, assigned by a hash of each record before any outcome was seen. Every night is tested once per baseline, by a model fitted on the other four folds of its own cohort.

### Metrics

Computed per night, then averaged so each night counts once, however long. Balanced accuracy is the mean recall over the stages a night contains; macro F1 the mean F1 over the stages it contains or the model predicts. Cohen’s kappa is zero by construction for a prior that gives every epoch the same answer.

### Uncertainty

Pointwise 95% whole-night bootstrap within each cohort: 10,000 draws (PCG64, seed 20261003), the same draws for both baselines and every metric. Conditional on the fixed cross-validation predictions — refitting and fold assignment are not included — and not adjusted for multiple comparisons.

Held · no score

## Neural-network and foundation models: held

No neural-network or foundation-model score exists for the Dreem cohorts. The stored signal metadata says millivolts, while the publisher’s own converter treats the same arrays as microvolts. Foundation models and other amplitude-sensitive encoders are held because the source’s physical units disagree: they are not run on a guessed unit. The two baselines above divide each epoch and channel by a positive gain before computing relative spectra, so they do not depend on that unit. The hold says nothing about how accurate those models would be.

[Its row in the holds register →](https://bci.report/releases/#hold-dreem-amplitude-sensitive-models)

Methods & limits

## What these baselines can and cannot say

Two fixed classical baselines on CPU, in two separate experiments. They set a floor and show what accuracy hides; they are not a ranking, a clinical tool or a model of expert scoring.

### Two separate experiments, not a comparison

DOD-H and DOD-O have separate models, folds and intervals, and were recorded at different centres with different equipment. Their figures are not a controlled comparison of health status, and nothing here measures transfer from one cohort to the other.

### Not clinical, not a diagnosis

The target is the consensus the publisher stores, not a reconstruction of each scorer’s vote. Nothing here is a diagnosis, a clinical tool or evidence of equivalence with human experts.

### Physical units unresolved

The stored metadata says millivolts; the upstream converter treats the arrays as microvolts. Dividing each epoch and channel by a positive gain makes the relative spectra independent of that scale, but it does not establish that calibration, reference, clipping or filtering are equivalent.

### Uncalibrated probabilities

The ridge’s scores are not calibrated probabilities: no calibration, confidence or abstention claim is made, and no probability score is published.

### A floor, not a chance level

Balanced accuracy averages only the stages a night contains. The prior predicts N2 everywhere, so on a night that lacks one of the five stages its balanced accuracy is one quarter rather than one fifth; that is why it sits slightly above one fifth. Do not read it as a chance line.

### Not comparable with other papers

Natural stage prevalence, no class balancing or tuning, five particular derivations and these folds. Do not compare these figures with papers that use other channels, cohorts, preprocessing or splits.

### Dreem Open Datasets (DOD-H and DOD-O)

Antoine Guillot, Fabien Sauvet, Emmanuel H. During and Valentin Thorey · Dreem Open Datasets: Multi-Scored Sleep Datasets to Compare Human and Automated Sleep Staging, IEEE Transactions on Neural Systems and Rehabilitation Engineering (2020), doi:10.1109/TNSRE.2020.3011181; preprint arXiv:1911.03221. Data: Dreem Open Datasets, Zenodo, doi:10.5281/zenodo.15900394. Official repository: github.com/Dreem-Organization/dreem-learning-open, revision 8b332d6827f5ae6a22f4bb97b4deef4273238ec3.

Deposit licence: MIT, as declared on the Zenodo record. It is recorded as the deposit’s declaration, not as a statement about every future use; BCI Report redistributes no upstream code or data.

[Paper ↗](https://doi.org/10.1109/TNSRE.2020.3011181) · [Preprint ↗](https://arxiv.org/abs/1911.03221) · [Zenodo deposit ↗](https://zenodo.org/records/15900394) · [Official repository at the pinned revision ↗](https://github.com/Dreem-Organization/dreem-learning-open/tree/8b332d6827f5ae6a22f4bb97b4deef4273238ec3) · [MIT](https://opensource.org/licenses/MIT)

### Consent and ethics, cohort by cohort

**DOD-H** · Ethics approval: Approved by the Committees of Protection of Persons (CPP); declared to the French National Agency for Medicines and Health Products Safety; carried out in compliance with the French Data Protection Act, International Conference on Harmonization (ICH) standards and the Declaration of Helsinki (1964, as revised in 2013). Informed consent: not stated. The paper's DOD-H description contains no informed-consent sentence.

**DOD-O** · Ethics approval: not stated. No ethics committee or institutional review board is named for DOD-O in the paper. Informed consent: All trial participants gave informed written consent before taking part.

Each cohort is published as it stands: one has an ethics approval and no consent sentence, the other written consent and no named committee. Source of these statements: arXiv:1911.03221v4 (27 April 2020), II. Materials and Methods, A. Datasets, read 2026-10-03.

Data source: [large-source-update.json](https://bci.report/data/large-source-update.json) · schema bci-report-large-source-update-v1 · generated 2026-10-03.

Keep exploring

Transfer

- [— Sensor transfer **Dry vs. wet electrodes**](https://bci.report/topics/dry-vs-wet/)
- [— Context transfer **Screen to VR**](https://bci.report/topics/screen-to-vr/)
- [— Montage **Fewer electrodes**](https://bci.report/topics/fewer-electrodes/)
- [— Motion robustness **On the move**](https://bci.report/topics/on-the-move/)

Adapting models

- [— Calibration budget **How much calibration?**](https://bci.report/topics/calibration-budget/)
- [— Model adaptation **Which part to update?**](https://bci.report/topics/model-adaptation/)
- [— Representation controls **Does pretraining help?**](https://bci.report/topics/does-pretraining-help/)

Reliability & clinical

- [— Abstention · Jev-style **When not to act**](https://bci.report/topics/when-not-to-act/)
- [— Sleep staging · simple baselines **Sleep-stage balance**](https://bci.report/topics/sleep-staging/)
- [— Clinical research **Clinical groups**](https://bci.report/topics/clinical-groups/)

[All questions, and the map of which kinds of transfer have been measured →](https://bci.report/topics/)

## Cite this page

BCI Report (2026). *Why can a sleep stager be right most of the time and still miss whole stages?* https://bci.report/topics/sleep-staging/

Figures from release [`large-source-update-20261003`](https://bci.report/releases/#large-source-update-20261003) (2026-10-03). Cite the upstream datasets as well: their credits are on this page.

Every release is archived on Zenodo: [doi:10.5281/zenodo.23123296](https://doi.org/10.5281/zenodo.23123296). [BibTeX for the site and its releases →](https://bci.report/api/#cite-heading) · [CITATION.cff ↗](https://raw.githubusercontent.com/twu3202/bci-report/main/CITATION.cff)

---
Markdown copy of https://bci.report/topics/sleep-staging/, generated from the published page. Figures are aggregate results; terms of use: https://bci.report/data-use/
