# Dreem Open Datasets (DOD-H and DOD-O): EEG decoding results

Five-stage sleep staging, healthy sleepers and people with obstructive sleep apnoea, kept apart

The Dreem Open Datasets are two overnight polysomnography cohorts released to compare human and automated sleep staging: DOD-H, 25 healthy volunteers recorded at the French Armed Forces Biomedical Research Institute (IRBA), and DOD-O, 55 patients with obstructive sleep apnoea recorded at the Stanford Sleep Medicine Center. Both were scored by five sleep technologists from different sleep centres. Described by Guillot, Sauvet, During and Thorey in IEEE Transactions on Neural Systems and Rehabilitation Engineering (2020), with the official dreem-learning-open code; archived on Zenodo (record 15900394, published July 2025).

**Also known as** Dreem Open Datasets · DOD-H · DOD-O · DOD · Zenodo 15900394 · Guillot et al. (2020)

Description sources [arxiv.org](https://arxiv.org/abs/1911.03221) · [doi.org](https://doi.org/10.1109/TNSRE.2020.3011181) · [zenodo.org](https://zenodo.org/records/15900394) · [github.com](https://github.com/Dreem-Organization/dreem-learning-open)

## Where it appears

- [Why can a sleep stager be right most of the time and still miss whole stages?](https://bci.report/topics/sleep-staging/)

## Published results

Every figure below is copied from a reviewed download, not recomputed for this page. Read each group with its protocol on the linked page.

### Sleep staging · DOD-H · healthy sleepers (a separate experiment)

Read with its protocol: [Sleep-stage balance](https://bci.report/topics/sleep-staging/) · [large-source-update.json](https://bci.report/data/large-source-update.json)

**Balanced accuracy, 2 configurations.** Dot: the mean over nights; line: 95% interval. No dashed line: the training prior is a floor this baseline sets, not a chance level.

- Training prior: 20.2% (20.0%–20.6%)
- Spectral ridge: 56.7% (53.1%–60.0%)

| Method | Condition | Metric | Value | People |
| --- | --- | --- | --- | --- |
| Training prior — reference or comparator, not a decoding model | DOD-H · healthy sleepers, every night held out once | Balanced accuracy | 20.2% (20.0%–20.6%) | 25 |
| Training prior — reference or comparator, not a decoding model | DOD-H · healthy sleepers, every night held out once | Accuracy | 48.0% (44.6%–51.4%) | 25 |
| Training prior — reference or comparator, not a decoding model | DOD-H · healthy sleepers, every night held out once | Macro F1 | 13.0% (12.3%–13.8%) | 25 |
| Training prior — reference or comparator, not a decoding model | DOD-H · healthy sleepers, every night held out once | Cohen’s kappa | 0.000 (0.000–0.000) | 25 |
| Spectral ridge | DOD-H · healthy sleepers, every night held out once | Balanced accuracy | 56.7% (53.1%–60.0%) | 25 |
| Spectral ridge | DOD-H · healthy sleepers, every night held out once | Accuracy | 72.3% (67.8%–76.3%) | 25 |
| Spectral ridge | DOD-H · healthy sleepers, every night held out once | Macro F1 | 53.3% (49.0%–57.3%) | 25 |
| Spectral ridge | DOD-H · healthy sleepers, every night held out once | Cohen’s kappa | 0.584 (0.526–0.636) | 25 |
| Spectral ridge | Spectral ridge minus training prior, the same nights | Paired difference, balanced accuracy | +36.5 pp (+32.9 pp–+39.8 pp) — 25 nights under both baselines; the interval excludes zero. | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage Wake | Recall | 64.8% (54.9%–73.9%) — 3,037 Wake epochs in the consensus. | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage Wake | Precision | 76.4% (70.1%–82.3%) | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage Wake | F1 | 66.5% (58.5%–73.8%) | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage N1 | Recall | 0.0% (0.0%–0.0%) — 1,505 N1 epochs in the consensus; never predicted on any night, so every one is missed and N1 precision is not defined (not zero). | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage N1 | F1 | 0.0% (0.0%–0.0%) | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage N2 | Recall | 85.4% (77.4%–91.8%) — 11,879 N2 epochs in the consensus. | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage N2 | Precision | 77.3% (72.8%–81.8%) | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage N2 | F1 | 78.9% (74.0%–83.1%) | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage N3 | Recall | 59.4% (46.8%–71.2%) — 3,514 N3 epochs in the consensus; defined on 24 of 25 nights. | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage N3 | Precision | 69.7% (55.7%–81.9%) — defined on 24 of 25 nights. | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage N3 | F1 | 56.2% (43.6%–67.7%) | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage REM | Recall | 73.6% (63.8%–82.6%) — 4,727 REM epochs in the consensus. | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage REM | Precision | 63.8% (55.4%–71.5%) | 25 |
| Spectral ridge | DOD-H · healthy sleepers · stage REM | F1 | 64.6% (56.1%–72.2%) | 25 |

### Sleep staging · DOD-O · people with obstructive sleep apnoea (a separate experiment)

Read with its protocol: [Sleep-stage balance](https://bci.report/topics/sleep-staging/) · [large-source-update.json](https://bci.report/data/large-source-update.json)

**Balanced accuracy, 2 configurations.** Dot: the mean over nights; line: 95% interval. No dashed line: the training prior is a floor this baseline sets, not a chance level.

- Training prior: 20.5% (20.1%–21.2%)
- Spectral ridge: 49.3% (47.3%–51.2%)

| Method | Condition | Metric | Value | People |
| --- | --- | --- | --- | --- |
| Training prior — reference or comparator, not a decoding model | DOD-O · people with obstructive sleep apnoea, every night held out once | Balanced accuracy | 20.5% (20.1%–21.2%) | 55 |
| Training prior — reference or comparator, not a decoding model | DOD-O · people with obstructive sleep apnoea, every night held out once | Accuracy | 49.2% (46.3%–52.2%) | 55 |
| Training prior — reference or comparator, not a decoding model | DOD-O · people with obstructive sleep apnoea, every night held out once | Macro F1 | 13.3% (12.8%–13.9%) | 55 |
| Training prior — reference or comparator, not a decoding model | DOD-O · people with obstructive sleep apnoea, every night held out once | Cohen’s kappa | 0.000 (0.000–0.000) | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea, every night held out once | Balanced accuracy | 49.3% (47.3%–51.2%) | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea, every night held out once | Accuracy | 72.5% (69.7%–75.0%) | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea, every night held out once | Macro F1 | 47.4% (44.9%–49.7%) | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea, every night held out once | Cohen’s kappa | 0.542 (0.504–0.578) | 55 |
| Spectral ridge | Spectral ridge minus training prior, the same nights | Paired difference, balanced accuracy | +28.8 pp (+26.7 pp–+30.7 pp) — 55 nights under both baselines; the interval excludes zero. | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage Wake | Recall | 80.4% (75.5%–85.1%) — 10,427 Wake epochs in the consensus. | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage Wake | Precision | 76.5% (72.0%–80.7%) | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage Wake | F1 | 75.6% (71.7%–79.3%) | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage N1 | Recall | 0.0% (0.0%–0.0%) — 2,860 N1 epochs in the consensus; never predicted on any night, so every one is missed and N1 precision is not defined (not zero). | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage N1 | F1 | 0.0% (0.0%–0.0%) | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage N2 | Recall | 92.8% (90.5%–94.7%) — 26,271 N2 epochs in the consensus. | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage N2 | Precision | 71.0% (67.9%–74.0%) | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage N2 | F1 | 79.8% (77.3%–82.0%) | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage N3 | Recall | 22.1% (15.4%–29.2%) — 5,500 N3 epochs in the consensus; defined on 52 of 55 nights. | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage N3 | Precision | 63.4% (52.4%–73.9%) — defined on 49 of 55 nights. | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage N3 | F1 | 27.9% (20.3%–35.8%) — defined on 53 of 55 nights. | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage REM | Recall | 49.5% (43.1%–55.9%) — 8,103 REM epochs in the consensus; defined on 53 of 55 nights. | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage REM | Precision | 68.2% (60.9%–74.9%) — defined on 54 of 55 nights. | 55 |
| Spectral ridge | DOD-O · people with obstructive sleep apnoea · stage REM | F1 | 52.5% (46.0%–58.7%) — defined on 54 of 55 nights. | 55 |

## Source and licence

### Dreem Open Datasets (DOD-H and DOD-O)

**Credit** Antoine Guillot, Fabien Sauvet, Emmanuel H. During and Valentin Thorey · Dreem Open Datasets: Multi-Scored Sleep Datasets to Compare Human and Automated Sleep Staging, IEEE Transactions on Neural Systems and Rehabilitation Engineering (2020), doi:10.1109/TNSRE.2020.3011181; preprint arXiv:1911.03221. Data: Dreem Open Datasets, Zenodo, doi:10.5281/zenodo.15900394. Official repository: github.com/Dreem-Organization/dreem-learning-open, revision 8b332d6827f5ae6a22f4bb97b4deef4273238ec3.

[MIT ↗](https://opensource.org/licenses/MIT) · [Dataset record ↗](https://zenodo.org/records/15900394)

BCI Report does not redistribute any recording. These are aggregate measurements computed by BCI Report under the licence above; the data belong to the people credited.

[All datasets →](https://bci.report/datasets/) · [All methods →](https://bci.report/methods/)

## Cite this page

BCI Report (2026). *Dreem Open Datasets (DOD-H and DOD-O): EEG decoding results*. https://bci.report/datasets/dreem-dod/

Figures from release [`large-source-update-20261003`](https://bci.report/releases/#large-source-update-20261003) (2026-10-03). Cite the upstream dataset as well: its credit is on this page.

Every release is archived on Zenodo: [doi:10.5281/zenodo.23123296](https://doi.org/10.5281/zenodo.23123296). [BibTeX for the site and its releases →](https://bci.report/api/#cite-heading) · [CITATION.cff ↗](https://raw.githubusercontent.com/twu3202/bci-report/main/CITATION.cff)

---
Markdown copy of https://bci.report/datasets/dreem-dod/, generated from the published page. Figures are aggregate results; terms of use: https://bci.report/data-use/
