# Can one EEG model answer several questions about the same data as well as separate models, and at what cost?

Shared encoders · route 2 of the decision-research plan

One encoder answering several questions about the same windows, against a separate model per question and against a head told which question it is answering, at matched data and compute. Motor imagery on OpenBMI and sleep on BOAS, with EESM19 as a crude replication.

- [Reviewed aggregate JSON ↓](https://bci.report/data/shared-representation-update.json)
- 

Measured on: [EESM19 scalp subset](https://bci.report/datasets/eesm19/); [OpenBMI motor imagery (Lee et al. 2019)](https://bci.report/datasets/openbmi/); [BOAS · Bitbrain Open Access Sleep dataset](https://bci.report/datasets/boas/) · Methods: [EEGNet](https://bci.report/methods/eegnet/); [CBraMod](https://bci.report/methods/cbramod/)

## Short answer

The cheaper set-up was not shown to do as well, on these data. A shared EEGNet with a fixed linear output per question cost less than a head told which question it was answering (**3,060** against **20,116** trainable parameters on motor imagery), but it was not shown to come within 2 pp of that head, the margin fixed before the run. On imagery against rest, in **51** people (OpenBMI), that head scored **+1.48 pp** higher (**+0.09** to **+2.89 pp**); on five-stage sleep scoring (BOAS) its whole interval lay above 2 pp. Naming the question added nothing measurable: the same hidden layer, not told the question, was equivalent to that head within ±2 pp on both motor-imagery questions and non-inferior at 2 pp on sleep. Against a separate EEGNet per question, the small shared EEGNet trunk scored **−1.43 pp** on which hand (**−2.57** to **−0.39 pp**): a difference, and a loss of 2 pp or more cannot be ruled out; on imagery or rest and on the sleep stage the comparison was inconclusive.

How the comparison works

## Four set-ups, the same questions

Every question is a label the dataset already has, asked by its identifier: the model is told which question it is answering, never given it in words. The four set-ups answer the same questions from the same windows.

- **A · Separate models**: one model per question.
- **B-lin · Fixed heads**: one shared encoder and one linear output per question; the fewest parameters.
- **B-sh · Shared hidden layer**: the shared encoder, then one hidden layer used by every question, then one output per question.
- **C1 · Question-conditioned head**: B-sh, with the hidden layer adjusted for each question by a learned code (FiLM).

B-lin, B-sh and C1 start from the same weights, see the same batches in the same order and train for the same number of steps. Two primary levels, three seeds each: EEGNet trained from scratch (E1), and heads trained on frozen CBraMod features (L1; the same set-ups, marked -fz). CBraMod adapted by LoRA (E2) ran with one seed, as a secondary level.

Three comparisons answer the question. P1, C1 − B-lin: do fixed heads do as well? P5, C1 − B-sh: does telling the head the question add anything? P2, B-lin − A: does sharing one encoder cost accuracy? P4 is P1 on frozen CBraMod features.

Each comparison is a paired difference in balanced accuracy, in percentage points (pp), with a 95% interval from resampling people (and, with three seeds, training runs). It gets two readings. A difference is shown when the interval excludes zero. And against a margin of ±2 pp, fixed by the owner before any result, it is equivalent when the whole interval lies inside the margin, non-inferior when the bound that matters does, and otherwise the margin is not met. “Fixed heads do as well for less” is said of a domain only if every counted question meets the margin on P1 and the fixed heads cost less. A question whose fixed-head accuracy is not clearly above chance + 5 pp is at floor: reported, never counted.

There are 21 primary comparisons and no multiplicity correction, so about one in 20 comparisons with no true difference may show one by chance. Balanced accuracy is averaged per person, then over people: it compares set-ups and is not a deployment error rate. This is route 2 of the research plan on [When not to act](https://bci.report/topics/when-not-to-act/#decision-research).

Primary · OpenBMI · EEGNet and CBraMod

## Motor imagery: imagery or rest, and which hand

Two questions about each 2-second window of OpenBMI’s motor-imagery runs and rest recordings: is the person imagining a hand movement or resting (MI-A), and which hand (MI-B, offline trials only)? 31,653 windows from 51 people; each person’s two sessions sit in one fold, so every test is on new people. Both questions passed the canary and the gate, so both count.

**+1.48 pp** C1 − B-lin on MI-A, EEGNet from scratch, three seeds

### Fixed heads did as well for less: not supported

On imagery or rest the conditioned head scored +1.48 pp above the fixed heads (+0.09 to +2.89 pp): a gap of 2 pp or more cannot be ruled out. On which hand the two were equivalent within ±2 pp (−0.09 pp, −1.24 to +1.04 pp). The fixed heads were the cheaper set-up: 3,060 against 20,116 trainable parameters, and 960 against 15,744 head multiply-accumulates per window.

Paired differences in balanced accuracy, in pp, with 95% intervals; 51 people, three seeds.

| Comparison | Question | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
| --- | --- | --- | --- | --- | --- |
| EEGNet trained from scratch (E1) |  |  |  |  |  |
| P1 · Conditioned head − fixed heads — C1 − B-lin | MI-A — imagery or rest | +1.48 pp (95% interval +0.09 to +2.89 pp) | C1 higher | 2 pp margin not met | Gate passed |
| P1 · Conditioned head − fixed heads — C1 − B-lin | MI-B — which hand | −0.09 pp (95% interval −1.24 to +1.04 pp) | No difference shown | Equivalent within ±2 pp | Gate passed |
| P5 · Conditioned head − the same layer, not told the question — C1 − B-sh | MI-A — imagery or rest | +0.31 pp (95% interval −0.91 to +1.54 pp) | No difference shown | Equivalent within ±2 pp | Gate passed |
| P5 · Conditioned head − the same layer, not told the question — C1 − B-sh | MI-B — which hand | +0.45 pp (95% interval −0.60 to +1.34 pp) | No difference shown | Equivalent within ±2 pp | Gate passed |
| P2 · Fixed heads on one small-CNN trunk − separate models — B-lin − A(MI-A) | MI-A — imagery or rest | −1.00 pp (95% interval −2.21 to +0.19 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size | Gate passed |
| P2 · Fixed heads on one small-CNN trunk − separate models — B-lin − A(MI-B) | MI-B — which hand | −1.43 pp (95% interval −2.57 to −0.39 pp) | A(MI-B) higher | 2 pp margin not met | Gate passed |
| Heads on frozen CBraMod features (L1) |  |  |  |  |  |
| P4 · Conditioned head − fixed heads — C1-fz − B-lin-fz-sgd | MI-A — imagery or rest | +1.54 pp (95% interval +0.86 to +2.30 pp) | C1-fz higher | 2 pp margin not met | Gate passed |
| P4 · Conditioned head − fixed heads — C1-fz − B-lin-fz-sgd | MI-B — which hand | +1.09 pp (95% interval +0.08 to +2.12 pp) | C1-fz higher | 2 pp margin not met — not read: at floor | At floor: reported, not counted |
| S1 · Conditioned head − the same layer, not told the question · secondary — C1-fz − B-sh-fz | MI-A — imagery or rest | −0.12 pp (95% interval −0.64 to +0.33 pp) | No difference shown | Equivalent within ±2 pp | Gate passed |
| S1 · Conditioned head − the same layer, not told the question · secondary — C1-fz − B-sh-fz | MI-B — which hand | +0.43 pp (95% interval −0.39 to +1.21 pp) | No difference shown | Equivalent within ±2 pp — not read: at floor | At floor: reported, not counted |

Telling the head the question added nothing measurable: C1 − B-sh was equivalent within ±2 pp on both questions from scratch, and on imagery or rest on frozen features. That does not show where the conditioned head’s lift over the fixed heads comes from: the same layer not told the question, against the fixed heads (B-sh − B-lin, secondary S2), was +1.17 pp on imagery or rest (−0.09 to +2.54 pp), inconclusive at this sample size. Against a separate EEGNet per question, sharing one small trunk (2,096 parameters) scored −1.43 pp on which hand (−2.57 to −0.39 pp): a difference, and a cost of 2 pp or more cannot be ruled out; on imagery or rest it was inconclusive. On frozen CBraMod features the conditioned head was again higher on imagery or rest (+1.54 pp, +0.86 to +2.30 pp); which hand is at floor there.

Balanced accuracy of each set-up: the mean over 51 people with its 95% interval. Chance is 50.0%; the gate’s floor is 55.0%. Overlapping intervals are not ranked: the paired differences above are the comparison.

| Set-up | MI-A · imagery or rest | MI-B · which hand |
| --- | --- | --- |
| EEGNet trained from scratch (E1) |  |  |
| Fixed heads — B-lin | 73.2% (95% interval 70.2%–76.2%) | 67.5% (95% interval 64.6%–70.5%) |
| Shared hidden layer — B-sh | 74.3% (95% interval 71.3%–77.3%) | 67.0% (95% interval 64.1%–69.9%) |
| Question-conditioned head — C1 | 74.6% (95% interval 71.9%–77.5%) | 67.4% (95% interval 64.4%–70.5%) |
| Separate models — A | 74.2% (95% interval 71.3%–76.9%) | 69.0% (95% interval 65.9%–72.0%) |
| Frozen CBraMod features (L1) |  |  |
| Fixed heads — B-lin-fz-sgd | 69.2% (95% interval 67.0%–71.5%) | 56.1% (95% interval 54.8%–57.4%) |
| Shared hidden layer — B-sh-fz | 70.9% (95% interval 68.4%–73.4%) | 56.7% (95% interval 55.4%–58.2%) |
| Question-conditioned head — C1-fz | 70.8% (95% interval 68.4%–73.2%) | 57.2% (95% interval 55.8%–58.6%) |

What each set-up costs on motor imagery (EEGNet, E1). Parameters, encoder passes, multiply-accumulates and steps are exact. The time per step was measured once, interleaved across set-ups on a shared GPU, and is indicative: in the cost rule a time counts only beyond 1.2×.

| Set-up | Trainable parameters | Encoder passes per window | Head multiply-accumulates per window | Training steps per fit | Time per step |
| --- | --- | --- | --- | --- | --- |
| Fixed heads — B-lin | 3,060 (encoder 2,096 · heads 964 · conditioning 0) | 1 | 960 | 7,928 | 7.71 ms |
| Shared hidden layer — B-sh | 18,402 (encoder 2,096 · heads 16,306 · conditioning 0) | 1 | 16,104 | 7,928 | 7.80 ms |
| Question-conditioned head — C1 | 20,116 (encoder 2,096 · heads 15,812 · conditioning 2,208) | 1 | 15,744 | 7,928 | 7.89 ms |
| Separate models — A | 5,156 (one EEGNet per question) | 2 | 960 | 10,492 | 7.55 ms |

### What the canary and the gate set aside

Before any fit, a ridge classifier on each window’s per-channel mean and log-variance, trained and tested within the first fold’s 40 training people, checked whether a question could be answered from those statistics alone; it demotes a question at 85.0% or more. It gave MI-A 66.1%, MI-B 55.9%, MI-C 60.1%, MI-D 47.3%, so none was demoted. The gate set one entry aside: on frozen CBraMod features the fixed heads scored 56.1% (54.8%–57.4%) on which hand; its interval does not lie wholly above the 55.0% floor, so that question is at floor there.

Primary · BOAS · EEGNet and CBraMod

## Sleep: the stage, the next change and the half of the night

Three questions about each 30-second epoch of the six polysomnography EEG channels, scored against the human consensus: the sleep stage (SL-A), whether the next epoch’s stage differs (SL-E), and the first or second half of the night (SL-F). 119,725 epochs from 100 people; all nights of a person sit in one fold. Stages at their natural mix: W 16.0%, N1 3.7%, N2 60.3%, N3 4.4%, REM 15.7%; the stage changes at the next epoch in 8.6% of epochs. SL-E and SL-F are largely predictable from the current stage: read off the true current stage alone, they reach 61.9% and 59.2%.

**BOAS is published with three stated gaps (owner decision, 7 October 2026):**

- The consent statement does not say whether participants agreed to public sharing or secondary use.
- The ethics and consent statements come from the publisher's dataset description and README. No peer-reviewed paper describes BOAS.
- The ethics reference was added to the release in version 1.1.1 (May 2025), and the release does not say when it was granted relative to the recordings.

Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.

**Credit:** Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

**+6.96 pp** C1 − B-lin on SL-A, EEGNet from scratch, three seeds

### Fixed heads did as well for less: not supported

Only the sleep stage counts. On EEGNet the fixed heads are at floor on the next change (54.4%, 53.5%–55.4%) and on the half of the night (54.2%, 53.0%–55.6%): reported, not counted. On the sleep stage the conditioned head scored +6.96 pp above the fixed heads (+4.92 to +9.09 pp): the whole interval lies above +2 pp. The fixed heads were cheaper: 34,905 against 243,817 trainable parameters, 33,696 against 240,384 head multiply-accumulates per epoch, and a shorter step.

Paired differences in balanced accuracy, in pp, with 95% intervals; 100 people, three seeds.

| Comparison | Question | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
| --- | --- | --- | --- | --- | --- |
| EEGNet trained from scratch (E1) |  |  |  |  |  |
| P1 · Conditioned head − fixed heads — C1 − B-lin | SL-A — sleep stage, five stages | +6.96 pp (95% interval +4.92 to +9.09 pp) | C1 higher | 2 pp margin not met | Gate passed |
| P1 · Conditioned head − fixed heads — C1 − B-lin | SL-E — next epoch’s stage differs | +3.02 pp (95% interval +1.55 to +4.35 pp) | C1 higher | 2 pp margin not met — not read: at floor | At floor: reported, not counted |
| P1 · Conditioned head − fixed heads — C1 − B-lin | SL-F — first or second half of the night | +1.42 pp (95% interval −0.25 to +2.90 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size — not read: at floor | At floor: reported, not counted |
| P5 · Conditioned head − the same layer, not told the question — C1 − B-sh | SL-A — sleep stage, five stages | −1.47 pp (95% interval −3.79 to +0.94 pp) | No difference shown | B-sh non-inferior at 2 pp | Gate passed |
| P5 · Conditioned head − the same layer, not told the question — C1 − B-sh | SL-E — next epoch’s stage differs | −1.41 pp (95% interval −2.40 to −0.50 pp) | B-sh higher | B-sh non-inferior at 2 pp — not read: at floor | At floor: reported, not counted |
| P5 · Conditioned head − the same layer, not told the question — C1 − B-sh | SL-F — first or second half of the night | −0.14 pp (95% interval −1.41 to +1.15 pp) | No difference shown | Equivalent within ±2 pp — not read: at floor | At floor: reported, not counted |
| P2 · Fixed heads on one small-CNN trunk − separate models — B-lin − A(SL-A) | SL-A — sleep stage, five stages | −0.84 pp (95% interval −2.73 to +0.91 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size | Gate passed |
| P2 · Fixed heads on one small-CNN trunk − separate models — B-lin − A(SL-E) | SL-E — next epoch’s stage differs | +1.20 pp (95% interval +0.07 to +2.44 pp) | B-lin higher | B-lin non-inferior at 2 pp — not read: at floor | At floor: reported, not counted |
| P2 · Fixed heads on one small-CNN trunk − separate models — B-lin − A(SL-F) | SL-F — first or second half of the night | +2.01 pp (95% interval +0.54 to +3.89 pp) | B-lin higher | B-lin non-inferior at 2 pp — not read: at floor | At floor: reported, not counted |
| Heads on frozen CBraMod features (L1) |  |  |  |  |  |
| P4 · Conditioned head − fixed heads — C1-fz − B-lin-fz-sgd | SL-A — sleep stage, five stages | +0.19 pp (95% interval −0.74 to +1.14 pp) | No difference shown | Equivalent within ±2 pp | Gate passed |
| P4 · Conditioned head − fixed heads — C1-fz − B-lin-fz-sgd | SL-E — next epoch’s stage differs | +1.75 pp (95% interval +0.92 to +2.63 pp) | C1-fz higher | 2 pp margin not met | Gate passed |
| P4 · Conditioned head − fixed heads — C1-fz − B-lin-fz-sgd | SL-F — first or second half of the night | +0.77 pp (95% interval −0.14 to +1.67 pp) | No difference shown | Equivalent within ±2 pp | Gate passed |
| S1 · Conditioned head − the same layer, not told the question · secondary — C1-fz − B-sh-fz | SL-A — sleep stage, five stages | −0.53 pp (95% interval −1.10 to +0.01 pp) | No difference shown | Equivalent within ±2 pp | Gate passed |
| S1 · Conditioned head − the same layer, not told the question · secondary — C1-fz − B-sh-fz | SL-E — next epoch’s stage differs | −0.22 pp (95% interval −0.76 to +0.30 pp) | No difference shown | Equivalent within ±2 pp | Gate passed |
| S1 · Conditioned head − the same layer, not told the question · secondary — C1-fz − B-sh-fz | SL-F — first or second half of the night | +0.05 pp (95% interval −0.57 to +0.70 pp) | No difference shown | Equivalent within ±2 pp | Gate passed |

On the sleep stage the gain is the hidden layer’s: the same layer not told the question scored +8.43 pp above the fixed heads (B-sh − B-lin, secondary S2; +5.83 to +11.12 pp), and C1 − B-sh was −1.47 pp (−3.79 to +0.94 pp), B-sh non-inferior at 2 pp; on frozen CBraMod features C1 − B-sh was equivalent within ±2 pp on all three questions. Sharing the trunk, against separate models, was inconclusive on the sleep stage. On frozen CBraMod features the fixed heads were equivalent within ±2 pp of the conditioned head on the sleep stage and the half of the night, not on the next change (+1.75 pp, +0.92 to +2.63 pp). Five-stage accuracy on the small EEGNet trunk is modest: 34.4% with fixed heads and 41.4% with the conditioned head, against a chance of 20.0%; fixed heads on frozen CBraMod features reach 70.7%.

Balanced accuracy of each set-up: the mean over 100 people with its 95% interval. Chance is 20.0% for the sleep stage and 50.0% for the other two; the gate’s floors are 25.0% and 55.0%. Overlapping intervals are not ranked.

| Set-up | SL-A · sleep stage, five stages | SL-E · next epoch’s stage differs | SL-F · first or second half of the night |
| --- | --- | --- | --- |
| EEGNet trained from scratch (E1) |  |  |  |
| Fixed heads — B-lin | 34.4% (95% interval 32.1%–37.0%) | 54.4% (95% interval 53.5%–55.4%) | 54.2% (95% interval 53.0%–55.6%) |
| Shared hidden layer — B-sh | 42.9% (95% interval 40.1%–45.7%) | 58.9% (95% interval 57.5%–60.2%) | 55.8% (95% interval 54.2%–57.4%) |
| Question-conditioned head — C1 | 41.4% (95% interval 38.6%–44.2%) | 57.5% (95% interval 56.0%–58.8%) | 55.6% (95% interval 54.2%–57.1%) |
| Separate models — A | 35.3% (95% interval 32.9%–37.8%) | 53.2% (95% interval 52.2%–54.3%) | 52.2% (95% interval 51.0%–53.7%) |
| Frozen CBraMod features (L1) |  |  |  |
| Fixed heads — B-lin-fz-sgd | 70.7% (95% interval 68.5%–72.7%) | 63.5% (95% interval 62.3%–64.7%) | 64.5% (95% interval 62.8%–66.1%) |
| Shared hidden layer — B-sh-fz | 71.4% (95% interval 69.1%–73.6%) | 65.5% (95% interval 64.1%–66.7%) | 65.2% (95% interval 63.6%–66.9%) |
| Question-conditioned head — C1-fz | 70.9% (95% interval 68.6%–73.0%) | 65.2% (95% interval 63.9%–66.5%) | 65.2% (95% interval 63.6%–66.9%) |

### A question that follows from another: does it need its own model?

Awake or asleep (SL-B) follows from the sleep stage. A separate wake-or-sleep model, A(SL-B), left less error than reading the answer off the five-stage model A(SL-A): its remaining error was 0.684 times the read-out’s. The interval does not rule out a reduction of 20% or more, the margin set before the run, so “a derivable question needs no model of its own” is not supported.

log R, the dedicated model’s remaining error (1 − AUROC) over the read-out’s, with its 95% interval; 100 people, three seeds.

| Comparison | log R | R | AUROC, dedicated / read-out | Difference shown? | Margin, 20% | Gate |
| --- | --- | --- | --- | --- | --- | --- |
| P3 · Its own model − reading it off the five-stage model — A(SL-B) / A(SL-A) · SL-B | −0.380 (95% interval −0.583 to −0.203) | 0.684 | 0.941 / 0.914 | Dedicated model leaves less error | 20% margin not met | Gate passed |

### CBraMod adapted by LoRA (E2): secondary, one seed, run last

This arm ran on 7 October 2026, after every other result of the run, and two independent audits, were known. Its design (35 LoRA fits, the seed, the recipe, the 30-second input as fifteen 2-second segments, the comparisons) and the condition for running it were fixed when the protocol was frozen, and nothing about it was chosen from results. It ran with temporary private checkpoints so that a fit could resume: during the run 39 were written for 33 fits, each replaced by its fit’s next one or deleted when its fit finished, so none outlived its fit, and 0 remain; a check of resuming, run before the fits, removed its own the same way. Its scoring ran the frozen code, and an independent audit recomputed all 9 entries.

Paired differences in balanced accuracy, in pp, with 95% intervals; 100 people, one seed: no training-run variance in these intervals.

| Comparison | Question | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
| --- | --- | --- | --- | --- | --- |
| S3-P1 · Conditioned head − fixed heads — C1 − B-lin | SL-A — sleep stage, five stages | +0.54 pp (95% interval −0.41 to +1.51 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the E2-sleep scorer, recomputed by its audit |
| S3-P1 · Conditioned head − fixed heads — C1 − B-lin | SL-E — next epoch’s stage differs | +3.00 pp (95% interval +2.20 to +3.81 pp) | C1 higher | 2 pp margin not met | Gate passed — attached by the E2-sleep scorer, recomputed by its audit |
| S3-P1 · Conditioned head − fixed heads — C1 − B-lin | SL-F — first or second half of the night | +0.80 pp (95% interval −0.19 to +1.79 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the E2-sleep scorer, recomputed by its audit |
| S3-P5 · Conditioned head − the same layer, not told the question — C1 − B-sh | SL-A — sleep stage, five stages | +0.15 pp (95% interval −0.15 to +0.44 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the E2-sleep scorer, recomputed by its audit |
| S3-P5 · Conditioned head − the same layer, not told the question — C1 − B-sh | SL-E — next epoch’s stage differs | +0.15 pp (95% interval −0.13 to +0.44 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the E2-sleep scorer, recomputed by its audit |
| S3-P5 · Conditioned head − the same layer, not told the question — C1 − B-sh | SL-F — first or second half of the night | +0.03 pp (95% interval −0.23 to +0.29 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the E2-sleep scorer, recomputed by its audit |
| S3-P2 · Fixed heads on one adapted CBraMod − separate models — B-lin − A(SL-A) | SL-A — sleep stage, five stages | +0.10 pp (95% interval −0.50 to +0.69 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the E2-sleep scorer, recomputed by its audit |
| S3-P2 · Fixed heads on one adapted CBraMod − separate models — B-lin − A(SL-E) | SL-E — next epoch’s stage differs | −0.70 pp (95% interval −1.51 to +0.10 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the E2-sleep scorer, recomputed by its audit |
| S3-P2 · Fixed heads on one adapted CBraMod − separate models — B-lin − A(SL-F) | SL-F — first or second half of the night | −0.31 pp (95% interval −1.59 to +0.98 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the E2-sleep scorer, recomputed by its audit |

With CBraMod adapted, the fixed heads were equivalent within ±2 pp of the conditioned head on the sleep stage and the half of the night; on the next change the conditioned head was +3.00 pp higher (+2.20 to +3.81 pp). Shared against separate models: equivalent within ±2 pp on all three. Gate arm, fixed heads at E2: SL-A 72.8% (70.9%–74.6%), SL-E 65.0% (63.8%–66.2%), SL-F 64.9% (63.3%–66.3%); all passed.

What each set-up costs on sleep (EEGNet, E1). Parameters, encoder passes, multiply-accumulates and steps are exact. The time per step was measured once, interleaved across set-ups on a shared GPU, and is indicative: in the cost rule a time counts only beyond 1.2×.

| Set-up | Trainable parameters | Encoder passes per epoch | Head multiply-accumulates per epoch | Training steps per fit | Time per step |
| --- | --- | --- | --- | --- | --- |
| Fixed heads — B-lin | 34,905 (encoder 1,200 · heads 33,705 · conditioning 0) | 1 | 33,696 | 19,982 | 3.24 ms |
| Shared hidden layer — B-sh | 241,593 (encoder 1,200 · heads 240,393 · conditioning 0) | 1 | 240,192 | 19,982 | 3.56 ms |
| Question-conditioned head — C1 | 243,817 (encoder 1,200 · heads 240,393 · conditioning 2,224) | 1 | 240,384 | 19,982 | 4.04 ms |
| Separate models — A | 37,305 (one EEGNet per question) | 3 | 33,696 | 59,916 | 2.60 ms |

### What the canary and the gate set aside

The canary, on the first fold’s 80 training people, gave SL-A 36.8%, SL-B 59.9%, SL-C 70.7%, SL-D 82.1%, SL-E 56.5%, SL-F 58.4%, all below 85.0%, so none was demoted. The gate set aside the next change and the half of the night on EEGNet; every question passed on frozen CBraMod features and with CBraMod adapted.

Secondary · crude replication · EESM19 · one seed

## The same comparison on EESM19, crude

Every stored 30-second epoch of the full EESM19 release (first scorer) that passes the published core rules, at its natural stage mix: 52,311 epochs from 20 people. EEGNet from scratch, one seed, and no shared-hidden-layer arm, so the question code cannot be told apart from the hidden layer here. A different protocol from the balanced scalp subset in the core matrix. SL-E and SL-F are largely predictable from the current stage; no stage-only read-out was computed here.

Paired differences in balanced accuracy, in pp, with 95% intervals; 20 people, one seed.

| Comparison | Question | The two arms’ means | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
| --- | --- | --- | --- | --- | --- | --- |
| S13-P1 · Conditioned head − fixed heads — C1 − B-lin | SL-A — sleep stage, five stages | 66.7% / 63.2% | +3.47 pp (95% interval +1.54 to +5.64 pp) | C1 higher | 2 pp margin not met | Gate passed — attached by the run |
| S13-P1 · Conditioned head − fixed heads — C1 − B-lin | SL-E — next epoch’s stage differs | 56.8% / 53.4% | +3.38 pp (95% interval +2.27 to +4.36 pp) | C1 higher | 2 pp margin not met — not read: at floor | At floor: reported, not counted — attached by the run |
| S13-P1 · Conditioned head − fixed heads — C1 − B-lin | SL-F — first or second half of the night | 63.8% / 61.8% | +2.09 pp (95% interval +0.38 to +3.74 pp) | C1 higher | 2 pp margin not met | Gate passed — attached by the run |
| S13-P2 · Fixed heads on one small-CNN trunk − separate models — B-lin − A(SL-A) | SL-A — sleep stage, five stages | 63.2% / 64.7% | −1.47 pp (95% interval −2.90 to −0.10 pp) | A(SL-A) higher | 2 pp margin not met | Gate passed — attached by the run |
| S13-P2 · Fixed heads on one small-CNN trunk − separate models — B-lin − A(SL-E) | SL-E — next epoch’s stage differs | 53.4% / 52.1% | +1.33 pp (95% interval +0.42 to +2.32 pp) | B-lin higher | B-lin non-inferior at 2 pp — not read: at floor | At floor: reported, not counted — attached by the run |
| S13-P2 · Fixed heads on one small-CNN trunk − separate models — B-lin − A(SL-F) | SL-F — first or second half of the night | 61.8% / 62.5% | −0.72 pp (95% interval −2.10 to +0.69 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size | Gate passed — attached by the run |

On the sleep stage the direction matches BOAS: the conditioned head above the fixed heads (+3.47 pp, +1.54 to +5.64 pp), margin not met. The canary, on 16 training people, demoted nothing; the gate set the next change aside.

Secondary · exploratory

## Secondary and exploratory results

Further arms on motor imagery and sleep, one seed unless marked: what training on every question at once changes, longer training, other encoders and other heads. Each carries its gate and who attached it.

Show the secondary results

Codes: K-all, every question of the dataset trained together (MI-C and MI-D added on OpenBMI; SL-B, SL-C and SL-D on BOAS); primary K, the primary questions only; 2x, twice the training epochs; step-matched, A(MI-B) trained for the shared set-ups’ number of steps; B-MLP64-fz, a separate 64-unit hidden layer per question; H, a hypernetwork over learned question codes (exploratory: equivalent to linear heads by construction); C2-fz, query tokens over CBraMod’s token features. Each gate says who attached it: the run, the level’s gate, or the independent secondary audit, which attached the gates the frozen scoring left out. S2 on frozen features was computed by the independent secondary audit, not by the run.

Motor imagery, OpenBMI

| Entry | Level | Question | Contrast | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
| --- | --- | --- | --- | --- | --- | --- | --- |
| S2 — three seeds | E1 | MI-A — imagery or rest | B-sh − B-lin | +1.17 pp (95% interval −0.09 to +2.54 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size | Gate passed — attached by the independent secondary audit |
| S2 — three seeds | E1 | MI-B — which hand | B-sh − B-lin | −0.54 pp (95% interval −1.65 to +0.74 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S10 | E1 | MI-A — imagery or rest | B-lin(K-all) − B-lin(primary) | −2.32 pp (95% interval −3.94 to −0.75 pp) | B-lin(primary) higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S10 | E1 | MI-B — which hand | B-lin(K-all) − B-lin(primary) | −1.84 pp (95% interval −3.26 to −0.41 pp) | B-lin(primary) higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S10 | E1 | MI-A — imagery or rest | C1(K-all) − C1(primary) | −1.12 pp (95% interval −2.47 to +0.21 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size | Gate passed — attached by the independent secondary audit |
| S10 | E1 | MI-B — which hand | C1(K-all) − C1(primary) | −2.44 pp (95% interval −4.75 to −0.24 pp) | C1(primary) higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S5 | E1 | MI-A — imagery or rest | C1(2x) − B-lin(2x) | +0.89 pp (95% interval −0.55 to +2.59 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size | Gate passed — attached by the independent secondary audit |
| S5 | E1 | MI-B — which hand | C1(2x) − B-lin(2x) | −0.24 pp (95% interval −1.43 to +0.98 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S6-P1 | E1 | MI-C — offline or online run | C1(K-all) − B-lin(K-all) | −0.15 pp (95% interval −2.86 to +2.50 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size | Gate passed — attached by the independent secondary audit |
| S6-P2 | E1 | MI-C — offline or online run | B-lin(K-all) − A(MI-C) | −2.75 pp (95% interval −5.43 to −0.22 pp) | A(MI-C) higher | 2 pp margin not met | Gate passed — attached by the run |
| S6-P1 | E1 | MI-D — day 1 or day 2 | C1(K-all) − B-lin(K-all) | −0.69 pp (95% interval −4.17 to +2.64 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S6-P2 | E1 | MI-D — day 1 or day 2 | B-lin(K-all) − A(MI-D) | +0.49 pp (95% interval −3.23 to +3.81 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size — not read: at floor | At floor: reported, not counted — attached by the run |
| S12 | E1 | MI-B — which hand | A(MI-B), step-matched − A(MI-B) | −0.03 pp (95% interval −1.63 to +1.44 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S12-P2 | E1 | MI-B — which hand | B-lin − A(MI-B), step-matched | −0.57 pp (95% interval −2.09 to +1.05 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size | Gate passed — attached by the independent secondary audit |
| S3-P1 | E2 | MI-A — imagery or rest | C1 − B-lin | −0.43 pp (95% interval −1.06 to +0.18 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S3-P2 | E2 | MI-A — imagery or rest | B-lin − A(MI-A) | −0.09 pp (95% interval −0.58 to +0.39 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S3-P5 | E2 | MI-A — imagery or rest | C1 − B-sh | −0.07 pp (95% interval −0.36 to +0.23 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S3-P1 | E2 | MI-B — which hand | C1 − B-lin | −1.07 pp (95% interval −2.08 to −0.08 pp) | B-lin higher | B-lin non-inferior at 2 pp | Gate passed — attached by the independent secondary audit |
| S3-P2 | E2 | MI-B — which hand | B-lin − A(MI-B) | +0.35 pp (95% interval −0.46 to +1.18 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S3-P5 | E2 | MI-B — which hand | C1 − B-sh | −0.16 pp (95% interval −0.71 to +0.42 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S7 | L1 · ST-EEGFormer Base, pooled | MI-A — imagery or rest | C1-fz − B-lin-fz-sgd | −1.11 pp (95% interval −1.89 to −0.32 pp) | B-lin-fz-sgd higher | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S7 | L1 · ST-EEGFormer Base, pooled | MI-B — which hand | C1-fz − B-lin-fz-sgd | −0.26 pp (95% interval −1.43 to +0.85 pp) | No difference shown | Equivalent within ±2 pp — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S14 | L1 · CBraMod, pooled | MI-A — imagery or rest | B-MLP64-fz − B-sh-fz | −0.08 pp (95% interval −0.66 to +0.49 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| X1 | L1 · CBraMod, pooled | MI-A — imagery or rest | H − B-lin-fz-sgd | −1.78 pp (95% interval −2.30 to −1.24 pp) | B-lin-fz-sgd higher | B-lin-fz-sgd non-inferior at 2 pp | Gate passed — attached by the independent secondary audit |
| S14 | L1 · CBraMod, pooled | MI-B — which hand | B-MLP64-fz − B-sh-fz | +0.59 pp (95% interval −0.32 to +1.57 pp) | No difference shown | Equivalent within ±2 pp — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| X1 | L1 · CBraMod, pooled | MI-B — which hand | H − B-lin-fz-sgd | −0.49 pp (95% interval −1.17 to +0.15 pp) | No difference shown | Equivalent within ±2 pp — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S8 | L1 · CBraMod tokens | MI-A — imagery or rest | C2-fz − B-lin-fz-sgd | −0.95 pp (95% interval −2.27 to +0.45 pp) | No difference shown | B-lin-fz-sgd non-inferior at 2 pp | Gate passed — attached by the independent secondary audit |
| S8 | L1 · CBraMod tokens | MI-B — which hand | C2-fz − B-lin-fz-sgd | −5.31 pp (95% interval −7.09 to −3.69 pp) | B-lin-fz-sgd higher | B-lin-fz-sgd non-inferior at 2 pp — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S2 — three seeds — computed by the independent audit | L1 | MI-A — imagery or rest | B-sh-fz − B-lin-fz-sgd | +1.66 pp (95% interval +0.95 to +2.44 pp) | B-sh-fz higher | 2 pp margin not met | Gate passed — the level’s gate |
| S2 — three seeds — computed by the independent audit | L1 | MI-B — which hand | B-sh-fz − B-lin-fz-sgd | +0.66 pp (95% interval −0.18 to +1.48 pp) | No difference shown | Equivalent within ±2 pp — not read: at floor | At floor: reported, not counted — the level’s gate |

**BOAS is published with three stated gaps (owner decision, 7 October 2026):**

- The consent statement does not say whether participants agreed to public sharing or secondary use.
- The ethics and consent statements come from the publisher's dataset description and README. No peer-reviewed paper describes BOAS.
- The ethics reference was added to the release in version 1.1.1 (May 2025), and the release does not say when it was granted relative to the recordings.

Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.

**Credit:** Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

Sleep, BOAS

| Entry | Level | Question | Contrast | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
| --- | --- | --- | --- | --- | --- | --- | --- |
| S2 — three seeds | E1 | SL-A — sleep stage, five stages | B-sh − B-lin | +8.43 pp (95% interval +5.83 to +11.12 pp) | B-sh higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S2 — three seeds | E1 | SL-E — next epoch’s stage differs | B-sh − B-lin | +4.43 pp (95% interval +3.10 to +5.71 pp) | B-sh higher | 2 pp margin not met — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S2 — three seeds | E1 | SL-F — first or second half of the night | B-sh − B-lin | +1.56 pp (95% interval −0.31 to +3.49 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S10 | E1 | SL-A — sleep stage, five stages | B-lin(K-all) − B-lin(primary) | −2.55 pp (95% interval −3.86 to −1.31 pp) | B-lin(primary) higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S10 | E1 | SL-E — next epoch’s stage differs | B-lin(K-all) − B-lin(primary) | +0.97 pp (95% interval +0.23 to +1.71 pp) | B-lin(K-all) higher | Equivalent within ±2 pp — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S10 | E1 | SL-F — first or second half of the night | B-lin(K-all) − B-lin(primary) | −0.61 pp (95% interval −1.37 to +0.16 pp) | No difference shown | Equivalent within ±2 pp — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S10 | E1 | SL-A — sleep stage, five stages | C1(K-all) − C1(primary) | +4.14 pp (95% interval +2.76 to +5.49 pp) | C1(K-all) higher | C1(K-all) non-inferior at 2 pp | Gate passed — attached by the independent secondary audit |
| S10 | E1 | SL-E — next epoch’s stage differs | C1(K-all) − C1(primary) | +0.96 pp (95% interval +0.11 to +1.79 pp) | C1(K-all) higher | Equivalent within ±2 pp — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S10 | E1 | SL-F — first or second half of the night | C1(K-all) − C1(primary) | +1.67 pp (95% interval +0.55 to +2.87 pp) | C1(K-all) higher | C1(K-all) non-inferior at 2 pp — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S5 | E1 | SL-A — sleep stage, five stages | C1(2x) − B-lin(2x) | +1.71 pp (95% interval +0.12 to +3.25 pp) | C1(2x) higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S5 | E1 | SL-E — next epoch’s stage differs | C1(2x) − B-lin(2x) | +2.13 pp (95% interval +1.27 to +2.93 pp) | C1(2x) higher | 2 pp margin not met — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S5 | E1 | SL-F — first or second half of the night | C1(2x) − B-lin(2x) | +2.78 pp (95% interval +1.72 to +3.94 pp) | C1(2x) higher | 2 pp margin not met — not read: at floor | At floor: reported, not counted — attached by the independent secondary audit |
| S6-P1 | E1 | SL-C — REM or NREM | C1(K-all) − B-lin(K-all) | +5.50 pp (95% interval +4.04 to +7.04 pp) | C1(K-all) higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S6-P2 | E1 | SL-C — REM or NREM | B-lin(K-all) − A(SL-C) | −9.59 pp (95% interval −11.77 to −7.58 pp) | A(SL-C) higher | 2 pp margin not met | Gate passed — attached by the run |
| S6-P1 | E1 | SL-D — N3 or not | C1(K-all) − B-lin(K-all) | +9.79 pp (95% interval +6.69 to +12.95 pp) | C1(K-all) higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S6-P2 | E1 | SL-D — N3 or not | B-lin(K-all) − A(SL-D) | −2.13 pp (95% interval −4.50 to +0.07 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size | Gate passed — attached by the run |
| S11 — three seeds | E1 | SL-E — next epoch’s stage differs | B-lin − read-out from B-lin | +4.35 pp (95% interval +3.46 to +5.28 pp) | Dedicated head higher | No margin: a descriptive contrast | No gate: descriptive — descriptive |
| S11 — three seeds | E1 | SL-E — next epoch’s stage differs | B-sh − read-out from B-sh | +8.26 pp (95% interval +6.91 to +9.52 pp) | Dedicated head higher | No margin: a descriptive contrast | No gate: descriptive — descriptive |
| S11 — three seeds | E1 | SL-E — next epoch’s stage differs | C1 − read-out from C1 | +6.67 pp (95% interval +5.35 to +7.90 pp) | Dedicated head higher | No margin: a descriptive contrast | No gate: descriptive — descriptive |
| S11 — three seeds | E1 | SL-E — next epoch’s stage differs | A(SL-E) − read-out from A(SL-A) | +3.09 pp (95% interval +2.02 to +4.14 pp) | Dedicated head higher | No margin: a descriptive contrast | No gate: descriptive — descriptive |
| S11 — three seeds | E1 | SL-F — first or second half of the night | B-lin − read-out from B-lin | −0.35 pp (95% interval −1.47 to +0.87 pp) | No difference shown | No margin: a descriptive contrast | No gate: descriptive — descriptive |
| S11 — three seeds | E1 | SL-F — first or second half of the night | B-sh − read-out from B-sh | −1.34 pp (95% interval −2.22 to −0.48 pp) | Read-out from the stage probabilities higher | No margin: a descriptive contrast | No gate: descriptive — descriptive |
| S11 — three seeds | E1 | SL-F — first or second half of the night | C1 − read-out from C1 | −1.24 pp (95% interval −2.03 to −0.48 pp) | Read-out from the stage probabilities higher | No margin: a descriptive contrast | No gate: descriptive — descriptive |
| S11 — three seeds | E1 | SL-F — first or second half of the night | A(SL-F) − read-out from A(SL-A) | −2.74 pp (95% interval −4.48 to −0.92 pp) | Read-out from the stage probabilities higher | No margin: a descriptive contrast | No gate: descriptive — descriptive |
| S7 | L1 · REVE Large, pooled | SL-A — sleep stage, five stages | C1-fz − B-lin-fz-sgd | −0.86 pp (95% interval −1.60 to −0.11 pp) | B-lin-fz-sgd higher | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S7 | L1 · REVE Large, pooled | SL-E — next epoch’s stage differs | C1-fz − B-lin-fz-sgd | +2.30 pp (95% interval +1.36 to +3.20 pp) | C1-fz higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S7 | L1 · REVE Large, pooled | SL-F — first or second half of the night | C1-fz − B-lin-fz-sgd | +0.08 pp (95% interval −0.79 to +0.98 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| S14 | L1 · CBraMod, pooled | SL-A — sleep stage, five stages | B-MLP64-fz − B-sh-fz | −0.17 pp (95% interval −0.89 to +0.54 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| X1 | L1 · CBraMod, pooled | SL-A — sleep stage, five stages | H − B-lin-fz-sgd | −2.72 pp (95% interval −3.40 to −2.09 pp) | B-lin-fz-sgd higher | B-lin-fz-sgd non-inferior at 2 pp | Gate passed — attached by the independent secondary audit |
| S14 | L1 · CBraMod, pooled | SL-E — next epoch’s stage differs | B-MLP64-fz − B-sh-fz | −0.96 pp (95% interval −1.71 to −0.23 pp) | B-sh-fz higher | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| X1 | L1 · CBraMod, pooled | SL-E — next epoch’s stage differs | H − B-lin-fz-sgd | −1.94 pp (95% interval −2.74 to −1.14 pp) | B-lin-fz-sgd higher | B-lin-fz-sgd non-inferior at 2 pp | Gate passed — attached by the independent secondary audit |
| S14 | L1 · CBraMod, pooled | SL-F — first or second half of the night | B-MLP64-fz − B-sh-fz | −0.33 pp (95% interval −1.13 to +0.53 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — attached by the independent secondary audit |
| X1 | L1 · CBraMod, pooled | SL-F — first or second half of the night | H − B-lin-fz-sgd | −1.50 pp (95% interval −2.29 to −0.73 pp) | B-lin-fz-sgd higher | B-lin-fz-sgd non-inferior at 2 pp | Gate passed — attached by the independent secondary audit |
| S8 | L1 · CBraMod tokens | SL-A — sleep stage, five stages | C2-fz − B-lin-fz-sgd | +1.64 pp (95% interval +0.58 to +2.70 pp) | C2-fz higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S8 | L1 · CBraMod tokens | SL-E — next epoch’s stage differs | C2-fz − B-lin-fz-sgd | +2.93 pp (95% interval +2.02 to +3.83 pp) | C2-fz higher | 2 pp margin not met | Gate passed — attached by the independent secondary audit |
| S8 | L1 · CBraMod tokens | SL-F — first or second half of the night | C2-fz − B-lin-fz-sgd | +1.05 pp (95% interval −0.03 to +2.11 pp) | No difference shown | 2 pp margin not met — inconclusive at this sample size | Gate passed — attached by the independent secondary audit |
| S2 — three seeds — computed by the independent audit | L1 | SL-A — sleep stage, five stages | B-sh-fz − B-lin-fz-sgd | +0.72 pp (95% interval −0.13 to +1.64 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — the level’s gate |
| S2 — three seeds — computed by the independent audit | L1 | SL-E — next epoch’s stage differs | B-sh-fz − B-lin-fz-sgd | +1.97 pp (95% interval +1.21 to +2.74 pp) | B-sh-fz higher | 2 pp margin not met | Gate passed — the level’s gate |
| S2 — three seeds — computed by the independent audit | L1 | SL-F — first or second half of the night | B-sh-fz − B-lin-fz-sgd | +0.72 pp (95% interval −0.01 to +1.51 pp) | No difference shown | Equivalent within ±2 pp | Gate passed — the level’s gate |

Questions that follow from the sleep stage (S4), log R

| Entry | Contrast | log R | R | AUROC, dedicated / read-out | Difference shown? | Margin, 20% | Gate |
| --- | --- | --- | --- | --- | --- | --- | --- |
| S4-SL-C — 100 people | SL-C · A(SL-C) / A(SL-A) | −0.824 (95% interval −0.980 to −0.675) | 0.438 | 0.908 / 0.789 | Dedicated model leaves less error | 20% margin not met | Gate passed — attached by the run |
| S4-SL-D — 71 people | SL-D · A(SL-D) / A(SL-A) | 0.126 (95% interval −0.002 to 0.251) | 1.134 | 0.975 / 0.978 | No difference shown | Read-out non-inferior at 20% | Gate passed — attached by the run |
| S4-within-B-lin(K-all) — 100 people | SL-B · B-lin(K-all) / B-lin(K-all) | −0.525 (95% interval −0.665 to −0.395) | 0.592 | 0.946 / 0.908 | Dedicated model leaves less error | 20% margin not met | Gate passed — attached by the run |
| S4-within-C1(K-all) — 100 people | SL-B · C1(K-all) / C1(K-all) | −0.088 (95% interval −0.144 to −0.031) | 0.916 | 0.946 / 0.941 | Dedicated model leaves less error | Equivalent within the 20% margin | Gate passed — attached by the run |

Coherence (descriptive): the share of test epochs where a dedicated wake, REM or N3 head contradicts the same set-up’s five-stage head — B-lin(K-all): SL-B 17.2%, SL-C 21.0%, SL-D 1.0%; C1(K-all): SL-B 9.8%, SL-C 12.9%, SL-D 1.2%; A(SL-B) against A(SL-A): 16.5%.

Methods & limits

## What these results can and cannot say

- The questions are label-backed classification targets from one dataset each, asked by identifier. They are not language prompts and say nothing about unseen questions (route 3).
- MI-A also separates a cue display and position in the run from imagery; MI-B is offline trials only; MI-C and MI-D are about the recording context, not brain states.
- SL-B, SL-C and SL-D are deterministic functions of the 5-stage label. Results on them are about head design, not about new information in the signal.
- SL-F is time of night, correlated with stage; SL-E concerns the next 30 s, not a prediction horizon in general.
- With identifier-only conditioning, FiLM can differ from the same head without it only by per-question thresholds and signs of the shared hidden units. A null result does not show that language conditioning cannot help.
- At E1 the shared encoder is a small CNN whose trunk has 1,200-2,096 parameters; P2 there is 'small-CNN trunk sharing', not foundation-model sharing; E2 on motor imagery ran at one seed, as a secondary arm.
- SL-E and SL-F are largely predictable from the current stage; their results are shown beside a stage-only readout.
- Questions answerable from window statistics or at floor or ceiling are reported but do not count for the route sentence.
- Costs were measured on a shared GPU and are indicative; parameter counts, encoder passes and multiply-accumulates are exact.
- Three datasets (EESM19 as a crude replication), one lab each, and one fixed recipe per encoder: this is not a ranking of encoders or of multi-task methods.
- BOAS has natural stage prevalence but is one population and one PSG system; balanced accuracy across people is not a deployment error rate.
- Per-person balanced accuracy averages people equally; it compares arms and is not a deployment error rate.
- The OpenBMI numbers here use route 2's own 2-s windows, label-free high-pass and person-disjoint folds; they are not comparable with the site's OpenBMI cross-session calibration figures.
- EESM19 full (S13) uses every stored epoch that passes the published core rules at its natural stage mix; it is a different protocol from the site's balanced sleep-scalp subset.
- On EESM19 (S13) no stage-only readout was computed (S11 is defined on BOAS only), so its SL-E and SL-F entries must carry the sentence that both questions are largely predictable from the current stage.
- E2 sleep ran after every other result was known; its design and run condition were fixed at the freeze, and nothing about it was chosen from results.

### Not run, and not reported

- **Ying 2025 multi-night sleep (S9)**: deferred by the owner before the freeze; nothing was run
- **E2 on motor imagery at three seeds**: not scheduled: E2 ran at one seed as a secondary arm, so the E1 sharing contrast keeps the label small-CNN trunk sharing
- **The ridge reference row on frozen CBraMod features (G-1)**: Declared in the frozen draft, never implemented, and not computed afterwards: it would be new fits, and each choice it leaves open would be made after every result is known. No decision uses it.
- **Arm-level results of the secondary arms (interval, AUROC, log loss, per-seed means)**: Not produced by the frozen stage 2; each secondary contrast carries the mean balanced accuracy of its two arms.
- **P3 at E2**: Not declared for the E2-sleep block, so its five wake-or-sleep fits feed no contrast.
- **Counts of people whose paired contrast is above or below zero**: Allowed by the BOAS review, not computed by the frozen stage 2, and not added afterwards.

### Checks

Three independent audits re-implemented the definitions without the study code and recomputed the figures from the frozen predictions: the primary audit passed 10 of 10 checks over 1,673 values, the secondary 10 of 10 over 971, and the E2-sleep audit 25 of 25; none found a mismatch.

### The recordings

### OpenBMI motor imagery (Lee et al. 2019)

Min-Ho Lee, O-Yeon Kwon, Yong-Jeong Kim, Hong-Kyung Kim, Young-Eun Lee, John Williamson, Siamac Fazli and Seong-Whan Lee · EEG dataset and OpenBMI toolbox for three BCI paradigms: an investigation into BCI illiteracy, GigaScience (2019), giz002, doi:10.1093/gigascience/giz002. Data: Supporting data, GigaScience Database, doi:10.5524/100542.

[Source ↗](https://doi.org/10.5524/100542) · [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/)

### BOAS, the Bitbrain Open Access Sleep dataset

Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

[Source ↗](https://openneuro.org/datasets/ds005555/versions/1.1.3) · [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/)

### EESM19, full release (first scorer)

Kaare B. Mikkelsen et al. · Accurate whole-night sleep monitoring with dry-contact ear-EEG (2019), doi:10.1038/s41598-019-53115-3; OpenNeuro ds005185 v1.0.2. Processed mirror: Zachary1150/EESM19-Processed.

[Source ↗](https://doi.org/10.18112/openneuro.ds005185.v1.0.2) · [CC0-1.0 declared by upstream and mirror](https://creativecommons.org/publicdomain/zero/1.0/)

### Methods these results draw on (titles checked against arXiv on 7 October 2026)

- [Yu & Yao, 2026 · Visual Jev: Accurate and Efficient Decisions from Shared Visual Context ↗](https://arxiv.org/abs/2609.25845)
- [Dai et al., 2026 · Towards Unified Multi-task EEG Analysis with Low-Rank Adaptation ↗](https://arxiv.org/abs/2604.25131)
- [Xiong et al., 2025 · EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models ↗](https://arxiv.org/abs/2508.17742)
- [Perez et al., 2017 · FiLM: Visual Reasoning with a General Conditioning Layer ↗](https://arxiv.org/abs/1709.07871)
- [Ha, Dai & Le, 2016 · HyperNetworks ↗](https://arxiv.org/abs/1609.09106)

Data source: [shared-representation-update.json](https://bci.report/data/shared-representation-update.json) · schema bci-report-shared-representation-update-v1 · generated 2026-10-07.

Keep exploring

Transfer

- [— Sensor transfer **Dry vs. wet electrodes**](https://bci.report/topics/dry-vs-wet/)
- [— Context transfer **Screen to VR**](https://bci.report/topics/screen-to-vr/)
- [— Montage **Fewer electrodes**](https://bci.report/topics/fewer-electrodes/)
- [— Motion robustness **On the move**](https://bci.report/topics/on-the-move/)

Adapting models

- [— Calibration budget **How much calibration?**](https://bci.report/topics/calibration-budget/)
- [— Model adaptation **Which part to update?**](https://bci.report/topics/model-adaptation/)
- [— Representation controls **Does pretraining help?**](https://bci.report/topics/does-pretraining-help/)
- [— Shared encoder · route 2 **One model, several questions**](https://bci.report/topics/shared-encoder/)

Reliability & clinical

- [— Abstention · Jev-style **When not to act**](https://bci.report/topics/when-not-to-act/)
- [— Sleep staging · simple baselines **Sleep-stage balance**](https://bci.report/topics/sleep-staging/)
- [— Clinical research **Clinical groups**](https://bci.report/topics/clinical-groups/)

[All questions, and the map of which kinds of transfer have been measured →](https://bci.report/topics/)

## Cite this page

BCI Report (2026). *Can one EEG model answer several questions about the same data as well as separate models, and at what cost?* https://bci.report/topics/shared-encoder/

Figures from release [`shared-representation-update-20261007`](https://bci.report/releases/#shared-representation-update-20261007) (2026-10-07). Cite the upstream datasets as well: their credits are on this page.

Every release is archived on Zenodo: [doi:10.5281/zenodo.23123296](https://doi.org/10.5281/zenodo.23123296). [BibTeX for the site and its releases →](https://bci.report/api/#cite-heading) · [CITATION.cff ↗](https://raw.githubusercontent.com/twu3202/bci-report/main/CITATION.cff)

---
Markdown copy of https://bci.report/topics/shared-encoder/, generated from the published page. Figures are aggregate results; terms of use: https://bci.report/data-use/
