# Can an EEG model answer questions asked in words?

Questions in language · route 3 of the decision-research plan

One small head over frozen EEG features, asked each question by a question number, a label template or a natural-language description: SSVEP on BETA, sleep on BOAS. Seen questions, rewordings and questions the EEG head was never trained on are reported apart, never as one zero-shot number.

- [Reviewed aggregate JSON ↓](https://bci.report/data/questions-in-language-update.json)
- 

Measured on: [BETA](https://bci.report/datasets/beta/); [BOAS · Bitbrain Open Access Sleep dataset](https://bci.report/datasets/boas/) · Methods: [CBraMod](https://bci.report/methods/cbramod/)

## Short answer

Only for questions it was trained on, and not on every dataset. A small head over frozen EEG features, asked a seen question with a label template or a description instead of a question number, was equivalent within the 2 pp margin on sleep (BOAS), but lost accuracy on SSVEP (BETA, **70** people): **−5.96 pp** with templates on a plain spectrum (**−6.84** to **−5.08 pp**), the margin not met. Rewording a seen question cost accuracy too, except a new sentence frame on sleep. Questions the EEG head was never trained on were not answered usefully: every unseen sleep question did worse than adding up the head’s own stage answers, and on unseen flicker frequencies the head reached **28.7%** on a plain spectrum where CCA, which needs no training, reached **80.9%** (chance **12.5%**). Explicit wordings such as “N1 or N2 sleep” used the link between words and stages; the everyday names “asleep” and “light sleep” ranked windows backwards, and in a boundary probe negated questions were answered as if they asked for what they negate.

How the comparison works

## One head, three ways to ask

The same small head answers every question; only the vector that tells it which question changes. Every test is on people the head never trained on.

- **ID · question number**: a one-hot code for each question the head was trained on. It cannot ask anything new.
- **TPL · label template**: a sentence frame filled with the label’s name, turned into a vector by a frozen text encoder.
- **DESC · description**: a sentence that describes the state without naming it, turned into a vector the same way.
- **SHUF · shuffled templates**: the template vectors attached to the wrong questions during training, by derangements declared in advance: a control for whether the head uses the link between a wording and its question.
- **NUM · numeric code**: for SSVEP, the frequency given as numbers, not words: a control without text.
- **B-sh · one output per question**: the same hidden layer with a fixed output per seen question, as in route 2; it cannot ask anything new either.
- **CCA · no training**: for SSVEP, canonical correlation analysis, which needs no training and can name any frequency: the reference an unseen frequency is held to.

A training frame for sleep, filled with a stage’s name: The scorer marked this 30-second epoch as N2 sleep.

SSVEP on BETA: a plain spectrum (L0) and frozen CBraMod features (L1). Sleep on BOAS: frozen CBraMod features (L1). Three training seeds each. The primary text encoder is bge-small-en-v1.5, frozen. The head is the same for every arm: standardised features, a small hidden layer adjusted by the question vector (FiLM), one output.

Three kinds of result, never pooled into one zero-shot number: a seen question in the wording the head trained with; the same question reworded; and a question the head never trained on. “Unseen” always means unseen by the EEG head, not by the text encoder: the frozen encoder has read the words, and for SSVEP every token of a held-out frequency also occurs in that rotation’s training texts. Unseen sleep questions are combinations of seen stages; unseen SSVEP questions are flicker frequencies held out in five rotations.

Each comparison is the language arm minus the reference, in percentage points (pp), with a 95% interval from resampling people, then training seeds, then held-out wordings. It gets two readings, both fixed before any fit. A difference is shown when the interval excludes zero. Against a margin of ±2 pp, it is equivalent when the whole interval lies inside, non-inferior when the bound against the language arm stays within the margin, and otherwise the margin is not met, which is not the same as a loss of 2 pp or more. For an unseen sleep question the comparison is R, the head’s remaining error (1 − AUROC) over that of the read-off from its own stage answers, with a margin of R < 1.25. The check of whether the head uses the link between wording and stage, the template arm against the shuffled templates, is a difference of AUROC. Among the primary comparisons, those with the numeric code, the neighbour average and the shuffled templates are read for a difference only; the secondary tables print both readings as the run computed them.

There are 35 pre-declared primary comparisons and no multiplicity correction, so at 95% about one in twenty comparisons with no true difference may show one by chance. Accuracy is averaged per person, then over people: it compares set-ups and is not a deployment error rate.

This is route 3 of the research plan on [When not to act](https://bci.report/topics/when-not-to-act/#decision-research). “Jev-style” here means the interface of a decision model like TypeSafe’s Jev, applied to EEG: encode the recording once, then answer several explicit, typed questions about it, each with a probability. The plan started from a vision paper, [Yu & Yao, 2026 · Visual Jev: Accurate and Efficient Decisions from Shared Visual Context ↗](https://arxiv.org/abs/2609.25845), which applies that idea to images. All three routes: [Jev-style questions on EEG: the evidence →](https://bci.report/jev-style/)

Primary · seen questions · BETA and BOAS

## Seen questions: asked in words, and reworded

Questions the head was trained on. On BETA, 70 people look at flickering targets; in each rotation the head trains on most of the frequencies and is asked about those (the 32-way task). On BOAS, 100 people, each half-minute epoch is scored into one of five stages at their natural mix, against the human consensus.

**BOAS is published with three stated gaps (owner decision, 7 October 2026):**

- The consent statement does not say whether participants agreed to public sharing or secondary use.
- The ethics and consent statements come from the publisher's dataset description and README. No peer-reviewed paper describes BOAS.
- The ethics reference was added to the release in version 1.1.1 (May 2025), and the release does not say when it was granted relative to the recordings.

Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.

**Credit:** Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

### In the training wording (P1)

**−5.96 pp** TPL − ID on SSVEP, a plain spectrum, BETA, three seeds

### Sleep: no measurable cost. SSVEP: a cost.

On sleep, a template or a description instead of a question number was equivalent within ±2 pp: −0.05 pp (−0.70 to +0.61 pp) and +0.02 pp (−0.59 to +0.62 pp). On SSVEP it cost accuracy at both levels, and the margin was not met: templates −5.96 pp on a plain spectrum and −3.49 pp on frozen CBraMod features; descriptions −5.46 pp and −3.37 pp.

Paired differences in balanced accuracy, language arm minus question number, in pp, with 95% intervals; BETA 70 people, BOAS 100 people, three seeds each.

| Comparison | Question | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
| --- | --- | --- | --- | --- | --- |
| SSVEP · BETA · a plain spectrum (L0) |  |  |  |  |  |
| P1 · Label template − question number — TPL − ID | seen, 32-way | −5.96 pp (95% interval −6.84 to −5.08 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 32-way | −5.46 pp (95% interval −6.26 to −4.66 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| SSVEP · BETA · frozen CBraMod features (L1) |  |  |  |  |  |
| P1 · Label template − question number — TPL − ID | seen, 32-way | −3.49 pp (95% interval −4.10 to −2.93 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 32-way | −3.37 pp (95% interval −3.98 to −2.80 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| Sleep · BOAS · frozen CBraMod features (L1) |  |  |  |  |  |
| P1 · Label template − question number — TPL − ID | seen, 5-way | −0.05 pp (95% interval −0.70 to +0.61 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 5-way | +0.02 pp (95% interval −0.59 to +0.62 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |

The seen SSVEP task is hard for every head on these features: with question numbers the head reached 32.1% (28.8%–35.2%) on a plain spectrum and 24.2% on frozen CBraMod features, against a chance of 3.1%; CCA, which needs no training, reached 65.5%. On sleep every arm reached about the same balanced accuracy, 71.5% with question numbers against a chance of 20.0%. The levels below are descriptive: overlapping intervals are not ranked, and the paired differences above are the comparison.

BETA, seen 32-way balanced accuracy (chance 3.1%): the mean over 70 people with its 95% interval, and the chance-corrected value, (level − chance) / (1 − chance). Three seeds; CCA needs no training, so it has none.

| Arm | A plain spectrum (L0) | chance-corrected | Frozen CBraMod features (L1) | chance-corrected |
| --- | --- | --- | --- | --- |
| ID — question number | 32.1% (95% interval 28.8%–35.2%) | 29.9% | 24.2% (95% interval 21.6%–26.9%) | 21.8% |
| TPL — label template | 26.1% (95% interval 23.5%–28.6%) | 23.7% | 20.7% (95% interval 18.5%–23.1%) | 18.2% |
| DESC — description | 26.6% (95% interval 23.9%–29.2%) | 24.2% | 20.9% (95% interval 18.7%–23.2%) | 18.3% |
| SHUF — shuffled templates | 25.1% (95% interval 22.5%–27.6%) | 22.7% | 20.2% (95% interval 18.1%–22.3%) | 17.6% |
| NUM — numeric code | 23.7% (95% interval 21.2%–26.1%) | 21.2% | 13.9% (95% interval 12.5%–15.2%) | 11.1% |
| B-sh — one output per question | 32.2% (95% interval 28.9%–35.3%) | 30.0% | 27.0% (95% interval 24.2%–30.1%) | 24.7% |
| CCA — no training | 65.5% (95% interval 59.7%–71.1%) | 64.3% | 65.5% (95% interval 59.7%–71.1%) | 64.3% |

BOAS, seen five-stage balanced accuracy (chance 20.0%), 100 people, three seeds, with the chance-corrected value and the log loss (mean binary cross-entropy, in nats). The shuffled-template arm’s seen score is read through its own derangement, and it has no log loss.

| Arm | Balanced accuracy | chance-corrected | Log loss |
| --- | --- | --- | --- |
| ID — question number | 71.5% (95% interval 69.1%–73.5%) | 64.3% | 0.267 (95% interval 0.242–0.296) |
| TPL — label template | 71.4% (95% interval 69.1%–73.4%) | 64.3% | 0.276 (95% interval 0.249–0.307) |
| DESC — description | 71.5% (95% interval 69.1%–73.5%) | 64.3% | 0.278 (95% interval 0.251–0.310) |
| SHUF — shuffled templates | 71.5% (95% interval 69.2%–73.5%) | 64.4% | — |
| B-sh — one output per question | 71.4% (95% interval 69.1%–73.4%) | 64.3% | 0.259 (95% interval 0.236–0.287) |

BOAS, AUROC of each seen stage, with 95% intervals; three seeds. N3 pools the 71 people with at least five N3 epochs; every other stage pools 100.

| Arm | W — 100 people | N1 — 100 people | N2 — 100 people | N3 — 71 people | REM — 100 people |
| --- | --- | --- | --- | --- | --- |
| ID — question number | 0.982 (0.976–0.987) | 0.876 (0.859–0.890) | 0.922 (0.904–0.937) | 0.978 (0.971–0.983) | 0.957 (0.945–0.967) |
| TPL — label template | 0.980 (0.972–0.987) | 0.872 (0.855–0.887) | 0.924 (0.909–0.937) | 0.979 (0.974–0.983) | 0.956 (0.944–0.967) |
| DESC — description | 0.981 (0.974–0.987) | 0.869 (0.851–0.884) | 0.923 (0.907–0.937) | 0.979 (0.974–0.983) | 0.956 (0.944–0.967) |

### Reworded (P2)

The same seen questions in wordings the head never trained with: a new sentence frame (a1, held-out frames resampled), a synonym of the label (a2: three per question, fixed, so these entries hold only for these three synonyms), or a new description (a3, held-out descriptions resampled). Reworded rows are never pooled with the unseen ones.

Paired differences in balanced accuracy, reworded minus the training wording, in pp, with 95% intervals; BETA 70 people, BOAS 100 people, three seeds each.

| Comparison | Question | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
| --- | --- | --- | --- | --- | --- |
| SSVEP · BETA · a plain spectrum (L0) |  |  |  |  |  |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −4.01 pp (95% interval −6.53 to −2.12 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −1.84 pp (95% interval −2.10 to −1.57 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −4.13 pp (95% interval −5.58 to −3.11 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| SSVEP · BETA · frozen CBraMod features (L1) |  |  |  |  |  |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −2.72 pp (95% interval −4.43 to −1.43 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −1.33 pp (95% interval −1.53 to −1.14 pp) | TPL(train) higher | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −2.95 pp (95% interval −3.81 to −2.26 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| Sleep · BOAS · frozen CBraMod features (L1) |  |  |  |  |  |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −0.46 pp (95% interval −1.33 to +0.16 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −12.60 pp (95% interval −13.74 to −11.43 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −20.87 pp (95% interval −29.49 to −13.71 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |

On sleep a new sentence frame cost nothing measurable: equivalent within ±2 pp (−0.46 pp, −1.33 to +0.16 pp). The three synonyms per stage scored −12.60 pp (−13.74 to −11.43 pp) and new descriptions −20.87 pp (−29.49 to −13.71 pp), a wide interval, because descriptions differ a lot from one another. On BETA every rewording scored lower, from −1.33 pp to −4.13 pp; only the synonyms on frozen CBraMod features stayed within the margin.

Each synonym on its own against the training wording, in pp with 95% intervals (descriptive); BETA 70 people, BOAS 100 people. On SSVEP the third synonym spells the number out.

| Synonym | SSVEP · BETA · a plain spectrum (L0) | SSVEP · BETA · frozen CBraMod features (L1) | Sleep · BOAS · frozen CBraMod features (L1) |
| --- | --- | --- | --- |
| Synonym 1 | −1.45 pp (95% interval −1.67 to −1.22 pp) | −1.09 pp (95% interval −1.27 to −0.90 pp) | −23.53 pp (95% interval −25.29 to −21.64 pp) |
| Synonym 2 | −0.87 pp (95% interval −1.04 to −0.70 pp) | −0.84 pp (95% interval −1.00 to −0.68 pp) | −10.85 pp (95% interval −12.59 to −9.17 pp) |
| Synonym 3 | −3.19 pp (95% interval −3.69 to −2.69 pp) | −2.06 pp (95% interval −2.44 to −1.73 pp) | −3.42 pp (95% interval −4.49 to −2.41 pp) |

Answer flip rate (descriptive): how often the head’s answer to the same window changes when only the wording changes; BETA 70 people, BOAS 100 people. On SSVEP even another training seed of the question-number head changes most answers, because the seen task is near its floor, so flip rates there mostly show uncertainty.

|  | New frame | Synonym | Between training frames | New description | Between training descriptions | Between seeds of ID |
| --- | --- | --- | --- | --- | --- | --- |
| BETA · L0 | 39.7% | 29.3% | 19.4% | 41.2% | 19.3% | 56.8% |
| BETA · L1 | 40.7% | 32.0% | 24.3% | 42.7% | 23.3% | 62.2% |
| BOAS · L1 | 6.4% | 20.0% | 3.3% | 45.4% | 3.1% | 7.4% |

Primary · unseen sleep questions · BOAS

## Sleep questions the EEG head was never trained on

Three questions the head never trained on, each a combination of the stages it did: asleep (N1, N2, N3 or REM), N1, N2 or N3 sleep, and light sleep (N1 or N2). Each was asked with templates, using an implicit name or an explicit one, and with new descriptions; 100 people, three seeds. Unseen by the EEG head, not by the text encoder: the encoder has read these words. Because each question is a combination, the head’s own stage answers can be added up into an answer, the read-off, which is what each wording is held to.

**BOAS is published with three stated gaps (owner decision, 7 October 2026):**

- The consent statement does not say whether participants agreed to public sharing or secondary use.
- The ethics and consent statements come from the publisher's dataset description and README. No peer-reviewed paper describes BOAS.
- The ethics reference was added to the release in version 1.1.1 (May 2025), and the release does not say when it was granted relative to the recordings.

Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.

**Credit:** Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

The names the template arm was asked with, as written before any score.

| Question | Implicit name | Explicit name |
| --- | --- | --- |
| Asleep | asleep | N1, N2, N3 or REM sleep |
| N1, N2 or N3 sleep | none | N1, N2 or N3 sleep |
| Light sleep | light sleep | N1 or N2 sleep |

### Can the head answer them? (P3)

**33.80** R for “asleep” with its implicit name: the head’s remaining error over the read-off’s

### Not answered: the read-off did better in every wording

In all eight conditions the interval lay above R = 1, in the read-off’s favour, and none came within the 1.25 margin: R ran from 2.42 (2.07–2.85; N1, N2 or N3 sleep, explicit name) to 33.80 (23.09–52.82; asleep, implicit name).

log R and R, the head’s remaining error (1 − AUROC) over the read-off’s, with 95% intervals; 100 people, three seeds. A positive log R, or R above 1, is against the language arm.

| Comparison | Question | log R, with R | Difference shown? | Margin, R < 1.25 | Gate |
| --- | --- | --- | --- | --- | --- |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | asleep · implicit name | +3.520 (95% interval +3.139 to +3.967) (R 33.80 (23.09–52.82)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | asleep · explicit name | +1.951 (95% interval +1.623 to +2.357) (R 7.04 (5.07–10.56)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | N1, N2 or N3 sleep · explicit name | +0.883 (95% interval +0.726 to +1.047) (R 2.42 (2.07–2.85)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | light sleep · implicit name | +2.267 (95% interval +2.072 to +2.466) (R 9.65 (7.94–11.77)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | light sleep · explicit name | +1.519 (95% interval +1.343 to +1.698) (R 4.57 (3.83–5.46)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | asleep · new descriptions — 10 held-out wordings, resampled | +3.444 (95% interval +2.913 to +3.956) (R 31.30 (18.41–52.26)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | N1, N2 or N3 sleep · new descriptions — 10 held-out wordings, resampled | +1.299 (95% interval +0.793 to +1.744) (R 3.67 (2.21–5.72)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | light sleep · new descriptions — 10 held-out wordings, resampled | +1.417 (95% interval +0.903 to +1.856) (R 4.12 (2.47–6.40)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |

AUROC of the template head on each unseen question, and of the read-off from its own stage answers (prior-corrected), with 95% intervals; 100 people, three seeds. Descriptive. An AUROC below 0.5 ranks windows in reverse.

| Question | Wording | Template head | Read-off from its stage answers |
| --- | --- | --- | --- |
| Asleep | implicit name | 0.317 (95% interval 0.294–0.341) | 0.980 (95% interval 0.971–0.987) |
| Asleep | explicit name | 0.858 (95% interval 0.839–0.875) | 0.980 (95% interval 0.971–0.987) |
| N1, N2 or N3 sleep | explicit name | 0.871 (95% interval 0.854–0.887) | 0.947 (95% interval 0.934–0.958) |
| Light sleep | implicit name | 0.244 (95% interval 0.222–0.267) | 0.922 (95% interval 0.907–0.935) |
| Light sleep | explicit name | 0.642 (95% interval 0.616–0.667) | 0.922 (95% interval 0.907–0.935) |

The implicit names ranked windows in reverse: AUROC 0.317 for “asleep” and 0.244 for “light sleep”, below 0.5, so windows that really were asleep, or in light sleep, scored lower on these questions than the others. The explicit names did better (0.858 and 0.642) but stayed short of the read-off (0.980 and 0.922). Descriptions have no alignment control: the shuffled control is built from template vectors. Held-out sentence frames (S13, in the secondary panel) give the same result for all five template cells.

### Is the link between wording and stage used? (P5)

The shuffled-template arm attaches the same template vectors to the wrong stages during training; the unseen questions keep their true vectors. If the head uses the link between a wording and its stage, the shuffled arm should do worse on them than the template arm.

Template arm minus shuffled templates: AUROC on the unseen questions, averaged over the cells of each kind of name, with 95% intervals; 100 people, three seeds. Difference only: no margin was set.

| Comparison | Question | Difference, AUROC | Difference shown? | Margin | Gate |
| --- | --- | --- | --- | --- | --- |
| P5 · Label template − shuffled templates — TPL − SHUF | unseen questions, implicit names | −0.378 (95% interval −0.501 to −0.253) | SHUF higher | Difference only: no margin | Headroom gate passed |
| P5 · Label template − shuffled templates — TPL − SHUF | unseen questions, explicit names | +0.629 (95% interval +0.603 to +0.654) | TPL higher | Difference only: no margin | Headroom gate passed |

AUROC on each unseen question, template arm and shuffled templates, with 95% intervals (descriptive); 100 people, three seeds.

| Question | Wording | Template arm | Shuffled templates |
| --- | --- | --- | --- |
| Asleep | implicit name | 0.317 (95% interval 0.294–0.341) | 0.676 (95% interval 0.529–0.812) |
| Asleep | explicit name | 0.858 (95% interval 0.839–0.875) | 0.137 (95% interval 0.109–0.168) |
| N1, N2 or N3 sleep | explicit name | 0.871 (95% interval 0.854–0.887) | 0.171 (95% interval 0.145–0.201) |
| Light sleep | implicit name | 0.244 (95% interval 0.222–0.267) | 0.641 (95% interval 0.537–0.738) |
| Light sleep | explicit name | 0.642 (95% interval 0.616–0.667) | 0.176 (95% interval 0.155–0.199) |

For explicit names, yes: attaching the templates to the wrong stages lowered the unseen AUROC, TPL − SHUF +0.629 (+0.603 to +0.654). For the implicit names it was not shown: the shuffled templates scored higher, −0.378 (−0.501 to −0.253).

Primary · unseen flicker frequencies · BETA

## Flicker frequencies the EEG head was never trained on (P4)

In each of five rotations, eight of BETA’s frequencies are held out: the head trains on the rest and is asked about the eight by stating the frequency in a template. 70 people, three seeds, at both levels. Unseen by the EEG head, not by the text encoder: every token of a held-out frequency also occurs in that rotation’s training texts, so only the combination is new. CCA answers the same question with no training at all.

**−52.21 pp** TPL − CCA, 8-way among unseen frequencies, a plain spectrum, BETA

### Far below CCA, which needs no training

In the 8-way among unseen frequencies the head reached 28.7% (26.9%–30.5%) on a plain spectrum and 19.0% on frozen CBraMod features; CCA reached 80.9%, and chance is 12.5%. The head also did worse than averaging its own answers for the two neighbouring seen frequencies, and worse than a numeric code of the frequency, at both levels.

Paired differences in accuracy, template head minus the reference, in pp, with 95% intervals; 70 people, three seeds. Comparisons with CCA carry the margin; with the neighbour average and the numeric code they are read for a difference only.

| Comparison | Question | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
| --- | --- | --- | --- | --- | --- |
| SSVEP · BETA · a plain spectrum (L0) |  |  |  |  |  |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, 8-way | −52.21 pp (95% interval −55.13 to −48.99 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequencies — TPL − NN(TPL) | unseen frequency, 8-way | −10.98 pp (95% interval −12.24 to −9.67 pp) | NN(TPL) higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, 8-way | −12.77 pp (95% interval −14.52 to −10.99 pp) | NUM higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, neighbour 3-way | −62.36 pp (95% interval −65.87 to −58.69 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, neighbour 3-way | −1.24 pp (95% interval −3.17 to +0.85 pp) | No difference shown | Difference only: no margin | Headroom gate passed |
| SSVEP · BETA · frozen CBraMod features (L1) |  |  |  |  |  |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, 8-way | −61.94 pp (95% interval −65.65 to −58.06 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequencies — TPL − NN(TPL) | unseen frequency, 8-way | −1.16 pp (95% interval −1.97 to −0.37 pp) | NN(TPL) higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, 8-way | −2.46 pp (95% interval −3.23 to −1.69 pp) | NUM higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, neighbour 3-way | −54.67 pp (95% interval −57.77 to −51.47 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, neighbour 3-way | +15.61 pp (95% interval +14.25 to +16.93 pp) | TPL higher | Difference only: no margin | Headroom gate passed |

In the neighbour 3-way the unseen frequency competes with the trained detectors of its two neighbours (seen bias, stated). On a plain spectrum the template head fell below the mean chance of that test, 19.3% against 34.2%, and showed no difference against the numeric code. On frozen CBraMod features it beat the numeric code by +15.61 pp, but those features keep the stimulus phase and neighbouring targets differ in phase, so that comparison tests frequency and phase together. The plain spectrum, which carries no phase, is the clean frequency test.

Accuracy on the unseen frequencies (descriptive): u8, the 8-way among the held-out frequencies (chance 12.5%); nn8, the read-off that averages the answers for the two neighbouring seen frequencies; u3, the neighbour 3-way (mean chance 34.2%, a 2-way at the two grid edges). 70 people; three seeds, CCA none.

| Arm | Test | A plain spectrum (L0) | chance-corrected | Frozen CBraMod features (L1) | chance-corrected |
| --- | --- | --- | --- | --- | --- |
| TPL — label template | u8 | 28.7% (95% interval 26.9%–30.5%) | 18.5% | 19.0% (95% interval 18.0%–19.9%) | 7.4% |
| DESC — description | u8 | 28.6% (95% interval 26.8%–30.4%) | 18.4% | 18.6% (95% interval 17.6%–19.6%) | 7.0% |
| SHUF — shuffled templates | u8 | 9.2% (95% interval 8.6%–9.7%) | -3.8% | 11.0% (95% interval 10.5%–11.5%) | -1.7% |
| NUM — numeric code | u8 | 41.5% (95% interval 38.2%–44.7%) | 33.1% | 21.4% (95% interval 20.2%–22.6%) | 10.2% |
| CCA — no training | u8 | 80.9% (95% interval 76.6%–84.8%) | 78.2% | 80.9% (95% interval 76.6%–84.8%) | 78.2% |
| TPL — label template | nn8 | 39.7% (95% interval 36.8%–42.5%) | 31.0% | 20.1% (95% interval 19.0%–21.3%) | 8.7% |
| ID — question number | nn8 | 44.3% (95% interval 40.8%–47.6%) | 36.4% | 18.3% (95% interval 17.3%–19.3%) | 6.6% |
| B-sh — one output per question | nn8 | 44.3% (95% interval 40.9%–47.5%) | 36.3% | 17.6% (95% interval 16.6%–18.8%) | 5.8% |
| TPL — label template | u3 | 19.3% (95% interval 18.2%–20.5%) | -22.5% | 27.0% (95% interval 26.0%–28.0%) | -10.8% |
| NUM — numeric code | u3 | 20.6% (95% interval 19.3%–21.8%) | -20.6% | 11.4% (95% interval 10.7%–12.1%) | -34.5% |
| CCA — no training | u3 | 81.7% (95% interval 78.5%–84.6%) | 72.2% | 81.7% (95% interval 78.5%–84.6%) | 72.2% |

Generalisation cost (descriptive): accuracy on the unseen frequencies minus accuracy on matched seen ones, in pp with 95% intervals; 70 people.

| Arm and test | A plain spectrum (L0) | Frozen CBraMod features (L1) |
| --- | --- | --- |
| ID, 8-way (neighbour read-off) | −14.26 pp (95% interval −15.50 to −13.06 pp) | −27.80 pp (95% interval −30.61 to −25.03 pp) |
| TPL, 8-way | −24.44 pp (95% interval −26.44 to −22.26 pp) | −24.27 pp (95% interval −26.72 to −21.85 pp) |
| TPL, 3-way | −27.83 pp (95% interval −30.18 to −25.37 pp) | −29.67 pp (95% interval −31.76 to −27.64 pp) |
| DESC, 8-way | −24.79 pp (95% interval −26.88 to −22.55 pp) | −24.81 pp (95% interval −27.31 to −22.39 pp) |
| DESC, 3-way | −28.95 pp (95% interval −31.33 to −26.46 pp) | −29.74 pp (95% interval −31.76 to −27.76 pp) |
| SHUF, 8-way | −40.63 pp (95% interval −44.16 to −36.98 pp) | −30.09 pp (95% interval −33.02 to −27.27 pp) |
| SHUF, 3-way | −34.96 pp (95% interval −37.50 to −32.43 pp) | −32.52 pp (95% interval −35.18 to −29.84 pp) |
| NUM, 8-way | −12.30 pp (95% interval −13.42 to −11.17 pp) | −16.16 pp (95% interval −17.88 to −14.47 pp) |
| NUM, 3-way | −7.27 pp (95% interval −8.49 to −6.02 pp) | −7.63 pp (95% interval −8.74 to −6.53 pp) |
| B-sh, 8-way (neighbour read-off) | −12.81 pp (95% interval −14.09 to −11.56 pp) | −31.19 pp (95% interval −34.09 to −28.32 pp) |
| CCA, 8-way | 0.00 pp (95% interval 0.00 to 0.00 pp) | 0.00 pp (95% interval 0.00 to 0.00 pp) |
| CCA, 3-way | +0.41 pp (95% interval +0.26 to +0.56 pp) | +0.41 pp (95% interval +0.26 to +0.56 pp) |

TPL − CCA in the 8-way, rotation by rotation, in pp with 95% intervals (secondary, recomputed by an independent audit); 70 people, three seeds. No rotation carries the result alone.

| Rotation | A plain spectrum (L0) | Frozen CBraMod features (L1) |
| --- | --- | --- |
| Rotation 0 | −49.88 pp (95% interval −53.16 to −46.16 pp) | −63.00 pp (95% interval −67.23 to −58.54 pp) |
| Rotation 1 | −52.96 pp (95% interval −57.02 to −48.99 pp) | −61.76 pp (95% interval −65.81 to −57.74 pp) |
| Rotation 2 | −48.74 pp (95% interval −51.97 to −45.42 pp) | −59.74 pp (95% interval −63.75 to −55.51 pp) |
| Rotation 3 | −57.34 pp (95% interval −60.84 to −53.73 pp) | −62.75 pp (95% interval −66.73 to −58.42 pp) |
| Rotation 4 | −52.14 pp (95% interval −55.80 to −48.20 pp) | −62.47 pp (95% interval −66.53 to −58.23 pp) |

Per-frequency and interior-only breakdowns were computed after the audits, and no independent audit covers them, so they are not published.

Before any result

## The pre-run checks, and how one was changed

A check was changed after it failed, twice. Everything about that is here, with the other checks that had to pass before any result was read.

A training-label permutation canary refits the heads on shuffled training labels and checks that test scores sit at chance: 17 configurations on the first fold, 38 intervals at 99.9%.

Revision 3 failed: 1 of 38 intervals excluded chance (sleep, frozen CBraMod features, the shuffled-template arm, unseen “light sleep” with its explicit name). Its labels had been permuted within each training person. The run stopped.

Revision 4, approved by the owner before any new fit, permuted labels across all training windows. It failed too: 3 of 38 intervals, all on BOAS, one of them below chance; every SSVEP interval included chance both times. Revision 4 said “no further revision”. The owner overrode that after seeing how it failed, and approved revision 5 before any new fit. Revision 4’s per-person centring diagnostic was not run: with labels permuted across all windows it no longer tests anything.

Revision 5 passed: 20 heads per configuration (340 fits), each metric compared with its exact per-person chance, and a two-level bootstrap over heads, then people; 0 of 38 intervals excluded zero. Both failed records are kept.

Direction counts, descriptive: of 760 single-head intervals in revision 5, 125 excluded the exact chance, 66 above and 59 below, and 18 cells failed in both directions. BOAS cells are published as pass or fail and counts only.

Why they failed, the reading revision 5 was built to test: a head trained on shuffled labels is still a random function of sleep features that are strongly organised by stage, so on new people its scores lean one way for everyone, up or down. A person bootstrap measures the spread between people, not between heads, so its 99.9% interval was too narrow for this question. The failures fit that reading: all were on BOAS, and one was below chance, which a leak cannot produce. It is not a proof that no leak exists: person-disjoint folds, no held-out label or wording in training and training-only scaling are asserted in the code, and both conformance audits check for leakage directly.

Decision S0b-6, approved by the owner: the SSVEP token rule says every token of a held-out frequency occurs in that rotation’s training texts. It stops the run for bge-small-en-v1.5 and all-MiniLM-L6-v2, and both passed in all 5 rotations. It cannot hold for multilingual-e5-small, whose tokenizer keeps one-decimal number pieces as single tokens while the held-out number itself is barred from training texts; for that encoder the rule was reported instead of stopping the run, and every Chinese-wording unseen-frequency entry says the frequency is a new token.

Engineering amendment EA-1, before the first fit: in the contiguous extrapolation band (S8) the neighbour average is not defined and the neighbour 3-way is relabelled, because its competitors are themselves held-out questions; failed fits are marked per cell, and none failed; the secondary scoring can resume after an interruption. No question, partition, wording, arm, metric, margin or gate changed.

The declared numeracy prediction was not activated. It said: if the primary text encoder orders frequencies poorly (Spearman ρ below 0.3 between frequency distance and embedding distance, for every training frame), the 8-way TPL − NUM flag will not read “TPL higher”. Its premise was false: ρ was 0.593, 0.563 and 0.526 for the three training frames of bge-small-en-v1.5, over 780 pairs. The flag reads “NUM higher” at both levels anyway; with the premise false, that is a description, not a confirmed prediction.

### The gates

- Headroom: the question-number head’s seen balanced accuracy must lie above chance + 5 pp and below 95%. It passed on all three primary blocks; the levels are in the tables above.
- Triviality canary: a ridge classifier on per-channel mean and log-variance must not answer the seen task by itself. It passed on BETA (on the first fold’s 60 training people) and on BOAS and EESM19 (carried from route 2); none is declared for motor imagery, so its gate reads “not run”.
- Read-off floor: an unseen sleep cell counts only if the read-off’s AUROC interval lies above 0.55. 8 of 8 primary and 34 of 34 secondary cells passed.
- Shuffled-EEG scorer check: each primary aggregate, recomputed with the test features shuffled within each person, must sit within three Monte-Carlo standard errors of chance. 51 of 51 cells did (26 SSVEP, 25 BOAS).
- Fit check, on the first fold’s training rows: every conditioned SSVEP arm trained more than 5 pp below the one-output-per-question head, so the declared fallback of 60 epochs applies to every SSVEP arm; sleep, EESM19 and motor imagery stay at 30. No fit failed.

Four independent audits passed. They re-implement the definitions without importing the study’s scoring code: the conformance audit of the primary run passed 106 checks with 0 failed; the numerical audit of the primary made 2,270 comparisons with 0 mismatches; the conformance audit of the secondary run passed 231 checks with 0 failed; the numerical audit of the secondary made 18,619 comparisons, 29 of them outside the tolerance, every one explained.

Secondary · beyond the primary comparisons

## Secondary results

Further datasets, features, a second text encoder, Chinese wordings and probes, none of them among the primary comparisons. One seed unless marked: a single seed carries no training-run variance. Every EESM19 and single-seed entry was declared likely inconclusive before the run, though none reached the “inconclusive” flags.

Show the secondary results

### S1 · EESM19 sleep, frozen CBraMod features · 20 people, three seeds · declared likely inconclusive

| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
| --- | --- | --- | --- | --- | --- |
| P1 · Label template − question number — TPL − ID | seen, 5-way | −0.04 pp (95% interval −0.47 to +0.39 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 5-way | −0.03 pp (95% interval −0.43 to +0.40 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −0.11 pp (95% interval −0.56 to +0.20 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −15.27 pp (95% interval −16.24 to −14.29 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −18.54 pp (95% interval −27.30 to −10.80 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | asleep · implicit name | +4.720 (95% interval +4.410 to +5.113) (R 112.17 (82.27–166.15)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | asleep · explicit name | +2.566 (95% interval +2.179 to +3.010) (R 13.02 (8.84–20.30)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | N1, N2 or N3 sleep · explicit name | +1.080 (95% interval +0.890 to +1.267) (R 2.95 (2.44–3.55)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | light sleep · implicit name | +2.409 (95% interval +2.208 to +2.609) (R 11.12 (9.10–13.58)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | light sleep · explicit name | +0.888 (95% interval +0.769 to +1.024) (R 2.43 (2.16–2.78)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | asleep · new descriptions — 10 held-out wordings, resampled | +4.707 (95% interval +4.279 to +5.154) (R 110.72 (72.15–173.16)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | N1, N2 or N3 sleep · new descriptions — 10 held-out wordings, resampled | +2.223 (95% interval +1.622 to +2.649) (R 9.24 (5.06–14.15)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | light sleep · new descriptions — 10 held-out wordings, resampled | +1.707 (95% interval +1.240 to +2.056) (R 5.51 (3.46–7.81)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P5 · Label template − shuffled templates — TPL − SHUF | unseen questions, implicit names | −0.391 (95% interval −0.470 to −0.311) | SHUF higher | Difference only: no margin | Headroom gate passed |
| P5 · Label template − shuffled templates — TPL − SHUF | unseen questions, explicit names | +0.667 (95% interval +0.634 to +0.698) | TPL higher | Difference only: no margin | Headroom gate passed |

### S2 · OpenBMI motor imagery, frozen CBraMod features · 51 people, three seeds

No triviality canary is declared for motor imagery, so its gate reads “not run” and excludes nothing. Motor imagery has only two three-cycles, so two of the three shuffled-template seeds share one derangement. The unseen “imagining a hand movement” came closest to its read-off: a difference, but within the ratio margin.

| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
| --- | --- | --- | --- | --- | --- |
| P1 · Label template − question number — TPL − ID | seen, 3-way | −0.49 pp (95% interval −1.07 to +0.02 pp) | No difference shown | Equivalent within ±2 pp | No canary declared for this task: not run |
| P1 · Description − question number — DESC − ID | seen, 3-way | −0.54 pp (95% interval −1.08 to +0.05 pp) | No difference shown | Equivalent within ±2 pp | No canary declared for this task: not run |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | +0.05 pp (95% interval −0.10 to +0.20 pp) | No difference shown | Equivalent within ±2 pp | No canary declared for this task: not run |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | +0.22 pp (95% interval −0.11 to +0.55 pp) | No difference shown | Equivalent within ±2 pp | No canary declared for this task: not run |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −3.49 pp (95% interval −6.63 to −1.42 pp) | DESC(train) higher | 2 pp margin not met | No canary declared for this task: not run |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | imagining a hand movement · implicit name | +0.076 (95% interval +0.050 to +0.107) (R 1.08 (1.05–1.11)) | Read-off leaves less error | Equivalent within the 1.25 ratio margin | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | imagining a hand movement · new descriptions — 10 held-out wordings, resampled | +0.088 (95% interval +0.048 to +0.138) (R 1.09 (1.05–1.15)) | Read-off leaves less error | Equivalent within the 1.25 ratio margin | Headroom gate passed |

### S3 · REVE-L features, BETA · 70 people, one seed · declared likely inconclusive

REVE-L features keep the stimulus phase, so their neighbour 3-way tests frequency and phase together.

| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
| --- | --- | --- | --- | --- | --- |
| P1 · Label template − question number — TPL − ID | seen, 32-way | −2.76 pp (95% interval −3.51 to −2.07 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 32-way | −2.56 pp (95% interval −3.24 to −1.90 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −4.48 pp (95% interval −7.66 to −2.16 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −2.16 pp (95% interval −2.44 to −1.88 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −5.05 pp (95% interval −6.50 to −3.77 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, 8-way | −65.32 pp (95% interval −69.18 to −61.17 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequencies — TPL − NN(TPL) | unseen frequency, 8-way | +5.81 pp (95% interval +4.54 to +7.04 pp) | TPL higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, 8-way | +1.92 pp (95% interval +0.85 to +2.92 pp) | TPL higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, neighbour 3-way | −38.75 pp (95% interval −41.62 to −35.60 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, neighbour 3-way | +35.22 pp (95% interval +32.91 to +37.42 pp) | TPL higher | Difference only: no margin | Headroom gate passed |

### S4 · text encoder all-MiniLM-L6-v2, BETA, a plain spectrum · 70 people, one seed · declared likely inconclusive

| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
| --- | --- | --- | --- | --- | --- |
| P1 · Label template − question number — TPL − ID | seen, 32-way | −6.73 pp (95% interval −7.75 to −5.68 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 32-way | −11.58 pp (95% interval −13.16 to −9.97 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −3.95 pp (95% interval −5.11 to −2.86 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −4.90 pp (95% interval −5.53 to −4.25 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −5.49 pp (95% interval −6.48 to −4.60 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, 8-way | −50.10 pp (95% interval −53.06 to −46.73 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequencies — TPL − NN(TPL) | unseen frequency, 8-way | −7.68 pp (95% interval −8.82 to −6.50 pp) | NN(TPL) higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, 8-way | −10.84 pp (95% interval −12.53 to −9.08 pp) | NUM higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, neighbour 3-way | −61.85 pp (95% interval −65.13 to −58.31 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, neighbour 3-way | −0.18 pp (95% interval −2.02 to +1.73 pp) | No difference shown | Difference only: no margin | Headroom gate passed |

### S4 · text encoder all-MiniLM-L6-v2, BETA, frozen CBraMod features · 70 people, one seed · declared likely inconclusive

| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
| --- | --- | --- | --- | --- | --- |
| P1 · Label template − question number — TPL − ID | seen, 32-way | −4.73 pp (95% interval −5.43 to −4.07 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 32-way | −8.84 pp (95% interval −10.12 to −7.60 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −2.89 pp (95% interval −3.73 to −2.08 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −3.09 pp (95% interval −3.57 to −2.63 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −2.98 pp (95% interval −3.67 to −2.34 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, 8-way | −61.55 pp (95% interval −65.21 to −57.49 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequencies — TPL − NN(TPL) | unseen frequency, 8-way | −1.09 pp (95% interval −1.69 to −0.48 pp) | NN(TPL) higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, 8-way | −2.03 pp (95% interval −2.86 to −1.19 pp) | NUM higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, neighbour 3-way | −54.63 pp (95% interval −57.51 to −51.29 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, neighbour 3-way | +15.57 pp (95% interval +14.27 to +16.75 pp) | TPL higher | Difference only: no margin | Headroom gate passed |

### S5 · Chinese wordings with multilingual-e5-small, BETA, a plain spectrum · 70 people, one seed · declared likely inconclusive

For this encoder each held-out frequency is a new token, not only a new combination of seen tokens (decision S0b-6).

| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
| --- | --- | --- | --- | --- | --- |
| P1 · Label template − question number — TPL − ID | seen, 32-way | −3.15 pp (95% interval −3.80 to −2.48 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 32-way | −4.94 pp (95% interval −5.69 to −4.18 pp) | ID higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −4.67 pp (95% interval −6.57 to −3.24 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −8.70 pp (95% interval −9.76 to −7.60 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −7.35 pp (95% interval −8.58 to −6.17 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, 8-way | −54.76 pp (95% interval −57.77 to −51.25 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequencies — TPL − NN(TPL) | unseen frequency, 8-way | −15.81 pp (95% interval −17.68 to −13.94 pp) | NN(TPL) higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, 8-way | −15.51 pp (95% interval −17.58 to −13.44 pp) | NUM higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no training — TPL − CCA | unseen frequency, neighbour 3-way | −59.98 pp (95% interval −63.74 to −55.86 pp) | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric code — TPL − NUM | unseen frequency, neighbour 3-way | +1.69 pp (95% interval −0.60 to +4.18 pp) | No difference shown | Difference only: no margin | Headroom gate passed |

### S8 · extrapolation to a contiguous held-out band at the top of the grid, BETA L0 · 70 people, one seed · declared likely inconclusive

| Comparison | Difference, pp | Difference shown? | Margin, ±2 pp |
| --- | --- | --- | --- |
| TPL - CCA (8-way) | −65.37 pp (95% interval −69.21 to −61.40 pp) | CCA higher | 2 pp margin not met |
| TPL - NUM (8-way) | +3.51 pp (95% interval +1.64 to +5.49 pp) | TPL higher | TPL non-inferior at 2 pp |
| TPL - NN(TPL) (8-way) | Not defined (EA-1) |  |  |
| TPL - CCA (3-way) | −50.19 pp (95% interval −54.05 to −46.25 pp) | CCA higher | 2 pp margin not met |

### S8 · extrapolation to a contiguous held-out band at the top of the grid, BETA L1 · 70 people, one seed · declared likely inconclusive

| Comparison | Difference, pp | Difference shown? | Margin, ±2 pp |
| --- | --- | --- | --- |
| TPL - CCA (8-way) | −67.47 pp (95% interval −71.22 to −63.56 pp) | CCA higher | 2 pp margin not met |
| TPL - NUM (8-way) | +0.88 pp (95% interval −0.70 to +2.47 pp) | No difference shown | TPL non-inferior at 2 pp |
| TPL - NN(TPL) (8-way) | Not defined (EA-1) |  |  |
| TPL - CCA (3-way) | −48.41 pp (95% interval −52.10 to −44.40 pp) | CCA higher | 2 pp margin not met |

In the band the neighbour 3-way competitors are themselves held-out questions, not trained detectors, and the neighbour average is not defined (EA-1). For bge-small-en-v1.5, many of these texts carry a number token never seen in training.

### S12 · Wearable-102, heads trained on all of BETA’s frequencies and asked about off-grid ones · 102 people, three seeds

The off-grid frequencies have two decimals, whose hundredths never occur in BETA’s training texts, so the text arms face new tokens; the numeric code is the clean comparison. Nine fits, three arms by three seeds, fixed before the freeze. Computed from the Tsinghua author mirror; no physical unit is claimed.

| Comparison | Difference, pp | Difference shown? | Margin, ±2 pp |
| --- | --- | --- | --- |
| Dry electrodes |  |  |  |
| TPL - CCA (12-way) | −48.18 pp (95% interval −52.36 to −43.90 pp) | CCA higher | 2 pp margin not met |
| NUM - CCA (12-way) | −42.16 pp (95% interval −46.05 to −38.38 pp) | CCA higher | 2 pp margin not met |
| TPL - NUM (12-way) | −6.02 pp (95% interval −7.32 to −4.73 pp) | NUM higher | 2 pp margin not met |
| DESC - CCA (12-way) | −48.52 pp (95% interval −52.99 to −44.02 pp) | CCA higher | 2 pp margin not met |
| Wet electrodes |  |  |  |
| TPL - CCA (12-way) | −62.92 pp (95% interval −66.16 to −59.66 pp) | CCA higher | 2 pp margin not met |
| NUM - CCA (12-way) | −56.11 pp (95% interval −59.16 to −52.83 pp) | CCA higher | 2 pp margin not met |
| TPL - NUM (12-way) | −6.82 pp (95% interval −8.31 to −5.33 pp) | NUM higher | 2 pp margin not met |
| DESC - CCA (12-way) | −63.56 pp (95% interval −66.68 to −60.31 pp) | CCA higher | 2 pp margin not met |

### S10, S11 and S13 · BETA L0: other arms against CCA and FBCCA, and held-out wordings · 70 people, three seeds

| Comparison | Difference, pp | Difference shown? | Margin, ±2 pp |
| --- | --- | --- | --- |
| S10 DESC - CCA (8-way) | −52.29 pp (95% interval −55.41 to −49.00 pp) | CCA higher | 2 pp margin not met |
| S10 DESC - CCA (3-way) | −62.72 pp (95% interval −66.34 to −58.98 pp) | CCA higher | 2 pp margin not met |
| S10 DESC - NN(DESC) (8-way) | −10.80 pp (95% interval −12.15 to −9.39 pp) | NN(DESC) higher | 2 pp margin not met |
| S10 NUM - CCA (8-way) | −39.44 pp (95% interval −42.25 to −36.48 pp) | CCA higher | 2 pp margin not met |
| S10 TPL - NN(ID) (8-way) | −15.63 pp (95% interval −17.48 to −13.62 pp) | NN(ID) higher | 2 pp margin not met |
| S10 TPL - SHUF (8-way) | +19.50 pp (95% interval +17.45 to +21.62 pp) | TPL higher | TPL non-inferior at 2 pp |
| S11 TPL - FBCCA (8-way) | −61.68 pp (95% interval −63.72 to −59.69 pp) | FBCCA higher | 2 pp margin not met |
| S11 TPL - FBCCA (3-way) | −73.35 pp (95% interval −75.61 to −70.93 pp) | FBCCA higher | 2 pp margin not met |
| S13 TPL held-out frames - CCA (8-way) | −53.71 pp (95% interval −56.73 to −50.27 pp) | CCA higher | 2 pp margin not met |
| S13 DESC held-out descriptions - CCA (8-way) | −54.01 pp (95% interval −57.07 to −50.83 pp) | CCA higher | 2 pp margin not met |

S9 and S11 · BETA L0: accuracy on the seen and unseen frequencies in one test, and FBCCA (descriptive)

| Arm | Seen frequencies | Unseen frequencies |
| --- | --- | --- |
| S9 TPL | 23.4% (95% interval 21.0%–25.7%) | 4.4% (95% interval 4.0%–4.8%) |
| S9 DESC | 24.0% (95% interval 21.5%–26.5%) | 4.2% (95% interval 3.8%–4.7%) |
| S9 SHUF | 23.4% (95% interval 20.9%–25.7%) | 1.1% (95% interval 1.0%–1.4%) |
| S9 NUM | 20.7% (95% interval 18.5%–22.9%) | 11.5% (95% interval 10.2%–12.7%) |
| One-output head, neighbour read-off (nn8) | 44.3% (95% interval 40.9%–47.5%) |  |
| FBCCA, seen 32-way | 81.1% (95% interval 77.0%–84.8%) |  |
| FBCCA, unseen 8-way | 90.4% (95% interval 87.7%–92.8%) |  |
| FBCCA, neighbour 3-way | 92.7% (95% interval 90.9%–94.4%) |  |

### S10, S11 and S13 · BETA L1: other arms against CCA and FBCCA, and held-out wordings · 70 people, three seeds

| Comparison | Difference, pp | Difference shown? | Margin, ±2 pp |
| --- | --- | --- | --- |
| S10 DESC - CCA (8-way) | −62.26 pp (95% interval −66.01 to −58.41 pp) | CCA higher | 2 pp margin not met |
| S10 DESC - CCA (3-way) | −54.93 pp (95% interval −58.01 to −51.72 pp) | CCA higher | 2 pp margin not met |
| S10 DESC - NN(DESC) (8-way) | −1.27 pp (95% interval −2.08 to −0.46 pp) | NN(DESC) higher | 2 pp margin not met |
| S10 NUM - CCA (8-way) | −59.48 pp (95% interval −62.92 to −55.87 pp) | CCA higher | 2 pp margin not met |
| S10 TPL - NN(ID) (8-way) | +0.68 pp (95% interval −0.25 to +1.61 pp) | No difference shown | Equivalent within ±2 pp |
| S10 TPL - SHUF (8-way) | +7.97 pp (95% interval +6.88 to +9.09 pp) | TPL higher | TPL non-inferior at 2 pp |
| S11 TPL - FBCCA (8-way) | −71.41 pp (95% interval −73.53 to −69.17 pp) | FBCCA higher | 2 pp margin not met |
| S11 TPL - FBCCA (3-way) | −65.66 pp (95% interval −67.61 to −63.62 pp) | FBCCA higher | 2 pp margin not met |
| S13 TPL held-out frames - CCA (8-way) | −62.31 pp (95% interval −66.06 to −58.27 pp) | CCA higher | 2 pp margin not met |
| S13 DESC held-out descriptions - CCA (8-way) | −62.58 pp (95% interval −66.15 to −58.59 pp) | CCA higher | 2 pp margin not met |

S9 and S11 · BETA L1: accuracy on the seen and unseen frequencies in one test, and FBCCA (descriptive)

| Arm | Seen frequencies | Unseen frequencies |
| --- | --- | --- |
| S9 TPL | 18.8% (95% interval 16.7%–21.1%) | 2.5% (95% interval 2.2%–2.8%) |
| S9 DESC | 18.9% (95% interval 16.9%–21.2%) | 2.5% (95% interval 2.3%–2.8%) |
| S9 SHUF | 18.3% (95% interval 16.4%–20.4%) | 1.6% (95% interval 1.3%–1.9%) |
| S9 NUM | 11.5% (95% interval 10.4%–12.7%) | 4.3% (95% interval 3.8%–4.8%) |
| One-output head, neighbour read-off (nn8) | 17.6% (95% interval 16.6%–18.8%) |  |
| FBCCA, seen 32-way | 81.1% (95% interval 77.0%–84.8%) |  |
| FBCCA, unseen 8-way | 90.4% (95% interval 87.7%–92.8%) |  |
| FBCCA, neighbour 3-way | 92.7% (95% interval 90.9%–94.4%) |  |

S9 asks about seen and unseen frequencies in one 40-way test; FBCCA, like CCA, needs no training. One seed for FBCCA.

### Secondary results on BOAS

**BOAS is published with three stated gaps (owner decision, 7 October 2026):**

- The consent statement does not say whether participants agreed to public sharing or secondary use.
- The ethics and consent statements come from the publisher's dataset description and README. No peer-reviewed paper describes BOAS.
- The ethics reference was added to the release in version 1.1.1 (May 2025), and the release does not say when it was granted relative to the recordings.

Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.

**Credit:** Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

### S3 · REVE-L features, BOAS · 100 people, one seed · declared likely inconclusive

| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
| --- | --- | --- | --- | --- | --- |
| P1 · Label template − question number — TPL − ID | seen, 5-way | +0.27 pp (95% interval −0.36 to +0.86 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 5-way | +0.12 pp (95% interval −0.48 to +0.71 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −0.39 pp (95% interval −1.04 to +0.05 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −13.54 pp (95% interval −14.57 to −12.49 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −21.86 pp (95% interval −30.62 to −14.12 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | asleep · implicit name | +4.149 (95% interval +3.753 to +4.614) (R 63.39 (42.66–100.85)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | asleep · explicit name | +2.647 (95% interval +2.274 to +3.070) (R 14.11 (9.72–21.55)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | N1, N2 or N3 sleep · explicit name | +1.317 (95% interval +1.163 to +1.483) (R 3.73 (3.20–4.41)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | light sleep · implicit name | +2.730 (95% interval +2.540 to +2.929) (R 15.33 (12.67–18.71)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | light sleep · explicit name | +1.872 (95% interval +1.691 to +2.057) (R 6.50 (5.43–7.82)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | asleep · new descriptions — 10 held-out wordings, resampled | +4.092 (95% interval +3.565 to +4.610) (R 59.84 (35.35–100.44)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | N1, N2 or N3 sleep · new descriptions — 10 held-out wordings, resampled | +1.829 (95% interval +1.250 to +2.292) (R 6.23 (3.49–9.90)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | light sleep · new descriptions — 10 held-out wordings, resampled | +1.774 (95% interval +1.162 to +2.236) (R 5.89 (3.20–9.35)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P5 · Label template − shuffled templates — TPL − SHUF | unseen questions, implicit names | −0.585 (95% interval −0.612 to −0.555) | SHUF higher | Difference only: no margin | Headroom gate passed |
| P5 · Label template − shuffled templates — TPL − SHUF | unseen questions, explicit names | +0.626 (95% interval +0.602 to +0.648) | TPL higher | Difference only: no margin | Headroom gate passed |

### S4 · text encoder all-MiniLM-L6-v2, BOAS, frozen CBraMod features · 100 people, one seed · declared likely inconclusive

| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
| --- | --- | --- | --- | --- | --- |
| P1 · Label template − question number — TPL − ID | seen, 5-way | +0.19 pp (95% interval −0.37 to +0.75 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 5-way | +0.01 pp (95% interval −0.64 to +0.72 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −1.72 pp (95% interval −4.31 to −0.09 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −14.04 pp (95% interval −15.29 to −12.68 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −25.74 pp (95% interval −34.29 to −16.26 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | asleep · implicit name | +3.798 (95% interval +3.385 to +4.273) (R 44.60 (29.53–71.76)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | asleep · explicit name | +1.325 (95% interval +0.999 to +1.715) (R 3.76 (2.71–5.56)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | N1, N2 or N3 sleep · explicit name | +0.882 (95% interval +0.717 to +1.047) (R 2.41 (2.05–2.85)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | light sleep · implicit name | +2.338 (95% interval +2.134 to +2.552) (R 10.36 (8.45–12.83)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | light sleep · explicit name | +1.796 (95% interval +1.598 to +2.000) (R 6.03 (4.94–7.39)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | asleep · new descriptions — 10 held-out wordings, resampled | +3.069 (95% interval +2.508 to +3.601) (R 21.52 (12.28–36.65)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | N1, N2 or N3 sleep · new descriptions — 10 held-out wordings, resampled | +1.118 (95% interval +0.793 to +1.449) (R 3.06 (2.21–4.26)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | light sleep · new descriptions — 10 held-out wordings, resampled | +1.070 (95% interval +0.840 to +1.333) (R 2.92 (2.32–3.79)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P5 · Label template − shuffled templates — TPL − SHUF | unseen questions, implicit names | −0.626 (95% interval −0.653 to −0.596) | SHUF higher | Difference only: no margin | Headroom gate passed |
| P5 · Label template − shuffled templates — TPL − SHUF | unseen questions, explicit names | +0.572 (95% interval +0.550 to +0.592) | TPL higher | Difference only: no margin | Headroom gate passed |

### S5 · Chinese wordings with multilingual-e5-small, BOAS, frozen CBraMod features · 100 people, one seed · declared likely inconclusive

| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
| --- | --- | --- | --- | --- | --- |
| P1 · Label template − question number — TPL − ID | seen, 5-way | +0.34 pp (95% interval −0.29 to +0.99 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P1 · Description − question number — DESC − ID | seen, 5-way | −0.28 pp (95% interval −0.93 to +0.42 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · New sentence frame − training wording — TPL(a1) − TPL(train) | seen, reworded — 12 held-out wordings, resampled | −0.70 pp (95% interval −1.83 to +0.08 pp) | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · Synonym − training wording — TPL(a2) − TPL(train) | seen, reworded — for these three synonyms: fixed, not resampled | −12.51 pp (95% interval −13.73 to −11.27 pp) | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptions — DESC(a3) − DESC(train) | seen, reworded — 10 held-out wordings, resampled | −25.81 pp (95% interval −33.97 to −17.73 pp) | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | asleep · implicit name | +3.826 (95% interval +3.424 to +4.277) (R 45.87 (30.69–72.04)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | asleep · explicit name | +1.558 (95% interval +1.216 to +1.950) (R 4.75 (3.37–7.03)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | N1, N2 or N3 sleep · explicit name | +0.737 (95% interval +0.584 to +0.898) (R 2.09 (1.79–2.45)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | light sleep · implicit name | +2.322 (95% interval +2.122 to +2.532) (R 10.19 (8.34–12.57)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answers — TPL / read-off | light sleep · explicit name | +1.615 (95% interval +1.424 to +1.809) (R 5.03 (4.15–6.11)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | asleep · new descriptions — 10 held-out wordings, resampled | +2.715 (95% interval +2.127 to +3.244) (R 15.11 (8.39–25.63)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | N1, N2 or N3 sleep · new descriptions — 10 held-out wordings, resampled | +0.798 (95% interval +0.599 to +1.039) (R 2.22 (1.82–2.83)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answers — DESC / read-off | light sleep · new descriptions — 10 held-out wordings, resampled | +0.943 (95% interval +0.773 to +1.123) (R 2.57 (2.17–3.07)) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |

### S13 · held-out sentence frames against the read-off, BOAS · 100 people, three seeds

| Comparison | log R, with R | Difference shown? | Margin, R < 1.25 |
| --- | --- | --- | --- |
| TPL asleep implicit (held-out frames) | +3.565 (95% interval +3.162 to +4.006) | Read-off leaves less error | 1.25 ratio margin not met |
| TPL asleep explicit (held-out frames) | +1.976 (95% interval +1.594 to +2.392) | Read-off leaves less error | 1.25 ratio margin not met |
| TPL NREM explicit (held-out frames) | +0.920 (95% interval +0.743 to +1.100) | Read-off leaves less error | 1.25 ratio margin not met |
| TPL light implicit (held-out frames) | +2.233 (95% interval +2.022 to +2.443) | Read-off leaves less error | 1.25 ratio margin not met |
| TPL light explicit (held-out frames) | +1.484 (95% interval +1.240 to +1.719) | Read-off leaves less error | 1.25 ratio margin not met |

### S6 · negation probe, BOAS: a boundary probe, never a route sentence

AUROC of each negated wording against the truth of the negated question, beside the AUROC of 1 − p(X), the head’s own answer to X turned around; three seeds. “Non-REM sleep” is scored against the truth of “not REM”.

| Wording | Against the truth of the negated question | AUROC of 1 − p(X) | People |
| --- | --- | --- | --- |
| not W | 0.022 (95% interval 0.014–0.030) | 0.980 (95% interval 0.972–0.987) | 100 |
| not N1 | 0.135 (95% interval 0.120–0.152) | 0.872 (95% interval 0.855–0.887) | 100 |
| not N2 | 0.124 (95% interval 0.110–0.139) | 0.924 (95% interval 0.909–0.937) | 100 |
| not N3 | 0.021 (95% interval 0.017–0.026) | 0.979 (95% interval 0.974–0.983) | 71 |
| not REM | 0.049 (95% interval 0.038–0.061) | 0.956 (95% interval 0.944–0.967) | 100 |
| non-REM sleep (implicit) | 0.048 (95% interval 0.037–0.060) | 0.956 (95% interval 0.944–0.967) | 100 |

Every negated wording was answered as if it asked for X itself: against the truth of the negated question its AUROC lay far below 0.5, while turning the head’s answer to X around would have answered it well. Writing more prompts does not supply signal-level ground truth.

Published with the results

## The wordings, the held-out frequencies and the shuffles

All 787 texts the questions were asked with, in English and Chinese, for sleep, SSVEP and motor imagery: training and held-out sentence frames, label names and synonyms, descriptions, the names and descriptions of the unseen questions, and the negation frames. Two agents wrote them from a fixed brief before any score, the second writing every held-out description without seeing the first’s; they are not users’ wordings and were not tuned. The file also holds the leakage rules they passed, the held-out frequencies of every rotation and the extrapolation band, the declared derangements of the shuffled control, the numeric code of a frequency and the pinned text encoders. Text only, [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/).

[The wordings, as JSON ↓](https://bci.report/data/questions-in-language-wordings.json)

The sleep stages’ names in the training frames, and the three synonyms each was reworded with:

| Stage | Training name | Synonyms (a2) |
| --- | --- | --- |
| W | wake | wakefulness, awake, stage W |
| N1 | N1 sleep | stage 1 sleep, sleep stage one, stage N1 |
| N2 | N2 sleep | stage 2 sleep, sleep stage two, stage N2 |
| N3 | N3 sleep | slow-wave sleep, deep sleep, stage N3 |
| REM | REM sleep | stage R sleep, paradoxical sleep, dream sleep |

Methods & limits

## What these results can and cannot say

- 'Unseen' means unseen by the EEG head; the text encoder has read these words, and every token of a held-out frequency also occurs in the training texts.
- Held-out sleep questions are combinations of seen stages ('asleep' is the complement of one), reported as compositional beside the read-off.
- Held-out SSVEP questions name a frequency, which CCA answers without training; interpolation tested, extrapolation secondary.
- In the neighbour 3-way the unseen frequency competes against trained detectors for its neighbours; at L1 that test also involves stimulus phase.
- Rewording and unseen results are separate rows and never pooled.
- Wordings were written by agents from a fixed brief before scoring, not by users and not tuned.
- Negation is a boundary probe; location and waveform-shape questions are not tested (no signal-level ground truth held).
- Frozen features and a small head; one recipe; not a ranking of text or EEG encoders.
- 35 pre-declared comparisons, not corrected for multiplicity.
- Questions about negation, location or waveform shape need signal-level ground truth. Writing more prompts does not supply it.
- Negation was not understood: the negated wordings ("not X", "non-REM sleep") scored as if they asked for X (AUROC 0.02-0.13 against the truth of the negated question; boundary probe S6, never a route sentence).
- Implicit names of the unseen sleep unions ("asleep", "light sleep") ranked windows in reverse (TPL AUROC 0.317 and 0.244, below 0.5); only the explicit names ("N1, N2, N3 or REM sleep") carried the composition.
- The neighbour 3-way at L1 involves stimulus phase as well as frequency; at L0 (a phase-free spectrum) it is the clean frequency test.
- Unseen SSVEP questions compete against trained detectors for their neighbours (seen bias, stated).
- The pre-run permutation canary failed twice (revisions 3 and 4) and was revised by the owner, who overrode revision 4's "no further revision"; revision 5 passed. Both failed records are kept and disclosed.
- For multilingual-e5-small each held-out frequency is a new token (S0b-6); S8 and S12 also involve new tokens.
- Secondary single-seed entries carry no training-run variance; EESM19 has 20 people; every EESM19 and single-seed entry was declared likely inconclusive.
- BETA numbers use the published 2-s windows and folds; Wearable-102 numbers are computed from the 2026-09-20 Tsinghua author mirror and claim no physical unit.
- BOAS: the three gaps and the attribution above; participants are pseudonymised in the public release; no result evaluates Bitbrain's headband or its automatic scoring.
- The head here is not a language model, and none of these results shows that EEG understands language.

### Not run, and not reported

- **revision-4 per-person centring diagnostic**: not run (revision 5 part A.2; no fit exists)
- **Ying scorer agreement; factorial cognitive sets**: deferred (D10)
- **E1 EEGNet trained jointly with the question vector**: deferred (D6)
- **third text encoder; LEAF-style Q-Former heads; route-1 calibration on language heads; user-written wordings**: deferred (protocol.deferred)
- **S8 TPL - NN(TPL)**: not defined (EA-1): the held-out band is contiguous, so no held-out frequency has two seen neighbours
- **location and waveform-shape questions**: not tested: no signal-level ground truth held
- **per-union paired contrasts beyond the P3 entries and the per-union levels**: not computed (W1)
- **counts of people above or below zero on paired contrasts**: allowed for BOAS but not computed by stage 2; not added
- **per-person distributions, per-fold values, per-window scores**: private by design

### Not published

- Per-person, per-night, per-recording and per-fold values, per-window scores, features, head weights and fold assignments.
- The per-frequency and interior-only SSVEP breakdowns: computed after the audits, which do not cover them.
- The values of the pre-run checks, which are published as pass or fail and as counts.
- Measured compute, timed on a shared GPU.
- Demographics, recording dates and clock times, and any image of a real recording.

### The text encoders

- [BAAI/bge-small-en-v1.5 ↗](https://huggingface.co/BAAI/bge-small-en-v1.5): primary · MIT · `5c38ec7`
- [sentence-transformers/all-MiniLM-L6-v2 ↗](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2): secondary, S4 · Apache-2.0 · `1110a24`
- [intfloat/multilingual-e5-small ↗](https://huggingface.co/intfloat/multilingual-e5-small): Chinese wordings, S5 · MIT · `614241f`

Frozen at pinned commits and never redistributed.

### The recordings

### BETA

Bingchuan Liu et al. · BETA: A Large Benchmark Database Toward SSVEP-BCI Application (2020), doi:10.3389/fnins.2020.00627. Figshare 12264401 v3; mirror Bingchuan/BETA.

[Source ↗](https://figshare.com/articles/dataset/The_BETA_database/12264401) · [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)

### BOAS, the Bitbrain Open Access Sleep dataset

Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

Published with three stated gaps, printed beside its results above.

[Source ↗](https://openneuro.org/datasets/ds005555/versions/1.1.3) · [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/)

### EESM19

Kaare B. Mikkelsen et al. · Accurate whole-night sleep monitoring with dry-contact ear-EEG (2019), doi:10.1038/s41598-019-53115-3; OpenNeuro ds005185 v1.0.2. Processed mirror: Zachary1150/EESM19-Processed.

[Source ↗](https://doi.org/10.18112/openneuro.ds005185.v1.0.2) · [CC0-1.0 declared by upstream and mirror](https://creativecommons.org/publicdomain/zero/1.0/)

### OpenBMI motor imagery

Min-Ho Lee, O-Yeon Kwon, Yong-Jeong Kim, Hong-Kyung Kim, Young-Eun Lee, John Williamson, Siamac Fazli and Seong-Whan Lee · EEG dataset and OpenBMI toolbox for three BCI paradigms: an investigation into BCI illiteracy, GigaScience (2019), giz002, doi:10.1093/gigascience/giz002. Data: Supporting data, GigaScience Database, doi:10.5524/100542.

[Source ↗](https://doi.org/10.5524/100542) · [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/)

### Wearable SSVEP BCI dataset (dry and wet electrodes)

Zhu, F., Jiang, L., Dong, G., Gao, X., & Wang, Y. (2021). An Open Dataset for Wearable SSVEP-Based Brain-Computer Interfaces (Version 4) [Data set]. Figshare. https://doi.org/10.6084/m9.figshare.13560281.v4

Zhu, F., Jiang, L., Dong, G., Gao, X., & Wang, Y. (2021). An Open Dataset for Wearable SSVEP-Based Brain-Computer Interfaces. Sensors, 21(4), 1256. https://doi.org/10.3390/s21041256

Computed from the official Tsinghua BCI Lab author mirror snapshot acquired 2026-09-20; the mirror itself has no separate DOI or version number.

[Source ↗](https://figshare.com/articles/dataset/An_Open_Dataset_for_Wearable_SSVEP-Based_Brain-Computer_Interfaces/13560281) · [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)

Features: frozen [CBraMod](https://bci.report/methods/cbramod/) (primary) and frozen [REVE](https://bci.report/methods/reve/) Large @ 317531c7, REVE-L (secondary).

Weights terms: REVE under the REVE Responsible Use License v1.0 (model versions: REVE Large @ 317531c7). No model’s authors endorse these results. [REVE, arXiv:2510.21585 ↗](https://arxiv.org/abs/2510.21585)

### Methods and sources these results draw on

- [Yu & Yao, 2026 · Visual Jev: Accurate and Efficient Decisions from Shared Visual Context ↗](https://arxiv.org/abs/2609.25845)
- [Perez et al., 2017 · FiLM: Visual Reasoning with a General Conditioning Layer ↗](https://arxiv.org/abs/1709.07871)
- [Lin, Zhang, Wu & Gao, 2006 · Frequency Recognition Based on CCA for SSVEP-Based BCIs, IEEE TBME 53:2610–2614 ↗](https://doi.org/10.1109/TBME.2006.886577)
- [Chen, Wang, Gao, Jung & Gao, 2015 · Filter bank CCA for implementing a high-speed SSVEP-based BCI, J. Neural Eng. 12:046008 ↗](https://doi.org/10.1088/1741-2560/12/4/046008)
- [Wallace et al., 2019 · Do NLP Models Know Numbers? Probing Numeracy in Embeddings, EMNLP 2019 ↗](https://aclanthology.org/D19-1534/)
- [Weller, Lawrie & Van Durme, 2024 · NevIR: Negation in Neural Information Retrieval, EACL 2024 ↗](https://aclanthology.org/2024.eacl-long.139/)
- [Xian, Lampert, Schiele & Akata · Zero-Shot Learning — A Comprehensive Evaluation of the Good, the Bad and the Ugly, TPAMI ↗](https://arxiv.org/abs/1707.00600)

Data source: [questions-in-language-update.json](https://bci.report/data/questions-in-language-update.json) · schema bci-report-questions-in-language-update-v1.

Wordings: [questions-in-language-wordings.json](https://bci.report/data/questions-in-language-wordings.json) · schema bci-report-questions-in-language-wordings-v1.

Keep exploring

Transfer

- [— Sensor transfer **Dry vs. wet electrodes**](https://bci.report/topics/dry-vs-wet/)
- [— Context transfer **Screen to VR**](https://bci.report/topics/screen-to-vr/)
- [— Montage **Fewer electrodes**](https://bci.report/topics/fewer-electrodes/)
- [— Motion robustness **On the move**](https://bci.report/topics/on-the-move/)
- [— Session transfer · same person **Later sessions**](https://bci.report/topics/later-sessions/)

Adapting models

- [— Calibration budget **How much calibration?**](https://bci.report/topics/calibration-budget/)
- [— Model adaptation **Which part to update?**](https://bci.report/topics/model-adaptation/)
- [— Representation controls **Does pretraining help?**](https://bci.report/topics/does-pretraining-help/)
- [— Shared encoder · route 2 **One model, several questions**](https://bci.report/topics/shared-encoder/)
- [— Questions in language · route 3 **Questions in language**](https://bci.report/topics/questions-in-language/)

Reliability & clinical

- [— Abstention · decision research **When not to act**](https://bci.report/topics/when-not-to-act/)
- [— Sleep staging · simple baselines **Sleep-stage balance**](https://bci.report/topics/sleep-staging/)
- [— Clinical research **Clinical groups**](https://bci.report/topics/clinical-groups/)

[All questions, and the map of which kinds of transfer have been measured →](https://bci.report/topics/)

## Cite this page

BCI Report (2026). *Can an EEG model answer questions asked in words?* https://bci.report/topics/questions-in-language/

Figures from release [`questions-in-language-update-20261008`](https://bci.report/releases/#questions-in-language-update-20261008) (2026-10-08). Cite the upstream datasets as well: their credits are on this page.

Every release is archived on Zenodo: [doi:10.5281/zenodo.23123296](https://doi.org/10.5281/zenodo.23123296). [BibTeX for the site and its releases →](https://bci.report/api/#cite-heading) · [CITATION.cff ↗](https://raw.githubusercontent.com/twu3202/bci-report/main/CITATION.cff)

---
Markdown copy of https://bci.report/topics/questions-in-language/, generated from the published page. Figures are aggregate results; terms of use: https://bci.report/data-use/
