BETA
Bingchuan Liu et al. · BETA: A Large Benchmark Database Toward SSVEP-BCI Application (2020), doi:10.3389/fnins.2020.00627. Figshare 12264401 v3; mirror Bingchuan/BETA.
Questions in language · route 3 of the decision-research plan
One small head over frozen EEG features, asked each question by a question number, a label template or a natural-language description: SSVEP on BETA, sleep on BOAS. Seen questions, rewordings and questions the EEG head was never trained on are reported apart, never as one zero-shot number.
Measured on: BETA; BOAS · Bitbrain Open Access Sleep dataset · Methods: CBraMod
Only for questions it was trained on, and not on every dataset. A small head over frozen EEG features, asked a seen question with a label template or a description instead of a question number, was equivalent within the 2 pp margin on sleep (BOAS), but lost accuracy on SSVEP (BETA, 70 people): −5.96 pp with templates on a plain spectrum (−6.84 to −5.08 pp), the margin not met. Rewording a seen question cost accuracy too, except a new sentence frame on sleep. Questions the EEG head was never trained on were not answered usefully: every unseen sleep question did worse than adding up the head’s own stage answers, and on unseen flicker frequencies the head reached 28.7% on a plain spectrum where CCA, which needs no training, reached 80.9% (chance 12.5%). Explicit wordings such as “N1 or N2 sleep” used the link between words and stages; the everyday names “asleep” and “light sleep” ranked windows backwards, and in a boundary probe negated questions were answered as if they asked for what they negate.
How the comparison works
The same small head answers every question; only the vector that tells it which question changes. Every test is on people the head never trained on.
A training frame for sleep, filled with a stage’s name: The scorer marked this 30-second epoch as N2 sleep.
SSVEP on BETA: a plain spectrum (L0) and frozen CBraMod features (L1). Sleep on BOAS: frozen CBraMod features (L1). Three training seeds each. The primary text encoder is bge-small-en-v1.5, frozen. The head is the same for every arm: standardised features, a small hidden layer adjusted by the question vector (FiLM), one output.
Three kinds of result, never pooled into one zero-shot number: a seen question in the wording the head trained with; the same question reworded; and a question the head never trained on. “Unseen” always means unseen by the EEG head, not by the text encoder: the frozen encoder has read the words, and for SSVEP every token of a held-out frequency also occurs in that rotation’s training texts. Unseen sleep questions are combinations of seen stages; unseen SSVEP questions are flicker frequencies held out in five rotations.
Each comparison is the language arm minus the reference, in percentage points (pp), with a 95% interval from resampling people, then training seeds, then held-out wordings. It gets two readings, both fixed before any fit. A difference is shown when the interval excludes zero. Against a margin of ±2 pp, it is equivalent when the whole interval lies inside, non-inferior when the bound against the language arm stays within the margin, and otherwise the margin is not met, which is not the same as a loss of 2 pp or more. For an unseen sleep question the comparison is R, the head’s remaining error (1 − AUROC) over that of the read-off from its own stage answers, with a margin of R < 1.25. The check of whether the head uses the link between wording and stage, the template arm against the shuffled templates, is a difference of AUROC. Among the primary comparisons, those with the numeric code, the neighbour average and the shuffled templates are read for a difference only; the secondary tables print both readings as the run computed them.
There are 35 pre-declared primary comparisons and no multiplicity correction, so at 95% about one in twenty comparisons with no true difference may show one by chance. Accuracy is averaged per person, then over people: it compares set-ups and is not a deployment error rate.
This is route 3 of the research plan on When not to act. “Jev-style” here means the interface of a decision model like TypeSafe’s Jev, applied to EEG: encode the recording once, then answer several explicit, typed questions about it, each with a probability. The plan started from a vision paper, Yu & Yao, 2026 · Visual Jev: Accurate and Efficient Decisions from Shared Visual Context ↗, which applies that idea to images. All three routes: Jev-style questions on EEG: the evidence →
Primary · seen questions · BETA and BOAS
Questions the head was trained on. On BETA, 70 people look at flickering targets; in each rotation the head trains on most of the frequencies and is asked about those (the 32-way task). On BOAS, 100 people, each half-minute epoch is scored into one of five stages at their natural mix, against the human consensus.
BOAS is published with three stated gaps (owner decision, 7 October 2026):
Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.
Credit: Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.
On sleep, a template or a description instead of a question number was equivalent within ±2 pp: −0.05 pp (−0.70 to +0.61 pp) and +0.02 pp (−0.59 to +0.62 pp). On SSVEP it cost accuracy at both levels, and the margin was not met: templates −5.96 pp on a plain spectrum and −3.49 pp on frozen CBraMod features; descriptions −5.46 pp and −3.37 pp.
| Comparison | Question | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
|---|---|---|---|---|---|
| SSVEP · BETA · a plain spectrum (L0) | |||||
| P1 · Label template − question numberTPL − ID | seen, 32-way | −5.96 pp95% interval −6.84 to −5.08 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 32-way | −5.46 pp95% interval −6.26 to −4.66 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| SSVEP · BETA · frozen CBraMod features (L1) | |||||
| P1 · Label template − question numberTPL − ID | seen, 32-way | −3.49 pp95% interval −4.10 to −2.93 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 32-way | −3.37 pp95% interval −3.98 to −2.80 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| Sleep · BOAS · frozen CBraMod features (L1) | |||||
| P1 · Label template − question numberTPL − ID | seen, 5-way | −0.05 pp95% interval −0.70 to +0.61 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 5-way | +0.02 pp95% interval −0.59 to +0.62 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
The seen SSVEP task is hard for every head on these features: with question numbers the head reached 32.1% (28.8%–35.2%) on a plain spectrum and 24.2% on frozen CBraMod features, against a chance of 3.1%; CCA, which needs no training, reached 65.5%. On sleep every arm reached about the same balanced accuracy, 71.5% with question numbers against a chance of 20.0%. The levels below are descriptive: overlapping intervals are not ranked, and the paired differences above are the comparison.
| Arm | A plain spectrum (L0) | chance-corrected | Frozen CBraMod features (L1) | chance-corrected |
|---|---|---|---|---|
| IDquestion number | 32.1%95% interval 28.8%–35.2% | 29.9% | 24.2%95% interval 21.6%–26.9% | 21.8% |
| TPLlabel template | 26.1%95% interval 23.5%–28.6% | 23.7% | 20.7%95% interval 18.5%–23.1% | 18.2% |
| DESCdescription | 26.6%95% interval 23.9%–29.2% | 24.2% | 20.9%95% interval 18.7%–23.2% | 18.3% |
| SHUFshuffled templates | 25.1%95% interval 22.5%–27.6% | 22.7% | 20.2%95% interval 18.1%–22.3% | 17.6% |
| NUMnumeric code | 23.7%95% interval 21.2%–26.1% | 21.2% | 13.9%95% interval 12.5%–15.2% | 11.1% |
| B-shone output per question | 32.2%95% interval 28.9%–35.3% | 30.0% | 27.0%95% interval 24.2%–30.1% | 24.7% |
| CCAno training | 65.5%95% interval 59.7%–71.1% | 64.3% | 65.5%95% interval 59.7%–71.1% | 64.3% |
| Arm | Balanced accuracy | chance-corrected | Log loss |
|---|---|---|---|
| IDquestion number | 71.5%95% interval 69.1%–73.5% | 64.3% | 0.26795% interval 0.242–0.296 |
| TPLlabel template | 71.4%95% interval 69.1%–73.4% | 64.3% | 0.27695% interval 0.249–0.307 |
| DESCdescription | 71.5%95% interval 69.1%–73.5% | 64.3% | 0.27895% interval 0.251–0.310 |
| SHUFshuffled templates | 71.5%95% interval 69.2%–73.5% | 64.4% | — |
| B-shone output per question | 71.4%95% interval 69.1%–73.4% | 64.3% | 0.25995% interval 0.236–0.287 |
| Arm | W100 people | N1100 people | N2100 people | N371 people | REM100 people |
|---|---|---|---|---|---|
| IDquestion number | 0.9820.976–0.987 | 0.8760.859–0.890 | 0.9220.904–0.937 | 0.9780.971–0.983 | 0.9570.945–0.967 |
| TPLlabel template | 0.9800.972–0.987 | 0.8720.855–0.887 | 0.9240.909–0.937 | 0.9790.974–0.983 | 0.9560.944–0.967 |
| DESCdescription | 0.9810.974–0.987 | 0.8690.851–0.884 | 0.9230.907–0.937 | 0.9790.974–0.983 | 0.9560.944–0.967 |
The same seen questions in wordings the head never trained with: a new sentence frame (a1, held-out frames resampled), a synonym of the label (a2: three per question, fixed, so these entries hold only for these three synonyms), or a new description (a3, held-out descriptions resampled). Reworded rows are never pooled with the unseen ones.
| Comparison | Question | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
|---|---|---|---|---|---|
| SSVEP · BETA · a plain spectrum (L0) | |||||
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −4.01 pp95% interval −6.53 to −2.12 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −1.84 pp95% interval −2.10 to −1.57 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −4.13 pp95% interval −5.58 to −3.11 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| SSVEP · BETA · frozen CBraMod features (L1) | |||||
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −2.72 pp95% interval −4.43 to −1.43 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −1.33 pp95% interval −1.53 to −1.14 pp | TPL(train) higher | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −2.95 pp95% interval −3.81 to −2.26 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| Sleep · BOAS · frozen CBraMod features (L1) | |||||
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −0.46 pp95% interval −1.33 to +0.16 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −12.60 pp95% interval −13.74 to −11.43 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −20.87 pp95% interval −29.49 to −13.71 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
On sleep a new sentence frame cost nothing measurable: equivalent within ±2 pp (−0.46 pp, −1.33 to +0.16 pp). The three synonyms per stage scored −12.60 pp (−13.74 to −11.43 pp) and new descriptions −20.87 pp (−29.49 to −13.71 pp), a wide interval, because descriptions differ a lot from one another. On BETA every rewording scored lower, from −1.33 pp to −4.13 pp; only the synonyms on frozen CBraMod features stayed within the margin.
| Synonym | SSVEP · BETA · a plain spectrum (L0) | SSVEP · BETA · frozen CBraMod features (L1) | Sleep · BOAS · frozen CBraMod features (L1) |
|---|---|---|---|
| Synonym 1 | −1.45 pp95% interval −1.67 to −1.22 pp | −1.09 pp95% interval −1.27 to −0.90 pp | −23.53 pp95% interval −25.29 to −21.64 pp |
| Synonym 2 | −0.87 pp95% interval −1.04 to −0.70 pp | −0.84 pp95% interval −1.00 to −0.68 pp | −10.85 pp95% interval −12.59 to −9.17 pp |
| Synonym 3 | −3.19 pp95% interval −3.69 to −2.69 pp | −2.06 pp95% interval −2.44 to −1.73 pp | −3.42 pp95% interval −4.49 to −2.41 pp |
| New frame | Synonym | Between training frames | New description | Between training descriptions | Between seeds of ID | |
|---|---|---|---|---|---|---|
| BETA · L0 | 39.7% | 29.3% | 19.4% | 41.2% | 19.3% | 56.8% |
| BETA · L1 | 40.7% | 32.0% | 24.3% | 42.7% | 23.3% | 62.2% |
| BOAS · L1 | 6.4% | 20.0% | 3.3% | 45.4% | 3.1% | 7.4% |
Primary · unseen sleep questions · BOAS
Three questions the head never trained on, each a combination of the stages it did: asleep (N1, N2, N3 or REM), N1, N2 or N3 sleep, and light sleep (N1 or N2). Each was asked with templates, using an implicit name or an explicit one, and with new descriptions; 100 people, three seeds. Unseen by the EEG head, not by the text encoder: the encoder has read these words. Because each question is a combination, the head’s own stage answers can be added up into an answer, the read-off, which is what each wording is held to.
BOAS is published with three stated gaps (owner decision, 7 October 2026):
Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.
Credit: Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.
| Question | Implicit name | Explicit name |
|---|---|---|
| Asleep | asleep | N1, N2, N3 or REM sleep |
| N1, N2 or N3 sleep | none | N1, N2 or N3 sleep |
| Light sleep | light sleep | N1 or N2 sleep |
In all eight conditions the interval lay above R = 1, in the read-off’s favour, and none came within the 1.25 margin: R ran from 2.42 (2.07–2.85; N1, N2 or N3 sleep, explicit name) to 33.80 (23.09–52.82; asleep, implicit name).
| Comparison | Question | log R, with R | Difference shown? | Margin, R < 1.25 | Gate |
|---|---|---|---|---|---|
| P3 · Label template head against the read-off from its own answersTPL / read-off | asleep · implicit name | +3.52095% interval +3.139 to +3.967R 33.80 (23.09–52.82) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | asleep · explicit name | +1.95195% interval +1.623 to +2.357R 7.04 (5.07–10.56) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | N1, N2 or N3 sleep · explicit name | +0.88395% interval +0.726 to +1.047R 2.42 (2.07–2.85) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | light sleep · implicit name | +2.26795% interval +2.072 to +2.466R 9.65 (7.94–11.77) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | light sleep · explicit name | +1.51995% interval +1.343 to +1.698R 4.57 (3.83–5.46) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | asleep · new descriptions10 held-out wordings, resampled | +3.44495% interval +2.913 to +3.956R 31.30 (18.41–52.26) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | N1, N2 or N3 sleep · new descriptions10 held-out wordings, resampled | +1.29995% interval +0.793 to +1.744R 3.67 (2.21–5.72) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | light sleep · new descriptions10 held-out wordings, resampled | +1.41795% interval +0.903 to +1.856R 4.12 (2.47–6.40) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| Question | Wording | Template head | Read-off from its stage answers |
|---|---|---|---|
| Asleep | implicit name | 0.31795% interval 0.294–0.341 | 0.98095% interval 0.971–0.987 |
| Asleep | explicit name | 0.85895% interval 0.839–0.875 | 0.98095% interval 0.971–0.987 |
| N1, N2 or N3 sleep | explicit name | 0.87195% interval 0.854–0.887 | 0.94795% interval 0.934–0.958 |
| Light sleep | implicit name | 0.24495% interval 0.222–0.267 | 0.92295% interval 0.907–0.935 |
| Light sleep | explicit name | 0.64295% interval 0.616–0.667 | 0.92295% interval 0.907–0.935 |
The implicit names ranked windows in reverse: AUROC 0.317 for “asleep” and 0.244 for “light sleep”, below 0.5, so windows that really were asleep, or in light sleep, scored lower on these questions than the others. The explicit names did better (0.858 and 0.642) but stayed short of the read-off (0.980 and 0.922). Descriptions have no alignment control: the shuffled control is built from template vectors. Held-out sentence frames (S13, in the secondary panel) give the same result for all five template cells.
The shuffled-template arm attaches the same template vectors to the wrong stages during training; the unseen questions keep their true vectors. If the head uses the link between a wording and its stage, the shuffled arm should do worse on them than the template arm.
| Comparison | Question | Difference, AUROC | Difference shown? | Margin | Gate |
|---|---|---|---|---|---|
| P5 · Label template − shuffled templatesTPL − SHUF | unseen questions, implicit names | −0.37895% interval −0.501 to −0.253 | SHUF higher | Difference only: no margin | Headroom gate passed |
| P5 · Label template − shuffled templatesTPL − SHUF | unseen questions, explicit names | +0.62995% interval +0.603 to +0.654 | TPL higher | Difference only: no margin | Headroom gate passed |
| Question | Wording | Template arm | Shuffled templates |
|---|---|---|---|
| Asleep | implicit name | 0.31795% interval 0.294–0.341 | 0.67695% interval 0.529–0.812 |
| Asleep | explicit name | 0.85895% interval 0.839–0.875 | 0.13795% interval 0.109–0.168 |
| N1, N2 or N3 sleep | explicit name | 0.87195% interval 0.854–0.887 | 0.17195% interval 0.145–0.201 |
| Light sleep | implicit name | 0.24495% interval 0.222–0.267 | 0.64195% interval 0.537–0.738 |
| Light sleep | explicit name | 0.64295% interval 0.616–0.667 | 0.17695% interval 0.155–0.199 |
For explicit names, yes: attaching the templates to the wrong stages lowered the unseen AUROC, TPL − SHUF +0.629 (+0.603 to +0.654). For the implicit names it was not shown: the shuffled templates scored higher, −0.378 (−0.501 to −0.253).
Primary · unseen flicker frequencies · BETA
In each of five rotations, eight of BETA’s frequencies are held out: the head trains on the rest and is asked about the eight by stating the frequency in a template. 70 people, three seeds, at both levels. Unseen by the EEG head, not by the text encoder: every token of a held-out frequency also occurs in that rotation’s training texts, so only the combination is new. CCA answers the same question with no training at all.
In the 8-way among unseen frequencies the head reached 28.7% (26.9%–30.5%) on a plain spectrum and 19.0% on frozen CBraMod features; CCA reached 80.9%, and chance is 12.5%. The head also did worse than averaging its own answers for the two neighbouring seen frequencies, and worse than a numeric code of the frequency, at both levels.
| Comparison | Question | Difference, pp | Difference shown? | Margin, ±2 pp | Gate |
|---|---|---|---|---|---|
| SSVEP · BETA · a plain spectrum (L0) | |||||
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, 8-way | −52.21 pp95% interval −55.13 to −48.99 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL) | unseen frequency, 8-way | −10.98 pp95% interval −12.24 to −9.67 pp | NN(TPL) higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, 8-way | −12.77 pp95% interval −14.52 to −10.99 pp | NUM higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, neighbour 3-way | −62.36 pp95% interval −65.87 to −58.69 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, neighbour 3-way | −1.24 pp95% interval −3.17 to +0.85 pp | No difference shown | Difference only: no margin | Headroom gate passed |
| SSVEP · BETA · frozen CBraMod features (L1) | |||||
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, 8-way | −61.94 pp95% interval −65.65 to −58.06 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL) | unseen frequency, 8-way | −1.16 pp95% interval −1.97 to −0.37 pp | NN(TPL) higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, 8-way | −2.46 pp95% interval −3.23 to −1.69 pp | NUM higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, neighbour 3-way | −54.67 pp95% interval −57.77 to −51.47 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, neighbour 3-way | +15.61 pp95% interval +14.25 to +16.93 pp | TPL higher | Difference only: no margin | Headroom gate passed |
In the neighbour 3-way the unseen frequency competes with the trained detectors of its two neighbours (seen bias, stated). On a plain spectrum the template head fell below the mean chance of that test, 19.3% against 34.2%, and showed no difference against the numeric code. On frozen CBraMod features it beat the numeric code by +15.61 pp, but those features keep the stimulus phase and neighbouring targets differ in phase, so that comparison tests frequency and phase together. The plain spectrum, which carries no phase, is the clean frequency test.
| Arm | Test | A plain spectrum (L0) | chance-corrected | Frozen CBraMod features (L1) | chance-corrected |
|---|---|---|---|---|---|
| TPLlabel template | u8 | 28.7%95% interval 26.9%–30.5% | 18.5% | 19.0%95% interval 18.0%–19.9% | 7.4% |
| DESCdescription | u8 | 28.6%95% interval 26.8%–30.4% | 18.4% | 18.6%95% interval 17.6%–19.6% | 7.0% |
| SHUFshuffled templates | u8 | 9.2%95% interval 8.6%–9.7% | -3.8% | 11.0%95% interval 10.5%–11.5% | -1.7% |
| NUMnumeric code | u8 | 41.5%95% interval 38.2%–44.7% | 33.1% | 21.4%95% interval 20.2%–22.6% | 10.2% |
| CCAno training | u8 | 80.9%95% interval 76.6%–84.8% | 78.2% | 80.9%95% interval 76.6%–84.8% | 78.2% |
| TPLlabel template | nn8 | 39.7%95% interval 36.8%–42.5% | 31.0% | 20.1%95% interval 19.0%–21.3% | 8.7% |
| IDquestion number | nn8 | 44.3%95% interval 40.8%–47.6% | 36.4% | 18.3%95% interval 17.3%–19.3% | 6.6% |
| B-shone output per question | nn8 | 44.3%95% interval 40.9%–47.5% | 36.3% | 17.6%95% interval 16.6%–18.8% | 5.8% |
| TPLlabel template | u3 | 19.3%95% interval 18.2%–20.5% | -22.5% | 27.0%95% interval 26.0%–28.0% | -10.8% |
| NUMnumeric code | u3 | 20.6%95% interval 19.3%–21.8% | -20.6% | 11.4%95% interval 10.7%–12.1% | -34.5% |
| CCAno training | u3 | 81.7%95% interval 78.5%–84.6% | 72.2% | 81.7%95% interval 78.5%–84.6% | 72.2% |
| Arm and test | A plain spectrum (L0) | Frozen CBraMod features (L1) |
|---|---|---|
| ID, 8-way (neighbour read-off) | −14.26 pp95% interval −15.50 to −13.06 pp | −27.80 pp95% interval −30.61 to −25.03 pp |
| TPL, 8-way | −24.44 pp95% interval −26.44 to −22.26 pp | −24.27 pp95% interval −26.72 to −21.85 pp |
| TPL, 3-way | −27.83 pp95% interval −30.18 to −25.37 pp | −29.67 pp95% interval −31.76 to −27.64 pp |
| DESC, 8-way | −24.79 pp95% interval −26.88 to −22.55 pp | −24.81 pp95% interval −27.31 to −22.39 pp |
| DESC, 3-way | −28.95 pp95% interval −31.33 to −26.46 pp | −29.74 pp95% interval −31.76 to −27.76 pp |
| SHUF, 8-way | −40.63 pp95% interval −44.16 to −36.98 pp | −30.09 pp95% interval −33.02 to −27.27 pp |
| SHUF, 3-way | −34.96 pp95% interval −37.50 to −32.43 pp | −32.52 pp95% interval −35.18 to −29.84 pp |
| NUM, 8-way | −12.30 pp95% interval −13.42 to −11.17 pp | −16.16 pp95% interval −17.88 to −14.47 pp |
| NUM, 3-way | −7.27 pp95% interval −8.49 to −6.02 pp | −7.63 pp95% interval −8.74 to −6.53 pp |
| B-sh, 8-way (neighbour read-off) | −12.81 pp95% interval −14.09 to −11.56 pp | −31.19 pp95% interval −34.09 to −28.32 pp |
| CCA, 8-way | 0.00 pp95% interval 0.00 to 0.00 pp | 0.00 pp95% interval 0.00 to 0.00 pp |
| CCA, 3-way | +0.41 pp95% interval +0.26 to +0.56 pp | +0.41 pp95% interval +0.26 to +0.56 pp |
| Rotation | A plain spectrum (L0) | Frozen CBraMod features (L1) |
|---|---|---|
| Rotation 0 | −49.88 pp95% interval −53.16 to −46.16 pp | −63.00 pp95% interval −67.23 to −58.54 pp |
| Rotation 1 | −52.96 pp95% interval −57.02 to −48.99 pp | −61.76 pp95% interval −65.81 to −57.74 pp |
| Rotation 2 | −48.74 pp95% interval −51.97 to −45.42 pp | −59.74 pp95% interval −63.75 to −55.51 pp |
| Rotation 3 | −57.34 pp95% interval −60.84 to −53.73 pp | −62.75 pp95% interval −66.73 to −58.42 pp |
| Rotation 4 | −52.14 pp95% interval −55.80 to −48.20 pp | −62.47 pp95% interval −66.53 to −58.23 pp |
Per-frequency and interior-only breakdowns were computed after the audits, and no independent audit covers them, so they are not published.
Before any result
A check was changed after it failed, twice. Everything about that is here, with the other checks that had to pass before any result was read.
A training-label permutation canary refits the heads on shuffled training labels and checks that test scores sit at chance: 17 configurations on the first fold, 38 intervals at 99.9%.
Revision 3 failed: 1 of 38 intervals excluded chance (sleep, frozen CBraMod features, the shuffled-template arm, unseen “light sleep” with its explicit name). Its labels had been permuted within each training person. The run stopped.
Revision 4, approved by the owner before any new fit, permuted labels across all training windows. It failed too: 3 of 38 intervals, all on BOAS, one of them below chance; every SSVEP interval included chance both times. Revision 4 said “no further revision”. The owner overrode that after seeing how it failed, and approved revision 5 before any new fit. Revision 4’s per-person centring diagnostic was not run: with labels permuted across all windows it no longer tests anything.
Revision 5 passed: 20 heads per configuration (340 fits), each metric compared with its exact per-person chance, and a two-level bootstrap over heads, then people; 0 of 38 intervals excluded zero. Both failed records are kept.
Direction counts, descriptive: of 760 single-head intervals in revision 5, 125 excluded the exact chance, 66 above and 59 below, and 18 cells failed in both directions. BOAS cells are published as pass or fail and counts only.
Why they failed, the reading revision 5 was built to test: a head trained on shuffled labels is still a random function of sleep features that are strongly organised by stage, so on new people its scores lean one way for everyone, up or down. A person bootstrap measures the spread between people, not between heads, so its 99.9% interval was too narrow for this question. The failures fit that reading: all were on BOAS, and one was below chance, which a leak cannot produce. It is not a proof that no leak exists: person-disjoint folds, no held-out label or wording in training and training-only scaling are asserted in the code, and both conformance audits check for leakage directly.
Decision S0b-6, approved by the owner: the SSVEP token rule says every token of a held-out frequency occurs in that rotation’s training texts. It stops the run for bge-small-en-v1.5 and all-MiniLM-L6-v2, and both passed in all 5 rotations. It cannot hold for multilingual-e5-small, whose tokenizer keeps one-decimal number pieces as single tokens while the held-out number itself is barred from training texts; for that encoder the rule was reported instead of stopping the run, and every Chinese-wording unseen-frequency entry says the frequency is a new token.
Engineering amendment EA-1, before the first fit: in the contiguous extrapolation band (S8) the neighbour average is not defined and the neighbour 3-way is relabelled, because its competitors are themselves held-out questions; failed fits are marked per cell, and none failed; the secondary scoring can resume after an interruption. No question, partition, wording, arm, metric, margin or gate changed.
The declared numeracy prediction was not activated. It said: if the primary text encoder orders frequencies poorly (Spearman ρ below 0.3 between frequency distance and embedding distance, for every training frame), the 8-way TPL − NUM flag will not read “TPL higher”. Its premise was false: ρ was 0.593, 0.563 and 0.526 for the three training frames of bge-small-en-v1.5, over 780 pairs. The flag reads “NUM higher” at both levels anyway; with the premise false, that is a description, not a confirmed prediction.
Four independent audits passed. They re-implement the definitions without importing the study’s scoring code: the conformance audit of the primary run passed 106 checks with 0 failed; the numerical audit of the primary made 2,270 comparisons with 0 mismatches; the conformance audit of the secondary run passed 231 checks with 0 failed; the numerical audit of the secondary made 18,619 comparisons, 29 of them outside the tolerance, every one explained.
Secondary · beyond the primary comparisons
Further datasets, features, a second text encoder, Chinese wordings and probes, none of them among the primary comparisons. One seed unless marked: a single seed carries no training-run variance. Every EESM19 and single-seed entry was declared likely inconclusive before the run, though none reached the “inconclusive” flags.
| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
|---|---|---|---|---|---|
| P1 · Label template − question numberTPL − ID | seen, 5-way | −0.04 pp95% interval −0.47 to +0.39 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 5-way | −0.03 pp95% interval −0.43 to +0.40 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −0.11 pp95% interval −0.56 to +0.20 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −15.27 pp95% interval −16.24 to −14.29 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −18.54 pp95% interval −27.30 to −10.80 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P3 · Label template head against the read-off from its own answersTPL / read-off | asleep · implicit name | +4.72095% interval +4.410 to +5.113R 112.17 (82.27–166.15) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | asleep · explicit name | +2.56695% interval +2.179 to +3.010R 13.02 (8.84–20.30) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | N1, N2 or N3 sleep · explicit name | +1.08095% interval +0.890 to +1.267R 2.95 (2.44–3.55) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | light sleep · implicit name | +2.40995% interval +2.208 to +2.609R 11.12 (9.10–13.58) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | light sleep · explicit name | +0.88895% interval +0.769 to +1.024R 2.43 (2.16–2.78) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | asleep · new descriptions10 held-out wordings, resampled | +4.70795% interval +4.279 to +5.154R 110.72 (72.15–173.16) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | N1, N2 or N3 sleep · new descriptions10 held-out wordings, resampled | +2.22395% interval +1.622 to +2.649R 9.24 (5.06–14.15) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | light sleep · new descriptions10 held-out wordings, resampled | +1.70795% interval +1.240 to +2.056R 5.51 (3.46–7.81) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P5 · Label template − shuffled templatesTPL − SHUF | unseen questions, implicit names | −0.39195% interval −0.470 to −0.311 | SHUF higher | Difference only: no margin | Headroom gate passed |
| P5 · Label template − shuffled templatesTPL − SHUF | unseen questions, explicit names | +0.66795% interval +0.634 to +0.698 | TPL higher | Difference only: no margin | Headroom gate passed |
No triviality canary is declared for motor imagery, so its gate reads “not run” and excludes nothing. Motor imagery has only two three-cycles, so two of the three shuffled-template seeds share one derangement. The unseen “imagining a hand movement” came closest to its read-off: a difference, but within the ratio margin.
| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
|---|---|---|---|---|---|
| P1 · Label template − question numberTPL − ID | seen, 3-way | −0.49 pp95% interval −1.07 to +0.02 pp | No difference shown | Equivalent within ±2 pp | No canary declared for this task: not run |
| P1 · Description − question numberDESC − ID | seen, 3-way | −0.54 pp95% interval −1.08 to +0.05 pp | No difference shown | Equivalent within ±2 pp | No canary declared for this task: not run |
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | +0.05 pp95% interval −0.10 to +0.20 pp | No difference shown | Equivalent within ±2 pp | No canary declared for this task: not run |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | +0.22 pp95% interval −0.11 to +0.55 pp | No difference shown | Equivalent within ±2 pp | No canary declared for this task: not run |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −3.49 pp95% interval −6.63 to −1.42 pp | DESC(train) higher | 2 pp margin not met | No canary declared for this task: not run |
| P3 · Label template head against the read-off from its own answersTPL / read-off | imagining a hand movement · implicit name | +0.07695% interval +0.050 to +0.107R 1.08 (1.05–1.11) | Read-off leaves less error | Equivalent within the 1.25 ratio margin | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | imagining a hand movement · new descriptions10 held-out wordings, resampled | +0.08895% interval +0.048 to +0.138R 1.09 (1.05–1.15) | Read-off leaves less error | Equivalent within the 1.25 ratio margin | Headroom gate passed |
REVE-L features keep the stimulus phase, so their neighbour 3-way tests frequency and phase together.
| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
|---|---|---|---|---|---|
| P1 · Label template − question numberTPL − ID | seen, 32-way | −2.76 pp95% interval −3.51 to −2.07 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 32-way | −2.56 pp95% interval −3.24 to −1.90 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −4.48 pp95% interval −7.66 to −2.16 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −2.16 pp95% interval −2.44 to −1.88 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −5.05 pp95% interval −6.50 to −3.77 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, 8-way | −65.32 pp95% interval −69.18 to −61.17 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL) | unseen frequency, 8-way | +5.81 pp95% interval +4.54 to +7.04 pp | TPL higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, 8-way | +1.92 pp95% interval +0.85 to +2.92 pp | TPL higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, neighbour 3-way | −38.75 pp95% interval −41.62 to −35.60 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, neighbour 3-way | +35.22 pp95% interval +32.91 to +37.42 pp | TPL higher | Difference only: no margin | Headroom gate passed |
| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
|---|---|---|---|---|---|
| P1 · Label template − question numberTPL − ID | seen, 32-way | −6.73 pp95% interval −7.75 to −5.68 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 32-way | −11.58 pp95% interval −13.16 to −9.97 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −3.95 pp95% interval −5.11 to −2.86 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −4.90 pp95% interval −5.53 to −4.25 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −5.49 pp95% interval −6.48 to −4.60 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, 8-way | −50.10 pp95% interval −53.06 to −46.73 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL) | unseen frequency, 8-way | −7.68 pp95% interval −8.82 to −6.50 pp | NN(TPL) higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, 8-way | −10.84 pp95% interval −12.53 to −9.08 pp | NUM higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, neighbour 3-way | −61.85 pp95% interval −65.13 to −58.31 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, neighbour 3-way | −0.18 pp95% interval −2.02 to +1.73 pp | No difference shown | Difference only: no margin | Headroom gate passed |
| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
|---|---|---|---|---|---|
| P1 · Label template − question numberTPL − ID | seen, 32-way | −4.73 pp95% interval −5.43 to −4.07 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 32-way | −8.84 pp95% interval −10.12 to −7.60 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −2.89 pp95% interval −3.73 to −2.08 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −3.09 pp95% interval −3.57 to −2.63 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −2.98 pp95% interval −3.67 to −2.34 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, 8-way | −61.55 pp95% interval −65.21 to −57.49 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL) | unseen frequency, 8-way | −1.09 pp95% interval −1.69 to −0.48 pp | NN(TPL) higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, 8-way | −2.03 pp95% interval −2.86 to −1.19 pp | NUM higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, neighbour 3-way | −54.63 pp95% interval −57.51 to −51.29 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, neighbour 3-way | +15.57 pp95% interval +14.27 to +16.75 pp | TPL higher | Difference only: no margin | Headroom gate passed |
For this encoder each held-out frequency is a new token, not only a new combination of seen tokens (decision S0b-6).
| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
|---|---|---|---|---|---|
| P1 · Label template − question numberTPL − ID | seen, 32-way | −3.15 pp95% interval −3.80 to −2.48 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 32-way | −4.94 pp95% interval −5.69 to −4.18 pp | ID higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −4.67 pp95% interval −6.57 to −3.24 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −8.70 pp95% interval −9.76 to −7.60 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −7.35 pp95% interval −8.58 to −6.17 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, 8-way | −54.76 pp95% interval −57.77 to −51.25 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL) | unseen frequency, 8-way | −15.81 pp95% interval −17.68 to −13.94 pp | NN(TPL) higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, 8-way | −15.51 pp95% interval −17.58 to −13.44 pp | NUM higher | Difference only: no margin | Headroom gate passed |
| P4 · Label template − CCA, no trainingTPL − CCA | unseen frequency, neighbour 3-way | −59.98 pp95% interval −63.74 to −55.86 pp | CCA higher | 2 pp margin not met | Headroom gate passed |
| P4 · Label template − numeric codeTPL − NUM | unseen frequency, neighbour 3-way | +1.69 pp95% interval −0.60 to +4.18 pp | No difference shown | Difference only: no margin | Headroom gate passed |
| Comparison | Difference, pp | Difference shown? | Margin, ±2 pp |
|---|---|---|---|
| TPL - CCA (8-way) | −65.37 pp95% interval −69.21 to −61.40 pp | CCA higher | 2 pp margin not met |
| TPL - NUM (8-way) | +3.51 pp95% interval +1.64 to +5.49 pp | TPL higher | TPL non-inferior at 2 pp |
| TPL - NN(TPL) (8-way) | Not defined (EA-1) | ||
| TPL - CCA (3-way) | −50.19 pp95% interval −54.05 to −46.25 pp | CCA higher | 2 pp margin not met |
| Comparison | Difference, pp | Difference shown? | Margin, ±2 pp |
|---|---|---|---|
| TPL - CCA (8-way) | −67.47 pp95% interval −71.22 to −63.56 pp | CCA higher | 2 pp margin not met |
| TPL - NUM (8-way) | +0.88 pp95% interval −0.70 to +2.47 pp | No difference shown | TPL non-inferior at 2 pp |
| TPL - NN(TPL) (8-way) | Not defined (EA-1) | ||
| TPL - CCA (3-way) | −48.41 pp95% interval −52.10 to −44.40 pp | CCA higher | 2 pp margin not met |
In the band the neighbour 3-way competitors are themselves held-out questions, not trained detectors, and the neighbour average is not defined (EA-1). For bge-small-en-v1.5, many of these texts carry a number token never seen in training.
The off-grid frequencies have two decimals, whose hundredths never occur in BETA’s training texts, so the text arms face new tokens; the numeric code is the clean comparison. Nine fits, three arms by three seeds, fixed before the freeze. Computed from the Tsinghua author mirror; no physical unit is claimed.
| Comparison | Difference, pp | Difference shown? | Margin, ±2 pp |
|---|---|---|---|
| Dry electrodes | |||
| TPL - CCA (12-way) | −48.18 pp95% interval −52.36 to −43.90 pp | CCA higher | 2 pp margin not met |
| NUM - CCA (12-way) | −42.16 pp95% interval −46.05 to −38.38 pp | CCA higher | 2 pp margin not met |
| TPL - NUM (12-way) | −6.02 pp95% interval −7.32 to −4.73 pp | NUM higher | 2 pp margin not met |
| DESC - CCA (12-way) | −48.52 pp95% interval −52.99 to −44.02 pp | CCA higher | 2 pp margin not met |
| Wet electrodes | |||
| TPL - CCA (12-way) | −62.92 pp95% interval −66.16 to −59.66 pp | CCA higher | 2 pp margin not met |
| NUM - CCA (12-way) | −56.11 pp95% interval −59.16 to −52.83 pp | CCA higher | 2 pp margin not met |
| TPL - NUM (12-way) | −6.82 pp95% interval −8.31 to −5.33 pp | NUM higher | 2 pp margin not met |
| DESC - CCA (12-way) | −63.56 pp95% interval −66.68 to −60.31 pp | CCA higher | 2 pp margin not met |
| Comparison | Difference, pp | Difference shown? | Margin, ±2 pp |
|---|---|---|---|
| S10 DESC - CCA (8-way) | −52.29 pp95% interval −55.41 to −49.00 pp | CCA higher | 2 pp margin not met |
| S10 DESC - CCA (3-way) | −62.72 pp95% interval −66.34 to −58.98 pp | CCA higher | 2 pp margin not met |
| S10 DESC - NN(DESC) (8-way) | −10.80 pp95% interval −12.15 to −9.39 pp | NN(DESC) higher | 2 pp margin not met |
| S10 NUM - CCA (8-way) | −39.44 pp95% interval −42.25 to −36.48 pp | CCA higher | 2 pp margin not met |
| S10 TPL - NN(ID) (8-way) | −15.63 pp95% interval −17.48 to −13.62 pp | NN(ID) higher | 2 pp margin not met |
| S10 TPL - SHUF (8-way) | +19.50 pp95% interval +17.45 to +21.62 pp | TPL higher | TPL non-inferior at 2 pp |
| S11 TPL - FBCCA (8-way) | −61.68 pp95% interval −63.72 to −59.69 pp | FBCCA higher | 2 pp margin not met |
| S11 TPL - FBCCA (3-way) | −73.35 pp95% interval −75.61 to −70.93 pp | FBCCA higher | 2 pp margin not met |
| S13 TPL held-out frames - CCA (8-way) | −53.71 pp95% interval −56.73 to −50.27 pp | CCA higher | 2 pp margin not met |
| S13 DESC held-out descriptions - CCA (8-way) | −54.01 pp95% interval −57.07 to −50.83 pp | CCA higher | 2 pp margin not met |
| Arm | Seen frequencies | Unseen frequencies |
|---|---|---|
| S9 TPL | 23.4%95% interval 21.0%–25.7% | 4.4%95% interval 4.0%–4.8% |
| S9 DESC | 24.0%95% interval 21.5%–26.5% | 4.2%95% interval 3.8%–4.7% |
| S9 SHUF | 23.4%95% interval 20.9%–25.7% | 1.1%95% interval 1.0%–1.4% |
| S9 NUM | 20.7%95% interval 18.5%–22.9% | 11.5%95% interval 10.2%–12.7% |
| One-output head, neighbour read-off (nn8) | 44.3%95% interval 40.9%–47.5% | |
| FBCCA, seen 32-way | 81.1%95% interval 77.0%–84.8% | |
| FBCCA, unseen 8-way | 90.4%95% interval 87.7%–92.8% | |
| FBCCA, neighbour 3-way | 92.7%95% interval 90.9%–94.4% | |
| Comparison | Difference, pp | Difference shown? | Margin, ±2 pp |
|---|---|---|---|
| S10 DESC - CCA (8-way) | −62.26 pp95% interval −66.01 to −58.41 pp | CCA higher | 2 pp margin not met |
| S10 DESC - CCA (3-way) | −54.93 pp95% interval −58.01 to −51.72 pp | CCA higher | 2 pp margin not met |
| S10 DESC - NN(DESC) (8-way) | −1.27 pp95% interval −2.08 to −0.46 pp | NN(DESC) higher | 2 pp margin not met |
| S10 NUM - CCA (8-way) | −59.48 pp95% interval −62.92 to −55.87 pp | CCA higher | 2 pp margin not met |
| S10 TPL - NN(ID) (8-way) | +0.68 pp95% interval −0.25 to +1.61 pp | No difference shown | Equivalent within ±2 pp |
| S10 TPL - SHUF (8-way) | +7.97 pp95% interval +6.88 to +9.09 pp | TPL higher | TPL non-inferior at 2 pp |
| S11 TPL - FBCCA (8-way) | −71.41 pp95% interval −73.53 to −69.17 pp | FBCCA higher | 2 pp margin not met |
| S11 TPL - FBCCA (3-way) | −65.66 pp95% interval −67.61 to −63.62 pp | FBCCA higher | 2 pp margin not met |
| S13 TPL held-out frames - CCA (8-way) | −62.31 pp95% interval −66.06 to −58.27 pp | CCA higher | 2 pp margin not met |
| S13 DESC held-out descriptions - CCA (8-way) | −62.58 pp95% interval −66.15 to −58.59 pp | CCA higher | 2 pp margin not met |
| Arm | Seen frequencies | Unseen frequencies |
|---|---|---|
| S9 TPL | 18.8%95% interval 16.7%–21.1% | 2.5%95% interval 2.2%–2.8% |
| S9 DESC | 18.9%95% interval 16.9%–21.2% | 2.5%95% interval 2.3%–2.8% |
| S9 SHUF | 18.3%95% interval 16.4%–20.4% | 1.6%95% interval 1.3%–1.9% |
| S9 NUM | 11.5%95% interval 10.4%–12.7% | 4.3%95% interval 3.8%–4.8% |
| One-output head, neighbour read-off (nn8) | 17.6%95% interval 16.6%–18.8% | |
| FBCCA, seen 32-way | 81.1%95% interval 77.0%–84.8% | |
| FBCCA, unseen 8-way | 90.4%95% interval 87.7%–92.8% | |
| FBCCA, neighbour 3-way | 92.7%95% interval 90.9%–94.4% | |
S9 asks about seen and unseen frequencies in one 40-way test; FBCCA, like CCA, needs no training. One seed for FBCCA.
BOAS is published with three stated gaps (owner decision, 7 October 2026):
Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.
Credit: Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.
| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
|---|---|---|---|---|---|
| P1 · Label template − question numberTPL − ID | seen, 5-way | +0.27 pp95% interval −0.36 to +0.86 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 5-way | +0.12 pp95% interval −0.48 to +0.71 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −0.39 pp95% interval −1.04 to +0.05 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −13.54 pp95% interval −14.57 to −12.49 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −21.86 pp95% interval −30.62 to −14.12 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P3 · Label template head against the read-off from its own answersTPL / read-off | asleep · implicit name | +4.14995% interval +3.753 to +4.614R 63.39 (42.66–100.85) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | asleep · explicit name | +2.64795% interval +2.274 to +3.070R 14.11 (9.72–21.55) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | N1, N2 or N3 sleep · explicit name | +1.31795% interval +1.163 to +1.483R 3.73 (3.20–4.41) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | light sleep · implicit name | +2.73095% interval +2.540 to +2.929R 15.33 (12.67–18.71) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | light sleep · explicit name | +1.87295% interval +1.691 to +2.057R 6.50 (5.43–7.82) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | asleep · new descriptions10 held-out wordings, resampled | +4.09295% interval +3.565 to +4.610R 59.84 (35.35–100.44) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | N1, N2 or N3 sleep · new descriptions10 held-out wordings, resampled | +1.82995% interval +1.250 to +2.292R 6.23 (3.49–9.90) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | light sleep · new descriptions10 held-out wordings, resampled | +1.77495% interval +1.162 to +2.236R 5.89 (3.20–9.35) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P5 · Label template − shuffled templatesTPL − SHUF | unseen questions, implicit names | −0.58595% interval −0.612 to −0.555 | SHUF higher | Difference only: no margin | Headroom gate passed |
| P5 · Label template − shuffled templatesTPL − SHUF | unseen questions, explicit names | +0.62695% interval +0.602 to +0.648 | TPL higher | Difference only: no margin | Headroom gate passed |
| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
|---|---|---|---|---|---|
| P1 · Label template − question numberTPL − ID | seen, 5-way | +0.19 pp95% interval −0.37 to +0.75 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 5-way | +0.01 pp95% interval −0.64 to +0.72 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −1.72 pp95% interval −4.31 to −0.09 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −14.04 pp95% interval −15.29 to −12.68 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −25.74 pp95% interval −34.29 to −16.26 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P3 · Label template head against the read-off from its own answersTPL / read-off | asleep · implicit name | +3.79895% interval +3.385 to +4.273R 44.60 (29.53–71.76) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | asleep · explicit name | +1.32595% interval +0.999 to +1.715R 3.76 (2.71–5.56) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | N1, N2 or N3 sleep · explicit name | +0.88295% interval +0.717 to +1.047R 2.41 (2.05–2.85) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | light sleep · implicit name | +2.33895% interval +2.134 to +2.552R 10.36 (8.45–12.83) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | light sleep · explicit name | +1.79695% interval +1.598 to +2.000R 6.03 (4.94–7.39) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | asleep · new descriptions10 held-out wordings, resampled | +3.06995% interval +2.508 to +3.601R 21.52 (12.28–36.65) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | N1, N2 or N3 sleep · new descriptions10 held-out wordings, resampled | +1.11895% interval +0.793 to +1.449R 3.06 (2.21–4.26) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | light sleep · new descriptions10 held-out wordings, resampled | +1.07095% interval +0.840 to +1.333R 2.92 (2.32–3.79) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P5 · Label template − shuffled templatesTPL − SHUF | unseen questions, implicit names | −0.62695% interval −0.653 to −0.596 | SHUF higher | Difference only: no margin | Headroom gate passed |
| P5 · Label template − shuffled templatesTPL − SHUF | unseen questions, explicit names | +0.57295% interval +0.550 to +0.592 | TPL higher | Difference only: no margin | Headroom gate passed |
| Comparison | Question | Difference: pp, log R or AUROC | Difference shown? | Margin | Gate |
|---|---|---|---|---|---|
| P1 · Label template − question numberTPL − ID | seen, 5-way | +0.34 pp95% interval −0.29 to +0.99 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P1 · Description − question numberDESC − ID | seen, 5-way | −0.28 pp95% interval −0.93 to +0.42 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · New sentence frame − training wordingTPL(a1) − TPL(train) | seen, reworded12 held-out wordings, resampled | −0.70 pp95% interval −1.83 to +0.08 pp | No difference shown | Equivalent within ±2 pp | Passes the triviality canary |
| P2 · Synonym − training wordingTPL(a2) − TPL(train) | seen, rewordedfor these three synonyms: fixed, not resampled | −12.51 pp95% interval −13.73 to −11.27 pp | TPL(train) higher | 2 pp margin not met | Passes the triviality canary |
| P2 · New description − training descriptionsDESC(a3) − DESC(train) | seen, reworded10 held-out wordings, resampled | −25.81 pp95% interval −33.97 to −17.73 pp | DESC(train) higher | 2 pp margin not met | Passes the triviality canary |
| P3 · Label template head against the read-off from its own answersTPL / read-off | asleep · implicit name | +3.82695% interval +3.424 to +4.277R 45.87 (30.69–72.04) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | asleep · explicit name | +1.55895% interval +1.216 to +1.950R 4.75 (3.37–7.03) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | N1, N2 or N3 sleep · explicit name | +0.73795% interval +0.584 to +0.898R 2.09 (1.79–2.45) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | light sleep · implicit name | +2.32295% interval +2.122 to +2.532R 10.19 (8.34–12.57) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Label template head against the read-off from its own answersTPL / read-off | light sleep · explicit name | +1.61595% interval +1.424 to +1.809R 5.03 (4.15–6.11) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | asleep · new descriptions10 held-out wordings, resampled | +2.71595% interval +2.127 to +3.244R 15.11 (8.39–25.63) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | N1, N2 or N3 sleep · new descriptions10 held-out wordings, resampled | +0.79895% interval +0.599 to +1.039R 2.22 (1.82–2.83) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| P3 · Description head against the read-off from its own answersDESC / read-off | light sleep · new descriptions10 held-out wordings, resampled | +0.94395% interval +0.773 to +1.123R 2.57 (2.17–3.07) | Read-off leaves less error | 1.25 ratio margin not met | Headroom gate passed |
| Comparison | log R, with R | Difference shown? | Margin, R < 1.25 |
|---|---|---|---|
| TPL asleep implicit (held-out frames) | +3.56595% interval +3.162 to +4.006 | Read-off leaves less error | 1.25 ratio margin not met |
| TPL asleep explicit (held-out frames) | +1.97695% interval +1.594 to +2.392 | Read-off leaves less error | 1.25 ratio margin not met |
| TPL NREM explicit (held-out frames) | +0.92095% interval +0.743 to +1.100 | Read-off leaves less error | 1.25 ratio margin not met |
| TPL light implicit (held-out frames) | +2.23395% interval +2.022 to +2.443 | Read-off leaves less error | 1.25 ratio margin not met |
| TPL light explicit (held-out frames) | +1.48495% interval +1.240 to +1.719 | Read-off leaves less error | 1.25 ratio margin not met |
| Wording | Against the truth of the negated question | AUROC of 1 − p(X) | People |
|---|---|---|---|
| not W | 0.02295% interval 0.014–0.030 | 0.98095% interval 0.972–0.987 | 100 |
| not N1 | 0.13595% interval 0.120–0.152 | 0.87295% interval 0.855–0.887 | 100 |
| not N2 | 0.12495% interval 0.110–0.139 | 0.92495% interval 0.909–0.937 | 100 |
| not N3 | 0.02195% interval 0.017–0.026 | 0.97995% interval 0.974–0.983 | 71 |
| not REM | 0.04995% interval 0.038–0.061 | 0.95695% interval 0.944–0.967 | 100 |
| non-REM sleep (implicit) | 0.04895% interval 0.037–0.060 | 0.95695% interval 0.944–0.967 | 100 |
Every negated wording was answered as if it asked for X itself: against the truth of the negated question its AUROC lay far below 0.5, while turning the head’s answer to X around would have answered it well. Writing more prompts does not supply signal-level ground truth.
Published with the results
All 787 texts the questions were asked with, in English and Chinese, for sleep, SSVEP and motor imagery: training and held-out sentence frames, label names and synonyms, descriptions, the names and descriptions of the unseen questions, and the negation frames. Two agents wrote them from a fixed brief before any score, the second writing every held-out description without seeing the first’s; they are not users’ wordings and were not tuned. The file also holds the leakage rules they passed, the held-out frequencies of every rotation and the extrapolation band, the declared derangements of the shuffled control, the numeric code of a frequency and the pinned text encoders. Text only, CC BY 4.0.
The sleep stages’ names in the training frames, and the three synonyms each was reworded with:
| Stage | Training name | Synonyms (a2) |
|---|---|---|
| W | wake | wakefulness, awake, stage W |
| N1 | N1 sleep | stage 1 sleep, sleep stage one, stage N1 |
| N2 | N2 sleep | stage 2 sleep, sleep stage two, stage N2 |
| N3 | N3 sleep | slow-wave sleep, deep sleep, stage N3 |
| REM | REM sleep | stage R sleep, paradoxical sleep, dream sleep |
Methods & limits
5c38ec71110a24614241fFrozen at pinned commits and never redistributed.
Bingchuan Liu et al. · BETA: A Large Benchmark Database Toward SSVEP-BCI Application (2020), doi:10.3389/fnins.2020.00627. Figshare 12264401 v3; mirror Bingchuan/BETA.
Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.
Published with three stated gaps, printed beside its results above.
Kaare B. Mikkelsen et al. · Accurate whole-night sleep monitoring with dry-contact ear-EEG (2019), doi:10.1038/s41598-019-53115-3; OpenNeuro ds005185 v1.0.2. Processed mirror: Zachary1150/EESM19-Processed.
Min-Ho Lee, O-Yeon Kwon, Yong-Jeong Kim, Hong-Kyung Kim, Young-Eun Lee, John Williamson, Siamac Fazli and Seong-Whan Lee · EEG dataset and OpenBMI toolbox for three BCI paradigms: an investigation into BCI illiteracy, GigaScience (2019), giz002, doi:10.1093/gigascience/giz002. Data: Supporting data, GigaScience Database, doi:10.5524/100542.
Zhu, F., Jiang, L., Dong, G., Gao, X., & Wang, Y. (2021). An Open Dataset for Wearable SSVEP-Based Brain-Computer Interfaces (Version 4) [Data set]. Figshare. https://doi.org/10.6084/m9.figshare.13560281.v4
Zhu, F., Jiang, L., Dong, G., Gao, X., & Wang, Y. (2021). An Open Dataset for Wearable SSVEP-Based Brain-Computer Interfaces. Sensors, 21(4), 1256. https://doi.org/10.3390/s21041256
Computed from the official Tsinghua BCI Lab author mirror snapshot acquired 2026-09-20; the mirror itself has no separate DOI or version number.
Features: frozen CBraMod (primary) and frozen REVE Large @ 317531c7, REVE-L (secondary).
Weights terms: REVE under the REVE Responsible Use License v1.0 (model versions: REVE Large @ 317531c7). No model’s authors endorse these results. REVE, arXiv:2510.21585 ↗
Data source: questions-in-language-update.json · schema bci-report-questions-in-language-update-v1.
Wordings: questions-in-language-wordings.json · schema bci-report-questions-in-language-wordings-v1.
BCI Report (2026). Can an EEG model answer questions asked in words? https://bci.report/topics/questions-in-language/
Figures from release questions-in-language-update-20261008 (2026-10-08). Cite the upstream datasets as well: their credits are on this page.
Every release is archived on Zenodo: doi:10.5281/zenodo.23123296. BibTeX for the site and its releases → · CITATION.cff ↗