Questions in language · route 3 of the decision-research plan

Can an EEG model answer questions asked in words?

One small head over frozen EEG features, asked each question by a question number, a label template or a natural-language description: SSVEP on BETA, sleep on BOAS. Seen questions, rewordings and questions the EEG head was never trained on are reported apart, never as one zero-shot number.

Measured on: BETA; BOAS · Bitbrain Open Access Sleep dataset · Methods: CBraMod

Short answer

Only for questions it was trained on, and not on every dataset. A small head over frozen EEG features, asked a seen question with a label template or a description instead of a question number, was equivalent within the 2 pp margin on sleep (BOAS), but lost accuracy on SSVEP (BETA, 70 people): −5.96 pp with templates on a plain spectrum (−6.84 to −5.08 pp), the margin not met. Rewording a seen question cost accuracy too, except a new sentence frame on sleep. Questions the EEG head was never trained on were not answered usefully: every unseen sleep question did worse than adding up the head’s own stage answers, and on unseen flicker frequencies the head reached 28.7% on a plain spectrum where CCA, which needs no training, reached 80.9% (chance 12.5%). Explicit wordings such as “N1 or N2 sleep” used the link between words and stages; the everyday names “asleep” and “light sleep” ranked windows backwards, and in a boundary probe negated questions were answered as if they asked for what they negate.

How the comparison works

One head, three ways to ask

The same small head answers every question; only the vector that tells it which question changes. Every test is on people the head never trained on.

A training frame for sleep, filled with a stage’s name: The scorer marked this 30-second epoch as N2 sleep.

SSVEP on BETA: a plain spectrum (L0) and frozen CBraMod features (L1). Sleep on BOAS: frozen CBraMod features (L1). Three training seeds each. The primary text encoder is bge-small-en-v1.5, frozen. The head is the same for every arm: standardised features, a small hidden layer adjusted by the question vector (FiLM), one output.

Three kinds of result, never pooled into one zero-shot number: a seen question in the wording the head trained with; the same question reworded; and a question the head never trained on. “Unseen” always means unseen by the EEG head, not by the text encoder: the frozen encoder has read the words, and for SSVEP every token of a held-out frequency also occurs in that rotation’s training texts. Unseen sleep questions are combinations of seen stages; unseen SSVEP questions are flicker frequencies held out in five rotations.

Each comparison is the language arm minus the reference, in percentage points (pp), with a 95% interval from resampling people, then training seeds, then held-out wordings. It gets two readings, both fixed before any fit. A difference is shown when the interval excludes zero. Against a margin of ±2 pp, it is equivalent when the whole interval lies inside, non-inferior when the bound against the language arm stays within the margin, and otherwise the margin is not met, which is not the same as a loss of 2 pp or more. For an unseen sleep question the comparison is R, the head’s remaining error (1 − AUROC) over that of the read-off from its own stage answers, with a margin of R < 1.25. The check of whether the head uses the link between wording and stage, the template arm against the shuffled templates, is a difference of AUROC. Among the primary comparisons, those with the numeric code, the neighbour average and the shuffled templates are read for a difference only; the secondary tables print both readings as the run computed them.

There are 35 pre-declared primary comparisons and no multiplicity correction, so at 95% about one in twenty comparisons with no true difference may show one by chance. Accuracy is averaged per person, then over people: it compares set-ups and is not a deployment error rate.

This is route 3 of the research plan on When not to act. “Jev-style” here means the interface of a decision model like TypeSafe’s Jev, applied to EEG: encode the recording once, then answer several explicit, typed questions about it, each with a probability. The plan started from a vision paper, Yu & Yao, 2026 · Visual Jev: Accurate and Efficient Decisions from Shared Visual Context ↗, which applies that idea to images. All three routes: Jev-style questions on EEG: the evidence →

Primary · seen questions · BETA and BOAS

Seen questions: asked in words, and reworded

Questions the head was trained on. On BETA, 70 people look at flickering targets; in each rotation the head trains on most of the frequencies and is asked about those (the 32-way task). On BOAS, 100 people, each half-minute epoch is scored into one of five stages at their natural mix, against the human consensus.

BOAS is published with three stated gaps (owner decision, 7 October 2026):

  1. The consent statement does not say whether participants agreed to public sharing or secondary use.
  2. The ethics and consent statements come from the publisher's dataset description and README. No peer-reviewed paper describes BOAS.
  3. The ethics reference was added to the release in version 1.1.1 (May 2025), and the release does not say when it was granted relative to the recordings.

Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.

Credit: Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

In the training wording (P1)

−5.96 ppTPL − ID on SSVEP, a plain spectrum, BETA, three seeds

Sleep: no measurable cost. SSVEP: a cost.

On sleep, a template or a description instead of a question number was equivalent within ±2 pp: −0.05 pp (−0.70 to +0.61 pp) and +0.02 pp (−0.59 to +0.62 pp). On SSVEP it cost accuracy at both levels, and the margin was not met: templates −5.96 pp on a plain spectrum and −3.49 pp on frozen CBraMod features; descriptions −5.46 pp and −3.37 pp.

Paired differences in balanced accuracy, language arm minus question number, in pp, with 95% intervals; BETA 70 people, BOAS 100 people, three seeds each.
ComparisonQuestionDifference, ppDifference shown?Margin, ±2 ppGate
SSVEP · BETA · a plain spectrum (L0)
P1 · Label template − question numberTPL − IDseen, 32-way−5.96 pp95% interval −6.84 to −5.08 ppID higher2 pp margin not metPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 32-way−5.46 pp95% interval −6.26 to −4.66 ppID higher2 pp margin not metPasses the triviality canary
SSVEP · BETA · frozen CBraMod features (L1)
P1 · Label template − question numberTPL − IDseen, 32-way−3.49 pp95% interval −4.10 to −2.93 ppID higher2 pp margin not metPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 32-way−3.37 pp95% interval −3.98 to −2.80 ppID higher2 pp margin not metPasses the triviality canary
Sleep · BOAS · frozen CBraMod features (L1)
P1 · Label template − question numberTPL − IDseen, 5-way−0.05 pp95% interval −0.70 to +0.61 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 5-way+0.02 pp95% interval −0.59 to +0.62 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary

The seen SSVEP task is hard for every head on these features: with question numbers the head reached 32.1% (28.8%–35.2%) on a plain spectrum and 24.2% on frozen CBraMod features, against a chance of 3.1%; CCA, which needs no training, reached 65.5%. On sleep every arm reached about the same balanced accuracy, 71.5% with question numbers against a chance of 20.0%. The levels below are descriptive: overlapping intervals are not ranked, and the paired differences above are the comparison.

BETA, seen 32-way balanced accuracy (chance 3.1%): the mean over 70 people with its 95% interval, and the chance-corrected value, (level − chance) / (1 − chance). Three seeds; CCA needs no training, so it has none.
ArmA plain spectrum (L0)chance-correctedFrozen CBraMod features (L1)chance-corrected
IDquestion number32.1%95% interval 28.8%–35.2%29.9%24.2%95% interval 21.6%–26.9%21.8%
TPLlabel template26.1%95% interval 23.5%–28.6%23.7%20.7%95% interval 18.5%–23.1%18.2%
DESCdescription26.6%95% interval 23.9%–29.2%24.2%20.9%95% interval 18.7%–23.2%18.3%
SHUFshuffled templates25.1%95% interval 22.5%–27.6%22.7%20.2%95% interval 18.1%–22.3%17.6%
NUMnumeric code23.7%95% interval 21.2%–26.1%21.2%13.9%95% interval 12.5%–15.2%11.1%
B-shone output per question32.2%95% interval 28.9%–35.3%30.0%27.0%95% interval 24.2%–30.1%24.7%
CCAno training65.5%95% interval 59.7%–71.1%64.3%65.5%95% interval 59.7%–71.1%64.3%
BOAS, seen five-stage balanced accuracy (chance 20.0%), 100 people, three seeds, with the chance-corrected value and the log loss (mean binary cross-entropy, in nats). The shuffled-template arm’s seen score is read through its own derangement, and it has no log loss.
ArmBalanced accuracychance-correctedLog loss
IDquestion number71.5%95% interval 69.1%–73.5%64.3%0.26795% interval 0.242–0.296
TPLlabel template71.4%95% interval 69.1%–73.4%64.3%0.27695% interval 0.249–0.307
DESCdescription71.5%95% interval 69.1%–73.5%64.3%0.27895% interval 0.251–0.310
SHUFshuffled templates71.5%95% interval 69.2%–73.5%64.4%—
B-shone output per question71.4%95% interval 69.1%–73.4%64.3%0.25995% interval 0.236–0.287
BOAS, AUROC of each seen stage, with 95% intervals; three seeds. N3 pools the 71 people with at least five N3 epochs; every other stage pools 100.
ArmW100 peopleN1100 peopleN2100 peopleN371 peopleREM100 people
IDquestion number0.9820.976–0.9870.8760.859–0.8900.9220.904–0.9370.9780.971–0.9830.9570.945–0.967
TPLlabel template0.9800.972–0.9870.8720.855–0.8870.9240.909–0.9370.9790.974–0.9830.9560.944–0.967
DESCdescription0.9810.974–0.9870.8690.851–0.8840.9230.907–0.9370.9790.974–0.9830.9560.944–0.967

Reworded (P2)

The same seen questions in wordings the head never trained with: a new sentence frame (a1, held-out frames resampled), a synonym of the label (a2: three per question, fixed, so these entries hold only for these three synonyms), or a new description (a3, held-out descriptions resampled). Reworded rows are never pooled with the unseen ones.

Paired differences in balanced accuracy, reworded minus the training wording, in pp, with 95% intervals; BETA 70 people, BOAS 100 people, three seeds each.
ComparisonQuestionDifference, ppDifference shown?Margin, ±2 ppGate
SSVEP · BETA · a plain spectrum (L0)
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−4.01 pp95% interval −6.53 to −2.12 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−1.84 pp95% interval −2.10 to −1.57 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−4.13 pp95% interval −5.58 to −3.11 ppDESC(train) higher2 pp margin not metPasses the triviality canary
SSVEP · BETA · frozen CBraMod features (L1)
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−2.72 pp95% interval −4.43 to −1.43 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−1.33 pp95% interval −1.53 to −1.14 ppTPL(train) higherEquivalent within ±2 ppPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−2.95 pp95% interval −3.81 to −2.26 ppDESC(train) higher2 pp margin not metPasses the triviality canary
Sleep · BOAS · frozen CBraMod features (L1)
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−0.46 pp95% interval −1.33 to +0.16 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−12.60 pp95% interval −13.74 to −11.43 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−20.87 pp95% interval −29.49 to −13.71 ppDESC(train) higher2 pp margin not metPasses the triviality canary

On sleep a new sentence frame cost nothing measurable: equivalent within ±2 pp (−0.46 pp, −1.33 to +0.16 pp). The three synonyms per stage scored −12.60 pp (−13.74 to −11.43 pp) and new descriptions −20.87 pp (−29.49 to −13.71 pp), a wide interval, because descriptions differ a lot from one another. On BETA every rewording scored lower, from −1.33 pp to −4.13 pp; only the synonyms on frozen CBraMod features stayed within the margin.

Each synonym on its own against the training wording, in pp with 95% intervals (descriptive); BETA 70 people, BOAS 100 people. On SSVEP the third synonym spells the number out.
SynonymSSVEP · BETA · a plain spectrum (L0)SSVEP · BETA · frozen CBraMod features (L1)Sleep · BOAS · frozen CBraMod features (L1)
Synonym 1−1.45 pp95% interval −1.67 to −1.22 pp−1.09 pp95% interval −1.27 to −0.90 pp−23.53 pp95% interval −25.29 to −21.64 pp
Synonym 2−0.87 pp95% interval −1.04 to −0.70 pp−0.84 pp95% interval −1.00 to −0.68 pp−10.85 pp95% interval −12.59 to −9.17 pp
Synonym 3−3.19 pp95% interval −3.69 to −2.69 pp−2.06 pp95% interval −2.44 to −1.73 pp−3.42 pp95% interval −4.49 to −2.41 pp
Answer flip rate (descriptive): how often the head’s answer to the same window changes when only the wording changes; BETA 70 people, BOAS 100 people. On SSVEP even another training seed of the question-number head changes most answers, because the seen task is near its floor, so flip rates there mostly show uncertainty.
New frameSynonymBetween training framesNew descriptionBetween training descriptionsBetween seeds of ID
BETA · L039.7%29.3%19.4%41.2%19.3%56.8%
BETA · L140.7%32.0%24.3%42.7%23.3%62.2%
BOAS · L16.4%20.0%3.3%45.4%3.1%7.4%

Primary · unseen sleep questions · BOAS

Sleep questions the EEG head was never trained on

Three questions the head never trained on, each a combination of the stages it did: asleep (N1, N2, N3 or REM), N1, N2 or N3 sleep, and light sleep (N1 or N2). Each was asked with templates, using an implicit name or an explicit one, and with new descriptions; 100 people, three seeds. Unseen by the EEG head, not by the text encoder: the encoder has read these words. Because each question is a combination, the head’s own stage answers can be added up into an answer, the read-off, which is what each wording is held to.

BOAS is published with three stated gaps (owner decision, 7 October 2026):

  1. The consent statement does not say whether participants agreed to public sharing or secondary use.
  2. The ethics and consent statements come from the publisher's dataset description and README. No peer-reviewed paper describes BOAS.
  3. The ethics reference was added to the release in version 1.1.1 (May 2025), and the release does not say when it was granted relative to the recordings.

Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.

Credit: Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

The names the template arm was asked with, as written before any score.
QuestionImplicit nameExplicit name
AsleepasleepN1, N2, N3 or REM sleep
N1, N2 or N3 sleepnoneN1, N2 or N3 sleep
Light sleeplight sleepN1 or N2 sleep

Can the head answer them? (P3)

33.80R for “asleep” with its implicit name: the head’s remaining error over the read-off’s

Not answered: the read-off did better in every wording

In all eight conditions the interval lay above R = 1, in the read-off’s favour, and none came within the 1.25 margin: R ran from 2.42 (2.07–2.85; N1, N2 or N3 sleep, explicit name) to 33.80 (23.09–52.82; asleep, implicit name).

log R and R, the head’s remaining error (1 − AUROC) over the read-off’s, with 95% intervals; 100 people, three seeds. A positive log R, or R above 1, is against the language arm.
ComparisonQuestionlog R, with RDifference shown?Margin, R < 1.25Gate
P3 · Label template head against the read-off from its own answersTPL / read-offasleep · implicit name+3.52095% interval +3.139 to +3.967R 33.80 (23.09–52.82)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offasleep · explicit name+1.95195% interval +1.623 to +2.357R 7.04 (5.07–10.56)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offN1, N2 or N3 sleep · explicit name+0.88395% interval +0.726 to +1.047R 2.42 (2.07–2.85)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offlight sleep · implicit name+2.26795% interval +2.072 to +2.466R 9.65 (7.94–11.77)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offlight sleep · explicit name+1.51995% interval +1.343 to +1.698R 4.57 (3.83–5.46)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offasleep · new descriptions10 held-out wordings, resampled+3.44495% interval +2.913 to +3.956R 31.30 (18.41–52.26)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offN1, N2 or N3 sleep · new descriptions10 held-out wordings, resampled+1.29995% interval +0.793 to +1.744R 3.67 (2.21–5.72)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offlight sleep · new descriptions10 held-out wordings, resampled+1.41795% interval +0.903 to +1.856R 4.12 (2.47–6.40)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
AUROC of the template head on each unseen question, and of the read-off from its own stage answers (prior-corrected), with 95% intervals; 100 people, three seeds. Descriptive. An AUROC below 0.5 ranks windows in reverse.
QuestionWordingTemplate headRead-off from its stage answers
Asleepimplicit name0.31795% interval 0.294–0.3410.98095% interval 0.971–0.987
Asleepexplicit name0.85895% interval 0.839–0.8750.98095% interval 0.971–0.987
N1, N2 or N3 sleepexplicit name0.87195% interval 0.854–0.8870.94795% interval 0.934–0.958
Light sleepimplicit name0.24495% interval 0.222–0.2670.92295% interval 0.907–0.935
Light sleepexplicit name0.64295% interval 0.616–0.6670.92295% interval 0.907–0.935

The implicit names ranked windows in reverse: AUROC 0.317 for “asleep” and 0.244 for “light sleep”, below 0.5, so windows that really were asleep, or in light sleep, scored lower on these questions than the others. The explicit names did better (0.858 and 0.642) but stayed short of the read-off (0.980 and 0.922). Descriptions have no alignment control: the shuffled control is built from template vectors. Held-out sentence frames (S13, in the secondary panel) give the same result for all five template cells.

Is the link between wording and stage used? (P5)

The shuffled-template arm attaches the same template vectors to the wrong stages during training; the unseen questions keep their true vectors. If the head uses the link between a wording and its stage, the shuffled arm should do worse on them than the template arm.

Template arm minus shuffled templates: AUROC on the unseen questions, averaged over the cells of each kind of name, with 95% intervals; 100 people, three seeds. Difference only: no margin was set.
ComparisonQuestionDifference, AUROCDifference shown?MarginGate
P5 · Label template − shuffled templatesTPL − SHUFunseen questions, implicit names−0.37895% interval −0.501 to −0.253SHUF higherDifference only: no marginHeadroom gate passed
P5 · Label template − shuffled templatesTPL − SHUFunseen questions, explicit names+0.62995% interval +0.603 to +0.654TPL higherDifference only: no marginHeadroom gate passed
AUROC on each unseen question, template arm and shuffled templates, with 95% intervals (descriptive); 100 people, three seeds.
QuestionWordingTemplate armShuffled templates
Asleepimplicit name0.31795% interval 0.294–0.3410.67695% interval 0.529–0.812
Asleepexplicit name0.85895% interval 0.839–0.8750.13795% interval 0.109–0.168
N1, N2 or N3 sleepexplicit name0.87195% interval 0.854–0.8870.17195% interval 0.145–0.201
Light sleepimplicit name0.24495% interval 0.222–0.2670.64195% interval 0.537–0.738
Light sleepexplicit name0.64295% interval 0.616–0.6670.17695% interval 0.155–0.199

For explicit names, yes: attaching the templates to the wrong stages lowered the unseen AUROC, TPL − SHUF +0.629 (+0.603 to +0.654). For the implicit names it was not shown: the shuffled templates scored higher, −0.378 (−0.501 to −0.253).

Primary · unseen flicker frequencies · BETA

Flicker frequencies the EEG head was never trained on (P4)

In each of five rotations, eight of BETA’s frequencies are held out: the head trains on the rest and is asked about the eight by stating the frequency in a template. 70 people, three seeds, at both levels. Unseen by the EEG head, not by the text encoder: every token of a held-out frequency also occurs in that rotation’s training texts, so only the combination is new. CCA answers the same question with no training at all.

−52.21 ppTPL − CCA, 8-way among unseen frequencies, a plain spectrum, BETA

Far below CCA, which needs no training

In the 8-way among unseen frequencies the head reached 28.7% (26.9%–30.5%) on a plain spectrum and 19.0% on frozen CBraMod features; CCA reached 80.9%, and chance is 12.5%. The head also did worse than averaging its own answers for the two neighbouring seen frequencies, and worse than a numeric code of the frequency, at both levels.

Paired differences in accuracy, template head minus the reference, in pp, with 95% intervals; 70 people, three seeds. Comparisons with CCA carry the margin; with the neighbour average and the numeric code they are read for a difference only.
ComparisonQuestionDifference, ppDifference shown?Margin, ±2 ppGate
SSVEP · BETA · a plain spectrum (L0)
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, 8-way−52.21 pp95% interval −55.13 to −48.99 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL)unseen frequency, 8-way−10.98 pp95% interval −12.24 to −9.67 ppNN(TPL) higherDifference only: no marginHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, 8-way−12.77 pp95% interval −14.52 to −10.99 ppNUM higherDifference only: no marginHeadroom gate passed
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, neighbour 3-way−62.36 pp95% interval −65.87 to −58.69 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, neighbour 3-way−1.24 pp95% interval −3.17 to +0.85 ppNo difference shownDifference only: no marginHeadroom gate passed
SSVEP · BETA · frozen CBraMod features (L1)
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, 8-way−61.94 pp95% interval −65.65 to −58.06 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL)unseen frequency, 8-way−1.16 pp95% interval −1.97 to −0.37 ppNN(TPL) higherDifference only: no marginHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, 8-way−2.46 pp95% interval −3.23 to −1.69 ppNUM higherDifference only: no marginHeadroom gate passed
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, neighbour 3-way−54.67 pp95% interval −57.77 to −51.47 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, neighbour 3-way+15.61 pp95% interval +14.25 to +16.93 ppTPL higherDifference only: no marginHeadroom gate passed

In the neighbour 3-way the unseen frequency competes with the trained detectors of its two neighbours (seen bias, stated). On a plain spectrum the template head fell below the mean chance of that test, 19.3% against 34.2%, and showed no difference against the numeric code. On frozen CBraMod features it beat the numeric code by +15.61 pp, but those features keep the stimulus phase and neighbouring targets differ in phase, so that comparison tests frequency and phase together. The plain spectrum, which carries no phase, is the clean frequency test.

Accuracy on the unseen frequencies (descriptive): u8, the 8-way among the held-out frequencies (chance 12.5%); nn8, the read-off that averages the answers for the two neighbouring seen frequencies; u3, the neighbour 3-way (mean chance 34.2%, a 2-way at the two grid edges). 70 people; three seeds, CCA none.
ArmTestA plain spectrum (L0)chance-correctedFrozen CBraMod features (L1)chance-corrected
TPLlabel templateu828.7%95% interval 26.9%–30.5%18.5%19.0%95% interval 18.0%–19.9%7.4%
DESCdescriptionu828.6%95% interval 26.8%–30.4%18.4%18.6%95% interval 17.6%–19.6%7.0%
SHUFshuffled templatesu89.2%95% interval 8.6%–9.7%-3.8%11.0%95% interval 10.5%–11.5%-1.7%
NUMnumeric codeu841.5%95% interval 38.2%–44.7%33.1%21.4%95% interval 20.2%–22.6%10.2%
CCAno trainingu880.9%95% interval 76.6%–84.8%78.2%80.9%95% interval 76.6%–84.8%78.2%
TPLlabel templatenn839.7%95% interval 36.8%–42.5%31.0%20.1%95% interval 19.0%–21.3%8.7%
IDquestion numbernn844.3%95% interval 40.8%–47.6%36.4%18.3%95% interval 17.3%–19.3%6.6%
B-shone output per questionnn844.3%95% interval 40.9%–47.5%36.3%17.6%95% interval 16.6%–18.8%5.8%
TPLlabel templateu319.3%95% interval 18.2%–20.5%-22.5%27.0%95% interval 26.0%–28.0%-10.8%
NUMnumeric codeu320.6%95% interval 19.3%–21.8%-20.6%11.4%95% interval 10.7%–12.1%-34.5%
CCAno trainingu381.7%95% interval 78.5%–84.6%72.2%81.7%95% interval 78.5%–84.6%72.2%
Generalisation cost (descriptive): accuracy on the unseen frequencies minus accuracy on matched seen ones, in pp with 95% intervals; 70 people.
Arm and testA plain spectrum (L0)Frozen CBraMod features (L1)
ID, 8-way (neighbour read-off)−14.26 pp95% interval −15.50 to −13.06 pp−27.80 pp95% interval −30.61 to −25.03 pp
TPL, 8-way−24.44 pp95% interval −26.44 to −22.26 pp−24.27 pp95% interval −26.72 to −21.85 pp
TPL, 3-way−27.83 pp95% interval −30.18 to −25.37 pp−29.67 pp95% interval −31.76 to −27.64 pp
DESC, 8-way−24.79 pp95% interval −26.88 to −22.55 pp−24.81 pp95% interval −27.31 to −22.39 pp
DESC, 3-way−28.95 pp95% interval −31.33 to −26.46 pp−29.74 pp95% interval −31.76 to −27.76 pp
SHUF, 8-way−40.63 pp95% interval −44.16 to −36.98 pp−30.09 pp95% interval −33.02 to −27.27 pp
SHUF, 3-way−34.96 pp95% interval −37.50 to −32.43 pp−32.52 pp95% interval −35.18 to −29.84 pp
NUM, 8-way−12.30 pp95% interval −13.42 to −11.17 pp−16.16 pp95% interval −17.88 to −14.47 pp
NUM, 3-way−7.27 pp95% interval −8.49 to −6.02 pp−7.63 pp95% interval −8.74 to −6.53 pp
B-sh, 8-way (neighbour read-off)−12.81 pp95% interval −14.09 to −11.56 pp−31.19 pp95% interval −34.09 to −28.32 pp
CCA, 8-way0.00 pp95% interval 0.00 to 0.00 pp0.00 pp95% interval 0.00 to 0.00 pp
CCA, 3-way+0.41 pp95% interval +0.26 to +0.56 pp+0.41 pp95% interval +0.26 to +0.56 pp
TPL − CCA in the 8-way, rotation by rotation, in pp with 95% intervals (secondary, recomputed by an independent audit); 70 people, three seeds. No rotation carries the result alone.
RotationA plain spectrum (L0)Frozen CBraMod features (L1)
Rotation 0−49.88 pp95% interval −53.16 to −46.16 pp−63.00 pp95% interval −67.23 to −58.54 pp
Rotation 1−52.96 pp95% interval −57.02 to −48.99 pp−61.76 pp95% interval −65.81 to −57.74 pp
Rotation 2−48.74 pp95% interval −51.97 to −45.42 pp−59.74 pp95% interval −63.75 to −55.51 pp
Rotation 3−57.34 pp95% interval −60.84 to −53.73 pp−62.75 pp95% interval −66.73 to −58.42 pp
Rotation 4−52.14 pp95% interval −55.80 to −48.20 pp−62.47 pp95% interval −66.53 to −58.23 pp

Per-frequency and interior-only breakdowns were computed after the audits, and no independent audit covers them, so they are not published.

Before any result

The pre-run checks, and how one was changed

A check was changed after it failed, twice. Everything about that is here, with the other checks that had to pass before any result was read.

A training-label permutation canary refits the heads on shuffled training labels and checks that test scores sit at chance: 17 configurations on the first fold, 38 intervals at 99.9%.

Revision 3 failed: 1 of 38 intervals excluded chance (sleep, frozen CBraMod features, the shuffled-template arm, unseen “light sleep” with its explicit name). Its labels had been permuted within each training person. The run stopped.

Revision 4, approved by the owner before any new fit, permuted labels across all training windows. It failed too: 3 of 38 intervals, all on BOAS, one of them below chance; every SSVEP interval included chance both times. Revision 4 said “no further revision”. The owner overrode that after seeing how it failed, and approved revision 5 before any new fit. Revision 4’s per-person centring diagnostic was not run: with labels permuted across all windows it no longer tests anything.

Revision 5 passed: 20 heads per configuration (340 fits), each metric compared with its exact per-person chance, and a two-level bootstrap over heads, then people; 0 of 38 intervals excluded zero. Both failed records are kept.

Direction counts, descriptive: of 760 single-head intervals in revision 5, 125 excluded the exact chance, 66 above and 59 below, and 18 cells failed in both directions. BOAS cells are published as pass or fail and counts only.

Why they failed, the reading revision 5 was built to test: a head trained on shuffled labels is still a random function of sleep features that are strongly organised by stage, so on new people its scores lean one way for everyone, up or down. A person bootstrap measures the spread between people, not between heads, so its 99.9% interval was too narrow for this question. The failures fit that reading: all were on BOAS, and one was below chance, which a leak cannot produce. It is not a proof that no leak exists: person-disjoint folds, no held-out label or wording in training and training-only scaling are asserted in the code, and both conformance audits check for leakage directly.

Decision S0b-6, approved by the owner: the SSVEP token rule says every token of a held-out frequency occurs in that rotation’s training texts. It stops the run for bge-small-en-v1.5 and all-MiniLM-L6-v2, and both passed in all 5 rotations. It cannot hold for multilingual-e5-small, whose tokenizer keeps one-decimal number pieces as single tokens while the held-out number itself is barred from training texts; for that encoder the rule was reported instead of stopping the run, and every Chinese-wording unseen-frequency entry says the frequency is a new token.

Engineering amendment EA-1, before the first fit: in the contiguous extrapolation band (S8) the neighbour average is not defined and the neighbour 3-way is relabelled, because its competitors are themselves held-out questions; failed fits are marked per cell, and none failed; the secondary scoring can resume after an interruption. No question, partition, wording, arm, metric, margin or gate changed.

The declared numeracy prediction was not activated. It said: if the primary text encoder orders frequencies poorly (Spearman ρ below 0.3 between frequency distance and embedding distance, for every training frame), the 8-way TPL − NUM flag will not read “TPL higher”. Its premise was false: ρ was 0.593, 0.563 and 0.526 for the three training frames of bge-small-en-v1.5, over 780 pairs. The flag reads “NUM higher” at both levels anyway; with the premise false, that is a description, not a confirmed prediction.

The gates

Four independent audits passed. They re-implement the definitions without importing the study’s scoring code: the conformance audit of the primary run passed 106 checks with 0 failed; the numerical audit of the primary made 2,270 comparisons with 0 mismatches; the conformance audit of the secondary run passed 231 checks with 0 failed; the numerical audit of the secondary made 18,619 comparisons, 29 of them outside the tolerance, every one explained.

Secondary · beyond the primary comparisons

Secondary results

Further datasets, features, a second text encoder, Chinese wordings and probes, none of them among the primary comparisons. One seed unless marked: a single seed carries no training-run variance. Every EESM19 and single-seed entry was declared likely inconclusive before the run, though none reached the “inconclusive” flags.

Show the secondary results

S1 · EESM19 sleep, frozen CBraMod features · 20 people, three seeds · declared likely inconclusive

ComparisonQuestionDifference: pp, log R or AUROCDifference shown?MarginGate
P1 · Label template − question numberTPL − IDseen, 5-way−0.04 pp95% interval −0.47 to +0.39 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 5-way−0.03 pp95% interval −0.43 to +0.40 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−0.11 pp95% interval −0.56 to +0.20 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−15.27 pp95% interval −16.24 to −14.29 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−18.54 pp95% interval −27.30 to −10.80 ppDESC(train) higher2 pp margin not metPasses the triviality canary
P3 · Label template head against the read-off from its own answersTPL / read-offasleep · implicit name+4.72095% interval +4.410 to +5.113R 112.17 (82.27–166.15)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offasleep · explicit name+2.56695% interval +2.179 to +3.010R 13.02 (8.84–20.30)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offN1, N2 or N3 sleep · explicit name+1.08095% interval +0.890 to +1.267R 2.95 (2.44–3.55)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offlight sleep · implicit name+2.40995% interval +2.208 to +2.609R 11.12 (9.10–13.58)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offlight sleep · explicit name+0.88895% interval +0.769 to +1.024R 2.43 (2.16–2.78)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offasleep · new descriptions10 held-out wordings, resampled+4.70795% interval +4.279 to +5.154R 110.72 (72.15–173.16)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offN1, N2 or N3 sleep · new descriptions10 held-out wordings, resampled+2.22395% interval +1.622 to +2.649R 9.24 (5.06–14.15)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offlight sleep · new descriptions10 held-out wordings, resampled+1.70795% interval +1.240 to +2.056R 5.51 (3.46–7.81)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P5 · Label template − shuffled templatesTPL − SHUFunseen questions, implicit names−0.39195% interval −0.470 to −0.311SHUF higherDifference only: no marginHeadroom gate passed
P5 · Label template − shuffled templatesTPL − SHUFunseen questions, explicit names+0.66795% interval +0.634 to +0.698TPL higherDifference only: no marginHeadroom gate passed

S2 · OpenBMI motor imagery, frozen CBraMod features · 51 people, three seeds

No triviality canary is declared for motor imagery, so its gate reads “not run” and excludes nothing. Motor imagery has only two three-cycles, so two of the three shuffled-template seeds share one derangement. The unseen “imagining a hand movement” came closest to its read-off: a difference, but within the ratio margin.

ComparisonQuestionDifference: pp, log R or AUROCDifference shown?MarginGate
P1 · Label template − question numberTPL − IDseen, 3-way−0.49 pp95% interval −1.07 to +0.02 ppNo difference shownEquivalent within ±2 ppNo canary declared for this task: not run
P1 · Description − question numberDESC − IDseen, 3-way−0.54 pp95% interval −1.08 to +0.05 ppNo difference shownEquivalent within ±2 ppNo canary declared for this task: not run
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled+0.05 pp95% interval −0.10 to +0.20 ppNo difference shownEquivalent within ±2 ppNo canary declared for this task: not run
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled+0.22 pp95% interval −0.11 to +0.55 ppNo difference shownEquivalent within ±2 ppNo canary declared for this task: not run
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−3.49 pp95% interval −6.63 to −1.42 ppDESC(train) higher2 pp margin not metNo canary declared for this task: not run
P3 · Label template head against the read-off from its own answersTPL / read-offimagining a hand movement · implicit name+0.07695% interval +0.050 to +0.107R 1.08 (1.05–1.11)Read-off leaves less errorEquivalent within the 1.25 ratio marginHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offimagining a hand movement · new descriptions10 held-out wordings, resampled+0.08895% interval +0.048 to +0.138R 1.09 (1.05–1.15)Read-off leaves less errorEquivalent within the 1.25 ratio marginHeadroom gate passed

S3 · REVE-L features, BETA · 70 people, one seed · declared likely inconclusive

REVE-L features keep the stimulus phase, so their neighbour 3-way tests frequency and phase together.

ComparisonQuestionDifference: pp, log R or AUROCDifference shown?MarginGate
P1 · Label template − question numberTPL − IDseen, 32-way−2.76 pp95% interval −3.51 to −2.07 ppID higher2 pp margin not metPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 32-way−2.56 pp95% interval −3.24 to −1.90 ppID higher2 pp margin not metPasses the triviality canary
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−4.48 pp95% interval −7.66 to −2.16 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−2.16 pp95% interval −2.44 to −1.88 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−5.05 pp95% interval −6.50 to −3.77 ppDESC(train) higher2 pp margin not metPasses the triviality canary
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, 8-way−65.32 pp95% interval −69.18 to −61.17 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL)unseen frequency, 8-way+5.81 pp95% interval +4.54 to +7.04 ppTPL higherDifference only: no marginHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, 8-way+1.92 pp95% interval +0.85 to +2.92 ppTPL higherDifference only: no marginHeadroom gate passed
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, neighbour 3-way−38.75 pp95% interval −41.62 to −35.60 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, neighbour 3-way+35.22 pp95% interval +32.91 to +37.42 ppTPL higherDifference only: no marginHeadroom gate passed

S4 · text encoder all-MiniLM-L6-v2, BETA, a plain spectrum · 70 people, one seed · declared likely inconclusive

ComparisonQuestionDifference: pp, log R or AUROCDifference shown?MarginGate
P1 · Label template − question numberTPL − IDseen, 32-way−6.73 pp95% interval −7.75 to −5.68 ppID higher2 pp margin not metPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 32-way−11.58 pp95% interval −13.16 to −9.97 ppID higher2 pp margin not metPasses the triviality canary
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−3.95 pp95% interval −5.11 to −2.86 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−4.90 pp95% interval −5.53 to −4.25 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−5.49 pp95% interval −6.48 to −4.60 ppDESC(train) higher2 pp margin not metPasses the triviality canary
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, 8-way−50.10 pp95% interval −53.06 to −46.73 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL)unseen frequency, 8-way−7.68 pp95% interval −8.82 to −6.50 ppNN(TPL) higherDifference only: no marginHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, 8-way−10.84 pp95% interval −12.53 to −9.08 ppNUM higherDifference only: no marginHeadroom gate passed
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, neighbour 3-way−61.85 pp95% interval −65.13 to −58.31 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, neighbour 3-way−0.18 pp95% interval −2.02 to +1.73 ppNo difference shownDifference only: no marginHeadroom gate passed

S4 · text encoder all-MiniLM-L6-v2, BETA, frozen CBraMod features · 70 people, one seed · declared likely inconclusive

ComparisonQuestionDifference: pp, log R or AUROCDifference shown?MarginGate
P1 · Label template − question numberTPL − IDseen, 32-way−4.73 pp95% interval −5.43 to −4.07 ppID higher2 pp margin not metPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 32-way−8.84 pp95% interval −10.12 to −7.60 ppID higher2 pp margin not metPasses the triviality canary
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−2.89 pp95% interval −3.73 to −2.08 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−3.09 pp95% interval −3.57 to −2.63 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−2.98 pp95% interval −3.67 to −2.34 ppDESC(train) higher2 pp margin not metPasses the triviality canary
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, 8-way−61.55 pp95% interval −65.21 to −57.49 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL)unseen frequency, 8-way−1.09 pp95% interval −1.69 to −0.48 ppNN(TPL) higherDifference only: no marginHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, 8-way−2.03 pp95% interval −2.86 to −1.19 ppNUM higherDifference only: no marginHeadroom gate passed
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, neighbour 3-way−54.63 pp95% interval −57.51 to −51.29 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, neighbour 3-way+15.57 pp95% interval +14.27 to +16.75 ppTPL higherDifference only: no marginHeadroom gate passed

S5 · Chinese wordings with multilingual-e5-small, BETA, a plain spectrum · 70 people, one seed · declared likely inconclusive

For this encoder each held-out frequency is a new token, not only a new combination of seen tokens (decision S0b-6).

ComparisonQuestionDifference: pp, log R or AUROCDifference shown?MarginGate
P1 · Label template − question numberTPL − IDseen, 32-way−3.15 pp95% interval −3.80 to −2.48 ppID higher2 pp margin not metPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 32-way−4.94 pp95% interval −5.69 to −4.18 ppID higher2 pp margin not metPasses the triviality canary
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−4.67 pp95% interval −6.57 to −3.24 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−8.70 pp95% interval −9.76 to −7.60 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−7.35 pp95% interval −8.58 to −6.17 ppDESC(train) higher2 pp margin not metPasses the triviality canary
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, 8-way−54.76 pp95% interval −57.77 to −51.25 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − average of the two neighbouring seen frequenciesTPL − NN(TPL)unseen frequency, 8-way−15.81 pp95% interval −17.68 to −13.94 ppNN(TPL) higherDifference only: no marginHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, 8-way−15.51 pp95% interval −17.58 to −13.44 ppNUM higherDifference only: no marginHeadroom gate passed
P4 · Label template − CCA, no trainingTPL − CCAunseen frequency, neighbour 3-way−59.98 pp95% interval −63.74 to −55.86 ppCCA higher2 pp margin not metHeadroom gate passed
P4 · Label template − numeric codeTPL − NUMunseen frequency, neighbour 3-way+1.69 pp95% interval −0.60 to +4.18 ppNo difference shownDifference only: no marginHeadroom gate passed

S8 · extrapolation to a contiguous held-out band at the top of the grid, BETA L0 · 70 people, one seed · declared likely inconclusive

ComparisonDifference, ppDifference shown?Margin, ±2 pp
TPL - CCA (8-way)−65.37 pp95% interval −69.21 to −61.40 ppCCA higher2 pp margin not met
TPL - NUM (8-way)+3.51 pp95% interval +1.64 to +5.49 ppTPL higherTPL non-inferior at 2 pp
TPL - NN(TPL) (8-way)Not defined (EA-1)
TPL - CCA (3-way)−50.19 pp95% interval −54.05 to −46.25 ppCCA higher2 pp margin not met

S8 · extrapolation to a contiguous held-out band at the top of the grid, BETA L1 · 70 people, one seed · declared likely inconclusive

ComparisonDifference, ppDifference shown?Margin, ±2 pp
TPL - CCA (8-way)−67.47 pp95% interval −71.22 to −63.56 ppCCA higher2 pp margin not met
TPL - NUM (8-way)+0.88 pp95% interval −0.70 to +2.47 ppNo difference shownTPL non-inferior at 2 pp
TPL - NN(TPL) (8-way)Not defined (EA-1)
TPL - CCA (3-way)−48.41 pp95% interval −52.10 to −44.40 ppCCA higher2 pp margin not met

In the band the neighbour 3-way competitors are themselves held-out questions, not trained detectors, and the neighbour average is not defined (EA-1). For bge-small-en-v1.5, many of these texts carry a number token never seen in training.

S12 · Wearable-102, heads trained on all of BETA’s frequencies and asked about off-grid ones · 102 people, three seeds

The off-grid frequencies have two decimals, whose hundredths never occur in BETA’s training texts, so the text arms face new tokens; the numeric code is the clean comparison. Nine fits, three arms by three seeds, fixed before the freeze. Computed from the Tsinghua author mirror; no physical unit is claimed.

ComparisonDifference, ppDifference shown?Margin, ±2 pp
Dry electrodes
TPL - CCA (12-way)−48.18 pp95% interval −52.36 to −43.90 ppCCA higher2 pp margin not met
NUM - CCA (12-way)−42.16 pp95% interval −46.05 to −38.38 ppCCA higher2 pp margin not met
TPL - NUM (12-way)−6.02 pp95% interval −7.32 to −4.73 ppNUM higher2 pp margin not met
DESC - CCA (12-way)−48.52 pp95% interval −52.99 to −44.02 ppCCA higher2 pp margin not met
Wet electrodes
TPL - CCA (12-way)−62.92 pp95% interval −66.16 to −59.66 ppCCA higher2 pp margin not met
NUM - CCA (12-way)−56.11 pp95% interval −59.16 to −52.83 ppCCA higher2 pp margin not met
TPL - NUM (12-way)−6.82 pp95% interval −8.31 to −5.33 ppNUM higher2 pp margin not met
DESC - CCA (12-way)−63.56 pp95% interval −66.68 to −60.31 ppCCA higher2 pp margin not met

S10, S11 and S13 · BETA L0: other arms against CCA and FBCCA, and held-out wordings · 70 people, three seeds

ComparisonDifference, ppDifference shown?Margin, ±2 pp
S10 DESC - CCA (8-way)−52.29 pp95% interval −55.41 to −49.00 ppCCA higher2 pp margin not met
S10 DESC - CCA (3-way)−62.72 pp95% interval −66.34 to −58.98 ppCCA higher2 pp margin not met
S10 DESC - NN(DESC) (8-way)−10.80 pp95% interval −12.15 to −9.39 ppNN(DESC) higher2 pp margin not met
S10 NUM - CCA (8-way)−39.44 pp95% interval −42.25 to −36.48 ppCCA higher2 pp margin not met
S10 TPL - NN(ID) (8-way)−15.63 pp95% interval −17.48 to −13.62 ppNN(ID) higher2 pp margin not met
S10 TPL - SHUF (8-way)+19.50 pp95% interval +17.45 to +21.62 ppTPL higherTPL non-inferior at 2 pp
S11 TPL - FBCCA (8-way)−61.68 pp95% interval −63.72 to −59.69 ppFBCCA higher2 pp margin not met
S11 TPL - FBCCA (3-way)−73.35 pp95% interval −75.61 to −70.93 ppFBCCA higher2 pp margin not met
S13 TPL held-out frames - CCA (8-way)−53.71 pp95% interval −56.73 to −50.27 ppCCA higher2 pp margin not met
S13 DESC held-out descriptions - CCA (8-way)−54.01 pp95% interval −57.07 to −50.83 ppCCA higher2 pp margin not met
S9 and S11 · BETA L0: accuracy on the seen and unseen frequencies in one test, and FBCCA (descriptive)
ArmSeen frequenciesUnseen frequencies
S9 TPL23.4%95% interval 21.0%–25.7%4.4%95% interval 4.0%–4.8%
S9 DESC24.0%95% interval 21.5%–26.5%4.2%95% interval 3.8%–4.7%
S9 SHUF23.4%95% interval 20.9%–25.7%1.1%95% interval 1.0%–1.4%
S9 NUM20.7%95% interval 18.5%–22.9%11.5%95% interval 10.2%–12.7%
One-output head, neighbour read-off (nn8)44.3%95% interval 40.9%–47.5%
FBCCA, seen 32-way81.1%95% interval 77.0%–84.8%
FBCCA, unseen 8-way90.4%95% interval 87.7%–92.8%
FBCCA, neighbour 3-way92.7%95% interval 90.9%–94.4%

S10, S11 and S13 · BETA L1: other arms against CCA and FBCCA, and held-out wordings · 70 people, three seeds

ComparisonDifference, ppDifference shown?Margin, ±2 pp
S10 DESC - CCA (8-way)−62.26 pp95% interval −66.01 to −58.41 ppCCA higher2 pp margin not met
S10 DESC - CCA (3-way)−54.93 pp95% interval −58.01 to −51.72 ppCCA higher2 pp margin not met
S10 DESC - NN(DESC) (8-way)−1.27 pp95% interval −2.08 to −0.46 ppNN(DESC) higher2 pp margin not met
S10 NUM - CCA (8-way)−59.48 pp95% interval −62.92 to −55.87 ppCCA higher2 pp margin not met
S10 TPL - NN(ID) (8-way)+0.68 pp95% interval −0.25 to +1.61 ppNo difference shownEquivalent within ±2 pp
S10 TPL - SHUF (8-way)+7.97 pp95% interval +6.88 to +9.09 ppTPL higherTPL non-inferior at 2 pp
S11 TPL - FBCCA (8-way)−71.41 pp95% interval −73.53 to −69.17 ppFBCCA higher2 pp margin not met
S11 TPL - FBCCA (3-way)−65.66 pp95% interval −67.61 to −63.62 ppFBCCA higher2 pp margin not met
S13 TPL held-out frames - CCA (8-way)−62.31 pp95% interval −66.06 to −58.27 ppCCA higher2 pp margin not met
S13 DESC held-out descriptions - CCA (8-way)−62.58 pp95% interval −66.15 to −58.59 ppCCA higher2 pp margin not met
S9 and S11 · BETA L1: accuracy on the seen and unseen frequencies in one test, and FBCCA (descriptive)
ArmSeen frequenciesUnseen frequencies
S9 TPL18.8%95% interval 16.7%–21.1%2.5%95% interval 2.2%–2.8%
S9 DESC18.9%95% interval 16.9%–21.2%2.5%95% interval 2.3%–2.8%
S9 SHUF18.3%95% interval 16.4%–20.4%1.6%95% interval 1.3%–1.9%
S9 NUM11.5%95% interval 10.4%–12.7%4.3%95% interval 3.8%–4.8%
One-output head, neighbour read-off (nn8)17.6%95% interval 16.6%–18.8%
FBCCA, seen 32-way81.1%95% interval 77.0%–84.8%
FBCCA, unseen 8-way90.4%95% interval 87.7%–92.8%
FBCCA, neighbour 3-way92.7%95% interval 90.9%–94.4%

S9 asks about seen and unseen frequencies in one 40-way test; FBCCA, like CCA, needs no training. One seed for FBCCA.

Secondary results on BOAS

BOAS is published with three stated gaps (owner decision, 7 October 2026):

  1. The consent statement does not say whether participants agreed to public sharing or secondary use.
  2. The ethics and consent statements come from the publisher's dataset description and README. No peer-reviewed paper describes BOAS.
  3. The ethics reference was added to the release in version 1.1.1 (May 2025), and the release does not say when it was granted relative to the recordings.

Participants are pseudonymised in the public release. No result here evaluates Bitbrain's headband or its automatic sleep scoring; neither is used.

Credit: Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

S3 · REVE-L features, BOAS · 100 people, one seed · declared likely inconclusive

ComparisonQuestionDifference: pp, log R or AUROCDifference shown?MarginGate
P1 · Label template − question numberTPL − IDseen, 5-way+0.27 pp95% interval −0.36 to +0.86 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 5-way+0.12 pp95% interval −0.48 to +0.71 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−0.39 pp95% interval −1.04 to +0.05 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−13.54 pp95% interval −14.57 to −12.49 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−21.86 pp95% interval −30.62 to −14.12 ppDESC(train) higher2 pp margin not metPasses the triviality canary
P3 · Label template head against the read-off from its own answersTPL / read-offasleep · implicit name+4.14995% interval +3.753 to +4.614R 63.39 (42.66–100.85)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offasleep · explicit name+2.64795% interval +2.274 to +3.070R 14.11 (9.72–21.55)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offN1, N2 or N3 sleep · explicit name+1.31795% interval +1.163 to +1.483R 3.73 (3.20–4.41)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offlight sleep · implicit name+2.73095% interval +2.540 to +2.929R 15.33 (12.67–18.71)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offlight sleep · explicit name+1.87295% interval +1.691 to +2.057R 6.50 (5.43–7.82)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offasleep · new descriptions10 held-out wordings, resampled+4.09295% interval +3.565 to +4.610R 59.84 (35.35–100.44)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offN1, N2 or N3 sleep · new descriptions10 held-out wordings, resampled+1.82995% interval +1.250 to +2.292R 6.23 (3.49–9.90)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offlight sleep · new descriptions10 held-out wordings, resampled+1.77495% interval +1.162 to +2.236R 5.89 (3.20–9.35)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P5 · Label template − shuffled templatesTPL − SHUFunseen questions, implicit names−0.58595% interval −0.612 to −0.555SHUF higherDifference only: no marginHeadroom gate passed
P5 · Label template − shuffled templatesTPL − SHUFunseen questions, explicit names+0.62695% interval +0.602 to +0.648TPL higherDifference only: no marginHeadroom gate passed

S4 · text encoder all-MiniLM-L6-v2, BOAS, frozen CBraMod features · 100 people, one seed · declared likely inconclusive

ComparisonQuestionDifference: pp, log R or AUROCDifference shown?MarginGate
P1 · Label template − question numberTPL − IDseen, 5-way+0.19 pp95% interval −0.37 to +0.75 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 5-way+0.01 pp95% interval −0.64 to +0.72 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−1.72 pp95% interval −4.31 to −0.09 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−14.04 pp95% interval −15.29 to −12.68 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−25.74 pp95% interval −34.29 to −16.26 ppDESC(train) higher2 pp margin not metPasses the triviality canary
P3 · Label template head against the read-off from its own answersTPL / read-offasleep · implicit name+3.79895% interval +3.385 to +4.273R 44.60 (29.53–71.76)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offasleep · explicit name+1.32595% interval +0.999 to +1.715R 3.76 (2.71–5.56)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offN1, N2 or N3 sleep · explicit name+0.88295% interval +0.717 to +1.047R 2.41 (2.05–2.85)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offlight sleep · implicit name+2.33895% interval +2.134 to +2.552R 10.36 (8.45–12.83)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offlight sleep · explicit name+1.79695% interval +1.598 to +2.000R 6.03 (4.94–7.39)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offasleep · new descriptions10 held-out wordings, resampled+3.06995% interval +2.508 to +3.601R 21.52 (12.28–36.65)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offN1, N2 or N3 sleep · new descriptions10 held-out wordings, resampled+1.11895% interval +0.793 to +1.449R 3.06 (2.21–4.26)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offlight sleep · new descriptions10 held-out wordings, resampled+1.07095% interval +0.840 to +1.333R 2.92 (2.32–3.79)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P5 · Label template − shuffled templatesTPL − SHUFunseen questions, implicit names−0.62695% interval −0.653 to −0.596SHUF higherDifference only: no marginHeadroom gate passed
P5 · Label template − shuffled templatesTPL − SHUFunseen questions, explicit names+0.57295% interval +0.550 to +0.592TPL higherDifference only: no marginHeadroom gate passed

S5 · Chinese wordings with multilingual-e5-small, BOAS, frozen CBraMod features · 100 people, one seed · declared likely inconclusive

ComparisonQuestionDifference: pp, log R or AUROCDifference shown?MarginGate
P1 · Label template − question numberTPL − IDseen, 5-way+0.34 pp95% interval −0.29 to +0.99 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P1 · Description − question numberDESC − IDseen, 5-way−0.28 pp95% interval −0.93 to +0.42 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P2 · New sentence frame − training wordingTPL(a1) − TPL(train)seen, reworded12 held-out wordings, resampled−0.70 pp95% interval −1.83 to +0.08 ppNo difference shownEquivalent within ±2 ppPasses the triviality canary
P2 · Synonym − training wordingTPL(a2) − TPL(train)seen, rewordedfor these three synonyms: fixed, not resampled−12.51 pp95% interval −13.73 to −11.27 ppTPL(train) higher2 pp margin not metPasses the triviality canary
P2 · New description − training descriptionsDESC(a3) − DESC(train)seen, reworded10 held-out wordings, resampled−25.81 pp95% interval −33.97 to −17.73 ppDESC(train) higher2 pp margin not metPasses the triviality canary
P3 · Label template head against the read-off from its own answersTPL / read-offasleep · implicit name+3.82695% interval +3.424 to +4.277R 45.87 (30.69–72.04)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offasleep · explicit name+1.55895% interval +1.216 to +1.950R 4.75 (3.37–7.03)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offN1, N2 or N3 sleep · explicit name+0.73795% interval +0.584 to +0.898R 2.09 (1.79–2.45)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offlight sleep · implicit name+2.32295% interval +2.122 to +2.532R 10.19 (8.34–12.57)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Label template head against the read-off from its own answersTPL / read-offlight sleep · explicit name+1.61595% interval +1.424 to +1.809R 5.03 (4.15–6.11)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offasleep · new descriptions10 held-out wordings, resampled+2.71595% interval +2.127 to +3.244R 15.11 (8.39–25.63)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offN1, N2 or N3 sleep · new descriptions10 held-out wordings, resampled+0.79895% interval +0.599 to +1.039R 2.22 (1.82–2.83)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed
P3 · Description head against the read-off from its own answersDESC / read-offlight sleep · new descriptions10 held-out wordings, resampled+0.94395% interval +0.773 to +1.123R 2.57 (2.17–3.07)Read-off leaves less error1.25 ratio margin not metHeadroom gate passed

S13 · held-out sentence frames against the read-off, BOAS · 100 people, three seeds

Comparisonlog R, with RDifference shown?Margin, R < 1.25
TPL asleep implicit (held-out frames)+3.56595% interval +3.162 to +4.006Read-off leaves less error1.25 ratio margin not met
TPL asleep explicit (held-out frames)+1.97695% interval +1.594 to +2.392Read-off leaves less error1.25 ratio margin not met
TPL NREM explicit (held-out frames)+0.92095% interval +0.743 to +1.100Read-off leaves less error1.25 ratio margin not met
TPL light implicit (held-out frames)+2.23395% interval +2.022 to +2.443Read-off leaves less error1.25 ratio margin not met
TPL light explicit (held-out frames)+1.48495% interval +1.240 to +1.719Read-off leaves less error1.25 ratio margin not met

S6 · negation probe, BOAS: a boundary probe, never a route sentence

AUROC of each negated wording against the truth of the negated question, beside the AUROC of 1 − p(X), the head’s own answer to X turned around; three seeds. “Non-REM sleep” is scored against the truth of “not REM”.
WordingAgainst the truth of the negated questionAUROC of 1 − p(X)People
not W0.02295% interval 0.014–0.0300.98095% interval 0.972–0.987100
not N10.13595% interval 0.120–0.1520.87295% interval 0.855–0.887100
not N20.12495% interval 0.110–0.1390.92495% interval 0.909–0.937100
not N30.02195% interval 0.017–0.0260.97995% interval 0.974–0.98371
not REM0.04995% interval 0.038–0.0610.95695% interval 0.944–0.967100
non-REM sleep (implicit)0.04895% interval 0.037–0.0600.95695% interval 0.944–0.967100

Every negated wording was answered as if it asked for X itself: against the truth of the negated question its AUROC lay far below 0.5, while turning the head’s answer to X around would have answered it well. Writing more prompts does not supply signal-level ground truth.

Published with the results

The wordings, the held-out frequencies and the shuffles

All 787 texts the questions were asked with, in English and Chinese, for sleep, SSVEP and motor imagery: training and held-out sentence frames, label names and synonyms, descriptions, the names and descriptions of the unseen questions, and the negation frames. Two agents wrote them from a fixed brief before any score, the second writing every held-out description without seeing the first’s; they are not users’ wordings and were not tuned. The file also holds the leakage rules they passed, the held-out frequencies of every rotation and the extrapolation band, the declared derangements of the shuffled control, the numeric code of a frequency and the pinned text encoders. Text only, CC BY 4.0.

The wordings, as JSON ↓

The sleep stages’ names in the training frames, and the three synonyms each was reworded with:

StageTraining nameSynonyms (a2)
Wwakewakefulness, awake, stage W
N1N1 sleepstage 1 sleep, sleep stage one, stage N1
N2N2 sleepstage 2 sleep, sleep stage two, stage N2
N3N3 sleepslow-wave sleep, deep sleep, stage N3
REMREM sleepstage R sleep, paradoxical sleep, dream sleep

Methods & limits

What these results can and cannot say

Not run, and not reported

Not published

The text encoders

Frozen at pinned commits and never redistributed.

The recordings

BETA

Bingchuan Liu et al. · BETA: A Large Benchmark Database Toward SSVEP-BCI Application (2020), doi:10.3389/fnins.2020.00627. Figshare 12264401 v3; mirror Bingchuan/BETA.

Source ↗ · CC BY 4.0

BOAS, the Bitbrain Open Access Sleep dataset

Eduardo López-Larraz, María Sierra-Torralba, Sergio Clemente, Galit Fierro, David Oriol, Javier Minguez, Luis Montesano and Jens G. Klinzing · The Bitbrain Open Access Sleep (BOAS) dataset, OpenNeuro ds005555, version 1.1.3 (2026), doi:10.18112/openneuro.ds005555.v1.1.3. Polysomnography EEG and the human-consensus stage labels only; the headband recordings and the publisher's automatic labels are not used.

Published with three stated gaps, printed beside its results above.

Source ↗ · CC0-1.0

EESM19

Kaare B. Mikkelsen et al. · Accurate whole-night sleep monitoring with dry-contact ear-EEG (2019), doi:10.1038/s41598-019-53115-3; OpenNeuro ds005185 v1.0.2. Processed mirror: Zachary1150/EESM19-Processed.

Source ↗ · CC0-1.0 declared by upstream and mirror

OpenBMI motor imagery

Min-Ho Lee, O-Yeon Kwon, Yong-Jeong Kim, Hong-Kyung Kim, Young-Eun Lee, John Williamson, Siamac Fazli and Seong-Whan Lee · EEG dataset and OpenBMI toolbox for three BCI paradigms: an investigation into BCI illiteracy, GigaScience (2019), giz002, doi:10.1093/gigascience/giz002. Data: Supporting data, GigaScience Database, doi:10.5524/100542.

Source ↗ · CC0-1.0

Wearable SSVEP BCI dataset (dry and wet electrodes)

Zhu, F., Jiang, L., Dong, G., Gao, X., & Wang, Y. (2021). An Open Dataset for Wearable SSVEP-Based Brain-Computer Interfaces (Version 4) [Data set]. Figshare. https://doi.org/10.6084/m9.figshare.13560281.v4

Zhu, F., Jiang, L., Dong, G., Gao, X., & Wang, Y. (2021). An Open Dataset for Wearable SSVEP-Based Brain-Computer Interfaces. Sensors, 21(4), 1256. https://doi.org/10.3390/s21041256

Computed from the official Tsinghua BCI Lab author mirror snapshot acquired 2026-09-20; the mirror itself has no separate DOI or version number.

Source ↗ · CC BY 4.0

Features: frozen CBraMod (primary) and frozen REVE Large @ 317531c7, REVE-L (secondary).

Weights terms: REVE under the REVE Responsible Use License v1.0 (model versions: REVE Large @ 317531c7). No model’s authors endorse these results. REVE, arXiv:2510.21585 ↗

Methods and sources these results draw on

Data source: questions-in-language-update.json · schema bci-report-questions-in-language-update-v1.

Wordings: questions-in-language-wordings.json · schema bci-report-questions-in-language-wordings-v1.

Cite this page

BCI Report (2026). Can an EEG model answer questions asked in words? https://bci.report/topics/questions-in-language/

Figures from release questions-in-language-update-20261008 (2026-10-08). Cite the upstream datasets as well: their credits are on this page.

Every release is archived on Zenodo: doi:10.5281/zenodo.23123296. BibTeX for the site and its releases → · CITATION.cff ↗