Session transfer · the same person, no labels from the later session

Does a decoder trained on an earlier session still work later?

Three datasets, four results, each read on its own. Motor imagery across recording sessions: frozen CBraMod with a readout fitted on the earlier session scores above a spectral baseline in both cohorts. Target detection at later visits: a first-visit decoder ranks targets less well at the last nominal visit, and its average precision is lower. Continuous cursor tracking: a fixed decoder does worse than a constant baseline for every admitted record.

Measured on: WBCIC-SHU motor imagery dataset; Longitudinal ERP dataset (RSVP); Forenzo & He continuous-tracking EEG-BCI dataset · Methods: CBraMod

Short answer

Partly, and each result stands on its own. On WBCIC-SHU motor imagery — trained on each person’s recording session 1 and tested on session 3 with no labels from it, 51 people in a two-class and 11 in a three-class cohort, kept apart — frozen CBraMod with a ridge readout fitted on session 1 reached 67.7% and 51.6% balanced accuracy, +13.8 pp and +14.0 pp above a relative spectral ridge, each paired interval above zero. It needs that person’s labelled earlier session, so it is same-person transfer, not decoding for someone new, and it does not show that pretraining caused the gain. On a longitudinal RSVP task, a decoder trained at each of 15 people’s first visit ranked targets with an AUROC of 0.888 at the publisher’s nominal Day 7 visit and 0.853 at nominal Day 200, a paired change of −0.035 (−0.065 to −0.006). With targets only about 2.5% of events, average precision fell from 0.359 to 0.240. The nominal visit labels do not tie either decline to elapsed time. In continuous cursor tracking, a fixed spectral ridge trained on each record’s earliest session had a higher overall error in its latest session than a constant source-mean comparator, for every admitted record in both cohorts and both response arms: a negative result, with a severe upper tail.

WBCIC-SHU · motor imagery · session 1 to session 3

Motor imagery two sessions later: two baselines and a frozen encoder

People imagined moving the left or the right hand (a two-class cohort, 51 people) or, in a separate three-class cohort (11 people), a foot movement as well, in three recording sessions. Each person’s complete session 1 trains the decoder and session 3 is the test; session 2 is unused, and no label from session 3 is used. Session numbers are recording-session ordinals, not a guaranteed time gap or distinct calendar days. The cohorts are kept apart and never averaged.

Participant-equal mean balanced accuracy on session 3, with a paired whole-participant bootstrap 95% interval. Chance is 50.0% for two classes and 33.3% for three, where the source-majority prior sits by construction. Accuracy and macro F1 are the secondary means.
DecoderTwo-class motor imagery · 51 peopleleft or right handThree-class motor imagery · 11 peopleleft hand, right hand or foot
Source-majority priorpredicts session 1’s most frequent class: a floor, not a model50.0%Accuracy 50.0% · macro F1 33.3%33.3%Accuracy 33.3% · macro F1 16.7%
Relative spectral power + ridgea fixed CPU baseline53.9%95% interval 52.7%–55.1%Accuracy 53.9% · macro F1 49.7%37.6%95% interval 35.4%–40.0%Accuracy 37.6% · macro F1 34.1%
CBraMod · frozen encoder, session-1 ridge readoutonly the readout is fitted; the encoder is not trained67.7%95% interval 64.8%–70.5%Accuracy 67.7% · macro F1 66.9%51.6%95% interval 45.2%–58.3%Accuracy 51.6% · macro F1 50.5%
Paired differences in balanced accuracy, in percentage points (pp), the same people and test trials. Each interval resamples matched per-person differences within its cohort, conditional on the fixed trained models and predictions.
DecoderTwo-class motor imagery · 51 peopleThree-class motor imagery · 11 people
Spectral ridge minus the source prior+3.9 pp95% interval +2.7 to +5.1 ppInterval above zero.+4.3 pp95% interval +2.1 to +6.6 ppInterval above zero.
Frozen CBraMod minus the spectral ridge+13.8 pp95% interval +11.3 to +16.4 ppInterval above zero.+14.0 pp95% interval +7.5 to +20.8 ppInterval above zero.
+13.8 ppFrozen CBraMod minus the spectral ridge, two-class cohort, 51 people

Above the spectral baseline in both cohorts — as a comparison of two pipelines

Paired interval +11.3 to +16.4 pp. In the small three-class cohort the gap was +14.0 pp (+7.5 to +20.8 pp). The CBraMod arm was frozen after the spectral results existed and reuses their test trials: a comparative follow-up, not a fresh test set. There is no matched random-weight control, the two feature spaces differ in size and preprocessing (the same ridge alpha does not equalize regularization across them), and whether this checkpoint saw WBCIC-SHU in pretraining is not established, so the gap compares the two pipelines; it does not show that pretraining caused it. Each person’s session-1 labels train the readout: zero labels from session 3 is not a decoder for someone new.

Trials: 10,199 training and 10,195 test trials in the two-class cohort, 3,299 and 3,300 in the three-class cohort; no person or trial was excluded. 6 of the 186 session records hold fewer trials than the publisher’s nominal count (40,490 delivered against 40,500); 3 of them are in the sessions used and are admitted under a reviewed variable-N rule, with no filling, truncation or performance-based exclusion; the original strict-count holds stay recorded. Why the trials are missing is unknown.

Longitudinal RSVP · target detection · a first-visit decoder

Does your decoder still work at later visits?

Same-person, offline RSVP classification; publisher nominal visit labels; no target-session adaptation.

In rapid serial visual presentation (RSVP), a stream of face images is shown and the decoder has to pick out the rare target. 15 people, the source’s Group A, were recorded at visits the publisher labels Day 1, 7, 80 and 200. Each person’s decoder is trained and calibrated on their first visit only, then scored unchanged at each later visit: 72,000 events per visit across the 15 people.

Decoder: Normalized ERP features + logistic/Platt (CPU baseline).

Nominal Day 795% interval 0.864–0.911
0.888
Nominal Day 8095% interval 0.854–0.903
0.879
Nominal Day 20095% interval 0.827–0.879
0.853
AUROC at each nominal later visit: the same people, one first-visit decoder each. Participant-equal mean AUROC. Dot: the mean; line: its participant-bootstrap 95% interval, printed under each label; dashed line: chance, a decoder that orders events at random. The visits are labels, evenly spaced, not a time axis, and no curve is fitted between them. AUROC ranks targets above non-targets; it is not accuracy.

Read the chart with these

  • Same person: each person’s first-visit labels train and calibrate their own decoder. Not a decoder for someone new.
  • Nominal visits: Day 7, Day 80 and Day 200 are the publisher’s labels, not verified participant-specific calendar intervals.
  • Rare targets: about 2.5% of events at every visit (the table gives each). AUROC ranks; average precision is not precision at a chosen threshold; neither is an online selection success rate. The target share describes the data; it is not a comparator.
  • 15 people and one fixed offline baseline: no clinical, real-time or causal claim, and no foundation model.
Participant-equal means at each nominal visit, 15 people; counts are event totals.
MeasureNominal Day 7Nominal Day 80Nominal Day 200
AUROCranks targets, 0 to 1; not accuracy0.88895% interval 0.864–0.9110.87995% interval 0.854–0.9030.85395% interval 0.827–0.879
Average precision0 to 1; not precision at a threshold0.3590.3100.240
Log lossper event, natural log; lower is better0.1540.1140.164
Brier scorelower is better0.0290.0270.038
ECE, 10 binsexpected calibration error; lower is better0.0230.0200.039
Target eventsof all events at the visit1,794 of 72,0002.5%1,786 of 72,0002.5%1,793 of 72,0002.5%
Delivered eventsnot available from the protected scoring interfacenot available (null, not zero)not available (null, not zero)not available (null, not zero)
−0.035AUROC, nominal Day 200 minus Day 7, the same 15 people

AUROC and average precision lower at nominal Day 200

Paired interval −0.065 to −0.006, below zero. 6 of 15 people lost 0.05 AUROC or more. Average precision is lower too: 0.359 at nominal Day 7, 0.240 at Day 200. This describes this cohort and this decoder; it does not show that elapsed time caused the change.

Label coding: the paper and the publisher’s sample script code 1 as non-target and 2 as target, and the release’s Trigger.txt reverses them. The paper and script were followed, a precedence decided before modelling; the contradiction stays visible.

Continuous cursor tracking · a negative result

A simple baseline exposes cross-session decoder failures

Offline, same-person, no target-session adaptation.

People used motor imagery to steer a cursor after a target moving continuously across a screen, in two cohorts the publisher calls Main and Transfer Learning. For each admitted record, a fixed spectral ridge is trained on every eligible non-chance run of the earliest complete recorded session and tested on the latest, with no label from it, against a source-mean comparator: the earliest session’s average response, repeated for every row. Can the decoder beat that constant in the later session? Here, it did not.

Cohort records admitted

23

Main 9 of 14, Transfer Learning 14 of 14. Records, not proven unique people.

Target rows per response arm

1,999,926

The same rows for both arms: two outcomes do not double the independent sample. No response row was removed.

Response fits

46

One per admitted record and response arm; every one succeeded.

The error is the joint source-SD-normalized RMSE: errors divided by the earliest session’s response standard deviation, averaged over both axes within each trial before the square root. Lower is better. It is an error, not a percentage or an accuracy. Each record averages its trials equally, and each cohort mean weights records equally; intervals resample records within a cohort.

Two response arms, each with its own table. They are different outcomes on the same rows: never combined, and their errors are not compared with each other.

Historical decoder velocity (historical decoder imitation)

The output of the decoder in use at recording time, as the publisher stored it, imitated: not intended hand motion or intended control.

Historical decoder velocity (historical decoder imitation) · error in the latest session, lower is better: the ridge’s mean with its interval and its median, the comparator’s mean, and the paired difference (positive: the ridge’s error is larger).
CohortSpectral ridge · meanSpectral ridge · medianSource-mean comparator · meanRidge minus comparator · same records
Main cohort9 of 14 records admitted192.29095% interval 1.292 to 574.2351.317the mean is 146 times the median1.26695% interval 1.214 to 1.316191.02495% interval 0.026 to 572.995ridge error higher on 9 of 9 records
Transfer Learning cohort14 of 14 records admitted44.17295% interval 0.955 to 130.4760.990the mean is 45 times the median0.86495% interval 0.801 to 0.92943.30895% interval 0.105 to 129.601ridge error higher on 14 of 14 records
Main · spectral ridge · mean95% interval 1.292–574.235
192.290
Main · spectral ridge · median
1.317
Main · source-mean comparator · mean95% interval 1.214–1.316
1.266
Transfer Learning · spectral ridge · mean95% interval 0.955–130.476
44.172
Transfer Learning · spectral ridge · median
0.990
Transfer Learning · source-mean comparator · mean95% interval 0.801–0.929
0.864
Historical decoder velocity: the ridge’s mean and median beside the comparator, on a log scale. Read within this arm only. Log scale, powers of ten from 10⁻² to 10⁶: each gridline is a hundredfold step, marking 1, 100 and 10,000. Dot: the value; line: its 95% interval; the ridge’s median has no interval here. The printed values are the published figures; the strip is their picture.

Constructed displacement proxy (constructed proxy)

Target position minus cursor position in the publisher’s screen coordinates, a separately fitted proxy.

Constructed displacement proxy (constructed proxy) · error in the latest session, lower is better: the ridge’s mean with its interval and its median, the comparator’s mean, and the paired difference (positive: the ridge’s error is larger).
CohortSpectral ridge · meanSpectral ridge · medianSource-mean comparator · meanRidge minus comparator · same records
Main cohort9 of 14 records admitted698.49995% interval 1.166 to 2,0931.225the mean is 570 times the median1.05695% interval 0.977 to 1.133697.44395% interval 0.121 to 2,092ridge error higher on 9 of 9 records
Transfer Learning cohort14 of 14 records admitted69.84495% interval 1.071 to 207.2871.182the mean is 59 times the median0.91695% interval 0.841 to 0.98768.92895% interval 0.143 to 206.410ridge error higher on 14 of 14 records
Main · spectral ridge · mean95% interval 1.166–2,093
698.499
Main · spectral ridge · median
1.225
Main · source-mean comparator · mean95% interval 0.977–1.133
1.056
Transfer Learning · spectral ridge · mean95% interval 1.071–207.287
69.844
Transfer Learning · spectral ridge · median
1.182
Transfer Learning · source-mean comparator · mean95% interval 0.841–0.987
0.916
Constructed displacement proxy: the ridge’s mean and median beside the comparator, on a log scale. Read within this arm only. Log scale, powers of ten from 10⁻² to 10⁶: each gridline is a hundredfold step, marking 1, 100 and 10,000. Dot: the value; line: its 95% interval; the ridge’s median has no interval here. The printed values are the published figures; the strip is their picture.

Across the four cohort-and-arm cells, the ridge’s mean error is 45 to 570 times its median: a few records carry very large errors. No scored record was removed, clipped or winsorized, and what caused the extreme errors — feature distribution shift, scaling sensitivity or something else — is not established. The every-record statement is about the overall primary error; it does not extend to the secondary metrics or to each recorded-decoder context.

Coverage. Main admits 9 of its 14 candidate records and is read conditional on that subset; 5 were held on metadata before any scoring and stay recorded: 2 because a session lacks its required Chance R01 record; 2 because the decoder-local run allocation differs from the documented contract; 1 because the sample-boundary geometry differs; its exact numeric cause was not independently reconstructed. Transfer Learning admits 14 of 14. That is 23 admitted cohort records, not 23 proven unique people: the cohorts may share people, and they are never pooled.

Transfer Learning is the publisher’s name for how that cohort’s data were collected; no transfer-learning model was trained here.

Secondary metrics per axis — raw RMSE, R² and Pearson correlation

Record-equal means with their 95% intervals. Raw RMSE is in each response’s stored units; R² and Pearson are unitless. A constant prediction has no correlation, so the comparator’s Pearson values are undefined: null, not zero. The every-record statement above does not extend to these metrics, and their paired differences are not published.

Historical decoder velocity · secondary metrics in the latest session.
Cohort and modelRaw RMSE, xRaw RMSE, yR², xR², yPearson, xPearson, y
Main cohortSpectral ridge38.88995% interval 0.309 to 116.03850.59995% interval 0.312 to 151.159−172,98395% interval −518,950 to −0.018−296,66395% interval −889,988 to −0.0460.02995% interval 0.00233 to 0.0610.03495% interval 0.00105 to 0.073
Main cohortSource-mean comparator0.30695% interval 0.290 to 0.3210.30295% interval 0.289 to 0.313−0.00049895% interval −0.000786 to −0.000255−0.00037795% interval −0.000659 to −0.000127undefined (null)undefined (null)
Transfer Learning cohortSpectral ridge4.44395% interval 0.212 to 12.87713.81595% interval 0.229 to 40.954−5,44095% interval −16,319 to −0.188−58,11795% interval −174,350 to −0.3390.04595% interval 0.00527 to 0.0960.01995% interval −0.022 to 0.075
Transfer Learning cohortSource-mean comparator0.19895% interval 0.184 to 0.2120.20195% interval 0.187 to 0.217−0.00018795% interval −0.000314 to −0.0000754−0.00093795% interval −0.00140 to −0.000531undefined (null)undefined (null)
Constructed displacement proxy · secondary metrics in the latest session.
Cohort and modelRaw RMSE, xRaw RMSE, yR², xR², yPearson, xPearson, y
Main cohortSpectral ridge124.82895% interval 0.412 to 373.616332.15295% interval 0.431 to 995.534−928,35095% interval −2,785,050 to −0.183−6,442,34195% interval −19,327,021 to −0.3880.05095% interval 0.031 to 0.0740.07295% interval 0.018 to 0.137
Main cohortSource-mean comparator0.38895% interval 0.366 to 0.4130.37695% interval 0.349 to 0.404−0.02695% interval −0.045 to −0.00771−0.02195% interval −0.039 to −0.00612undefined (null)undefined (null)
Transfer Learning cohortSpectral ridge9.77295% interval 0.415 to 28.42338.15195% interval 0.445 to 113.515−21,50495% interval −64,511 to −0.318−337,64095% interval −1,012,920 to −0.3910.03595% interval 0.00847 to 0.0670.03495% interval 0.00753 to 0.062
Transfer Learning cohortSource-mean comparator0.36695% interval 0.327 to 0.4020.38095% interval 0.341 to 0.414−0.0066495% interval −0.010 to −0.00352−0.01695% interval −0.030 to −0.00526undefined (null)undefined (null)

Methods & limits

What these results can and cannot say

How each result was produced

WBCIC-SHU baselines

The publisher’s processed derivative: 58 anonymous channel indices at 250 Hz in four-second epochs. Welch relative band power in 4–8, 8–13, 13–30 and 30–40 Hz per channel, as log ratios: 232 values per trial. A StandardScaler fitted on session 1 only, then a RidgeClassifier (alpha 1, no class weights). No hyperparameter search, target normalization, calibration or abstention.

WBCIC-SHU, frozen CBraMod

The same trials, resampled to 200 Hz and zero-padded to 800 samples; each trial and channel divided by its own standard deviation, a dimensionless input, not the checkpoint’s native amplitude, so upstream benchmark numbers are not reproduced. CBraMod in evaluation mode with every weight frozen; its four temporal patch embeddings averaged into 11,600 values per trial; a separate StandardScaler and RidgeClassifier fitted on each person’s session 1 only. No backbone training, LoRA or target-session calibration.

Longitudinal RSVP

57 EEG channels, −200 to +700 ms around each event, baseline-normalized; six 100 ms averages from +100 to +700 ms, 342 values per event. Per person, first-visit blocks 1–3 fit a fixed L2 logistic regression and block 4 fits Platt calibration; later visits are scored on blocks 2–4. Group B is not used. No hyperparameter search, target-label calibration, fine-tuning or LoRA.

Continuous cursor tracking

62 channels at 1,000 Hz, already band-passed and notch-filtered by the publisher. Five bands from 4 to 30 Hz through a per-trial FIR filter, the preceding second of band power as channel-relative log power: 310 values per row, from a trial-contained 2-second past window, every 40 ms. Per response arm, a StandardScaler and response scaling fitted on the earliest session only, then Ridge (alpha 1). No tuning, LoRA, target-session calibration or foundation-model inference.

What the intervals cover

Every interval is a pointwise 95% percentile interval from 10,000 participant (or, for cursor tracking, record) bootstrap resamples within its cohort, with paired draws for every difference. Each is conditional on the fixed trained models and predictions: refitting, preprocessing, model and protocol selection — and, for CBraMod, pretraining overlap — are not included, and there is no multiple-comparison adjustment.

Limits that hold for all four results

Not published here

Per-person and per-record values of any kind — every minimum and percentile, and every median other than the four cursor-tracking ridge medians printed beside their means; which people declined or which records did worse; WBCIC-SHU confusion matrices; the cursor-tracking results split by recorded decoder; measured compute; and any participant, record or file identifier.

Credits, licences and consent

WBCIC-SHU motor imagery (Yang et al. 2025)

Banghua Yang, Fenqi Rong, Yunlong Xie, Du Li, Jiayang Zhang, Fu Li, Guangming Shi and Xiaorong Gao · A multi-day and high-quality EEG dataset for motor imagery brain-computer interface, Scientific Data 12, 488 (2025), doi:10.1038/s41597-025-04826-y. Data: Banghua Yang and Fenqi Rong, WBCIC-SHU Motor Imagery Dataset, Figshare+, doi:10.25452/figshare.plus.22671172.v5, CC BY 4.0. Derived analysis by BCI Report; not endorsed by the authors.

Ethics and consent, as the paper states them: Approved by the Tsinghua University Medical Ethics Committee (approval number 20190002); carried out in line with the Declaration of Helsinki. Written informed consent was obtained from the participants after they were informed about the procedure, purpose, requirements and motor-imagery techniques.

Dataset record ↗ · Paper ↗ · CC BY 4.0 · Read from the paper’s full text ↗

Longitudinal ERP dataset, RSVP face task (Yang et al. 2025)

Chen Yang, Yufeng Zhang, Hongxin Zhang, Yixuan Li, Yijun Wang and Xiaorong Gao · A longitudinal EEG dataset of event-related potential, figshare, doi:10.6084/m9.figshare.27201003.v1, CC0. Paper: Yufeng Zhang, Hongxin Zhang, Yixuan Li, Yijun Wang, Xiaorong Gao and Chen Yang, A longitudinal EEG dataset of event-related potential, Scientific Data 12, 1069 (2025), doi:10.1038/s41597-025-05378-x. Derived analysis by BCI Report; not endorsed by the authors.

Ethics and consent, as the paper states them: Approved by the Tsinghua Institutional Review Board (No. 20230058). Each participant read and signed an informed consent form before the experiment.

Dataset record ↗ · Paper ↗ · CC0-1.0 · Read from the paper’s full text ↗

Forenzo & He continuous-tracking EEG-BCI dataset (KiltHub)

Dylan Forenzo and Bin He · EEG-BCI Dataset for “Continuous Tracking using Deep Learning-based Decoding for Non-invasive Brain-Computer Interface”, KiltHub, Carnegie Mellon University, doi:10.1184/R1/25360300.v1, CC BY 4.0. Citation the record requests: Dylan Forenzo, Hao Zhu, Jenn Shanahan, Jaehyun Lim and Bin He, Continuous tracking using deep learning-based decoding for noninvasive brain–computer interface, PNAS Nexus 3(4), pgae145 (2024), doi:10.1093/pnasnexus/pgae145. Derived analysis by BCI Report; not endorsed by the authors and not a reproduction of the paper's online results.

Ethics and consent, as the paper states them: Approved by the Institutional Review Board at Carnegie Mellon University; no approval number is given. Each subject provided written consent to the protocol before participating.

Dataset record ↗ · Paper ↗ · CC BY 4.0 · Read from the paper’s full text ↗

CBraMod

CBraMod, the frozen encoder: its paper and the pinned checkpoint’s SHA-256 0792cb808c14… (revision b9e961003214…). Its pretraining exposure for WBCIC-SHU is not established: the model paper describes pretraining on the TUH EEG corpus and the pinned repository lists SHU as a downstream task, and neither gives a checkpoint-specific file manifest.

CBraMod paper ↗ · CBraMod’s page

Derived analyses by BCI Report, aggregates only; not endorsed by the data’s authors, and no raw data are redistributed.

Data source: later-sessions-update.json · schema bci-report-later-sessions-update-v1.

Cite this page

BCI Report (2026). Does a decoder trained on an earlier session still work later? https://bci.report/topics/later-sessions/

Figures from release later-sessions-update-20261008 (2026-10-08). Cite the upstream datasets as well: their credits are on this page.

Every release is archived on Zenodo: doi:10.5281/zenodo.23123296. BibTeX for the site and its releases → · CITATION.cff ↗