Cohort records admitted
Main 9 of 14, Transfer Learning 14 of 14. Records, not proven unique people.
Session transfer · the same person, no labels from the later session
Three datasets, four results, each read on its own. Motor imagery across recording sessions: frozen CBraMod with a readout fitted on the earlier session scores above a spectral baseline in both cohorts. Target detection at later visits: a first-visit decoder ranks targets less well at the last nominal visit, and its average precision is lower. Continuous cursor tracking: a fixed decoder does worse than a constant baseline for every admitted record.
Measured on: WBCIC-SHU motor imagery dataset; Longitudinal ERP dataset (RSVP); Forenzo & He continuous-tracking EEG-BCI dataset · Methods: CBraMod
Partly, and each result stands on its own. On WBCIC-SHU motor imagery — trained on each person’s recording session 1 and tested on session 3 with no labels from it, 51 people in a two-class and 11 in a three-class cohort, kept apart — frozen CBraMod with a ridge readout fitted on session 1 reached 67.7% and 51.6% balanced accuracy, +13.8 pp and +14.0 pp above a relative spectral ridge, each paired interval above zero. It needs that person’s labelled earlier session, so it is same-person transfer, not decoding for someone new, and it does not show that pretraining caused the gain. On a longitudinal RSVP task, a decoder trained at each of 15 people’s first visit ranked targets with an AUROC of 0.888 at the publisher’s nominal Day 7 visit and 0.853 at nominal Day 200, a paired change of −0.035 (−0.065 to −0.006). With targets only about 2.5% of events, average precision fell from 0.359 to 0.240. The nominal visit labels do not tie either decline to elapsed time. In continuous cursor tracking, a fixed spectral ridge trained on each record’s earliest session had a higher overall error in its latest session than a constant source-mean comparator, for every admitted record in both cohorts and both response arms: a negative result, with a severe upper tail.
WBCIC-SHU · motor imagery · session 1 to session 3
People imagined moving the left or the right hand (a two-class cohort, 51 people) or, in a separate three-class cohort (11 people), a foot movement as well, in three recording sessions. Each person’s complete session 1 trains the decoder and session 3 is the test; session 2 is unused, and no label from session 3 is used. Session numbers are recording-session ordinals, not a guaranteed time gap or distinct calendar days. The cohorts are kept apart and never averaged.
| Decoder | Two-class motor imagery · 51 peopleleft or right hand | Three-class motor imagery · 11 peopleleft hand, right hand or foot |
|---|---|---|
| Source-majority priorpredicts session 1’s most frequent class: a floor, not a model | 50.0%Accuracy 50.0% · macro F1 33.3% | 33.3%Accuracy 33.3% · macro F1 16.7% |
| Relative spectral power + ridgea fixed CPU baseline | 53.9%95% interval 52.7%–55.1%Accuracy 53.9% · macro F1 49.7% | 37.6%95% interval 35.4%–40.0%Accuracy 37.6% · macro F1 34.1% |
| CBraMod · frozen encoder, session-1 ridge readoutonly the readout is fitted; the encoder is not trained | 67.7%95% interval 64.8%–70.5%Accuracy 67.7% · macro F1 66.9% | 51.6%95% interval 45.2%–58.3%Accuracy 51.6% · macro F1 50.5% |
| Decoder | Two-class motor imagery · 51 people | Three-class motor imagery · 11 people |
|---|---|---|
| Spectral ridge minus the source prior | +3.9 pp95% interval +2.7 to +5.1 ppInterval above zero. | +4.3 pp95% interval +2.1 to +6.6 ppInterval above zero. |
| Frozen CBraMod minus the spectral ridge | +13.8 pp95% interval +11.3 to +16.4 ppInterval above zero. | +14.0 pp95% interval +7.5 to +20.8 ppInterval above zero. |
Paired interval +11.3 to +16.4 pp. In the small three-class cohort the gap was +14.0 pp (+7.5 to +20.8 pp). The CBraMod arm was frozen after the spectral results existed and reuses their test trials: a comparative follow-up, not a fresh test set. There is no matched random-weight control, the two feature spaces differ in size and preprocessing (the same ridge alpha does not equalize regularization across them), and whether this checkpoint saw WBCIC-SHU in pretraining is not established, so the gap compares the two pipelines; it does not show that pretraining caused it. Each person’s session-1 labels train the readout: zero labels from session 3 is not a decoder for someone new.
Trials: 10,199 training and 10,195 test trials in the two-class cohort, 3,299 and 3,300 in the three-class cohort; no person or trial was excluded. 6 of the 186 session records hold fewer trials than the publisher’s nominal count (40,490 delivered against 40,500); 3 of them are in the sessions used and are admitted under a reviewed variable-N rule, with no filling, truncation or performance-based exclusion; the original strict-count holds stay recorded. Why the trials are missing is unknown.
Longitudinal RSVP · target detection · a first-visit decoder
Same-person, offline RSVP classification; publisher nominal visit labels; no target-session adaptation.
In rapid serial visual presentation (RSVP), a stream of face images is shown and the decoder has to pick out the rare target. 15 people, the source’s Group A, were recorded at visits the publisher labels Day 1, 7, 80 and 200. Each person’s decoder is trained and calibrated on their first visit only, then scored unchanged at each later visit: 72,000 events per visit across the 15 people.
Decoder: Normalized ERP features + logistic/Platt (CPU baseline).
| Measure | Nominal Day 7 | Nominal Day 80 | Nominal Day 200 |
|---|---|---|---|
| AUROCranks targets, 0 to 1; not accuracy | 0.88895% interval 0.864–0.911 | 0.87995% interval 0.854–0.903 | 0.85395% interval 0.827–0.879 |
| Average precision0 to 1; not precision at a threshold | 0.359 | 0.310 | 0.240 |
| Log lossper event, natural log; lower is better | 0.154 | 0.114 | 0.164 |
| Brier scorelower is better | 0.029 | 0.027 | 0.038 |
| ECE, 10 binsexpected calibration error; lower is better | 0.023 | 0.020 | 0.039 |
| Target eventsof all events at the visit | 1,794 of 72,0002.5% | 1,786 of 72,0002.5% | 1,793 of 72,0002.5% |
| Delivered eventsnot available from the protected scoring interface | not available (null, not zero) | not available (null, not zero) | not available (null, not zero) |
Paired interval −0.065 to −0.006, below zero. 6 of 15 people lost 0.05 AUROC or more. Average precision is lower too: 0.359 at nominal Day 7, 0.240 at Day 200. This describes this cohort and this decoder; it does not show that elapsed time caused the change.
Label coding: the paper and the publisher’s sample script code 1 as non-target and 2 as target, and the release’s Trigger.txt reverses them. The paper and script were followed, a precedence decided before modelling; the contradiction stays visible.
Continuous cursor tracking · a negative result
Offline, same-person, no target-session adaptation.
People used motor imagery to steer a cursor after a target moving continuously across a screen, in two cohorts the publisher calls Main and Transfer Learning. For each admitted record, a fixed spectral ridge is trained on every eligible non-chance run of the earliest complete recorded session and tested on the latest, with no label from it, against a source-mean comparator: the earliest session’s average response, repeated for every row. Can the decoder beat that constant in the later session? Here, it did not.
Cohort records admitted
Main 9 of 14, Transfer Learning 14 of 14. Records, not proven unique people.
Target rows per response arm
The same rows for both arms: two outcomes do not double the independent sample. No response row was removed.
Response fits
One per admitted record and response arm; every one succeeded.
The error is the joint source-SD-normalized RMSE: errors divided by the earliest session’s response standard deviation, averaged over both axes within each trial before the square root. Lower is better. It is an error, not a percentage or an accuracy. Each record averages its trials equally, and each cohort mean weights records equally; intervals resample records within a cohort.
Two response arms, each with its own table. They are different outcomes on the same rows: never combined, and their errors are not compared with each other.
The output of the decoder in use at recording time, as the publisher stored it, imitated: not intended hand motion or intended control.
| Cohort | Spectral ridge · mean | Spectral ridge · median | Source-mean comparator · mean | Ridge minus comparator · same records |
|---|---|---|---|---|
| Main cohort9 of 14 records admitted | 192.29095% interval 1.292 to 574.235 | 1.317the mean is 146 times the median | 1.26695% interval 1.214 to 1.316 | 191.02495% interval 0.026 to 572.995ridge error higher on 9 of 9 records |
| Transfer Learning cohort14 of 14 records admitted | 44.17295% interval 0.955 to 130.476 | 0.990the mean is 45 times the median | 0.86495% interval 0.801 to 0.929 | 43.30895% interval 0.105 to 129.601ridge error higher on 14 of 14 records |
Target position minus cursor position in the publisher’s screen coordinates, a separately fitted proxy.
| Cohort | Spectral ridge · mean | Spectral ridge · median | Source-mean comparator · mean | Ridge minus comparator · same records |
|---|---|---|---|---|
| Main cohort9 of 14 records admitted | 698.49995% interval 1.166 to 2,093 | 1.225the mean is 570 times the median | 1.05695% interval 0.977 to 1.133 | 697.44395% interval 0.121 to 2,092ridge error higher on 9 of 9 records |
| Transfer Learning cohort14 of 14 records admitted | 69.84495% interval 1.071 to 207.287 | 1.182the mean is 59 times the median | 0.91695% interval 0.841 to 0.987 | 68.92895% interval 0.143 to 206.410ridge error higher on 14 of 14 records |
Across the four cohort-and-arm cells, the ridge’s mean error is 45 to 570 times its median: a few records carry very large errors. No scored record was removed, clipped or winsorized, and what caused the extreme errors — feature distribution shift, scaling sensitivity or something else — is not established. The every-record statement is about the overall primary error; it does not extend to the secondary metrics or to each recorded-decoder context.
Coverage. Main admits 9 of its 14 candidate records and is read conditional on that subset; 5 were held on metadata before any scoring and stay recorded: 2 because a session lacks its required Chance R01 record; 2 because the decoder-local run allocation differs from the documented contract; 1 because the sample-boundary geometry differs; its exact numeric cause was not independently reconstructed. Transfer Learning admits 14 of 14. That is 23 admitted cohort records, not 23 proven unique people: the cohorts may share people, and they are never pooled.
Transfer Learning is the publisher’s name for how that cohort’s data were collected; no transfer-learning model was trained here.
Record-equal means with their 95% intervals. Raw RMSE is in each response’s stored units; R² and Pearson are unitless. A constant prediction has no correlation, so the comparator’s Pearson values are undefined: null, not zero. The every-record statement above does not extend to these metrics, and their paired differences are not published.
| Cohort and model | Raw RMSE, x | Raw RMSE, y | R², x | R², y | Pearson, x | Pearson, y |
|---|---|---|---|---|---|---|
| Main cohortSpectral ridge | 38.88995% interval 0.309 to 116.038 | 50.59995% interval 0.312 to 151.159 | −172,98395% interval −518,950 to −0.018 | −296,66395% interval −889,988 to −0.046 | 0.02995% interval 0.00233 to 0.061 | 0.03495% interval 0.00105 to 0.073 |
| Main cohortSource-mean comparator | 0.30695% interval 0.290 to 0.321 | 0.30295% interval 0.289 to 0.313 | −0.00049895% interval −0.000786 to −0.000255 | −0.00037795% interval −0.000659 to −0.000127 | undefined (null) | undefined (null) |
| Transfer Learning cohortSpectral ridge | 4.44395% interval 0.212 to 12.877 | 13.81595% interval 0.229 to 40.954 | −5,44095% interval −16,319 to −0.188 | −58,11795% interval −174,350 to −0.339 | 0.04595% interval 0.00527 to 0.096 | 0.01995% interval −0.022 to 0.075 |
| Transfer Learning cohortSource-mean comparator | 0.19895% interval 0.184 to 0.212 | 0.20195% interval 0.187 to 0.217 | −0.00018795% interval −0.000314 to −0.0000754 | −0.00093795% interval −0.00140 to −0.000531 | undefined (null) | undefined (null) |
| Cohort and model | Raw RMSE, x | Raw RMSE, y | R², x | R², y | Pearson, x | Pearson, y |
|---|---|---|---|---|---|---|
| Main cohortSpectral ridge | 124.82895% interval 0.412 to 373.616 | 332.15295% interval 0.431 to 995.534 | −928,35095% interval −2,785,050 to −0.183 | −6,442,34195% interval −19,327,021 to −0.388 | 0.05095% interval 0.031 to 0.074 | 0.07295% interval 0.018 to 0.137 |
| Main cohortSource-mean comparator | 0.38895% interval 0.366 to 0.413 | 0.37695% interval 0.349 to 0.404 | −0.02695% interval −0.045 to −0.00771 | −0.02195% interval −0.039 to −0.00612 | undefined (null) | undefined (null) |
| Transfer Learning cohortSpectral ridge | 9.77295% interval 0.415 to 28.423 | 38.15195% interval 0.445 to 113.515 | −21,50495% interval −64,511 to −0.318 | −337,64095% interval −1,012,920 to −0.391 | 0.03595% interval 0.00847 to 0.067 | 0.03495% interval 0.00753 to 0.062 |
| Transfer Learning cohortSource-mean comparator | 0.36695% interval 0.327 to 0.402 | 0.38095% interval 0.341 to 0.414 | −0.0066495% interval −0.010 to −0.00352 | −0.01695% interval −0.030 to −0.00526 | undefined (null) | undefined (null) |
Methods & limits
The publisher’s processed derivative: 58 anonymous channel indices at 250 Hz in four-second epochs. Welch relative band power in 4–8, 8–13, 13–30 and 30–40 Hz per channel, as log ratios: 232 values per trial. A StandardScaler fitted on session 1 only, then a RidgeClassifier (alpha 1, no class weights). No hyperparameter search, target normalization, calibration or abstention.
The same trials, resampled to 200 Hz and zero-padded to 800 samples; each trial and channel divided by its own standard deviation, a dimensionless input, not the checkpoint’s native amplitude, so upstream benchmark numbers are not reproduced. CBraMod in evaluation mode with every weight frozen; its four temporal patch embeddings averaged into 11,600 values per trial; a separate StandardScaler and RidgeClassifier fitted on each person’s session 1 only. No backbone training, LoRA or target-session calibration.
57 EEG channels, −200 to +700 ms around each event, baseline-normalized; six 100 ms averages from +100 to +700 ms, 342 values per event. Per person, first-visit blocks 1–3 fit a fixed L2 logistic regression and block 4 fits Platt calibration; later visits are scored on blocks 2–4. Group B is not used. No hyperparameter search, target-label calibration, fine-tuning or LoRA.
62 channels at 1,000 Hz, already band-passed and notch-filtered by the publisher. Five bands from 4 to 30 Hz through a per-trial FIR filter, the preceding second of band power as channel-relative log power: 310 values per row, from a trial-contained 2-second past window, every 40 ms. Per response arm, a StandardScaler and response scaling fitted on the earliest session only, then Ridge (alpha 1). No tuning, LoRA, target-session calibration or foundation-model inference.
Every interval is a pointwise 95% percentile interval from 10,000 participant (or, for cursor tracking, record) bootstrap resamples within its cohort, with paired draws for every difference. Each is conditional on the fixed trained models and predictions: refitting, preprocessing, model and protocol selection — and, for CBraMod, pretraining overlap — are not included, and there is no multiple-comparison adjustment.
Per-person and per-record values of any kind — every minimum and percentile, and every median other than the four cursor-tracking ridge medians printed beside their means; which people declined or which records did worse; WBCIC-SHU confusion matrices; the cursor-tracking results split by recorded decoder; measured compute; and any participant, record or file identifier.
Banghua Yang, Fenqi Rong, Yunlong Xie, Du Li, Jiayang Zhang, Fu Li, Guangming Shi and Xiaorong Gao · A multi-day and high-quality EEG dataset for motor imagery brain-computer interface, Scientific Data 12, 488 (2025), doi:10.1038/s41597-025-04826-y. Data: Banghua Yang and Fenqi Rong, WBCIC-SHU Motor Imagery Dataset, Figshare+, doi:10.25452/figshare.plus.22671172.v5, CC BY 4.0. Derived analysis by BCI Report; not endorsed by the authors.
Ethics and consent, as the paper states them: Approved by the Tsinghua University Medical Ethics Committee (approval number 20190002); carried out in line with the Declaration of Helsinki. Written informed consent was obtained from the participants after they were informed about the procedure, purpose, requirements and motor-imagery techniques.
Dataset record ↗ · Paper ↗ · CC BY 4.0 · Read from the paper’s full text ↗
Chen Yang, Yufeng Zhang, Hongxin Zhang, Yixuan Li, Yijun Wang and Xiaorong Gao · A longitudinal EEG dataset of event-related potential, figshare, doi:10.6084/m9.figshare.27201003.v1, CC0. Paper: Yufeng Zhang, Hongxin Zhang, Yixuan Li, Yijun Wang, Xiaorong Gao and Chen Yang, A longitudinal EEG dataset of event-related potential, Scientific Data 12, 1069 (2025), doi:10.1038/s41597-025-05378-x. Derived analysis by BCI Report; not endorsed by the authors.
Ethics and consent, as the paper states them: Approved by the Tsinghua Institutional Review Board (No. 20230058). Each participant read and signed an informed consent form before the experiment.
Dataset record ↗ · Paper ↗ · CC0-1.0 · Read from the paper’s full text ↗
Dylan Forenzo and Bin He · EEG-BCI Dataset for “Continuous Tracking using Deep Learning-based Decoding for Non-invasive Brain-Computer Interface”, KiltHub, Carnegie Mellon University, doi:10.1184/R1/25360300.v1, CC BY 4.0. Citation the record requests: Dylan Forenzo, Hao Zhu, Jenn Shanahan, Jaehyun Lim and Bin He, Continuous tracking using deep learning-based decoding for noninvasive brain–computer interface, PNAS Nexus 3(4), pgae145 (2024), doi:10.1093/pnasnexus/pgae145. Derived analysis by BCI Report; not endorsed by the authors and not a reproduction of the paper's online results.
Ethics and consent, as the paper states them: Approved by the Institutional Review Board at Carnegie Mellon University; no approval number is given. Each subject provided written consent to the protocol before participating.
Dataset record ↗ · Paper ↗ · CC BY 4.0 · Read from the paper’s full text ↗
CBraMod, the frozen encoder: its paper and the pinned checkpoint’s SHA-256 0792cb808c14… (revision b9e961003214…). Its pretraining exposure for WBCIC-SHU is not established: the model paper describes pretraining on the TUH EEG corpus and the pinned repository lists SHU as a downstream task, and neither gives a checkpoint-specific file manifest.
Derived analyses by BCI Report, aggregates only; not endorsed by the data’s authors, and no raw data are redistributed.
Data source: later-sessions-update.json · schema bci-report-later-sessions-update-v1.
BCI Report (2026). Does a decoder trained on an earlier session still work later? https://bci.report/topics/later-sessions/
Figures from release later-sessions-update-20261008 (2026-10-08). Cite the upstream datasets as well: their credits are on this page.
Every release is archived on Zenodo: doi:10.5281/zenodo.23123296. BibTeX for the site and its releases → · CITATION.cff ↗