People
each recorded once on each display
Context transfer · P300 on a screen and in VR
21 people used the same P300 interface on a PC screen and in a virtual-reality headset. Calibrated on one display and tested on the other, a simple logistic regression reached 66.0% balanced accuracy, where chance is 50%. That is how well a calibration carries over. It is not how much the change costs: this protocol did not also test each person on the display they were calibrated on.
Interval 61.6%–70.2%; AUROC 0.732. Each person was calibrated on one display and tested on the other in both directions, and the two directions were averaged within the person. Without a same-display score for the same people, this number cannot say whether the display change lowered performance, or by how much.
Two fixed baselines
Target against non-target epochs, 5 non-targets to every target, so accuracy would reward always answering “non-target”; balanced accuracy and AUROC do not. Primary timing.
| Baseline | Balanced accuracy | AUROC |
|---|---|---|
| Mean-window logistic regression | 66.0%61.6%–70.2%66.0% balanced accuracy | 0.7320.671–0.785 |
| Spatiotemporal shrinkage LDA | 63.6%59.6%–67.7%63.6% balanced accuracy | 0.7130.652–0.771 |
Prespecified sensitivity
The headset and the screen report stimulus onset with different delays. The primary analysis shifts each epoch by the source-reported offset — 19 samples on the PC and 60 in VR, at 512 Hz. The sensitivity analysis leaves epochs at the recorded tag. Neither scheme was chosen from test scores.
| Baseline | Onset corrected (primary) | At the recorded tag |
|---|---|---|
| Mean-window logistic regression | 66.0%61.6%–70.2% | 63.8%61.0%–66.5% |
| Spatiotemporal shrinkage LDA | 63.6%59.6%–67.7% | 57.2%55.1%–59.4% |
Spatiotemporal shrinkage LDA depends on timing most: 63.6% with the correction, 57.2% without. A mean offset is not the same as correct timing, but a result reported without saying which timing it used is incomplete.
What the scores are computed on
Every label comes from the source; none was interpolated. One malformed row at the very end of one recording was discarded; it followed every possible epoch.
People
each recorded once on each display
Labelled events
1 target to 5 non-targets
Epochs kept, primary timing
139 rejected at 500 µV peak to peak
Methods & limits
A within-person, offline transfer test with two simple baselines. It narrows one question; it does not rank displays or headsets.
Nobody was tested on the display they were calibrated on in this protocol, so the cost of changing display is not measured here — only the level reached after the change.
Each person supplies both their calibration and their test data, on the same day. Nothing here speaks to a model that works for a new person or a later session.
Two simple CPU classifiers with no tuning. No foundation model and no fine-tuned model was run in this batch.
Correcting the average onset delay does not remove event-to-event jitter, which a headset can add.
Grégoire Cattan, Anton Andreev, Pedro L. C. Rodrigues and Marco Congedo · Dataset of an EEG-based BCI experiment in Virtual Reality and on a Personal Computer, Zenodo, doi:10.5281/zenodo.2605205; documentation arXiv:1903.11297.
All participants gave written informed consent covering the experimental process, the data management procedures and the right to withdraw at any moment; the study was approved by the ethics committee of the University of Grenoble Alpes (CERNI). Only cohort aggregates appear here.
Data source: reviewed aggregate JSON · schema bci-report-context-update-v1 · generated 2026-09-27.