Head only · encoder frozen
Macro F1 54.6% · 402 trainable parameters
95% interval 54.2%–59.0%Model adaptation · new people; next day held
Which part of a pretrained encoder to update when it meets people it has not seen, on a task it is adapted to: the head only, the last block, or a low-rank adapter — with the same checkpoint, folds, starting heads, batches and five-epoch recipe for all three. The next-day version of the question has run too; its source is held, so it is listed without a figure.
Measured on: EEGMAT · Methods: LaBraM
On new people, same task, zero labels from the test person — LaBraM Base on mental arithmetic versus rest, 36 people in five participant-disjoint folds — updating the last block with the head reached 65.7% balanced accuracy and rank-4 LoRA 64.1%, against 56.6% for a short head-only fit; LoRA minus last block crosses zero, so neither is shown to be better. That head-only fit is not the best a frozen encoder can do: the core matrix's frozen LaBraM with a ridge head scored 64.6% on the same people and folds, a different head and not a paired comparison. The next day is not answered here: that experiment ran, but its source is held, so no figure from it is published.
New people, same task, zero labels from the test person · reviewed 1 October 2026
LaBraM Base on EEGMAT, mental arithmetic versus rest: 36 people in five participant-disjoint folds, so every score is on people the model was not adapted on, and no label from a test person is used. The three update rules start from the same head, see the same batches in the same order and train for the same five epochs. Three seeds are averaged within each person before the cohort mean.
Head only · encoder frozen
Macro F1 54.6% · 402 trainable parameters
95% interval 54.2%–59.0%Last block + head
Macro F1 62.3% · 482,882 trainable parameters
95% interval 61.6%–69.7%LoRA rank 4 + head
Macro F1 60.5% · 38,802 trainable parameters
95% interval 60.2%–67.8%For scale: the core matrix's frozen LaBraM, on the same people and folds but read out by a closed-form ridge head on standardized features, scored 64.6% (core matrix release, experiments.json). The head-only arm here is a short five-epoch head on unscaled features — not the best a frozen encoder can do — and the two are not a paired comparison.
30 of 36 people improved and 6 got worse; paired interval +4.8 to +10.1 pp. Updating the last block instead gained +9.1 pp (+6.0 to +12.2 pp). LoRA minus last block was −1.6 pp, interval −4.4 to +1.0 pp: it crosses zero, so neither is shown to be better. 17 people did better with LoRA, 19 with the last block.
| Comparison | Balanced accuracy | Macro F1 |
|---|---|---|
| LoRA − head only | +7.5 pp+4.8 to +10.1 pp30 helped · 6 harmed · 0 unchanged | +5.9 pp+2.7 to +9.0 pp22 helped · 14 harmed · 0 unchanged |
| Last block − head only | +9.1 pp+6.0 to +12.2 pp30 helped · 5 harmed · 1 unchanged | +7.7 pp+4.4 to +11.2 pp27 helped · 9 harmed · 0 unchanged |
| LoRA − last block | −1.6 pp−4.4 to +1.0 pp17 helped · 19 harmed · 0 unchanged | −1.9 pp−5.0 to +1.1 pp15 helped · 21 harmed · 0 unchanged |
On macro F1 the gains are smaller, and LoRA helps fewer people over head only than it does on balanced accuracy: the metric matters.
| Update rule | Trainable parameters | Training time, 15 fits | Balanced accuracy, seed by seed |
|---|---|---|---|
| Head only · encoder frozen | 402 | 51.5 s | 56.9% · 56.4% · 56.5% |
| Last block + head | 482,882 | 65.3 s | 65.7% · 65.7% · 65.7% |
| LoRA rank 4 + head | 38,802 | 128.2 s | 64.4% · 65.1% · 62.7% |
LoRA trains far fewer weights than the last block but took longer here: its adapters sit in all twelve blocks, so gradients flow back through most of the encoder, and this implementation is an effective-weight parametrization, not an optimized low-rank kernel. Fewer trainable weights do not by themselves mean less time or memory.
Promised on the roadmap and not delivered by this batch: peak memory (only lower-bound samples were recorded, so none is published) and a comparison at a fixed computation budget (every rule trained for the same five epochs, not the same compute).
Next day · status only
Whether a model fitted on one day still works on a later one, and which part to update when it does not, has no published figure here. Two sources were prepared for it; for different reasons, neither can be scored in public.
BNCI2015-001: 12 people recorded on two days. Each person's model is trained on day A, calibrated with the first 0, 10, 20 or 40 labeled trials of day B under the same three update rules, and tested on the same later day-B trials. Unlike the result above, this design uses the test person's own labels: their day-A trials to train and, except at the zero budget, their first day-B trials to calibrate. It ran on 22 September and passed an independent replay. The source's editorial hold stands — its catalogue licence is CC BY-NC-ND and its description names no ethics approval — so no figure from it is published.
Stieger longitudinal BCI: one person's 11 sessions were acquired, not the full release. What is published is that the loader and a chronological split work on it — earlier sessions train, later sessions test. With one person, every score would be that person's, so none is published; the source is listed as status only in the release of 27 September.
A new person, a new day, a new dataset and a new device are different questions, answered on different data. A task adapter trained on several people is not evidence that it transfers across devices.
A related question now has a measured answer, for two classical baselines on other data: the same person’s next session, with and without a few labeled trials from it. It is on How much calibration data does an EEG decoder need?.
Methods & limits
One fixed recipe, three update rules and three seeds, on new people for a task the model is adapted to. It describes that recipe on these people; it is not a ranking of update rules.
Five epochs, AdamW at learning rate 1e-4, batch 32, with no search, validation split, early stopping or checkpoint selection. Another learning rate or training budget could reorder the rules.
The folds hold out people on a task the model is adapted to. Nothing here measures a new recording day, a new headset or a new dataset.
EEGMAT is public, and whether it was in LaBraM's pretraining data is not certified.
Person bootstrap after averaging the three seeds within each person. It ignores the dependence that shared cross-validation models create.
The LaBraM base encoder has 5,819,936 parameters. Rank-4 updates to its twelve fused QKV weights add 38,400 trainable parameters (0.66%); with the 402-parameter head that is the 38,802 above. The adapter tensors take 150 KiB in float32 before metadata.
Engineering checks passed before any real-EEG run: exact adapter save and reload; exact merged-weight equivalence; finite, nonzero adapter gradients in every block; frozen encoder weights unchanged after updates; identical output at zero adapter initialization.
These are parameter counts — not accuracy, speed, memory-saving or privacy claims.
Other groups already benchmark EEG adapters: OpenEEGBench covers parameter-efficient fine-tuning of EEG foundation models; a systematic study of test-time adaptation reports inconsistent gains under distribution shift; a recent preprint compares self-supervised adaptation under fixed computation budgets. Our emphasis is the calibration cost, the cases that get worse, and whether recordings are compatible at all.
Igor Zyma, Ivan Seleznov, Anton Popov, Mariia Chernykh, Oleksii Shpenkov · EEG During Mental Arithmetic Tasks 1.0.0, PhysioNet, doi:10.13026/C2JQ1P. Study: Zyma et al. (2019), doi:10.3390/data4010014. PhysioNet platform: Pollard et al. (2026), doi:10.1038/s44360-026-00096-z.
Open Data Commons Attribution License 1.0 · aggregate results only; the adaptation reuses the 2026-09-20 review of these recordings.
James R. Stieger, Stephen A. Engel and Bin He · Continuous sensorimotor rhythm based brain computer interface learning in a large population, Scientific Data 8, 98 (2021); data at doi:10.6084/m9.figshare.13123148.v1.
CC-BY-4.0 · status only: no figure from this source is published.
Data source: adaptation-update.json · schema bci-report-adaptation-update-v1 · generated 2026-10-01.
Data source: evidence-update.json · the adapter counts and checks, from the engineering check released on 22 September.
BCI Report (2026). New people, next day: which part of a pretrained model should you update? https://bci.report/topics/model-adaptation/
Figures from releases adaptation-update-20261001 (2026-10-01), evidence-update-20260922 (2026-09-22), research-preview-20260920 (2026-09-20). Cite the upstream datasets as well: their credits are on this page.
Every release is archived on Zenodo: doi:10.5281/zenodo.23123296. BibTeX for the site and its releases → · CITATION.cff ↗