# New people, next day: which part of a pretrained model should you update?

Model adaptation · new people; next day held

Which part of a pretrained encoder to update when it meets people it has not seen, on a task it is adapted to: the head only, the last block, or a low-rank adapter — with the same checkpoint, folds, starting heads, batches and five-epoch recipe for all three. The next-day version of the question has run too; its source is held, so it is listed without a figure.

- [Reviewed aggregate JSON ↓](https://bci.report/data/adaptation-update.json)
- 

Measured on: [EEGMAT](https://bci.report/datasets/eegmat/) · Methods: [LaBraM](https://bci.report/methods/labram/)

## Short answer

On new people, same task, zero labels from the test person — LaBraM Base on mental arithmetic versus rest, **36** people in five participant-disjoint folds — updating the last block with the head reached **65.7%** balanced accuracy and rank-4 LoRA **64.1%**, against **56.6%** for a short head-only fit; LoRA minus last block crosses zero, so neither is shown to be better. That head-only fit is not the best a frozen encoder can do: the core matrix's frozen LaBraM with a ridge head scored **64.6%** on the same people and folds, a different head and not a paired comparison. The next day is not answered here: that experiment ran, but its source is held, so no figure from it is published.

New people, same task, zero labels from the test person · reviewed 1 October 2026

## Adapting a foundation model: head only, last block or LoRA?

LaBraM Base on EEGMAT, mental arithmetic versus rest: 36 people in five participant-disjoint folds, so every score is on people the model was not adapted on, and no label from a test person is used. The three update rules start from the same head, see the same batches in the same order and train for the same five epochs. Three seeds are averaged within each person before the cohort mean.

Head only · encoder frozen

56.6%

Macro F1 54.6% · 402 trainable parameters

— 95% interval 54.2%–59.0%

Last block + head

65.7%

Macro F1 62.3% · 482,882 trainable parameters

— 95% interval 61.6%–69.7%

LoRA rank 4 + head

64.1%

Macro F1 60.5% · 38,802 trainable parameters

— 95% interval 60.2%–67.8%

For scale: the core matrix's frozen LaBraM, on the same people and folds but read out by a closed-form ridge head on standardized features, scored 64.6% (core matrix release, [experiments.json](https://bci.report/data/experiments.json)). The head-only arm here is a short five-epoch head on unscaled features — not the best a frozen encoder can do — and the two are not a paired comparison.

**+7.5 pp** LoRA minus head only, same people

### LoRA beat the short head-only fit; it did not beat the last block

30 of 36 people improved and 6 got worse; paired interval +4.8 to +10.1 pp. Updating the last block instead gained +9.1 pp (+6.0 to +12.2 pp). LoRA minus last block was −1.6 pp, interval −4.4 to +1.0 pp: it crosses zero, so neither is shown to be better. 17 people did better with LoRA, 19 with the last block.

Change in participant-mean score for the same 36 people under both rules · descriptive person-bootstrap 95% interval · people helped, harmed and unchanged.

| Comparison | Balanced accuracy | Macro F1 |
| --- | --- | --- |
| LoRA − head only | +7.5 pp (+4.8 to +10.1 pp) (30 helped · 6 harmed · 0 unchanged) | +5.9 pp (+2.7 to +9.0 pp) (22 helped · 14 harmed · 0 unchanged) |
| Last block − head only | +9.1 pp (+6.0 to +12.2 pp) (30 helped · 5 harmed · 1 unchanged) | +7.7 pp (+4.4 to +11.2 pp) (27 helped · 9 harmed · 0 unchanged) |
| LoRA − last block | −1.6 pp (−4.4 to +1.0 pp) (17 helped · 19 harmed · 0 unchanged) | −1.9 pp (−5.0 to +1.1 pp) (15 helped · 21 harmed · 0 unchanged) |

On macro F1 the gains are smaller, and LoRA helps fewer people over head only than it does on balanced accuracy: the metric matters.

### What each rule costs

Training time is the sum over 15 fits (5 folds × 3 seeds) on one Apple-silicon Mac, PyTorch MPS backend, excluding loading and evaluation.

| Update rule | Trainable parameters | Training time, 15 fits | Balanced accuracy, seed by seed |
| --- | --- | --- | --- |
| Head only · encoder frozen | 402 | 51.5 s | (56.9% · 56.4% · 56.5%) |
| Last block + head | 482,882 | 65.3 s | (65.7% · 65.7% · 65.7%) |
| LoRA rank 4 + head | 38,802 | 128.2 s | (64.4% · 65.1% · 62.7%) |

LoRA trains far fewer weights than the last block but took longer here: its adapters sit in all twelve blocks, so gradients flow back through most of the encoder, and this implementation is an effective-weight parametrization, not an optimized low-rank kernel. Fewer trainable weights do not by themselves mean less time or memory.

Promised on the roadmap and not delivered by this batch: peak memory (only lower-bound samples were recorded, so none is published) and a comparison at a fixed computation budget (every rule trained for the same five epochs, not the same compute).

Next day · status only

## Same person, the next day: run, not published

Whether a model fitted on one day still works on a later one, and which part to update when it does not, has no published figure here. Two sources were prepared for it; for different reasons, neither can be scored in public.

### Two recording days: run and replayed, held

BNCI2015-001: 12 people recorded on two days. Each person's model is trained on day A, calibrated with the first 0, 10, 20 or 40 labeled trials of day B under the same three update rules, and tested on the same later day-B trials. Unlike the result above, this design uses the test person's own labels: their day-A trials to train and, except at the zero budget, their first day-B trials to calibrate. It ran on 22 September and passed an independent replay. The source's editorial hold stands — its catalogue licence is CC BY-NC-ND and its description names no ethics approval — so no figure from it is published.

### Same person, later sessions: status only

Stieger longitudinal BCI: one person's 11 sessions were acquired, not the full release. What is published is that the loader and a chronological split work on it — earlier sessions train, later sessions test. With one person, every score would be that person's, so none is published; the source is listed as status only in the [release of 27 September](https://bci.report/releases/#context-update-20260927).

### Four questions kept apart

A new person, a new day, a new dataset and a new device are different questions, answered on different data. A task adapter trained on several people is not evidence that it transfers across devices.

A related question now has a measured answer, for two classical baselines on other data: the same person’s next session, with and without a few labeled trials from it. It is on [How much calibration data does an EEG decoder need?](https://bci.report/topics/calibration-budget/#next-session).

Methods & limits

## What this result is not

One fixed recipe, three update rules and three seeds, on new people for a task the model is adapted to. It describes that recipe on these people; it is not a ranking of update rules.

### One recipe, not a tuned ranking

Five epochs, AdamW at learning rate 1e-4, batch 32, with no search, validation split, early stopping or checkpoint selection. Another learning rate or training budget could reorder the rules.

### New people, not a new day or device

The folds hold out people on a task the model is adapted to. Nothing here measures a new recording day, a new headset or a new dataset.

### A known dataset

EEGMAT is public, and whether it was in LaBraM's pretraining data is not certified.

### Descriptive intervals

Person bootstrap after averaging the three seeds within each person. It ignores the dependence that shared cross-validation models create.

### Adapter size

The LaBraM base encoder has 5,819,936 parameters. Rank-4 updates to its twelve fused QKV weights add 38,400 trainable parameters (0.66%); with the 402-parameter head that is the 38,802 above. The adapter tensors take 150 KiB in float32 before metadata.

**Engineering checks passed before any real-EEG run:** exact adapter save and reload; exact merged-weight equivalence; finite, nonzero adapter gradients in every block; frozen encoder weights unchanged after updates; identical output at zero adapter initialization.

These are parameter counts — not accuracy, speed, memory-saving or privacy claims.

### Not the first to ask

Other groups already benchmark EEG adapters: [OpenEEGBench](https://github.com/braindecode/OpenEEGBench) covers parameter-efficient fine-tuning of EEG foundation models; a [systematic study of test-time adaptation](https://proceedings.mlr.press/v340/lee26a.html) reports inconsistent gains under distribution shift; a recent [preprint](https://arxiv.org/abs/2608.24727v1) compares self-supervised adaptation under fixed computation budgets. Our emphasis is the calibration cost, the cases that get worse, and whether recordings are compatible at all.

### EEGMAT · LaBraM adaptation, three matched update rules

Igor Zyma, Ivan Seleznov, Anton Popov, Mariia Chernykh, Oleksii Shpenkov · EEG During Mental Arithmetic Tasks 1.0.0, PhysioNet, doi:10.13026/C2JQ1P. Study: Zyma et al. (2019), doi:10.3390/data4010014. PhysioNet platform: Pollard et al. (2026), doi:10.1038/s44360-026-00096-z.

Open Data Commons Attribution License 1.0 · aggregate results only; the adaptation reuses the 2026-09-20 review of these recordings.

[Dataset record ↗](https://physionet.org/content/eegmat/1.0.0/)

### Stieger longitudinal BCI · one-person pilot

James R. Stieger, Stephen A. Engel and Bin He · Continuous sensorimotor rhythm based brain computer interface learning in a large population, Scientific Data 8, 98 (2021); data at doi:10.6084/m9.figshare.13123148.v1.

CC-BY-4.0 · status only: no figure from this source is published.

[Dataset record ↗](https://doi.org/10.6084/m9.figshare.13123148.v1)

Data source: [adaptation-update.json](https://bci.report/data/adaptation-update.json) · schema bci-report-adaptation-update-v1 · generated 2026-10-01.

Data source: [evidence-update.json](https://bci.report/data/evidence-update.json) · the adapter counts and checks, from the engineering check released on 22 September.

Keep exploring

Transfer

- [— Sensor transfer **Dry vs. wet electrodes**](https://bci.report/topics/dry-vs-wet/)
- [— Context transfer **Screen to VR**](https://bci.report/topics/screen-to-vr/)
- [— Montage **Fewer electrodes**](https://bci.report/topics/fewer-electrodes/)
- [— Motion robustness **On the move**](https://bci.report/topics/on-the-move/)

Adapting models

- [— Calibration budget **How much calibration?**](https://bci.report/topics/calibration-budget/)
- [— Model adaptation **Which part to update?**](https://bci.report/topics/model-adaptation/)
- [— Representation controls **Does pretraining help?**](https://bci.report/topics/does-pretraining-help/)

Reliability & clinical

- [— Abstention · Jev-style **When not to act**](https://bci.report/topics/when-not-to-act/)
- [— Sleep staging · simple baselines **Sleep-stage balance**](https://bci.report/topics/sleep-staging/)
- [— Clinical research **Clinical groups**](https://bci.report/topics/clinical-groups/)

[All questions, and the map of which kinds of transfer have been measured →](https://bci.report/topics/)

## Cite this page

BCI Report (2026). *New people, next day: which part of a pretrained model should you update?* https://bci.report/topics/model-adaptation/

Figures from releases [`adaptation-update-20261001`](https://bci.report/releases/#adaptation-update-20261001) (2026-10-01), [`evidence-update-20260922`](https://bci.report/releases/#evidence-update-20260922) (2026-09-22), [`research-preview-20260920`](https://bci.report/releases/#research-preview-20260920) (2026-09-20). Cite the upstream datasets as well: their credits are on this page.

Every release is archived on Zenodo: [doi:10.5281/zenodo.23123296](https://doi.org/10.5281/zenodo.23123296). [BibTeX for the site and its releases →](https://bci.report/api/#cite-heading) · [CITATION.cff ↗](https://raw.githubusercontent.com/twu3202/bci-report/main/CITATION.cff)

---
Markdown copy of https://bci.report/topics/model-adaptation/, generated from the published page. Figures are aggregate results; terms of use: https://bci.report/data-use/
