LaBraM
5.82M parameters
A fixed configuration has been evaluated. See each task for channels and training mode.
Open EEG evaluation · Snapshot 2026-09-20
The core matrix covers 9 decoding methods under 8 fixed protocols on 7 public datasets. Four deployment topics add 82 measurements on sensors, movement, calibration and pretraining.
Deployment questions
82 additional aggregate measurements, organized by the decision they can inform. They are repeated conditions within protocols, not 82 independent experiments and not an overall ranking.
What changes when a decoder crosses between two native eight-channel recordings from the same 102 people?
2 s SSVEP · 12 targets · balanced accuracyRead the evidence →02Motion robustnessStanding, walking and running results, with scalp and ear recordings and incompatible time windows kept apart.
SSVEP balanced accuracy · ERP ROC AUCRead the evidence →03Adaptation budgetTwelve, 24 or 48 labeled target trials help some methods more than others—and trial count is not elapsed time.
Common future blocks · target-only fittingRead the evidence →04Representation controlsMatched pretrained and constructor-random encoders under fixed and train-selected readout settings.
Two tasks · two encoders · three random initializationsRead the evidence →Full-precision proportions, paired contrasts, descriptive intervals, citations and five three-seed sensitivity groups.
Download reviewed topic data · JSON ↓Core benchmark matrix
The original eight-protocol snapshot, separate from the deployment topics above. Blank cells are protocols a method has not been run on — not failures.
| Method | Motor imagery & restn=10 · chance 50% | Idle & commandn=4 · detection | SSVEP · 8 channelsn=70 · chance 2.5% | SSVEP · 4 channelsn=70 · chance 2.5% | Arithmetic & restn=36 · chance 50% | P300 target ERPn=21 · chance 50% | Semantic target ERPn=30 · chance 50% | Sleep stagingn=20 · chance 20% |
|---|---|---|---|---|---|---|---|---|
| CBraModFoundation model · 8 of 8 | 62.9 | 18.3 | 33.7 | 27.6 | 62.3 | 55.0 | 55.8 | 71.3 |
| EEGNetCompact model · 8 of 8 | 71.0 | 1.7 | 55.8 | 44.1 | 67.6 | 53.3 | 61.4 | 56.5 |
| LaBraMFoundation model · 8 of 8 | 53.5 | 18.3 | 10.8 | 12.9 | 64.6 | 49.4≤ chance | 53.2 | 71.1 |
| Spectral ridgeClassical method · 7 of 8 | 54.9 | · | 50.7 | 48.2 | 56.8 | 51.0 | 51.0 | 72.2 |
| CSP+LDAClassical method · 2 of 8 | 47.1≤ chance | 23.3 | · | · | · | · | · | · |
| Standard CCAClassical method · 2 of 8 | · | · | 63.1 | 57.6 | · | · | · | · |
| Temporal ridgeClassical method · 2 of 8 | · | · | · | · | · | 53.8 | 56.3 | · |
| Deep4NetCompact model · 1 of 8 | · | 3.3 | · | · | · | · | · | · |
| ShallowFBCSPNetCompact model · 1 of 8 | · | 43.3 | · | · | · | · | · | · |
Best on that protocolBar = position between that protocol's chance level and 100%Not evaluatedCore matrix · JSON ↓
Bars are comparable down a column and not across one: chance level is 50% for the two-class tasks, 20% for sleep staging and 2.5% for 40-class SSVEP, and every protocol uses a different cohort and electrode layout. There is no overall score. The idle column reports command detection only — read it with the false-activation rate in its protocol.
One protocol at a time
Cohort, electrode layout, training budget, source terms and known limitations, alongside the numbers they belong to.
ds003810 · 10 subjects · 1,200 epochs · 5 participant-disjoint folds
| Model / configuration | Balanced accuracy | Macro F1 | Compute time |
|---|---|---|---|
| Supervised fit · 15 ch | 54.9%Descriptive 95% interval: 53.2–56.7% | 0.544Mean across held-out participants | 0.2s |
| Supervised fit · 15 ch | 47.1%Descriptive 95% interval: 43.8–50.2%At or below the 50% chance level | 0.406Mean across held-out participants | 8.8s |
| Frozen encoder + ridge head · 15 ch | 53.5%Descriptive 95% interval: 51.4–55.4% | 0.509Mean across held-out participants | 2.2s |
| Frozen encoder + ridge head · 15 ch | 62.9%Descriptive 95% interval: 59.6–66.5% | 0.617Mean across held-out participants | 5.0s |
| Scratch · 20 epochs · 15 ch | 71.0%Descriptive 95% interval: 65.0–76.3%Single seed — the highest of 3 run (68.17–71.00%, mean 69.47%) | 0.680Mean across held-out participants | 55.5s |
Scroll the table sideways for the remaining columns.
Fixed budgets and adapters; no universal ranking. Scoring time includes fitting and prediction, and may include accelerator waiting. It excludes data preparation.
Ten-person laboratory task with prompted rest. This is not continuous-idle monitoring or a physical low-channel headset test.
Methods →Model directory
Availability and measured performance are separate. Parameter counts depend on the backbone and task configuration.
5.82M parameters
A fixed configuration has been evaluated. See each task for channels and training mode.
25.29M parameters
Evaluated locally, but no score is published in this release: the checkpoint license is unresolved. Results are withheld pending that review, not because the model failed to run.
4.88M parameters
A fixed configuration has been evaluated. See each task for channels and training mode.
1.9K parameters
A fixed configuration has been evaluated. See each task for channels and training mode.
30.5K parameters
A fixed configuration has been evaluated. See each task for channels and training mode.
274.6K parameters
A fixed configuration has been evaluated. See each task for channels and training mode.
— parameters
A fixed configuration has been evaluated. See each task for channels and training mode.
3.19M parameters
Pretrained bipolar tokens do not directly match the unipolar recordings. A separately validated adapter is required.
— parameters
Code and electrode positions are accessible. Base weights require accepting the official access agreement; they have not been downloaded.
3.46M parameters
Public, ungated weights; research candidate only; not downloaded or executed for this release. The official model card provides a 13,847,488-byte safetensors checkpoint trained on Lee2019 at 128 Hz. All 17 ds005342 channel names match its 62-channel pretraining layout, but a separate locked adaptation protocol is still required.
157.00M parameters
Public, ungated weights; research candidate only; not downloaded or executed in this expansion. The Braindecode checkpoint is 628,580,476 bytes and expects 20 channels at 250 Hz. A documented montage or input adapter is required for 17-channel data.
— parameters
Public source and an official Google Drive weight link are listed; anonymous weight download, file size, and execution remain unverified. The backbone accepts caller-supplied brain regions and channel ordering, while the published BCIC head is fixed to 22 channels. A new adapter and a checkpoint-key loading audit are required.
— parameters
Public code and Hugging Face weights; research candidate only; not downloaded or executed for this release. The model page shows an approximately 318 MB checkpoint. The published EEGConformer input path is fixed to 22 channels, so using 17 channels requires an explicit mapping or replacement of the spatial layer.
— parameters
Public source; no official pretrained checkpoint was found; not executed for this release. This is a from-scratch baseline rather than a foundation-model checkpoint. The author script fixes the spatial convolution to 22 channels and must be parameterized for other montages.
— parameters
Public implementation; no pretrained weights; not executed for this release. The maintained Braindecode implementation accepts the channel count and sampling rate directly, making it a practical lightweight from-scratch comparator.
— parameters
Fixed reference method; inspect each task for input features and fit protocol.
— parameters
Fixed reference method; inspect each task for input features and fit protocol.
— parameters
Fixed reference method; inspect each task for input features and fit protocol.
Public data
Source terms are checked separately from numerical results. Unresolved data remain outside this release.
Clear CC0 snapshot plus dataset-specific ethics and signed-consent evidence; aggregate-only publication materially limits privacy exposure.
Attribution license and study-specific approval/consent are documented; eligibility assumes aggregate-only output and exclusion of subject-info fields.
Clear upstream CC0 and study ethics evidence; use only the pinned scalp-signal derivative and publish aggregates, with provenance disclosed.
Eligible only under the stated personal, research-led, noncommercial operation. Any ads, sponsorship, fees, or use directed toward commercial advantage requires fresh review or permission.
Small clinical cohort and linked health attributes create disproportionate reidentification risk; aggregate-only output does not resolve the missing secondary-publication consent evidence.
License permits only noncommercial unadapted redistribution, and the consent/ethics chain for public secondary results remains unverified.
Ethics evidence is positive, but upstream license and the processed mirror's authority to apply CC BY remain unresolved.
Strong study-specific open-sharing evidence. Aggregate metrics avoid redistribution of stimulus text and participant metadata; credit under CC BY.
Pinned CC0 release with dataset-specific IRB and consent evidence and comparatively low metadata exposure.
Reuse license is clear, but the dataset-specific human-subject consent and ethics chain is not evidenced strongly enough for a new public benchmark release.
Per-dataset license is verified, but consent evidence and the application of ND to the planned publication pipeline should be clarified.
Per-dataset license is verified, but consent evidence and the application of ND to the planned publication pipeline should be clarified.
Eligible only as an aggregate-only case study: remove all participant IDs/results and condition breakdowns, label n=4 prominently, avoid subgroup claims, and make no claim that the cohort is anonymous.
Field notes
Curated research updates, linked to original sources.
Researchers surveyed 55 representative models and evaluated 12 open-source foundation models with task-specific baselines across 13 datasets and nine BCI paradigms. The study reports that linear probing is often insufficient and that larger models do not automatically generalize better, reinforcing the need for fixed cross-subject and few-shot calibration protocols.
Original arXiv paperREVE reports pretraining on 92 datasets, about 60,000 hours of EEG, and 25,000 participants, with position encodings designed for different electrode arrangements. Its public code, weights, and tutorials make cross-montage transfer an auditable evaluation target. Access to the Base checkpoint is gated; it is not yet available in our local test pool.
Original arXiv paper and NeurIPS 2025 paper pageNeurIPT models homogeneous and heterogeneous spatiotemporal relationships across BCI tasks and electrode layouts and appeared as a NeurIPS 2025 poster. Its official repository still says that code restoration is in progress, so it is a research update rather than a reproducible model card.
OpenReview / NeurIPS 2025CSBrain combines local temporal windows, brain regions, and structured sparse attention, with reported experiments spanning 16 datasets and 11 task types. The official repository provides pretraining and downstream scripts plus a public weight entry point, making it a candidate for a later unified-protocol evaluation.
Original arXiv paper and NeurIPS 2025 paper pageThis study applies HuBERT-style self-supervised learning to eight-channel scalp EEG and examines motor imagery, P300, subject differences, and alpha rhythms. It motivates a separate deployment-oriented track for low channel counts, limited preprocessing, and runtime budgets.
Original arXiv paperEvidence before ranking
Only explicitly reviewed cohort-level research results are exported. Raw EEG, individual results and model weights stay outside the website. Each comparison states its source, license, protocol and limitations. No cross-task overall score.
Download all results · JSON ↓Downloads, loading checks and scored evaluations have distinct status labels. Models awaiting an adapter have no result.
Report missed commands, false activations and abstention. Zero errors in a short session do not establish all-day reliability.
Splits, channels, preprocessing and training modes travel with each protocol. Frozen encoders and scratch baselines are labeled separately.
Experiments run offline on a Mac or local GPU workstation. Timings are configuration-specific. This website performs no EEG inference or diagnosis.