BBCI Report Research preview

Open EEG evaluation · Snapshot 2026-09-20

Every EEG score, with the protocol that produced it.

The core matrix covers 9 decoding methods under 8 fixed protocols on 7 public datasets. Four deployment topics add 82 measurements on sensors, movement, calibration and pretraining.

Protocols
8
Datasets
7
Comparisons
39
Methods
9
Research preview

Deployment questions

Four ways to read the new evidence

82 additional aggregate measurements, organized by the decision they can inform. They are repeated conditions within protocols, not 82 independent experiments and not an overall ranking.

01Sensor transfer

Dry vs. wet electrodes

What changes when a decoder crosses between two native eight-channel recordings from the same 102 people?

2 s SSVEP · 12 targets · balanced accuracyRead the evidence →
02Motion robustness

On the move

Standing, walking and running results, with scalp and ear recordings and incompatible time windows kept apart.

SSVEP balanced accuracy · ERP ROC AUCRead the evidence →
03Adaptation budget

How much calibration?

Twelve, 24 or 48 labeled target trials help some methods more than others—and trial count is not elapsed time.

Common future blocks · target-only fittingRead the evidence →
04Representation controls

Does pretraining help?

Matched pretrained and constructor-random encoders under fixed and train-selected readout settings.

Two tasks · two encoders · three random initializationsRead the evidence →

Full-precision proportions, paired contrasts, descriptive intervals, citations and five three-seed sensitivity groups.

Download reviewed topic data · JSON ↓

Core benchmark matrix

9 methods × 8 protocols

The original eight-protocol snapshot, separate from the deployment topics above. Blank cells are protocols a method has not been run on — not failures.

Balanced accuracy of each method on each protocol. Chance level differs by protocol and is given in each column heading.
MethodMotor imagery & restn=10 · chance 50%Idle & commandn=4 · detectionSSVEP · 8 channelsn=70 · chance 2.5%SSVEP · 4 channelsn=70 · chance 2.5%Arithmetic & restn=36 · chance 50%P300 target ERPn=21 · chance 50%Semantic target ERPn=30 · chance 50%Sleep stagingn=20 · chance 20%
CBraModFoundation model · 8 of 862.918.333.727.662.355.055.871.3
EEGNetCompact model · 8 of 871.01.755.844.167.653.361.456.5
LaBraMFoundation model · 8 of 853.518.310.812.964.649.4≤ chance53.271.1
Spectral ridgeClassical method · 7 of 854.9·50.748.256.851.051.072.2
CSP+LDAClassical method · 2 of 847.1≤ chance23.3······
Standard CCAClassical method · 2 of 8··63.157.6····
Temporal ridgeClassical method · 2 of 8·····53.856.3·
Deep4NetCompact model · 1 of 8·3.3······
ShallowFBCSPNetCompact model · 1 of 8·43.3······

Best on that protocolBar = position between that protocol's chance level and 100%Not evaluatedCore matrix · JSON ↓

Bars are comparable down a column and not across one: chance level is 50% for the two-class tasks, 20% for sleep staging and 2.5% for 40-class SSVEP, and every protocol uses a different cohort and electrode layout. There is no overall score. The idle column reports command detection only — read it with the false-activation rate in its protocol.

One protocol at a time

The evidence behind each score

Cohort, electrode layout, training budget, source terms and known limitations, alongside the numbers they belong to.

Rest versus right-hand imagery

ds003810 · 10 subjects · 1,200 epochs · 5 participant-disjoint folds

Measured here

Results at a glance

CSV ↓
Model / configurationBalanced accuracyMacro F1Compute time
Supervised fit · 15 ch54.9%Descriptive 95% interval: 53.2–56.7%0.544Mean across held-out participants0.2s
Supervised fit · 15 ch47.1%Descriptive 95% interval: 43.8–50.2%At or below the 50% chance level0.406Mean across held-out participants8.8s
Frozen encoder + ridge head · 15 ch53.5%Descriptive 95% interval: 51.4–55.4%0.509Mean across held-out participants2.2s
Frozen encoder + ridge head · 15 ch62.9%Descriptive 95% interval: 59.6–66.5%0.617Mean across held-out participants5.0s
Scratch · 20 epochs · 15 ch71.0%Descriptive 95% interval: 65.0–76.3%Single seed — the highest of 3 run (68.17–71.00%, mean 69.47%)0.680Mean across held-out participants55.5s

Scroll the table sideways for the remaining columns.

Fixed budgets and adapters; no universal ranking. Scoring time includes fitting and prediction, and may include accelerator waiting. It excludes data preparation.

Read with care

Ten-person laboratory task with prompted rest. This is not continuous-idle monitoring or a physical low-channel headset test.

Methods →

Model directory

From compact baselines to foundation models

Availability and measured performance are separate. Parameter counts depend on the backbone and task configuration.

Foundation modelEvaluated

LaBraM

5.82M parameters

A fixed configuration has been evaluated. See each task for channels and training mode.

Foundation modelRights review pending

EEGPT

25.29M parameters

Evaluated locally, but no score is published in this release: the checkpoint license is unresolved. Results are withheld pending that review, not because the model failed to run.

Foundation modelEvaluated

CBraMod

4.88M parameters

A fixed configuration has been evaluated. See each task for channels and training mode.

Compact modelEvaluated

EEGNet

1.9K parameters

A fixed configuration has been evaluated. See each task for channels and training mode.

Compact modelEvaluated

ShallowFBCSPNet

30.5K parameters

A fixed configuration has been evaluated. See each task for channels and training mode.

Compact modelEvaluated

Deep4Net

274.6K parameters

A fixed configuration has been evaluated. See each task for channels and training mode.

Classical methodEvaluated

CSP+LDA

parameters

A fixed configuration has been evaluated. See each task for channels and training mode.

Foundation modelAdapter needed

BIOT

3.19M parameters

Pretrained bipolar tokens do not directly match the unipolar recordings. A separately validated adapter is required.

Foundation modelAccess gated

REVE Base

parameters

Code and electrode positions are accessible. Base weights require accepting the official access agreement; they have not been downloaded.

Foundation modelResearch candidate

Signal-JEPA

3.46M parameters

Public, ungated weights; research candidate only; not downloaded or executed for this release. The official model card provides a 13,847,488-byte safetensors checkpoint trained on Lee2019 at 128 Hz. All 17 ds005342 channel names match its 62-channel pretraining layout, but a separate locked adaptation protocol is still required.

Foundation modelResearch candidate

BENDR

157.00M parameters

Public, ungated weights; research candidate only; not downloaded or executed in this expansion. The Braindecode checkpoint is 628,580,476 bytes and expects 20 channels at 250 Hz. A documented montage or input adapter is required for 17-channel data.

Foundation modelAccess unverified

CSBrain

parameters

Public source and an official Google Drive weight link are listed; anonymous weight download, file size, and execution remain unverified. The backbone accepts caller-supplied brain regions and channel ordering, while the published BCIC head is fixed to 22 channels. A new adapter and a checkpoint-key loading audit are required.

Foundation modelResearch candidate

Neuro-GPT

parameters

Public code and Hugging Face weights; research candidate only; not downloaded or executed for this release. The model page shows an approximately 318 MB checkpoint. The published EEGConformer input path is fixed to 22 channels, so using 17 channels requires an explicit mapping or replacement of the spatial layer.

Compact modelResearch candidate

EEG-Conformer

parameters

Public source; no official pretrained checkpoint was found; not executed for this release. This is a from-scratch baseline rather than a foundation-model checkpoint. The author script fixes the spatial convolution to 22 channels and must be parameterized for other montages.

Compact modelResearch candidate

EEGSimpleConv

parameters

Public implementation; no pretrained weights; not executed for this release. The maintained Braindecode implementation accepts the channel count and sampling rate directly, making it a practical lightweight from-scratch comparator.

Classical methodEvaluated

Standard CCA

parameters

Fixed reference method; inspect each task for input features and fit protocol.

Classical methodEvaluated

Spectral ridge

parameters

Fixed reference method; inspect each task for input features and fit protocol.

Classical methodEvaluated

Temporal ridge

parameters

Fixed reference method; inspect each task for input features and fit protocol.

Public data

Public data, documented experiments

Source terms are checked separately from numerical results. Unresolved data remain outside this release.

Motor imagery / rest

ds003810 ↗

Clear CC0 snapshot plus dataset-specific ethics and signed-consent evidence; aggregate-only publication materially limits privacy exposure.

Subjects
10
Channels
15
Aggregate resultsCC0-1.0 ↗Peterson et al. · OpenNeuro ds003810, version 2.0.2. Study: https://pmc.ncbi.nlm.nih.gov/articles/PMC9114495/
Arithmetic / rest

EEGMAT ↗

Attribution license and study-specific approval/consent are documented; eligibility assumes aggregate-only output and exclusion of subject-info fields.

Subjects
36
Channels
19
Aggregate resultsOpen Data Commons Attribution License 1.0 ↗Igor Zyma, Ivan Seleznov, Anton Popov, Mariia Chernykh, Oleksii Shpenkov · EEG During Mental Arithmetic Tasks 1.0.0, PhysioNet, doi:10.13026/C2JQ1P. Study: Zyma et al. (2019), doi:10.3390/data4010014. PhysioNet platform: Pollard et al. (2026), doi:10.1038/s44360-026-00096-z.
Five-stage sleep

EESM19 scalp subset ↗

Clear upstream CC0 and study ethics evidence; use only the pinned scalp-signal derivative and publish aggregates, with provenance disclosed.

Subjects
20
Channels
6
Aggregate resultsCC0-1.0 declared by upstream and mirror ↗Kaare B. Mikkelsen et al. · Accurate whole-night sleep monitoring with dry-contact ear-EEG (2019), doi:10.1038/s41598-019-53115-3; OpenNeuro ds005185 v1.0.2. Processed mirror: Zachary1150/EESM19-Processed.
40-target SSVEP

BETA ↗

Eligible only under the stated personal, research-led, noncommercial operation. Any ads, sponsorship, fees, or use directed toward commercial advantage requires fresh review or permission.

Subjects
70
Channels
4 / 8 selected
Aggregate resultsCC BY 4.0 ↗Bingchuan Liu et al. · BETA: A Large Benchmark Database Toward SSVEP-BCI Application (2020), doi:10.3389/fnins.2020.00627. Figshare 12264401 v3; mirror Bingchuan/BETA.
Research source

BNCI2014-008 ↗

Small clinical cohort and linked health attributes create disproportionate reidentification risk; aggregate-only output does not resolve the missing secondary-publication consent evidence.

Subjects
Channels
Not in this releaseCC BY-NC-ND 4.0 ↗BNCI Horizon 2020 catalog · Original researchers credited at source.
Research source

BNCI2014-009 ↗

License permits only noncommercial unadapted redistribution, and the consent/ethics chain for public secondary results remains unverified.

Subjects
Channels
Not in this releaseCC BY-NC-ND 4.0 ↗BNCI Horizon 2020 catalog · Original researchers credited at source.
Semantic target ERP

TMNRED / ds005383 ↗

Strong study-specific open-sharing evidence. Aggregate metrics avoid redistribution of stimulus text and participant metadata; credit under CC BY.

Subjects
30
Channels
30
Aggregate resultsOpenNeuro metadata says CC0; accompanying publication/GitHub says CC BY 4.0 ↗Yanru Bai, Qi Tang et al. · TMNRED, A Chinese Language EEG Dataset for Fuzzy Semantic Target Identification in Natural Reading Environments (2025), doi:10.1038/s41597-025-05036-2. OpenNeuro ds005383 v1.0.0.
P300 target ERP

ds006593 ↗

Pinned CC0 release with dataset-specific IRB and consent evidence and comparatively low metadata exposure.

Subjects
21
Channels
19
Aggregate resultsCC0-1.0 ↗OpenNeuro ds006593 contributors · version 1.0.0, doi:10.18112/openneuro.ds006593.v1.0.0; original author credits retained at the linked source.
Research source

physionet-eegmmidb-1.0.0 ↗

Reuse license is clear, but the dataset-specific human-subject consent and ethics chain is not evidenced strongly enough for a new public benchmark release.

Subjects
Channels
Not in this releaseOpen Data Commons Attribution License 1.0 ↗PhysioNet EEG Motor Movement/Imagery Dataset 1.0.0 · Original researchers credited at source.
Research source

BNCI2014-001 ↗

Per-dataset license is verified, but consent evidence and the application of ND to the planned publication pipeline should be clarified.

Subjects
Channels
Not in this releaseCC BY-ND 4.0 ↗BNCI Horizon 2020 catalog · Original researchers credited at source.
Research source

BNCI2014-004 ↗

Per-dataset license is verified, but consent evidence and the application of ND to the planned publication pipeline should be clarified.

Subjects
Channels
Not in this releaseCC BY-ND 4.0 ↗BNCI Horizon 2020 catalog · Original researchers credited at source.
Cue-gated idle / command

ds005342 ↗

Eligible only as an aggregate-only case study: remove all participant IDs/results and condition breakdowns, label n=4 prominently, avoid subgroup claims, and make no claim that the cohort is anonymous.

Subjects
4 evaluated / 32 prepared
Channels
17
Aggregate resultsCC0-1.0 ↗OpenNeuro ds005342 contributors · version 1.0.3, doi:10.18112/openneuro.ds005342.v1.0.3; associated study doi:10.3389/fninf.2022.961089. Original author credits are retained at the linked source.

Field notes

Notes from the field

Curated research updates, linked to original sources.

Benchmark

A new benchmark compares EEG foundation models with task-specific decoders ↗

Researchers surveyed 55 representative models and evaluated 12 open-source foundation models with task-specific baselines across 13 datasets and nine BCI paradigms. The study reports that linear probing is often insufficient and that larger models do not automatically generalize better, reinforcing the need for fixed cross-subject and few-shot calibration protocols.

Original arXiv paper
Model release

REVE releases an EEG foundation model for varying electrode layouts ↗

REVE reports pretraining on 92 datasets, about 60,000 hours of EEG, and 25,000 participants, with position encodings designed for different electrode arrangements. Its public code, weights, and tutorials make cross-montage transfer an auditable evaluation target. Access to the Base checkpoint is gated; it is not yet available in our local test pool.

Original arXiv paper and NeurIPS 2025 paper page
Model release

CSBrain adds cross-scale spatiotemporal structure to general EEG decoding ↗

CSBrain combines local temporal windows, brain regions, and structured sparse attention, with reported experiments spanning 16 datasets and 11 task types. The official repository provides pretraining and downstream scripts plus a public weight entry point, making it a candidate for a later unified-protocol evaluation.

Original arXiv paper and NeurIPS 2025 paper page

Evidence before ranking

A score is only useful with its conditions.

Only explicitly reviewed cohort-level research results are exported. Raw EEG, individual results and model weights stay outside the website. Each comparison states its source, license, protocol and limitations. No cross-task overall score.

Download all results · JSON ↓
01

A successful forward pass is not a benchmark

Downloads, loading checks and scored evaluations have distinct status labels. Models awaiting an adapter have no result.

02

Show the failure modes

Report missed commands, false activations and abstention. Zero errors in a short session do not establish all-day reliability.

03

Compare the same task

Splits, channels, preprocessing and training modes travel with each protocol. Frozen encoders and scratch baselines are labeled separately.

04

Start with fixed, audited local runs

Experiments run offline on a Mac or local GPU workstation. Timings are configuration-specific. This website performs no EEG inference or diagnosis.

Experiment details