Open BCI benchmark infrastructure

BCI-Bench

Benchmarking Generalizable and Deployment-Ready Brain-Computer Interfaces

An open evaluation platform for testing BCI algorithms across users, sessions, devices, calibration settings, and deployment constraints.

Website MVP in active development

Independent, standardized evaluation of whether BCI algorithms generalize beyond the datasets on which they were developed.

7benchmark tracks
4release stages
Opencommunity infrastructure

Why BCI-Bench?

From dataset scores to evidence of real generalization.

BCI research needs shared protocols that reveal how algorithms behave beyond familiar data and ideal laboratory conditions.

  1. 01
    Comparable evaluation

    Standardized preprocessing, splits, metrics, and reporting make results easier to interpret and reproduce.

  2. 02
    Independent verification

    Hidden test sets and controlled protocols reduce overfitting to public benchmark data.

  3. 03
    Deployment-oriented evidence

    Reliability, calibration cost, latency, and worst-subject behavior matter alongside average accuracy.

Core benchmark tracks

Evaluation built around the conditions that make BCI difficult.

Each track asks a focused scientific question and reports more than one aggregate score.

01

Cross-Subject Generalization

Can a model work on users it has never seen?

Evaluation focus

Held-out participants with no subject-specific training.

02

Cross-Session Stability

Does performance persist across days and neural drift?

Evaluation focus

Future sessions and longitudinal reliability.

03

Cross-Device Transfer

Can a method transfer across acquisition systems and sites?

Evaluation focus

Devices, montages, sampling rates, and centers.

04

Few-Shot Calibration

How much user-specific data is needed to become useful?

Evaluation focus

Zero-shot, 1-shot, 5-shot, and calibration curves.

05

Robustness and Reliability

How does a model fail under noisy or incomplete signals?

Evaluation focus

Missing channels, noise, and worst-subject outcomes.

06

Deployment Readiness

Can the method meet practical runtime constraints?

Evaluation focus

Latency, memory, efficiency, and real-time feasibility.

Platform workflow

A transparent path from task selection to a verified result.

1Select a benchmark task
2Download data and starter kit
3Develop locally
4Submit outputs or model
5Run standard evaluation
6Receive detailed report
7Publish verified result

Roadmap

Building the infrastructure in verifiable stages.

The first release establishes the public platform and scientific direction. Evaluation services follow in measured increments.

Current

v0.1 Website MVP

Public site, benchmark descriptions, roadmap, and early documentation.

Planned

v0.2 Static Release

Starter kits, baselines, validation splits, and example leaderboards.

Planned

v0.3 Submission Prototype

Prediction uploads, automatic scoring, and private reports.

Planned

v1.0 Hidden Evaluation

Container submission, hidden data, and verified readiness metrics.

Get involved

Help shape open, reproducible evaluation standards for BCI research.

We welcome feedback on benchmark definitions, datasets, evaluation metrics, baselines, and community governance.