Cross-Subject Generalization
Can a model work on users it has never seen?
Open BCI benchmark infrastructure
Benchmarking Generalizable and Deployment-Ready Brain-Computer Interfaces
An open evaluation platform for testing BCI algorithms across users, sessions, devices, calibration settings, and deployment constraints.
Website MVP in active development
Independent, standardized evaluation of whether BCI algorithms generalize beyond the datasets on which they were developed.
Why BCI-Bench?
BCI research needs shared protocols that reveal how algorithms behave beyond familiar data and ideal laboratory conditions.
Standardized preprocessing, splits, metrics, and reporting make results easier to interpret and reproduce.
Hidden test sets and controlled protocols reduce overfitting to public benchmark data.
Reliability, calibration cost, latency, and worst-subject behavior matter alongside average accuracy.
Core benchmark tracks
Each track asks a focused scientific question and reports more than one aggregate score.
Can a model work on users it has never seen?
Does performance persist across days and neural drift?
Can a method transfer across acquisition systems and sites?
How much user-specific data is needed to become useful?
How does a model fail under noisy or incomplete signals?
Can the method meet practical runtime constraints?
Platform workflow
Roadmap
The first release establishes the public platform and scientific direction. Evaluation services follow in measured increments.
Public site, benchmark descriptions, roadmap, and early documentation.
Starter kits, baselines, validation splits, and example leaderboards.
Prediction uploads, automatic scoring, and private reports.
Container submission, hidden data, and verified readiness metrics.
Get involved
We welcome feedback on benchmark definitions, datasets, evaluation metrics, baselines, and community governance.