Note
Go to the end to download the full example code.
Track 2 – BCI decoding (cross-session)¶
Given short EEG windows recorded while a user performs one of three cued mental commands (kinesthetic motor imagery, mental calculation, or word association), decode the active command. The competition tests cross-session generalisation: models train on a user’s earlier sessions and are scored on their later ones, with no per-session recalibration allowed.
Shift: earlier sessions -> later sessions (Graz + BrainHero contexts), within the same user.
Headline metric: balanced accuracy averaged over subject-session-context cells (higher is better).
Data: 20 subjects, 6 sessions each, 47 channels (43 EEG, 2 EMG, 2 EOG) at 500 Hz, ~80 hours in total. Sessions 1-3 of the 10 evaluation subjects are released as labelled calibration; sessions 4-6 are the hidden test set. The 10 training subjects have all 6 sessions released.
Note
The official Track 2 corpus (Graz / BrainHero, 3 classes: MI / Calc / Word) is released through NeuralBench when submissions open. Until then, Starter-kit analogs below lists the NeuralBench tasks that come closest, one per axis of the official task.
New to NeuralBench? Start here¶
This page is a track guide, not an introduction to the ecosystem. If you arrived straight from the competition website:
Challenge overview – what NeuralBench is, how it relates to the competition, and the baseline numbers for all four tracks.
Installation and the quickstart – get a task running on a 1.5 GB dataset before you download anything large.
Official Track 2 guide – registration, rules, data access, prizes, leaderboard. Authoritative on every competition matter; this page only covers the code.
Where to find this task in NeuralBench¶
The starter-kit baseline is Motor imagery classification.
The task’s own default dataset is Stieger2021Continuous: 62 subjects
of 4-class MI, and the corpus the published NeuralBench Track 2 baseline
numbers come from. It is also a ~640 GB download (~940 GB once MOABB has
converted it), and it is not what Codabench scores against.
So these pages build against the recommended warm-up configuration instead, selected with an explicit flag:
CLI:
neuralbench eeg motor_imagery --dataset dreyer2023Dataset:
Dreyer2023Large(87 subjects, 27-channel EEG, 2-class motor imagery – left hand / right hand, ~19 GB). This is the corpus Codabench scores Track 2 against during the warm-up phase, so it is the one to build against first.Shift: held-out subjects, not the cross-session shift of the competition. Use it to validate the training pipeline and architecture choice.
Headline metric key:
test/bal_acc.
Important
--dataset dreyer2023 is required. Dropping it does not fall back
to the warm-up corpus – it runs Stieger2021Continuous, a
different task (4 classes, not 2) on a very much larger download.
Every command on this page carries the flag for that reason.
What the config is. A NeuralBench task is one config.yaml, and
nothing else: a YAML overlay on neuralbench/defaults/config.yaml naming
the study to load, how to split it, what the target is, the loss, and the
metrics. Reading it is the fastest way to know exactly what the baseline
does.
Show tasks/eeg/motor_imagery/config.yaml
# Copyright (c) Meta Platforms, Inc. and affiliates.
# All rights reserved.
#
# This source code is licensed under the license found in the
# LICENSE file in the root directory of this source tree.
data:
study:
source:
name: Stieger2021Continuous
split:
name: SklearnSplit
split_by: subject
valid_split_ratio: 0.2
test_split_ratio: 0.2
valid_random_state: 33
test_random_state: 33
target:
=replace=: true
name: LabelEncoder
event_types: Stimulus
event_field: code
return_one_hot: true
aggregation: trigger
trigger_event_type: Stimulus
start: 0.0
duration: 4.0
summary_columns: [code]
compute_class_weights: true
brain_model_output_size: &brain_model_output_size 4
trainer_config.monitor: val/bal_acc
trainer_config.mode: max
loss:
name: CrossEntropyLoss
kwargs:
label_smoothing: 0.1
metrics: !!python/object/apply:neuralbench.defaults.metrics.get_classification_metric_configs
- *brain_model_output_size
Show tasks/eeg/motor_imagery/datasets/dreyer2023.yaml
# Copyright (c) Meta Platforms, Inc. and affiliates.
# All rights reserved.
#
# This source code is licensed under the license found in the
# LICENSE file in the root directory of this source tree.
# Dreyer2023Large combines parts A (subjects 1-60), B (61-81), and C (82-87).
# 6 subjects in part C are re-recordings of subjects from A or B under
# different IDs (exact mapping unknown). We assign part B (61-81) to test
# so that train/test populations are disjoint. Part C subjects stay in
# train only, where overlap with A is harmless.
data:
study:
source:
name: Dreyer2023Large
split:
=replace=: true
name: PredefinedSplit
test_split_query: "subject.isin(['Dreyer2023Large/61', 'Dreyer2023Large/62', 'Dreyer2023Large/63', 'Dreyer2023Large/64', 'Dreyer2023Large/65', 'Dreyer2023Large/66', 'Dreyer2023Large/67', 'Dreyer2023Large/68', 'Dreyer2023Large/69', 'Dreyer2023Large/70', 'Dreyer2023Large/71', 'Dreyer2023Large/72', 'Dreyer2023Large/73', 'Dreyer2023Large/74', 'Dreyer2023Large/75', 'Dreyer2023Large/76', 'Dreyer2023Large/77', 'Dreyer2023Large/78', 'Dreyer2023Large/79', 'Dreyer2023Large/80', 'Dreyer2023Large/81'])"
col_name: split
valid_split_by: subject
valid_split_ratio: 0.2
valid_random_state: 33
brain_model_output_size: &brain_model_output_size 2
metrics: !!python/object/apply:neuralbench.defaults.metrics.get_classification_metric_configs
- *brain_model_output_size
How to change it, in increasing order of effort:
--dataset <name>mergestasks/eeg/motor_imagery/datasets/<name>.yamlover the base config. Seventeen MI corpora ship that way, and the competition corpus will too once it lands.-m <model>and-w <preset>swap the architecture and the adaptation strategy (frozen probe, LoRA, full fine-tuning) without touching any file.Anything else – window length, learning rate, split – is a config edit. From Python, pass dotted keys to
evaluate_model()(overrides={"data.duration": 3.0}). From a source checkout (pip install -e), editconfig.yamldirectly, or add your owndatasets/*.yamlvariant beside the existing ones and select it with--dataset. Both routes are described in Adding a New Task. The split is the field to reach for here: see Adapting to the competition setup.
Split and model selection¶
Split (``–dataset dreyer2023``). Subject-level and predefined, not
random. Dreyer2023Large pools the study’s parts A (subjects 1-60),
B (61-81) and C (82-87); PredefinedSplit assigns all 21 subjects of
part B to test and everything else to train, then holds out 20 % of the
remaining subjects as validation (valid_split_by: subject, seed 33).
Part B is used for test because six part-C subjects are re-recordings of
part-A subjects under different IDs, so carving test out of A or C could
leak a person across folds. Every subject therefore appears in exactly one
fold, and the partition is identical on every machine.
That part-B test partition is the current Codabench warm-up evaluation
set, so a test/bal_acc from this configuration and a warm-up
leaderboard score measure the same thing.
Split (task default ``Stieger2021Continuous``). Subject-level
60 / 20 / 20 drawn by SklearnSplit with split_by: subject and both
seeds fixed at 33, which on 62 subjects resolves to 36 train /
13 validation / 13 test.
Either way the starter-kit shift is cross-subject, while the sealed phase’s is cross-session within subject. See Adapting to the competition setup for what changes.
Model selection. The checkpoint with the highest ``val/bal_acc``
is kept – validation balanced (macro-averaged) accuracy, the same
quantity as the headline test/bal_acc, just on the validation fold.
Training runs for at most 40 epochs and stops early after 5 epochs
without improvement; only that single best checkpoint is scored on test.
The warm-up scorer also ranks on balanced accuracy, but pooled over all evaluation windows. The sealed phase instead averages it over subject-session-context cells, so that every cell counts equally regardless of how many windows it holds – a different number from the same predictions.
Reproducing the baseline¶
Dreyer2023Large and the alternative MI datasets are served by
MOABB, which the base install does not pull, so install it first:
pip install 'moabb>=1.7.1'.
# 1. Download Dreyer2023Large into DATA_DIR: ~19 GB. One-off per
# machine, and safe to interrupt and re-run.
neuralbench eeg motor_imagery --dataset dreyer2023 --download
# 2. Preprocess into CACHE_DIR -- resample, filter, scale and window
# every recording once, so each later run reads the cache instead.
# No GPU needed, and it fans out over SLURM when one is configured.
neuralbench eeg motor_imagery --dataset dreyer2023 --prepare
# 3. Sanity check before you queue anything: 2 epochs, a data subset, one
# seed, always in-process, so progress lands in your terminal. Name
# the model you actually plan to run -- a bare --debug takes the
# config default, which is EEGNet.
neuralbench eeg motor_imagery --dataset dreyer2023 -m eegnet --debug
# 4. Same check for the foundation model. The first build pulls REVE's
# weights from the HuggingFace Hub, which needs network access; doing
# it here rather than in a queued run keeps any failure in your
# terminal instead of a job log.
neuralbench eeg motor_imagery --dataset dreyer2023 -m reve --debug
# 5. Full baseline -- task-specific model (EEGNet). The default grid is
# three seeds (concurrent on SLURM).
neuralbench eeg motor_imagery --dataset dreyer2023 -m eegnet
# 6. Full baseline -- foundation model (REVE), fine-tuned end to end.
# ~69M parameters against EEGNet's ~1.5k, all of them trainable here,
# so this one wants a datacentre GPU rather than a laptop; it also
# preprocesses at 200 Hz against the 120 Hz default, warming a second
# cache.
neuralbench eeg motor_imagery --dataset dreyer2023 -m reve
Tip
Smaller still: --dataset tangermann2012 is BCI Competition IV-2a,
9 subjects of 22-channel four-class MI in under 1 GB, with the whole
download-prepare-train loop in well under an hour (~15 min to prepare,
~2 min per training seed) against a well-known published baseline.
Dropping --dataset dreyer2023 from any of the commands above runs
Stieger2021Continuous instead. Budget ~940 GB on disk (~640 GB, plus
~300 GB for the copy MOABB converts on first read), ~96 GB of cache and
~65 min of preparation over 75 SLURM jobs, then ~30 min per EEGNet seed
and ~2.5 h per REVE seed. Do not skip --prepare there: a cold-cache
--debug on that corpus spent ~45 min doing the same work serially
in-process.
Pretrained model weights covers the hub cache, and no run has a CPU
fallback – --debug included.
Steps 5 and 6 cache the test-metric dictionary under SAVE_DIR, and
re-running the same command with --plot-cached turns those cached
metrics into comparison plots and CSV tables without retraining. On SLURM
they return as soon as the grid is queued, so the numbers appear in the job
logs rather than your terminal; see Collecting and plotting your results.
Evaluating a model of your own¶
A model that lives in your own codebase needs no YAML here:
evaluate_model() takes the built instance, wraps it in a
probe sized to the task, and returns the scores as a DataFrame.
from neuralbench import check_model, evaluate_model
print(check_model(my_model, "eeg", "motor_imagery")) # shapes only, seconds
scores = evaluate_model(my_model, "eeg", "motor_imagery", name="my-fm", debug=True)
See Evaluating your own model for what
forward has to accept, the adaptation presets, and how to fan the runs
out to SLURM.
Starter-kit analogs¶
No public dataset has all of the official task at once – three mental commands, six sessions, one user at a time – so the four public corpora the competition points to are split across three NeuralBench tasks. Each covers a different axis of Track 2, and all four are worth training on:
Command |
Dataset |
What it gives you |
|---|---|---|
|
|
87 subjects of 2-class MI, split on held-out subjects, and the corpus Codabench scores against during warm-up. |
|
|
The most data by far (62 subjects, 615 h) for the motor-imagery class, cross-subject, and the source of the published baseline numbers. |
|
|
The closest paradigm match: five cued mental tasks including arithmetic and letter association, split cross-session (session 1 held out as test). |
|
|
Mental calculation against a rest baseline, 36 subjects. |
Every other MI corpus registered in
Motor imagery classification (BCI Competition IV, Cho2017,
Lee2019, …) can be selected the same way with --dataset <name>.
Adapting to the competition setup¶
To match the official Track 2 evaluation regime, two pieces need to change once the official dataset is released:
Dataset source: register the new MI / Calc / Word study and set
data.study.source.nameto it, withbrain_model_output_size: 3.Split: replace the default
SklearnSplitwith a predefined per-subject split where sessions 1-3 are train and sessions 4-6 are test. Theneuralbench.transforms.PredefinedSplitalready used bymental_imagery,reaction_timeandpsychopathologyis the right primitive – thetest_split_querybecomes"subject in evaluation_subjects and session in [4, 5, 6]".
Submissions may dispatch internally to per-subject sub-models using
the per-example metadata dictionary m (subject, session, run,
paradigm). Re-training on the hidden later sessions is forbidden.
Total running time of the script: (0 minutes 0.000 seconds)