I came to BCI through time-series work rather than neuroscience, and the failure modes looked familiar: leaky splits, preprocessing that erases the signal it was meant to expose, and scores that do not survive the next recording session.

Explore this article visually3 figures

01

Audit one window across the split

02

What the electrode actually sees

A scalp EEG channel at C3 records a voltage on the order of 10–100 µV: the summed activity of a large population of neurons under the electrode, attenuated by bone and tissue, plus the electrical signature of every muscle nearby. A blink is around 100 µV. The mains supply contributes a 50 Hz line (60 Hz in North America) that is often larger than the brain signal. The useful pattern for a motor-imagery task—a drop in 8–13 Hz mu and 13–30 Hz beta power over the motor cortex when a person imagines moving a hand—is a small change on top of all of that.

That is the engineering reality behind the phrase “brain–computer interface”: a low signal-to-noise, non-stationary time series with events of known timing, and a question that is much narrower than “reading thoughts”. Can a 2-second window after a cue be classified as left-hand versus right-hand imagery well enough to move a cursor? Framed that way, the problem is recognisable to anyone who has worked with sensor data.

03

Windows are a modelling decision

A recording is not a table of independent rows. Its meaning depends on sampling rate (250 Hz is typical for consumer and research headsets), channel identity and montage, the reference used, the event markers, and the exact preprocessing applied. A model input is a window cut from that stream, and the cut is a choice with consequences.

For motor imagery a 2-second window starting 0.5 s after the cue is a common default; the response takes a few hundred milliseconds to develop and fades after a couple of seconds. A 0.5-second stride gives the interface responsiveness. It also means every sample belongs to four windows, and that is where evaluations go wrong: if windows are shuffled and split, the test set contains near-copies of the training set, and the score is a measurement of memory, not of decoding.

Diagram 01

Windows, markers, and a split that respects time

C3 · 250 Hz · avg ref · 50 Hz notch · 8–30 Hz band-passcue · t = 0+1.0 s2 s windows · 0.5 s stride · every sample belongs to 4 windowsSplit A — shuffled windowsneighbours on both sides → inflated score (leakage)Split B — by session, in ordersession 1 · Mondaysession 2 · Wednesdaysession 3 · Friday · testthe test set contains the electrode that shifted and the end-of-week fatigueSame model, same data: only the split changes, and the split decides whether the number means anything.
Top: one filtered channel with a cue marker and overlapping 2-second windows. Bottom: two ways to split the same recordings. Only the second one measures what a person will experience next Friday.

The rule is the same as for any time series: split by the unit that will be new at inference time. For a BCI that is the session, and beyond that the participant. A held-out Friday session contains the electrode that shifted slightly and the fatigue of the end of the week. A shuffled split contains neither, and reports a number nobody will see in practice.

04

The preprocessing you did not log is the one that fooled you

Figure 02Data flow

Keep preprocessing inside the evaluation contract

Select an element to explore its role.
Read every explanation
Record
Keep sampling rate, channel identities and event timing with the raw recording. These define what a window means.
Split
Choose the unit that will be new at inference time. Keep overlapping windows from that unit on the same side of the split.
Fit
Fit CSP, scaling and other learned transformations on training data only. Apply the fitted transforms to the held-out data without refitting.
Evaluate
Report per-person outcomes and rejected decisions. For online use, account for buffering and filter delay instead of reusing an offline score.
01 / 04
Record

Keep sampling rate, channel identities and event timing with the raw recording. These define what a window means.

An evaluation schema, not a universal EEG preprocessing recipe. Timing and filters depend on the experimental task.

Before a decoder sees a feature the signal is typically notch-filtered at line frequency, band-passed to the range of interest, re-referenced (common average is a frequent choice), and checked for bad channels and artifacts. Each step has parameters, and each parameter can quietly help or hurt. A band-pass that starts at 8 Hz is right for mu but throws away the slow potentials another paradigm depends on. An aggressive artifact rejection can remove exactly the trials where the participant was concentrating hardest.

pythonA pipeline whose settings are data. The dict is saved beside every derived file; the raw file is never modified.
PREPROC = {
    "version": "2026.09.1",
    "notch_hz": 50.0,
    "bandpass_hz": (8.0, 30.0),
    "reference": "average",
    "epoch_s": (0.5, 2.5),          # relative to cue
    "reject_uv": 150.0,             # peak-to-peak, after filtering
}

raw = mne.io.read_raw_fif(path, preload=True)          # immutable source
raw.notch_filter(PREPROC["notch_hz"]).filter(*PREPROC["bandpass_hz"])
raw.set_eeg_reference(PREPROC["reference"])
epochs = mne.Epochs(raw, events, tmin=PREPROC["epoch_s"][0], tmax=PREPROC["epoch_s"][1],
                    baseline=None, reject=dict(eeg=PREPROC["reject_uv"] * 1e-6), preload=True)
epochs.info["description"] = json.dumps(PREPROC)

The habit that matters is that the settings travel with the data. When a result changes between two runs, the first question is whether the preprocessing changed, and that question should be answerable by diffing two small dictionaries rather than by reading a notebook's history.

05

Earn the deep model

A useful baseline in motor imagery is decades old: common spatial patterns to find the channel combinations that separate the two classes, log-variance of the filtered signal as features, and a linear discriminant classifier. It trains in seconds on a few dozen trials, is interpretable (the spatial filters should look like motor cortex), and is hard to beat on small per-participant datasets.

ApproachData neededStrengthFailure to watch
Band power / CSP + LDATens of trials per classFast, interpretable, robust when calibrated per sessionDegrades as electrodes shift; needs recalibration
Riemannian (covariance + tangent space)SimilarLess sensitive to scaling and small shiftsHarder to explain to a clinician
Compact CNN (EEGNet-style)Hundreds of trials or transferLearns spatial-spectral filters jointly; cross-subject transferOverfits gloriously on a leaky split
Sequence / attention modelsThousands of trialsLonger dependencies, multi-paradigmRarely justified by the data available per person

The comparison that should decide is not aggregate accuracy on a benchmark. It is accuracy per participant on a held-out session, calibration time before the interface is usable, latency per decision, and the cost of a wrong command. A deep model that wins by three points on a pooled benchmark and loses per participant has not won.

06

Abstain is a class

Figure 03Comparison

A decoder needs a no-decision path

Select an element to explore its role.
Read every explanation
Usable signal
Combine signal quality and task-specific validation before emitting a command. Consequential actions may require an explicit confirmation beyond the decoder.
Unresolved signal
Do not force a left/right answer when signal quality is poor. Show the actionable issue and let the person recover rather than treating abstention as an invisible failure.
01 / 02
Usable signal

Combine signal quality and task-specific validation before emitting a command. Consequential actions may require an explicit confirmation beyond the decoder.

Illustrative decision paths. Agreement between overlapping windows is correlated evidence, not independent confirmation.

A decoder that must output left or right on every window will output nonsense during the windows where the person sneezed, looked away or simply did not try. The safer design treats “no decision” as an output with its own rules: a posterior below a threshold, disagreement between the last three overlapping windows, or a signal-quality flag on the relevant channels all route to abstain. An assistive cursor can afford a low threshold and reversible moves; anything that triggers a consequential action should demand agreement across windows and an explicit confirmation.

Abstaining also gives the interface something honest to show. A quality indicator and a “recalibrate” control are not admissions of weakness; they are the parts of the system the person can act on.

07

Session three is the benchmark

The result that matters in BCI work is not a high score inside one controlled recording. It is a system that is still useful when the person comes back on another day, with the cap seated a little differently, more tired, in a noisier room. Report held-out-session and held-out-participant results separately, show the class balance, inspect the confusion matrix per person, and keep the raw recordings immutable so the evaluation can be re-run when the pipeline changes.

Everything upstream of the model—acquisition, context, windows, splits, abstention—decides whether the number at the end means anything. The model is the last thing to improve, not the first.

08

An offline pipeline is not an online decoder

The snippet illustrates offline epoch construction: mark bad channels explicitly before re-referencing and keep the original recording on disk. With epochs starting at 0.5 seconds, use baseline=None unless you deliberately include a baseline interval. Zero-phase filtering can use future samples, so online claims require a causal or buffered pipeline with its delay measured. Fit CSP, scaling and other learned transforms inside each training fold. Agreement between overlapping windows is correlated evidence, not three independent confirmations.

Discuss this article

More articles

09Retrying is easy. Avoiding duplicate actions is harder.Architecture · Reliability · 5 min read08RAG: evaluate the evidence before the answerRetrieval · Evaluation · 6 min read01Building agentic AI beyond the demoAgentic AI · 9 min read