Validating an EEG analysis chain before the first real recording
A continuation of the neuro entry: before any real recording, the analysis chain was tested on synthetic data with known ground truth and negative controls. Negative results on stimulus and signal, four methodological audit findings and their corrections, a norm restricted to amplitudes — and corrections to the entry of 2026-08-12, including a withdrawn intended-purpose sentence.
Koch Laboratory — validating an EEG analysis chain before the first real recording
This entry continues the “neuro” entry (dated 2026-08-10 in the notebook, stamped in its present wording on 2026-08-12), which set out research directions for cognitive screening on four channels of consumer EEG. That entry said what we are looking for; this one says what was checked before the first person sits down for a measurement intended for the database: whether the analysis chain recovers an effect that was put into it, and whether it produces no effect where there is none. Added to this are negative results on stimulus and signal, the findings of a methodological audit together with their corrections, and design changes at a publishable level. Corrections to the entry of 2026-08-12 are collected at the end; that entry itself remains unchanged.
Method and intended purpose. As in the rest of the lab: problem → state of the art → falsifiable criterion → method → result with boundary conditions → next step. “Confirmed” means met within the criterion, “refuted” means tested and rejected with the reason given, “partial” means with the missing part named. Where “Construction: withheld” appears, the technical detail remains unpublished. The platform’s intended purpose: a measuring instrument for research and wellness use; it is not a medical device and does not provide information for diagnostic or therapeutic decisions.
1. Direction 1 — the ERP chain on synthetic ground truth, with a negative control
Research question. How do we know that the analysis chain — from Bluetooth packets through signal reconstruction and epoching to the statistics — finds an event-related potential that is in the data, and reports none where there is none? On a recording of a real person this cannot be checked, because the ground truth there is unknown. State of the art (published). Simulation of event-related activity with known ground truth as a way to validate analysis methods (e.g. SEREEGA, Krol et al. 2018); negative controls as a condition for the credibility of an effect; validation of the consumer headband for ERP research (the Krigolson group). Success criterion. (a) At low noise the chain recovers the injected amplitude and latency; (b) in a recording without an injected response no channel is significant; (c) the same recording yields the same numbers in the local analysis tool and in processing on the platform. Method. The synthetic recordings reproduce the headband’s data transport (packets with sequence numbers, variable arrival times, lost packets), plus blinks, alpha rhythm and 50 Hz mains interference. Injected truth: a difference wave of +6.5 µV at a latency of 350 ms between two event classes. The variants with low noise, with noise typical of the headband and with noise alone ran through the local analysis tool (11 August 2026); the recording with a response and the negative control also went through the platform’s full path: upload of the recording → processing queue → analysis → result (13 August 2026). Result. Confirmed within the criterion. (a) At low noise, recovered +6.3 µV at 336 ms and +6.5 µV at 355 ms on two frontal channels, against a truth of +6.5 µV at 350 ms. (b) Negative control: no significant channel (p from 0.09 to 0.42); on the platform’s full path the injected effect was detected on one channel out of four (uncorrected p = 0.015), and the negative control produced no significant channel (p from 0.28 to 0.89). (c) The local tool and the platform processing give identical numbers on the same recording — it is one implementation, not two. After the corrections from direction 5 the chain processed synthetic recordings again (23 August 2026), but a complete repetition of both controls, with comparison against the ground truth, is not documented. Boundary conditions. At noise typical of the headband, an effect that really is in the data comes out significant on only one or two channels out of four — that is the realistic sensitivity of a single session. Peak latency picks up noise when trials are few (with ten epochs the peak fell at 441 ms against a truth of 350 ms), so the primary measure is the mean amplitude in a fixed post-stimulus window, with the peak reported as auxiliary. Under the Holm correction introduced after the audit (direction 5), the single-channel result from the platform’s full path would no longer be significant at the 0.05 level (adjusted value 4 × 0.015 = 0.06). Next step. Quantify the sensitivity of a single session under the corrected statistics — how often the known effect reaches significance at realistic noise — and repeat the negative control; then the first recording of a real person, against the criterion from the previous entry: a visible averaged event-related potential after at least 30 repetitions. Status: confirmed on synthetic data; the criterion on real data is open — there are no such data in the analysis database yet.
2. Direction 2 — inter-subject correlation on synthetic pairs, and shared sessions
Research question. Does the ISC measure distinguish a pair of recordings with a shared, stimulus-driven modulation from a pair without it — with a null distribution derived from the same data and a result that is identical at every recomputation? And can a dyad be recorded at all: two or more people watching the same film in different places? State of the art (published). ISC in EEG with naturalistic stimuli (Dmochowski et al. 2012; Ki, Kelly, Parra 2016); surrogate data — circular shifts, phase randomisation (Theiler et al. 1992) — as a null distribution; remote hyperscanning as an established current of social neuroscience (previous entry). Success criterion. Correlated pair: high r and p < 0.05; uncorrelated pair: r close to zero and p > 0.05; recomputation: identical p. Method. Synthetic recordings with and without a shared amplitude modulation; correlation of amplitude envelopes between people after alignment to the stimulus; p from circular shifts. Result. Confirmed on synthetic data (14 August 2026): correlated pair r = 0.93, p = 0.005 (the smallest value attainable with the number of permutations used); negative control r = −0.001, p = 0.98. Methodological finding: a purely periodic modulation yields a conservative p under a circular-shift null, because such a null preserves autocorrelation — a property of correct statistics, not a fault; test signals must be spectrally rich, like a real film. After the audit (direction 5) the computation was rebuilt; both controls are part of the automated tests. Shared sessions were implemented on 14 August 2026: a joint start for participants in different places, recordings assigned to a common session slot, and ISC for every pair, recomputed after each session. Construction (synchronisation of start and clocks across locations): withheld. Boundary conditions. The lab has a single headband — no real dyad has been recorded, so the hyperscanning criterion from the previous entry remains untested. Of the controls named in that criterion, circular shifts are in place; sham pairs (people who were never connected) and correction for multiple comparisons across pairs are not yet implemented. Next step. A second headband and the first real dyad; sham pairs and correction across pairs before any interpretation of a result. Status: partial — method confirmed on synthetic data, shared sessions implemented, criterion on real dyads open.
3. Direction 3 — the stimulus: a rejected generative model and a one-frame offset
Research question. ERP averaging requires physically identical repetitions that differ only in the intended dimension, with stimulus onset in a known frame. Where does such a stimulus come from — and how do we know that the marked frame is the one in which the stimulus actually appears? State of the art (published). In experimental psychology, stimuli are generated programmatically with controlled parameters, and their timing is checked by hardware measurement (photodiode), as in comparative studies of experiment software (Bridges et al. 2020). Success criterion. (a) Repetitions identical to the frame; (b) conditions differ only in the intended dimension; (c) every marked onset falls in the frame in which the stimulus first appears — checked independently of the generator; (d) documented provenance and licence that allow a stimulus version to be frozen for years. Method. An attempt to build a stimulus film with a generative video model; an independent audit of every stimulus version — a program that reads only the finished video file and the marker list and does not trust what the generator declares. Result. (1) Generative video model rejected as a stimulus source (11 August 2026). The reason was not aesthetic but methodological: the model does not repeat a shot to the frame (a); between clips, luminance, contrast, motion and pace change, and each of these parameters evokes a brain response of its own (b); an event cannot be ordered for a specific frame (c); provenance and licence of the material are unclear against the requirement of a freeze lasting years (d). (2) A systematic one-frame offset in our own generator — caught by the independent audit. On part of the trials the marker fell one frame before the frame in which the stimulus actually appeared; the latencies of those trials would have come out one frame — several tens of milliseconds — too long, on the scale of the latency differences we are looking for. Corrected in the stimulus generators; the lab’s stimulus versions currently in use passed this audit before publication. Boundary conditions. The audit checks the file, not the screen: what happens between the file and the eye (decoder, compositor, display) is covered only by the hardware measurement — not yet performed. Next step. Measurement of the timing of the presentation chain with a photodiode. Status: negative result (generative model rejected) and an error of our own caught and corrected.
4. Direction 4 — the signal: blinks rather than trial count, ICA on four channels, pilot recordings
Research question. With four dry electrodes, what limits the detectability of an ERP in a single session — the number of trials or artefacts? Do standard artefact-removal tools work on four channels? And are the existing pilot recordings usable for ERP? State of the art (published). Ocular artefacts dominate the frontal channels; ICA is the standard for removing blinks in multichannel EEG, but its separating power depends on the number of channels; the alternatives are epoch rejection and regression-based correction (Gratton, Coles, Donchin 1983). Success criterion. An intervention counts as helpful only if, in simulation, it improves detection of the known injected effect and produces no effect in the negative control. A recording counts as usable for ERP only if it has the nominal sampling rate and carries stimulus markers synchronised with the presentation. Method. The simulation from direction 1 at noise typical of the headband, with blinks; the number of trials, the epoch rejection threshold and ICA (on or off) were varied. For the pilot recordings — a check of the effective sampling rate and of the markers. Result. (1) The hypothesis “more trials = better detectability” refuted in simulation: more than tripling the number of events did not improve detectability; only stricter rejection of epochs contaminated by blinks did. The variance comes from blinks, not from the trial count — electrode contact and participant instructions take priority, not a longer stimulus. (2) ICA on four channels unreliable: in simulation it did not reliably isolate the blink component; it remains an option, and the main defence is epoch rejection. (3) Pilot recordings unusable for ERP: none carries stimulus markers, and the most recent export from a third-party recording app had an effective sampling rate more than two orders of magnitude below the nominal one — a component lasting about 100 ms fits between two samples. Such files are good at most for band trends and electrode-contact checks, never for ERP and never for the normative database. Next step. Regression-based blink correction as an alternative to ICA — tested on the same synthetic recordings, with a negative control. Status: two hypotheses refuted, pilot data excluded from ERP — negative results, published deliberately.
5. Direction 5 — methodological audit: four findings and their corrections
Research question. Does the analysis code do what the methodology claims — including in cases nobody tested by hand? State of the art (published). Control of the family-wise error rate across many tests (Holm 1979); pseudoreplication, that is, counting repeated measurements of the same person as independent observations (Hurlbert 1984; Lazic 2010); reproducible computation as a requirement for permutation tests. Success criterion. A finding counts as closed only when the correction is in the code and a regression test covers it. Method. An audit of the project (15 August 2026) that covered, among other things, the analysis code and the methodology; corrections the following day; regression tests in continuous integration; a second audit (September 2026) that checked, among other things, whether the earlier findings are closed. Result. Four methodological findings:
- Four channels tested without correction. Each channel was assessed separately at p < 0.05 — with independent channels, that is about an 18.5% chance of at least one false positive in a session with no effect at all. Correction: the Holm–Bonferroni procedure; the raw p stays in the result, and significance is decided by the adjusted value.
- Corrupted packets could produce a fictitious signal. Validation let through packets with malformed fields, and large gaps were silently interpolated — the chain could “fill in” signal where there was none. Correction: strict structural validation, ordering of packets and discarding of duplicates; a recording with a continuous gap longer than 2 s or with more than 20% of samples missing is rejected rather than interpolated.
- ISC: compressed time and a non-repeatable seed. The time axis was assembled from consecutive packets, so gaps vanished and the axes of two people with different losses slid against each other; the permutation seed depended on a hash function randomised in every process, so p changed at every recomputation. Correction (revised ISC): time axis from the sequence numbers, with gaps explicitly marked; correlation only on windows present for both people; a stable seed derived from the pair’s identifiers with a cryptographic hash function; hashes of the input data stored with the result.
- The norm: repetitions as independent observations, missing quality control as a pass. Several sessions (and recomputations) of one person entered the norm as if they were further people, and missing presentation-quality metrics were treated as a positive result. Correction: one result per session and one session per person in the reference group — the size of the norm counts people, not measurements; missing quality control means exclusion.
After the audits, the handling of consent and deletion was also corrected — before any data of real participants existed. The corrections went into the code on 16 August 2026. Findings 2 and 3 are covered by regression tests in continuous integration (corrupted, duplicated and out-of-order packets, wrap-around of the sequence counter, gaps, ISC reproducibility and gap marking), and the second audit confirmed them as closed. The corrections for findings 1 and 4 are in the code but without dedicated regression tests; the second audit rated finding 4 as only partially closed: admission of a session to the norm does not yet check the device and firmware version and has no packet-loss threshold of its own. Next step. Regression tests for findings 1 and 4; admission rules for the norm that cover the device and firmware version (cf. direction 6 of the previous entry) and a packet-loss threshold. Status: two findings closed; two corrected in the code but without a regression test — one of them only partially closed.
6. Direction 6 — the norm: amplitudes only, latencies after calibration, percentiles from N ≥ 20
Research question. When does a number shown against a comparison group mean what it suggests? A percentile looks like a fact even when the group consists of a few people, or when the measured quantity carries a constant, unmeasured offset. State of the art (published). Test norms are reported against defined reference groups of sufficient size; latency measurements require hardware verification of stimulus timing (Bridges et al. 2020). Success criterion. A percentile is shown only if (a) the quantity is free of unmeasured systematic offsets and (b) the reference group contains at least 20 people. Method. Absolute latency contains a constant, unmeasured delay of the chain Bluetooth + compositor + screen; differences between conditions within one session are free of it. Hence three rules: only amplitudes are normed; latencies stay out of the norm until the hardware timing measurement; with fewer than 20 people in the reference group the group size is shown instead of a percentile — counted in people, not in sessions (direction 5). Result. Implemented (15 and 16 August 2026) and checked by tests: latencies do not enter the norm, and below the threshold the group size appears instead of a percentile. Boundary condition: the analysis database contains no recordings of real people yet, so no reference group exists — the rules are in force but have not yet had anything to act on. Next step. After the timing measurement: a decision on whether absolute latencies can enter the norm — with the measured offset and its spread stated openly. Status: implemented; rules in force, not yet exercised on real data.
Corrections to the entry of 2026-08-12
The entry of 2026-08-12 stands in the wording in which it was stamped. The corrections below do not replace it; they correct it in the light of what is known today.
- Direction 1 (event marking in the browser) — the status was understated. It listed only acquisition, although on the day of the entry the path from recording to ERP result already existed (in the local analysis tool; on the platform from the following day) and had been validated the day before on a synthetic recording with known ground truth and a negative control (direction 1 of this entry). The ERP criterion on a recording of a real person was not met then and is not met today. The timing reference is a direct hardware measurement of the presentation chain (photodiode), not yet performed; the parallel laboratory pipeline named in that criterion (a desktop acquisition path) was not built.
- Direction 2 — the “main gap” has moved. The entry named the lack of event markers in the pilot recordings as the main gap. The markers were no longer the bottleneck even then — every recording from our own player contains them — and the most recent pilot recording also proved unusable for ERP for a second, independent reason (direction 4 of this entry). The main gap today is data: there is no recording of a real person in the analysis database, there is no normative cohort, and hardware timing remains unmeasured.
- Direction 3 — shared sessions. The entry stated: “the shared-session architecture is designed, implementation not started”. That was true on the day of the entry and was out of date two days later: shared sessions were implemented on 14 August 2026 (direction 2 of this entry). The criterion of that direction remains untested — no real dyad has been recorded so far.
- Direction 5 — operations log and consents. The status described the log as “append-only” and consents as events carrying the version of the text. In reality both properties were weaker at the time than the description suggested: they were conventions of the application, not guarantees enforced by the system. They have been enforced only since the corrections that followed the audits of August and September 2026. Anchoring and the resolution of the conflict with the right to erasure remain open, as in that entry.
- Intended purpose — sentence withdrawn. We withdraw the sentence from the “Method” paragraph: “the tool provides information supporting a physician’s decision and never adjudicates; the medical-device path (MDR) is deliberately a separate, later stage”. The wording that applies: the platform is a measuring instrument for research and wellness use; it is not a medical device and does not provide information for diagnostic or therapeutic decisions. The word “screening” in that entry denotes a research question, not a use.
Program status: active. The analysis chain is validated on synthetic data with negative controls, the findings of the methodological audit are corrected (two fully closed), shared sessions are implemented, and the rules of the norm are narrowed. Boundary conditions, stated plainly: there is so far no recording of a real person in the analysis database; the lab has a single headband, so no real dyad has been recorded; hardware timing (photodiode) has not been measured; browser acquisition works only in Chromium-based browsers (Web Bluetooth) — not in Firefox, not in Safari and not on iPhone or iPad. Next stage: the timing measurement with a photodiode and the first recordings of real people on a frozen stimulus version.
Proof of time
This text was hashed and stamped on publication. Verifying it takes both files: the stamp proves that one exact byte string existed at that moment, and only the source below is that byte string. Any later edit changes the hash and breaks the match — which is the point.
SHA-256 of the stamped text: e7441c099ae1ab27de6158664ee9801819c3a89d0cc248885811ce63bd8ce292
- Sigelith anchor portfolio-neuro-validation.public.md.beattime.json
- OpenTimestamps proof portfolio-neuro-validation.public.md.ots
- Stamped source text portfolio-neuro-validation.public.md
Verify with: ots verify <proof> --file <source>