Measurement is next
Seventeen sessions under REB 2025-8251 closed the gap between an approved design and an interpretable result. De-identified logs, questionnaires, and interviews are on record. Architecture branch remains inferred from the recorded host; GPU frame time and state-to-photons delay were not logged.
Qualitative rigor already named in the packet—codebook, memo trail, mapping from quote to theme, treatment of disconfirming cases—was used. Effect sizes and uncertainty sit beside a limited N. The loop is closed; it is not an invitation to pretend five elite participants proved equivalence on engagement.
The interactive edition therefore labels implemented claims, empirical claims, and remaining measurement gaps differently. GPU timings remain absent. Every later vocabulary expansion depends on knowing what this single-condition study can actually say.
Evidence first
The loop is closed.
01 · ProtocolREB 2025-8251 and the single-condition design ran: seventeen sessions on one floor.
ASCII field for the empirical loop: the protocol ran, seventeen sessions are on record, measurement is next.
Hold water, colour, duration, and a better question about “lost”
Five people predicted they would get bored inside ten minutes. A longer session is the only way to know. P12 pointed at winter pieces of forty minutes. Holding the mapping fixed while time stretches is the test.
Commandable arrest is cheap to act on and already has a name in rowing: hold water. A colour condition — abstract against naturalistic — is the experiment this protocol chose not to run. Palettes exist in the source; the next study can use them as a factor.
The “lost myself” item needs to ask whether lost was good. Peripheral uptake was reported, not shown. Those are measurement problems, not feature requests.
Benchmark before expanding vocabulary
Add reproducible GPU timings and end-to-end input latency measurements before making performance comparisons across devices or software revisions. As noted in the software chapter, pressure loops and bloom chains dominate cost on the compute path; optional particle advection is the other structural heavyweight. PM5 ingestion latency depends on BLE notification cadence and main-thread handoff; Python bridge timing depends on poll-and-flush cadence and UDP scheduling.
Next-study design should optionally instrument per-pass frame timing and transport latency for causal mediation analyses, and stratify by architecture branch or enforce homogeneous hardware per study block. Formal pre-registered hypotheses belong in confirmatory follow-up work after the single-condition baseline exists. The roadmap therefore puts instrumentation ahead of feature accretion: know the frame, then grow the vocabulary.
No in-repo benchmark dataset currently reports measured occupancy, bandwidth counters, or per-pass GPU timings. Filling that gap is a prerequisite for honest cross-revision claims, not a polish task after new sensors arrive. Measurement is how the seventeen-stage graph and the branch matrix become comparable across machines without leaning on structural cost stories alone.
Measurement
Benchmark before vocabulary.
01 · InstrumentAdd reproducible GPU timings and end-to-end input latency before comparing devices or revisions.
Bayer dither of the critical path — instrument timing before expanding features.
Make capabilities explicit
A future renderer could expose device-tiered visual quality while preserving the same interaction semantics, so quality changes never look like interaction changes. Bloom, particles, and ripple coupling constants already diverge by architecture; making those tiers explicit in UI and session metadata would turn an honest limitation into a controllable experimental factor rather than a surprise that appears only when an Intel fallback machine enters the lab.
New sensing or projection-correction techniques should enter the architecture as documented providers and transforms, not as hidden special cases. As noted in the sensing chapter, the InputEvent vocabulary already makes that extension path concrete: a new IMU stream or bridge protocol enters as another provider, logs through the same sink, and obeys the same refractory and ownership rules. Controlled multi-condition comparisons belong after this single-condition baseline.
The contribution guides and architecture verification anchors already tell rebuilders where to attach such work: Renderer and shader modules for passes, InputEvent and providers for sensing, RowingConstants for coupling knobs, and the reproducibility packet for anything that later claims to be an experimental condition. Evolution should widen the inspectable surface, not the undocumented one.
Extend the packet before widening claims
Additional confounds named for future protocols include visualization mode order effects, hardware thermal throttling over long sessions, PM5 characteristic drop-rate variability, and calibration quality variance across participants. Where possible, those fields should enter analysis datasets as covariates or block-level controls. Tracking them is cheaper than discovering them after a vocabulary expansion has already mixed several uncontrolled factors into one wake.
Existing mitigations—architecture metadata, PM5 liveness logging, IMU hysteresis, timestamp plausibility gates, and interview framing for novelty—should remain in the packet rather than being reinvented when the interaction vocabulary grows. Expanding sensing without expanding metadata would recreate the measurement gap the roadmap is trying to close. The confound table already pairs mechanism with mitigation; future work should extend that table rather than ignore it.
The point is not to freeze the system. It is to ensure that every new wake-revealing feature arrives with a way to say which machine, which branch, which provider, and which coupling produced it. Without that, future chapters risk becoming denser while becoming less interpretable.
Keep the seam visible
As the interactive edition sits beside the thesis PDF, it should continue to label implemented claims and empirical claims differently. The archive that supports the thesis—source-grounded architecture, REB-grounded method, and named hybrids—should remain inspectable. Future chapters can deepen without rewriting history, and they can add multi-condition designs without pretending the current single-condition study already isolated every causal factor.
The current system and the completed study already define a clear seam for what comes next. Measurement first, vocabulary second. That order is the roadmap: more measurable before more complex, and always honest about which wake was painted on which path. The next wake should reveal more because the packet got stronger, not because the prose got louder.
Roadmap
More measurable before more complex.
01 · EvidenceThe empirical loop is closed under REB 2025-8251. Next come GPU timings, then vocabulary.
8-bit ROWSIM title card for the roadmap order: loop closed, then measurement, then vocabulary.