Trace AI — Paired-counterfactual collection protocol

Paired-counterfactual collection protocol (v0.1)

Status: protocol, 2026-08-03. Design document for the corpus-collection operation that unlocks TRB Tracks 1 and 3. Reads on: whitepaper §4.4 (sources of counterfactual continuations, esp. source #2); BENCHMARK_SPEC.md §3 (Track 1 governance, Track 3 self-attribution); AGENCY_AND_SELECTION.md §4 (the six functional criteria).

The benchmark’s central number — the governance gap Δgov\Delta_{\text{gov}} — cannot be computed on synthetic counterfactuals alone. It requires paired continuations produced by two humans (or the same human under two conditions) who differ only in a labeled slot of rtr_t. This document specifies how those sessions are collected, at a level a research assistant can execute without further design work.


1. What we are collecting, in one sentence

Pairs of Reasoning Trace sessions in which two subjects meet the same stimulus with a deliberately different value of one rtr_t slot, holding every other variable — stimulus, modality, elicitation prompt, session length — as constant as the format allows.

2. The unit: a counterfactual pair

A counterfactual pair is a 2-tuple of sessions (SA,SB)(S_A, S_B) satisfying:

Every pair is annotated with a pair_id that both sessions carry in their provenance block (pair_id, pair_role: "A" | "B", manipulation: {slot, value_A, value_B, method}).

3. Cause-labeled outcomes for Track 3 (self-attribution)

Track 3 requires sessions where an outcome is caused either by the agent’s own intervention or by an exogenous environment change, indistinguishable from the outcome alone. We collect these as an additional slot on the same pair schema:

Post-session, the subject is asked to attribute each of KK salient events to self or environment — this gives us the label for Track 3’s balanced-accuracy score. The subject’s attribution is elicited before they see the app’s log, and the app’s log is the ground truth.

The germ of this in the current corpus is the two natural instances in Session 7f3a (the 2+2=2 → 2+2=4 demonstration-vs-endorsement, and the 3:44 pm → 3:46 AM clock-check reversal). Those are the intuition; the protocol above operationalizes them.

4. Session recipe

A single session in a pair, ~45–75 minutes:

  1. Onboarding (5 min). Consent script (see §7); short priming statement that is the manipulation for this pair’s role; instructions to open the recorder.
  2. Stimulus presentation (1 min). Byte-identical stimulus is displayed. Timer starts.
  3. Free work (30–60 min). Subject produces the output the stimulus asks for, self-annotating in the trace-capture app as they go. The app schedules 3–6 elicitation probes at intervals sampled from the trace-length distribution of the existing corpus.
  4. Cause events (2–4 during free work). The app injects the pair’s cause-labeled events (§3) at pre-computed timestamps. Injection is silent — no visual cue that this is an experimental event.
  5. Post-hoc attribution (10 min). Subject reviews their trace timeline (without the intervention labels) and marks each highlighted event as self-caused or environment-caused. This produces the Track 3 label file.
  6. Second consent (2 min). Subject re-consents to the trace for the specific release scope. Withdrawal here still pays the collection fee and destroys the trace.

5. What lands on disk (per pair)

data/pair_<pair_id>/
├── A/
│   ├── trace_<session_id_A>.json         # standard v0.2 trace + provenance.pair_id, pair_role
│   ├── audio/…                           # per the schema's channels block
│   ├── attribution_labels.json           # {event_id: "self" | "environment"} produced in §4.6
│   └── consent/                          # signed second-consent + release-scope record
├── B/                                    # same layout, role "B"
└── pair.json                             # {pair_id, stimulus_sha256, manipulation, collected_at, collector}

pair.json is the single source of truth about what was manipulated; the two per-session files carry a reference to it so a session cannot be interpreted out of pair context.

6. Sample-size targets

8. What the pilot buys, concretely

The pilot (n=20n=20) is not a training set. It is the artifact that turns TRB from a spec into a document a partner can read and a number they can put next to a checkpoint:

9. What this does not do


v0.1, 2026-08-03. See DECISIONS.md (D27) for provenance.