Draft v0.1 — 2026-08-04. Jawaun Brown¹, with Claude
Opus 4.7 as co-drafter. ¹ Trace AI. Correspondence:
hello@trace.ai.
The Trace AI whitepaper (v0.1) proposes a stacked- architecture (§8 L8) in which a single turn emits multiple parallel self-representations, each along a distinct lens (task / self / meta / perception / correction). Reflect v0.3 implements this at runtime — but treats the heads as independent projections: each head is scored, and turn-level discipline is over head scores. In this note we add an overlap-repair layer on top of the stack. We define a per-slot cross-head agreement measure, an induced public consisting of exactly the slots that survive agreement above threshold, and a governance-under-consensus score that combines the existing discipline score with the fraction of that survives comparison across lenses. The formulation is a direct application of an idea from a different research area — the observer-overlap mechanism that has been used to define public physical facts in observer-first physics programs — to the cognitive-substrate level at which Trace AI operates. The construction ships in Reflect v0.4 and adds no new training signal; it exposes a signal that was already latent in the multi-head envelope.
The stacked- architecture (whitepaper §8 L8) posits that different lenses of a model’s self-representation should be independent projections of the same underlying computation:
For a well-formed turn, these lenses should agree about facts about the model itself (uncertainty, self-belief, planned next action) while legitimately differing about facts about the ask (goal, belief about task). The current Reflect envelope makes both kinds of variation visible but does not measure them. In particular it does not compute the public : the subset of the self-representation that is stable across lenses, and that a downstream consumer could safely treat as the model’s lens-independent commitment for the turn.
The proposal here is minimal: add that measurement.
Fix a turn with an emitted stack of heads, each with a lens label . For each pair and each slot define a per-slot agreement :
goal, belief_about_task,
belief_about_self, planned_next):
,
the Jaccard index over lowercased word sets with words of length
filtered out. This matches the _jaccard surrogate that
already stands in for TV distance in the local governance path
(server.py:_jaccard).uncertainty (a real in
):
.For heads, define by convention (a single-head turn is trivially public).
Fix per-slot thresholds . The current defaults are:
| Slot | |
|---|---|
goal |
0.30 |
belief_about_task |
0.30 |
belief_about_self |
0.35 |
uncertainty |
0.70 |
planned_next |
0.30 |
These are calibration values, not theorems. The rationale: the
task slots (goal, belief_about_task,
planned_next) are expected to differ across lenses (each
head is looking at a different thing), so their thresholds are lower;
the self slots (belief_about_self,
uncertainty) are about the same underlying entity — the
model — so their thresholds are higher.
Slot is public on the turn iff . The public is the map from public slots to their (any) shared value; when the slot passed on Jaccard, the public value is taken from the first head; when the slot passed on uncertainty overlap, the public value is the mean.
The consensus score is the fraction of slots that are public: .
The governance-under-consensus score is
where
is the existing per-turn discipline score (min over head scores of the
six-rule audit; server.py:_run_discipline_checks).
For : , . The extension is inert for single-head turns, as it should be.
A turn with high but low is the informative failure case: the model is internally disciplined per lens but the lenses disagree — the model holds inconsistent beliefs about itself across projections of the same turn. That is a failure mode the existing envelope hides.
§8 L8 states the stacked- direction: “a stack of -heads , jointly compressed and jointly supervised, one per interpretive lens… , and the causal-governance condition must hold jointly over the stack.” That formulation gave the architecture but not the readout: how does one summarize what the stack says publicly?
The overlap-repair layer answers that. The public is a natural readout of the stack:
It is a strict subset of every individual . It carries no slot that any pair of lenses disagreed about. The complement — the private slots — is where the lenses interpret the turn differently, and that disagreement is itself information.
For the training objective: the causal-governance loss (whitepaper §4.3) is defined per head; a stack-level extension would apply it to the public (so that intervention on a slot that survived overlap moves the action at least as much as intervention on any single head). This note does not develop that; it exposes the object.
The mechanism — public facts as what survives overlap comparison across observer perspectives — is imported from an unrelated area. Recent work on observer-first physics (e.g. Observer Patch Holography, 2025) posits bounded systems that read part of themselves, keep records, and repair disagreement, with objective facts emerging only from what survives overlap across observers. That program uses the mechanism to reconstruct spacetime and gauge structure from an observer patch net. We use exactly the mechanism, at the cognitive-substrate level, on the multi-lens -stack of a single agent.
The mathematical content transferred is small: it is the shape of the construction (per-slot agreement, threshold, public subset) rather than any specific theorem. We do not claim any of that program’s physics results follow. What we claim is that the architecture — treat parallel perspectives, compare them per-slot, keep only what agrees — is a general pattern that applies wherever multiple projections of a single underlying representation are available.
Reflect v0.4 emits a new consensus field alongside the
existing discipline field on the
/api/reflect/turn result envelope:
{
"r_t_stack": [ … ],
"discipline": { "score": d, "checks": [ … ], "per_head": [ … ] },
"consensus": {
"score": c,
"governance": G_cons = d * c,
"public_r_t": { … slots that survived overlap },
"private_slots": [ list of slot names where lenses disagreed ],
"per_slot_agreement": {
"goal": α_goal_min,
"belief_about_task": α_task_min,
"belief_about_self": α_self_min,
"uncertainty": α_unc_min,
"planned_next": α_next_min
},
"thresholds": { … per-slot τ_s (defaults above) },
"single_head": (bool — c is 1 by convention if only 1 head)
}
}
The client renders a Public
panel below the head tabs, showing the slots in public_r_t
and marking each slot in private_slots with a short reason
(e.g. "goal: lenses split, min α = 0.14"). A
badge sits next to the existing
badge on the turn card.
Nothing in the training pipeline changes; the field is a readout of what the model already emitted.
belief_about_self more than on what belongs
in goal?r_overrided on one head, the stack’s other heads should
still update their agreement with it. We currently compute agreement
post-override on all heads uniformly; the interventional variant is
future work.v0.1, 2026-08-04. Shipped as Reflect v0.4 (server.py + reflect.html). Cross-references: WHITEPAPER §8 L8; BENCHMARK_SPEC §5; DECISIONS D36.