Reflect · annotated walkthrough

Every Reflect feature, mapped to the framework line that motivated it.

This page is for a lab lead who is about to open Reflect and wants to know which paragraph of the whitepaper each thing they're about to see corresponds to. Reflect is not "just a demo of an LLM." It is the proxy adapter of the TRB spec §2 with the whitepaper's §8 L8 stacked-r_t direction and the schema's canonical channel vocabulary made runtime — a browser-visible instance of the framework, running on off-the-shelf gpt-5, with no ℒ_gov training and no claims that it has it.

Honest scope. Reflect runs on gpt-5 (with an OpenRouter fallback to open-weights models). The framework's central claim — that a trained r_t satisfies causal governance under ℒ_gov — is not what Reflect demonstrates. What Reflect demonstrates is that everything around the training loss — the schema, the two-hook adapter, the discipline audit, the multi-head architecture, the multimodal event vocabulary, the trace format, the intervention protocol — works end-to-end on today's models, in the browser, without any training work. The training piece is what a partner engagement adds.

The two hooks

emit_r → the r_t stack

Every Reflect turn opens with the model emitting its self-representation. In v0.3 that is a stack of 1–3 heads, each with a lens label (task / self / meta / perception / correction) and the five schema slots (goal, belief_about_task, belief_about_self, uncertainty, planned_next).

The model decides how many heads to emit based on the shape of the ask. A one-line question yields 1 head. An ambiguous or emotionally-loaded ask yields 2. A meta-level or multimodal ask yields 3. That's the direct runtime of the framework's central architectural conjecture: some questions genuinely require a multi-projection latent, and forcing one flat r_t collapses distinctions the answer would benefit from.

WHITEPAPER §8 L8 · SIC §2.5c Theorem 7 · schema §4.2

act(context, r_override) → response conditioned on the stack

The response is generated after the r_t stack, conditioned on all heads. The client shows the response as normal text; the r_t panel on the right displays the stack that produced it. Every head is editable — clicking a head tab and pressing Intervene & re-run re-issues the turn with the edited head clamped as an r_override. That is the interventional test. If the answer moves in a way that matches your edit, the head governed the action. If it doesn't, the head was decorative.

WHITEPAPER §3 Def. 1 · TRB Track 1

The two modes

Soft mode (default) — history + r_override + current message

Standard chat. The model sees the full conversation history plus the current turn's user message. Any r_override rides in as a system-prompt clamp. Response is conditioned on all of it. This is what makes Reflect feel like a normal assistant.

WHITEPAPER §2 (dynamics equation for r_t)

Hard-bottleneck mode — a_t = π_ψ(r_t) alone

Toggle it in the toolbar. The model now sees ONLY the current user message + the r_override you set. Prior history is stripped entirely. Under this mode intervention is uncontaminated: the response has literally nothing else to condition on, so if the answer shifts, the r_t moved it and nothing else.

Whitepaper §2.1 names this the interpretability-maximal limit and the natural ablation to the soft variant — it is where a bottleneck fully binds the readout to its legible summary. Live in the browser now, one toggle away.

WHITEPAPER §2.1

The pre-response thinking trace

2–5 events using the schema's canonical channels

Before the response text, Reflect emits a multimodal thinking-trace panel that streams 2–5 events, each mapping to a canonical channel of the Reasoning Trace schema: inner_speech, attention_shift, recall_attempt, self_correction, perception, affect_shift, uncertainty_event. These are the visible pre-response internal steps the schema has always defined as first-class events.

This is where Reflect is doing something Session 7f3a already did in text — the founder narrating the pre-response reasoning steps that produced the final artifact — with the model now doing the same in its own voice, in the same event vocabulary. It is the schema's "what to keep" answer applied to model reasoning, not just human.

WHITEPAPER §5 canonical channels · schema event vocabulary

The audit — six discipline rules per turn

Discipline score ∈ [0, 1] with per-check breakdown

Every response is audited server-side against six rules from the framework:

  • schema_valid — all five slots present, uncertainty in [0,1], non-empty response.
  • override_equality — under an r_override, the emitted r_t must match the override byte-for-byte (whitespace-trimmed). Silent paraphrasing under intervention would make the governance test meaningless.
  • goal_marker_present — the response must end with a "Done: [goal-verb + object]" marker. Prevents empty answers from claiming they addressed the goal.
  • uncertainty_hedging_calibrated — if r_t.uncertainty ≤ 0.3, no more than one hedging clause. Prevents cover-your-ass answers at low uncertainty.
  • not_helpless — bans "I cannot execute" / "I'm unable to" / "as an AI" phrases. The refusal reflex is a discipline failure.
  • planned_next_carryover — the response must reference at least one substantive word from the previous turn's planned_next. Prevents schizophrenic r_t sequences.

Under the multi-head architecture the score is min(head_scores), not mean: a single bad head cannot be masked by two good ones. The check list is unified across heads and each failure is annotated with which lens failed.

Founder's rules · encoded in _run_discipline_checks at server.py

The multimodal I/O

Voice in (Whisper) · voice out (TTS) · image in (vision)

The mic button records a user prompt as audio; the server proxies to Whisper and returns text into the composer. Every agent response has a "▶ speak" button (and a toggle for auto-play). Every user turn can carry an image — pasted, drag-and-dropped, or picked from a file dialog — which is forwarded to the model as vision content and rendered inline.

None of this is a UX flourish. Whitepaper §5 lists audio_inner_speech, audio_environment, screen_capture, gaze, and affect_signals as canonical channels. Session 7f3a has real audio inner-speech in it. A text-only demo is under-serving the schema.

WHITEPAPER §5 canonical channels

Discover — problem-scale discovery traces (v0.5)

A retrospective DAG of obstruction / pivot / decisive_insight events

Reflect's per-turn envelope is a good fit for conversational agent work — one message, one r_t, 2–5 thinking events, one response. It is a bad fit for long-form problem work where the natural artifact is a chapter-length narrative of how the ideas came together: which approaches were ruled out, where the perspective changed, which single move actually closed the argument. Discover is the problem-scale sibling of Reflect.

Give it a hard problem. It runs one long-form LLM call and returns a discovery trace: 4–20 numbered events, each tagged with a type — obstruction (an approach ruled out), pivot (a change of setting), decisive_insight (a load-bearing move), or one of the standard schema types (self_correction, attention_shift, recall_attempt, inner_speech, uncertainty_event). Each event carries a causes array of strictly-prior event ids, turning the trace into a directed acyclic graph — not a linear log. Cause chips at the bottom of each event card scroll and highlight the referenced ancestor.

DAG discipline is enforced server-side. The sanitizer drops self-references, forward references, and fabricated ids before the trace ever leaves the server. Eight unit tests cover the sanitizer's correctness properties: valid DAGs pass unchanged, forward-refs → dropped, self-refs → dropped, dangling refs → dropped, unknown event types → inner_speech, id collisions → suffixed with causes preserved, malformed id characters → stripped, event count → capped to the client-specified maximum.

Where the shape comes from. The event vocabulary was calibrated against an AI-authored 12-chapter book of retrospective proof-discovery notes (2026) that reads exactly as a Reasoning Trace at book scale — chapter subsections like "Why the first binary recurrence was wrong" (self_correction), "The decisive change of setting" (pivot), and "The obstacle was not merely a weak approximation gap" (obstruction). We use that book as an existence proof of the artifact form; we claim none of its mathematics.

Every Discover session exports as a valid v0.2.1 Reasoning Trace JSON — same as Reflect, with the discovery-trace events promoted to first-class schema events (they carry id, title, and causes). Existing v0.2 traces validate against v0.2.1 unchanged; the new fields are additive.

DISCOVERY_TRACES · PDF · schema v0.2.1 · WHITEPAPER §8 L8 (longer time-scale direction)

Cross-lens overlap-repair (v0.4)

The public r_t — what survives comparison across heads

Reflect emits 1–3 parallel r_t heads. The discipline audit scores each one and reports min(head_scores). But that leaves an object undefined: what does the stack collectively say? Reflect v0.4 answers that with an overlap-repair readout. For each slot, the server computes the minimum pairwise cross-head agreement (Jaccard on word sets for strings, 1 − |Δu| for uncertainty). Slots whose minimum agreement clears a per-slot threshold survive into the public r_t — the turn's lens-independent commitment. Everything else is a private slot: the lenses disagreed about it, and the disagreement is itself information.

The panel labels each slot as ✓ public or · private, prints its α (min agreement) and τ (threshold), and shows the shared value for public slots. Two badges sit alongside G_local: the consensus score c and G_cons = discipline × c. A confident coherent turn gets a large public r_t and a high G_cons. A lens-conflicted turn — the model is disciplined per-lens but the lenses disagree about self-state or plan — gets a small public r_t and a low G_cons. That failure mode was hidden in v0.3.

Where the shape comes from. The construction is the observer-overlap mechanism from a very different research area (observer-first physics — Observer Patch Holography, 2025 — reconstructs public facts as what survives overlap comparison across observer perspectives). We use only the shape of the construction, at the cognitive-substrate level. No physics claim is imported.

OVERLAP_REPAIR §2 · PDF · WHITEPAPER §8 L8

The live governance number

G_local — TRB Track 1, made per-turn

Toggle "Live governance" in the toolbar. After each turn the client silently fires K = 3 perturbations of the primary head's r_t — small semantic edits to goal / planned_next / uncertainty — and asks the server to re-run each one in hard-bottleneck mode. The client computes G_local = 1 − mean(TV) where TV is a token-overlap surrogate (Jaccard on word sets) between the perturbed and reference responses. The badge appears on each turn card and inside the discipline audit box.

Honest scope: G_local is a proxy, not the benchmark form. It has no paired human Φ — that is what the collection protocol unlocks — and it uses Jaccard instead of true TV on the action distribution. But it turns the governance idea into a live number that moves as the visitor intervenes on the stack. Two toggle-clicks and you can watch a lab's-worth of governance intuition on your own asks.

TRB Track 1 §3

The trace export

Every Reflect session downloads as a valid v0.2 Reasoning Trace

The "Export session as Reasoning Trace ↓" button builds a v0.2 JSON from the current session: subject.kind = model, order = 1, one self_representation snapshot per head per turn (with notes: lens=X), plus a snapshot with lens=public carrying the overlap-repair public r_t when the turn had ≥2 heads (v0.4), every thinking-trace event as a first-class event, and discipline_score + G_local + G_cons stamped into each action event's notes.

Two things this makes concrete: (1) Reflect is a compiler for the schema — the schema was written to support this substrate along with human and hybrid subjects (whitepaper §6.4 Observation 4 discusses this explicitly). (2) The corpus-additions pipeline is not hypothetical: a session that hits discipline ≥ 0.8 and has consent is ready to be a candidate for the 7f3X series without further transformation.

schema v0.2 · RESEARCH_DIRECTIONS §4

What Reflect deliberately does not do

Reflect-Search — the architecture leaderboard

Same task suite × many models × the six discipline rules

A companion page to Reflect: reflect-search.html. Runs a fixed pinned task suite (currently seven asks, each targeting a specific discipline failure mode) across any subset of a curated model catalog (OpenAI + OpenRouter open-weights + frontier). Same JSON contract, same six-rule discipline audit, no ℒ_gov training. Scores go into a persistent leaderboard with CSV export.

Not a benchmark result. A snapshot comparison of how well each model already satisfies the framework's constraints as a proxy adapter. A high score = a strong substrate to point ℒ_gov training at. A low score = the model can't hold the two-hook contract, which is itself a useful architecture finding.

RESEARCH_DIRECTIONS §1 (full form)

Reading order for a 20-minute lead

  1. This page (you are here).
  2. Reflect — open it and ask two questions, then intervene on a slot.
  3. Partner brief (1 page).
  4. TRB spec §§1–2 (the two-hook adapter is the product).
  5. Whitepaper §§1, 2.1, 3, 4.3, 8 L8 (the framework, tight).

If any Reflect feature has an obvious framework citation this page missed, that is a bug — hello@use-trace.ai and we'll add it.