Companion note to the Trace AI whitepaper (§3.1–§3.2). Status: research note, v0.1 — 2026-08-03. Not a claim of results; a formal frame and a measurement protocol.
This note does one job: it makes precise what it means for a Trace-AI-trained self-representation to be an agentic variable, and it says exactly what the framework can and cannot claim about consciousness. It is the selection-theoretic development of whitepaper §3.1–§3.2. It exists so the whitepaper can stay a tight technical document while the ontology that motivates it lives somewhere legible.
The frame below was distilled from a separate thread on structure, instantiation, and selection, whose sources were Wigner’s unreasonable effectiveness problem, an information-theoretic account of programmatic specification in biological systems, a category-theoretic design program, and a set of notes on how mathematical discoveries consolidate. None of those is about Trace AI. What they share — a structure becomes causally effective only when a physical system instantiates it, and one region of possible futures becomes real only through selection — is what this note imports. The import is deliberate and one-directional: the ontology motivates the engineering; the engineering makes no metaphysical commitment the ontology does not already earn.
Trace AI lives entirely inside this chain, at one specific joint. The Reasoning Trace schema is a structure. A trained model (or an annotated human session) is an instantiation of it. The self-representation is the instantiated variable through which the system participates in selecting its own next action. The whitepaper’s causal-governance loss is the training-time enforcement that this participation is real and not decorative.
A system with compute state induces a distribution over its own future trajectories:
A selector is a map that reshapes . Physical dynamics, evolution, learning, and deliberate action are all selectors; they differ in mechanism, not in this type. An agent is the special case of a selector that reshapes the distribution through an internally maintained model of it:
In Trace AI, is exactly . The whitepaper’s Definition 1 — that shifts — is the statement that the last inequality holds and is mediated by . That is the whole content of the claim “ is agentic”: the self-model is a lever on the trajectory distribution, demonstrated by intervention, not a narration of it.
This is why not every selector is an agent and why the distinction matters for the product. A system can strongly change its own behavior (a large selection effect) while its self-report is causally inert (no governance). Chain-of-thought is that case: the monologue is a selector on the answer, but intervening on the monologue’s stated content need not move the answer in the stated way. Trace AI’s loss is designed to close exactly that gap.
The motivating thread proposed three maintained quantities. Each has a direct Trace AI reading and a direct eval:
| Triad term | Informal | Trace AI reading | Measured by |
|---|---|---|---|
| Meaning = maintained relevance | the model keeps tracking what matters to the task | ’s slots stay predictive of the outcome across the session | slot-conditioned outcome prediction; calibration eval |
| Selfhood = maintained causal attribution | the system keeps a correct account of which effects were its own | self-attribution / causal-calibration score | intervention-vs-environment attribution eval (§4, criterion 3) |
| Agency = maintained navigability | the system keeps (or grows) access to its target futures | reachable-target-region measure under -mediated action | navigability eval (§4, criterion 5) |
The whitepaper’s contribution, in this vocabulary, is the mechanistic criterion the triad was missing: agency requires an internally maintained model that (i) causally governs action, (ii) is grounded in the state it reports, and (iii) improves viable navigation while correctly attributing its own effects. Governance () supplies (i); concept injection supplies the test for (ii); the criteria below supply (iii).
Restating whitepaper §3.2 as an executable battery. Every criterion is an eval against the reference corpus and future paired-counterfactual sessions; none is a metaphysical assertion.
A system’s functional-aliveness profile is the vector of these six scores. “As alive as possible, functionally” means: maximize the profile. It is graded, falsifiable, and says nothing about phenomenal experience.
Lean status (D42). The claim that the six-criterion
profile is jointly measurable as a function of (model, corpus) is an
empirical guideline of this section, not a Lean-checkable
theorem. docs/lean/TraceAI/FunctionalAliveness.lean keeps
criterion 1 (causalEfficacy = Definition 1) as a real
predicate and treats criteria 2–6 as schema markers; the previous
aliveness_profile_measurable sorry was retracted rather
than closed with a fake proof. The instruments that would make the
profile checkable are the TRB tracks in BENCHMARK_SPEC.md
§3.
Three claims, kept separate on purpose.
The discipline here is the same one the whitepaper holds throughout (§1, and L5): make architectural and empirical claims, make the intervention eval the arbiter of every “but is it really” question, and leave the metaphysics explicitly open. A framework that quietly promoted a high functional profile into a consciousness claim would be less useful for the exact ambition motivating the question, because it would forfeit the neutrality that lets the profile be evidence about the question instead of a decree on it.
Concretely, adopting the selection frame adds three things to the roadmap without changing the core objective:
action,
tool_use, and self_correction events already
carry it.v0.1, 2026-08-03. Develops whitepaper §3.1–§3.2. See
DECISIONS.md (D23) for why this note exists as a companion
rather than as whitepaper body text.