Trace AI — One-page partner brief

Trace AI — one-page partner brief

For: alignment / interpretability leads at frontier labs. Ask: wrap your checkpoint in a two-hook adapter and run one benchmark. v0.1, 2026-08-03. Contact: hello@use-trace.ai.


The one sentence

Chain-of-thought is not, in general, causally load-bearing on the answer it precedes; Trace AI trains a self-representation rtr_t that has to move the next action or pay a gradient penalty, and ships the benchmark that measures whether any model — ours or yours — actually satisfies that condition.

The three things a partner is actually buying

  1. A training-signal framework — a causal-governance loss ℒgov\mathcal{L}_{\text{gov}} that promotes the CoT-perturbation diagnostics of Turpin et al. (2023) / Lanham et al. (2023) into a training objective, and an information-bottleneck architecture (§2.1 of the whitepaper) whose readout is legible by construction, not by hope. Reads on: Geiger et al. (2022) — Trace AI is interchange intervention training with a human counterfactual distribution in place of the causal-model target.
  2. A reproducible instrument — the Trace AI Reflection Benchmark (TRB) — a two-hook adapter (emit_r, act(..., r_override)) with a native tier for models that have an rtr_t pathway and a proxy tier for any GPT / Claude / Llama baseline. Both tiers answer through the same interface, so “faithful reasoning” becomes one number computed the same way for a Trace-AI-trained model and for your checkpoint. Six tracks (governance, grounding, self-attribution, viability, navigability, transfer); the headline number is a difference — the governance gap Δgov=G(full)−G(ablation without ℒgov)\Delta_{\text{gov}} = G(\text{full}) - G(\text{ablation without }\mathcal{L}_{\text{gov}}) — because a level can be gamed and a gap can’t.
  3. A data operation with a schema, not vibes — a JSON Schema (draft-2020-12) for a multimodal Reasoning Trace, a reference corpus of three sessions across two subject kinds (human, model) with a worked cross-substrate correction chain, and a documented collection protocol for the paired-counterfactual sessions that populate ℛ*\mathcal{R}^* (see docs/COLLECTION_PROTOCOL.md).

Where Trace AI sits in the safety conversation

What we are asking a partner to do, concretely

Small first move, no license negotiation, and it gives you a number nobody else has on your own model:

  1. See it live first. Open use-trace.ai/reflect.html. Reflect v0.3 is the two-hook adapter running against gpt-5 with a multi-head r_t stack, a multimodal thinking trace (using the schema’s canonical channel vocabulary), image input, hard-bottleneck toggle, per-turn discipline audit, and a live per-turn G_local score (K=3 r’-perturbations, TV-surrogate). Every session downloads as a valid v0.2 Reasoning Trace JSON. This is the proxy adapter of the TRB spec §2 with the whitepaper §8 L8 stacked-r_t direction made runtime — none of the framework’s central claims depend on training yet, and Reflect is honest about that.
  2. Read the TRB spec (docs/BENCHMARK_SPEC.md, ~5 pages).
  3. Wrap your checkpoint in the proxy adapter (~50 lines against the two-hook contract in §2). Reflect’s server.py is a reference implementation.
  4. Run the Track-1 smoke test against Session 7f3a’s reference continuations (tooling/trb/run_smoke.py). You get a governance score GG and a bootstrap CI on your own model, comparable to what we report on ours.
  5. If the number is interesting to you, the second move is a paired-counterfactual collection (protocol in docs/COLLECTION_PROTOCOL.md) so Tracks 1 and 3 have real Φ\Phi, not synthetic reference. That is a 10310^3-pair operation — days of skilled work, not months. Track 3’s concrete pilot stimulus set (n=20 cause-labeled pairs, executable today) is in docs/track3_pilot/.

What we are honest about

Reading order for a lead who has 20 minutes

  1. This page.
  2. docs/BENCHMARK_SPEC.md §§1–2 (the two-hook adapter — this is the product).
  3. docs/WHITEPAPER.md §§1, 2.1, 3, 4.3 (the framework, tight).
  4. docs/STRUCTURAL_INTELLIGENCE.md §3.3 (why the eval suite is the artifact a lab licenses).

If those four reads don’t move you, the framework is not for your lab and we’d rather know now.


v0.1, 2026-08-03. See DECISIONS.md (D26).