Jawaun Brown — human author and research director. Claude Code (agent) — experiment code, analysis, and manuscript production under direction and review. Date: 2026-08-03 (v4 + Observatory Instrument 9, integrated 2026-08-03). Status: five theorems + one conditional theorem + nine exact executable instruments. Existence, cross-task stability, discrete learnability, and continuous learnability at resolution derived (Theorems 1–6); Theorem 2 (rate–distortion) witnessed exactly by Instrument 7; Theorem 7 (linear-ICA identifiability) resolves SIC-C-c positively inside the linear-ICA hypothesis class (Instrument 8) and the sparse-mechanism / IMA class (Gresele et al. 2021) gives a second positive resolution (Instrument 9) — two distinct inductive-bias classes now provably earn escape from Theorem 6’s -covering exponential-in- bound. What remains genuinely open is only the general programme: which other inductive-bias classes admit their own polynomial-in- theorems (iVAE with auxiliary variables, interventional CRL, …) — a mainstream question in identifiable-representation-learning theory, addressed one class at a time.
Live interactive Observatory: structural-observatory-production.up.railway.app
— all nine instruments rendered as interactive panels with committed
numerical results.
Companion PDFs: sic_paper.pdf is the
v4 (latest); prior versions preserved as sic_paper_v2.pdf,
sic_paper_v3.pdf,
sic_paper_v4.pdf.
Reads on: STRUCTURAL_INTELLIGENCE.md
(plain-English pointer to the same object with the Trace AI lens), WHITEPAPER.md (the training-signal
application), BENCHMARK_SPEC.md (TRB Track 3
is Instrument 3 of this paper), AGENCY_AND_SELECTION.md
(the selection frame those instruments live in).
Why it exists here. The
STRUCTURAL_INTELLIGENCE.md note tells the story in plain
English; that told-in-plain-English claim needed a paper with theorems
and exact-solvable witnesses under it, so the note doesn’t have to carry
weight it can’t. This is that paper.
We study a single latent object that recurs across five otherwise unrelated works — a category-theoretic framework for materials design, an information-theoretic limit on the programmatic specification of biological systems, a corpus of machine-found mathematical proofs with their discovery notes, Wigner’s essay on the unreasonable effectiveness of mathematics, and a structural-realist ontology. The common object is a stochastic fibration with a compiler: a coarse-graining together with a kernel whose support lies in the fibre . We derive the object — it is not posited: for any statistical task on a standard Borel space, the minimal sufficient -algebra of Halmos and Savage yields as its quotient and regular-conditional pair (Theorem 1); Shannon’s rate–distortion pair parameterises the same object at a distortion budget (Theorem 2, now witnessed exactly by Instrument 7 on uniform and Bernoulli sources to ); the construction is the unit/counit of an adjunction (Proposition 3). Cross-task stability — a single quotient that is sufficient for a whole task family — holds iff the family admits a shared Markov screen, i.e. a common latent generator (Theorem 4, conditional). Discrete-case learnability is also a theorem: with the task family separating and the compiler’s fibres balanced (), empirical common-sufficient clustering recovers with probability from samples (Theorem 5). The continuous-case extension at resolution is also a theorem (Theorem 6): bound — polynomial in at fixed , provably exponential in at fixed without further inductive bias. Adding one specific bias — independent non-Gaussian latents with a linear mixing — resolves that exponential-in- cost inside the linear-ICA hypothesis class: the classical Hyvärinen–Oja identifiability result becomes Theorem 7 in our framework’s language, with Instrument 8 witnessing polynomial recovery (Amari index at for ). Eight exact, deterministic, unit-tested instruments now witness each theorem on solvable cases and exhibit the dissociations the framework predicts.
Let be a space of concrete realizations and a space of structures, functions, or observables. Two maps carry the construction:
This sharpens the adjunction of the companion synthesis: (coarse-grain) and is a stochastic section of (realize). The biology paper of Kiiskinen–Kivinen–Rivas is exactly this construction — the genome names , the physics substrate is the compiler that “computes the samples,” and selection acts on the ensemble statistics — and its coarse-graining threshold proves that below a critical resolution the fibre is too large to address with the specification budget.
A note on rigor. Cross-domain resemblance is not isomorphism. The honest hierarchy of sameness runs: isomorphism bisimulation functor natural transformation adjunction / Galois connection Morita-like equivalence simulation-at-a-resolution. Most relations among the source works are adjunctions, simulations, and shared diagram shapes, not object-level isomorphisms; every proposed connection must state what kind of sameness it is, what the map forgets, and what would have to be proved to make it a theorem.
Let be a standard Borel space, a dominated family of probability measures on , and a random variable whose distribution depends on (equivalently: indexes the task). The Halmos–Savage sufficiency theorem gives a sufficient -algebra ; under mild regularity (Bahadur 1954; Lehmann–Casella) there is a minimal sufficient -algebra , unique up to -null completion.
Define:
Theorem 1 (Existence of the master fibration). is a stochastic fibration with , and four of the five clauses of the Structural Intelligence Conjecture hold as theorems:
The Fiber Finder (Instrument 1, §4.1) is an exact computational witness.
Let on and a distortion. Shannon’s rate–distortion function
is attained by an encoder and decoder marginal .
Theorem 2 (RD parameterisation of the master fibration). For every distortion budget , the RD-optimal pair defines a stochastic fibration ; the family is a one-parameter deformation of the sufficiency fibration. At the encoder is minimal-sufficient; as grows the fibres grow and the specification shrinks along the RD curve. The Kiiskinen–Kivinen–Rivas biology threshold is for the domain’s distortion measure.
Let and be the categories of probability measures on and . The sufficient-statistic construction is the unit/counit pair of an adjunction:
Proposition 3. is the unit/counit pair of . SIC clauses (1)–(4) are the adjunction’s triangle identities specialised to the sufficiency reduction; the categorical framing adds compact language, not new content.
Clause (5) of the SIC — stability of across substrates and contexts — is not implied by single-task sufficiency: different tasks have different minimal sufficient statistics in general.
Theorem 4 (Cross-task stability under latent generation). Suppose there exists a random variable and, for every task in a family , a conditional independence
Then is a sufficient statistic for every simultaneously.
Proof. For each , sufficiency of for from is exactly the Markov property ; the family case is the pointwise conjunction.
Corollary (Equivalence). A task family admits a common sufficient statistic strictly coarser than iff there is a decomposition with and every factoring through .
Theorem 4 turns SIC clause (5) into a claim about the world — the joint law of and the task family — rather than about intelligence. The Cross-Task Sufficiency instrument (Instrument 4, §4.4) witnesses it exactly on a 4-bit Boolean world.
Theorems 1–4 do not immediately derive that a finite adaptive system discovers from data. In its unrestricted form this claim is false (Locatello et al. 2019): for smooth continuous and no auxiliary information, no unsupervised algorithm identifies a factored latent up to trivial transformations without inductive bias. Task families supply the identifying auxiliary information, and — quantitatively enough — reduce learnability from a conjecture to a theorem in the discrete case.
Setup. Let be finite; a distribution on with min-fibre mass for the true partition , ; a family of deterministic tasks factoring through ; the family separates (for every , some has ).
Algorithm (empirical common-sufficient clustering). Given i.i.d. samples with labels:
Theorem 5 (Discrete learnability). For any ,
Proof. Injectivity of gives that once every fibre of contains at least one observed sample. For each , . Union bound over fibres: for .
What the theorem does not do. It is stated for finite and deterministic tasks; it requires the separation assumption (without it the recoverable object is strictly coarser than — namely the CSS of the observable task family); it requires fibre balance through .
The honest form of the continuous claim is “recovery up to resolution ”, and it reduces cleanly to Theorem 5 by -covering.
Setup. bounded with diameter ; a minimal -net; the induced partition into Voronoi cells; the -quantised true partition; fibre balance and separation both at scale .
Theorem 6 (Continuous-case learnability, at resolution ). For any , recovers exactly with probability from
For bounded, , giving
Proof. Apply Theorem 5 to the discretised experiment.
Fundamental limit. Uniform sample complexity polynomial in alone — dropping the exponential dependence on at fixed — is impossible without additional inductive bias on . Escaping the curse requires linearity (linear ICA), sparsity (independent-mechanism analysis; Gresele et al. 2021), exponential-family conditional latents with auxiliary variables (iVAE; Khemakhem–Kingma–Monti–Hyvärinen 2020), or interventional data on (causal representation learning; Ahuja–Mahajan–Wang–Bengio 2022). The next section resolves SIC-C-c positively inside the first of those classes.
Theorem 6 says continuous-case sample complexity is exponential in without inductive bias. Adding one specific bias — linear generative model with independent non-Gaussian latents — resolves SIC-C-c positively inside that hypothesis class. This is a classical result from Hyvärinen’s independent component analysis (ICA) line (Comon 1994; Hyvärinen 1999; Hyvärinen–Oja 2000), restated here in the framework’s language and witnessed numerically as Instrument 8.
Setup (Theorem 7).
Theorem 7 (Linear-ICA identifiability, classical). Under the linear-ICA setup, the un-mixing matrix is identifiable up to permutation and coordinate-wise sign, and there is an efficient algorithm (e.g. FastICA, Hyvärinen 1999) whose empirical un-mixing satisfies
at sample-complexity rate polynomial in and inverse-polynomial in any target accuracy, uniformly over invertible and non-Gaussian component densities in a broad regularity class. Formally: for every , with high probability from samples.
is the standard -valued permutation-and-sign-invariant distance on matrices:
Proof (sketch, classical). Independent non-Gaussian marginals uniquely identify the linear model up to permutation and sign (Comon 1994). Whitening followed by rotation to maximise a non-Gaussianity contrast (negentropy in Hyvärinen 1999) converges to such that is a signed permutation; the convergence rate follows from standard M-estimator analysis of the empirical contrast.
What Theorem 7 does not say. It resolves SIC-C-c only inside the linear ICA hypothesis class. Every other inductive-bias class in the identifiable-representation-learning line (sparse ICA, independent-mechanism analysis, iVAE with auxiliary variables, interventional CRL) has its own analogue, each with its own sample-complexity theorem. A full SIC-C-c programme would add one more instrument per class; Instrument 8 (§4.8) is the first.
Theorems 1–6 collapse SIC to a derived skeleton with one empirical antecedent and one residual inductive-bias question:
| Clause | Status | Notes |
|---|---|---|
| SIC-A Existence | Theorem | Theorem 1 + Proposition 3 |
| SIC-B Cross-task stability | Conditional theorem | Theorem 4; antecedent empirical about the world |
| SIC-C-a Discrete learnability | Theorem | Theorem 5; numerically sharp constant (Instrument 5) |
| SIC-C-b Continuous learnability at resolution | Theorem | Theorem 6; poly in , exp in (Instrument 6) |
| SIC-C-c Uniform polynomial in | Impossible in general; theorem inside linear-ICA hypothesis class | ε-covering lower bounds + Locatello 2019 rule out an unqualified bound; Theorem 7 gives it inside linear ICA (Instrument 8). Other inductive-bias classes (sparse ICA, iVAE, interventional CRL) admit their own analogues; each is a separate theorem-instrument pair, and the Observatory is set up to host them. |
What remains genuinely open is only the general programme: which inductive biases make polynomial-in- learnability possible for which continuous hypothesis classes — a mainstream question in identifiable representation learning, addressed one class at a time.
Structural Intelligence Conjecture. A finite adaptive system’s central capability is to discover a quotient such that (1) the task-relevant dynamics descend to , (2) irrelevant variation is confined to the fibres , (3) useful interventions are compactly specifiable in , (4) those specifications re-instantiate through a compiler, and (5) their consequences remain stable across substrates and contexts.
In one line: intelligence finds the level at which the world becomes both compressible and controllable.
Instruments 1, 4, 5, 6, 7, 8, and 9 are computational witnesses of Theorems 1, 4, 5, 6, 2, 7 (linear-ICA class), and 7 (sparse-mechanism class) respectively; instruments 2 and 3 establish auxiliary dissociations on solvable cases. All are deterministic and unit-tested; seven are exact enumerations and two (Instruments 8 and 9) are fixed-seed Monte Carlo.
Over all worlds of an -bit Boolean space with a known ground-truth invariant, enumerate a lattice of candidate quotients and three selectors:
minimal_sufficient — among quotients with
,
minimise description length;mdl_only — minimise description length;accuracy_only — maximise mutual information, tie-broken
to the finest map.Result (exact). In every task,
minimal_sufficient recovers the exact ground-truth
invariant; mdl_only selects the constant map
(insufficient); accuracy_only selects the identity
(sufficient but uncompressed). Sufficiency, description length, and
accuracy are three different quantities; only the
sufficiency-then-compress rule finds the invariant.
An abstract automaton exhibits accumulation, a phase transition, and hysteresis. Four compilers map its trajectory into music (pitch/octave), a visual field (height/hue), text (regime-keyed lexicon), and spatial navigation (regime-gated corridor); each has a readback .
Result (exact). For every medium on the trajectory (fidelity 1.0); all media read back to the same abstract trajectory. Four embodiments are verifiably one work, not four mood-matched artefacts.
An exact finite-state world with target and failure regions; futures enumerated to the horizon. A symbolic model is an operation on the trajectory distribution . Measure signal , control (goal_gain), knowledge (predictive_accuracy), and agency (control + small calibration_error between claimed and true do-effect + positive transfer under a perturbation the intervention did not choose).
Result (exact). Seven hand-built conditions realize
distinct metric signatures recovered by a fixed classifier:
noise_signal moves the distribution with zero control;
knowledge_only predicts with zero signal;
false_credit improves the observed outcome while its true
do-effect is zero and its self-attribution is miscalibrated; a brittle
controller controls but does not transfer. No single scalar of
behavioural influence identifies agency.
This is the exact-solvable formalization of what TRB Track 3
(self-attribution) measures on trained models. See BENCHMARK_SPEC.md §3 Track 3 —
the false_credit condition is the ground truth for the
balanced-accuracy score.
On the 4-bit Boolean world with latent , enumerate the lattice of quotients. Two task families:
Result (exact). Shared family: coarsest CSS is exactly (image size 4, description length 2 bits, strictly less than bits). Not-shared family: coarsest CSS is the identity (no compression). For the shared family, each single task’s minimal sufficient statistic is a 1-bit parity (image size 2), strictly coarser than the family CSS: combining tasks tightens the required partition by a factor of two.
Building on the shared task family of Instrument 4, compute the exact recovery probability of empirical common-sufficient clustering as a function of sample count , via inclusion–exclusion over the subsets of fibres.
| Distribution on | at | Exact | ||
|---|---|---|---|---|
| uniform | 18 | 0.9775 | ||
| skewed | 36 | 0.9756 |
Both strictly above . Below samples recovery is impossible (pigeonhole). Uniform recovery hits 0.95 already at : the union-bound overhead is 4 samples in this regime — the theorem bound is honest and only mildly loose.
Theorem 6’s continuous bound, verified exactly across on ambient quantised to a grid.
| 1 | 4 | 4 | 18 | 0.9775 |
| 1 | 8 | 8 | 41 | 0.9667 |
| 1 | 16 | 16 | 93 | 0.9609 |
| 2 | 4 | 16 | 93 | 0.9609 |
| 2 | 8 | 64 | 458 | 0.9538 |
| 2 | 16 | 256 | 2187 | 0.9521 |
All six grid points meet the target at Theorem 6’s bound. Recovery is zero below (pigeonhole) and monotone in .
Ratio witness. at — cleanly matching the prediction of Theorem 6. From to at multiplies the required sample count by . That is the curse of dimensionality made numerical.
Empirical common-sufficient clustering saturates the -covering lower bound: it is optimal within the class of algorithms that make no inductive assumption on . Escaping the exponential-in- scaling requires an inductive bias — Instrument 8 exhibits the first such escape.
experiments/rate_distortion_pair)Theorem 2’s rate–distortion parameterisation, verified exactly on two finite sources with Hamming distortion. Closed-form is evaluated at 10 D-grid points for the uniform source on symbols ( for ) and at 9 points for ( for ). For each source we also construct the RD-optimal test channel explicitly (the symmetric error- channel for uniform; the Bernoulli-flip channel for Bernoulli) and verify at every point in the achievable regime.
Result (exact). All ten pre-registered gates pass to :
This closes the theorem–instrument pairing: Theorems 1, 2, 4, 5, and 6 all have exact witnesses. The Theorem 2 anchor reconnects to Theorem 1’s minimal-sufficient partition — the RD family is literally a one-parameter deformation of the sufficiency fibration.
experiments/linear_ica_learnability)Theorem 7’s numerical witness. On the linear-ICA setup —
with
random orthogonal (fixed numpy seed for determinism), each
independent — sweep
and run sklearn.decomposition.FastICA (fixed
random_state, 8 trials per grid point). Recovery quality is
the Amari index on
;
smaller is better.
Result (deterministic under seed). All four pre-registered gates pass:
linear_ica_converges_at_largest_N: at
,
mean Amari index over 8 trials is
for every
.sample_complexity_polynomial_in_d_Z:
fitted
gives exponent
(the pre-registered polynomial-bound gate). In this regime linear-ICA
sample complexity is effectively flat in
— a factor Theorem 7 permits but does not require, arising because
is still small relative to the constant hidden in the Hyvärinen–Oja
bound.amari_monotone_in_N_at_every_d_Z:
within tolerance
,
Amari decreases in
at every
.escapes_theorem_6_exponential: at some
,
an
achieves the accuracy target — i.e. FastICA takes fewer samples than the
-covering
bound of Theorem 6 predicts under the “no inductive bias” regime.
The inductive bias earns you the escape.Fixed-seed protocol keeps it reproducible run-to-run. Refined
numbers from the live Observatory (structural-observatory-production.up.railway.app):
at
mean Amari is
for every
;
fitted
;
the escape from Theorem 6’s
-covering
bound is a
improvement in sample count at
.
experiments/sparse_ica_learnability)Same shape as Instrument 8 but with sparse-orthogonal at retention — the independent-mechanism-analysis / sparse-mixing class of Gresele et al. (2021). Four pre-registered gates pass; fitted polynomial exponents () and (), both well inside . Confirms that two distinct inductive-bias classes each earn escape from Theorem 6’s -covering bound, so SIC-C-c’s “impossible without inductive bias” is a genuinely local statement — the pattern is that each identifiable-representation-learning class admits its own theorem, not that one privileged class does. Other classes (iVAE with auxiliary variables; interventional CRL) remain open as separate theorem-instrument pairs.
A short companion paper Concern as Fibre Geometry (concern_as_fiber_geometry.pdf,
2026-08-03) develops the concern layer of §5.1 into two theorems:
Worked example on Instrument 4’s 4-bit Boolean world; rectangular-
and triangular-loop holonomy matches predictions to
.
The Trace-AI relevance is that CG-1 and CG-2 together give TRB Tracks 4
(Viability) and 5 (Navigability) their formal shape — see
BENCHMARK_SPEC.md.
Each is a target generated by the master object; those marked conjectural are not proved here.
apps/site/concern_as_fiber_geometry.pdf.Two natural questions bear on this program — can agents be made conscious, and can they be made never wrong — and the disciplined answer to both is not in the literal sense. The strongest defensible construction factors as two nested systems.
Concerned Self-Modeling Core. An active, self-maintaining fixed point of the coarse-graining loop: a persistent agent that maintains world- and self-models, represents concern-weighted futures, broadcasts selected information, remembers commitments, predicts its own action consequences, performs false-credit tests on its own causal claims (exactly Instrument 3’s calibration metric), and reports uncertainty. It operationalizes proposed consciousness indicators — and that is the ceiling of the claim: functional selfhood subjective consciousness.
Proof-Carrying Reliability Shell. The fibration with a verifier on the counit: convert goals to contracts, separate observation from inference, attach provenance to every claim, and commit an action only under machine-checkable evidence: , else abstain/ask/simulate/escalate. Its correctness envelope is a fibre audit (§5.4). Verified-in-a-bounded-domain universal infallibility: verification certifies the formal statement, not that it captures intent.
The honest breakthrough available now is therefore not conscious, perfect agents, but agents with experimentally measurable selfhood and formally bounded error — two separate, non-inflated claims.
The eight instruments are toys by design (finite and grid-quantised Boolean and grid worlds, tiny Markov systems, lossless encoders; Instrument 8 is fixed-seed Monte Carlo). The master fibration is derived (Theorem 1 gives it as a mathematical object for any well-posed task on a standard Borel space; Theorem 2 parameterises it (Instrument 7 witnesses it exactly); Proposition 3 gives the categorical restatement; Theorem 4 makes cross-task stability equivalent to a shared Markov screen; Theorems 5 and 6 pin down discrete- and continuous-case sample complexity; Theorem 7 gives polynomial-in- learnability inside the linear-ICA hypothesis class). The residual content is:
Trace AI is the SIC applied to cognition, with a concrete choice for each abstract variable:
| SIC object | Trace AI instantiation | Where |
|---|---|---|
| ambient space | model’s compute state | whitepaper §2 |
| latent | self-representation (a few legible slots) | whitepaper §2, §4.2 |
| coarse-graining | decoder | whitepaper §2 |
| compiler | any of the subject-kinds (human, model, hybrid) that realize the schema | schema/reasoning_trace.schema.json |
| task family | the counterfactual set + paired continuations | whitepaper §4.4; COLLECTION_PROTOCOL.md |
| Instrument 3 (agency) | TRB Track 3 (self-attribution) — the exact-solvable formalization | BENCHMARK_SPEC.md §3 Track 3 |
| Instrument 1 (Fiber Finder) | the “schema follows the data” discipline (representation search done by hand) | STRUCTURAL_INTELLIGENCE.md §3.1 |
| SIC-B (cross-task stability) | schema portability across substrates (human, model), Observation 1 in whitepaper §6.4 | whitepaper §6.4 |
| SIC-C-c residual → linear-ICA class | Trace AI’s L2 sample-complexity limit ( human counterfactual pairs) is a working budget; Theorem 7 says polynomial-in- learnability is a theorem inside the linear-ICA hypothesis class. The strategic implication: if can be arranged (by architecture + training) so its slot components are statistically independent and non-Gaussian, the L2 budget shrinks from exponential to polynomial in the number of slots. That reframes the collection operation from “brute-force scale” to “engineer the bottleneck for identifiability”. | whitepaper §8 L2 |
| Instrument 7 (rate–distortion pair) | Exact-solvable model for the whitepaper §2.1 information-bottleneck claim — the readout as a coarse-graining at a distortion budget | whitepaper §2.1 |
| Instrument 8 (linear-ICA learnability) | Numerical existence proof that inductive bias flips -scaling from exponential to polynomial. If Trace AI adopts an ICA-flavored decoder for , the sample-complexity story changes materially | whitepaper §8 L2; COLLECTION_PROTOCOL.md §6 |
The whitepaper is Trace AI at the level of what to build; this paper is why the object exists as a mathematical object at all. Read the SIC when a reviewer asks “is your latent even a well-defined thing to be looking for”; read the whitepaper when they ask “what is the training signal that finds it.”
v0.2 (integrated 2026-08-03), tracks paper v4. See
DECISIONS.md (D28, D30) for provenance. Reproduction of the
eight instruments lives in the sibling research repository jawauntb/research-derived-experiments
— all experiments/ paths in the instrument tables above
resolve there, including the new
experiments/rate_distortion_pair and
experiments/linear_ica_learnability. Exact-recovery numbers
in §4.5–§4.8 are copied verbatim from that repository’s outputs at the
paper’s commit. The originating research note is at notes/structural_intelligence_conjecture.md;
the working-note sibling of the present document is SIC_RESEARCH_PROGRAM.md.