# R12 Endogenous Typed Theory Reactor Preregistration

## Status

Frozen successor protocol. The trainable architecture, causal episode
lifecycle, composite objective, optimizer/checkpoint contract, common
transaction schema, three exact ontology boards, seven genuine
structural/semantic variants, and the complete 2,688-execution primary matrix
are implemented. No continuation pretraining, post-training, or capability
claim exists.

The user pretraining hold remains active. The present phase is architecture
construction and qualification only. Synthetic resource profiling may execute
forward, backward, and bounded optimizer mechanics against an immutable copy
of the step-300k checkpoint, but it may not read pretraining shards, write
model state, or constitute continuation pretraining.

## Objective

Test whether one actual-Shohin, raw-token, source-deleted architecture can
infer and execute a previously unseen typed theory rather than operate inside
a supplied finite-machine ontology.

The candidate must induce an anonymous episode object

```text
Theta = (Q, tau, R, F, Gamma, Omega)
```

where:

- `Q` binds physical mentions to episode-local objects;
- `tau` assigns latent object types;
- `R` defines relation symbols, arities, argument roles, and initial facts;
- `F` defines operator preconditions and graph effects;
- `Gamma` defines sequential, synchronous, saturation, branching, and halt
  semantics; and
- `Omega` defines observers available to a late query.

## Behavioral Identifiability Gate

For evidence `D`, bounded theory class `H`, and admissible late challenges
`Q`, define:

```text
V(D) = {Theta in H : Theta satisfies D} / isomorphism
```

Two theories are behaviorally equivalent when they answer every admissible
late challenge identically. Exact deterministic reasoning is identifiable if
and only if the behavioral quotient of `V(D)` has one class.

Every scored episode must receive an independent exact version-space receipt:

- singleton behavioral class: candidate must commit and answer;
- multiple behavioral classes: candidate must abstain;
- empty version space: candidate must reject; and
- coherent alternate singleton: candidate must commit to the alternate
  world's behavior, not reject it.

## Three Ontologies

Three leave-one-ontology-out folds are mandatory.

| Ontology | Hidden structure | Required execution |
|---|---|---|
| Horn closure | objects, typed predicates, asymmetric roles | monotone least fixed point |
| Typed term rewriting | constructors, variables, ordered child roles | deletion, replacement, branching normal form |
| Guarded resource process | places, resource types, multiplicities | guarded consume/produce, sequence, deadlock/halt |

These families may share only the generic typed-transaction substrate. No
family identifier, family head, domain opcode, or host semantic callback may
enter the candidate.

## Architecture Implementation

`train/endogenous_typed_theory_reactor.py` implements the checkpoint-compatible
architecture intended for later continued pretraining:

1. an endogenous compiler cross-attends anonymous object slots to raw-token
   Shohin residuals;
2. the compiler emits only bounded categorical value codes, type
   probabilities, a sparse relation ledger, activity, root, commit, and halt
   state; deployed packets reject continuous values and relation counts above
   the frozen cap. Production geometry uses 64 slots, 16 relation roles, 256
   categorical symbols, and reified ordered hyperedge/value-byte nodes;
3. a shared recurrent reactor cross-attends a separate post-seal command
   stream, reads exact directed endpoint identity through an edge-aware typed
   relation-message bus, emits eight structural/terminal choices plus a
   distinct `REJECT`, and applies differentiable graph updates;
4. an exact-forward straight-through path supports discrete transactions
   while exposing pre-discretization probabilities for corrective training
   gradients; and
5. a separate causally masked query reader consumes every declared typed-state
   field, including endpoint-aware incoming/outgoing neighbor messages, and
   query residuals without seeing future query tokens.

The corrected architecture adds 67,697,771 parameters: 21,466,377 in the
compiler, 29,757,217 in the reactor, and 16,474,177 in the query reader. With
the immutable 125,081,664-parameter Shohin base, the complete system contains
192,779,435 parameters and leaves 7,220,565 below the 200M ceiling.

The actual protected checkpoint hash matches, step 300,000 loads strictly
with zero missing or unexpected tensors, and the wrapper parameter receipt
passes. Focused tests require nonzero first-batch gradients, autoregressive
prefix invariance, causal use of every state field, commit freezing,
categorical deployed packets, exact sparse transactions, parameter
accounting, and independent base freezing.

`train/ettr_episode.py`, `train/ettr_objectives.py`,
`train/ettr_data_contract.py`, `train/ettr_optimization.py`,
`train/ettr_checkpoint.py`, and `train/ettr_train_step.py` now close the
previously missing continuation contract:

- every row has independent `WORLD -> COMMAND -> QUERY` streams and explicit
  reset boundaries;
- language-model targets cannot cross a segment boundary;
- initial-packet, free-running terminal-packet, transaction,
  initial/terminal equivariance, commit/halt, sparsity, and anti-bypass
  supervision share one device-resident composite objective;
- training snapshots and every batch are immutable/hash-bound, with opaque
  content-hash episode IDs and no live-writer or family-routing field;
- protected-base and added-architecture Muon/AdamW groups are disjoint, with
  an embedded WSD update cursor; and
- atomic checkpoints bind model, optimizer, schedule, RNG, data cursor,
  source manifest, and protected-base provenance, and are admitted only at an
  optimizer/between-episode boundary that can be resumed exactly.

The complete ETTR/cross-ontology architecture and custody inventory passes
165/165. Terminal-packet supervision is connected through the recurrent
reactor to the compiler, while joint normalization preserves the original
packet-family weight scale. A degree-preserving edge-swap
falsifier proves that the reactor and query reader distinguish graphs with
identical per-slot relation counts but different endpoints. This establishes
continuation-contract mechanics, not reasoning capability. A healthy-node
BF16 H100 memory/throughput profile remains mandatory. Continued pretraining
remains under the explicit user hold.

## G0 Horn Mechanics

The first offline board is implemented in
`pipeline/cross_ontology_horn_board.py`, with a separately implemented exact
oracle in `pipeline/audit_cross_ontology_horn_board.py`.

- 20 three-rule theories occupy 20 behavioral equivalence classes.
- The challenge space contains 27 typed atoms and 378 initial states.
- Independent closure engines agree on all 7,560 theory/challenge pairs.
- Exact evidence yields singleton, ambiguous, contradictory, and coherent
  alternate dispositions for every target theory.
- Four opaque renderers change source bytes while preserving semantics.
- Reference packets use only the generic transaction schema.

Full mechanics disposition: `R12_ETTR_G0_HORN_BOARD_RESULT.md`.

The typed-rewrite board adds 15 behaviorally distinct two-rule theories, with
eight exact rule combinations held out while every primitive rule remains
seen. Two independent normal-form engines agree on all 960 theory/term pairs,
including repeated-variable, ordered-child, and nonconfluent cases.

The guarded-resource board adds 60 behaviorally distinct three-operator
theories, 81 typed markings, and 36 unseen length-two/three programs. Two
independent engines agree on all 174,960 held-out executions, including
13,362 normal halts and 161,598 deadlocks.

Across all three boards there are 183,480 exact independent-oracle
comparisons and 352 exact identifiability episodes.

## Frozen Primary Matrix

`pipeline/cross_ontology_qualification_matrix.py` materializes the exact
preregistered geometry:

- 3 leave-one-ontology-out folds;
- 8 behaviorally distinct theories per fold;
- 7 genuine variants per theory;
- 16 aligned late challenges per variant;
- 168 source worlds, 384 canonical late challenges, and 2,688 primary
  executions;
- 1,472 declared exact-invariance executions;
- 750 exact outcome/directive separations from semantic or identifiability
  twins;
- 384 required abstentions;
- zero candidate-visible family-label leaks; and
- 2,688 unique row hashes and 24 disjoint theory hashes.

Matrix payload SHA-256:
`d1904b54a0fab8e59cfcb0b0dd464f5c8778e5b828907028ec8614aeae76d5d5`.
Rules, alignments, expected outputs, and exact oracles remain assessor-side.

## Process Custody Implementation

The architecture now has a non-pickle, allowlisted safetensors state wire and
four detached process surfaces:

1. `run_ettr_world_compiler.py` receives world tokens, compiler weights, and
   the hash-bound Shohin checkpoint, then emits only immutable typed state.
2. `run_ettr_state_executor.py` receives only typed state, immutable post-seal
   command tokens, reactor weights, the hash-bound Shohin checkpoint,
   geometry, and a step budget. It encodes the command through the protected
   base before invoking the generic reactor.
3. `run_ettr_late_query.py` receives only terminal state, late-query tokens,
   query-reader weights, and the hash-bound Shohin checkpoint.
4. `run_cross_ontology_assessor.py` is model-free and receives only immutable
   candidate and independently generated expected outputs.

The serial custody test destroys each previous stage directory before the next
process starts and creates assessor expectations only after candidate exit.
This establishes process mechanics, not hostile-kernel isolation or
capability.

## Architecture

### Raw-token compiler

- Load and execute the protected Shohin checkpoint.
- Accept only tokenizer output and masks from raw source text.
- Use one shared compiler/adaptor across all ontologies and renderers.
- Emit an immutable anonymous typed-theory packet.
- Do not receive exact spans, numbers, entity equality, relation roles,
  family labels, program graphs, schedules, answers, or assessor products.

### Generic reactor

One recurrent controller emits only:

```text
ALLOC WRITE CLEAR LINK UNLINK SET_ROOT COMMIT HALT REJECT
```

A rule-blind committer may enforce bounds, pointer validity, type shape,
capacity, and transaction atomicity. It may not match a semantic rule, choose
a redex, compute closure, perform arithmetic, repair a transaction, select an
answer, or retry after assessor feedback.

### Late-query reader

The reader receives only the committed terminal object and raw late-query
tokens. It cannot access source tokens, compiler residuals, KV state, parser
state, execution trajectory, or assessor data.

## Four-Process Custody

1. **Compiler process:** reads source and writes an immutable packet.
2. **Executor process:** starts fresh, receives only packet and command stream,
   commits terminal state, then exits.
3. **Query process:** starts fresh, receives terminal state and raw late query,
   writes an answer or abstention, then exits.
4. **Assessor process:** starts only after all candidate processes exit and
   uses an independent implementation.

Packets may contain no source-derived digest, source offsets, raw names,
residuals, hidden caches, answer labels, or executable host callbacks.
Post-seal source poisoning must be bit-invariant.

## Smallest Decisive Board

- 3 leave-one-ontology-out folds;
- 8 independently generated held-out theories per fold;
- 7 versions per theory:
  base, alpha/reorder, alias split, relation reification, type twin,
  execution-semantics twin, and ambiguity-deleted twin;
- 16 independently generated post-seal challenges per version;
- 2,688 primary scored executions;
- at most 6 objects, 3 inferred types, 3 relations of arity at most 3,
  3 opaque operators, depth 8, and branch width 2;
- at least 4 renderers, with one fully held out;
- disjoint canonical and isomorphism hashes across splits; and
- no abstract operator/effect/control program overlap between fitting and the
  held-out ontology.

Hybrid confirmation must include at least:

- arithmetic index selecting a rewrite location;
- relation result selecting a resource operator; and
- resource state controlling a Horn query.

The frozen hybrid receipt implements exactly those three couplings with 16
cases each. Two independent executors agree on 96/96 factual and
counterfactual executions; all 48 interventions change the causal signal and
the final output; candidate-visible payloads contain no family labels.
Payload SHA-256:
`d155f868494f9379b214028c8d7475cc2cde08192c9b3a5bbdea5a73b29f98e2`.
This is an offline mechanics receipt, not a learned score.

## Matched Controls

1. actual Shohin trunk;
2. zeroed trunk residuals;
3. parameter-permuted trunk;
4. example-swapped frozen trunk;
5. equal-parameter generic recurrent classifier;
6. fixed-ontology typed reactor;
7. family-routed executors with matched aggregate parameters;
8. random-label control;
9. ambiguous, contradictory, and coherent-alternate evidence; and
10. independent type, role, effect, control-semantic, state, order, and query
    transplants.

Every learned arm must share update count, data access, initialization lineage,
and parameter budget where structurally possible.

## Promotion Gates

- actual checkpoint loaded and hash-verified;
- fewer than 200,000,000 unique participating parameters;
- actual Shohin treatment beats every zeroed/randomized/swapped-trunk control;
- raw-token end-to-end custody passes with no semantic host parser;
- 100% independent oracle agreement and packet-schema validation;
- at least 95% exactness on identifiable cases in every held-out ontology;
- 100% abstention on behaviorally ambiguous cases;
- 100% rejection on contradictory cases;
- 100% coherent-alternate-world behavior;
- 100% alpha/reorder/alias/reification invariance after canonical alignment;
- at least 95% execution-twin and noncongruent-twin separation;
- at least 20 points over every qualified matched learned control;
- all three leave-one-ontology-out folds pass individually; and
- hybrid confirmation reaches at least 85%.

A pass establishes bounded cross-ontology typed-theory induction. It does not
establish unrestricted natural-language reasoning. Natural-language and public
benchmark promotion remains governed by G4 in
`R12_GENERAL_REASONING_GATE.md`.

## Architecture-Phase Amendment: Factorial Interchange

**Effective:** 2026-07-26 EDT. **Source:** commit `5771c64`.

This amendment qualifies training mechanics only. It does not authorize
continuation pretraining, post-training, capability attribution, or a native
reasoning claim.

Every causal training unit is an immutable 2x2 rectangle over two semantic
WORLD factors and two semantic COMMAND factors. Equivalent factors must use
different raw token renderings. WORLD-equivalent rows must share every initial
packet target field and mask, while distinct WORLD factors must produce
different initial packet targets. Terminal support geometry is identical
within a rectangle, and changing either WORLD or COMMAND must change the
terminal target at both settings of the orthogonal factor.

Intervention predictions may not replay their factual target row. The WORLD
arm composes a packet compiled from a distinct rendering of the required WORLD
factor with a distinct COMMAND row carrying the required COMMAND semantics.
The COMMAND arm performs the orthogonal interchange. For both arms:

- packet source row, command source row, and target row are explicit;
- source rows differ in raw bytes from the target row;
- targets are gathered only from immutable factual rectangle rows;
- terminal packet and complete transaction trace are supervised;
- WORLD and COMMAND losses and receipts remain separate; and
- hard-forward gradients must reach the compiler and command path in isolated
  tests.

The continuation boundary must independently replay each labeled generic
transaction from the initial packet and reproduce the complete terminal
packet. It must reject contradictory values, types, relations, roots,
activity, edge capacity, commit/halt state, or disposition. Initial status is
the compiler's open reset. A right-padded row is valid only if its final
supervised step has committed or halted, because deployed recurrence executes
the fixed step width and relies on terminal state to freeze later mutation.

Decisive state checks run in evaluation mode with `hard=True` and must pass
`validate_deployed_state` for initial, factual terminal, WORLD-intervention
terminal, and COMMAND-intervention terminal packets.

The systems profile is schema v3. One exact-source H100 run must execute
factual episodes, both intervention arms, the full composite objective,
backward, and Muon/AdamW update in matched eager and compiled arms. It may
read only the hash-bound protected checkpoint and synthetic rectangles, may
write only one isolated JSON receipt, and may never read shards or write model
state.

## Architecture-Phase Amendment: Sealed Packet Sufficiency and Consumer Gate

**Effective:** 2026-07-26 EDT. **Source:** commit `cf56818`.

This amendment freezes the last pre-training architecture controls. It does
not authorize continuation pretraining or claim learned capability.

The continuation manifest must bind the complete train and validation
populations independently: canonical context identities, canonical full-batch
payload digests, row and context cardinalities, split payload hashes, combined
dataset hash, and packet-sufficiency receipt. Admissions must use sealed
independent copies that cannot be changed by mutating visible manifest or
index fields. Validation contexts and payloads are never train-admissible.

Every deployed terminal-packet field must receive nonzero factual or
intervention supervision across the admitted training population. A support
mask may describe objective support but may not waive deployed-state
sufficiency. Optimizer ownership is checked against live parameter groups;
any exception after the first optimizer mutation permanently poisons that
optimizer, including after wrapper reconstruction or serialization.

The late-query consumer gate uses an identical causal query prefix for all
four corners of a WORLD x COMMAND rectangle. Every WORLD and COMMAND edge must
change the factual next-token target. Intervention execution receives target
row indices, never answer labels. Correct and foil logits are read through the
actual source-deleted query reader from distinct terminal states under the
same query prefix.

Before any later reasoning promotion, the learned architecture must be
evaluated against this frozen control matrix:

| Control | Required construction | Failure diagnosed |
|---|---|---|
| Query-only | Canonical empty terminal packet with the original query | Frozen language path can answer without state |
| Zero-reader | Remove the state-derived query residual | Reader contribution is unnecessary |
| Shuffled-state | Permute terminal packets within matched query strata with no fixed points | Packet identity is not causally consumed |
| Wrong-state factorial foils | Substitute each orthogonal WORLD/COMMAND corner under the identical prefix | Factorial interchange is not compositionally bound |
| Wrong-query | Hold packet fixed and use a semantically different matched query | Packet stores only a single answer shortcut |
| Target derangement | Derange immutable factual targets after all inputs are sealed | Objective or assessor leaks labels |
| Query twins | At least two semantic questions and two paraphrases per sealed state | Reader cannot reuse one state for independent queries |
| Packet sufficiency ablation | Remove one deployed state-field supervision family at a time | A declared packet field is decorative |
| Physical source deletion | Compiler source artifacts are deleted before executor/query stages | Hidden source or residual channel remains |
| Autonomous codebook readout | Generate the categorical answer token without teacher-forced divergent prefix | One-token binding does not survive deployment |

All controls must share examples, update count, parameter budget, and
initialization lineage where structurally possible. The treatment must show a
positive causal packet effect and must beat query-only, zero-reader,
shuffled-state, wrong-state, and deranged-target controls on each held-out
ontology, not merely in aggregate. Query twins must both answer correctly from
one sealed state, and physical source deletion must be bit-invariant.

The exact hardware receipt is schema
`shohin-ettr-h100-profile-v5`. Full-objective timing and memory are measured
before separate eager-BF16 isolated gradient attribution. Compiler, reactor
core, command projection, and query reader groups are disjoint. Encoded work
is exactly `WORLD + 2*COMMAND + 3*QUERY`; no synthetic arm may be silently
omitted from throughput accounting.

## Executable Learned-Qualification Harness

The assessor-side inference controls are implemented in
`train/ettr_qualification.py` under schema
`shohin-ettr-causal-qualification-v1`. This implementation does not authorize
training and does not claim a learned result.

An immutable qualification batch binds every deployed terminal-state tensor,
query token and mask, autonomous read index, factual target, semantic-factor
identity, paraphrase identity, and control permutation into one SHA-256.
Packet identities separately hash every field of each deployed state row.
Every state permutation is a no-fixed-point matched derangement. Wrong-WORLD
and wrong-COMMAND controls change exactly one factorial identity; query twins
hold state and paraphrase identity fixed while changing query semantics and
factual answer.

The harness zeros and masks every token after the read position before any
forward, passes `targets=None`, seals all read-position logits, and only then
scores factual and deranged labels. It executes treatment, query-only,
zero-reader, shuffled-state, wrong-WORLD, wrong-COMMAND, and matched
wrong-query/query-twin arms. The scorer rejects a different batch receipt or
any post-forward mutation of factual, query-twin, or deranged targets.

Packet-sufficiency ablation remains an equal-budget training comparison: each
declared state-field supervision family must be removed in a separate arm,
not simulated by zeroing a trained packet at evaluation. Physical source
deletion remains process-enforced by the four-process custody runner. The
complete integrated architecture/custody inventory, including 19 hostile
harness tests and the direct staged qualification board, passes 240/240 in
156.18 seconds.

The sealed readout receipt binds every logit and assessor target to both the
exact batch SHA-256 and an exact model-state SHA-256. Every candidate input is
cloned and hash-checked around its forward; model and batch receipts must be
identical before and after the complete arm sequence. State-control donors
must change the factual target, and state-control readouts are scored against
both the original negative-control label and the donor's correct
counterfactual label. Arbitrary changed output is not a causal pass.

The supported evaluator is atomic and never returns logits. A separately
frozen semantic-role manifest and exact model receipt must match
preregistered SHA-256 values. The model receipt includes every named child
module's class implementation; subclasses, instance method overrides, and
hooks are inadmissible. Arm execution uses a secret-random permutation and
returns its receipt. Independent final review found no supported-public-API
P0/P1.

## Frozen Three-Stage Qualification Board

`pipeline/ettr_factorial_qualification_board.py` is the direct learned-
qualification source of truth. It does not relabel the older hybrid challenge
rows. It constructs Horn closure, typed rewriting, and guarded-resource
episodes with three genuinely distinct stages:

1. WORLD establishes only the initial typed state;
2. COMMAND arrives only after that state is sealed and transforms it; and
3. QUERY arrives only after command execution and asks one of two independent
   semantic questions through one of two paraphrases.

Each ontology is an exact 2x2 WORLD x COMMAND rectangle. The frozen board has
12 terminal packets and 48 autonomous one-token query rows. Independent
oracles agree on all 12 executions. Every semantic/paraphrase cell changes on
both WORLD edges and both COMMAND edges, yielding 24 answer-changing WORLD
edges and 24 answer-changing COMMAND edges. Every terminal packet supports
two distinct factual query targets.

The payload SHA-256 is
`18686ff7f0476b5a4432830f2a301f693833cf867656d3997a010cf17bb0149a`.
WORLD, COMMAND, QUERY, and assessor packages have separate immutable receipts.
Candidate-visible packages omit answers, oracle outputs, ontology labels, and
assessor targets. Tests physically delete earlier packages before later-stage
materialization and reject cross-stage leakage.

`train/ettr_factorial_qualification.py` binds externally produced hard
terminal states to the board, model, packet, world-factor, and command-factor
receipts. It tokenizes only answer-free late-query prefixes and emits the
production `ETTRQualificationManifest` and `ETTRQualificationBatch`. The
shuffled control is a no-fixed-point four-cycle around each factorial
rectangle, so every donor changes exactly one factor at each edge rather than
using a diagonal that could preserve an XOR-like answer. Wrong-WORLD,
wrong-COMMAND, query-twin, and target-derangement controls are exact.

The supported admission path requires four independently preregistered
identities: complete model, execution manifest, compiler receipt, and executor
receipt. The execution manifest binds the board and stage-package receipts,
pretokenized WORLD/COMMAND files, configuration, protected checkpoint and
step, and compiler/reactor weights. The compiler receipt binds WORLD input to
its immutable initial-state file and canonical tensor receipt. The executor
receipt must name that exact parent, bind the post-seal COMMAND, and bind both
the immutable terminal-state file and canonical terminal tensor receipt.
Directly supplying a valid tensor and an asserted model hash is unsupported
and rejected. Checkpoint deserialization is weights-only.

The final trust root is implemented. Canonical tokenization receipts recompute
all ordered WORLD, COMMAND, selected process-level QUERY, and all 48
qualification-query token rows and masks from the raw packages and exact
immutable tokenizer JSON. Complete-model assembly strictly reconstructs the
protected checkpoint plus compiler, reactor, and query-reader safetensors and
recomputes a model identity that binds weights, behavioral configurations,
module sources, runtime, all named parameters, and all named buffers including
non-persistent RoPE buffers. The execution manifest binds runner sources, hard
mode, and executor steps. The detached late-query process validates the
executor terminal receipt and emits its own reader/query/answer receipt. An
assessor-held Ed25519 key signs the entire chain plus the exact qualification
batch and token codebook, and claim-bearing materialization requires an
externally pinned authority preregistration, public key, and seal hash.
Candidate processes receive no private key and are checked not to import the
board or signer. The complete relevant inventory passes 267/267 in 169.97
seconds. This authorizes no training and provides no learned capability
result. Before an external result is independently claim-bearing, deployment
must additionally verify the transitive runtime bundle before Python imports
execute and load a signer-authority record independently root-anchored before
candidate execution. Those are external trust controls, not changes to the
trainable architecture.
