# R12 Projected SD-CST Fresh-Board Preregistration

**Status:** first board closed before training or scored access; one narrow
derived-buffer contract repair is frozen below before a new source commit and
fresh board

## Pre-score infrastructure amendment

Source `76a183df6eb0a47c0b06a5db9ba4079e6399f4b6`, board seed
`3099288459709017829`, training seed `235733286388889829`, and job `693998`
failed inside `initialize_model` before either arm trained, before a checkpoint
or gate configuration existed, and before development or confirmation access.
Both access counts remained zero. The exact byte parent lacks ten newly learned
projected parameters and one deterministic non-trainable `permutations` lookup
buffer. The initializer incorrectly required all missing state-dict keys to be
trainable parameters.

The sole permitted repair separates these categories: the missing-key set must
equal `PROJECTED_TRAINABLE_NAMES` union exactly
`PARENT_DERIVED_BUFFER_NAMES = {"permutations"}`; the derived buffer is fixed by
the architecture, receives no gradients, and remains covered by full-state and
frozen-state digests. Any additional or absent key fails. No architecture,
parameter count, optimizer, data schema, evaluator, threshold, control, or
custody rule changes. The first board is closed and cannot be scored under the
new source. A new clean source commit must precede fresh independent board and
training seeds.

## Question

Can a newly initialized projected binding path, trained only on compiler fields,
compile unseen programs and a separately disclosed late query into a categorical
packet that an unchanged source-blind recurrent core executes exactly?

The prior projected compiler and job `693986` establish mechanics only on a
consumed training board. Their fitted projected weights are diagnostic controls
and are forbidden from the primary treatment.

## Fixed claim boundary

A pass establishes fresh-distribution source-deleted execution for this bounded
three-entity state-transport language. It does not establish broad natural-
language reasoning. The 125,081,664-parameter Shohin trunk is included in the
nominal system count but is inactive in the projected compiler forward path.
Both nominal and active parameter counts must be reported.

## Primary system

- Load the exact byte-addressed parent checkpoint with SHA-256
  `e5f87a1d5b22d24250a6aac6fb7c70b4a77dbdf01bd5f5c509020a3584dfa6f9`.
- Freeze every parent tensor.
- Freshly initialize only `PROJECTED_TRAINABLE_NAMES` after the source commit.
- Train exactly 6,748,897 projected binding parameters.
- Use the exact mechanics execution core with SHA-256
  `166ca6f81dd962b06a94f7a3661921a410760090ed1b750d78ec1b0f610113f1`.
- Do not refit the motor or reader from episode data.
- Nominal complete-system size is 146,057,595, below both the frozen 150M
  comparison cap and the global strict-below-200M cap.

The consumed projected checkpoint with SHA-256
`f347d1aea90dd3c60f7500167c7c22884451b365880259698306c6fce8ab10f3`
is a zero-shot diagnostic only. It cannot become the treatment or select a
checkpoint, threshold, renderer, or seed.

## Fresh board

After a clean source commit, draw one board seed. Build:

- 48,000 training rows with compiler/query targets only;
- 288 development families times eight paired variants = 2,304 rows; and
- 288 sealed-confirmation families times eight paired variants = 2,304 rows.

Every evaluation split has 48 families at each active depth one through six.
The eight variants are canonical, query swap, paraphrase, binding recode, order
counterfactual, stop shift, storage-order shuffle, and post-halt suffix.
Operation sequences, opaque names, exact normalized prompts, and 13-grams are
disjoint across splits. Opaque names retain the existing fixed-width contract;
name-length OOD is not introduced in this experiment. The independent audit
must additionally bind normalized renderer grammar and lexical inventories,
rather than trusting template identifiers alone.

The new board must also reserve all 48,000 operation sequences from the exact
training board that produced the frozen byte-addressed parent. Across that
inherited train split, every new split must have zero sequence, opaque-name, and
exact-prompt overlap, and both scored splits must have zero 13-gram overlap.
The new training split intentionally retains the inherited renderer grammar, so
its inherited grammar-level 13-gram overlap is measured and hash-bound rather
than falsely claimed to be zero.

Training rows contain no final state, answer, trajectory, execution trace,
paired-family outcome, or evaluator result.

## Learned arms

Both learned arms use fresh initialization, the same parent, architecture,
minibatch order, update budget, and optimizer settings.

1. `treatment`: true binding roles.
2. `row_shuffled_labels`: one deterministic post-commit role permutation is
   selected independently for every training row from all six permutations.
   Binding pointer slots, initial-state categories, and event identities are
   permuted consistently within that row; program bytes are unchanged. The
   aggregate row-to-permutation digest is frozen. Unlike a single global
   relabeling, this destroys the cross-row role function and is corrupted
   supervision rather than an equivalent coordinate system.

The shared-key predecessor and consumed projected checkpoint are diagnostic
only and need not receive equal training compute. The binding-source-free
compiler preserves the frozen parent's real line/kind/amount/query path while
replacing projected binding pointers and fingerprints with uniform evidence.
Uniform, source-free-packet, shuffled, reset, freeze, state/query/suffix,
post-STOP, and force-alive controls act on sealed packets or compiler inputs and
do not receive outcome labels.

## Training contract

For each learned arm:

| Field | Frozen value |
|---|---:|
| train rows | 48,000 |
| epochs | 4 |
| batch size | 64 |
| updates | 3,000 |
| learning rate | 0.0003 |
| warmup | 100 |
| schedule | cosine to zero |
| AdamW betas | 0.9, 0.95 |
| weight decay | 0.01 |
| gradient clipping | 1.0 |

Losses are initial state, active-event identity, declaration pointer,
initial-occurrence pointer, and event-occurrence pointer. The frozen parent
supplies line, kind, amount, and query fields. No final state, answer, executor,
variant, depth, or confirmation information is available to optimization.

The parent digest, frozen-tensor digest, trainable names/count, initialization
seed, minibatch-order digest, and final full-state digest are recorded. No
checkpoint selection is allowed: epoch four is the sole score-bearing state.

## Source-deleted score path

For each evaluation batch:

1. compile the program;
2. argmax and seal exactly 25 CPU `uint8` program categories;
3. poison and destroy every GPU-side program tensor and compiler output;
4. disclose and compile the late query;
5. seal one CPU `uint8` query category and destroy query tensors;
6. serialize only typed packet tensors plus the certified motor/reader tensors;
7. run a separate source-blind process; and
8. let the independent assessor compare predictions with the oracle.

The host evaluator retains hash-bound row evidence and the oracle so that it can
score the sealed result; this is not claimed to be physically erased. The
separate executor receives no source bytes, name, row/family ID, target, oracle,
variant, depth, split metadata, or projected compiler. Serialization rejects
unexpected keys, objects, dtypes, ranks, and shapes.

## Frozen development gates

All gates below are conjunctions. `exact_packet` means exact initial, kind,
active identity, active amount, and late query.

### Treatment accuracy

- exact packet overall at least 95%;
- exact packet at least 90% in every variant and every depth;
- initial, kind, identity, amount, and query at least 98% each overall;
- declaration, initial-occurrence, and event-occurrence pointer localization at
  least 90% overall and 80% in every variant;
- autonomous final state, answer, and joint result at least 90% overall;
- state, answer, and joint result at least 90% in every variant;
- state, answer, and joint result at least 85% at every depth;
- execution conditional on an exact packet exactly 100%; and
- every observed STOP-position bucket exact conditional on an exact packet.

### Matched attribution

- treatment exceeds row-shuffled-label exact packet by at least 50 percentage
  points; and
- treatment exceeds row-shuffled-label autonomous joint result by at least 50
  percentage points.

### Paired consistency and interventions

- at least 85% of required paired families are eligible;
- binding-recode and paraphrase state/answer consistency is 100% when both
  packets are exact;
- query swap preserves state and follows the changed query on 100% of eligible
  pairs;
- storage-order shuffle preserves state/answer on 100% of eligible pairs;
- post-halt suffix preserves normal state/answer on 100% of eligible pairs;
- force-alive, state swap, reset, freeze, and suffix-operand interventions match
  their independently simulated changed oracle on 100% of eligible cases; and
- each causal intervention has at least 15% separating opportunities, except
  query swap, which must separate at least 85%.

### Negative controls and custody

- uniform/binding-source-free/source-free-packet/shuffled state at most 35% and
  answer at most 45%;
  the higher state ceiling is frozen because the board balances answers exactly
  but its six terminal-state classes are intentionally not uniform;
- reset/freeze state and answer at most 75% against the canonical oracle;
- source poisoning and deletion preserve sealed packets bit-for-bit on 100%;
- motor certificate 78/78 and reader certificate 18/18;
- exact frozen parent/core/source/board/checkpoint/evaluator/assessor hashes;
- nominal complete system strictly below 150M and global system strictly below
  200M; and
- development/confirmation access exactly `1/0`.

Confirmation uses the exact same treatment checkpoint, executor, evaluator,
assessor, thresholds, and absolute gates. It is opened once only if every
development gate passes. No rescore, alternate decode, threshold change,
checkpoint choice, renderer edit, or source repair is allowed after a scored
split is opened.

## Custody order

1. Commit all scientific source, tests, renderer inventories, schemas, and this
   preregistration from a clean tracked worktree.
2. Draw board, treatment initialization, label permutation, and training seeds.
3. Build all splits once; chmod confirmation `0600`.
4. Run independent board and compiler-target audits before GPU access.
5. Train all learned arms without reading development.
6. Freeze a hash-bound gate configuration.
7. Atomically consume the development ledger with `O_EXCL` before opening its
   bytes and evaluate all arms in one job.
8. If development bytes are opened and any infrastructure or scientific gate
   fails, close the board without retry.
9. Open confirmation only on assessor authorization.
