← Complete research archive
Plans & synthesisPreregistered89 lines

R12 Causal Grammar Firewall Plan

S7 proves exact bounded execution under a strong cyclic substrate. S9.1 proves that model-produced graph fields can control that executor on a templated language board. Neither result separates semantic grounding from recognition of fixed section order, ontology phrases, punctuat…

R12_CAUSAL_GRAMMAR_FIREWALL_PLAN.mdOpen original Markdown ↗

R12 Causal Grammar Firewall Plan

Status: ordered successor to S9.2; design only, no board or score access

Why this stage exists

S7 proves exact bounded execution under a strong cyclic substrate. S9.1 proves that model-produced graph fields can control that executor on a templated language board. Neither result separates semantic grounding from recognition of fixed section order, ontology phrases, punctuation, and field position. Closing S9.1's final abstentions would improve the parser but would not answer that question.

The firewall is the first stage whose positive result could justify a broader language-to-machine grounding claim. It is deliberately hostile to layout and template shortcuts while preserving the same underlying graph semantics.

Factorized source interventions

Each latent graph must be rendered into paired sources along independently sampled axes:

  1. Clause topology: roster, cards, entry, events, and query may be interleaved or reordered while explicit references preserve meaning.
  2. Same-layout counterfactuals: two sources have identical token counts, punctuation, clause positions, and nonce widths but differ in exactly one operation, entity, entry, successor, nil, or query binding.
  3. Decoys: quoted, negated, superseded, or explicitly inactive cards and events are syntactically plausible but absent from the target graph.
  4. Argument order: active and passive clauses, fronted objects, and reverse mention order express the same typed relation.
  5. Ontology removal: held-out renderers omit words such as card, event, entry, control, registry, and query while retaining ordinary natural descriptions.
  6. Coreference: repeated exact surfaces, non-identical aliases, and local pronouns are independently varied so exact-byte equality is useful but not sufficient.

Training, development, and sealed confirmation must use disjoint renderer generators, nonce pools, graph instances, and combinations of these axes. Counterfactual pairs stay in the same split.

Matched systems

  • treatment compiler and unchanged bounded executor;
  • equal-budget no-occurrence-class compiler;
  • oracle-masked retrained layout-only compiler;
  • a surface/position-only finite classifier with matched label access;
  • paired-consistent shuffled supervision;
  • oracle graph upper bound; and
  • source-free and uniform-logit inference controls.

No system may inspect execution, state, answer, graph validity, or retries while choosing source fields. A compile failure is final.

Required diagnostics

  • exact graph/state/answer by intervention axis and depth;
  • exact changed-edge response on each same-layout counterfactual pair;
  • false inclusion rate for quoted, negated, superseded, and inactive decoys;
  • root, child, binding, link, entry, nil, and query error decomposition retained even when graph compilation fails;
  • operation/entity/position/event alpha invariance;
  • layout-only and surface-only train fit as well as held-out scores; and
  • manual source, emitted graph, recurrent trace, and answer transcripts sampled before aggregate interpretation.

Provisional admission floors

These numbers must be frozen with an exact board design before any score read:

  • at least 90% exact graphs on every held-out grammar family;
  • at least 98% exact changed-edge response on counterfactual pairs;
  • at least 95% exact rejection of inactive decoys;
  • 100% graph/state/answer invariance under all-symbol alpha recoding;
  • every valid emitted graph semantically exact;
  • layout-only and surface/position-only controls below 10% exact graph; and
  • treatment at least 20 points above every learned shortcut control.

Decision boundary

A pass would support robust semantic compilation across a bounded graph language. It still would not establish arbitrary program induction, a generic executor, self-generated decomposition, or general reasoning. A failure means further standard-board anchor optimization is parser engineering and must not be reported as a reasoning breakthrough. The next theory would then have to change the language representation or training identifiability, not merely add search or width.