← Complete research archive
Architecture researchAudit125 lines

DIVERGE CWC1 - EWC1 - NPL2 Integration

The confirmed CWC1 whole-world selector can remove EWC1's semantic shortcut by choosing one complete candidate before structural extraction. Frozen EWC1 then acts only as a transcription layer inside that selected candidate. If the selected typed WORLD is exact, unchanged confirm…

docs/research/DIVERGE_CWC1_EWC1_NPL2_INTEGRATION.mdOpen original Markdown ↗

DIVERGE CWC1 -> EWC1 -> NPL2 Integration

Status: confirmed across development and five fixed confirmation seeds.

Capability hypothesis

The confirmed CWC1 whole-world selector can remove EWC1's semantic shortcut by choosing one complete candidate before structural extraction. Frozen EWC1 then acts only as a transcription layer inside that selected candidate. If the selected typed WORLD is exact, unchanged confirmed NPL2 should retain its late-query reasoning score.

No component is retrained. This is one bounded composition test, not an EWC1 retry and not authorization for continuation pretraining.

Frozen path

two complete WORLD candidates + natural directive
  -> confirmed CWC1 involution selector
  -> one physical complete candidate
  -> frozen EWC1 structural extractor
  -> typed WORLD; source deleted
  -> unchanged NVE1/EIC1/NPL2/executor/verifier
  -> late QUERY

CWC1 checkpoint cae4d896...ab1a and confirmation result f42802ce...5a9d are immutable. EWC1 checkpoint 0816ed1c...d1ed is used only under its documented boundary: normal structure was 4096/4096, but it failed semantic source scrub and is not a semantic owner. The confirmed NPL2, EIC1, NVE1, STI1, executor, verifier, base checkpoint, and tokenizer remain unchanged and hash-checked.

Wrapper board

Every existing NPL2 WORLD program is paired with a deterministic decoy using the same aliases, registers, depth, and surface grammar but different initial state and operation sequence. The true candidate position is exactly balanced within each of all 64 positive/negative CWC renderer pairs. Candidate labels, sources, and identities are unique. The original program and all downstream assessor data are unchanged.

Development wraps the fixed 256-episode NPL1/NPL2 development split, yielding 7,168 candidate decisions. Five conditional confirmation wrappers use the five already-frozen NPL2 confirmation seeds. All wrapper JSONL files and input hashes must be materialized and audited before the first integrated score.

Revision 0077e78-r1 is admitted before scoring. Every split has 7,168 rows, all 64 renderer pairs, exact 3584/3584 target balance, and 56/56 targets inside every renderer. Across six splits there are 43,008 unique sources, 43,008 unique identities, and 86,016 unique candidate labels with zero cross-split overlap. Independent clean regeneration is byte-identical. The aggregate audit SHA-256 is 23ca6485b3c0b7ce7e7073a5460a726d6739969ffe06281b4ea822baf33ff4a3. Wrapper SHA-256 values are development cd5e53e7...e98a9 and confirmation seeds 7fde2eb1...cb64, f94a63e0...c369f, e03c407e...f9c83, 22923c76...08f2, and 04f6c9cc...f972.

Conjunctive development gate

  • CWC normal and mapped-counterfactual selection at least 99%;
  • mapped counterfactual selects the original decoy at least 99%;
  • directive scrub between 49% and 51%, maximum absolute margin at most 1e-6, exact projection;
  • selected EWC typed WORLD at least 99% joint exact;
  • forced-opposite typed WORLD at most 1% exact against the true WORLD while EWC transcribes the decoy itself at least 99%;
  • NPL2 late-query exactness at least 80%, within five points of PL1 oracle, and at least ten points above every non-oracle arm;
  • unchanged NPL2 reset, shuffled-credit, wrong-branch, transplant, eligibility, rollback, source-deletion, and protected-owner gates;
  • all source, data, checkpoint, runtime, and result hashes exact.

Only a conjunctive pass opens the five fixed confirmation jobs. A failure closes this exact composition without seed, width, threshold, renderer, duration, parser, or loss variants.

Claim boundary

A confirmation pass would establish controlled end-to-end composition of a learned natural directive selector, learned structural transcription, and confirmed source-deleted NPL2 reasoning. Candidate generation, the mini-language, execution, and verification remain engineered. It would not establish unrestricted language reasoning, an involution score advantage, or permission for a long pretraining run.

Frozen result

Development job 744665 passes the unchanged NPL2 conjunctive gate at 84.8145%, exactly equal to PL1 oracle and above the strongest non-oracle arm at 3.4668%. Primary result SHA-256 is b0fffe1a...ff24d.

The frozen base NPL2 report format did not serialize the wrapper's WORLD receipt. A reporting-only patch added a separate atomic compilation report; no weights, data, prompts, thresholds, evaluator, or control changed. Exact isolated replay 744666 reproduces every score and records:

  • CWC normal selection 7168/7168;
  • mapped counterfactual selection 7168/7168 and 0/7168 against the original physical world;
  • directive scrub 3584/7168 with exactly zero margin;
  • selected EWC typed WORLD 7168/7168;
  • forced opposite WORLD 0/7168 against true while EWC transcribes that decoy 7168/7168;
  • exact zero projection residual.

Replay result and compilation SHA-256 values are 19155055...3e83 and ceee236a...e4b.

Five concurrent confirmation jobs 744668--744672 pass, and CPU aggregate 744673 passes every original NPL2 and new WORLD-composition condition. Aggregate NPL2 is 35,066/40,960 = 85.6104%, exactly equal to oracle. The strongest non-oracle arm is 3.9185%; EVIDENCE is 614,321/614,400 and QUERY is 40,960/40,960. Per-seed NPL2/oracle rates are 87.0972%, 82.6416%, 85.7178%, 87.3169%, and 85.2783%. Every seed has 100% selector and selected-structure exactness, chance/zero-margin directive scrub, and complete causal flip under mapped counterfactual. Aggregate SHA-256 is fb84120ea92aec7b6b44e833ef3603c23b767eea53954c178741f98b92fd5c33.

This confirms the bounded composition. It does not change the earlier CWC1 control finding: standard augmentation selected the controlled candidates equally well. It also does not remove engineered candidate construction, mini-language execution, or verification.