Generate one complete internal trajectory
A pinned Qwen3.5-9B B1 owner receives only the source problem and produces a full first attempt.
Same-family temporal reasoning
Shohin's current product architecture gives one model family multiple temporally distinct roles: produce a complete internal draft, revise it with a trained later owner, compare the coherent trajectories, and commit one final answer. The mechanism now has a confirmed broad-board result.
Reasoning can unfold across learned roles
A standard generation commits tokens while it is still deciding what the answer should be. Shohin preserves that first trajectory, gives a later owner the explicit job of revision, and delays the final commitment until both coherent attempts exist.
Model-owned temporal workspace
Both generative owners use the same pinned model family with different small adapter states. The commit owner sees complete candidate trajectories, and all inference-time supervision channels used during training are absent.
A pinned Qwen3.5-9B B1 owner receives only the source problem and produces a full first attempt.
A distinct small LoRA state on the same pinned backbone sees source plus internal draft and emits a complete replacement solution.
A learned commit owner evaluates the original and revised trajectories without candidate provenance, task routing, or correctness labels.
Selection never averages tokens or answer fields. The output is one intact model-generated trajectory.
source → internal draft → trained revision → coherent comparison → one final trajectoryProtected product board · seven task families
The matched original-B1 control gets the same internal draft and a second generation pass. Trained revision still adds 58 product answers by itself; commitment raises the final result to 383/538.
This is the project's first protected broad product pass. It is a post-trained architecture on a Qwen3.5-9B host—not yet a result from the 125M scratch checkpoint and not a claim of frontier-model parity.
DIVERGE · independently falsified owners
Before the product pass, the project isolated source-deleted state, learned execution, anonymous identity, raw command and value reading, and unseen-law synthesis in controlled worlds. These are mechanism results, not the source of the 9B public benchmark claim.
Natural verifier transactions drive branch-local learning. After source deletion, transfer exactly matches the typed oracle on all five seeds; the strongest non-oracle reaches 3.9185%.
Outcome-only training recovers every transition and composes the learned laws through depths 4, 8, 16, and 32 while preserving NPL2 exactly.
Natural before/after evidence identifies eight fresh laws per episode, then source-deleted execution answers every late confirmation query.
Raw variable-length commands compile into ordered episode-local operation pointers without typed symbols, target length, or inference-time alignment.
Unseen names and coherent table reindexing remain exact. Breaking occurrence identity or crossing owner bases collapses the result.
Raw bytes become complete value-bearing evidence with no numeric-span scanner or integer parser; all 40,960 late answers remain exact.
The qualified readers feed a neural law synthesizer without exact support intersection: 1,280/1,280 laws and 20,480/20,480 terminal states across five boards.
Five source-disjoint seeds preserve exact operation binding, terminal state, and late answers under renaming and reindexing; scrub and decoy controls collapse to zero.
Same-family owners separate proposal, correction, and commitment while preserving complete trajectories and inference-time autonomy.
383/538 · 75.815% macroSuccessive gates remove typed commands, lexical identity, numeric spans, host parsing, exact support search, and explicit operation binding while maintaining destructive controls.
Five source-disjoint confirmation boards per ownerWhat is established
What remains engineered
Evidence before adjectives
The first sparse-MoE source-disjoint screen is positive: 143/256 versus 111/256 unchanged. The next falsifiable step is to test the same owner-to-revision temporal intervention on larger sparse hosts with matched unchanged and self-refinement controls.