EvidenceUpdated 19 August 2026

The live ledger

Compare the same model.
Change one mechanism.

Shohin now has a protected broad-board pass. The evidence chain separates the value of an extra inference pass, trained internal revision, and whole-trajectory commitment—and preserves the failed gates that prevented earlier improvements from being overclaimed.

01Protected product result

Qwen3.5-9B · seven task families

Draft, revision, and commitment pass together.

The first owner generates a complete internal draft. A trained same-family owner revises it. A final learned policy selects one intact trajectory without external tools or inference-time labels.

Confirmed · 08 August 2026
ModelScaleCheckpointMacroGSM8KMATH-500Executable codeGPQABBH logicAIME 2024
Shohin · draft/revise/commitConfirmed product resultQwen3.5-9B · same-family owners383 / 538 solved · protected board75.815%87 / 10072 / 10035 / 40114 / 19875 / 1006 / 30
Trained internal revisionRevision resultQwen3.5-9B · IDR1374 / 538 solved75.005%88 / 10069 / 10035 / 40104 / 19878 / 1003 / 30
Matched original second passMatched controlQwen3.5-9B · B1316 / 538 solved67.263%85 / 10060 / 10034 / 4062 / 19875 / 1004 / 30
Coherent trajectory oracleSelection ceilingNot deployable · upper bound399 / 538 solved78.619%90 / 10077 / 10035 / 40118 / 19879 / 1006 / 30
MeasurementResultStatus
Protected seven-task product macrosame-family draft/revision/commit versus the matched original-B1 second pass67.263% → 75.815%Confirmed
Total solved problems+67 solved on the exact protected board316 → 383 / 538Confirmed
MATH / GPQA movementthe largest gains occur in mathematics and graduate-level science60 → 72 · 62 → 114Confirmed
Executable code retentionthe broad gain does not trade away the protected code floor34 → 35 / 40Confirmed
AIME 2024small absolute sample; reported separately from the 538-problem solved count4 → 6 / 30Confirmed
Independent commit controlnear-identical performance means antisymmetry itself is not the established cause382 / 538 · 75.813%Confirmed

The 538-problem solved count covers GSM8K, MATH-500, executable code, GPQA, and BBH. AIME is reported separately inside the same frozen evaluation. The coherent oracle is a diagnostic ceiling, not a deployable system.

02Sparse-MoE transfer

Qwen · 3B active / Mixtral · 39B active

Temporal revision transfers across two sparse-MoE families.

Qwen's causal gate solves 143 versus 111 on 256 rows. On an independent 1,023-row validation, Mixtral's trained revision solves 448 versus 147 unchanged and 356 under matched self-refinement while every native router and expert remains frozen.

MeasurementResultStatus
Temporal-causal gate+32 correct and +12.50 percentage points on the exact source-disjoint Q36 screen111 → 143 / 256Measured
Paired capabilitytwo-sided exact McNemar p = 9.43 × 10⁻⁷ versus the unchanged host38 wins · 6 lossesMeasured
Unchanged-case retentionthe intervention preserves most identities the unchanged host already solved105 / 111 · 94.59%Measured
Domain movement86/128 BBH logic, 46/117 MATH-500, and 11/11 MBPP after interventionBBH +15 · MATH +15 · MBPP +2Measured
Routing-only ablationselector supervision alone gains +27; causal response training adds five net answers over that ablation138 / 256Measured
Native router / expert weights32,784 gate parameters blend owner and revision states across the final 16 layers0 trainableAudited
Mixtral unchanged
147
of 1,023
Self-refinement
356
+209 answers
Trained revision
448
+301 answers
Selective commit
287
93.2% retention

Boundary: Mixtral's trained revision is statistically decisive over both controls, but it scores 0/22 on executable code. The selective commit restores code yet misses the 95% retention gate by three baseline answers. This is measured transfer, not a qualified release.

03Source-disjoint attribution

IDR1 + learned commitment

The learned policy—not merely pass two—produces the gain.

Both arms receive the same source, internal draft, decoding budget, evaluator, and second generation pass. Swapping only the revision LoRA state changes holdout from 495 to 625 correct. Commitment then reaches 652 while remaining invariant to candidate order.

Qwen3.5 scale curve · identical holdout

The architecture lifts every measured scale.

OriginalWith architecture
IntelligenceModel parameters · logarithmic spacing

0.8B parameters: original 18.92%; with Shohin architecture 25.65%; improvement +6.72 percentage points.

4B parameters: original 29.71%; with Shohin architecture 43.32%; improvement +13.61 percentage points.

9B parameters: original 38.70%; with Shohin architecture 48.87%; improvement +10.16 percentage points.

“Intelligence” is represented here by exact accuracy on the same 1,279-example source-disjoint MATH, logic/science, and executable-code holdout. Hollow points are unchanged same-model second passes; solid points are trained internal-draft revision. Lines show the causal lift, not a fitted scaling law.
MeasurementResultStatus
IDR1 source-disjoint holdouttrained revision versus 495/1,279 for the matched original-B1 second pass625 / 1,279Confirmed
Learned revision contribution+83 MATH, +46 logic/science, and +1 execution-verified code+130 answersConfirmed
Whole-trajectory commit holdout27 answers over IDR1 and seven over the frozen metadata selector652 / 1,279Confirmed
Matched independent committhe one-answer difference rejects a distinct antisymmetric-mechanism claim651 / 1,279Confirmed
Candidate-order consistencythe commit decision is invariant to presentation order100% · zero swap errorConfirmed
Inference-time external supervisionno tool, task router, answer label, correctness bit, or verifier feedbackNoneAudited
04Failed product gates

Useful movement · rejected promotion

Every regression remains part of the result.

QPT1, SAG1, VCR1, and source-only distillation each revealed a useful mechanism or capability direction. Each also failed a preregistered conjunctive gate, so none was promoted as the final architecture.

MeasurementResultStatus
QPT1 on Qwen3.5-4B+69 solved overall, but executable code regressed from 30/40 to 26/4055.630% → 62.588%Failed
SAG1 on Qwen3.5-4Bretained 30/40 code but lost 12 MATH answers versus the continued control62.253% · 292 / 538Failed
VCR1 external-candidate revisionlarge multistage gain, but missed its strict code promotion floor by one answer72.302% · 368 / 538Failed
Source-only distillation controlremoving candidate trajectories loses 153 answers versus candidate-conditioned revision490 / 1,279 holdoutFailed
05Controlled source-deleted reasoning

NPL2 · five source-disjoint seeds

The retained session state—not the source—solves transfer.

In the controlled DIVERGE line, natural verified transactions update a 64-scalar session policy. Demonstrations, verifier messages, transcripts, and source bytes are deleted before held late queries.

MeasurementResultStatus
NPL2 source-deleted late-query transferexactly equal to the typed PL1 oracle on every one of five source-disjoint seeds35,066 / 40,960 · 85.6104%Confirmed
Exact terminal state-pair transferretained after demonstrations, verifier messages, and source bytes were deleted85.3223%Confirmed
Episode-local operation-map recoverythe mutable memory is only a 64-scalar session policy85.0781%Confirmed
Strongest non-oracle armmore than 81 points below the verified-plasticity treatment3.9185%Confirmed
Reset / wrong branch / transplantthe acquired episode state—not a static answer shortcut—causes the transfer1.057% / 1.067% / 1.028%Confirmed
Poison and exact rollbackpoison changes every session and rollback restores every output and receipt1,280 / 1,280 sessionsConfirmed

Qualified claim: controlled natural mini-language reasoning with a structural world interface, exact executor/verifier, bounded algebra, and typed transfer programs. It is not the cause asserted for the 9B product score.

06Controlled architecture owners

Scaffold removal · destructive controls

State, identity, language, laws, and execution were isolated.

Frozen gates replaced hard-coded transition semantics, typed command programs, lexical identity, numeric spans, host integer parsing, and exact law-support intersection. OPB1 then closed operation binding.

MeasurementResultStatus
MZE1 learned recurrent executor400 learned parameters; exact at held depths 4, 8, 16, and 32; shifted control 0.2657%75,272 / 75,272 transitionsConfirmed
EAL2 unseen episode-law inductioneight fresh laws per episode inferred from natural before/after outcomes across five boards40,960 / 40,960 queriesConfirmed
NCP1 raw command compilationnormal, renamed, and reverse-order commands; source scrub and shuffled control are zero20,480 / 20,480 programsConfirmed
OQB1 occurrence-quotient identity busunseen renaming and coherent reindexing remain exact; broken identity controls collapse40,960 / 40,960 answersConfirmed
SVE1 spanless value-event readerno numeric-span scanner or raw-integer parser; 40,960/40,960 late answers30,720 / 30,720 sequencesConfirmed
SNL1 neural law composition1,280/1,280 laws and 20,480/20,480 states without exact support intersection40,960 / 40,960 answersConfirmed
OPB1 evidence-to-operation binding20,480/20,480 states and 40,960/40,960 answers; source scrub and decoy controls are zero30,720 / 30,720 bindingsConfirmed

Established

  • Trained same-family revision adds 130 source-disjoint answers over a matched second pass.
  • The complete product system adds 67 solved problems and +8.552 macro points.
  • Commit decisions are exactly invariant to candidate presentation order.
  • Code remains at 35/40 while mathematics and GPQA improve strongly.
  • Controlled DIVERGE mechanisms retain their source-deletion and causal results.

Not established

  • A unique causal advantage from antisymmetric scoring over independent candidate scoring.
  • A monotonic sparse-MoE scaling law or capability-preserving release across every host.
  • The same capability in the original 125M scratch checkpoint.
  • Frontier-model parity or unrestricted open-domain reasoning.
  • Permission to tune on the protected product board.
07Original Shohin foundation

From scratch · July 2026

The 125M model established the full training stack.

The scratch checkpoint proved the independent tokenizer, corpus, distributed training, checkpointing, contamination audit, and evaluation systems. Its weak reasoning performance motivated the architecture program.

MeasurementResultStatus
Trained parametersdense decoder-only baseline trained from scratch125,081,664Audited
Optimizer stepsraw pretraining checkpoint300,000Audited
Corpus capacitydecontaminated mounted corpus; nominal exposure included replay57.826B tokensAudited
Raw public reasoningthe reason the project moved to architecture and post-trainingLow single digitsConfirmed
08Earlier bounded precursor

Historical composition result

Six primitives. Thirty-six unseen pairs.

The early transition generator showed that tiny learned laws could compose. It did not establish broad reasoning, but supplied one of the project's first clean execution signals.

MeasurementResultStatus
Learned transition parameterstrained on six primitive one-step cells4,934Confirmed
Held-out two-step transitionsnone appeared during training36 / 36Confirmed
End-to-end exact statebounded precursor protocol97.607%Confirmed
End-to-end answerssame sole frozen confirmation protocol98.096%Confirmed