The live ledger
Compare the same model.
Change one mechanism.
Shohin now has a protected broad-board pass. The evidence chain separates the value of an extra inference pass, trained internal revision, and whole-trajectory commitment—and preserves the failed gates that prevented earlier improvements from being overclaimed.
Qwen3.5-9B · seven task families
Draft, revision, and commitment pass together.
The first owner generates a complete internal draft. A trained same-family owner revises it. A final learned policy selects one intact trajectory without external tools or inference-time labels.
Confirmed · 08 August 2026| Model | Scale | Checkpoint | Macro | GSM8K | MATH-500 | Executable code | GPQA | BBH logic | AIME 2024 |
|---|---|---|---|---|---|---|---|---|---|
| Shohin · draft/revise/commitConfirmed product result ↗ | Qwen3.5-9B · same-family owners | 383 / 538 solved · protected board | 75.815% | 87 / 100 | 72 / 100 | 35 / 40 | 114 / 198 | 75 / 100 | 6 / 30 |
| Trained internal revisionRevision result ↗ | Qwen3.5-9B · IDR1 | 374 / 538 solved | 75.005% | 88 / 100 | 69 / 100 | 35 / 40 | 104 / 198 | 78 / 100 | 3 / 30 |
| Matched original second passMatched control ↗ | Qwen3.5-9B · B1 | 316 / 538 solved | 67.263% | 85 / 100 | 60 / 100 | 34 / 40 | 62 / 198 | 75 / 100 | 4 / 30 |
| Coherent trajectory oracleSelection ceiling ↗ | Not deployable · upper bound | 399 / 538 solved | 78.619% | 90 / 100 | 77 / 100 | 35 / 40 | 118 / 198 | 79 / 100 | 6 / 30 |
The 538-problem solved count covers GSM8K, MATH-500, executable code, GPQA, and BBH. AIME is reported separately inside the same frozen evaluation. The coherent oracle is a diagnostic ceiling, not a deployable system.
Qwen · 3B active / Mixtral · 39B active
Temporal revision transfers across two sparse-MoE families.
Qwen's causal gate solves 143 versus 111 on 256 rows. On an independent 1,023-row validation, Mixtral's trained revision solves 448 versus 147 unchanged and 356 under matched self-refinement while every native router and expert remains frozen.
- Mixtral unchanged
- 147 of 1,023
- Self-refinement
- 356 +209 answers
- Trained revision
- 448 +301 answers
- Selective commit
- 287 93.2% retention
Boundary: Mixtral's trained revision is statistically decisive over both controls, but it scores 0/22 on executable code. The selective commit restores code yet misses the 95% retention gate by three baseline answers. This is measured transfer, not a qualified release.
IDR1 + learned commitment
The learned policy—not merely pass two—produces the gain.
Both arms receive the same source, internal draft, decoding budget, evaluator, and second generation pass. Swapping only the revision LoRA state changes holdout from 495 to 625 correct. Commitment then reaches 652 while remaining invariant to candidate order.
Qwen3.5 scale curve · identical holdout
The architecture lifts every measured scale.
0.8B parameters: original 18.92%; with Shohin architecture 25.65%; improvement +6.72 percentage points.
4B parameters: original 29.71%; with Shohin architecture 43.32%; improvement +13.61 percentage points.
9B parameters: original 38.70%; with Shohin architecture 48.87%; improvement +10.16 percentage points.
Useful movement · rejected promotion
Every regression remains part of the result.
QPT1, SAG1, VCR1, and source-only distillation each revealed a useful mechanism or capability direction. Each also failed a preregistered conjunctive gate, so none was promoted as the final architecture.
NPL2 · five source-disjoint seeds
The retained session state—not the source—solves transfer.
In the controlled DIVERGE line, natural verified transactions update a 64-scalar session policy. Demonstrations, verifier messages, transcripts, and source bytes are deleted before held late queries.
Qualified claim: controlled natural mini-language reasoning with a structural world interface, exact executor/verifier, bounded algebra, and typed transfer programs. It is not the cause asserted for the 9B product score.
Scaffold removal · destructive controls
State, identity, language, laws, and execution were isolated.
Frozen gates replaced hard-coded transition semantics, typed command programs, lexical identity, numeric spans, host integer parsing, and exact law-support intersection. OPB1 then closed operation binding.
Established
- Trained same-family revision adds 130 source-disjoint answers over a matched second pass.
- The complete product system adds 67 solved problems and +8.552 macro points.
- Commit decisions are exactly invariant to candidate presentation order.
- Code remains at 35/40 while mathematics and GPQA improve strongly.
- Controlled DIVERGE mechanisms retain their source-deletion and causal results.
Not established
- A unique causal advantage from antisymmetric scoring over independent candidate scoring.
- A monotonic sparse-MoE scaling law or capability-preserving release across every host.
- The same capability in the original 125M scratch checkpoint.
- Frontier-model parity or unrestricted open-domain reasoning.
- Permission to tune on the protected product board.
From scratch · July 2026
The 125M model established the full training stack.
The scratch checkpoint proved the independent tokenizer, corpus, distributed training, checkpointing, contamination audit, and evaluation systems. Its weak reasoning performance motivated the architecture program.
Historical composition result
Six primitives. Thirty-six unseen pairs.
The early transition generator showed that tiny learned laws could compose. It did not establish broad reasoning, but supplied one of the project's first clean execution signals.