Shohin MTR1 Small-MoE Transfer Gate
Status: closed negative on development, 2026-08-09. Holdout remains sealed. The result motivates direct route attribution and an MoE-native successor rather than a nearby shared-attention retry.
Question
Does trained same-family temporal revision transfer to a small open mixture of experts while its router and experts remain frozen?
MTR1 is the next scale/family point after the Qwen dense scale curve, SmolLM3-3B aggregate transfer, and the OLMo2-7B negative. It is not scratch pretraining and does not authorize the larger Qwen3.6 MoE campaign.
Host
- model:
allenai/OLMoE-1B-7B-0125-Instruct; - revision:
b89a7c4bc24fb9e55ce2543c9458ce0ca5c4650e; - license: Apache-2.0;
- topology: 16 decoder layers, width 2,048, 64 experts, 8 active experts per token, 7B total parameters and approximately 1B active parameters;
- context: 4,096 tokens.
The pinned snapshot, its config hash, and its complete file-manifest hash must be recorded before capability work starts.
Changed Factor
- The unchanged OLMoE owner writes one complete greedy draft from the source.
- A separate state of the same OLMoE model reads the identical source and exact internal draft.
- Only rank-8 LoRA projections in shared attention token mixers of the final four layers are trained for 256 updates. Router and expert parameters remain frozen and their names/hashes are checked.
- The reviser emits one coherent complete trajectory. No fieldwise averaging, external solver, verifier, answer router, or teacher runs at inference.
Training uses batch 1, accumulation 8, maximum sequence length 4,096, AdamW,
learning rate 2e-5, alpha 16, and the existing data/order seeds. Draft and
final generation are greedy with a 768-token budget and batch 4.
Data And Evaluator
MTR1 reuses the exact source-disjoint IDR/TTR geometry:
- 4,096 MATH identities, SHA-256
e0ede832...dbe5; - 4,096 logic/science identities, SHA-256
5a96859f...017; - 200 execution-verified MBPP identities, SHA-256
0b6d068b...398; - 9,655 training presentations, 1,289 development identities, and 1,279 sealed holdout identities after model-owned drafts are joined;
- unchanged exact-answer and executable-code assessors.
No target or assessor field is visible to the runtime. Holdout remains sealed until the development conjunction passes.
Matched Arms
All arms use the same source, internal draft, final prompt, evaluator, decoding budget, and inference accounting where applicable.
- trained temporal revision;
- unchanged second pass;
- generic self-refinement;
- longer source-only generation with the same two-pass token ceiling;
- best-of-two with deterministic tie handling;
- parameter/update-matched independent commitment with the draft span masked.
The report must include prompt/generated tokens, wall time, peak memory, trainable and active parameters, estimated FLOPs, router load distribution, expert utilization, and accuracy per generated token and per estimated FLOP.
Pass And Stop Rules
Mechanics must first prove a finite backward update, exact trainable-name inventory, zero trainable router/expert parameters, nonempty model-owned drafts, and complete provenance.
Development passes only if all conditions hold across all 1,289 identities:
- treatment exceeds unchanged second pass by at least 5 absolute points;
- treatment exceeds the strongest fully matched standard control by at least 3 absolute points;
- MATH, logic/science, and executable-code correct counts are each nondecreasing versus unchanged;
- every arm has complete generation and compute receipts;
- router/expert accounting is complete and no protected parameter changed.
A development miss closes exact MTR1 without seed, rank, layer, duration, prompt, threshold, or decoding rescue. A pass opens the one sealed holdout evaluation. Only a development-plus-holdout pass can authorize the larger Qwen3.6 MoE campaign.
Frozen result
All six arms completed the 1,289-identity development board:
| Arm | Correct | Accuracy |
|---|---|---|
| trained shared-attention revision | 204 | 15.8262% |
| unchanged second pass | 191 | 14.8177% |
| generic self-refinement | 169 | 13.111% |
| long single generation | 167 | 12.956% |
| best-of-two | 134 | 10.396% |
| draft-masked independent training | 189 | 14.6625% |
Treatment gains +13 answers / +1.0085 points over unchanged and +15
answers / +1.1637 points over the strongest other trained control. Both
required magnitude gates fail. Broad-domain deltas versus unchanged are math
+1, logic/science +12, and executable code 0, so retention passes but is
not enough to promote the mechanism.
Router accounting over 87 rows and 74,935 tokens uses all 64 experts in every
layer with normalized entropy near 0.932. Trained-versus-base route-count L1
drift is zero in layers 0--11 and 0.00184, 0.00381, 0.00709, and
0.01955 in layers 12--15, for an all-layer mean of 0.002018. The adapter
therefore changes a small number of answers while leaving the sparse expert
program almost unchanged.
Exact MTR1 is closed. This result rejects final-four-layer shared-attention adaptation as a sufficient MoE revision mechanism; it does not reject temporal revision on MoE in general.