Research archiveComplete tracked record · 2026

Results and dead ends belong together

The whole search,
not only the wins.

Every tracked research document is published here: theories, preregistrations, implementation receipts, positive results, negative results, corrections, audits, and current work. The main site explains today; this archive preserves how Shohin got here.

Documents
564
Readable pages
563
Source record
8.1 MB
Source revision
73ffaf8
01Research map

15 stages of the search.

This is the navigational summary. It does not replace the source record below, and it does not turn a bounded success into a broad reasoning claim.

01 · Foundation

Train the language model from scratch

A custom tokenizer, 125M-parameter decoder, decontaminated corpus, distributed training stack, 300,000-step checkpoint, benchmark harness, and the diagnosis of monodomain loss cliffs.

Training system established; raw reasoning remained weak.
02 · Behavior

Test post-training, traces, and latent interventions

Broad SFT, typed-state supervision, operator traces, semantic bridges, recurrent memory, Jacobian probes, carry motors, and direct adaptive interviews tested whether useful computation already existed in the base.

Local skills moved; broad autonomous reasoning did not.
03 · Algebra

Learn laws that compose

The S3–S9 program studied categorical registers, generators, affine laws, Cayley composition, nil-linked graphs, occurrence quotients, and global anchor closure.

Execution and bounded composition succeeded; language binding remained the bottleneck.
04 · Compiler

Bind language to physical state

Referential literal pointers, witness routes, renderer orbits, relation tensors, record buses, source deletion, and variable-cardinality state tested renderer-invariant compilation.

Many mechanics passed; robust unseen-renderer compilation repeatedly failed.
05 · Machines

Separate compiler, executor, and late reader

Episodic Functor Compiler, S4-Tied Particle Transport, quotient-algebra, and source-sealed machine protocols attacked causal custody and host-side shortcut risks.

Strong finite mechanics and no-go theorems narrowed what can count as model-owned reasoning.
06 · ETTR

Build an endogenous typed theory reactor

A 67.7M-parameter typed workspace compiled WORLD state, executed generic transactions, deleted source access, and answered late queries through a model-owned reader.

Architecture and custody passed; stable broad autonomous learning did not.
07 · Product

Measure answer quality on capable backbones

Qwen and SmolLM3 experiments compared LoRA, tied recurrent workspaces, and untied dense workspaces on math, code, science, logic, AIME, and manual composition.

T2 showed hard-math signals; C2 became a math specialist but lost code and logic.
08 · Product

Turn verified model work into better weights

The selected V12 checkpoint generated fresh math and science solutions, exact verifiers admitted 4,113 trajectories, and a 1:1 protective replay warm start trained for 400 updates.

Matched macro rose from 45.4% to 48.9%; later product work superseded this checkpoint.
09 · DIVERGE

Preserve coherent alternatives until evidence resolves them

Hard HSC1 packets failed language shift even though the correct semantics remained in their score tensor. ULC1 lifted the top two coherent interpretations into source-sealed guarded state and refined them only with verified evidence.

The learned-source mechanism recovered 256/256 fresh episodes and exposed the ownership interfaces that had to become causal.
10 · Plasticity

Learn an episode, then delete every lesson

EIC1, NVE1, and NPL2 bind natural verifier transactions to a protected 64-scalar session policy. MZE1 replaces hard-coded operation semantics with a learned recurrent law owner.

Across five source-disjoint boards, late-query transfer reaches 85.6104%—exactly the typed oracle—while the strongest non-oracle reaches 3.9185%.
11 · Compiler

Remove typed programs, names, spans, and integer parsing

EAL2 induces unseen episode laws; NCP1 compiles raw commands; OQB1 creates anonymous occurrence addresses; SVE1 emits complete value events directly from bytes.

Each owner confirms on five fixed boards. SNL1 then confirms their no-retraining composition without the exact support-intersection solver: 40,960/40,960 late answers.
12 · Binding

Learn which episode law owns each operation

OPB1 replaces explicit evidence-to-operation attachment and tests rename, coherent reindex, source scrub, and decoy interventions across five source-disjoint seeds.

30,720/30,720 bindings, 20,480/20,480 terminal states, and 40,960/40,960 answers; scrub and decoy controls fall to zero.
13 · Broad gates

Reject gains that hide domain regressions

QPT1 raised Qwen3.5-4B macro by 6.958 points but lost code. SAG1 protected code but lost MATH. VCR1 delivered a larger revision gain but missed its strict code floor by one answer.

All three directions remain evidence, but none was promoted past its frozen conjunctive gate.
14 · Temporal architecture

Draft internally, revise with a trained owner, then commit

IDR1 removes the external proposal model: one Qwen3.5-9B owner drafts, a later same-family owner revises, and a learned policy selects one complete trajectory without inference-time labels or tools.

The protected product board reaches 383/538 and 75.815% macro—67 more solved and +8.552 points over the matched original second pass.
15 · Sparse transfer

Carry temporal revision through expert routing

A learned 32,784-parameter temporal-causal gate blends owner and revision states across the final 16 layers of Qwen3.6-35B-A3B while native router and expert weights remain frozen.

On 256 source-disjoint rows: 143 correct versus 111 unchanged, +12.50 points, 38 paired wins and 6 losses. Larger-host scaling remains unmeasured.
02Source documents

Searchable ledger

564 documents, without selective memory.

Large master ledgers remain available as original Markdown; normal research documents also have formatted reading pages. Filters are descriptive and never change the underlying source.

564 documents match
Product reasoningClosed / no-go

Capability Diagnosis: 2026-07-12

Shohin is healthy as a training run but is not yet an intelligent general reasoner. At raw step 166,250 it behaves primarily as a text-completion model: it can emit fragments of familiar templates but does not reliably execute arithmetic, preserve equation invariants, apply trans…

Pretraining & dataClosed / no-go

Data, teacher, and the reasoning-distillation recipe

⤴ Superseded by MASTER PLAN.md §4 (tokenizer) + §6 (data plan) (the plan of record). Kept as background; the master plan adds the Reasoning-Gym procedural corpus, concrete token counts, and the verifier.

Architecture researchPreregistered

Shohin pretrain divergence — diagnosis & fix

Symptom During the 135M flagship pretrain, training loss held at ~1.0–1.2 for hundreds of steps, then cliffed from ~1.1 to ~3.0 within a single logging interval and pinned at a finite ~3.0 (perplexity ~20) forever after. Not NaN, not a slow drift — a step function. Recurred acros…

Architecture researchClosed / no-go

Shohin ETTR Experiment History

This file is the index for experimental Git histories that have been folded into the private repository's canonical main history. It exists so rejected, superseded, incomplete, and custody-only work remains discoverable without reactivating obsolete source files in the live archi…

Plans & synthesisPreregistered

SHOHIN-135M — Master Plan

A 72-hour, 8×H100 run to take the open ≤150M reasoning crown

Plans & synthesisPreregistered

The Shohin build plan (stage by stage)

⤴ Superseded by MASTER PLAN.md (the plan of record). Kept as background; where numbers differ (token budget ~100–200B here vs ~580B there, targets, optimizer), the master plan wins.

Pretraining & dataClosed / no-go

Shohin pretraining data admission plan

Research snapshot: 2026-07-21 . This is a historical source survey. The active admission policy is docs/research/PHASE2 DATA SELECTION STANDARD.md, and the machine-readable candidate registry is pipeline/pretrain sources.json. No source or percentage in this historical survey aut…

Architecture researchClosed / no-go

R11a v3 Preregistration: Minimal Internal Causal-Mediator Canary

R11a is the smallest internal causal-mediator canary admitted after review. A source-only pass writes six private 96-wide slots. One tied writer cell updates them for four rounds. A structurally separate source-free query pass reads the same state after decoder blocks 12, 18, and…

Architecture researchClosed / no-go

R11 Preregistration: Interchange-Trained Internal Broadcast Workspace

This broad draft proposed an Interchange-Trained Recurrent Broadcast Workspace (IT-RBW) inside the exact 30-layer, 576-wide Shohin decoder. A source-only pass writes six private 96-wide latent slots. One weight-tied neural cell refines those slots four times. A source-free query …

Theory & no-go resultsPreregistered

Apical-Basal Critical Resonance

Shohin's current reasoning experiments expose three separable bottlenecks:

Theory & no-go resultsClosed / no-go

R12 Active Verifier Query No-Go

Let a candidate be theta=(M,s), a compact residual machine and current state. An experiment supplies an event word, query, and proposed witness. The verifier returns one bit

Language compilerClosed / no-go

R12 Active Witness Allocation No-Go

Let a target threshold be theta in {1,...,N} and let an ordinary answer query at x in {1,...,N-1} return

Architecture researchPreregistered

R12 Addressed Categorical Workspace Preregistration

Working name: Addressed Categorical Workspace (ACW), trained with Counterexample-Guided Broadcast Refinement (CGBR).

Architecture researchAudit

R12 Autocatalytic Hysteretic Relation Field Preregistration

Pre-artifact launch b4dcbf0 exhausted local MPS memory in a dense parent-by-child membrane expansion before an optimizer update or output write. Source 4fc5a11 replaces it with an exact gather after adding a fail-closed one-child-per-typed-role validator. The full-geometry MPS ca…

Language compilerClosed / no-go

R12 Atomic COMMAND-Binding Oracle-Initial Diagnostic

Completed negative on H100 job 725517. This remains an isolation experiment, never a native-reasoning claim.

Architecture researchClosed / no-go

R12 Atomic Typed-Edit Algebra Preregistration

Preregistered before inspecting the sealed sparse-residual result. That result failed both required state-isolated axes. The atomic architecture is now implemented and locally qualified before its first H100 launch. Fifty-five focused and related tests pass with clean Ruff, byte …

Theory & no-go resultsClosed / no-go

R12 Axiomatic Presentation Identifiability No-Go

The candidate attempted to teach a small set of typed generators and axioms, hold out long compositions, and use relation-equivalent words plus source-deleted state interchanges to force a learned compositional action.

Architecture researchClosed / no-go

R12 Canonical Residual Naming Control

Let a deterministic Moore system have at most n reachable states, a known reset state s0, finite event alphabet Sigma, and exact observable outputs. For a history u and continuation v, write

Architecture researchClosed / no-go

R12 Causal-Address Revelation

This document preserves the useful theorem as a diagnostic for physically isolated memory banks and records the exact collapse that rejected it.

Architecture researchPreregistered

R12 causal carry motor preregistration

The post-DRS causal cycle isolated a narrow failure. On the frozen 50-case boundary board, the native model produced the exact first state in 38/50 cases. Every native failure first diverged at the serialized c= value. At a teacher-forced prefix, the newly written result digit wa…

Architecture researchClosed / no-go

R12 Dual-Provenance Carry-Motor Recovery Preregistration

The sole purpose of this protocol is to recover the already preregistered carry motor fit from a mechanical Python/JSON representation defect without claiming that new executor code produced the upstream plan or feature tensors.

Architecture researchPreregistered

R12 Causal-Delta Terminal-State Preregistration

Implemented and locally qualified on 2026-08-01. Both preregistered doses are complete on H100 and independently replicated on V100. The result is negative.

Plans & synthesisPreregistered

R12 Causal Grammar Firewall Plan

S7 proves exact bounded execution under a strong cyclic substrate. S9.1 proves that model-produced graph fields can control that executor on a templated language board. Neither result separates semantic grounding from recognition of fixed section order, ontology phrases, punctuat…

Evaluation & auditsPreregistered

R12 Causal Result-Digit Motor — Prereg (sibling to carry motor)

Parameter budget (2026-07-17): the frozen flagship has exactly 125,081,664 unique parameters (verified by instantiating the 300k checkpoint configuration with tied embeddings counted once). The system must remain strictly below 150,000,000 total parameters . Default DigitMotor is…

Evaluation & auditsResult

R12 Wide Result-Digit Motor Result

Decision: REJECT WIDE RESULT DIGIT MOTOR AS AUTONOMOUS ACTUATOR

Architecture researchClosed / no-go

R12 CDRL Neural Optimization Preregistration

Cluster: Newton H100 (normal, one GPU). Isolated output under artifacts/r12/cdrl neural v1/. Must not write flagship, ACW, or shared SFT paths.

Evaluation & auditsClosed / no-go

R12 CDRL Neural Optimization Result

Job: Newton 691750 on evc22, exit 0:0, elapsed 00:03:25. Decision SHA-256: ad94ac15ca17eaa2c5381aa0a3f94fc60a49dbbf2a528552a1212b3ecf1cabdb

Theory & no-go resultsClosed / no-go

R12 Certified Language Bridge Boundary

Flattened reasoning rows retain a question, trace, answer, and family, but not the semantic transition states needed to prove future equivalence or derive distinguishing continuations. Answer verification for OpenMath and bounded unit tests for code likewise do not prove residual…

Theory & no-go resultsClosed / no-go

R12 Closed Deliberation No-Go

A learner observes a training object Z, then spends T internal rounds choosing questions, answering them from its own state, and updating that state before emitting a hypothesis. The hoped-for claim was that this deliberation could discover target information unavailable to a one…

Theory & no-go resultsClosed / no-go

R12 Closed Late-Query Information No-Go

Implementation authority: none. This document authorizes no data, code, fit, score, CPU board, or GPU job.

Architecture researchPreregistered

R12 Closure-Tied Action Algebra

Pre-neural CPU mechanics and theory. No board seed, learned model, H100 job, or capability result is authorized by this document.

Theory & no-go resultsClosed / no-go

R12 Coherent Action Extension Audit

Implementation authority: none. This document authorizes no data build, model change, fit, score, CPU board, or GPU job.

Theory & no-go resultsClosed / no-go

R12 Commutator Factorization No-Go

Let X be N finite residual states, A be m labeled events with permutation transitions T a, and let joint late-query signatures separate states. Define

Theory & no-go resultsClosed / no-go

R12 Compiler-Prior No-Go

Let theta contain p learned bits and let a recurrent evaluator process a length-L input by

Architecture researchClosed / no-go

R12 Contractive Packet Recurrence CPU Preregistration

Decision: NO-GO as a new reasoning primitive. A source-independent projection can eliminate bounded off-manifold packet noise, but it cannot strictly contract a wrong valid semantic packet. On the frozen finite board, the favorable mechanism is exactly a five-lane repetition code…

Theory & no-go resultsPreregistered

R12 Counterfactual Conjugate Commit hypothesis

The post-DRS evidence localizes a transaction failure rather than a missing local arithmetic rule:

Evaluation & auditsResult

R12 Counterfactual Cursor-Action Mechanics Result

The immutable board contains 600 cells, 120 sources, 180 adjacent-order pairs, and 24 renderer groups. The independent audit reproduced the frozen symbolic scores exactly:

Evaluation & auditsClosed / no-go

R12 Counterfactual Cursor-Action Neural Result

Claim boundary: This is a negative result for the frozen final-block, head-zero Q-sidecar realization and its orbit-interchange training protocol. It does not prove that a learned cursor controller is impossible. It does reject promotion of this adapter, this placement, and this …

Theory & no-go resultsResult

R12 Counterfactual Cursor-Action Theory

Claim class: a bounded causal-controller training hypothesis. The cursor is not a new computational primitive. At fixed maximum depth it is exactly a finite-state transducer and, under fixed-duration steps, a positional table.

Theory & no-go resultsClosed / no-go

R12 Cross-Domain Fault-Channel No-Go

CPU protocol: R12-CROSS-DOMAIN-FAULT-CHANNEL-NO-GO-v1

Evaluation & auditsClosed / no-go

R12 Cross-Domain Fault-Channel Review Result

Decision: the combined 0/3 no-go is rejected. Replication and pure invertible transport retain narrow no-go results; globally enforced algebraic relations uniquely repair the frozen missing transition and reopen only a resource-counted learnability hypothesis.

Architecture researchClosed / no-go

R12 CTAA Neural Falsifier

Revision 2 draft architecture and custody contract. Not source-frozen. No production board seed, training seed, H100 job, development access, confirmation access, or reasoning claim is authorized by this document.

Architecture researchAudit

R12 Cursor Readout / Actuation Diagnostic

Cursor-action v1 failed because its final-block, head-zero Q intervention moved the five action logits by about 0.10 while the median gap to the full-vocabulary winner was about 2.54. The next experiment must distinguish two possibilities without tuning on the exposed v1 confirma…