Build LogOpen development record

Architecture, answers, failures

Build the reasoner in public.

A chronology of the decisions that changed the project—from the original scratch model, through the first recurrent-workspace lift, through verified source-deleted mechanism tests, failed broad gates, and the first protected product pass from trained internal revision and coherent whole-trajectory commitment.

01
2026Pretraining

Completing a 300K-step model from scratch

A complete 125M-parameter language-model stack: tokenizer, data, model, long-run training, checkpoints, and evaluation.

02
2026Systems

Tracing loss cliffs to monodomain boundaries

Recurring instability was isolated to long homogeneous data blocks and removed with deterministic domain-interleaved batches.

03
2026Generalization

Composing every unseen two-step transition

A 4,934-parameter module learned six primitives and composed all 36 withheld two-step transitions on a bounded protocol.

04
2026Architecture

Opening a tied recurrent workspace

Sixteen learned slots, one tied cross-attentive cell, eight internal updates, and a gated residual back into the prompt are now public.

05
2026Early result

Finding a short-run recurrent-workspace lift

T2 initially improved small Qwen reasoning boards and produced a coherent AIME-only solution, establishing useful specialist computation.

06
2026Correction

Rejecting T2 as a broad general adapter

Longer Qwen and SmolLM3 comparisons erased the broad lead. T2 retained hard-math evidence but lost enough general capability to stop scaling.

07
2026Scale result

Selecting the SmolLM3 dense workspace

At 3,000 updates C2 became the strongest single arm, adding 58 GSM8K and 41 MATH answers over B1 while losing code and logic.

08
2026Full evaluation

Discovering specialization, not broad improvement

On 3,392 previously unopened problems, B1 and C2 solve 1,782 and 1,781. The workspace shifts capability toward math and science rather than raising the total.

09
2026Router failure

Closing domain-label expert routing

A prompt-only domain router reached 1,790 unopened answers—just eight above B1. Subject identity is not a reliable proxy for per-problem expert advantage.

10
2026Router closure

Rejecting paired-outcome routing

On 200 fresh verified prompts B1 solves 32, C2 solves 15, and a perfect selector solves only 38. The complementary-expert hypothesis is closed.

11
2026Now

Selecting the V12 late-layer parent

A verified 16M-target comparison selected the 1,000-update late-layer checkpoint; simply training longer degraded its broad capability.

12
2026Self-improvement

Generating 4,113 verified solutions from the model

Long math rollouts and exact answer verification produced 2,313 math and 1,800 science traces from fresh, evaluation-disjoint prompts.

13
2026Training result

Turning verified rollouts into a broad model gain

A 400-update warm start with equal protective replay raised the matched five-domain macro from 45.4% to 48.9% while preserving code.

14
2026Confirmation

Advancing the replay checkpoint to expanded boards

Expanded GPQA improves from 27/198 to 34/198 and HumanEval holds at 45/164; the remaining large boards are still running.

15
2026Now

Mining only the failures for a second improvement round

The active hard-negative wave targets 1,783 unsolved math and 2,296 unsolved science prompts, admitting only newly verified solutions.

16
2026Architecture

Passing the exact DIVERGE mechanics gate

A source-sealed factorized packet preserved 252 coherent worlds with zero support loss, zero false certificates, and a 16.90× aggregate storage advantage.

17
2026Compiler failure

Closing both pooled DIVERGE compilers

The tiny compiler won only two of five seeds; the frozen SmolLM2 residual successor missed exact-packet promotion by orders of magnitude. The next design must preserve token roles and source order.

18
2026Compiler failure

Closing hard HSC1 parsing under language shift

A hierarchical structured compiler reached 96.09% exact train packets but fell to 8.59% lexical and zero renderer/composition exactness when forced to emit one irreversible parse.

19
2026Diagnosis

Finding the valid interpretation inside K=2 support

A frozen 1,024-episode audit found that semantic templates and alignments survived shift; retaining only two coherent interpretations preserved the valid world in every episode.

20
2026Architecture result

Passing the uncertainty-lifted DIVERGE gate

ULC1 recovered 256/256 fresh episodes across train, lexical, renderer, and composition cohorts while matched top-1, particle, recurrence, and soft-mixture controls remained far behind.

21
2026Causal binding

Confirming equivariant identity under a mapped mention swap

EIC1 reaches 768/768 on normal and mapped-swap boards while an equal-FLOP duplicate control falls to 384/768, qualifying the typed identity owner.

22
2026Reasoning result

Matching the oracle after every demonstration is deleted

NPL2 reaches 85.6104% late-query exactness across five source-disjoint seeds—exactly the typed oracle—while the strongest non-oracle arm reaches 3.9185%.

23
2026Execution

Replacing hard-coded transitions with a learned recurrent law owner

MZE1 learns all 75,272 finite transitions with 400 parameters, executes held programs through depth 32, and preserves the full NPL2 confirmation score.

24
2026System identification

Inducing fresh episode-local laws from natural outcomes

EAL2 reads natural before/after evidence, identifies eight unseen laws per episode, deletes the demonstrations, and answers all 40,960 confirmation queries.

25
2026Language compiler

Removing typed command programs

NCP1 maps raw variable-length commands into ordered episode-local operation pointers and stays exact under renaming and reversed command order across five boards.

26
2026Identity

Turning repeated surface identity into an anonymous address

OQB1 quotients repeated occurrences into an episode-local basis while a neural owner attaches values and queries; coherent reindexing remains exact and broken identity collapses.

27
2026Language compiler

Removing numeric spans and host integer parsing

SVE1 reads raw bytes into complete value events with no numeric-span scanner, confirming 30,720/30,720 event sequences and 40,960/40,960 answers.

28
2026Architecture result

Confirming spanless neural episode-law composition

SNL1 composes the frozen byte readers, identity bus, command compiler, and neural law synthesizer: 1,280/1,280 laws, 20,480/20,480 states, and 40,960/40,960 answers.

29
2026Compiler frontier

Removing exact evidence-to-operation binding

OPB1 later confirmed exact evidence-to-operation binding over five source-disjoint seeds, closing another controlled compiler interface.

30
2026Controlled architecture

Confirming operation binding without source leakage

OPB1 confirms 30,720 operation bindings, 20,480 terminal states, and 40,960 answers; source scrub and decoy controls fall to zero state and answer accuracy.

31
2026Product failure

Learning from QPT1's broad gain and code regression

The Qwen3.5-4B QPT1 intervention raised macro accuracy from 55.630% to 62.588% and solved 69 more problems, but its code score fell from 30/40 to 26/40, closing the exact gate.

32
2026Product failure

Preserving code with SAG1 but losing mathematics

SAG1 retained 30/40 executable-code problems and improved broad performance over B1, yet a 12-answer MATH regression versus the continued control rejected promotion.

33
2026Temporal reasoning

Teaching a model to revise its own internal draft

IDR1 uses one Qwen3.5-9B owner to draft and a later same-family owner to revise. On source-disjoint holdout, trained revision solves 625/1,279 versus 495 for the matched untrained second pass.

34
2026Architecture result

Committing one coherent whole trajectory

A model-owned commit policy raises source-disjoint holdout to 652/1,279 with exact order consistency. Its one-answer edge over an independent scorer is too small to attribute the gain specifically to antisymmetry.

35
2026Product confirmation

Passing the protected seven-task product board

Same-family draft, trained revision, and whole-trajectory commitment solve 383/538 at 75.815% macro—67 more problems and +8.552 points over the matched original second pass while retaining code.

36
2026Now

Measuring temporal revision across sparse-MoE families

Qwen3.6-35B-A3B reaches 143/256 versus 111 unchanged. Mixtral-8x22B reaches 448/1,023 versus 147 unchanged and 356 matched self-refinement; its release gate remains closed because code and baseline retention regress.

Active

The Qwen3.5-9B draft/revision/commit system is confirmed at 383/538 and 75.815% macro on the protected product board. The distinct antisymmetric-scoring claim is closed because its matched control is only one answer behind. Sparse-MoE transfer now reaches 143/256 on a source-disjoint Qwen3.6-35B-A3B screen versus 111/256 unchanged. On Mixtral-8x22B it reaches 448/1,023 versus 147 unchanged and 356 matched self-refinement. That gain is real, but the release remains closed because executable code and baseline retention regress.