Error-Syndrome Revision (ESR1)
Status: closed negative on development; sealed holdout remains unopened.
Motivation
SCTR1 showed that commitment is not the current bottleneck on OLMo2-7B. Always-revise preserves 221 of 222 correct internal drafts and repairs 38 of 1,067 incorrect drafts. The union of draft and revision contains only one additional correct answer beyond always-revise, so a perfect selector has a measured ceiling of 0.0776 percentage points. ESR1 therefore targets the ability to construct a correction, not the decision to keep one.
Mechanism
The source-only prompt contains the problem and the model's own internal draft. An eight-step tied recurrent workspace produces 16 model-owned soft prefix states. Standard causal language-model loss trains the full corrected response. ESR1 adds one fixed auxiliary target derived only during training:
s = mean(E[verified response tokens]) - mean(E[internal draft tokens])
L_syndrome = 1 - cosine(mean(workspace prefix), s)
L = L_LM + 0.01 L_halt + 0.25 L_syndrome
The embedding table is frozen. The exact internal-draft span is located with token offsets; truncation that removes it fails closed. At inference the verified response and syndrome target are absent. The workspace receives only the same problem-plus-draft prompt as always-revise and emits a soft prefix before one coherent replacement trajectory.
This is a bounded capability experiment, not a novelty claim. It tests whether explicitly supervising the direction of correction makes a recurrent latent workspace materially better at revision.
Matched Arms
syndrome: recurrent workspace plus fixed syndrome loss.ettr: identical recurrent workspace, parameters, recurrent depth, LoRA, data, update count, seed, and evaluation, without syndrome loss.always_revise: completed standard LoRA reviser from the same OLMo2-7B, data, update count, and evaluator.
Both new arms use 9,655 frozen training examples, 256 optimizer updates, batch size 1, accumulation 8, sequence length 4,096, LoRA rank 8 on four late layers, learning rate 2e-5, 16 workspace slots, width 512, and eight recurrent steps. Development has 1,289 identities spanning MATH-500, BBH logic/science, and executable MBPP. Eight exact evaluation shards are merged before scoring.
Frozen Gate
ESR1 passes only if all conditions hold:
- syndrome accuracy is at least 5.0 points above always-revise;
- syndrome accuracy is at least 3.0 points above the same-workspace control;
- correct-count deltas versus always-revise are nonnegative in math, logic, and code;
- all 1,289 development identities are covered by eight merged shards.
Only a conjunctive pass authorizes the sealed holdout. A failure closes this exact objective without weight, width, duration, seed, prompt, or threshold variants.
Result
Jobs 746561/746562 completed all 256 matched updates. Sixteen development
shards completed after exact replacements 746597/746598 for two jobs that
landed on evc46, where no CUDA device was exposed. The replacements changed
no scientific setting. Exact comparison 746600 produced:
| Arm | Correct | Accuracy |
|---|---|---|
| ESR1 syndrome | 255/1,289 | 19.7828% |
| Same-workspace control | 239/1,289 | 18.5415% |
| Standard always-revise | 259/1,289 | 20.0931% |
The syndrome objective adds 16 answers, or 1.2413 points, over the identical
workspace. It nevertheless trails the simpler always-revise model by four
answers, or 0.3103 points. Versus always-revise, domain correct-count deltas
are math -9, logic/science +9, and executable code -4. All substantive
promotion gates fail. Final syndrome loss is 0.95396; the correction target
did not become a sharply aligned latent direction.
The qualified finding is narrow: embedding-residual supervision causes a
small improvement over an otherwise identical recurrent workspace, but the
workspace itself is inferior to direct trained revision and ESR1 is not a
transferable reasoning architecture. Holdout was not opened. Comparison
SHA-256 is
de34596bf5efe0c76674a2c44fa9bcbd4e5d42911a45f8d5559dc6332a710ab4.
The two fits, sixteen successful shards, and two four/21-second infrastructure
failures charged 0.763 aggregate H100-hours.