The validation result
8 unseen books // 40 windows // 40,920 next-token targets
Eight books // eight gains
The aggregate is not being dragged by one lucky record
| Record | Book | Baseline PPL | RC PPL | Reduction | Improved |
|---|---|---|---|---|---|
| 2 | Travels in Morocco, Vol. 2 | 39.6202 | 33.0576 | 16.564% | 4 / 5 |
| 3 | Impressions of Theophrastus Such | 51.0193 | 42.4333 | 16.829% | 4 / 5 |
| 4 | Odd Craft, Part 4 | 19.8284 | 18.3424 | 7.494% | 4 / 5 |
| 5 | The S. W. F. Club | 19.3801 | 17.7407 | 8.459% | 4 / 5 |
| 6 | Frank Merriwell Down South | 24.8161 | 21.4131 | 13.713% | 4 / 5 |
| 7 | Critical Miscellanies, Vol. 2 | 31.0531 | 25.0763 | 19.247% | 4 / 5 |
| 8 | From Sand Hill to Pine | 36.0620 | 32.2304 | 10.625% | 4 / 5 |
| 9 | Child's Health Primer | 14.0101 | 11.9957 | 14.378% | 4 / 5 |
Clean pattern, sharp exception: the first window of every book worsened by 1.46%–3.97%; all 32 later windows improved. The locked protocol kept every unfavorable window. The two largest book gains account for only 34.3% of positive NLL gain.
The mechanism
A measured leak from deep state to shallow state, one token later
t
layer 4
layer 11
first pass
α = 0.15 // 10-token ramp // one additional iteration // hidden-axis L2 norm match
Mechanical sequence
- 01 // warmup
- No K/V committed on the first rolling step.
- 02 // capture
- Take residual output after destination layer 4.
- 03 // normalize
- Rescale layer-11 source to the destination vector's L2 magnitude.
- 04 // mix
- Convex blend at αt/βt, then inject as input to layer 5.
- 05 // advance
- Commit only the rolling pair's left token; the newest token remains first-pass readout.
Fixed calibration
- Source → target
- 11 → 4, zero-based
- Mix
- α = 0.15 // β = 0.85
- Ramp
- αt = min(t / 10, 1) × 0.15
- Recurrence
- 1 additional iteration
- Weights
- 999,885,952 // frozen // unchanged
Matched protocol
Nothing changes between lanes except recirculation state flow
Model lane
- Checkpoint
- google/gemma-3-1b-pt
- Revision
- fcf18a2…d8eb29
- Artifact hash
- ee5250f6…69cb27
- Numerics
- bfloat16 model // float32 NLL
- Backend
- PyTorch 2.10 // MPS // eager attention
Data lane
- Corpus
- PG-19 test, pinned parquet mirror
- Selection
- First 5 complete windows × unseen records 2–9
- Window
- 1,024 tokens // no padding
- Scored
- 1,023 targets per window
- Identity
- Ordered per-window SHA-256 verified in both lanes
Safety interlocks
The positive number only counts if the machine is doing the intended computation
Vendored baseline and official Transformers logits are bitwise equal on the real checkpoint.
No trainable parameters; version counters remain unchanged after evaluation.
530 mixes for 530 fixture tokens, each with destination/source shape [1, 2, 1152].
Source is rescaled per token on the hidden axis before α/β mixing.
Changing tokens 514–529 produces zero change through position 513, crossing the 512-token sliding-attention boundary.
Native cache lengths match: 511 positions in sliding layers and 529 in full-attention layers.
Records 2–9 and all 40 ordered hashes are locked; no record or token hash overlaps the exploratory sample.
First-64-token logits and a full 1,024-token target-loss vector exactly match the reference-derived evaluator.
Shop cost
Quality went up; serial prefill throughput went through the floor
The confirmatory reference-style path walks the full prefill token by token: 10.585 seconds for ordinary forward versus 1,240.263 seconds for recirculation. It validates the method, not production latency.
+2.340 GB sampled after synchronized windows
6.137 GB → 8.478 GB
The loop hits the road
Locked Gemma 3 1B IT transfer // two signals, two regressions
Capability status: MIXED. The exact PT-calibrated mechanism moved MMLU-Pro and normalized HellaSwag upward, but GSM8K and IFEval downward. Lower PG-19 perplexity did not transfer uniformly into useful-task accuracy.
| Locked primary metric | Baseline | RC | Delta | Flips + / − | Paired p | Call |
|---|---|---|---|---|---|---|
| MMLU-Pro exact match | 9.524% | 11.905% | +2.381 pp | 2 / 1 | 1.000 | IMPROVED |
| GSM8K flexible exact match | 40.000% | 30.000% | −10.000 pp | 3 / 8 | 0.227 | REGRESSED |
| IFEval prompt strict | 58.000% | 54.000% | −4.000 pp | 3 / 5 | 0.727 | REGRESSED |
| HellaSwag normalized | 41.000% | 43.000% | +2.000 pp | 6 / 4 | 0.754 | IMPROVED |
What changed
- MMLU
- One gain was stricter answer formatting: baseline knew D but the extractor rejected bold D; recirculation emitted (D).
- GSM8K
- Three wrong→right flips were outweighed by eight right→wrong flips, including double-counting dozen prices.
- IFEval
- Extra text sometimes completed structure, but also added forbidden capitals, surplus answers, or a postamble.
- HellaSwag
- Normalized choice scoring gained two net items; raw accuracy stayed exactly 43% → 43%.
Capability shop cost
- Baseline
- 1,690.40 s evaluator time
- Recirculation
- 6,111.64 s evaluator time
- Combined
- 2 h 10 m 2 s // 3.62× lane ratio
- Output length
- Longer on 91 / 142 generative pairs; +5,868 tokens net
- Verification
- All prompt, token, config, score, frozen-weight, cache, and causal checks passed
Capability uncertainty // no clean transfer claim
Every paired-bootstrap interval includes zero and every exact paired p-value is at least 0.227. These fixed subsets are directional evidence, not full-benchmark estimates. The two positive primary movements are small; the 10-point GSM8K regression is the largest capability delta observed. Read the full paired report or inspect the machine-readable comparison.
Dust in the bearings
What this result does and does not establish
Locked confirmation // broader, still bounded
The current confirmatory run evaluates 40,920 predicted tokens across 40 windows from eight previously unseen books—about 0.379% of the paper's 10.8M-token PG-19 evaluation. Its 13.503% reduction is close to the paper's 14.41%, but eight sequential records are not a substitute for the full corpus.
Earlier exploratory run // preserved
The earlier exploratory run used 10 windows and 10,230 targets from records 0–1, reducing perplexity by 13.676% with 8 of 10 windows improved. Its initial two-window pilot worsened by 0.396%; it was extended once by a content-blind rule and then stopped. Those artifacts and caveats remain intact and are not presented as the current confirmatory sample.
Systematic window boundary exception
In the locked confirmatory run, window 0 of all eight books worsened by 1.46%–3.97%, while every one of the 32 later windows improved. The protocol retained all eight unfavorable windows. This pattern is real in the sample and unexplained; it is a reason to avoid overgeneralizing the aggregate.
Earlier exploratory mask failure // caught and rejected
During the earlier exploratory work, one diagnostic incorrectly gave sliding layers the full-attention mask, inflating pilot baseline PPL to 75.636. An official-forward equivalence check exposed it. Native 512-token sliding masks restored pilot baseline PPL to 19.201; only corrected numbers appear in either experiment's final result files.
Paper ambiguities and uncertainty
The paper does not pin software, model, tokenizer, data revisions, dtype, or device. It introduces the 1B ten-token ramp but does not explicitly state whether Table 1 includes it. The confirmatory whole-book bootstrap is 10.70%–16.12%, but these are sequential rather than randomly sampled books, so it is descriptive stability evidence—not population-level confidence.
Perplexity did not transfer cleanly
The PT artifact establishes lower perplexity on its specified evaluations—not an equivalent gain in intelligence. The new IT capability milestone is mixed: two small positive primary deltas coexist with GSM8K and IFEval regressions. It still does not establish broad gains in reasoning, coding, instruction following, agentic task completion, or general model intelligence. The paper's near-zero generation-latency claim is also not tested here: this implementation measures serial reference computation.