Snapshot-bracket evaluation
Snapshot-bracket evaluation is the project's experimental control protocol for comparing mutable-model learning methods.
Project status: Required by the research contract; implementation and historical failures are documented. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Record a baseline and trajectory cursor, save a snapshot, perform one experimental intervention, measure learning and retention, then restore the baseline and record the final cursor. For strict comparisons, verify restored outputs/state and control background workers. Train submissions return commits; evaluation must wait until the intended update is visible.
Implementation and controls
Trainfer exposes snapshot save/load and trajectory APIs; cont runners place restoration in a finally block. Current snapshot loading resets optimizer state and reapplies or clears residual hooks. The protocol records model, quantization, adapter, tokenizer/template, seeds, sampling, reference, optimizer, verifier, actual training count, wall time, and artifact paths. One heavy GPU experiment is allowed at a time under the repository contract.
Evidence and evaluation
R-004 first reported a byte-exact save/load result; a later save→train→load probe found output drift. Separate residual-hook and optimizer-reset bugs were subsequently addressed. The V-KTO sprint was curtailed to one salvage seed and does not satisfy its original full multi-seed/ICL/SFT protocol. These distinctions must accompany reported numbers.
Limitations and interpretation
Restoring weight bytes alone does not demonstrate identical effective inference, optimizer history, RNG state, background replay state, or every kernel path. A commit cursor is an ordering barrier, not a scientific validity certificate. Cold anchors taken under different models/templates cannot be pooled; even identical pass rates can conceal opposite per-item changes.
Sources
- cont: AGENTS.md — checkout audited
87946914c7b9. - cont: docs/research/JOURNAL.md — checkout audited
87946914c7b9. - trainfer: trainfer/snapshot.py — checkout audited
1c6391f3773b. - trainfer: trainfer/engine/train.py — checkout audited
1c6391f3773b. - agi: autoresearch/SPRINT_may24_RESULTS.md — historical revision
3842fd8875ca.