Greedy memorization
Greedy memorization is Trainfer's iterative context-to-weights recipe for one prompt/response pair.
Project status: Implemented in memorize.py and used by the autoresearch runner. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Measure the fraction of supervised target response tokens that are argmax under teacher forcing. Submit weighted SFT on the pair, wait for the returned commit token, remeasure, and repeat. Stop on target fraction, no-improvement patience, or maximum steps. The baseline can short-circuit training if the target fraction is already met.
Implementation and controls
greedy_rank_fraction evaluates under the model mode lock. iterate_memorize defaults to max_steps=30, threshold=0.95, plateau_patience=3, and weight=1. The training spec uses a singleton weighted-SFT sample; optional learning rate is passed as effective_lr. Per-step history records rank, matched/total counts, and commits. The loop itself does not establish a snapshot bracket; experiment callers must do so.
Evidence and evaluation
R-001/R-001b used 100 templated mythical facts on Qwen3-8B 4-bit. The strong arm (lr=5e-4, patience=10, cap=100) averaged 8 steps and achieved mean inserted-pair rank 0.957; pair zero finished at 0.889 versus baseline 0.556. This did not show pair-zero forgetting on that template family. Historical Track B Qwen3-0.6B runs instead showed roughly −3.7 to −3.8 percentage points. The early GSM8K five-example experiment reported cold 32%, memorize-based fine-tuning 44%, and matched five-shot ICL 96% over 50 heldout tasks. Its own journal withdrew the sample-efficiency SOTA claim.
Limitations and interpretation
A 0.95 teacher-forced token fraction is not 95% exact-answer accuracy or proof of free-running reproduction. Earlier tokens supplied by the target hide autoregressive error propagation. A singleton normalized sample weight cancels; it is not an independent learning-rate multiplier. Monitoring only one earlier templated fact is not a broad retention guarantee. The journal's “structurally safe” extrapolation exceeds its measured scope.
Sources
- trainfer: trainfer/memorize.py — checkout audited
1c6391f3773b. - trainfer: trainfer/objectives/sft.py — checkout audited
1c6391f3773b. - cont: docs/research/JOURNAL.md — checkout audited
87946914c7b9. - agi: autoresearch/TRACK_B_06B_RESULTS.md — historical revision
3842fd8875ca.