Autoresearch recipe optimization
Autoresearch recipe optimization is cont's constrained loop for changing a training recipe and scoring its behavior against the shared daemon.
Project status: Implemented workflow; individual configured branches have different implementation/completion states. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Read config, evaluate a cold preservation probe, save a baseline, run the selected training pulse, evaluate heldout tasks and the preservation probe, and restore the baseline in a finally block. The agent keeps or discards a config change according to the run's scalar score.
Implementation and controls
autoresearch/program.md defines the workflow and config.json is the intended lever. The actual runner emits score = heldout_rate − lambda × max(0, probe_cold − probe_post), with component rates also logged. This is more precise than the charter shorthand “single heldout_pass_rate metric.” Implemented modes include memorize, hybrid pre-sampling/unlike, self-synthesis, entity-masked SFT, and stacked phases. The rlvr option in this runner returns a not-wired note; the standalone RLVR scheduler is separate.
Evidence and evaluation
C-002 and the May 16 journal document repeated trials, background replay contamination, CoT evaluation effects, and the eventual matched ICL comparison. A 0.70 small-suite score was a local recipe result, not published SOTA. Changing evaluation reasoning mode materially affected scores independently of changing the update recipe.
Limitations and interpretation
A small fixed heldout set can be overfit by repeated agent selection. A scalar hides which task or prior skill regressed; retain the components and per-task records. Snapshot restore is necessary but requires independent verification, including optimizer resets and asynchronous commit barriers. New daemon functionality belongs upstream in Trainfer rather than being smuggled into a recipe-only comparison.
Sources
- cont: autoresearch/program.md — checkout audited
87946914c7b9. - cont: autoresearch/experiment.py — checkout audited
87946914c7b9. - cont: autoresearch/config.json — checkout audited
87946914c7b9. - cont: docs/research/JOURNAL.md — checkout audited
87946914c7b9. - cont: AGENTS.md — checkout audited
87946914c7b9.