Idle feedback replay

From The Hei Canon

Idle feedback replay reuses logged user feedback as low-priority training when Trainfer's compute queue is idle.

Project status: Implemented scheduler, controlled by idle_replay (default false). This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Choose an eligible historical feedback record using recency weight 2^(−age_hours / half_life_hours), reconstruct its training spec, and enqueue a normal update. Caps prevent one record from being replayed indefinitely during the same process.

Implementation and controls

engine/replay.py::IdleReplayScheduler uses trajectory byte offsets as stable record IDs. Defaults: idle threshold 30 seconds, poll every two seconds, minimum three feedback records, half-life 24 hours, maximum three replays per record. The in-memory cap map resets on restart. Controller.feedback_to_batch reconstructs binary→KTO, rewrite/preferred→weighted SFT, critique→CoH, critique-with-rewrite→CoH. These replay routes are concrete and need not match every rich live-feedback mode.

Evidence and evaluation

Replay tests check eligibility, caps, and routing. The May 16 research journal records a real methodological problem: idle replay changed cold probes between experiments; subsequent arms disabled it and used a fresh daemon. This is evidence that background training matters to experiment state, not evidence of a measured retention benefit.

Limitations and interpretation

Recency bias does not become uniform merely because every record gets older: pairwise weight ratios stay constant under a common time shift. Replay reproduces feedback quality problems and can overfit a small corpus across restarts. It is not the richer contradiction-aware/self-distillation replay envisioned in the plan.

Sources

See also