Low-confidence replay weighting
Low-confidence replay weighting is the proposal to prioritize uncertain examples for additional learning.
Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Use the policy's average response log probability as a cheap confidence proxy and raise replay priority for low-confidence events. This aims to spend updates on informative, difficult examples instead of only repeating easy ones.
Implementation and controls
Roadmap PR O proposes a twofold replay-priority multiplier and a 500-event mixed-feedback comparison. Its gate asks for gains on at least two harness tasks without a measurable queue-latency regression. This is a scheduling/weighting change, not a new objective.
Evidence and evaluation
The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.
Limitations and interpretation
Low confidence is not synonymous with correct, difficult, or informative: it can flag noise and errors. Response length normalization and cross-objective calibration matter. It pulls in a different direction from FLOW's preference for base-familiar examples; neither should be assumed superior without matched testing.
Sources
- cont: docs/research/production-implementation-roadmap.md — checkout audited
87946914c7b9.