Active preference-query selection
Active preference-query selection is the PILAF-inspired proposal to choose which prompts should receive scarce human feedback.
Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Rank recent prompts by a policy/reference disagreement proxy and ask for feedback on the most informative candidates. This changes the distribution of labelled training examples while reusing existing downstream objectives.
Implementation and controls
Roadmap PR M proposes a feedback-solicitation endpoint and a Studio suggestion view. The test would compare learning after 50 actively chosen events against 75 uniformly chosen events, with a kill point at 150 if no efficiency benefit appears. A frozen reference supplies the comparison policy; the exact disagreement statistic needs implementation definition.
Evidence and evaluation
The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.
Limitations and interpretation
Disagreement can mean ambiguity or broken outputs rather than useful learning potential. Selection must account for user burden, annotation quality, and selection bias. The named endpoint is a design target, not an endpoint found in the audited implementation.
Sources
- cont: docs/research/production-implementation-roadmap.md — checkout audited
87946914c7b9. - cont: docs/research/surveys/icml-2025.md — checkout audited
87946914c7b9.