KL anchor
KL anchor is the reference-distribution regularizer composable with Trainfer training objectives.
Project status: Registered as kl_anchor, including prompt, full-sequence, and target-position scopes. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Penalize KL(live policy || reference policy). The scope selects where distributions are compared:
prompt: prompt positions only (default).full_sequence: prompt plus the first available response-like field.target_position: the last real token position of each sample; exclude specified vocabulary IDs so corrected tokens can move.
A coefficient in batch_objectives scales the sidecar contribution.
Implementation and controls
objectives/kl.py resolves prompt or prefix, and response fields response, good, chosen, better_response. Target-position exclusions combine caller exclusions with per-sample surgery-token information. Reference selection supports an external model or adapter-disabled forward. Vocabulary KL is chunked to reduce transient memory.
This is a separate primitive from CCPD's internal prompt-only KL. Supplying a sidecar to one recipe does not automatically change CCPD's internal scope.
Evidence and evaluation
The research synthesis identified prompt-only anchoring as a response-drift blind spot and proposed full-sequence scope. The current code now exposes that scope, so the old roadmap's “in-flight” status is historical. V-KTO supplied target-position anchoring but still showed severe regression; a nonzero KL coefficient alone did not establish preservation.
Limitations and interpretation
The anchor constrains only represented positions and its chosen reference. Full-sequence concatenation is not automatically identical to the SFT chat-template/mask geometry. Excluding correction tokens deliberately removes constraints there. Adapter-disabled reference can move after residual merges. KL regularization is a soft penalty, not a hard retention or safety certificate.
Sources
- trainfer: trainfer/objectives/kl.py — checkout audited
1c6391f3773b. - trainfer: trainfer/engine/train.py — checkout audited
1c6391f3773b. - cont: docs/research/sample-efficiency-synthesis.md — checkout audited
87946914c7b9. - agi: autoresearch/experiment_humaneval.py — historical revision
3842fd8875ca.