KTO

From The Hei Canon

KTO is the binary-preference objective used by Trainfer. The project implements its own per-step loss rather than calling a multi-epoch KTOTrainer.

Project status: Registered as kto; also used by historical verifier-driven experiments. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

For each response, compute its mean log probability per supervised response token under the live policy and a reference. Let r = β(log π − log π_ref), using those per-token means. With a detached drift baseline z:

  • Desirable: λ_D [1 − sigmoid(r − z)].
  • Undesirable: λ_U [1 − sigmoid(z − r)].

Average the sample losses. Unlike V-KTO, ordinary KTO does not specify how responses are generated or how labels are obtained.

Implementation and controls

objectives/kto.py::kto_loss accepts samples with prompt, response, and label. Defaults: β=0.1, λ_D=1.0, λ_U=1.5. An explicit frozen model can supply the reference; otherwise the default is a no-gradient forward under model.disable_adapter(). The fallback when no reference is available uses zero reference log probabilities. For batches of at least two samples, z is the detached mean of the same batch's r values. For singletons, an EMA keyed by model identity and β is used after initialization; default EMA coefficient is 0.9. Optimizer reset clears this EMA. Although the docstring describes mismatched pairs, the inspected code does not shuffle prompts/responses to construct mismatched pairs. Binary feedback routes up to desirable and other binary values to undesirable in Controller.feedback_to_batch.

Evidence and evaluation

The historical Track B Qwen3-0.6B Arm 3 used every passing rollout as desirable and every failing rollout as undesirable, skipping all-pass/all-fail groups. The lattice reports −2.5 percentage points relative to cold, classified as preservation within its campaign tolerance. This is separate from the planned HumanEval pure-KTO control and the damaging V-KTO salvage run. Neither demonstrates a universal KTO improvement.

Limitations and interpretation

Adapter-disabled reference means base plus any separately applied residual, not necessarily the original session checkpoint. Length normalization is a project choice. The code's singleton docstring calls a loss value of 0.5 a zero-signal update; because z is detached, a numerical value of 0.5 does not imply zero gradient. EMA behavior should be described as implemented, not justified by that claim. KTO has no general pointwise preservation guarantee; labels, imbalance, reference drift, and repeated updates all matter.

Sources

See also