EOS preference regularization

From The Hei Canon

EOS preference regularization is the project's proposed response-length shortcut control for preference objectives.

Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Add a penalty on differences in end-of-sequence or end-of-turn log probability between preferred and rejected sequences, using a coefficient lambda_eos. The intended effect is to stop the objective from exploiting termination behavior instead of substantive correctness.

Implementation and controls

Roadmap PR K names KTO, hinge, and optionally CoH; proposes default lambda_eos=0 and an experimental value 0.1. It requires identifying each tokenizer's effective end token and evaluating response-length variance alongside task performance. The proposed squared log-probability-difference term is separate from the length normalization already present in KTO/hinge.

Evidence and evaluation

The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.

Limitations and interpretation

Length variance reduction alone is not evidence of better answers. Matching EOS probabilities at mismatched positions may be ill-defined, and different chat templates use different stopping rules. The source proposal says “eliminates” the shortcut, but no local evidence establishes that strong claim; the current registry contains no separate EOS regularizer.

Sources

See also