Per-objective optimizer isolation

From The Hei Canon

Per-objective optimizer isolation separates adaptive optimizer histories for different Trainfer objectives while keeping the same model parameters.

Project status: Implemented optional mode; default off. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Maintain one optimizer instance per objective name so momentum and second-moment estimates from one loss do not automatically become another loss's history. Separate learning-rate groups alone would not provide independent histories for the same parameters.

Implementation and controls

TrainEngine lazily stores optimizer instances; cfg.per_objective_optim=True enables separate plain torch AdamW optimizers and per_objective_lr supplies rates. Shared mode defaults to AdamW8bit when available. Separate plain optimizers avoid the bitsandbytes global-manager complication identified in the design. Important boundary: _step_multi sums multiple primary losses and uses a shared optimizer slot for the combined step. The flag does not split a multi-objective batch into separate gradient updates. Reset drops every cached optimizer and KTO EMA.

Evidence and evaluation

The research note hypothesizes that differences in CoH/KTO/SFT gradient scale can make mixed feedback unstable and prioritizes isolation. Tests establish optimizer identity and reset behavior. A controlled stream-level semantic-feedback benefit remains distinct from those structural checks.

Limitations and interpretation

Extra optimizer states cost memory. Separate histories do not separate parameters: objectives still overwrite shared capabilities. The older note's numerical example claiming a smaller KTO gradient divided by a large SFT second moment yields an LR boost has its direction wrong; that arithmetic describes suppression relative to a matched smaller denominator. Treat scale-mixing as a testable concern rather than inheriting that calculation.

Sources

See also