Optimizer research candidates
Optimizer research candidates collects the optimizer alternatives discussed in the project beyond the implemented AdamW/Lion selection and per-objective isolation.
Project status: Mixed proposal/deferred status; none of the alternatives below is a registered choice in the audited Trainfer optimizer selector. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
The optimizer notes distinguish several levers:
- Schedule-free optimization: remove dependence on a hand-designed finite training schedule for a continuing stream.
- AdEMAMix8bit: add a slower momentum timescale as a candidate response to nonstationary training.
- Muon/Riemannion: matrix-update geometry alternatives requiring care around parameter shape and model family.
- Gradient surgery (PCGrad, CAGrad, GradVac, MGDA): compute and reconcile separate task gradients rather than merely rescale a summed loss.
- EMA loss normalization/DB-MTL-style log transforms: normalize objective magnitudes before mixing them, without storing separate per-task gradients.
Implementation and controls
The prioritized optimizer note and production roadmap propose wrappers and controlled 500-event stream comparisons. The AdEMAMix spike sets a wall-time regression gate of 5%. The notes defer expensive gradient surgery because it needs separate backward passes and gradient storage. Current TrainEngine supports AdamW8bit/Lion8bit selection with plain AdamW fallbacks, plus separate plain AdamW instances.
Evidence and evaluation
These are research candidates and structural arguments in the cited notes. Their external benchmarks are not local efficacy evidence. No completed, comparable multi-arm result for all of these optimizers was found in the inspected reports.
Limitations and interpretation
Do not equate loss normalization with optimizer-history isolation or gradient-direction conflict resolution. A streaming daemon lacks a natural end-of-training checkpoint; methods with train/eval parameter averaging need an explicit serving policy. Model family, tensor shapes, memory and latency can rule out an otherwise plausible optimizer.
Sources
- cont: docs/research/optimizer-sample-efficiency.md — checkout audited
87946914c7b9. - cont: docs/research/sample-efficiency-lit-review.md — checkout audited
87946914c7b9. - cont: docs/research/production-implementation-roadmap.md — checkout audited
87946914c7b9. - trainfer: trainfer/engine/train.py — checkout audited
1c6391f3773b.