Progressive LoRA residual consolidation
Progressive LoRA residual consolidation is Trainfer's mechanism for accumulating completed adapter updates without repeatedly requantizing the base model.
Project status: Implemented in ModelState.merge_active_into_residual. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
For each LoRA site compute delta = (alpha/rank) × B A, add it to the stored residual, then reset the active adapter. Effective inference uses base + accumulated residual + current active adapter. The stored base weights remain unchanged.
Implementation and controls
state.py accumulates deltas in float32 arithmetic and stores canonical CPU bfloat16 residual tensors, then binds them into the live forward path. It uses hooks and an Unsloth matmul-LoRA integration so fast inference sees the residual. Snapshot state contains active adapter and residual. This actual path supersedes the plan's earlier dequantize–merge–requantize sketch.
Evidence and evaluation
The source/tests cover residual application, null-merge behavior, and snapshot rebinding. Historical experiments exposed post-merge VRAM growth and snapshot-load hooks not restoring effective inference state; later fixes explicitly bind/clear residuals on load. The cont research contract retains a conservative merge/reset rule for heavy experiments.
Limitations and interpretation
Consolidation stores existing learned changes; it does not undo forgetting or create a new learning signal. Rounding, hook coverage, fused-kernel behavior, and memory consumption require validation. Adapter-disabled references may still see residuals. “Always serving” describes one shared service, not simultaneous unrestricted GPU training and inference.
Sources
- trainfer: trainfer/state.py — checkout audited
1c6391f3773b. - trainfer: trainfer/snapshot.py — checkout audited
1c6391f3773b. - trainfer: trainfer/PLAN.md — checkout audited
1c6391f3773b. - cont: AGENTS.md — checkout audited
87946914c7b9.