Progressive LoRA residual consolidation

From The Hei Canon

Progressive LoRA residual consolidation is Trainfer's mechanism for accumulating completed adapter updates without repeatedly requantizing the base model.

Project status: Implemented in ModelState.merge_active_into_residual. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

For each LoRA site compute delta = (alpha/rank) × B A, add it to the stored residual, then reset the active adapter. Effective inference uses base + accumulated residual + current active adapter. The stored base weights remain unchanged.

Implementation and controls

state.py accumulates deltas in float32 arithmetic and stores canonical CPU bfloat16 residual tensors, then binds them into the live forward path. It uses hooks and an Unsloth matmul-LoRA integration so fast inference sees the residual. Snapshot state contains active adapter and residual. This actual path supersedes the plan's earlier dequantize–merge–requantize sketch.

Evidence and evaluation

The source/tests cover residual application, null-merge behavior, and snapshot rebinding. Historical experiments exposed post-merge VRAM growth and snapshot-load hooks not restoring effective inference state; later fixes explicitly bind/clear residuals on load. The cont research contract retains a conservative merge/reset rule for heavy experiments.

Limitations and interpretation

Consolidation stores existing learned changes; it does not undo forgetting or create a new learning signal. Rounding, hook coverage, fused-kernel behavior, and memory consumption require validation. Adapter-disabled references may still see residuals. “Always serving” describes one shared service, not simultaneous unrestricted GPU training and inference.

Sources

See also