Trainfer learning methods

From The Hei Canon
(Redirected from Trainfer methods)

Trainfer learning methods is the source-audited catalogue of methods introduced, implemented, tested or concretely proposed in agi/lile, Trainfer and Cont. It covers the current objective registry, every training-mechanism branch in cont's main autoresearch runner, the recovered May–June 2026 experimental arms, the method tier menu and roadmap candidates. Trainfer research literature preserves the complete project survey accounts for further paper-only methods.

Audit date: 14 September 2026. This catalogue documents research; it does not establish that continual learning has been solved. Each method page separates mechanism, implementation/controls, measured evidence, and limitations. “Implemented” means found in the audited source, not necessarily enabled in a running daemon; “proposed” does not mean shipped. Standard methods are adaptations, not claims of project invention.

Reading the evidence

  • V-KTO means verifier-graded KTO with adaptive rollouts; its June pilot regressed severely after 51 updates. The implementation survives in agi history.
  • CCPD and CCD are different objectives: critique-based candidate ranking versus context teacher/student matching.
  • KTO is the loss; V-KTO and the binary-verifier Arm 3 are data/sampling recipes using it.
  • The early GSM8K five-example fine-tune reached 44%, but matched five-shot ICL reached 96%. The journal withdrew its SOTA claim.
  • “Razin-safe” in older source comments is not a certificate for actual AdamW/LoRA behavior. See Razin safety and Razin safety monitor.
  • Reported numbers belong to their model, suite, seed and date. Pending table cells in old templates are not completed experiments; prose hypotheses are not measured mechanisms.

Implemented objective primitives

Learning and replay recipes

Historical experimental arms

  • V-KTO — Implemented in historical agi code; the June 2026 pilot was abandoned after severe regression.
  • Rejection fine-tuning — Implemented in historical experiment_pi_old_synth.py; tested on Track B.
  • STaR-style canonical fallback — Implemented as historical --mode star; tested on Track B.

Optimization, state and evaluation

Concrete proposals and comparison methods

  • AdaSTaR-RFT — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Intruder-dimension damping — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • FLOW weighting — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • EOS preference regularization — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Active preference-query selection — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Pretraining replay injection — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Low-confidence replay weighting — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Hypernetwork LoRA generation — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Streaming DataInf — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • SEAL-style self-editing — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Difficulty-targeted rollout replay — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Hint-RFT — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Function-vector anchoring — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • SCoRe-style correction — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Rehearsal and snapshot self-distillation — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Deferred feedback batching — Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified.
  • Optimizer research candidates — Mixed proposal/deferred status; none of the alternatives below is a registered choice in the audited Trainfer optimizer selector.
  • RFT research variants — Literature/design candidates, distinct from the measured local RFT, STaR, KTO and V-KTO arms.
  • DPO and GRPO in Trainfer research — Background/comparison methods; no dpo, ipo, ppo, or grpo key in the audited objective registry.

Coverage and provenance

The objective coverage is sft, weighted_sft, ntp, kto, coh, hinge, kl_anchor, safety_monitor, unlike, ccd and conditionally ccpd_v2. SFT and weighted SFT share an article; trace infilling is a mask capability rather than a registry key. DPO/IPO/PPO/GRPO are discussed in the plan but are not keys in this registry.

Current code sources are pinned to their audited revisions. Historical articles cite agi revision 3842fd8; locally that history was reachable through origin/feat/lion8bit-optimizer. A branch name does not establish which optimizer an experiment used. A reproducible source archive, SHA-256 manifest, wikitext pages, and seed script live in ~/ht/wiki/learning_methods_2026_09_14/ and ~/ht/wiki/seed_learning_methods_2026_09_14.py.

Sources