Trainfer research survey/lit-review-may17

From The Hei Canon

Literature scan — may17 wrap, may18 prep (2026-05-17) is an archived project literature survey, reproduced here to preserve the full project account of its candidate methods.

Evidence status: This is the source document's historical literature assessment, not a fresh verification of the cited papers and not evidence that their algorithms were implemented or evaluated in Trainfer. Status labels, priorities, future-work statements and numerical literature claims belong to the original note. The source snapshot was taken on 14 September 2026. Current project implementations and qualified local results are documented at Trainfer learning methods.

Literature scan — may17 wrap, may18 prep (2026-05-17)

Scan covered three threads bearing directly on may17 findings:

  1. Track A failure — K=5 memorize at LoRA r=16 / lr=2e-3 cannot replace
  strong factual priors (Suva for Tarawa, Delacroix for Géricault). What's
  the *established* way to do fact replacement?
  1. Track B "ICL wins" — 5-shot ICL beat fine-tune by +7.5pp mean across
  4 seeds. Is this universal, or is K=5 ICL-vs-FT regime-dependent?
  1. LoRA rank + KL anchoring — REPORT_may17 next-tag direction speculated
  r=64 + KL anchor might unblock Track A. Is that grounded?

Thread 1 — fact replacement (Track A asymmetry)

Established method: ROME / MEMIT (rank-one MLP edits, not SGD fine-tune).

  • ROME treats an MLP layer as a key-value store, applies a rank-one weight
 update to insert a new (subject → object) association. Reference hyperparameters
 for GPT-J / Llama3-8B: lr=0.5, max_steps=35–50, layer 6, KL factor 0.0625
 (Meng et al. 2022, MAKE / TACL 2025).
  • We were doing the *opposite* regime: lr=2e-3 over 100 steps via LoRA r=16 SGD
 on Q→A pairs. That's a tiny-step, distributed-update path; ROME is a
 one-shot, surgical update. Different tool entirely.
 "model edit success is essentially unrelated to where factual information
 is stored" — Causal Tracing's correlation with edit success is near zero.
 So we don't need a perfect causal trace to make ROME work; the rank-one
 surgical edit at *any* mid-layer MLP can succeed.
  • Failure mode of ROME we should expect: "ROME conditions the model to
 process information in the COUNTERFACT prompt format. ... particularly
 vulnerable to format mismatches during downstream fine-tuning"
 (Retention of Edited Knowledge, 2025).
 This aligns with Mei's may17 Thread #2 template-echo finding — surface
 form matters.
  • 2025 improvements (AlphaEdit, NAMET, UltraEdit, LKS, MAKE) target *sequential*
 multi-thousand edits, where ROME's catastrophic forgetting kicks in after
 ~10 edits. For K=5 single-pass we don't need them.

Track A v2 mechanism candidate: ROME-style rank-one edit, applied per factoid, mid-layer MLP, ~35–50 steps at lr=0.5. May16's experiment.py has an entity_masked_sft stub; ROME is a *different* mechanism that may live alongside it.

Thread 2 — ICL vs fine-tune regime

T-Few precedent: parameter-efficient fine-tune CAN beat ICL at few-shot.

 parameter-efficient fine-tuning yields 6% higher accuracy and 1000× less
 FLOPs than in-context learning on few-shot tasks." This is the
 counter-example to our may17 ICL-beats-FT result.
 K ∈ {1, 4, 8, ..., 1024} on GLUE / HotpotQA / Multi-News with FT vs ICL
 vs LoRA. "In-context learning, while efficient in terms of parameters,
 underperforms across all tasks; LoRA offers a compelling compromise,
 achieving competitive performance with theoretically lower resource
 demands." Reverses our Track B finding qualitatively.
 hybrid — fine-tune the model on task-specific data *augmented with
 in-context examples, mimicking the structure of k-shot prompts*. The
 hybrid combines ICL's sample efficiency with FT's persistence. **This is
 the natural next move for Track B v2: train on prompts that already
 contain the K-shot demos, then eval without demos.**

Why our may17 Track B result might be corpus-specific:

  • Cited regime ICL excels in: parametric arithmetic on small models with
 pretrained competence. Cold rate already ~58%; ICL has lots to extract.
  • Cited regime FT excels in: classification / QA / multi-step where ICL
 cannot fit the chain in context (the Bornschein paper's claim).
  • HumanEval (C-001) is closer to the FT-wins regime — multi-step code
 generation where 5 demos in prompt may not fit the diversity. *This is
 why C-001 matters: it tests whether the "ICL wins" pattern is corpus-bound.*

Thread 3 — LoRA rank and forgetting

Strong empirical anchor for our REPORT_may17 hypothesis:

 LoRA-rank-tradeoffs 2025): "LoRA learns less and forgets less. ... SFT
 often possesses substantially higher intrinsic rank and thus greater
 capacity to both specialize and overwrite pretrained knowledge."
 Translation: higher rank → more capacity to overwrite priors → better
 shot at Track A fact replacement.
 "Higher LoRA rank leads to increased forgetting during instruction tuning.
 ... Higher rank correlates with forgetting at low angles (similar tasks)
 but not at high angles." Predicts: r=64 will help fact-replacement but
 hurt probe — *exactly the trade-off REPORT_may17 named*. So our hypothesis
 is grounded but the trade-off is real, not just our implementation.
 [LoRA-Null / Tang et al. 2025], [MiLoRA, CLoRA]: 2025 methods that
 *initialize the LoRA adapter in directions orthogonal to pretrained
 knowledge* (null space of pretrained activations). Decouples
 fact-acquisition (in the adapter subspace) from forgetting (which
 requires modifying pretrained-relevant directions). **This is the
 publishable lever for "memorize without breaking probe."**

Track A v2 mechanism rank candidate: instead of "just bump to r=64", try OPLoRA / LoRA-Null at r=16. Same capacity, but adapter forced into non-interfering directions. Lower forgetting risk.

Thread 4 — measurement design (McNemar power on N=64)

Critical finding that mei's C-001 SCOPE needs to absorb:

Lachin 1992; real-statistics.com / McNemar power:

  • Detecting 75% → 85% paired with McNemar at 80% power needs **~113
 patients**. N=64 is underpowered for a +10pp effect at typical
 discordance.
  • "The power of McNemar's test depends on the discordant cells (b and c),
 not the total proportions." If most heldout items are concordant
 (cold-right & post-right, or both-wrong), the effective sample size
 shrinks below 64. With realistic 30% discordance, effective N ≈ 19.
  • Recommendations for N=64:
  • Use the exact (binomial) or mid-p McNemar, not the asymptotic
   chi-square. Continuity-corrected is "considerably less powerful."
  • One-sided if directional ("post is better than cold") — gains power.
  • Pre-register expected discordance to size the effect we can actually
   detect.

Implication for C-001 SCOPE addendum:

  • Currently SCOPE targets "post ≥ cold + 10pp at p<0.05" on N=64 paired.
 At ~30% discordance this is at or below the detectable boundary.
  • Either widen N (use full 164 split, do not subset to heldout-only —
 acceptable since we evaluate cold-vs-post under the same recipe), or
 raise the effect target to the +15pp detectable band, or run more
 seeds and aggregate paired counts across seeds (Bonferroni-corrected).

Recommended may18 campaign shape

Ordered by leverage on the SOTA bar:

  1. Track A v2 — ROME mechanism arm (highest novelty, directly tests
  the K=5 fact-replacement asymmetry). Apply per-factoid rank-one MLP
  edit at the lr=0.5 / steps=35 / KL=0.0625 regime, with entity_masked_sft
  as a complementary arm. Pre-condition: lile needs a ROME engine —
  nontrivial daemon-side work.
  1. OPLoRA / LoRA-Null at r=16 instead of "bump to r=64." Cheaper than
  ROME (same lile machinery, just adapter init change); addresses the
  "memorize without breaking probe" failure mode directly. Lower-risk
  intermediate step.
  1. Track B v2 — Fine-Tuned-In-Context-Learner hybrid. Train on prompts
  that *already contain* the K-shot demos. Eval without demos. Tests
  whether the persisting-ICL-in-parameters claim survives. Doesn't
  require new lile mechanism; just a prompt-format change in the
  training data builder.
  1. C-001 statistical preregistration. Before training arm: compute
  power for the cold pass rate × expected discordance, pick the right
  McNemar variant (mid-p one-sided), and either bump N (use full 164)
  or adjust the effect target. Without this, even a recipe that works
  won't reach p<0.05.

What did NOT survive the scan

  • "More steps / lower lr" as Track A rescue. Already falsified by run
 02 (lr/20 × 2× steps gave same forward_paraphrased and worse probe
 degradation). Literature confirms: SGD-based LoRA at any lr is the
 wrong tool for surgical fact replacement; ROME is.
  • "Single-seed N=50 cold-eval" as a baseline. May17 multi-seed lesson
 + McNemar power finding → multi-seed AND wider-N are both load-bearing.

Sources

Archived source