Trainfer research survey/lit-review-may17
Literature scan — may17 wrap, may18 prep (2026-05-17) is an archived project literature survey, reproduced here to preserve the full project account of its candidate methods.
Evidence status: This is the source document's historical literature assessment, not a fresh verification of the cited papers and not evidence that their algorithms were implemented or evaluated in Trainfer. Status labels, priorities, future-work statements and numerical literature claims belong to the original note. The source snapshot was taken on 14 September 2026. Current project implementations and qualified local results are documented at Trainfer learning methods.
Literature scan — may17 wrap, may18 prep (2026-05-17)
Scan covered three threads bearing directly on may17 findings:
- Track A failure — K=5 memorize at LoRA r=16 / lr=2e-3 cannot replace
strong factual priors (Suva for Tarawa, Delacroix for Géricault). What's the *established* way to do fact replacement?
- Track B "ICL wins" — 5-shot ICL beat fine-tune by +7.5pp mean across
4 seeds. Is this universal, or is K=5 ICL-vs-FT regime-dependent?
- LoRA rank + KL anchoring — REPORT_may17 next-tag direction speculated
r=64 + KL anchor might unblock Track A. Is that grounded?
Thread 1 — fact replacement (Track A asymmetry)
Established method: ROME / MEMIT (rank-one MLP edits, not SGD fine-tune).
- ROME treats an MLP layer as a key-value store, applies a rank-one weight
update to insert a new (subject → object) association. Reference hyperparameters for GPT-J / Llama3-8B: lr=0.5, max_steps=35–50, layer 6, KL factor 0.0625 (Meng et al. 2022, MAKE / TACL 2025).
- We were doing the *opposite* regime: lr=2e-3 over 100 steps via LoRA r=16 SGD
on Q→A pairs. That's a tiny-step, distributed-update path; ROME is a one-shot, surgical update. Different tool entirely.
- Critical finding (Hase et al.):
"model edit success is essentially unrelated to where factual information is stored" — Causal Tracing's correlation with edit success is near zero. So we don't need a perfect causal trace to make ROME work; the rank-one surgical edit at *any* mid-layer MLP can succeed.
- Failure mode of ROME we should expect: "ROME conditions the model to
process information in the COUNTERFACT prompt format. ... particularly vulnerable to format mismatches during downstream fine-tuning" (Retention of Edited Knowledge, 2025). This aligns with Mei's may17 Thread #2 template-echo finding — surface form matters.
- 2025 improvements (AlphaEdit, NAMET, UltraEdit, LKS, MAKE) target *sequential*
multi-thousand edits, where ROME's catastrophic forgetting kicks in after ~10 edits. For K=5 single-pass we don't need them.
Track A v2 mechanism candidate: ROME-style rank-one edit, applied per
factoid, mid-layer MLP, ~35–50 steps at lr=0.5. May16's experiment.py has
an entity_masked_sft stub; ROME is a *different* mechanism that may live
alongside it.
Thread 2 — ICL vs fine-tune regime
T-Few precedent: parameter-efficient fine-tune CAN beat ICL at few-shot.
parameter-efficient fine-tuning yields 6% higher accuracy and 1000× less FLOPs than in-context learning on few-shot tasks." This is the counter-example to our may17 ICL-beats-FT result.
- PEARC 2025: tested
K ∈ {1, 4, 8, ..., 1024} on GLUE / HotpotQA / Multi-News with FT vs ICL
vs LoRA. "In-context learning, while efficient in terms of parameters,
underperforms across all tasks; LoRA offers a compelling compromise,
achieving competitive performance with theoretically lower resource
demands." Reverses our Track B finding qualitatively.
hybrid — fine-tune the model on task-specific data *augmented with in-context examples, mimicking the structure of k-shot prompts*. The hybrid combines ICL's sample efficiency with FT's persistence. **This is the natural next move for Track B v2: train on prompts that already contain the K-shot demos, then eval without demos.**
Why our may17 Track B result might be corpus-specific:
- Cited regime ICL excels in: parametric arithmetic on small models with
pretrained competence. Cold rate already ~58%; ICL has lots to extract.
- Cited regime FT excels in: classification / QA / multi-step where ICL
cannot fit the chain in context (the Bornschein paper's claim).
- HumanEval (C-001) is closer to the FT-wins regime — multi-step code
generation where 5 demos in prompt may not fit the diversity. *This is why C-001 matters: it tests whether the "ICL wins" pattern is corpus-bound.*
Thread 3 — LoRA rank and forgetting
Strong empirical anchor for our REPORT_may17 hypothesis:
- Biderman et al. 2024 (cited in
LoRA-rank-tradeoffs 2025): "LoRA learns less and forgets less. ... SFT often possesses substantially higher intrinsic rank and thus greater capacity to both specialize and overwrite pretrained knowledge." Translation: higher rank → more capacity to overwrite priors → better shot at Track A fact replacement.
"Higher LoRA rank leads to increased forgetting during instruction tuning. ... Higher rank correlates with forgetting at low angles (similar tasks) but not at high angles." Predicts: r=64 will help fact-replacement but hurt probe — *exactly the trade-off REPORT_may17 named*. So our hypothesis is grounded but the trade-off is real, not just our implementation.
[LoRA-Null / Tang et al. 2025], [MiLoRA, CLoRA]: 2025 methods that *initialize the LoRA adapter in directions orthogonal to pretrained knowledge* (null space of pretrained activations). Decouples fact-acquisition (in the adapter subspace) from forgetting (which requires modifying pretrained-relevant directions). **This is the publishable lever for "memorize without breaking probe."**
Track A v2 mechanism rank candidate: instead of "just bump to r=64", try OPLoRA / LoRA-Null at r=16. Same capacity, but adapter forced into non-interfering directions. Lower forgetting risk.
Thread 4 — measurement design (McNemar power on N=64)
Critical finding that mei's C-001 SCOPE needs to absorb:
Lachin 1992; real-statistics.com / McNemar power:
- Detecting 75% → 85% paired with McNemar at 80% power needs **~113
patients**. N=64 is underpowered for a +10pp effect at typical discordance.
- "The power of McNemar's test depends on the discordant cells (b and c),
not the total proportions." If most heldout items are concordant (cold-right & post-right, or both-wrong), the effective sample size shrinks below 64. With realistic 30% discordance, effective N ≈ 19.
- Recommendations for N=64:
- Use the exact (binomial) or mid-p McNemar, not the asymptotic
chi-square. Continuity-corrected is "considerably less powerful."
- One-sided if directional ("post is better than cold") — gains power.
- Pre-register expected discordance to size the effect we can actually
detect.
Implication for C-001 SCOPE addendum:
- Currently SCOPE targets "post ≥ cold + 10pp at p<0.05" on N=64 paired.
At ~30% discordance this is at or below the detectable boundary.
- Either widen N (use full 164 split, do not subset to heldout-only —
acceptable since we evaluate cold-vs-post under the same recipe), or raise the effect target to the +15pp detectable band, or run more seeds and aggregate paired counts across seeds (Bonferroni-corrected).
Recommended may18 campaign shape
Ordered by leverage on the SOTA bar:
- Track A v2 — ROME mechanism arm (highest novelty, directly tests
the K=5 fact-replacement asymmetry). Apply per-factoid rank-one MLP
edit at the lr=0.5 / steps=35 / KL=0.0625 regime, with entity_masked_sft
as a complementary arm. Pre-condition: lile needs a ROME engine —
nontrivial daemon-side work.
- OPLoRA / LoRA-Null at r=16 instead of "bump to r=64." Cheaper than
ROME (same lile machinery, just adapter init change); addresses the "memorize without breaking probe" failure mode directly. Lower-risk intermediate step.
- Track B v2 — Fine-Tuned-In-Context-Learner hybrid. Train on prompts
that *already contain* the K-shot demos. Eval without demos. Tests whether the persisting-ICL-in-parameters claim survives. Doesn't require new lile mechanism; just a prompt-format change in the training data builder.
- C-001 statistical preregistration. Before training arm: compute
power for the cold pass rate × expected discordance, pick the right McNemar variant (mid-p one-sided), and either bump N (use full 164) or adjust the effect target. Without this, even a recipe that works won't reach p<0.05.
What did NOT survive the scan
- "More steps / lower lr" as Track A rescue. Already falsified by run
02 (lr/20 × 2× steps gave same forward_paraphrased and worse probe degradation). Literature confirms: SGD-based LoRA at any lr is the wrong tool for surgical fact replacement; ROME is.
- "Single-seed N=50 cold-eval" as a baseline. May17 multi-seed lesson
+ McNemar power finding → multi-seed AND wider-N are both load-bearing.
Sources
- ROME / Meng et al. 2022 — https://rome.baulab.info/
- MEMIT — https://memit.baulab.info/
- MAKE (TACL 2025) — https://direct.mit.edu/tacl/article/doi/10.1162/TACL.a.26/132652/MAKE-Memory-Associated-Knowledge-Editing
- Retention of Edited Knowledge — https://arxiv.org/pdf/2507.14198
- Does Localization Inform Editing? (Hase et al.) — https://asmadotgh.github.io/assets/pdf/13353_does_localization_inform_editi.pdf
- T-Few / IA³ — https://arxiv.org/pdf/2205.05638
- PEARC 2025 FT-vs-ICL-vs-LoRA — https://dl.acm.org/doi/full/10.1145/3708035.3736091
- Fine-Tuned In-Context Learners — https://arxiv.org/abs/2512.19879
- LoRA Rank Trade-offs — https://arxiv.org/html/2512.15634v1
- Subspace Geometry Governs Catastrophic Forgetting — https://arxiv.org/html/2603.02224
- OPLoRA — https://arxiv.org/html/2510.13003v2
- Qwen3 Technical Report — https://arxiv.org/pdf/2505.09388
- Lachin 1992, McNemar sample size — https://onlinelibrary.wiley.com/doi/abs/10.1002/sim.4780110909
- McNemar mid-p PMC — https://pmc.ncbi.nlm.nih.gov/articles/PMC3716987/
Archived source
- agi: autoresearch/LIT_REVIEW_may17.md — historical revision
3842fd8875ca.