Trainfer research literature
Trainfer research literature indexes the complete archived project surveys accompanying Trainfer learning methods. These pages preserve the local introductions to externally proposed methods, including candidates that never became project code. They are historical research notes, not independently revalidated literature reviews.
The method articles identify the current implementation, experiment or proposal separately. No paper-only method should be inferred to be installed, benchmarked or original to the project from its presence here.
`bitsandbytes` 8-bit optimizer stability — version guidance
- Complete project survey
- Version pin recommendation [STRONG]
- Block size caveat [RELEVANT]
- Stability guardrails [STRONG]
- Fallback path [RELEVANT]
- Sources
ACL / NAACL / TACL 2025 — selective loss, personalization, continual SFT
- Complete project survey
- Selective / token-level / sequence-level loss weighting
- 1. WIT — On the Effect of Instruction Tuning Loss on Generalization [STRONG]
- 2. S3FT — Selective Self-to-Supervised Fine-Tuning [STRONG]
- 3. AlignDistil — Token-Level LM Alignment as Adaptive Policy Distillation [RELEVANT]
- Multi-objective / SFT+DPO balancing
- 4. Balancing the Budget: SFT vs PFT Trade-offs [STRONG]
- Critique-and-revise / self-correction
- 5. Confidence vs Critique [RELEVANT]
- 6. S²R — Self-verify + Self-correct via RL [BACKGROUND]
- Online / continual fine-tuning with forgetting mitigation
- 7. GORP — Continual Gradient Low-Rank Projection Fine-Tuning [STRONG]
- 8. HFT — Half Fine-Tuning [STRONG]
- 9. SEE — Sequential Ensemble of Experts [BACKGROUND]
- 10. COPR — Continual Preference Learning with Optimal Policy Regularization [STRONG]
- Personalization / few-shot user adaptation
- 11. PROPER — Progressive Learning for Personalized LLMs [RELEVANT]
- 12. CHAMELEON — Personalize Your LLM [RELEVANT]
- 13. PLUM — On the Way to LLM Personalization [STRONG]
- Debiasing / targeted correction
- 14. FairSteer — Inference-Time Debiasing via Dynamic Steering [RELEVANT]
- Synthesis
COLM 2024 / 2025 — language-modeling specialist venue
- Complete project survey
- COLM 2025
- 1. LoRI — Reducing Cross-Task Interference in Multi-Task LoRA [STRONG]
- 2. OCRM — Off-policy Corrected Reward Modeling [STRONG]
- 3. REFA — Token-Level EOS Regularization for Preference Optimization [RELEVANT]
- 4. Active Exploration as Contextual Dueling Bandit for RLHF [RELEVANT]
- 5. Don't lie to your friends — Collaborative self-play [STRONG]
- 6. Bayesian Scaling Laws for ICL [BACKGROUND]
- 7. VaPR — Vision-language Preference Alignment for Reasoning [BACKGROUND]
- 8. PersonaMem — Know Me, Respond to Me [BACKGROUND]
- COLM 2024
- 9. D2PO — Discriminator-Guided DPO with Response Evaluation Models [STRONG]
- 10. Self-Rewarding Language Models [RELEVANT]
- 11. LoraHub — Cross-Task Generalization via Dynamic LoRA Composition [BACKGROUND]
- 12. Negative Preference Optimization — From Collapse to Effective Unlearning [BACKGROUND]
- 13. Instruction Mining [BACKGROUND]
- Sources
Context Distillation (2025)
- Complete project survey
- Overview
- Literature
- Implementation in Trainfer
Influence functions for streaming LoRA adaptation
- Complete project survey
- Current methods (all batch, all offline)
- DataInf [RELEVANT]
- LESS [RELEVANT]
- GREATS [BACKGROUND]
- Metagradient Descent / REPLAY (2025) [RELEVANT]
- The open gap [STRONG]
- Sketch: a streaming DataInf probe for trainfer
- Sources
EMNLP 2025 — user feedback, personalization, RLVR for instruction following
- Complete project survey
- Directly actionable for live-training daemon
- 1. User Feedback in Human-LLM Dialogues: A Lens to Understand Users but Noisy as a Learning Signal [STRONG]
- 2. FaST — Feature-aware Sampling and Tuning for Personalized Preference Alignment [STRONG]
- 3. Drift — Decoding-time Personalized Alignments with Implicit User Preferences [RELEVANT]
- 4. PRIME — Cognitive Dual-Memory and Personalized Thought Process [STRONG]
- 5. pFedGPT — Hierarchically Optimizing LoRA Aggregation Weights [BACKGROUND]
- Sample efficiency / selective loss
- 6. LimaCost — Data Valuation for Instruction Tuning [RELEVANT]
- 7. Low-Confidence Gold — Refining Low-Confidence Samples [STRONG]
- 8. MaZO — Masked Zeroth-Order Optimization for Multi-Task FT [BACKGROUND]
- Critique / RLVR / reasoning
- 9. MultiCritique — Training LMs to Critique With Multi-agent Feedback [RELEVANT]
- 10. No Need for Explanations — LLMs Learn from Mistakes In-Context [STRONG]
- 11. VerIF — Verification Engineering for RL in Instruction Following [RELEVANT]
- Memory / retrieval
- 12. Awesome-RAG-Reasoning (resource paper) [BACKGROUND]
- Negative findings
ICLR 2025 / 2026 — test-time adaptation, continual LoRA, sparse-feedback RL
- Complete project survey
- ICLR 2026 — highest priority
- 1. In-Place Test-Time Training (Oral) [STRONG]
- 2. Reward Is Enough: LLMs Are In-Context RL Learners [RELEVANT]
- 3. Meta-UCF — Unified Task-Conditioned LoRA Generation for Continual Learning [STRONG]
- 4. Doc-to-LoRA — Instantly Internalize Contexts [STRONG]
- 5. Uni-DPO — Unified Dynamic Preference Optimization [RELEVANT]
- 6. Data Selection for Efficient Preference Alignment [RELEVANT]
- 7. ∇-Reasoner — Test-Time Gradient Descent in Latent Space [BACKGROUND]
- 8. Navigating the Cost-Performance Pareto Frontier of Test-Time LLM Agent Adaptation [STRONG]
- 9. Aligner, Diagnose Thyself — Meta-Learning for Intrinsic Feedback [RELEVANT]
- 10. Alignment through Meta-Weighted Online Sampling [RELEVANT]
- 11. In-Context Adaptation (ICA) [BACKGROUND]
- ICLR 2025 — load-bearing keepers
- 12. Spurious Forgetting in Continual Learning of Language Models [STRONG]
- 13. SD-LoRA — Scalable Decoupled LoRA for Class Incremental Learning [RELEVANT]
- 14. Unlocking Function Vectors for Catastrophic Forgetting in Continual Instruction Tuning [STRONG]
- 15. HMoRA — Hierarchical Mixture of LoRA Experts [RELEVANT]
- Skipped
- Sources
ICML 2025 — sample-efficient LLM learning
- Complete project survey
- Tier 1 — Directly load-bearing
- 1. FLOW — Upweighting Easy Samples Mitigates Forgetting [STRONG]
- 2. ConfPO — Policy Confidence for Critical Token Selection [STRONG]
- 3. EXPO — Explicit Preference Optimization [RELEVANT]
- 4. Theoretical Analysis of KL-regularized RLHF with Multiple Reference Models [STRONG]
- 5. PF-PPO — Policy Filtration for RLHF [RELEVANT]
- 6. PILAF — Policy-Interpolated Learning for Aligned Feedback [STRONG]
- Tier 2 — Worth citing / implementing
- 7. TLM — Test-Time Learning for LLMs [RELEVANT]
- 8. Surprising Effectiveness of Test-Time Training for Few-Shot Learning [BACKGROUND]
- 9. Flat-LoRA — LoRA over a Flat Loss Landscape [STRONG]
- 10. TSAM — Tilted Sharpness-Aware Minimization [BACKGROUND]
- 11. Scaling Laws for Forgetting with Pretraining Data Injection [STRONG]
- 12. From RAG to Memory + Cut-and-Replay [BACKGROUND]
- Skipped (off-scope)
- Synthesis
Meta-learning for LLM adaptation, 2024–2026
- Complete project survey
- 1. MAML / Reptile / first-order meta-learning at LLM scale
- MAML-en-LLM [BACKGROUND]
- ABMLL — Low-Rank Amortized Bayesian Meta-Learning [BACKGROUND]
- ReptiLoRA [STRONG]
- 2. ICL as implicit meta-gradient descent
- Metagradient Descent / REPLAY [STRONG]
- COLD-Steer [SKIP]
- 3–4. Learning-to-update / learned optimizers
- μLO — Compute-Efficient Meta-Generalization of Learned Optimizers [BACKGROUND]
- Negative result: no 2025–26 "Meta-SGD for LLMs" paper. Research gap.
- 5. Fast-adaptation LoRA (hypernetwork-generated)
- HyperLoRA [STRONG]
- Text-to-LoRA (T2L) [STRONG]
- Meta-LoRA [STRONG]
- HyperAdaLoRA [BACKGROUND]
- 6. Task clustering + per-cluster LoRA heads
- TC-LoRA [STRONG]
- K-Merge / K-Merge++ [STRONG]
- MoLE-CIE, D-MoLE, HMoRA, LoRA-Mixer [BACKGROUND]
- 7. Test-time training for LLMs
- TTRL — Test-Time RL [STRONG]
- TTC-RL — Learning on the Job [STRONG]
- TLM — Test-Time Learning for LLMs [BACKGROUND]
- qTTT, TTT-E2E [SKIP]
- Few-shot personalization (meta-learning flavor)
- FSPO — Few-Shot Preference Optimization [STRONG]
- Meta Reward Modeling (MRM) [STRONG]
- Fermi [BACKGROUND]
- Synthesis — four design-level implications beyond the existing lit review
NeurIPS 2025 — sample-efficient LLM learning
- Complete project survey
- Test-time training / test-time RL
- 1. SEAL — Self-Adapting Language Models [STRONG]
- 2. TTRL — Test-Time Reinforcement Learning [STRONG]
- Self-improvement / self-play
- 3. Absolute Zero — Reinforced Self-play Reasoning with Zero Data [RELEVANT]
- 4. ExIt — Bootstrapping Task Spaces for Self-Improvement [RELEVANT]
- 5. CoVo — Consistent Paths Lead to Truth [STRONG]
- Continual learning / forgetting
- 6. Nested Learning — Hope architecture [BACKGROUND]
- 7. SuRe — Surprise-Driven Prioritised Replay for Continual LLM Learning [STRONG]
- 8. GainLoRA — Gated Integration of LoRA for Continual Learning [RELEVANT]
- 9. LoRA vs Full Fine-tuning: An Illusion of Equivalence [STRONG]
- Preference learning
- 10. RePO — Preference Learning through ReLU Optimization [RELEVANT]
- Off-policy / data-efficient RL
- 11. Difficulty-Targeted Online Data Selection + Rollout Replay [STRONG]
- 12. Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization [RELEVANT]
- Synthesis
- Sources
Optimizer landscape, 2025–2026 — LoRA + online finetuning
- Complete project survey
- Key 2025 results
- AdEMAMix — Better, Faster, Older [STRONG]
- Benchmarking Optimizers for LLM Pretraining — small-batch caveats [STRONG]
- AdamW8bit + param_groups — the correct idiom [STRONG]
- Lion8bit at LoRA scale [RELEVANT]
- Muon / Riemannion call for Qwen3 [RELEVANT]
- ScheduleFree-AdamW [RELEVANT]
- Open citations (added to optimizer-sample-efficiency.md References)
2026 Synthesis: Test-Time RL, Self-Improvement, o1/R1 Lineage, PRM vs ORM
- Complete project survey
- 1. Test-Time RL / RL at inference
- 2. Self-improvement loops (generate-then-train)
- 3. o1 / DeepSeek-R1 lineage — sample-efficient variants
- 4. PRMs vs ORMs — sample-efficiency empirics
- TL;DR for the daemon
Sample-Efficiency Literature Review for `trainfer`
- Complete project survey
- 1. Replay buffer design for continual LoRA
- 2. Catastrophic forgetting in LoRA adapters
- 3. Per-sample loss weighting for SFT-from-feedback
- 4. Objective mixing (SFT + KTO + CoH + hinge + contrastive at heterogeneous scales)
- 5. Verifiable-reward online learning
- 6. Learning from natural-language critique (semantic feedback)
- 7. Sample efficiency of QLoRA / LoRA continual updates
- Cross-cutting honest call-outs
- Citations I could not confirm and have NOT included
RFT family — what the literature implies for our next mechanism arms
- Complete project survey
- Starting position
- The five variants, ordered by relevance to our failure modes
- AdaSTaR (adaptive sampling) — highest relevance
- Hint-RFT (in-prompt hints) — high relevance
- RIFT (signed-weighted) — medium relevance
- Prefix-RFT (demo-anchored) — low-medium relevance, high risk
- STARS (block-rejection at inference) — out of scope for training arms
- Recommendation for after may18
- Open questions for prophet
- Honesty caveats
Literature scan — may17 wrap, may18 prep (2026-05-17)
- Complete project survey
- Thread 1 — fact replacement (Track A asymmetry)
- Thread 2 — ICL vs fine-tune regime
- Thread 3 — LoRA rank and forgetting
- Thread 4 — measurement design (McNemar power on N=64)
- Recommended may18 campaign shape
- What did NOT survive the scan
- Sources