Trainfer research survey/rft-literature-notes
RFT family — what the literature implies for our next mechanism arms is an archived project literature survey, reproduced here to preserve the full project account of its candidate methods.
Evidence status: This is the source document's historical literature assessment, not a fresh verification of the cited papers and not evidence that their algorithms were implemented or evaluated in Trainfer. Status labels, priorities, future-work statements and numerical literature claims belong to the original note. The source snapshot was taken on 14 September 2026. Current project implementations and qualified local results are documented at Trainfer learning methods.
RFT family — what the literature implies for our next mechanism arms
Working notes from mei, 2026-05-18. Audience: prophet for may18+ scoping, Hei for direction-setting. Aim is to translate the five RFT-family variants prophet flagged in his lit scan into experimental implications for our lattice, not to summarize each paper exhaustively.
Starting position
The C-001 lattice (1.7B HumanEval) + Track B 0.6B lattice currently shows:
- Pure SFT (no rejection): self-distillation drift, monotone
damage with step count — (a) at 1.7B, lr=2e-3 / 1e-4 memorize at 0.6B.
- Pure RAFT (rejection but no negatives): sparse-data
starvation — Arm 1 seed=0 at 0.6B, -45pp damage from training 5 epochs on 1 qualifying prompt.
- ICL (no training): does not transfer (HumanEval) or matches
cold within noise (Track B 0.6B).
Both training failure modes share a root cause: at K=5 prompts, the rollout distribution at T=0.8 produces either too few quality positives (Arm 1) or too many low-quality positives (a). The K=5 sample budget is structurally the problem.
The 2025 RFT variants each address some aspect of this. Below: my read on what each variant's experimental design implies for OUR lattice.
The five variants, ordered by relevance to our failure modes
AdaSTaR (adaptive sampling) — highest relevance
Core idea: increase rollout budget per prompt for prompts that fail at the default N. The model gets MORE chances on hard prompts; training-data quality stays high without sparse-data starvation.
Why relevant: directly addresses Arm 1's failure. If 4 of 5 training prompts fail all rollouts at N=10, AdaSTaR would re-roll those 4 at N=30 or N=100 until at least one passes. The qualifying training set grows from 1 prompt to (likely) 4-5.
Cost: compute scales linearly with extra rollouts. At K=5 train, worst case 5×N_max rollouts per arm. For N_max=50 that's 250 rollouts per training pass — ~3x our current cost. Still tractable at 0.6B in <60 min.
Falsification we'd want: AdaSTaR-RFT vs vanilla RFT (Arm 1) on Track B 0.6B at same K=5 train. If AdaSTaR beats Arm 1 by >5pp on ≥2 of 4 seeds, the sparse-data hypothesis is confirmed and the sampling strategy is load-bearing.
Risk: hardest training prompts may still all-fail even at N=1000. AdaSTaR doesn't help when the model is simply incapable of the task. We'd need a fallback (Prefix-RFT or canonical-injection).
Hint-RFT (in-prompt hints) — high relevance
Core idea: add a hint (partial solution sketch, hypothesis, relevant fact) to the rollout prompt. INCREASES rollout-pass rate without requiring more rollouts.
Why relevant: orthogonal lever to AdaSTaR. Addresses sparse-data starvation by making each rollout MORE likely to pass, not by running more rollouts. Cost is constant in N.
Why interesting for HumanEval specifically: a HumanEval prompt already includes a docstring with usage examples. The "hint" is already there. So Hint-RFT may not add much on HumanEval but could be huge on Track B parametric where the prompt is bare arithmetic.
Falsification: Hint-RFT on Track B 0.6B vs vanilla RFT. If Hint-RFT achieves higher rollout-pass-rate AND higher post-train heldout, it's the right mechanism for short-prompt corpora. If it helps rollout-pass but not heldout, the hint is leaking the answer (corpus-design issue).
Risk: the hint may collapse the model onto a single canonical solution path, hurting generalization on heldout instances with different surface structure.
RIFT (signed-weighted) — medium relevance
Core idea: weight training examples by signed reward magnitude. Positives get +weight, negatives get -weight, magnitude reflects confidence/margin.
Why relevant: this is essentially what our unlike + kto
branches do, but explicit and more disciplined. The (b) arm we're
running is closer to RIFT than (a) is (it has both signed signals
firing). If RIFT lands a clean result in the lit, our (b) is
testing roughly the right family.
Why I rank it medium: RIFT alone doesn't address the K=5 sample budget. It changes the LOSS, not the data. Our (b) result will inform whether the loss shape is the lever; AdaSTaR-style data diversification is a more leveraged next move.
Falsification: covered by our (b) result. If (b) lands ≥ cold, the signed-weighted signal is doing real work. If not, RIFT is unlikely to help either.
Prefix-RFT (demo-anchored) — low-medium relevance, high risk
Core idea: include a canonical-solution prefix in the rollout prompt — partial demonstration. Model completes the rest.
Why low-medium: this is essentially "teacher-forcing on prefixes." It increases rollout-pass rate dramatically but risks the same demo-distraction failure mode we saw in our ICL probe (coherent-but-wrong-hybrid output). On HumanEval 1.7B, ICL with 5 full canonical solutions REDUCED pass rate from 37.5% to 32.3%. Prefix-RFT is a less-extreme version of the same intervention; may hurt for the same reason.
When it would shine: corpora where the rollout-pass rate is so low that ANY positive signal is preferable to none. Our Track B 0.6B sparse case (1 of 5 prompts passing) fits. HumanEval 1.7B does not.
Falsification: Prefix-RFT vs Hint-RFT on Track B 0.6B at same K. If Prefix-RFT wins where the model can't otherwise rollout-pass at all, it has a niche. If both lose to vanilla cold, the corpus is too hard for FT at this scale regardless.
STARS (block-rejection at inference) — out of scope for training arms
Core idea: at INFERENCE time, sample multiple rollouts and pick the best by verifier. NOT a training algorithm; a test-time-compute scaling technique.
Why out of scope: doesn't change the trained model. Equivalent to "best-of-N sampling" at deployment. Could be a parallel campaign ("does test-time scaling beat training?") but doesn't compete with the RFT-family training arms.
Worth noting: prophet's may18 design treats this as Arm 3 indirectly via KTO pass/fail. Test-time best-of-N IS the strongest inference-only baseline; if our training arms don't beat best-of-N they're not adding deployment value.
Recommendation for after may18
If may18 Arms 1, 2, 3 close with cold ≈ FT or below, the next campaign should be AdaSTaR. Reasons:
- Directly addresses the Arm 1 failure mode (sparse data).
- Cheap to wire — adaptive rollout budget is an enhancement to the
existing RFT runner, not a new objective.
- Orthogonal to (b) — if (b)'s unlike+kto is the loss-shape lever,
AdaSTaR's adaptive sampling is the data-quality lever. Both could compose.
If AdaSTaR also closes negative, the next move is **Hint-RFT on Track B parametric specifically** (HumanEval has docstring hints already; parametric needs explicit ones). If that closes negative too, the conclusion is "FT on these corpora at this scale does not work; need bigger model OR different objective family entirely (e.g. DPO at the response-pair level, or process reward models)."
Open questions for prophet
- AdaSTaR's adaptive budget — do we cap at N_max=50 or let it run
unbounded? Compute cost asymmetric across prompts.
- Hint-RFT for HumanEval — would the existing docstring count as
the hint, or do we need explicit additional hints?
- RIFT vs our (b) — is RIFT a different parameterization of the
same gradient or genuinely different? Worth quickly reading the paper before may19 design.
- Test-time scaling (STARS, best-of-N at deployment) — should
this be its own campaign, or treated as a pure baseline we benchmark against rather than measure as a training arm?
Honesty caveats
Knowledge cutoff is Jan 2026; some of these "2025 variants" are at or beyond it. My specifics on each paper's empirical findings may be imprecise. The EXPERIMENTAL IMPLICATIONS for our lattice are based on the variant's design principle (what kind of signal it adds), which is more robust than specific number claims. prophet's direct paper reading should be the source of truth for empirical specifics.
For our planning purposes the relative-ranking (AdaSTaR > Hint-RFT
- RIFT > Prefix-RFT > STARS-as-training) is what matters; the rank
order is principled even where the per-paper details are loose.
Archived source
- agi: autoresearch/RFT_LITERATURE_NOTES.md — historical revision
3842fd8875ca.