SEAL-style self-editing
SEAL-style self-editing is the project's proposal to let a model propose synthetic training material and, more broadly, changes to its own learning process.
Project status: Research proposal/design in the inspected project sources; no implementation or completed local efficacy run identified. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Generate candidate self-edits or training examples, screen them, apply accepted changes, and use downstream behavior to evaluate their utility. The literature-inspired outer-loop framing is richer than merely paraphrasing an existing question.
Implementation and controls
Roadmap Spike 4 proposes cont/teach/self_edit.py: after a number of feedback events, propose a synthetic example, verify it, and add it to replay. Its gate compares forgetting across self-edits with forgetting across the same number of real feedback events. The larger SEAL description includes hyperparameter edits and outer RL; those stages are not supplied by this proposed minimal module.
Evidence and evaluation
The cited project document records this candidate and its intended experiment. It does not provide a completed local result for this method. Published-paper results mentioned by that document are background, not Trainfer measurements.
Limitations and interpretation
A verifier can filter obvious failures but may miss semantic or retention damage. Self-generated training data can reinforce errors. This design must be distinguished from the already implemented self_synth paraphrase branch and from an actual trained outer-loop controller.
Sources
- cont: docs/research/production-implementation-roadmap.md — checkout audited
87946914c7b9. - cont: docs/research/surveys/neurips-2025-sample-efficient.md — checkout audited
87946914c7b9.