Chain of Hindsight

From The Hei Canon

Chain of Hindsight is the project's feedback-to-text SFT objective, abbreviated CoH.

Project status: Registered as coh; pure-CoH historical ablation was pre-registered. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Create an assistant target containing “Previous response: …”, “Feedback: …”, and “Revised response: …”. Train ordinary token-normalized SFT on the entire formatted body. If no good response exists, the target ends immediately after the revised-response heading; it still supervises the bad response and critique.

Implementation and controls

objectives/coh.py::_format_coh_sample accepts prompt, bad, critique, and optional good. This implementation does not train only the correction. The replay router uses CoH for nl_critique and nl_critique_with_rewrite. The RLVR recipe emits CoH only when the judge supplies the needed critique. Target metadata permits SFT-family monitoring.

Evidence and evaluation

The historical SCOPE_may20_coh_dichotomy_test.md proposes a pure-CoH test of the campaign's surface-form/preference distinction. The audited lattice still marks the arm TBD. A scoped experiment and a runnable objective are evidence of design and implementation, not a completed positive ablation.

Limitations and interpretation

Without a supplied revision, there is no newly verified correct response in the training target. Reinforcing the whole body can reinforce undesirable text in its hindsight framing. The “SFT-family” classification is about loss form, not a guarantee of semantic correction or stable retention. Distinguish CoH's text association from CCPD, which actually samples and ranks candidates.

Sources

See also