Next-token prediction (Trainfer)

From The Hei Canon

Next-token prediction (Trainfer) is raw-text language-model training without the user/assistant chat frame.

Project status: Registered as ntp. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Tokenize a raw passage with special tokens, supervise shifted next-token prediction across that passage, and average per-sample token-normalized negative log likelihood. There is no prompt masking or supplied preference label.

Implementation and controls

objectives/ntp.py::ntp_loss accepts {"text": "..."} samples and defaults to max_len=2048. Empty text raises an error; longer text is truncated. It returns the same position metadata used by the safety sidecar for SFT. The trainer can compose it with a KL anchor.

Evidence and evaluation

The source establishes a callable primitive for imprinting domain text or style. No isolated NTP retention or efficiency result was found in the inspected campaign reports. Pretraining replay injection and FLOW weighting discuss possible data/weighting combinations; those proposals should not be represented as NTP benchmark results.

Limitations and interpretation

Predicting a passage is not equivalent to answering questions about it. Truncation, formatting, data quality, and repeated exposure affect what is learned. A plain likelihood objective does not prevent catastrophic forgetting under the production optimizer.

Sources

See also