Surgical unlikelihood

From The Hei Canon

Surgical unlikelihood is a token-position correction primitive that reduces the probability of a specified bad next token, optionally teaching a good token at the same position.

Project status: Registered as unlike. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

At a triggered prefix position, L = −log(1 − p_bad) + positive_weight × [−log(p_good)] when a good token is provided. Without a positive teacher only the first term applies. Triggers use probability/rank conditions rather than requiring the bad token to be argmax; non-triggered samples contribute zero unlikelihood loss.

Implementation and controls

objectives/unlike.py::unlike_loss takes prefix, bad_token_id, optional good_token_id, and trigger overrides. Defaults include positive_weight=1, rank threshold 20, probability threshold 0.05, epsilon 1e-6. The dispatched preconditions reject pure unlike without a KL sidecar unless explicitly overridden for research. They warn about non-target-position anchors, omitted surgery exclusions, and effective learning rate below the heuristic 5e-5 floor. Checks depend on trainer-plumbed metadata; a bare function call without that metadata does not run every dispatch check. The objective itself does not compute the anchor.

Evidence and evaluation

The HumanEval historical unlike-enabled RLVR seed-0 arm lost 18.8 percentage points relative to its cold anchor. The logical-reasoning pre-sample hybrid tied memorize rather than demonstrating an additional gain. Neither is an isolated token-correction efficacy proof. The repository contains exact-logit analyses and calibration proposals, which must be distinguished from production measurements.

Limitations and interpretation

The small-step positive-teacher counterexample shows why “lower learning rate is always safer” is unjustified. But optimizer LR is not the theorem's logit-step eta, and 5e-5 is not a universal safe floor. The composite KL/unlike operational bound remains unproved in the reviewed wiki safety doctrine. Negative mass can move to other unwanted tokens; validate both the intended correction and unrelated behavior.

Sources

See also