CCPD

From The Hei Canon
(Redirected from CCPD v2)

CCPD is Critique-Conditional Policy Distillation, the project's proposal for using the policy itself to turn a natural-language critique into a training signal.

Project status: CCPD v2 is conditionally registered as ccpd_v2; v1 is a superseded design. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

The detached scoring function is r_c(y) = β[mean log π_old(y given x,c) − mean log π_old(y given x)]. The critique changes context; the scoring path has no gradient. V2 assembles auxiliary candidates, computes centered rank advantages, then minimizes:

L = −mean(A_i log π_theta(y_i given x))
    + alpha × SFT(top-m candidates)
    + gamma × KL anchor

Here response log probabilities are length-normalized and A is detached. The policy learns without the critique in its inference prompt. The original v1 proposal used a differentiable critique-conditioned log-ratio construction; the plan revised it to auxiliary sampling plus detached rewards after its likelihood-displacement review.

Implementation and controls

objectives/ccpd.py implements score_rc, rank_advantages, and ccpd_v2_loss. One event per call is required. Fields: prompt, optional bad, critique, preferred, and pre-sampled aux_candidates. Defaults: k_aux=4, top_m=2, β=0.1, alpha=0.3, gamma=0.05, tau=0.5, max_new_tokens=256. Auxiliary generation uses critique context when available. Preferred-only scoring uses old-policy likelihood and pins preferred/bad examples to the ends; with a critique it uses r_c instead. π_old is a no-gradient view of current weights during the event, not a separately stored historical model. The built-in KL is over prompt positions. Registry import failure can omit CCPD while leaving other objectives available.

Evidence and evaluation

The status report's matched-k ranking benchmark used N=20, k=6: Qwen3-8B mean Spearman +0.207 (60% positive), Qwen3-0.6B +0.231 (65% positive). The 0.6B k=8 run gave +0.183. The recorded decision was a cautious opt-in/hinge-primary posture, not a reliable universal reward model. Tests established finite losses and gradients; the r_c training test checked movement, not improvement direction.

Limitations and interpretation

Implementation caveat: rank_advantages assigns ordinal ranks using double argsort, including on ties. The skip test compares the range of these rank advantages to tau, not the range of raw critique scores. For N candidates the rank range is normally N−1 even when raw scores tie, so the default is not a meaningful low-discrimination gate. The tau=10 test proves a high threshold skips, not that tied critiques are detected. Detached rewards remove gradients through scoring but do not remove reward misspecification, negative-advantage effects, or optimizer-induced forgetting. Prompt-only KL does not anchor full response behavior. The plan's stronger “safe by construction” and low-side-effect assertions are not end-to-end guarantees.

Sources

See also