In-context learning baseline

From The Hei Canon

In-context learning baseline is the project's no-weight-update comparison using demonstrations in the inference prompt.

Project status: Measured baseline in the logical, GSM8K, and HumanEval campaigns. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.

Mechanism

Provide the same available demonstrations as prompt context and generate an answer without training. Compare with context-to-weights methods at matched demonstration count, while also reporting the extra inference tokens and repeated prompt cost.

Implementation and controls

Historical autoresearch includes dedicated ICL probes. The essential controls are identical examples, heldout tasks, extraction/verification rules, model, and decoding mode. CoT-enabled evaluation is another inference-side lever and must be held constant or separately ablated; an evaluation-mode improvement is not necessarily learned weight improvement.

Evidence and evaluation

The May 16 GSM8K comparison used the same five examples and 50 heldout tasks: zero-shot cold 16/50 (32%), five-example fine-tuning 22/50 (44%), and five-shot ICL 48/50 (96%). The journal explicitly withdrew a sample-efficiency SOTA claim. Historical small-model/HumanEval lattice results differ by task and model and are not interchangeable with this comparison.

Limitations and interpretation

ICL pays context cost per query, while weight updates persist without demonstrations. Equal example count is not equal total compute, privacy properties, latency, or retention. Strong ICL at one setting does not imply it always dominates training; it establishes the necessary baseline for that claim.

Sources

See also