In-context learning baseline
In-context learning baseline is the project's no-weight-update comparison using demonstrations in the inference prompt.
Project status: Measured baseline in the logical, GSM8K, and HumanEval campaigns. This entry describes the source audit of 14 September 2026; historical measurements retain their original dates.
Mechanism
Provide the same available demonstrations as prompt context and generate an answer without training. Compare with context-to-weights methods at matched demonstration count, while also reporting the extra inference tokens and repeated prompt cost.
Implementation and controls
Historical autoresearch includes dedicated ICL probes. The essential controls are identical examples, heldout tasks, extraction/verification rules, model, and decoding mode. CoT-enabled evaluation is another inference-side lever and must be held constant or separately ablated; an evaluation-mode improvement is not necessarily learned weight improvement.
Evidence and evaluation
The May 16 GSM8K comparison used the same five examples and 50 heldout tasks: zero-shot cold 16/50 (32%), five-example fine-tuning 22/50 (44%), and five-shot ICL 48/50 (96%). The journal explicitly withdrew a sample-efficiency SOTA claim. Historical small-model/HumanEval lattice results differ by task and model and are not interchangeable with this comparison.
Limitations and interpretation
ICL pays context cost per query, while weight updates persist without demonstrations. Equal example count is not equal total compute, privacy properties, latency, or retention. Strong ICL at one setting does not imply it always dominates training; it establishes the necessary baseline for that claim.
Sources
- cont: docs/research/JOURNAL.md — checkout audited
87946914c7b9. - agi: autoresearch/HUMANEVAL_ICL_RESULTS.md — historical revision
3842fd8875ca. - agi: autoresearch/LATTICE.md — historical revision
3842fd8875ca.