heldout is an honest search-policy harness: a hypothesis-tree ratchet where the anti-overfitting rules are enforced in executable code, not prompts. A dev-slice win can't be kept unless an untouched held-out slice confirms it; the baseline only ratchets up on a fresh re-measure; every killed branch leaves a lesson; and no LLM ever grades its own work. Proven (Phase 1 silver on MLE-bench for $3.13) and honestly nulled (two pre-registered equal-budget A/Bs vs a dumb retry loop). MIT, zero runtime deps, works as a Claude Code skill.
Built with