Define non-overlapping discovery, validation, optional pre-confirmation, and blind-test ranges before calculating candidate outcomes. Purge labels that cross a boundary and apply an embargo before model updates. The framework does not prescribe calendar dates.
- Validate unique asset-time keys, point-in-time availability, universe construction and adjustment policy.
- Freeze the expression vocabulary, proposal budgets, seeds, resource limits and candidate count.
- Generate typed expressions from mutation, crossover, perturbation, random and LLM-guided sources.
- Normalize and review expressions before any expensive evaluation.
- Evaluate discovery Rank IC and cross-sectional spreads without compounding overlapping labels.
- Correct the complete candidate family for multiple testing.
- Freeze direction and preprocessing from discovery only.
- Validate sign, stability, trimmed economics and regime behavior on untouched periods.
- Remove redundant factors by a predeclared correlation and mechanism rule.
- Lock survivors, then open pre-confirmation and blind partitions in order.
| Layer | Required evidence |
|---|---|
| Data | unique keys, available_at <= signal_time, declared universe, survivorship and adjustment audits |
| Expression | safe typed AST, allowed fields/operators, bounded depth/nodes/lookback, no future label access |
| Discovery | finite coverage, Rank IC, mean and remove-best-tail spread, FDR-adjusted significance |
| Validation | frozen direction, cross-period sign stability, tail robustness, non-redundancy |
| Portfolio | only when downstream construction is defined: turnover, costs, capacity and non-overlapping return path |
| Blind test | immutable candidate and execution lock; one opening; no reselection |
Store expression, source, parents, seed, data/config/code hashes, metrics, FDR family, correlations, rejection reasons and checkpoint index for every attempt. LLM prompts and responses are provenance only.
- Annualized returns or drawdowns made by compounding overlapping forward labels.
- Results using values whose release/availability time is after the signal.
- Thresholds selected after viewing validation or blind outcomes.
- A small accepted subset whose p-values ignore the discarded search population.
- Economic explanations without independently verified out-of-sample evidence.