Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CONTEXT.md
Original file line number Diff line number Diff line change
Expand Up @@ -292,6 +292,11 @@ Repository checkouts live under ignored `benchmarks/.cache/repos/`; generated ar
ignored `benchmarks/results/`. Benchmark Dense and Sparse vectors and channel rankings live under the
ignored `benchmarks/.cache/retrieval/v1/` cache; production indexes remain separate. See
`benchmarks/README.md` and `benchmarks/BASELINE.md`.
Schema 25 makes promotion an excluded-evidence decision: grouped intent folds and repository holdouts
must all be present and pass aggregate, query-form, and repository guardrails. Fit-all candidates are
diagnostic and inherit that decision. The last grouped fold is the explicit untouched final test and
must pass independently. Artifacts retain exact blocker values, deterministic grouped bootstrap
intervals, and fold-level selection/local-perturbation stability diagnostics.

### Scorer

Expand Down
29 changes: 20 additions & 9 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -211,10 +211,12 @@ Weight selection uses two grouped strategies:
- Leave-one-repository-out: calibrate on two repositories and validate on the third. This is emitted
only when multiple repositories are selected in the same run.

Candidate selection is development-only. Grouped and repository holdouts are reported separately; the
artifact records nested cross-validation as the final promotion plan rather than presenting fit-all
quality as an untouched final test. The Markdown report also emits unweighted query-form and repository
holdout rows with candidate-versus-production guardrail metrics.
Candidate selection is development-only inside each outer fold. Grouped intent folds and repository
holdouts are the outer evaluation. Schema 25 reserves the last grouped fold as the explicit untouched
final test; its candidate is selected without those samples, and a missing or failing final fold blocks
promotion. Fit-all quality remains diagnostic: its promotion status is copied from the aggregated
excluded-fold evidence and can never make itself eligible. The Markdown report emits aggregate,
unweighted query-form, and repository holdout rows with candidate-versus-production guardrails.

The grid optimizes development `Recall@20`, then `Recall@10`, 4k context recall, and MRR. Each fold's
weights are evaluated unchanged on its excluded samples. Only after cross-validation does the report
Expand All @@ -238,9 +240,17 @@ Identifier-shape and query-length slopes range from -1 to 1 so one query signal
while damping another. This bounded coarse-to-fine search avoids the combinatorial explosion of a full
20-parameter product; it is deterministic but does not claim a global optimum. Exact metric ties prefer
the lower-coefficient candidate. Each objective must remain within the configured 1% development
guardrail against the current Production router before it can be recommended.
If no candidate satisfies those guardrails, the artifact records `promotionStatus: no-eligible-candidate`;
the best non-eligible candidate remains diagnostic only and must not be promoted.
guardrail against the current Production router before it can be evaluated. Final promotion additionally
requires every configured outer strategy and every aggregate, query-form, and repository partition to
pass. Every failure records its strategy, fold, partition, metric, candidate value, baseline, tolerance,
and delta. If no candidate satisfies those guardrails, the artifact records
`promotionStatus: no-eligible-candidate`; the best non-eligible candidate remains diagnostic only.

Schema 25 adds deterministic grouped bootstrap intervals for paired candidate-minus-baseline deltas.
It also records exact-selection frequency across outer folds, the number and width of observed
single-coordinate perturbations, epsilon-neighbor fraction, and median/worst holdout drop. The search is
currently deterministic with one seed and restart; those counts are explicit rather than implying
unmeasured restart stability.

Each router result also records a deterministic `random-scout` baseline using the same parameter grid and
global-scout budget. `Random R@20` and `Random Ctx@4k` show whether the structured search beats that
Expand Down Expand Up @@ -350,8 +360,9 @@ output size without introducing an LLM or provider-specific tokenizer.

Each run writes ignored JSON and Markdown artifacts under `benchmarks/results`. JSON rows retain the
repository, revision, language, size, category, difficulty, query form, grouped fold, model, variant,
individual gold ranks, timing, and every metric. Schema 24 adds selectable router strategies while
retaining shared candidate-queue lifecycle and
individual gold ranks, timing, and every metric. Schema 25 derives promotion from excluded folds and adds
exact blockers, grouped bootstrap uncertainty, and selection-stability evidence. It retains schema 24's
selectable router strategies, shared candidate-queue lifecycle, and
per-router candidate-pool initialization timings. Each artifact stores each authored query and its exact
file-qualified ground truth once, records productive Sparse timings, and adds
the fixed equal-weight RRF baseline. The Markdown report includes quality by query form,
Expand Down
262 changes: 262 additions & 0 deletions benchmarks/retrieval/evaluation/promotion-evidence.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,262 @@
import type { EvidenceRouterConfig } from "../../../src/domain/retrieval.js"
import type {
CandidateStability,
EvidenceRouterSearchResult,
GuardrailBlocker,
HoldoutUncertainty,
PromotionEvidence,
QualityMetric,
QualitySummary,
RouterObjective,
ValidationStrategy,
} from "./types.js"

const BOOTSTRAP_SAMPLES = 1_000
const STABILITY_EPSILON = 0.005
const QUALITY_METRICS: readonly QualityMetric[] = [
"recallAt5",
"recallAt10",
"recallAt20",
"recallAt50",
"contextRecallAt4096",
"meanReciprocalRank",
]

/** Excluded-fold fields needed to derive promotion evidence. */
export type PromotionHoldoutRow = Pick<
EvidenceRouterSearchResult,
| "model"
| "fusion"
| "objective"
| "strategy"
| "fold"
| "validation"
| "productionValidation"
| "config"
| "holdoutBreakdown"
>

const precise = (value: number): number => Number(value.toFixed(12))

/** Describe every metric that falls below its baseline after applying the configured tolerance. */
export const buildGuardrailBlockers = (
partition: GuardrailBlocker["partition"],
name: string,
candidate: QualitySummary,
baseline: QualitySummary,
metrics: readonly QualityMetric[],
tolerance: number,
): readonly GuardrailBlocker[] =>
metrics.flatMap((metric) => {
const delta = precise(candidate[metric] - baseline[metric])
return delta < -tolerance
? [
{
partition,
name,
metric,
candidateValue: candidate[metric],
baselineValue: baseline[metric],
tolerance,
delta,
},
]
: []
})

const percentile = (values: readonly number[], fraction: number): number => {
if (values.length === 0) return 0
const sorted = [...values].sort((left, right) => left - right)
return sorted[Math.min(sorted.length - 1, Math.floor(fraction * sorted.length))] ?? 0
}

const bootstrapInterval = (
deltas: readonly number[],
): Pick<HoldoutUncertainty, "meanDelta" | "lowerBound" | "upperBound" | "bootstrapSamples"> => {
if (deltas.length === 0)
return { meanDelta: 0, lowerBound: 0, upperBound: 0, bootstrapSamples: 0 }
let state = 1
const means = Array.from({ length: BOOTSTRAP_SAMPLES }, () => {
let total = 0
for (let index = 0; index < deltas.length; index++) {
state = (Math.imul(state, 1_664_525) + 1_013_904_223) >>> 0
total += deltas[state % deltas.length] ?? 0
}
return total / deltas.length
})
return {
meanDelta: precise(deltas.reduce((sum, delta) => sum + delta, 0) / deltas.length),
lowerBound: precise(percentile(means, 0.025)),
upperBound: precise(percentile(means, 0.975)),
bootstrapSamples: BOOTSTRAP_SAMPLES,
}
}

const buildUncertainty = (rows: readonly PromotionHoldoutRow[]): readonly HoldoutUncertainty[] => {
const observations = new Map<
string,
{
strategy: ValidationStrategy
partition: GuardrailBlocker["partition"]
name: string
metric: QualityMetric
deltas: number[]
}
>()
for (const row of rows)
for (const holdout of row.holdoutBreakdown)
for (const metric of QUALITY_METRICS) {
const key = `${row.strategy}\0${holdout.dimension}\0${holdout.name}\0${metric}`
const observation = observations.get(key) ?? {
strategy: row.strategy,
partition: holdout.dimension,
name: holdout.name,
metric,
deltas: [],
}
observation.deltas.push(holdout.candidate[metric] - holdout.baseline[metric])
observations.set(key, observation)
}
return [...observations.values()].map(({ deltas, ...observation }) => ({
...observation,
...bootstrapInterval(deltas),
}))
}

const objectiveMetric = (objective: RouterObjective): QualityMetric => {
if (objective === "direct") return "recallAt5"
return objective === "reranker-top20" ? "recallAt20" : "recallAt50"
}

const configValues = (config: EvidenceRouterConfig): readonly number[] => [
...Object.values(config.baseWeights),
...Object.values(config.scoreInfluence),
...Object.values(config.geometryInfluence),
...Object.values(config.termCoverageInfluence),
...Object.values(config.pairwiseAgreementInfluence),
...Object.values(config.denseConfidenceInfluence),
...Object.values(config.identifierInfluence),
...Object.values(config.queryLengthInfluence),
]

const differingCoordinates = (
left: EvidenceRouterConfig,
right: EvidenceRouterConfig,
): { readonly count: number; readonly width: number } => {
const leftValues = configValues(left)
const rightValues = configValues(right)
let count = 0
let width = 0
for (let index = 0; index < leftValues.length; index++) {
const difference = Math.abs((leftValues[index] ?? 0) - (rightValues[index] ?? 0))
if (difference > 0) {
count++
width = Math.max(width, difference)
}
}
return { count, width }
}

const selectionFrequency = (
rows: readonly PromotionHoldoutRow[],
): Pick<CandidateStability, "distinctSelections" | "selectionFrequency"> => {
const frequencies = new Map<string, number>()
for (const row of rows) {
const key = JSON.stringify(row.config)
frequencies.set(key, (frequencies.get(key) ?? 0) + 1)
}
return {
distinctSelections: frequencies.size,
selectionFrequency: rows.length === 0 ? 0 : Math.max(0, ...frequencies.values()) / rows.length,
}
}

const localPerturbations = (
rows: readonly PromotionHoldoutRow[],
): Pick<CandidateStability, "localPerturbations" | "plateauWidth" | "epsilonNeighborFraction"> => {
const neighbors: Array<{ readonly qualityDifference: number; readonly width: number }> = []
for (let leftIndex = 0; leftIndex < rows.length; leftIndex++)
for (let rightIndex = leftIndex + 1; rightIndex < rows.length; rightIndex++) {
const left = rows[leftIndex]
const right = rows[rightIndex]
if (left === undefined || right === undefined) continue
const distance = differingCoordinates(left.config, right.config)
if (distance.count !== 1) continue
const metric = objectiveMetric(left.objective)
neighbors.push({
qualityDifference: Math.abs(left.validation[metric] - right.validation[metric]),
width: distance.width,
})
}
const epsilonNeighbors = neighbors.filter(
({ qualityDifference }) => qualityDifference <= STABILITY_EPSILON,
)
return {
localPerturbations: neighbors.length,
plateauWidth: precise(Math.max(0, ...epsilonNeighbors.map(({ width }) => width))),
epsilonNeighborFraction:
neighbors.length === 0 ? 0 : epsilonNeighbors.length / neighbors.length,
}
}

const buildStability = (rows: readonly PromotionHoldoutRow[]): CandidateStability => {
const metric = rows[0] === undefined ? "recallAt20" : objectiveMetric(rows[0].objective)
const holdoutDrops = rows.map((row) => row.productionValidation[metric] - row.validation[metric])
return {
folds: rows.length,
...selectionFrequency(rows),
...localPerturbations(rows),
medianHoldoutDrop: precise(percentile(holdoutDrops, 0.5)),
worstCaseHoldoutDrop: precise(Math.max(0, ...holdoutDrops)),
seeds: 1,
restarts: 1,
}
}

const promotionKey = (row: PromotionHoldoutRow): string =>
`${row.model}\0${row.fusion}\0${row.objective}`

/** Derive promotion decisions exclusively from excluded-fold results, never fit-all quality. */
export const derivePromotionEvidence = (
rows: readonly PromotionHoldoutRow[],
expectedStrategies: readonly ValidationStrategy[],
finalTest: { readonly strategy: ValidationStrategy; readonly fold: string },
): readonly PromotionEvidence[] => {
const groups = new Map<string, PromotionHoldoutRow[]>()
for (const row of rows)
groups.set(promotionKey(row), [...(groups.get(promotionKey(row)) ?? []), row])
return [...groups.values()].map((group) => {
const first = group[0]
if (first === undefined) throw new Error("Promotion evidence group cannot be empty")
const strategies = new Set(group.map((row) => row.strategy))
const missingStrategies = expectedStrategies.filter((strategy) => !strategies.has(strategy))
const finalTestRow = group.find(
(row) => row.strategy === finalTest.strategy && row.fold === finalTest.fold,
)
const finalTestGuardrailsMet =
finalTestRow?.holdoutBreakdown.every((holdout) => holdout.guardrailsMet) ?? false
const blockers = group.flatMap((row) =>
row.holdoutBreakdown.flatMap((holdout) =>
holdout.blockers.map((blocker) => ({ ...blocker, strategy: row.strategy, fold: row.fold })),
),
)
return {
model: first.model,
fusion: first.fusion,
objective: first.objective,
promotionStatus:
missingStrategies.length === 0 && finalTestGuardrailsMet && blockers.length === 0
? "eligible"
: "no-eligible-candidate",
missingStrategies,
finalTest: {
...finalTest,
present: finalTestRow !== undefined,
guardrailsMet: finalTestGuardrailsMet,
},
blockers,
uncertainty: buildUncertainty(group),
stability: buildStability(group),
}
})
}
49 changes: 48 additions & 1 deletion benchmarks/retrieval/evaluation/report.ts
Original file line number Diff line number Diff line change
Expand Up @@ -62,6 +62,51 @@ const formatRouterWeightColumns = (result: {
influences: formatInfluences(result.config),
})

const renderPromotionEvidence = (artifact: BenchmarkArtifact): readonly string[] => {
const summaries = artifact.promotionEvidence.map(
(evidence) =>
`| ${evidence.model} | ${evidence.fusion} | ${evidence.objective} | ${promotionLabel(evidence.promotionStatus)} | ${evidence.missingStrategies.join(", ") || "none"} | ${evidence.finalTest.strategy}:${evidence.finalTest.fold} (${evidence.finalTest.guardrailsMet ? "pass" : "fail"}) | ${evidence.stability.folds} | ${evidence.stability.distinctSelections} | ${percent(evidence.stability.selectionFrequency)} | ${evidence.stability.localPerturbations} | ${evidence.stability.plateauWidth.toFixed(3)} | ${percent(evidence.stability.epsilonNeighborFraction)} | ${percent(evidence.stability.medianHoldoutDrop)} | ${percent(evidence.stability.worstCaseHoldoutDrop)} |`,
)
const blockers = artifact.promotionEvidence.flatMap((evidence) =>
evidence.blockers.map(
(blocker) =>
`| ${evidence.model} | ${evidence.fusion} | ${evidence.objective} | ${blocker.strategy} | ${blocker.fold} | ${blocker.partition}:${blocker.name} | ${blocker.metric} | ${percent(blocker.candidateValue)} | ${percent(blocker.baselineValue)} | ${percent(blocker.tolerance)} | ${percent(blocker.delta)} |`,
),
)
const uncertainty = artifact.promotionEvidence.flatMap((evidence) =>
evidence.uncertainty.map(
(interval) =>
`| ${evidence.model} | ${evidence.fusion} | ${evidence.objective} | ${interval.strategy} | ${interval.partition}:${interval.name} | ${interval.metric} | ${percent(interval.meanDelta)} | ${percent(interval.lowerBound)} | ${percent(interval.upperBound)} | ${interval.bootstrapSamples} |`,
),
)
return [
"",
"## Promotion Evidence",
"",
"Fit-all quality is diagnostic. Promotion status below is derived only from excluded grouped and repository holdouts; missing required strategies block promotion.",
"",
"| Model | Fusion | Objective | Promotion | Missing strategies | Final test | Folds | Distinct selections | Selection frequency | Local perturbations | Plateau width | Epsilon neighbors | Median drop | Worst drop |",
"| --- | --- | --- | --- | --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |",
...summaries,
"",
"### Guardrail Blockers",
"",
"| Model | Fusion | Objective | Strategy | Fold | Partition | Metric | Candidate | Baseline | Tolerance | Delta |",
"| --- | --- | --- | --- | --- | --- | --- | ---: | ---: | ---: | ---: |",
...(blockers.length === 0
? ["| - | - | - | - | - | - | - | - | - | - | no blockers |"]
: blockers),
"",
"### Holdout Uncertainty",
"",
"Deterministic grouped bootstrap intervals resample excluded folds and report paired candidate-minus-baseline deltas.",
"",
"| Model | Fusion | Objective | Strategy | Partition | Metric | Mean delta | 95% lower | 95% upper | Bootstrap samples |",
"| --- | --- | --- | --- | --- | --- | ---: | ---: | ---: | ---: |",
...uncertainty,
]
}

/** Render quality and marginal channel contribution grouped by query representation. */
export const renderMarkdownReport = (artifact: BenchmarkArtifact): string => {
const groups = new Map<string, QueryMeasurement[]>()
Expand Down Expand Up @@ -92,7 +137,7 @@ export const renderMarkdownReport = (artifact: BenchmarkArtifact): string => {
.map(([kind, weight]) => `${kind}=${weight}`)
.join(", ")}.`,
"",
`Validation protocol: candidates use ${artifact.validationProtocol.selection}; holdouts are ${artifact.validationProtocol.holdouts.join(" and ")}; final promotion requires the recorded ${artifact.validationProtocol.finalTest}.`,
`Validation protocol: candidates use ${artifact.validationProtocol.selection}; holdouts are ${artifact.validationProtocol.holdouts.join(" and ")}; final promotion requires untouched ${artifact.validationProtocol.finalTest.strategy} fold ${artifact.validationProtocol.finalTest.fold}.`,
"",
`Search strategy: \`${artifact.searchStrategy.algorithm}\` (${artifact.searchStrategy.globalScouts} global scouts, beam ${artifact.searchStrategy.beamWidth}, ${artifact.searchStrategy.coordinatePasses} coordinate passes, ${artifact.searchStrategy.proxySampleFraction * 100}% proxy with minimum ${artifact.searchStrategy.proxyMinimumSamples}, ${strategyFactorLabel} factor ${strategyFactor}x).`,
"",
Expand Down Expand Up @@ -323,6 +368,8 @@ export const renderMarkdownReport = (artifact: BenchmarkArtifact): string => {
)
}

lines.push(...renderPromotionEvidence(artifact))

lines.push(
"",
"## Recommended Evidence Router",
Expand Down
Loading
Loading