You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Issue #163 established the production fusion seam, the five retrieval channels, the compatibility RRF path, and explicit runtime profile selection. The remaining work is benchmark-only: complete the optimization evidence, choose candidates robustly, and make promotion decisions explainable.
This issue must not add schema versioning to the production evidence-router configuration and must not change the Production compatibility default. Benchmark artifacts may remain versioned.
Current Baseline
The repository already contains the main benchmark plumbing:
FUSION_METHODS includes RRF, Relative Score fusion, and DBSF.
OPTIMIZATION_PROFILES includes search-priority, balanced, code-navigation, basic-exploration, and natural-language.
ROUTER_OBJECTIVES includes direct, reranker-top20, and reranker-top50.
The full benchmark profile is configured to evaluate all supported fusion methods.
PIX_BENCH_OPTIMIZATION_PROFILE selects one benchmark optimization profile per run.
The evidence-router search records parameter levels, candidate counts, proxy/full evaluations, cache hits, protected elites, and proxy agreement.
A deterministic random-scout baseline is implemented and reported beside router results.
Grouped and leave-one-repository-out holdout code exists, but full current-schema matrix evidence has not been completed because runs are long.
What To Build
Complete and harden the benchmark optimization and promotion-evidence track without changing Production query behavior.
Acceptance Criteria
Execute or shard the complete matrix across every supported fusion method, benchmark optimization profile, router objective, repository, and model; mergeable artifacts must retain the exact matrix coordinates.
Keep benchmark artifact versioning, but keep the Production evidence-router configuration unversioned and validated by its structural schema.
Record the effective search dimensionality, parameter levels, normalization, budget, seeds, tie-breaking, raw/unique candidates, cache hits, proxy/full evaluations, promotions, protected elites, and best observed candidates.
Compare the Halton/beam search against the deterministic random-scout baseline on the same holdouts; report quality deltas and candidate counts. A Sobol baseline is optional if it adds useful coverage.
Keep the search explicitly derivative-free and heuristic; do not claim global optimality.
If the benchmark labels support it, add an optional supervised/Learning-to-Rank comparison track. If they do not, document why it is not valid for this corpus.
If soft-rank, soft-top-K, surrogate, or other proxy objectives are used, report the proxy-to-exact metric gap; exact Recall@K and context recall remain the final decision metrics.
Define guardrail tiers explicitly: hard deployment guardrails, diagnostic metrics, and operational guardrails.
Decide and document whether MRR, Success@K, additional context budgets, latency, and memory are hard guardrails or diagnostics; do not add them implicitly.
For every rejected candidate, report the exact blocking guardrails: aggregate/query-form/repository partition, metric, candidate value, baseline value, tolerance, and delta. A boolean yes/no is insufficient.
Preserve explicit no-eligible-candidate outcomes; a guardrail-failing diagnostic candidate must never be presented as promotable.
Select candidates only from development data, evaluate grouped and leave-one-repository-out holdouts, and implement an untouched final test or actual nested cross-validation rather than only recording a plan.
Add grouped bootstrap or equivalent uncertainty estimates for query-form and repository comparisons.
Add robust-selection diagnostics: discrete local perturbations, plateau width, epsilon-neighbor fraction, median/worst-case drop, selection frequency across folds/seeds/restarts, rank-20 versus rank-21 cutoff margins, and complexity-aware tie-breaking.
Extend Shapley and marginal channel diagnostics with inactive/redundant dimension detection, one-at-a-time ablations, and a documented sensitivity method such as Morris, Sobol, or fANOVA; retain important interaction effects.
Add explicit edge-case tests for every supported fusion method covering empty, one-element, constant, short, tied, translated, rescaled, outlier, out-of-range, missing-channel, zero-weight, candidate-depth, and evidence-signal behavior.
Reports must retain per-query-form, per-repository, per-model, aggregate, uncertainty, guardrail-blocker, and promotion-status results.
Document why the selected candidate is preferred over the highest-fit candidate and why rejected candidates are not eligible for Production promotion.
Re-run the project quality gates and document any inherited Fallow findings separately from benchmark correctness.
Out Of Scope
Changing the Production compatibility default.
Adding schema versioning to the Production evidence-router configuration.
Using authored benchmark labels as implicit runtime query intent.
Blocked By
None - the current benchmark seam and search implementation already exist. Production promotion remains a separate decision in #163 after this issue produces sufficient evidence.
Code-Verified Review (HEAD 78bc1ef)
This section reconciles the retrieval/fusion review with the current implementation. It is intentionally explicit about what is already present and what still needs work for reliable retrieval evidence.
Already present, do not reimplement
src/lib/retrieval/fusion.ts already caches prepared rankings, per-method normalization, channel positions, and reusable typed-array work state. The earlier generic cache concern is addressed by preparedRankingsCache; a full feature-matrix or partial top-K optimization is optional follow-up work only if profiling still shows fusion evaluation as the bottleneck.
Benchmark optimization profiles are selected through PIX_BENCH_OPTIMIZATION_PROFILE, carried into runner.ts, used for weighted query-form summaries, and persisted in schema-20 artifacts.
Router diagnostics already record the declared parameter count and levels, raw/unique candidates, proxy/full evaluations and cache hits, proxy promotions, proxy/full agreement, and protected elites. A deterministic random-scout comparison and explicit no-eligible-candidate status also exist.
The production compatibility default remains separate from benchmark profiles and remains unchanged.
Code-verified gaps that must be resolved
Align the selection objective with the selected optimization profile.optimizeWeights and optimizeFusionWeights still call selectBestWeights(..., reranker-top20, ...) explicitly. For the default search-priority profile, whose declared objective is direct, this can select static candidates with the fallback reranker-top20 priority instead of the profile priority. The objective used for candidate selection must be the declared profile objective, unless the distinction is intentional and recorded.
Make promotion status holdout-derived.promotionStatus is currently assigned from development-time or fit-all guardrail checks. holdoutBreakdown is emitted later but does not determine promotion status. A candidate that passes development and fails an excluded fold can therefore still be labelled eligible. Fit-all candidates must be diagnostic-only; only aggregated grouped and repository holdouts, plus the final test protocol, may produce a promotable status.
Implement the final test rather than recording only a plan. Schema 20 records finalTest: nested-cross-validation-plan, but no nested outer/inner evaluation is executed. Either implement nested cross-validation or keep promotion blocked until an untouched final test exists.
Measure the candidate-set ceiling. Add recall of the gold targets in the union of all channel top-candidateDepth lists, before fusion. Without unionRecallAtCandidateDepth, a fusion result cannot distinguish a bad ordering from a missing candidate and reranker decisions remain under-specified.
Make the global scout comparison objective-equivalent.buildGlobalRouterSeeds samples only the influence coordinates (parameters.slice(CHANNELS.length)); base weights are inherited from a small seed set. The random scout samples all coordinates, but selectRandomRouter ranks it with the hard-coded reranker-top20 objective and reuses one random candidate for all router objectives. Global exploration and the random baseline must cover the same coordinates and use the same objective, guardrails, and holdout protocol.
Strengthen proxy validation.buildProxySamples stratifies by repository and query form only. Include category and difficulty, record per-stratum counts, and measure false-pruning of candidates that are strong on the full development set. Set-overlap of the proxy and full top candidate sets is useful but is not a rank correlation or a guarantee that the full-development winner survived.
Report effective dimensionality. The router declares 40 coordinates, but current evidence makes six coordinates structurally neutral: non-Dense denseConfidenceInfluence values and Dense/Sparse termCoverageInfluence values do not affect routeWithEvidence. Record active, inactive, and data-constant coordinates separately; use the effective dimension for search-budget interpretation.
Emit exact guardrail blockers.HoldoutQuality currently stores candidate values, baseline values, and one boolean. Reports must additionally identify every failing metric with candidate value, baseline value, tolerance, and delta, including the query-form or repository partition that failed.
Add uncertainty and selection stability before promotion. The current benchmark has no grouped bootstrap or equivalent uncertainty estimates, no local perturbation/plateau diagnostics, and no selection-frequency report across folds, seeds, or restarts. These are required to distinguish a robust parameter family from a discrete metric tie.
Provide a mergeable matrix workflow. One runner invocation selects one optimization profile and one model selection. A complete matrix therefore requires multiple runs, but the current benchmark does not provide artifact merge/coverage validation. Add a manifest/merge command that rejects missing or duplicate coordinates and preserves profile, fusion, model, repository, objective, and validation strategy.
Code-retrieval-specific fusion checks
Keep the current one-element/constant-list score of 0.5 as the compatibility baseline, but benchmark a separate exact-presence variant. For an exact Identity hit, 0.5 is intentionally neutral for score fusion but may understate the strongest code-navigation signal described by ADR-0013. Do not silently change the production compatibility behavior without an ablation.
Add an ablation for missing-channel agreement. channelPairwiseAgreement currently averages over four possible peers even when peer lists are empty, so unavailable evidence can look like disagreement. Compare this with an available-peer denominator and document the chosen semantics.
Treat identifierLikelihood as a measured hypothesis, not a semantic fact. The current implementation is essentially a non-empty, no-whitespace test. Add code-shape variants such as namespace/path separators, CamelCase, and snake_case only as benchmark experiments with per-form and per-repository holdouts.
Add a candidate-union diagnostic before an optional candidate-level Learning-to-Rank track. The exact file/symbol labels support such a comparison, but the four query forms of one intent are correlated and must remain grouped. A linear or pairwise model should be regularized and evaluated with the same repository holdouts; it must not replace the core fusion validation prematurely.
Prioritized optional fusion experiments
After the core validation protocol is correct, add benchmark-only candidates in this order:
CombMNZ or an equivalent consensus bonus on Relative Score and DBSF.
Interpolation between normalized RRF and score fusion.
Robust score normalization using quantiles or median/MAD in addition to mean/standard-deviation DBSF.
A rank-segment probability baseline such as ProbFuse if the corpus has enough independent relevance judgments.
Condorcet/Kemeny, Markov-chain aggregation, and high-dimensional CMA-ES are not required for the first reliable benchmark. They add complexity without addressing the current holdout, candidate-ceiling, and parameter-stability gaps.
Parallel Search Execution
The router search is CPU-bound and currently evaluates candidates synchronously. Effect.forEach with a concurrency option does not move synchronous CPU work to additional cores: it schedules Effects/fibers, but the synchronous evaluator still blocks the Node event loop.
Decision
Use a small native node:worker_threads pool for the benchmark's CPU-bound candidate evaluation. Effect Worker is technically viable in the installed effect@4.0.0-beta.98 and @effect/platform-node packages, but it is an unstable worker/spawner protocol around the same Node worker threads. This benchmark has a narrow internal protocol and no reusable application worker service, so native workers have the smaller surface and lower integration risk. Effect may still wrap pool acquisition, shutdown, failure propagation, and interruption at the benchmark boundary; it is not required to perform the CPU parallelism.
Do not split individual parameter axes. Beam coordinates, protected elites, the archive, and objective-specific selection are stateful and order-sensitive. Split independent candidate evaluations into batches instead, while keeping beam/archive/cache updates and deterministic tie-breaking in the main thread.
Additional Acceptance Criteria
Add a configurable native worker pool for proxy and full candidate evaluation. Default to max(1, availableParallelism - 1) workers, with an environment override and a serial fallback for one worker or unsupported runtimes.
Send a compact immutable prepared search snapshot to workers. Do not clone full source text for every candidate task; measure snapshot serialization and transfer time separately from evaluation time.
Keep beam selection, protected-elite handling, archive updates, cache ownership, and final objective selection deterministic in the main thread. Worker result ordering must not affect the selected candidate.
Batch candidate configurations per worker to amortize message and startup overhead. Reuse a fixed pool during one router search rather than spawning one worker per candidate.
Compare serial and worker-pool runs on identical snapshots and seeds. Their selected configurations and exact quality summaries must match; report wall time, worker count, batch size, worker startup, serialization, proxy, and full-evaluation time.
Handle worker failure, interruption, and scope cleanup explicitly. A failed parallel run must either fail the benchmark clearly or use the documented serial fallback; it must never silently produce partial search evidence.
Keep the worker pool benchmark-only. Production query behavior and the Production compatibility default must not depend on worker availability.
Missing Router-Model And Search Optimizations
The earlier design review also proposed the following items. They were not yet captured above and are added here so the issue covers the complete optimization path rather than only execution parallelism and promotion evidence.
Compare the current multiplicative router kernel with a log-linear gate. Evaluate a bounded model such as w_c(q) = b_c * exp(beta_c dot evidence(q)), followed by the existing scale normalization. Record the same holdout, complexity, and stability diagnostics. The current product of score, geometry, coverage, agreement, identifier, and length factors can amplify interactions; the log-linear form provides an additive, regularizable comparison.
Separate static fusion fitting from evidence fitting. Evaluate a staged search that first selects static fusion weights, then freezes them while fitting evidence influences. Also compare alternating block-coordinate fitting with the current joint search of base weights and influences. Report whether dynamic routing adds value beyond a better static baseline.
Add search restarts and a local-search comparison. Compare the current Halton/beam path with deterministic random-restart coordinate descent or Sobol/global-scout plus local refinement. Keep the best candidate only through the same development-only selection and holdout protocol. Do not claim a global optimum.
Add a scalar or Pareto utility with explicit regularization. The current lexicographic metric priorities create broad discrete plateaus. Compare the profile priority with an explicit utility that can penalize router complexity, context regressions, and unstable coefficients. The selected candidate should be the simplest candidate within the configured quality/uncertainty tolerance, not automatically the highest fit-all row.
Prune structurally inactive parameters before search. Reporting the six neutral coordinates is insufficient. The optimizer should be able to remove or lock parameters that cannot affect the current evidence representation, while preserving an opt-in full-dimensional diagnostic mode for detecting future signal changes.
Complete the prepared-evaluator optimization if profiling justifies it. The existing ranking cache removes repeated normalization and position work, but each candidate still rebuilds and sorts fused results. Benchmark a compact per-sample feature matrix and top-K/partial-selection evaluator against the current implementation. Keep exact metric equality as the correctness requirement.
Report both weighted and macro aggregates. Preserve the configured query-form weights, but also report macro averages across query forms and repositories. Candidate selection and guardrails must not be able to hide a repository or query-form regression through sample-count weighting alone.
Add a one-standard-error or equivalent simplicity rule. When several candidates are statistically indistinguishable, prefer the lower-complexity candidate with fewer active influences, smaller coefficient magnitude, and more stable selection frequency.
Verify score semantics at the fusion boundary. Every channel must declare that higher scores mean better matches before Relative Score or DBSF normalization. Add tests for similarity versus distance orientation, translation/rescaling invariance where intended, and outlier behavior; do not infer comparability from numeric ranges alone.
Rank-Sensitive Quality Metric: NDCG
Add NDCG@5, NDCG@10, NDCG@20, and NDCG@50 to the benchmark measurements and quality summaries. NDCG (Normalized Discounted Cumulative Gain) complements Recall@K and MRR by measuring the ordering of all relevant chunks within the cutoff rather than only whether a target was found.
Define the gain function explicitly for the existing exact file + symbol labels. At minimum, a chunk matching one or more resolved gold targets receives binary relevance; if multi-target questions are retained, document whether gain is capped at one or proportional to the number of distinct gold targets represented by that chunk.
Define how overlapping chunks, duplicate matches, missing targets, and zero-ideal-DCG queries are handled. Keep the four query forms of one intent grouped during validation.
Report NDCG by query form, repository, model, fusion, aggregate, and holdout partition. Use it as a diagnostic or selectable objective first; promote it to a hard guardrail only after its behavior is stable and its relationship to context recall is documented.
Direct Retrieval Objective
Make NDCG@5 the primary rank-sensitive optimization metric for the direct objective because Production returns five results by default and direct retrieval exposes the fused order to the user. Report NDCG@10 and NDCG@20 alongside it for larger requested result sets.
Keep Recall@20/Recall@50 and ContextRecall@4096 as hard coverage and context-budget guardrails for direct retrieval. NDCG must not improve the first few ranks by silently losing additional gold symbols or useful context.
Keep MRR as a diagnostic rather than the sole direct objective. With binary relevance and one gold target, NDCG and MRR are closely related; with multiple exact file + symbol targets, NDCG additionally rewards placing all relevant chunks early.
Keep reranker-top20 and reranker-top50 primarily recall-oriented because their purpose is candidate coverage before a later reranker. Use their NDCG values diagnostically unless a reranker-specific ordering objective is explicitly introduced.
Run an objective ablation comparing the current direct priority (Recall@5/Recall@10 first) with NDCG@5 first under identical folds, profiles, guardrails, and seeds. Select the objective from excluded-fold evidence rather than fit-all quality.
Parent
#163
Why This Is Separate
Issue #163 established the production fusion seam, the five retrieval channels, the compatibility RRF path, and explicit runtime profile selection. The remaining work is benchmark-only: complete the optimization evidence, choose candidates robustly, and make promotion decisions explainable.
This issue must not add schema versioning to the production evidence-router configuration and must not change the Production compatibility default. Benchmark artifacts may remain versioned.
Current Baseline
The repository already contains the main benchmark plumbing:
FUSION_METHODSincludes RRF, Relative Score fusion, and DBSF.OPTIMIZATION_PROFILESincludes search-priority, balanced, code-navigation, basic-exploration, and natural-language.ROUTER_OBJECTIVESincludes direct, reranker-top20, and reranker-top50.PIX_BENCH_OPTIMIZATION_PROFILEselects one benchmark optimization profile per run.What To Build
Complete and harden the benchmark optimization and promotion-evidence track without changing Production query behavior.
Acceptance Criteria
no-eligible-candidateoutcomes; a guardrail-failing diagnostic candidate must never be presented as promotable.Out Of Scope
Blocked By
None - the current benchmark seam and search implementation already exist. Production promotion remains a separate decision in #163 after this issue produces sufficient evidence.
Code-Verified Review (HEAD
78bc1ef)This section reconciles the retrieval/fusion review with the current implementation. It is intentionally explicit about what is already present and what still needs work for reliable retrieval evidence.
Already present, do not reimplement
src/lib/retrieval/fusion.tsalready caches prepared rankings, per-method normalization, channel positions, and reusable typed-array work state. The earlier generic cache concern is addressed bypreparedRankingsCache; a full feature-matrix or partial top-K optimization is optional follow-up work only if profiling still shows fusion evaluation as the bottleneck.PIX_BENCH_OPTIMIZATION_PROFILE, carried intorunner.ts, used for weighted query-form summaries, and persisted in schema-20 artifacts.no-eligible-candidatestatus also exist.Code-verified gaps that must be resolved
optimizeWeightsandoptimizeFusionWeightsstill callselectBestWeights(..., reranker-top20, ...)explicitly. For the defaultsearch-priorityprofile, whose declared objective isdirect, this can select static candidates with the fallbackreranker-top20priority instead of the profile priority. The objective used for candidate selection must be the declared profile objective, unless the distinction is intentional and recorded.promotionStatusis currently assigned from development-time or fit-all guardrail checks.holdoutBreakdownis emitted later but does not determine promotion status. A candidate that passes development and fails an excluded fold can therefore still be labelledeligible. Fit-all candidates must be diagnostic-only; only aggregated grouped and repository holdouts, plus the final test protocol, may produce a promotable status.finalTest: nested-cross-validation-plan, but no nested outer/inner evaluation is executed. Either implement nested cross-validation or keep promotion blocked until an untouched final test exists.candidateDepthlists, before fusion. WithoutunionRecallAtCandidateDepth, a fusion result cannot distinguish a bad ordering from a missing candidate and reranker decisions remain under-specified.buildGlobalRouterSeedssamples only the influence coordinates (parameters.slice(CHANNELS.length)); base weights are inherited from a small seed set. The random scout samples all coordinates, butselectRandomRouterranks it with the hard-codedreranker-top20objective and reuses one random candidate for all router objectives. Global exploration and the random baseline must cover the same coordinates and use the same objective, guardrails, and holdout protocol.buildProxySamplesstratifies by repository and query form only. Include category and difficulty, record per-stratum counts, and measure false-pruning of candidates that are strong on the full development set. Set-overlap of the proxy and full top candidate sets is useful but is not a rank correlation or a guarantee that the full-development winner survived.denseConfidenceInfluencevalues and Dense/SparsetermCoverageInfluencevalues do not affectrouteWithEvidence. Record active, inactive, and data-constant coordinates separately; use the effective dimension for search-budget interpretation.HoldoutQualitycurrently stores candidate values, baseline values, and one boolean. Reports must additionally identify every failing metric with candidate value, baseline value, tolerance, and delta, including the query-form or repository partition that failed.Code-retrieval-specific fusion checks
0.5as the compatibility baseline, but benchmark a separate exact-presence variant. For an exact Identity hit,0.5is intentionally neutral for score fusion but may understate the strongest code-navigation signal described by ADR-0013. Do not silently change the production compatibility behavior without an ablation.channelPairwiseAgreementcurrently averages over four possible peers even when peer lists are empty, so unavailable evidence can look like disagreement. Compare this with an available-peer denominator and document the chosen semantics.identifierLikelihoodas a measured hypothesis, not a semantic fact. The current implementation is essentially a non-empty, no-whitespace test. Add code-shape variants such as namespace/path separators, CamelCase, and snake_case only as benchmark experiments with per-form and per-repository holdouts.Prioritized optional fusion experiments
After the core validation protocol is correct, add benchmark-only candidates in this order:
Condorcet/Kemeny, Markov-chain aggregation, and high-dimensional CMA-ES are not required for the first reliable benchmark. They add complexity without addressing the current holdout, candidate-ceiling, and parameter-stability gaps.
Parallel Search Execution
The router search is CPU-bound and currently evaluates candidates synchronously.
Effect.forEachwith a concurrency option does not move synchronous CPU work to additional cores: it schedules Effects/fibers, but the synchronous evaluator still blocks the Node event loop.Decision
Use a small native
node:worker_threadspool for the benchmark's CPU-bound candidate evaluation. Effect Worker is technically viable in the installedeffect@4.0.0-beta.98and@effect/platform-nodepackages, but it is an unstable worker/spawner protocol around the same Node worker threads. This benchmark has a narrow internal protocol and no reusable application worker service, so native workers have the smaller surface and lower integration risk. Effect may still wrap pool acquisition, shutdown, failure propagation, and interruption at the benchmark boundary; it is not required to perform the CPU parallelism.Do not split individual parameter axes. Beam coordinates, protected elites, the archive, and objective-specific selection are stateful and order-sensitive. Split independent candidate evaluations into batches instead, while keeping beam/archive/cache updates and deterministic tie-breaking in the main thread.
Additional Acceptance Criteria
max(1, availableParallelism - 1)workers, with an environment override and a serial fallback for one worker or unsupported runtimes.Missing Router-Model And Search Optimizations
The earlier design review also proposed the following items. They were not yet captured above and are added here so the issue covers the complete optimization path rather than only execution parallelism and promotion evidence.
w_c(q) = b_c * exp(beta_c dot evidence(q)), followed by the existing scale normalization. Record the same holdout, complexity, and stability diagnostics. The current product of score, geometry, coverage, agreement, identifier, and length factors can amplify interactions; the log-linear form provides an additive, regularizable comparison.Rank-Sensitive Quality Metric: NDCG
NDCG@5,NDCG@10,NDCG@20, andNDCG@50to the benchmark measurements and quality summaries.NDCG(Normalized Discounted Cumulative Gain) complements Recall@K and MRR by measuring the ordering of all relevant chunks within the cutoff rather than only whether a target was found.file + symbollabels. At minimum, a chunk matching one or more resolved gold targets receives binary relevance; if multi-target questions are retained, document whether gain is capped at one or proportional to the number of distinct gold targets represented by that chunk.Direct Retrieval Objective
NDCG@5the primary rank-sensitive optimization metric for thedirectobjective because Production returns five results by default and direct retrieval exposes the fused order to the user. ReportNDCG@10andNDCG@20alongside it for larger requested result sets.Recall@20/Recall@50andContextRecall@4096as hard coverage and context-budget guardrails for direct retrieval. NDCG must not improve the first few ranks by silently losing additional gold symbols or useful context.file + symboltargets, NDCG additionally rewards placing all relevant chunks early.reranker-top20andreranker-top50primarily recall-oriented because their purpose is candidate coverage before a later reranker. Use their NDCG values diagnostically unless a reranker-specific ordering objective is explicitly introduced.Recall@5/Recall@10first) withNDCG@5first under identical folds, profiles, guardrails, and seeds. Select the objective from excluded-fold evidence rather than fit-all quality.