Repository navigation
fix(benchmark): align metric comparisons with their scale - #5887
Conversation
Signed-off-by: Jim-jimu <49069997+bmh201708@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; declaration_source=runtime_reported; model=gpt-6.1-sol; provider=OpenAI; reasoning_effort=xhigh; execution_observation_id=c430c8f62a9d1b345da84c47d35504743748d25e7a2b16e11bcde076d5b7e9d9.
Exact-head review: APPROVE
Head: 3eb5137
动机
维护者在比较同一测试用例的两个运行结果时,需要按通过比例识别改善或退步。
例如 80/100 → 90/200,旧实现按通过数量增加 10 判为改善,实际通过率从 80% 降到 45%;错误方向和排序会把维护者带到错误的分析对象。
本次验证确认比例比较按下降 35 个百分点解释,最大差异排序和页面显示同样使用比例;原始数量差仍保留供追溯。
本 PR 只修复已有指标比较与展示,不修改评分器、运行任务、已记录原始指标或上传权限。
改动思路
一个已有 benchmark reducer 统一判断可比较性与比较尺度,board、study 和页面消费其派生结果,避免三处各算一套方向。
当前 PR 同时覆盖 Python reducer、真实 CLI 输出、TS 数据解码和已有 Cases 页面,不新增运行或配置入口。
最强的反对理由是仅让一个局部输出变绿,而其它真实消费者仍按旧规则执行。因而本次从同一输入的 base/head 对照出发,再追踪所有受影响的共享消费者和恢复入口。未要求本 PR 完成无关的语言迁移、部署或上层路线图;原始记录/权限和失败后的当前 owner 都保留。
具体改动
spec_ref docs/architecture/rfcs/benchmark-study-upload-dashboard-v0.md;spec_revision 4d254fb。以下使用该不可变 pre-change 版本的实际章节标识;PR 自己改写后的文档没有被用作独立 oracle。
4. Dashboard projection:implemented;Every aggregate names its denominator; formal comparisons include only rows accepted by existing matched-pair or factorial contracts. 验证:303 benchmark tests and same frozen base/head input; real upload → study-dashboard CLI。
4.1 Benchmark-specific metric mapping:implemented;Other benchmark families keep their native metric names, units, directions and success thresholds. 验证:higher/lower direction, scalar/ratio, equal rates, unavailable auxiliary metric, fixed factorial total cases。
5. Privacy and provider boundary:implemented;Installing the capability or rendering a dashboard grants no provider permissions; raw/private material remains outside the contract. 验证:303 tests, 12-file boundary scan, real local synthetic CLI packet with no network/write authority。
关键代码讲解
build_benchmark_metric_delta()(factorial_contrast.py:43)先检查单位、方向和分母形状,再计算比例;原始 delta 保留,方向由 delta_rate 决定。不能比较时留下端点和具体原因,不产生数值结论。benchmark_metric_comparison_value()(同文件:104)为已有 board/study 消费者提供同一个有效比较值,避免按 raw 数量排名。- study_projection.py:1245 的 Cases 排序只比较相同尺度的有效主要指标;标量/比例混合时不编造最大差异。
largestContrastLabel()(benchmark-study-page.tsx:67)把率差显示为百分点,正负/flat 颜色与 reducer 方向一致;TS parser 同步允许 unavailable,但拒绝不可用原因与数值结论并存。
全部 12 文件、+517/-33 已按 production、tests、docs、两张截图分类读审。Python 是已有专门 benchmark 计算 provider,并未新增通用控制面 Python 决策源。原始评分、上传 allowlist 和 factorial 资格 owner 保留。
对主干的风险
303 项 benchmark 测试通过、零跳过;TS 数据契约 smoke、浏览器 smoke 和生产 build 通过。相同冻结输入独立跑 base/head:80/100 → 90/200 在 base 为 +10/improved,在 head 为 raw+10、rate−0.35/regressed;另一个 1/1 候选为 +0.20,最大幅度正确选择前者。再通过真实本地 upload JSONL 和 CLI 产生 packet,构建页面显示 −35 pp,点击进入对应 run;390px 无页面横向溢出。无效 packet 显示不可用,恢复有效 packet 可重新读取。浏览器 smoke 的路由数据是合成 mock;这次额外页面的数据来自真实 owning CLI,未将 mock 当 reducer 证明。
零分母、单位/方向不一致、单边分母、同率、lower-is-better、标量兼容、非有限差值、四臂 auxiliary mismatch 和 mixed-scale ranking 都覆盖。semantic advisory 先跑、支持范围内零新词载体;全树 semantic smoke 通过。5 direct + 19 selected 全通过,12 文件边界扫描零命中。premerge 的非零退出是 benchmark-sensitive 维护者审查 hold,保留该合并门禁。未查询或等待 CI。
语义与 CI 对齐
比较方向是已有语义的正确派生;unavailable reason 是 benchmark 局部诊断集合,未创造跨领域 actor/权限协议。字段新增与 TS 解码、README 和回归测试一起交付;全树语义检查没有被缩小或修改。
低分母、元数据不一致和非有限差值已独立检查;四臂主要指标仍要求固定分母,未借本次修复放宽设计资格。未启动真实 benchmark/模型/外部上传,未验证旧独立 dashboard 的混合部署。
初次验证命令使用不存在的测试路径、canary 参数和缺少测试依赖的 interpreter;保留失败诊断后,使用实际测试清单、正确参数和 worktree environment 完成验证,未改变产品或放宽 oracle。
我的整体评价
APPROVE。修复了可复现的比例方向与排序错误,long_horizon 改善后续分析可靠性,user_experience 改善了数据到页面的一致解释;标量兼容和原始指标仍保留。
未来维护梳理已在当前边界落实:Applied bounded shared comparison scale helper。现有兼容路径有真实 reader/历史记录依据,未另造第二套决策源,也未引入未用框架。当前 PR 同时覆盖 Python reducer、真实 CLI 输出、TS 数据解码和已有 Cases 页面,不新增运行或配置入口。 本结论只针对上述精确 head;后续 head/base 或共享 caller 变化须重新判断。独立测试支持审查结论,合并仍由现有仓库维护者门禁和当前授权决定,未执行部署或合并。
English verdict: APPROVE — the exact head corrects the reproduced rate-direction/ranking defect across the affected real callers. Same-input immutable base/head probes distinguish the historical failure from the fix, with native local validation and a built frontend recovery/readback. Original metrics/receipts, scoped authority and legitimate recovery remain intact. 低分母、元数据不一致和非有限差值已独立检查;四臂主要指标仍要求固定分母,未借本次修复放宽设计资格。未启动真实 benchmark/模型/外部上传,未验证旧独立 dashboard 的混合部署。
Goal And Delivered Outcome
50/100 → 60/200previously reported improvement fromdelta=10; it now reports regression fromdelta_rate=-0.2, displayed as-20 pp. Against50/100, the dashboard now ranks9/10(+40 pp) ahead of600/1000(+10 pp). The regression and browser checks below cover these cases.loopx-project/loopx:main.Implemented against
study_projection.py,experiment_board.pytest_ratio_comparisons_round_trip_and_rank_on_the_same_scale; primary incompatibility regressionsfactorial_contrast.pyScope And Continuation
deltaand raw interaction differences retain their original meaning. Ratio labels use percentage points in Markdown and the dashboard.comparison_unavailable_reason, without numeric conclusions. An unavailable primary metric excludes the comparison; an unavailable auxiliary metric preserves primary eligibility and is omitted from paired binary transition counts. Consumers of invalid comparisons must handle this unavailable form. Both-missing optional metadata remains supported for historical board rows. The existing fixed-total factorial primary-metric contract remains enforced; mixed scalar/rate comparisons are not magnitude-ranked together.Validation
3eb513743ac629609f2bd58e919ec02261be4c47; baseline4d254fb6f5cca703c31a18d87488bf489ca486ef. The benchmark suite was rerun on the committed head. Build/browser/static checks used the same source content before commit. The broad Python run preceded the final unavailable-interaction guard and Markdown fallback; affected benchmark tests were rerun afterward.regression_parityunitpython -m pytest -q tests/capabilities/test_benchmark*.py: 303 passed. Covers direction in both optimization senses, equal rates, incompatible metadata, unavailable auxiliary metrics, factorial qualification, and scalar compatibility.real_entrypointbenchmark study-dashboard, and verifies API/CLI parity and rate-based selection.integrationsmoke:benchmark-study, andsmoke:benchmark-study-browser. A separate built-browser check rendered real CLI-generated packets via a substituted local JSON response; it did not use an external backend.statictypecheck:control-plane; semantic diff advisory; diff hygiene; isolated public-boundary scan (3,591 files).integrationunitpython -m pytest -q: 1,728 passed, 124 skipped, 2 failed before interruption; not a complete suite pass. Both failures reproduce on clean baseline: source-session registry-loader allowlist drift, and top-level module budget 148 versus 147.integrationtest:control-plane: 4,174 passed, 32 skipped, 1 failed. The SQLite capacity rehearsal fails atmatched_fill; the same test fails on baseline with Node 22.22.3 / SQLite 3.51.3.integrationquota_should_run/small/jsonemits 21,871 characters against a 20,000 ceiling. Baseline reproduces the same result.See validation disclosure guidance.
Frontend / Visual Evidence
Type of Change
LoopX Area
Technical Direction
Shared-authority RFC fixture impact
Boundary Checklist
none.Signed-off-bytrailer (git commit -s).