Repository navigation
feat(pr-review): reuse repository-owned review experience with Reward Memory - #5956
Conversation
…memory Signed-off-by: huangruiteng <huangrt01@163.com>
…ent evidence Signed-off-by: huangruiteng <huangrt01@163.com>
Signed-off-by: huangruiteng <huangrt01@163.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent · gpt-6.1-sol · OpenAI · runtime_reported · reasoning_effort=xhigh
阻塞发现:R1 [P2] 真实 normalized PR packet 的 changed paths 未进入检索;R2 [P2] 安装 skill 的既有 180 行 native smoke 从 base pass 退化为 head 184 行 fail。下面评审覆盖整个 PR,未把经验投递等同于模型采纳或正确 verdict。
动机
启用既有 memory recall 的 PR 审查者,希望在当前评审中收到适用的公开经验,同时仍按当前契约独立判断。 例如审查标题为 Adjust wording、改动 recovery RFC 的 PR:预期可按修改路径召回恢复设计经验;当前真实 CLI 把文件归一化为 key_files,新增 reader 却读取 files,因此只用标题检索,返回空 guidance。
独立复现确认:常规标题且相关修改路径时,原生 PR packet 的经验召回仍为空;同一输入只向 reader 目前读取的字段提供这条路径,现有文件 reader/SDK 立即返回一条经验,定位到路径接续缺陷。 本 PR 不默认启用 memory、不改变历史 verdict、provider ranking、merge 或额度权限,也不证明经验被模型采纳或提升 review 质量。 当前 changed-path 检索接续和安装 skill 的现有预算回归未修;同 frame/head/model/budget 的模型采纳与质量对照保留后续 pilot。
改动思路
复用既有 memory SDK、typed consumption 与 review 配置是合适的;新 Python 模块只做公开文件和 BM25 provider IO,不应另建 utility 或权限判断源。 当前 PR 的有用切片是显式 opt-in 的公开上下文投递;先修 canonical changed-path 接续和 skill 预算,模型采纳及因果质量对照保留 §9.1 后续 pilot。
最强反对理由是单次维护者纠正可能被升格成普遍 gold label,或让模型继承历史判决;本版明确区分 retrieval、delivery、application、outcome、utility,source 不带奖励标签,不新增排名/执行 owner,因此 bounded readonly reader 的位置和默认策略成立。只留文档不能进入当前调用;另造 evaluator/utility 框架没有必要。不能要求尚未测量的因果质量实验先完成,也不能因 reader 测试通过就认为当前路径可靠。
独立 spec_ref docs/architecture/rfcs/post-outcome-memory-utility-attribution-v0.md,spec_revision 4e9a4fb04a62dca72462d071e91fba3b9e0b2f7b:§3 五种事实分离;§5 原权限、主工作和 provider mutation 边界保留;§9 Stage 4 controlled pilot deferred;§10 ranking 前质量/成本/错误衰减等 qualification 不在当前 context-only 交付范围。新增 §9.1 是后续 pilot 提案,不是已通过的独立验收。
具体改动
Exact head 626185b31d653a34a14d31d798b37f8f4e539127,base 4e9a4fb04a62dca72462d071e91fba3b9e0b2f7b;完整18路径 +725/-5。新增178行 repository_experience.py 作为 review owner 下的只读公开文件/BM25 provider,限制24文件、每文件16K、返回最多3条,调用既有 schema qualifier 和 typed SDK,准确记录 context delivery,semantic_disposition/utility 仍未知。CLI、Markdown 和条件式安装 skill 消费该行;package-data 包含公开经验;中英文操作文档写明原配置 opt-in、readback、disable 与权限边界。
公开 #5944 三份历史 review 和 #5940 issue 已直接读取核对:API commit 与旧 review 正文声明的 head 不同;第二份旧批准当前已 DISMISSED,第三份为维护者引导的设计修改请求。经验正确保留 frame/来源差别,不能从这段历史推断同 head 因果对照、原模型采用过 memory 或谁恒定正确。新的质量实验只是计划。
R1 [P2] repository_experience.py:160–161:消费 canonical changed paths。 build_pr_review_packet 把实际文件放入 key_files(pr_review.py:978),但新 query 读取 item.get("files", []),真实 CLI 没有这个字段。独立用真实 CLI normalization、当前公开文件、BM25 和 SDK 复现:标题 Adjust wording、路径 docs/architecture/rfcs/recovery.md、相同 head/配置/source,实际是 empty/provider_returned_no_items、guidance 空;只把同一路径传入 reader 目前读取的字段的对照立刻得到 context_delivered 和一条经验。14个现有 focused tests 都可通过,因为标题中的 recovery 已足够匹配。请直接读 canonical key_files 或有界 normalized path projection;加真实 normalized CLI 的普通标题/相关路径正例与无关路径负例,不要求作者改标题。
R2 [P2] skills/loopx-pr-review/SKILL.md:77–80:恢复现有安装 skill 预算。 不可变 base 的同一 examples/pr-review-command-smoke.py 通过,skill180行;本head184行,在原第124行 assert <=180 失败,四新增行就是本diff。预算保护 host skill 的薄适配职责;这段条件 advice 有价值,但可与既有 frame 指令合并,保留当前 head、应用/效用分离和 authority 的全部含义。若要调整预算,按已有 budget decision 做同工作量价值/成本论证,不能仅提高断言以变绿。
对主干的风险
14个当前 focused tests 通过,覆盖真实 CLI、启用/关闭、错 agent/未配 surface、no-action/readiness、坏配置/源变更/不合格经验、readonly/public rendering。独立生产路径反例证明 R1,原生 required smoke base/head 证明 R2;两项均为本PR回归。原生premerge5 direct通过、19 selected中17通过;初次 semantic 检查因这个新验证工作区缺 TypeScript dev dependency 失败,执行原 npm ci --ignore-scripts 后,相同不可变 base/head 全树 semantic smoke 均通过,故是已恢复的环境前提,不归为语义代码缺陷。Ruff通过,advisory无支持的新carrier(动态构造未覆盖);其它风险、public boundary 与 maintainability检查通过。未查询、轮询或等待CI;未自行安装 wheel,也未测真实模型采纳/质量。
default-off 保持:只有原实验资格、automatic_recall 和显式 review surface 都满足才调用source;安装文件、可用 provider 和 docs 不构成启用。条件 skill 仅在行存在时适用。Reader不能写memory/provider、选择Todo、消耗其它额度、继承merge权限;当前经验是 guidance,existing machine evidence义务仍独立 enforced。公开 recovery 词汇限定在provider-owned案例,不侵入generic quota/settlement;没有新增substring denylist或第二typed决策源。
我的整体评价
REQUEST_CHANGES。机制归属与规模本身合理,当前 source/context 切片不应被迫扩展成 utility framework;但真实 caller 丢失当前 paths、安装 skill 预算回归使本head的用户结果 not_yet_proven。请在同一有界修改中修复并重跑上面的判别用例。
有限未来重构检查落在 normalized packet/provider seam:保留单一 canonical path 形状,避免 raw GH 字段和 renderer knowledge 重复;安装 skill 压缩当前说明,不减原义。模型采纳和同 frame/head/model/budget 的质量对照仍留给已有 pilot,不创建新的并行任务或将这次修复视为因果资格。
English verdict: REQUEST_CHANGES — 626185b; the real normalized PR caller emits key_files but the new query reads files, so a generic-title/relevant-path review falsely receives no advice despite passing title-based tests. The unchanged native skill-budget smoke also regresses from base180/pass to head184/fail. Reuse and default-off readonly boundaries are sound; repair the canonical path handoff and thin the conditional skill before re-review. Context delivery, semantic adoption and causal quality remain separate.
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh
动机
长期评审不同领域 PR 的 Agent,需要记住哪些检查曾漏掉真实用户结果,同时判断这条经验是否适用于当前改动。 以前只在案例文档里保存评审分歧;现在明确开启的 Agent 能在当前提交的评审 packet 中收到适用条件、验证方法和停止条件。普通标题也能根据真实文件路径召回。
真实 CLI 与独立安装 wheel 都在普通标题、recovery 路径时投递一条建议;无关路径不投递。关闭能力的完整 packet 与不可变基线相同。 不授予评审、合并或跨 Agent 权限,不把历史人工指引设为金标准,也不宣称长期评审质量或成本已改善。 自动采用、held-out 误阻断率、净质量、成本和人工注意力仍需现有 RFC Stage4 的受控试验;此 PR 只完成可用的上下文切片。
改动思路
The existing typed Reward Memory owner governs qualification, scope and context-delivery semantics. A local read-only file/lexical adapter makes Git advice usable without a parallel store or verdict authority. This PR completes bounded repository advice at a real CLI/installed packet. Automatic model adoption and controlled quality/cost utility remain the existing RFC Stage4 gap.
独立依据:docs/architecture/rfcs/post-outcome-memory-utility-attribution-v0.md,不可变修订 4e9a4fb04a62dca72462d071e91fba3b9e0b2f7b。逐项判断:§3 分离召回、采用、结果与效用;§5 保持权限边界;§9-Stage1 不将上下文投递归因为效用;§9-Stage4 的 held-out/成本/误阻断 pilot 保留为现有后续,不能由新增 §9.1 自证完成。
最强反对理由是只有一个案例,词汇相关不一定有帮助,也可能诱发误阻断。明确 opt-in、最多三条、适用与停止条件,加上当前 head 的单独判断,使它成为有界且可回滚的可用切片;无需另建 store、调度器或 verdict owner。
具体改动
- 仓库保存 #5944 的评审口径、actor/API head/body head 和未解决结果;召回 JSON 不包含旧 APPROVE/REQUEST_CHANGES 标签。原始评审 API 的不同 body head 已读回,历史意见不继承为当前验收。
- 原始
626185b31d653a34a14d31d798b37f8f4e539127在普通标题Adjust wording下,会因读取不存在的files而漏掉 recovery 路径;原 14 测试仍可通过。已用规范化key_files修复,并加生产 CLI 的相关路径正例、无关路径负例。当前 source16项及独立 wheel 实际 CLI 都通过。 - 原 smoke 明确失败在184行。已整合现有 review-frame 指令与条件建议,保留 unverified/metadata/CI 限制、真实反例执行、Reviewer/body-marker 来源、先 spec 后 diff、当前 head 的采用/拒绝/不适用,以及不继承 verdict、不混淆投递/语义应用/效用。当前180行,原预算没有提高。
- 11种 off/未资格化/错 surface/其它仓库/inactive/readiness 条件保持整个 packet;不可变 base 与当前 head 的同一 fixture、仅固定时钟的完整 packet 相等。无关路径为空;坏源或源在读回前变化仍为空,不形成用户 gate。SDK/TS 仍记录 delivery=true、semantic_disposition=null、utility=false。
验证:303项相关 Python 测试;10项既有 TS 决策测试;Ruff、focused mypy、docs-governance、Chat bundle/wheel、完整 semantic smoke;原生 canary 5项直接检查+19项风险检查全部通过,manual holds0,精确 quality receipt cqr_f5814cd6c32810f50267 有效。没有查询、轮询或等待 CI。
语义与 CI 对齐
沿用既有 procedural experience、context delivery 和 TypeScript decision owner,没有增加持久分类或授予操作权限。changed-diff advisory 与完整 semantic smoke 均通过。按当前评审契约不等待 CI;以上判断来自明确执行的本机验证,不把未查询的 CI 当作已通过。
对主干的风险
最终 head 240f10e7eb4463bbd4e94101d98525ef8679b24e,18文件完整 PR 已复核,已集成最新已抓取 main 的 5f51559dc93b43d023a9c915c52ec1dc409e3e65;main 历史没有作为本 PR 新增范围。所有5个本 PR commit 均有 DCO。公开增量凭据/私有路径扫描通过。
BM25 只提供相关候选,不能证明适用性;全量模型采用、真实 interactive editor 操作、长期成本/质量/人工注意力未在本 PR 验证。文档中的 #5944 是有限历史证据,不是普遍 gold label。其它 Goal、配置过期/不可用、provider 写入、权限与工作选择不由这个 reader 接管。显式停用 automatic_recall 或移除 review surface 可回滚,无持久格式迁移。
我的整体评价
两个原 blocker 均已在真实入口验证修复,没有剩余阻塞性发现。长期方向有结构性的正向价值:经验终于可到达工作中的 reviewer,普通标题不需要配合特定词,同时仍允许拒绝不适用经验;尚不能从投递成功或测试数推断净效用。未来改动准备已落实为复用规范化 packet seam 和既有 TS owner,无需附加框架。
This PR completes bounded repository advice at a real CLI/installed packet. Automatic model adoption and controlled quality/cost utility remain the existing RFC Stage4 gap.
English verdict: APPROVE
|
Review-frame alignment and merge decision for The independent pre-change specification is post-outcome-memory-utility-attribution-v0.md. Criteria §3, §5 and §9-Stage1 are satisfied for this bounded context-delivery slice; §9-Stage4 controlled adoption/quality/cost attribution remains explicitly deferred in that existing RFC. The newly added §9.1 is assessed implementation scope, not an independent acceptance oracle. Both original blockers are resolved: the production path query now uses canonical normalized The maintainer explicitly authorized repairing, self-merging and upgrading this PR in the current request. This is the separate authority for admin bypass; the read-only readiness packet and successful tests do not themselves grant it. Merge only after native exact-head readiness returns ready for this unchanged head. The bounded future-facing pass reused the canonical normalization seam and existing typed decision owner; no additional framework was introduced. |
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh
动机
长期评审不同领域 PR 的 Agent,需要记住哪些检查曾漏掉真实用户结果,同时判断这条经验是否适用于当前改动。 以前只在案例文档里保存评审分歧;现在明确开启的 Agent 能在当前提交的评审 packet 中收到适用条件、验证方法和停止条件。普通标题也能根据真实文件路径召回。
真实 CLI 与独立安装 wheel 都在普通标题、recovery 路径时投递一条建议;无关路径不投递。关闭能力的完整 packet 与不可变基线相同。 不授予评审、合并或跨 Agent 权限,不把历史人工指引设为金标准,也不宣称长期评审质量或成本已改善。 自动采用、held-out 误阻断率、净质量、成本和人工注意力仍需现有 RFC Stage4 的受控试验;此 PR 只完成可用的上下文切片。
改动思路
The existing typed Reward Memory owner governs qualification, scope and context-delivery semantics. A local read-only file/lexical adapter makes Git advice usable without a parallel store or verdict authority. This PR completes bounded repository advice at a real CLI/installed packet. Automatic model adoption and controlled quality/cost utility remain the existing RFC Stage4 gap.
独立依据:docs/architecture/rfcs/post-outcome-memory-utility-attribution-v0.md,不可变修订 4e9a4fb04a62dca72462d071e91fba3b9e0b2f7b。逐项判断:§3 分离召回、采用、结果与效用;§5 保持权限边界;§9-Stage1 不将上下文投递归因为效用;§9-Stage4 的 held-out/成本/误阻断 pilot 保留为现有后续,不能由新增 §9.1 自证完成。
最强反对理由是只有一个案例,词汇相关不一定有帮助,也可能诱发误阻断。明确 opt-in、最多三条、适用与停止条件,加上当前 head 的单独判断,使它成为有界且可回滚的可用切片;无需另建 store、调度器或 verdict owner。
具体改动
- 仓库保存 #5944 的评审口径、actor/API head/body head 和未解决结果;召回 JSON 不包含旧 APPROVE/REQUEST_CHANGES 标签。原始评审 API 的不同 body head 已读回,历史意见不继承为当前验收。
- 原始
626185b31d653a34a14d31d798b37f8f4e539127在普通标题Adjust wording下,会因读取不存在的files而漏掉 recovery 路径;原 14 测试仍可通过。已用规范化key_files修复,并加生产 CLI 的相关路径正例、无关路径负例。当前 source16项及独立 wheel 实际 CLI 都通过。 - 原 smoke 明确失败在184行。已整合现有 review-frame 指令与条件建议,保留 unverified/metadata/CI 限制、真实反例执行、Reviewer/body-marker 来源、先 spec 后 diff、当前 head 的采用/拒绝/不适用,以及不继承 verdict、不混淆投递/语义应用/效用。当前180行,原预算没有提高。
- 11种 off/未资格化/错 surface/其它仓库/inactive/readiness 条件保持整个 packet;不可变 base 与当前 head 的同一 fixture、仅固定时钟的完整 packet 相等。无关路径为空;坏源或源在读回前变化仍为空,不形成用户 gate。SDK/TS 仍记录 delivery=true、semantic_disposition=null、utility=false。
验证:303项相关 Python 测试;10项既有 TS 决策测试;Ruff、focused mypy、docs-governance、Chat bundle/wheel、完整 semantic smoke;原生 canary 5项直接检查+19项风险检查全部通过,manual holds0,精确 quality receipt cqr_f5814cd6c32810f50267 有效。没有查询、轮询或等待 CI。
语义与 CI 对齐
沿用既有 procedural experience、context delivery 和 TypeScript decision owner,没有增加持久分类或授予操作权限。changed-diff advisory 与完整 semantic smoke 均通过。按当前评审契约不等待 CI;以上判断来自明确执行的本机验证,不把未查询的 CI 当作已通过。
对主干的风险
最终 head 240f10e7eb4463bbd4e94101d98525ef8679b24e,18文件完整 PR 已复核,已集成最新已抓取 main 的 5f51559dc93b43d023a9c915c52ec1dc409e3e65;main 历史没有作为本 PR 新增范围。所有5个本 PR commit 均有 DCO。公开增量凭据/私有路径扫描通过。
BM25 只提供相关候选,不能证明适用性;全量模型采用、真实 interactive editor 操作、长期成本/质量/人工注意力未在本 PR 验证。文档中的 #5944 是有限历史证据,不是普遍 gold label。其它 Goal、配置过期/不可用、provider 写入、权限与工作选择不由这个 reader 接管。显式停用 automatic_recall 或移除 review surface 可回滚,无持久格式迁移。
我的整体评价
两个原 blocker 均已在真实入口验证修复,没有剩余阻塞性发现。长期方向有结构性的正向价值:经验终于可到达工作中的 reviewer,普通标题不需要配合特定词,同时仍允许拒绝不适用经验;尚不能从投递成功或测试数推断净效用。未来改动准备已落实为复用规范化 packet seam 和既有 TS owner,无需附加框架。
This PR completes bounded repository advice at a real CLI/installed packet. Automatic model adoption and controlled quality/cost utility remain the existing RFC Stage4 gap.
English verdict: APPROVE
This is the same model-authored self-review recorded through the already-authorized PR owner account, not a second independent reviewer or a human review. Maintainer authority for this PR explicitly covers self-merge/admin bypass.
Goal And Delivered Outcome
PR #5944 received technical approvals and a later maintainer-directed Request changes that made user value and delivery order explicit. That comparison had no reusable repository experience or path into an enabled review Agent's context.
This change preserves the public comparison and a distilled procedural experience in Git, then delivers matching advice through the existing PR-review CLI and Reward Memory decision boundary. It also refines roadmap S6/S11 and the existing utility RFC to separate this first usable phase from later evidence of improved reviews.
main. This does not close the recovery design issue.Author Declaration
procedural_experience_contract_v0, decision-consumption contract and post-outcome utility RFC at base4e9a4fb04a62dca72462d071e91fba3b9e0b2f7b.pr_review_queue/experiences/pr-5944-review-frame.mdexperiences/loopx-project/loopx/pr-5944-v1.jsonrepository_experience.py; existing Reward Memory/TS decision ownersSelf-check: full changed-source read, public-boundary scan, focused runtime/contract tests, documentation checks, packaged frontend build and isolated installed-wheel readback. Historical verdict labels remain in the comparison Markdown; retrieval loads only the qualified JSON procedure.
Scope And Continuation
pull_request_review.reviewsurface. The reader uses only the selected repository's bundled sources, respects the configured result limit up to three, verifies exact readback and performs no external provider call/write.Validation
626185b31d653a34a14d31d798b37f8f4e539127.8133093; all 14 new delivery/off/failure/CLI cases rerun on the final head after a behavior-preserving identifier renameloopx checkpublic-boundary scan; final full-tree vocabulary smoke passed with the original budgetsCoverage: the changed backend is a read-only filesystem adapter over existing SDK owners, with real production CLI and installed package evidence. Provider enablement/preflight behavior is unchanged and retains its existing tests. Runtime output distinguishes delivered context from semantic adoption and utility. No PostgreSQL authority-store path changes.
Frontend / Visual Evidence
Type of Change
LoopX Area
Technical Direction
S6/S11 review learning through existing Reward Memory; S8/S12 PR-review caller. Runtime changes remain for maintainer review and merge.