Skip to content

chore(release): prepare 1.3.0 long-horizon semantic control quality - #5682

Merged
loopx-agent merged 3 commits into
mainfrom
codex/release-1.3.0
Oct 5, 2026
Merged

loopx-agent merged 3 commits into
mainfrom
codex/release-1.3.0

Conversation

@loopx-agent

@loopx-agent loopx-agent commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Prepare LoopX 1.3.0: long-horizon work and stronger semantic control

Prepare the 1.3.0 version and documentation contracts around the latest long-horizon benchmark diagnosis and Astra-assisted semantic review. The existing regression repair stays in #5533; it is reviewed at 5b177a2 and awaits maintainer integration. This release PR is reviewed at 5c80a25.

The preparation source passes version/help/book contracts, packaged Chat source verification and the corrected locale contract. Its initial CLI-budget and architecture checks retain independently reproduced frozen-main failures; #5533 supplies the separately reviewed repairs. On that repair head, 19 selected risk checks plus five direct checks, 143 main-adoption tests, 52 configuration/API tests, 143 budget tests, 96 matched output measurements and three real host installers pass. One existing Pydantic warning remains.

Publication remains held for maintainer integration and final frozen-source qualification. The release includes a preliminary isolated DS V4.1 Flash/high native Goal run, not a final model-portfolio qualification. No benchmark jobs were launched. The preparation PR merge gate is ready. The regression PR has an APPROVED exact-head review and clear closeout, but its merge gate remains blocked by GitHub REVIEW_REQUIRED. These read-only gates grant no merge authority.

Complete bilingual candidate release notes

LoopX v1.3.0 — Long-horizon work, stronger semantic control

Release candidate — publication held until final exact-source qualification and maintainer review are complete.

Personal illustrated candidate guide — verified user owner, text, two image blocks and exported rendering; remains a candidate until release artifacts are verified.

Long-horizon benchmark diagnosis and Astra-assisted engineering shape this release's theme: retain the right work and evidence, make recovery actionable, and preserve the authority behind every next step. The latest research blog explains the failures that motivated the work. Astra describes the development and semantic-review effort; the reported EdgeBench worker experiment uses gpt-6.1-sol / xhigh.

Release Decision

Who should upgrade: Operators of continuing Codex work, evidence-driven replanning, and workspace Chat should consider this version once it is published. Users satisfied with v1.2.4 can remain there while the candidate is qualified.

What this release solves: Useful history could be crowded out by repeated observations; an already-known replan could demand a failed round trip; task steps, permission failures and settlement hints could point to the wrong recovery. This release collects bounded fixes to those paths, alongside clearer workspace Chat and explicit configuration recovery.

Breaking changes: No intentional breaking migration. Ordinary host-declared project Chat now defaults to workspace_write; choose workspace_read explicitly when needed. Existing read-only App bindings retain their grant. New generated settlement commands carry correctly placed global --format json; direct CLI defaults and stored receipts are unchanged. Managed execution names deepseek-flash; explicit historical model settings remain respected. Canonical new-Goal creation and Explore execution remain opt-in. Existing Goals are not migrated by a device preference.

How to verify: Expect loopx 1.3.0, a healthy owning installation, and a current scoped status/diagnostic readback after upgrading. Investigate unavailable or blocked results before starting work.

loopx --version
loopx doctor
loopx --format json status
loopx diagnose --goal-id "$GOAL_ID"

Contributors: Release maintainer @huangruiteng, with maintainer automation through @loopx-agent. The community contributions listed below are derived from v1.2.4 → candidate merged PRs.

State Kernel & Control Plane

  • Continue the correct task: Next Action writeback binds to the selected Todo, Agent and current basis; accountable Goal-level writes remain valid. Old or unrelated steps cannot silently take over the next turn. #5531, #5588

  • Use evidence before repeating a route: Dense typed replan context retains outcome and route diversity, with exact omitted-history recovery. Explicit successor selection can finish the known replan admission in one CLI call; genuinely fresh vision gaps can rearm it. #5536, #5624, #5628, #5646

  • Settle with the original authority: Checkpoint recovery, unchanged-artifact/negative-evidence guidance and vision closeout follow the canonical settlement plan. Full and compact packets preserve conditional steps and identities; generated JSON commands run as returned. #5573, #5629, #5633, #5640, #5636, #5667, #5672

  • Separate denial from failure: Host permission diagnostics retain actionable recovery; isolated edits preserve the complete evidence basis. Delegated stop acknowledges the request separately from proven native-process drain and lease settlement. #5644, #5587, #5308

Capabilities & Workflows

  • Explore Harness becomes usable in an ordinary turn: Evidence-only and planning modes share the existing owner and a turn-start context; finding titles survive projection. Configuration proves availability, while reading and changed decisions still need evidence. #5610, #5658

  • Work in an ordinary workspace conversation: The App Scope picker opens a granted workspace through the existing composer without synthesizing a Goal or borrowing portfolio authority. Writes stay inside the host grant; revoked or changed contexts reject new work. #5540, #5555

  • Recover configuration explicitly: Export and verify complete stored configuration, then restore to an isolated checkpoint. Device settings can opt future empty Goals into canonical File/SQLite authority and a frozen soft-claim/hard-lease policy; existing Goals retain their owner. #5557, #5569

  • Inspect public GitHub evidence: A SHA-pinned anonymous provider retrieves bounded sources; parent admission and downstream ledger readback remain separate. CLI/provider and conversation readback are shipped; complete frontend initiation/admission remains a staged boundary. #5459

  • Keep desktop updates and recovery durable: Update-channel pointers reject rollback, and interrupted backup rotation can recover through the existing owner. #5645, #5647

Quality & Testing

  • Test semantics and retire duplicate owners: Native child receipts preserve resumed result correlation; stop fixtures retain real drain obligations. Redundant smokes, Python Todo transition/admission adapters and an unused shadow producer are retired; four document versions and the Lark visibility pair each use their defining owner. #5455, #5539, #5613, #5621, #5597, #5664, #5666, #5670

  • Reduce bounded preflight cost: The affected Turn/MCP path reuses typed projections; checkpoint context resolves in one TypeScript request. These measured engineering slices do not close sustained-operation or model-efficiency acceptance. #5283, #5585

Benchmarks & Integrations

  • Diagnose long-horizon work with controlled profiles: Native EdgeBench supports official, single-call, native Goal, heartbeat-resume and heartbeat-Explore workers. Trial-wide deadlines, terminal visualizer results and runnable successor admission use their intended lifecycle; feedback policy is distinguished from evaluator isolation. #5591, #5600, #5618, #5617, #5635

  • Private Lark Chat composes with the same workspace owner: Explicit App/source bindings support ordinary Chat, confirmed steward commissions, scoped status/help and authorized attached-Agent selection. One listener and durable Core requests preserve original audiences. Installed/live Lark and mobile qualification remain separate. #5541, #5542, #5544, #5546, #5550, #5637

  • More explicit provider observations: The managed model alias names DeepSeek V4.1 Flash, with operator-configurable Ark rows. Finance 0.8.3 distinguishes producer accuracy, encoded periods and parent-declared economic periods; the earlier position guard reads filled protection separately from trade authority. #5561, #5631, #5521

Documentation & Compatibility

  • Explain what the experiment establishes: The latest bilingual blog separates prompt confounds, task/judge accounting, genuine strategy failure, control-protocol recovery and context cost. It retains failed samples and open causal questions. No matched v1.2.4-versus-v1.3.0 outcome trial has been qualified; this release makes no benchmark-score, statistical superiority or long-horizon success-rate claim. #5649, #5635

  • Keep entrypoints current: Wheels ship project-scoped skill sources; current authority-format reads avoid migration locks; refresh authoring help explains the 1,200-character step limit and in-flight versus replan closeout. PR review supports explicit additional owner accounts and per-Agent direction without granting merge authority. #5560, #5661, #5653, #5672, #5608, #5656

  • Make Chat failures observable: Session input and delivery-time payload conflicts remain explicit without changing replay identity; canonical SQLite alignment keeps its original basis. The DSH plug-in is independently released as 0.1.1-beta.6 with a qualified 0.2.0-rc.2 host boundary; install/readback/removal is documented below. #5554, #5686, #5589, #5687

  • Route ordinary work through its actual context: Native-steward provenance reaches context delivery; owner discovery defaults to the registered portfolio. Ordinary project prompts stay focused on workspace work, local Goal admission reuses configured defaults and scope matching, and blocked research plans return structured errors. #5615, #5683, #5659, #5676, #5681

Community Contributors

  • @Inference1 — first-time external contributor: four document-version owners and one Lark visibility owner (#5597, #5621).
  • @mikamikasuki — first-time external contributor: refresh no-write mutation regression, exact source-session observation and retirement of completed contributor/RFC checkpoints (#5574, #5576, #5577, #5579, #5595, #5601, #5613).
  • @jackie-cqz — external contributor: native child result attribution, public GitHub evidence, shared validation, Windows paths/configuration publication and Chat delivery retry (#5455, #5459, #5526, #5602, #5603, #5606).
  • @Duang777 — external contributor: Goal recreation fences, validated inbox read commits and accountable Goal-level writeback, desktop update-pointer rollback protection and interrupted backup rotation recovery (#5338, #5563, #5588, #5645, #5647).
  • @hhyykk — external contributor: one-request TypeScript checkpoint context (#5585).
  • @catwithtudou — external contributor: repository identity normalization for leases and preserved diagnostic chronology (#5572, #5584).
  • @songoow — external contributor: governed delegation stop, native fixture cleanup, exact-target CI recovery and focused smoke governance (#5308, #5534, #5535, #5539, #5567).
  • @BigDataDZ — external contributor: the committed goal-direction F2 revision-drift fixture (#5549).
  • @maxliux5 — external contributor: personal follow-through documentation consolidated into the existing manager profile (#5384).
  • @AronSwan — external contributor: explicit session input failures and delivery-time payload/transport conflicts, preserving replay identity (#5554).

Optional Capability Activation & Use

Explore Harness

Activation: In Goal settings → Capability Center choose evidence or planning. CLI planning opt-in is shown below; the default is off.

Validation: Read turn context and summary for the same Goal/Agent; recorded nodes alone do not prove adoption.

Disable / rollback: Set --explore-mode off --execute; retained evidence stays readable.

Authority boundary: Analysis and planning grant no worker spawning, claims, quota spend, merge or external publication.

Docs: Versioned guide

loopx configure-goal --goal-id "$GOAL_ID" --explore-mode planning --explore-harness-profile adaptive-resilient --execute
loopx explore turn-context --goal-id "$GOAL_ID" --agent-id "$AGENT_ID"
loopx explore summary --goal-id "$GOAL_ID"
loopx configure-goal --goal-id "$GOAL_ID" --explore-mode off --execute

TurnEnvelope and captured decisions

Activation: Opt in per guard invocation with --turn-envelope; add --decision-output-dir only with an explicit Turn id and a new directory whose parent exists.

Validation: Read the returned capture and verify Goal/Agent/Turn, original source hash and ok; observe rejection and incomplete publication honestly.

Disable / rollback: Omit both options. Delete only no-longer-needed private captures through ordinary file management.

Authority boundary: Saved decisions are private observations, not fresh admission; selection, leases, cancellation and mutation-time checks remain mandatory. Frontend/Lark do not consume these files.

Docs: Versioned guide

loopx --format json quota should-run --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --turn-instance-id "$TURN_ID" --turn-envelope --decision-output-dir ./guard-001
cat ./guard-001/decision.json

Ordinary workspace Chat

Activation: Run loopx chat; in the App steward conversation choose an already granted workspace in Scope. Host-declared roots now default to workspace_write; use the explicit read-only command below when required.

Validation: Read the selected scope, effective grant and returned result in the same conversation. A changed grant creates a new context; old history remains.

Disable / rollback: Restart the service with --project-workspace-grant workspace_read, or revoke the host workspace grant; new work on the old context is rejected.

Authority boundary: Workspace writes follow AGENTS.md and the actual sandbox. No hidden Goal, portfolio grant, peer delegation or external-send authority is created.

Docs: Versioned guide

loopx chat --project-workspace-grant workspace_read --no-open

Owner private Lark conversations

Activation: In Settings → Lark explicitly verify the selected App and personal owner, choose a workspace, executor, grant and ordinary Chat or steward role, then connect. /delegate --tokens N objective requires an original-source confirmation. /agents, /agent TARGET_REF and /project use separately authorized attached-Agent targets.

Validation: In the bound private conversation send /status and /help; verify the App/source, role, workspace grant, original Session/Turn and queue. Inspect the same binding in Settings.

Disable / rollback: Disconnect that exact App binding in Settings → Lark; /stop targets the current ordinary Turn and /stop-commission targets the bound commission. Remove exact Agent target grants separately.

Authority boundary: Personal credentials, source/owner and listener identity remain explicit. Ordinary Chat creates no Goal; commission confirmation does not grant arbitrary writes, automatic heartbeat or canonical task acceptance. Live Lark/mobile qualification is separate.

Docs: Versioned guide

loopx chat --project-workspace-grant workspace_read --no-open
loopx chat --help

Configuration checkpoints

Activation: Settings → Capability Center → Configuration backup and recovery downloads a private checkpoint. CLI export uses a new destination, preview first, then --execute.

Validation: Run configuration-backup verify on the exact file; compare source scope and digest. Integrity verification does not certify privacy.

Disable / rollback: An isolated restored checkpoint can be removed without affecting live settings. Revert any adopted setting through its existing revision-checked editor or machine-config/configure-goal transaction.

Authority boundary: Capture and isolated restore copy no credential store, Host session, grant, provider selection, live registry, fence, lease or scheduler. Backups remain private, including secrets already embedded in configuration.

Docs: Versioned guide

loopx --format json configuration-backup export --goal-id "$GOAL_ID" --output "$NEW_CHECKPOINT_FILE"
loopx --format json configuration-backup export --goal-id "$GOAL_ID" --output "$NEW_CHECKPOINT_FILE" --execute
loopx --format json configuration-backup verify --input "$NEW_CHECKPOINT_FILE"

Canonical new-Goal creation

Activation: In Device defaults → New Goal authority opt in to canonical creation and choose File/SQLite plus soft_claim/hard_lease. The same v1 document uses the existing preview/apply transaction; default remains off.

Validation: Inspect machine settings, bootstrap a new empty project, then read its Todos and native authority receipt; existing Goals are not retargeted.

Disable / rollback: Preview/apply canonical_creation=false to disable future creation, or remove the goal_storage namespace with its exact removal-plan revision. Existing Goals retain storage and fences.

Authority boundary: Storage and execution policy are independent of tools, accounts, network, scheduling and migration authority. Missing authority cannot be recreated as empty by forced bootstrap.

Docs: Versioned guide

Save / 保存为 goal-storage.json:

{"schema_version":"loopx_goal_storage_defaults_v1","new_goal_provider":"sqlite","canonical_creation":true,"new_goal_handoff_mode":"hard_lease"}
loopx machine-config preview --namespace goal_storage --config-json goal-storage.json
loopx machine-config apply --namespace goal_storage --config-json goal-storage.json --expected-plan-revision "$PLAN_REVISION" --execute
loopx machine-config inspect
loopx machine-config remove --namespace goal_storage
#Review removal, then use its returned revision.
loopx machine-config remove --namespace goal_storage --expected-plan-revision "$REMOVAL_REVISION" --execute

Governed delegation stop

Activation: Use an already configured exact binding, registered requester and operation; delegation stop --execute explicitly requests stop. CLI/MCP and the App team surface share that owner.

Validation: delegation read distinguishes acknowledgement, proved native process drain and settlement. Missing supervision remains unknown and requires the returned recovery.

Disable / rollback: Remove the exact operator binding to revoke new execution; use stop on an active operation. A stopped/settled receipt is not a resume grant.

Authority boundary: Do not infer Host exit, lease release or non-execution from absent records. Windows/unsupported stopping refuses before launch-side cancellation effects.

Docs: Versioned guide

loopx delegation stop --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --execution-config "$DELEGATION_CONFIG" --operation-id "$OPERATION_ID" --execute
loopx delegation read --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --execution-config "$DELEGATION_CONFIG" --operation-id "$OPERATION_ID"

PR review queue ownership and direction

Activation: Capability Center configures additional owner logins and forward/reverse direction; Goal CLI can set exact per-Agent direction. Current-session request intake precedes generic queue discovery.

Validation: Inspect configure-goal and the read-only pr-review packet; compare effective accounts, Agent and direction. Queue ownership is separate from GitHub author identity.

Disable / rollback: Use --clear-pr-review-owner-logins --execute, --pr-review-agent-order AGENT=inherit, or --clear-pr-review-configuration --execute. Remove owner_logins from Goal/device settings before downgrading.

Authority boundary: These settings grant no GitHub review, comment, dismissal, merge, cross-Agent write or scheduler authority; review depth and CI policy retain their owner.

Docs: Versioned guide

loopx configure-goal --goal-id "$GOAL_ID" --pr-review-owner-login maintainer --pr-review-agent-order reviewer-a=forward --execute
loopx configure-goal --goal-id "$GOAL_ID"
loopx pr-review --goal-id "$GOAL_ID" --agent-id reviewer-a --repo "$REPOSITORY" --format json
loopx configure-goal --goal-id "$GOAL_ID" --clear-pr-review-configuration --execute

Public GitHub evidence

Activation: Opt in per plan with --public-github and a full commit SHA source; execute --execute authorizes anonymous source reads.

Validation: Read back the exact plan and execution receipt. Parent admit/reject and downstream research-ledger coverage are separate explicit operations.

Disable / rollback: Omit --public-github and --execute; no persistent provider switch is installed. Retire evidence only under the admitted downstream-coverage rules.

Authority boundary: No token, cookie, private repository, mutable branch source, raw-page persistence, automatic admission or outbound message permission. Complete frontend initiation remains open.

Docs: Versioned guide

loopx external-evidence plan --public-github --objective "Inspect public source" --user-activity "Choose a source" --decision "Whether a literal is present" --evidence-kind literal_match --source "$SHA_PINNED_PUBLIC_URL" --search-term LoopX --format json > plan.json
loopx external-evidence execute --plan-json plan.json --execute --format json > execution.json
loopx external-evidence readback --plan-json plan.json --receipt-json execution.json

Managed DeepSeek model selection

Activation: Managed execution now names deepseek-flash@high. Explicit LOOPX_TURN_MODEL/DSH_MODEL or --dsh-model overrides are still honored; configure credentials through the owning provider.

Validation: Read turn run-once --help and the returned runtime profile/actual provider identity; a model alias alone is not live qualification.

Disable / rollback: Set LOOPX_TURN_MODEL to the previous explicitly selected model and restart the owning host. Do not change an existing session silently; omitted configuration restores the shipped managed default.

Authority boundary: A model setting grants no API credential, paid invocation, workspace, permission, Goal or provider promotion authority. CPA routing remains separately installed/operator-owned.

Docs: Versioned guide

LOOPX_TURN_MODEL=deepseek-flash loopx turn run-once --help
#Explicit rollback selection for the next owning host invocation:
export LOOPX_TURN_MODEL="$PREVIOUS_MODEL"

Finance evidence assessment

Activation: Finance Value Discovery is a separately installed optional extension, version 0.8.3; use the pinned source package and register/enable its manifest. The additive assess-period input is finance_period_comparison_input_v1.

Validation: Run extension doctor, then assess-period on a reviewed local input. Producer accuracy, period eligibility, source authenticity and investment truth remain distinct.

Disable / rollback: Disable loopx-finance-value-discovery with the command below. Retain private account/position observations outside public research projections.

Authority boundary: The package performs deterministic evidence assessments; it grants no investment advice, order, transfer, signing or execution authority. Missing source/timezone/binding evidence stays explicit.

Docs: Versioned guide

#In a v1.3.0 source checkout:
python3 -m pip install ./packages/loopx-finance-value-discovery
loopx extension install --manifest packages/loopx-finance-value-discovery/extension.toml --execute --format json
loopx extension enable loopx-finance-value-discovery --execute
loopx extension doctor loopx-finance-value-discovery --execute
loopx-finance-value-discovery assess-period --input-json period.json
loopx extension disable loopx-finance-value-discovery --execute

EdgeBench native trials

Activation: Explicit research-only invocation from a v1.3.0 checkout; Linux, Docker, pinned EdgeBench/SForge/Harbor, selected Codex binary and authorized model access are prerequisites. Use the selected worker and feedback profile.

Validation: Run --help before the reviewed trial; read session/model, native terminal result, evaluator completion and isolation receipts. Registration and sampling counts do not prove countable outcomes.

Disable / rollback: Do not start another trial; stop the exact owned trial/controller and TLS relay through their existing cancellation lifecycle. Preserve private logs and incomplete results.

Authority boundary: A release launches no benchmark jobs. The native judge owns task/scoring; blind policy requires separately qualified credential/network/submission isolation. Raw evidence and secrets remain private.

Docs: Versioned guide

python -m benchmark.edgebench.run --help
#Paid execution only after the operator reviews task, pin, isolation and budget:
python -m benchmark.edgebench.run --task "$TASK_ID" --tasks-dir "$TASKS_DIR" --log-dir "$PRIVATE_LOG_DIR" --run-id "$UNIQUE_ATTEMPT" --worker heartbeat-resume --model "$MODEL" --effort xhigh --judge-url "$JUDGE_URL"

DSH LoopX plug-in

Activation: Install the separately published 0.1.1-beta.6 into the web profile below. Check DSH compatibility first; the qualified current host is 0.2.0-rc.2. Loading prepares the isolated CLI/skills and GoalBar; the passive Driver activates only after this exact Session invokes the installed loopx skill.

Validation: Check dsh --version, then resolve the exact live Session binding. Require status=bound and one Goal/Agent pair before using GoalBar. Installed files alone do not establish a live binding.

Disable / rollback: Remove dsh-loopx-plugin from the same profile and restart DSH. To roll back, install a retained previous tarball supported by that host; preserve LoopX state.

Authority boundary: Installation may prepare the isolated CLI and skills, but grants no Goal binding, quota spend, model call or execution by itself. LoopX owns Goal/Todo/Agent decisions; no credentials are bundled. GoalBar and Driver use the authenticated local Session boundary.

Docs: Versioned guide

dsh --version
dsh plugin --profile web add "https://github.com/loopx-project/loopx/releases/download/dsh-loopx-plugin-v0.1.1-beta.6/dsh-loopx-plugin-0.1.1-beta.6.tgz"
dsh --profile web --port 0
#After invoking the installed loopx skill in the exact DSH Session:
loopx --registry .loopx/registry.json --format json resolve-agent-thread --host-surface deepseek-harness-native --thread-id "$DSH_SESSION_ID"
dsh plugin --profile web remove dsh-loopx-plugin

Install / Update

When published, the package version/tag will be 1.3.0 / v1.3.0. Python 3.11+ and Node.js 22.22.3+ remain required. New PyPI installs use:

python3 -m pip install --upgrade loopx
loopx workflow-skills --install
loopx slash-commands --install
loopx doctor

Existing installs preserve their acquisition owner:

loopx update check
loopx update plan
loopx update apply
loopx doctor
loopx extension doctor --all-enabled --execute

Do not install this candidate from a named stable channel yet. After publication, inspect the package/assets, source manifest and update feed independently; an uploaded source package does not prove a signed desktop build or installed runtime activation.

Qualification

Pending release gates: final exact-commit qualification, maintainer integration, published artifacts and final guide readback remain open. Preliminary real-model execution and personal guide ownership/rendering have been verified on the stated candidate source; no prior-source result is relabeled as final-release evidence.

  • Passed on the candidate lineage: release version/readiness contracts, generated help/manpage (after explicit advanced-command classification), bilingual Developer Book checkpoint, required-scope Ruff, Mypy (19 files), packaged Chat build, 42 installed-bundle dashboard command tests, and the eight changed public paths' boundary scan.

  • Regression repair remains in existing #5533, at 5b177a2. Exact-head review is published. On this exact source, 19/19 selected risk checks plus five direct checks, 143 main-adoption tests, 52 configuration/API tests, 143 CLI-budget tests and three real host installers pass; 96 matched base/head output measurements raise no review signals. One existing Pydantic warning is retained. Ruff, 19-file Mypy and an actual Chat rebuild/source verification pass. The 17-path source-invalidation record separates unchanged older evidence from three files changed by main integration. The pinned 91666 diagnostic completed with 16,748 passes, 15 failures, 118 skips and 453 passed subtests. Five stale-bundle failures were independently rebuilt/qualified; the other ten reproduced on pinned main and are repaired in current fixtures/docs with refusal/no-effects assertions retained. This failed run is not relabeled as passing. The repair is not yet integrated into the release source; Maintainer integration and final exact-source qualification remain open; the managed merge gate does not consult CI.

  • Final frozen risk canary: 12 selected / 12 executed, no blocking failures; 5 direct checks passed. CLI output-budget tests: 22 passed. One pre-existing maintainability advisory remains. Earlier failed attempts are retained separately and are superseded only by this unchanged-head rerun.

  • Broad scripts-inclusive Ruff reports six unchanged E402 import-placement findings; the actual CI lint scope passes. Initial no-clone smoke failed during disposable-home cleanup; fresh-clone and update smokes passed separately. Repository-hygiene reports symbolic Basic-auth construction and synthetic private-network fixture literals, not a leaked credential; the changed-path scan is clean.

  • Output budgets preserve useful detail: identical crowded fixtures produce 14,647 JSON characters for Turn planning and 44,126 for diagnosis on both frozen main and candidate. The regression ceilings are calibrated from 14,600 to 15,000 and from 44,000 to 45,000; complete Vision, executable recovery, line limits and semantic/settlement assertions remain. The monitor ceiling remains 2,000 with its existing stable-path fixture. These are presentation-regression budgets, not token, fee or experiment-promotion limits.

  • Preliminary native Codex Goal on candidate 8056cc6 passed with DeepSeek V4.1 Flash / high: two real model Turns, two Todo settlements and independent acceptance. Final exact-source qualification will be repeated. One alternate-provider probe was authentication-blocked; it is not a model-semantic failure. No benchmark trial was launched.

  • Verified personal user ownership and bounded guide history: the last substantive guide covers v1.2.4 and was updated on October 3, 2026. The new candidate guide has verified owner, complete text, two image blocks and nine-page exported rendering. Final released artifacts and guide status still need readback.

中文摘要

主题:长程 benchmark 与 Astra 辅助工程推动语义控制面质量提升。 重点是续接正确工作、让证据进入重规划,以及按原权限恢复和结算。最新 blog 保留实验混杂、评分修复、真实策略失败及开放问题。Astra 是开发与语义审查主线;本文 EdgeBench worker 实验使用的基础模型仍为 gpt-6.1-sol / xhigh。

候选版本,尚未发布。 个人图文候选指南已验证个人归属、正文、两张图与导出排版;正式发布后按实际产物原地更新。

升级决策

**谁需要升级:**需要持续 Codex 工作、证据驱动重规划或普通 workspace Chat 的用户,在正式发布后可考虑升级;当前满足需求的 v1.2.4 用户可等待候选验证完成。

**解决了什么:**早期有效结果被重复观察淹没、已知 replan 仍要求失败重入、任务步骤与权限/结算提示错位;本次集合这些有界修复,并改善 workspace 对话与配置恢复。

**是否有破坏性变更:**无主动 breaking migration。host-declared project Chat 新默认 workspace_write,需只读时显式选择 workspace_read,旧 App 只读 grant 保留。新生成结算命令自带位置正确的 JSON 参数,历史命令与直接 CLI 默认不变。managed 模型命名改为 deepseek-flash,明确旧配置保持优先。canonical 新 Goal 与 Explore 执行保持 opt-in,不迁移已有 Goal。

**如何验证:**正式升级后期望 loopx 1.3.0、健康安装及正确作用域状态。使用上方 loopx --version、loopx doctor、loopx --format json status 与 loopx diagnose --goal-id "$GOAL_ID",遇到 unavailable/blocked 先恢复再工作。

**贡献者:**维护者 @huangruiteng,自动化维护账号 @loopx-agent;以下社区贡献来自 v1.2.4 到候选的真实合并范围。

状态内核与控制面

  • **续接正确任务:**下一步写回绑定 Todo、Agent 与当前依据,同时保留真实的 Goal 级责任写回;旧步骤和无关任务不能覆盖当前工作。#5531, #5588

  • **让证据改变下一步:**重规划保留结果与路线差异,并可恢复被省略历史;显式后继选择可在一次 CLI 调用内完成已知重规划准入,新 vision gap 仍会重新触发。#5536, #5624, #5628, #5646

  • **按原权限结算:**checkpoint 恢复、未改产物/负证据与 vision 收口采用规范计划;完整包和短包保留条件、身份及可直接执行的 JSON 命令。#5573, #5629, #5633, #5640, #5636, #5667, #5672

  • **区分拒绝与故障:**宿主权限错误保留可操作恢复,隔离编辑保留完整依据;委派停止区分已接收、真实进程退出与 lease 结算。#5644, #5587, #5308

能力与工作流

  • **普通轮可使用 Explore Harness:**证据与规划模式复用已有 owner 和轮前入口,finding 标题保留;开启、读取与决策改变仍是独立证据。#5610, #5658

  • **直接处理工作区请求:**App 的 Scope 选择器沿用原 composer,无须创建 Goal;写入受宿主 grant 约束,撤权或上下文变化后拒绝新工作。#5540, #5555

  • **显式恢复配置:**导出并验证完整存储配置,只恢复到隔离 checkpoint;设备设置可为未来空 Goal 选择规范 File/SQLite 与冻结策略,已有 Goal 保持归属。#5557, #5569

  • **读取公开 GitHub 证据:**匿名 provider 按 SHA 取有限来源,父 Agent 采纳与下游账本回读分开;已交付 CLI/provider 和会话回读,完整前端发起/采纳仍为阶段边界。#5459

  • **保持桌面更新与恢复可靠:**更新 channel pointer 防止回退,已中断的备份轮换通过已有 owner 恢复。#5645, #5647

质量与测试

  • **验证语义并退役重复规则:**原生 child 回执保持续接结果归属,停止验证保留真实 drain 义务;删除冗余 smoke、Python Todo 适配与无用 shadow producer,文档版本和 Lark visibility 复用各自 owner。#5455, #5539, #5613, #5621, #5597, #5664, #5666, #5670

  • **减少有界预检开销:**Turn/MCP 复用类型化投影,checkpoint context 合并为一次 TS 请求;这不等于多日稳定性或模型效率验收完成。#5283, #5585

基准与集成

  • **用受控 profile 诊断长程工作:**EdgeBench 接入五种 worker,整场 deadline、终态显示与可运行后继各守生命周期;反馈政策与 evaluator 隔离明确区分。#5591, #5600, #5618, #5617, #5635

  • **私聊 Lark 复用工作区 owner:**明确 App/source 绑定支持普通对话、确认委托、作用域 status/help 与授权 Agent 选择;单 listener 与耐久请求保持原受众,安装/真实飞书和手机验收另行判断。#5541, #5542, #5544, #5546, #5550, #5637

  • **更明确的 provider 观察:**managed 模型使用 DeepSeek V4.1 Flash 名称,Ark 行可由 operator 配置;Finance 0.8.3 区分 producer 精度、编码期间与父调用者声明的经济期间,持仓守护的保护读回不授予交易权限。#5561, #5631, #5521

文档与兼容性

  • **说明实验到底证明了什么:**最新中英 blog 区分提示混杂、评分账本、策略失败、协议恢复和上下文成本,保留失败与因果问题。本次未完成 v1.2.4/v1.3.0 配对 outcome 验证,不声明 benchmark 涨分、统计优越性或长程成功率提升。#5649, #5635

  • **保持入口可用:**wheel 交付项目 skill,当前 authority 格式读取无需迁移锁;refresh help 解释 1,200 字符限制与 in-flight/replan 边界;PR review 支持 owner 列表与 Agent 方向,但不增加合并权限。#5560, #5661, #5653, #5672, #5608, #5656

  • **让 Chat 失败可见:**会话输入和交付时 payload 冲突保持明确,原 replay 身份不变;规范 SQLite alignment 保留原始 basis。DSH 插件独立发布为 0.1.1-beta.6,当前宿主资格边界为 0.2.0-rc.2;启用、读回和移除见下文。#5554, #5686, #5589, #5687

  • **让普通工作使用实际上下文:**native steward 的 provenance 进入上下文交付;owner discovery 默认读取已登记 portfolio。普通项目提示聚焦 workspace 工作,本地 Goal 准入复用既有默认配置与 scope matching,被阻塞的 research plan 返回结构化错误。#5615, #5683, #5659, #5676, #5681

社区贡献者

  • @Inference1 — 首次外部贡献者:四个文档版本 owner 与单一 Lark visibility owner(#5597, #5621)。
  • @mikamikasuki — 首次外部贡献者:refresh 无写入 mutation 回归、精确 source-session 观察及已交付贡献/RFC 检查点退役(#5574, #5576, #5577, #5579, #5595, #5601, #5613)。
  • @jackie-cqz — 外部贡献者:原生 child 结果归属、公开 GitHub 证据、共享验证、Windows 路径/配置发布及 Chat 交付重试(#5455, #5459, #5526, #5602, #5603, #5606)。
  • @Duang777 — 外部贡献者:Goal 重建 fence、校验后的 inbox 读取提交与 Goal 级责任写回、桌面更新指针防回退与中断备份轮换恢复(#5338, #5563, #5588, #5645, #5647)。
  • @hhyykk — 外部贡献者:单次 TS 请求的 checkpoint context(#5585)。
  • @catwithtudou — 外部贡献者:lease 仓库身份归一及诊断时间顺序(#5572, #5584)。
  • @songoow — 外部贡献者:受治理委派停止、原生 fixture 清理、精确目标 CI 恢复与 smoke 治理(#5308, #5534, #5535, #5539, #5567)。
  • @BigDataDZ — 外部贡献者:Goal direction F2 revision-drift fixture(#5549)。
  • @maxliux5 — 外部贡献者:将个人 follow-through 文档归并到已有 manager profile(#5384)。
  • @AronSwan — 外部贡献者:会话输入失败及交付时 payload/transport 冲突的明确处理,保持 replay 身份(#5554)。

可选能力启用与使用

Explore Harness

**启用:**在 Goal 设置 → 能力中心选择 evidence 或 planning;下方命令启用 planning,默认关闭。

**验证:**读取同一 Goal/Agent 的 turn-context 与 summary;节点存在不等于已采用。

**停用 / 回退:**执行 --explore-mode off --execute;保留既有证据。

**权限边界:**分析与规划不授予 spawn、claim、扣额、合并或外发权限。

文档:版本固定指南

loopx configure-goal --goal-id "$GOAL_ID" --explore-mode planning --explore-harness-profile adaptive-resilient --execute
loopx explore turn-context --goal-id "$GOAL_ID" --agent-id "$AGENT_ID"
loopx explore summary --goal-id "$GOAL_ID"
loopx configure-goal --goal-id "$GOAL_ID" --explore-mode off --execute

TurnEnvelope and captured decisions

**启用:**每次 guard 显式加 --turn-envelope;保存完整 decision 还需明确 Turn id、已存在父目录与全新目标目录。

**验证:**读取 capture,核对 Goal/Agent/Turn、原始 source hash 与 ok;拒绝或不完整保存不可当成功。

**停用 / 回退:**省略两个参数;仅通过普通文件管理删除不再需要的私有 capture。

**权限边界:**旧观察不提供新准入;选择、lease、取消与写入时校验仍有效。frontend/Lark 不消费这些文件。

文档:版本固定指南

loopx --format json quota should-run --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --turn-instance-id "$TURN_ID" --turn-envelope --decision-output-dir ./guard-001
cat ./guard-001/decision.json

Ordinary workspace Chat

**启用:**运行 loopx chat,在管家对话的 Scope 选择已授权 workspace。宿主声明的目录默认 workspace_write;需要只读时用下方命令。

**验证:**核对 Scope、生效 grant 与同一对话的结果;grant 改变创建新上下文,旧历史保留。

**停用 / 回退:**以 --project-workspace-grant workspace_read 重启服务,或撤销宿主 workspace grant;旧上下文的新工作被拒绝。

**权限边界:**写入遵循 AGENTS.md 与实际 sandbox;不创建隐藏 Goal,不借 portfolio、peer 委派或外发权限。

文档:版本固定指南

loopx chat --project-workspace-grant workspace_read --no-open

Owner private Lark conversations

**启用:**在设置 → Lark 核验 App 与个人 owner,选择 workspace、executor、grant 和普通 Chat/管家角色后连接。/delegate --tokens N objective 需原来源确认;/agents、/agent TARGET_REF、/project 使用另行授权目标。

**验证:**在绑定私聊发送 /status、/help,核对 App/source、角色、workspace grant、原 Session/Turn 与队列;设置回读同一绑定。

**停用 / 回退:**在设置 → Lark 断开精确 App;/stop 停止普通 Turn,/stop-commission 停止绑定委托;另行撤销精确 Agent grant。

**权限边界:**个人凭据、source/owner、listener 身份仍明确。普通 Chat 不创建 Goal,委托确认不授予任意写入、自动 heartbeat 或规范任务验收。真实 Lark/手机资格另行验收。

文档:版本固定指南

loopx chat --project-workspace-grant workspace_read --no-open
loopx chat --help

Configuration checkpoints

**启用:**设置 → 能力中心 → 配置备份与恢复可下载私有 checkpoint;CLI export 使用新目标,先预览再 --execute。

**验证:**对精确文件运行 configuration-backup verify,核对来源范围与摘要;完整性不证明隐私安全。

**停用 / 回退:**未采用的隔离 checkpoint 可删除而不影响 live settings;已采用设置通过已有 revision 校验 editor 或 machine-config/configure-goal 回退。

**权限边界:**不复制 credential store、Host session、grant、live registry、fence、lease 或 scheduler;配置本来包含的秘密仍属私有。

文档:版本固定指南

loopx --format json configuration-backup export --goal-id "$GOAL_ID" --output "$NEW_CHECKPOINT_FILE"
loopx --format json configuration-backup export --goal-id "$GOAL_ID" --output "$NEW_CHECKPOINT_FILE" --execute
loopx --format json configuration-backup verify --input "$NEW_CHECKPOINT_FILE"

Canonical new-Goal creation

**启用:**设备默认 → 新 Goal 的权威存储显式启用 canonical creation,选择 File/SQLite 及 soft_claim/hard_lease;v1 document 采用已有 preview/apply,默认关闭。

**验证:**inspect 后 bootstrap 新空项目,再读 Todos 和原生 authority receipt;已有 Goal 不改目标。

**停用 / 回退:**预览/应用 canonical_creation=false 关闭未来创建,或按精确 removal-plan revision 删除 goal_storage namespace;已有存储与 fence 保持。

**权限边界:**存储/执行策略不授权工具、账号、网络、调度或迁移;丢失的 authority 不能以 forced bootstrap 当空数据重建。

文档:版本固定指南

Save / 保存为 goal-storage.json:

{"schema_version":"loopx_goal_storage_defaults_v1","new_goal_provider":"sqlite","canonical_creation":true,"new_goal_handoff_mode":"hard_lease"}
loopx machine-config preview --namespace goal_storage --config-json goal-storage.json
loopx machine-config apply --namespace goal_storage --config-json goal-storage.json --expected-plan-revision "$PLAN_REVISION" --execute
loopx machine-config inspect
loopx machine-config remove --namespace goal_storage
#Review removal, then use its returned revision.
loopx machine-config remove --namespace goal_storage --expected-plan-revision "$REMOVAL_REVISION" --execute

Governed delegation stop

**启用:**使用已配置的精确 binding、注册 requester 与 operation,delegation stop --execute 显式请求停止;CLI/MCP 与 App 团队表面复用 owner。

验证:delegation read 区分确认、已证明的原生进程 drain 与结算;监督缺失保持 unknown,执行返回的恢复路径。

**停用 / 回退:**移除精确 operator binding 撤销新执行,对 active operation 请求 stop;停止/结算回执不是 resume grant。

**权限边界:**缺少记录不证明 Host 退出、lease 释放或未执行;不支持的停止平台在 launch-side 取消效果之前拒绝。

文档:版本固定指南

loopx delegation stop --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --execution-config "$DELEGATION_CONFIG" --operation-id "$OPERATION_ID" --execute
loopx delegation read --goal-id "$GOAL_ID" --agent-id "$AGENT_ID" --execution-config "$DELEGATION_CONFIG" --operation-id "$OPERATION_ID"

PR review queue ownership and direction

**启用:**能力中心配置额外 owner 登录名与 forward/reverse;Goal CLI 可设置精确 Agent 方向。当前会话请求先于普通队列发现。

**验证:**读取 configure-goal 与只读 pr-review packet,核对实际账号、Agent 和方向;queue owner 不等于 GitHub author。

**停用 / 回退:**用 --clear-pr-review-owner-logins --execute、--pr-review-agent-order AGENT=inherit 或 --clear-pr-review-configuration --execute;降级前移除 Goal/设备 owner_logins。

**权限边界:**不授予 GitHub review、comment、dismiss、merge、跨 Agent 写入或调度权限;review 深度和 CI policy 沿用已有 owner。

文档:版本固定指南

loopx configure-goal --goal-id "$GOAL_ID" --pr-review-owner-login maintainer --pr-review-agent-order reviewer-a=forward --execute
loopx configure-goal --goal-id "$GOAL_ID"
loopx pr-review --goal-id "$GOAL_ID" --agent-id reviewer-a --repo "$REPOSITORY" --format json
loopx configure-goal --goal-id "$GOAL_ID" --clear-pr-review-configuration --execute

Public GitHub evidence

**启用:**每个 plan 使用 --public-github 和完整 SHA 来源;execute --execute 只授权匿名来源读取。

**验证:**回读精确 plan/execution receipt;父采纳/拒绝与研究账本覆盖另行显式操作。

**停用 / 回退:**省略 --public-github 与 --execute;没有持久 provider switch。证据退休仍遵循采纳与下游覆盖规则。

**权限边界:**不使用 token、cookie、私有仓库、可变分支、原始页面持久化、自动采纳或外发权限;完整 frontend 发起仍开放。

文档:版本固定指南

loopx external-evidence plan --public-github --objective "Inspect public source" --user-activity "Choose a source" --decision "Whether a literal is present" --evidence-kind literal_match --source "$SHA_PINNED_PUBLIC_URL" --search-term LoopX --format json > plan.json
loopx external-evidence execute --plan-json plan.json --execute --format json > execution.json
loopx external-evidence readback --plan-json plan.json --receipt-json execution.json

Managed DeepSeek model selection

**启用:**managed execution 使用 deepseek-flash@high;显式 LOOPX_TURN_MODEL/DSH_MODEL 或 --dsh-model 仍优先,凭据由 provider 配置。

**验证:**读 turn run-once --help 及返回的 runtime profile/实际 provider 身份;alias 本身不证明 live qualification。

**停用 / 回退:**将 LOOPX_TURN_MODEL 设置为之前明确选择的模型,重启 owning host;不要静默修改旧 session;省略配置恢复 shipped managed default。

**权限边界:**模型设置不提供 API 凭据、付费调用、workspace、Goal 或 provider 晋升权限;CPA route 仍为独立 operator-owned 部署。

文档:版本固定指南

LOOPX_TURN_MODEL=deepseek-flash loopx turn run-once --help
#Explicit rollback selection for the next owning host invocation:
export LOOPX_TURN_MODEL="$PREVIOUS_MODEL"

Finance evidence assessment

**启用:**Finance Value Discovery 是独立安装的可选 extension 0.8.3;从版本固定源码安装,再注册/启用 manifest。assess-period 输入为 finance_period_comparison_input_v1。

**验证:**运行 extension doctor,再对已审阅本地输入 assess-period;producer 精度、期间资格、来源真实性与投资真值分开。

**停用 / 回退:**用下方命令 disable loopx-finance-value-discovery;私有账户/持仓观察不放公开研究投影。

**权限边界:**只做确定性证据评估,不授予投资建议、下单、转账、签名或执行权限;来源/timezone/binding 缺失保持明确。

文档:版本固定指南

#In a v1.3.0 source checkout:
python3 -m pip install ./packages/loopx-finance-value-discovery
loopx extension install --manifest packages/loopx-finance-value-discovery/extension.toml --execute --format json
loopx extension enable loopx-finance-value-discovery --execute
loopx extension doctor loopx-finance-value-discovery --execute
loopx-finance-value-discovery assess-period --input-json period.json
loopx extension disable loopx-finance-value-discovery --execute

EdgeBench native trials

**启用:**仅在 v1.3.0 checkout 中显式启动研究:需要 Linux、Docker、固定 EdgeBench/SForge/Harbor、选定 Codex 与已授权模型。明确 worker 和反馈 profile。

**验证:**先 --help 再审阅试验命令;回读 session/model、终态、评测完成与隔离回执。注册/采样计数不证明有效 outcome。

**停用 / 回退:**不再启动新试验;按已有 cancellation lifecycle 停止精确 trial/controller 与 TLS relay,保留私有日志和未完成结果。

**权限边界:**发布不会启动 benchmark;judge 拥有任务/评分,blind 还需凭据、网络、提交资格隔离验收;原始证据与秘密保持私有。

文档:版本固定指南

python -m benchmark.edgebench.run --help
#Paid execution only after the operator reviews task, pin, isolation and budget:
python -m benchmark.edgebench.run --task "$TASK_ID" --tasks-dir "$TASKS_DIR" --log-dir "$PRIVATE_LOG_DIR" --run-id "$UNIQUE_ATTEMPT" --worker heartbeat-resume --model "$MODEL" --effort xhigh --judge-url "$JUDGE_URL"

DSH LoopX plug-in

**启用:**按下方命令将独立发布的 0.1.1-beta.6 安装到 web profile,先核对 DSH 兼容范围;当前已资格验证宿主为 0.2.0-rc.2。加载准备隔离 CLI/skills 和 GoalBar;Driver 仅在这个精确 Session 调用已安装 loopx skill 后激活。

**验证:**读 dsh --version,再解析精确 live Session 的绑定,要求 status=bound 且唯一 Goal/Agent 对。文件已安装不等于 Session 已绑定。

**停用 / 回退:**从同一 profile remove dsh-loopx-plugin 后重启 DSH。回退时安装预先保留且与宿主兼容的旧 tarball,保留 LoopX 状态。

**权限边界:**安装可以准备隔离 CLI 与 skills,本身不授予 Goal 绑定、扣额、模型调用或执行权限。Goal/Todo/Agent 决策仍由 LoopX 管理;包不含凭据,GoalBar/Driver 使用已认证本地 Session 边界。

文档:版本固定指南

dsh --version
dsh plugin --profile web add "https://github.com/loopx-project/loopx/releases/download/dsh-loopx-plugin-v0.1.1-beta.6/dsh-loopx-plugin-0.1.1-beta.6.tgz"
dsh --profile web --port 0
#After invoking the installed loopx skill in the exact DSH Session:
loopx --registry .loopx/registry.json --format json resolve-agent-thread --host-surface deepseek-harness-native --thread-id "$DSH_SESSION_ID"
dsh plugin --profile web remove dsh-loopx-plugin

发布验证

包/tag 为 1.3.0 / v1.3.0,维护者集成、最终冻结提交、产物及正式发布时间仍待门禁完成。候选已通过版本/readiness、help/manual、双语书籍检查点、必需范围 Ruff、19 文件 Mypy、打包 Chat、42 项 dashboard 命令及改动路径公开边界检查。先前 174 通过/6 失败和任务板 1 通过/22 失败保留为历史证据。

主线回归继续复用已有 #5533,当前 head 为 5b177a260b5b417cf5317c76b5ebdc9937dfa8a9,精确提交评审已发布。当前精确源码的 19/19 项 risk 检查及 5 项直接检查、143 项主线兼容性测试、52 项配置/API 测试、143 项 CLI 预算测试和 3 项真实宿主安装测试通过;96 项相同 base/head 输出测量无 review signal,保留一项既存 Pydantic warning。Ruff、19 文件 Mypy 与真实 Chat 重建/源码校验通过;17 路径的源码变更记录区分旧证据可复用边界和主线集成改变的三个文件。冻结的 91666 完整诊断结束:16,748 passed、15 failed、118 skipped、453 passed subtests。五项旧 bundle 失败已独立构建/复验;其余十项在固定主线同样复现,当前 fixture/文档修复保留拒绝与无副作用断言。原失败运行不改称通过;修复尚未纳入发布源码,维护者合并与最终精确源码资格仍待完成;此 managed 合并门不查询 CI。

相同 crowded fixture 在冻结主线与候选分别测得 Turn plan JSON 14,647 字符、diagnose 44,126 字符;回归阈值按证据从 14,600 调为 15,000、44,000 调为 45,000。完整 Vision、可执行恢复、行数和语义/结算断言保留;monitor 使用已有稳定路径 fixture,仍保留 2,000 字符阈值。这些是展示回归预算,不是 token、费用或实验晋升限额。

DS V4.1 Flash/high 已在 8056cc61d 候选隔离原生 Goal 中真实通过:2 Turns、2 Todo 结算及独立验收。最终发布提交仍须重跑。一次备用 provider 探测鉴权失败,不将它当作模型语义失败。没有启动新的 benchmark。

个人飞书账号权限及旧指南历史已核验;1.3.0 候选指南已读回个人归属、完整正文、两张图和九页导出排版,正式发布后原地更新实际产物与状态。首次 no-clone smoke 在隔离 HOME 清理失败,fresh-clone/update 分别通过;最终 full-public、完整 pytest 和精确 commit 资格仍须完成,真实 PostgreSQL 未被称为通过。

Compare: v1.2.4...v1.3.0

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; gpt-6.1-sol; OpenAI; runtime_reported; reasoning_effort=xhigh

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

动机

维护者准备命名版本,用户通过版本、帮助和开发者书确认安装来源。准备 1.3.0 时,旧元数据仍标为 1.2.4,帮助分类检查遗漏现有配置备份命令;修改后版本对齐且帮助预检通过。干净候选报告 1.3.0,版本、manpage、开发者书及错误 tag 拒绝检查通过。本 PR 不发布 tag、不移动 stable、不替代完整资格验证,也不声称 benchmark 收益已成立。回归修复仍在 #5533,最终合并提交的完整验证、真实模型资格、个人飞书指南和发布产物 readback 尚待完成。

改动思路

本次评审基于完整 head 5c80a2551d36fb5fe6a0f466537c2da997490b6d,对照不可变主线 aada23d751c2a7e352d97430cd951e08335bcd8e。只复用既有版本源、帮助分类和文档生成路径,不引入另一份发布状态或新 CLI。版本只改 package identity;配置备份原本已经注册,加入既有命令专属帮助集合,让 parser 与 manpage 的集合检查重新闭合。同步主线解决旧界面语言断言,合并后 PR 差异仍为八个文件、九行新增和八行删除。

具体改动

规范依据:docs/product/release-readiness.md,spec_revision aada23d751c2a7e352d97430cd951e08335bcd8e。named-version-contract 已实现:loopx.__version__ 与 pyproject.toml 同为 1.3.0,错误 tag 会拒绝。documentation-preflight 已实现:manpage、四处开发者书版本锚点和帮助分类对齐,实际预检通过。compatibility-gate 属于最终晋级的 deferred 项,由维护者合并 #5533 后冻结提交执行,不能拿这次准备检查替代。

关键代码讲解

loopx/__init__.py::__version__ 是公开版本来源,pyproject.toml 镜像该值,既有 release artifact validator 负责匹配 tag;实际传入 v1.2.4 得到预期 exit 2,没有创建任何 ref。loopx/help_surface.py::MANPAGE_COMMAND_HELP_ONLY 在原集合中添加现有 configuration-backup,既有 parser census 检查所有顶层命令必须属于手册组或该集合。没有删命令、改参数或打开备份执行权限。man/loopx.1 只改变生成版本,四处 book checkpoint 只改变当前版本基线;历史示例没有批量重写。语义 advisory 识别了该集合扩展,决定复用既有本地 owner;全树语义检查通过。

对主干的风险

版本/help/manpage/developer-book smokes、source-built Chat verify,以及此前失败的界面语言检查均在此 head 通过。实际默认 CLI 的同 fixture 主线/候选对照有 96 行,零 candidate-only,动作签名与结构无差异;诊断与 Turn plan 的旧字符上限失败在两边一致,当前预算修正在 #5533,不能称这个 head 的完整 premerge 全绿。真实 Git baseline 与此 head 的 maintainability 检查还出现同样三个非本 PR 改动的 debt finding,保留为主线资格缺口。归因采用精确源码与失败身份,不以相同测试数量替代证据。同步后的新 GUI 包已按源码重建;没有宣称正式发布包、PyPI、Windows 或真实模型资格已通过。

语义与 CI 对齐

既有集合采用精确 membership,不是文字猜测或 substring denylist。CLI 可见性不授予执行、访问、费用或 actor 生命周期权限。本次没有 optional-capability 行为改变,实际同输入对照覆盖默认路径;并未借“未执行某功能”证明隔离。当前配置的评审依赖本地证据,不等待 CI;原构建失败和独立基线失败仍保留,最终发布门不因此豁免。

我的整体评价

APPROVE,交付判断为 justified_increment:这个独立、可回滚的准备步骤解决版本和文档预检一致性;long_horizon 与 user_experience 为 preserved,执行、调度、结算和授权路径没有改变。未来重构检查无需新增抽象,现有 owner 已是最小完整边界。最大的剩余风险是把候选准备误称已发布;因此本次不自合并,不宣称完成 release。最终合并、完整资格和发布 readback 仍由当前发布工作继续处理。

English verdict: APPROVE - 5c80a25; coherent 1.3.0 metadata/help/documentation preparation, with focused checks and paired CLI evidence. Baseline qualification failures and final release gates remain explicit; no release or maintainer merge is performed.

@loopx-agent

Copy link
Copy Markdown
Collaborator Author

Owner-authorized release preparation integration at 5c80a2551d36fb5fe6a0f466537c2da997490b6d.

The exact-head COMMENTED approval conclusion remains valid. The refreshed LoopX gate returns ready=true for the unchanged head; transient unknown mergeability was independently resolved through GitHub REST readback and a conflict-free native merge-tree check. The owner explicitly authorized admin bypass for release integration. Version/help/book contracts, packaged Chat source verification, locale checks and the 13-surface bilingual body validator passed; unchanged earlier baseline failures and their independently reviewed #5533 repair remain in the review. #5533 is now merged.

Final integrated-source full regression, install/upgrade/host, live-model qualification and artifact/guide readback remain separate from this version/preparation merge. Existing failed receipts are retained, and no tag, release or stable promotion is claimed. CI is not consulted under the resolved managed policy. The adjacent refactor pass uses the existing help-only command classification; no new owner or capability is introduced.

@loopx-agent
loopx-agent merged commit 5981489 into main Oct 5, 2026
24 of 30 checks passed
@loopx-agent
loopx-agent deleted the codex/release-1.3.0 branch October 5, 2026 17:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant