Skip to content

TRACKING: work plan from 2026-07-29 — run 2 against meta, and the block-2 judgement calls #25

Description

@mmcky

Work plan opened 2026-07-28, kept current since. Revised 2026-08-07 against main at 3dc9aeb, with every claim below re-checked rather than carried forward — the comment thread holds the running narrative, this body holds the state.

Blocks 1, 3 and 4 are done. The front of the plan is now the maintainer read of run 2 — the one thing no self-audit can supply — and block 2, which is the same kind of work. Program state is in #16; defects in #21, plus #57/#58/#59 from run 2.

Targeted revision 2026-08-13 — three facts moved; blocks 2 and 4 are unchanged and block 4 is still the front. (1) #33 merged as f035a51, benchmark 0.4.0, tagged: the benchmark-thread row and its Parallel box below are stale where they say it is open — Kenko signed off 08-07, and his queue is now the #10 checklist, then the plain-NumPy ge_arrow follow-up. (2) #44 merged (maintenance round): FUTURE-IDEAS.md retired into issues #36#42, the NEXT-SESSION.md scratch-file practice retired — this issue is now the sole carrier of cross-session state — and the namespace-consolidation proposal is #43, unblocked by #33's merge and awaiting a decision. (3) The pre-existing open issues now carry QEP-2 type labels (this one stays untyped pending a QuantEcon/qeps field report on plan/tracking issues).

Revised 2026-08-20 07:30Z (times now carried on stamps — same-day sessions have begun to collide) — a fourth thread arrived; the front did not move. (1) The qe workplan family shipped: #46 (/qe:workplan 0.3.0 — report bundle → tracking issue with sub-issues) then #47 (0.4.0 — renamed /qe:workplan-project, joined by /qe:workplan-issue and /qe:workplan-update), with #48 catching the repo-level docs up. The convention the family operates is headed for a QEP: the discussion the 2026-08-13 note above was pending now exists at QuantEcon/qeps#15. This revision is itself /qe:workplan-update's first run from an installed plugin. (2) Block 4's two preconditions were re-verified at this stamp and both still stand: the token still lacks read:project, and the enabled audit is still 0.1.4 — 0.2.0 sits in the plugin cache but is not the active version, so claude plugin update audit@quantecon is still required before run 2. Block 4 remains the front. (3) Third-party movement since 2026-08-13, none of it this session's: Kenko's #10 checklist answers (2026-08-14/16) — see the rewritten benchmark box.

Re-verified 2026-08-20 07:50Z on resume by a fresh session — /qe:workplan-update's first resume run, and the second half of its acceptance test: the body alone was enough to relocate the anchor, the link graph and the front. Nothing moved in the 20 minutes since the stamp above. Block 4's preconditions re-measured and both still stand (token still lacks read:project; enabled audit still 0.1.4). Off-plan parallel-session trace, not this session's: QuantEcon/action-translation#283, #284 and #285 filed 07:33–07:45Z — review-tooling work this plan does not carry.

Revised 2026-08-25 02:30Z — the plugin surface consolidated under one namespace; the front did not move, and one of its preconditions got easier. (1) Three merges since the last stamp: #51 (2026-08-24, qe 0.5.0 — workplan-issue and workplan-update consolidated into /qe:workplan with a new read verb; this revision is that consolidated skill's first update run), #52 (qe 0.6.0 — the seven check-* scaffolding skills removed from the tree; the policy is now that no scaffolding ships and an unbuilt skill lives only in its tracking issue, with the plan and the preserved frontmatter descriptions on #3), and #53 (qe 0.7.0 — #43 implemented and auto-closed: the benchmark and audit plugins fold into qe as /qe:benchmark and /qe:audit-issues, engines at qe/scripts/{benchmark,audit}/, references at qe/references/{benchmark,audit}/, retired changelogs preserved as historical sections; tags qe--v0.6.0 and qe--v0.7.0 pushed, old stream tags remain archaeology). (2) Block 4's plugin precondition inverted, favourably: there is no audit@quantecon to update any more — this machine's enabled audit 0.1.4 is uninstalled and qe updated 0.5.0 → 0.7.0, which carries the post-#34 procedure, so the undetectable stale-procedure hazard that box warned about is closed; the residual precondition is a session restart so 0.7.0's skills load, and the run-2 invocation is now /qe:audit-issues. The token scope was re-measured at this stamp: still no read:project, so that precondition holds. (3) Blocks 2 and 4 are otherwise unchanged and block 4 is still the front.

Revised 2026-08-25 17:10Zblock 4 is done: run 2 executed, interrupted, and resumed. The front of the plan moves for the first time since this issue was written. (1) Run 2 complete against QuantEcon/meta — 317 items / 138 open, ~53 min wall clock including the interrupt. Record in #56 (merged 2026-08-25 as bcfdca5); results on #16; three follow-ups filed as #57 (defect), #58, #59. (2) Resumability held — killed at 68/138 open, fresh session resumed at #261, the point predicted and recorded beforehand; no re-walk, no skip, no duplicates. That was the claim the whole validation program existed to check. (3) Read-only moved from asserted to measured — a fingerprint of all 317 issues before and after hashed identically. (4) One real defect: the closed pass is not incrementally checkpointed (one write for 179 items against 15+ for 138 open), so SKILL.md's "loses one item rather than the phase" fails there — #21's defect 2 in a new form, #57. (5) #23 is now answerable and the question has moved: not "shrink the apparatus" but "should the closed pass scale, and on what?" — it is 56% of the run for the cheapest checks, and worse on older trackers.

Blocks are ordered by dependency, not by size. Block 3 before block 4 was the one ordering that mattered, and it has now been honoured.

Where this sits in the repo

Added 2026-08-07, because this plan is audit-shaped and had never said so. It grew out of the 2026-07-28 audit session: blocks 1 to 4 are all audit, and until this section the word "style" did not appear anywhere in it. The repo has four live threads, and they are independent — none blocks another.

Thread Plan State Waiting on
audit — validation program #12, program in #16 Run 2 done 2026-08-25 — resumability validated, read-only measured, one defect (#57). Record merged as bcfdca5. Two of four matrix runs complete You: the maintainer read of run 2's tiering (record §9) — tiering, the 19 drafted closing comments, and where the bundle lives
benchmark — triage-first #4, rubric calls in #7 #33 merged 2026-08-13 — 0.4.0, tagged; now /qe:benchmark in qe 0.7.0 You: read Kenko's #10 checklist answers (posted 2026-08-14/16), then his ge_arrow plain-NumPy follow-up. #7 is the one untouched thread
qe — the style surface #3 Scaffold removed from the tree in 0.6.0 (#52); the 19-item plan lives solely in #3, none landed Nothing, for items 1 and 3 — see below
qe — the workplan family #3; convention QEP discussion in QuantEcon/qeps#15 0.3.0 + 0.4.0 shipped 2026-08-20 (#46#48): workplan-project / workplan-issue / workplan-update Consolidated 2026-08-24 to two skills in 0.5.0 (#51): /qe:workplan (lifecycle, incl. the new read verb) + /qe:workplan-project. Validated runs: the 2026-08-20 update/resume pair, and this issue's 2026-08-25 revision as the consolidated skill's first update; workplan-project is unrun

The validation-program gate does not reach qe, and it is worth being exact about why. The "Explicitly not doing yet" section below holds the /audit:prs, /audit:tech-debt and /audit:translations skills back until the shared method is proven. That argument is scoped to the audit family — skills that would inherit the audit method docs (qe/references/audit/doctrine.md since 0.7.0; previously audit/references/doctrine.md), which is what "the shared method" means. A qe style skill inherits none of it: different references, a different upstream source of truth in style-guide, and different blockers — the 0.7.0 namespace merge changes nothing here, because the method boundary was never the plugin boundary. Run 2 could succeed or fail and would say nothing about whether the style preflight engine is sound.

So qe items 1 and 3 are unblocked today — the preflight engine and the umbrella skill body. Verified 2026-08-07 that nothing had silently landed; 2026-08-25: the scaffold itself is now gone from the tree (#52 deleted the seven check-* skills and qe/references/rules/), so #3 is the sole home of the plan and the preserved skill descriptions, and re-landing goes straight to the one-umbrella shape #43 settled. The upstream blockers gating items 2, 4 and 5 have not moved — project-style-guide#6 since 2026-07-21, project-style-guide#2 since 2026-06-11.

Recorded so it can be questioned rather than inferred: style is the stated flagship, the largest recurring theme across the ~630 merged lecture PRs that justify the qe plugin at all. It currently has nineteen unlanded items beside three audit releases in a week. That is not a decision this plan made — it is where the sessions went. The only real coupling between the threads is maintainer attention, which is a scheduling fact and not a dependency, so none is asserted here.

Block 1 — merge what is waiting ✅

This block has refilled and re-emptied twice since it was written, which is why it is worth keeping rather than deleting: #26, #27, #28 and #29 landed the same day as its own contents, then #30, #31 and #32 on 2026-08-03, then #34 and #35 on 2026-08-07. Nothing is open in this repo as of the revision except #33, which belongs to the benchmark thread below.

Block 2 — the judgement only you can supply

Unchanged since this plan was written, and the only part of it that cannot be delegated.

  • Checks 9 and 10 on the run-1 bundle: does the tiering in 01-issue-triage-report.md match what .dev/PLAN.md actually says, and would you act on it? Both are still marked Maintainer's call — pending in the merged record (reviews/audit-run-action-translation-2026-07-28.md, lines 47–48), so it merged honestly as-is — append afterwards, or comment on TESTING: validation program for /audit:issues — run it before generalising the method #16.
  • Decide the bundle's fate. It is in ~/work/quantecon/action-translation/.dev/scratch/audit-2026-07-28/ — matched by that repo's .gitignore rule .dev/scratch/*, so it survives but cannot be committed where it sits, and .dev/audits/ still does not exist there. Options: move it to .dev/audits/2026-07-28-issues/ now (pre-empting the convention), or leave it pending QuantEcon/QuantEcon.manual#140 — still open, last updated 2026-08-02, so waiting remains available.
  • One finding in that bundle is still unactioned. feat(glossary): add Japanese (ja) translation glossary action-translation#69 has been open since June and re-creates a directory that Wave 1 deliberately deleted; still open, last updated 2026-08-01. Corrected 2026-08-07: this box originally named two findings. The other — the hold on Apply localisation rules on the sync path for newly-created files action-translation#225, measured in #227 to regress the path it touches (93.3% → 33.3%, p = 0.002) — needs nothing. That PR already carries two comments from you dated 2026-07-27, the do-not-merge hold and the confirmed regression, both of which predate the 2026-07-28 run, so the bundle's "unsent drafted comment" was already redundant when it was drafted. #225 is still open and untouched since 2026-07-27.

Block 3 — the two severity-1 defects (#21) ✅

Done — shipped as audit 0.2.0 in #34, merged 2026-08-07 (d4e8df8) and tagged audit--v0.2.0. They went in as the single doctrine-level PR this block proposed, on the grounds it gave: both concern whether an audit's own record can be trusted. #35 then brought the tutorial and CATALOG in line, since step 4 had been teaching the resume rule #34 replaced — in the very document used to test resume.

  • Defect 1 — [verified] citations are not checked for reachability. doctrine §2 required file:line, a merged PR, or a tag, but never that a cited commit be an ancestor of the ref the audit names. Run 1's headline finding cited a commit that exists only on an unmerged branch. Fixed, and broader than this box described: reachability is now required of every citation form, because file:line and PR citations are ref-relative too and fail the same way; phase 2 runs git merge-base --is-ancestor <sha> <ref> before tagging a commit. Copilot's review then found a third instance of the same class — §1 rule 1 still carrying its own copy of the accepted-forms list — fixed at the cause in 27a5db6 by deleting the duplicate rather than syncing it.
  • Defect 2 — the closed side is never checkpointed. findings.md held the 56 open issues; the 62 closed ones went straight to the catalog. Fixed: findings.md now carries ## Open and ## Closed, the resume rule partitions issues.json by state and resumes each side independently, and the skill says explicitly not to infer progress from a single block or from the file's length.

Block 4 — run 2 ✅

  • gh auth refresh -s read:project first. Done 2026-08-25, and verified by capability rather than by the scope string: organization(login:"QuantEcon"){projectsV2{nodes{title}}} returns project titles. Worth recording how the check can lie — projectsV2{totalCount} alone succeeds without the scope, so a totalCount probe reads as success while the gap that changed run 1's conclusions is still open; query a real project field. Note also that gh auth status lagged the grant by minutes here, so trust the API, not the status line.
  • claude plugin update before starting. Added 2026-08-07; superseded 2026-08-25. The hazard this box guarded — the installed audit 0.1.4 silently executing the pre-fix procedure — is closed by the qe: one namespace — benchmark and audit fold into qe (0.7.0) #53 consolidation: audit@quantecon is uninstalled, qe is at 0.7.0, and /qe:audit-issues carries the post-audit: a citation must resolve on the ref the audit named, and both passes get checkpointed (0.2.0) #34 procedure. Residual: start run 2 from a fresh session, so the 0.7.0 skills are actually loaded.
  • Run 2 against QuantEcon/metadone 2026-08-25, 317 items / 138 open, ~53 min wall clock including the interrupt, bundle in the untracked .audit/2026-08-25/. Six of TESTING: validation program for /audit:issues — run it before generalising the method #16's seven claims marked HELD with evidence; tiering is the seventh and is yours to judge. Per TESTING: validation program for /audit:issues — run it before generalising the method #16's matrix it is the sharpest test of what generalises: org-wide issues with no code to verify against, which directly stresses doctrine rule 1 and may reveal that [verified] means little for a decision record.
  • Interrupt it deliberatelydone: killed at 68/138 open with the session closed outright, and a fresh session resumed at #261, the item predicted and written down before the resume. No re-walk, no skip, no duplicates. Resumability is validated. Two gaps remain for run 3, recorded on TESTING: validation program for /audit:issues — run it before generalising the method #16: the kill left a clean write so the truncation guard was never exercised, and the interrupt landed in the open pass rather than the closed one.

Why block 3 came first, and what it bought. Run 2's headline purpose is testing resume, and defect 2 meant the checkpoint covered only half of phase 2. Testing a resume path against a checkpoint already known to be incomplete would have produced a failure that said nothing about resume logic and everything about a missing block. That is now fixed, so an interrupt tests the claim rather than a known hole.

Run 2 also produces the second data point that #23 (does the phase apparatus need to shrink?) is explicitly waiting on.

Parallel — no dependencies, any time

Explicitly not doing yet

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions