Skip to content

capt-hook helper install cannot converge on a busy host #148

Description

@yasyf

capt-hook helper install cannot converge on a busy host: it took nine attempts to land 12.50.1, and succeeded only by winning a race.

Three separate defects, one observed symptom each. Parts 1 and 2 are diagnosed; part 3 is not, and I have said so rather than guessed.

Host: Darwin 25.6.0, arm64. Homebrew formula yasyf/tap/captain-hook, deployment at ~/Applications/Captain Hook.app. Roughly twenty agent sessions were running concurrently, each firing hooks through the plugin's bin/hook shim.

1. The quiesce gate cannot hold the launchd agent down across the swap

brew upgrade captain-hook moved the Cellar to 12.50.1. Every subsequent capt-hook helper install then failed with:

Error: Homebrew holds 12.50.1 but the running helper reports 12.50.0 — the deployment did not converge. The Cellar copy is not the deployment; nothing is installed until the host answers.

~/.claude/state/update/update.log records the reason. Seven of the eight failures are the same refusal, with a different capt-hookd pid every time:

2026-09-21T03:59:56Z package-install exit 1: captain package: land delivered app: deploy: live processes remain on the deployment's executables: pid 8603 …/Contents/MacOS/Captain Hook, pid 8757 /Users/yas
2026-09-21T04:42:14Z package-install exit 1: captain package: land delivered app: deploy: live processes remain on the deployment's executables: pid 71321 …/Contents/Helpers/capt-hookd
2026-09-21T04:43:22Z package-install exit 1: captain package: land delivered app: deploy: live processes remain on the deployment's executables: pid 96880 …/Contents/Helpers/capt-hookd
2026-09-21T04:45:53Z package-install exit 1: captain package: land delivered app: deploy: live processes remain on the deployment's executables: pid 43553 …/Contents/Helpers/capt-hookd
2026-09-21T04:46:17Z package-install exit 1: captain package: land delivered app: deploy: live processes remain on the deployment's executables: pid 49795 …/Contents/Helpers/capt-hookd, pid 49811 /Users/y
2026-09-21T04:46:48Z package-install exit 1: captain package: land delivered app: deploy: live processes remain on the deployment's executables: pid 56875 …/Contents/Helpers/capt-hookd, pid 56885 /Users/y
2026-09-21T04:47:27Z package-install exit 1: captain package: land delivered app: deploy: live processes remain on the deployment's executables: pid 59679 …/Contents/Helpers/capt-hookd, pid 59715 /Users/y

quiesceInstalledApplication (internal/hookd/deployment.go:281) stops the installed generation and then daemonkit's executable-scoped inventory gate (deploy.ErrLive) asserts nothing live remains. Nothing keeps the host from coming back between those two steps. On this host com.yasyf.captain-hook.host.v1 is a launchd agent, and every hook call across every session also cold-starts capt-hookd — so a fresh process is essentially always live by the time the inventory runs. The rising pids show a new process each attempt, not a survivor that refused to die.

The refusal itself is right; a half-quiesced daemon should never be swapped under. The gap is that nothing holds the agent down for the duration of the land. The caller is left to supply that, which is not something a caller can do correctly.

What converged it here was running the sanctioned command in a bounded retry loop until one attempt happened to fall in a quiet window — the fifth of that loop, ninth overall. That is luck, not convergence, and it is the behavior the installer should be providing internally: bootout (or otherwise hold) the launchd agent, land, bring it back.

2. The tool-env build and the ownership step share one deadline

One failure is different:

2026-09-21T04:40:30Z package-install exit 1: captain package: own deployment processes: daemonkit: OwnProcesses was handed a context with no budget left: context deadline exceeded; nothing was applied and /Users/yasyf/Applications/Captain Hook.a

That run was the first helper install after the upgrade, and it spent roughly eight minutes materializing the 12.50.1 tool environment under ~/.daemonkit/tools/capt-hook/12.50.1 — roughly 500 MB of spaCy, wasmtime and their dependencies — before reaching the ownership step. By then the shared deadline was gone.

This is a budget-allocation bug rather than a slow machine: the first install of any new version pays the env build, so the step that most needs its budget is the one the shared deadline starves. The env build and the process-ownership and land step need separate deadlines.

Minor, same area: the breadcrumbs are truncated to 200 characters (updater.py, stderr.strip()[:200]), which cuts every one of the lines above mid-path — the second capt-hookd pid and the trailing explanation are lost. A longer cap, or writing the full stderr, would make this far easier to diagnose from the log alone.

3. The stale tool-env prune is not running — undiagnosed

The 12.50.0 changelog says the generated install scripts pin binrun v0.8.0, "whose installer prunes a distribution's stale tool environments beyond the newest three while skipping any a live worker still runs from."

On this host that prune is not happening. Twelve environments survive, back to 12.28.0 from 2026-09-14, totaling 4.6 GB:

255M  ~/.daemonkit/tools/capt-hook/12.28.0/
256M  ~/.daemonkit/tools/capt-hook/12.30.11/
256M  ~/.daemonkit/tools/capt-hook/12.30.8/
292M  ~/.daemonkit/tools/capt-hook/12.38.0/
292M  ~/.daemonkit/tools/capt-hook/12.40.0/
292M  ~/.daemonkit/tools/capt-hook/12.41.1/
511M  ~/.daemonkit/tools/capt-hook/12.42.0/
504M  ~/.daemonkit/tools/capt-hook/12.45.0/
504M  ~/.daemonkit/tools/capt-hook/12.48.0/
504M  ~/.daemonkit/tools/capt-hook/12.49.0/
504M  ~/.daemonkit/tools/capt-hook/12.50.0/
504M  ~/.daemonkit/tools/capt-hook/12.50.1/

The 12.50.1 environment was created by capt-hook helper install during this session, so that path reached the installer and still pruned nothing. I have not worked out why — whether the prune lives only on the generated install-script path and not on package-install, whether the live-worker skip is over-matching on a host with many sessions, or something else. Someone who knows that installer will see it faster than I would by guessing.

Why this matters beyond one slow upgrade

While the Cellar held 12.50.1 and the deployment held 12.50.0, sessions kept running the old code. The fix in #145 is a concrete example: probing through the plugin's own shim, git stash list still returned the 12.50.0 refusal long after brew reported the upgrade done. The version string and the behavior disagreed, and only the behavior was true.

The same drift class produced a different symptom earlier in the same session — a PreCompact hook failing outright:

captain: the capt-hook 12.48.0 tool env is not installed; run 'capt-hook helper install'

so this is systemic rather than a one-off. The two together suggest the deployment and its tool environments can each drift from the package independently, and that helper install is currently the only reconciler while also being the step that cannot reliably run.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions