Skip to content

fix(runtime): treat a reused PID as a stale lock - #915

Open
edwardtoday wants to merge 1 commit into
kxn:masterfrom
edwardtoday:fix/relayd-lock-pid-reuse
Open

edwardtoday wants to merge 1 commit into
kxn:masterfrom
edwardtoday:fix/relayd-lock-pid-reuse

Conversation

@edwardtoday

@edwardtoday edwardtoday commented Sep 16, 2026 •

Copy link
Copy Markdown

Fixes #913.

问题

clearStaleLock 判断锁是否仍被持有时,只检查了 processAlive(record.PID)。持有者被强杀(SIGKILL、断电、launchd 退出超时)后 relayd.lock 会残留下来;一旦该 PID 被系统复用给任意无关进程,processAlive 就恒为真,daemon 每次启动都报

service error: acquire service runtime lock: lock held by another process

配合 KeepAlive.SuccessfulExit=false,launchd 把它放大成崩溃循环,服务在人工删掉锁文件之前一直不可用(本机 runs 达到 7668,而实际上没有任何进程在运行)。出问题的锁记录的是 pid=991,该 PID 随后被 iCloudDriveService 的 ContainerMetadataExtractor.xpc 复用。

同一个函数还有第二条失效路径:锁文件无法解析时(例如写锁过程中被强杀导致 JSON 截断),它返回 JSON 错误而不是清理文件,同样是永久卡死。

修复

写锁的进程,其启动时间不可能晚于锁的创建时间;所以把持有者的进程启动时间和 record.CreatedAt 比较,启动时间更晚即说明该 PID 已被复用,锁是陈旧的。

  • 启动时间读取:darwin 用 unix.SysctlKinfoProc,linux 用 /proc/<pid>/stat + btime,windows 用 GetProcessTimes;其他平台保持原有行为
  • 留 1 秒容差吸收时钟抖动;取不到启动时间时保持原有行为,避免误清仍然有效的锁
  • 无法解析的锁文件现在会被清理

测试

  • TestAcquireLockClearsLockWhosePIDWasReused —— 复现所报的卡死;在 master 上以 lock held by another process 失败
  • TestAcquireLockClearsCorruptLockFile —— 在 master 上以 unexpected end of JSON input 失败
  • TestAcquireLockFailsWhileLiveOwnerHoldsIt(既有测试)确认真正在运行的持有者仍被尊重

已用 go test ./internal/runtime/ 验证,并额外通过 GOOS=linux / GOOS=windows 执行 go vet ./internal/runtime/。

备注

CI 里的 scripts/check/no-local-paths.sh 在干净的 master 上同样失败(internal/core/orchestrator/service_target_picker_polluted_workspace_test.go 中的 /home/qagent/... 不在 allowlist 内),与本 PR 无关。

clearStaleLock only checked processAlive on the recorded PID. After the
lock holder is killed without a chance to clean up (SIGKILL, power loss,
launchd exit timeout), the lock file stays behind. As soon as that PID is
recycled by any unrelated process, processAlive stays true forever and the
daemon can never acquire the lock again: it exits with "lock held by
another process" on every start, which launchd's KeepAlive amplifies into
a crash loop until the lock file is deleted by hand.

The same function also returned an error instead of clearing the file when
the lock JSON could not be parsed, so a lock truncated mid-write had the
same permanent effect.

Compare the holder's process start time with the lock creation time: the
process that wrote the lock cannot have started after the lock was
created, so a later start time means the PID was reused. Keep the previous
behaviour when the start time is unavailable, and clear unparseable lock
files. Start time lookups are implemented for darwin, linux and windows.

Fixes kxn#913
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] relayd.lock 只按 PID 判活:PID 被复用后 daemon 永久卡在 lock held,launchd 崩溃循环拖垮服务

1 participant