Summary
Follow-up from the round-3 review of #141 (head 3ff60504). If the ingestor's stats temp file (<stats path>.tmp) belongs to another user, every write fails. This typically happens after the service user changes outside Docker. The ingestor then logs the same error every second without saying how to fix it, and the server keeps serving the last good stats with no sign that they are stale. Low severity: it is an operability issue, not a leak. This exists only on the #141 branch, which added the owner check.
Relates to #118.
Where (#141 head 3ff60504)
cmd/ingestor/stats_file.go:192: checkStatsTmpOwner returns "%s belongs to uid %d, not %d".
cmd/ingestor/stats_file.go:352: the writer loop logs [stats-file] write %s: %v on every tick (1 Hz) with no rate limit.
/api/mqtt/status (and whatever else reads the stats file) shows the frozen data as if it were current.
Proposed fix
- Log the owner error once (or at a low rate, for example once a minute while it persists) and add a hint, for example
remove <path>.tmp or fix its owner.
- Log once when writes succeed again.
- Make staleness visible: when the stats file has not been updated for N seconds,
/api/mqtt/status should say so (for example a stale/age field, or a lastError). Reuse the staleness rule that /api/perf/io already applies, if possible.
Acceptance criteria
- Test: a foreign-owned tmp file gives at most one log line per interval, and the message contains the hint.
- Test: when the stats file is stale, the status response marks it as stale.
- No change in behavior when the tmp file is owned correctly.
Summary
Follow-up from the round-3 review of #141 (head
3ff60504). If the ingestor's stats temp file (<stats path>.tmp) belongs to another user, every write fails. This typically happens after the service user changes outside Docker. The ingestor then logs the same error every second without saying how to fix it, and the server keeps serving the last good stats with no sign that they are stale. Low severity: it is an operability issue, not a leak. This exists only on the #141 branch, which added the owner check.Relates to #118.
Where (#141 head
3ff60504)cmd/ingestor/stats_file.go:192:checkStatsTmpOwnerreturns"%s belongs to uid %d, not %d".cmd/ingestor/stats_file.go:352: the writer loop logs[stats-file] write %s: %von every tick (1 Hz) with no rate limit./api/mqtt/status(and whatever else reads the stats file) shows the frozen data as if it were current.Proposed fix
remove <path>.tmp or fix its owner./api/mqtt/statusshould say so (for example astale/agefield, or a lastError). Reuse the staleness rule that/api/perf/ioalready applies, if possible.Acceptance criteria