Skip to content

fix(ingestor): rate-limited stats tmp errors, FIFO-safe writer, stale status (#160, #161) - #216

Merged
dborup merged 11 commits into
masterfrom
codex/issue-160-161-stats-tmp
Oct 4, 2026
Merged

dborup merged 11 commits into
masterfrom
codex/issue-160-161-stats-tmp

Conversation

@dborup-agent

@dborup-agent dborup-agent commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Relates to #160, #161

Problem

Both issues concern the ingestor's stats temp file (<stats path>.tmp, cmd/ingestor/stats_file.go).

Plan / what changed

  1. Tests first (commit test(...), red on master): FIFO without a reader, FIFO with a reader, stop() with a FIFO at tmp, a broken tmp in the writer loop, the owner-error hint, and the stale marking in /api/mqtt/status.
  2. Fix (commit fix(...)):
    • writeStatsAtomic opens with O_NONBLOCK. A FIFO without a reader fails at once with ENXIO instead of blocking. Regular-file I/O ignores the flag. On Windows oNonBlock is 0, like oNoFollow.
    • After the open, the existing f.Stat() is used to refuse anything that is not a regular file, for example a FIFO that has a reader or a device. The error is <tmp>: not a regular file (mode …); remove it. When the open itself fails and an Lstat shows a non-regular entry, the open error comes first: <tmp>: open: <errno>; not a regular file (mode …); remove it. The entry is left in place for the operator to remove. The FileInfo is passed to checkStatsTmpOwner, so there is no second fstat.
    • The owner error is <tmp>: owned by uid X, ingestor uid Y; remove it or fix its owner.
    • When open itself fails with EACCES/EPERM, the usual fix(ingestor): foreign-owned stats .tmp logs an error every second with no hint, and stats go stale silently #160 case (a 0600 tmp of another service user and a non-root ingestor), an Lstat of the tmp gives the same hint with the owner, <tmp>: open: permission denied; owned by uid X, ingestor uid Y; remove it or fix its owner. For an unopenable tmp of the ingestor's own user, it gives …; mode -r--------; remove it or fix its permissions instead. The Lstat runs only after a failed open.
    • Every writeStatsAtomic error is a statsWriteError. It names the tmp path once, at the start, followed by the failed step with its cause (the path stripped from the *PathError/*LinkError), what is wrong with the tmp, and what to do. Unwrap returns the cause, so errors.Is still sees the errno.
    • New statsWriteLog, owned by the writer goroutine:
      • The first failure is logged at once, as [stats-file] write failed: <error>.
      • After that, a failure line comes at most once a minute, whether the failure persists or alternates with successes (flapping), with (N more failed writes since the last report).
      • The first success after a logged failure line logs [stats-file] write <path>: ok again after N failed writes, once. A failure episode inside the interval is not logged on its own: the next failure line counts it. So a flapping failure gives at most one failure line and one recovery line per minute.
      • The interval runs on the monotonic clock: the writer passes time.Now(), not the UTC tick time. A negative difference, from a wall-clock step back, counts as an interval passed, so a step cannot silence a persisting failure. SampledAt keeps the UTC tick time.
      • A healthy writer logs nothing, as before.
    • Server, read-only: /api/mqtt/status gains stale (bool) and sampleAgeSec (int, omitted when there is no readable sampleAt).
      • The rule is the one from /api/perf/io, now extracted as ingestorStatsStale(ts, now) (> IngestorStatsStaleThreshold, 5s) and used at all three sites.
      • An unparseable sampleAt counts as stale.
      • With no stats file, stale is false and sampleAt is empty, as before (no data). This also happens when the writer has failed from its very first tick, because then no stats file ever existed. A consumer must read an empty sampleAt as "nothing to trust", not as fresh data.
      • The OpenAPI description is updated.

The server only reads the stats file, and all writes stay in cmd/ingestor. No new map[string]interface{}; the new tests use map[string]any only for fixture assertions. The 9 fork guards in deploy.yml are unchanged.

Perf

This is not a hot path. The writer ticks at 1 Hz, and the change adds no syscall per tick: it adds a flag to an open, which already happened, and reuses the fstat that the owner check already did. The limiter costs O(1) per tick. /api/mqtt/status adds one time.Parse per request.

Tests

  • TestWriteStatsAtomicFIFOWithoutReaderFailsFast_161 (Unix only, 2s timeout; the FIFO gets a reader if the call blocks, so no goroutine leaks)
  • TestWriteStatsAtomicFIFOWithReaderRefused_161
  • TestStatsFileWriterStopsWithFIFOAtTmp_161 (stop within 3s)
  • TestStatsFileWriterLogsWriteFailureOnce_160 (writer loop at 5 ms: one failure line, then one recovery line)
  • TestStatsWriteLogAtMostOncePerInterval_160 (injected clock: 150 s of 1 Hz failures give 3 lines)
  • TestWriteStatsAtomicForeignTmpErrorNamesTheFix_160
  • TestMqttStatusMarksStaleStatsFile_160 (hour-old, 2× threshold, fresh, unparseable, no file)

Round 2 (review findings F1–F4):

  • TestWriteStatsAtomicUnopenableForeignTmpNamesTheFix_160 and TestWriteStatsAtomicUnopenableOwnTmpNamesTheFix_160 (F1: a 0400 tmp, EACCES from open)
  • TestStatsWriteLogFlappingAtMostOncePerInterval_160 (F2: 120 s of alternating failure/success at 1 Hz give 4 lines; before, 120)
  • TestStatsWriteLogBackwardClockStepStillLogs_160 (F3: a 1 h step back; before, 1 line in 190 s)
  • TestStatsWriteErrorNamesThePathOnce_160 and TestStatsFileWriterFailureLineNamesThePathOnce_160 (F4: FIFO, directory, foreign owner, unopenable foreign tmp, and the writer's line)

Mutants are listed in the report comment.

Not in scope

  • A hard link at <tmp> to another file of the ingestor's own user passes the regular-file and owner checks, and that file is then overwritten. This is pre-existing on master and was found in review (F6). A follow-up issue is proposed in the round-2 report.

  • The Observers panel (public/mqtt-status-panel.js) does not render stale yet. That is a UI follow-up that needs browser validation.

  • /api/perf/write-sources and /api/healthz read the same file. The liveness data in /api/healthz carries its own unix timestamps, and neither endpoint was asked for here.

🤖 Generated with Claude Code

dborup and others added 2 commits October 4, 2026 08:24
…t staleness (#160, #161)

Red on master:
- a FIFO at <stats>.tmp blocks writeStatsAtomic and the writer's stop;
- a FIFO with a reader is not refused with a clear error;
- a broken tmp path logs a failure on every tick and no recovery line;
- the foreign-owner error carries no hint;
- /api/mqtt/status has no stale/sampleAgeSec marking.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… status (#160, #161)

#161: writeStatsAtomic opens the tmp with O_NONBLOCK, so a FIFO without a
reader fails with ENXIO instead of blocking the writer (and its stop,
which waits on the goroutine) forever. Anything that is not a regular
file is then refused via the fstat the owner check already did, with
"<tmp> is not a regular file (mode ...); remove it".

#160: the owner error names the fix ("remove <tmp> or fix its owner").
statsWriteLog logs the first failure at once, a persisting one at most
once a minute with a count, and the first success after it once. A
healthy writer still logs nothing.

/api/mqtt/status (read-only) gains stale and sampleAgeSec, using the
/api/perf/io rule, now shared as ingestorStatsStale.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dborup-agent

Copy link
Copy Markdown
Collaborator Author

Rapport — CS-pve-agent1 PR#216 #160+#161 — head e3789e4

Status: All acceptance criteria are met and every CI job that ran is green. The PR is a draft, ready for review. It has not been merged or marked ready.

Evidence tags: [T] test or CI, [A] analysis, [K] known, not re-run.

Commits

  1. 2892c609 tests: these are red on master. Six new tests failed against master code [T]:
    • the FIFO without a reader blocked for 2s;
    • stop hung;
    • the writer logged 59 failure lines in 300 ms;
    • the owner error had no hint;
    • /api/mqtt/status had no stale field.
  2. e3789e4c fix, plus the unit test for the new limiter type.

#161: FIFO at <stats>.tmp

Criterion Result
A FIFO without a reader returns an error quickly instead of blocking (Unix-only test with a timeout) Met [T]. TestWriteStatsAtomicFIFOWithoutReaderFailsFast_161 has a 2s guard. If the call blocks, the test adds a reader to release the goroutine and then fails. With the fix the call returns in under 1 ms.
O_NONBLOCK, then Fstat, refusing anything that is not a regular file with a clear error Met [T]. The open fails with ENXIO and the message is <tmp> is not a regular file (mode p…); remove it: …. A FIFO with a reader opens but is refused after Fstat: TestWriteStatsAtomicFIFOWithReaderRefused_161. Nothing is published.
Shutdown hang (stop waits on <-done) Confirmed on master [T]: stop hung for more than 3s. Fixed: TestStatsFileWriterStopsWithFIFOAtTmp_161 stops in about 0.1s.
A regular tmp still works, including the #118 owner check Met [T]. The existing tests …RefusesForeignTmp_118, …TruncatesOwnStaleTmp_118, …SymlinkAtDestIsReplaced and the other stats-file tests are green.
Windows build oNonBlock = 0, like oNoFollow [A]. GOOS=windows go build and GOOS=freebsd go build of the ingestor pass, and so does GOOS=darwin go vet [T].

#160: foreign-owned or unusable tmp, and staleness

Criterion Result
At most one log line per interval, with a hint Met [T]. The writer loop at 5 ms logs exactly one failure line, then exactly one recovery line (TestStatsFileWriterLogsWriteFailureOnce_160). The limiter with an injected clock turns 150 s of 1 Hz failures into 3 lines (at 0, 60 and 120 s, each with "59 more failed writes"), and the recovery is logged once (TestStatsWriteLogAtMostOncePerInterval_160). The interval is statsWriteErrLogEvery = 1m.
The hint names the fix Met [T]. The message is <tmp> belongs to uid X, not Y; remove <tmp> or fix its owner (TestWriteStatsAtomicForeignTmpErrorNamesTheFix_160).
One log line when writes work again Met [T]. The line is [stats-file] write <path>: ok again after N failed writes.
A stale stats file is marked in /api/mqtt/status Met [T]. The response gains stale and sampleAgeSec. The rule is the one /api/perf/io uses, > IngestorStatsStaleThreshold (5s), extracted as ingestorStatsStale and used at all three call sites. An unparseable sampleAt counts as stale. With no stats file, stale is false, because there is no data at all (TestMqttStatusMarksStaleStatsFile_160).
No change when the tmp is owned correctly Met [T][A]. The success path logs nothing, as before. The open adds one flag, and the existing fstat is reused, so the owner check has no second fstat.
Server is read-only Met [A]. Only os.ReadFile and time.Parse were added under cmd/server, and all writes stay in cmd/ingestor.

Tests

  • [T] cd cmd/ingestor && go test ./... passed (ok, 510s), and so did cd cmd/server && go test ./... (ok, 34s). Both ran with TMPDIR on tmpfs locally. With /tmp on the local disk, both suites hit the timeout, with no failing test and a different test running at each timeout. The cause is SQLite fsync IO on this host, not this change. CI ran the suites normally.
  • [T] -race -count=3 passed on the affected tests: in the ingestor the _160/_161/stats-file/ingestor: assign explicit, collision-resistant MQTT client IDs #118 tmp tests, in the server the MqttStatus/PerfIO/_160 tests.
  • [T] sh test-all.sh: 214 passed, 0 failed.
  • [T] Mutants: all 7 were red with the mutant and green on the code.
    Mutant Killed by
    M1: drop O_NONBLOCK the FIFO-without-reader test and the stop test
    M2: drop the regular-file Fstat check the FIFO-with-reader test
    M3: no rate limit the writer-loop test and the limiter unit test
    M4: no recovery line the writer-loop test and the limiter unit test
    M5: no owner hint the hint test
    M6: stale threshold ×1000 the mqtt stale test
    M7: an unparseable sampleAt counted as fresh the mqtt stale test
  • [A] No new map[string]interface{}; the tests use map[string]any only for fixture assertions. The 9 fork guards in deploy.yml are unchanged (counted).

CI (run on head e3789e4)

  • [T] Go Build & Test: pass
  • [T] Playwright E2E Tests: pass
  • [T] Build & Publish Docker Image: pass
  • [T] Release Artifacts, Publish Badges & Summary, Deploy Staging: skipped. This is expected on a PR and on a fork.

Remaining

  • The Observers panel (public/mqtt-status-panel.js) does not render stale yet. That is a UI follow-up that needs browser validation; there was no browser validation here because there is no UI change.
  • /api/perf/write-sources serves the same file without a stale marker. /api/healthz's ingest_liveness carries its own unix timestamps. Neither endpoint was in scope.
  • A non-regular tmp is refused and left in place for the operator; there is no automatic removal. This is the conservative option from fix(ingestor): a FIFO at the stats .tmp path blocks the stats writer forever #161.

@adminopenclaw8-sketch

Copy link
Copy Markdown
Collaborator

Review — CS-Minimax PR#216 stats-tmp — head e3789e4

Dom: APPROVE med nits

Independent, read-only review of head e3789e4c and of its merge into origin/master 0f88865b (merged tree 093a1d62, from git merge-tree --write-tree). Both trees were unpacked with git archive into scratch. Platform: go1.27.0 darwin/arm64.

Evidence tags: [T] run here, [A] analysis of the source, [K] taken from the author's report or CI, not re-run.

Findings

# Severity Finding Evidence
F1 should-fix, non-blocking The #160 hint is missing when open itself fails with EACCES, and this is probably the common #160 case. A tmp left by the previous service user has mode 0600, because the writer creates it that way. A non-root ingestor then cannot open it, so checkStatsTmpOwner never runs and the log line is the bare open <tmp>: permission denied. The hint appears only when the open succeeds: the ingestor runs as root, or the foreign tmp happens to be writable. The rate limit does apply to this case. [T][A]
F2 nit A flapping failure is not rate-limited. After a success, the next failure is logged at once. Failures that alternate with successes at 1 Hz give 120 lines in 120 s (a failure line and an "ok again" line each time) [T]. Master would log the 60 failures only, since it logs every failure and nothing on success [A]. [T][A]
F3 nit The limiter is driven by tickAt := time.Now().UTC(), and UTC() strips the monotonic reading. A backward wall-clock step therefore silences a persisting failure for the size of the step plus 1 min. With a 1 h step back, 3600 s of failures gave only the first line. Passing a plain time.Now() to writeLog.failed fixes this; SampledAt can keep tickAt. [T]
F4 nit A failure line repeats the full path up to three times, for example [stats-file] write <path>: <path>.tmp is not a regular file (mode prw-------); remove it: open <path>.tmp: device not configured. The owner error also names <tmp> twice. This is cosmetic: the path is the operator's own stats path and no secret, and the rate limit keeps the volume down. [T]
F5 info Staleness is visible only in the JSON of /api/mqtt/status (stale, sampleAgeSec). No UI renders it yet, which the PR states. /api/perf/io already dropped stale samples and is unchanged. /api/healthz and /api/perf/write-sources are unchanged. One edge case: if the writer fails from the very first tick, no stats file ever exists, and the response is stale:false with empty sampleAt and no rows. This is documented as "no data", but a consumer must read an empty sampleAt as "nothing to trust". [T][A]
F6 info, pre-existing, outside the delta A hard link at <path>.tmp to another file of the ingestor's own user passes the regular-file and owner checks. That file is then truncated and overwritten with the stats JSON; this probe behaves the same on master. Planting it needs write access to the stats directory (by default /tmp) and, on Linux with fs.protected_hardlinks=1, write access to the target, so it is low risk. Refusing nlink > 1 would close it. This is a follow-up candidate, not something for this PR. [T][A]

1. #160: rate-limited error, hint, stale status

  • Rate limit. statsWriteLog is owned by the writer goroutine.
    • The first failure is logged at once.
    • A persisting failure is logged at most once per statsWriteErrLogEvery (1 min), with the count of suppressed failures.
    • Recovery is logged once.
    • The healthy path logs nothing, as before.
    • The codebase has no shared log-throttle helper for this, and the pattern (first line at once, then a counted line per interval) matches the watchdog's style. [A]
    • The injected-clock unit test and the writer-loop test pin all of this. [T]
    • Gaps: flapping (F2) and wall-clock steps (F3).
  • Hint. The owner error now ends ; remove <tmp> or fix its owner. The not-regular error ends ; remove it. Neither contains a secret: only the stats path and two uids. [T] The hint is missing on the EACCES path (F1).
  • Where staleness shows. Only in /api/mqtt/status; see F5. The rule is the /api/perf/io one, extracted into ingestorStatsStale and used at all three call sites. A mutant that multiplies the threshold by 1000 fails both TestMqttStatusMarksStaleStatsFile_160 and the existing TestReadIngestorIOSample_StaleBeyondThreshold. [T] The writer interval is fixed at 1 s (main.go), so a healthy sampleAt is at most about 2 s old, well under the 5 s threshold, and stale does not flicker. [A]
  • Back to normal. succeeded() logs ok again after N failed writes once and clears failing, failures and suppressed. A new failure after that is logged at once and counted afresh (unit test). [T] On the server side, stale is computed per request from sampleAt, so it clears as soon as the file is fresh again. [A]

2. #161: FIFO-safe writer

  • No blocking.
    • The writer opens with O_CREATE|O_WRONLY|O_NOFOLLOW|O_NONBLOCK. A FIFO without a reader fails at once with ENXIO.
    • f.Stat() on the open descriptor refuses anything that is not a regular file, such as a FIFO that has a reader, or a device.
    • The same FileInfo feeds checkStatsTmpOwner, so there is no extra fstat. [T][A]
  • TOCTOU.
    • The regular-file and owner checks use the open descriptor, so no window opens between check and use for the data written. [A]
    • The Lstat in the open-failure branch only chooses the error text. [A]
    • The rename(tmp, path) by name after the checks is unchanged from master. Swapping the tmp in that window needs the right to replace the ingestor's own file; with the sticky bit on /tmp, other users do not have it. [A]
  • Symlinks.
    • A symlink at the tmp gives ELOOP, reported as … is not a regular file (mode L…); remove it: …, and the target is unchanged. [T]
    • A symlink at the destination is replaced by the rename (existing TestWriteStatsAtomic_SymlinkAtDestIsReplaced, green). [T]
    • The hard-link case is F6.
  • Non-unix.
    • oNonBlock is syscall.O_NONBLOCK under //go:build !windows, and 0 on Windows (like oNoFollow). [A]
    • Cross-builds of the ingestor (CGO_ENABLED=0) give identical results on master and on the merged tree:
      • ok: linux/amd64, linux/arm64, darwin/arm64, freebsd, openbsd, windows;
      • fails on both, because of modernc.org/sqlite/libc: netbsd, solaris, illumos, aix, plan9, js/wasm, wasip1. [T]
    • GOOS=windows go test -c fails on both trees in channel_proposals_test.go and hot_reload_test.go (syscall.Kill). That is pre-existing and not in this PR. [T]
  • Atomic rename and cleanup. Unchanged: write → close → rename, and the tmp is removed if write, close or rename fails. A non-regular or foreign tmp is deliberately left for the operator. A probe of three normal writes gives identical bytes, mode 0600, a regular file, and no tmp left behind, on both master and the merged tree. [T]

3. Normal operation

  • Snapshot, encoding (trailing-newline strip), interval (StartStatsFileWriter(store, time.Second)), path (CORESCOPE_INGESTOR_STATS or /tmp/corescope-ingestor-stats.json) and file mode are unchanged. [A]
  • The write probe is byte-identical on master and the PR. [T]
  • The only visible change on the healthy path is the two new additive fields in /api/mqtt/status. [A]

4. Concurrency and shutdown

  • stop() still closes quit once and waits on done. The goroutine can now only wait in select or in a non-blocking-on-FIFO write, so stop() cannot hang on a FIFO. [T][A]
  • statsWriteLog is used only by the writer goroutine; log.Printf is safe for concurrent use. [A]
  • The new and nearby tests are clean under -race -count=5. [T]

5. Performance

  • This is not a hot path: 1 tick per second. [A]
  • Syscalls per healthy tick are unchanged: open (one more flag), the same single fstat, fchmod, ftruncate, write, close, rename. Lstat runs only after a failed open. [A]
  • The limiter does not allocate: testing.AllocsPerRun gives 0 for a suppressed failed() and 0 for a healthy succeeded(). [T]
  • The server adds one time.Parse per /api/mqtt/status request. [A]

6. Tests, mutants and FIFO reproduction

Run Result
cd cmd/ingestor && go test -race -count=1 ./... on the merged tree ok (506 s) [T]
go test -race -count=5 on the 6 new ingestor tests plus the stats-file, WriteStatsAtomic and _118 tests (39 tests), merged tree ok [T]
go test -race -count=5 on _160, MqttStatus, PerfIO and Ingestor tests in cmd/server (19 tests), merged tree ok [T]
cd cmd/server && go test -count=1 ./... on the merged tree (no -race) ok (45 s) [T]
go vet for ingestor (darwin and linux) and server ok [T]
gofmt -l on the changed files clean, except cmd/server/perf_io.go, which is unformatted on master too (pre-existing; the PR's hunk is formatted) [T]
CI on head (Go Build & Test, Playwright, Docker) success [K]

FIFO reproduction. A probe test, not committed, on a master copy and on the merged tree. It used a real mkfifo at <tmpdir>/stats.json.tmp, a 2 s guard for writeStatsAtomic and a 3 s guard for stop().

master 0f88865b merged tree
writeStatsAtomic with a FIFO and no reader blocked > 2 s. When a reader appeared it returned nil, and the FIFO itself was renamed into place as the stats file (prw-------), so a server reading it would block in turn [T] returned at once: … is not a regular file (mode prw-------); remove it: open …: device not configured, and nothing was published [T]
StartStatsFileWriter (20 ms) with a FIFO at the tmp, then stop() hung > 3 s; returned only after a reader was attached [T] returned at once [T]

Mutants. Each was applied to a copy of the merged tree and run under -race -count=2 against the ingestor set above, or the server set for M6–M7.

# Mutant Result
M1 regular-file check after fstat removed killed: FIFOWithReaderRefused_161 [T]
M2 oNonBlock = 0 on unix killed: FIFOWithoutReaderFailsFast_161, StopsWithFIFOAtTmp_161 [T]
M3 rate limit removed killed: LogsWriteFailureOnce_160, AtMostOncePerInterval_160 [T]
M4 recovery does not clear failing killed: LogsWriteFailureOnce_160, AtMostOncePerInterval_160 [T]
M5 open-failure Lstat diagnosis removed killed: FIFOWithoutReaderFailsFast_161 [T]
M6 /api/mqtt/status never marks stale killed: MqttStatusMarksStaleStatsFile_160 [T]
M7 shared stale threshold × 1000 killed: MqttStatusMarksStaleStatsFile_160 and ReadIngestorIOSample_StaleBeyondThreshold [T]

The author's mutants M1–M7 were not re-run. [K]

7. Rules

  • Changed files against master: 5 in cmd/ingestor and 4 in cmd/server (mqtt_status.go, openapi.go, perf_io.go, plus a test). [T]
    • The server changes are the expected read-only stale marker: the non-test hunks add no file writes, renames, removes or SQL writes. [T]
    • .github is unchanged. [T]
  • No new map[string]interface{} (0 added lines), and no map[string]any outside tests. [T]
  • Fork guards: deploy.yml 9 and release-fast-path.yml 1 matches of github.repository == 'Kpa-clawbot/CoreScope', identical to master. [T]
  • No closing keywords in the title, body or the 2 commit messages ("Relates to" only). [T]
  • Both commits have author and committer dborup <kontakt@meshview.dk>. [T]

Not verified

  • Linux behaviour: all runs were on darwin/arm64. ENXIO for a FIFO without a reader is POSIX and the same on Linux [A], but it was not run there. CI covers the new tests on Linux. [K]
  • A tmp that is really foreign-owned (another uid). That needs root here. F1 is shown with a tmp of our own that we cannot write (0400), which takes the same EACCES path. [T][A]
  • Device files at the tmp path, which also need root to create.
  • Staging and prod, browser and UI (the PR has no UI change).

The head was e3789e4c936ade274ba872ad5ab349af7c8ca199 before this review (git ls-remote) and still e3789e4c936ade274ba872ad5ab349af7c8ca199 after it. The PR is still a draft and was not modified.


Generated by Claude Code

dborup and others added 9 commits October 4, 2026 11:16
…160)

The usual #160 case is a 0600 tmp left by another service user and a
non-root ingestor: open(2) fails with EACCES before checkStatsTmpOwner
runs, so the log line was the bare "open <tmp>: permission denied". Pin
the hint and the owner uids for a foreign tmp, and a permissions hint
for an unopenable tmp of the ingestor's own user. (PR #216 review, F1)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…#160)

When open(2) on the tmp fails with EACCES/EPERM, an Lstat of the tmp
chooses the hint: "remove <tmp> or fix its owner (owned by uid X,
ingestor uid Y)" for another user's file, "... or fix its permissions
(mode ...)" for the ingestor's own. The permission error stays in the
%w chain. The Lstat runs only after a failed open; a healthy tick is
unchanged. (PR #216 review, F1)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…160)

After a success the next failure was logged at once, so failures that
alternate with successes at 1 Hz gave 120 lines in 120 s (a failure
line and an "ok again" line each time). Pin at most one failure line
and one recovery line per interval, with the failures in between
counted, and update the limiter test's second episode to match.
(PR #216 review, F2)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
statsWriteLog now keeps the last failure line's time across episodes:
a failure line comes at most once per interval, whether the failure
persists or alternates with successes, and the "ok again" line follows
only a logged failure line. A failure episode inside the interval is
counted and reported with the next failure line. A logged recovery
reports the whole episode and clears that count. Persistent failures
log as before (150 s at 1 Hz: 3 lines). (PR #216 review, F2)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ilures (#160)

The writer drove the limiter with time.Now().UTC(), which drops the
monotonic reading, and the limiter took a negative difference to the
last failure line as "inside the interval": after a 1 h step back a
persisting failure was not logged for an hour. Pin a line at the step
and then one per interval. (PR #216 review, F3)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e a step back (#160)

The writer passes time.Now() to statsWriteLog.failed instead of the
UTC tick time, so the interval runs on the monotonic clock; SampledAt
and the source statuses keep the UTC tick time. The limiter also treats
a negative difference to the last failure line as an interval passed,
so even a wall-clock time cannot silence a persisting failure after a
step back. (PR #216 review, F3)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A failure line named the path up to three times, e.g. "[stats-file]
write <path>: <tmp> is not a regular file (mode ...); remove it: open
<tmp>: ...". Pin that each writeStatsAtomic error (FIFO, directory,
foreign owner, unopenable foreign tmp) starts with the tmp path and
names it once, and that the writer's failure line adds no second copy.
The foreign-owner hint is now "remove it or fix its owner".
(PR #216 review, F4)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
writeStatsAtomic now returns a statsWriteError for every failed step.
Its message starts with the tmp path and does not repeat it: the failed
step with its cause (the path stripped from the *PathError/*LinkError),
then what is wrong with the tmp and what to do, e.g.

  <tmp>: open: permission denied; owned by uid 1000, ingestor uid 1001; remove it or fix its owner
  <tmp>: open: no such device or address; not a regular file (mode prw-------); remove it

Unwrap returns the cause, so errors.Is still sees the errno. The
writer's line for such an error is "[stats-file] write failed: <err>",
without the path in front; the "ok again" line is unchanged. This
replaces errStatsTmpNotRegular and statsTmpPermissionHint. Only the
message text changes; what is refused, removed or published is the
same. (PR #216 review, F4)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dborup-agent

Copy link
Copy Markdown
Collaborator Author

Rapport — CS-pve-agent3 PR#216 runde 2 — head 0bc53bc

Status: F1–F4 are fixed, each with a red test first and killed mutants. F5 is documented in the PR description, and F6 is proposed as a follow-up issue below. CI is green on 0bc53bc4. The PR stays a draft.

Evidence tags: [T] test or run output, [K] checked in code, diff or CI log, [A] assessment or inference.

Review feedback addressed (commit 0bc53bc4)

Branch: origin/master (a0086bdd) was merged in with a merge commit (6cc3f652). Then one test commit (red) and one fix commit per finding followed. There was no rebase, amend or force-push. [K]

  1. F1: the fix(ingestor): foreign-owned stats .tmp logs an error every second with no hint, and stats go stale silently #160 hint when open fails with EACCES/EPERM. Commits 29c794ed (test) → e191d6d6 (fix).
    • After a failed open with a permission error, an Lstat of the tmp chooses the hint:
      • for another user's file: <tmp>: open: permission denied; owned by uid X, ingestor uid Y; remove it or fix its owner;
      • for the ingestor's own unopenable file: …; mode -r--------; remove it or fix its permissions.
    • The Lstat runs only after a failed open, and the permission error stays reachable with errors.Is. [K]
    • Tests: TestWriteStatsAtomicUnopenableForeignTmpNamesTheFix_160 and …OwnTmpNamesTheFix_160. Both use a 0400 tmp, so open fails with EACCES, and skip when running as root. Before the fix they failed with the bare open <tmp>: permission denied. [T]
    • Mutants: F1a (no hint on ErrPermission) and F1b (owner comparison inverted) are both killed. [T]
  2. F2: flapping is rate-limited. Commits b55511ab → 1474234b.
    • The limiter keeps the time of the last failure line across episodes. A failure line comes at most once per interval, and "ok again" follows only a logged failure line. An episode inside the interval is counted into the next failure line. A logged recovery reports the whole episode and clears the count. [K]
    • Test: TestStatsWriteLogFlappingAtMostOncePerInterval_160 with an injected clock, failure and success alternating at 1 Hz for 120 s. It gave 120 lines before and 4 now:
      • failure;
      • ok;
      • failure with "29 more failed writes";
      • ok. [T]
    • The existing limiter test's second episode was updated to the new rule; a persisting failure still gives 3 lines in 150 s. [T]
    • Mutants: F2a (lastLog reset on success), F2b (recovery line always logged) and F2c (count not cleared by a logged recovery, which double-counts) are all killed. [T]
  3. F3: monotonic clock and backward steps. Commits 050a858d → e0684783.
    • The writer passes time.Now() to writeLog.failed; SampledAt and the source statuses keep the UTC tick time. [K]
    • The limiter also treats a negative difference to the last failure line as an interval passed, so even a wall-clock time cannot be silenced by a step back. [K]
    • Test: TestStatsWriteLogBackwardClockStepStillLogs_160, a 1 h step back with an injected clock. It gave 1 line in 190 s before, and now a line at the step and then one per minute. [T]
    • Mutant F3a (the d >= 0 guard removed) is killed. [T]
    • Mutant F3b (passing tickAt again) survives. With the guard, the monotonic reading changes no observable behaviour, and a test cannot see it without adding a clock seam to the writer. [T][A]
  4. F4: the path is named once. Commits 90bcd67c → 0bc53bc4.
    • Every writeStatsAtomic error is a statsWriteError: the tmp path first, then the failed step with its cause (the path stripped from the *PathError/*LinkError), what is wrong, and what to do. Unwrap returns the cause. [K]
    • The writer's line is [stats-file] write failed: <err>; the "ok again" line is unchanged. [K]
    • Examples from a run [T]:
      • <tmp>: open: no such device or address; not a regular file (mode prw-------); remove it
      • <tmp>: owned by uid 1000, ingestor uid 1001; remove it or fix its owner
    • Tests: TestStatsWriteErrorNamesThePathOnce_160 (FIFO, directory, foreign owner, unopenable foreign tmp) and TestStatsFileWriterFailureLineNamesThePathOnce_160. Before the fix the path appeared 2–3 times. [T]
    • The owner-hint test now expects remove it or fix its owner, because the tmp path leads the message. [K]
    • Mutants: F4a (path not stripped from the cause), F4b (path put in front of the line again) and F4c (Unwrap returns nil) are all killed. [T]
    • Two of the original FIFO mutants were re-checked against the new error type: the regular-file check removed, and oNonBlock = 0. Both are killed. [T]
  5. F5: no code change. The PR description now says that an empty sampleAt (no stats file, including a writer that has failed since its first tick) means "nothing to trust", not fresh data. [K]
  6. F6: no code change; follow-up proposed, not created.
    • Title: fix(ingestor): refuse a hard-linked stats .tmp (nlink > 1)
    • Text: A hard link at <stats path>.tmp to another file of the ingestor's own user passes the regular-file and owner checks in writeStatsAtomic. That file is then truncated and overwritten with the stats JSON, and the rename publishes it. This is pre-existing on master and was found in the fix(ingestor): rate-limited stats tmp errors, FIFO-safe writer, stale status (#160, #161) #216 review (F6). Planting it needs write access to the stats directory (by default /tmp) and, with fs.protected_hardlinks=1, access to the target, so the risk is low. Proposed fix: after the open, refuse a tmp whose Fstat shows nlink > 1, leave it in place, and use the same statsWriteError shape (…; hard-linked (nlink N); remove it). Acceptance: a Unix test with a hard link at the tmp gets an error, the link target is unchanged, and nothing is published; the existing stats-file tests stay green. [A]

Behaviour outside F1–F4 is unchanged: the stats file content, format, 1 s interval, path and mode, and what is refused, removed or published. Only log and error text and the limiter's timing changed. [K]

Tests

  • [T] cd cmd/ingestor && go test -race -count=1 -timeout 20m ./...: ok (763 s; TMPDIR on tmpfs).
  • [T] -race -count=5 on the new and affected ingestor tests (_160, _161, _118, StatsWriteLog, StatsAtomic, StatsFile, NamesThePathOnce, Unopenable, SymlinkAtDest): ok.
  • [T] cmd/server is untouched in round 2 (0 changed lines since the merge commit). After the merge, go test ./... passes (35 s), and -race on the _160/MqttStatus/PerfIO/stale tests passes too.
  • [T] go vet (ingestor) passes, and GOOS=windows go build (ingestor) passes.
  • [T] gofmt -l on the changed files flags only cmd/server/perf_io.go, which is unformatted on master too and not touched in this round.
  • [T] sh test-all.sh: 214 passed, 0 failed.
  • [K] Fork guards are unchanged: deploy.yml 9, release-fast-path.yml 1. There are no new map[string]interface{}, and cmd/server stays read-only.

CI (run 37199526893, head 0bc53bc4)

Job Result
Go Build & Test pass [T]
Playwright E2E Tests pass [T]
Build & Publish Docker Image pass [T]; local image build only, no push [K]
Release Artifacts / Deploy Staging / Publish Badges & Summary skipped (PR) [T]

Remaining

  • [A] F3b: the monotonic reading itself is not pinned by a test (see item 3).
  • [A] F6 is still open as a proposed follow-up; it is not part of this PR.
  • [A] As before: the Observers panel does not render stale yet, and /api/perf/write-sources and /api/healthz carry no stale marker.

@dborup-agent

Copy link
Copy Markdown
Collaborator Author

Review — CS-pve-agent2 PR#216 runde 2 — head 0bc53bc

Dom: APPROVE med nits

Independent, read-only re-review of head 0bc53bc464a7311c131b60db9470f26535421ae7 and of its merge into origin/master 376d51c8. The merged tree is e854386d, from git merge-tree --write-tree (clean). Head, the merged tree and master were unpacked with git archive into scratch. No worktree was made on the branch.

Platform: go1.27.1 linux/amd64, run as a non-root user (uid 1000), with TMPDIR on tmpfs. The host has fs.protected_regular=2, fs.protected_fifos=1 and fs.protected_hardlinks=1. To create real foreign-owned files (uid 65534) inside my own scratch directories, I used local sudo chown. The ingestor itself always ran as uid 1000.

Evidence tags: [T] run here, [A] analysis of the source, [K] taken from the author's report or CI, not re-run.

Round 1 findings

# Status Evidence
F1: hint when open fails with EACCES/EPERM Fixed. Verified with a real foreign-owned 0600 tmp and with a foreign 0666 tmp in a sticky directory, run as non-root. [T][A]
F2: flapping is rate-limited Fixed. I found no silent persisting failure and no double or lost count; there is one bounded caveat (N4). [T][A]
F3: monotonic clock and backward steps Fixed. The d >= 0 guard cannot cause a burst, and survivor F3b is acceptable. [T][A]
F4: the path is named once Fixed on every error path, including rename, write and close. The rename path has no test (N2). [T][A]
F5: staleness only in JSON, and the empty-sampleAt edge Documented in the PR description. There is no code change, and the server is untouched since round 1. [A]
F6: hard-linked tmp Out of scope, now tracked as #228. [K]

New findings

# Severity Finding Evidence
N1 nit (tests) Two new F4 tests drive a FIFO without the release guard that the _161 tests use (writeStatsAtomicOrRelease): the FIFO subtest of TestStatsWriteErrorNamesThePathOnce_160 calls writeStatsAtomic directly, and TestStatsFileWriterFailureLineNamesThePathOnce_160 calls stop() unguarded. If #161 regresses (my mutant R12, oNonBlock dropped from the open), the package therefore hangs in TestStatsWriteErrorNamesThePathOnce_160/FIFO_without_a_reader until the package timeout, instead of failing in 2–3 s. The _161 tests alone catch R12 in 2.0 s and 3.1 s. Suggested fix: route the FIFO case through writeStatsAtomicOrRelease, and the writer test through a guarded stop. [T]
N2 nit (test gap) F4's rename, write and close paths are implemented correctly but not tested. Mutant R10, which stops stripping the *os.LinkError, survives. A rename-failure probe (a non-empty directory at the destination) gives <tmp>: rename to stats.json: file exists on the PR code. With R10 it gives <tmp>: rename to stats.json: rename <tmp> <path>: file exists, which names the path three times. Adding that case to TestStatsWriteErrorNamesThePathOnce_160 would pin it. Write and close are hard to fail in a test and are covered by the same code path. [T]
N3 info Mutant R7, which gives the permission hint for any open error on a regular tmp, survives. The current code is right: an existing regular tmp that fails with something other than a permission error (such as EROFS or ETXTBSY) gets no misleading owner or mode hint. This is not pinned, and a test would need a read-only mount, so I would not require one. [T][A]
N4 info This is a consequence of F2's design. After a logged ok again, a new failure that starts within one interval of the last failure line is counted, not logged, so the log's last line can read ok again for up to about 1 min while writes fail again. The window is bounded; the next line reports the failures, and /api/mqtt/status shows stale after 5 s. Related wording: the recovery count covers only the current episode, which is the run of failures since the last success. In the flapping test, a line (29 more failed writes since the last report) is followed by ok again after 1 failed writes. That is consistent with the doc, but an operator may read it as a contradiction. [T][A]
N5 nit (DRY) The owner detail and hint (owned by uid %d, ingestor uid %d / remove it or fix its owner) are built in two places, setPermissionHint and checkStatsTmpOwner. checkStatsTmpOwner could reuse one helper. [A]
N6 info An unwritable stats directory with no tmp gives a bare <tmp>: open: permission denied without a hint, because Lstat finds nothing. That is correct, since the hint is about the tmp file, and it is outside #160. [T]

None of these block. N1 and N2 are small test hardenings that could also go to a follow-up.

1. F1: hint after a failed open

  • Lstat only after a failed open. strace of three healthy writes shows openat(tmp, O_WRONLY|O_CREAT|O_NONBLOCK|O_NOFOLLOW|O_CLOEXEC) and renameat per write. It shows no lstat/newfstatat of the tmp. The only path newfstatat is Go's os.Rename checking the destination, which is stdlib and pre-existing. The tmp's fstat runs on the descriptor. [T] Mutant R2 (Lstat before open) survives, as expected, because it is not observable in the output; the property holds by trace and by reading. [T][A]
  • A real foreign owner, run as non-root [T]:
    • A 0600 tmp owned by uid 65534 gives <tmp>: open: permission denied; owned by uid 65534, ingestor uid 1000; remove it or fix its owner.
    • A 0666 tmp owned by uid 65534 in a sticky world-writable directory: protected_regular makes O_CREAT fail with EACCES even though the mode allows writing, and the hint is the same. The foreign file's content is unchanged ("old").
    • A foreign FIFO in a sticky directory: protected_fifos gives EACCES, and the error is <tmp>: open: permission denied; not a regular file (mode prw-rw-rw-); remove it. The non-regular branch wins, which is right.
  • An unreadable tmp of the ingestor's own user. TestWriteStatsAtomicUnopenableOwnTmpNamesTheFix_160 (0400) gives …; mode -r--------; remove it or fix its permissions with no owner text. It passes here, run as non-root. [T]
  • errors.Is. errors.Is(err, fs.ErrPermission) and errors.Is(err, syscall.EACCES) are both true on the hint error with a real foreign owner. [T] Mutant R3 (Unwrap returns nil) is killed by TestWriteStatsAtomicUnopenableForeignTmpNamesTheFix_160. [T]
  • Leaks. The message holds only the tmp path, the errno text, two uids and a file mode. It contains no content, user names or environment. [T][A]
  • Branches. Mutant R4 (always the owner hint) is killed by the OwnTmp test. [T] A rarer case: an own tmp with the immutable flag (EPERM with a normal mode) would get mode -rw-------; …fix its permissions. That is vague but not wrong, so I take it as acceptable. [A]

2. F2: the limiter

  • Logic [A]:
    • failed counts the failure. It logs at once when no failure line was logged before, when a full interval has passed since the last failure line, or when the clock stepped back. Otherwise it adds to suppressed.
    • succeeded logs ok again only if the episode had a logged failure line (reported). It then clears failing, reported and failures. It clears suppressed only after a logged recovery, so an unlogged episode is carried into the next failure line.
  • Can a persisting failure go silent? No. lastLog changes only when a line is logged, and the monotonic now keeps growing, so d reaches every within one minute. With wall-clock times, a step back logs at once (the guard), and a step forward makes d larger. [A][T]
  • Double or lost counts. I walked through persisting failures, flapping, a failure line inside a later episode, and a recovery after one. The failure line's (N more …) counts exactly the failures that were not logged since the last failure line or logged recovery. ok again after N counts the current episode, including its logged failure. Nothing is counted twice or lost. [A] The mutants confirm it [T]:
    • R1 (lastLog reset on success) is killed.
    • R5 (suppressed cleared on every success, which loses unlogged episodes) is killed.
    • R9 (recovery logged after an unlogged episode) is killed.
    • The caveat is N4.

3. F3: monotonic clock and steps

  • The writer passes time.Now(), and the stored lastLog is that same value, so now.Sub(lastLog) uses the monotonic readings. Wall-clock steps in either direction do not affect the limiter inside a running process. [A]
  • Forward step. It cannot cause too many lines. With the monotonic clock it has no effect. With wall-clock values (tests, or F3b), a forward step makes at most one line come early, and lastLog then moves to the new time. A step back logs one line, after which the normal interval applies. Only a clock that stepped back on every tick could produce a line per tick, and the writer's clock cannot do that. [A]
  • d >= 0: mutant R6 (guard removed) is killed by TestStatsWriteLogBackwardClockStepStillLogs_160. [T]
  • F3b survives (R8: the writer passes tickAt again). [T] I find that acceptable: thanks to the guard, the worst a wall-clock step can do is move one line earlier. It cannot silence anything or cause a burst. Pinning the monotonic reading would need a clock seam in the writer only for this. [A]

4. F4: path named once

  • Every return in writeStatsAtomic is a *statsWriteError, or nil: open (with an optional not-regular or permission detail), stat, not regular, owner, chmod, truncate, write, close and rename. checkStatsTmpOwner returns an untyped nil on success, so there is no typed-nil trap. [A]
  • newStatsWriteError strips the path from *fs.PathError and *os.LinkError. Rename shows only filepath.Base(path). Unwrap gives the errno. [A][T]
  • The writer's line reads [stats-file] write failed: <err> for a statsWriteError, and mutant R11 (path put back) is killed. The ok again line names the path once. [T]
  • The rename text is verified by probe; its test is missing (N2). [T]

5. #161: FIFO safety (own reproduction with mkfifo)

Probe master 376d51c8 merged tree e854386d
writeStatsAtomic with a FIFO at the tmp and no reader (2 s guard) blocked > 2 s; after a reader appeared it failed in truncate (invalid argument), and the FIFO was left in place [T] returned at once: <tmp>: open: no such device or address; not a regular file (mode prw-------); remove it, with nothing published and the FIFO left in place [T]
StartStatsFileWriter (20 ms) with a FIFO at the tmp, then stop() (3 s guard) hung > 3 s [T] returned at once, with one line [stats-file] write failed: <tmp>: open: …; not a regular file …; remove it [T]

The rewrite has not weakened anything [A][T]:

  • O_NONBLOCK and O_NOFOLLOW are still in the open, as the strace above shows.
  • The regular-file check still runs on the descriptor, before the owner check, chmod, truncate or write.
  • A FIFO with a reader is still refused (TestWriteStatsAtomicFIFOWithReaderRefused_161).
  • Mutant R12 (oNonBlock dropped) is killed by both _161 tests in 2–3 s. The hang it causes in an unguarded new test is N1.

6. Normal operation unchanged

  • On the healthy path the only differences from master are the O_NONBLOCK flag and the fstat that moved from checkStatsTmpOwner into writeStatsAtomic, called once as before. Chmod 0600, truncate, write, close and rename are unchanged. [A]
  • A write probe on master and on the merged tree gives identical content, mode -rw-------, and no tmp left behind. [T]
  • IngestorStatsSnapshot, the encoding (trailing-newline strip), statsFilePath() (CORESCOPE_INGESTOR_STATS, default unchanged) and the interval passed by main.go are not touched by the PR's diff. Only the log and error text and the limiter changed in round 2. [A]

7. Rules

  • Round 2 changes only cmd/ingestor: three files since the merge commit. cmd/server was changed in round 1 only, and that change is the read-only stale marker: ingestorStatsStale, mqttStatusStaleness, OpenAPI text, and a test that writes a fixture stats file. [T]
  • No new map[string]interface{} (0 added lines), and no new map[string]any outside tests. [T]
  • The fork guard github.repository == 'Kpa-clawbot/CoreScope' appears 9 times in deploy.yml and 1 in release-fast-path.yml, at head and on the merged tree, and the PR does not touch .github/. [T]
  • There are no closing keywords in the title, the body or the 10 commit messages ("Relates to" only). [T]
  • All commits on the branch since master have author and committer dborup <kontakt@meshview.dk>. [T]
  • Merge 6cc3f652 (parents e3789e4c and a0086bdd) is a clean merge commit. Its tree f752089c is identical to git merge-tree --write-tree e3789e4c a0086bdd, so it carries no extra edits. [T]

Tests and mutants

Run Result
cd cmd/ingestor && go test -race -count=1 -timeout 20m ./..., merged tree, TMPDIR on tmpfs ok (780 s), 0 DATA RACE [T]
go test -race -count=5 -run '_160|_161', merged tree, ingestor (12 tests, none skipped as non-root) ok [T]
go test -race -count=5 -run '_160|MqttStatus|PerfIO|Stale', merged tree, server ok [T]
go vet (ingestor), GOOS=windows go build (ingestor) ok [T]
gofmt -l on the changed ingestor files clean [T]
CI run 37199526893 on head success [K]

The mutants were applied one at a time to a copy of the merged tree and run under -race -count=2 against _160|_161|_118|StatsWrite|StatsFile|WriteStatsAtomic:

# Mutant Result
R1 limiter resets lastLog on success killed: AtMostOncePerInterval_160, FlappingAtMostOncePerInterval_160
R2 Lstat before the open survives; not observable, and strace confirms the real code has no Lstat on the healthy path
R3 Unwrap returns nil killed: UnopenableForeignTmpNamesTheFix_160
R4 permission hint always names the owner killed: UnopenableOwnTmpNamesTheFix_160
R5 suppressed cleared on every success killed: AtMostOncePerInterval_160, FlappingAtMostOncePerInterval_160
R6 d >= 0 guard removed killed: BackwardClockStepStillLogs_160
R7 permission hint for any open error survives (N3)
R8 writer passes tickAt (the author's F3b) survives; acceptable (§3)
R9 recovery logged after an unlogged episode killed: AtMostOncePerInterval_160, FlappingAtMostOncePerInterval_160
R10 *os.LinkError path not stripped survives (N2)
R11 writer's line repeats the path killed: FailureLineNamesThePathOnce_160
R12 oNonBlock dropped from the open killed by FIFOWithoutReaderFailsFast_161 (2.0 s) and StopsWithFIFOAtTmp_161 (3.1 s); the full set hangs in the unguarded FIFO subtest (N1)

The author's mutants were not re-run. [K]

Not verified

  • Device files at the tmp path, and a tmp on a read-only filesystem (R7/N3), because they need mounts or device nodes.
  • An ingestor actually running as root: by analysis it reaches the owner check after a successful open, which round 1 already covered.
  • Platforms other than linux/amd64 at run time, although the Windows build compiles. macOS was covered in round 1. [K]
  • The full cmd/server suite and sh test-all.sh. The server is unchanged in round 2, and CI is green. [K]
  • Staging, prod, browser and UI (the PR has no UI change).

The head was 0bc53bc464a7311c131b60db9470f26535421ae7 before this review (git ls-remote) and still 0bc53bc464a7311c131b60db9470f26535421ae7 after it. The PR is still a draft and was not modified.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants