Skip to content

build(docker): Alpine 3.24.1 runtime with Mosquitto 2.1 compatibility - #57

Merged
dborup merged 4 commits into
masterfrom
codex/alpine-3241-runtime
Sep 16, 2026
Merged

dborup merged 4 commits into
masterfrom
codex/alpine-3241-runtime

Conversation

@adminopenclaw8-sketch

Copy link
Copy Markdown
Collaborator

Summary

Moves the Go runtime image (Dockerfile, Dockerfile.go) from alpine:3.20 to
alpine:3.24.1@sha256:28bd5fe8b56d1bd048e5babf5b10710ebe0bae67db86916198a6eec434943f8b,
and adds the two Mosquitto 2.1 compatibility fixes the upgrade needs. It also adds a
regression test for both Mosquitto behaviours.

The Go builder is unchanged: golang:1.27.1-alpine3.24@sha256:cf6fca6641884b8433441b2b0652976f975e1d0fdd26d177eaaf8596087f3125.

Relation to #22: #22 bumps only the Alpine base in Dockerfile. On its own that brings in
Mosquitto 2.1.2, which changes two defaults this image depends on (see below). This PR is meant as
the safer alternative: the same base bump, pinned by digest, plus the compatibility fixes and a
test. #22 is left untouched.

Changes (6 files, 4 commits)

File Change
Dockerfile, Dockerfile.go runtime FROM alpine:3.20 → alpine:3.24.1@sha256:28bd5fe8…
docker/mosquitto.conf allow_duplicate_messages false
docker/supervisord-go.conf, docker/supervisord-go-no-caddy.conf broker only: /usr/bin/env -u PUID -u PGID /usr/sbin/mosquitto …
scripts/test-docker-mosquitto-compat.sh new Docker regression test (not wired into CI)

Commits: 079a96b0 (runtime and config), then 3c1b7ca1, dcbdd90b and 2f3f03dd, which change
only the test script. The production files are unchanged since 079a96b0.

Mosquitto 2.1 compatibility

  1. Duplicate delivery. Mosquitto 2.1.0 changed the default of allow_duplicate_messages to
    true. The ingestor's overlapping default subscriptions (meshcore/+/+/packets and
    meshcore/#) then receive each message twice. In a controlled arm64 run with one publish,
    the unfixed candidate compared with baseline showed:

    • broker deliveries: 2 instead of 1
    • observations: 2 instead of 1
    • WebSocket broadcasts: 2 instead of 1
    • ingestor tx_dupes: 0 → 1

    DB dedup did not catch the second delivery. The two observation rows got observer_idx NULL
    and 1, which the unique index treats as distinct. Setting allow_duplicate_messages false
    restores the 2.0 behaviour for the built-in broker.

  2. PUID/PGID. Mosquitto 2.1 drops privileges to PUID/PGID when they are set.
    entrypoint-go.sh exports /app/data/.env, so a .env containing PUID=1000 made the broker
    run as 1000:1000. It then failed to write /var/lib/mosquitto/mosquitto.db (Permission denied), and retained messages were lost across restarts. The variables are now removed for
    the broker process only. Caddy, the ingestor and the server still inherit them.

Relevant package changes (Alpine 3.20.10 → 3.24.1, 51 → 62 packages)

Package Before After
mosquitto / -libs / -clients 2.0.18-r0 2.1.2-r1
caddy 2.7.6-r7 (built with go1.22.5) 2.11.4-r1 (built with go1.26.8)
supervisor 4.2.5-r5 4.3.0-r1
python3 3.12.13-r0 3.14.7-r1
libssl3 / libcrypto3 3.3.7-r0 3.5.7-r0
ca-certificates(-bundle) 20260413-r0 (145 roots) 20260611-r0 (119 roots)
musl 1.2.5-r3 1.2.6-r2
busybox 1.36.1-r31 1.37.0-r31
wget 1.24.5-r0 1.25.0-r3
libwebsockets 4.3.3-r0 4.3.5-r2
cjson 1.7.19-r0 1.7.19-r1

New packages: alpine-release, brotli-libs, gmp, gnutls, libapk, libedit, libmicrohttpd, libtasn1,
nettle, p11-kit, zstd-libs.

Unchanged: the corescope-server, corescope-ingestor and corescope-decrypt binaries are
byte-identical between the baseline and candidate images (go1.27.1, CGO disabled).

Evidence

A. Earlier image evidence (images built from the original base c5d472ba)

These images were built from the original application base c5d472ba, not from the
integration with current master. Image IDs:

  • baseline-amd64 bdf96341027e…, alpine 3.20.10
  • candfix-amd64 f3f0042d5abd…, this PR's runtime change
  • candfix-arm64 342f149535e9…
  • candidate (base bump without the fixes): amd64 bc398d24dfca…, arm64 2009c2e6a722…

Build logs record alpine:3.24.1@sha256:28bd5fe8… and the builder digest above.

  • Regression test, fail-before/pass-after

    • arm64, script from 079a96b0: candidate FAIL, candfix PASS.
    • Native amd64, final script 2f3f03dd (sha256 b6058a17…):
      • baseline PASS, 28/28 checks
      • candidate FAIL, 11 of 28 failed (broker ran as 1000:1000, retained message lost, persistence write error)
      • candfix PASS, 28/28
    • The final script also passed on candfix-arm64, 28/28, including padded-listing cleanup runs.
  • Caddy init, native amd64: 30 starts each for baseline, candidate and candfix, plus 100
    extra candfix starts. That is 130 native candfix starts with 0 non-zero exits and 0
    panics/fatal errors.

  • Full stack, native amd64: 20 baseline + 20 candfix starts, each observed for 60 s.
    All 40 were ready in about 3.7–3.9 s and exited 0. No OOM, 0 panics. Every run had
    tx_ins=1 tx_dup=0 obs_ins=1, and the broker ran as 101:102.

  • No-Caddy variant, native amd64: baseline and candfix both exit 0 with
    tx_inserted=1 obs_inserted=1 tx_dupes=0 write_errors=0.

  • Persistence: retained messages survive a broker restart, and the persistence DB is owned
    by the broker user. Both are covered by the regression test.

  • Synthetic TLS (isolated network, test CA): on baseline and candfix the ingestor subscribes
    with the test CA. It does not connect with an untrusted CA or a wrong hostname (0 subscriptions,
    0 rows). A separate probe confirmed the rejection reasons (verify codes 19 and 62).

  • Rollback baseline → candfix → baseline on a shared data volume: schema and migration hashes
    are stable, and rows are kept in both directions (counts 2/3/3 as expected, integrity ok).

  • Real endpoints, 6 credential-free TLS handshakes from native amd64 (2026-09-16 11:51 UTC):
    one per endpoint per image, all verified with both images.

    • Setup: DNS, TCP and TLS only, no application data. RootCAs=nil, so each image's own trust
      store was used, with full chain and hostname verification.
    • The probe was built with the pinned go1.27.1 builder and has the same DefaultGODEBUG as the
      ingestor (go directive 1.22).
    Endpoint TLS Verified roots (identical in both images)
    mqtt.meshview.dk:8883 1.3 ISRG Root X2 69729b8e…, ISRG Root X1 96bcec06…
    acme-v02.api.letsencrypt.org:443 1.3 ISRG Root X2, ISRG Root X1
    acme.zerossl.com:443 1.2 Sectigo Public Server Authentication Root R46 7bb647a6…, USERTrust RSA e793c9b0…

    Intermediates and roots were identical between baseline and candfix. None of the 26 roots
    removed in the new CA bundle appear in these chains.

B. Fresh integration check against current master (2026-09-16)

  • Master 12cc30ffdb1bd46cf847da263ca4e92f6b5d935b is 8 commits ahead of c5d472ba. Those
    commits touch public/app.js, two JS tests, test-all.sh and one added test line in
    .github/workflows/deploy.yml. There is no overlap with this PR's files and nothing in build,
    runtime, config or TLS.
  • git merge-tree --write-tree: no conflicts. Against master the diff is exactly the 6 files,
    byte-identical to the feature diff. Master's own changes, including deploy.yml, are preserved
    unchanged.
  • In the integration tree:
    • Dockerfile and Dockerfile.go are byte-identical to the build context of the tested
      candidate images.
    • The three config files are byte-identical to the files inside candfix-amd64 and candfix-arm64.
  • Checks from a full export of the integration tree:
    • bash -n on the test script: ok
    • git diff --check: clean
    • scripts/check-dockerfile-internal-pkgs.sh: ok
    • go test -run TestForkGuard ./cmd/server (go1.27.1, offline): 3/3 PASS
  • Config checks in the existing candfix images (arm64 native, amd64 same ID as tested):
    • All four supervisord-go*.conf files parse with supervisor 4.3.0.
    • env -u PUID -u PGID removes both variables.
    • Mosquitto 2.1.2 loads the new mosquitto.conf with 0 errors and runs as mosquitto, even
      with PUID/PGID=1000 set.
  • No new image was built from the integration tree, and no new broad test matrix was run.

Limitations

  • Emulated Caddy crashes: Caddy init crashed under amd64 emulation on arm64. QEMU is not
    proven
    to be the cause. Native amd64 (above) showed no crashes.
  • Server startup race: corescope-server exits once with db-not-ready and is restarted by
    supervisor. This happens in all stack runs, baseline included, so it is not introduced here.
  • allow_duplicate_messages: deprecated upstream (removal planned for Mosquitto 3.0; the
    broker logs a warning at startup). It only affects the built-in broker.
  • External broker and DB dedup: staging uses an external broker, which this setting does not
    cover. The underlying DB dedup gap (observer_idx NULL vs 1 in idx_observations_dedup) is
    not fixed here.
  • CA bundle: goes from 145 to 119 roots; the new set is a strict subset (26 removed, including
    DigiCert Global Root CA, DigiCert High Assurance EV Root CA, Entrust, GTS Root R2 and COMODO).
    The handshake results are a snapshot; certificate chains can change.
  • What the TLS probe does not cover: Caddy's own TLS stack (different Go versions) and actual
    ACME account registration, issuance or renewal. A verified handshake only shows the endpoint's
    certificate is trusted.
  • Production configuration is unknown, e.g. rejectUnauthorized and broker URLs.
  • Not fixed here: the existing Dockerfile.go build failure and the Dockerfile.node
    limitation.
  • CI: scripts/test-docker-mosquitto-compat.sh is not registered in CI.
  • Staging: deploy needs separate approval and a decision on disk space.

🤖 Generated with Claude Code

Openclaw and others added 4 commits September 14, 2026 19:06
Move the runtime stage of Dockerfile and Dockerfile.go from the EOL
alpine:3.20 to alpine:3.24.1, pinned by digest. The Go builder
(golang:1.27.1-alpine3.24, digest-pinned) is unchanged, and the Go
binaries are byte-identical to the Alpine 3.20 image per architecture.

Alpine 3.24 ships Mosquitto 2.1.2, which changed two behaviours the
in-container broker relied on:

- allow_duplicate_messages now defaults to true. The ingestor can hold
  overlapping subscriptions (meshcore/+/+/packets + meshcore/#), so each
  packet was delivered twice and packetsTotal, tx_dupes, observer upserts
  and observation rows doubled. docker/mosquitto.conf sets it back to
  false. Mosquitto 2.0.18 and 2.1.2 both honour it and both log a
  deprecation warning; upstream plans to remove it in Mosquitto 3.0.
- PUID/PGID are honoured for privilege dropping. entrypoint-go.sh exports
  /app/data/.env, so a PUID there made the broker run as that uid and
  fail to write /var/lib/mosquitto/mosquitto.db. The supervisord programs
  that start the broker now strip PUID/PGID with `env -u` for mosquitto
  only; every other process keeps its environment and user.

scripts/test-docker-mosquitto-compat.sh is a behavioural regression test
for both. It runs a built image with synthetic data and --network none,
publishes to the in-container broker and checks deliveries, ingestor
counters, the broker's uid/gid and persistence across a broker restart.
It fails on the unfixed Alpine 3.24.1 image and passes with this change.
It is not registered in CI.

Dockerfile.go still fails at `go mod download` because of missing
COPY internal/ lines; that pre-existing issue is out of scope here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Follow-up to the independent review of scripts/test-docker-mosquitto-compat.sh.
Only the test changes; the image, broker config and supervisord files are
untouched and the existing 28 functional checks and their expected values
stay as they were.

- Cleanup: containers get a unique com.corescope.test.run=<RUN_ID> label and
  are registered right after `docker create`. On pass, failure, start errors
  and SIGINT/SIGTERM/SIGHUP the test keeps their logs, removes only
  containers carrying its own run label together with their anonymous
  volumes (docker rm -f -v, with a checked fallback) and verifies both are
  gone. A failing run keeps its non-zero status; a cleanup failure turns a
  passing run into exit 3.
- No registry contact: the image reference must exist locally, is resolved
  once to its image ID, and every container is created from that ID with
  --pull never. A missing image fails before any container is created.
- Wall-clock limits: TOTAL_S (whole run, with a cleanup reserve), WAIT_S
  (each wait) and EXEC_S (each docker command) replace iteration-counted
  loops; wget uses -T. SIGKILL still cannot run the cleanup, so the header
  keeps recommending an outer timeout.
- Completion: every former fire-and-forget wait is now a precondition that
  stops the run when it does not hold. Exit 0 requires all 28 checks and all
  15 preconditions to have run and passed; the image ID and the Mosquitto
  version are printed.

The test still covers only the default supervisord start path and says
nothing about Caddy start-up crashes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…osed

Follow-up to the independent review of 3c1b7ca (N1, N3, N5, N8). Only
scripts/test-docker-mosquitto-compat.sh changes; the 28 checks, the 15
preconditions and their expected values are unchanged.

- N1: stdout/stderr are saved on fds 3/4 at start and restored at the top
  of the signal and exit traps and of finish, so signal, RESULT and cleanup
  lines are printed even when a signal lands inside a redirected docker
  call. Child processes get fds 3/4 closed.
- N3: if the case 1 container cannot be verifiably stopped and removed, the
  run stops with a non-zero status before case 2 is created. A later
  successful cleanup pass never clears that earlier failure.
- N5: the exit path no longer waits unbounded on a background child. Own
  children get TERM, then KILL after a short grace period, and only while
  their PID still has the parent PID and start time recorded at spawn, so
  a recycled PID is never signalled.
- N8: containers and volumes are present, absent or unknown. Absence is
  only concluded from a successful docker listing that no longer contains
  the object; failed or timed-out listings, inspects and post-removal
  checks are reported as CLEANUP UNRESOLVED and make the run non-zero,
  without any broader fallback deletion.

Cleanup now runs before the RESULT line, and a verified cleanup is part of
exit 0. SIGKILL still cannot run any of this, so the header keeps
recommending an outer timeout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tate

Follow-up to the independent review of dcbdd90 (blocker B1). Only
scripts/test-docker-mosquitto-compat.sh changes; the 28 checks, the 15
preconditions and their expected values are unchanged.

container_state and volume_state matched with printf ... | grep -qxF. grep -q
exits at the first hit, the writing printf then gets SIGPIPE, and pipefail
turns that hit into a non-zero pipeline, so on a host whose docker ps -a -q or
docker volume ls -q listing is large enough an existing container or volume was
reported as absent: cleanup skipped it, "already gone" was printed for a
container that was still running, and case 2 could be started on top of case 1.

The match now reads the whole listing, and the two status sources are kept
apart: writer 0 and grep 0 is present, writer 0 and grep 1 is absent, and
anything else - a failed writer or a grep error - is unknown, which stays
CLEANUP UNRESOLVED and makes the run non-zero. A failed or timed-out listing is
still unknown as before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@dborup
dborup merged commit aac4fae into master Sep 16, 2026
5 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants