Sibling fork OKI-Mesh/CoreScope (forked from Kpa-clawbot/CoreScope on 2026-06-24, three weeks before ours) has landed two small operational fixes that address failure modes our own staging path is exposed to. This issue is to decide whether and when to port them — no work is proposed yet.
Note on references: all references to the sibling fork are deliberately written in code format, not as links, to avoid creating cross-reference notifications in their repository.
1. Staging has no memory bound — directly applicable
Their OKI-Mesh/CoreScope#124 adds two settings to docker-compose.staging.yml:
mem_limit: 5g # hard cap
- GOMEMLIMIT=${GOMEMLIMIT:-2700MiB} # Go soft GC target
The pair works together: GOMEMLIMIT keeps the heap bounded via GC, mem_limit is the backstop that lets Docker OOM-kill the container instead of letting the host thrash. Their stated trigger: an unbounded corescope-serve exhausted the host and drove a neighbouring box into a watchdog reboot loop.
Verified gap on our side. None of our three compose files sets either value:
| File |
mem_limit |
GOMEMLIMIT |
docker-compose.staging.yml |
absent |
absent |
docker-compose.yml |
absent |
absent |
docker-compose.no-mosquitto.yml |
absent |
absent |
Our staging-go runs with restart: unless-stopped and no memory ceiling, so a memory-hungry container restarts and starts over.
Caveat: their 5g / 2700MiB are sized to their prod host. Ours need to be chosen for the staging VM, not copied.
2. Health-check window — the lesson ports, the diff does not
Their OKI-Mesh/CoreScope#128 widens the prod deploy health gate from 120s to 300s, and the rollback gate from 90s to 240s. Cause: an ~8 GB meshcore.db cold start (neighbour-graph load + hash migration + cache rebuild) takes ~3 minutes, so a healthy release was marked failed and rolled back — and the rollback then tripped the same too-short window.
We have no deploy-prod.yml, so that file cannot be ported. But the same number exists in our staging path, .github/workflows/deploy.yml:
for i in $(seq 1 120); do ... "Staging failed health check after 120s"
Our staging is not a toy dataset — measurements recorded during #93 were a cold staging copy with index build ~16.6 s and backfill ~9 min.
Difference in our favour: our staging gate has no rollback branch. On failure it dumps 50 log lines and exits 1; rollback is a manual procedure with preserved containers and snapshots. The second half of their failure (rollback hitting the same short window) cannot happen here.
What is actually needed: not a ported number, but a measurement. Time a real cold start of staging against production-scale data and size the window from that, with margin. Their own comment is worth carrying over in spirit:
Dev/offband boots in seconds only because its DB is ~25 MB and unseeded; it is not representative of prod startup.
Proposed decision
Out of scope
Their larger structural work — single Go module (OKI-Mesh/CoreScope#116), Goose migrations (OKI-Mesh/CoreScope#121, OKI-Mesh/CoreScope#131) and migration in the deploy path (OKI-Mesh/CoreScope#138) — is roughly 140 files and ~6,400 changed lines, and rewrites schema ownership. That needs its own epic with staging verification and a DB backup, and must not be mixed into current route_mask work.
Sibling fork
OKI-Mesh/CoreScope(forked fromKpa-clawbot/CoreScopeon 2026-06-24, three weeks before ours) has landed two small operational fixes that address failure modes our own staging path is exposed to. This issue is to decide whether and when to port them — no work is proposed yet.1. Staging has no memory bound — directly applicable
Their
OKI-Mesh/CoreScope#124adds two settings todocker-compose.staging.yml:The pair works together:
GOMEMLIMITkeeps the heap bounded via GC,mem_limitis the backstop that lets Docker OOM-kill the container instead of letting the host thrash. Their stated trigger: an unboundedcorescope-serveexhausted the host and drove a neighbouring box into a watchdog reboot loop.Verified gap on our side. None of our three compose files sets either value:
mem_limitGOMEMLIMITdocker-compose.staging.ymldocker-compose.ymldocker-compose.no-mosquitto.ymlOur
staging-goruns withrestart: unless-stoppedand no memory ceiling, so a memory-hungry container restarts and starts over.Caveat: their 5g / 2700MiB are sized to their prod host. Ours need to be chosen for the staging VM, not copied.
2. Health-check window — the lesson ports, the diff does not
Their
OKI-Mesh/CoreScope#128widens the prod deploy health gate from 120s to 300s, and the rollback gate from 90s to 240s. Cause: an ~8 GBmeshcore.dbcold start (neighbour-graph load + hash migration + cache rebuild) takes ~3 minutes, so a healthy release was marked failed and rolled back — and the rollback then tripped the same too-short window.We have no
deploy-prod.yml, so that file cannot be ported. But the same number exists in our staging path,.github/workflows/deploy.yml:Our staging is not a toy dataset — measurements recorded during #93 were a cold staging copy with index build ~16.6 s and backfill ~9 min.
Difference in our favour: our staging gate has no rollback branch. On failure it dumps 50 log lines and exits 1; rollback is a manual procedure with preserved containers and snapshots. The second half of their failure (rollback hitting the same short window) cannot happen here.
What is actually needed: not a ported number, but a measurement. Time a real cold start of staging against production-scale data and size the window from that, with margin. Their own comment is worth carrying over in spirit:
Proposed decision
docker-compose.staging.yml, with values measured for our staging VM rather than copieddeploy.ymlbased on that measurementdeploy-prod.ymlis introduced — the rollback-window half of their fix becomes relevant thenOut of scope
Their larger structural work — single Go module (
OKI-Mesh/CoreScope#116), Goose migrations (OKI-Mesh/CoreScope#121,OKI-Mesh/CoreScope#131) and migration in the deploy path (OKI-Mesh/CoreScope#138) — is roughly 140 files and ~6,400 changed lines, and rewrites schema ownership. That needs its own epic with staging verification and a DB backup, and must not be mixed into currentroute_maskwork.