Repository navigation
fix(qa): bind blacklist pubkey in SQLite query (port of upstream #1982) - #82
Merged
Merged
Conversation
The §10.2 probe in blacklist-test.sh counted `transmissions WHERE from_node = :pubkey`, but no CoreScope database has a from_node column: the ingestor's CREATE TABLE, the from_pubkey_v1 migration and internal/dbschema's AssertReady all define transmissions.from_pubkey. Against a real target the probe could only fail with "no such column". The unit test hid this by building its own table with an invented from_node column. It now takes the transmissions DDL straight from cmd/ingestor/db.go, also runs the query against the committed staging-captured fixture, and keeps the old from_node table only as a negative case that must error rather than count 0. from_pubkey is written only for ADVERTs, as hex.EncodeToString output (lowercase). The hex gate and the server's nodeBlacklist accept any case, so the bound value is lowercased in SQL; binding, stdin transport, the capability probe and the container->host fallback are unchanged. New coverage: exact-match and case handling, prefix non-match, more injection payloads, empty/whitespace/multibyte/10k values, missing table/column surfacing, and a stubbed ssh_t that runs the remote command through a real bash -c to prove the fallback, the rejection of a sqlite3 that cannot bind, and that SQL and pubkey travel on stdin, never argv. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
teardown was installed as `trap teardown EXIT INT TERM` and took its exit status from $?. On SIGINT/SIGTERM that is the status of whatever ran before the signal, usually 0, so an aborted run tore down correctly and then exited 0 — indistinguishable from a pass. Found by forcing a SIGTERM mid-run in the disposable SSH/Docker QA environment. Signal traps now pass 130 (INT) / 143 (TERM) to teardown, which still adds 1 if teardown itself fails. The traps live in install_teardown_traps so the unit test drives the real installation: clean exit, failure counts, teardown failure, and SIGTERM with and without a failing teardown. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Independent review of the previous commit found that a second Ctrl-C or TERM arriving while teardown restores the target re-entered teardown through the TEARDOWN_DONE branch and exited at once: the node could stay blacklisted on the target, with no teardown-failed line and the earlier status lost. teardown now ignores INT/TERM once it starts; its steps are already bounded by CURL_TIMEOUT / RESTART_WAIT_S and SIGKILL still works. Tests: SIGINT maps to 130 and still tears down; a signal raised during teardown neither aborts it nor replaces the exit status. In CI a missing sqlite3 now fails instead of skipping the binding and error groups, and the fixture pubkey is validated before the reference count uses it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Follow-up review of fd47977: `trap '' INT TERM` is inherited by every child, so once teardown started, the ssh steps ignored Ctrl-C too, and only curl is actually time-bounded (ConnectTimeout covers connection setup, not a hung session). A stuck `docker restart` or a dropped connection left the operator with SIGQUIT/SIGKILL, both of which skip the rest of teardown. teardown now installs a handler that only reports the signal. Children keep the default disposition, so a terminal Ctrl-C stops the step in progress; that step fails and teardown reports teardown-failed instead of exiting silently. SSH gains ServerAliveInterval=15/CountMax=4 so a dead session fails within about a minute. Tests: a child started during teardown must still die on SIGINT (it signals itself; trap -p cannot show an inherited ignore), and SIGINT with a failing teardown must exit 131, which distinguishes the INT trap from bash's own default 130. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 23, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
TEST_PUBKEYinto the blacklist retention query; bind it as an SQLite parameter via a byte-safe hex literalsqlite3for working parameter binding instead of silently falling back; no interpolating fallback existstransmissions.from_pubkey— the query previously namedfrom_node, which no CoreScope database hasThe
from_nodeblockerThe §10.2 probe counted
SELECT COUNT(*) FROM transmissions WHERE from_node = :pubkey. The column does not exist, so against any real CoreScope database the probe could only fail withno such column: from_node. The unit test hid this because it created its own table with an inventedfrom_nodecolumn. (The column name predates this PR; the old interpolated query used it too.)Schema evidence (master
f52bf7d5)transmissions(… from_pubkey TEXT …); nofrom_nodeanywhere in the schemacmd/ingestor/db.go:348-360from_pubkey_v1adds the column +idx_transmissions_from_pubkeyon older DBscmd/ingestor/db.go:803-818,internal/dbschema/dbschema.go:419-432internal/dbschema/dbschema.go:158(mustCol("transmissions", "from_pubkey"))cmd/ingestor/db.go:2576-2580; backfillcmd/ingestor/maintenance.go:188-310hex.EncodeToStringof the 32-byte key → 64 lowercase hex charscmd/ingestor/decoder.go:329cmd/server/db.go:866, 915, 1197, 4355,cmd/server/advert_stats.go:66nodeBlacklistmatching is case-insensitive (ToLower+TrimSpace)cmd/server/config.go:927-941, 977-996packets_v) does not altertransmissionscmd/ingestor/db.go:456Since stored values are lowercase but the script's hex gate and the blacklist accept any case, the query is now
WHERE from_pubkey = lower(:pubkey)— the value is still a bound parameter; only its normalisation happens in SQL. The count is "ADVERT transmissions attributed to this node", which is what the ingestor can attribute (encrypted payloads carry no sender).Changes in this update
5ec3873c— queryfrom_pubkey; the unit-test fixture is now thetransmissionsDDL extracted fromcmd/ingestor/db.go, plus a run against the committed staging-capturedtest-fixtures/e2e-fixture.db; the old inventedfrom_nodetable is kept only as a negative case that must error15f6c597—install_teardown_traps: INT/TERM pass 130/143 toteardown(still +1 if teardown fails)fd479771,8a0a420e— review fixes: a second INT/TERM during teardown no longer aborts it (it is reported, not ignored, so a hung ssh step can still be interrupted and is then reported asteardown-failed); SSH gainsServerAliveInterval=15/ServerAliveCountMax=442482fd1— merge of master (brings in test(e2e): pin munger packets window (port of upstream #1942) #78, ci: run eight previously orphaned E2E suites (port of upstream #2045) #80, fix(channels): cancel deferred color-swatch focus (port of upstream #1945) #81);deploy.ymldiffers from master only by this PR's unit-test stepBinding, stdin transport, the capability probe, the container→host fallback, visible SQL errors and token-via-stdin handling are unchanged. No production code or schema is touched.
Verification
Unit tests —
bash qa/scripts/test-blacklist-sql.shCI=true/bin/sh= dash, no sqlite3CI=truewith no sqlite3 onPATHCoverage added: query text names
from_pubkeyand notfrom_node; DDL comes from the ingestor; legitimate pubkey → its rows (also upper-case input); absent pubkey and pubkey prefix → 0; seven injection payloads bind literally (the interpolated form leaks all 6 rows); empty, whitespace, multibyte and 10,000-byte values; real captured fixture count equals a direct count; missing column / missing table → non-zero exit + named on stderr + no count; a stubbedssh_tthat runs the remote command through a realbash -cproves container→host fallback, container runner, rejection of ansqlite3that cannot bind, a remote schema error kept as a failure, and that SQL and the (hex-encoded) pubkey travel on stdin and never in the remote argv; teardown exit codes for clean, failing, failed-teardown and SIGTERM runs.Mutation checks (each must make the suite fail)
from_node = :pubkeylower('%s')interpolation') OR 1=1 --leaks all 6 rows)0from_pubkeytrap teardown EXIT INT TERM(old trap)trap teardown INT(no 130)trap '' INT TERMin teardown (inherited ignore)All ten were run against the final test file; the unmutated suite passes.
Full SSH/Docker run in an isolated, disposable environment
Run from a disposable environment on the shared
demohost with its own resource prefix; nothing belonging to other work was touched, and no staging or production system was contacted.Dockerfile(nosqlite3, as in production), plus a test-only variant with thesqlitepackage added for the container-runner case--internalDocker network (outbound blocked, verified),DISABLE_MOSQUITTO/DISABLE_CADDY, no MQTT sources, its own data directory andconfig.json; reachable only through a proxy published on the host's loopbackblacklist-test.shran on the demo host and reached a disposable sshd "target host" container over real OpenSSH (ephemeral key, isolatedknown_hosts); that container drove the real app container through the Docker socket (docker restart,docker exec -i … sqlite3) and edited the real bind-mountedconfig.jsonFinal pass on
8a0a420e:sqlite3in the app image or on the targetretain-failed: no sqlite3 able to bind, teardown oksqlite3, target with ithost, 3 rows retained, teardown oksqlite3container, 3 rows retained, teardown okFor every run: before and after, the node was visible (detail 200, in list),
nodeBlacklistwas[]with an identical config hash, the app container was running, and the transmission counts were unchanged (4 total, 3 for the target). The unrelated node's row was never counted.Process arguments were captured completely with
strace -f -e trace=execveover the script's process tree (plus a host-widepssampler): the SQL text, the column name and the token appeared in 0 argv entries in all runs. Logs were sanitised; the synthetic pubkey does not occur in them. The environment (stopped containers, networks, images, data and logs) is kept for inspection.Other checks
bash -non both scripts,git diff --check.github/workflows/deploy.ymlparsed with Ruby Psych; against master it differs only by this PR's unit-test stepcd cmd/server && go test -run 'ForkGuard|Workflow' .— passfd479771and8a0a420eKnown limitations
ssh(PK=… bash -s), thecurlURLs (/api/nodes/<pk>, and the admin probe below) and twogreppatterns — 9execveentries per full run. It is the node's public identifier, served by the API itself; the SQL path, which is what this PR changes, does not expose it. Moving these to stdin would be a separate change.ADMIN_API_TOKENpath queries/api/admin/transmissions?from_node=…. No such endpoint exists incmd/server, so it always falls through to the SQLite probe. Pre-existing; the token itself stays out of argv.TARGET_DB_PATHis used for both runners. In the container runner it is a container path; for the host fallback it must be a host path. In the demo both views mounted the data at the same path. When they differ, the host fallback fails asretain-failed(never a false pass), andsqlite3may create an empty file at that path if its directory exists on the host (-readonlywas considered but not adopted here).from_pubkey; a node that has never sent an ADVERT reports 0 and fails §10.2, as intended.RESTART_WAIT_S.sqlite3, so the operator's target host needs thesqlite3CLI for §10.2.shellcheckwas not available locally.Upstream
Fork port of upstream CoreScope PR
1982(issue1977), applied to fork master and extended here with thefrom_pubkeyand exit-status fixes, which are fork-local.Workflow safety
The workflow diff against master is only the local SQL unit-test step. Fork guards for GHCR publishing, releases, staging deployment and badge writes are unchanged. No staging, production or upstream system was contacted.
🤖 Generated with Claude Code