test: make the Docker harness work on a NixOS host - #52
Conversation
Both Docker integration targets failed on this host: 15 of 34 natmap_docker tests, and every auto_discover test that execs the binary. The host-built lab-ops links libssl.so.3 from a Nix store path, and its RUNPATH points at a toolchain directory that does not contain OpenSSL, so the binary cannot resolve the library on its own. The /nix mount added in 7c943da fixed the ELF interpreter, not the library. Set the loader path for the binary alone, via a two-line sh wrapper bind-mounted over /usr/local/bin/lab-ops that execs the real binary. Setting LD_LIBRARY_PATH for the whole container is the obvious fix and is wrong: nix libcrypto carries a RUNPATH into nix glibc, so it drags nix libdl into the image's own curl, which then fails against the image's glibc. That attempt turned 34 passing auto_discover tests into 34 failures. Scoping it to one process leaves everything else alone, and exec keeps the wrapper out of the process tree so the test scripts' kill %1 cleanup still works. The store path is read from OPENSSL_LIB_DIR rather than hardcoded, so a nix upgrade does not silently break it, and the whole Nix block stays behind one check that is inert off NixOS. Full suite green: 436 passed, 0 failed across 10 targets. Also found, not fixed: the Docker harness leaks containers on failure. teardown() only runs on the success path, so any failure leaves containers named it-* behind and the next run fails at name conflict while masking the original error. A Drop guard or a startup cleanup would make these self-healing.
Release PreviewNext version: 0.1.29 (2026-09-30)Added
Changed
Fixed
|
The test-strategy section told the agent never to run ./dev.sh all or ./dev.sh test without asking, which made every session ship with the Docker integration suite unverified. The maintainer now allows it. The preference for targeted runs stays, since it is still the faster path and most changes do not touch the Docker paths. The same section's worked example was `cargo test -p natmap`, which fails: the package is lab-ops_natmap. The file contradicted itself fifteen lines later. Fixed. Adds a Build environment section, because openssl-sys cannot find OpenSSL on this host and every cargo command dies before it reaches the code. Both store paths are required, and they are not interchangeable: OPENSSL_DIR alone fails, because the -dev path holds the headers and the other holds the .so files.
teardown() returns a shell fragment that the tests embed at the end of their script, so a test that hits `exit 1` never reaches it. The test container mounts /var/run/docker.sock, so the it-* container the test created lives on the host and survives the run. The next run then dies at `docker run --name` with "Conflict. The container name ... is already in use", which reports a name collision instead of the real failure. That is not hypothetical. A 34-failure run left 21 containers behind and every later run failed the same way until they were removed by hand; afterwards the same code passed 47/47 untouched. Sweep anchored `^it-` containers once, in the existing INIT.call_once block, before the image build. Once blocks every other test until the closure returns, so the sweep finishes before any test container starts and can only ever remove a leftover from a previous run. It must not run per test: this suite is not single-threaded, and a per-test sweep would delete containers belonging to tests running at that moment, which is worse than the leak it fixes. The anchor matters too. Docker's name filter is a substring match, so a loose `it-` would also reach names like `audit-it-decoy`. The sweep is best effort and stays quiet when there is nothing to remove, so it can never be the reason the suite goes red. natmap_docker needs no equivalent: it never passes docker --name, so it has no leak to sweep. Verified by seeding it-nodomain, the name that actually masks a real failure: the sweep removed it and the suite passed 47/47. With that same leftover on the pristine tree, registration::service_id_no_domain_falls_back _to_name reports the Conflict error instead. Breaking that test's expected prefix on purpose still surfaces the original assertion failure, not a Conflict. Full suite: 442 passed, 0 failed across 10 targets.
2ec4709 to
7e850e3
Compare
|
Folded a one-line correction into the docs commit
Nothing enforces it.
Full suite after the rebase: 442 passed, 0 failed across 10 targets. The rebase touched one line of markdown and no code; |
What
The Docker integration suites could not run on this NixOS host at all. Every
test that execs the host-built
lab-opsbinary inside the container failed atstartup:
15 of 34
natmap_dockertests failed, and everyauto_discovertest that execsthe binary. The 19 that passed were the ones that never launch the binary.
This is a follow-up to #39. That PR mounted
/nixto fix the ELF interpreter,which was the first of the two problems. The library itself is the second: the
binary's
RUNPATHis a toolchain directory that does not contain OpenSSL, somounting the store makes the library reachable without making it searched.
The fix
LD_LIBRARY_PATHis set for the binary alone, via a two-lineshwrapperbind-mounted over
/usr/local/bin/lab-opsthatexecs the real binary.The obvious version — setting it for the whole container — is wrong, and I
measured that rather than assuming it. It takes
natmap_dockerfrom 19/34 to34/34 and simultaneously turns
auto_discoverfrom green to 34 failures:nix
libcrypto.so.3carries aRUNPATHinto nix glibc, so putting itsdirectory on the container's global loader path drags nix
libdlinto theimage's own
curl, which then fails against the image's glibc 2.39. Scopingthe variable to one process leaves everything else alone.
execalso keeps thewrapper out of the process tree, so the test scripts'
kill %1cleanup stillworks.
The store path is read from
OPENSSL_LIB_DIRrather than hardcoded, so a nixupgrade does not silently break it, and the whole Nix block stays behind one
/nixcheck that is inert off NixOS.Also in this PR
Sweep leaked
it-*containers before the suite starts.teardown()returns ashell fragment that tests embed at the end of their script, so a test hitting
exit 1never reaches it. The test container mounts the Docker socket, so thatcontainer lives on the host and survives. The next run then dies at
docker run --namewith a name conflict that reports the wrong problem.Not hypothetical: a 34-failure run left 21 containers behind, and every later run
failed identically until they were cleared by hand. Afterwards the same code
passed 47/47 untouched.
The sweep runs once, inside the existing
INIT.call_onceblock, not per test.That is deliberate: this suite is not single-threaded.
AGENTS.mdclaimed--test-threads=1was enforced by.cargo/config.toml; it is not — that filecontains only
rustflags. A per-test sweep would delete containers belonging totests running at that moment, which is worse than the leak.
Onceblocks everyother test until the closure returns, so a one-shot sweep finishes before any
test container starts and can only remove a leftover from a previous run.
natmap_dockerneeds no equivalent; it never passes docker--name.Verification
Full suite, NixOS host, with two seeded leftovers (
it-nodomain, the name thatactually masks a real failure, plus a decoy):
No
it-*containers left behind. The seededit-nodomainwas swept.Control on the pristine tree with the same leftover:
registration::service_id_no_domain_falls_back_to_namereports theConflicterror instead of running. Breaking that test's expected prefix on purpose, on this branch, surfaces the original assertion failure with zeroConflictoccurrences — which is the acceptance criterion, since a failure reporting as an unrelated failure is the exact problem being fixed.cargo +nightly fmt --allclean, clippy 0 warnings / 0 errors.Not included
refactor/audit-cut, deliberately unreviewed and off this branch.AGENTS.mdstill carries two false claims about this suite, both reportedseparately rather than fixed here: the
--test-threads=1enforcement, and--all-featuresbeing redundant (it is not, on this host — see below).