Skip to content

test: make the Docker harness work on a NixOS host - #52

Merged
FAZuH merged 3 commits into
mainfrom
test/nix-docker-harness
Sep 30, 2026
Merged

FAZuH merged 3 commits into
mainfrom
test/nix-docker-harness

Conversation

@FAZuH

@FAZuH FAZuH commented Sep 30, 2026

Copy link
Copy Markdown
Owner

What

The Docker integration suites could not run on this NixOS host at all. Every
test that execs the host-built lab-ops binary inside the container failed at
startup:

lab-ops: error while loading shared libraries: libssl.so.3: cannot open shared object file

15 of 34 natmap_docker tests failed, and every auto_discover test that execs
the binary. The 19 that passed were the ones that never launch the binary.

This is a follow-up to #39. That PR mounted /nix to fix the ELF interpreter,
which was the first of the two problems. The library itself is the second: the
binary's RUNPATH is a toolchain directory that does not contain OpenSSL, so
mounting the store makes the library reachable without making it searched.

The fix

LD_LIBRARY_PATH is set for the binary alone, via a two-line sh wrapper
bind-mounted over /usr/local/bin/lab-ops that execs the real binary.

The obvious version — setting it for the whole container — is wrong, and I
measured that rather than assuming it. It takes natmap_docker from 19/34 to
34/34 and simultaneously turns auto_discover from green to 34 failures:

curl: /lib/x86_64-linux-gnu/libc.so.6: version `GLIBC_ABI_DT_X86_64_PLT' not found
(required by /nix/store/...-glibc-2.42-84/lib/libdl.so.2)

nix libcrypto.so.3 carries a RUNPATH into nix glibc, so putting its
directory on the container's global loader path drags nix libdl into the
image's own curl, which then fails against the image's glibc 2.39. Scoping
the variable to one process leaves everything else alone. exec also keeps the
wrapper out of the process tree, so the test scripts' kill %1 cleanup still
works.

The store path is read from OPENSSL_LIB_DIR rather than hardcoded, so a nix
upgrade does not silently break it, and the whole Nix block stays behind one
/nix check that is inert off NixOS.

Also in this PR

Sweep leaked it-* containers before the suite starts. teardown() returns a
shell fragment that tests embed at the end of their script, so a test hitting
exit 1 never reaches it. The test container mounts the Docker socket, so that
container lives on the host and survives. The next run then dies at
docker run --name with a name conflict that reports the wrong problem.

Not hypothetical: a 34-failure run left 21 containers behind, and every later run
failed identically until they were cleared by hand. Afterwards the same code
passed 47/47 untouched.

The sweep runs once, inside the existing INIT.call_once block, not per test.
That is deliberate: this suite is not single-threaded. AGENTS.md claimed
--test-threads=1 was enforced by .cargo/config.toml; it is not — that file
contains only rustflags. A per-test sweep would delete containers belonging to
tests running at that moment, which is worse than the leak. Once blocks every
other test until the closure returns, so a one-shot sweep finishes before any
test container starts and can only remove a leftover from a previous run.

natmap_docker needs no equivalent; it never passes docker --name.

Verification

Full suite, NixOS host, with two seeded leftovers (it-nodomain, the name that
actually masks a real failure, plus a decoy):

test result: ok. 65 passed    (lab-ops lib)
test result: ok. 0  passed    (lab-ops bin)
test result: ok. 47 passed    (auto_discover, 91.5s)
test result: ok. 4  passed    (cf2ansible)
test result: ok. 34 passed    (natmap_docker)
test result: ok. 67 passed    (auto-discover unit)
test result: ok. 34 passed    (lab-lib)
test result: ok. 163 passed   (natmap unit)
test result: ok. 21 passed    (cli)
test result: ok. 7  passed    (model)

No it-* containers left behind. The seeded it-nodomain was swept.

Control on the pristine tree with the same leftover: registration::service_id_no_domain_falls_back_to_name reports the Conflict error instead of running. Breaking that test's expected prefix on purpose, on this branch, surfaces the original assertion failure with zero Conflict occurrences — which is the acceptance criterion, since a failure reporting as an unrelated failure is the exact problem being fixed.

cargo +nightly fmt --all clean, clippy 0 warnings / 0 errors.

Not included

  • The refactor and comment/doc passes from the audit are parked on
    refactor/audit-cut, deliberately unreviewed and off this branch.
  • AGENTS.md still carries two false claims about this suite, both reported
    separately rather than fixed here: the --test-threads=1 enforcement, and
    --all-features being redundant (it is not, on this host — see below).

Both Docker integration targets failed on this host: 15 of 34
natmap_docker tests, and every auto_discover test that execs the binary.
The host-built lab-ops links libssl.so.3 from a Nix store path, and its
RUNPATH points at a toolchain directory that does not contain OpenSSL,
so the binary cannot resolve the library on its own. The /nix mount
added in 7c943da fixed the ELF interpreter, not the library.

Set the loader path for the binary alone, via a two-line sh wrapper
bind-mounted over /usr/local/bin/lab-ops that execs the real binary.
Setting LD_LIBRARY_PATH for the whole container is the obvious fix and
is wrong: nix libcrypto carries a RUNPATH into nix glibc, so it drags
nix libdl into the image's own curl, which then fails against the
image's glibc. That attempt turned 34 passing auto_discover tests into
34 failures. Scoping it to one process leaves everything else alone,
and exec keeps the wrapper out of the process tree so the test scripts'
kill %1 cleanup still works.

The store path is read from OPENSSL_LIB_DIR rather than hardcoded, so a
nix upgrade does not silently break it, and the whole Nix block stays
behind one check that is inert off NixOS.

Full suite green: 436 passed, 0 failed across 10 targets.

Also found, not fixed: the Docker harness leaks containers on failure.
teardown() only runs on the success path, so any failure leaves
containers named it-* behind and the next run fails at name conflict
while masking the original error. A Drop guard or a startup cleanup
would make these self-healing.
@github-actions

Copy link
Copy Markdown

Release Preview

Next version: v0.1.29

0.1.29 (2026-09-30)

Added

  • Added GET /rules endpoint to natmap daemon that returns live NAT rules.
  • Added value 0 to host_port, which asks the natmap daemon for a free host port.

Changed

  • Changed host-port allocation so the natmap daemon assigns the ports.

Fixed

  • Fixed error logs hiding the real cause behind a generic message
  • Fixed port reservation failing when a port is released and immediately re-allocated.
  • Fixed full sync failure removing services that were already registered.

The test-strategy section told the agent never to run ./dev.sh all or
./dev.sh test without asking, which made every session ship with the
Docker integration suite unverified. The maintainer now allows it. The
preference for targeted runs stays, since it is still the faster path
and most changes do not touch the Docker paths.

The same section's worked example was `cargo test -p natmap`, which
fails: the package is lab-ops_natmap. The file contradicted itself
fifteen lines later. Fixed.

Adds a Build environment section, because openssl-sys cannot find
OpenSSL on this host and every cargo command dies before it reaches the
code. Both store paths are required, and they are not interchangeable:
OPENSSL_DIR alone fails, because the -dev path holds the headers and
the other holds the .so files.
teardown() returns a shell fragment that the tests embed at the end of
their script, so a test that hits `exit 1` never reaches it. The test
container mounts /var/run/docker.sock, so the it-* container the test
created lives on the host and survives the run. The next run then dies
at `docker run --name` with "Conflict. The container name ... is already
in use", which reports a name collision instead of the real failure.

That is not hypothetical. A 34-failure run left 21 containers behind and
every later run failed the same way until they were removed by hand;
afterwards the same code passed 47/47 untouched.

Sweep anchored `^it-` containers once, in the existing INIT.call_once
block, before the image build. Once blocks every other test until the
closure returns, so the sweep finishes before any test container starts
and can only ever remove a leftover from a previous run. It must not run
per test: this suite is not single-threaded, and a per-test sweep would
delete containers belonging to tests running at that moment, which is
worse than the leak it fixes.

The anchor matters too. Docker's name filter is a substring match, so a
loose `it-` would also reach names like `audit-it-decoy`. The sweep is
best effort and stays quiet when there is nothing to remove, so it can
never be the reason the suite goes red.

natmap_docker needs no equivalent: it never passes docker --name, so it
has no leak to sweep.

Verified by seeding it-nodomain, the name that actually masks a real
failure: the sweep removed it and the suite passed 47/47. With that same
leftover on the pristine tree, registration::service_id_no_domain_falls_back
_to_name reports the Conflict error instead. Breaking that test's expected
prefix on purpose still surfaces the original assertion failure, not a
Conflict.

Full suite: 442 passed, 0 failed across 10 targets.
@FAZuH
FAZuH force-pushed the test/nix-docker-harness branch from 2ec4709 to 7e850e3 Compare September 30, 2026 14:31
@FAZuH

FAZuH commented Sep 30, 2026

Copy link
Copy Markdown
Owner Author

Folded a one-line correction into the docs commit cf70f33 (history rewritten, hence the force-push).

AGENTS.md:39 claimed:

Docker tests require --test-threads=1 (enforced by .cargo/config.toml)

Nothing enforces it. .cargo/config.toml contains only rustflags, so the Docker suites run with cargo's default parallelism. That claim is what misdirected the sweep design in the first draft, so leaving it would misdirect the next reader the same way. It now says the setting is not enforced, tells you to pass -- --test-threads=1 yourself when you need determinism, and points at #54 for the flakiness.

docs/dev/testing.md:129 and :139 are unchanged. They are advice ("Docker tests must run single-threaded", with the flag in the command) and never claimed config enforcement, so they were not in conflict — and #54's 44-47/47 evidence is consistent with that advice being right.

Full suite after the rebase: 442 passed, 0 failed across 10 targets. The rebase touched one line of markdown and no code; git diff between the pre- and post-rebase trees is that line alone.

@FAZuH
FAZuH merged commit 1028196 into main Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant