Skip to content

Latest commit

 

History

History
2158 lines (1930 loc) · 139 KB

File metadata and controls

2158 lines (1930 loc) · 139 KB

Testing

Every gate is cheap to run, and each is run on its own: there is no aggregate runner. The lists below are the statement of record; no count of them is written in this file, because a number beside a list is a second claim about one population with nothing checking the two against each other, which this file has already recorded happening twice.

gate reads needs
test-inventory.sh the lists that claim to cover src/, every file under test/, the runner behind each runnable file under test/, the DEFER ledger and the XXX table sh, grep, git
test-citations.sh every file:line in a table under doc/ or doc/history/ resolves, and the named symbol is on the line sh, grep
test-history.sh every changelog row's commit resolves with a matching subject, no two rows share a version, and every document stating the tree's current version names the newest row git with full history
test-provenance.sh every file under src/ has an origin row, and a carried file is re-checked with cmp sh, git
test-absence.sh every "X() is not carried" and "->method is not written" claim resolves against src/ sh, grep
test-shim.sh the shim parses against test/stub in both knob positions, with a control that must fail a C compiler
test-syntax.sh the files it names compile under clang and gcc at the kernel of record, warnings are failures, with a control that kernel's tree, clang and gcc
test-checkpatch.sh the style deviation set against doc/checkpatch-baseline.txt checkpatch.pl from the kernel of record: found by the same search test-syntax.sh uses, KDIR first, the baseline's checker by sha256 winning wherever it sits; only CHECKPATCH overrides
test-vectors-contract.sh the exit status, output wording and constant spelling of the two vector files a consumer compiles a C compiler
test-posix.sh the gates declaring #!/bin/sh parse under dash and busybox ash dash and busybox
test-doc-prose.sh vale over every tracked .md; any finding is a failure vale
test-fixtures.sh every fixture mounted on a guest, every file compared to its manifest a guest, the fixture images, KDIR matching the guest's kernel
test-enospc.sh a volume filled as root or as a user, what it kept and what each writer was told the same

Exit 2 from any of them means the instrument could not run, which is neither pass nor fail and is never recorded as either.

The scripts below need the guest fleet and are not gates. Each builds the module itself, exits 2 without the fleet, and produces a reading rather than a verdict; the readings are in doc/history/verification-record.md.

script does
fuzz-mount.sh puts mutated images through the mount path and, with H2_FUZZ_WRITE=1, the write path
f4-roundtrip.sh writes a tree here and checks it on DragonFly, then the other way
cut-flush.sh cuts DragonFly off mid-write and mounts the result here
crash-matrix.sh runs the 0.6 crash matrix against the FreeBSD port
root-boot.sh boots a kernel whose root filesystem is a HAMMER2 volume, with the mapping and exec checks that let it be one
pfs-domains.sh creates PFS roots here, mounts each by label on both sides, and has DragonFly check what was written in each
million-tree.sh writes a million-file tree here and has both sides count it
throughput.sh times one large file against ext4 and reads its allocation order from the image beside DragonFly's
latency.sh times random 4 KiB reads and write-then-fsync per operation on this port and on ext4 and btrfs in the same guest, from test/hammer2-latency.c, with the page cache dropped first
nix-closure.sh copies a real Nix closure in through the port and reads it cold beside squashfs and erofs
bulkfree.sh writes a set, removes it, and runs the bulkfree scan that frees it
hpanic-contain.sh reads what a device in error keeps off the media
cluster-sync.sh runs the cluster's synchronization thread on two volumes: a lone SLAVE's thread passing and stopping on unmount and pfs-delete, then a SLAVE created after its MASTER's files and compared with it by fsck_hammer2. Two of its checks read fsck_hammer2's exit status, which until 0.9.66 a command substitution in the same command had replaced with its own, so both passed whatever fsck did
cluster-quorum.sh two MASTERs of one cluster id on two volumes, which is the configuration hammer2_cluster_check() decides for: with one master it agrees with the only chain, and with two the quorum is pfs_nmasters / 2 + 1, so both must agree before a lookup is answered. A set made off the volume and summed first is written through the cluster's own mount, read back after an unmount and a cache drop, and compared sum for sum with what was written, and each volume's media is fingerprinted and fsck'd separately. Until 0.9.67 the read-back was a count of the sums and the read came from the page cache the write left, so a file reading back wrong would have passed; H2_QUORUM_CONTROL=1 changes one file after its sum is taken and must fail that check. The fake pass it guards is a second volume silently unseen, which would leave an ordinary single-master mount passing every other check: the support-thread count distinguishes them, since hammer2_vfsops.c skips that thread for a MASTER element only when pfs_nmasters is 1
analyze.sh clang's static analyzer over the port's own seven files with the syntax gate's kernel flag set, a path-sensitive reading no compiler or sparse pass takes; a host instrument, not a gate, since the carried core carries upstream's dead stores and the reading is a candidate list to triage against the origin tree
dfly-enospc.sh fills a volume to capacity on the DragonFly guest under an allocator that refuses on a count, the reproducer behind the staged unmount patch
make compile_commands.json the kernel of record's own gen_compile_commands.py over the .cmd files a build leaves, so a semantic index reads every file in the module. Not a gate and not checkable by one: the database is untracked, and a source file absent from it returns no callers for a symbol rather than an error, which reads as a negative result. hammer2_synchro.c was outside it for a day and a caller search over hammer2_xop_start_except() returned one of its four callers

test-absence.sh resolves a claim rather than a citation. Where a document says a named function is not carried, it asks src/ whether that function is defined and fails when it is. The vocabulary is the origin table's: carried means imported substantially unchanged, and a function this port rewrote is rewritten, so a symbol that is present here has not been "not carried" whatever else is true of it. Its first run found one, in hammer2_inode.c, where hammer2_igetv() was called uncarried after this port rewrote it on iget5_locked(). It reads only the claims that name a symbol, which is a fraction of the class it belongs to, and it says so on every run rather than leaving the rest to inherit its credibility.

It reads a second claim shape since 2026-09-04: ->method is not written, resolved against the operations tables by asking whether that member is initialized at the start of a line under src/. That shape was added because the first one could not see the defect that kept recurring. Three documents described ->iterate_shared as unwritten after it was, and the README's opening paragraph said the port does not mount anything for four days after it began mounting. Its first run found ->reconfigure described as unwritten while it is wired into hammer2_fs_context_ops and deliberately returns -EROFS, which is a stronger statement than the prose was making.

Only the arrow form is read, a bare method name not being distinguishable from ordinary prose, and only the present tense. A claim written in the past with the commit it was true at is a dated observation and cannot go stale, which is why doc/history/verification-record.md records the readdir floor as "was not written at 1f025fe". The gate carries a --selftest driving six directions, including that one, on a fixture tree rather than on this repository's own prose.

test-shim.sh compiles hammer2_os.h and hammer2_compat.h against the stubs in test/stub, in both positions of the HAMMER2_INVARIANTS knob, plus a negative control: the header is broken on a copy and the compile must fail. Without that control a gate whose healthy signature is silence cannot be told from a gate that never opened the file.

The gates run against the built tree

make was first run on 2026-09-02. It put thirteen objects and their .cmd files beside the sources, and test-provenance.sh and test-inventory.sh went red on the spot: the first asked for an origin row for each of thirty files kbuild had just written, the second read XXX out of the strings inside hammer2.o and asked the status table for a row. Neither had a bug that could pass something wrong. Both enumerated src/ and had never seen anything there that was not source.

test-checkpatch.sh and test-citations.sh had the same shape without having tripped. The one that bites latest is the *.c glob: kbuild writes hammer2.mod.c, which no build has reached, because modpost stops first.

All four now exclude kbuild's output, and the patterns match .gitignore's. One suffix was missing from both lists until 2026-09-26: .o.d, kbuild's dependency file. It went unnoticed because every build that reached this tree had succeeded, and a successful build's .o.d files are covered by the *.cmd entries beside them in the reader's eye but not in a matcher. A make against the host's kernel, which the version #error stops before the compile, still writes the .d for whichever object it reached, so a failed build left src/sys/fs/hammer2/.hammer2_io.o.d behind and the next test-provenance.sh run failed asking for a row for it. That had been hand-deleted once as a stray, which treated the symptom. .o.d is in .gitignore and in all four exclusion lists now, driven both ways: with the real artifact present the gate passes, and a planted non-source file in src/ still fails it with exit 1. The permanent guard is not a new gate but an ordering: the pre-push hook builds the module before it runs any gate, so every gate runs against the tree a developer actually has. Until 2026-09-03 this paragraph also said that step asserts the undefined set is exactly the four named in doc/README.status.md. Nothing asserted that. The step read modinfo and counted warnings, and a fifth undefined reference would have been a failed link with no list, which is a red run but not the one described here.

What it does assert is in script/build-check.sh, which is not a gate and is named build-check rather than test- for that reason: the build fails, or the build is not warning-clean, or the build reports success and there is no hammer2.ko. It lives in a script because it has two callers now, the kernel of record and the floor, and a check copied into a second caller can rot in one copy while the other stays right. Its three failing directions were driven on 2026-09-02 and 2026-09-03, by giving hammer2_io.c a call to an undefined function, then an unused static, and by pointing it at a directory holding no kernel, which is COULD-NOT-RUN and not a pass.

Its warning pattern requires file:line:column, because a bare warning: also matches kbuild's banner about the runner's compiler differing from the kernel's, which is a fact about the machine. That failed the step for two runs while the build was clean. The pattern is therefore checked against a line built to match before it is trusted on a log that should have none, since no warnings and a pattern that stopped matching print the same number.

One kernel, one tree

The floor and the kernel of record are the same release, 7.3, built on the maintainer's machine, so the syntax gate and build-check.sh against that tree, both run by the pre-push hook on every push, are the only builds there are. Hosted CI builds nothing: the runner has neither that tree nor headers at the floor, and fetching and building a kernel there only repeated what the push had done. From 2026-09-03 to 2026-09-05 a second CI job fetched a 6.15 tarball, built the kernel and linked the module against it, because the floor was 6.15 and nothing had ever compiled there. It found two spellings the floor lacked, then a type rename, then a codec defconfig leaves out, then a config edit olddefconfig silently undid, then the ->write_begin signature, and the floor moved to 7.3 with the job deleted; README.porting.md has the ruling. What that job taught survives it: build-check.sh takes a KDIR, so a build against any tree is one command, and a build that has never been run against the tree the #error names is an assertion and not a constraint. The constraint is measured from the other side too: KDIR pointed at a mainline 7.2 tree with H2_KERNEL_REF=7.2 fails 44 of 46 checks, and the errors under the #error are the two facilities the floor exists for; README.status.md quotes them.

Between 2026-08-29 and 2026-09-03 this repository sent twenty-six failed CI runs, counted by asking the API which step failed in each rather than by remembering: twenty-one at a module build, three at the repository gates, and two at the floor job's own assertion that its kernel tree carries a symbol table. Two of the twenty-six were the gate's fault, a bare warning: matching kbuild's compiler banner; every other one was a real defect, in the tree or in the workflow being written at the time. So the gates were right and the volume was a working habit rather than a defect rate: CI was being used as a compiler, one push per question, and each answer arrived as a failure notification to the maintainer.

test-doc-prose.sh runs vale over every tracked .md file with the styles in styles/, which are house YAML rather than a downloaded package so the gate needs no network and no vale sync. It arrived on 2026-08-29 with doc/research/, which had been governed in Saxum since 2026-08-25 and was moved here on the rule that a component owns its own development.

Two things about it were wrong until 2026-09-02 and are worth recording, because both are the shape where a gate prints and still passes. Vale's own exit status is nonzero for errors only, and every rule in styles/Hammer2 is a warning, so the gate printed twelve findings and exited 0 on every run it ever made. It now counts the findings itself and fails on any of them; the twelve, eleven British spellings and one wordy phrase, were fixed in the same change. And CI never installed vale, so the gate reported COULD-NOT-RUN on every push, which is the same defect the move was meant to close, one layer out. A third was wrong until 2026-09-04: the population was find doc, which is every document except the four a reader meets first, so README.md, CONTRIBUTING.md, CHANGELOG.md and the pull request template were ungoverned. That is the same defect the gate's own header records about doc/research/, applied to the directory that prompted it rather than to the tree. The README's opening paragraph said this port does not mount anything for four days after it began mounting, and when the population widened those four files held seven British spellings. The population is now git ls-files, so a new document is governed the day it is committed, and the three root files are asserted by name rather than counted: a population that narrowed back to doc/ would still be non-empty and would still pass, which is how this gate read past the README. CI now installs vale pinned by version and sha256, for the reason checkpatch is pinned: a different checker reports a different finding set on unchanged prose. The version of record is 3.18.0.

The gate asserts a non-empty population before it reads anything, so a doc/ that has moved fails rather than passing on an empty sweep, and it carries no negative control of its own because the failing direction was driven by hand: a one-line document containing a British spelling turns the run red and its removal turns it green again.

test-provenance.sh reads doc/provenance.csv and asks three things: that no file under src/ lacks a row, that no row names a file that is gone, and that every row claiming a byte-for-byte carry still IS one. Only the third asks a question this repository cannot answer alone, and it is the reason the gate exists: an origin, commit and license claim is the first thing an upstream reviewer checks and the last thing anyone can reconstruct afterwards. So it is re-run with cmp against the origin clone rather than read. Where no clone is on the machine, nothing was verified that this tree could not verify about itself, and the gate exits 2 rather than passing on a table that only agrees with itself; CI clones the origin at the commit the CSV names so that check runs on every push. What it cannot do is in its own header: derived and ours rows have no mechanical test, so they are counted in the summary rather than checked.

test-syntax.sh compiles hammer2.h and hammer2_io.c against the real kernel headers with two compilers, clang and gcc, under a W=1-class warning set. Two compilers because they disagree about what is worth saying, and a single one is a single opinion: both independently reported the LIST_HEAD and RB_ROOT redefinitions, which is what made those credible rather than stylistic. A warning in a file under src/ fails the gate; one in a kernel header does not, since we do not own those and cannot fix them. gcc is optional and the gate says so when it is absent. Its header line names WHICH resolution it took - KDIR, /lib/modules/$(uname -r)/build, or the nix-store fallback - because a fallback that has never fired is indistinguishable from one that works. That is not hypothetical here: IO_MODEL.md described the nix branch as the source of the kernel of record while the /lib/modules path was present on every run, so the document and the script agreed in wording and disagreed in behavior, and nothing could notice. Point KDIR at a path that does not exist to exercise the fallback: it then resolves nothing and returns COULD-NOT-RUN naming itself.

It also refuses a kernel that is not the one of record: this tree compiles against the latest Linux, the pin is KERNEL_REF in the script, and a tree of any other version is COULD-NOT-RUN. H2_KERNEL_REF checks another version deliberately, which is the only way that reads as a pass. It carries two more controls: a wrong folio call that the same headers must refuse, and the 64KB ceiling guard, which must fire when the ceiling is shrunk. Set KDIR to test against a tree other than the running kernel's.

test-checkpatch.sh is the odd one: it does not ask for silence, it asks that the recorded deviation set has not grown. It identifies its checker by sha256 against the baseline's second line, and prints where every value came from, because an ASSERTED version and a DERIVED one used to render identically: pointed at the v6.15 checker with CHECKPATCH_REF=v7.2, it reported a real style regression, the assertion having laundered a wrong checker into a verdict about this code. A content mismatch is COULD-NOT-RUN now whatever names the checker carries. It finds that checker by the same search test-syntax.sh uses when neither CHECKPATCH nor KDIR names one, so the unattended run reaches the tree of record wherever the syntax gate does; before 2026-09-29 it stopped at the host tree and the store and skipped $HOME/kernels, which is where this file says the trees live, so on that machine the syntax gate passed and this one refused. See README.kernel-style.md for why this tree is BSD style on purpose and what that means for mainline. Both of its sorts are LC_ALL=C, because the baseline is compared byte for byte and glibc collation differs between machines; the first CI run failed with every count identical and four lines in a different order.

None of the compile gates runs anything. -fsyntax-only compiles nothing and links nothing, which is the honest limit of what can be checked before a module builds.

The repository gates check the documentation against the tree rather than the tree against a compiler. test-inventory.sh reads the three hand-maintained lists that claim to cover src/sys/fs/hammer2/, reports a file missing from any of them, and compares the origin table's line count against the file it names. test-citations.sh reads the file:line citations that sit in a doc/ table row against the line each names, comparing against the source rather than a stored baseline, and grades each pass by how specific its anchor is, so a row anchored on a common token is reported as weak rather than counted with the strong ones. It checks table rows only, and prints how many citation-shaped tokens it left alone: 33 on 2026-09-05, of which 31 are in doc/research/ and name line numbers in DragonFly's own tree, which this repository cannot resolve. That is why the exclusion exists, and the count is printed rather than assumed so a prose citation into src/ cannot hide in it. The one that is such a citation, hammer2_inode.c:1664 in doc/README.status.md, was read by hand on the same day and lands on the size comparison the sentence describes. This paragraph's own example is the thirty-third. test-history.sh checks that every changelog row's commit hash resolves with a matching subject, and names any deliverable commit that has no row. It reads one thing no other gate reads: the documents that state which version the tree is on. A document naming an older one describes the state before the milestone it sits above, and that had happened four times in two days (0.2.116, 8d2495a, 0.2.130, and the roadmap and status openings after 0.9.25), so the newest row's number is now checked against every sentence in a state document that names one. CHANGELOG.md and doc/history/ are excluded as measurements of the version they were written at. The check asserts it found at least one such sentence and carries a control that must read 0.0.0 out of a planted sentence on every run, since a matcher that never fires reports clean; both directions were driven on 2026-09-26 by planting a stale sentence and by breaking the matcher. test-inventory.sh also reads the DEFER ledger in doc/README.status.md against src/ in both directions: a marker with no row, and a row whose marker the source no longer holds. The second is the one nothing else would catch, since a deleted marker leaves a row reading as outstanding work forever. The match is on the marker text verbatim, so rewording a trigger in one place and not the other is a failure rather than a drift. Both directions were driven on 2026-08-26 by making each break in turn.

test-inventory.sh has a second population, test/, where every file must either be named by a gate or be listed as staged below. It also checks two DIFFERENT claims about the gates themselves: that no document states a wrong COUNT of them, and that the three documents printing runnable command lists NAME every one. Those are not the same check - The agent instructions file, since untracked, said "eight" correctly on 2026-08-26 while listing seven, so the count passed and the list a future reader would run was short by the newest gate. What none of them can check is whether a row's CLAIM is true; that takes a person reading the artifact the row names.

It has a third population, added because the second could not see this one: the RUNNER behind each runnable file under test/. Being named by a document and being run are two properties, and the second population only gates the first, which is how three exercisers in three milestones came to have a reading published with nothing running the file that produced it: the fallocate one at 0.9.37, the fiemap one at 0.9.39, and test/getdents-resume.c at 0.9.40. Each was found by a hand sweep, so the class kept recurring; the gate is that sweep. Every .c under test/ must appear in script/*.sh outside a comment, which is where a run, a compile or a read of it lives. The two files compiled by Saxum through LINUX_HAMMER2 are named in the gate and skipped, because neither a search of this tree nor the gate can see that consumer. The population is asserted at eight or more, so a sweep that matched almost nothing cannot report clean, and a planted .c with no runner fails it in both directions.

test-syntax.sh --selftest and test-checkpatch.sh --selftest check the two prints that separate a loosened run from a real one: the override warning and the checker's sha256 provenance. Both prints were added on 2026-08-26 to fix the class where output nobody reads is trusted, and neither was read by anything, which is that same defect arriving inside its own repair. The syntax selftest failed on its first run because the warning WRAPS and the matcher read one line at a time - a rule about matching wrapped prose not firing while writing a matcher for wrapped prose. CI runs both, and derives which gates have a --selftest rather than naming them. They re-invoke their own gate with bash, never sh: the first version used sh "$0", which works on a machine whose /bin/sh is bash and is a syntax error under dash, so it passed here and failed on the runner. That is the class of defect a local run cannot reach, and it is what CI is for.

The syntax selftest's unoverridden direction is exercised only where the kernel of record is present. Elsewhere it prints a note saying it was not exercised, rather than failing: on a hosted runner that tree is absent, and failing there would turn an environment difference into a red gate.

The syntax selftest's third check is a designed guard replacing an accidental one. The gate reads VERSION/PATCHLEVEL from a build tree's own Makefile, so linux-api-headers cannot satisfy it - and that immunity was luck of construction, not intent, until the check existed. A guard nobody designed is a guard nobody maintains. The specimen is a directory holding nothing but an include/linux/version.h claiming 7.2, which must be COULD-NOT-RUN. Falsified with the plausible improvement a later maintainer makes, a version.h fallback when the Makefile is missing: the gate then accepts the fake and prints 7 check(s), 5 failed against the kernel of record (7.2), charging five failures to this code on behalf of a kernel that does not exist.

Every gate that uses a toolchain names the one it used, and every gate that resolves a tree names how it resolved it. The reason is a shape worth recognizing: where the DEFAULT invocation and a deliberate one answer different questions, the unattended run and the careful run disagree and only the careful one is ever right, while both print the same summary. It was live on 2026-08-26 in the syntax gate, which needed KDIR typed to reach the kernel of record, and in the style gate, which read the host's build tree while the record's own checkpatch.pl sat in the store. It was live in test-shim.sh too, though not in the direction that sentence guessed. It said a reviewer reaching for clang would get a different opinion; what happened on 2026-08-26 is that two versions of the same compiler disagreed. GCC 13 on the runner accepted a struct file the stub tree never declared and GCC 16 here warned about it, so 346dac6 was green in CI and red on the workstation. The gate names its compiler in its header line and does not require one, which is the right trade for a gate whose whole point is needing nothing but cc: two opinions are worth more than one pinned opinion, as long as a disagreement is read as a finding rather than a flake. It is still true of the vectors contract.

That sentence opened with a count instead of naming them, and script/test-inventory.sh failed it: a number word immediately before "gates" is compared against how many gates exist, so a PARTIAL count in that shape is a finding. The gate was right, and naming them costs nothing. The first attempt to document this failed the same check again, because writing the offending phrase in an explanation is still writing it - the check reads the file, not the intent.

On a hosted runner test-syntax.sh is the only gate that declines, because the kernel of record is the latest release and ubuntu-latest ships headers years behind it. That is recorded as a skip and never as a pass. Everything else runs there, including the vectors contract's behavioral half, which had been declining for want of an xxHash to link until libxxhash-dev was added to the runner on 2026-08-26.

test-posix.sh parses the gates declaring #!/bin/sh with dash and busybox ash. It exists because every gate here is normally run by bash, so a bash-only construct in such a script runs forever and breaks the day something honors the shebang - which happened on 2026-08-26, when both selftests re-invoked their gate with sh "$0" and failed on a runner whose /bin/sh is dash.

It MEASURES its own reach on every run and prints it as observed, rather than asserting a table. A gate stating its own coverage is a claim nothing checks, sitting in the one place a reader uses to decide whether a clean run means anything. The first two versions of this gate carried such a table, and both were wrong: one named a single construct, the next three, where there are four. An under-claim is as false as an over-claim and is re-checked less often.

What it asserts instead are two properties of a working checker, which hold whatever the reach turns out to be: each shell must reject at least one probe, or it is inert here and a clean result means nothing; and a plain POSIX script must be accepted, or the instrument refuses everything and a clean result is unreachable rather than earned. Falsified both ways - making every probe inert, and making the plain script unparseable, each fail naming which property broke. With no shell realized the gate exits 2.

Since 2026-09-05 it also reads the remote blocks the fleet scripts hand to ssh, which are shell that no gate had ever parsed. Those blocks are passed as one single-quoted string, so a single quote anywhere inside a block ENDS that string: everything after it is expanded by the workstation instead of the guest. A shell parse cannot see that, because the text left behind is still valid shell, so the gate checks the property the quoting actually rests on and counts the single quotes in each block, which must be zero. It parses the blocks as well, and asserts the population, since a pattern that stopped matching would check nothing and still print a clean count. The check was written after an awk '{print $3}' added to test-enospc.sh aborted a run on the workstation with $3: unbound variable before the guest was reached, and it reads 2 against that block and 0 against the repaired one.

Observed on 2026-08-26: dash rejects process substitution, arrays, the function keyword and here-strings; busybox ash rejects arrays and here-strings; both accept [[ ]], declare, local, +=, arithmetic, ANSI-C quoting and brace expansion as ordinary words. So a clean run means "no bash SYNTAX" and never "no bashisms". The hosted runner reproduced those figures exactly, on a different distribution.

The shells are found in the nix store, not on PATH. An earlier version of this section said no POSIX shell existed on the development machine, on the strength of command -v, while both were already installed under /nix/store. On a machine whose software lives in a store, PATH answers for the current session and not for the machine.

Run what CI runs, before pushing

CI was being used as a compiler: one push per question, each answer arriving as a failure notification. Twenty-six such runs between 2026-08-29 and 2026-09-03, and two more on 2026-09-04 when the checkpatch pin moved and the gate selftests were not run locally, which CI runs as a step of its own.

script/pre-push-check.sh runs every script/test-*.sh and every gate selftest, the selftests enumerated by implementation rather than by name the way CI enumerates them. Install it:

$ ln -sf ../../script/pre-push-check.sh .git/hooks/pre-push

It is not named test-*.sh and does not live among the gates as an equal: the thirteen are run individually on purpose, and one exit status for thirteen questions is the thing this repository does not want. H2_SKIP_PREPUSH=1 overrides it for a push that is deliberately ahead of a green tree.

It runs what CI runs, and CI's result still has to be read on the forge after the push: three pushes on 2026-09-04 were red there on one undefined symbol while every local gate was green, and each got its changelog row. gh run list after a push is part of the push.

It fetches the checker the baseline records and caches it under $XDG_CACHE_HOME/linux_hammer2, keyed by the sha256 the baseline names. Without that the style gate finds whatever checker the machine has, cannot attribute a moved deviation set to this code, and reports COULD-NOT-RUN, which this check would report as a warning: a real style regression would then pass a push. Both directions were driven before it was committed. On a clean tree it exits 0; with one category count altered in the baseline it names the moved line and exits 1.

Making a fixture to mount

There is a Linux-native HAMMER2 writer on this machine and there has been since 2026-08-25: Kusumi's makefs port, packaged by the distribution and built in the store. doc/research/HAMMER2_TEST_FIXTURE_PLAN.md names it and every other writer in the fleet. Nothing in this repository has to be written to produce media, and a claim that media is unavailable is a claim about a search, not about the machine.

makefs -t hammer2 -o Label=TEST,MountLabel=TEST <image> <tree>

Both options or it fails: Label= creates the PFS and MountLabel= picks where the tree lands, and with only the first the tool creates the PFS and then tries to mount the default DATA.

Two costs measured rather than assumed. The image is fully allocated and not sparse, so it occupies its full size on the filesystem holding it, and the default size is 8 GiB. -s 1g is refused, and the tool reports trying default image size 8.00GB, which does not agree with the toolchain document's statement that images size in 1 GiB chunks and the packaged test takes one. Whichever is right, an 8 GiB write per fixture is what this machine does today, so fixtures belong on real storage and want caching rather than regenerating.

The mount names the PFS, and a mount without one asks for DATA:

mount -t hammer2 -o ro /dev/loop0@TEST /mnt/h2

An image made by makefs is a Linux tool's output. The milestone's claim is about media DragonFly wrote, which is the dragonflybsd642 guest in the fleet, and the two are different measurements.

The two fixtures

f1.img holds five paths and its largest file is 16 bytes, which is under HAMMER2_EMBEDDED_BYTES and so lives in the inode. It proves the directory operations and nothing about the on-media read path, because every file in it takes the embedded branch.

f2.img exists for that reason and is built on the boundaries the read completion branches on: 511 bytes, 512, 4096, 65536 and 200000. The pairs matter more than the sizes. 511 and 512 straddle the embedded limit, so one reaches the inode and the next reaches a data block, and 200000 is the only one where a folio can begin partway through a block, which is the arithmetic upstream has no counterpart for.

makefs -t hammer2 -o Label=TEST2,MountLabel=TEST2 f2.img tree2

Compare with md5sum against the tree the image was made from rather than by reading the files, and read at an unaligned offset inside the largest with dd skip=, which no whole-file compare exercises.

f3.img and f4.img are the same tree written with CompressionType=lz4 and CompressionType=zlib. That option is why the compressed floors stopped being unmeasurable: they had been recorded as needing a fixture that did not exist, when the fixture was one flag away. Before writing either decoder, run the floors against these images. The compressible file must fail with EIO and name its method in dmesg while the incompressible one reads, which proves in one run both that the floor refuses rather than corrupts and that the image holds what its name claims. A decoder written first would have had neither guarantee.

tree3 is chosen so that one file compresses and one cannot, since a compressor that falls back on incompressible data writes a raw block inside a compressed volume, and that mixed case is the one a single compressible file would miss.

makefs -t hammer2 -o Label=TEST3,MountLabel=TEST3,CompressionType=lz4 \
    f3.img tree3

Check these two with the page cache dropped as well as cold, and scan kmemleak twice: both decompression paths allocate per folio and the ZLIB one also allocates an inflate workspace, so a leak here is per read rather than per mount.

The fixture gate

script/test-fixtures.sh is every read-path measurement in this document run without a person typing them. It builds the module, starts the guest if it is not already running, attaches each image in H2_FIXTURE_DIR whose manifest is committed under test/fixtures/, mounts it read-only by the label the manifest names, and compares every file with md5sum -c.

The manifests are here and the images are not, because an image is 2 to 8 GiB and fully allocated while the manifest is the part that constitutes the claim. f5.manifest holds the checksums DragonFly itself reported before unmounting, which is the only form that measurement can be kept in: regenerating them on Linux would compare this module against itself.

Each manifest also records the on-media block count for every file, which i_blocks carries, and the gate compares those too. A # link target relpath row records a symlink, which md5sum follows and so never reads; the gate compares readlink against it. That is an assertion about the fixture rather than about the code: a set of matching checksums cannot tell you that an image still holds compressed blocks, so an image regenerated without compression would pass every checksum while the compressed path silently stopped being exercised. What the counts do not distinguish is one compressor from another, f3 at LZ4 and f4 at ZLIB reporting the same numbers because both compress below the smallest allocation; the floors run before either decoder existed are what established which image is which.

The counts corrected a claim in README.status.md on their first run. d512.bin was described there as the first file on media, and it reports zero blocks: hammer2_inode.c compares size > HAMMER2_EMBEDDED_BYTES, so the bound is inclusive and 512 bytes is still embedded. Two fixture files were testing the same branch while the document said they straddled it.

It also exercises the ioctl surface on every image that verifies, through test/hammer2-ioctl-exercise.c, which the gate compiles statically here and runs on the guest twice: once as root and once under setpriv --reuid=65534. The program includes the driver's own hammer2_ioctl.h rather than copying the command numbers, so a renumbered command breaks its build instead of passing against a stale constant, and it prints one label value line per call which the gate compares literally. Ten results per image, a hundred over the fixture set.

It covers the read-only subset and says so in the summary line. The fixtures are attached read-only on purpose, so snapshot creation, PFS create and delete, growfs and bulkfree cannot be reached from it and stay hand-verified on a guest; "the ioctls are gated" would otherwise stand for commands nothing here runs. What it does cover is the version, PFS and inode queries, the super-root scan, the volume list, and three refusals: an unknown command under HAMMER2's own type letter, a command belonging to another driver, and a zero-size command, which is reachable only as root because the entry point checks the capability before the size.

test/hammer2-seek.c is run by test-enospc.sh, not by the fixture gate: it has to write the file whose holes it asks about, and the fixture images are attached read-only because the fixture is the claim. It builds a file whose shape it knows, one block of data, one block of hole, one block of data and a truncate into a fourth, and asks where the data is at twelve offsets. SEEK_DATA and SEEK_HOLE are the one read-path facility whose failure is a wrong answer rather than a refusal: a filesystem registering no ->llseek of its own is answered by generic_file_llseek(), which treats the whole file as data, so SEEK_HOLE reports end-of-file for a file whose middle is a hole and a sparse file copied with cp --sparse=always or archived with tar -S comes out dense and looks correct. The test asserts the hole is real on media first, via the block count, because a file whose middle block was written as zeroes is not sparse and the run would be measuring a different file from the one it built. Against the module before this work it failed six of twelve checks, SEEK_HOLE at 0 answering the file's size where the hole ends one block in; doc/history/verification-record.md has both runs.

That first version of the exerciser also passed two checks that were wrong, on this port and on no other filesystem: SEEK_HOLE inside a hole and at i_size, both of which tmpfs and btrfs answer differently. The expectations came from this driver's behavior instead of from lseek(2), so the test agreed with the bug it was written to catch and reported a clean run. Corrected, the same twelve checks pass on these filesystems, which need no filesystem under test to run against and are the control to reach for first: cc -O2 -o /tmp/h2seek test/hammer2-seek.c && /tmp/h2seek /tmp/probe (tmpfs) and the same against a path on btrfs. The record has the detail.

Two defects came out of its first two runs, one in the driver and one in the exerciser. An unrecognized command returned EOPNOTSUPP, which a BSD's ioctl layer maps to ENOTTY and Linux does not, so userland read "Operation not supported" where every other driver says "Inappropriate ioctl for device"; the dispatch's default arm now returns ENOTTY and the deliberate refusals above it stay EOPNOTSUPP. The exerciser then reported zero volumes from a call that had succeeded, because nvolumes is the caller's capacity going in and it had been zeroed. That second one is the reason the gate checks the two counts for being non-zero rather than only checking the status: a scan that copies nothing returns success.

test/hammer2-dedup.c runs on the same volume, by the same gate and on the same terms. Deduplication is on by default, hammer2_dedup_enable being 1, and the write path asks hammer2_dedup_lookup() before allocating every data block, so a second copy of a block should point at the first one's media. Nothing here had ever written a duplicate block: throughput.sh is the only script that names dedup and it draws fresh data every pass so a run cannot read as a dedup hit, and README.md's opening paragraph names block-level deduplication among the format's features. The exerciser is black-box: it writes one file, reads statfs, writes a second file holding the same bytes, and reads statfs again, so a shared block shows as a small second cost against the first file's. It asserts the first file moved the count before it interprets the second, since a duplicate that costs nothing and a run that wrote nothing look the same in the difference alone.

That exerciser's control is a filesystem with no dedup, and there are two on any machine that has them: on tmpfs the duplicate costs its full allocation, and on btrfs 1032 blocks against the first file's 1024, so both fail the third check. The control needs no guest and is the check that says the test can detect the absence of the feature at all. It is run by copying the unit out of statfs rather than assuming 64 KiB, because btrfs reports 4 KiB and a test that only runs on this port cannot be controlled by a filesystem that is not this port.

test/hammer2-fallocate.c runs on the same volume, by the same gate and on the same terms, and it is the instrument for the operation the readiness audit called the largest real gap. No BSD port carries a fallocate vop, so before this the call fell through to the kernel's default and every caller was told EOPNOTSUPP, which an installer and a package manager both meet. The format makes the modes converge in a way worth stating: a write whose block is all zeros is not stored, so a hole and a zeroed range are the same thing on this media and PUNCH_HOLE cannot be told from ZERO_RANGE afterwards. Preallocation as ext4 means it, blocks reserved and readable as zeros, does not exist here, because the media never holds a zero block.

The exerciser therefore checks what is observable and true on any filesystem rather than that a preallocation reserved media: a punched or zeroed range reads back as zeros with the bytes outside it compared byte for byte, a punch leaves a hole and never changes the size, KEEP_SIZE never changes it either, and a plain allocate extends the file. The file is written from a pseudo-random stream with no zero block in it, because a file that was already zeros would read as a hole and a punch of it would prove nothing, and its size is asserted to have moved before any punched reading is believed.

Two things make it a control and not merely a check. Its punch range is one whole block taken from statvfs, because the first version punched 16 KiB from a fixed offset: on the 4 KiB-block reference that covers whole blocks, and on this port's 64 KiB block it covers none, so nothing could be freed and the exerciser reported a failure of a correct operation. And whether the punch freed anything is asked of SEEK_HOLE, not of st_blocks, since this port's block accounting moves at allocation and at the bulkfree pass rather than at an elision. btrfs accepts every mode and runs all sixteen checks, and tmpfs accepts PUNCH_HOLE while refusing ZERO_RANGE, so the same binary on both shows what a refusal looks like against this port's answer; a mode that is refused prints falloc-skip and is counted, and the gate fails a run in which no check ran.

The exerciser also punches a range spanning SEVERAL whole blocks, on a file of its own, and asserts that exactly the requested bytes read as zeros and nothing outside them moved. That range is the shape a walk defect lives in: the one-block punch above enters the page cache walk once, so it could not see hammer2_fallocate() advancing by PAGE_SIZE over read_mapping_folio(), which returns the folio containing the index, and taking a 64 KiB folio sixteen times. The extra check is a correctness guard for that range and not a discriminator for the walk, passing on the pre-fix build as well, because a byte-compare cannot observe how many times a folio was visited; the walk's own reading is the timing in doc/history/verification-record.md.

test/hammer2-fiemap.c runs on the same volume, by the same gate and on the same terms, and it exists because the two operations it covers had their numbers published with nothing to reproduce them. ->fiemap and ->freeze_fs/->unfreeze_fs landed together as 0.9.38 with eleven checks and zero failures written into the changelog row and the readiness audit, and no script in this repository ran the file that produced that count: test-inventory.sh knows a file is NAMED and never that anything runs it, which is the gap this wiring closes. FIEMAP is the read-path operation whose failure is a wrong answer rather than a refusal, and in the same direction as SEEK_HOLE: a filesystem carrying no ->fiemap is answered EOPNOTSUPP from fs/ioctl.c and filefrag(8) is told nothing, but a stub returning one extent covering the whole file answers success and describes a sparse file as dense. The exerciser opens the hole in the middle of the file itself with a sparse seek, asserts against st_blocks that the hole is real before asking, and compares the returned map against the shape it built, so the first fake pass is caught by the gap check and the second, a map whose physical addresses are all zero, by a check that at least one extent carries an address.

Its control is the file rather than a reference filesystem, and the reason is in the file: FIEMAP and the seek whences both describe this media's own layout, so a control on btrfs would compare two different layouts rather than check an answer. tmpfs is not the control either, because it refuses FIEMAP outright, measured EOPNOTSUPP. The consumer that shows the map is good for something is filefrag -v, which reads the same two extents at the same blocks and reports the hole as a gap. The freeze half of the same file measured the operation the first version of the exerciser wrongly condemned: a write while frozen blocks and is released by the thaw, so every blocking call in the test is in a child and every thaw is reached by the parent that cannot block. A run in which FIEMAP or FIFREEZE is refused prints fiemap-skip and runs fewer checks rather than failing, so the gate pins the count at eleven: a volume that refused everything prints one check and zero failures, because the assertion that the file is really sparse runs before the call is made. That is measured rather than reasoned about, being the whole output against tmpfs on the host, and it is why the count rather than the failure number carries the verdict here. Any other count is a failure, in either direction.

test-fixtures.sh carries a negative control per image rather than only in the selftest. After a manifest verifies, one hash in it is altered and the same mount is compared again, which must fail. Without that, an empty sums file, a silent md5sum and a mount that landed somewhere else all read as a pass.

It leaves the machine as it found it: what it attached is detached, and the guest is shut down only if the gate started it. It will not start one unless H2_FIXTURE_START=1 says so, because script/pre-push-check.sh runs every gate on every push and a gate that boots a 4 GiB domain when it finds one stopped spends that on every push, on a machine whose memory somebody else is using. H2_FIXTURE_MODARGS is handed to the guest's insmod; H2_FIXTURE_MODARGS=io_buf_only=1 reads every fixture through the block buffer the DIO layer otherwise takes only when the page cache cannot make a 64 KiB folio, so that path is read by the same manifests as the page cache path rather than waited for.

Both fleet gates wait five minutes for a guest that does not answer ssh before giving up on it, which is what a boot takes and also what a guest that is up but wedged costs a caller who meant only to check the tree. H2_GUEST_WAIT is the number of five-second rounds to wait, defaulting to sixty, and it bounds the waiting rather than the check: the pre-push hook sets it to two, so a push against a wedged guest ends in a report in seconds instead of holding the connection open, and a hand run keeps the five minutes that distinguishes a boot from a hang. When a guest is running and still silent at the end of the wait, the gate prints every vCPU's instruction pointer, the disk requests over three seconds and whether the guest agent answers, read from outside before the exit trap shuts the guest down. All twelve vCPUs of artix-s6-kde were at one address in module space with no disk requests on 2026-09-26, which is the deadlock doc/history/verification-record.md describes, read from outside the guest.

The push has a bound of its own for the same reason. script/pre-push-check.sh runs every gate in turn, and a gate that outlives the window a remote keeps an idle connection open is worse than one that fails, because the push dies with nothing sent and no gate having returned a failure. H2_PREPUSH_BUDGET is the number of seconds the whole hook may take, defaulting to 400. A gate past the budget is reported COULD-NOT-RUN with the fact that it was not run, which is what exit 2 already means, and a gate that outlives the seconds remaining is killed and reported the same way. The eleven gates needing only this machine take about forty-five seconds together, so a healthy tree never reaches it, and H2_PREPUSH_BUDGET=<seconds> raises it for a run that means to wait.

Exit 2 is COULD-NOT-RUN and is reported for a missing guest, a stopped one without that variable, missing images, no KDIR and a guest that does not answer ssh, because most machines have none of these and CI has none at all. A gate that passed there would make the whole read path look covered by CI when nothing ran.

The module has one build-time control of its own, never installed. make HAMMER2_FOLIO_CONTROL=1 produces a module whose mount-time folio-size check asks for twice what the kernel offers, so it must refuse every mount and name both numbers. Build, load on the guest, run, read dmesg. A second control, HAMMER2_RW_EXPERIMENT, lifted the read-write mount refusal for measurement from the first read-write mount to the crash matrix, and retired with the refusal; the runs doc/history/verification-record.md records under that name were made with it.

One of those is worth its own line. KDIR defaults to the host's own build tree, so the first pre-push run of this gate built for the host and reported the guest refusing to load it as a failure. insmod rejects a module on vermagic, which is knowable before the attempt, so the gate now compares the module's vermagic with the guest's release and reports COULD-NOT-RUN naming both. A verdict reached against the wrong kernel is an artifact of the setup and not a finding about the code.

KDIR=~/kernels/linux-7.3-rc5 H2_FIXTURE_START=1 \
    bash script/test-fixtures.sh

Measured on 2026-09-04: eleven images, 43 files, 34 stat rows, 5 statfs rows, 2 symlinks, one corrupt file refused, one mount refused, 0 failures. f12 is f1's tree written by DragonFly's kernel, for the listing in README.status.md. A # stat mode nlink uid gid inode relpath row and a # statfs size used free inodes-used row carry what DragonFly's own stat and df reported, and f11 exists to hold hard links, a setuid bit, an owner and a 0750 directory. The ten the gate mounts are f11, makefs output, makefs at LZ4 and at ZLIB, the boundary tree, media DragonFly wrote at its LZ4 default, media DragonFly wrote after hammer2 setcomp zlib on the mount root, a device carrying two PFSes of which the gate mounts ROOT, and two copies of the LZ4 image altered on purpose: f9 with one data byte flipped, whose manifest carries # corrupt random128k.bin and whose other files must still verify, and f10 with one volume-header bit flipped, whose manifest is # refuse and no file rows. f8, the installed DragonFly root, has no manifest here: it is read by Saxum's walker as root and compared against two other readers, recorded in doc/history/verification-record.md. Mounting both of f7's PFSes at once is a measurement recorded there too, since a manifest names one label.

The gate starts its guest only when no other domain is running, since each holds 4 GiB and the host is shared with other sessions' benches; H2_FIXTURE_SHARE=1 overrides that. It attaches one image at a time, always as vdb, and releases it before the next. It used to hold every image attached and the guest ran out of virtio slots at the eighth, which the gate reported as an attach failure of its own making. For f6 the block counts in the manifest are DragonFly's own stat output, so that image compares this reader against the writer rather than against itself.

Media DragonFly wrote

Everything above is makefs output, which is a Linux tool. For the milestone's own claim the writer has to be DragonFly:

virsh start dragonflybsd642
virsh attach-disk dragonflybsd642 <image> vdb --targetbus virtio
virsh reboot dragonflybsd642        # no virtio-blk hotplug there
newfs_hammer2 -L DFLY /dev/vbd1
mount -t hammer2 /dev/vbd1@DFLY /mnt/h2w

To ask whether a device callback reaches every PFS mounted on it, wrap the device in a linear dm target on the Linux guest, mount two PFSes from /dev/mapper/<name>@<label>, dmsetup suspend it, and fsfreeze -u each mount: the thaw exits 0 on a frozen superblock and fails with EINVAL on one the freeze never reached. doc/history/verification-record.md records the result for f7 with and without the per-mount claim.

Attach without --mode readonly on the DragonFly side: its HAMMER2 opens the device for writing whatever the mount asks, and a read-only attachment fails the mount with EINVAL.

For f6 the same, with the disk attached to the shut-off domain under --config so no reboot is needed, and hammer2 setcomp zlib /mnt/h2w run before the first file is written, since the setting is inherited by new inodes and does not rewrite existing ones. Root over ssh works with the key; the unprivileged user's doas asks for a password.

Two things about that guest cost time. Its root shell is csh, where 2>&1 is a syntax error rather than a redirect, so run anything with redirection through sh -c or copy a script over. And it does not hotplug virtio-blk, so a disk attached to the running domain needs a reboot to be enumerated; virsh reboot keeps the same QEMU process, so a live attachment survives it where a shutdown would lose it.

Take the checksums on DragonFly before unmounting. They are the ground truth, and taking them afterwards on Linux would compare this module against itself.

Run one guest at a time. Each holds 4 GiB, both together are most of what this machine has spare, and the two halves of the test never overlap: DragonFly writes, then is shut down, then Linux reads.

The Linux guest's kernel is plain mainline 7.3 at the candidate the kernel of record names, built on the host and copied into the guest by hand: 7.3.0-rc1 from ~/kernels/linux-7.3-rc1 until 2026-09-26, then 7.3.0-rc4 from ~/kernels/linux-7.3-rc4, and 7.3.0-rc5 from ~/kernels/linux-7.3-rc5, the tree KDIR names, since 2026-09-29. All stay installed and the grub default chooses between them, so a reading in doc/history/verification-record.md taken on that guest names the build it ran on, and one dated before 2026-09-26 was taken on rc1. Its configuration is the tarball's default plus the debug and instrument options the port's readings depend on: PROVE_LOCKING with DEBUG_RWSEMS, PROVE_RCU, DEBUG_KMEMLEAK, BLK_DEV_IO_TRACE, DWARF 5 debug information so addr2line resolves module offsets, PREEMPT under PREEMPT_DYNAMIC, LOCKDEP_CHAINS_BITS at 20 since the million-file runs filled the default table of 16 bits at 172 s and switched the validator off, and SQUASHFS and EROFS_FS as modules for the closure reference reads. A change to that configuration changes the guest every reading after it is taken on, so it is recorded here and the build number uname -v prints is recorded beside the readings it first appears in.

A second build of the same source sits beside it in the guest since 2026-09-07, the release build: 7.3.0-rc1-release from ~/kernels/linux-7.3-rc1-release, 7.3.0-rc4-release from ~/kernels/linux-7.3-rc4-release since 2026-09-26, and 7.3.0-rc5-release from ~/kernels/linux-7.3-rc5-release since 2026-09-29, the same configuration with every debug option off, chosen at boot by the grub default, which setkernel.sh on the guest rewrites. A third build, 7.3.0-rc5-kasan from ~/kernels/linux-7.3-rc5-kasan, sits beside them: the debug configuration plus KASAN inline with KASAN_VMALLOC, UBSAN with the bounds, shift, bool and enum checks, DEBUG_ATOMIC_SLEEP, DEBUG_LIST, DEBUG_VM, DEBUG_OBJECTS and the fault injection framework with FAILSLAB, FAIL_PAGE_ALLOC and FAIL_MAKE_REQUEST behind debugfs. The debug kernel's lockdep and kmemleak see locks and leaks and nothing else; an out-of-bounds read inside a folio, a use after free the allocator has not yet reused, a sleep under a spinlock, or a shift past the type are this build's reading and no other's, and it is where a fleet gate runs when the question is memory rather than order. It answers a second question by accident: it runs at about half the debug kernel's speed, so a window two tasks have to land in together is wider, and its first fill found a folio-lock against inode-lock cycle between a writer and the sync loop that twenty-odd fills on the debug kernel had never hit and lockdep cannot see, a folio lock being a page bit and not a class. A gate that has passed on the debug kernel has not been run slow. The release build exists because a throughput number is a claim about the port and a lockdep kernel charges the port for every lock it takes per block: a profile of the 512 MiB read on the debug kernel put a third of the reader's samples in lock bookkeeping, more in kmemleak's object tracking and page zeroing on allocation, and 3% in the module, and the same read on the release build ran four times faster. throughput.sh prints the guest kernel and whether it carries PROVE_LOCKING beside its numbers, and refuses to be read as a rate without that line. A defect run stays on the debug kernel; a rate is taken on the release one and says so. The release build has no function tracer, so the read_folio count that gate prints reads as unavailable there.

Read stat -c %b on the result as well as the checksum. i_blocks is the on-media count, so it says which branch of the read completion each file took, and a set of matching checksums proves nothing about which paths ran.

Writing to a fixture, and reading the write back on DragonFly

The write path is exercised on f13.img, a byte copy of f5 made with cp before every run, so the untouched f5 is the baseline every comparison is against:

cp f5.img f13.img
virsh attach-disk artix-s6-kde f13.img vdb --targetbus virtio --config
virsh start artix-s6-kde

On the guest, mount without ro, write, sync, umount, remount ro and read the file back; then power the guest off, detach the image, and attach it to dragonflybsd642 the same way. DragonFly's cat and stat, then fsck_hammer2 /dev/vbd1, are the verdict, and the host's hammer2 show from hammer2-utils over f5.img and f13.img, diffed, says which chains the flush rewrote and with what transaction ids. doc/history/verification-record.md records the first such run.

Three things about a write test that a read test never needed:

  • Capture the serial console before the write. A carried hpanic is panic() here, and that guest sits in a panic with nothing written to its disk, so the message exists only if it left the machine. The guest's command line carries console=ttyS0,115200, and on the host

    script -q -f -c "virsh --connect qemu:///system console artix-s6-kde" \
        serial.log </dev/null
    

    records the line into serial.log until it is stopped. The first flush panicked, and without this the panic was an ssh connection reset and a guest that came back with an empty log. A panic and a hung-task report reach the line at the default console loglevel; the module's own hprintf lines are KERN_INFO and do not, so write 8 to /proc/sys/kernel/printk first if the transcript is to carry them. Two runs recorded only the shutdown for want of that.

  • Detach the test from the ssh session. Run it under setsid with its output on the guest's root disk, then read the file after; a guest that resets takes the session with it, and a session that ends kills a test still running.

  • A hung sync is not a hung guest. The hung-task detector reports it at kernel.hung_task_timeout_secs, lowered to 20 for these runs, and names the lock and the holder, which is how the first deadlock was read. virsh destroy is then the only way out, and it costs a core like a shutdown does.

Tracing what the flush writes, and in what order

The guest kernel carries CONFIG_BLK_DEV_IO_TRACE but no blktrace binary, and tracefs is not mounted at boot, so the trace is taken with the block tracepoints directly:

mount -t tracefs nodev /sys/kernel/tracing
T=/sys/kernel/tracing
echo > $T/trace
echo 1 > $T/events/block/block_rq_issue/enable
echo 1 > $T/events/block/block_rq_complete/enable
# ... the writes ...
echo "written, syncing" > $T/trace_marker; sync; echo synced > $T/trace_marker
echo 0 > $T/events/block/block_rq_issue/enable
echo 0 > $T/events/block/block_rq_complete/enable
dn=$(lsblk -nd -o MAJ:MIN /dev/vdb | tr -d ' ' | tr : ,)
grep -E "$dn |tracing_mark" $T/trace

The tracepoints name a device by major and minor with a comma between, 254,16, not by its name, and a filter written for vdb matches nothing: the first two runs of this reported no events and looked like an empty write. Print the per-device counts alongside, so an empty filter shows against the root disk's hundreds. doc/README.status.md carries the trace for one write and sync, in which the volume header at sector 0 is the last request and follows a completed flush.

Mutated media against the mount path

script/fuzz-mount.sh N SEED is the corpus 0.5 asks for: N copies of a seed image, each with a few bytes changed at recorded offsets, hot-plugged read-only into the running guest one after another, mounted, listed and read end to end under the shipped build. A mount may succeed or be refused and a file may read or fail with EIO; what fails the run is a WARNING, BUG, oops, hung task or lockdep report in the guest's log, or a guest that stops answering. The corpus is the generator and the seed number: every image's mutations are written to the log as offset:old>new, so a finding reproduces from its seed and index and no image is kept. Two controls run before the corpus, the seed itself which must mount with every file readable, and the seed with one bit of its volume header crc changed which must be refused; a run whose controls fail is a run whose reader or whose refusal detection is broken, and its counts mean nothing.

The seed is a small volume, because the mutator samples until it hits a byte that is not zero and a 2 GiB fixture is almost entirely zero. The mutator also redraws until the new byte differs from the old, because a byte replaced by its own value is no mutation and one recorded mutation in five was, over two seeds, before that check; doc/README.status.md carries the figures and which generator each recorded seed reproduces from. It is made on the host by hammer2-utils and populated through the write path:

truncate -s 64M /mnt/storage/hammer2-fixtures/fz-seed.img
newfs_hammer2 -L FUZZ /mnt/storage/hammer2-fixtures/fz-seed.img
# on the guest, with the module loaded and the image attached:
# mount, create a few directories, files at several
# sizes, a symlink and a hard link, sync, unmount

H2_FUZZ_WRITE=1 is the write side of the same corpus: each image is attached read-write, mounted read-write, read as above, and then written into, a new file, 256 KiB of random data, a directory and one unlink, followed by sync and umount, which is what reaches the block-table, freemap and check-method sites a read never does. Since hpanic marks the device in error and returns, a fault is its own count in the verdict rather than a kernel report, the module is reloaded after every image because the mark is module-wide, and an umount or rmmod that does not return 0 fails the run; a BUG, an oops or a hung task is still a report. The seed control must take every write. Read on 2026-09-07 at ce59742, seed 1, fifty images: forty mounted and ten were refused, as the read side reads them; thirty of the forty refused all four writes with EIO and dirtied nothing, ten took all four and synced, none faulted, none reported, none stuck. A mutation in the first 4 MiB lands in the freemap's reserved zone, and a freemap leaf whose check no longer matches is refused at the allocation rather than allocated over, which is the refusal those thirty read.

The image is copied under /var/tmp/hammer2-fuzz for each mutation, because libvirt takes ownership of a file it attaches and the next copy over it in the fixtures directory is refused. The script builds the module against KDIR, exits 2 without a guest, a seed or a kernel tree, and starts the guest only under H2_FIXTURE_START=1, as the fixture gate does.

The round trip both ways, from the tree

script/f4-roundtrip.sh is F4 as a script: it formats a 2 GiB image on the host with hammer2-utils' newfs_hammer2, builds the experimental module, boots the Linux guest to write a tree and its manifest, boots the DragonFly guest to check that manifest and write a tree of its own, and boots the Linux guest again to check DragonFly's manifest and what is left of its own. Every checksum is the writer's, so neither reader is compared against itself. It refuses to run beside a running guest, exits 2 without both guests, the tools or a kernel tree, and shuts each guest down when its turn is over. KDIR names the kernel tree, and H2_NEWFS and H2_FSCK name the tools when they are not on PATH.

script/cut-flush.sh SECONDS is the interrupted-flush fixture: DragonFly writes small files to a copy of f5 with a sync every two hundred, the host destroys the domain after SECONDS, and the cut-off image is copied. This port mounts one copy read-write, which runs the carried hammer2_recovery(), reads every file, writes one more and syncs; DragonFly mounts that result and then recovers the other copy itself. The header's two tids are printed at each stage, so a run that cut inside the window where freemap_tid lags is visible; none has yet. So the script's fourth stage makes that state on purpose: DragonFly's recovered copy has its header's freemap_tid lowered by H2_CUT_LAG transactions, four by default, and the sector's two CRC32C checksums recomputed, and both recoveries run on it; the run fails unless this port's mount announces freemap recovery over those transactions and both checkers are clean afterwards. The checksum routine is checked against the stored values before the rewrite, so a wrong header layout stops the stage rather than making a corrupt image that would be refused for the wrong reason.

script/crash-matrix.sh SECONDS is the crash matrix, calibrated against Kusumi's FreeBSD port on the freebsd15 guest. An 8 GiB volume is made on the host by newfs_hammer2, so all four volume header zones exist, and a copy of it is attached to a writer, which mounts it read-write and runs the same loop as the cut-flush fixture. After SECONDS one of four things happens: the writing process is killed and the volume unmounted (kill), the guest kernel is made to panic through sysrq here and debug.kdb.panic there (panic), the host destroys the domain (power), or it destroys the domain and then zeroes the second half of the newest valid header, which is a 64 KiB header write that reached the media only in part (torn). The cut-off image is judged three ways: the host's fsck_hammer2, this port mounting one copy read-write, reading every file, writing one more and syncing, and the FreeBSD port doing the same to the other copy and running its own fsck_hammer2; the host checks both results again. The FreeBSD port writes every cell first, so its recovery of its own crash is on record before this port is judged against it, and then this port writes the same cells. H2_CRASH_REPS runs each cell, two by default; the summary reports a cell green only when every run of it produced the same verdicts, and names one that did not. The FreeBSD guest is reached through the freebsd ssh alias as a wheel user with passwordless doas, and needs the port built from freebsd-hammer2-upstream and installed; H2_CRASH_CELLS and H2_CRASH_WRITERS narrow a run. The mirror_tid of every valid header is printed at the cut, so the torn cell shows which header it destroyed and which one the recoveries fell back to.

Every script that drives a guest waits for it the same way: a guest that answers ssh is used whatever the domain says, one listed running that does not answer is waited on, bounded, because a booting guest answers within the wait and one shutting down turns to shut off inside it and is then started, and only one that stays listed running and silent for the whole wait is given up on, reported as COULD-NOT-RUN with the host's load average beside it. Reading the domain state alone was wrong both ways in one day: a guest the fixture gate was still shutting down read as usable and cost the fuzzer its whole wait, and a guest a batch reset had just started read as shutting down and cost the batch both its runs. A guest a script started is shut down from an exit trap, not from its last line, since a COULD-NOT-RUN after the start had left one running for the next script to refuse. The long guest runs are bounded from the host with timeout, H2_RUN_TIMEOUT seconds and 1800 by default, because a guest whose task hangs keeps sshd answering and the ssh open; the bound expiring is reported as the guest hanging, which is a failure and not a skip.

A guest whose task has hung is read before it is reset, not after, because the reset destroys the only report. ssh is often gone by then: sshd's fork touches the wedged mount and hangs with it, as it did on the million-file deadlock. The QEMU guest agent does not, and virsh qemu-agent-command with guest-exec and capture-output runs a shell in the guest and returns its output; dmesg, w and t into /proc/sysrq-trigger, ps with wchan, and /proc/lockdep_stats all came out of that guest in seconds while ssh had been silent for minutes. script/guest-dmesg.sh <domain> is that capture, and million-tree.sh runs it where its run timeout expires, before destroying the guest, so the report a hang leaves is in the log whether or not anyone was watching. Anything that touches the mount hangs the probe too, so the script touches none. Two samples of virsh domstats --cpu-total --block a few seconds apart tell a hang from a slow guest first: no CPU time and no writes across the gap is a hang, and the boot-under-load story is the wrong one. The trace offsets resolve against the module the run built, which carries debug lines, with addr2line -i. After virsh destroy the fixture disk is still attached in the persistent configuration and wants a detach-disk --config of its own.

script/pfs-domains.sh is the half of 0.8 that belongs to this side. PFS roots are the port's storage domains, and the milestone checks them by mounting each by label; which labels a consumer lays down and its installer are the consumer's, and what a volume written here has to satisfy is this: the host formats a 2 GiB image with one root PFS, the Linux guest creates three more through this port's own ioctl, SYSTEM, STORE and CACHE by default, mounts each by label as a filesystem of its own, writes a tree and a manifest in each and unmounts; DragonFly mounts each by label, checks every manifest with its own md5, lists the PFSes from a mounted one and runs its checker; the host's checker runs after each side with its negative control. H2_PFS_DOMAINS names the roots. Between the writes and the unmount the Linux guest takes a snapshot of the first root, mounts the snapshot read-write by its label, changes one file in it and adds another, and reads the live root's file back unchanged; DragonFly then checks the snapshot's manifest as it checks the others, and with the live root and the snapshot mounted side by side reports the changed file reading apart and the added file absent from the live root. H2_PFS_SNAP names the snapshot. H2_PFS_VOLUMES=2 formats the filesystem across two 1 GiB images instead of one 2 GiB image, attaches both to each guest, mounts by the colon-separated device pair on both sides, fills the first root with H2_PFS_FILL MB, 1200 by default, so the writes cross into the second volume, and reads the volume count from volume-list on each side; the host's checker and its control run over the pair. f7 covered PFSes DragonFly made; this covers PFSes this port made, which nothing had mounted on DragonFly before, and a snapshot written into, which the capability declaration's snapshot rows stand on.

script/million-tree.sh is the first of 0.9's criteria that needs only the fleet. The host formats an 8 GiB image; the Linux guest writes H2_TREE_FILES files, a million by default, under H2_TREE_FANOUT directories, each file holding its own path so the tree is data as well as inodes, with the shell's builtins so the rate is the driver's and not fork's; then syncs, unmounts, remounts, drops the page cache and counts, spot-checks two hundred files by content, and prints the create rate, the sync and unmount times, MemAvailable at each step, the module's slab, the cold walk time, lockdep's state and the kernel warnings since the module loaded. DragonFly mounts the same volume, counts it, spot-checks the same two hundred files and runs its checker; the host's checker runs after each side with its negative control. Every reading is a number, which is what 0.9 asks of each of its rows, and the count on each side must be the count written. H2_TREE_MODARGS passes module parameters to the guest's insmod, which is how a control run loads with nofs_scope=0, the shim's lock scope off; on a write refusal the script keeps the reclaim and compaction counters, the buddy lists, the largest slabs and the inode cache size from that moment, which is what told the order-4 folio limit apart from the scope and the mask. H2_TREE_GUESTPRE is a command the guest runs before insmod, for a control that changes the guest rather than the module; the run records the guest's memory size and whether kmemleak is on, read by a write to its debugfs node that is refused once it is off. H2_TREE_WRITERS is the number of shells writing at once, each taking the directories congruent to its number, so 1 is the serial tree and more is 0.9's parallel build; each reports its own count and its own first refusal. H2_TREE_CHURN=1 adds three phases: a tenth of the tree deleted and written again with hammer2 snapshot taken through it, the snapshot mounted read-only after the remount, counted and spot-checked by content on both sides, since what it caught is whatever the churn had written and the reading is that every file in it holds its own path; and the whole tree deleted, timed, with the blocks the volume gave back, which is the store garbage collection reading. DragonFly then expects the live tree empty and the snapshot at the count this side found. The delete pass found the directory link-count defect on its first run. H2_REPEAT=n runs the whole script n times and tallies the outcomes as test-enospc.sh does, keeping each run's log under H2_LOGDIR; a lock cycle is a race, so one clean run after a change to the shim's locking is a run and the tally is the rate. The tally refuses to report if it saw fewer outcomes than it ran.

script/nix-closure.sh is F6's harness, the first of 0.9's criteria: a real Nix closure read at a measured cost beside the same read on squashfs and erofs. The closure is one the host's store already holds, named by its top-level path in H2_CLOSURE; the host writes its paths into a squashfs image and, where mkfs.erofs is found, an erofs image, each cached by the store hash since a closure never changes. The Linux guest mounts both beside an empty HAMMER2 volume, copies the closure in with cp -a so hard links, symlinks and modes travel through the write path, syncs, unmounts and remounts, and takes three cold readings on each filesystem with the page cache dropped between: a metadata walk, a full read, and a hashed read that is also the content check, every file's SHA-256 in path order compared between the copy and its source, with the symlink targets and the hard-linked file count compared the same way. DragonFly mounts the volume, counts it and runs its checker; the host's checker runs after each side with its negative control. Rates are printed and never judged; a run fails on a count or hash that differs, a kernel warning, lockdep turning itself off during the run (a lockdep report prints no cut here, so the warning count alone passed one), a checker verdict or a missing reading, and a guest whose run times out is read through guest-dmesg.sh before it is reset. The guest's whole dmesg is saved beside the log (H2_CLOSURE_DMESG), because the second warning of a run is the one that says why lockdep went off, and H2_NC_GUESTPRE runs a command on the guest before the module loads, for the control that turns kmemleak off. The same cp -a goes into ext4 on a fifth disk and ext4 is read cold with the others (H2_NC_EXT4=0 skips it); the two copy times side by side are the reading the XOP pool decision turns on. H2_NC_JOBS=n deals the store paths round n writers that copy at once, for the parallel-build row (a hard link across two writers' shares arrives as two files, so that count is reported under more than one writer and not judged), and H2_NC_GC=0 skips the collection that otherwise follows the reads: every other store path removed while a reader walks the ones that stay, what stays hashed against its source, and the count DragonFly must then see. The guest kernel needs squashfs and erofs as modules, which the debug guest's did not until F6 asked.

script/latency.sh takes the reading throughput.sh cannot: what one operation costs, rather than what a large sequential stream averages. Every performance number in this tree before it was sequential, which is the half where a copy-on-write filesystem with 64 KiB blocks and a checksum per block is expected to look good; the cost of a random small read and of making one write durable is the other half, and nothing measured it. The measurement is test/hammer2-latency.c, because percentiles over per-operation timings cannot be taken honestly in shell. Three passes: random 4 KiB reads, random 4 KiB writes each followed by fsync, and fsync alone over a batch of eight writes. The two write forms are separate because a bare pwrite returns when the page cache has the bytes, so its latency is the page cache's and would flatter every filesystem equally, while a database or an installer pays the commit. O_DIRECT is not used and cannot be: this port carries no ->direct_IO, so open(O_DIRECT) returns EINVAL.

It runs on HAMMER2 and on ext4 and btrfs in the same guest, because the number is only interesting as a comparison. The failure it is built around is a reading served from the page cache: that reports hundreds of nanoseconds and would otherwise look like an excellent result, so the file is sized above the guest's RAM, the driving side drops the caches before the read pass, and the exerciser asserts its own median is above a microsecond and prints the raw minimum beside every percentile. The host runs the same binary on tmpfs first as the cache control, which must report cache speed and fail that check, so the check is known to be live before the guest's numbers are believed. H2_LAT_MIB may not be set below 256, since a file smaller than a guest's cache cannot produce a media reading whatever else is done.

script/throughput.sh takes the two readings that decide whether the port adds ->readahead and changes its writeback order, the services DragonFly's buffer cache gives HAMMER2 through cluster_readx() and cluster_write(). The Linux guest writes one random file, 512 MiB by default (H2_TP_MIB), from memory to a HAMMER2 volume, to ext4 and to btrfs on two more disks, btrfs being the checksummed copy-on-write filesystem a fair comparison needs, times each write with fsync and two reads at 1 MiB and 64 KiB requests after a remount and a dropped cache, and checks the HAMMER2 copy by hash. Each file is read once unmeasured first: the images are files on the host, and the first read after a write goes to the host's disk while the next hits the host's cache, a difference of ten times that was read as the driver's before the priming read was added. Every timed read is therefore cold in the guest and warm on the host. DragonFly then writes a file of the same size to the same volume and reads both files cold the same way, which is the reference for the read rate. The host reads both layouts from the image with hammer2 show, which prints every data blockref in key order with its media offset, and reports the count of steps that are contiguous, forward or backward, DragonFly's file being the reference the core's allocation comment was written against. A name longer than the inode's inline field lives in a directory entry, so the parser resolves the name to the inode number first. Throughput is printed and never judged; the run fails on a hash mismatch, a kernel warning, a checker verdict, a wrong block count or a missing reading. The parser was checked against a fixture file whose five blocks the tool lists as four contiguous steps, and its first two runs found the tool called without its subcommand and the name looked up in the wrong block, each reading zero blocks for both files and failing on the count.

H2_TP_WRITERS=n splits the same total between n writers running at once, each on its own random source and its own file, and times the admission, the syncfs and the write with fsync the same way, for all three filesystems. It exists because the single-writer reading does not reach what the write path costs when several writers dirty the volume together: the flush is where a per-block wait is paid, and one writer rarely makes the kernel wait. The file names change with the writer count (big.0 upward rather than one big), so the read phase, the post-remount hash check, the DragonFly leg and the layout parser all walk the list the run actually wrote, and each file is hashed against its own source; a run fails on any file read back wrong. Every writer's refusal is counted and printed, and the run fails if any writer was refused, so a full volume cannot read as a fast one.

The same script carries the race trigger for the write path's one invariant that no check code can see, a file folio the core is reading must not be modified. test/hammer2-mmap-exercise.c takes a race argument beside the two modes the fixture and full-volume gates already use: it opens eight files, maps each shared and writable, and stores through the mappings in a loop with no bound while the parent drives writeback over the same ranges with sync_file_range(SYNC_FILE_RANGE_WRITE), which starts writeback and does not wait for it. That is the writer that can reach a folio the core holds, since write(2) cannot (it takes i_rwsem exclusively and a second writer waits) and a write fault reaches the folio through ->page_mkwrite, which takes the folio lock and not i_rwsem. The module counts what changed under the core: the write XOP hashes every folio it reads with XXH64 immediately before and after and increments the read-only parameter folio_changed on any difference. The script runs the trigger before it reads the counter and treats a non-zero count as a failure, with unavailable read as the instrument saying it cannot answer rather than a pass, since the counter is on the debug kernel's module.

All four judge their image by the host's fsck_hammer2 exiting zero, and each of those verdicts now carries its negative control beside it, on the image it judged rather than in a selftest: the same checker is run on a sparse copy with one volume header byte complemented, at an offset inside the first CRC section and clear of the magic and the CRC, and must fail naming the header CRC. A checker that accepts anything, a wrong binary on the path or a copy that landed elsewhere all read as a pass without it. The byte is complemented rather than set, for the reason the fixture gate's manifest control flips rather than sets. The matrix's torn cell is the same control in live form and keeps it. Three pass strings were also found to match over an empty population, zero files checked with zero mismatches and zero unreadable entries out of zero, and the counts are now asserted with the verdicts.

Space a remove does not free, until the scan runs

script/bulkfree.sh measures the one thing a store's garbage collection depends on that no other run had exercised: on HAMMER2 a remove frees nothing. The freemap is rebuilt by a scan, HAMMER2IOC_BULKFREE_SCAN, which the port carried and answered and no run had asked for, and the scan is two passes by design: the first moves a block nothing references from allocated to staged, the next frees what stayed staged, which is what lets a pass run beside writers without a transaction. The Linux guest mounts a fresh 2G volume, writes H2_BF_MB one-megabyte random files, removes them, reads the free count after each step, runs hammer2 bulkfree twice with a sync between and captures the pass statistics the kernel prints for each, then writes the set a second time and removes it. The remove leaving the count where it was is the run's own control: if the count had moved at the remove, the scan would not be what is measured. DragonFly mounts the result and runs one pass, which stages the second set and should free nothing if this side's two passes were complete; its checker and the host's run after each side with the usual negative control. The first run asked for one pass, read nothing freed, and watched DragonFly's pass free the set, which is the second pass doing its job and the harness's misreading, so the script says two passes where it says scan. README.status.md has the readings.

Listing a fixture, and what a clean run does not say

The fixture under /mnt/storage/hammer2-fixtures/tree is five paths: a file, a symlink, a directory, a directory inside it and a file inside that. It is deliberately deeper than one level, because the defect that ->iterate_shared found was invisible at the root: the mount root holds a reference through pmp->iroot that a subdirectory does not, so an unbalanced reference count only shows below it.

Attach the image as a disk rather than through a loop device. The Artix guest's kernel is configured by localmodconfig against its own loaded modules, which trims CONFIG_BLK_DEV_LOOP, and a virtio disk exercises the 7.3 device-open shim as well:

virsh attach-disk artix-s6-kde /mnt/storage/hammer2-fixtures/f1.img vdb \
    --targetbus virtio
mount -t hammer2 -o ro /dev/vdb@TEST /mnt/h2
find /mnt/h2

The attachment is live and not persistent, which virsh dumpxml artix-s6-kde --inactive reports by not naming the image at all. It survives virsh reset, which is why a run of resets can pass without the disk ever being in the domain's stored configuration, and it is gone after a shutdown. Add --persistent to keep it, or expect to attach it again.

A find is the check worth running rather than one ls, since it walks every directory and reaches each one through a lookup on its parent. At e76ad21 it returns all five paths and exits 0, umount and rmmod both return 0, and kmemleak reports nothing after a scan.

Two results in that run were floors and not failures: reading a file returned EINVAL, ->read_folio not being written, and readlink on the symlink returned EINVAL, ->get_link not being written, so ls -l on a directory holding a symlink exited 1 while listing correctly. Both are written since. A symlink's target is file data on this filesystem, so ->get_link is page_get_link() over the same ->read_folio, which is how the DragonFly and NetBSD ports read it too, through hammer2_read_file(). The fixture gate's # link target relpath rows are the check, f1 carrying the one symlink the fixtures hold.

A clean lockdep run on this meant nothing until 0.4.3: every chain lock took its class from one init_rwsem() call site, so lockdep reported recursion at the first mount and cleared debug_locks. Every lock now carries a class and a nesting level, doc/history/verification-record.md records the measurement, and the fixture gate reads debug_locks before its first mount and after its last unmount and fails if it dropped.

Build against mainline, test against the kernel that ships

The port claims to build against an unpatched Linux, and it runs on the kernel the consuming distribution actually ships, which is built with its own configuration and optimization. Those are two claims and the version pin cannot separate them, since it compares VERSION and PATCHLEVEL and a patched tree satisfies it exactly as mainline does. Run both and record both lines:

bash script/test-syntax.sh
KDIR=<the shipping kernel's build tree> bash script/test-syntax.sh

The first takes the unpatched tree, preferring it over a patched one at the same version, and searches /lib/modules/$(uname -r)/build, then anything in H2_KERNEL_TREES, then $HOME/kernels/*, then the store. The second names the shipping kernel deliberately.

The shipping kernel is Saxum's own build, 7.3.0-rc1-saxum, compiled from CachyOS 7.3-rc1 with -march=znver4 and BBR3. A stock distribution kernel from a binary cache is not a substitute for it: it measures a configuration nobody here runs. That kernel's -dev output is in the store and the module builds against it:

make KDIR=/nix/store/<hash>-linux-x86_64-unknown-linux-gnu-7.3-rc1-dev/lib/modules/7.3.0-rc1-saxum/build

It is built with clang 22 and thin LTO, so kbuild has to be told LLVM=1 or gcc rejects four of the kernel's own flags; the module Makefile reads CONFIG_CC_IS_CLANG from the tree's config and adds it, so the line above is enough. The result carries vermagic 7.3.0-rc1-saxum and loads on nothing else. No guest boots that kernel yet; the Saxum server edition, which will, was not built when this was written, so the fixture gate still runs on artix-s6-kde at plain 7.3.0-rc1. The store's 7.3.0-rc1-cachyos figures are a superseded measurement rather than a standing requirement.

Both the header line and the summary line carry the release string and either mainline or patched, read from the tree's EXTRAVERSION: anything left after stripping a leading -rcN was added by whoever built the tree. A kernel built here with its own optimization carries a suffix too and so classifies without being named.

The summary line is the one that gets quoted into a document, which is why it carries the tree rather than the release series alone. A quotation that named only the series was written into README.status.md describing a mainline tree while the run behind it had used the store's patched one.

The same distinction applies to the module. make KDIR=<tree> writes the tree's release into vermagic, and a module built against one kernel is refused by another before any of its code runs, so the runtime test needs a module built against the kernel it will be loaded on. The shipping kernel's module directory is its release string, 7.3.0-rc1-saxum, and installing modules anywhere else leaves them where the running kernel does not search.

Building against that kernel inherits its flags through kbuild, including -march=znver4, so the resulting module requires a Zen 4 host.

Getting a kernel newer than the distribution ships

The kernel of record moves faster than any guest in the fleet, so testing against it means installing a kernel rather than finding one. Two routes are known to work and neither needs a kernel build.

Fedora carries the current stable series and the development series side by side, and its kernel packages have shallow enough dependencies to install across a release. Both of the kernels this port has been loaded on came from there, into a guest that was running Fedora 44:

# dnf --releasever=45 --enablerepo=updates-testing -y \
    install kernel-7.2.3-300.fc45 kernel-devel-7.2.3-300.fc45
# dnf --repofrompath=raw,https://dl.fedoraproject.org/pub/fedora/linux/development/rawhide/Everything/x86_64/os/ \
    --repo=raw --nogpgcheck -y install kernel-<exact-nevr> kernel-devel-<exact-nevr>
# grubby --set-default /boot/vmlinuz-<version> && reboot

Name the exact version. dnf install kernel against a repository that already has some kernel installed reports Nothing to do and exits 0, which reads as success and installs nothing. List first, with list --showduplicates, and install what the listing names.

Installing across a release upgrades what the kernel package depends on. The 7.2.3 install above pulled Fedora 45's glibc, gcc and openssl into a Fedora 44 guest. That is fine for a disposable test guest and is worth knowing before doing it to one that is not.

The second route is the chaotic-cx/nyx nix flake, which packages the CachyOS kernels and had a cached 7.3-rc1 build on 2026-09-03. It is a substitution rather than a build. The CachyOS pacman repository is a different channel with its own cadence and had no 7.3 kernel on the same day, so a reading of one says nothing about the other.

Every COULD-NOT-RUN branch, driven

An error path nobody has driven is an untested branch wearing the costume of a safety net: it reads as defensive prose rather than as code, so it is the last thing anyone thinks to exercise. Every such branch was driven on 2026-08-26, in a scratch copy of the tree, by removing the input each one names. The count is deliberately not written here: it would be a second claim about the same population as the table below, with nothing checking it, and this file has already recorded one instance of a count and a list disagreeing. Read the table.

gate branch how it was driven
test-citations.sh no doc/*.md the directory moved aside
test-provenance.sh no origin clone, so no carry re-verified H2_CLONE_DIR pointed at a path that does not exist
test-history.sh not a repository .git moved aside
test-history.sh no changelog the file moved aside
test-history.sh no current-state document every .md outside CHANGELOG.md and doc/history/ moved aside
test-history.sh a document names a version that is not the newest row's 0.9.25 changed to 0.9.24 in README.status.md
test-history.sh the state-sentence matcher stops matching its pattern altered in a scratch copy of the gate, which the built-in control catches
test-inventory.sh no src/sys/fs/hammer2 moved aside
test-absence.sh the population is empty doc/ and src/ moved aside
test-absence.sh no claim matched, so the pattern has stopped the phrase it matches renamed in a scratch copy of the gate
test-absence.sh a claim naming a symbol that IS defined hammer2_chain_lookup() and hammer2_chain_scan() appended to README.porting.md, the second wrapped across two lines
test-absence.sh a claim naming a ->method that IS wired --selftest, on a fixture tree rather than on this repository
test-absence.sh no ->method claim matched, so that pattern has stopped --selftest
test-inventory.sh no test/ moved aside
test-doc-prose.sh a finding in a root document a British spelling appended to README.md
test-doc-prose.sh the population narrowed back to doc/ the file list filtered in a scratch copy of the gate
test-checkpatch.sh no baseline moved aside
test-checkpatch.sh no checkpatch.pl CHECKPATCH at a path that does not exist
test-checkpatch.sh no perl a PATH assembled from store paths holding none
test-vectors-contract.sh a vector file missing moved aside
test-shim.sh no compiler CC naming one that does not exist
test-syntax.sh no kernel build dir KDIR at a path that does not exist
test-posix.sh no shell realized H2_DASH and H2_BUSYBOX at paths that do not exist
test-doc-prose.sh no vale a PATH holding none, driven 2026-09-02
test-fixtures.sh no image, no guest, no KDIR each driven by pointing the variable at a path that does not exist
test-fixtures.sh a manifest that does not match the media one hash altered in f5.manifest, which failed the image and named it
test-fixtures.sh the comparison itself cannot fail --selftest, and a per-image control on every run
test-fixtures.sh a module built for another kernel the default KDIR, which is the host's, against a guest at the kernel of record, 7.3.0-rc1 when driven
test-fixtures.sh a sanitizer report that turns lockdep off its first run on the KASAN kernel, which failed on lockdep's state and printed nothing, since the report it looked for was lockdep's own; it now prints a KASAN or UBSAN report beside the two lockdep forms
test-enospc.sh a lockdep shutdown that no captured banner attributes nothing; the run counted every shutdown as the cycle it was written for, and now reports the banner and exits 2 when none names it
test-enospc.sh a run against a guest still holding a wedged module its own second run, which reported five failures about a filesystem that had never mounted; the setup steps now report themselves and exit 2
test-enospc.sh a guest listed running that never answered ssh the host load average, twice, with nothing about the guest; the refusal now prints every vCPU's instruction pointer, the disk requests over three seconds and whether the guest agent answers, read from outside before the exit trap shuts the guest down
root-boot.sh a boot that never mounted the volume a second boot against a label the volume does not carry, which stops at the mount and does not reach PID 1
root-boot.sh a checksum comparison that cannot fail its first run, where an unanchored sed made the two sides unequal by construction
test-fixtures.sh an ioctl that answers with the wrong errno the recorded results, which caught EOPNOTSUPP where Linux wants ENOTTY on the first run
test-fixtures.sh a scan that returns success having read nothing the PFS and volume counts, which must be non-zero, and which caught a zeroed capacity
test-fixtures.sh more images than there are target names 26 manifests with an image beside each, which reports the ceiling rather than attaching over the last

All exit 2 and name what was missing. Two defects fell out of driving them: test-shim.sh was the only gate whose message omitted the COULD-NOT-RUN prefix, so anything scanning output rather than status would have missed it; and test-posix.sh had no way to reach its own no-shell branch, because the store lookup finds a shell on any machine that has one, which is why the override exists.

An absent tool must decline, never pass. A probe whose success is cheap to satisfy trivially - an absent binary above all - reports a clean run having examined nothing, and the summary looks identical either way. Measured 2026-08-26 by naming a compiler that does not exist: test-shim.sh, test-syntax.sh and test-vectors-contract.sh each exit 2 and name what was missing, and the vectors gate says which half it still completed. test-checkpatch.sh exits 2 with no perl under a PATH assembled from store paths for coreutils, sed, grep, diff, awk and bash, which contains no perl.

That fourth one was first recorded as UNVERIFIED, on the grounds that emptying PATH breaks the shell and perl cannot be hidden by directory because it shares one with everything else the gate needs. Both facts are true and the conclusion was wrong: a nix machine keeps each tool in its own store path, so a PATH without perl is assembled rather than subtracted. A record that something cannot be checked is a claim about an instrument and decays like any other. The form that survives is UNVERIFIED BY THIS ROUTE with the route named, so that the next reader can see which route was asked and whether another exists.

Exit 2 from any gate here means the instrument could not run: no compiler, no kernel headers, no checkpatch.pl, or a population that came back empty. That is not a verdict on the code, and it should not be recorded as a failure.

test-vectors-contract.sh belongs to neither group. It asserts that this repository still keeps the promises the next section describes: the -DXXH_VECTORS_CONTROL hook, the uppercase hex constants, the printf that writes Castagnoli ... MATCH, and the printf that opens a line with XXH64 and carries want. Where an xxHash is available to link against it also runs the vectors and requires exit 0, then compiles with the control define and requires nonzero, because a status only ever observed as 0 is not tested. That run is also where the wording is read out of the program's own output rather than only out of its source, since stdout is the surface the consumer reads: a printf left in the file but reached under a condition that never holds passes every source check and prints nothing. It carries a negative control that runs every time: a lowercased copy of the file must fail the comparison, since a case-sensitive check and a case-insensitive one look identical while both are passing, and only the case-sensitive one catches the defect that actually happened.

Every pattern in it is anchored on the code rather than on a token. These files describe their own contract in comments, so a check for Castagnoli.*MATCH matched the comment quoting it and stayed green after the printf was deleted. That was found by running the control, not by reading the gate.

A full volume

script/test-enospc.sh fills a 2 GiB volume until the first write fails, calls sync(2), and reads debug_locks on both sides of each step. The circular lock dependency it was written to reproduce is fixed, so is the fault that left the module unremovable after a fill, and so is the held lock freed during the fill that it captured whole once it kept its log. It records what fsync(2) on the last written file and syncfs(2) on the volume return after the fill, with the error text kept, because an exit status of 1 from a missing path and one from a failed sync read the same. It hashes every file as it writes it, off the volume, drops the page cache after the sync and checks every file from the media, printing the intact and damaged counts together so a check that read nothing cannot pass as one that found nothing wrong; that check is what found a fill losing nearly all of itself. H2_ENOSPC_FILES=n caps the fill short of full, which is the control: a volume with room must read back everything it accepted, or the check is what is broken. Free space is read after the sync, because the write path refuses a fill while statfs still shows the space its dirty pages will take; the free count is printed beside it, since statfs subtracts the whole reserve and reads zero for anything under a twentieth of the volume. Around that sync it reads the freemap's two allocation counters, the bytes handed out for data and for everything else, which the module exports as alloc_data_bytes and alloc_meta_bytes, and after it prints the last refusals the write entry put in the debug log, each with the free count and the dirty bytes it judged by, the module's debug prints being on for the run; what the sync took against what the count promised is read from those lines, beside the count of data blocks given new media, the blocks assembled around a folio smaller than the block, the block folios the write entry could not allocate, and the guest's free pages by order. The image of a failed run is kept beside the next run's as enospc.img.failed, since its freemap, read on the host with hammer2 freemap, is what a loss is diagnosed from. After the fill it writes 128 KiB through a shared mapping of a file sized while there was room, with test/hammer2-mmap-exercise.c in its existing mode, and 128 KiB through write(2) into another such file, and fails the run if write(2) is refused and the mapping is not; the refusal at the fault is a SIGBUS, 135 from the guest's shell. The fill ends in 64 KiB pieces so that less than either probe is left above the threshold. The reserve refuses a user with twice the free space it refuses root at, so H2_ENOSPC_USER=1 runs the fill and both probes as nobody under setpriv, H2_ENOSPC_MODARGS is handed to the guest's insmod so io_buf_only=1 puts the whole fill through the DIO layer's block buffer and its bio writes, and, with H2_LOCKDEBUG=1, fail_alloc_after=N has the freemap allocator refuse every allocation past the Nth, which is how the paths behind an allocation failure are run now that the reserve keeps a fill from reaching one, and that build prints every chain still allocated at the unload, with its type, key, references, flags and parent; debug_hpanic=1 has the mount helper call hpanic so the fault path is exercised on demand, and debug_hpanic=2, on the same debug build and writable after loading under /sys/module/hammer2/parameters/, has sync_fs call it on the second call after the knob is set, which is the first sync(2) once the knob is set after an earlier sync, so what that earlier sync wrote is on the media and what comes after the fault is not: script/hpanic-contain.sh is that reading, twenty files synced, the knob set, twenty more written, a hard stop, fsck_hammer2 on the host and the two counts at a remount, run once with the knob and once with H2_KNOB=0 as the control, and on 2026-09-07 it read 20 and 0 with the knob and 20 and 20 without; debug_hpanic=3 fires once from hammer2_base_insert on the next block-table insert, and the same script with H2_KNOB=3 is the acceptance reading for the returning hpanic: the guest pauses after its first sync so the host can copy the image, then the knob, one file, a sync that must fail with EIO, a create that must be refused, the mount options, umount and rmmod, the device hashed through O_DIRECT before and after, the log counted for BUG, oops and WARNING, and the host names each 64 KiB block that differs from its copy; H2_ACCEPT_CONTROL=1 H2_KNOB=0 runs that sequence without the fault. On 2026-09-07 the reading was EIO, refused, ro, 0, 0, identical, fsck_hammer2 clean and no block changed, against the control's 0, 0, rw, five blocks changed; the two runs before it, which read the device changed and the freemap leaf's check bad, are what found the block device's own writeback and the header write, recorded in README.porting.md; and every run reads both thresholds after the sync with one 64 KiB write as the user and one as root: the user is refused in either mode, root is accepted after a user's fill and refused after its own, and a run where root reads the same either way has one threshold where the driver claims two. The readings sit after the sync because the refusal counts dirty pages, and writeback between two writes moves that count by more than the gap between the thresholds; after the sync the pages are allocated blocks and what is free is under the fill's threshold plus one step. They take fresh names: the first run reused the name the refused user write had left behind, and the guest's fs.protected_regular had the VFS refuse root that open, EACCES on a file another user owns in a sticky world-writable directory, which read as the reserve refusing root. Every run prints the source hash it was built from, with a dirty mark. doc/history/verification-record.md carries the account.

It became a gate after ten clean runs on the build that passes it, across three shapes of the instrument, all in the account. It exits 2 without a guest, like test-fixtures.sh, so a push from a machine without the fleet reports it as not run rather than as passed. Its first run as a gate found that its readings named only the faults it was written for, so the warning from the compaction daemon that the file mapping did not implement folio migration sat in 39 of 62 kept logs and every one read as clean. It now counts every cut here line in the capture as a kernel warning, prints the first, and fails on one; the selftest holds the pattern against the line the kernel prints. Every WARN prints that line, whichever subsystem it is from, so a warning the readings do not name is still counted. The count is scoped to the capture after the module loaded, because the capture is the ring from boot and the first run of the count found the guest kernel warning for itself there, a DMA allocation in the USB host controller during boot, untainted and before the module. When the count is not zero the first warning's trace is kept in the log, which the run that found that one had not done.

A debug build logs each PFS as it is freed and each one the teardown syncs, unhashed with %px so the addresses can be compared against the one an oops prints. That is what turned "a chain points at a freed PFS" from a reading of the code into a matched address. It is behind HAMMER2_LOCKDEBUG, so a normal build carries neither the print nor the raw pointer.

Two readings around that fault were wrong in the same way and are worth naming, because the shape recurs: they took the first match in the whole log rather than the one belonging to the oops, so a healthy run reported a faulting instruction and a faulting register out of a page-allocator warning printed long before, and did it while reporting no oops at all. They are scoped to the oops now.

The reproducer carries its own control now, which every gate in this tree had and it did not. sh script/test-enospc.sh --selftest checks each reading in three directions against a line the kernel really prints: the pattern must match that line, must not fire on a healthy run's log, and must appear in the half of the file the run actually uses. That last one matters because the remote block is one quoted string and cannot share a variable with the host half, so the patterns are copies and copies drift. The check searches the file with its own block stripped out, since a search whose pattern sits in its own command line finds itself: the first version did exactly that and would have passed for ever. The control runs at the start of every real run rather than when someone remembers it, because a reading that has quietly stopped matching is the failure this script has actually had, three times.

A batch also refuses to continue if the script changed underneath it. Every iteration re-reads the file, so an edit made while a batch is running silently changes the runs after it, and a half-written file gives them a syntax error that the tally scores as a driver failure. Two runs of a six-run batch were lost that way before the checksum existed.

The capture window used to close before the thing being measured. The run streams /dev/kmsg from before the module loads, and it stopped that stream immediately after the sync(2), which is several steps before the unmount: every message the unmount produced was therefore invisible, and readings taken after it were reading a log that had already ended. The first batch to ask what the module still held reported, on four runs whose rmmod had succeeded, that ->kill_sb never ran, which cannot be true and is what exposed it. Reading the log never required stopping it, so the capture now runs on until the unmount is over, and it runs unbuffered, because cat block-buffers to a file and a line printed during the unmount can sit in that buffer while a grep for it reads as a line never printed. Which form ran is reported per run.

What the module still holds is now printed by the driver rather than inferred. hammer2_assert_clean() reads the four allocation counters on the unload path, which a module that will not unload never reaches, so the one failure they would explain is the one they could not report. ->kill_sb prints them too, unconditionally: that function runs only when the superblock is actually destroyed, so a failing run with no such line says the superblock is pinned rather than that the counters were clean. Printing only a nonzero count would have made those two cases produce the same silence, which is what the first version of it did.

Both of those are intermittent, which makes a single run the wrong instrument: it answers about itself and nothing else. H2_REPEAT=n runs the whole thing n times and tallies pass, fail and could-not-run, printing per run the readings that tell the two apart, so an intermittent fault is a rate rather than an anecdote. Each iteration resets the guest first, because a guest left wedged by one run costs the next one as well: three runs were spent that way before this existed. The tally asserts it counted as many outcomes as it ran, so a loop that lost one cannot report a clean rate.

One reading it no longer takes on trust is the unmount. For several runs the script judged the unmount by the exit status of the umount process, which is killed on these runs by something outside the script, and read that as an unmount that never finished. It asks whether the filesystem went away instead, and on the runs where it has asked, the filesystem went away every time.

Every other write measurement in this tree was taken on a volume with room in it, which is why this went unfound through the whole write path, the crash matrix and the round trip. A filesystem's behavior when it runs out of space is its own surface.

Reading the report is the part the script had to be rebuilt for. It streams /dev/kmsg to a file for the whole run and prints what it derives from the log before it attempts the unmount, because the ring wraps before a fill ends and a bad unmount is where the evidence is lost. Five reproductions before that produced no readable report. The capture is doc/enospc-lockdep.txt, which records a defect since fixed and says so in its own first line, because a file holding a lockdep report reads as current behavior to anyone who opens it.

Two things it now refuses to assume. Reports are counted by banner occurrences rather than by lines matching DEADLOCK or circular, which counted one report as three. And debug_locks reading 0 is not taken for the cycle on its own, since an unlock imbalance, a held lock freed and three lockdep resource ceilings read the same there: the run reports which banner named the shutdown, and a shutdown it cannot attribute is COULD-NOT-RUN rather than a confirmation.

The instrument the defect needed is in the tree rather than rebuilt each time. HAMMER2_LOCKDEBUG=1 on the make line compiles hammer2_dbg_held_chains() live, which walks the calling task's held locks, resolves each chain lock back to its chain and prints the type, key, class, recursion depth and lockcnt. Unset, every call compiles away and the module carries no symbol for it, which is checked by building both ways. H2_LOCKDEBUG=1 bash script/test-enospc.sh builds with it and summarizes what it printed.

Two designs it deliberately is not, both of which were built and both of which read as evidence while carrying none: one recording the acquire on every lock, which names where a chain was last locked rather than where the live one came from, and one clearing that record on unlock, which erases it for a counted lock that is still held. This one reads CONFIG_LOCKDEP's own held-lock records, so it cannot disagree with the report it is being used to explain, and it prints nothing when nothing is held: a quiet run is the absence of a finding rather than a passing one.

Its own first two runs failed on the script rather than on the driver, and the second is the interesting one. A guest left wedged by a previous run of this same script cannot load the module, and every check after that reported a failure describing a filesystem that had never been mounted. The setup steps now report themselves and the run exits 2, so the state this script leaves behind is told apart from the defect it exists to find.

The full volume on DragonFly

script/dfly-enospc.sh runs the same fill on the DragonFly guest, against the tree the patches under doc/upstream/ are written for, so a defect this port finds in carried code can be shown on the code's own kernel before it is filed. It needs a kernel built from DragonFly source with two sysctls added to hammer2_freemap.c, vfs.hammer2.fail_alloc_after and vfs.hammer2.alloc_count, which doc/upstream/README-provenance.md describes; on the release kernel the allocator never refuses and the run says so. The refusal is lifted after H2_KNOB_LIFT seconds, 120 by default, because a refusal left in place on DragonFly wedges the syncer on the buffers it cannot write and the fill never reaches the unmount. The reading is the writers' last error, the file count before and after a remount, the kernel messages the fill produced, and fsck; its two runs so far are recorded in the provenance document beside the patch they verified.

The volume as a root filesystem, from the tree

script/root-boot.sh is the boot as a script. It formats a 4 GiB image with newfs_hammer2, builds the module, and uses the guest to put a static init on the volume and to run two checks the guest is needed for: a binary copied onto the volume and executed from it, and a 128 KiB file, two of this port's 64 KiB folios, written entirely through a shared writable mapping and msynced, then compared on media after drop_caches and again after a fresh mount. It then boots qemu-system-x86_64 directly, with an initramfs holding the module and that same init, which reads root= from the kernel command line, mounts the volume, moves the mount over / and executes /sbin/init from it. Six lines of the boot's own transcript are required, from the module load to PID 1 writing and reading back and remounting read-only.

The control is a second boot against a label the volume does not carry, which must fail at the mount and must not reach PID 1. It is placed there because every claim the first boot makes rests on that mount, and because the boot narrates itself: a check that grepped only for the last line would pass against a kernel that mounted nothing.

Its first run failed on a defect in the script rather than in the driver. sed -n 's/^mapped sum //p' also matches the mapped sum after remount line, so the first checksum came back as two lines and could never equal the second. Both patterns are anchored on the whole line now. This is the shape rule 24 names: a matcher that reads more than it means, in a comparison that can only fail.

It exits 2 without qemu-system-x86_64, /dev/kvm, cpio, newfs_hammer2, a kernel image or the fixture directory, and it names which. KDIR supplies the kernel image by default, so it wants the kernel of record and not the host's headers; H2_BZIMAGE overrides.

What it does not show is a distribution. Nothing on the volume but the init has ever run: no service manager, no package manager, no shared-library loader.

What was falsified, and when

A fixture is not shown to work by its own green run, and reading a fixture you just wrote is the least reliable way to answer whether it tests anything. So each of these was run against the defect it exists for rather than merely run. This is a LIST OF WHAT WAS DONE, not a claim that everything has been done: an "every check has been falsified" sentence is a claim about a population that grows, so it would be false the moment a check is added rather than eventually, and nothing would notice. A new check comes with the run that showed it failing, added here.

falsified on check how
2026-08-26 xxh64: -DXXH_VECTORS_CONTROL hook present the #ifdef replaced with #if 0, comment left in place
2026-08-26 xxh64: constants uppercase the constant lowercased
2026-08-26 xxh64: HAMMER2 seed uppercase that constant lowercased separately, because sharing a code path with something falsified is not being falsified
2026-08-26 crc32c: a printf writes 'Castagnoli' then 'MATCH' the printf reworded, comment left in place
2026-08-26 xxh64: a printf opens with XXH64 and carries 'want' the prefix renamed to XXHASH64 in the source
2026-08-26 the wording negative control the printf line stripped from a copy, which must stop the pattern matching
2026-08-26 the output opens lines with XXH64 and carries 'want' the printf guarded by if (0), so the source check passes and stdout is empty. Falsified separately from the source check for that reason: under a rename the gate exits on the source failure and never reaches this one
2026-08-26 the vectors negative control the shared matches() made case-insensitive
2026-08-26 an overridden run says so in its summary the override warning deleted, and again partially
2026-08-26 a UAPI-shaped tree claiming 7.2 is COULD-NOT-RUN a version.h fallback added, the improvement a later maintainer plausibly writes
2026-08-26 the checkpatch selftest the sha256 mismatch text deleted
2026-08-26 posix: each shell rejects a probe every probe body replaced with echo, which reports the shell inert
2026-08-26 posix: a plain script is accepted the plain script made unparseable, which reports a clean result unreachable
2026-08-26 test-checkpatch.sh declines without perl run under a PATH assembled from store paths holding no perl

A DATE REPAIRS AGING AND NEVER FALSIFICATION, and the two get treated as one thing. A dated completeness claim still says something false the moment the population grows: the date stays true while the sentence stops being, so nothing about it looks stale. That is why this is a table of what was done rather than a dated sentence about everything.

No gate derives this table, and that is deliberate. The population is checks inside scripts. Counting invocations by grep is the hand-maintained-list defect one level up wearing a regex, and running every gate from inside another gate to read its printed count couples the gates for a claim that is documentation rather than behavior. Written here rather than settled in conversation, because a decision not to build something is invisible to the next reader unless it is in the tree they grep. The specification repository reached the same answer about its own typed list for a better reason: extensionless documents have no extension to enumerate on.

The syntax selftest's second direction, that an unoverridden run carries no override warning, was left out of that table while it was vacuous: with no kernel of record installed, an unoverridden run is COULD-NOT-RUN and carries no warning either way. It stopped being vacuous on 2026-08-26, when such a run began really compiling.

Run from outside this tree, by a gate in another repository

A test file nothing runs reads exactly like a test file that passes, so these two were written up on 2026-08-26 as staged and unrun. That was wrong within the hour, and wrongly reassuring in the direction that costs most: no gate HERE runs them, and Saxum's scripts/test-hammer2-checkalg.sh compiles both, against the FreeBSD port's vendored xxhash and its icrc32.c, reaching this tree through LINUX_HAMMER2. The sweep that concluded "run by nothing" searched this repository only, which is the whole of the mistake: a consumer one directory over answers a question no local grep can.

What that makes them. Four things are an interface with a consumer that cannot be seen from here: the exit status of each file, the wording of the Castagnoli ... MATCH line, the XXH64 prefix and want that the consumer counts its vectors by, and the uppercase hex of the xxh64 constants. The rewrite that fixed their logic lowercased one constant, Saxum's negative control seds on that literal, and its gate spent an hour reporting correctly that it was comparing nothing. -DXXH_VECTORS_CONTROL exists so that control never has to depend on this file's text again.

No gate here reaches into another repository to check any of that, and none should: the port stands on its own. The contract is written down instead, in the table below and in each file's header, and script/test-inventory.sh fails on any file under test/ that neither a gate names nor this table lists.

file run locally when state on 2026-08-26
test/crc32c-vectors.c iscsi_crc32(), which arrives with the check algorithms in 0.2 the exit status accepted either CRC-32C or CRC-32 IEEE, so the one question it exists to ask went unanswered while it reported success. Now it accepts Castagnoli only, and names IEEE when it sees it
test/xxh64-vectors.c an xxhash.h in this tree, same import two of three cases asserted nothing, and the seeded case used xxHash's golden-ratio prime where HAMMER2 seeds with 0x4d617474446c6c6e. Four vectors now, all four measured against xxhsum 0.8.3 and libxxhash 0.8.3, and compiled and run green against the system xxHash on 2026-08-26
test/getdents-resume.c a directory can be listed, which it can, on a guest with a mount ->iterate_shared resumes across calls, which one ls cannot exercise: a 32 KiB buffer takes a small directory in a single call, so the branch that stops mid-directory never runs. It reads with a 64-byte buffer, one or two entries at a time, and fails on a runaway rather than hanging. Against the five-path fixture on 2026-09-04: the root gives 5 entries over 3 calls and the subdirectory 3 over 2, each name exactly once. Run by test-fixtures.sh on the first manifest that verifies; the count is not the exerciser's, which prints its names and its two totals and has no check protocol, so the gate asserts the two properties itself: the calls exceeded one, so the mid-directory stop was reached at all, and no name repeated, which is the signature of a read restarting from offset zero instead of resuming

The first two are not wired into a gate HERE, because there is nothing in this tree to link either against; Saxum links them against the BSD tree instead. They get a local gate the day 0.2 imports the algorithms, and the inventory gate is what remembers to ask. The third needs a booted guest holding a mount, and it is wired into test-fixtures.sh on the first manifest that verifies: it ran by hand until that gate existed, which is how its row came to carry a reading from an invocation nothing recorded. Until then, changing either file's output shape or exit status breaks a gate in a repository this one does not reference.

The one that drives a fix in src/ without compiling src/

Saxum's scripts/test-hammer2-reptrack.sh compiles tests/storage/hammer2/reptrack_harness.c, which is not a test file of this tree and not a copy of one. It reimplements hammer2_chain_repparent() and hammer2_chain_repchange() as two threads with hammer2_spin_ex as an ownerless non-recursive mutex, DragonFly's spin_lock, and it carries both controls: the stock protocol must deadlock, the fixed protocol must complete, and both wrong means the harness is wrong. It is the only instrument on this disk that reaches the unreleased reptrack->spin found in H1_READING_1_SPIN_AUDIT.md, and it is green on 2026-09-27, 2 checks 0 failed.

It does NOT reference LINUX_HAMMER2, unlike the three Saxum gates that do, so a green run is evidence about the protocol and not about this tree's file. The link to src/sys/fs/hammer2/hammer2_chain.c is that the harness body and the carried function agree hunk for hunk, which is a reading taken on a date and not a property anything enforces. It is recorded here so the next sweep for what runs this tree's code does not conclude "run by nothing" from a search of this repository alone, which is the mistake this section already documents twice.

The reference control every kernel-facing test declares

This is the readiness audit's item 5, and the reason it generalizes is already in this repository's history: test-seek.sh shipped a check that asserted the implementation rather than the contract, and the only thing that caught it was running the same question against tmpfs and btrfs on the same machine. A test written from the implementation passes against the implementation. A test with a reference beside it cannot.

So every file under test/ that asserts a kernel-facing contract declares the reference it is checked against, or declares that it has none and why. The declaration is a line in the table below, which test-inventory.sh reads: a file under test/ with no row is a finding, and so is a row naming a file that does not exist. Nothing is inferred from the file's text, because a lexical rule was tried and got this wrong in both directions in one pass, calling hammer2-dedup.c pure when it measures a kernel behavior and missing hammer2-dedup.c in the same breath.

test file reference control what it asserts
test/hammer2-seek.c tmpfs and btrfs on the same guest SEEK_DATA/SEEK_HOLE answer the offsets the media holds. The control is what found the original defect: both references disagreed with the port.
test/hammer2-dedup.c tmpfs and btrfs on the same guest a duplicate block is not charged twice. Both controls charge full price, which is what makes the reading specific to this port.
test/hammer2-fallocate.c tmpfs and btrfs on the same guest a punched or zeroed range reads back as zeros with the bytes outside it intact, and the punch leaves a hole. Its ranges come from the filesystem's own block size, so the same binary is a real check on a 4 KiB control and on this port's 64 KiB block. It also punches a range spanning several whole blocks and asserts the boundary, which is the range a walk that visits each folio more than once lives in. Measured 2026-10-02, after that check was added: 16 checks and 0 failures on this port and 14 with 1 refusal on tmpfs, which accepts PUNCH_HOLE and refuses ZERO_RANGE. The 2026-09-28 reading was 12 on this port and on btrfs, 10 and 1 on tmpfs. btrfs is not re-measured here: the guest has no loop device, and a run pointed at an unmounted directory reads the guest's own root filesystem and reports a clean control beside this port's answer.
test/hammer2-fiemap.c none, and it says why: the control is the file FIEMAP describes this filesystem's own layout, so a reference filesystem would compare two layouts rather than check an answer. The file's shape is built by the test and known before the call, and the map is compared against it: a hole opened at block 1 must come back as a gap, and the written blocks at blocks 0 and 2 must come back as data with a nonzero physical address. tmpfs is not the control because it refuses FIEMAP outright (measured 2026-09-29: EOPNOTSUPP). The consumer that proves the map useful is filefrag -v, which reads the same two extents at the same blocks. The same file also covers ->freeze_fs/->unfreeze_fs on the same terms, with every blocking call in a child so the thaw is always reached, which is the structure the first version of the test lacked when it deadlocked itself and got the vop wrongly withdrawn. It is run by test-enospc.sh on the volume with room, before the fill, because it writes the file whose hole it asks about; every write it makes is unlinked before the fill measures what it did.
test/hammer2-fh.c none, and it says why: the control is the handle it forges a file handle names the object it was taken from. A reference filesystem would compare its own handle format, not this port's answer, and the operation has no wrong-answer-that-looks-right mode the way the seeks do: a decoder either finds the inode the handle names or refuses. So the force is applied inside the file, by the checks that make a wrong object detectable: two files with different contents whose handles must open two different inodes, and a handle with a bit flipped in its inode number that must be REFUSED rather than resolved to a neighbor. A decoder that ignored the inode number would pass "the handle reopens the file" and fail these. What it does not cover, and a reader would otherwise assume it did: the acceptance callback. open_by_handle_at(2) passes vfs_dentry_acceptable, which returns 1 without looking when ctx->flags is zero, and that is zero for every handle opened without O_DIRECTORY (fs/fhandle.c), so the regular-file checks prove the decode returns the object named and not that the kernel would have rejected a wrong one. The directory check does exercise the callback, O_DIRECTORY being what sets the flags. nfsd passes nfsd_acceptable and walks parents, which is the only caller that reaches ->fh_to_parent, so that member and the acceptance path are the two things this file cannot see; doc/README.capabilities.md names the same gap. It is run by test-enospc.sh on the volume with room, before the fill, because it writes, unlinks and renames its own files; it makes one directory and removes it, so the fill measures what it did.
test/hammer2-latency.c ext4 and btrfs on the same guest per-operation latency is a property of the port, not of the machine.
test/hammer2-mmap-exercise.c tmpfs on the same guest a shared writable mapping reaches the media.
test/hammer2-ioctl-exercise.c tmpfs on the same guest the ioctls refuse what they should, and the refusals are the filesystem's and not the VFS's.
test/hammer2-header.c none, and it says why the on-disk header layout, which is a format fact with no other filesystem to compare against. DragonFly is the reference and it is the cross-side runs in the fleet.
test/getdents-resume.c none, and it says why ->iterate_shared resumes across calls, which is a VFS contract and not a per-filesystem choice; test-fixtures.sh runs it on the first manifest that verifies, and asserts resumption and uniqueness itself, since the exerciser reports totals rather than checks.
test/crc32c-vectors.c reference vectors, not a filesystem the CRC-32C digest against the published vector.
test/xxh64-vectors.c reference vectors, not a filesystem xxHash64 against xxhsum's own output.
test/syntax-check.c the compiler is the reference the carried headers define no name the kernel defines.
test/contract/ctl-sparse-user.c sparse is the reference, and it must refuse this the negative control for the syntax gate's sparse pass: a kernel pointer handed to copy_to_user() has to be refused, or that pass is blind to its own class.
test/rootfs/h2root-init.c the same file in both roles, which is the control PID 1 runs off a HAMMER2 volume: it writes, syncs, reads back and powers off. The initramfs role and the on-volume role are the same binary, so a failure in either is not a difference between two builds.

A row that says "none" is a decision recorded, not a gap: what remains open is whether the decision is right, which the row states. What the gate enforces is that the decision was made at all.

What the real test will be

A volume created by DragonFly's newfs_hammer2, mounted here, compared file by file; then the reverse. HAMMER2 writes an XXH64 digest into every blockref and every implementation verifies it, so a subtly wrong port produces volumes that read as corrupt on DragonFly rather than as buggy. The cross-implementation round trip is the only test that catches that, and no amount of self-consistency substitutes for it.

The gates that will watch a running system

Every gate here reads a file that is not changing while it reads. The gates 0.3 and after do not: a build-and-load script watches a guest boot, and the crash matrix watches a filesystem being cut off mid-write. Those gates can fail in a way none of the current ones can, by sampling before the phase they are about begins. A guest that has not yet loaded the module reads clean for the same reason a healthy one does, and the sample carries no timestamp saying which it was, so nothing in the output distinguishes them afterwards.

It does not feel like guessing, because the activity is measuring, and it is measuring, of the wrong phase. The reading that survives review is the one that happened to be right anyway, which is invisible to anyone checking answers rather than method.

So a gate that samples a running system carries a positive control: something that MUST move during the phase under test, sampled in the same pass as the measurement. A module load that has begun has a dmesg line; a write that has begun has a growing device. An unchanging number with no such control is unproven rather than reassuring, and a gate resting on one is asserting its own conclusion.

This is the same requirement as the negative controls above, one axis over. Those ask whether the instrument can fail at all. This asks whether it was looking while there was anything to see.