Repository navigation
Upstream unit tests in the image build; patches 0044–0054, cbt history windows, tier read errors - #27
Merged
Merged
Conversation
…eries patches The image configures --disable-unit-tests, and no patch had ever run the upstream suites of the code it changes. A first attempt ran them in a separate job that built SPDK a second time. This one runs them on the image's own tree instead. - Dockerfile, unit-tests stage: FROM the builder, it compiles the owed suites against the libraries the builder made, and runs them. The collector (release image) and the debug image both derive from it. The suites run in the debug build, on each arch: asserts are on there, as upstream runs them. The release build passes through at no cost. A failing suite fails its debug build, and with it ci-gate and the release publish, which needs every build of the matrix. - images/spdk/unit-tests.map says which leaf suites of test/unit/ a patched path owes: longest prefix wins, and "-" says upstream has no suite. A patched .c under lib/, module/ or app/ that matches no line fails the stage. A patched file under test/unit/ adds its own suite. - images/spdk/unit-tests.sh selects, builds (lib/ut, then each leaf directory) and runs them. It builds from the leaf because parent Makefiles skip some suites with a mere warning (blob.c needs CUnit 2.1-3, which Fedora 43 ships). A missing binary fails, like a failing suite. The script is plain bash 3, so its selection can be checked anywhere; today it selects 18 suites. - Layer order: the build dependencies (pip pins, pkgdep.sh, liburing and libaio) now come before the modules and patches. They depend on the upstream tree alone, so a patch change no longer rebuilds them. patches/README.md, "Unit tests", documents it. The first run is expected red. The separate job's runs already showed that vbdev_lvol_ut no longer compiles against the series: 0005 makes vbdev_lvol.c include vbdev_tier.h, and the unit test's Makefile has no path to it. What the suites find is fixed in the patches that caused it, in the commits that follow.
… stage reports every broken suite in one pass vbdev_lvol_ut includes vbdev_lvol.c, which the series changed twice without the unit test following: - 0009 completes an ENOSPC with spdk_bdev_io_complete_nvme_status. The unit test gets a stub. - 0041 includes vbdev_tier.h and registers a per-band usage provider on a bdev_tier composite. The unit test's Makefile gets the tier module on its include path. Its stubs give no composite, so no provider is registered: vbdev_tier_get_by_name returns NULL, and set/clear usage provider and spdk_bs_count_allocated_clusters_in_lba_range are stubs. Both edits are folded into the patch that made them necessary (fixup commits autosquashed in a series worktree, then scripts/patches.sh regen). Only these two patch files change, and the 42 still apply. unit-tests.sh now builds every owed suite even after one fails, and runs the ones that built. A single pass reports every broken suite, instead of stopping at the first one in the list.
The previous run failed three suites. Each fix is in the patch that broke the suite. - bdev_raid_ut did not link. Ten patches call functions the suite neither includes nor stubs, and the calls pulled objects of libspdk_json and libspdk_jsonrpc whose other definitions clash with the suite's own. Each patch now stubs what it calls: 0007 (heat), 0009 (NVMe status completion), 0013 (RPC audit, two json calls), 0015 (outcome registry), 0017 (CBT epoch query, a json call), 0018 (verifying outcome), 0019 (CBT auto epoch), 0021 (envelopes), 0025 (superblock clear), 0033 (a json call). - bdev_raid_ut then failed on 0014: a raid is refused without an incarnation, and the suite's delete requests carried an uninitialised expected_incarnation. The suite now creates under an incarnation and deletes unchecked. Two tests cover 0014: no creation without an identity, and a delete that expects another incarnation is refused. - raid1_ut did not link: 0009 reads the NVMe status of a failed member write. The stub reports a device fault. A new test covers 0009's thin-exhausted member, which is kept and marks the raid I/O. - rpc_ut failed on 0010: the listen chmods a socket file that the suite's stubbed bind never creates. The suite records the chmod and the umask the bind runs under, and checks both (0600, 077). A new test checks that a failed chmod refuses the listen. All 18 suites pass in a local debug build of the unit-tests stage (arm64).
…tests only) Patch 0044 starts as tests only, and the unit-tests stage is expected to fail on blob_ut: - blob_thin_prov_new_cluster_reads_zeroes: every free cluster is filled with 0xAA, as a device whose unmap does not zero leaves them. A write of one io_unit goes into a new cluster, and the rest of the cluster must read zeroes. Upstream returns the 0xAA. - blob_thin_prov_full_cluster_write_is_not_cleared_first: a write that covers the whole cluster costs the payload and the metadata page(s), and nothing more. It passes before the fix too: it guards the fix's cost. - blob_thin_prov_rw, blob_thin_prov_write_count_io and blob_thin_prov_rle count the cleared cluster in their write accounting. Locally (debug, arm64): blob_ut fails 4 tests in each of its 5 suites. The fix follows in the next commit.
A cluster that a write allocates in a blob backed by zeroes (a thin blob with no parent, or a clone over a range its snapshot never allocated) was inserted as the device left it. The write filled its own range, and the rest of the cluster returned what was stored there before. The free path counts on unmap to leave zeroes, which a SATA SSD does not guarantee and a bdev without unmap support never provides. When the new cluster is not copied from a parent, bs_allocate_and_copy_cluster now clears it with write_zeroes before inserting it. A write (or write_zeroes) that covers the whole cluster skips the clear, so a rebuild or a copy does not write twice. A clear that fails returns the cluster to the pool. The inflate path goes through the same branch. Locally (debug, arm64): blob_ut passes, 495 tests.
…ndows they need (0045, tests only) Three cases for bdev_raid_ut, red on the series as it stands: - two seed ranges inside one window: 12 blocks to copy under one lock; the window is copied whole (128 blocks); - one range starting mid-window across three 4-block windows: 9 blocks under three locks, the first at the range; the windows are copied whole (12); - one range at the end of the raid: 4 blocks under one lock, at the range; the clean windows before it are each locked and skipped. The stubbed range quiesce counts the windows the copy phase locks.
…dows they need (0045) A window that a seed range touched was copied whole, and every clean window was locked and skipped in turn. Measured on a three-node cluster under a random-write load: 5 GiB re-copied for 395 MiB of dirty chunks (4794 ranges), 110 s. Copying only the ranges, one window per range part, still took 100 s: the cost is the window cycle (quiesce, copy, channel update, unquiesce), not the bytes. - A seeded process jumps over what its ranges leave clean, with no window. Only a copy needs a window's quiesce (a write landing between the read of the source and the write of the target), and the seeded target is a member of every channel for writes: the routing by process offset sends a write to it on either side. - A window copies every range part it holds, under its one lock, and ends where everything before it is copied or clean. Out of requests, it ends at the first block not submitted and the next window resumes there. The verify phase and the copied-bytes accounting are unchanged. bdev_raid_ut: the three cases of the previous commit pass, and every raid suite stays green (bdev_raid 20, bdev_raid_sb 9, concat 3, raid0 8, raid1 5, raid5f 8).
A JSON-RPC request carries ~200 ranges (SPDK_JSONRPC_MAX_VALUES = 1024
parsed values, a 32 KiB receive buffer), so a caller had to fold a larger
delta to fit, and folding a scattered delta recopies every clean block in
between: measured on a three-node cluster under random writes, 4485 dirty
ranges (368 MiB) folded to 128 recopied 3.2 GiB of clean blocks, and the
seeded rebuild of a 5 GiB volume took as long as a full one (95-110 s).
bdev_raid_start_seeded_rebuild_from_epoch {name, base_bdev, cbt_bdev,
epoch_id, expected_incarnation, rebuild_token} seeds the same rebuild with
the dirty ranges of a frozen epoch, read in-process: the cbt module answers
vbdev_cbt_query_epoch_ranges (the ranges bdev_cbt_epoch_get_dirty_ranges
walks, up to the caller's cap; more is -E2BIG, since a truncated delta is a
wrong one, and the caller rebuilds in full). The two RPCs share their checks
(raid, incarnation, running process, write_only member, range bounds, token)
and their audit.
bdev_raid_ut fakes the query: 64 one-block ranges a block apart copy their
64 blocks, and a delta the query refuses starts nothing. Every raid suite
stays green (bdev_raid 22, bdev_raid_sb 9, concat 3, raid0 8, raid1 5,
raid5f 8).
… (0047, tests only) Two cases, red on the series as it stands: - raid1_ut: a 64-block window offered as the generic layer offers it (the rest of the window to the next request) is copied by four requests of at most the part, every part read before any completes, the reads tiling the window and spread evenly over the members in sync, and each part's read completing writes that part to the target. raid1 takes the window whole in one request, and on three members the second one serves no read; - bdev_raid_ut: raid_bdev_process_request_max_part offers parts that are whole write units and all fit RAID_BDEV_PROCESS_MAX_QD requests, smaller than any window with room to split. It offers the window whole. raid1_ut now inits a process request's raid_io for real (its target names the raid) and records every I/O the module submits.
…(0047) raid_bdev_process_request_max_part offers a window to the process's requests in RAID_BDEV_PROCESS_MAX_QD parts, whole write units, and raid1 copies one part per request. A window's parts are now in flight together: each part's write starts as its own read completes, and the reads spread over the members in sync by their outstanding blocks. Copied as one request, a window was read whole from one member, then written whole: read and write never overlapped, and the first member served every read. Measured on a 3-node bench, a full rebuild of a 5 GiB mirror3 under a verified write load ran at 36 MiB/s for 141 s, every byte read across the link from a remote member while the nexus's local member served none. raid5f, bound to its stripe, does not call it; the seeded copy already takes what the module returns and offers the rest to the next request.
…s only) raid1_ut, red on the series as it stands: four parts of a rebuild window, two read back holding data and two holding nothing but zeroes. With a target that takes write-zeroes, the data parts are written as data and the zero parts as write-zeroes; with a target that cannot, every part is written as data. The zero parts go out as data writes. The window-in-parts test (0047) now gives its requests real buffers: the copy reads a part back before choosing how to write it. The suite records the write-zeroes it is sent, and the write-zeroes support of the members is a knob of the suite.
A part the source read back as nothing but zeroes goes to the target as write_zeroes (raid_bdev_write_zeroes_blocks, the data offset applied like the other wrappers): the same blocks, read back the same, and no payload on the wire. A full rebuild of a sparse volume no longer ships its never-written clusters across the link — measured on a 3-node bench, about 3 of the 5 GiB of a mirror3 full rebuild, all of it zero, all of it through the nexus node's egress, which 0047 had made the bound. Not when the target cannot take write-zeroes (the bdev layer would then send a zero buffer anyway), nor when the blocks carry metadata. The check reads the part already in memory, and a data part stops at its first nonzero word.
…ed (0049, tests only) blob_ut, red on the series as it stands: write-zeroes over half of a cluster a thin blob has not allocated allocates the cluster and writes to the device; it is to allocate nothing, write nothing, and read back zeroes. The same test pins what stays: write-zeroes over an allocated cluster holding data zeroes what it covers, and on a clone, whose unallocated cluster reads its parent's data, it allocates the cluster and zeroes it.
…d (0049) On a blob backed by zeroes, an unallocated cluster already reads zeroes: write-zeroes over it now completes at once, with no cluster allocated and nothing written. Before, it allocated the cluster (cleared, since 0044) and zeroed it. A full rebuild sends every never-written part of its source as write-zeroes (0048), so a thin volume rebuilt in full took its whole size on the target, and the zeroes it carried no longer on the wire still went to the target's disk: measured on a 3-node bench, the SATA disk of the rebuilt replica wrote 94 MiB/s through a 5 GiB rebuild with about 3 GiB never written. An allocated cluster is still zeroed. A clone or any blob with another back device reads its unallocated clusters from it, and keeps the allocation.
…er (0050, tests only) subsystem_ut, red on the series as it stands. A Write Exclusive - All Registrants reservation acquired by A, then A preempted by B: the reservation stays under A's key, which no registrant holds, and the state persisted for it fails to restore (-EINVAL). The same once B unregisters after C registered. The restore test now refuses a missing key for a single holder only, and takes it for a reservation held by all registrants, under the key of its first registrant.
…r (0050) Every registrant holds a Write Exclusive or Exclusive Access - All Registrants reservation, so when the registrant that acquired it leaves, preempted by another or unregistering, the reservation stays. It stayed under the key of the registrant that left: the state persisted for it named a reservation key no registrant held, and the restore refused it (-EINVAL), failing the add of the namespace. After a writer handover, the next restart of a target left its namespace out, and the target unexported. The reservation now stands under the key of the registrant that holds it next. The restore takes a state persisted before this, under the key of its first registrant; a single holder's key that no registrant holds is still refused.
…sts only) bdev_nvme_ut, red on the series as it stands. The suite now compiles nvme_rpc.c in, so the I/O passthrough can be driven against a real controller channel. With the channel's I/O qpair cleared, as bdev_nvme clears it from a disconnect until the reconnect, the command is still handed to the NVMe library (rc 0, the channel kept): in a real process that NULL qpair is dereferenced. The test expects -ENXIO with the channel given back, and the command to go through once connected again.
bdev_nvme_send_cmd with an I/O command took the controller channel's qpair (bdev_nvme_get_io_qpair) without a check and handed it to spdk_nvme_ctrlr_cmd_io_raw_with_md. bdev_nvme frees that qpair from a disconnect until the reconnect (bdev_nvme_disconnected_qpair_cb), so a reservation (reservations are I/O commands) sent to a member whose path had just gone down dereferenced NULL and killed the process, every raid it carried with it. The command is now refused with -ENXIO and the channel given back; a command for which no channel can be had is refused with -ENOMEM instead of reaching the assert-only accessor.
… tests only) bdev_raid_ut, red on the series as it stands. A current member that leaves a raid1, removed or failed, should open one epoch, on the cbt stacked on the raid and named after the member's UUID. 0019 asks each survivor's own name instead (31 calls here, none to the raid), a bdev no cbt sits on, so no epoch ever opens and a member lost without notice is always rebuilt in full. The suite also pins five ejections that must open none: a member being rebuilt in full (its superblock slot not configured), one behind the generation its peers carry, a write-only joiner, a raid without a superblock, and a failed member an earlier removal failed to take.
0019 opened the epoch at a member's ejection under each survivor's own name. The cbt sits on the raid, not on its members: every call answered -ENODEV, no epoch ever opened, and a member lost without notice was always rebuilt in full. The raid now asks once, for the raid's name and the member's UUID, the identity the control plane derives from the replica it placed there. vbdev_cbt_auto_epoch_open finds the cbt stacked on that bdev and opens the epoch there. An epoch holds the writes a member misses from its departure on, so it opens only for a member that left current: its superblock slot configured at the generation its peers carry, not a write-only joiner, and not a failed member an earlier removal failed to take (a new removal_failed flag, which unlike remove_error survives the retry). The cbt also refuses while it carries any live epoch, frozen or rebuilding included: a freeze empties the live bitmap, and an epoch opened after one may start short of writes the member missed just before it left. Each of those members is rebuilt in full, as before.
… epoch is a delta (bdev_cbt_rotate) The live bitmap is cleared only by bdev_cbt_reset, which needs a moment when every backend is known in sync. A member that leaves without notice gives none, so the epoch the raid opens at its ejection took the whole live bitmap: every write since the last reset. On a volume that went days without an epoch, that is the whole device, and the "delta" rebuild copied all of it (measured on a lab cluster: 81920 of 81920 chunks). bdev_cbt_rotate, called periodically while no epoch is live, moves the live bitmap to a previous-window bitmap (the window before is dropped) and starts it again empty. vbdev_cbt_auto_epoch_open ORs the previous window back before it opens, so its delta holds one to two windows of writes before the departure and everything after. Two rotations are at least CBT_ROTATE_MIN_INTERVAL_US (10 s) apart whoever asks: a write the member missed a few milliseconds before its ejection is always in one of the two bitmaps. A rotation is refused (-EBUSY) while an epoch is live, since that epoch reads the live bitmap whole; a reset clears both.
…-use fix: blob thin clusters, raid rebuild copies and ejection epoch, cbt history windows, nvmf reservations, bdev_nvme passthrough (0044–0052)
… superblocks (0053) A raid create took every listed leg as a configured member. The control plane had to clear the legs' superblocks first, since a leg stamped by another raid is refused, and clearing erased the only evidence of which leg holds the volume's last acknowledged writes: a leg left behind by an ejection came back as a configured member, and a creator that lost its place could take legs from a raid still serving. bdev_raid_create with a view_epoch (a raid1 with a superblock, the volume's lineage and the creating incarnation) now reads every listed leg's superblock, opened read-only, before it claims any: - A superblock carries a stamp, minor 2, carved out of the header's reserved bytes: the lineage, the owner incarnation and the last view epoch it stamped. A minor-1 superblock reads as unstamped. - A leg stamped for another volume is refused (-EEXIST). A leg another owner stamped in the view's epoch or a newer one is refused (-EBUSY), nothing claimed or written, unless force_restamp names that exact epoch (expected_view_epoch; -ESTALE for any other). - Among the legs stamped for the volume, only those at the highest content generation, configured in their own superblock, are configured, at that generation. Every other leg (behind, a copy never completed, unstamped, or stamped before minor 2) stays out, in an empty slot where the superblock records it MISSING at its own generation, and the control plane rebuilds it. No stamped leg holding a configured copy is -ENODATA. When no leg carries a stamp of the volume, every leg is configured, as before. - Each member is read again as the raid claims it and must still be what the arbitration read (-ESTALE). A leg the survey cannot read refuses the create: read as blank, the one leg at the highest generation would be left out, never read again, and the raid would come up on the legs behind it. - An add into an online raid of a lineage checks the leg the same way and takes it as a new member, rebuilt like any other, where upstream refuses any leg another raid stamped. - A raid reassembled from a stamped superblock keeps its view, and get_bdevs reports lineage and view_epoch. declared_slots creates a raid wider than its legs, the slots past them empty and sized like the members, for members added later: the raid comes online degraded instead of through null bdevs added and removed. A level that operates only whole refuses it. bdev_raid_ut gets a raid1-level module and a superblock per leg, and pins the arbitration, each refusal, the claim-time re-read, the empty slots, the online add and the reassembly (14 tests); bdev_raid_sb_ut pins the stamp and the slots init_superblock writes. Each rule was taken out in turn and its tests went red: the generation (3 tests), the admission (3), the claim-time re-read (1), the online add (1), the empty slot size (3, and an online add hits the data_size assert), the reassembly (1), the stamp and the MISSING slots (bdev_raid_sb_ut), a read error read as blank (1), the view parameters (1). Every raid suite passes: bdev_raid 39, bdev_raid_sb 15, raid0 8, raid1 7, raid5f 8, concat 3.
…(0054)
0053 refuses a create on legs another owner stamped in the create's view
epoch or a newer one. A raid stamps its legs once, at its create, so an
owner that stays in the view while the view moves on held them with an
old epoch, and a create from a view the owner already left behind took
them.
bdev_raid_journal_view_epoch {name, view_epoch, expected_incarnation,
lineage} stamps a new epoch on an online raid: the superblock's
view_epoch, its owner (the raid's incarnation, which takes over the
stamp of a raid it claimed after a reassembly), and the view_epoch of
each member it serves, then writes the superblock to them. An owner
that stays in the view journals each epoch, so a create from a view no
more recent finds the legs held for as long as the owner serves.
- Only the raid's incarnation journals (-ESTALE for another).
- The epoch never goes back (-ESTALE). The same epoch is written again,
so a journal whose write failed is retried as it was.
- A raid created outside the view protocol joins its volume's view when
the journal names the lineage (-EINVAL without one, -EEXIST for
another lineage).
- A superblock an older image wrote is raised to minor 2, where the
stamp is read.
bdev_raid_ut pins the stamp, each refusal, the join and a failed write.
With superblock writes landing on the legs, a create by another
incarnation is refused at the journaled epoch and at the one before,
and takes the legs at the next. Taken out in turn, the stamp (3 tests,
the end-to-end one included), the monotonic epoch (1), the minor 2
upgrade (2) and the join (1) each turned their tests red. Every raid
suite passes: bdev_raid 44, bdev_raid_sb 15, raid0 8, raid1 7, raid5f 8,
concat 3.
…ut a superblock bdev_tier_read_sb answered valid:false both for a disk that carries no valid superblock (-EILSEQ) and for a read that failed (-EIO, -ENOMEM). The superblock module tells the two apart; the RPC folded them. A caller that reads valid:false as "no superblock" then lets an older generation win the reassembly when the read of the disk holding the newest one fails, and, when the reads fail on every band, lays a fresh layout over disks that hold data. The RPC now answers an error for a failed read, with its errno, and valid:false only for a disk without a valid superblock. The verdict is tier_sb_read_answer() in vbdev_tier.h, which test_tier_units pins: a failed read, and a read with neither a superblock nor an error, is FAILED. With a failed read mapped to "none" the test goes red (2 assertions). The tier and lvol modules build against the series.
A member copied from another member of its raid while no raid was assembled carries the source's superblock, where its own entry says what the source knew of it: MISSING when it was behind. The next create in a view leaves it out and rebuilds it in full, and with the source gone refuses the volume although the copy holds the content. bdev_raid_restamp_copied_member gives its own entry the source's standing: CONFIGURED at the source's content generation and view epoch.
Examine assembled a raid of a lineage from whichever member it read first: the raid came online on that member's record of the others, wrote its superblock and rebuilt the others in full before the control plane's fence, and the members' own standings were never arbitrated. A member stamped for a lineage now stays unclaimed until a create in its view assembles it; an unstamped one is reassembled as upstream does.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR now carries four things:
bdev_cbt_rotate), merged here from fix: blob thin clusters, raid rebuild copies and ejection epoch, cbt history windows, nvmf reservations, bdev_nvme passthrough (0044–0052) #26 (the second section). Each patch comes as two commits: its tests alone, red on the series below it, then the fix.scripts/patches.sh checkpasses on a clean checkout at the pinned commit (55 patches), andregenfrom the series leaves every patch byte-identical to this branch. In the image pipeline, the 18 suites pass on amd64 and arm64.More of the series lands on this branch next (see the last section).
The image build runs the upstream unit tests
Why
The image configured
--disable-unit-tests, and no patch of the series had ever run the upstream suites of the code it changes. That covers the blobstore, raid, lvol, nvmf, the bdev layer and util.A first attempt, in #26, ran them in a separate job that built SPDK a second time. It was dropped: tests should share the image's build.
What
One build, which the tests and the release both use
unit-testsstage inimages/spdk/Dockerfilecomes after the builder. It compiles the owed suites against the libraries the builder made, and runs them.ci-gateand the release publish, which needs every build of the matrix.Tests that follow the series
images/spdk/unit-tests.mapmaps a patched path to the leaf suites oftest/unit/that cover it. The longest prefix wins, and-says upstream has no suite..cunderlib/,module/orapp/that matches no line fails the stage, so a component is mapped when its first patch lands.test/unit/adds its own suite, so a patch's own tests always run.vbdev_lvol,bdev_nvme, 3 bdev, 2 blob,jsonrpc_server,nvme_tcp, nvmfsubsystem,rpc,bit_array.No silent skip
blob.cneeds CUnit 2.1-3, which Fedora 43 ships (CUnit-2.1.3-35.fc43, built from the2.1-3tarball).Faster rebuilds
pkgdep.sh, liburing and libaio) now come before the modules and patches. They depend on the upstream tree alone, so a patch change no longer rebuilds them.patches/README.md, "Unit tests", documents it. Locally:docker build -f images/spdk/Dockerfile --target unit-tests --build-arg BUILD_TYPE=debug .What the suites found
Each failure is fixed in the patch that caused it, through
scripts/patches.sh regen.Before this pipeline (the separate job's runs):
vbdev_lvol_utno longer compiled. 0005 makesvbdev_lvol.cincludevbdev_tier.h, and the unit test's Makefile had no path to it. Fixed in 0009 and 0041 (155bd2d).First run (155bd2d): 15 of 18 suites passed. Three failed (fixed in 0f74e3f):
bdev_raid_utdid not link. Ten patches call functions the suite neither includes nor stubs. The calls also pulled objects oflibspdk_jsonandlibspdk_jsonrpcwhose other definitions clash with the suite's own. Each patch now stubs what it calls: 0007, 0009, 0013, 0015, 0017, 0018, 0019, 0021, 0025, 0033.bdev_raid_utthen failed on 0014. A raid is refused without an incarnation, and the suite's delete requests carried an uninitialisedexpected_incarnation. The suite now creates its raids under an incarnation, and deletes them unchecked.raid1_utdid not link: 0009 reads the NVMe status of a failed member write. The stub reports a device fault.rpc_utfailed on 0010. The listen chmods a socket file that the suite's stubbed bind never creates. The suite now records the chmod instead.Second run (0f74e3f): all 18 suites pass on amd64 and arm64.
New tests of the patches' own behaviour, added where the suite was already being fixed:
bdev_raid_ut): no creation without an incarnation. A delete that expects another incarnation is refused, and the raid stays.raid1_ut): a thin-exhausted member is kept, and the raid I/O is marked to complete withCAPACITY_EXCEEDED.rpc_ut): the bind runs under umask 077, the socket is chmod'ed to 0600, and the process umask is restored. A failed chmod refuses the listen.Patches 0044 to 0052, and cbt history windows
bdev_cbt_rotate)0044: a thin cluster is cleared before use
Problem. A cluster that a write allocates in a blob backed by zeroes (a thin blob with no parent, or a clone over a range its snapshot never allocated) is inserted as the device left it. The write fills its own range, and the rest of the cluster returns whatever was stored there before.
The free path counts on unmap to leave zeroes (see
spdk_free_cluster_unmap_complete: "so if the cluster is reclaimed in the future, it won't leak old data"). A SATA SSD does not guarantee that, and a bdev without unmap support never provides it. So:Remanence across blobs. The unwritten part of a new cluster returns another blob's data. A filesystem on top does not expose it, since it reads unwritten extents as zeroes, but anything that reads the raw device does.
Mirror legs differ in bytes nobody wrote. Each device keeps its own residue, and a raid1 verify over such lvols (
bdev_raid_verify_ranges, or any block compare of the legs) reports a content divergence.mkfs.xfsalone leaves such bytes:WRITE_ZEROESon an unallocated cluster allocates it too);On a lab cluster, a fresh 5 GiB three-leg mirror was reported divergent in 1198 blocks of 4 KiB right after
mkfs.xfs, with no application data, one leg in the minority everywhere.Fix.
bs_allocate_and_copy_clusterclears it withwrite_zeroesbefore inserting it.write_zeroes) that covers the whole cluster skips the clear: it overwrites everything, and a rebuild or a copy, which write whole clusters, does not write twice.Cost: one extra cluster write (1 MiB by default) on the first partial write into a cluster, once per cluster lifetime. Not covered: thick blobs get their clusters at creation and still count on the free path.
Tests.
blob_ut:blob_thin_prov_new_cluster_reads_zeroes: every free cluster is filled with0xAA, then a one-io_unit write goes into a new cluster. The rest of the cluster reads zeroes.blob_thin_prov_full_cluster_write_is_not_cleared_first: a whole-cluster write costs the payload and the metadata page(s), and no clear.blob_thin_prov_rw,blob_thin_prov_write_count_ioandblob_thin_prov_rlenow count the cleared cluster in their write accounting.In the image pipeline, tests only (e539f7e): red on amd64 and arm64,
FAILED: 1 of 18 suites: lib/blob/blob.c, four tests in each of the 5 suites ofblob_ut. With the fix (e94d862): all 18 suites pass on amd64 and arm64.0045 and 0046: a seeded rebuild copies its delta, and only its delta
A seeded rebuild (a member that missed some writes, with the dirty ranges known) copied every window a range touched in full, and still locked and skipped every clean window. Measured on a three-node lab cluster under a random-write load: 5 GiB re-copied for 395 MiB of dirty chunks (4794 ranges), in 110 s. Copying only the ranges, one window per range part, still took 100 s. The cost was the window cycle (quiesce, copy, channel update, unquiesce), not the bytes.
0045. A seeded process now jumps over what its ranges leave clean, with no window at all: only a copy needs a window's quiesce, and the seeded target already receives every write on either side of the process offset. A window copies every range part it holds under its one lock. The verify phase and the copied-bytes accounting are unchanged.
0046. A JSON-RPC request carries about 200 ranges, so a caller had to fold a larger delta to fit it, and folding a scattered delta recopies every clean block in between (4485 dirty ranges folded to 128 recopied 3.2 GiB of clean blocks: as slow as a full rebuild).
bdev_raid_start_seeded_rebuild_from_epochseeds the same rebuild with the dirty ranges of a frozen CBT epoch, read in-process throughvbdev_cbt_query_epoch_ranges. A delta larger than the caller's cap is refused (-E2BIG) rather than truncated, and the caller rebuilds in full. Both RPCs share their checks and their audit.Tests.
bdev_raid_ut: two seed ranges inside one window copy 12 blocks under one lock (the window was copied whole); a range across three windows copies 9 blocks under three locks; a range at the end of the raid takes one lock (the clean windows before it were each locked); 64 one-block ranges from a faked epoch copy their 64 blocks, and a delta the query refuses starts nothing. Every raid suite stays green.0047 and 0048: a full rebuild keeps the link busy, and only with data
0047. A rebuild window was copied as one request: read whole from one member, then written whole. Read and write never overlapped, and the first member served every read. Measured on a 3-node bench, a full rebuild of a 5 GiB three-leg mirror under a verified write load ran at 36 MiB/s for 141 s, every byte read across the link from a remote member while the local one sat idle.
raid_bdev_process_request_max_partnow offers a window inRAID_BDEV_PROCESS_MAX_QDparts of whole write units, and raid1 copies one part per request: each part's write starts as its own read completes, and the reads spread over the members in sync. raid5f, bound to its stripe, does not call it.0048. A part the source read back as nothing but zeroes goes to the target as write-zeroes: the same blocks, read back the same, and no payload on the wire. On the same bench, about 3 of the 5 GiB of a full rebuild were zeroes, all of them through the nexus node's egress. Not when the target cannot take write-zeroes, nor when the blocks carry metadata. The check reads the part already in memory and stops at the first nonzero word.
Tests.
raid1_ut: a 64-block window is copied by four requests in flight together, the reads tiling the window and spread evenly over the members (before: one request, and on three members the second served no read); four parts read back as data and zeroes go out as data and write-zeroes, or all as data when the target cannot take write-zeroes.bdev_raid_ut: the parts are whole write units and fit the queue depth.0049: write-zeroes leaves an unallocated thin cluster unallocated
On a blob backed by zeroes, an unallocated cluster already reads zeroes, so write-zeroes over it now completes at once: no cluster allocated, nothing written. Before, it allocated the cluster (cleared, since 0044) and zeroed it. Since 0048 sends every never-written part of a rebuild as write-zeroes, a thin volume rebuilt in full took its whole size on the target, and its zeroes still went to the target's disk (94 MiB/s of writes through a 5 GiB rebuild with about 3 GiB never written). An allocated cluster is still zeroed, and a clone, which reads its unallocated clusters from its parent, still allocates and zeroes.
Tests.
blob_ut: write-zeroes over half of an unallocated cluster of a thin blob allocates nothing, writes nothing and reads back zeroes; over an allocated cluster holding data it zeroes what it covers; on a clone it allocates the cluster and zeroes it.0050: a reservation held by all registrants outlives its acquirer
Every registrant holds a Write Exclusive or Exclusive Access – All Registrants reservation, so it stays when the registrant that acquired it leaves (preempted, or unregistering). It stayed under the key of the registrant that left: the persisted state named a reservation key no registrant held, and the restore refused it (
-EINVAL), failing the add of the namespace. After a writer handover, the next restart of a target left the namespace out, unexported. The reservation now stands under the key of the registrant that holds it next. The restore takes a state persisted before this fix under the key of its first registrant; a single holder's key that no registrant holds is still refused.Tests.
subsystem_ut: an All Registrants reservation acquired by A, then A preempted by B, keeps a key B holds and restores; the same once B unregisters after C registered. The restore test refuses a missing key for a single holder only.0051: an I/O passthrough needs a connected qpair
bdev_nvme_send_cmdwith an I/O command took the controller channel's qpair (bdev_nvme_get_io_qpair) without a check and handed it tospdk_nvme_ctrlr_cmd_io_raw_with_md. bdev_nvme frees that qpair from a disconnect until the reconnect (bdev_nvme_disconnected_qpair_cb), so a reservation command (reservations are I/O commands) sent to a member whose path had just gone down dereferenced NULL and killed the process, every raid it carried with it. The command is now refused with-ENXIOand the channel given back; a command for which no channel can be had is refused with-ENOMEM.Tests.
bdev_nvme_utnow compilesnvme_rpc.cin, so the passthrough is driven against a real controller channel. With the channel's I/O qpair cleared, as a disconnect clears it, the command was still handed to the NVMe library (rc 0, channel kept). It is now refused with-ENXIO, the channel given back, and it goes through once connected again. 55/55 with the fix.0052: a raid1 ejection opens its epoch on the raid's cbt
0019 opens a CBT epoch when a member leaves a raid1, so the member's return is a delta instead of a full rebuild. It asked each survivor's own bdev name for that epoch, but the cbt sits on the raid, not on its members: every call answered
-ENODEV, no epoch ever opened, and a member lost without notice was always rebuilt in full.The raid now asks once, with its own name and the member's UUID.
vbdev_cbt_auto_epoch_openfinds the cbt stacked on that bdev and opens the epoch there. The UUID is the identity a control plane can derive for the member: an NVMe-oF initiator bdev inherits the UUID of the namespace it attaches.An epoch holds the writes a member misses from its departure on, so two guards keep the delta whole:
removal_failedflag which, unlikeremove_error, survives the retry). A member being rebuilt in full, or seated behind, owes more than the writes it misses from now on.Every member these guards turn away is rebuilt in full, as before.
Tests.
bdev_raid_utdrives real ejections of a superblock raid (the suite's module stands in for raid1). A current member, removed or failed, opens one epoch, on the raid's name and under the member's UUID. Before the fix it opened 31, one per survivor name. Five ejections open none: a member being rebuilt in full, one behind its peers' generation, a write-only joiner, a raid without a superblock, and a failed member whose first removal failed (driven through a quiesce that fails once). Tests only: 2 of 25 red, 11 assertions. With the fix: 25/25, and every raid suite stays green. The cbt side has no suite in the image pipeline (the module's tests run against a standalone model).cbt: the live bitmap in rotating windows (
bdev_cbt_rotate)With 0052 the epoch at ejection opens, but on a lab cluster its delta resolved to 81920 of 81920 chunks: the live bitmap holds every write since its last
bdev_cbt_reset, and nothing can reset it at the ejection, since the writes the member just missed are in it. On a volume that goes days without an epoch, the "delta" is the whole device.bdev_cbt_rotate, called periodically while no epoch is live, moves the live bitmap to a previous-window bitmap (dropping the window before) and starts it again empty.vbdev_cbt_auto_epoch_openORs the previous window back before it opens, so its delta holds one to two windows of writes before the departure and everything after. Two rotations are at leastCBT_ROTATE_MIN_INTERVAL_US(10 s) apart whoever asks, so a write the member missed a few milliseconds before its ejection is always in one of the two bitmaps. A rotation is refused with-EBUSYwhile an epoch is open, frozen or rebuilding, since that epoch reads the live bitmap whole, and with-EAGAINwhen it comes too soon. A reset clears both bitmaps. The module README describes it.On the lab cluster, with a rotation every 15 s, the epoch opened at a member's silent loss froze 32852 of 81920 chunks (the writes of the outage) instead of all of them, and the member's repair took 43 to 45 s instead of 65 s. A volume younger than two windows still carries its first writes in the delta (the discard of a fresh
mkfs, for one).Raid membership read from the superblocks
0053: a raid create in a view arbitrates its members from their superblocks
Problem.
bdev_raid_createtakes every listed leg as a configured member, and refuses a leg another raid stamped (-EEXIST). A control plane that recreates a raid over legs an earlier raid stamped (a failover, a republish on another node, a restart) has to clear their superblocks first (0025). That erases the only record of which leg holds the last acknowledged writes: a leg left behind by an ejection is taken back as configured, with no rebuild, and a creator that lost its place takes legs from a raid that is still serving.Fix. A create with a
view_epoch(a raid1 with a superblock, the volume'slineageand the creatingincarnation) reads every listed leg's superblock, opened read-only, before it claims any.lineage, theownerincarnation and the lastview_epochit stamped: minor 2, carved out of the header's reserved bytes, offsets pinned by static asserts. A minor-1 superblock reads as unstamped.-EEXIST). A leg another owner stamped in the view's epoch or a newer one belongs to a raid that a view at least as recent still counts, and is refused too (-EBUSY). Nothing is claimed or written.force_restampoverwrites such stamps only at the epoch named byexpected_view_epoch(-ESTALEat any other). A leg stamped by the same incarnation, or at an older epoch, is free.-ENODATA). When no leg carries a stamp of the volume (blank legs, or legs stamped by an older image), every leg is configured, as before.-ESTALE). A leg the survey cannot read refuses the create: read as blank, the one leg at the highest generation would be left out, never read again, and the raid would come up on the legs behind it.bdev_raid_get_bdevsreportslineageandview_epoch.declared_slots. A raid can be created wider than its legs, the slots past them empty and sized like the members, for members added later. It comes online degraded, with no null bdevs added and removed to open the slots. A level that operates only whole refuses it.Tests.
bdev_raid_utgets a raid1-level module and a superblock per leg, with a read error (transient or not) and a later superblock for a leg that changes between the survey and the claim. 14 tests: blank legs, the highest generation, legs without a configured stamp, another owner's legs (same and newer epoch), an older epoch and the create's own incarnation, another volume, restamping, no configured copy, a read error, a leg that changed (its generation, a stamp appearing), the parameters a create in a view needs, declared slots, the online add, and the reassembly.bdev_raid_sb_utpins the stamp and the slot entriesinit_superblockwrites, read back after a write. Each rule was taken out in turn and its tests went red. Every raid suite passes: bdev_raid 39, bdev_raid_sb 15, raid0 8, raid1 7, raid5f 8, concat 3.0054: the raid's owner journals each view epoch on its members
Problem. 0053 refuses a create on legs another owner stamped in the create's view epoch or a newer one. A raid stamps its legs once, at its create: an owner that stays in the view while the view moves on holds them with an old epoch, and a create from a view the owner already left behind takes them.
Fix.
bdev_raid_journal_view_epoch {name, view_epoch, expected_incarnation, lineage}stamps a new epoch on an online raid: the superblock'sview_epoch, itsowner(the raid's incarnation, which takes over the stamp of a raid it claimed after a reassembly), and theview_epochof each member it serves, then writes the superblock to them. An owner that stays in the view journals each epoch, so a create from a view no more recent finds the legs held for as long as the owner serves.-ESTALEfor another).-ESTALE). The same epoch is written again, so a journal whose write failed is retried as it was.lineage(-EINVALwithout one,-EEXISTfor another).Tests.
bdev_raid_ut(5 tests): the stamp, each refusal, the join, a failed write, and, with superblock writes landing on the legs, a create by another incarnation refused at the journaled epoch and at the one before, then taking the legs at the next. Taken out in turn, the stamp, the monotonic epoch, the minor 2 upgrade and the join each turned their tests red. Every raid suite passes: bdev_raid 44, bdev_raid_sb 15, raid0 8, raid1 7, raid5f 8, concat 3.0055: a copied member takes its source's standing
Problem. A control plane that repairs a member while no raid is assembled (a cold repair) copies another member's content onto it, superblock included. The copy then carries the source's superblock, where the copy's own entry says what the source knew of it: MISSING, at an older generation, since it was behind. The next create in a view (0053) reads that entry, leaves the copy out and rebuilds it in full, although it holds the source's content. With the source gone by then, the create is refused (
-ENODATA) and the volume cannot come up, although a complete copy exists.Fix.
bdev_raid_restamp_copied_member {base_bdev, member_uuid, source_uuid, lineage}gives the copy's own entry the source's standing: CONFIGURED at the source's content generation and view epoch. The superblock'sseq_numberand crc are renewed, and the superblock is written back at the head of the bdev. No other entry and no stamp field changes. The pure function israid_bdev_sb_restamp_copied_member()inbdev_raid_sb.c.-EINVAL), of another lineage (-EEXIST), without an entry for either member or naming the same one twice (-ENOENT), or whose source entry is not CONFIGURED (-ESTALE). Copied from a member that was itself behind, the copy would be behind.restamped: false). The response carries the copy'scontent_generation.bdev_raid_clear_superblock: a member of an assembled raid, or a namespace a subsystem exposes, is claimed and cannot be restamped.-ENOTSUP).The restamp claims that the copy holds what the source held when it was copied. The caller has to know that no raid wrote to the source since then. A raid that left the copy out and kept writing does not move the source's generation, since the copy was already behind.
Tests.
bdev_raid_sb_ut(2 tests, in each of its 3 suites): the restamp (state, generation, epoch, sequence number, crc, a third member left as it was) and the no-op that follows it, then each refusal, with the superblock left as it was. With the state change and the crc renewal taken out, the restamp test went red in all three suites.bdev_raid_utstubs the pure function. Every raid suite passes: bdev_raid 44, bdev_raid_sb 21, raid0 8, raid1 7, raid5f 8, concat 3.0056: discovery leaves a stamped member to a create in view
Problem. A control plane that attaches a volume's legs before it fences, and creates the raid in a view after (0053), never got to arbitrate: SPDK's examine assembled the raid as soon as the legs appeared, from the superblock of the first member it read. On the lab cluster, at each republish, the raid came online on the first member's record of the others. It wrote its superblock and started rebuilding the second member 200 ms before the writer's reservations were acquired. The member it rebuilt in full was a copy whose own superblock recorded it current (0055).
Fix. Examine no longer assembles a raid from a member whose superblock is stamped for a lineage (minor 2 with a lineage). The member stays unclaimed, and a create in its view assembles it, from every member's own standing. A member without a stamp is still reassembled as upstream does. An add the control plane issues still reads the member's superblock (0053).
Tests.
bdev_raid_ut(45 tests): a stamped member left unclaimed, with no raid created and its superblock read once. As the control, the same member below minor 2 has discovery create its raid. With the check taken out, the test went red. Every raid suite passes.tier: a superblock read that fails is an error
bdev_tier_read_sbansweredvalid: falseboth for a disk that carries no valid superblock (-EILSEQ) and for a read that failed (-EIO,-ENOMEM). The superblock module tells the two apart, and the RPC folded them. A caller that readsvalid: falseas "no superblock" then lets an older generation win the reassembly when the read of the disk holding the newest one fails, and, when the reads fail on every band, lays a fresh layout over disks that hold data.The RPC now answers an error for a failed read, with its errno, and
valid: falseonly for a disk without a valid superblock. The verdict istier_sb_read_answer()invbdev_tier.h, pinned bytest_tier_units(red with a failed read mapped to "none": 2 assertions).docs/RPC-CONTRACT.mdsays it.Next on this branch
bdev_raid_rebuild_rangesreads through the raid, so on a raid1 the read can come from the member whose clusters were just remapped without a copy: they read back without an error, and the write-back puts what they held on every mirror. A member that lost a tier band should leave the raid before the remap and come back through a seeded rebuild of the remapped ranges and the writes it missed, which needsbdev_raid_start_seeded_rebuild_from_epochto take extra ranges.