Skip to content

Upstream unit tests in the image build; patches 0044–0054, cbt history windows, tier read errors - #27

Merged
rducom merged 27 commits into
mainfrom
unit-tests-in-the-image-build
Oct 6, 2026
Merged

rducom merged 27 commits into
mainfrom
unit-tests-in-the-image-build

Conversation

@rducom

@rducom rducom commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

This PR now carries four things:

  1. The image build runs the upstream unit suites of everything the series patches (the first section).
  2. Nine patches of the series, 0044 to 0052, and one change to the cbt module (bdev_cbt_rotate), merged here from fix: blob thin clusters, raid rebuild copies and ejection epoch, cbt history windows, nvmf reservations, bdev_nvme passthrough (0044–0052) #26 (the second section). Each patch comes as two commits: its tests alone, red on the series below it, then the fix.
  3. Patches 0053 to 0056: raid membership read from the superblocks (the third section). A raid create in a view arbitrates its members from their superblocks, the raid's owner journals each view epoch on its members, a member copied while no raid was assembled takes its source's standing, and discovery leaves a stamped member to a create in view. Each comes as one commit, and each of their rules was seen to fail its tests when taken out.
  4. A tier fix: a superblock read that fails is an error, not a disk without a superblock (the fourth section).

scripts/patches.sh check passes on a clean checkout at the pinned commit (55 patches), and regen from the series leaves every patch byte-identical to this branch. In the image pipeline, the 18 suites pass on amd64 and arm64.

More of the series lands on this branch next (see the last section).

The image build runs the upstream unit tests

Why

The image configured --disable-unit-tests, and no patch of the series had ever run the upstream suites of the code it changes. That covers the blobstore, raid, lvol, nvmf, the bdev layer and util.

A first attempt, in #26, ran them in a separate job that built SPDK a second time. It was dropped: tests should share the image's build.

What

One build, which the tests and the release both use

  • A unit-tests stage in images/spdk/Dockerfile comes after the builder. It compiles the owed suites against the libraries the builder made, and runs them.
  • The collector (release image) and the debug image both derive from it. No job is added to the workflow.
  • The suites run in the debug build, on each arch: asserts are on there, which is how upstream runs them. The release build passes through at no cost.
  • A failing suite fails its debug build, and with it ci-gate and the release publish, which needs every build of the matrix.

Tests that follow the series

  • images/spdk/unit-tests.map maps a patched path to the leaf suites of test/unit/ that cover it. The longest prefix wins, and - says upstream has no suite.
  • A patched .c under lib/, module/ or app/ that matches no line fails the stage, so a component is mapped when its first patch lands.
  • A patched file under test/unit/ adds its own suite, so a patch's own tests always run.
  • Today the map selects 18 suites: 6 raid, vbdev_lvol, bdev_nvme, 3 bdev, 2 blob, jsonrpc_server, nvme_tcp, nvmf subsystem, rpc, bit_array.

No silent skip

  • Each suite is built from its leaf directory, because parent Makefiles skip some suites with a warning. blob.c needs CUnit 2.1-3, which Fedora 43 ships (CUnit-2.1.3-35.fc43, built from the 2.1-3 tarball).
  • A missing binary fails the stage, as does a failing suite.

Faster rebuilds

  • The build dependencies (pip pins, pkgdep.sh, liburing and libaio) now come before the modules and patches. They depend on the upstream tree alone, so a patch change no longer rebuilds them.

patches/README.md, "Unit tests", documents it. Locally:

docker build -f images/spdk/Dockerfile --target unit-tests --build-arg BUILD_TYPE=debug .

What the suites found

Each failure is fixed in the patch that caused it, through scripts/patches.sh regen.

Before this pipeline (the separate job's runs): vbdev_lvol_ut no longer compiled. 0005 makes vbdev_lvol.c include vbdev_tier.h, and the unit test's Makefile had no path to it. Fixed in 0009 and 0041 (155bd2d).

First run (155bd2d): 15 of 18 suites passed. Three failed (fixed in 0f74e3f):

  • bdev_raid_ut did not link. Ten patches call functions the suite neither includes nor stubs. The calls also pulled objects of libspdk_json and libspdk_jsonrpc whose other definitions clash with the suite's own. Each patch now stubs what it calls: 0007, 0009, 0013, 0015, 0017, 0018, 0019, 0021, 0025, 0033.
  • bdev_raid_ut then failed on 0014. A raid is refused without an incarnation, and the suite's delete requests carried an uninitialised expected_incarnation. The suite now creates its raids under an incarnation, and deletes them unchecked.
  • raid1_ut did not link: 0009 reads the NVMe status of a failed member write. The stub reports a device fault.
  • rpc_ut failed on 0010. The listen chmods a socket file that the suite's stubbed bind never creates. The suite now records the chmod instead.

Second run (0f74e3f): all 18 suites pass on amd64 and arm64.

New tests of the patches' own behaviour, added where the suite was already being fixed:

  • 0014 (bdev_raid_ut): no creation without an incarnation. A delete that expects another incarnation is refused, and the raid stays.
  • 0009 (raid1_ut): a thin-exhausted member is kept, and the raid I/O is marked to complete with CAPACITY_EXCEEDED.
  • 0010 (rpc_ut): the bind runs under umask 077, the socket is chmod'ed to 0600, and the process umask is restored. A failed chmod refuses the listen.

Patches 0044 to 0052, and cbt history windows

Patch Area What it changes
0044 blob a thin cluster is cleared before use
0045 raid a seeded rebuild copies its ranges, and only under the windows they need
0046 raid a seeded rebuild reads its ranges from a frozen CBT epoch, in-process
0047 raid1 a rebuild copies its window in parts, all in flight together
0048 raid1 a rebuild writes an all-zero part as write-zeroes
0049 blob write-zeroes leaves an unallocated thin cluster unallocated
0050 nvmf a reservation held by all registrants outlives its acquirer
0051 bdev_nvme an I/O passthrough needs a connected qpair
0052 raid1, cbt a raid1 ejection opens its epoch on the raid's cbt
cbt cbt the live bitmap in rotating history windows (bdev_cbt_rotate)

0044: a thin cluster is cleared before use

Problem. A cluster that a write allocates in a blob backed by zeroes (a thin blob with no parent, or a clone over a range its snapshot never allocated) is inserted as the device left it. The write fills its own range, and the rest of the cluster returns whatever was stored there before.

The free path counts on unmap to leave zeroes (see spdk_free_cluster_unmap_complete: "so if the cluster is reclaimed in the future, it won't leak old data"). A SATA SSD does not guarantee that, and a bdev without unmap support never provides it. So:

  • Remanence across blobs. The unwritten part of a new cluster returns another blob's data. A filesystem on top does not expose it, since it reads unwritten extents as zeroes, but anything that reads the raw device does.

  • Mirror legs differ in bytes nobody wrote. Each device keeps its own residue, and a raid1 verify over such lvols (bdev_raid_verify_ranges, or any block compare of the legs) reports a content divergence. mkfs.xfs alone leaves such bytes:

    • the 2 KiB after the headers of each allocation group;
    • most of the last cluster, since it zeroes only the device's last 128 KiB (and WRITE_ZEROES on an unallocated cluster allocates it too);
    • the rest of the cluster at the end of its log.

    On a lab cluster, a fresh 5 GiB three-leg mirror was reported divergent in 1198 blocks of 4 KiB right after mkfs.xfs, with no application data, one leg in the minority everywhere.

Fix.

  • When the new cluster is not copied from a parent, bs_allocate_and_copy_cluster clears it with write_zeroes before inserting it.
  • A write (or write_zeroes) that covers the whole cluster skips the clear: it overwrites everything, and a rebuild or a copy, which write whole clusters, does not write twice.
  • A clear that fails returns the cluster to the pool.
  • The inflate path, whose clusters over zero ranges were also handed out uninitialized, goes through the same branch, so an inflated cluster reads zeroes too.

Cost: one extra cluster write (1 MiB by default) on the first partial write into a cluster, once per cluster lifetime. Not covered: thick blobs get their clusters at creation and still count on the free path.

Tests. blob_ut:

  • blob_thin_prov_new_cluster_reads_zeroes: every free cluster is filled with 0xAA, then a one-io_unit write goes into a new cluster. The rest of the cluster reads zeroes.
  • blob_thin_prov_full_cluster_write_is_not_cleared_first: a whole-cluster write costs the payload and the metadata page(s), and no clear.
  • blob_thin_prov_rw, blob_thin_prov_write_count_io and blob_thin_prov_rle now count the cleared cluster in their write accounting.

In the image pipeline, tests only (e539f7e): red on amd64 and arm64, FAILED: 1 of 18 suites: lib/blob/blob.c, four tests in each of the 5 suites of blob_ut. With the fix (e94d862): all 18 suites pass on amd64 and arm64.

0045 and 0046: a seeded rebuild copies its delta, and only its delta

A seeded rebuild (a member that missed some writes, with the dirty ranges known) copied every window a range touched in full, and still locked and skipped every clean window. Measured on a three-node lab cluster under a random-write load: 5 GiB re-copied for 395 MiB of dirty chunks (4794 ranges), in 110 s. Copying only the ranges, one window per range part, still took 100 s. The cost was the window cycle (quiesce, copy, channel update, unquiesce), not the bytes.

0045. A seeded process now jumps over what its ranges leave clean, with no window at all: only a copy needs a window's quiesce, and the seeded target already receives every write on either side of the process offset. A window copies every range part it holds under its one lock. The verify phase and the copied-bytes accounting are unchanged.

0046. A JSON-RPC request carries about 200 ranges, so a caller had to fold a larger delta to fit it, and folding a scattered delta recopies every clean block in between (4485 dirty ranges folded to 128 recopied 3.2 GiB of clean blocks: as slow as a full rebuild). bdev_raid_start_seeded_rebuild_from_epoch seeds the same rebuild with the dirty ranges of a frozen CBT epoch, read in-process through vbdev_cbt_query_epoch_ranges. A delta larger than the caller's cap is refused (-E2BIG) rather than truncated, and the caller rebuilds in full. Both RPCs share their checks and their audit.

Tests. bdev_raid_ut: two seed ranges inside one window copy 12 blocks under one lock (the window was copied whole); a range across three windows copies 9 blocks under three locks; a range at the end of the raid takes one lock (the clean windows before it were each locked); 64 one-block ranges from a faked epoch copy their 64 blocks, and a delta the query refuses starts nothing. Every raid suite stays green.

0047 and 0048: a full rebuild keeps the link busy, and only with data

0047. A rebuild window was copied as one request: read whole from one member, then written whole. Read and write never overlapped, and the first member served every read. Measured on a 3-node bench, a full rebuild of a 5 GiB three-leg mirror under a verified write load ran at 36 MiB/s for 141 s, every byte read across the link from a remote member while the local one sat idle. raid_bdev_process_request_max_part now offers a window in RAID_BDEV_PROCESS_MAX_QD parts of whole write units, and raid1 copies one part per request: each part's write starts as its own read completes, and the reads spread over the members in sync. raid5f, bound to its stripe, does not call it.

0048. A part the source read back as nothing but zeroes goes to the target as write-zeroes: the same blocks, read back the same, and no payload on the wire. On the same bench, about 3 of the 5 GiB of a full rebuild were zeroes, all of them through the nexus node's egress. Not when the target cannot take write-zeroes, nor when the blocks carry metadata. The check reads the part already in memory and stops at the first nonzero word.

Tests. raid1_ut: a 64-block window is copied by four requests in flight together, the reads tiling the window and spread evenly over the members (before: one request, and on three members the second served no read); four parts read back as data and zeroes go out as data and write-zeroes, or all as data when the target cannot take write-zeroes. bdev_raid_ut: the parts are whole write units and fit the queue depth.

0049: write-zeroes leaves an unallocated thin cluster unallocated

On a blob backed by zeroes, an unallocated cluster already reads zeroes, so write-zeroes over it now completes at once: no cluster allocated, nothing written. Before, it allocated the cluster (cleared, since 0044) and zeroed it. Since 0048 sends every never-written part of a rebuild as write-zeroes, a thin volume rebuilt in full took its whole size on the target, and its zeroes still went to the target's disk (94 MiB/s of writes through a 5 GiB rebuild with about 3 GiB never written). An allocated cluster is still zeroed, and a clone, which reads its unallocated clusters from its parent, still allocates and zeroes.

Tests. blob_ut: write-zeroes over half of an unallocated cluster of a thin blob allocates nothing, writes nothing and reads back zeroes; over an allocated cluster holding data it zeroes what it covers; on a clone it allocates the cluster and zeroes it.

0050: a reservation held by all registrants outlives its acquirer

Every registrant holds a Write Exclusive or Exclusive Access – All Registrants reservation, so it stays when the registrant that acquired it leaves (preempted, or unregistering). It stayed under the key of the registrant that left: the persisted state named a reservation key no registrant held, and the restore refused it (-EINVAL), failing the add of the namespace. After a writer handover, the next restart of a target left the namespace out, unexported. The reservation now stands under the key of the registrant that holds it next. The restore takes a state persisted before this fix under the key of its first registrant; a single holder's key that no registrant holds is still refused.

Tests. subsystem_ut: an All Registrants reservation acquired by A, then A preempted by B, keeps a key B holds and restores; the same once B unregisters after C registered. The restore test refuses a missing key for a single holder only.

0051: an I/O passthrough needs a connected qpair

bdev_nvme_send_cmd with an I/O command took the controller channel's qpair (bdev_nvme_get_io_qpair) without a check and handed it to spdk_nvme_ctrlr_cmd_io_raw_with_md. bdev_nvme frees that qpair from a disconnect until the reconnect (bdev_nvme_disconnected_qpair_cb), so a reservation command (reservations are I/O commands) sent to a member whose path had just gone down dereferenced NULL and killed the process, every raid it carried with it. The command is now refused with -ENXIO and the channel given back; a command for which no channel can be had is refused with -ENOMEM.

Tests. bdev_nvme_ut now compiles nvme_rpc.c in, so the passthrough is driven against a real controller channel. With the channel's I/O qpair cleared, as a disconnect clears it, the command was still handed to the NVMe library (rc 0, channel kept). It is now refused with -ENXIO, the channel given back, and it goes through once connected again. 55/55 with the fix.

0052: a raid1 ejection opens its epoch on the raid's cbt

0019 opens a CBT epoch when a member leaves a raid1, so the member's return is a delta instead of a full rebuild. It asked each survivor's own bdev name for that epoch, but the cbt sits on the raid, not on its members: every call answered -ENODEV, no epoch ever opened, and a member lost without notice was always rebuilt in full.

The raid now asks once, with its own name and the member's UUID. vbdev_cbt_auto_epoch_open finds the cbt stacked on that bdev and opens the epoch there. The UUID is the identity a control plane can derive for the member: an NVMe-oF initiator bdev inherits the UUID of the namespace it attaches.

An epoch holds the writes a member misses from its departure on, so two guards keep the delta whole:

  • The member left current. Its superblock slot is configured at the generation its peers carry, it is not a write-only joiner, and it is not a failed member that an earlier removal failed to take (a new removal_failed flag which, unlike remove_error, survives the retry). A member being rebuilt in full, or seated behind, owes more than the writes it misses from now on.
  • The cbt carries no live epoch. Open, frozen and rebuilding ones all refuse. An open one bounds a round its caller started. A freeze empties the live bitmap, so after one, an epoch opened now may start short of writes the member missed just before it left.

Every member these guards turn away is rebuilt in full, as before.

Tests. bdev_raid_ut drives real ejections of a superblock raid (the suite's module stands in for raid1). A current member, removed or failed, opens one epoch, on the raid's name and under the member's UUID. Before the fix it opened 31, one per survivor name. Five ejections open none: a member being rebuilt in full, one behind its peers' generation, a write-only joiner, a raid without a superblock, and a failed member whose first removal failed (driven through a quiesce that fails once). Tests only: 2 of 25 red, 11 assertions. With the fix: 25/25, and every raid suite stays green. The cbt side has no suite in the image pipeline (the module's tests run against a standalone model).

cbt: the live bitmap in rotating windows (bdev_cbt_rotate)

With 0052 the epoch at ejection opens, but on a lab cluster its delta resolved to 81920 of 81920 chunks: the live bitmap holds every write since its last bdev_cbt_reset, and nothing can reset it at the ejection, since the writes the member just missed are in it. On a volume that goes days without an epoch, the "delta" is the whole device.

bdev_cbt_rotate, called periodically while no epoch is live, moves the live bitmap to a previous-window bitmap (dropping the window before) and starts it again empty. vbdev_cbt_auto_epoch_open ORs the previous window back before it opens, so its delta holds one to two windows of writes before the departure and everything after. Two rotations are at least CBT_ROTATE_MIN_INTERVAL_US (10 s) apart whoever asks, so a write the member missed a few milliseconds before its ejection is always in one of the two bitmaps. A rotation is refused with -EBUSY while an epoch is open, frozen or rebuilding, since that epoch reads the live bitmap whole, and with -EAGAIN when it comes too soon. A reset clears both bitmaps. The module README describes it.

On the lab cluster, with a rotation every 15 s, the epoch opened at a member's silent loss froze 32852 of 81920 chunks (the writes of the outage) instead of all of them, and the member's repair took 43 to 45 s instead of 65 s. A volume younger than two windows still carries its first writes in the delta (the discard of a fresh mkfs, for one).

Raid membership read from the superblocks

0053: a raid create in a view arbitrates its members from their superblocks

Problem. bdev_raid_create takes every listed leg as a configured member, and refuses a leg another raid stamped (-EEXIST). A control plane that recreates a raid over legs an earlier raid stamped (a failover, a republish on another node, a restart) has to clear their superblocks first (0025). That erases the only record of which leg holds the last acknowledged writes: a leg left behind by an ejection is taken back as configured, with no rebuild, and a creator that lost its place takes legs from a raid that is still serving.

Fix. A create with a view_epoch (a raid1 with a superblock, the volume's lineage and the creating incarnation) reads every listed leg's superblock, opened read-only, before it claims any.

  • The stamp. A superblock now carries the volume's lineage, the owner incarnation and the last view_epoch it stamped: minor 2, carved out of the header's reserved bytes, offsets pinned by static asserts. A minor-1 superblock reads as unstamped.
  • Who may take a leg. A leg stamped for another volume is refused (-EEXIST). A leg another owner stamped in the view's epoch or a newer one belongs to a raid that a view at least as recent still counts, and is refused too (-EBUSY). Nothing is claimed or written. force_restamp overwrites such stamps only at the epoch named by expected_view_epoch (-ESTALE at any other). A leg stamped by the same incarnation, or at an older epoch, is free.
  • Which legs are configured. The generations of one lineage compare: an ejection advances the members that stay, in the superblock transaction that records it, and a create carries the generation it configures its members at. Only the legs at the highest generation, configured in their own superblock, are configured, at that generation. Every other leg (behind, a copy never completed, unstamped, or stamped before minor 2) stays out, in an empty slot the superblock records MISSING at the leg's own generation, for the control plane to rebuild. With no stamped leg holding a configured copy, the create is refused (-ENODATA). When no leg carries a stamp of the volume (blank legs, or legs stamped by an older image), every leg is configured, as before.
  • The medium can move under the create. Each member is read again as the raid claims it and must still be what the arbitration read (-ESTALE). A leg the survey cannot read refuses the create: read as blank, the one leg at the highest generation would be left out, never read again, and the raid would come up on the legs behind it.
  • Adds. An add into an online raid of a lineage checks the leg by the same rules and takes it as a new member, rebuilt like any other, where upstream refuses any leg another raid stamped. A raid reassembled from a stamped superblock keeps its view, and bdev_raid_get_bdevs reports lineage and view_epoch.
  • declared_slots. A raid can be created wider than its legs, the slots past them empty and sized like the members, for members added later. It comes online degraded, with no null bdevs added and removed to open the slots. A level that operates only whole refuses it.

Tests. bdev_raid_ut gets a raid1-level module and a superblock per leg, with a read error (transient or not) and a later superblock for a leg that changes between the survey and the claim. 14 tests: blank legs, the highest generation, legs without a configured stamp, another owner's legs (same and newer epoch), an older epoch and the create's own incarnation, another volume, restamping, no configured copy, a read error, a leg that changed (its generation, a stamp appearing), the parameters a create in a view needs, declared slots, the online add, and the reassembly. bdev_raid_sb_ut pins the stamp and the slot entries init_superblock writes, read back after a write. Each rule was taken out in turn and its tests went red. Every raid suite passes: bdev_raid 39, bdev_raid_sb 15, raid0 8, raid1 7, raid5f 8, concat 3.

0054: the raid's owner journals each view epoch on its members

Problem. 0053 refuses a create on legs another owner stamped in the create's view epoch or a newer one. A raid stamps its legs once, at its create: an owner that stays in the view while the view moves on holds them with an old epoch, and a create from a view the owner already left behind takes them.

Fix. bdev_raid_journal_view_epoch {name, view_epoch, expected_incarnation, lineage} stamps a new epoch on an online raid: the superblock's view_epoch, its owner (the raid's incarnation, which takes over the stamp of a raid it claimed after a reassembly), and the view_epoch of each member it serves, then writes the superblock to them. An owner that stays in the view journals each epoch, so a create from a view no more recent finds the legs held for as long as the owner serves.

  • Only the raid's incarnation journals (-ESTALE for another).
  • The epoch never goes back (-ESTALE). The same epoch is written again, so a journal whose write failed is retried as it was.
  • A raid created outside the view protocol joins its volume's view when the journal names the lineage (-EINVAL without one, -EEXIST for another).
  • A superblock an older image wrote is raised to minor 2, where the stamp is read.

Tests. bdev_raid_ut (5 tests): the stamp, each refusal, the join, a failed write, and, with superblock writes landing on the legs, a create by another incarnation refused at the journaled epoch and at the one before, then taking the legs at the next. Taken out in turn, the stamp, the monotonic epoch, the minor 2 upgrade and the join each turned their tests red. Every raid suite passes: bdev_raid 44, bdev_raid_sb 15, raid0 8, raid1 7, raid5f 8, concat 3.

0055: a copied member takes its source's standing

Problem. A control plane that repairs a member while no raid is assembled (a cold repair) copies another member's content onto it, superblock included. The copy then carries the source's superblock, where the copy's own entry says what the source knew of it: MISSING, at an older generation, since it was behind. The next create in a view (0053) reads that entry, leaves the copy out and rebuilds it in full, although it holds the source's content. With the source gone by then, the create is refused (-ENODATA) and the volume cannot come up, although a complete copy exists.

Fix. bdev_raid_restamp_copied_member {base_bdev, member_uuid, source_uuid, lineage} gives the copy's own entry the source's standing: CONFIGURED at the source's content generation and view epoch. The superblock's seq_number and crc are renewed, and the superblock is written back at the head of the bdev. No other entry and no stamp field changes. The pure function is raid_bdev_sb_restamp_copied_member() in bdev_raid_sb.c.

  • Nothing is written for a superblock without a stamp (-EINVAL), of another lineage (-EEXIST), without an entry for either member or naming the same one twice (-ENOENT), or whose source entry is not CONFIGURED (-ESTALE). Copied from a member that was itself behind, the copy would be behind.
  • A copy that already has that standing is not written again (restamped: false). The response carries the copy's content_generation.
  • Opening for write is the gate, as for bdev_raid_clear_superblock: a member of an assembled raid, or a namespace a subsystem exposes, is claimed and cannot be restamped.
  • A bdev with interleaved metadata is refused (-ENOTSUP).

The restamp claims that the copy holds what the source held when it was copied. The caller has to know that no raid wrote to the source since then. A raid that left the copy out and kept writing does not move the source's generation, since the copy was already behind.

Tests. bdev_raid_sb_ut (2 tests, in each of its 3 suites): the restamp (state, generation, epoch, sequence number, crc, a third member left as it was) and the no-op that follows it, then each refusal, with the superblock left as it was. With the state change and the crc renewal taken out, the restamp test went red in all three suites. bdev_raid_ut stubs the pure function. Every raid suite passes: bdev_raid 44, bdev_raid_sb 21, raid0 8, raid1 7, raid5f 8, concat 3.

0056: discovery leaves a stamped member to a create in view

Problem. A control plane that attaches a volume's legs before it fences, and creates the raid in a view after (0053), never got to arbitrate: SPDK's examine assembled the raid as soon as the legs appeared, from the superblock of the first member it read. On the lab cluster, at each republish, the raid came online on the first member's record of the others. It wrote its superblock and started rebuilding the second member 200 ms before the writer's reservations were acquired. The member it rebuilt in full was a copy whose own superblock recorded it current (0055).

Fix. Examine no longer assembles a raid from a member whose superblock is stamped for a lineage (minor 2 with a lineage). The member stays unclaimed, and a create in its view assembles it, from every member's own standing. A member without a stamp is still reassembled as upstream does. An add the control plane issues still reads the member's superblock (0053).

Tests. bdev_raid_ut (45 tests): a stamped member left unclaimed, with no raid created and its superblock read once. As the control, the same member below minor 2 has discovery create its raid. With the check taken out, the test went red. Every raid suite passes.

tier: a superblock read that fails is an error

bdev_tier_read_sb answered valid: false both for a disk that carries no valid superblock (-EILSEQ) and for a read that failed (-EIO, -ENOMEM). The superblock module tells the two apart, and the RPC folded them. A caller that reads valid: false as "no superblock" then lets an older generation win the reassembly when the read of the disk holding the newest one fails, and, when the reads fail on every band, lays a fresh layout over disks that hold data.

The RPC now answers an error for a failed read, with its errno, and valid: false only for a disk without a valid superblock. The verdict is tier_sb_read_answer() in vbdev_tier.h, pinned by test_tier_units (red with a failed read mapped to "none": 2 assertions). docs/RPC-CONTRACT.md says it.

Next on this branch

  • A raid1 range rebuild after a tier remap. bdev_raid_rebuild_ranges reads through the raid, so on a raid1 the read can come from the member whose clusters were just remapped without a copy: they read back without an error, and the write-back puts what they held on every mirror. A member that lost a tier band should leave the raid before the remap and come back through a seeded rebuild of the remapped ranges and the writes it missed, which needs bdev_raid_start_seeded_rebuild_from_epoch to take extra ranges.

…eries patches

The image configures --disable-unit-tests, and no patch had ever run the
upstream suites of the code it changes. A first attempt ran them in a
separate job that built SPDK a second time. This one runs them on the
image's own tree instead.

- Dockerfile, unit-tests stage: FROM the builder, it compiles the owed
  suites against the libraries the builder made, and runs them. The
  collector (release image) and the debug image both derive from it.
  The suites run in the debug build, on each arch: asserts are on there,
  as upstream runs them. The release build passes through at no cost. A
  failing suite fails its debug build, and with it ci-gate and the
  release publish, which needs every build of the matrix.
- images/spdk/unit-tests.map says which leaf suites of test/unit/ a
  patched path owes: longest prefix wins, and "-" says upstream has no
  suite. A patched .c under lib/, module/ or app/ that matches no line
  fails the stage. A patched file under test/unit/ adds its own suite.
- images/spdk/unit-tests.sh selects, builds (lib/ut, then each leaf
  directory) and runs them. It builds from the leaf because parent
  Makefiles skip some suites with a mere warning (blob.c needs CUnit
  2.1-3, which Fedora 43 ships). A missing binary fails, like a failing
  suite. The script is plain bash 3, so its selection can be checked
  anywhere; today it selects 18 suites.
- Layer order: the build dependencies (pip pins, pkgdep.sh, liburing
  and libaio) now come before the modules and patches. They depend on
  the upstream tree alone, so a patch change no longer rebuilds them.

patches/README.md, "Unit tests", documents it.

The first run is expected red. The separate job's runs already showed
that vbdev_lvol_ut no longer compiles against the series: 0005 makes
vbdev_lvol.c include vbdev_tier.h, and the unit test's Makefile has no
path to it. What the suites find is fixed in the patches that caused
it, in the commits that follow.
@rducom
rducom requested a review from a team as a code owner September 30, 2026 20:30
… stage reports every broken suite in one pass

vbdev_lvol_ut includes vbdev_lvol.c, which the series changed twice
without the unit test following:
- 0009 completes an ENOSPC with spdk_bdev_io_complete_nvme_status. The
  unit test gets a stub.
- 0041 includes vbdev_tier.h and registers a per-band usage provider on
  a bdev_tier composite. The unit test's Makefile gets the tier module on
  its include path. Its stubs give no composite, so no provider is
  registered: vbdev_tier_get_by_name returns NULL, and set/clear usage
  provider and spdk_bs_count_allocated_clusters_in_lba_range are stubs.

Both edits are folded into the patch that made them necessary (fixup
commits autosquashed in a series worktree, then scripts/patches.sh
regen). Only these two patch files change, and the 42 still apply.

unit-tests.sh now builds every owed suite even after one fails, and
runs the ones that built. A single pass reports every broken suite,
instead of stopping at the first one in the list.
The previous run failed three suites. Each fix is in the patch that broke
the suite.

- bdev_raid_ut did not link. Ten patches call functions the suite neither
  includes nor stubs, and the calls pulled objects of libspdk_json and
  libspdk_jsonrpc whose other definitions clash with the suite's own. Each
  patch now stubs what it calls: 0007 (heat), 0009 (NVMe status
  completion), 0013 (RPC audit, two json calls), 0015 (outcome registry),
  0017 (CBT epoch query, a json call), 0018 (verifying outcome), 0019 (CBT
  auto epoch), 0021 (envelopes), 0025 (superblock clear), 0033 (a json
  call).
- bdev_raid_ut then failed on 0014: a raid is refused without an
  incarnation, and the suite's delete requests carried an uninitialised
  expected_incarnation. The suite now creates under an incarnation and
  deletes unchecked. Two tests cover 0014: no creation without an
  identity, and a delete that expects another incarnation is refused.
- raid1_ut did not link: 0009 reads the NVMe status of a failed member
  write. The stub reports a device fault. A new test covers 0009's
  thin-exhausted member, which is kept and marks the raid I/O.
- rpc_ut failed on 0010: the listen chmods a socket file that the suite's
  stubbed bind never creates. The suite records the chmod and the umask
  the bind runs under, and checks both (0600, 077). A new test checks that
  a failed chmod refuses the listen.

All 18 suites pass in a local debug build of the unit-tests stage (arm64).
…tests only)

Patch 0044 starts as tests only, and the unit-tests stage is expected to
fail on blob_ut:

- blob_thin_prov_new_cluster_reads_zeroes: every free cluster is filled
  with 0xAA, as a device whose unmap does not zero leaves them. A write
  of one io_unit goes into a new cluster, and the rest of the cluster
  must read zeroes. Upstream returns the 0xAA.
- blob_thin_prov_full_cluster_write_is_not_cleared_first: a write that
  covers the whole cluster costs the payload and the metadata page(s),
  and nothing more. It passes before the fix too: it guards the fix's
  cost.
- blob_thin_prov_rw, blob_thin_prov_write_count_io and
  blob_thin_prov_rle count the cleared cluster in their write
  accounting.

Locally (debug, arm64): blob_ut fails 4 tests in each of its 5 suites.
The fix follows in the next commit.
A cluster that a write allocates in a blob backed by zeroes (a thin
blob with no parent, or a clone over a range its snapshot never
allocated) was inserted as the device left it. The write filled its own
range, and the rest of the cluster returned what was stored there
before. The free path counts on unmap to leave zeroes, which a SATA SSD
does not guarantee and a bdev without unmap support never provides.

When the new cluster is not copied from a parent,
bs_allocate_and_copy_cluster now clears it with write_zeroes before
inserting it. A write (or write_zeroes) that covers the whole cluster
skips the clear, so a rebuild or a copy does not write twice. A clear
that fails returns the cluster to the pool. The inflate path goes
through the same branch.

Locally (debug, arm64): blob_ut passes, 495 tests.
rducom added 17 commits October 1, 2026 04:58
…ndows they need (0045, tests only)

Three cases for bdev_raid_ut, red on the series as it stands:
- two seed ranges inside one window: 12 blocks to copy under one lock; the
  window is copied whole (128 blocks);
- one range starting mid-window across three 4-block windows: 9 blocks under
  three locks, the first at the range; the windows are copied whole (12);
- one range at the end of the raid: 4 blocks under one lock, at the range;
  the clean windows before it are each locked and skipped.
The stubbed range quiesce counts the windows the copy phase locks.
…dows they need (0045)

A window that a seed range touched was copied whole, and every clean window
was locked and skipped in turn. Measured on a three-node cluster under a
random-write load: 5 GiB re-copied for 395 MiB of dirty chunks (4794 ranges),
110 s. Copying only the ranges, one window per range part, still took 100 s:
the cost is the window cycle (quiesce, copy, channel update, unquiesce), not
the bytes.

- A seeded process jumps over what its ranges leave clean, with no window.
  Only a copy needs a window's quiesce (a write landing between the read of
  the source and the write of the target), and the seeded target is a member
  of every channel for writes: the routing by process offset sends a write to
  it on either side.
- A window copies every range part it holds, under its one lock, and ends
  where everything before it is copied or clean. Out of requests, it ends at
  the first block not submitted and the next window resumes there.

The verify phase and the copied-bytes accounting are unchanged. bdev_raid_ut:
the three cases of the previous commit pass, and every raid suite stays green
(bdev_raid 20, bdev_raid_sb 9, concat 3, raid0 8, raid1 5, raid5f 8).
A JSON-RPC request carries ~200 ranges (SPDK_JSONRPC_MAX_VALUES = 1024
parsed values, a 32 KiB receive buffer), so a caller had to fold a larger
delta to fit, and folding a scattered delta recopies every clean block in
between: measured on a three-node cluster under random writes, 4485 dirty
ranges (368 MiB) folded to 128 recopied 3.2 GiB of clean blocks, and the
seeded rebuild of a 5 GiB volume took as long as a full one (95-110 s).

bdev_raid_start_seeded_rebuild_from_epoch {name, base_bdev, cbt_bdev,
epoch_id, expected_incarnation, rebuild_token} seeds the same rebuild with
the dirty ranges of a frozen epoch, read in-process: the cbt module answers
vbdev_cbt_query_epoch_ranges (the ranges bdev_cbt_epoch_get_dirty_ranges
walks, up to the caller's cap; more is -E2BIG, since a truncated delta is a
wrong one, and the caller rebuilds in full). The two RPCs share their checks
(raid, incarnation, running process, write_only member, range bounds, token)
and their audit.

bdev_raid_ut fakes the query: 64 one-block ranges a block apart copy their
64 blocks, and a delta the query refuses starts nothing. Every raid suite
stays green (bdev_raid 22, bdev_raid_sb 9, concat 3, raid0 8, raid1 5,
raid5f 8).
… (0047, tests only)

Two cases, red on the series as it stands:
- raid1_ut: a 64-block window offered as the generic layer offers it (the
  rest of the window to the next request) is copied by four requests of at
  most the part, every part read before any completes, the reads tiling the
  window and spread evenly over the members in sync, and each part's read
  completing writes that part to the target. raid1 takes the window whole in
  one request, and on three members the second one serves no read;
- bdev_raid_ut: raid_bdev_process_request_max_part offers parts that are
  whole write units and all fit RAID_BDEV_PROCESS_MAX_QD requests, smaller
  than any window with room to split. It offers the window whole.
raid1_ut now inits a process request's raid_io for real (its target names
the raid) and records every I/O the module submits.
…(0047)

raid_bdev_process_request_max_part offers a window to the process's requests
in RAID_BDEV_PROCESS_MAX_QD parts, whole write units, and raid1 copies one
part per request. A window's parts are now in flight together: each part's
write starts as its own read completes, and the reads spread over the members
in sync by their outstanding blocks.

Copied as one request, a window was read whole from one member, then written
whole: read and write never overlapped, and the first member served every
read. Measured on a 3-node bench, a full rebuild of a 5 GiB mirror3 under a
verified write load ran at 36 MiB/s for 141 s, every byte read across the
link from a remote member while the nexus's local member served none.

raid5f, bound to its stripe, does not call it; the seeded copy already takes
what the module returns and offers the rest to the next request.
…s only)

raid1_ut, red on the series as it stands: four parts of a rebuild window, two
read back holding data and two holding nothing but zeroes. With a target that
takes write-zeroes, the data parts are written as data and the zero parts as
write-zeroes; with a target that cannot, every part is written as data. The
zero parts go out as data writes.

The window-in-parts test (0047) now gives its requests real buffers: the copy
reads a part back before choosing how to write it. The suite records the
write-zeroes it is sent, and the write-zeroes support of the members is a
knob of the suite.
A part the source read back as nothing but zeroes goes to the target as
write_zeroes (raid_bdev_write_zeroes_blocks, the data offset applied like the
other wrappers): the same blocks, read back the same, and no payload on the
wire. A full rebuild of a sparse volume no longer ships its never-written
clusters across the link — measured on a 3-node bench, about 3 of the 5 GiB of
a mirror3 full rebuild, all of it zero, all of it through the nexus node's
egress, which 0047 had made the bound.

Not when the target cannot take write-zeroes (the bdev layer would then send
a zero buffer anyway), nor when the blocks carry metadata. The check reads
the part already in memory, and a data part stops at its first nonzero word.
…ed (0049, tests only)

blob_ut, red on the series as it stands: write-zeroes over half of a cluster
a thin blob has not allocated allocates the cluster and writes to the device;
it is to allocate nothing, write nothing, and read back zeroes. The same test
pins what stays: write-zeroes over an allocated cluster holding data zeroes
what it covers, and on a clone, whose unallocated cluster reads its parent's
data, it allocates the cluster and zeroes it.
…d (0049)

On a blob backed by zeroes, an unallocated cluster already reads zeroes:
write-zeroes over it now completes at once, with no cluster allocated and
nothing written. Before, it allocated the cluster (cleared, since 0044) and
zeroed it. A full rebuild sends every never-written part of its source as
write-zeroes (0048), so a thin volume rebuilt in full took its whole size on
the target, and the zeroes it carried no longer on the wire still went to the
target's disk: measured on a 3-node bench, the SATA disk of the rebuilt
replica wrote 94 MiB/s through a 5 GiB rebuild with about 3 GiB never written.

An allocated cluster is still zeroed. A clone or any blob with another back
device reads its unallocated clusters from it, and keeps the allocation.
…er (0050, tests only)

subsystem_ut, red on the series as it stands. A Write Exclusive - All
Registrants reservation acquired by A, then A preempted by B: the
reservation stays under A's key, which no registrant holds, and the state
persisted for it fails to restore (-EINVAL). The same once B unregisters
after C registered. The restore test now refuses a missing key for a single
holder only, and takes it for a reservation held by all registrants, under
the key of its first registrant.
…r (0050)

Every registrant holds a Write Exclusive or Exclusive Access - All
Registrants reservation, so when the registrant that acquired it leaves,
preempted by another or unregistering, the reservation stays. It stayed
under the key of the registrant that left: the state persisted for it
named a reservation key no registrant held, and the restore refused it
(-EINVAL), failing the add of the namespace. After a writer handover, the
next restart of a target left its namespace out, and the target
unexported.

The reservation now stands under the key of the registrant that holds it
next. The restore takes a state persisted before this, under the key of its
first registrant; a single holder's key that no registrant holds is still
refused.
…sts only)

bdev_nvme_ut, red on the series as it stands. The suite now compiles
nvme_rpc.c in, so the I/O passthrough can be driven against a real
controller channel. With the channel's I/O qpair cleared, as bdev_nvme
clears it from a disconnect until the reconnect, the command is still
handed to the NVMe library (rc 0, the channel kept): in a real process
that NULL qpair is dereferenced. The test expects -ENXIO with the
channel given back, and the command to go through once connected again.
bdev_nvme_send_cmd with an I/O command took the controller channel's
qpair (bdev_nvme_get_io_qpair) without a check and handed it to
spdk_nvme_ctrlr_cmd_io_raw_with_md. bdev_nvme frees that qpair from a
disconnect until the reconnect (bdev_nvme_disconnected_qpair_cb), so a
reservation (reservations are I/O commands) sent to a member whose path
had just gone down dereferenced NULL and killed the process, every raid
it carried with it. The command is now refused with -ENXIO and the
channel given back; a command for which no channel can be had is
refused with -ENOMEM instead of reaching the assert-only accessor.
… tests only)

bdev_raid_ut, red on the series as it stands. A current member that
leaves a raid1, removed or failed, should open one epoch, on the cbt
stacked on the raid and named after the member's UUID. 0019 asks each
survivor's own name instead (31 calls here, none to the raid), a bdev
no cbt sits on, so no epoch ever opens and a member lost without notice
is always rebuilt in full.

The suite also pins five ejections that must open none: a member being
rebuilt in full (its superblock slot not configured), one behind the
generation its peers carry, a write-only joiner, a raid without a
superblock, and a failed member an earlier removal failed to take.
0019 opened the epoch at a member's ejection under each survivor's own
name. The cbt sits on the raid, not on its members: every call answered
-ENODEV, no epoch ever opened, and a member lost without notice was
always rebuilt in full.

The raid now asks once, for the raid's name and the member's UUID, the
identity the control plane derives from the replica it placed there.
vbdev_cbt_auto_epoch_open finds the cbt stacked on that bdev and opens
the epoch there.

An epoch holds the writes a member misses from its departure on, so it
opens only for a member that left current: its superblock slot
configured at the generation its peers carry, not a write-only joiner,
and not a failed member an earlier removal failed to take (a new
removal_failed flag, which unlike remove_error survives the retry).
The cbt also refuses while it carries any live epoch, frozen or
rebuilding included: a freeze empties the live bitmap, and an epoch
opened after one may start short of writes the member missed just
before it left. Each of those members is rebuilt in full, as before.
… epoch is a delta (bdev_cbt_rotate)

The live bitmap is cleared only by bdev_cbt_reset, which needs a moment
when every backend is known in sync. A member that leaves without notice
gives none, so the epoch the raid opens at its ejection took the whole
live bitmap: every write since the last reset. On a volume that went
days without an epoch, that is the whole device, and the "delta" rebuild
copied all of it (measured on a lab cluster: 81920 of 81920 chunks).

bdev_cbt_rotate, called periodically while no epoch is live, moves the
live bitmap to a previous-window bitmap (the window before is dropped)
and starts it again empty. vbdev_cbt_auto_epoch_open ORs the previous
window back before it opens, so its delta holds one to two windows of
writes before the departure and everything after. Two rotations are at
least CBT_ROTATE_MIN_INTERVAL_US (10 s) apart whoever asks: a write the
member missed a few milliseconds before its ejection is always in one
of the two bitmaps. A rotation is refused (-EBUSY) while an epoch is
live, since that epoch reads the live bitmap whole; a reset clears both.
…-use

fix: blob thin clusters, raid rebuild copies and ejection epoch, cbt history windows, nvmf reservations, bdev_nvme passthrough (0044–0052)
@rducom rducom changed the title ci(image): the image build runs the upstream unit tests of what the series patches Upstream unit tests in the image build; patches 0044–0052 and cbt history windows Oct 3, 2026
… superblocks (0053)

A raid create took every listed leg as a configured member. The control
plane had to clear the legs' superblocks first, since a leg stamped by
another raid is refused, and clearing erased the only evidence of which
leg holds the volume's last acknowledged writes: a leg left behind by an
ejection came back as a configured member, and a creator that lost its
place could take legs from a raid still serving.

bdev_raid_create with a view_epoch (a raid1 with a superblock, the
volume's lineage and the creating incarnation) now reads every listed
leg's superblock, opened read-only, before it claims any:

- A superblock carries a stamp, minor 2, carved out of the header's
  reserved bytes: the lineage, the owner incarnation and the last view
  epoch it stamped. A minor-1 superblock reads as unstamped.
- A leg stamped for another volume is refused (-EEXIST). A leg another
  owner stamped in the view's epoch or a newer one is refused (-EBUSY),
  nothing claimed or written, unless force_restamp names that exact
  epoch (expected_view_epoch; -ESTALE for any other).
- Among the legs stamped for the volume, only those at the highest
  content generation, configured in their own superblock, are
  configured, at that generation. Every other leg (behind, a copy never
  completed, unstamped, or stamped before minor 2) stays out, in an
  empty slot where the superblock records it MISSING at its own
  generation, and the control plane rebuilds it. No stamped leg holding
  a configured copy is -ENODATA. When no leg carries a stamp of the
  volume, every leg is configured, as before.
- Each member is read again as the raid claims it and must still be
  what the arbitration read (-ESTALE). A leg the survey cannot read
  refuses the create: read as blank, the one leg at the highest
  generation would be left out, never read again, and the raid would
  come up on the legs behind it.
- An add into an online raid of a lineage checks the leg the same way
  and takes it as a new member, rebuilt like any other, where upstream
  refuses any leg another raid stamped.
- A raid reassembled from a stamped superblock keeps its view, and
  get_bdevs reports lineage and view_epoch.

declared_slots creates a raid wider than its legs, the slots past them
empty and sized like the members, for members added later: the raid
comes online degraded instead of through null bdevs added and removed.
A level that operates only whole refuses it.

bdev_raid_ut gets a raid1-level module and a superblock per leg, and
pins the arbitration, each refusal, the claim-time re-read, the empty
slots, the online add and the reassembly (14 tests); bdev_raid_sb_ut
pins the stamp and the slots init_superblock writes. Each rule was taken
out in turn and its tests went red: the generation (3 tests), the
admission (3), the claim-time re-read (1), the online add (1), the empty
slot size (3, and an online add hits the data_size assert), the
reassembly (1), the stamp and the MISSING slots (bdev_raid_sb_ut), a
read error read as blank (1), the view parameters (1). Every raid suite
passes: bdev_raid 39, bdev_raid_sb 15, raid0 8, raid1 7, raid5f 8,
concat 3.
@rducom rducom changed the title Upstream unit tests in the image build; patches 0044–0052 and cbt history windows Upstream unit tests in the image build; patches 0044–0053 and cbt history windows Oct 3, 2026
rducom added 2 commits October 3, 2026 14:28
…(0054)

0053 refuses a create on legs another owner stamped in the create's view
epoch or a newer one. A raid stamps its legs once, at its create, so an
owner that stays in the view while the view moves on held them with an
old epoch, and a create from a view the owner already left behind took
them.

bdev_raid_journal_view_epoch {name, view_epoch, expected_incarnation,
lineage} stamps a new epoch on an online raid: the superblock's
view_epoch, its owner (the raid's incarnation, which takes over the
stamp of a raid it claimed after a reassembly), and the view_epoch of
each member it serves, then writes the superblock to them. An owner
that stays in the view journals each epoch, so a create from a view no
more recent finds the legs held for as long as the owner serves.

- Only the raid's incarnation journals (-ESTALE for another).
- The epoch never goes back (-ESTALE). The same epoch is written again,
  so a journal whose write failed is retried as it was.
- A raid created outside the view protocol joins its volume's view when
  the journal names the lineage (-EINVAL without one, -EEXIST for
  another lineage).
- A superblock an older image wrote is raised to minor 2, where the
  stamp is read.

bdev_raid_ut pins the stamp, each refusal, the join and a failed write.
With superblock writes landing on the legs, a create by another
incarnation is refused at the journaled epoch and at the one before,
and takes the legs at the next. Taken out in turn, the stamp (3 tests,
the end-to-end one included), the monotonic epoch (1), the minor 2
upgrade (2) and the join (1) each turned their tests red. Every raid
suite passes: bdev_raid 44, bdev_raid_sb 15, raid0 8, raid1 7, raid5f 8,
concat 3.
…ut a superblock

bdev_tier_read_sb answered valid:false both for a disk that carries no
valid superblock (-EILSEQ) and for a read that failed (-EIO, -ENOMEM).
The superblock module tells the two apart; the RPC folded them. A
caller that reads valid:false as "no superblock" then lets an older
generation win the reassembly when the read of the disk holding the
newest one fails, and, when the reads fail on every band, lays a fresh
layout over disks that hold data.

The RPC now answers an error for a failed read, with its errno, and
valid:false only for a disk without a valid superblock. The verdict is
tier_sb_read_answer() in vbdev_tier.h, which test_tier_units pins: a
failed read, and a read with neither a superblock nor an error, is
FAILED. With a failed read mapped to "none" the test goes red (2
assertions). The tier and lvol modules build against the series.
@rducom rducom changed the title Upstream unit tests in the image build; patches 0044–0053 and cbt history windows Upstream unit tests in the image build; patches 0044–0054, cbt history windows, tier read errors Oct 3, 2026
rducom added 2 commits October 4, 2026 14:28
A member copied from another member of its raid while no raid was assembled carries the source's superblock, where its own entry says what the source knew of it: MISSING when it was behind. The next create in a view leaves it out and rebuilds it in full, and with the source gone refuses the volume although the copy holds the content. bdev_raid_restamp_copied_member gives its own entry the source's standing: CONFIGURED at the source's content generation and view epoch.
Examine assembled a raid of a lineage from whichever member it read first: the raid came online on that member's record of the others, wrote its superblock and rebuilt the others in full before the control plane's fence, and the members' own standings were never arbitrated. A member stamped for a lineage now stays unclaimed until a create in its view assembles it; an unstamped one is reassembled as upstream does.
@rducom
rducom merged commit 0c8213f into main Oct 6, 2026
12 checks passed
@rducom
rducom deleted the unit-tests-in-the-image-build branch October 6, 2026 13:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant