Repository navigation
fix: blob thin clusters, raid rebuild copies and ejection epoch, cbt history windows, nvmf reservations, bdev_nvme passthrough (0044–0052) - #26
Merged
rducom merged 18 commits intoOct 3, 2026
Conversation
…tests only) Patch 0044 starts as tests only, and the unit-tests stage is expected to fail on blob_ut: - blob_thin_prov_new_cluster_reads_zeroes: every free cluster is filled with 0xAA, as a device whose unmap does not zero leaves them. A write of one io_unit goes into a new cluster, and the rest of the cluster must read zeroes. Upstream returns the 0xAA. - blob_thin_prov_full_cluster_write_is_not_cleared_first: a write that covers the whole cluster costs the payload and the metadata page(s), and nothing more. It passes before the fix too: it guards the fix's cost. - blob_thin_prov_rw, blob_thin_prov_write_count_io and blob_thin_prov_rle count the cleared cluster in their write accounting. Locally (debug, arm64): blob_ut fails 4 tests in each of its 5 suites. The fix follows in the next commit.
A cluster that a write allocates in a blob backed by zeroes (a thin blob with no parent, or a clone over a range its snapshot never allocated) was inserted as the device left it. The write filled its own range, and the rest of the cluster returned what was stored there before. The free path counts on unmap to leave zeroes, which a SATA SSD does not guarantee and a bdev without unmap support never provides. When the new cluster is not copied from a parent, bs_allocate_and_copy_cluster now clears it with write_zeroes before inserting it. A write (or write_zeroes) that covers the whole cluster skips the clear, so a rebuild or a copy does not write twice. A clear that fails returns the cluster to the pool. The inflate path goes through the same branch. Locally (debug, arm64): blob_ut passes, 495 tests.
rducom
force-pushed
the
blob-thin-cluster-cleared-before-use
branch
from
September 30, 2026 21:20
c377fee to
e539f7e
Compare
rducom
marked this pull request as ready for review
September 30, 2026 21:33
…ndows they need (0045, tests only) Three cases for bdev_raid_ut, red on the series as it stands: - two seed ranges inside one window: 12 blocks to copy under one lock; the window is copied whole (128 blocks); - one range starting mid-window across three 4-block windows: 9 blocks under three locks, the first at the range; the windows are copied whole (12); - one range at the end of the raid: 4 blocks under one lock, at the range; the clean windows before it are each locked and skipped. The stubbed range quiesce counts the windows the copy phase locks.
…dows they need (0045) A window that a seed range touched was copied whole, and every clean window was locked and skipped in turn. Measured on a three-node cluster under a random-write load: 5 GiB re-copied for 395 MiB of dirty chunks (4794 ranges), 110 s. Copying only the ranges, one window per range part, still took 100 s: the cost is the window cycle (quiesce, copy, channel update, unquiesce), not the bytes. - A seeded process jumps over what its ranges leave clean, with no window. Only a copy needs a window's quiesce (a write landing between the read of the source and the write of the target), and the seeded target is a member of every channel for writes: the routing by process offset sends a write to it on either side. - A window copies every range part it holds, under its one lock, and ends where everything before it is copied or clean. Out of requests, it ends at the first block not submitted and the next window resumes there. The verify phase and the copied-bytes accounting are unchanged. bdev_raid_ut: the three cases of the previous commit pass, and every raid suite stays green (bdev_raid 20, bdev_raid_sb 9, concat 3, raid0 8, raid1 5, raid5f 8).
A JSON-RPC request carries ~200 ranges (SPDK_JSONRPC_MAX_VALUES = 1024
parsed values, a 32 KiB receive buffer), so a caller had to fold a larger
delta to fit, and folding a scattered delta recopies every clean block in
between: measured on a three-node cluster under random writes, 4485 dirty
ranges (368 MiB) folded to 128 recopied 3.2 GiB of clean blocks, and the
seeded rebuild of a 5 GiB volume took as long as a full one (95-110 s).
bdev_raid_start_seeded_rebuild_from_epoch {name, base_bdev, cbt_bdev,
epoch_id, expected_incarnation, rebuild_token} seeds the same rebuild with
the dirty ranges of a frozen epoch, read in-process: the cbt module answers
vbdev_cbt_query_epoch_ranges (the ranges bdev_cbt_epoch_get_dirty_ranges
walks, up to the caller's cap; more is -E2BIG, since a truncated delta is a
wrong one, and the caller rebuilds in full). The two RPCs share their checks
(raid, incarnation, running process, write_only member, range bounds, token)
and their audit.
bdev_raid_ut fakes the query: 64 one-block ranges a block apart copy their
64 blocks, and a delta the query refuses starts nothing. Every raid suite
stays green (bdev_raid 22, bdev_raid_sb 9, concat 3, raid0 8, raid1 5,
raid5f 8).
… (0047, tests only) Two cases, red on the series as it stands: - raid1_ut: a 64-block window offered as the generic layer offers it (the rest of the window to the next request) is copied by four requests of at most the part, every part read before any completes, the reads tiling the window and spread evenly over the members in sync, and each part's read completing writes that part to the target. raid1 takes the window whole in one request, and on three members the second one serves no read; - bdev_raid_ut: raid_bdev_process_request_max_part offers parts that are whole write units and all fit RAID_BDEV_PROCESS_MAX_QD requests, smaller than any window with room to split. It offers the window whole. raid1_ut now inits a process request's raid_io for real (its target names the raid) and records every I/O the module submits.
…(0047) raid_bdev_process_request_max_part offers a window to the process's requests in RAID_BDEV_PROCESS_MAX_QD parts, whole write units, and raid1 copies one part per request. A window's parts are now in flight together: each part's write starts as its own read completes, and the reads spread over the members in sync by their outstanding blocks. Copied as one request, a window was read whole from one member, then written whole: read and write never overlapped, and the first member served every read. Measured on a 3-node bench, a full rebuild of a 5 GiB mirror3 under a verified write load ran at 36 MiB/s for 141 s, every byte read across the link from a remote member while the nexus's local member served none. raid5f, bound to its stripe, does not call it; the seeded copy already takes what the module returns and offers the rest to the next request.
…s only) raid1_ut, red on the series as it stands: four parts of a rebuild window, two read back holding data and two holding nothing but zeroes. With a target that takes write-zeroes, the data parts are written as data and the zero parts as write-zeroes; with a target that cannot, every part is written as data. The zero parts go out as data writes. The window-in-parts test (0047) now gives its requests real buffers: the copy reads a part back before choosing how to write it. The suite records the write-zeroes it is sent, and the write-zeroes support of the members is a knob of the suite.
A part the source read back as nothing but zeroes goes to the target as write_zeroes (raid_bdev_write_zeroes_blocks, the data offset applied like the other wrappers): the same blocks, read back the same, and no payload on the wire. A full rebuild of a sparse volume no longer ships its never-written clusters across the link — measured on a 3-node bench, about 3 of the 5 GiB of a mirror3 full rebuild, all of it zero, all of it through the nexus node's egress, which 0047 had made the bound. Not when the target cannot take write-zeroes (the bdev layer would then send a zero buffer anyway), nor when the blocks carry metadata. The check reads the part already in memory, and a data part stops at its first nonzero word.
…ed (0049, tests only) blob_ut, red on the series as it stands: write-zeroes over half of a cluster a thin blob has not allocated allocates the cluster and writes to the device; it is to allocate nothing, write nothing, and read back zeroes. The same test pins what stays: write-zeroes over an allocated cluster holding data zeroes what it covers, and on a clone, whose unallocated cluster reads its parent's data, it allocates the cluster and zeroes it.
…d (0049) On a blob backed by zeroes, an unallocated cluster already reads zeroes: write-zeroes over it now completes at once, with no cluster allocated and nothing written. Before, it allocated the cluster (cleared, since 0044) and zeroed it. A full rebuild sends every never-written part of its source as write-zeroes (0048), so a thin volume rebuilt in full took its whole size on the target, and the zeroes it carried no longer on the wire still went to the target's disk: measured on a 3-node bench, the SATA disk of the rebuilt replica wrote 94 MiB/s through a 5 GiB rebuild with about 3 GiB never written. An allocated cluster is still zeroed. A clone or any blob with another back device reads its unallocated clusters from it, and keeps the allocation.
…er (0050, tests only) subsystem_ut, red on the series as it stands. A Write Exclusive - All Registrants reservation acquired by A, then A preempted by B: the reservation stays under A's key, which no registrant holds, and the state persisted for it fails to restore (-EINVAL). The same once B unregisters after C registered. The restore test now refuses a missing key for a single holder only, and takes it for a reservation held by all registrants, under the key of its first registrant.
…r (0050) Every registrant holds a Write Exclusive or Exclusive Access - All Registrants reservation, so when the registrant that acquired it leaves, preempted by another or unregistering, the reservation stays. It stayed under the key of the registrant that left: the state persisted for it named a reservation key no registrant held, and the restore refused it (-EINVAL), failing the add of the namespace. After a writer handover, the next restart of a target left its namespace out, and the target unexported. The reservation now stands under the key of the registrant that holds it next. The restore takes a state persisted before this, under the key of its first registrant; a single holder's key that no registrant holds is still refused.
…sts only) bdev_nvme_ut, red on the series as it stands. The suite now compiles nvme_rpc.c in, so the I/O passthrough can be driven against a real controller channel. With the channel's I/O qpair cleared, as bdev_nvme clears it from a disconnect until the reconnect, the command is still handed to the NVMe library (rc 0, the channel kept): in a real process that NULL qpair is dereferenced. The test expects -ENXIO with the channel given back, and the command to go through once connected again.
bdev_nvme_send_cmd with an I/O command took the controller channel's qpair (bdev_nvme_get_io_qpair) without a check and handed it to spdk_nvme_ctrlr_cmd_io_raw_with_md. bdev_nvme frees that qpair from a disconnect until the reconnect (bdev_nvme_disconnected_qpair_cb), so a reservation (reservations are I/O commands) sent to a member whose path had just gone down dereferenced NULL and killed the process, every raid it carried with it. The command is now refused with -ENXIO and the channel given back; a command for which no channel can be had is refused with -ENOMEM instead of reaching the assert-only accessor.
… tests only) bdev_raid_ut, red on the series as it stands. A current member that leaves a raid1, removed or failed, should open one epoch, on the cbt stacked on the raid and named after the member's UUID. 0019 asks each survivor's own name instead (31 calls here, none to the raid), a bdev no cbt sits on, so no epoch ever opens and a member lost without notice is always rebuilt in full. The suite also pins five ejections that must open none: a member being rebuilt in full (its superblock slot not configured), one behind the generation its peers carry, a write-only joiner, a raid without a superblock, and a failed member an earlier removal failed to take.
0019 opened the epoch at a member's ejection under each survivor's own name. The cbt sits on the raid, not on its members: every call answered -ENODEV, no epoch ever opened, and a member lost without notice was always rebuilt in full. The raid now asks once, for the raid's name and the member's UUID, the identity the control plane derives from the replica it placed there. vbdev_cbt_auto_epoch_open finds the cbt stacked on that bdev and opens the epoch there. An epoch holds the writes a member misses from its departure on, so it opens only for a member that left current: its superblock slot configured at the generation its peers carry, not a write-only joiner, and not a failed member an earlier removal failed to take (a new removal_failed flag, which unlike remove_error survives the retry). The cbt also refuses while it carries any live epoch, frozen or rebuilding included: a freeze empties the live bitmap, and an epoch opened after one may start short of writes the member missed just before it left. Each of those members is rebuilt in full, as before.
… epoch is a delta (bdev_cbt_rotate) The live bitmap is cleared only by bdev_cbt_reset, which needs a moment when every backend is known in sync. A member that leaves without notice gives none, so the epoch the raid opens at its ejection took the whole live bitmap: every write since the last reset. On a volume that went days without an epoch, that is the whole device, and the "delta" rebuild copied all of it (measured on a lab cluster: 81920 of 81920 chunks). bdev_cbt_rotate, called periodically while no epoch is live, moves the live bitmap to a previous-window bitmap (the window before is dropped) and starts it again empty. vbdev_cbt_auto_epoch_open ORs the previous window back before it opens, so its delta holds one to two windows of writes before the departure and everything after. Two rotations are at least CBT_ROTATE_MIN_INTERVAL_US (10 s) apart whoever asks: a write the member missed a few milliseconds before its ejection is always in one of the two bitmaps. A rotation is refused (-EBUSY) while an epoch is live, since that epoch reads the live bitmap whole; a reset clears both.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR carries nine patches of the series, 0044 to 0052, and one change to the cbt module (
bdev_cbt_rotate). Each patch comes as two commits: its tests alone, red on the series below it, then the fix. It is stacked on #27, whose image pipeline builds and runs the upstream unit suites the series touches.scripts/patches.sh checkpasses on a clean checkout at the pinned commit (51 patches), andregenfrom the series leaves every patch byte-identical to this branch.0044: a thin cluster is cleared before use
Problem
A cluster that a write allocates in a blob backed by zeroes (a thin blob with no parent, or a clone over a range its snapshot never allocated) is inserted as the device left it. The write fills its own range, and the rest of the cluster returns whatever was stored there before.
The free path counts on unmap to leave zeroes (see
spdk_free_cluster_unmap_complete: "so if the cluster is reclaimed in the future, it won't leak old data"). A SATA SSD does not guarantee that, and a bdev without unmap support never provides it. So:Remanence across blobs. The unwritten part of a new cluster returns another blob's data. A filesystem on top does not expose it, since it reads unwritten extents as zeroes, but anything that reads the raw device does.
Mirror legs differ in bytes nobody wrote. Each device keeps its own residue, and a raid1 verify over such lvols (
bdev_raid_verify_ranges, or any block compare of the legs) reports a content divergence.mkfs.xfsalone leaves such bytes:WRITE_ZEROESon an unallocated cluster allocates it too);On a lab cluster, a fresh 5 GiB three-leg mirror was reported divergent in 1198 blocks of 4 KiB right after
mkfs.xfs, with no application data, one leg in the minority everywhere.Fix
bs_allocate_and_copy_clusterclears it withwrite_zeroesbefore inserting it.write_zeroes) that covers the whole cluster skips the clear: it overwrites everything, and a rebuild or a copy, which write whole clusters, does not write twice.Cost: one extra cluster write (1 MiB by default) on the first partial write into a cluster, once per cluster lifetime.
Not covered: thick blobs get their clusters at creation and still count on the free path.
Tests
blob_ut:blob_thin_prov_new_cluster_reads_zeroes: every free cluster is filled with0xAA, then a one-io_unit write goes into a new cluster. The rest of the cluster reads zeroes.blob_thin_prov_full_cluster_write_is_not_cleared_first: a whole-cluster write costs the payload and the metadata page(s), and no clear.blob_thin_prov_rw,blob_thin_prov_write_count_ioandblob_thin_prov_rlenow count the cleared cluster in their write accounting.In the image pipeline:
FAILED: 1 of 18 suites: lib/blob/blob.c. Four tests fail in each of the 5 suites ofblob_ut:blob_thin_prov_new_cluster_reads_zeroesand the three accounting tests.blob_thin_prov_full_cluster_write_is_not_cleared_firstalso passes before the fix: it guards the fix's cost.A first version of this PR ran
blob_utin a separate job that built SPDK a second time. That job was dropped in favour of #27.0045 and 0046: a seeded rebuild copies its delta, and only its delta
A seeded rebuild (a member that missed some writes, with the dirty ranges known) copied every window a range touched in full, and still locked and skipped every clean window. Measured on a three-node lab cluster under a random-write load: 5 GiB re-copied for 395 MiB of dirty chunks (4794 ranges), in 110 s. Copying only the ranges, one window per range part, still took 100 s. The cost was the window cycle (quiesce, copy, channel update, unquiesce), not the bytes.
0045. A seeded process now jumps over what its ranges leave clean, with no window at all: only a copy needs a window's quiesce, and the seeded target already receives every write on either side of the process offset. A window copies every range part it holds under its one lock. The verify phase and the copied-bytes accounting are unchanged.
0046. A JSON-RPC request carries about 200 ranges, so a caller had to fold a larger delta to fit it, and folding a scattered delta recopies every clean block in between (4485 dirty ranges folded to 128 recopied 3.2 GiB of clean blocks: as slow as a full rebuild).
bdev_raid_start_seeded_rebuild_from_epochseeds the same rebuild with the dirty ranges of a frozen CBT epoch, read in-process throughvbdev_cbt_query_epoch_ranges. A delta larger than the caller's cap is refused (-E2BIG) rather than truncated, and the caller rebuilds in full. Both RPCs share their checks and their audit.Tests.
bdev_raid_ut: two seed ranges inside one window copy 12 blocks under one lock (the window was copied whole); a range across three windows copies 9 blocks under three locks; a range at the end of the raid takes one lock (the clean windows before it were each locked); 64 one-block ranges from a faked epoch copy their 64 blocks, and a delta the query refuses starts nothing. Every raid suite stays green.0047 and 0048: a full rebuild keeps the link busy, and only with data
0047. A rebuild window was copied as one request: read whole from one member, then written whole. Read and write never overlapped, and the first member served every read. Measured on a 3-node bench, a full rebuild of a 5 GiB three-leg mirror under a verified write load ran at 36 MiB/s for 141 s, every byte read across the link from a remote member while the local one sat idle.
raid_bdev_process_request_max_partnow offers a window inRAID_BDEV_PROCESS_MAX_QDparts of whole write units, and raid1 copies one part per request: each part's write starts as its own read completes, and the reads spread over the members in sync. raid5f, bound to its stripe, does not call it.0048. A part the source read back as nothing but zeroes goes to the target as write-zeroes: the same blocks, read back the same, and no payload on the wire. On the same bench, about 3 of the 5 GiB of a full rebuild were zeroes, all of them through the nexus node's egress. Not when the target cannot take write-zeroes, nor when the blocks carry metadata. The check reads the part already in memory and stops at the first nonzero word.
Tests.
raid1_ut: a 64-block window is copied by four requests in flight together, the reads tiling the window and spread evenly over the members (before: one request, and on three members the second served no read); four parts read back as data and zeroes go out as data and write-zeroes, or all as data when the target cannot take write-zeroes.bdev_raid_ut: the parts are whole write units and fit the queue depth.0049: write-zeroes leaves an unallocated thin cluster unallocated
On a blob backed by zeroes, an unallocated cluster already reads zeroes, so write-zeroes over it now completes at once: no cluster allocated, nothing written. Before, it allocated the cluster (cleared, since 0044) and zeroed it. Since 0048 sends every never-written part of a rebuild as write-zeroes, a thin volume rebuilt in full took its whole size on the target, and its zeroes still went to the target's disk (94 MiB/s of writes through a 5 GiB rebuild with about 3 GiB never written). An allocated cluster is still zeroed, and a clone, which reads its unallocated clusters from its parent, still allocates and zeroes.
Tests.
blob_ut: write-zeroes over half of an unallocated cluster of a thin blob allocates nothing, writes nothing and reads back zeroes; over an allocated cluster holding data it zeroes what it covers; on a clone it allocates the cluster and zeroes it.0050: a reservation held by all registrants outlives its acquirer
Every registrant holds a Write Exclusive or Exclusive Access – All Registrants reservation, so it stays when the registrant that acquired it leaves (preempted, or unregistering). It stayed under the key of the registrant that left: the persisted state named a reservation key no registrant held, and the restore refused it (
-EINVAL), failing the add of the namespace. After a writer handover, the next restart of a target left the namespace out, unexported. The reservation now stands under the key of the registrant that holds it next. The restore takes a state persisted before this fix under the key of its first registrant; a single holder's key that no registrant holds is still refused.Tests.
subsystem_ut: an All Registrants reservation acquired by A, then A preempted by B, keeps a key B holds and restores; the same once B unregisters after C registered. The restore test refuses a missing key for a single holder only.0051: an I/O passthrough needs a connected qpair
bdev_nvme_send_cmdwith an I/O command took the controller channel's qpair (bdev_nvme_get_io_qpair) without a check and handed it tospdk_nvme_ctrlr_cmd_io_raw_with_md. bdev_nvme frees that qpair from a disconnect until the reconnect (bdev_nvme_disconnected_qpair_cb), so a reservation command (reservations are I/O commands) sent to a member whose path had just gone down dereferenced NULL and killed the process, every raid it carried with it. The command is now refused with-ENXIOand the channel given back; a command for which no channel can be had is refused with-ENOMEM.Tests.
bdev_nvme_utnow compilesnvme_rpc.cin, so the passthrough is driven against a real controller channel. With the channel's I/O qpair cleared, as a disconnect clears it, the command was still handed to the NVMe library (rc 0, channel kept). It is now refused with-ENXIO, the channel given back, and it goes through once connected again. 55/55 with the fix.0052: a raid1 ejection opens its epoch on the raid's cbt
0019 opens a CBT epoch when a member leaves a raid1, so the member's return is a delta instead of a full rebuild. It asked each survivor's own bdev name for that epoch, but the cbt sits on the raid, not on its members: every call answered
-ENODEV, no epoch ever opened, and a member lost without notice was always rebuilt in full.The raid now asks once, with its own name and the member's UUID.
vbdev_cbt_auto_epoch_openfinds the cbt stacked on that bdev and opens the epoch there. The UUID is the identity a control plane can derive for the member: an NVMe-oF initiator bdev inherits the UUID of the namespace it attaches.An epoch holds the writes a member misses from its departure on, so two guards keep the delta whole:
removal_failedflag which, unlikeremove_error, survives the retry). A member being rebuilt in full, or seated behind, owes more than the writes it misses from now on.Every member these guards turn away is rebuilt in full, as before.
Tests.
bdev_raid_utdrives real ejections of a superblock raid (the suite's module stands in for raid1). A current member, removed or failed, opens one epoch, on the raid's name and under the member's UUID. Before the fix it opened 31, one per survivor name. Five ejections open none: a member being rebuilt in full, one behind its peers' generation, a write-only joiner, a raid without a superblock, and a failed member whose first removal failed (driven through a quiesce that fails once). Tests only: 2 of 25 red, 11 assertions. With the fix: 25/25, and every raid suite stays green. The cbt side has no suite in the image pipeline (the module's tests run against a standalone model).cbt: the live bitmap in rotating windows (
bdev_cbt_rotate)With 0052 the epoch at ejection opens, but on a lab cluster its delta resolved to 81920 of 81920 chunks: the live bitmap holds every write since its last
bdev_cbt_reset, and nothing can reset it at the ejection, since the writes the member just missed are in it. On a volume that goes days without an epoch, the "delta" is the whole device.bdev_cbt_rotate, called periodically while no epoch is live, moves the live bitmap to a previous-window bitmap (dropping the window before) and starts it again empty.vbdev_cbt_auto_epoch_openORs the previous window back before it opens, so its delta holds one to two windows of writes before the departure and everything after. Two rotations are at leastCBT_ROTATE_MIN_INTERVAL_US(10 s) apart whoever asks, so a write the member missed a few milliseconds before its ejection is always in one of the two bitmaps. A rotation is refused with-EBUSYwhile an epoch is open, frozen or rebuilding, since that epoch reads the live bitmap whole, and with-EAGAINwhen it comes too soon. A reset clears both bitmaps. The module README describes it.