Skip to content

fix: blob thin clusters, raid rebuild copies and ejection epoch, cbt history windows, nvmf reservations, bdev_nvme passthrough (0044–0052) - #26

Merged
rducom merged 18 commits into
unit-tests-in-the-image-buildfrom
blob-thin-cluster-cleared-before-use
Oct 3, 2026
Merged

rducom merged 18 commits into
unit-tests-in-the-image-buildfrom
blob-thin-cluster-cleared-before-use

Conversation

@rducom

@rducom rducom commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

This PR carries nine patches of the series, 0044 to 0052, and one change to the cbt module (bdev_cbt_rotate). Each patch comes as two commits: its tests alone, red on the series below it, then the fix. It is stacked on #27, whose image pipeline builds and runs the upstream unit suites the series touches.

Patch Area What it changes
0044 blob a thin cluster is cleared before use
0045 raid a seeded rebuild copies its ranges, and only under the windows they need
0046 raid a seeded rebuild reads its ranges from a frozen CBT epoch, in-process
0047 raid1 a rebuild copies its window in parts, all in flight together
0048 raid1 a rebuild writes an all-zero part as write-zeroes
0049 blob write-zeroes leaves an unallocated thin cluster unallocated
0050 nvmf a reservation held by all registrants outlives its acquirer
0051 bdev_nvme an I/O passthrough needs a connected qpair
0052 raid1, cbt a raid1 ejection opens its epoch on the raid's cbt

scripts/patches.sh check passes on a clean checkout at the pinned commit (51 patches), and regen from the series leaves every patch byte-identical to this branch.

0044: a thin cluster is cleared before use

Problem

A cluster that a write allocates in a blob backed by zeroes (a thin blob with no parent, or a clone over a range its snapshot never allocated) is inserted as the device left it. The write fills its own range, and the rest of the cluster returns whatever was stored there before.

The free path counts on unmap to leave zeroes (see spdk_free_cluster_unmap_complete: "so if the cluster is reclaimed in the future, it won't leak old data"). A SATA SSD does not guarantee that, and a bdev without unmap support never provides it. So:

  • Remanence across blobs. The unwritten part of a new cluster returns another blob's data. A filesystem on top does not expose it, since it reads unwritten extents as zeroes, but anything that reads the raw device does.

  • Mirror legs differ in bytes nobody wrote. Each device keeps its own residue, and a raid1 verify over such lvols (bdev_raid_verify_ranges, or any block compare of the legs) reports a content divergence. mkfs.xfs alone leaves such bytes:

    • the 2 KiB after the headers of each allocation group;
    • most of the last cluster, since it zeroes only the device's last 128 KiB (and WRITE_ZEROES on an unallocated cluster allocates it too);
    • the rest of the cluster at the end of its log.

    On a lab cluster, a fresh 5 GiB three-leg mirror was reported divergent in 1198 blocks of 4 KiB right after mkfs.xfs, with no application data, one leg in the minority everywhere.

Fix

  • When the new cluster is not copied from a parent, bs_allocate_and_copy_cluster clears it with write_zeroes before inserting it.
  • A write (or write_zeroes) that covers the whole cluster skips the clear: it overwrites everything, and a rebuild or a copy, which write whole clusters, does not write twice.
  • A clear that fails returns the cluster to the pool.
  • The inflate path, whose clusters over zero ranges were also handed out uninitialized, goes through the same branch, so an inflated cluster reads zeroes too.

Cost: one extra cluster write (1 MiB by default) on the first partial write into a cluster, once per cluster lifetime.

Not covered: thick blobs get their clusters at creation and still count on the free path.

Tests

blob_ut:

  • blob_thin_prov_new_cluster_reads_zeroes: every free cluster is filled with 0xAA, then a one-io_unit write goes into a new cluster. The rest of the cluster reads zeroes.
  • blob_thin_prov_full_cluster_write_is_not_cleared_first: a whole-cluster write costs the payload and the metadata page(s), and no clear.
  • blob_thin_prov_rw, blob_thin_prov_write_count_io and blob_thin_prov_rle now count the cleared cluster in their write accounting.

In the image pipeline:

  • Tests only (e539f7e): red on amd64 and arm64, FAILED: 1 of 18 suites: lib/blob/blob.c. Four tests fail in each of the 5 suites of blob_ut: blob_thin_prov_new_cluster_reads_zeroes and the three accounting tests. blob_thin_prov_full_cluster_write_is_not_cleared_first also passes before the fix: it guards the fix's cost.
  • With the fix (e94d862): green, all 18 suites pass on amd64 and arm64.

A first version of this PR ran blob_ut in a separate job that built SPDK a second time. That job was dropped in favour of #27.

0045 and 0046: a seeded rebuild copies its delta, and only its delta

A seeded rebuild (a member that missed some writes, with the dirty ranges known) copied every window a range touched in full, and still locked and skipped every clean window. Measured on a three-node lab cluster under a random-write load: 5 GiB re-copied for 395 MiB of dirty chunks (4794 ranges), in 110 s. Copying only the ranges, one window per range part, still took 100 s. The cost was the window cycle (quiesce, copy, channel update, unquiesce), not the bytes.

0045. A seeded process now jumps over what its ranges leave clean, with no window at all: only a copy needs a window's quiesce, and the seeded target already receives every write on either side of the process offset. A window copies every range part it holds under its one lock. The verify phase and the copied-bytes accounting are unchanged.

0046. A JSON-RPC request carries about 200 ranges, so a caller had to fold a larger delta to fit it, and folding a scattered delta recopies every clean block in between (4485 dirty ranges folded to 128 recopied 3.2 GiB of clean blocks: as slow as a full rebuild). bdev_raid_start_seeded_rebuild_from_epoch seeds the same rebuild with the dirty ranges of a frozen CBT epoch, read in-process through vbdev_cbt_query_epoch_ranges. A delta larger than the caller's cap is refused (-E2BIG) rather than truncated, and the caller rebuilds in full. Both RPCs share their checks and their audit.

Tests. bdev_raid_ut: two seed ranges inside one window copy 12 blocks under one lock (the window was copied whole); a range across three windows copies 9 blocks under three locks; a range at the end of the raid takes one lock (the clean windows before it were each locked); 64 one-block ranges from a faked epoch copy their 64 blocks, and a delta the query refuses starts nothing. Every raid suite stays green.

0047 and 0048: a full rebuild keeps the link busy, and only with data

0047. A rebuild window was copied as one request: read whole from one member, then written whole. Read and write never overlapped, and the first member served every read. Measured on a 3-node bench, a full rebuild of a 5 GiB three-leg mirror under a verified write load ran at 36 MiB/s for 141 s, every byte read across the link from a remote member while the local one sat idle. raid_bdev_process_request_max_part now offers a window in RAID_BDEV_PROCESS_MAX_QD parts of whole write units, and raid1 copies one part per request: each part's write starts as its own read completes, and the reads spread over the members in sync. raid5f, bound to its stripe, does not call it.

0048. A part the source read back as nothing but zeroes goes to the target as write-zeroes: the same blocks, read back the same, and no payload on the wire. On the same bench, about 3 of the 5 GiB of a full rebuild were zeroes, all of them through the nexus node's egress. Not when the target cannot take write-zeroes, nor when the blocks carry metadata. The check reads the part already in memory and stops at the first nonzero word.

Tests. raid1_ut: a 64-block window is copied by four requests in flight together, the reads tiling the window and spread evenly over the members (before: one request, and on three members the second served no read); four parts read back as data and zeroes go out as data and write-zeroes, or all as data when the target cannot take write-zeroes. bdev_raid_ut: the parts are whole write units and fit the queue depth.

0049: write-zeroes leaves an unallocated thin cluster unallocated

On a blob backed by zeroes, an unallocated cluster already reads zeroes, so write-zeroes over it now completes at once: no cluster allocated, nothing written. Before, it allocated the cluster (cleared, since 0044) and zeroed it. Since 0048 sends every never-written part of a rebuild as write-zeroes, a thin volume rebuilt in full took its whole size on the target, and its zeroes still went to the target's disk (94 MiB/s of writes through a 5 GiB rebuild with about 3 GiB never written). An allocated cluster is still zeroed, and a clone, which reads its unallocated clusters from its parent, still allocates and zeroes.

Tests. blob_ut: write-zeroes over half of an unallocated cluster of a thin blob allocates nothing, writes nothing and reads back zeroes; over an allocated cluster holding data it zeroes what it covers; on a clone it allocates the cluster and zeroes it.

0050: a reservation held by all registrants outlives its acquirer

Every registrant holds a Write Exclusive or Exclusive Access – All Registrants reservation, so it stays when the registrant that acquired it leaves (preempted, or unregistering). It stayed under the key of the registrant that left: the persisted state named a reservation key no registrant held, and the restore refused it (-EINVAL), failing the add of the namespace. After a writer handover, the next restart of a target left the namespace out, unexported. The reservation now stands under the key of the registrant that holds it next. The restore takes a state persisted before this fix under the key of its first registrant; a single holder's key that no registrant holds is still refused.

Tests. subsystem_ut: an All Registrants reservation acquired by A, then A preempted by B, keeps a key B holds and restores; the same once B unregisters after C registered. The restore test refuses a missing key for a single holder only.

0051: an I/O passthrough needs a connected qpair

bdev_nvme_send_cmd with an I/O command took the controller channel's qpair (bdev_nvme_get_io_qpair) without a check and handed it to spdk_nvme_ctrlr_cmd_io_raw_with_md. bdev_nvme frees that qpair from a disconnect until the reconnect (bdev_nvme_disconnected_qpair_cb), so a reservation command (reservations are I/O commands) sent to a member whose path had just gone down dereferenced NULL and killed the process, every raid it carried with it. The command is now refused with -ENXIO and the channel given back; a command for which no channel can be had is refused with -ENOMEM.

Tests. bdev_nvme_ut now compiles nvme_rpc.c in, so the passthrough is driven against a real controller channel. With the channel's I/O qpair cleared, as a disconnect clears it, the command was still handed to the NVMe library (rc 0, channel kept). It is now refused with -ENXIO, the channel given back, and it goes through once connected again. 55/55 with the fix.

0052: a raid1 ejection opens its epoch on the raid's cbt

0019 opens a CBT epoch when a member leaves a raid1, so the member's return is a delta instead of a full rebuild. It asked each survivor's own bdev name for that epoch, but the cbt sits on the raid, not on its members: every call answered -ENODEV, no epoch ever opened, and a member lost without notice was always rebuilt in full.

The raid now asks once, with its own name and the member's UUID. vbdev_cbt_auto_epoch_open finds the cbt stacked on that bdev and opens the epoch there. The UUID is the identity a control plane can derive for the member: an NVMe-oF initiator bdev inherits the UUID of the namespace it attaches.

An epoch holds the writes a member misses from its departure on, so two guards keep the delta whole:

  • The member left current. Its superblock slot is configured at the generation its peers carry, it is not a write-only joiner, and it is not a failed member that an earlier removal failed to take (a new removal_failed flag which, unlike remove_error, survives the retry). A member being rebuilt in full, or seated behind, owes more than the writes it misses from now on.
  • The cbt carries no live epoch. Open, frozen and rebuilding ones all refuse. An open one bounds a round its caller started. A freeze empties the live bitmap, so after one, an epoch opened now may start short of writes the member missed just before it left.

Every member these guards turn away is rebuilt in full, as before.

Tests. bdev_raid_ut drives real ejections of a superblock raid (the suite's module stands in for raid1). A current member, removed or failed, opens one epoch, on the raid's name and under the member's UUID. Before the fix it opened 31, one per survivor name. Five ejections open none: a member being rebuilt in full, one behind its peers' generation, a write-only joiner, a raid without a superblock, and a failed member whose first removal failed (driven through a quiesce that fails once). Tests only: 2 of 25 red, 11 assertions. With the fix: 25/25, and every raid suite stays green. The cbt side has no suite in the image pipeline (the module's tests run against a standalone model).

cbt: the live bitmap in rotating windows (bdev_cbt_rotate)

With 0052 the epoch at ejection opens, but on a lab cluster its delta resolved to 81920 of 81920 chunks: the live bitmap holds every write since its last bdev_cbt_reset, and nothing can reset it at the ejection, since the writes the member just missed are in it. On a volume that goes days without an epoch, the "delta" is the whole device.

bdev_cbt_rotate, called periodically while no epoch is live, moves the live bitmap to a previous-window bitmap (dropping the window before) and starts it again empty. vbdev_cbt_auto_epoch_open ORs the previous window back before it opens, so its delta holds one to two windows of writes before the departure and everything after. Two rotations are at least CBT_ROTATE_MIN_INTERVAL_US (10 s) apart whoever asks, so a write the member missed a few milliseconds before its ejection is always in one of the two bitmaps. A rotation is refused with -EBUSY while an epoch is open, frozen or rebuilding, since that epoch reads the live bitmap whole, and with -EAGAIN when it comes too soon. A reset clears both bitmaps. The module README describes it.

…tests only)

Patch 0044 starts as tests only, and the unit-tests stage is expected to
fail on blob_ut:

- blob_thin_prov_new_cluster_reads_zeroes: every free cluster is filled
  with 0xAA, as a device whose unmap does not zero leaves them. A write
  of one io_unit goes into a new cluster, and the rest of the cluster
  must read zeroes. Upstream returns the 0xAA.
- blob_thin_prov_full_cluster_write_is_not_cleared_first: a write that
  covers the whole cluster costs the payload and the metadata page(s),
  and nothing more. It passes before the fix too: it guards the fix's
  cost.
- blob_thin_prov_rw, blob_thin_prov_write_count_io and
  blob_thin_prov_rle count the cleared cluster in their write
  accounting.

Locally (debug, arm64): blob_ut fails 4 tests in each of its 5 suites.
The fix follows in the next commit.
A cluster that a write allocates in a blob backed by zeroes (a thin
blob with no parent, or a clone over a range its snapshot never
allocated) was inserted as the device left it. The write filled its own
range, and the rest of the cluster returned what was stored there
before. The free path counts on unmap to leave zeroes, which a SATA SSD
does not guarantee and a bdev without unmap support never provides.

When the new cluster is not copied from a parent,
bs_allocate_and_copy_cluster now clears it with write_zeroes before
inserting it. A write (or write_zeroes) that covers the whole cluster
skips the clear, so a rebuild or a copy does not write twice. A clear
that fails returns the cluster to the pool. The inflate path goes
through the same branch.

Locally (debug, arm64): blob_ut passes, 495 tests.
@rducom
rducom changed the base branch from main to unit-tests-in-the-image-build September 30, 2026 21:20
@rducom
rducom force-pushed the blob-thin-cluster-cleared-before-use branch from c377fee to e539f7e Compare September 30, 2026 21:20
@rducom
rducom marked this pull request as ready for review September 30, 2026 21:33
rducom added 13 commits October 1, 2026 04:58
…ndows they need (0045, tests only)

Three cases for bdev_raid_ut, red on the series as it stands:
- two seed ranges inside one window: 12 blocks to copy under one lock; the
  window is copied whole (128 blocks);
- one range starting mid-window across three 4-block windows: 9 blocks under
  three locks, the first at the range; the windows are copied whole (12);
- one range at the end of the raid: 4 blocks under one lock, at the range;
  the clean windows before it are each locked and skipped.
The stubbed range quiesce counts the windows the copy phase locks.
…dows they need (0045)

A window that a seed range touched was copied whole, and every clean window
was locked and skipped in turn. Measured on a three-node cluster under a
random-write load: 5 GiB re-copied for 395 MiB of dirty chunks (4794 ranges),
110 s. Copying only the ranges, one window per range part, still took 100 s:
the cost is the window cycle (quiesce, copy, channel update, unquiesce), not
the bytes.

- A seeded process jumps over what its ranges leave clean, with no window.
  Only a copy needs a window's quiesce (a write landing between the read of
  the source and the write of the target), and the seeded target is a member
  of every channel for writes: the routing by process offset sends a write to
  it on either side.
- A window copies every range part it holds, under its one lock, and ends
  where everything before it is copied or clean. Out of requests, it ends at
  the first block not submitted and the next window resumes there.

The verify phase and the copied-bytes accounting are unchanged. bdev_raid_ut:
the three cases of the previous commit pass, and every raid suite stays green
(bdev_raid 20, bdev_raid_sb 9, concat 3, raid0 8, raid1 5, raid5f 8).
A JSON-RPC request carries ~200 ranges (SPDK_JSONRPC_MAX_VALUES = 1024
parsed values, a 32 KiB receive buffer), so a caller had to fold a larger
delta to fit, and folding a scattered delta recopies every clean block in
between: measured on a three-node cluster under random writes, 4485 dirty
ranges (368 MiB) folded to 128 recopied 3.2 GiB of clean blocks, and the
seeded rebuild of a 5 GiB volume took as long as a full one (95-110 s).

bdev_raid_start_seeded_rebuild_from_epoch {name, base_bdev, cbt_bdev,
epoch_id, expected_incarnation, rebuild_token} seeds the same rebuild with
the dirty ranges of a frozen epoch, read in-process: the cbt module answers
vbdev_cbt_query_epoch_ranges (the ranges bdev_cbt_epoch_get_dirty_ranges
walks, up to the caller's cap; more is -E2BIG, since a truncated delta is a
wrong one, and the caller rebuilds in full). The two RPCs share their checks
(raid, incarnation, running process, write_only member, range bounds, token)
and their audit.

bdev_raid_ut fakes the query: 64 one-block ranges a block apart copy their
64 blocks, and a delta the query refuses starts nothing. Every raid suite
stays green (bdev_raid 22, bdev_raid_sb 9, concat 3, raid0 8, raid1 5,
raid5f 8).
… (0047, tests only)

Two cases, red on the series as it stands:
- raid1_ut: a 64-block window offered as the generic layer offers it (the
  rest of the window to the next request) is copied by four requests of at
  most the part, every part read before any completes, the reads tiling the
  window and spread evenly over the members in sync, and each part's read
  completing writes that part to the target. raid1 takes the window whole in
  one request, and on three members the second one serves no read;
- bdev_raid_ut: raid_bdev_process_request_max_part offers parts that are
  whole write units and all fit RAID_BDEV_PROCESS_MAX_QD requests, smaller
  than any window with room to split. It offers the window whole.
raid1_ut now inits a process request's raid_io for real (its target names
the raid) and records every I/O the module submits.
…(0047)

raid_bdev_process_request_max_part offers a window to the process's requests
in RAID_BDEV_PROCESS_MAX_QD parts, whole write units, and raid1 copies one
part per request. A window's parts are now in flight together: each part's
write starts as its own read completes, and the reads spread over the members
in sync by their outstanding blocks.

Copied as one request, a window was read whole from one member, then written
whole: read and write never overlapped, and the first member served every
read. Measured on a 3-node bench, a full rebuild of a 5 GiB mirror3 under a
verified write load ran at 36 MiB/s for 141 s, every byte read across the
link from a remote member while the nexus's local member served none.

raid5f, bound to its stripe, does not call it; the seeded copy already takes
what the module returns and offers the rest to the next request.
…s only)

raid1_ut, red on the series as it stands: four parts of a rebuild window, two
read back holding data and two holding nothing but zeroes. With a target that
takes write-zeroes, the data parts are written as data and the zero parts as
write-zeroes; with a target that cannot, every part is written as data. The
zero parts go out as data writes.

The window-in-parts test (0047) now gives its requests real buffers: the copy
reads a part back before choosing how to write it. The suite records the
write-zeroes it is sent, and the write-zeroes support of the members is a
knob of the suite.
A part the source read back as nothing but zeroes goes to the target as
write_zeroes (raid_bdev_write_zeroes_blocks, the data offset applied like the
other wrappers): the same blocks, read back the same, and no payload on the
wire. A full rebuild of a sparse volume no longer ships its never-written
clusters across the link — measured on a 3-node bench, about 3 of the 5 GiB of
a mirror3 full rebuild, all of it zero, all of it through the nexus node's
egress, which 0047 had made the bound.

Not when the target cannot take write-zeroes (the bdev layer would then send
a zero buffer anyway), nor when the blocks carry metadata. The check reads
the part already in memory, and a data part stops at its first nonzero word.
…ed (0049, tests only)

blob_ut, red on the series as it stands: write-zeroes over half of a cluster
a thin blob has not allocated allocates the cluster and writes to the device;
it is to allocate nothing, write nothing, and read back zeroes. The same test
pins what stays: write-zeroes over an allocated cluster holding data zeroes
what it covers, and on a clone, whose unallocated cluster reads its parent's
data, it allocates the cluster and zeroes it.
…d (0049)

On a blob backed by zeroes, an unallocated cluster already reads zeroes:
write-zeroes over it now completes at once, with no cluster allocated and
nothing written. Before, it allocated the cluster (cleared, since 0044) and
zeroed it. A full rebuild sends every never-written part of its source as
write-zeroes (0048), so a thin volume rebuilt in full took its whole size on
the target, and the zeroes it carried no longer on the wire still went to the
target's disk: measured on a 3-node bench, the SATA disk of the rebuilt
replica wrote 94 MiB/s through a 5 GiB rebuild with about 3 GiB never written.

An allocated cluster is still zeroed. A clone or any blob with another back
device reads its unallocated clusters from it, and keeps the allocation.
…er (0050, tests only)

subsystem_ut, red on the series as it stands. A Write Exclusive - All
Registrants reservation acquired by A, then A preempted by B: the
reservation stays under A's key, which no registrant holds, and the state
persisted for it fails to restore (-EINVAL). The same once B unregisters
after C registered. The restore test now refuses a missing key for a single
holder only, and takes it for a reservation held by all registrants, under
the key of its first registrant.
…r (0050)

Every registrant holds a Write Exclusive or Exclusive Access - All
Registrants reservation, so when the registrant that acquired it leaves,
preempted by another or unregistering, the reservation stays. It stayed
under the key of the registrant that left: the state persisted for it
named a reservation key no registrant held, and the restore refused it
(-EINVAL), failing the add of the namespace. After a writer handover, the
next restart of a target left its namespace out, and the target
unexported.

The reservation now stands under the key of the registrant that holds it
next. The restore takes a state persisted before this, under the key of its
first registrant; a single holder's key that no registrant holds is still
refused.
…sts only)

bdev_nvme_ut, red on the series as it stands. The suite now compiles
nvme_rpc.c in, so the I/O passthrough can be driven against a real
controller channel. With the channel's I/O qpair cleared, as bdev_nvme
clears it from a disconnect until the reconnect, the command is still
handed to the NVMe library (rc 0, the channel kept): in a real process
that NULL qpair is dereferenced. The test expects -ENXIO with the
channel given back, and the command to go through once connected again.
bdev_nvme_send_cmd with an I/O command took the controller channel's
qpair (bdev_nvme_get_io_qpair) without a check and handed it to
spdk_nvme_ctrlr_cmd_io_raw_with_md. bdev_nvme frees that qpair from a
disconnect until the reconnect (bdev_nvme_disconnected_qpair_cb), so a
reservation (reservations are I/O commands) sent to a member whose path
had just gone down dereferenced NULL and killed the process, every raid
it carried with it. The command is now refused with -ENXIO and the
channel given back; a command for which no channel can be had is
refused with -ENOMEM instead of reaching the assert-only accessor.
@rducom rducom changed the title fix(blob): a thin cluster is cleared before use (0044) fix: blob thin clusters, raid rebuild copies, nvmf reservations, bdev_nvme passthrough (0044–0051) Oct 2, 2026
rducom added 2 commits October 2, 2026 09:12
… tests only)

bdev_raid_ut, red on the series as it stands. A current member that
leaves a raid1, removed or failed, should open one epoch, on the cbt
stacked on the raid and named after the member's UUID. 0019 asks each
survivor's own name instead (31 calls here, none to the raid), a bdev
no cbt sits on, so no epoch ever opens and a member lost without notice
is always rebuilt in full.

The suite also pins five ejections that must open none: a member being
rebuilt in full (its superblock slot not configured), one behind the
generation its peers carry, a write-only joiner, a raid without a
superblock, and a failed member an earlier removal failed to take.
0019 opened the epoch at a member's ejection under each survivor's own
name. The cbt sits on the raid, not on its members: every call answered
-ENODEV, no epoch ever opened, and a member lost without notice was
always rebuilt in full.

The raid now asks once, for the raid's name and the member's UUID, the
identity the control plane derives from the replica it placed there.
vbdev_cbt_auto_epoch_open finds the cbt stacked on that bdev and opens
the epoch there.

An epoch holds the writes a member misses from its departure on, so it
opens only for a member that left current: its superblock slot
configured at the generation its peers carry, not a write-only joiner,
and not a failed member an earlier removal failed to take (a new
removal_failed flag, which unlike remove_error survives the retry).
The cbt also refuses while it carries any live epoch, frozen or
rebuilding included: a freeze empties the live bitmap, and an epoch
opened after one may start short of writes the member missed just
before it left. Each of those members is rebuilt in full, as before.
@rducom rducom changed the title fix: blob thin clusters, raid rebuild copies, nvmf reservations, bdev_nvme passthrough (0044–0051) fix: blob thin clusters, raid rebuild copies and ejection epoch, nvmf reservations, bdev_nvme passthrough (0044–0052) Oct 2, 2026
… epoch is a delta (bdev_cbt_rotate)

The live bitmap is cleared only by bdev_cbt_reset, which needs a moment
when every backend is known in sync. A member that leaves without notice
gives none, so the epoch the raid opens at its ejection took the whole
live bitmap: every write since the last reset. On a volume that went
days without an epoch, that is the whole device, and the "delta" rebuild
copied all of it (measured on a lab cluster: 81920 of 81920 chunks).

bdev_cbt_rotate, called periodically while no epoch is live, moves the
live bitmap to a previous-window bitmap (the window before is dropped)
and starts it again empty. vbdev_cbt_auto_epoch_open ORs the previous
window back before it opens, so its delta holds one to two windows of
writes before the departure and everything after. Two rotations are at
least CBT_ROTATE_MIN_INTERVAL_US (10 s) apart whoever asks: a write the
member missed a few milliseconds before its ejection is always in one
of the two bitmaps. A rotation is refused (-EBUSY) while an epoch is
live, since that epoch reads the live bitmap whole; a reset clears both.
@rducom rducom changed the title fix: blob thin clusters, raid rebuild copies and ejection epoch, nvmf reservations, bdev_nvme passthrough (0044–0052) fix: blob thin clusters, raid rebuild copies and ejection epoch, cbt history windows, nvmf reservations, bdev_nvme passthrough (0044–0052) Oct 2, 2026
@rducom
rducom merged commit 520e829 into unit-tests-in-the-image-build Oct 3, 2026
11 checks passed
@rducom
rducom deleted the blob-thin-cluster-cleared-before-use branch October 3, 2026 11:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant