From 2f34ba83ecfaf22a39c684e9933d6abd5f600d31 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Wed, 30 Sep 2026 22:30:07 +0200 Subject: [PATCH 01/26] ci(image): the image build runs the upstream unit tests of what the series patches The image configures --disable-unit-tests, and no patch had ever run the upstream suites of the code it changes. A first attempt ran them in a separate job that built SPDK a second time. This one runs them on the image's own tree instead. - Dockerfile, unit-tests stage: FROM the builder, it compiles the owed suites against the libraries the builder made, and runs them. The collector (release image) and the debug image both derive from it. The suites run in the debug build, on each arch: asserts are on there, as upstream runs them. The release build passes through at no cost. A failing suite fails its debug build, and with it ci-gate and the release publish, which needs every build of the matrix. - images/spdk/unit-tests.map says which leaf suites of test/unit/ a patched path owes: longest prefix wins, and "-" says upstream has no suite. A patched .c under lib/, module/ or app/ that matches no line fails the stage. A patched file under test/unit/ adds its own suite. - images/spdk/unit-tests.sh selects, builds (lib/ut, then each leaf directory) and runs them. It builds from the leaf because parent Makefiles skip some suites with a mere warning (blob.c needs CUnit 2.1-3, which Fedora 43 ships). A missing binary fails, like a failing suite. The script is plain bash 3, so its selection can be checked anywhere; today it selects 18 suites. - Layer order: the build dependencies (pip pins, pkgdep.sh, liburing and libaio) now come before the modules and patches. They depend on the upstream tree alone, so a patch change no longer rebuilds them. patches/README.md, "Unit tests", documents it. The first run is expected red. The separate job's runs already showed that vbdev_lvol_ut no longer compiles against the series: 0005 makes vbdev_lvol.c include vbdev_tier.h, and the unit test's Makefile has no path to it. What the suites find is fixed in the patches that caused it, in the commits that follow. --- images/spdk/Dockerfile | 73 +++++++++++++++------ images/spdk/unit-tests.map | 30 +++++++++ images/spdk/unit-tests.sh | 127 +++++++++++++++++++++++++++++++++++++ patches/README.md | 17 +++++ 4 files changed, 227 insertions(+), 20 deletions(-) create mode 100644 images/spdk/unit-tests.map create mode 100755 images/spdk/unit-tests.sh diff --git a/images/spdk/Dockerfile b/images/spdk/Dockerfile index 4d22cc2..4074d49 100644 --- a/images/spdk/Dockerfile +++ b/images/spdk/Dockerfile @@ -4,9 +4,11 @@ # ============================================================================ # # Multi-stage build: -# Stage 1 (builder) — Compiles SPDK from source on Fedora 43 -# Stage 2 (collector) — Collects exact runtime closure (binary + shared libs) -# Stage 3 (final) — Scratch-based, no shell, no package manager +# Stage 1 (builder) — Compiles SPDK from source on Fedora 43 +# Stage 1b (unit-tests) — Runs the upstream unit tests of what the series +# patches, on the tree built in stage 1 +# Stage 2 (collector) — Collects exact runtime closure (binary + shared libs) +# Stage 3 (final) — Scratch-based, no shell, no package manager # # Build: # docker build -t ghcr.io/evariops/spdk:v1.0.0 . @@ -59,18 +61,10 @@ RUN git clone --branch "${SPDK_VERSION}" --depth 1 --recurse-submodules \ WORKDIR /build/spdk -# ── Inject CBT + tier modules and patches ── -COPY module/bdev/cbt/ module/bdev/cbt/ -COPY module/bdev/tier/ module/bdev/tier/ -COPY patches/ /build/patches/ - -# Register CBT + tier in the bdev module build, link list, and deps & apply patches -# hadolint ignore=SC2016 -RUN sed -i '/^DIRS-y += delay/s/$/ cbt tier/' module/bdev/Makefile && \ - sed -i '/^BLOCKDEV_MODULES_LIST += bdev_zone_block/a BLOCKDEV_MODULES_LIST += bdev_cbt\nBLOCKDEV_MODULES_LIST += bdev_tier' mk/spdk.modules.mk && \ - sed -i '/^DEPDIRS-bdev_passthru/a DEPDIRS-bdev_cbt := $(BDEV_DEPS_THREAD)\nDEPDIRS-bdev_tier := $(BDEV_DEPS_THREAD)' mk/spdk.lib_deps.mk && \ - sed -i '/^DEPDIRS-bdev_lvol/s/$/ bdev_tier/' mk/spdk.lib_deps.mk && \ - for p in /build/patches/*.patch; do echo "Applying ${p}" && git apply "${p}" || exit 1; done +# ── Build dependencies — BEFORE the modules and patches ── +# They depend on the upstream tree alone (pkgdep.sh is upstream's), so a change +# to a patch or a module no longer invalidates these layers: only the apply, +# configure and make below run again. # Fedora's CMake rejects cmake_minimum_required(<3.5) in SPDK subprojects (ISA-L, etc.) ENV CMAKE_POLICY_VERSION_MINIMUM=3.5 @@ -88,6 +82,19 @@ RUN ./scripts/pkgdep.sh -d RUN dnf install -y --setopt=install_weak_deps=False --nodocs liburing-devel libaio-devel && \ dnf clean all && rm -rf /var/cache/dnf +# ── Inject CBT + tier modules and patches ── +COPY module/bdev/cbt/ module/bdev/cbt/ +COPY module/bdev/tier/ module/bdev/tier/ +COPY patches/ /build/patches/ + +# Register CBT + tier in the bdev module build, link list, and deps & apply patches +# hadolint ignore=SC2016 +RUN sed -i '/^DIRS-y += delay/s/$/ cbt tier/' module/bdev/Makefile && \ + sed -i '/^BLOCKDEV_MODULES_LIST += bdev_zone_block/a BLOCKDEV_MODULES_LIST += bdev_cbt\nBLOCKDEV_MODULES_LIST += bdev_tier' mk/spdk.modules.mk && \ + sed -i '/^DEPDIRS-bdev_passthru/a DEPDIRS-bdev_cbt := $(BDEV_DEPS_THREAD)\nDEPDIRS-bdev_tier := $(BDEV_DEPS_THREAD)' mk/spdk.lib_deps.mk && \ + sed -i '/^DEPDIRS-bdev_lvol/s/$/ bdev_tier/' mk/spdk.lib_deps.mk && \ + for p in /build/patches/*.patch; do echo "Applying ${p}" && git apply "${p}" || exit 1; done + # Configure: static SPDK libs + minimal module set. # # --without-shared All SPDK code linked statically into spdk_tgt. @@ -138,10 +145,35 @@ RUN if [ "${BUILD_TYPE}" != "debug" ]; then \ fi +# ============================================================================ +# Stage 1b: Upstream unit tests of what the series patches +# ============================================================================ +# On the tree just built, not a second build: the suites are compiled against +# the libraries above and run here. Both images derive from this stage (the +# collector below, and the debug image's COPY). The suites run in the DEBUG +# build (asserts on, the configuration upstream runs them in), on each arch; a +# failing suite fails that build, and the release publish needs every build of +# the matrix. The release build passes through at no cost. +# images/spdk/unit-tests.map says which suites a patched file owes; see +# patches/README.md ("Unit tests"). +FROM builder AS unit-tests + +SHELL ["/bin/bash", "-o", "pipefail", "-c"] + +ARG BUILD_TYPE=release + +COPY images/spdk/unit-tests.sh images/spdk/unit-tests.map /build/unit-tests/ +RUN if [ "${BUILD_TYPE}" = "debug" ]; then \ + /build/unit-tests/unit-tests.sh /build/patches /build/unit-tests/unit-tests.map; \ + else \ + echo "Unit tests run in the debug build of this pipeline (asserts on)."; \ + fi + + # ============================================================================ # Stage 2: Collect the exact runtime closure # ============================================================================ -FROM builder AS collector +FROM unit-tests AS collector SHELL ["/bin/bash", "-o", "pipefail", "-c"] @@ -221,16 +253,17 @@ RUN dnf install -y --setopt=install_weak_deps=False --nodocs \ jq \ && dnf clean all && rm -rf /var/cache/dnf -# Binary only (runtime .so provided by Fedora packages above) -COPY --from=builder /build/spdk/build/bin/spdk_tgt /usr/local/bin/spdk_tgt +# Binary only (runtime .so provided by Fedora packages above). Copied from the +# unit-tests stage, so the debug image too is built only once its suites pass. +COPY --from=unit-tests /build/spdk/build/bin/spdk_tgt /usr/local/bin/spdk_tgt # spdk_dd is the raid5f bench's full-stripe writer: raid5f only accepts full-stripe # writes, and remote initiators are MDTS-bounded, so the writer must run INSIDE the # target pod. Patch 0024 makes its exit code honest. -COPY --from=builder /build/spdk/build/bin/spdk_dd /usr/local/bin/spdk_dd +COPY --from=unit-tests /build/spdk/build/bin/spdk_dd /usr/local/bin/spdk_dd # SPDK RPC scripts (useful for live debugging: rpc.py, spdkcli, etc.) -COPY --from=builder /build/spdk/scripts/ /usr/local/share/spdk/scripts/ +COPY --from=unit-tests /build/spdk/scripts/ /usr/local/share/spdk/scripts/ LABEL org.opencontainers.image.title="SPDK Data Engine (debug)" \ org.opencontainers.image.description="Fedora-based SPDK debug image with gdb, strace, perf, and RPC scripts" \ diff --git a/images/spdk/unit-tests.map b/images/spdk/unit-tests.map new file mode 100644 index 0000000..edc6d6c --- /dev/null +++ b/images/spdk/unit-tests.map @@ -0,0 +1,30 @@ +# Which upstream unit-test suites a patched source file owes (see +# unit-tests.sh, and patches/README.md, "Unit tests"). +# +# One line: [ ...] +# - the prefix is matched against the files the series patches (the +# `+++ b/` lines of patches/*.patch); the LONGEST matching prefix wins; +# - a suite is a leaf directory under test/unit/ holding one *_ut.c, whose +# binary is that file without its .c; +# - "-" says upstream has no suite for it: said, not guessed. +# +# Only .c files under lib/, module/ and app/ are looked up. Headers, symbol +# maps and Makefiles are covered by the suites of the .c files beside them. A +# patched .c file that matches no line fails the stage: a component is mapped +# when its first patch lands. A file the series patches under test/unit/ adds +# its own suite, so a patch's own tests always run. + +lib/bdev/ lib/bdev/bdev.c lib/bdev/mt/bdev.c lib/bdev/part.c +lib/blob/ lib/blob/blob.c lib/blob/blob_bdev.c +lib/jsonrpc/ lib/jsonrpc/jsonrpc_server.c +lib/nvme/nvme_tcp.c lib/nvme/nvme_tcp.c +lib/nvmf/subsystem.c lib/nvmf/subsystem.c +lib/nvmf/nvmf_rpc.c - +lib/nvmf/nvmf_pause_rpc.c - +lib/rpc/ lib/rpc/rpc.c +lib/util/bit_array.c lib/util/bit_array.c +module/bdev/lvol/ lib/bdev/vbdev_lvol.c +module/bdev/nvme/ lib/bdev/nvme/bdev_nvme.c +module/bdev/raid/ lib/bdev/raid/bdev_raid.c lib/bdev/raid/bdev_raid_sb.c lib/bdev/raid/raid0.c lib/bdev/raid/raid1.c lib/bdev/raid/raid5f.c lib/bdev/raid/concat.c +module/bdev/uring/ - +app/spdk_dd/ - diff --git a/images/spdk/unit-tests.sh b/images/spdk/unit-tests.sh new file mode 100755 index 0000000..058aafc --- /dev/null +++ b/images/spdk/unit-tests.sh @@ -0,0 +1,127 @@ +#!/usr/bin/env bash +# SPDX-License-Identifier: BSD-3-Clause +# Copyright (c) 2026 Evariops. +# +# unit-tests.sh +# +# Run from the root of an SPDK tree that is already built (the image's builder +# stage). Builds and runs the upstream unit-test suites the patch series owes: +# unit-tests.map says which suites a patched file names. +# +# The suites are compiled against the libraries of that build: SPDK is not +# built a second time. Each suite is built from its own leaf directory, because +# the parent Makefiles skip some suites with a mere warning (blob.c wants +# CUnit 2.1-3), and a skip must not read as a pass. A suite whose binary is +# missing after its build fails the stage, as does a suite that fails. +set -euo pipefail + +patches_dir="${1:?usage: unit-tests.sh }" +map="${2:?usage: unit-tests.sh }" + +# Plain bash 3 on purpose (no mapfile, no associative arrays): the selection is +# checkable on any workstation, not only in the builder. + +# ── The files the series patches ── +touched=() +while IFS= read -r f; do + touched+=("${f}") +done < <(sed -n 's|^+++ b/||p' "${patches_dir}"/[0-9]*.patch | sort -u) + +# ── The map ── +prefixes=() +suites_of=() +while read -r prefix rest; do + [[ -z "${prefix}" || "${prefix}" == \#* ]] && continue + prefixes+=("${prefix}") + suites_of+=("${rest}") +done < "${map}" + +# ── Suites owed ── +owed=() +unmapped=() +for f in "${touched[@]}"; do + case "${f}" in + test/unit/*) + # A patch's own tests: the leaf directory of the file it touches. + owed+=("$(dirname "${f#test/unit/}")") + continue + ;; + lib/*.c | module/*.c | app/*.c) ;; + *) continue ;; + esac + + best=-1 + best_len=0 + for i in "${!prefixes[@]}"; do + p="${prefixes[$i]}" + if [[ "${f}" == "${p}"* && ${#p} -gt ${best_len} ]]; then + best=${i} + best_len=${#p} + fi + done + if ((best < 0)); then + unmapped+=("${f}") + continue + fi + for s in ${suites_of[$best]}; do + [[ "${s}" == "-" ]] || owed+=("${s}") + done +done + +if ((${#unmapped[@]} > 0)); then + echo "FATAL: patched file(s) with no line in unit-tests.map (map them; \"-\" when upstream has no suite):" >&2 + printf ' %s\n' "${unmapped[@]}" >&2 + exit 1 +fi + +suites=() +while IFS= read -r s; do + suites+=("${s}") +done < <(printf '%s\n' "${owed[@]}" | sort -u) +echo "Unit-test suites owed by the series (${#suites[@]}):" +printf ' %s\n' "${suites[@]}" + +# ── Build ── +# The ut library is built with the tests only (lib/Makefile), and the image +# configures them off. +make -C lib/ut -j"$(nproc)" +for s in "${suites[@]}"; do + if [[ ! -d "test/unit/${s}" ]]; then + echo "FATAL: unit-tests.map names test/unit/${s}, which does not exist" >&2 + exit 1 + fi + make -C "test/unit/${s}" -j"$(nproc)" +done + +# ── Run ── +logs="$(mktemp -d)" +failed=() +for s in "${suites[@]}"; do + src="$(find "test/unit/${s}" -maxdepth 1 -name '*_ut.c' | head -n 1)" + if [[ -z "${src}" ]]; then + echo "FATAL: test/unit/${s} holds no *_ut.c" >&2 + exit 1 + fi + bin="${src%.c}" + log="${logs}/${s//\//_}.log" + if [[ ! -x "${bin}" ]]; then + echo "FAIL ${s}: ${bin} was not built" + failed+=("${s}") + continue + fi + start=${SECONDS} + if "${bin}" > "${log}" 2>&1; then + echo "PASS ${s} ($((SECONDS - start)) s)" + grep -E '^\s+(suites|tests|asserts)\s' "${log}" | sed 's/^/ /' || true + else + echo "FAIL ${s} ($((SECONDS - start)) s)" + tail -n 80 "${log}" | sed 's/^/ /' + failed+=("${s}") + fi +done + +if ((${#failed[@]} > 0)); then + echo "FAILED: ${#failed[@]} of ${#suites[@]} suites: ${failed[*]}" >&2 + exit 1 +fi +echo "All ${#suites[@]} suites passed." diff --git a/patches/README.md b/patches/README.md index 2d95913..eb82251 100644 --- a/patches/README.md +++ b/patches/README.md @@ -94,6 +94,23 @@ scripts/patches.sh regen /path/to/spdk-worktree **Series format: RAW `git diff` output, no mail header.** `git apply` (Dockerfile + `check`/`apply`) is the only consumer; **`git am` is NOT part of the contract** — it chokes on raw diffs. `regen` emits raw per-commit diffs with the commit subject as filename, so a regen of an untouched series is byte-stable. Never hand-edit hunks. +## Unit tests + +The image build configures `--disable-unit-tests`, and until the `unit-tests` stage of `images/spdk/Dockerfile` existed, no patch had run the upstream suites of the code it changes. That stage runs them now: + +- **On the image's own tree.** It comes after the builder and compiles the suites against the libraries the builder made. SPDK is built once. The collector (release image) and the debug image both derive from it. +- **In the debug build, on each arch.** Asserts are on there, which is the configuration upstream runs its suites in. The release build passes through at no cost. A failing suite fails the debug build of the matrix, and with it `ci-gate` and the release publish, which needs every build. +- **The suites a patch owes.** `images/spdk/unit-tests.map` maps a path prefix to the leaf directories of `test/unit/` that cover it. The longest prefix wins, and `-` says upstream has no suite. Every `.c` under `lib/`, `module/` or `app/` that the series patches must match a line, or the stage fails: a component is mapped when its first patch lands. A patch that touches a file under `test/unit/` adds that suite, so a patch's own tests always run. +- **No silent skip.** Each suite is built from its leaf directory: parent Makefiles skip some suites with a warning (`blob.c` needs CUnit 2.1-3). A suite whose binary is missing fails, as does a suite that fails. + +A patch that changes behavior an upstream suite pins updates that suite in the same patch. A PR's build, and so its unit tests, runs with the `build-images` label. + +Locally, the same stage (`docker build` is enough; no registry): + +```sh +docker build -f images/spdk/Dockerfile --target unit-tests --build-arg BUILD_TYPE=debug . +``` + ## Upstreaming Candidates, easiest first: 0006 (degraded-read, small/general), 0003 (pause/resume), 0002 (allocated_ranges). 0004 (blob relocate) needs an RFC. Each patch merged upstream removes rebase surface here. From 155bd2dbd1f69943458379d40ee991f21a002fbf Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Wed, 30 Sep 2026 22:34:38 +0200 Subject: [PATCH 02/26] fix(0009, 0041): vbdev_lvol_ut compiles against the series again; the stage reports every broken suite in one pass vbdev_lvol_ut includes vbdev_lvol.c, which the series changed twice without the unit test following: - 0009 completes an ENOSPC with spdk_bdev_io_complete_nvme_status. The unit test gets a stub. - 0041 includes vbdev_tier.h and registers a per-band usage provider on a bdev_tier composite. The unit test's Makefile gets the tier module on its include path. Its stubs give no composite, so no provider is registered: vbdev_tier_get_by_name returns NULL, and set/clear usage provider and spdk_bs_count_allocated_clusters_in_lba_range are stubs. Both edits are folded into the patch that made them necessary (fixup commits autosquashed in a series worktree, then scripts/patches.sh regen). Only these two patch files change, and the 42 still apply. unit-tests.sh now builds every owed suite even after one fails, and runs the ones that built. A single pass reports every broken suite, instead of stopping at the first one in the list. --- images/spdk/unit-tests.sh | 21 +++++++++---- ...9-lvol-raid-enospc-capacity-exceeded.patch | 14 +++++++++ .../0041-lvol-tier-band-usage-provider.patch | 31 +++++++++++++++++++ 3 files changed, 60 insertions(+), 6 deletions(-) diff --git a/images/spdk/unit-tests.sh b/images/spdk/unit-tests.sh index 058aafc..8a978ec 100755 --- a/images/spdk/unit-tests.sh +++ b/images/spdk/unit-tests.sh @@ -83,19 +83,24 @@ printf ' %s\n' "${suites[@]}" # ── Build ── # The ut library is built with the tests only (lib/Makefile), and the image -# configures them off. +# configures them off. Every suite is built even after one fails, and so run +# below: one pass reports every broken suite, not the first one in the list. make -C lib/ut -j"$(nproc)" +logs="$(mktemp -d)" +failed=() for s in "${suites[@]}"; do if [[ ! -d "test/unit/${s}" ]]; then echo "FATAL: unit-tests.map names test/unit/${s}, which does not exist" >&2 exit 1 fi - make -C "test/unit/${s}" -j"$(nproc)" + if ! make -C "test/unit/${s}" -j"$(nproc)" > "${logs}/${s//\//_}.build.log" 2>&1; then + echo "FAIL ${s}: does not build" + { grep -E 'error|undefined reference' "${logs}/${s//\//_}.build.log" | head -n 40 | sed 's/^/ /'; } || true + failed+=("${s}") + fi done # ── Run ── -logs="$(mktemp -d)" -failed=() for s in "${suites[@]}"; do src="$(find "test/unit/${s}" -maxdepth 1 -name '*_ut.c' | head -n 1)" if [[ -z "${src}" ]]; then @@ -105,8 +110,12 @@ for s in "${suites[@]}"; do bin="${src%.c}" log="${logs}/${s//\//_}.log" if [[ ! -x "${bin}" ]]; then - echo "FAIL ${s}: ${bin} was not built" - failed+=("${s}") + # A build that failed is reported above; a build that "succeeded" + # without its binary (a skip) is reported here. + [[ " ${failed[*]} " == *" ${s} "* ]] || { + echo "FAIL ${s}: ${bin} was not built" + failed+=("${s}") + } continue fi start=${SECONDS} diff --git a/patches/0009-lvol-raid-enospc-capacity-exceeded.patch b/patches/0009-lvol-raid-enospc-capacity-exceeded.patch index 9959f86..79100e9 100644 --- a/patches/0009-lvol-raid-enospc-capacity-exceeded.patch +++ b/patches/0009-lvol-raid-enospc-capacity-exceeded.patch @@ -119,3 +119,17 @@ index 4b24829..4fc593f 100644 } } +diff --git a/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c b/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c +index d0e3cbb..339dd32 100644 +--- a/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c ++++ b/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c +@@ -42,6 +42,9 @@ bool g_bdev_is_missing = false; + + DEFINE_STUB_V(spdk_bdev_module_fini_start_done, (void)); + DEFINE_STUB_V(spdk_bdev_update_bs_blockcnt, (struct spdk_bs_dev *bs_dev)); ++/* Evariops 0009: an ENOSPC completes the I/O with an NVMe status. */ ++DEFINE_STUB_V(spdk_bdev_io_complete_nvme_status, (struct spdk_bdev_io *bdev_io, uint32_t cdw0, ++ int sct, int sc)); + DEFINE_STUB_V(spdk_lvs_grow_live, (struct spdk_lvol_store *lvs, + spdk_lvs_op_complete cb_fn, void *cb_arg)); + DEFINE_STUB(spdk_bdev_get_memory_domains, int, (struct spdk_bdev *bdev, diff --git a/patches/0041-lvol-tier-band-usage-provider.patch b/patches/0041-lvol-tier-band-usage-provider.patch index f6b7ab4..289d654 100644 --- a/patches/0041-lvol-tier-band-usage-provider.patch +++ b/patches/0041-lvol-tier-band-usage-provider.patch @@ -402,3 +402,34 @@ index 3673b97..bb9e82a 100644 SPDK_INFOLOG(vbdev_lvol, "Lvol store found on %s - begin parsing\n", req->base_bdev->name); +diff --git a/test/unit/lib/bdev/vbdev_lvol.c/Makefile b/test/unit/lib/bdev/vbdev_lvol.c/Makefile +index 5e8f2e1..42e4a31 100644 +--- a/test/unit/lib/bdev/vbdev_lvol.c/Makefile ++++ b/test/unit/lib/bdev/vbdev_lvol.c/Makefile +@@ -7,4 +7,7 @@ SPDK_ROOT_DIR := $(abspath $(CURDIR)/../../../../../) + + TEST_FILE = vbdev_lvol_ut.c + ++# Evariops 0041: vbdev_lvol.c includes vbdev_tier.h (the injected bdev_tier module). ++CFLAGS += -I$(SPDK_ROOT_DIR)/module/bdev/tier ++ + include $(SPDK_ROOT_DIR)/mk/spdk.unittest.mk +diff --git a/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c b/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c +index 339dd32..de98175 100644 +--- a/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c ++++ b/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c +@@ -60,6 +60,14 @@ DEFINE_STUB(spdk_lvs_esnap_missing_add, int, + DEFINE_STUB(spdk_blob_get_esnap_bs_dev, struct spdk_bs_dev *, (const struct spdk_blob *blob), NULL); + DEFINE_STUB(spdk_lvol_is_degraded, bool, (const struct spdk_lvol *lvol), false); + DEFINE_STUB(spdk_blob_get_num_allocated_clusters, uint64_t, (struct spdk_blob *blob), 0); ++/* Evariops 0041: an lvolstore on a bdev_tier composite registers a per-band usage provider. ++ * No composite here: the lookup finds none, so no provider is registered or counted. */ ++DEFINE_STUB(vbdev_tier_get_by_name, struct vbdev_tier *, (const char *name), NULL); ++DEFINE_STUB_V(vbdev_tier_set_usage_provider, (struct vbdev_tier *t, vbdev_tier_usage_fn fn, ++ void *ctx)); ++DEFINE_STUB_V(vbdev_tier_clear_usage_provider, (struct vbdev_tier *t)); ++DEFINE_STUB(spdk_bs_count_allocated_clusters_in_lba_range, uint64_t, ++ (struct spdk_blob_store *bs, uint64_t lba_start, uint64_t lba_count), 0); + + struct spdk_blob { + uint64_t id; From 0f74e3f5aaae745697d94364296c547934f7f246 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Wed, 30 Sep 2026 23:12:39 +0200 Subject: [PATCH 03/26] fix(patches): bdev_raid_ut, raid1_ut and rpc_ut pass against the series The previous run failed three suites. Each fix is in the patch that broke the suite. - bdev_raid_ut did not link. Ten patches call functions the suite neither includes nor stubs, and the calls pulled objects of libspdk_json and libspdk_jsonrpc whose other definitions clash with the suite's own. Each patch now stubs what it calls: 0007 (heat), 0009 (NVMe status completion), 0013 (RPC audit, two json calls), 0015 (outcome registry), 0017 (CBT epoch query, a json call), 0018 (verifying outcome), 0019 (CBT auto epoch), 0021 (envelopes), 0025 (superblock clear), 0033 (a json call). - bdev_raid_ut then failed on 0014: a raid is refused without an incarnation, and the suite's delete requests carried an uninitialised expected_incarnation. The suite now creates under an incarnation and deletes unchecked. Two tests cover 0014: no creation without an identity, and a delete that expects another incarnation is refused. - raid1_ut did not link: 0009 reads the NVMe status of a failed member write. The stub reports a device fault. A new test covers 0009's thin-exhausted member, which is kept and marks the raid I/O. - rpc_ut failed on 0010: the listen chmods a socket file that the suite's stubbed bind never creates. The suite records the chmod and the umask the bind runs under, and checks both (0600, 077). A new test checks that a failed chmod refuses the listen. All 18 suites pass in a local debug build of the unit-tests stage (arm64). --- patches/0007-raid-nexus-heat.patch | 15 +++ ...9-lvol-raid-enospc-capacity-exceeded.patch | 94 +++++++++++++ patches/0010-rpc-socket-chmod.patch | 123 +++++++++++++++++ patches/0013-raid1-seeded-rebuild.patch | 35 +++++ patches/0014-raid-incarnation-estale.patch | 124 ++++++++++++++++++ .../0015-raid-rebuild-outcome-registry.patch | 21 +++ patches/0017-raid-member-observation.patch | 25 ++++ patches/0018-raid-integrated-verify.patch | 13 ++ .../0019-raid-auto-epoch-at-ejection.patch | 15 +++ patches/0021-raid-envelopes.patch | 17 +++ patches/0025-raid-clear-superblock.patch | 18 +++ ...ev-inflight-distinct-failure-visible.patch | 14 ++ 12 files changed, 514 insertions(+) diff --git a/patches/0007-raid-nexus-heat.patch b/patches/0007-raid-nexus-heat.patch index d96921e..c85f1ff 100644 --- a/patches/0007-raid-nexus-heat.patch +++ b/patches/0007-raid-nexus-heat.patch @@ -387,3 +387,18 @@ index 0000000..2eda221 + free(req.name); +} +SPDK_RPC_REGISTER("vbdev_nexus_get_heat", rpc_vbdev_nexus_get_heat, SPDK_RPC_RUNTIME) +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index c4b478a..ce3af2f 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -200,6 +200,10 @@ DEFINE_STUB(spdk_bdev_writev_blocks_ext, int, (struct spdk_bdev_desc *desc, + struct spdk_io_channel *ch, struct iovec *iov, int iovcnt, uint64_t offset_blocks, + uint64_t num_blocks, spdk_bdev_io_completion_cb cb, void *cb_arg, + struct spdk_bdev_ext_io_opts *opts), 0); ++/* Evariops 0007: a raid records its I/O heat once heat is enabled, which no test ++ * here does, and destroys it with the raid. */ ++DEFINE_STUB_V(raid_tier_heat_record, (void *heat, uint64_t offset_blocks, bool is_write)); ++DEFINE_STUB_V(raid_tier_heat_destroy, (void *heat)); + + uint32_t + spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) diff --git a/patches/0009-lvol-raid-enospc-capacity-exceeded.patch b/patches/0009-lvol-raid-enospc-capacity-exceeded.patch index 79100e9..0d58eaf 100644 --- a/patches/0009-lvol-raid-enospc-capacity-exceeded.patch +++ b/patches/0009-lvol-raid-enospc-capacity-exceeded.patch @@ -119,6 +119,100 @@ index 4b24829..4fc593f 100644 } } +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index ce3af2f..d47d6c2 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -204,6 +204,10 @@ DEFINE_STUB(spdk_bdev_writev_blocks_ext, int, (struct spdk_bdev_desc *desc, + * here does, and destroys it with the raid. */ + DEFINE_STUB_V(raid_tier_heat_record, (void *heat, uint64_t offset_blocks, bool is_write)); + DEFINE_STUB_V(raid_tier_heat_destroy, (void *heat)); ++/* Evariops 0009: a raid I/O that a thin-exhausted member marked completes with an ++ * NVMe status. Only raid1.c marks one, and it is not part of this suite. */ ++DEFINE_STUB_V(spdk_bdev_io_complete_nvme_status, (struct spdk_bdev_io *bdev_io, uint32_t cdw0, ++ int sct, int sc)); + + uint32_t + spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) +diff --git a/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c b/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c +index 7e93c16..9c3eb8f 100644 +--- a/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c ++++ b/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c +@@ -68,6 +68,19 @@ raid_bdev_fail_base_bdev(struct raid_base_bdev_info *base_info) + base_info->is_failed = true; + } + ++/* Evariops 0009: a failed member write is told apart by its NVMe status: a device ++ * fault fails the member, thin-pool exhaustion keeps it. The writes this suite ++ * fails are device faults, unless a test sets another status. */ ++static int g_failed_write_sc = SPDK_NVME_SC_INTERNAL_DEVICE_ERROR; ++ ++void ++spdk_bdev_io_get_nvme_status(const struct spdk_bdev_io *bdev_io, uint32_t *cdw0, int *sct, int *sc) ++{ ++ *cdw0 = 0; ++ *sct = SPDK_NVME_SCT_GENERIC; ++ *sc = g_failed_write_sc; ++} ++ + static int + test_setup(void) + { +@@ -327,6 +340,46 @@ test_raid1_write_error(void) + run_for_each_raid1_config(_test_raid1_write_error); + } + ++static void ++_test_raid1_write_capacity_exceeded(struct raid_bdev *raid_bdev, ++ struct raid_bdev_io_channel *raid_ch) ++{ ++ struct raid1_info *r1_info = raid_bdev->module_private; ++ struct raid_bdev_io *raid_io; ++ struct raid_base_bdev_info *base_info; ++ struct spdk_bdev_io bdev_io = {}; ++ bool exhausted; ++ ++ /* Evariops 0009: the first member's thin pool is exhausted. The member is ++ * kept, and the raid I/O is marked, so that it completes with ++ * CAPACITY_EXCEEDED (bdev_raid.c, outside this suite). */ ++ g_failed_write_sc = SPDK_NVME_SC_CAPACITY_EXCEEDED; ++ g_io_status = SPDK_BDEV_IO_STATUS_PENDING; ++ raid_io = get_raid_io(r1_info, raid_ch, SPDK_BDEV_IO_TYPE_WRITE, 64); ++ raid1_submit_write_request(raid_io); ++ ++ RAID_FOR_EACH_BASE_BDEV(raid_bdev, base_info) { ++ exhausted = raid_bdev_base_bdev_slot(base_info) == 0; ++ base_info->is_failed = false; ++ bdev_io.bdev = base_info->desc->bdev; ++ raid1_write_bdev_io_completion(&bdev_io, !exhausted, raid_io); ++ CU_ASSERT(base_info->is_failed == false); ++ if (exhausted) { ++ /* Still in flight: the other members have not completed. */ ++ CU_ASSERT(raid_io->enospc == true); ++ } ++ } ++ CU_ASSERT(g_io_status != SPDK_BDEV_IO_STATUS_PENDING); ++ ++ g_failed_write_sc = SPDK_NVME_SC_INTERNAL_DEVICE_ERROR; ++} ++ ++static void ++test_raid1_write_capacity_exceeded(void) ++{ ++ run_for_each_raid1_config(_test_raid1_write_capacity_exceeded); ++} ++ + static void + _test_raid1_read_error(struct raid_bdev *raid_bdev, struct raid_bdev_io_channel *raid_ch) + { +@@ -520,6 +573,7 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_raid1_start); + CU_ADD_TEST(suite, test_raid1_read_balancing); + CU_ADD_TEST(suite, test_raid1_write_error); ++ CU_ADD_TEST(suite, test_raid1_write_capacity_exceeded); + CU_ADD_TEST(suite, test_raid1_read_error); + + allocate_threads(1); diff --git a/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c b/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c index d0e3cbb..339dd32 100644 --- a/test/unit/lib/bdev/vbdev_lvol.c/vbdev_lvol_ut.c diff --git a/patches/0010-rpc-socket-chmod.patch b/patches/0010-rpc-socket-chmod.patch index 8f86d14..c3b5ffd 100644 --- a/patches/0010-rpc-socket-chmod.patch +++ b/patches/0010-rpc-socket-chmod.patch @@ -57,3 +57,126 @@ index ead7c66..46446e2 100644 return server; ret: +diff --git a/test/unit/lib/rpc/rpc.c/Makefile b/test/unit/lib/rpc/rpc.c/Makefile +index 5801065..641a9e9 100644 +--- a/test/unit/lib/rpc/rpc.c/Makefile ++++ b/test/unit/lib/rpc/rpc.c/Makefile +@@ -8,10 +8,12 @@ include $(SPDK_ROOT_DIR)/mk/spdk.common.mk + + TEST_FILE = rpc_ut.c + ++# Evariops 0010: chmod, which the listen applies to the socket file. + SPDK_MOCK_SYMBOLS = \ + open \ + close \ + flock \ +- unlink ++ unlink \ ++ chmod + + include $(SPDK_ROOT_DIR)/mk/spdk.unittest.mk +diff --git a/test/unit/lib/rpc/rpc.c/rpc_ut.c b/test/unit/lib/rpc/rpc.c/rpc_ut.c +index 1512b10..b438ab6 100644 +--- a/test/unit/lib/rpc/rpc.c/rpc_ut.c ++++ b/test/unit/lib/rpc/rpc.c/rpc_ut.c +@@ -24,9 +24,6 @@ DEFINE_STUB_V(spdk_jsonrpc_end_result, (struct spdk_jsonrpc_request *request, + DEFINE_STUB(spdk_jsonrpc_begin_result, struct spdk_json_write_ctx *, + (struct spdk_jsonrpc_request *request), (void *)1); + DEFINE_STUB(spdk_json_decode_bool, int, (const struct spdk_json_val *val, void *out), 0); +-DEFINE_STUB(spdk_jsonrpc_server_listen, struct spdk_jsonrpc_server *, (int domain, int protocol, +- struct sockaddr *listen_addr, socklen_t addrlen, spdk_jsonrpc_handle_request_fn handle_request), +- (struct spdk_jsonrpc_server *)0Xdeaddead); + DEFINE_STUB(spdk_jsonrpc_server_poll, int, (struct spdk_jsonrpc_server *server), 0); + DEFINE_STUB_V(spdk_jsonrpc_server_shutdown, (struct spdk_jsonrpc_server *server)); + +@@ -37,6 +34,33 @@ DEFINE_WRAPPER(open, int, (const char *pathname, int flags, mode_t mode), (pathn + DEFINE_WRAPPER(close, int, (int fd), (fd)); + DEFINE_WRAPPER(flock, int, (int fd, int operation), (fd, operation)); + ++/* Evariops 0010: the listen binds the socket under umask 077, then chmods it to ++ * 0600. The bind is stubbed: the umask it runs under is recorded, and the chmod ++ * of a socket file it did not create is recorded rather than performed. */ ++static mode_t g_listen_umask; ++static mode_t g_chmod_mode; ++static int g_chmod_rc; ++ ++struct spdk_jsonrpc_server * ++spdk_jsonrpc_server_listen(int domain, int protocol, struct sockaddr *listen_addr, ++ socklen_t addrlen, spdk_jsonrpc_handle_request_fn handle_request) ++{ ++ g_listen_umask = umask(0); ++ umask(g_listen_umask); ++ ++ return (struct spdk_jsonrpc_server *)0Xdeaddead; ++} ++ ++int __wrap_chmod(const char *pathname, mode_t mode); ++ ++int ++__wrap_chmod(const char *pathname, mode_t mode) ++{ ++ g_chmod_mode = mode; ++ ++ return g_chmod_rc; ++} ++ + int + spdk_json_decode_object(const struct spdk_json_val *values, + const struct spdk_json_object_decoder *decoders, size_t num_decoders, void *out) +@@ -217,14 +241,22 @@ test_spdk_rpc_listen_close(void) + struct spdk_rpc_server *server; + const char listen_addr[128] = "/var/tmp/spdk-rpc-ut.sock"; + char rpc_lock_path[128] = {}; ++ mode_t saved_umask; + + MOCK_SET(open, 1); + MOCK_SET(close, 0); + MOCK_SET(flock, 0); + ++ saved_umask = umask(022); + server = spdk_rpc_server_listen(listen_addr); + SPDK_CU_ASSERT_FATAL(server != NULL); + ++ /* Evariops 0010: bound owner-only, re-asserted 0600, and the process umask ++ * given back. */ ++ CU_ASSERT(g_listen_umask == 077); ++ CU_ASSERT(g_chmod_mode == 0600); ++ CU_ASSERT(umask(saved_umask) == 022); ++ + snprintf(rpc_lock_path, sizeof(server->lock_path), "%s.lock", + server->listen_addr_unix.sun_path); + +@@ -240,6 +272,25 @@ test_spdk_rpc_listen_close(void) + MOCK_CLEAR(flock); + } + ++static void ++test_spdk_rpc_listen_chmod_failure(void) ++{ ++ const char listen_addr[128] = "/var/tmp/spdk-rpc-ut.sock"; ++ ++ MOCK_SET(open, 1); ++ MOCK_SET(close, 0); ++ MOCK_SET(flock, 0); ++ ++ /* Evariops 0010: a socket that cannot be made owner-only is not served. */ ++ g_chmod_rc = -1; ++ CU_ASSERT(spdk_rpc_server_listen(listen_addr) == NULL); ++ g_chmod_rc = 0; ++ ++ MOCK_CLEAR(open); ++ MOCK_CLEAR(close); ++ MOCK_CLEAR(flock); ++} ++ + static void + test_rpc_run_multiple_servers(void) + { +@@ -283,6 +334,7 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_rpc_get_methods); + CU_ADD_TEST(suite, test_rpc_spdk_get_version); + CU_ADD_TEST(suite, test_spdk_rpc_listen_close); ++ CU_ADD_TEST(suite, test_spdk_rpc_listen_chmod_failure); + CU_ADD_TEST(suite, test_rpc_run_multiple_servers); + + num_failures = spdk_ut_run_tests(argc, argv, NULL); diff --git a/patches/0013-raid1-seeded-rebuild.patch b/patches/0013-raid1-seeded-rebuild.patch index 52fc898..03a3beb 100644 --- a/patches/0013-raid1-seeded-rebuild.patch +++ b/patches/0013-raid1-seeded-rebuild.patch @@ -630,3 +630,38 @@ index 4fc593f..f1d4f0f 100644 raid1_ch->read_blocks_outstanding[i] < read_blocks_min) { read_blocks_min = raid1_ch->read_blocks_outstanding[i]; idx = i; +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index d47d6c2..4b33262 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -156,6 +156,8 @@ DEFINE_STUB_V(spdk_jsonrpc_send_bool_response, (struct spdk_jsonrpc_request *req + bool value)); + DEFINE_STUB(spdk_json_decode_string, int, (const struct spdk_json_val *val, void *out), 0); + DEFINE_STUB(spdk_json_decode_uint32, int, (const struct spdk_json_val *val, void *out), 0); ++/* Evariops 0013: the seeded-rebuild RPC decodes 64-bit ranges. */ ++DEFINE_STUB(spdk_json_decode_uint64, int, (const struct spdk_json_val *val, void *out), 0); + DEFINE_STUB(spdk_json_decode_uuid, int, (const struct spdk_json_val *val, void *out), 0); + DEFINE_STUB(spdk_json_decode_array, int, (const struct spdk_json_val *values, + spdk_json_decode_fn decode_func, +@@ -174,6 +176,9 @@ DEFINE_STUB(spdk_json_write_named_array_begin, int, (struct spdk_json_write_ctx + DEFINE_STUB(spdk_json_write_null, int, (struct spdk_json_write_ctx *w), 0); + DEFINE_STUB(spdk_json_write_named_uint64, int, (struct spdk_json_write_ctx *w, const char *name, + uint64_t val), 0); ++/* Evariops 0013: the seeded-rebuild RPC answers with a formatted rebuild id. */ ++DEFINE_STUB(spdk_json_write_named_string_fmt, int, (struct spdk_json_write_ctx *w, ++ const char *name, const char *fmt, ...), 0); + DEFINE_STUB(spdk_strerror, const char *, (int errnum), NULL); + DEFINE_STUB(spdk_bdev_queue_io_wait, int, (struct spdk_bdev *bdev, struct spdk_io_channel *ch, + struct spdk_bdev_io_wait_entry *entry), 0); +@@ -208,6 +213,11 @@ DEFINE_STUB_V(raid_tier_heat_destroy, (void *heat)); + * NVMe status. Only raid1.c marks one, and it is not part of this suite. */ + DEFINE_STUB_V(spdk_bdev_io_complete_nvme_status, (struct spdk_bdev_io *bdev_io, uint32_t cdw0, + int sct, int sc)); ++/* Evariops 0013: an RPC that changes state is audited by the jsonrpc server. Its ++ * object also defines the error responses, which this suite defines itself: the ++ * audit is stubbed rather than linked. */ ++DEFINE_STUB_V(spdk_jsonrpc_request_audit, (struct spdk_jsonrpc_request *request, ++ const char *method, const char *detail)); + + uint32_t + spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) diff --git a/patches/0014-raid-incarnation-estale.patch b/patches/0014-raid-incarnation-estale.patch index ff2e3a2..7267938 100644 --- a/patches/0014-raid-incarnation-estale.patch +++ b/patches/0014-raid-incarnation-estale.patch @@ -446,3 +446,127 @@ index 830a42d..3de537f 100644 } SPDK_RPC_REGISTER("bdev_raid_remove_base_bdev", rpc_bdev_raid_remove_base_bdev, SPDK_RPC_RUNTIME) +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index 4b33262..c262806 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -18,6 +18,8 @@ + #define MAX_RAIDS 2 + #define BLOCK_CNT (1024ul * 1024ul * 1024ul * 1024ul) + #define MD_SIZE 8 ++/* Evariops 0014: the incarnation the suite creates its raids under. */ ++#define UT_INCARNATION "ut-incarnation" + + struct spdk_bdev_channel { + struct spdk_io_channel *channel; +@@ -552,6 +554,10 @@ spdk_json_decode_object(const struct spdk_json_val *values, + _out->strip_size_kb = req->strip_size_kb; + _out->level = req->level; + _out->superblock_enabled = req->superblock_enabled; ++ if (req->incarnation != NULL) { ++ _out->incarnation = strdup(req->incarnation); ++ SPDK_CU_ASSERT_FATAL(_out->incarnation != NULL); ++ } + _out->base_bdevs.num_base_bdevs = req->base_bdevs.num_base_bdevs; + for (i = 0; i < req->base_bdevs.num_base_bdevs; i++) { + _out->base_bdevs.base_bdevs[i] = strdup(req->base_bdevs.base_bdevs[i]); +@@ -856,6 +862,8 @@ create_test_req(struct rpc_bdev_raid_create *r, const char *raid_name, + r->strip_size_kb = (g_strip_size * g_block_len) / 1024; + r->level = 123; + r->superblock_enabled = superblock_enabled; ++ r->incarnation = strdup(UT_INCARNATION); ++ SPDK_CU_ASSERT_FATAL(r->incarnation != NULL); + r->base_bdevs.num_base_bdevs = g_max_base_drives; + for (i = 0; i < g_max_base_drives; i++, bbdev_idx++) { + snprintf(name, 16, "%s%u%s", "Nvme", bbdev_idx, "n1"); +@@ -888,6 +896,7 @@ free_test_req(struct rpc_bdev_raid_create *r) + uint8_t i; + + free(r->name); ++ free(r->incarnation); + for (i = 0; i < r->base_bdevs.num_base_bdevs; i++) { + free(r->base_bdevs.base_bdevs[i]); + } +@@ -899,6 +908,8 @@ create_raid_bdev_delete_req(struct rpc_bdev_raid_delete *r, const char *raid_nam + { + r->name = strdup(raid_name); + SPDK_CU_ASSERT_FATAL(r->name != NULL); ++ /* Evariops 0014: unchecked unless a test names the incarnation it expects. */ ++ r->expected_incarnation = NULL; + + g_rpc_req = r; + g_rpc_req_size = sizeof(*r); +@@ -974,6 +985,64 @@ test_delete_raid(void) + reset_globals(); + } + ++/* Evariops 0014: no creation without an incarnation identity. */ ++static void ++test_create_raid_without_incarnation(void) ++{ ++ struct rpc_bdev_raid_create req; ++ ++ set_globals(); ++ CU_ASSERT(raid_bdev_init() == 0); ++ ++ create_raid_bdev_create_req(&req, "raid1", 0, true, 0, false); ++ free(req.incarnation); ++ req.incarnation = NULL; ++ rpc_bdev_raid_create(NULL, NULL); ++ CU_ASSERT(g_rpc_err == 1); ++ verify_raid_bdev_present("raid1", false); ++ free_test_req(&req); ++ ++ raid_bdev_exit(); ++ base_bdevs_cleanup(); ++ reset_globals(); ++} ++ ++/* Evariops 0014: a delete that expects another incarnation is refused (-ESTALE), ++ * and the raid stays. The one that expects the raid's own deletes it. */ ++static void ++test_delete_raid_stale_incarnation(void) ++{ ++ struct rpc_bdev_raid_create construct_req; ++ struct rpc_bdev_raid_delete delete_req; ++ ++ set_globals(); ++ CU_ASSERT(raid_bdev_init() == 0); ++ ++ create_raid_bdev_create_req(&construct_req, "raid1", 0, true, 0, false); ++ rpc_bdev_raid_create(NULL, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ verify_raid_bdev_present("raid1", true); ++ free_test_req(&construct_req); ++ ++ create_raid_bdev_delete_req(&delete_req, "raid1", 0); ++ delete_req.expected_incarnation = strdup("another-incarnation"); ++ SPDK_CU_ASSERT_FATAL(delete_req.expected_incarnation != NULL); ++ rpc_bdev_raid_delete(NULL, NULL); ++ CU_ASSERT(g_rpc_err == 1); ++ verify_raid_bdev_present("raid1", true); ++ ++ create_raid_bdev_delete_req(&delete_req, "raid1", 0); ++ delete_req.expected_incarnation = strdup(UT_INCARNATION); ++ SPDK_CU_ASSERT_FATAL(delete_req.expected_incarnation != NULL); ++ rpc_bdev_raid_delete(NULL, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ verify_raid_bdev_present("raid1", false); ++ ++ raid_bdev_exit(); ++ base_bdevs_cleanup(); ++ reset_globals(); ++} ++ + static void + test_create_raid_invalid_args(void) + { +@@ -1832,6 +1901,8 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_create_raid); + CU_ADD_TEST(suite, test_create_raid_superblock); + CU_ADD_TEST(suite, test_delete_raid); ++ CU_ADD_TEST(suite, test_create_raid_without_incarnation); ++ CU_ADD_TEST(suite, test_delete_raid_stale_incarnation); + CU_ADD_TEST(suite, test_create_raid_invalid_args); + CU_ADD_TEST(suite, test_delete_raid_invalid_args); + CU_ADD_TEST(suite, test_io_channel); diff --git a/patches/0015-raid-rebuild-outcome-registry.patch b/patches/0015-raid-rebuild-outcome-registry.patch index 15b43b5..5ebc2d7 100644 --- a/patches/0015-raid-rebuild-outcome-registry.patch +++ b/patches/0015-raid-rebuild-outcome-registry.patch @@ -584,3 +584,24 @@ index 3de537f..31896d5 100644 static const struct spdk_json_object_decoder rpc_bdev_raid_set_options_decoders[] = { {"process_window_size_kb", offsetof(struct spdk_raid_bdev_opts, process_window_size_kb), spdk_json_decode_uint32, true}, {"process_max_bandwidth_mb_sec", offsetof(struct spdk_raid_bdev_opts, process_max_bandwidth_mb_sec), spdk_json_decode_uint32, true}, +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index c262806..c538a3c 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -220,6 +220,16 @@ DEFINE_STUB_V(spdk_bdev_io_complete_nvme_status, (struct spdk_bdev_io *bdev_io, + * audit is stubbed rather than linked. */ + DEFINE_STUB_V(spdk_jsonrpc_request_audit, (struct spdk_jsonrpc_request *request, + const char *method, const char *detail)); ++/* Evariops 0015: a rebuild records its outcome in a registry. None is opened here, ++ * which the raid allows: a rebuild without an entry is only unobservable. */ ++DEFINE_STUB(raid_rebuild_outcome_open, struct raid_rebuild_outcome *, (const char *token, ++ const char *raid_name, const char *base_bdev_name), NULL); ++DEFINE_STUB_V(raid_rebuild_outcome_add_bytes, (struct raid_rebuild_outcome *outcome, ++ uint64_t bytes)); ++DEFINE_STUB_V(raid_rebuild_outcome_finish, (struct raid_rebuild_outcome *outcome, ++ enum raid_rebuild_outcome_state terminal, bool verified)); ++DEFINE_STUB_V(raid_rebuild_outcomes_write_json, (struct spdk_json_write_ctx *w, ++ const char *token_filter)); + + uint32_t + spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) diff --git a/patches/0017-raid-member-observation.patch b/patches/0017-raid-member-observation.patch index 5af8fed..9ff9a77 100644 --- a/patches/0017-raid-member-observation.patch +++ b/patches/0017-raid-member-observation.patch @@ -192,3 +192,28 @@ index 43841bc..ebe39cc 100644 }; struct raid_bdev_io; +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index c538a3c..ff51486 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -176,6 +176,9 @@ DEFINE_STUB(spdk_json_write_array_end, int, (struct spdk_json_write_ctx *w), 0); + DEFINE_STUB(spdk_json_write_named_array_begin, int, (struct spdk_json_write_ctx *w, + const char *name), 0); + DEFINE_STUB(spdk_json_write_null, int, (struct spdk_json_write_ctx *w), 0); ++/* Evariops 0017: a member is reported with the time it entered its state. */ ++DEFINE_STUB(spdk_json_write_named_int64, int, (struct spdk_json_write_ctx *w, const char *name, ++ int64_t val), 0); + DEFINE_STUB(spdk_json_write_named_uint64, int, (struct spdk_json_write_ctx *w, const char *name, + uint64_t val), 0); + /* Evariops 0013: the seeded-rebuild RPC answers with a formatted rebuild id. */ +@@ -230,6 +233,10 @@ DEFINE_STUB_V(raid_rebuild_outcome_finish, (struct raid_rebuild_outcome *outcome + enum raid_rebuild_outcome_state terminal, bool verified)); + DEFINE_STUB_V(raid_rebuild_outcomes_write_json, (struct spdk_json_write_ctx *w, + const char *token_filter)); ++/* Evariops 0017: a member is reported with its CBT epoch. The CBT module is not part ++ * of this suite: no member has one (-ENODEV). */ ++DEFINE_STUB(vbdev_cbt_query_latest_epoch, int, (const char *bdev_name, ++ struct vbdev_cbt_epoch_facts *out), -ENODEV); + + uint32_t + spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) diff --git a/patches/0018-raid-integrated-verify.patch b/patches/0018-raid-integrated-verify.patch index 67887bf..cafe429 100644 --- a/patches/0018-raid-integrated-verify.patch +++ b/patches/0018-raid-integrated-verify.patch @@ -551,3 +551,16 @@ index fe71954..f2611d8 100644 free(process); } +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index ff51486..aab5d12 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -237,6 +237,8 @@ DEFINE_STUB_V(raid_rebuild_outcomes_write_json, (struct spdk_json_write_ctx *w, + * of this suite: no member has one (-ENODEV). */ + DEFINE_STUB(vbdev_cbt_query_latest_epoch, int, (const char *bdev_name, + struct vbdev_cbt_epoch_facts *out), -ENODEV); ++/* Evariops 0018: a rebuild's outcome turns to verifying before its verify pass. */ ++DEFINE_STUB_V(raid_rebuild_outcome_set_verifying, (struct raid_rebuild_outcome *outcome)); + + uint32_t + spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) diff --git a/patches/0019-raid-auto-epoch-at-ejection.patch b/patches/0019-raid-auto-epoch-at-ejection.patch index 09e28ec..05e453c 100644 --- a/patches/0019-raid-auto-epoch-at-ejection.patch +++ b/patches/0019-raid-auto-epoch-at-ejection.patch @@ -40,3 +40,18 @@ index f2611d8..0a1cdee 100644 if (raid_bdev->sb) { struct raid_bdev_superblock *sb = raid_bdev->sb; uint8_t slot = raid_bdev_base_bdev_slot(base_info); +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index aab5d12..43d6c52 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -239,6 +239,10 @@ DEFINE_STUB(vbdev_cbt_query_latest_epoch, int, (const char *bdev_name, + struct vbdev_cbt_epoch_facts *out), -ENODEV); + /* Evariops 0018: a rebuild's outcome turns to verifying before its verify pass. */ + DEFINE_STUB_V(raid_rebuild_outcome_set_verifying, (struct raid_rebuild_outcome *outcome)); ++/* Evariops 0019: a raid1 ejection opens a CBT epoch on each survivor. No survivor ++ * has CBT here (-ENODEV), which the raid passes over. */ ++DEFINE_STUB(vbdev_cbt_auto_epoch_open, int, (const char *bdev_name, const char *stale_backend_id), ++ -ENODEV); + + uint32_t + spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) diff --git a/patches/0021-raid-envelopes.patch b/patches/0021-raid-envelopes.patch index 858a6b7..f4e2a54 100644 --- a/patches/0021-raid-envelopes.patch +++ b/patches/0021-raid-envelopes.patch @@ -376,3 +376,20 @@ index 31896d5..5d9917f 100644 /* Evariops 0015: bdev_raid_get_rebuild_outcomes {token?} — read the rebuild * outcome registry. Entries outlive the raid bdev they describe and are matched * by token, so an attempt can still be resolved after its raid is gone. See +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index 43d6c52..55bdca4 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -243,6 +243,12 @@ DEFINE_STUB_V(raid_rebuild_outcome_set_verifying, (struct raid_rebuild_outcome * + * has CBT here (-ENODEV), which the raid passes over. */ + DEFINE_STUB(vbdev_cbt_auto_epoch_open, int, (const char *bdev_name, const char *stale_backend_id), + -ENODEV); ++/* Evariops 0021: the envelopes pace and admit the rebuilds. At their defaults here: ++ * no cap (0), and every rebuild admitted. */ ++DEFINE_STUB(raid_envelope_active_mb_sec, uint32_t, (enum raid_envelope_class klass), 0); ++DEFINE_STUB(raid_envelopes_admit_rebuild, bool, (void), true); ++DEFINE_STUB_V(raid_envelopes_get, (struct raid_envelopes *out)); ++DEFINE_STUB_V(raid_envelopes_set, (const struct raid_envelopes *envelopes)); + + uint32_t + spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) diff --git a/patches/0025-raid-clear-superblock.patch b/patches/0025-raid-clear-superblock.patch index eb3485c..a17f6f5 100644 --- a/patches/0025-raid-clear-superblock.patch +++ b/patches/0025-raid-clear-superblock.patch @@ -171,3 +171,21 @@ index 5d9917f..720a6cb 100644 + } +} +SPDK_RPC_REGISTER("bdev_raid_clear_superblock", rpc_bdev_raid_clear_superblock, SPDK_RPC_RUNTIME) +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index 55bdca4..806c8d3 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -249,6 +249,13 @@ DEFINE_STUB(raid_envelope_active_mb_sec, uint32_t, (enum raid_envelope_class kla + DEFINE_STUB(raid_envelopes_admit_rebuild, bool, (void), true); + DEFINE_STUB_V(raid_envelopes_get, (struct raid_envelopes *out)); + DEFINE_STUB_V(raid_envelopes_set, (const struct raid_envelopes *envelopes)); ++/* Evariops 0025: clearing a member's superblock writes a zeroed block over it. No ++ * test here clears one. */ ++DEFINE_STUB(spdk_bdev_get_block_size, uint32_t, (const struct spdk_bdev *bdev), 0); ++DEFINE_STUB(spdk_bdev_get_buf_align, size_t, (const struct spdk_bdev *bdev), 0); ++DEFINE_STUB(spdk_bdev_write, int, (struct spdk_bdev_desc *desc, struct spdk_io_channel *ch, ++ void *buf, uint64_t offset, uint64_t nbytes, spdk_bdev_io_completion_cb cb, ++ void *cb_arg), 0); + + uint32_t + spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) diff --git a/patches/0033-raid-remove-base-bdev-inflight-distinct-failure-visible.patch b/patches/0033-raid-remove-base-bdev-inflight-distinct-failure-visible.patch index 159e4d9..8520f41 100644 --- a/patches/0033-raid-remove-base-bdev-inflight-distinct-failure-visible.patch +++ b/patches/0033-raid-remove-base-bdev-inflight-distinct-failure-visible.patch @@ -190,3 +190,17 @@ index d137c1a..44066af 100644 int raid_bdev_remove_base_bdev(struct spdk_bdev *base_bdev, const char *expected_incarnation, raid_base_bdev_cb cb_fn, void *cb_ctx); +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index 806c8d3..9efc3c3 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -170,6 +170,9 @@ DEFINE_STUB(spdk_json_write_object_begin, int, (struct spdk_json_write_ctx *w), + DEFINE_STUB(spdk_json_write_named_object_begin, int, (struct spdk_json_write_ctx *w, + const char *name), 0); + DEFINE_STUB(spdk_json_write_string, int, (struct spdk_json_write_ctx *w, const char *val), 0); ++/* Evariops 0033: a member is reported with the errno of its failed removal. */ ++DEFINE_STUB(spdk_json_write_named_int32, int, (struct spdk_json_write_ctx *w, const char *name, ++ int32_t val), 0); + DEFINE_STUB(spdk_json_write_object_end, int, (struct spdk_json_write_ctx *w), 0); + DEFINE_STUB(spdk_json_write_array_begin, int, (struct spdk_json_write_ctx *w), 0); + DEFINE_STUB(spdk_json_write_array_end, int, (struct spdk_json_write_ctx *w), 0); From e539f7e0cdde403ddb61a12d7e38c725f82c156e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Wed, 30 Sep 2026 23:14:15 +0200 Subject: [PATCH 04/26] test(blob): a thin cluster reads zeroes where no write landed (0044, tests only) Patch 0044 starts as tests only, and the unit-tests stage is expected to fail on blob_ut: - blob_thin_prov_new_cluster_reads_zeroes: every free cluster is filled with 0xAA, as a device whose unmap does not zero leaves them. A write of one io_unit goes into a new cluster, and the rest of the cluster must read zeroes. Upstream returns the 0xAA. - blob_thin_prov_full_cluster_write_is_not_cleared_first: a write that covers the whole cluster costs the payload and the metadata page(s), and nothing more. It passes before the fix too: it guards the fix's cost. - blob_thin_prov_rw, blob_thin_prov_write_count_io and blob_thin_prov_rle count the cleared cluster in their write accounting. Locally (debug, arm64): blob_ut fails 4 tests in each of its 5 suites. The fix follows in the next commit. --- ...blob-thin-cluster-cleared-before-use.patch | 207 ++++++++++++++++++ 1 file changed, 207 insertions(+) create mode 100644 patches/0044-blob-thin-cluster-cleared-before-use.patch diff --git a/patches/0044-blob-thin-cluster-cleared-before-use.patch b/patches/0044-blob-thin-cluster-cleared-before-use.patch new file mode 100644 index 0000000..d980cb0 --- /dev/null +++ b/patches/0044-blob-thin-cluster-cleared-before-use.patch @@ -0,0 +1,207 @@ +diff --git a/test/unit/lib/blob/blob.c/blob_ut.c b/test/unit/lib/blob/blob.c/blob_ut.c +index c168e08..3bb3f5c 100644 +--- a/test/unit/lib/blob/blob.c/blob_ut.c ++++ b/test/unit/lib/blob/blob.c/blob_ut.c +@@ -4628,8 +4628,9 @@ blob_thin_prov_rw(void) + CU_ASSERT(free_clusters - 1 == spdk_bs_free_cluster_count(bs)); + CU_ASSERT(spdk_blob_get_num_allocated_clusters(blob) == 1); + /* For thin-provisioned blob we need to write 20 io_units plus one page metadata and +- * read 0 bytes */ +- expected_bytes = 20 * io_unit_size + spdk_bs_get_page_size(bs); ++ * read 0 bytes. Each thread allocated a cluster and cleared it before inserting it, ++ * the one that lost the race included. */ ++ expected_bytes = 20 * io_unit_size + spdk_bs_get_page_size(bs) + 2 * spdk_bs_get_cluster_size(bs); + if (g_use_extent_table) { + /* Add one more page for EXTENT_PAGE write */ + expected_bytes += spdk_bs_get_page_size(bs); +@@ -4731,17 +4732,18 @@ blob_thin_prov_write_count_io(void) + CU_ASSERT(free_clusters - (2 * i + 1) == spdk_bs_free_cluster_count(bs)); + + CU_ASSERT(g_dev_read_bytes == read_bytes); ++ /* The write covers one io_unit of a new cluster: the cluster is cleared first. */ + if (!g_use_extent_table) { + /* For legacy metadata, we should have written the io_unit for + * the write I/O, plus the blob's primary metadata page + */ +- expected_bytes = io_unit_size + spdk_bs_get_page_size(bs); ++ expected_bytes = CLUSTER_SZ + io_unit_size + spdk_bs_get_page_size(bs); + } else { + /* For extent table metadata, we should have written the io_unit for + * the write I/O, plus 2 metadata pages - the extent page and the + * blob's primary metadata page + */ +- expected_bytes = io_unit_size + 2 * spdk_bs_get_page_size(bs); ++ expected_bytes = CLUSTER_SZ + io_unit_size + 2 * spdk_bs_get_page_size(bs); + } + CU_ASSERT((g_dev_write_bytes - write_bytes) == expected_bytes); + +@@ -4774,8 +4776,9 @@ blob_thin_prov_write_count_io(void) + /* + * For legacy metadata, we should have written the I/O and the primary metadata page. + * For extent table metadata, we should have written the I/O and the extent metadata page. ++ * Both clear the new cluster first. + */ +- expected_bytes = io_unit_size + spdk_bs_get_page_size(bs); ++ expected_bytes = CLUSTER_SZ + io_unit_size + spdk_bs_get_page_size(bs); + CU_ASSERT((g_dev_write_bytes - write_bytes) == expected_bytes); + + /* Send unmap aligned to the whole cluster - should free it up */ +@@ -4809,6 +4812,138 @@ blob_thin_prov_write_count_io(void) + g_bs = NULL; + } + ++/* ++ * Fills every free cluster of the device with a pattern: what a device whose unmap does not ++ * zero keeps in the clusters of blobs deleted earlier. ++ */ ++static void ++ut_fill_free_clusters(struct spdk_blob_store *bs, uint8_t pattern) ++{ ++ uint32_t i; ++ ++ for (i = 0; i < bs->total_clusters; i++) { ++ if (!spdk_bit_pool_is_allocated(bs->used_clusters, i)) { ++ memset(&g_dev_buffer[bs_cluster_to_lba(bs, i) * bs->dev->blocklen], pattern, ++ bs->cluster_sz); ++ } ++ } ++} ++ ++/* ++ * A cluster allocated by a write that covers part of it is cleared first: the rest of it ++ * reads zeroes, not what the device held there, which is another blob's data where the ++ * device's unmap does not zero (and on each leg of a mirror, something different). ++ */ ++static void ++blob_thin_prov_new_cluster_reads_zeroes(void) ++{ ++ struct spdk_blob_store *bs = g_bs; ++ struct spdk_blob *blob; ++ struct spdk_io_channel *channel; ++ struct spdk_blob_opts opts; ++ uint64_t io_unit_size = spdk_bs_get_io_unit_size(bs); ++ uint64_t cluster_size = spdk_bs_get_cluster_size(bs); ++ uint64_t io_units_per_cluster = cluster_size / io_unit_size; ++ uint8_t payload_write[BLOCKLEN]; ++ uint8_t *payload_read; ++ ++ channel = spdk_bs_alloc_io_channel(bs); ++ SPDK_CU_ASSERT_FATAL(channel != NULL); ++ ++ ut_fill_free_clusters(bs, 0xAA); ++ ++ ut_spdk_blob_opts_init(&opts); ++ opts.thin_provision = true; ++ opts.num_clusters = 1; ++ blob = ut_blob_create_and_open(bs, &opts); ++ CU_ASSERT(spdk_blob_get_num_allocated_clusters(blob) == 0); ++ ++ /* One io_unit at the start of the cluster: the rest of it is never written. */ ++ memset(payload_write, 0xE5, sizeof(payload_write)); ++ spdk_blob_io_write(blob, channel, payload_write, 0, 1, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ CU_ASSERT(spdk_blob_get_num_allocated_clusters(blob) == 1); ++ ++ payload_read = calloc(1, cluster_size); ++ SPDK_CU_ASSERT_FATAL(payload_read != NULL); ++ spdk_blob_io_read(blob, channel, payload_read, 0, io_units_per_cluster, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ CU_ASSERT(memcmp(payload_read, payload_write, io_unit_size) == 0); ++ CU_ASSERT(spdk_mem_all_zero(payload_read + io_unit_size, cluster_size - io_unit_size)); ++ ++ free(payload_read); ++ ut_blob_close_and_delete(bs, blob); ++ spdk_bs_free_io_channel(channel); ++ poll_threads(); ++ g_blob = NULL; ++ g_blobid = 0; ++} ++ ++/* ++ * A write that covers the whole new cluster leaves nothing of the device's old content, so the ++ * cluster is not cleared first: a rebuild or a copy writes whole clusters, and clearing each of ++ * them would double what it writes. ++ */ ++static void ++blob_thin_prov_full_cluster_write_is_not_cleared_first(void) ++{ ++ struct spdk_blob_store *bs = g_bs; ++ struct spdk_blob *blob; ++ struct spdk_io_channel *channel; ++ struct spdk_blob_opts opts; ++ uint64_t io_unit_size = spdk_bs_get_io_unit_size(bs); ++ uint64_t cluster_size = spdk_bs_get_cluster_size(bs); ++ uint64_t io_units_per_cluster = cluster_size / io_unit_size; ++ uint64_t write_bytes; ++ uint64_t expected_bytes; ++ uint8_t *payload_write; ++ uint8_t *payload_read; ++ ++ channel = spdk_bs_alloc_io_channel(bs); ++ SPDK_CU_ASSERT_FATAL(channel != NULL); ++ ++ ut_fill_free_clusters(bs, 0xAA); ++ ++ ut_spdk_blob_opts_init(&opts); ++ opts.thin_provision = true; ++ opts.num_clusters = 1; ++ blob = ut_blob_create_and_open(bs, &opts); ++ ++ payload_write = malloc(cluster_size); ++ payload_read = calloc(1, cluster_size); ++ SPDK_CU_ASSERT_FATAL(payload_write != NULL && payload_read != NULL); ++ memset(payload_write, 0xE5, cluster_size); ++ ++ write_bytes = g_dev_write_bytes; ++ spdk_blob_io_write(blob, channel, payload_write, 0, io_units_per_cluster, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ CU_ASSERT(spdk_blob_get_num_allocated_clusters(blob) == 1); ++ ++ /* The payload and the blob's metadata page; no cleared cluster before them. */ ++ expected_bytes = cluster_size + spdk_bs_get_page_size(bs); ++ if (g_use_extent_table) { ++ /* Add one more page for EXTENT_PAGE write */ ++ expected_bytes += spdk_bs_get_page_size(bs); ++ } ++ CU_ASSERT(g_dev_write_bytes - write_bytes == expected_bytes); ++ ++ spdk_blob_io_read(blob, channel, payload_read, 0, io_units_per_cluster, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ CU_ASSERT(memcmp(payload_read, payload_write, cluster_size) == 0); ++ ++ free(payload_write); ++ free(payload_read); ++ ut_blob_close_and_delete(bs, blob); ++ spdk_bs_free_io_channel(channel); ++ poll_threads(); ++ g_blob = NULL; ++ g_blobid = 0; ++} ++ + static void + blob_thin_prov_unmap_cluster(void) + { +@@ -5124,8 +5259,8 @@ blob_thin_prov_rle(void) + CU_ASSERT(g_bserrno == 0); + CU_ASSERT(free_clusters - 1 == spdk_bs_free_cluster_count(bs)); + /* For thin-provisioned blob we need to write 10 pages plus one page metadata and +- * read 0 bytes */ +- expected_bytes = 10 * io_unit_size + spdk_bs_get_page_size(bs); ++ * read 0 bytes, and clear the new cluster before the write lands in it. */ ++ expected_bytes = 10 * io_unit_size + spdk_bs_get_page_size(bs) + spdk_bs_get_cluster_size(bs); + if (g_use_extent_table) { + /* Add one more page for EXTENT_PAGE write */ + expected_bytes += spdk_bs_get_page_size(bs); +@@ -10364,6 +10499,8 @@ main(int argc, char **argv) + CU_ADD_TEST(suite_bs, blob_insert_cluster_msg_test); + CU_ADD_TEST(suite_bs, blob_thin_prov_rw); + CU_ADD_TEST(suite, blob_thin_prov_write_count_io); ++ CU_ADD_TEST(suite_bs, blob_thin_prov_new_cluster_reads_zeroes); ++ CU_ADD_TEST(suite_bs, blob_thin_prov_full_cluster_write_is_not_cleared_first); + CU_ADD_TEST(suite, blob_thin_prov_unmap_cluster); + CU_ADD_TEST(suite_bs, blob_thin_prov_rle); + CU_ADD_TEST(suite_bs, blob_thin_prov_rw_iov); From e94d8627ea70237be0e2a22e6f5307f283cfa8c8 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Wed, 30 Sep 2026 23:15:21 +0200 Subject: [PATCH 05/26] fix(blob): a thin cluster is cleared before use (0044) A cluster that a write allocates in a blob backed by zeroes (a thin blob with no parent, or a clone over a range its snapshot never allocated) was inserted as the device left it. The write filled its own range, and the rest of the cluster returned what was stored there before. The free path counts on unmap to leave zeroes, which a SATA SSD does not guarantee and a bdev without unmap support never provides. When the new cluster is not copied from a parent, bs_allocate_and_copy_cluster now clears it with write_zeroes before inserting it. A write (or write_zeroes) that covers the whole cluster skips the clear, so a rebuild or a copy does not write twice. A clear that fails returns the cluster to the pool. The inflate path goes through the same branch. Locally (debug, arm64): blob_ut passes, 495 tests. --- ...blob-thin-cluster-cleared-before-use.patch | 66 +++++++++++++++++++ patches/README.md | 3 +- 2 files changed, 68 insertions(+), 1 deletion(-) diff --git a/patches/0044-blob-thin-cluster-cleared-before-use.patch b/patches/0044-blob-thin-cluster-cleared-before-use.patch index d980cb0..d8e702b 100644 --- a/patches/0044-blob-thin-cluster-cleared-before-use.patch +++ b/patches/0044-blob-thin-cluster-cleared-before-use.patch @@ -1,3 +1,69 @@ +diff --git a/lib/blob/blobstore.c b/lib/blob/blobstore.c +index f774f99..4d656c9 100644 +--- a/lib/blob/blobstore.c ++++ b/lib/blob/blobstore.c +@@ -2827,6 +2827,40 @@ blob_copy(struct spdk_blob_copy_cluster_ctx *ctx, spdk_bs_user_op_t *op, uint64_ + blob_write_copy_cpl, ctx); + } + ++/* ++ * The write that triggered the allocation overwrites the whole cluster: no byte of what the ++ * device held there survives it, so clearing the cluster first would write it twice (a rebuild ++ * or a copy writes whole clusters). ++ */ ++static bool ++blob_op_covers_cluster(struct spdk_blob *blob, spdk_bs_user_op_t *op, uint64_t cluster_start_io_unit) ++{ ++ struct spdk_bs_user_op_args *args = &((struct spdk_bs_request_set *)op)->u.user_op; ++ ++ if (args->type != SPDK_BLOB_WRITE && args->type != SPDK_BLOB_WRITEV && ++ args->type != SPDK_BLOB_WRITE_ZEROES) { ++ return false; ++ } ++ ++ return args->offset <= cluster_start_io_unit && ++ args->offset + args->length >= cluster_start_io_unit + bs_io_units_per_cluster(blob); ++} ++ ++static void ++blob_zero_fill_cpl(spdk_bs_sequence_t *seq, void *cb_arg, int bserrno) ++{ ++ struct spdk_blob_copy_cluster_ctx *ctx = cb_arg; ++ ++ if (bserrno) { ++ /* The cluster was never handed to the blob: it goes back to the pool. */ ++ blob_insert_cluster_revert(ctx); ++ bs_sequence_finish(seq, bserrno); ++ return; ++ } ++ ++ blob_write_copy_cpl(seq, cb_arg, 0); ++} ++ + static void + bs_allocate_and_copy_cluster(struct spdk_blob *blob, + struct spdk_io_channel *_ch, +@@ -2937,9 +2971,19 @@ bs_allocate_and_copy_cluster(struct spdk_blob *blob, + blob_write_copy, ctx); + } + +- } else { ++ } else if (blob_op_covers_cluster(blob, op, cluster_start_io_unit)) { + blob_insert_cluster_on_md_thread(ctx->blob, cluster_number, ctx->new_cluster, + ctx->new_extent_page, ctx->new_cluster_page, blob_insert_cluster_cpl, ctx); ++ } else { ++ /* ++ * A cluster whose backing reads as zeroes is handed out as the device left it: the part ++ * the triggering write does not cover would read whatever an earlier blob wrote there, ++ * wherever the device's unmap does not zero (the free path counts on it to). That leaks ++ * one blob's data into another, and makes the legs of a mirror differ in bytes nobody ++ * wrote. The cluster is cleared before the blob sees it. ++ */ ++ bs_sequence_write_zeroes_dev(ctx->seq, bs_cluster_to_lba(blob->bs, ctx->new_cluster), ++ bs_cluster_to_lba(blob->bs, 1), blob_zero_fill_cpl, ctx); + } + } + diff --git a/test/unit/lib/blob/blob.c/blob_ut.c b/test/unit/lib/blob/blob.c/blob_ut.c index c168e08..3bb3f5c 100644 --- a/test/unit/lib/blob/blob.c/blob_ut.c diff --git a/patches/README.md b/patches/README.md index eb82251..1e590d8 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0041`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0044`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -49,6 +49,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0041`) | 0039 | an armed delayed reconnect owns the continuation — the path returns instead of re-driving a failed disconnect, ending a mutual recursion that overflows the reactor stack | module/bdev/nvme (bdev_nvme.c) | 0032 | | 0040 | the delete-stop only MARKS — 0035's stop drove the window machinery from a foreign thread, unquiescing under in-flight requests; it now sets STOPPING with `-ECANCELED` and each resting state concludes | module/bdev/raid (bdev_raid.c) | 0035, 0039 | | 0041 | lvol → tier band usage provider — when an lvolstore loads on a `bdev_tier` composite the lvol layer registers a per-band fill-accounting provider, gated on the tier grain being a whole multiple of the lvolstore cluster (else no provider, honest unknown), and **unregisters it at teardown INITIATION** (before `spdk_lvs_unload`/`destroy`, whose synchronous store free would otherwise race a `get_bands` poll — UAF); `bdev_tier_get_bands` then fills `used_blocks` by counting allocated clusters of the blobstore's `used_clusters` pool in the band's LBA range (`spdk_bs_count_allocated_clusters_in_lba_range`, new in lib/blob — exclusive cluster-aligned bounds, one masked popcount per 64 clusters via `spdk_bit_pool_count_allocated_in_range` / `spdk_bit_array_count_set_in_range`, new in lib/util) | lib/blob, lib/util (ranged popcount), bdev_lvol; `#include`s `vbdev_tier.h` (requires the tier module, like 0005) | 0004 (bit-pool accessors), 0005 (lvol `-I` tier CFLAGS), tier module | +| 0044 | a thin cluster is cleared before use — a cluster allocated by a write (or write_zeroes, or an inflate) that covers only part of it is zeroed before the blob sees it. Its unwritten part returned what the device held there: another blob's data wherever the device's unmap does not zero (the free path counts on unmap to), and different bytes on each leg of a mirror, which a verify of the mirror reads as a content divergence. A write that covers the whole cluster skips the clear, so a rebuild or a copy does not write twice. `blob_ut` pins both, and its thin-provisioning I/O accounting now counts the cleared cluster | lib/blob (blobstore.c), test/unit/lib/blob | — | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From d42771cd7cdd494bdc4b67452314722d273cfc9d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 04:58:02 +0200 Subject: [PATCH 06/26] test(raid): a seeded rebuild copies its ranges, and only under the windows they need (0045, tests only) Three cases for bdev_raid_ut, red on the series as it stands: - two seed ranges inside one window: 12 blocks to copy under one lock; the window is copied whole (128 blocks); - one range starting mid-window across three 4-block windows: 9 blocks under three locks, the first at the range; the windows are copied whole (12); - one range at the end of the raid: 4 blocks under one lock, at the range; the clean windows before it are each locked and skipped. The stubbed range quiesce counts the windows the copy phase locks. --- ...aid-seeded-rebuild-copies-its-ranges.patch | 157 ++++++++++++++++++ 1 file changed, 157 insertions(+) create mode 100644 patches/0045-raid-seeded-rebuild-copies-its-ranges.patch diff --git a/patches/0045-raid-seeded-rebuild-copies-its-ranges.patch b/patches/0045-raid-seeded-rebuild-copies-its-ranges.patch new file mode 100644 index 0000000..1a78986 --- /dev/null +++ b/patches/0045-raid-seeded-rebuild-copies-its-ranges.patch @@ -0,0 +1,157 @@ +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index 9efc3c3..f5a439b 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -667,11 +667,23 @@ spdk_bdev_unquiesce(struct spdk_bdev *bdev, struct spdk_bdev_module *module, + return 0; + } + ++/* Evariops 0045: the windows the copy phase of a process locks, the verify ++ * phase's aside, and the lowest of them. */ ++static uint32_t g_copy_windows_locked; ++static uint64_t g_lowest_copy_window = UINT64_MAX; ++ + int + spdk_bdev_quiesce_range(struct spdk_bdev *bdev, struct spdk_bdev_module *module, + uint64_t offset, uint64_t length, + spdk_bdev_quiesce_cb cb_fn, void *cb_arg) + { ++ struct raid_bdev *raid_bdev = SPDK_CONTAINEROF(bdev, struct raid_bdev, bdev); ++ ++ if (raid_bdev->process != NULL && raid_bdev->process->verify == NULL) { ++ g_copy_windows_locked++; ++ g_lowest_copy_window = spdk_min(g_lowest_copy_window, offset); ++ } ++ + if (cb_fn) { + cb_fn(cb_arg, 0); + } +@@ -1547,6 +1559,118 @@ test_raid_process(void) + reset_globals(); + } + ++/* Evariops 0045: a seeded rebuild copies its seed ranges and nothing else. A ++ * window that intersected a range used to be copied whole: scattered writes ++ * across a volume dirty every window, and the "seeded" pass copied the volume. ++ * Runs the rebuild with the given window and returns the blocks copied. */ ++static uint64_t ++run_seeded_rebuild(uint32_t window_size_kb, const struct raid_bdev_seed_range *seed, uint32_t num_ranges) ++{ ++ struct rpc_bdev_raid_create req; ++ struct rpc_bdev_raid_delete destroy_req; ++ struct raid_bdev *pbdev; ++ struct spdk_bdev *base_bdev; ++ struct spdk_thread *process_thread; ++ struct raid_bdev_seed_range *ranges; ++ struct spdk_raid_bdev_opts opts = {}; ++ uint64_t num_blocks_processed = 0; ++ ++ set_globals(); ++ CU_ASSERT(raid_bdev_init() == 0); ++ opts.process_window_size_kb = window_size_kb; ++ CU_ASSERT(raid_bdev_set_opts(&opts) == 0); ++ g_copy_windows_locked = 0; ++ g_lowest_copy_window = UINT64_MAX; ++ ++ create_raid_bdev_create_req(&req, "raid1", 0, true, 0, false); ++ TAILQ_FOREACH(base_bdev, &g_bdev_list, internal.link) { ++ base_bdev->blockcnt = 128; ++ } ++ rpc_bdev_raid_create(NULL, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ free_test_req(&req); ++ ++ TAILQ_FOREACH(pbdev, &g_raid_bdev_list, global_link) { ++ if (strcmp(pbdev->bdev.name, "raid1") == 0) { ++ break; ++ } ++ } ++ SPDK_CU_ASSERT_FATAL(pbdev != NULL); ++ pbdev->module_private = &num_blocks_processed; ++ pbdev->min_base_bdevs_operational = 0; ++ ++ ranges = calloc(num_ranges, sizeof(*ranges)); ++ SPDK_CU_ASSERT_FATAL(ranges != NULL); ++ memcpy(ranges, seed, num_ranges * sizeof(*ranges)); ++ pbdev->base_bdev_info[0].write_only_pending = true; ++ CU_ASSERT(raid_bdev_start_seeded_rebuild(&pbdev->base_bdev_info[0], ranges, num_ranges, NULL) == 0); ++ poll_app_thread(); ++ SPDK_CU_ASSERT_FATAL(pbdev->process != NULL); ++ ++ process_thread = g_latest_thread; ++ while (spdk_thread_poll(process_thread, 0, 0) > 0) { ++ poll_app_thread(); ++ } ++ CU_ASSERT(pbdev->process == NULL); ++ poll_app_thread(); ++ ++ create_raid_bdev_delete_req(&destroy_req, "raid1", 0); ++ rpc_bdev_raid_delete(NULL, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ ++ /* The window size is global: put the default back for the next tests. */ ++ opts.process_window_size_kb = 1024; ++ CU_ASSERT(raid_bdev_set_opts(&opts) == 0); ++ raid_bdev_exit(); ++ base_bdevs_cleanup(); ++ reset_globals(); ++ ++ return num_blocks_processed; ++} ++ ++static void ++test_raid_process_seeded_copies_only_its_ranges(void) ++{ ++ /* One 1 MiB window holds the whole 128-block raid: both ranges fall in it, ++ * and are copied under that one window's lock. */ ++ const struct raid_bdev_seed_range seed[] = { ++ { .offset_blocks = 10, .num_blocks = 4 }, ++ { .offset_blocks = 40, .num_blocks = 8 }, ++ }; ++ ++ CU_ASSERT(run_seeded_rebuild(1024, seed, 2) == 12); ++ CU_ASSERT(g_copy_windows_locked == 1); ++} ++ ++static void ++test_raid_process_seeded_range_across_windows(void) ++{ ++ /* 16 KiB windows of 4 blocks: [6, 15) starts mid-window and spans three, ++ * the first of them at the range. */ ++ const struct raid_bdev_seed_range seed[] = { ++ { .offset_blocks = 6, .num_blocks = 9 }, ++ }; ++ ++ CU_ASSERT(run_seeded_rebuild(16, seed, 1) == 9); ++ CU_ASSERT(g_copy_windows_locked == 3); ++ CU_ASSERT(g_lowest_copy_window == 6); ++} ++ ++static void ++test_raid_process_seeded_skips_a_clean_span_without_a_window(void) ++{ ++ /* A clean span needs no copy, hence no lock: the 25 windows before the ++ * one range used to be locked and fast-skipped one after the other — on a ++ * 1 TiB volume, 262144 window cycles for an empty delta. */ ++ const struct raid_bdev_seed_range seed[] = { ++ { .offset_blocks = 100, .num_blocks = 4 }, ++ }; ++ ++ CU_ASSERT(run_seeded_rebuild(16, seed, 1) == 4); ++ CU_ASSERT(g_copy_windows_locked == 1); ++ CU_ASSERT(g_lowest_copy_window == 100); ++} ++ + static void + test_raid_process_with_qos(void) + { +@@ -1953,6 +2077,9 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_raid_level_conversions); + CU_ADD_TEST(suite, test_raid_io_split); + CU_ADD_TEST(suite, test_raid_process); ++ CU_ADD_TEST(suite, test_raid_process_seeded_copies_only_its_ranges); ++ CU_ADD_TEST(suite, test_raid_process_seeded_range_across_windows); ++ CU_ADD_TEST(suite, test_raid_process_seeded_skips_a_clean_span_without_a_window); + CU_ADD_TEST(suite, test_raid_process_with_qos); + + spdk_thread_lib_init(test_new_thread_fn, 0); From 604f3ef3bc826f637641b96afba329d7e0f3c8ad Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 04:58:30 +0200 Subject: [PATCH 07/26] fix(raid): a seeded rebuild copies its ranges, and only under the windows they need (0045) A window that a seed range touched was copied whole, and every clean window was locked and skipped in turn. Measured on a three-node cluster under a random-write load: 5 GiB re-copied for 395 MiB of dirty chunks (4794 ranges), 110 s. Copying only the ranges, one window per range part, still took 100 s: the cost is the window cycle (quiesce, copy, channel update, unquiesce), not the bytes. - A seeded process jumps over what its ranges leave clean, with no window. Only a copy needs a window's quiesce (a write landing between the read of the source and the write of the target), and the seeded target is a member of every channel for writes: the routing by process offset sends a write to it on either side. - A window copies every range part it holds, under its one lock, and ends where everything before it is copied or clean. Out of requests, it ends at the first block not submitted and the next window resumes there. The verify phase and the copied-bytes accounting are unchanged. bdev_raid_ut: the three cases of the previous commit pass, and every raid suite stays green (bdev_raid 20, bdev_raid_sb 9, concat 3, raid0 8, raid1 5, raid5f 8). --- ...aid-seeded-rebuild-copies-its-ranges.patch | 99 +++++++++++++++++++ patches/README.md | 1 + 2 files changed, 100 insertions(+) diff --git a/patches/0045-raid-seeded-rebuild-copies-its-ranges.patch b/patches/0045-raid-seeded-rebuild-copies-its-ranges.patch index 1a78986..b919eb4 100644 --- a/patches/0045-raid-seeded-rebuild-copies-its-ranges.patch +++ b/patches/0045-raid-seeded-rebuild-copies-its-ranges.patch @@ -1,3 +1,102 @@ +diff --git a/module/bdev/raid/bdev_raid.c b/module/bdev/raid/bdev_raid.c +index 5064454..583707e 100644 +--- a/module/bdev/raid/bdev_raid.c ++++ b/module/bdev/raid/bdev_raid.c +@@ -3332,6 +3332,71 @@ raid_bdev_process_window_seeded_clean(struct raid_bdev_process *process) + return true; + } + ++/* Evariops 0045: the copy of a seeded window — every seed range part the window ++ * holds, all under its one lock. The window then ends where everything before it ++ * is either copied or clean; out of requests, it ends at the first block not ++ * submitted, and the next window resumes there. Copying the window whole as soon ++ * as one range touched it made scattered writes copy the whole volume (5 GiB for ++ * 395 MiB of dirty chunks), and one window per range part made them pay a lock ++ * cycle per part (100 s for the same 395 MiB). Process thread only. */ ++static void ++raid_bdev_process_window_seeded_copy(struct raid_bdev_process *process, uint64_t offset_end) ++{ ++ const uint64_t offset = process->window_offset; ++ uint64_t covered = offset_end; ++ uint32_t idx; ++ int ret; ++ ++ for (idx = process->seed_range_idx; idx < process->num_seed_ranges; idx++) { ++ const struct raid_bdev_seed_range *range = &process->seed_ranges[idx]; ++ uint64_t start = spdk_max(range->offset_blocks, offset); ++ const uint64_t end = spdk_min(range->offset_blocks + range->num_blocks, offset_end); ++ ++ if (start >= offset_end) { ++ break; ++ } ++ while (start < end) { ++ ret = raid_bdev_submit_process_request(process, start, end - start); ++ if (ret <= 0) { ++ covered = start; ++ goto submitted; ++ } ++ process->window_remaining += ret; ++ start += ret; ++ } ++ } ++ ++submitted: ++ if (process->window_remaining > 0) { ++ process->window_size = covered - offset; ++ } else { ++ raid_bdev_process_finish(process, process->window_status); ++ } ++} ++ ++/* Evariops 0045: a seeded process jumps, with no window, over what its ranges ++ * leave clean. Only a copy needs a window's quiesce (a write landing between the ++ * read of the source and the write of the target); the seeded target is a member ++ * of every channel for writes, so the routing by process.offset sends a write to ++ * it on either side of the offset. Without the jump, every clean window was ++ * locked and fast-skipped in turn: 262144 cycles on a 1 TiB volume, whatever its ++ * delta. Process thread, no window locked. */ ++static void ++raid_bdev_process_seeded_skip_clean(struct raid_bdev_process *process) ++{ ++ while (process->seed_range_idx < process->num_seed_ranges) { ++ const struct raid_bdev_seed_range *range = &process->seed_ranges[process->seed_range_idx]; ++ ++ if (range->offset_blocks + range->num_blocks <= process->window_offset) { ++ process->seed_range_idx++; ++ continue; ++ } ++ process->window_offset = spdk_max(process->window_offset, range->offset_blocks); ++ return; ++ } ++ process->window_offset = process->raid_bdev->bdev.blockcnt; ++} ++ + static void + _raid_bdev_process_thread_run(struct raid_bdev_process *process) + { +@@ -3351,6 +3416,11 @@ _raid_bdev_process_thread_run(struct raid_bdev_process *process) + return; + } + ++ if (spdk_unlikely(process->seeded)) { ++ raid_bdev_process_window_seeded_copy(process, offset_end); ++ return; ++ } ++ + while (offset < offset_end) { + ret = raid_bdev_submit_process_request(process, offset, offset_end - offset); + if (ret <= 0) { +@@ -3907,6 +3977,10 @@ raid_bdev_process_thread_run(struct raid_bdev_process *process) + return; + } + ++ if (spdk_unlikely(process->seeded)) { ++ raid_bdev_process_seeded_skip_clean(process); ++ } ++ + if (process->window_offset == raid_bdev->bdev.blockcnt) { + SPDK_DEBUGLOG(bdev_raid, "process completed on %s\n", raid_bdev->bdev.name); + /* Evariops 0018: run the verify phase before concluding. The diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c index 9efc3c3..f5a439b 100644 --- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c diff --git a/patches/README.md b/patches/README.md index 1e590d8..69551c6 100644 --- a/patches/README.md +++ b/patches/README.md @@ -50,6 +50,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0044`) | 0040 | the delete-stop only MARKS — 0035's stop drove the window machinery from a foreign thread, unquiescing under in-flight requests; it now sets STOPPING with `-ECANCELED` and each resting state concludes | module/bdev/raid (bdev_raid.c) | 0035, 0039 | | 0041 | lvol → tier band usage provider — when an lvolstore loads on a `bdev_tier` composite the lvol layer registers a per-band fill-accounting provider, gated on the tier grain being a whole multiple of the lvolstore cluster (else no provider, honest unknown), and **unregisters it at teardown INITIATION** (before `spdk_lvs_unload`/`destroy`, whose synchronous store free would otherwise race a `get_bands` poll — UAF); `bdev_tier_get_bands` then fills `used_blocks` by counting allocated clusters of the blobstore's `used_clusters` pool in the band's LBA range (`spdk_bs_count_allocated_clusters_in_lba_range`, new in lib/blob — exclusive cluster-aligned bounds, one masked popcount per 64 clusters via `spdk_bit_pool_count_allocated_in_range` / `spdk_bit_array_count_set_in_range`, new in lib/util) | lib/blob, lib/util (ranged popcount), bdev_lvol; `#include`s `vbdev_tier.h` (requires the tier module, like 0005) | 0004 (bit-pool accessors), 0005 (lvol `-I` tier CFLAGS), tier module | | 0044 | a thin cluster is cleared before use — a cluster allocated by a write (or write_zeroes, or an inflate) that covers only part of it is zeroed before the blob sees it. Its unwritten part returned what the device held there: another blob's data wherever the device's unmap does not zero (the free path counts on unmap to), and different bytes on each leg of a mirror, which a verify of the mirror reads as a content divergence. A write that covers the whole cluster skips the clear, so a rebuild or a copy does not write twice. `blob_ut` pins both, and its thin-provisioning I/O accounting now counts the cleared cluster | lib/blob (blobstore.c), test/unit/lib/blob | — | +| 0045 | a seeded rebuild copies its seed ranges, and only under the windows they need — a seeded process jumps over what its ranges leave clean with no window (only a copy needs a window's quiesce, and the seeded target takes writes on either side of the process offset), and a window copies every range part it holds under its one lock. A window that a range touched used to be copied whole (scattered writes: 5 GiB re-copied for 395 MiB of dirty chunks), and every clean window was locked and skipped in turn (262144 cycles on a 1 TiB volume, whatever its delta). `bdev_raid_ut` pins the blocks copied and the windows locked | bdev_raid, test/unit/lib/bdev/raid | 0013 | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From 42b0e9c467cd29dbf0eedcec3b29878503583e47 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 05:20:53 +0200 Subject: [PATCH 08/26] feat(raid): a seeded rebuild reads its ranges from the CBT epoch (0046) A JSON-RPC request carries ~200 ranges (SPDK_JSONRPC_MAX_VALUES = 1024 parsed values, a 32 KiB receive buffer), so a caller had to fold a larger delta to fit, and folding a scattered delta recopies every clean block in between: measured on a three-node cluster under random writes, 4485 dirty ranges (368 MiB) folded to 128 recopied 3.2 GiB of clean blocks, and the seeded rebuild of a 5 GiB volume took as long as a full one (95-110 s). bdev_raid_start_seeded_rebuild_from_epoch {name, base_bdev, cbt_bdev, epoch_id, expected_incarnation, rebuild_token} seeds the same rebuild with the dirty ranges of a frozen epoch, read in-process: the cbt module answers vbdev_cbt_query_epoch_ranges (the ranges bdev_cbt_epoch_get_dirty_ranges walks, up to the caller's cap; more is -E2BIG, since a truncated delta is a wrong one, and the caller rebuilds in full). The two RPCs share their checks (raid, incarnation, running process, write_only member, range bounds, token) and their audit. bdev_raid_ut fakes the query: 64 one-block ranges a block apart copy their 64 blocks, and a delta the query refuses starts nothing. Every raid suite stays green (bdev_raid 22, bdev_raid_sb 9, concat 3, raid0 8, raid1 5, raid5f 8). --- module/bdev/cbt/vbdev_cbt.c | 47 ++ module/bdev/cbt/vbdev_cbt_query.h | 23 + ...uild-reads-its-ranges-from-the-epoch.patch | 522 ++++++++++++++++++ patches/README.md | 1 + 4 files changed, 593 insertions(+) create mode 100644 patches/0046-raid-seeded-rebuild-reads-its-ranges-from-the-epoch.patch diff --git a/module/bdev/cbt/vbdev_cbt.c b/module/bdev/cbt/vbdev_cbt.c index 13f0adb..b4221f2 100644 --- a/module/bdev/cbt/vbdev_cbt.c +++ b/module/bdev/cbt/vbdev_cbt.c @@ -573,6 +573,53 @@ vbdev_cbt_query_latest_epoch(const char *bdev_name, struct vbdev_cbt_epoch_facts return 0; } +/* Cross-module query — the delta of an epoch for the raid module's seeded + * rebuild (see vbdev_cbt_query.h): the ranges bdev_cbt_epoch_get_dirty_ranges + * walks, in the query's own type, refused whole when they do not all fit. */ +int +vbdev_cbt_query_epoch_ranges(const char *bdev_name, const char *epoch_id, uint32_t max, + struct vbdev_cbt_range **out, uint32_t *count) +{ + struct cbt_dirty_range *ranges = NULL; + struct vbdev_cbt_range *copy = NULL; + uint64_t dirty_chunks, total_chunks; + uint32_t num_ranges = 0, chunk_size_kb, i; + bool truncated = false; + int rc; + + assert(spdk_get_thread() == spdk_thread_get_app_thread()); + + *out = NULL; + *count = 0; + + rc = bdev_cbt_epoch_get_dirty_ranges(bdev_name, epoch_id, max, &ranges, &num_ranges, + &dirty_chunks, &total_chunks, &chunk_size_kb, &truncated); + if (rc != 0) { + return rc; + } + if (truncated) { + free(ranges); + return -E2BIG; + } + + if (num_ranges > 0) { + copy = calloc(num_ranges, sizeof(*copy)); + if (copy == NULL) { + free(ranges); + return -ENOMEM; + } + for (i = 0; i < num_ranges; i++) { + copy[i].offset_blocks = ranges[i].offset_blocks; + copy[i].num_blocks = ranges[i].length_blocks; + } + } + free(ranges); + + *out = copy; + *count = num_ranges; + return 0; +} + static int vbdev_cbt_dump_info_json(void *ctx, struct spdk_json_write_ctx *w) { diff --git a/module/bdev/cbt/vbdev_cbt_query.h b/module/bdev/cbt/vbdev_cbt_query.h index cfd29ec..0053d7c 100644 --- a/module/bdev/cbt/vbdev_cbt_query.h +++ b/module/bdev/cbt/vbdev_cbt_query.h @@ -45,4 +45,27 @@ int vbdev_cbt_query_latest_epoch(const char *bdev_name, struct vbdev_cbt_epoch_f */ int vbdev_cbt_auto_epoch_open(const char *bdev_name, const char *stale_backend_id); +/* One dirty range of an epoch, in blocks of the cbt bdev: the raid's own, since + * a cbt bdev passes its LBAs through. */ +struct vbdev_cbt_range { + uint64_t offset_blocks; + uint64_t num_blocks; +}; + +/** + * The dirty ranges of a frozen (or rebuilding) epoch of the cbt bdev named + * \c bdev_name, sorted and non-overlapping, for the raid module to seed a + * rebuild with in-process. A JSON-RPC request carries ~200 ranges at most + * (1024 parsed values, a 32 KiB receive buffer), and folding a scattered delta + * down to that many recopies every clean block in between. + * + * \param max the most ranges the caller takes. A delta with more is -E2BIG: + * truncated, it would not be a smaller delta but a wrong one. + * \return 0 with \c *out allocated (the caller frees it; NULL when \c *count is + * 0); -ENODEV, -ENOENT or -EINVAL as bdev_cbt_epoch_get_dirty_ranges; + * -E2BIG; -ENOMEM. App thread only. + */ +int vbdev_cbt_query_epoch_ranges(const char *bdev_name, const char *epoch_id, uint32_t max, + struct vbdev_cbt_range **out, uint32_t *count); + #endif /* SPDK_VBDEV_CBT_QUERY_H */ diff --git a/patches/0046-raid-seeded-rebuild-reads-its-ranges-from-the-epoch.patch b/patches/0046-raid-seeded-rebuild-reads-its-ranges-from-the-epoch.patch new file mode 100644 index 0000000..f7213c7 --- /dev/null +++ b/patches/0046-raid-seeded-rebuild-reads-its-ranges-from-the-epoch.patch @@ -0,0 +1,522 @@ +diff --git a/module/bdev/raid/bdev_raid_rpc.c b/module/bdev/raid/bdev_raid_rpc.c +index 6ded0ca..d4127d2 100644 +--- a/module/bdev/raid/bdev_raid_rpc.c ++++ b/module/bdev/raid/bdev_raid_rpc.c +@@ -8,6 +8,8 @@ + #include "bdev_raid.h" + #include "bdev_raid_outcomes.h" + #include "bdev_raid_envelopes.h" ++/* Evariops 0046: the seed of a rebuild read from the CBT epoch in-process. */ ++#include "../cbt/vbdev_cbt_query.h" + #include "spdk/util.h" + #include "spdk/string.h" + #include "spdk/log.h" +@@ -670,9 +672,9 @@ SPDK_RPC_REGISTER("bdev_raid_add_base_bdev", rpc_bdev_raid_add_base_bdev, SPDK_R + * progress shows up as the raid's process in bdev_raid_get_bdevs, and + * completion makes the member read-eligible and CONFIGURED in the superblock. */ + +-/* The rebuild engine works in 1 MiB windows, so finer ranges buy nothing: the +- * caller coalesces its dirty set down to this many ranges. Same cap as +- * bdev_raid_rebuild_ranges. */ ++/* A request decodes this many ranges at most; the JSON-RPC server parses ~200 ++ * anyway. A larger delta goes through bdev_raid_start_seeded_rebuild_from_epoch, ++ * which reads it in-process. Same cap as bdev_raid_rebuild_ranges. */ + #define RAID_SEEDED_REBUILD_MAX_RANGES 4096u + + struct rpc_bdev_raid_start_seeded_rebuild { +@@ -759,51 +761,40 @@ seed_range_cmp(const void *a, const void *b) + return 0; + } + +-static void +-rpc_bdev_raid_start_seeded_rebuild(struct spdk_jsonrpc_request *request, +- const struct spdk_json_val *params) ++/* The raid and its write_only-attached member a seeded rebuild targets, or ++ * false with the request answered: unknown raid, stale incarnation, a process ++ * already running, a member that is not write_only-attached. */ ++static bool ++seeded_rebuild_target(struct spdk_jsonrpc_request *request, const char *name, ++ const char *base_bdev, const char *expected_incarnation, ++ struct raid_bdev **raid_out, struct raid_base_bdev_info **base_out) + { +- struct rpc_bdev_raid_start_seeded_rebuild req = {}; + struct raid_bdev *raid_bdev; + struct raid_base_bdev_info *base_info = NULL, *iter; +- struct raid_bdev_seed_range *ranges; +- struct spdk_json_write_ctx *w; +- uint64_t blockcnt; +- char detail[192]; +- uint32_t i; +- int rc; +- +- if (spdk_json_decode_object(params, rpc_bdev_raid_start_seeded_rebuild_decoders, +- SPDK_COUNTOF(rpc_bdev_raid_start_seeded_rebuild_decoders), +- &req)) { +- spdk_jsonrpc_send_error_response(request, SPDK_JSONRPC_ERROR_PARSE_ERROR, +- "spdk_json_decode_object failed"); +- goto cleanup; +- } + +- raid_bdev = raid_bdev_find_by_name(req.name); ++ raid_bdev = raid_bdev_find_by_name(name); + if (raid_bdev == NULL) { + spdk_jsonrpc_send_error_response_fmt(request, -ENODEV, "raid bdev %s is not found in config", +- req.name); +- goto cleanup; ++ name); ++ return false; + } + +- if (raid_bdev_check_incarnation(raid_bdev, req.expected_incarnation) != 0) { ++ if (raid_bdev_check_incarnation(raid_bdev, expected_incarnation) != 0) { + spdk_jsonrpc_send_error_response_fmt(request, -ESTALE, + "stale incarnation for raid bdev %s", +- req.name); +- goto cleanup; ++ name); ++ return false; + } + + if (raid_bdev->process != NULL) { + spdk_jsonrpc_send_error_response_fmt(request, -EBUSY, + "raid bdev %s already has a background process", +- req.name); +- goto cleanup; ++ name); ++ return false; + } + + RAID_FOR_EACH_BASE_BDEV(raid_bdev, iter) { +- if (iter->name != NULL && strcmp(iter->name, req.base_bdev) == 0) { ++ if (iter->name != NULL && strcmp(iter->name, base_bdev) == 0) { + base_info = iter; + break; + } +@@ -811,24 +802,41 @@ rpc_bdev_raid_start_seeded_rebuild(struct spdk_jsonrpc_request *request, + if (base_info == NULL) { + spdk_jsonrpc_send_error_response_fmt(request, -ENODEV, + "base bdev %s is not a member of raid bdev %s", +- req.base_bdev, req.name); +- goto cleanup; ++ base_bdev, name); ++ return false; + } + + if (!base_info->write_only_pending) { + spdk_jsonrpc_send_error_response_fmt(request, -EINVAL, + "base bdev %s is not write_only-attached", +- req.base_bdev); +- goto cleanup; ++ base_bdev); ++ return false; + } + +- /* Sort, then validate bounds and non-overlap (the engine's fast-skip +- * cursor depends on both). Touching ranges are fine. */ +- qsort(req.ranges, req.num_ranges, sizeof(*req.ranges), seed_range_cmp); ++ *raid_out = raid_bdev; ++ *base_out = base_info; ++ return true; ++} + +- blockcnt = raid_bdev->bdev.blockcnt; +- for (i = 0; i < req.num_ranges; i++) { +- const struct raid_bdev_seed_range *range = &req.ranges[i]; ++/* Validates sorted ranges against the raid, audits the RPC, starts the seeded ++ * rebuild and answers the request. Takes ownership of ranges, freed on every ++ * path the rebuild does not take them. */ ++static void ++seeded_rebuild_start(struct spdk_jsonrpc_request *request, const char *method, ++ struct raid_bdev *raid_bdev, struct raid_base_bdev_info *base_info, ++ struct raid_bdev_seed_range *ranges, uint32_t num_ranges, ++ const char *rebuild_token, const char *audit_extra) ++{ ++ struct spdk_json_write_ctx *w; ++ uint64_t blockcnt = raid_bdev->bdev.blockcnt; ++ char detail[256]; ++ uint32_t i; ++ int rc; ++ ++ /* The engine's cursor depends on sorted, non-overlapping ranges; touching ++ * ranges are fine. */ ++ for (i = 0; i < num_ranges; i++) { ++ const struct raid_bdev_seed_range *range = &ranges[i]; + + if (range->num_blocks == 0 || range->offset_blocks >= blockcnt || + range->num_blocks > blockcnt - range->offset_blocks) { +@@ -837,55 +845,85 @@ rpc_bdev_raid_start_seeded_rebuild(struct spdk_jsonrpc_request *request, + "exceeds the raid size (%" PRIu64 " blocks)", + i, range->offset_blocks, range->num_blocks, + blockcnt); +- goto cleanup; ++ goto err; + } +- if (i > 0 && req.ranges[i - 1].offset_blocks + req.ranges[i - 1].num_blocks > ++ if (i > 0 && ranges[i - 1].offset_blocks + ranges[i - 1].num_blocks > + range->offset_blocks) { + spdk_jsonrpc_send_error_response_fmt(request, SPDK_JSONRPC_ERROR_INVALID_PARAMS, + "ranges %u and %u overlap", i - 1, i); +- goto cleanup; ++ goto err; + } + } + +- if (req.rebuild_token != NULL && +- strnlen(req.rebuild_token, RAID_REBUILD_TOKEN_MAX) == RAID_REBUILD_TOKEN_MAX) { ++ if (rebuild_token != NULL && ++ strnlen(rebuild_token, RAID_REBUILD_TOKEN_MAX) == RAID_REBUILD_TOKEN_MAX) { + spdk_jsonrpc_send_error_response_fmt(request, -ENAMETOOLONG, + "rebuild_token exceeds %d characters", + RAID_REBUILD_TOKEN_MAX - 1); +- goto cleanup; ++ goto err; + } + + /* Evariops 0011: audit every RPC that writes to the array — this one arms + * a background writer. */ +- snprintf(detail, sizeof(detail), "raid=%s base=%s ranges=%zu", +- req.name, req.base_bdev, req.num_ranges); +- spdk_jsonrpc_request_audit(request, "bdev_raid_start_seeded_rebuild", detail); ++ snprintf(detail, sizeof(detail), "raid=%s base=%s ranges=%u%s", ++ raid_bdev->bdev.name, base_info->name, num_ranges, audit_extra); ++ spdk_jsonrpc_request_audit(request, method, detail); + + /* An empty set is legal: nothing was dirty, the backfill skips the whole + * device, and the completion still promotes the member. */ +- if (req.num_ranges == 0) { +- free(req.ranges); ++ if (num_ranges == 0) { ++ free(ranges); + ranges = NULL; +- } else { +- ranges = req.ranges; + } +- req.ranges = NULL; + +- rc = raid_bdev_start_seeded_rebuild(base_info, ranges, (uint32_t)req.num_ranges, +- req.rebuild_token); ++ rc = raid_bdev_start_seeded_rebuild(base_info, ranges, num_ranges, rebuild_token); + if (rc != 0) { +- free(ranges); + spdk_jsonrpc_send_error_response_fmt(request, rc, + "Failed to start seeded rebuild on %s/%s: %s", +- req.name, req.base_bdev, spdk_strerror(-rc)); +- goto cleanup; ++ raid_bdev->bdev.name, base_info->name, ++ spdk_strerror(-rc)); ++ goto err; + } + + w = spdk_jsonrpc_begin_result(request); + spdk_json_write_object_begin(w); +- spdk_json_write_named_string_fmt(w, "rebuild_id", "%s:%s", req.name, req.base_bdev); ++ spdk_json_write_named_string_fmt(w, "rebuild_id", "%s:%s", raid_bdev->bdev.name, ++ base_info->name); + spdk_json_write_object_end(w); + spdk_jsonrpc_end_result(request, w); ++ return; ++ ++err: ++ free(ranges); ++} ++ ++static void ++rpc_bdev_raid_start_seeded_rebuild(struct spdk_jsonrpc_request *request, ++ const struct spdk_json_val *params) ++{ ++ struct rpc_bdev_raid_start_seeded_rebuild req = {}; ++ struct raid_bdev *raid_bdev; ++ struct raid_base_bdev_info *base_info; ++ struct raid_bdev_seed_range *ranges; ++ ++ if (spdk_json_decode_object(params, rpc_bdev_raid_start_seeded_rebuild_decoders, ++ SPDK_COUNTOF(rpc_bdev_raid_start_seeded_rebuild_decoders), ++ &req)) { ++ spdk_jsonrpc_send_error_response(request, SPDK_JSONRPC_ERROR_PARSE_ERROR, ++ "spdk_json_decode_object failed"); ++ goto cleanup; ++ } ++ ++ if (!seeded_rebuild_target(request, req.name, req.base_bdev, req.expected_incarnation, ++ &raid_bdev, &base_info)) { ++ goto cleanup; ++ } ++ ++ qsort(req.ranges, req.num_ranges, sizeof(*req.ranges), seed_range_cmp); ++ ranges = req.ranges; ++ req.ranges = NULL; ++ seeded_rebuild_start(request, "bdev_raid_start_seeded_rebuild", raid_bdev, base_info, ++ ranges, (uint32_t)req.num_ranges, req.rebuild_token, ""); + + cleanup: + free_rpc_bdev_raid_start_seeded_rebuild(&req); +@@ -893,6 +931,107 @@ cleanup: + SPDK_RPC_REGISTER("bdev_raid_start_seeded_rebuild", rpc_bdev_raid_start_seeded_rebuild, + SPDK_RPC_RUNTIME) + ++/* Evariops 0046: bdev_raid_start_seeded_rebuild_from_epoch {name, base_bdev, ++ * cbt_bdev, epoch_id} — the seeded rebuild of bdev_raid_start_seeded_rebuild, ++ * seeded with the dirty ranges of a frozen CBT epoch read in-process. A request ++ * carries ~200 ranges at most (1024 parsed values, a 32 KiB receive buffer): ++ * folding a scattered delta down to that many recopied the clean blocks in ++ * between, 3.2 GiB of them for 4485 dirty ranges (368 MiB), measured. A delta of ++ * more ranges than this is refused (-E2BIG): the caller rebuilds in full. */ ++#define RAID_SEEDED_REBUILD_MAX_EPOCH_RANGES 65536u ++ ++struct rpc_bdev_raid_start_seeded_rebuild_from_epoch { ++ char *name; ++ char *base_bdev; ++ /* The cbt bdev over this raid, and the frozen epoch of the member's outage. */ ++ char *cbt_bdev; ++ char *epoch_id; ++ char *expected_incarnation; ++ char *rebuild_token; ++}; ++ ++static void ++free_rpc_bdev_raid_start_seeded_rebuild_from_epoch( ++ struct rpc_bdev_raid_start_seeded_rebuild_from_epoch *req) ++{ ++ free(req->name); ++ free(req->base_bdev); ++ free(req->cbt_bdev); ++ free(req->epoch_id); ++ free(req->expected_incarnation); ++ free(req->rebuild_token); ++} ++ ++static const struct spdk_json_object_decoder rpc_bdev_raid_start_seeded_rebuild_from_epoch_decoders[] = { ++ {"name", offsetof(struct rpc_bdev_raid_start_seeded_rebuild_from_epoch, name), spdk_json_decode_string}, ++ {"base_bdev", offsetof(struct rpc_bdev_raid_start_seeded_rebuild_from_epoch, base_bdev), spdk_json_decode_string}, ++ {"cbt_bdev", offsetof(struct rpc_bdev_raid_start_seeded_rebuild_from_epoch, cbt_bdev), spdk_json_decode_string}, ++ {"epoch_id", offsetof(struct rpc_bdev_raid_start_seeded_rebuild_from_epoch, epoch_id), spdk_json_decode_string}, ++ {"expected_incarnation", offsetof(struct rpc_bdev_raid_start_seeded_rebuild_from_epoch, expected_incarnation), spdk_json_decode_string, true}, ++ {"rebuild_token", offsetof(struct rpc_bdev_raid_start_seeded_rebuild_from_epoch, rebuild_token), spdk_json_decode_string, true}, ++}; ++ ++static void ++rpc_bdev_raid_start_seeded_rebuild_from_epoch(struct spdk_jsonrpc_request *request, ++ const struct spdk_json_val *params) ++{ ++ struct rpc_bdev_raid_start_seeded_rebuild_from_epoch req = {}; ++ struct raid_bdev *raid_bdev; ++ struct raid_base_bdev_info *base_info; ++ struct vbdev_cbt_range *epoch_ranges = NULL; ++ struct raid_bdev_seed_range *ranges = NULL; ++ uint32_t num_ranges = 0, i; ++ char audit_extra[96]; ++ int rc; ++ ++ if (spdk_json_decode_object(params, rpc_bdev_raid_start_seeded_rebuild_from_epoch_decoders, ++ SPDK_COUNTOF(rpc_bdev_raid_start_seeded_rebuild_from_epoch_decoders), ++ &req)) { ++ spdk_jsonrpc_send_error_response(request, SPDK_JSONRPC_ERROR_PARSE_ERROR, ++ "spdk_json_decode_object failed"); ++ goto cleanup; ++ } ++ ++ if (!seeded_rebuild_target(request, req.name, req.base_bdev, req.expected_incarnation, ++ &raid_bdev, &base_info)) { ++ goto cleanup; ++ } ++ ++ rc = vbdev_cbt_query_epoch_ranges(req.cbt_bdev, req.epoch_id, ++ RAID_SEEDED_REBUILD_MAX_EPOCH_RANGES, ++ &epoch_ranges, &num_ranges); ++ if (rc != 0) { ++ spdk_jsonrpc_send_error_response_fmt(request, rc, ++ "epoch %s of %s gives no seed for %s/%s: %s", ++ req.epoch_id, req.cbt_bdev, req.name, req.base_bdev, ++ spdk_strerror(-rc)); ++ goto cleanup; ++ } ++ ++ if (num_ranges > 0) { ++ ranges = calloc(num_ranges, sizeof(*ranges)); ++ if (ranges == NULL) { ++ free(epoch_ranges); ++ spdk_jsonrpc_send_error_response(request, -ENOMEM, spdk_strerror(ENOMEM)); ++ goto cleanup; ++ } ++ for (i = 0; i < num_ranges; i++) { ++ ranges[i].offset_blocks = epoch_ranges[i].offset_blocks; ++ ranges[i].num_blocks = epoch_ranges[i].num_blocks; ++ } ++ } ++ free(epoch_ranges); ++ ++ snprintf(audit_extra, sizeof(audit_extra), " epoch=%s", req.epoch_id); ++ seeded_rebuild_start(request, "bdev_raid_start_seeded_rebuild_from_epoch", raid_bdev, ++ base_info, ranges, num_ranges, req.rebuild_token, audit_extra); ++ ++cleanup: ++ free_rpc_bdev_raid_start_seeded_rebuild_from_epoch(&req); ++} ++SPDK_RPC_REGISTER("bdev_raid_start_seeded_rebuild_from_epoch", ++ rpc_bdev_raid_start_seeded_rebuild_from_epoch, SPDK_RPC_RUNTIME) ++ + struct rpc_bdev_raid_remove_base_bdev { + /* Base bdev name */ + char *name; +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index f5a439b..11b46f2 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -246,6 +246,40 @@ DEFINE_STUB_V(raid_rebuild_outcome_set_verifying, (struct raid_rebuild_outcome * + * has CBT here (-ENODEV), which the raid passes over. */ + DEFINE_STUB(vbdev_cbt_auto_epoch_open, int, (const char *bdev_name, const char *stale_backend_id), + -ENODEV); ++ ++/* Evariops 0046: a seeded rebuild reads its ranges from a CBT epoch. The CBT ++ * module is not part of this suite: the epoch's delta is this fake's. */ ++static const struct raid_bdev_seed_range *g_epoch_seed; ++static uint32_t g_epoch_seed_count; ++static int g_epoch_query_rc; ++static char g_epoch_queried[64]; ++ ++int ++vbdev_cbt_query_epoch_ranges(const char *bdev_name, const char *epoch_id, uint32_t max, ++ struct vbdev_cbt_range **out, uint32_t *count) ++{ ++ uint32_t i; ++ ++ snprintf(g_epoch_queried, sizeof(g_epoch_queried), "%s/%s", bdev_name, epoch_id); ++ *out = NULL; ++ *count = 0; ++ if (g_epoch_query_rc != 0) { ++ return g_epoch_query_rc; ++ } ++ if (g_epoch_seed_count > max) { ++ return -E2BIG; ++ } ++ if (g_epoch_seed_count > 0) { ++ *out = calloc(g_epoch_seed_count, sizeof(**out)); ++ SPDK_CU_ASSERT_FATAL(*out != NULL); ++ for (i = 0; i < g_epoch_seed_count; i++) { ++ (*out)[i].offset_blocks = g_epoch_seed[i].offset_blocks; ++ (*out)[i].num_blocks = g_epoch_seed[i].num_blocks; ++ } ++ } ++ *count = g_epoch_seed_count; ++ return 0; ++} + /* Evariops 0021: the envelopes pace and admit the rebuilds. At their defaults here: + * no cap (0), and every rebuild admitted. */ + DEFINE_STUB(raid_envelope_active_mb_sec, uint32_t, (enum raid_envelope_class klass), 0); +@@ -1562,9 +1596,12 @@ test_raid_process(void) + /* Evariops 0045: a seeded rebuild copies its seed ranges and nothing else. A + * window that intersected a range used to be copied whole: scattered writes + * across a volume dirty every window, and the "seeded" pass copied the volume. +- * Runs the rebuild with the given window and returns the blocks copied. */ ++ * Runs the rebuild with the given window and returns the blocks copied — ++ * started directly, or (Evariops 0046) through the RPC that reads the seed from ++ * a CBT epoch, whose query answers query_rc. A refused start copies nothing. */ + static uint64_t +-run_seeded_rebuild(uint32_t window_size_kb, const struct raid_bdev_seed_range *seed, uint32_t num_ranges) ++seeded_rebuild_run(uint32_t window_size_kb, const struct raid_bdev_seed_range *seed, uint32_t num_ranges, ++ bool via_epoch, int query_rc) + { + struct rpc_bdev_raid_create req; + struct rpc_bdev_raid_delete destroy_req; +@@ -1599,20 +1636,43 @@ run_seeded_rebuild(uint32_t window_size_kb, const struct raid_bdev_seed_range *s + pbdev->module_private = &num_blocks_processed; + pbdev->min_base_bdevs_operational = 0; + +- ranges = calloc(num_ranges, sizeof(*ranges)); +- SPDK_CU_ASSERT_FATAL(ranges != NULL); +- memcpy(ranges, seed, num_ranges * sizeof(*ranges)); + pbdev->base_bdev_info[0].write_only_pending = true; +- CU_ASSERT(raid_bdev_start_seeded_rebuild(&pbdev->base_bdev_info[0], ranges, num_ranges, NULL) == 0); ++ if (via_epoch) { ++ struct rpc_bdev_raid_start_seeded_rebuild_from_epoch epoch_req = { ++ .name = strdup("raid1"), ++ .base_bdev = strdup(pbdev->base_bdev_info[0].name), ++ .cbt_bdev = strdup("cbt_raid1"), ++ .epoch_id = strdup("epoch-0046"), ++ }; ++ ++ g_epoch_seed = seed; ++ g_epoch_seed_count = num_ranges; ++ g_epoch_query_rc = query_rc; ++ g_json_decode_obj_create = 0; ++ g_rpc_req = &epoch_req; ++ g_rpc_req_size = sizeof(epoch_req); ++ rpc_bdev_raid_start_seeded_rebuild_from_epoch(NULL, NULL); ++ g_epoch_query_rc = 0; ++ CU_ASSERT(strcmp(g_epoch_queried, "cbt_raid1/epoch-0046") == 0); ++ CU_ASSERT(g_rpc_err == (query_rc != 0)); ++ } else { ++ ranges = calloc(num_ranges, sizeof(*ranges)); ++ SPDK_CU_ASSERT_FATAL(ranges != NULL); ++ memcpy(ranges, seed, num_ranges * sizeof(*ranges)); ++ CU_ASSERT(raid_bdev_start_seeded_rebuild(&pbdev->base_bdev_info[0], ranges, num_ranges, NULL) == 0); ++ } + poll_app_thread(); +- SPDK_CU_ASSERT_FATAL(pbdev->process != NULL); + +- process_thread = g_latest_thread; +- while (spdk_thread_poll(process_thread, 0, 0) > 0) { +- poll_app_thread(); ++ if (query_rc == 0) { ++ SPDK_CU_ASSERT_FATAL(pbdev->process != NULL); ++ process_thread = g_latest_thread; ++ while (spdk_thread_poll(process_thread, 0, 0) > 0) { ++ poll_app_thread(); ++ } + } + CU_ASSERT(pbdev->process == NULL); + poll_app_thread(); ++ g_rpc_err = 0; + + create_raid_bdev_delete_req(&destroy_req, "raid1", 0); + rpc_bdev_raid_delete(NULL, NULL); +@@ -1628,6 +1688,40 @@ run_seeded_rebuild(uint32_t window_size_kb, const struct raid_bdev_seed_range *s + return num_blocks_processed; + } + ++static uint64_t ++run_seeded_rebuild(uint32_t window_size_kb, const struct raid_bdev_seed_range *seed, uint32_t num_ranges) ++{ ++ return seeded_rebuild_run(window_size_kb, seed, num_ranges, false, 0); ++} ++ ++/* Evariops 0046: the epoch's ranges, read in-process, are the seed as they are — ++ * here 64 one-block ranges a block apart, each copied alone. A request carries ++ * ~200 ranges, and a scattered delta folded to fit one recopied the clean blocks ++ * in between. */ ++static void ++test_raid_seeded_rebuild_from_epoch_copies_every_range(void) ++{ ++ struct raid_bdev_seed_range seed[64]; ++ uint32_t i; ++ ++ for (i = 0; i < 64; i++) { ++ seed[i].offset_blocks = 2 * i; ++ seed[i].num_blocks = 1; ++ } ++ ++ CU_ASSERT(seeded_rebuild_run(16, seed, 64, true, 0) == 64); ++} ++ ++static void ++test_raid_seeded_rebuild_from_epoch_refuses_a_delta_it_cannot_take_whole(void) ++{ ++ const struct raid_bdev_seed_range seed[] = { ++ { .offset_blocks = 10, .num_blocks = 4 }, ++ }; ++ ++ CU_ASSERT(seeded_rebuild_run(16, seed, 1, true, -E2BIG) == 0); ++} ++ + static void + test_raid_process_seeded_copies_only_its_ranges(void) + { +@@ -2080,6 +2174,8 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_raid_process_seeded_copies_only_its_ranges); + CU_ADD_TEST(suite, test_raid_process_seeded_range_across_windows); + CU_ADD_TEST(suite, test_raid_process_seeded_skips_a_clean_span_without_a_window); ++ CU_ADD_TEST(suite, test_raid_seeded_rebuild_from_epoch_copies_every_range); ++ CU_ADD_TEST(suite, test_raid_seeded_rebuild_from_epoch_refuses_a_delta_it_cannot_take_whole); + CU_ADD_TEST(suite, test_raid_process_with_qos); + + spdk_thread_lib_init(test_new_thread_fn, 0); diff --git a/patches/README.md b/patches/README.md index 69551c6..4e5d631 100644 --- a/patches/README.md +++ b/patches/README.md @@ -51,6 +51,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0044`) | 0041 | lvol → tier band usage provider — when an lvolstore loads on a `bdev_tier` composite the lvol layer registers a per-band fill-accounting provider, gated on the tier grain being a whole multiple of the lvolstore cluster (else no provider, honest unknown), and **unregisters it at teardown INITIATION** (before `spdk_lvs_unload`/`destroy`, whose synchronous store free would otherwise race a `get_bands` poll — UAF); `bdev_tier_get_bands` then fills `used_blocks` by counting allocated clusters of the blobstore's `used_clusters` pool in the band's LBA range (`spdk_bs_count_allocated_clusters_in_lba_range`, new in lib/blob — exclusive cluster-aligned bounds, one masked popcount per 64 clusters via `spdk_bit_pool_count_allocated_in_range` / `spdk_bit_array_count_set_in_range`, new in lib/util) | lib/blob, lib/util (ranged popcount), bdev_lvol; `#include`s `vbdev_tier.h` (requires the tier module, like 0005) | 0004 (bit-pool accessors), 0005 (lvol `-I` tier CFLAGS), tier module | | 0044 | a thin cluster is cleared before use — a cluster allocated by a write (or write_zeroes, or an inflate) that covers only part of it is zeroed before the blob sees it. Its unwritten part returned what the device held there: another blob's data wherever the device's unmap does not zero (the free path counts on unmap to), and different bytes on each leg of a mirror, which a verify of the mirror reads as a content divergence. A write that covers the whole cluster skips the clear, so a rebuild or a copy does not write twice. `blob_ut` pins both, and its thin-provisioning I/O accounting now counts the cleared cluster | lib/blob (blobstore.c), test/unit/lib/blob | — | | 0045 | a seeded rebuild copies its seed ranges, and only under the windows they need — a seeded process jumps over what its ranges leave clean with no window (only a copy needs a window's quiesce, and the seeded target takes writes on either side of the process offset), and a window copies every range part it holds under its one lock. A window that a range touched used to be copied whole (scattered writes: 5 GiB re-copied for 395 MiB of dirty chunks), and every clean window was locked and skipped in turn (262144 cycles on a 1 TiB volume, whatever its delta). `bdev_raid_ut` pins the blocks copied and the windows locked | bdev_raid, test/unit/lib/bdev/raid | 0013 | +| 0046 | a seeded rebuild reads its ranges from the CBT epoch — `bdev_raid_start_seeded_rebuild_from_epoch {name, base_bdev, cbt_bdev, epoch_id}` seeds the rebuild of `bdev_raid_start_seeded_rebuild` with the dirty ranges of a frozen epoch, read in-process through `vbdev_cbt_query_epoch_ranges` (up to 65536; more is -E2BIG, the caller rebuilds in full). A request carries ~200 ranges (1024 parsed values, a 32 KiB receive buffer): a scattered delta folded to fit recopied every clean block in between, 3.2 GiB for 4485 dirty ranges (368 MiB), measured. Both RPCs share their checks. `bdev_raid_ut` fakes the query | bdev_raid (rpc), module/bdev/cbt (query), test/unit/lib/bdev/raid | 0013, 0045, cbt module | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From f82970b1fcb0be2135a89bee9232f4e36ec000bb Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 06:30:20 +0200 Subject: [PATCH 09/26] test(raid1): a rebuild copies its window in parts, in flight together (0047, tests only) Two cases, red on the series as it stands: - raid1_ut: a 64-block window offered as the generic layer offers it (the rest of the window to the next request) is copied by four requests of at most the part, every part read before any completes, the reads tiling the window and spread evenly over the members in sync, and each part's read completing writes that part to the target. raid1 takes the window whole in one request, and on three members the second one serves no read; - bdev_raid_ut: raid_bdev_process_request_max_part offers parts that are whole write units and all fit RAID_BDEV_PROCESS_MAX_QD requests, smaller than any window with room to split. It offers the window whole. raid1_ut now inits a process request's raid_io for real (its target names the raid) and records every I/O the module submits. --- ...1-rebuild-copies-its-window-in-parts.patch | 304 ++++++++++++++++++ patches/README.md | 3 +- 2 files changed, 306 insertions(+), 1 deletion(-) create mode 100644 patches/0047-raid1-rebuild-copies-its-window-in-parts.patch diff --git a/patches/0047-raid1-rebuild-copies-its-window-in-parts.patch b/patches/0047-raid1-rebuild-copies-its-window-in-parts.patch new file mode 100644 index 0000000..d469caf --- /dev/null +++ b/patches/0047-raid1-rebuild-copies-its-window-in-parts.patch @@ -0,0 +1,304 @@ +diff --git a/module/bdev/raid/bdev_raid.c b/module/bdev/raid/bdev_raid.c +index 583707e..d34fbf9 100644 +--- a/module/bdev/raid/bdev_raid.c ++++ b/module/bdev/raid/bdev_raid.c +@@ -3266,6 +3266,12 @@ raid_bdev_process_request_complete(struct raid_bdev_process_request *process_req + } + } + ++uint32_t ++raid_bdev_process_request_max_part(const struct raid_bdev_process_request *process_req) ++{ ++ return process_req->process->max_window_size; ++} ++ + static int + raid_bdev_submit_process_request(struct raid_bdev_process *process, uint64_t offset_blocks, + uint32_t num_blocks) +diff --git a/module/bdev/raid/bdev_raid.h b/module/bdev/raid/bdev_raid.h +index 280d61c..f584d19 100644 +--- a/module/bdev/raid/bdev_raid.h ++++ b/module/bdev/raid/bdev_raid.h +@@ -481,6 +481,10 @@ void *raid_bdev_channel_get_module_ctx(struct raid_bdev_io_channel *raid_ch); + struct raid_base_bdev_info *raid_bdev_channel_get_base_info(struct raid_bdev_io_channel *raid_ch, + struct spdk_bdev *base_bdev); + void raid_bdev_process_request_complete(struct raid_bdev_process_request *process_req, int status); ++/* Evariops 0047: the largest part of its window one process request should copy, ++ * for a module that copies any range (raid1). A module bound to a stripe (raid5f) ++ * ignores it. */ ++uint32_t raid_bdev_process_request_max_part(const struct raid_bdev_process_request *process_req); + void raid_bdev_io_init(struct raid_bdev_io *raid_io, struct raid_bdev_io_channel *raid_ch, + enum spdk_bdev_io_type type, uint64_t offset_blocks, + uint64_t num_blocks, struct iovec *iovs, int iovcnt, void *md_buf, +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index 11b46f2..a488d6d 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -1765,6 +1765,35 @@ test_raid_process_seeded_skips_a_clean_span_without_a_window(void) + CU_ASSERT(g_lowest_copy_window == 100); + } + ++/* Evariops 0047: a window is offered to its requests in parts that all fit at once, ++ * so its copy is in flight together, and a window with room for it is split. The ++ * largest part is whole write units; the window's last part is what remains. */ ++static void ++test_raid_process_request_max_part(void) ++{ ++ struct raid_bdev raid_bdev = {}; ++ struct raid_bdev_process process = { .raid_bdev = &raid_bdev }; ++ struct raid_bdev_process_request process_req = { .process = &process }; ++ const uint64_t windows[] = { 8192, 2048, 1000, 17, 3, 1 }; ++ const uint32_t write_units[] = { 1, 8 }; ++ uint32_t w, u, part; ++ ++ for (u = 0; u < SPDK_COUNTOF(write_units); u++) { ++ raid_bdev.bdev.write_unit_size = write_units[u]; ++ for (w = 0; w < SPDK_COUNTOF(windows); w++) { ++ process.max_window_size = windows[w]; ++ part = raid_bdev_process_request_max_part(&process_req); ++ ++ CU_ASSERT(part > 0); ++ CU_ASSERT(part % write_units[u] == 0); ++ CU_ASSERT((uint64_t)part * RAID_BDEV_PROCESS_MAX_QD >= windows[w]); ++ if (windows[w] >= 2ULL * RAID_BDEV_PROCESS_MAX_QD * write_units[u]) { ++ CU_ASSERT(part < windows[w]); ++ } ++ } ++ } ++} ++ + static void + test_raid_process_with_qos(void) + { +@@ -2176,6 +2205,7 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_raid_process_seeded_skips_a_clean_span_without_a_window); + CU_ADD_TEST(suite, test_raid_seeded_rebuild_from_epoch_copies_every_range); + CU_ADD_TEST(suite, test_raid_seeded_rebuild_from_epoch_refuses_a_delta_it_cannot_take_whole); ++ CU_ADD_TEST(suite, test_raid_process_request_max_part); + CU_ADD_TEST(suite, test_raid_process_with_qos); + + spdk_thread_lib_init(test_new_thread_fn, 0); +diff --git a/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c b/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c +index 9c3eb8f..88f4295 100644 +--- a/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c ++++ b/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c +@@ -23,11 +23,6 @@ DEFINE_STUB_V(raid_bdev_queue_io_wait, (struct raid_bdev_io *raid_io, struct spd + struct spdk_io_channel *ch, spdk_bdev_io_wait_cb cb_fn)); + DEFINE_STUB_V(raid_bdev_process_request_complete, (struct raid_bdev_process_request *process_req, + int status)); +-DEFINE_STUB_V(raid_bdev_io_init, (struct raid_bdev_io *raid_io, +- struct raid_bdev_io_channel *raid_ch, +- enum spdk_bdev_io_type type, uint64_t offset_blocks, +- uint64_t num_blocks, struct iovec *iovs, int iovcnt, void *md_buf, +- struct spdk_memory_domain *memory_domain, void *memory_domain_ctx)); + DEFINE_STUB(raid_bdev_remap_dix_reftag, int, (void *md_buf, uint64_t num_blocks, + struct spdk_bdev *bdev, uint32_t remapped_offset), -1); + DEFINE_STUB(spdk_bdev_notify_blockcnt_change, int, (struct spdk_bdev *bdev, uint64_t size), 0); +@@ -38,14 +33,41 @@ DEFINE_STUB(spdk_bdev_unmap_blocks, int, (struct spdk_bdev_desc *desc, struct sp + uint64_t offset_blocks, uint64_t num_blocks, spdk_bdev_io_completion_cb cb, + void *cb_arg), 0); + ++/* Evariops 0047: the I/O the module submitted, in order, so a test reads back which ++ * member served a read and where a write landed, and completes each in turn. */ ++struct ut_io { ++ bool write; ++ struct spdk_bdev_desc *desc; ++ uint64_t offset_blocks; ++ uint64_t num_blocks; ++ spdk_bdev_io_completion_cb cb; ++ void *cb_arg; ++}; ++static struct ut_io g_ios[32]; ++static int g_num_ios; ++ ++static void ++ut_record_io(bool write, struct spdk_bdev_desc *desc, uint64_t offset_blocks, uint64_t num_blocks, ++ spdk_bdev_io_completion_cb cb, void *cb_arg) ++{ ++ g_last_io_desc = desc; ++ g_last_io_cb = cb; ++ ++ if (g_num_ios < (int)SPDK_COUNTOF(g_ios)) { ++ g_ios[g_num_ios++] = (struct ut_io) { ++ .write = write, .desc = desc, .offset_blocks = offset_blocks, ++ .num_blocks = num_blocks, .cb = cb, .cb_arg = cb_arg, ++ }; ++ } ++} ++ + int + spdk_bdev_readv_blocks_ext(struct spdk_bdev_desc *desc, + struct spdk_io_channel *ch, + struct iovec *iov, int iovcnt, uint64_t offset_blocks, uint64_t num_blocks, + spdk_bdev_io_completion_cb cb, void *cb_arg, struct spdk_bdev_ext_io_opts *opts) + { +- g_last_io_desc = desc; +- g_last_io_cb = cb; ++ ut_record_io(false, desc, offset_blocks, num_blocks, cb, cb_arg); + + return 0; + } +@@ -56,12 +78,34 @@ spdk_bdev_writev_blocks_ext(struct spdk_bdev_desc *desc, + struct iovec *iov, int iovcnt, uint64_t offset_blocks, uint64_t num_blocks, + spdk_bdev_io_completion_cb cb, void *cb_arg, struct spdk_bdev_ext_io_opts *opts) + { +- g_last_io_desc = desc; +- g_last_io_cb = cb; ++ ut_record_io(true, desc, offset_blocks, num_blocks, cb, cb_arg); + + return 0; + } + ++/* raid1 inits a raid_io only for a process request, whose target names the raid: ++ * enough to init it for real, so the copy it submits runs through the module. */ ++void ++raid_bdev_io_init(struct raid_bdev_io *raid_io, struct raid_bdev_io_channel *raid_ch, ++ enum spdk_bdev_io_type type, uint64_t offset_blocks, ++ uint64_t num_blocks, struct iovec *iovs, int iovcnt, void *md_buf, ++ struct spdk_memory_domain *memory_domain, void *memory_domain_ctx) ++{ ++ struct raid_bdev_process_request *process_req = ++ SPDK_CONTAINEROF(raid_io, struct raid_bdev_process_request, raid_io); ++ ++ raid_test_bdev_io_init(raid_io, process_req->target->raid_bdev, raid_ch, type, ++ offset_blocks, num_blocks, iovs, iovcnt, md_buf); ++} ++ ++static uint32_t g_process_max_part = UINT32_MAX; ++ ++uint32_t ++raid_bdev_process_request_max_part(const struct raid_bdev_process_request *process_req) ++{ ++ return g_process_max_part; ++} ++ + void + raid_bdev_fail_base_bdev(struct raid_base_bdev_info *base_info) + { +@@ -274,6 +318,118 @@ test_raid1_read_balancing(void) + run_for_each_raid1_config(_test_raid1_read_balancing); + } + ++static int ++ut_member_of(struct raid_bdev *raid_bdev, struct spdk_bdev_desc *desc) ++{ ++ uint8_t i; ++ ++ for (i = 0; i < raid_bdev->num_base_bdevs; i++) { ++ if (raid_bdev->base_bdev_info[i].desc == desc) { ++ return i; ++ } ++ } ++ return -1; ++} ++ ++/* Evariops 0047: a rebuild window is copied in parts, all in flight together, and ++ * the parts' reads spread over the members in sync. Copied as ONE request, the ++ * window was read whole from the first member before any of it was written: read ++ * and write never overlapped, and on a 3-node bench every byte of a 5 GiB rebuild crossed ++ * the link from a remote member (36 MiB/s) while the nexus's local member idled. ++ * The target is the last member: the process channel has no read channel for it. */ ++static void ++_test_raid1_process_window_in_parts(struct raid_bdev *raid_bdev, struct raid_bdev_io_channel *raid_ch) ++{ ++ struct raid1_io_channel *raid1_ch = raid_bdev_channel_get_module_ctx(raid_ch); ++ struct raid_bdev_process_request reqs[4] = {}; ++ const uint8_t target = raid_bdev->num_base_bdevs - 1; ++ const uint64_t window = 64; ++ uint32_t served[8] = {}; ++ uint64_t offset = 0; ++ int num_reqs = 0, num_reads, i; ++ ++ if (raid_bdev->bdev.blockcnt < window) { ++ return; ++ } ++ ++ g_process_max_part = window / SPDK_COUNTOF(reqs); ++ g_num_ios = 0; ++ raid_ch->_base_channels[target] = NULL; ++ ++ /* As the generic layer offers it: the rest of the window, to the next request. */ ++ while (offset < window && num_reqs < (int)SPDK_COUNTOF(reqs)) { ++ struct raid_bdev_process_request *req = &reqs[num_reqs++]; ++ int ret; ++ ++ req->target = &raid_bdev->base_bdev_info[target]; ++ req->target_ch = (void *)1; ++ req->offset_blocks = offset; ++ req->num_blocks = window - offset; ++ req->iov.iov_len = req->num_blocks * raid_bdev->bdev.blocklen; ++ ++ ret = raid1_submit_process_request(req, raid_ch); ++ CU_ASSERT(ret > 0 && (uint32_t)ret <= g_process_max_part); ++ CU_ASSERT(req->num_blocks == (uint32_t)ret); ++ CU_ASSERT(req->iov.iov_len == (uint64_t)ret * raid_bdev->bdev.blocklen); ++ if (ret <= 0) { ++ break; ++ } ++ offset += ret; ++ } ++ CU_ASSERT(offset == window); ++ CU_ASSERT(num_reqs == (int)SPDK_COUNTOF(reqs)); ++ ++ /* Every part is read before any completes, each from a member in sync, and the ++ * reads tile the window; members in sync serve as many parts each. */ ++ num_reads = g_num_ios; ++ CU_ASSERT(num_reads == num_reqs); ++ offset = 0; ++ for (i = 0; i < num_reads; i++) { ++ int member = ut_member_of(raid_bdev, g_ios[i].desc); ++ ++ CU_ASSERT(!g_ios[i].write); ++ CU_ASSERT(member >= 0 && member != target); ++ CU_ASSERT(g_ios[i].offset_blocks == offset); ++ offset += g_ios[i].num_blocks; ++ if (member >= 0) { ++ served[member]++; ++ } ++ } ++ CU_ASSERT(offset == window); ++ for (i = 0; i < target; i++) { ++ CU_ASSERT(served[i] == (uint32_t)num_reqs / target); ++ } ++ ++ /* Each part's read completing writes that part to the target, and only it. */ ++ for (i = 0; i < num_reads; i++) { ++ struct spdk_bdev_io bdev_io = {}; ++ ++ g_ios[i].cb(&bdev_io, true, g_ios[i].cb_arg); ++ CU_ASSERT(g_num_ios == num_reads + i + 1); ++ CU_ASSERT(g_ios[num_reads + i].write); ++ CU_ASSERT(ut_member_of(raid_bdev, g_ios[num_reads + i].desc) == target); ++ CU_ASSERT(g_ios[num_reads + i].offset_blocks == g_ios[i].offset_blocks); ++ CU_ASSERT(g_ios[num_reads + i].num_blocks == g_ios[i].num_blocks); ++ } ++ for (i = num_reads; i < g_num_ios; i++) { ++ struct spdk_bdev_io bdev_io = {}; ++ ++ g_ios[i].cb(&bdev_io, true, g_ios[i].cb_arg); ++ } ++ for (i = 0; i < raid_bdev->num_base_bdevs; i++) { ++ CU_ASSERT(raid1_ch->read_blocks_outstanding[i] == 0); ++ } ++ ++ raid_ch->_base_channels[target] = (void *)1; ++ g_process_max_part = UINT32_MAX; ++} ++ ++static void ++test_raid1_process_window_in_parts(void) ++{ ++ run_for_each_raid1_config(_test_raid1_process_window_in_parts); ++} ++ + static void + _test_raid1_write_error(struct raid_bdev *raid_bdev, struct raid_bdev_io_channel *raid_ch) + { +@@ -572,6 +728,7 @@ main(int argc, char **argv) + suite = CU_add_suite("raid1", test_setup, test_cleanup); + CU_ADD_TEST(suite, test_raid1_start); + CU_ADD_TEST(suite, test_raid1_read_balancing); ++ CU_ADD_TEST(suite, test_raid1_process_window_in_parts); + CU_ADD_TEST(suite, test_raid1_write_error); + CU_ADD_TEST(suite, test_raid1_write_capacity_exceeded); + CU_ADD_TEST(suite, test_raid1_read_error); diff --git a/patches/README.md b/patches/README.md index 4e5d631..561f704 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0044`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0047`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -52,6 +52,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0044`) | 0044 | a thin cluster is cleared before use — a cluster allocated by a write (or write_zeroes, or an inflate) that covers only part of it is zeroed before the blob sees it. Its unwritten part returned what the device held there: another blob's data wherever the device's unmap does not zero (the free path counts on unmap to), and different bytes on each leg of a mirror, which a verify of the mirror reads as a content divergence. A write that covers the whole cluster skips the clear, so a rebuild or a copy does not write twice. `blob_ut` pins both, and its thin-provisioning I/O accounting now counts the cleared cluster | lib/blob (blobstore.c), test/unit/lib/blob | — | | 0045 | a seeded rebuild copies its seed ranges, and only under the windows they need — a seeded process jumps over what its ranges leave clean with no window (only a copy needs a window's quiesce, and the seeded target takes writes on either side of the process offset), and a window copies every range part it holds under its one lock. A window that a range touched used to be copied whole (scattered writes: 5 GiB re-copied for 395 MiB of dirty chunks), and every clean window was locked and skipped in turn (262144 cycles on a 1 TiB volume, whatever its delta). `bdev_raid_ut` pins the blocks copied and the windows locked | bdev_raid, test/unit/lib/bdev/raid | 0013 | | 0046 | a seeded rebuild reads its ranges from the CBT epoch — `bdev_raid_start_seeded_rebuild_from_epoch {name, base_bdev, cbt_bdev, epoch_id}` seeds the rebuild of `bdev_raid_start_seeded_rebuild` with the dirty ranges of a frozen epoch, read in-process through `vbdev_cbt_query_epoch_ranges` (up to 65536; more is -E2BIG, the caller rebuilds in full). A request carries ~200 ranges (1024 parsed values, a 32 KiB receive buffer): a scattered delta folded to fit recopied every clean block in between, 3.2 GiB for 4485 dirty ranges (368 MiB), measured. Both RPCs share their checks. `bdev_raid_ut` fakes the query | bdev_raid (rpc), module/bdev/cbt (query), test/unit/lib/bdev/raid | 0013, 0045, cbt module | +| 0047 | a raid1 rebuild copies its window in parts — `raid_bdev_process_request_max_part` offers a window to the process's requests in `RAID_BDEV_PROCESS_MAX_QD` parts (whole write units), and raid1 copies one part per request, so a window's parts are in flight together: each part's write starts as its own read completes, and the reads spread over the members in sync. Copied as one request, a window was read whole from one member, then written whole: read and write never overlapped, and the first member served every read — a 5 GiB rebuild ran at 36 MiB/s, all of it read across the link from a remote member while the nexus's local member idled. raid5f, bound to its stripe, ignores the part. `raid1_ut` pins the parts, their spread and their writes; `bdev_raid_ut` the part's bounds | bdev_raid, raid1, test/unit/lib/bdev/raid | 0013 | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From 3ec19eec2a94ecbb889eeeaace47e944e925d8a1 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 06:31:05 +0200 Subject: [PATCH 10/26] fix(raid1): a rebuild copies its window in parts, in flight together (0047) raid_bdev_process_request_max_part offers a window to the process's requests in RAID_BDEV_PROCESS_MAX_QD parts, whole write units, and raid1 copies one part per request. A window's parts are now in flight together: each part's write starts as its own read completes, and the reads spread over the members in sync by their outstanding blocks. Copied as one request, a window was read whole from one member, then written whole: read and write never overlapped, and the first member served every read. Measured on a 3-node bench, a full rebuild of a 5 GiB mirror3 under a verified write load ran at 36 MiB/s for 141 s, every byte read across the link from a remote member while the nexus's local member served none. raid5f, bound to its stripe, does not call it; the seeded copy already takes what the module returns and offers the rest to the next request. --- ...1-rebuild-copies-its-window-in-parts.patch | 33 +++++++++++++++++-- 1 file changed, 30 insertions(+), 3 deletions(-) diff --git a/patches/0047-raid1-rebuild-copies-its-window-in-parts.patch b/patches/0047-raid1-rebuild-copies-its-window-in-parts.patch index d469caf..c0d04b6 100644 --- a/patches/0047-raid1-rebuild-copies-its-window-in-parts.patch +++ b/patches/0047-raid1-rebuild-copies-its-window-in-parts.patch @@ -1,15 +1,24 @@ diff --git a/module/bdev/raid/bdev_raid.c b/module/bdev/raid/bdev_raid.c -index 583707e..d34fbf9 100644 +index 583707e..da919f1 100644 --- a/module/bdev/raid/bdev_raid.c +++ b/module/bdev/raid/bdev_raid.c -@@ -3266,6 +3266,12 @@ raid_bdev_process_request_complete(struct raid_bdev_process_request *process_req +@@ -3266,6 +3266,21 @@ raid_bdev_process_request_complete(struct raid_bdev_process_request *process_req } } ++/* Evariops 0047: a window is offered to its requests in RAID_BDEV_PROCESS_MAX_QD ++ * parts, so all of them are in flight together: each part's write starts as its ++ * own read completes, and the reads spread over the members that serve them. A ++ * window copied as one request was read whole, then written whole — read and ++ * write never overlapped, and one member served every read. */ +uint32_t +raid_bdev_process_request_max_part(const struct raid_bdev_process_request *process_req) +{ -+ return process_req->process->max_window_size; ++ const struct raid_bdev_process *process = process_req->process; ++ const uint64_t write_unit = spdk_max(process->raid_bdev->bdev.write_unit_size, 1); ++ const uint64_t part = spdk_divide_round_up(process->max_window_size, RAID_BDEV_PROCESS_MAX_QD); ++ ++ return spdk_divide_round_up(part, write_unit) * write_unit; +} + static int @@ -30,6 +39,24 @@ index 280d61c..f584d19 100644 void raid_bdev_io_init(struct raid_bdev_io *raid_io, struct raid_bdev_io_channel *raid_ch, enum spdk_bdev_io_type type, uint64_t offset_blocks, uint64_t num_blocks, struct iovec *iovs, int iovcnt, void *md_buf, +diff --git a/module/bdev/raid/raid1.c b/module/bdev/raid/raid1.c +index f1d4f0f..51efb51 100644 +--- a/module/bdev/raid/raid1.c ++++ b/module/bdev/raid/raid1.c +@@ -590,6 +590,13 @@ raid1_submit_process_request(struct raid_bdev_process_request *process_req, + struct raid_bdev_io *raid_io = &process_req->raid_io; + int ret; + ++ /* Evariops 0047: one part of the window, not the window whole: the next ++ * requests take the rest, in flight beside this one. */ ++ process_req->num_blocks = spdk_min(process_req->num_blocks, ++ raid_bdev_process_request_max_part(process_req)); ++ process_req->iov.iov_len = (size_t)process_req->num_blocks * ++ process_req->target->raid_bdev->bdev.blocklen; ++ + raid_bdev_io_init(raid_io, raid_ch, SPDK_BDEV_IO_TYPE_READ, + process_req->offset_blocks, process_req->num_blocks, + &process_req->iov, 1, process_req->md_buf, NULL, NULL); diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c index 11b46f2..a488d6d 100644 --- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c From e961ea2d0368990984092ff0876bf8036f195eab Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 07:49:40 +0200 Subject: [PATCH 11/26] test(raid1): a rebuild writes a zero part as write-zeroes (0048, tests only) raid1_ut, red on the series as it stands: four parts of a rebuild window, two read back holding data and two holding nothing but zeroes. With a target that takes write-zeroes, the data parts are written as data and the zero parts as write-zeroes; with a target that cannot, every part is written as data. The zero parts go out as data writes. The window-in-parts test (0047) now gives its requests real buffers: the copy reads a part back before choosing how to write it. The suite records the write-zeroes it is sent, and the write-zeroes support of the members is a knob of the suite. --- ...d-writes-a-zero-part-as-write-zeroes.patch | 156 ++++++++++++++++++ patches/README.md | 3 +- 2 files changed, 158 insertions(+), 1 deletion(-) create mode 100644 patches/0048-raid1-rebuild-writes-a-zero-part-as-write-zeroes.patch diff --git a/patches/0048-raid1-rebuild-writes-a-zero-part-as-write-zeroes.patch b/patches/0048-raid1-rebuild-writes-a-zero-part-as-write-zeroes.patch new file mode 100644 index 0000000..51d5d0d --- /dev/null +++ b/patches/0048-raid1-rebuild-writes-a-zero-part-as-write-zeroes.patch @@ -0,0 +1,156 @@ +diff --git a/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c b/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c +index 88f4295..cad171b 100644 +--- a/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c ++++ b/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c +@@ -37,6 +37,7 @@ DEFINE_STUB(spdk_bdev_unmap_blocks, int, (struct spdk_bdev_desc *desc, struct sp + * member served a read and where a write landed, and completes each in turn. */ + struct ut_io { + bool write; ++ bool zeroes; + struct spdk_bdev_desc *desc; + uint64_t offset_blocks; + uint64_t num_blocks; +@@ -83,6 +84,26 @@ spdk_bdev_writev_blocks_ext(struct spdk_bdev_desc *desc, + return 0; + } + ++/* Evariops 0048: whether the members take write-zeroes, and the write-zeroes they were sent. */ ++static bool g_write_zeroes_supported = true; ++ ++bool ++spdk_bdev_io_type_supported(struct spdk_bdev *bdev, enum spdk_bdev_io_type io_type) ++{ ++ return io_type == SPDK_BDEV_IO_TYPE_WRITE_ZEROES ? g_write_zeroes_supported : true; ++} ++ ++int ++spdk_bdev_write_zeroes_blocks(struct spdk_bdev_desc *desc, struct spdk_io_channel *ch, ++ uint64_t offset_blocks, uint64_t num_blocks, ++ spdk_bdev_io_completion_cb cb, void *cb_arg) ++{ ++ ut_record_io(true, desc, offset_blocks, num_blocks, cb, cb_arg); ++ g_ios[g_num_ios - 1].zeroes = true; ++ ++ return 0; ++} ++ + /* raid1 inits a raid_io only for a process request, whose target names the raid: + * enough to init it for real, so the copy it submits runs through the module. */ + void +@@ -361,6 +382,9 @@ _test_raid1_process_window_in_parts(struct raid_bdev *raid_bdev, struct raid_bde + struct raid_bdev_process_request *req = &reqs[num_reqs++]; + int ret; + ++ /* A real buffer: the copy reads it back before choosing how to write it (0048). */ ++ req->iov.iov_base = calloc(1, window * raid_bdev->bdev.blocklen); ++ SPDK_CU_ASSERT_FATAL(req->iov.iov_base != NULL); + req->target = &raid_bdev->base_bdev_info[target]; + req->target_ch = (void *)1; + req->offset_blocks = offset; +@@ -419,6 +443,9 @@ _test_raid1_process_window_in_parts(struct raid_bdev *raid_bdev, struct raid_bde + for (i = 0; i < raid_bdev->num_base_bdevs; i++) { + CU_ASSERT(raid1_ch->read_blocks_outstanding[i] == 0); + } ++ for (i = 0; i < num_reqs; i++) { ++ free(reqs[i].iov.iov_base); ++ } + + raid_ch->_base_channels[target] = (void *)1; + g_process_max_part = UINT32_MAX; +@@ -430,6 +457,88 @@ test_raid1_process_window_in_parts(void) + run_for_each_raid1_config(_test_raid1_process_window_in_parts); + } + ++/* Evariops 0048: a rebuild part the source read back as nothing but zeroes is written to the ++ * target as write-zeroes, so no payload crosses the link for it; a part holding data is written ++ * as data. A target that cannot take write-zeroes gets every part as data. */ ++static void ++_test_raid1_process_zero_part_written_as_zeroes(struct raid_bdev *raid_bdev, struct raid_bdev_io_channel *raid_ch) ++{ ++ struct raid_bdev_process_request reqs[4]; ++ const uint8_t target = raid_bdev->num_base_bdevs - 1; ++ const uint64_t window = 64; ++ const uint64_t part = window / SPDK_COUNTOF(reqs); ++ const size_t part_len = part * raid_bdev->bdev.blocklen; ++ int pass, i; ++ ++ if (raid_bdev->bdev.blockcnt < window) { ++ return; ++ } ++ ++ for (pass = 0; pass < 2; pass++) { ++ const bool supported = pass == 0; ++ ++ g_write_zeroes_supported = supported; ++ g_process_max_part = part; ++ g_num_ios = 0; ++ raid_ch->_base_channels[target] = NULL; ++ ++ for (i = 0; i < (int)SPDK_COUNTOF(reqs); i++) { ++ struct raid_bdev_process_request *req = &reqs[i]; ++ ++ memset(req, 0, sizeof(*req)); ++ req->iov.iov_base = calloc(1, window * raid_bdev->bdev.blocklen); ++ SPDK_CU_ASSERT_FATAL(req->iov.iov_base != NULL); ++ req->target = &raid_bdev->base_bdev_info[target]; ++ req->target_ch = (void *)1; ++ req->offset_blocks = i * part; ++ req->num_blocks = window - i * part; ++ req->iov.iov_len = req->num_blocks * raid_bdev->bdev.blocklen; ++ CU_ASSERT(raid1_submit_process_request(req, raid_ch) == (int)part); ++ } ++ CU_ASSERT(g_num_ios == (int)SPDK_COUNTOF(reqs)); ++ ++ /* Parts 0 and 2 read data back, parts 1 and 3 nothing but zeroes. */ ++ for (i = 0; i < (int)SPDK_COUNTOF(reqs); i++) { ++ struct spdk_bdev_io bdev_io = {}; ++ ++ if (i % 2 == 0) { ++ memset(reqs[i].iov.iov_base, 0xa5, part_len); ++ } ++ g_ios[i].cb(&bdev_io, true, g_ios[i].cb_arg); ++ } ++ ++ CU_ASSERT(g_num_ios == 2 * (int)SPDK_COUNTOF(reqs)); ++ for (i = 0; i < (int)SPDK_COUNTOF(reqs); i++) { ++ struct ut_io *written = &g_ios[SPDK_COUNTOF(reqs) + i]; ++ ++ CU_ASSERT(written->write); ++ CU_ASSERT(ut_member_of(raid_bdev, written->desc) == target); ++ CU_ASSERT(written->offset_blocks == i * part); ++ CU_ASSERT(written->num_blocks == part); ++ CU_ASSERT(written->zeroes == (supported && i % 2 == 1)); ++ } ++ ++ for (i = SPDK_COUNTOF(reqs); i < g_num_ios; i++) { ++ struct spdk_bdev_io bdev_io = {}; ++ ++ g_ios[i].cb(&bdev_io, true, g_ios[i].cb_arg); ++ } ++ for (i = 0; i < (int)SPDK_COUNTOF(reqs); i++) { ++ free(reqs[i].iov.iov_base); ++ } ++ raid_ch->_base_channels[target] = (void *)1; ++ } ++ ++ g_write_zeroes_supported = true; ++ g_process_max_part = UINT32_MAX; ++} ++ ++static void ++test_raid1_process_zero_part_written_as_zeroes(void) ++{ ++ run_for_each_raid1_config(_test_raid1_process_zero_part_written_as_zeroes); ++} ++ + static void + _test_raid1_write_error(struct raid_bdev *raid_bdev, struct raid_bdev_io_channel *raid_ch) + { +@@ -729,6 +838,7 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_raid1_start); + CU_ADD_TEST(suite, test_raid1_read_balancing); + CU_ADD_TEST(suite, test_raid1_process_window_in_parts); ++ CU_ADD_TEST(suite, test_raid1_process_zero_part_written_as_zeroes); + CU_ADD_TEST(suite, test_raid1_write_error); + CU_ADD_TEST(suite, test_raid1_write_capacity_exceeded); + CU_ADD_TEST(suite, test_raid1_read_error); diff --git a/patches/README.md b/patches/README.md index 561f704..13772c3 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0047`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0048`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -53,6 +53,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0047`) | 0045 | a seeded rebuild copies its seed ranges, and only under the windows they need — a seeded process jumps over what its ranges leave clean with no window (only a copy needs a window's quiesce, and the seeded target takes writes on either side of the process offset), and a window copies every range part it holds under its one lock. A window that a range touched used to be copied whole (scattered writes: 5 GiB re-copied for 395 MiB of dirty chunks), and every clean window was locked and skipped in turn (262144 cycles on a 1 TiB volume, whatever its delta). `bdev_raid_ut` pins the blocks copied and the windows locked | bdev_raid, test/unit/lib/bdev/raid | 0013 | | 0046 | a seeded rebuild reads its ranges from the CBT epoch — `bdev_raid_start_seeded_rebuild_from_epoch {name, base_bdev, cbt_bdev, epoch_id}` seeds the rebuild of `bdev_raid_start_seeded_rebuild` with the dirty ranges of a frozen epoch, read in-process through `vbdev_cbt_query_epoch_ranges` (up to 65536; more is -E2BIG, the caller rebuilds in full). A request carries ~200 ranges (1024 parsed values, a 32 KiB receive buffer): a scattered delta folded to fit recopied every clean block in between, 3.2 GiB for 4485 dirty ranges (368 MiB), measured. Both RPCs share their checks. `bdev_raid_ut` fakes the query | bdev_raid (rpc), module/bdev/cbt (query), test/unit/lib/bdev/raid | 0013, 0045, cbt module | | 0047 | a raid1 rebuild copies its window in parts — `raid_bdev_process_request_max_part` offers a window to the process's requests in `RAID_BDEV_PROCESS_MAX_QD` parts (whole write units), and raid1 copies one part per request, so a window's parts are in flight together: each part's write starts as its own read completes, and the reads spread over the members in sync. Copied as one request, a window was read whole from one member, then written whole: read and write never overlapped, and the first member served every read — a 5 GiB rebuild ran at 36 MiB/s, all of it read across the link from a remote member while the nexus's local member idled. raid5f, bound to its stripe, ignores the part. `raid1_ut` pins the parts, their spread and their writes; `bdev_raid_ut` the part's bounds | bdev_raid, raid1, test/unit/lib/bdev/raid | 0013 | +| 0048 | a raid1 rebuild writes a zero part as write-zeroes — a part the source read back as nothing but zeroes goes to the target as `write_zeroes` (`raid_bdev_write_zeroes_blocks`): the same blocks, read back the same, and no payload on the wire, so a full rebuild of a sparse volume no longer ships its never-written clusters across the link (about 3 of the 5 GiB of the rebuild measured for 0047). Not when the target cannot take write-zeroes (the bdev layer would send a zero buffer anyway), nor when the blocks carry metadata. `raid1_ut` pins both kinds of write, and a target without write-zeroes getting data | raid1, bdev_raid (header), test/unit/lib/bdev/raid | 0047 | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From 8ed1d2a872f7dcfdc56197458a8ece9d060529df Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 07:49:57 +0200 Subject: [PATCH 12/26] fix(raid1): a rebuild writes a zero part as write-zeroes (0048) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A part the source read back as nothing but zeroes goes to the target as write_zeroes (raid_bdev_write_zeroes_blocks, the data offset applied like the other wrappers): the same blocks, read back the same, and no payload on the wire. A full rebuild of a sparse volume no longer ships its never-written clusters across the link — measured on a 3-node bench, about 3 of the 5 GiB of a mirror3 full rebuild, all of it zero, all of it through the nexus node's egress, which 0047 had made the bound. Not when the target cannot take write-zeroes (the bdev layer would then send a zero buffer anyway), nor when the blocks carry metadata. The check reads the part already in memory, and a data part stops at its first nonzero word. --- ...d-writes-a-zero-part-as-write-zeroes.patch | 83 +++++++++++++++++++ 1 file changed, 83 insertions(+) diff --git a/patches/0048-raid1-rebuild-writes-a-zero-part-as-write-zeroes.patch b/patches/0048-raid1-rebuild-writes-a-zero-part-as-write-zeroes.patch index 51d5d0d..3b36cf4 100644 --- a/patches/0048-raid1-rebuild-writes-a-zero-part-as-write-zeroes.patch +++ b/patches/0048-raid1-rebuild-writes-a-zero-part-as-write-zeroes.patch @@ -1,3 +1,86 @@ +diff --git a/module/bdev/raid/bdev_raid.h b/module/bdev/raid/bdev_raid.h +index f584d19..947359c 100644 +--- a/module/bdev/raid/bdev_raid.h ++++ b/module/bdev/raid/bdev_raid.h +@@ -560,6 +560,18 @@ raid_bdev_unmap_blocks(struct raid_base_bdev_info *base_info, struct spdk_io_cha + num_blocks, cb, cb_arg); + } + ++/** ++ * Raid bdev I/O wrapper for spdk_bdev_write_zeroes_blocks function. ++ */ ++static inline int ++raid_bdev_write_zeroes_blocks(struct raid_base_bdev_info *base_info, struct spdk_io_channel *ch, ++ uint64_t offset_blocks, uint64_t num_blocks, ++ spdk_bdev_io_completion_cb cb, void *cb_arg) ++{ ++ return spdk_bdev_write_zeroes_blocks(base_info->desc, ch, base_info->data_offset + offset_blocks, ++ num_blocks, cb, cb_arg); ++} ++ + /** + * Raid bdev I/O read/write wrapper for spdk_bdev_flush_blocks function. + */ +diff --git a/module/bdev/raid/raid1.c b/module/bdev/raid/raid1.c +index 51efb51..388b9a9 100644 +--- a/module/bdev/raid/raid1.c ++++ b/module/bdev/raid/raid1.c +@@ -7,6 +7,7 @@ + + #include "spdk/likely.h" + #include "spdk/log.h" ++#include "spdk/string.h" /* Evariops 0048: spdk_mem_all_zero */ + #include "spdk/nvme_spec.h" /* Evariops 0009: CAPACITY_EXCEEDED discrimination */ + + struct raid1_info { +@@ -547,6 +548,24 @@ _raid1_process_submit_write(void *ctx) + raid1_process_submit_write(process_req); + } + ++/* Evariops 0048: a part the source read back as nothing but zeroes goes to the target as ++ * write-zeroes — the same blocks, read back the same, and no payload on the wire, so a full ++ * rebuild of a sparse volume no longer ships its never-written clusters across the link. Not ++ * when the target cannot take write-zeroes (the bdev layer would send a zero buffer anyway), nor ++ * when the blocks carry metadata. */ ++static bool ++raid1_process_part_is_zeroes(struct raid_bdev_process_request *process_req) ++{ ++ struct raid_base_bdev_info *target = process_req->target; ++ ++ if (target->raid_bdev->bdev.md_len != 0 || ++ !spdk_bdev_io_type_supported(spdk_bdev_desc_get_bdev(target->desc), SPDK_BDEV_IO_TYPE_WRITE_ZEROES)) { ++ return false; ++ } ++ ++ return spdk_mem_all_zero(process_req->iov.iov_base, process_req->iov.iov_len); ++} ++ + static void + raid1_process_submit_write(struct raid_bdev_process_request *process_req) + { +@@ -554,11 +573,17 @@ raid1_process_submit_write(struct raid_bdev_process_request *process_req) + struct spdk_bdev_ext_io_opts io_opts; + int ret; + +- raid1_init_ext_io_opts(&io_opts, raid_io); +- ret = raid_bdev_writev_blocks_ext(process_req->target, process_req->target_ch, +- raid_io->iovs, raid_io->iovcnt, +- raid_io->offset_blocks, raid_io->num_blocks, +- raid1_process_write_completed, process_req, &io_opts); ++ if (raid1_process_part_is_zeroes(process_req)) { ++ ret = raid_bdev_write_zeroes_blocks(process_req->target, process_req->target_ch, ++ raid_io->offset_blocks, raid_io->num_blocks, ++ raid1_process_write_completed, process_req); ++ } else { ++ raid1_init_ext_io_opts(&io_opts, raid_io); ++ ret = raid_bdev_writev_blocks_ext(process_req->target, process_req->target_ch, ++ raid_io->iovs, raid_io->iovcnt, ++ raid_io->offset_blocks, raid_io->num_blocks, ++ raid1_process_write_completed, process_req, &io_opts); ++ } + if (spdk_unlikely(ret != 0)) { + if (ret == -ENOMEM) { + raid_bdev_queue_io_wait(raid_io, spdk_bdev_desc_get_bdev(process_req->target->desc), diff --git a/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c b/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c index 88f4295..cad171b 100644 --- a/test/unit/lib/bdev/raid/raid1.c/raid1_ut.c From fe0903e66b997c9bc85448e206550f481951627a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 08:24:19 +0200 Subject: [PATCH 13/26] test(blob): write-zeroes leaves an unallocated thin cluster unallocated (0049, tests only) blob_ut, red on the series as it stands: write-zeroes over half of a cluster a thin blob has not allocated allocates the cluster and writes to the device; it is to allocate nothing, write nothing, and read back zeroes. The same test pins what stays: write-zeroes over an allocated cluster holding data zeroes what it covers, and on a clone, whose unallocated cluster reads its parent's data, it allocates the cluster and zeroes it. --- ...unallocated-thin-cluster-unallocated.patch | 119 ++++++++++++++++++ patches/README.md | 3 +- 2 files changed, 121 insertions(+), 1 deletion(-) create mode 100644 patches/0049-blob-write-zeroes-leaves-an-unallocated-thin-cluster-unallocated.patch diff --git a/patches/0049-blob-write-zeroes-leaves-an-unallocated-thin-cluster-unallocated.patch b/patches/0049-blob-write-zeroes-leaves-an-unallocated-thin-cluster-unallocated.patch new file mode 100644 index 0000000..c603d81 --- /dev/null +++ b/patches/0049-blob-write-zeroes-leaves-an-unallocated-thin-cluster-unallocated.patch @@ -0,0 +1,119 @@ +diff --git a/test/unit/lib/blob/blob.c/blob_ut.c b/test/unit/lib/blob/blob.c/blob_ut.c +index 3bb3f5c..cdf55b1 100644 +--- a/test/unit/lib/blob/blob.c/blob_ut.c ++++ b/test/unit/lib/blob/blob.c/blob_ut.c +@@ -4944,6 +4944,106 @@ blob_thin_prov_full_cluster_write_is_not_cleared_first(void) + g_blobid = 0; + } + ++/* ++ * Evariops 0049: write-zeroes on a cluster that a thin blob backed by zeroes has not allocated ++ * changes nothing it reads, so it allocates nothing and writes nothing. A full rebuild sends every ++ * never-written part of its source as write-zeroes (0048): allocating the cluster of each made the ++ * rebuilt replica fully allocated, and wrote zeroes across the target's disk. An allocated cluster ++ * is still zeroed, and a clone, which reads its unallocated clusters from its parent, still ++ * allocates and zeroes. ++ */ ++static void ++blob_thin_prov_write_zeroes_on_unallocated_cluster_allocates_nothing(void) ++{ ++ struct spdk_blob_store *bs = g_bs; ++ struct spdk_blob *blob, *clone; ++ struct spdk_io_channel *channel; ++ struct spdk_blob_opts opts; ++ uint64_t io_unit_size = spdk_bs_get_io_unit_size(bs); ++ uint64_t cluster_size = spdk_bs_get_cluster_size(bs); ++ uint64_t io_units_per_cluster = cluster_size / io_unit_size; ++ uint64_t half = io_units_per_cluster / 2; ++ uint64_t write_bytes; ++ spdk_blob_id snapshotid; ++ uint8_t *payload; ++ ++ channel = spdk_bs_alloc_io_channel(bs); ++ SPDK_CU_ASSERT_FATAL(channel != NULL); ++ payload = malloc(cluster_size); ++ SPDK_CU_ASSERT_FATAL(payload != NULL); ++ ++ ut_fill_free_clusters(bs, 0xAA); ++ ++ ut_spdk_blob_opts_init(&opts); ++ opts.thin_provision = true; ++ opts.num_clusters = 2; ++ blob = ut_blob_create_and_open(bs, &opts); ++ ++ /* Half of cluster 0, never written: nothing allocated, nothing written, zeroes read back. */ ++ write_bytes = g_dev_write_bytes; ++ spdk_blob_io_write_zeroes(blob, channel, 0, half, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ CU_ASSERT(spdk_blob_get_num_allocated_clusters(blob) == 0); ++ CU_ASSERT(g_dev_write_bytes == write_bytes); ++ memset(payload, 0xFF, cluster_size); ++ spdk_blob_io_read(blob, channel, payload, 0, io_units_per_cluster, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ CU_ASSERT(spdk_mem_all_zero(payload, cluster_size)); ++ ++ /* An allocated cluster holding data: write-zeroes still zeroes what it covers. */ ++ memset(payload, 0xE5, cluster_size); ++ spdk_blob_io_write(blob, channel, payload, io_units_per_cluster, io_units_per_cluster, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ CU_ASSERT(spdk_blob_get_num_allocated_clusters(blob) == 1); ++ spdk_blob_io_write_zeroes(blob, channel, io_units_per_cluster, half, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ spdk_blob_io_read(blob, channel, payload, io_units_per_cluster, io_units_per_cluster, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ CU_ASSERT(spdk_mem_all_zero(payload, cluster_size / 2)); ++ CU_ASSERT(payload[cluster_size / 2] == 0xE5 && payload[cluster_size - 1] == 0xE5); ++ ++ /* A clone reads its unallocated cluster 1 from the snapshot, whose second half holds data: ++ * write-zeroes there allocates the cluster and zeroes it. */ ++ spdk_bs_create_snapshot(bs, spdk_blob_get_id(blob), NULL, blob_op_with_id_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ snapshotid = g_blobid; ++ spdk_bs_create_clone(bs, snapshotid, NULL, blob_op_with_id_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ spdk_bs_open_blob(bs, g_blobid, blob_op_with_handle_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ SPDK_CU_ASSERT_FATAL(g_blob != NULL); ++ clone = g_blob; ++ ++ spdk_blob_io_write_zeroes(clone, channel, io_units_per_cluster + half, half, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ CU_ASSERT(spdk_blob_get_num_allocated_clusters(clone) == 1); ++ memset(payload, 0xFF, cluster_size); ++ spdk_blob_io_read(clone, channel, payload, io_units_per_cluster, io_units_per_cluster, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ CU_ASSERT(spdk_mem_all_zero(payload, cluster_size)); ++ ++ free(payload); ++ ut_blob_close_and_delete(bs, clone); ++ ut_blob_close_and_delete(bs, blob); ++ spdk_bs_delete_blob(bs, snapshotid, blob_op_complete, NULL); ++ poll_threads(); ++ CU_ASSERT(g_bserrno == 0); ++ spdk_bs_free_io_channel(channel); ++ poll_threads(); ++ g_blob = NULL; ++ g_blobid = 0; ++} ++ + static void + blob_thin_prov_unmap_cluster(void) + { +@@ -10501,6 +10601,7 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, blob_thin_prov_write_count_io); + CU_ADD_TEST(suite_bs, blob_thin_prov_new_cluster_reads_zeroes); + CU_ADD_TEST(suite_bs, blob_thin_prov_full_cluster_write_is_not_cleared_first); ++ CU_ADD_TEST(suite_bs, blob_thin_prov_write_zeroes_on_unallocated_cluster_allocates_nothing); + CU_ADD_TEST(suite, blob_thin_prov_unmap_cluster); + CU_ADD_TEST(suite_bs, blob_thin_prov_rle); + CU_ADD_TEST(suite_bs, blob_thin_prov_rw_iov); diff --git a/patches/README.md b/patches/README.md index 13772c3..619c3cb 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0048`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0049`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -54,6 +54,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0048`) | 0046 | a seeded rebuild reads its ranges from the CBT epoch — `bdev_raid_start_seeded_rebuild_from_epoch {name, base_bdev, cbt_bdev, epoch_id}` seeds the rebuild of `bdev_raid_start_seeded_rebuild` with the dirty ranges of a frozen epoch, read in-process through `vbdev_cbt_query_epoch_ranges` (up to 65536; more is -E2BIG, the caller rebuilds in full). A request carries ~200 ranges (1024 parsed values, a 32 KiB receive buffer): a scattered delta folded to fit recopied every clean block in between, 3.2 GiB for 4485 dirty ranges (368 MiB), measured. Both RPCs share their checks. `bdev_raid_ut` fakes the query | bdev_raid (rpc), module/bdev/cbt (query), test/unit/lib/bdev/raid | 0013, 0045, cbt module | | 0047 | a raid1 rebuild copies its window in parts — `raid_bdev_process_request_max_part` offers a window to the process's requests in `RAID_BDEV_PROCESS_MAX_QD` parts (whole write units), and raid1 copies one part per request, so a window's parts are in flight together: each part's write starts as its own read completes, and the reads spread over the members in sync. Copied as one request, a window was read whole from one member, then written whole: read and write never overlapped, and the first member served every read — a 5 GiB rebuild ran at 36 MiB/s, all of it read across the link from a remote member while the nexus's local member idled. raid5f, bound to its stripe, ignores the part. `raid1_ut` pins the parts, their spread and their writes; `bdev_raid_ut` the part's bounds | bdev_raid, raid1, test/unit/lib/bdev/raid | 0013 | | 0048 | a raid1 rebuild writes a zero part as write-zeroes — a part the source read back as nothing but zeroes goes to the target as `write_zeroes` (`raid_bdev_write_zeroes_blocks`): the same blocks, read back the same, and no payload on the wire, so a full rebuild of a sparse volume no longer ships its never-written clusters across the link (about 3 of the 5 GiB of the rebuild measured for 0047). Not when the target cannot take write-zeroes (the bdev layer would send a zero buffer anyway), nor when the blocks carry metadata. `raid1_ut` pins both kinds of write, and a target without write-zeroes getting data | raid1, bdev_raid (header), test/unit/lib/bdev/raid | 0047 | +| 0049 | write-zeroes leaves an unallocated thin cluster unallocated — on a blob backed by zeroes, an unallocated cluster already reads zeroes, so write-zeroes over it completes with no cluster allocated and nothing written. A full rebuild sends every never-written part of its source as write-zeroes (0048), and allocating each made the rebuilt replica fully allocated — a thin volume rebuilt in full took its whole size on the target — and wrote zeroes across the target's disk. An allocated cluster is still zeroed; a clone, which reads its unallocated clusters from its parent, still allocates and zeroes. `blob_ut` pins the three | lib/blob (blobstore.c), test/unit/lib/blob | 0044 | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From f8ba4e77781d50ba797fff4bef65f3a81c454443 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 08:24:37 +0200 Subject: [PATCH 14/26] fix(blob): write-zeroes leaves an unallocated thin cluster unallocated (0049) On a blob backed by zeroes, an unallocated cluster already reads zeroes: write-zeroes over it now completes at once, with no cluster allocated and nothing written. Before, it allocated the cluster (cleared, since 0044) and zeroed it. A full rebuild sends every never-written part of its source as write-zeroes (0048), so a thin volume rebuilt in full took its whole size on the target, and the zeroes it carried no longer on the wire still went to the target's disk: measured on a 3-node bench, the SATA disk of the rebuilt replica wrote 94 MiB/s through a 5 GiB rebuild with about 3 GiB never written. An allocated cluster is still zeroed. A clone or any blob with another back device reads its unallocated clusters from it, and keeps the allocation. --- ...unallocated-thin-cluster-unallocated.patch | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) diff --git a/patches/0049-blob-write-zeroes-leaves-an-unallocated-thin-cluster-unallocated.patch b/patches/0049-blob-write-zeroes-leaves-an-unallocated-thin-cluster-unallocated.patch index c603d81..da4a95b 100644 --- a/patches/0049-blob-write-zeroes-leaves-an-unallocated-thin-cluster-unallocated.patch +++ b/patches/0049-blob-write-zeroes-leaves-an-unallocated-thin-cluster-unallocated.patch @@ -1,3 +1,22 @@ +diff --git a/lib/blob/blobstore.c b/lib/blob/blobstore.c +index 4d656c9..b3105f4 100644 +--- a/lib/blob/blobstore.c ++++ b/lib/blob/blobstore.c +@@ -3255,6 +3255,14 @@ blob_request_submit_op_single(struct spdk_io_channel *_ch, struct spdk_blob *blo + } + + bs_batch_close(batch); ++ } else if (op_type == SPDK_BLOB_WRITE_ZEROES && blob_backed_with_zeroes_dev(blob)) { ++ /* Evariops 0049: an unallocated cluster of a blob backed by zeroes already reads ++ * zeroes, so zeroing it changes nothing: no cluster is allocated and nothing is ++ * written. A full rebuild sends every never-written part of its source as ++ * write-zeroes, and allocating each made the rebuilt replica fully allocated. A ++ * clone reads its unallocated clusters from its parent and keeps the allocation. */ ++ cb_fn(cb_arg, 0); ++ return; + } else { + /* Queue this operation and allocate the cluster */ + spdk_bs_user_op_t *op; diff --git a/test/unit/lib/blob/blob.c/blob_ut.c b/test/unit/lib/blob/blob.c/blob_ut.c index 3bb3f5c..cdf55b1 100644 --- a/test/unit/lib/blob/blob.c/blob_ut.c From a55abb1cf3237046d3b4c9b9657a7b2b6dace7b4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 09:56:28 +0200 Subject: [PATCH 15/26] test(nvmf): a reservation held by all registrants outlives its acquirer (0050, tests only) subsystem_ut, red on the series as it stands. A Write Exclusive - All Registrants reservation acquired by A, then A preempted by B: the reservation stays under A's key, which no registrant holds, and the state persisted for it fails to restore (-EINVAL). The same once B unregisters after C registered. The restore test now refuses a missing key for a single holder only, and takes it for a reservation held by all registrants, under the key of its first registrant. --- ...ts-reservation-outlives-its-acquirer.patch | 143 ++++++++++++++++++ patches/README.md | 3 +- 2 files changed, 145 insertions(+), 1 deletion(-) create mode 100644 patches/0050-nvmf-all-registrants-reservation-outlives-its-acquirer.patch diff --git a/patches/0050-nvmf-all-registrants-reservation-outlives-its-acquirer.patch b/patches/0050-nvmf-all-registrants-reservation-outlives-its-acquirer.patch new file mode 100644 index 0000000..53703e3 --- /dev/null +++ b/patches/0050-nvmf-all-registrants-reservation-outlives-its-acquirer.patch @@ -0,0 +1,143 @@ +diff --git a/test/unit/lib/nvmf/subsystem.c/subsystem_ut.c b/test/unit/lib/nvmf/subsystem.c/subsystem_ut.c +index d53d45d..28622b8 100644 +--- a/test/unit/lib/nvmf/subsystem.c/subsystem_ut.c ++++ b/test/unit/lib/nvmf/subsystem.c/subsystem_ut.c +@@ -1779,6 +1779,98 @@ test_reservation_acquire_release_with_ptpl(void) + ut_reservation_deinit(); + } + ++/* Restores the persisted reservation of g_ns onto a fresh namespace of the same bdev, as a ++ * target restart does, and checks it stands under the expected key. */ ++static void ++ut_reservation_check_restores_under(uint64_t crkey) ++{ ++ struct spdk_nvmf_reservation_info info; ++ struct spdk_nvmf_ns ns = {}; ++ int rc; ++ ++ rc = nvmf_ns_update_reservation_info(&g_ns); ++ SPDK_CU_ASSERT_FATAL(rc == 0); ++ memset(&info, 0, sizeof(info)); ++ rc = nvmf_ns_reservation_load(&g_ns, &info); ++ SPDK_CU_ASSERT_FATAL(rc == 0); ++ CU_ASSERT(info.rtype == SPDK_NVME_RESERVE_WRITE_EXCLUSIVE_ALL_REGS); ++ CU_ASSERT(info.crkey == crkey); ++ ++ ns.bdev = g_ns.bdev; ++ TAILQ_INIT(&ns.registrants); ++ rc = nvmf_ns_reservation_restore(&ns, &info); ++ CU_ASSERT(rc == 0); ++ CU_ASSERT(ns.rtype == SPDK_NVME_RESERVE_WRITE_EXCLUSIVE_ALL_REGS); ++ CU_ASSERT(ns.crkey == crkey); ++ nvmf_ns_reservation_clear_all_registrants(&ns); ++} ++ ++/* Evariops 0050: a reservation held by all registrants outlives the registrant that acquired ++ * it. When that registrant leaves, preempted by another or unregistering, the reservation ++ * stands under the key of a registrant that stays, and its persisted state restores. */ ++static void ++test_reservation_all_regs_acquirer_leaves_with_ptpl(void) ++{ ++ struct spdk_nvmf_request *req; ++ struct spdk_nvme_cpl *rsp; ++ struct spdk_nvmf_registrant *reg; ++ bool update_sgroup; ++ ++ ut_reservation_init(); ++ ++ req = ut_reservation_build_req(16); ++ rsp = &req->rsp->nvme_cpl; ++ SPDK_CU_ASSERT_FATAL(req != NULL); ++ ++ g_ns.ptpl_file = "/tmp/Ns1PR.cfg"; ++ ut_reservation_build_register_request(req, SPDK_NVME_RESERVE_REGISTER_KEY, 0, ++ SPDK_NVME_RESERVE_PTPL_PERSIST_POWER_LOSS, 0, 0xa1); ++ update_sgroup = nvmf_ns_reservation_register(&g_ns, &g_ctrlr1_A, req); ++ SPDK_CU_ASSERT_FATAL(update_sgroup == true); ++ SPDK_CU_ASSERT_FATAL(rsp->status.sc == SPDK_NVME_SC_SUCCESS); ++ ut_reservation_build_register_request(req, SPDK_NVME_RESERVE_REGISTER_KEY, 0, 0, 0, 0xb1); ++ nvmf_ns_reservation_register(&g_ns, &g_ctrlr_B, req); ++ SPDK_CU_ASSERT_FATAL(rsp->status.sc == SPDK_NVME_SC_SUCCESS); ++ ++ ut_reservation_build_acquire_request(req, SPDK_NVME_RESERVE_ACQUIRE, 0, ++ SPDK_NVME_RESERVE_WRITE_EXCLUSIVE_ALL_REGS, 0xa1, 0x0); ++ nvmf_ns_reservation_acquire(&g_ns, &g_ctrlr1_A, req); ++ SPDK_CU_ASSERT_FATAL(rsp->status.sc == SPDK_NVME_SC_SUCCESS); ++ SPDK_CU_ASSERT_FATAL(g_ns.crkey == 0xa1); ++ ++ /* TEST CASE: B preempts A, the registrant that acquired the reservation */ ++ ut_reservation_build_acquire_request(req, SPDK_NVME_RESERVE_PREEMPT, 0, ++ SPDK_NVME_RESERVE_WRITE_EXCLUSIVE_ALL_REGS, 0xb1, 0xa1); ++ nvmf_ns_reservation_acquire(&g_ns, &g_ctrlr_B, req); ++ SPDK_CU_ASSERT_FATAL(rsp->status.sc == SPDK_NVME_SC_SUCCESS); ++ SPDK_CU_ASSERT_FATAL(nvmf_ns_reservation_get_registrant(&g_ns, &g_ctrlr1_A.hostid) == NULL); ++ reg = nvmf_ns_reservation_get_registrant(&g_ns, &g_ctrlr_B.hostid); ++ SPDK_CU_ASSERT_FATAL(reg != NULL); ++ CU_ASSERT(g_ns.rtype == SPDK_NVME_RESERVE_WRITE_EXCLUSIVE_ALL_REGS); ++ CU_ASSERT(g_ns.holder == reg); ++ CU_ASSERT(g_ns.crkey == 0xb1); ++ ut_reservation_check_restores_under(0xb1); ++ ++ /* TEST CASE: C registers, then B, whose key the reservation now stands under, unregisters */ ++ ut_reservation_build_register_request(req, SPDK_NVME_RESERVE_REGISTER_KEY, 0, 0, 0, 0xc1); ++ nvmf_ns_reservation_register(&g_ns, &g_ctrlr_C, req); ++ SPDK_CU_ASSERT_FATAL(rsp->status.sc == SPDK_NVME_SC_SUCCESS); ++ ut_reservation_build_register_request(req, SPDK_NVME_RESERVE_UNREGISTER_KEY, 0, 0, 0xb1, 0); ++ nvmf_ns_reservation_register(&g_ns, &g_ctrlr_B, req); ++ SPDK_CU_ASSERT_FATAL(rsp->status.sc == SPDK_NVME_SC_SUCCESS); ++ SPDK_CU_ASSERT_FATAL(nvmf_ns_reservation_get_registrant(&g_ns, &g_ctrlr_B.hostid) == NULL); ++ reg = nvmf_ns_reservation_get_registrant(&g_ns, &g_ctrlr_C.hostid); ++ SPDK_CU_ASSERT_FATAL(reg != NULL); ++ CU_ASSERT(g_ns.rtype == SPDK_NVME_RESERVE_WRITE_EXCLUSIVE_ALL_REGS); ++ CU_ASSERT(g_ns.holder == reg); ++ CU_ASSERT(g_ns.crkey == 0xc1); ++ ut_reservation_check_restores_under(0xc1); ++ ++ unlink(g_ns.ptpl_file); ++ ut_reservation_free_req(req); ++ ut_reservation_deinit(); ++} ++ + static void + test_reservation_release(void) + { +@@ -3346,12 +3438,30 @@ test_nvmf_ns_reservation_restore(void) + spdk_uuid_fmt_lower(uuid, sizeof(uuid), &s_uuid); + snprintf(info.registrants[1].host_uuid, SPDK_UUID_STRING_LEN, "%s", uuid); + +- /* info->rkey not exist in registrants */ ++ /* info->rkey not exist in registrants: the key of a single holder that no registrant holds */ ++ info.rtype = SPDK_NVME_RESERVE_WRITE_EXCLUSIVE; + info.crkey = 0xa; + + rc = nvmf_ns_reservation_restore(&ns, &info); + CU_ASSERT(rc == -EINVAL); + ++ /* Evariops 0050: info->rkey not exist in registrants of a reservation held by all registrants: ++ * the registrant that acquired it left, which a state persisted before 0050 still records. It ++ * restores, under the key of the registrant that holds it now. */ ++ info.rtype = SPDK_NVME_RESERVE_WRITE_EXCLUSIVE_ALL_REGS; ++ ++ rc = nvmf_ns_reservation_restore(&ns, &info); ++ CU_ASSERT(rc == 0); ++ CU_ASSERT(ns.rtype == SPDK_NVME_RESERVE_WRITE_EXCLUSIVE_ALL_REGS); ++ reg0 = TAILQ_FIRST(&ns.registrants); ++ SPDK_CU_ASSERT_FATAL(reg0 != NULL); ++ CU_ASSERT(ns.holder == reg0); ++ CU_ASSERT(ns.crkey == reg0->rkey); ++ ++ rc = nvmf_ns_reservation_clear_all_registrants(&ns); ++ CU_ASSERT(rc == 2); ++ CU_ASSERT(TAILQ_EMPTY(&ns.registrants)); ++ + /* info->rkey exists in registrants */ + info.crkey = 0xb; + +@@ -3630,6 +3740,7 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_reservation_register_with_ptpl); + CU_ADD_TEST(suite, test_reservation_acquire_preempt); + CU_ADD_TEST(suite, test_reservation_acquire_release_with_ptpl); ++ CU_ADD_TEST(suite, test_reservation_all_regs_acquirer_leaves_with_ptpl); + CU_ADD_TEST(suite, test_reservation_release); + CU_ADD_TEST(suite, test_reservation_unregister_notification); + CU_ADD_TEST(suite, test_reservation_release_notification); diff --git a/patches/README.md b/patches/README.md index 619c3cb..750e515 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0049`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0050`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -55,6 +55,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0049`) | 0047 | a raid1 rebuild copies its window in parts — `raid_bdev_process_request_max_part` offers a window to the process's requests in `RAID_BDEV_PROCESS_MAX_QD` parts (whole write units), and raid1 copies one part per request, so a window's parts are in flight together: each part's write starts as its own read completes, and the reads spread over the members in sync. Copied as one request, a window was read whole from one member, then written whole: read and write never overlapped, and the first member served every read — a 5 GiB rebuild ran at 36 MiB/s, all of it read across the link from a remote member while the nexus's local member idled. raid5f, bound to its stripe, ignores the part. `raid1_ut` pins the parts, their spread and their writes; `bdev_raid_ut` the part's bounds | bdev_raid, raid1, test/unit/lib/bdev/raid | 0013 | | 0048 | a raid1 rebuild writes a zero part as write-zeroes — a part the source read back as nothing but zeroes goes to the target as `write_zeroes` (`raid_bdev_write_zeroes_blocks`): the same blocks, read back the same, and no payload on the wire, so a full rebuild of a sparse volume no longer ships its never-written clusters across the link (about 3 of the 5 GiB of the rebuild measured for 0047). Not when the target cannot take write-zeroes (the bdev layer would send a zero buffer anyway), nor when the blocks carry metadata. `raid1_ut` pins both kinds of write, and a target without write-zeroes getting data | raid1, bdev_raid (header), test/unit/lib/bdev/raid | 0047 | | 0049 | write-zeroes leaves an unallocated thin cluster unallocated — on a blob backed by zeroes, an unallocated cluster already reads zeroes, so write-zeroes over it completes with no cluster allocated and nothing written. A full rebuild sends every never-written part of its source as write-zeroes (0048), and allocating each made the rebuilt replica fully allocated — a thin volume rebuilt in full took its whole size on the target — and wrote zeroes across the target's disk. An allocated cluster is still zeroed; a clone, which reads its unallocated clusters from its parent, still allocates and zeroes. `blob_ut` pins the three | lib/blob (blobstore.c), test/unit/lib/blob | 0044 | +| 0050 | a reservation held by all registrants outlives the registrant that acquired it — when that registrant leaves (preempted by another, or unregistering), the reservation, which every registrant holds, now stands under the key of the registrant that holds it next. It kept the key of the one that left, and the persisted state (`ptpl_file`) recorded a reservation key no registrant held: the restore refused it, so after a writer handover, the next restart of a target failed to add the namespace back, and the target stayed unexported. The restore also takes a state persisted before this patch, under the key of its first registrant; a single holder's key that no registrant holds is still refused. `subsystem_ut` pins the preempt, the unregister and the restore | lib/nvmf (subsystem.c), test/unit/lib/nvmf/subsystem.c | — | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From d7b99d82c0530fcad7204a8d8a1d01508e268245 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 09:57:32 +0200 Subject: [PATCH 16/26] fix(nvmf): a reservation held by all registrants outlives its acquirer (0050) Every registrant holds a Write Exclusive or Exclusive Access - All Registrants reservation, so when the registrant that acquired it leaves, preempted by another or unregistering, the reservation stays. It stayed under the key of the registrant that left: the state persisted for it named a reservation key no registrant held, and the restore refused it (-EINVAL), failing the add of the namespace. After a writer handover, the next restart of a target left its namespace out, and the target unexported. The reservation now stands under the key of the registrant that holds it next. The restore takes a state persisted before this, under the key of its first registrant; a single holder's key that no registrant holds is still refused. --- ...ts-reservation-outlives-its-acquirer.patch | 47 +++++++++++++++++++ 1 file changed, 47 insertions(+) diff --git a/patches/0050-nvmf-all-registrants-reservation-outlives-its-acquirer.patch b/patches/0050-nvmf-all-registrants-reservation-outlives-its-acquirer.patch index 53703e3..2340a37 100644 --- a/patches/0050-nvmf-all-registrants-reservation-outlives-its-acquirer.patch +++ b/patches/0050-nvmf-all-registrants-reservation-outlives-its-acquirer.patch @@ -1,3 +1,50 @@ +diff --git a/lib/nvmf/subsystem.c b/lib/nvmf/subsystem.c +index 356361f..3d53309 100644 +--- a/lib/nvmf/subsystem.c ++++ b/lib/nvmf/subsystem.c +@@ -2919,6 +2919,7 @@ nvmf_ns_reservation_restore(struct spdk_nvmf_ns *ns, struct spdk_nvmf_reservatio + struct spdk_nvmf_registrant *reg, *holder = NULL; + struct spdk_uuid bdev_uuid, holder_uuid; + bool rkey_flag = false; ++ bool all_regs; + + SPDK_DEBUGLOG(nvmf, "NSID %u, PTPL %u, Number of registrants %u\n", + ns->nsid, info->ptpl_activated, info->num_regs); +@@ -2934,7 +2935,12 @@ nvmf_ns_reservation_restore(struct spdk_nvmf_ns *ns, struct spdk_nvmf_reservatio + rkey_flag = true; + } + } +- if (!rkey_flag && info->crkey != 0) { ++ /* Evariops 0050: a reservation held by all registrants outlives the registrant that ++ * acquired it, and a state persisted before 0050 still records the key of that registrant ++ * after it left: it restores, under the key of the registrant that holds it (below). */ ++ all_regs = (info->rtype == SPDK_NVME_RESERVE_WRITE_EXCLUSIVE_ALL_REGS || ++ info->rtype == SPDK_NVME_RESERVE_EXCLUSIVE_ACCESS_ALL_REGS); ++ if (!rkey_flag && info->crkey != 0 && !all_regs) { + return -EINVAL; + } + +@@ -2972,6 +2978,9 @@ nvmf_ns_reservation_restore(struct spdk_nvmf_ns *ns, struct spdk_nvmf_reservatio + + if (nvmf_ns_reservation_all_registrants_type(ns)) { + ns->holder = TAILQ_FIRST(&ns->registrants); ++ if (!rkey_flag) { ++ ns->crkey = ns->holder->rkey; ++ } + } else { + ns->holder = holder; + } +@@ -3267,6 +3276,10 @@ nvmf_ns_reservation_check_release_on_remove_registrant(struct spdk_nvmf_ns *ns, + if (next_reg && nvmf_ns_reservation_all_registrants_type(ns)) { + /* the next valid registrant is the new holder now */ + ns->holder = next_reg; ++ /* Evariops 0050: every registrant holds the reservation, so it outlives the one that ++ * acquired it, and now stands under the key of the registrant that holds it. Left under ++ * the key of a registrant that is gone, its persisted state fails to restore. */ ++ ns->crkey = next_reg->rkey; + } else if (nvmf_ns_reservation_registrant_is_holder(ns, reg)) { + /* release the reservation */ + nvmf_ns_reservation_release_reservation(ns); diff --git a/test/unit/lib/nvmf/subsystem.c/subsystem_ut.c b/test/unit/lib/nvmf/subsystem.c/subsystem_ut.c index d53d45d..28622b8 100644 --- a/test/unit/lib/nvmf/subsystem.c/subsystem_ut.c From f74e7564ef1a36f0fd67d97bfbee7e9f79df4b38 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 22:11:53 +0200 Subject: [PATCH 17/26] test(bdev-nvme): an I/O passthrough needs a connected qpair (0051, tests only) bdev_nvme_ut, red on the series as it stands. The suite now compiles nvme_rpc.c in, so the I/O passthrough can be driven against a real controller channel. With the channel's I/O qpair cleared, as bdev_nvme clears it from a disconnect until the reconnect, the command is still handed to the NVMe library (rc 0, the channel kept): in a real process that NULL qpair is dereferenced. The test expects -ENXIO with the channel given back, and the command to go through once connected again. --- ...-passthrough-needs-a-connected-qpair.patch | 101 ++++++++++++++++++ patches/README.md | 3 +- 2 files changed, 103 insertions(+), 1 deletion(-) create mode 100644 patches/0051-bdev-nvme-io-passthrough-needs-a-connected-qpair.patch diff --git a/patches/0051-bdev-nvme-io-passthrough-needs-a-connected-qpair.patch b/patches/0051-bdev-nvme-io-passthrough-needs-a-connected-qpair.patch new file mode 100644 index 0000000..7365c45 --- /dev/null +++ b/patches/0051-bdev-nvme-io-passthrough-needs-a-connected-qpair.patch @@ -0,0 +1,101 @@ +diff --git a/test/unit/lib/bdev/nvme/bdev_nvme.c/bdev_nvme_ut.c b/test/unit/lib/bdev/nvme/bdev_nvme.c/bdev_nvme_ut.c +index c64dc7c..bd022fe 100644 +--- a/test/unit/lib/bdev/nvme/bdev_nvme.c/bdev_nvme_ut.c ++++ b/test/unit/lib/bdev/nvme/bdev_nvme.c/bdev_nvme_ut.c +@@ -15,6 +15,19 @@ + + #include "unit/lib/json_mock.c" + ++/* nvme_rpc.c is compiled in so its I/O passthrough can be driven against a real controller ++ * channel (Evariops 0051); its JSON decoding and RPC registration are not exercised here. */ ++#include "spdk/rpc.h" ++DEFINE_STUB(spdk_json_decode_object, int, (const struct spdk_json_val *values, ++ const struct spdk_json_object_decoder *decoders, size_t num_decoders, void *out), 0); ++DEFINE_STUB(spdk_json_decode_string, int, (const struct spdk_json_val *val, void *out), 0); ++DEFINE_STUB(spdk_json_decode_uint32, int, (const struct spdk_json_val *val, void *out), 0); ++DEFINE_STUB(spdk_json_strequal, bool, (const struct spdk_json_val *val, const char *str), false); ++DEFINE_STUB_V(spdk_rpc_register_method, (const char *method, spdk_rpc_method_handler func, ++ uint32_t state_mask)); ++ ++#include "bdev/nvme/nvme_rpc.c" ++ + #include "bdev/nvme/bdev_mdns_client.c" + + static void *g_accel_p = (void *)0xdeadbeaf; +@@ -8490,6 +8503,68 @@ test_race_between_ctrlr_loss_timeout_and_pending_failover(void) + CU_ASSERT(nvme_ctrlr_get_by_name("nvme0") == NULL); + } + ++/* bdev_nvme frees a controller channel's I/O qpair from a disconnect until the reconnect ++ * (reset, failed controller, delayed reconnect). bdev_nvme_send_cmd handed that NULL qpair to ++ * the NVMe library and the process died on it (SIGSEGV), taking every raid of the process down: ++ * a reservation command is an I/O command. Refused instead, with the channel given back. */ ++static void ++test_nvme_rpc_io_cmd_needs_a_connected_qpair(void) ++{ ++ struct spdk_nvme_transport_id trid = {}; ++ struct spdk_nvme_ctrlr ctrlr = {}; ++ struct nvme_ctrlr *nvme_ctrlr; ++ struct rpc_bdev_nvme_send_cmd_ctx ctx = {}; ++ struct spdk_nvme_cmd cmd = {}; ++ struct spdk_io_channel *ch; ++ struct nvme_ctrlr_channel *ctrlr_ch; ++ struct spdk_nvme_qpair *qpair; ++ int rc; ++ ++ ut_init_trid(&trid); ++ TAILQ_INIT(&ctrlr.active_io_qpairs); ++ ++ set_thread(0); ++ ++ rc = nvme_ctrlr_create(&ctrlr, "nvme0", &trid, NULL); ++ CU_ASSERT(rc == 0); ++ ++ nvme_ctrlr = nvme_ctrlr_get_by_name("nvme0"); ++ SPDK_CU_ASSERT_FATAL(nvme_ctrlr != NULL); ++ ++ ch = spdk_get_io_channel(nvme_ctrlr); ++ SPDK_CU_ASSERT_FATAL(ch != NULL); ++ ctrlr_ch = spdk_io_channel_get_ctx(ch); ++ qpair = ctrlr_ch->qpair->qpair; ++ SPDK_CU_ASSERT_FATAL(qpair != NULL); ++ ++ /* Disconnected: no qpair to carry the command. */ ++ ctrlr_ch->qpair->qpair = NULL; ++ ctx.nvme_ctrlr = nvme_ctrlr; ++ rc = nvme_rpc_io_cmd_bdev_nvme(&ctx, &cmd, NULL, 0, NULL, 0, 0); ++ CU_ASSERT(rc == -ENXIO); ++ CU_ASSERT(ctx.ctrlr_io_ch == NULL); ++ ++ /* Connected again: the command goes through, holding the channel until it completes. */ ++ ctrlr_ch->qpair->qpair = qpair; ++ rc = nvme_rpc_io_cmd_bdev_nvme(&ctx, &cmd, NULL, 0, NULL, 0, 0); ++ CU_ASSERT(rc == 0); ++ CU_ASSERT(ctx.ctrlr_io_ch == ch); ++ spdk_put_io_channel(ctx.ctrlr_io_ch); ++ ++ spdk_put_io_channel(ch); ++ ++ poll_threads(); ++ ++ rc = spdk_bdev_nvme_delete("nvme0", &g_any_path, NULL, NULL); ++ CU_ASSERT(rc == 0); ++ ++ poll_threads(); ++ spdk_delay_us(1000); ++ poll_threads(); ++ ++ CU_ASSERT(nvme_ctrlr_get_by_name("nvme0") == NULL); ++} ++ + int + main(int argc, char **argv) + { +@@ -8554,6 +8629,7 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_bdev_reset_abort_io); + CU_ADD_TEST(suite, test_race_between_clear_pending_resets_and_reset_ctrlr_complete); + CU_ADD_TEST(suite, test_race_between_ctrlr_loss_timeout_and_pending_failover); ++ CU_ADD_TEST(suite, test_nvme_rpc_io_cmd_needs_a_connected_qpair); + + allocate_threads(3); + set_thread(0); diff --git a/patches/README.md b/patches/README.md index 750e515..f4aef02 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0050`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0051`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -56,6 +56,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0050`) | 0048 | a raid1 rebuild writes a zero part as write-zeroes — a part the source read back as nothing but zeroes goes to the target as `write_zeroes` (`raid_bdev_write_zeroes_blocks`): the same blocks, read back the same, and no payload on the wire, so a full rebuild of a sparse volume no longer ships its never-written clusters across the link (about 3 of the 5 GiB of the rebuild measured for 0047). Not when the target cannot take write-zeroes (the bdev layer would send a zero buffer anyway), nor when the blocks carry metadata. `raid1_ut` pins both kinds of write, and a target without write-zeroes getting data | raid1, bdev_raid (header), test/unit/lib/bdev/raid | 0047 | | 0049 | write-zeroes leaves an unallocated thin cluster unallocated — on a blob backed by zeroes, an unallocated cluster already reads zeroes, so write-zeroes over it completes with no cluster allocated and nothing written. A full rebuild sends every never-written part of its source as write-zeroes (0048), and allocating each made the rebuilt replica fully allocated — a thin volume rebuilt in full took its whole size on the target — and wrote zeroes across the target's disk. An allocated cluster is still zeroed; a clone, which reads its unallocated clusters from its parent, still allocates and zeroes. `blob_ut` pins the three | lib/blob (blobstore.c), test/unit/lib/blob | 0044 | | 0050 | a reservation held by all registrants outlives the registrant that acquired it — when that registrant leaves (preempted by another, or unregistering), the reservation, which every registrant holds, now stands under the key of the registrant that holds it next. It kept the key of the one that left, and the persisted state (`ptpl_file`) recorded a reservation key no registrant held: the restore refused it, so after a writer handover, the next restart of a target failed to add the namespace back, and the target stayed unexported. The restore also takes a state persisted before this patch, under the key of its first registrant; a single holder's key that no registrant holds is still refused. `subsystem_ut` pins the preempt, the unregister and the restore | lib/nvmf (subsystem.c), test/unit/lib/nvmf/subsystem.c | — | +| 0051 | an I/O passthrough needs a connected qpair — `bdev_nvme_send_cmd` with an I/O command (NVMe reservations are I/O commands) took the controller channel's qpair without a check and handed it to the NVMe library. bdev_nvme frees that qpair from a disconnect until the reconnect (reset, failed controller, delayed reconnect), so a reservation sent to a member whose path had just gone down dereferenced NULL and killed the process, and every raid it carried with it. The command is now refused (`-ENXIO`, the channel given back), as is one for which no channel can be had. `bdev_nvme_ut` pins the disconnected and the connected case | module/bdev/nvme (nvme_rpc.c), test/unit/lib/bdev/nvme | — | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From 8a3f3f1bf2f715f22545f00139bed20687cd7db6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Thu, 1 Oct 2026 22:12:06 +0200 Subject: [PATCH 18/26] fix(bdev-nvme): an I/O passthrough needs a connected qpair (0051) bdev_nvme_send_cmd with an I/O command took the controller channel's qpair (bdev_nvme_get_io_qpair) without a check and handed it to spdk_nvme_ctrlr_cmd_io_raw_with_md. bdev_nvme frees that qpair from a disconnect until the reconnect (bdev_nvme_disconnected_qpair_cb), so a reservation (reservations are I/O commands) sent to a member whose path had just gone down dereferenced NULL and killed the process, every raid it carried with it. The command is now refused with -ENXIO and the channel given back; a command for which no channel can be had is refused with -ENOMEM instead of reaching the assert-only accessor. --- ...-passthrough-needs-a-connected-qpair.patch | 24 +++++++++++++++++++ 1 file changed, 24 insertions(+) diff --git a/patches/0051-bdev-nvme-io-passthrough-needs-a-connected-qpair.patch b/patches/0051-bdev-nvme-io-passthrough-needs-a-connected-qpair.patch index 7365c45..3f0026d 100644 --- a/patches/0051-bdev-nvme-io-passthrough-needs-a-connected-qpair.patch +++ b/patches/0051-bdev-nvme-io-passthrough-needs-a-connected-qpair.patch @@ -1,3 +1,27 @@ +diff --git a/module/bdev/nvme/nvme_rpc.c b/module/bdev/nvme/nvme_rpc.c +index bbcb903..d43a9ec 100644 +--- a/module/bdev/nvme/nvme_rpc.c ++++ b/module/bdev/nvme/nvme_rpc.c +@@ -162,7 +162,19 @@ nvme_rpc_io_cmd_bdev_nvme(struct rpc_bdev_nvme_send_cmd_ctx *ctx, struct spdk_nv + int ret; + + ctx->ctrlr_io_ch = spdk_get_io_channel(_nvme_ctrlr); ++ if (ctx->ctrlr_io_ch == NULL) { ++ return -ENOMEM; ++ } ++ ++ /* bdev_nvme frees a channel's I/O qpair from a disconnect until the reconnect (reset, ++ * failed controller, delayed reconnect). Submitting on it dereferenced NULL and took the ++ * whole process down, every raid with it; refuse the command instead. */ + io_qpair = bdev_nvme_get_io_qpair(ctx->ctrlr_io_ch); ++ if (io_qpair == NULL) { ++ spdk_put_io_channel(ctx->ctrlr_io_ch); ++ ctx->ctrlr_io_ch = NULL; ++ return -ENXIO; ++ } + + ret = spdk_nvme_ctrlr_cmd_io_raw_with_md(_nvme_ctrlr->ctrlr, io_qpair, + cmd, buf, nbytes, md_buf, nvme_rpc_bdev_nvme_cb, ctx); diff --git a/test/unit/lib/bdev/nvme/bdev_nvme.c/bdev_nvme_ut.c b/test/unit/lib/bdev/nvme/bdev_nvme.c/bdev_nvme_ut.c index c64dc7c..bd022fe 100644 --- a/test/unit/lib/bdev/nvme/bdev_nvme.c/bdev_nvme_ut.c From 350d9cf1f208b204bf658e49df4d71989c07cfda Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Fri, 2 Oct 2026 09:12:06 +0200 Subject: [PATCH 19/26] test(raid): a raid1 ejection opens its epoch on the raid's cbt (0052, tests only) bdev_raid_ut, red on the series as it stands. A current member that leaves a raid1, removed or failed, should open one epoch, on the cbt stacked on the raid and named after the member's UUID. 0019 asks each survivor's own name instead (31 calls here, none to the raid), a bdev no cbt sits on, so no epoch ever opens and a member lost without notice is always rebuilt in full. The suite also pins five ejections that must open none: a member being rebuilt in full (its superblock slot not configured), one behind the generation its peers carry, a write-only joiner, a raid without a superblock, and a failed member an earlier removal failed to take. --- ...tion-opens-an-epoch-on-the-raids-cbt.patch | 220 ++++++++++++++++++ patches/README.md | 3 +- 2 files changed, 222 insertions(+), 1 deletion(-) create mode 100644 patches/0052-raid1-ejection-opens-an-epoch-on-the-raids-cbt.patch diff --git a/patches/0052-raid1-ejection-opens-an-epoch-on-the-raids-cbt.patch b/patches/0052-raid1-ejection-opens-an-epoch-on-the-raids-cbt.patch new file mode 100644 index 0000000..394ae22 --- /dev/null +++ b/patches/0052-raid1-ejection-opens-an-epoch-on-the-raids-cbt.patch @@ -0,0 +1,220 @@ +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index a488d6d..c359160 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -242,10 +242,22 @@ DEFINE_STUB(vbdev_cbt_query_latest_epoch, int, (const char *bdev_name, + struct vbdev_cbt_epoch_facts *out), -ENODEV); + /* Evariops 0018: a rebuild's outcome turns to verifying before its verify pass. */ + DEFINE_STUB_V(raid_rebuild_outcome_set_verifying, (struct raid_rebuild_outcome *outcome)); +-/* Evariops 0019: a raid1 ejection opens a CBT epoch on each survivor. No survivor +- * has CBT here (-ENODEV), which the raid passes over. */ +-DEFINE_STUB(vbdev_cbt_auto_epoch_open, int, (const char *bdev_name, const char *stale_backend_id), +- -ENODEV); ++/* Evariops 0019, 0052: a raid1 ejection opens a CBT epoch for the member that ++ * leaves. The CBT module is not part of this suite: this fake records what the ++ * raid asks it for. */ ++static uint32_t g_ejection_epochs; ++static char g_ejection_epoch_over[64]; ++static char g_ejection_epoch_member[64]; ++ ++int ++vbdev_cbt_auto_epoch_open(const char *bdev_name, const char *member_id) ++{ ++ g_ejection_epochs++; ++ snprintf(g_ejection_epoch_over, sizeof(g_ejection_epoch_over), "%s", bdev_name); ++ snprintf(g_ejection_epoch_member, sizeof(g_ejection_epoch_member), "%s", member_id); ++ ++ return 0; ++} + + /* Evariops 0046: a seeded rebuild reads its ranges from a CBT epoch. The CBT + * module is not part of this suite: the epoch's delta is this fake's. */ +@@ -679,12 +691,15 @@ spdk_bdev_get_by_name(const char *bdev_name) + return NULL; + } + ++/* Evariops 0052: the status a quiesce of a whole bdev completes with. */ ++static int g_quiesce_status; ++ + int + spdk_bdev_quiesce(struct spdk_bdev *bdev, struct spdk_bdev_module *module, + spdk_bdev_quiesce_cb cb_fn, void *cb_arg) + { + if (cb_fn) { +- cb_fn(cb_arg, 0); ++ cb_fn(cb_arg, g_quiesce_status); + } + + return 0; +@@ -2156,6 +2171,162 @@ test_raid_io_split(void) + reset_globals(); + } + ++/* Evariops 0052: how the member in slot 0 leaves the raid. */ ++enum ut_ejection { ++ /* removed by the control plane, or hot-removed with its bdev */ ++ UT_EJECTION_REMOVED, ++ /* failed by an I/O error */ ++ UT_EJECTION_FAILED, ++ /* failed again after its first removal failed */ ++ UT_EJECTION_FAILED_AFTER_A_FAILED_REMOVAL, ++}; ++ ++/* Evariops 0052: creates a superblock raid1 whose members all sit CONFIGURED at ++ * one generation, lets setup change that, ejects the member in slot 0 and returns ++ * how many epochs the ejection opened; member receives that member's UUID. The ++ * suite stubs the superblock's allocation, so the raid gets one here, and the ++ * suite's own module stands in for raid1. */ ++static uint32_t ++ejection_epochs_run(enum ut_ejection ejection, ++ void (*setup)(struct raid_bdev *raid_bdev, struct raid_bdev_superblock *sb), ++ char *member, size_t member_len) ++{ ++ struct rpc_bdev_raid_create req; ++ struct rpc_bdev_raid_delete destroy_req; ++ struct raid_bdev_superblock *sb; ++ struct raid_base_bdev_info *base_info; ++ struct raid_bdev *pbdev; ++ uint8_t i; ++ ++ set_globals(); ++ CU_ASSERT(raid_bdev_init() == 0); ++ create_raid_bdev_create_req(&req, "raid1", 0, true, 0, true); ++ rpc_bdev_raid_create(NULL, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ free_test_req(&req); ++ ++ pbdev = raid_bdev_find_by_name("raid1"); ++ SPDK_CU_ASSERT_FATAL(pbdev != NULL); ++ pbdev->level = RAID1; ++ pbdev->min_base_bdevs_operational = 1; ++ sb = calloc(1, sizeof(*sb) + pbdev->num_base_bdevs * sizeof(sb->base_bdevs[0])); ++ SPDK_CU_ASSERT_FATAL(sb != NULL); ++ sb->base_bdevs_size = pbdev->num_base_bdevs; ++ for (i = 0; i < pbdev->num_base_bdevs; i++) { ++ sb->base_bdevs[i].slot = i; ++ sb->base_bdevs[i].state = RAID_SB_BASE_BDEV_CONFIGURED; ++ sb->base_bdevs[i].content_generation = 7; ++ pbdev->base_bdev_info[i].content_generation = 7; ++ } ++ pbdev->sb = sb; ++ if (setup != NULL) { ++ setup(pbdev, sb); ++ } ++ ++ base_info = &pbdev->base_bdev_info[0]; ++ spdk_uuid_fmt_lower(member, member_len, &base_info->uuid); ++ g_ejection_epochs = 0; ++ g_ejection_epoch_over[0] = '\0'; ++ g_ejection_epoch_member[0] = '\0'; ++ ++ switch (ejection) { ++ case UT_EJECTION_REMOVED: ++ CU_ASSERT(raid_bdev_remove_base_bdev(spdk_bdev_desc_get_bdev(base_info->desc), NULL, ++ NULL, NULL) == 0); ++ break; ++ case UT_EJECTION_FAILED_AFTER_A_FAILED_REMOVAL: ++ g_quiesce_status = -EIO; ++ raid_bdev_fail_base_bdev(base_info); ++ poll_app_thread(); ++ g_quiesce_status = 0; ++ CU_ASSERT(base_info->is_configured); ++ CU_ASSERT(base_info->remove_error == -EIO); ++ /* fallthrough */ ++ case UT_EJECTION_FAILED: ++ raid_bdev_fail_base_bdev(base_info); ++ break; ++ } ++ poll_app_thread(); ++ CU_ASSERT(!base_info->is_configured); ++ ++ pbdev->sb = NULL; ++ free(sb); ++ create_raid_bdev_delete_req(&destroy_req, "raid1", 0); ++ rpc_bdev_raid_delete(NULL, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev_exit(); ++ base_bdevs_cleanup(); ++ reset_globals(); ++ ++ return g_ejection_epochs; ++} ++ ++/* Evariops 0052: a raid1 member that leaves current, removed or failed, opens one ++ * epoch, on the cbt bdev stacked on the raid and named after the member's UUID: ++ * the identity the control plane derives from the replica it placed there. 0019 ++ * opened one under each survivor's own name, a bdev no cbt sits on (-ENODEV), so ++ * a member lost without notice was always rebuilt in full. */ ++static void ++test_raid1_ejection_opens_an_epoch_on_the_raids_cbt(void) ++{ ++ const enum ut_ejection ejections[] = { UT_EJECTION_REMOVED, UT_EJECTION_FAILED }; ++ char member[SPDK_UUID_STRING_LEN]; ++ size_t i; ++ ++ for (i = 0; i < SPDK_COUNTOF(ejections); i++) { ++ CU_ASSERT(ejection_epochs_run(ejections[i], NULL, member, sizeof(member)) == 1); ++ CU_ASSERT(strcmp(g_ejection_epoch_over, "raid1") == 0); ++ CU_ASSERT(strcmp(g_ejection_epoch_member, member) == 0); ++ } ++} ++ ++static void ++ut_member_rebuilt_in_full(struct raid_bdev *raid_bdev, struct raid_bdev_superblock *sb) ++{ ++ sb->base_bdevs[0].state = RAID_SB_BASE_BDEV_MISSING; ++} ++ ++static void ++ut_member_behind(struct raid_bdev *raid_bdev, struct raid_bdev_superblock *sb) ++{ ++ sb->base_bdevs[0].content_generation = 6; ++ raid_bdev->base_bdev_info[0].content_generation = 6; ++} ++ ++static void ++ut_member_joining_write_only(struct raid_bdev *raid_bdev, struct raid_bdev_superblock *sb) ++{ ++ raid_bdev->base_bdev_info[0].write_only_pending = true; ++ raid_bdev->base_bdev_info[0].read_excluded = true; ++} ++ ++static void ++ut_no_superblock(struct raid_bdev *raid_bdev, struct raid_bdev_superblock *sb) ++{ ++ raid_bdev->sb = NULL; ++} ++ ++/* Evariops 0052: an epoch holds the writes a member misses from its departure on. ++ * A member that already lacked some (being rebuilt in full, behind the generation ++ * its peers carry, or joining write-only) would be seeded with less than it owes, ++ * and so would one a removal already failed to take, which may have missed writes ++ * since that attempt. None of their ejections opens an epoch, nor one with no ++ * superblock to tell: each of them is rebuilt in full. */ ++static void ++test_raid1_ejection_of_a_member_behind_opens_no_epoch(void) ++{ ++ char member[SPDK_UUID_STRING_LEN]; ++ ++ CU_ASSERT(ejection_epochs_run(UT_EJECTION_REMOVED, ut_member_rebuilt_in_full, ++ member, sizeof(member)) == 0); ++ CU_ASSERT(ejection_epochs_run(UT_EJECTION_REMOVED, ut_member_behind, member, sizeof(member)) == 0); ++ CU_ASSERT(ejection_epochs_run(UT_EJECTION_REMOVED, ut_member_joining_write_only, ++ member, sizeof(member)) == 0); ++ CU_ASSERT(ejection_epochs_run(UT_EJECTION_REMOVED, ut_no_superblock, member, sizeof(member)) == 0); ++ CU_ASSERT(ejection_epochs_run(UT_EJECTION_FAILED_AFTER_A_FAILED_REMOVAL, NULL, ++ member, sizeof(member)) == 0); ++} ++ + static int + test_new_thread_fn(struct spdk_thread *thread) + { +@@ -2207,6 +2378,8 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_raid_seeded_rebuild_from_epoch_refuses_a_delta_it_cannot_take_whole); + CU_ADD_TEST(suite, test_raid_process_request_max_part); + CU_ADD_TEST(suite, test_raid_process_with_qos); ++ CU_ADD_TEST(suite, test_raid1_ejection_opens_an_epoch_on_the_raids_cbt); ++ CU_ADD_TEST(suite, test_raid1_ejection_of_a_member_behind_opens_no_epoch); + + spdk_thread_lib_init(test_new_thread_fn, 0); + g_app_thread = spdk_thread_create("app_thread", NULL); diff --git a/patches/README.md b/patches/README.md index f4aef02..8035ceb 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0051`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0052`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -57,6 +57,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0051`) | 0049 | write-zeroes leaves an unallocated thin cluster unallocated — on a blob backed by zeroes, an unallocated cluster already reads zeroes, so write-zeroes over it completes with no cluster allocated and nothing written. A full rebuild sends every never-written part of its source as write-zeroes (0048), and allocating each made the rebuilt replica fully allocated — a thin volume rebuilt in full took its whole size on the target — and wrote zeroes across the target's disk. An allocated cluster is still zeroed; a clone, which reads its unallocated clusters from its parent, still allocates and zeroes. `blob_ut` pins the three | lib/blob (blobstore.c), test/unit/lib/blob | 0044 | | 0050 | a reservation held by all registrants outlives the registrant that acquired it — when that registrant leaves (preempted by another, or unregistering), the reservation, which every registrant holds, now stands under the key of the registrant that holds it next. It kept the key of the one that left, and the persisted state (`ptpl_file`) recorded a reservation key no registrant held: the restore refused it, so after a writer handover, the next restart of a target failed to add the namespace back, and the target stayed unexported. The restore also takes a state persisted before this patch, under the key of its first registrant; a single holder's key that no registrant holds is still refused. `subsystem_ut` pins the preempt, the unregister and the restore | lib/nvmf (subsystem.c), test/unit/lib/nvmf/subsystem.c | — | | 0051 | an I/O passthrough needs a connected qpair — `bdev_nvme_send_cmd` with an I/O command (NVMe reservations are I/O commands) took the controller channel's qpair without a check and handed it to the NVMe library. bdev_nvme frees that qpair from a disconnect until the reconnect (reset, failed controller, delayed reconnect), so a reservation sent to a member whose path had just gone down dereferenced NULL and killed the process, and every raid it carried with it. The command is now refused (`-ENXIO`, the channel given back), as is one for which no channel can be had. `bdev_nvme_ut` pins the disconnected and the connected case | module/bdev/nvme (nvme_rpc.c), test/unit/lib/bdev/nvme | — | +| 0052 | a raid1 ejection opens its epoch on the raid's cbt — when a current member leaves a raid1 (removed, hot-removed or failed), the cbt stacked on the raid opens an epoch named after the member's UUID, the identity the control plane derives from the replica it placed there, and records from then on every write the member misses, so its return is a delta instead of a full rebuild. 0019 asked each survivor's own name for it, a bdev no cbt sits on, and no epoch ever opened. Only for a member that left current (its superblock slot configured at its peers' generation, not a write-only joiner, not a failed member an earlier removal failed to take), and only while the cbt carries no live epoch: a freeze empties the live bitmap, and an epoch opened after one may start short of writes the member missed just before it left. `bdev_raid_ut` pins the removed and the failed member, and the five ejections that open none | bdev_raid, cbt module (`vbdev_cbt_auto_epoch_open`), test/unit/lib/bdev/raid | 0016, 0019, 0033 | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From eac15e9404d05e9dea6ce4fee355cb6d6a97864e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Fri, 2 Oct 2026 09:12:34 +0200 Subject: [PATCH 20/26] fix(raid): a raid1 ejection opens its epoch on the raid's cbt (0052) 0019 opened the epoch at a member's ejection under each survivor's own name. The cbt sits on the raid, not on its members: every call answered -ENODEV, no epoch ever opened, and a member lost without notice was always rebuilt in full. The raid now asks once, for the raid's name and the member's UUID, the identity the control plane derives from the replica it placed there. vbdev_cbt_auto_epoch_open finds the cbt stacked on that bdev and opens the epoch there. An epoch holds the writes a member misses from its departure on, so it opens only for a member that left current: its superblock slot configured at the generation its peers carry, not a write-only joiner, and not a failed member an earlier removal failed to take (a new removal_failed flag, which unlike remove_error survives the retry). The cbt also refuses while it carries any live epoch, frozen or rebuilding included: a freeze empties the live bitmap, and an epoch opened after one may start short of writes the member missed just before it left. Each of those members is rebuilt in full, as before. --- module/bdev/cbt/vbdev_cbt.c | 43 ++++-- module/bdev/cbt/vbdev_cbt_query.h | 22 +-- ...tion-opens-an-epoch-on-the-raids-cbt.patch | 139 ++++++++++++++++++ patches/README.md | 2 +- 4 files changed, 183 insertions(+), 23 deletions(-) diff --git a/module/bdev/cbt/vbdev_cbt.c b/module/bdev/cbt/vbdev_cbt.c index b4221f2..0c11661 100644 --- a/module/bdev/cbt/vbdev_cbt.c +++ b/module/bdev/cbt/vbdev_cbt.c @@ -503,44 +503,61 @@ cbt_epoch_state_name(enum cbt_epoch_state state) } } -/* Epoch-at-ejection — see vbdev_cbt_query.h. Refuses (-EEXIST) when an OPEN - * epoch already bounds the round: an implicit takeover would steal the nonce - * from under the controller. FROZEN/REBUILDING epochs do not block — those are - * prior rounds being digested while the live bitmap keeps tracking. */ +/* The cbt bdev stacked on the bdev named base_bdev_name, if any. */ +static struct vbdev_cbt * +cbt_find_by_base_name(const char *base_bdev_name) +{ + struct vbdev_cbt *node; + + TAILQ_FOREACH(node, &g_cbt_nodes, link) { + if (node->base_bdev != NULL && + strcmp(spdk_bdev_get_name(node->base_bdev), base_bdev_name) == 0) { + return node; + } + } + return NULL; +} + +/* Epoch-at-ejection — see vbdev_cbt_query.h. Refuses (-EEXIST) while the cbt + * carries a live epoch. An open one bounds a round the controller started, and + * taking it over would steal its nonce. A frozen or rebuilding one may have taken + * out of the live bitmap, at its freeze, writes the member missed just before it + * left: an epoch opened now would start short of them. */ int -vbdev_cbt_auto_epoch_open(const char *bdev_name, const char *stale_backend_id) +vbdev_cbt_auto_epoch_open(const char *base_bdev_name, const char *member_id) { struct vbdev_cbt *cbt; struct cbt_epoch *ep; char epoch_id[CBT_EPOCH_ID_MAX]; char nonce[CBT_NONCE_MAX]; + const char *cbt_name; uint64_t max_gen = 0; uint64_t ticks = spdk_get_ticks(); assert(spdk_get_thread() == spdk_thread_get_app_thread()); - cbt = cbt_find_by_name(bdev_name); + cbt = cbt_find_by_base_name(base_bdev_name); if (cbt == NULL) { return -ENODEV; } + if (cbt_any_epoch_open(cbt)) { + return -EEXIST; + } TAILQ_FOREACH(ep, &cbt->epochs, link) { - if (ep->state == CBT_EPOCH_OPEN) { - return -EEXIST; - } if (ep->generation > max_gen) { max_gen = ep->generation; } } + cbt_name = spdk_bdev_get_name(&cbt->cbt_bdev); snprintf(epoch_id, sizeof(epoch_id), "auto-%016" PRIx64, ticks); snprintf(nonce, sizeof(nonce), "auto%08x", (uint32_t)ticks); - SPDK_NOTICELOG("CBT: auto epoch '%s' (nonce %s) on '%s' — member '%s' ejected\n", - epoch_id, nonce, bdev_name, stale_backend_id); + SPDK_NOTICELOG("CBT: auto epoch '%s' (nonce %s) on '%s' — member '%s' left '%s'\n", + epoch_id, nonce, cbt_name, member_id, base_bdev_name); - return bdev_cbt_epoch_open(bdev_name, epoch_id, stale_backend_id, - max_gen + 1, nonce); + return bdev_cbt_epoch_open(cbt_name, epoch_id, member_id, max_gen + 1, nonce); } /* Cross-module query — the raid module publishes these facts per member in ITS diff --git a/module/bdev/cbt/vbdev_cbt_query.h b/module/bdev/cbt/vbdev_cbt_query.h index 0053d7c..b7571d3 100644 --- a/module/bdev/cbt/vbdev_cbt_query.h +++ b/module/bdev/cbt/vbdev_cbt_query.h @@ -33,17 +33,21 @@ struct vbdev_cbt_epoch_facts { int vbdev_cbt_query_latest_epoch(const char *bdev_name, struct vbdev_cbt_epoch_facts *out); /** - * Open a delta epoch bounding an unplanned member loss. The raid module calls - * this on each surviving member's cbt when a member leaves, so the missing - * writes stay a delta instead of degrading to a full rebuild. The generated - * epoch id and nonce are reported through get_bdevs, which is how the - * control-plane adopts the round. + * Open a delta epoch bounding an unplanned member loss, on the cbt bdev stacked + * on the bdev named \c base_bdev_name. The raid module calls this when a current + * member leaves a raid1, with the raid's name and the member's UUID as + * \c member_id: from then on the cbt records every write the member misses, so + * its return is a delta instead of a full rebuild. The generated epoch id and + * nonce are reported through get_bdevs, which is how the control-plane adopts + * the round. * - * \return 0 on success; -EEXIST if an OPEN epoch already tracks the round - * (never take over implicitly); -ENODEV if \c bdev_name is not a cbt - * bdev; other negative errno from the epoch machinery. App thread only. + * \return 0 on success; -EEXIST if the cbt carries a live epoch (open, frozen or + * rebuilding): never take a round over, nor start one short of writes a + * freeze took out of the live bitmap; -ENODEV if no cbt bdev sits on + * \c base_bdev_name; other negative errno from the epoch machinery. App + * thread only. */ -int vbdev_cbt_auto_epoch_open(const char *bdev_name, const char *stale_backend_id); +int vbdev_cbt_auto_epoch_open(const char *base_bdev_name, const char *member_id); /* One dirty range of an epoch, in blocks of the cbt bdev: the raid's own, since * a cbt bdev passes its LBAs through. */ diff --git a/patches/0052-raid1-ejection-opens-an-epoch-on-the-raids-cbt.patch b/patches/0052-raid1-ejection-opens-an-epoch-on-the-raids-cbt.patch index 394ae22..f383feb 100644 --- a/patches/0052-raid1-ejection-opens-an-epoch-on-the-raids-cbt.patch +++ b/patches/0052-raid1-ejection-opens-an-epoch-on-the-raids-cbt.patch @@ -1,3 +1,142 @@ +diff --git a/module/bdev/raid/bdev_raid.c b/module/bdev/raid/bdev_raid.c +index da919f1..22ea08b 100644 +--- a/module/bdev/raid/bdev_raid.c ++++ b/module/bdev/raid/bdev_raid.c +@@ -516,6 +516,7 @@ raid_bdev_free_base_bdev_resource(struct raid_base_bdev_info *base_info) + /* Evariops 0033: reset remove_error with is_failed, so a reused slot never + * inherits the previous incarnation's removal failure. */ + base_info->remove_error = 0; ++ base_info->removal_failed = false; + raid_bdev_base_bdev_state_touch(base_info); + + /* clear `data_offset` to allow it to be recalculated during configuration */ +@@ -2271,12 +2272,69 @@ raid_bdev_remove_base_bdev_write_sb_cb(int status, struct raid_bdev *raid_bdev, + raid_bdev_remove_base_bdev_cont(base_info); + } + ++static uint64_t raid_bdev_sb_max_content_generation(const struct raid_bdev_superblock *sb); ++ ++/* Evariops 0052: whether the member leaving held every write the raid acknowledged ++ * before it left: its superblock slot CONFIGURED at the generation its peers carry ++ * (read before the ejection rewrites the slot), not a write-only joiner, and not a ++ * failed member whose earlier removal failed, which may have missed writes since. */ ++static bool ++raid_bdev_base_bdev_left_current(struct raid_base_bdev_info *base_info) ++{ ++ const struct raid_bdev_superblock *sb = base_info->raid_bdev->sb; ++ uint8_t slot = raid_bdev_base_bdev_slot(base_info); ++ uint8_t i; ++ ++ if (sb == NULL || base_info->write_only_pending || base_info->read_excluded || ++ base_info->removal_failed) { ++ return false; ++ } ++ ++ for (i = 0; i < sb->base_bdevs_size; i++) { ++ const struct raid_bdev_sb_base_bdev *sb_base_bdev = &sb->base_bdevs[i]; ++ ++ if (sb_base_bdev->state == RAID_SB_BASE_BDEV_CONFIGURED && sb_base_bdev->slot == slot) { ++ return sb_base_bdev->content_generation == raid_bdev_sb_max_content_generation(sb); ++ } ++ } ++ ++ return false; ++} ++ ++/* Evariops 0019, 0052: bound the delta the moment a raid1 member leaves. The cbt ++ * bdev stacked on the raid opens an epoch named after the member's UUID, which the ++ * control plane derives from the replica it placed there, and records from then on ++ * every write the member misses. Only for a member that left current: the epoch ++ * holds what it misses from now on, not what it never had. -EEXIST means the cbt ++ * already carries a live epoch, -ENODEV that no cbt sits on the raid; the member is ++ * then rebuilt in full, as it is when it leaves behind. */ ++static void ++raid_bdev_open_ejection_epoch(struct raid_base_bdev_info *base_info) ++{ ++ struct raid_bdev *raid_bdev = base_info->raid_bdev; ++ char member_id[SPDK_UUID_STRING_LEN]; ++ int rc; ++ ++ if (raid_bdev->level != RAID1 || spdk_uuid_is_null(&base_info->uuid) || ++ !raid_bdev_base_bdev_left_current(base_info)) { ++ return; ++ } ++ ++ spdk_uuid_fmt_lower(member_id, sizeof(member_id), &base_info->uuid); ++ rc = vbdev_cbt_auto_epoch_open(raid_bdev->bdev.name, member_id); ++ if (rc != 0 && rc != -EEXIST && rc != -ENODEV) { ++ SPDK_WARNLOG("raid bdev '%s': no epoch for ejected member '%s' (%s): %s; it will be " ++ "rebuilt in full\n", raid_bdev->bdev.name, ++ base_info->name != NULL ? base_info->name : "(null)", member_id, ++ spdk_strerror(-rc)); ++ } ++} ++ + static void + raid_bdev_remove_base_bdev_on_quiesced(void *ctx, int status) + { + struct raid_base_bdev_info *base_info = ctx; + struct raid_bdev *raid_bdev = base_info->raid_bdev; +- struct raid_base_bdev_info *survivor; + + if (status != 0) { + SPDK_ERRLOG("Failed to quiesce raid bdev %s: %s\n", +@@ -2285,28 +2343,7 @@ raid_bdev_remove_base_bdev_on_quiesced(void *ctx, int status) + return; + } + +- /* Evariops 0019: bound the delta the moment the member leaves, by opening +- * an auto epoch on every surviving member's cbt — each survivor then records +- * every post-ejection write. -EEXIST means a live round already bounds it, +- * -ENODEV that the member is not cbt-wrapped; both are expected. */ +- if (raid_bdev->level == RAID1 && base_info->name != NULL) { +- RAID_FOR_EACH_BASE_BDEV(raid_bdev, survivor) { +- int rc; +- +- if (survivor == base_info || !survivor->is_configured || +- survivor->name == NULL) { +- continue; +- } +- rc = vbdev_cbt_auto_epoch_open(survivor->name, base_info->name); +- if (rc != 0 && rc != -EEXIST && rc != -ENODEV) { +- SPDK_WARNLOG("raid bdev '%s': auto epoch on '%s' for " +- "ejected '%s' failed: %s — the debt path " +- "will degrade to a full rebuild (D14)\n", +- raid_bdev->bdev.name, survivor->name, +- base_info->name, spdk_strerror(-rc)); +- } +- } +- } ++ raid_bdev_open_ejection_epoch(base_info); + + if (raid_bdev->sb) { + struct raid_bdev_superblock *sb = raid_bdev->sb; +@@ -2599,6 +2636,9 @@ raid_bdev_fail_base_remove_cb(void *ctx, int status) + * local reflex. Service I/O is unaffected either way: raid1 routes on the + * slot's channel, never on is_failed. */ + base_info->remove_error = status; ++ /* Evariops 0052: unlike remove_error, which the next attempt clears, this ++ * stays until the slot is freed. */ ++ base_info->removal_failed = true; + raid_bdev_base_bdev_state_touch(base_info); + + SPDK_NOTICELOG("Failed to remove failed base bdev '%s' from raid bdev '%s' slot %d: %s (%d) — the member " +diff --git a/module/bdev/raid/bdev_raid.h b/module/bdev/raid/bdev_raid.h +index 947359c..8f5d692 100644 +--- a/module/bdev/raid/bdev_raid.h ++++ b/module/bdev/raid/bdev_raid.h +@@ -116,6 +116,11 @@ struct raid_base_bdev_info { + * the only thing that re-arms the local fail path. */ + int remove_error; + ++ /* Evariops 0052: a removal of this failed member has itself failed. The member ++ * may have missed writes since, so the removal that finally takes it opens no ++ * CBT epoch. Cleared with remove_error when the slot is freed. */ ++ bool removal_failed; ++ + /* callback for base bdev configuration */ + raid_base_bdev_cb configure_cb; + diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c index a488d6d..c359160 100644 --- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c diff --git a/patches/README.md b/patches/README.md index 8035ceb..41421bd 100644 --- a/patches/README.md +++ b/patches/README.md @@ -26,7 +26,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0052`) | 0016 | raid extended superblock — per-member `content_generation`/`view_epoch` persisted in the same transaction | bdev_raid (SB minor 0→1, carved from reserved bytes) | 0013 | | 0017 | raid per-member observation — `state`/`since`/generations plus live cbt epoch facts in `get_bdevs` | bdev_raid; `#include`s `../cbt/vbdev_cbt_query.h` (requires the cbt module, like 0005 requires tier) | 0013, 0016, cbt module | | 0018 | raid integrated verify — sampled windows compared post-copy under `quiesce_range`, re-copied before declaring DIVERGENT, with a `verified` seal gating `epoch_close(consumed)` | bdev_raid | 0013, 0015, 0016 | -| 0019 | raid auto epoch at member ejection — opens a cbt epoch on the survivors so their tracking bounds the later delta; the nonce is reported in `get_bdevs` | bdev_raid; calls `vbdev_cbt_auto_epoch_open` (cbt module) | 0013, cbt module | +| 0019 | raid auto epoch at member ejection — opens a cbt epoch so its tracking bounds the later delta; the nonce is reported in `get_bdevs`. Asked on the survivors' names, it never opened one: 0052 opens it on the cbt stacked on the raid | bdev_raid; calls `vbdev_cbt_auto_epoch_open` (cbt module) | 0013, cbt module | | 0020 | nvmf audited force-resume — lets the fence path break a standing pause barrier deliberately, leaving an audit record | lib/nvmf (nvmf_pause_rpc.c) | 0003, 0011 | | 0021 | raid envelopes — per-class bandwidth caps × (nominal, maintenance) and a rebuild concurrency bound, via `bdev_raid_set/get_envelopes` | bdev_raid (adds `bdev_raid_envelopes.{c,h}`) | 0013 | | 0022 | raid `verify_ranges` — exhaustive divergence detector, LBA-locked and envelope-paced, which reports and never repairs; raid1 compares copies, raid5f checks the XOR syndrome per stripe | bdev_raid (adds `bdev_raid_verify_ranges.c`) | 0008, 0014, 0015, 0021 | From 34a51639b7fe0ff34fbeba6455d52f896331902f Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Fri, 2 Oct 2026 10:17:29 +0200 Subject: [PATCH 21/26] feat(cbt): keep the live bitmap in rotating windows, so an ejection's epoch is a delta (bdev_cbt_rotate) The live bitmap is cleared only by bdev_cbt_reset, which needs a moment when every backend is known in sync. A member that leaves without notice gives none, so the epoch the raid opens at its ejection took the whole live bitmap: every write since the last reset. On a volume that went days without an epoch, that is the whole device, and the "delta" rebuild copied all of it (measured on a lab cluster: 81920 of 81920 chunks). bdev_cbt_rotate, called periodically while no epoch is live, moves the live bitmap to a previous-window bitmap (the window before is dropped) and starts it again empty. vbdev_cbt_auto_epoch_open ORs the previous window back before it opens, so its delta holds one to two windows of writes before the departure and everything after. Two rotations are at least CBT_ROTATE_MIN_INTERVAL_US (10 s) apart whoever asks: a write the member missed a few milliseconds before its ejection is always in one of the two bitmaps. A rotation is refused (-EBUSY) while an epoch is live, since that epoch reads the live bitmap whole; a reset clears both. --- module/bdev/cbt/README.md | 2 ++ module/bdev/cbt/vbdev_cbt.c | 52 ++++++++++++++++++++++++++-- module/bdev/cbt/vbdev_cbt.h | 19 ++++++++++ module/bdev/cbt/vbdev_cbt_internal.h | 9 +++++ module/bdev/cbt/vbdev_cbt_rpc.c | 32 +++++++++++++++++ 5 files changed, 112 insertions(+), 2 deletions(-) diff --git a/module/bdev/cbt/README.md b/module/bdev/cbt/README.md index 3d4f7d7..67381b5 100644 --- a/module/bdev/cbt/README.md +++ b/module/bdev/cbt/README.md @@ -55,6 +55,8 @@ The module cannot converge on its own when the write rate approaches rebuild ban Clearing is reset-driven: there is no automatic clear. Once the backend is re-added and all backends are synchronized, the orchestrator MUST call `bdev_cbt_reset`, or the bitmap grows monotonically and "partial" rebuilds degrade toward full-surface copies. +A reset needs a moment when every backend is known in sync, and a member that leaves without notice gives none: the epoch the raid opens at its ejection (`vbdev_cbt_auto_epoch_open`) cannot clear anything, since the writes the member just missed are in the live bitmap. So the live bitmap is also kept in windows: `bdev_cbt_rotate`, called periodically while no epoch is live, moves it to a previous-window bitmap and starts it again empty, and the epoch at ejection takes the previous window back before it opens. Its delta then holds one to two windows of writes before the departure, and everything after, instead of every write since the last reset. Two rotations are at least `CBT_ROTATE_MIN_INTERVAL_US` (10 s) apart whoever asks, so a write missed a few milliseconds before the ejection is always in one of the two bitmaps. A reset clears both. + ## RAID integration The companion patch (`patches/0001-raid-add-skip_rebuild-parameter.patch`) adds a `skip_rebuild` boolean to `bdev_raid_add_base_bdev`. When true, the RAID module skips its full surface rebuild and instead quiesces the raid, opens `base_channel[slot]` on every existing I/O channel for the re-added bdev — without this, existing channels would never write to the backend — then unquiesces and writes the superblock. diff --git a/module/bdev/cbt/vbdev_cbt.c b/module/bdev/cbt/vbdev_cbt.c index 0c11661..359f33b 100644 --- a/module/bdev/cbt/vbdev_cbt.c +++ b/module/bdev/cbt/vbdev_cbt.c @@ -550,6 +550,15 @@ vbdev_cbt_auto_epoch_open(const char *base_bdev_name, const char *member_id) } } + /* The member may have missed writes from just before a rotation: those moved + * to the previous window, and come back into the delta here. With no live + * epoch, no rotation runs until this epoch closes. */ + for (uint64_t i = 0; i < cbt->bitmap_size_bytes; i++) { + if (cbt->bitmap_prev[i] != 0) { + __atomic_fetch_or(&cbt->bitmap[i], cbt->bitmap_prev[i], __ATOMIC_ACQ_REL); + } + } + cbt_name = spdk_bdev_get_name(&cbt->cbt_bdev); snprintf(epoch_id, sizeof(epoch_id), "auto-%016" PRIx64, ticks); snprintf(nonce, sizeof(nonce), "auto%08x", (uint32_t)ticks); @@ -693,6 +702,7 @@ _cbt_device_unregister_cb(void *io_device) struct vbdev_cbt *cbt_node = io_device; free(cbt_node->bitmap); + free(cbt_node->bitmap_prev); free(cbt_node->cbt_bdev.name); struct cbt_epoch *ep; @@ -895,14 +905,18 @@ vbdev_cbt_register(const char *bdev_name) cbt_node->total_blocks = bdev->blockcnt; cbt_node->bitmap = calloc(1, cbt_node->bitmap_size_bytes); - if (!cbt_node->bitmap) { - SPDK_ERRLOG("CBT: bitmap allocation failed (%lu bytes)\n", + cbt_node->bitmap_prev = calloc(1, cbt_node->bitmap_size_bytes); + if (!cbt_node->bitmap || !cbt_node->bitmap_prev) { + SPDK_ERRLOG("CBT: bitmap allocation failed (2 x %lu bytes)\n", (unsigned long)cbt_node->bitmap_size_bytes); spdk_bdev_close(cbt_node->base_desc); + free(cbt_node->bitmap); + free(cbt_node->bitmap_prev); free(cbt_node->cbt_bdev.name); free(cbt_node); return -ENOMEM; } + cbt_node->rotated_at = spdk_get_ticks(); /* ── Copy geometry from base bdev ── */ cbt_node->cbt_bdev.write_cache = bdev->write_cache; @@ -950,6 +964,7 @@ vbdev_cbt_register(const char *bdev_name) TAILQ_REMOVE(&g_cbt_nodes, cbt_node, link); spdk_io_device_unregister(cbt_node, NULL); free(cbt_node->bitmap); + free(cbt_node->bitmap_prev); free(cbt_node->cbt_bdev.name); free(cbt_node); return rc; @@ -963,6 +978,7 @@ vbdev_cbt_register(const char *bdev_name) TAILQ_REMOVE(&g_cbt_nodes, cbt_node, link); spdk_io_device_unregister(cbt_node, NULL); free(cbt_node->bitmap); + free(cbt_node->bitmap_prev); free(cbt_node->cbt_bdev.name); free(cbt_node); return rc; @@ -2542,6 +2558,38 @@ bdev_cbt_reset(const char *cbt_name) __atomic_thread_fence(__ATOMIC_ACQUIRE); memset(cbt->bitmap, 0, cbt->bitmap_size_bytes); + /* The caller vouches every backend in sync through now: the history window + * goes with the rest. */ + memset(cbt->bitmap_prev, 0, cbt->bitmap_size_bytes); + return 0; +} + +int +bdev_cbt_rotate(const char *cbt_name) +{ + struct vbdev_cbt *cbt = cbt_find_by_name(cbt_name); + uint64_t now = spdk_get_ticks(); + + assert(spdk_get_thread() == spdk_thread_get_app_thread()); + + if (!cbt) { + return -ENODEV; + } + /* A live epoch reads the live bitmap whole: moving bits out of it would take + * them out of that epoch's delta. */ + if (cbt_any_epoch_open(cbt)) { + return -EBUSY; + } + if (now - cbt->rotated_at < CBT_ROTATE_MIN_INTERVAL_US * spdk_get_ticks_hz() / SPDK_SEC_TO_USEC) { + return -EAGAIN; + } + + /* Per-byte exchange, as at freeze: an I/O thread OR-ing a bit between a copy + * and a clear would lose it. */ + for (uint64_t i = 0; i < cbt->bitmap_size_bytes; i++) { + cbt->bitmap_prev[i] = __atomic_exchange_n(&cbt->bitmap[i], 0, __ATOMIC_ACQ_REL); + } + cbt->rotated_at = now; return 0; } diff --git a/module/bdev/cbt/vbdev_cbt.h b/module/bdev/cbt/vbdev_cbt.h index e072c62..735a82a 100644 --- a/module/bdev/cbt/vbdev_cbt.h +++ b/module/bdev/cbt/vbdev_cbt.h @@ -24,6 +24,10 @@ extern "C" { #define CBT_EPOCH_ID_MAX 64 #define CBT_BACKEND_ID_MAX 128 #define CBT_NONCE_MAX 32 /* control-plane epoch nonce */ +/* Two rotations of the history window are at least this far apart, whoever asks: + * a write a member missed just before it left must still be in one of the two + * bitmaps when the epoch for that member opens, a few milliseconds later. */ +#define CBT_ROTATE_MIN_INTERVAL_US 10000000 #define CBT_REBUILD_TOKEN_MAX 192 /* consumed-close proof token */ #define CBT_REBUILD_DEFAULT_QD 16 #define CBT_REBUILD_MAX_QD 128 @@ -260,6 +264,21 @@ int bdev_cbt_epoch_get_dirty_ranges(const char *cbt_name, const char *epoch_id, uint32_t *out_chunk_size_kb, bool *out_truncated); +/** + * Start a new history window: the live bitmap moves to the previous-window bitmap + * (dropping the window before it) and starts again empty. Without it, the live + * bitmap holds every write since the last reset, and the epoch opened when a + * member leaves (vbdev_cbt_auto_epoch_open) would copy all of them, on a volume + * that went days without an epoch the whole device. That epoch takes the previous + * window back, so it covers one to two windows before the departure and + * everything after. + * + * \return 0; -EBUSY while an epoch is open, frozen or rebuilding (it reads the + * live bitmap whole); -EAGAIN less than CBT_ROTATE_MIN_INTERVAL_US after + * the previous rotation; -ENODEV if \c cbt_name is not a cbt bdev. + */ +int bdev_cbt_rotate(const char *cbt_name); + /* ── Legacy aliases (deprecated, will be removed in v2) ────────────── */ int bdev_cbt_start_tracking(const char *cbt_name); diff --git a/module/bdev/cbt/vbdev_cbt_internal.h b/module/bdev/cbt/vbdev_cbt_internal.h index 126c0ee..3bdc535 100644 --- a/module/bdev/cbt/vbdev_cbt_internal.h +++ b/module/bdev/cbt/vbdev_cbt_internal.h @@ -33,6 +33,15 @@ struct vbdev_cbt { * bdev_cbt_reset, refused while any epoch is active. Nothing inside the * target can know that every distributed backend is in sync. */ + /* ── History window (bdev_cbt_rotate) ── */ + /* The writes of the previous window, rotated out of the live bitmap while no + * epoch was active. An epoch opened when a member leaves takes them back: + * the writes the member missed just before it left are in one of the two. */ + uint8_t *bitmap_prev; + /* Ticks of the last rotation; rotations are at least + * CBT_ROTATE_MIN_INTERVAL_US apart. */ + uint64_t rotated_at; + /* ── Epoch management ── */ TAILQ_HEAD(, cbt_epoch) epochs; uint64_t epoch_count; diff --git a/module/bdev/cbt/vbdev_cbt_rpc.c b/module/bdev/cbt/vbdev_cbt_rpc.c index 8c8f2d5..148b2ed 100644 --- a/module/bdev/cbt/vbdev_cbt_rpc.c +++ b/module/bdev/cbt/vbdev_cbt_rpc.c @@ -702,6 +702,38 @@ rpc_bdev_cbt_reset(struct spdk_jsonrpc_request *request, } SPDK_RPC_REGISTER("bdev_cbt_reset", rpc_bdev_cbt_reset, SPDK_RPC_RUNTIME) +/* bdev_cbt_rotate: start a new history window (see bdev_cbt_rotate). Answers the + * errno as the JSON-RPC error code, as bdev_cbt_reset does: -EBUSY while an epoch + * is live, -EAGAIN when called again too soon. */ +static void +rpc_bdev_cbt_rotate(struct spdk_jsonrpc_request *request, + const struct spdk_json_val *params) +{ + struct rpc_cbt_name_only req = {NULL}; + int rc; + + if (spdk_json_decode_object(params, rpc_cbt_name_only_decoders, + SPDK_COUNTOF(rpc_cbt_name_only_decoders), &req)) { + spdk_jsonrpc_send_error_response(request, SPDK_JSONRPC_ERROR_INTERNAL_ERROR, + "Failed to decode parameters"); + goto cleanup; + } + + rc = bdev_cbt_rotate(req.name); + if (rc != 0) { + spdk_jsonrpc_send_error_response(request, rc, spdk_strerror(-rc)); + goto cleanup; + } + + struct spdk_json_write_ctx *w = spdk_jsonrpc_begin_result(request); + spdk_json_write_bool(w, true); + spdk_jsonrpc_end_result(request, w); + +cleanup: + free_rpc_cbt_name_only(&req); +} +SPDK_RPC_REGISTER("bdev_cbt_rotate", rpc_bdev_cbt_rotate, SPDK_RPC_RUNTIME) + /* ================================================================== */ /* bdev_cbt_partial_rebuild (async — deferred RPC response) */ /* ================================================================== */ From b6fbbcac331d6d96a5fbfd58b78846ef4d861755 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Sat, 3 Oct 2026 14:20:16 +0200 Subject: [PATCH 22/26] feat(raid): a raid create in a view arbitrates its members from their superblocks (0053) A raid create took every listed leg as a configured member. The control plane had to clear the legs' superblocks first, since a leg stamped by another raid is refused, and clearing erased the only evidence of which leg holds the volume's last acknowledged writes: a leg left behind by an ejection came back as a configured member, and a creator that lost its place could take legs from a raid still serving. bdev_raid_create with a view_epoch (a raid1 with a superblock, the volume's lineage and the creating incarnation) now reads every listed leg's superblock, opened read-only, before it claims any: - A superblock carries a stamp, minor 2, carved out of the header's reserved bytes: the lineage, the owner incarnation and the last view epoch it stamped. A minor-1 superblock reads as unstamped. - A leg stamped for another volume is refused (-EEXIST). A leg another owner stamped in the view's epoch or a newer one is refused (-EBUSY), nothing claimed or written, unless force_restamp names that exact epoch (expected_view_epoch; -ESTALE for any other). - Among the legs stamped for the volume, only those at the highest content generation, configured in their own superblock, are configured, at that generation. Every other leg (behind, a copy never completed, unstamped, or stamped before minor 2) stays out, in an empty slot where the superblock records it MISSING at its own generation, and the control plane rebuilds it. No stamped leg holding a configured copy is -ENODATA. When no leg carries a stamp of the volume, every leg is configured, as before. - Each member is read again as the raid claims it and must still be what the arbitration read (-ESTALE). A leg the survey cannot read refuses the create: read as blank, the one leg at the highest generation would be left out, never read again, and the raid would come up on the legs behind it. - An add into an online raid of a lineage checks the leg the same way and takes it as a new member, rebuilt like any other, where upstream refuses any leg another raid stamped. - A raid reassembled from a stamped superblock keeps its view, and get_bdevs reports lineage and view_epoch. declared_slots creates a raid wider than its legs, the slots past them empty and sized like the members, for members added later: the raid comes online degraded instead of through null bdevs added and removed. A level that operates only whole refuses it. bdev_raid_ut gets a raid1-level module and a superblock per leg, and pins the arbitration, each refusal, the claim-time re-read, the empty slots, the online add and the reassembly (14 tests); bdev_raid_sb_ut pins the stamp and the slots init_superblock writes. Each rule was taken out in turn and its tests went red: the generation (3 tests), the admission (3), the claim-time re-read (1), the online add (1), the empty slot size (3, and an online add hits the data_size assert), the reassembly (1), the stamp and the MISSING slots (bdev_raid_sb_ut), a read error read as blank (1), the view parameters (1). Every raid suite passes: bdev_raid 39, bdev_raid_sb 15, raid0 8, raid1 7, raid5f 8, concat 3. --- ...s-its-members-from-their-superblocks.patch | 2120 +++++++++++++++++ patches/README.md | 3 +- 2 files changed, 2122 insertions(+), 1 deletion(-) create mode 100644 patches/0053-raid-create-arbitrates-its-members-from-their-superblocks.patch diff --git a/patches/0053-raid-create-arbitrates-its-members-from-their-superblocks.patch b/patches/0053-raid-create-arbitrates-its-members-from-their-superblocks.patch new file mode 100644 index 0000000..f3e6048 --- /dev/null +++ b/patches/0053-raid-create-arbitrates-its-members-from-their-superblocks.patch @@ -0,0 +1,2120 @@ +diff --git a/module/bdev/raid/bdev_raid.c b/module/bdev/raid/bdev_raid.c +index 22ea08b..c18bde7 100644 +--- a/module/bdev/raid/bdev_raid.c ++++ b/module/bdev/raid/bdev_raid.c +@@ -1225,6 +1225,11 @@ raid_bdev_write_info_json(struct raid_bdev *raid_bdev, struct spdk_json_write_ct + if (raid_bdev->incarnation != NULL) { + spdk_json_write_named_string(w, "incarnation", raid_bdev->incarnation); + } ++ /* Evariops 0053: absent for a raid created outside the view protocol. */ ++ if (!spdk_uuid_is_null(&raid_bdev->lineage)) { ++ spdk_json_write_named_uuid(w, "lineage", &raid_bdev->lineage); ++ spdk_json_write_named_uint64(w, "view_epoch", raid_bdev->view_epoch); ++ } + spdk_json_write_named_string(w, "state", raid_bdev_state_to_str(raid_bdev->state)); + /* Evariops 0034: the verify holds an LBA range lock for a long time. Publish + * who holds it and since when, so a stalled I/O can be attributed. */ +@@ -2023,6 +2028,33 @@ raid_bdev_configure_write_sb_cb(int status, struct raid_bdev *raid_bdev, void *c + } + } + ++/* Evariops 0053: a slot its create left empty takes the data size of the members ++ * it configured: the level sizes the raid over every slot, and a member added ++ * there later is checked against it. Upstream configures a raid with every slot ++ * filled, or from a superblock that records the size of each. */ ++static void ++raid_bdev_size_empty_slots(struct raid_bdev *raid_bdev) ++{ ++ struct raid_base_bdev_info *base_info; ++ uint64_t data_size = UINT64_MAX; ++ ++ RAID_FOR_EACH_BASE_BDEV(raid_bdev, base_info) { ++ if (base_info->is_configured) { ++ data_size = spdk_min(data_size, base_info->data_size); ++ } ++ } ++ ++ if (data_size == UINT64_MAX) { ++ return; ++ } ++ ++ RAID_FOR_EACH_BASE_BDEV(raid_bdev, base_info) { ++ if (!base_info->is_configured && base_info->data_size == 0) { ++ base_info->data_size = data_size; ++ } ++ } ++} ++ + /* + * brief: + * If raid bdev config is complete, then only register the raid bdev to +@@ -2054,6 +2086,11 @@ raid_bdev_configure(struct raid_bdev *raid_bdev, raid_bdev_configure_cb cb, void + } + raid_bdev->strip_size_shift = spdk_u32log2(raid_bdev->strip_size); + ++ /* Evariops 0053: every member the create arbitrated has been re-read as ++ * it was claimed. */ ++ raid_bdev->arbitrating = false; ++ raid_bdev_size_empty_slots(raid_bdev); ++ + rc = raid_bdev->module->start(raid_bdev); + if (rc != 0) { + SPDK_ERRLOG("raid module startup callback failed\n"); +@@ -4696,6 +4733,420 @@ raid_bdev_configure_base_bdev_cont(struct raid_base_bdev_info *base_info) + } + } + ++/* Evariops 0053: whether a superblock carries a stamp: minor 2 or later, with a ++ * lineage. Below minor 2 the stamp's bytes were reserved. */ ++static bool ++raid_bdev_sb_stamped(const struct raid_bdev_superblock *sb) ++{ ++ return sb->version.minor >= 2 && !spdk_uuid_is_null(&sb->lineage); ++} ++ ++/* Evariops 0053: reads into stamp what sb says of the leg member: the stamp of ++ * the raid that wrote it, and the leg's own slot, the first entry naming it as ++ * examine finds it. */ ++static void ++raid_bdev_read_stamp(const struct raid_bdev_superblock *sb, const struct spdk_uuid *member, ++ struct raid_bdev_stamp *stamp) ++{ ++ uint8_t i; ++ ++ stamp->stamped = raid_bdev_sb_stamped(sb); ++ if (stamp->stamped) { ++ spdk_uuid_copy(&stamp->lineage, &sb->lineage); ++ snprintf(stamp->owner, sizeof(stamp->owner), "%.*s", (int)sizeof(sb->owner) - 1, ++ (const char *)sb->owner); ++ stamp->epoch = sb->view_epoch; ++ } ++ ++ for (i = 0; i < sb->base_bdevs_size; i++) { ++ const struct raid_bdev_sb_base_bdev *sb_base_bdev = &sb->base_bdevs[i]; ++ ++ if (spdk_uuid_compare(&sb_base_bdev->uuid, member) == 0) { ++ stamp->own_configured = sb_base_bdev->state == RAID_SB_BASE_BDEV_CONFIGURED; ++ stamp->own_generation = sb_base_bdev->content_generation; ++ stamp->own_view_epoch = sb_base_bdev->view_epoch; ++ return; ++ } ++ } ++} ++ ++/* Evariops 0053: whether a raid of view may take a leg so stamped. A leg of ++ * another lineage is another volume's (-EEXIST). A leg another owner stamped in ++ * the view's epoch or a newer one is held by a raid that a view at least as ++ * recent still counts, since an owner that stays in the view journals each new ++ * epoch on its legs (-EBUSY). Such a stamp is overwritten only at the epoch the ++ * view restamps, which the control plane names after reading it (-ESTALE at ++ * any other). An unstamped leg, or one stamped by this owner or at an older ++ * epoch, is free. */ ++static int ++raid_bdev_stamp_admits(const struct raid_bdev_view *view, const struct raid_bdev_stamp *stamp) ++{ ++ if (!stamp->stamped) { ++ return 0; ++ } ++ ++ if (spdk_uuid_compare(&stamp->lineage, &view->lineage) != 0) { ++ return -EEXIST; ++ } ++ ++ if (stamp->epoch < view->epoch || ++ (view->incarnation != NULL && strcmp(stamp->owner, view->incarnation) == 0)) { ++ return 0; ++ } ++ ++ if (view->restamp_epoch == 0) { ++ return -EBUSY; ++ } ++ ++ return stamp->epoch == view->restamp_epoch ? 0 : -ESTALE; ++} ++ ++/* Evariops 0053: the view a raid stamps its members in. Restamping is the ++ * create's alone. */ ++static void ++raid_bdev_view_of(const struct raid_bdev *raid_bdev, struct raid_bdev_view *view) ++{ ++ spdk_uuid_copy(&view->lineage, &raid_bdev->lineage); ++ view->epoch = raid_bdev->view_epoch; ++ view->incarnation = raid_bdev->incarnation; ++ view->restamp_epoch = raid_bdev->arbitrating ? raid_bdev->restamp_epoch : 0; ++} ++ ++/* Evariops 0053: whether a leg claimed by this raid is checked by its stamp: ++ * during the create that arbitrated the members, and on any add into an online ++ * raid of a lineage. A raid still assembling from its superblock keeps ++ * upstream's rules, as does a raid created outside the view protocol. */ ++static bool ++raid_bdev_checks_stamps(const struct raid_bdev *raid_bdev) ++{ ++ return raid_bdev->arbitrating || ++ (!spdk_uuid_is_null(&raid_bdev->lineage) && raid_bdev->state == RAID_BDEV_STATE_ONLINE); ++} ++ ++/* Evariops 0053: checks the superblock read from a leg as the raid claims it, ++ * sb NULL when the leg carries none. The view's rules first. During the create, ++ * the leg must still be what the arbitration read, or the create is stale ++ * (-ESTALE): a stamp of the lineage, configured in its own superblock at the ++ * arbitrated generation, or no stamp of it when the create migrated. An add ++ * into an online raid takes the leg as a new member, rebuilt like any other, ++ * whatever raid its superblock names, this one included. */ ++static int ++raid_bdev_check_stamp(struct raid_bdev *raid_bdev, const struct raid_bdev_superblock *sb, ++ const struct spdk_uuid *member) ++{ ++ struct raid_bdev_stamp stamp = {}; ++ struct raid_bdev_view view; ++ int rc; ++ ++ if (sb != NULL) { ++ raid_bdev_read_stamp(sb, member, &stamp); ++ } ++ ++ raid_bdev_view_of(raid_bdev, &view); ++ rc = raid_bdev_stamp_admits(&view, &stamp); ++ if (rc != 0 || !raid_bdev->arbitrating) { ++ return rc; ++ } ++ ++ if (raid_bdev->arbitration_migrates) { ++ return stamp.stamped ? -ESTALE : 0; ++ } ++ ++ if (!stamp.stamped || !stamp.own_configured || ++ stamp.own_generation != raid_bdev->arbitrated_generation) { ++ return -ESTALE; ++ } ++ ++ return 0; ++} ++ ++/* Evariops 0053: why a raid of a lineage refuses a leg, for its logs. */ ++static const char * ++raid_bdev_stamp_refusal(int rc) ++{ ++ switch (rc) { ++ case -EEXIST: ++ return "stamped for another volume"; ++ case -EBUSY: ++ return "stamped by another owner in the view's epoch or a newer one"; ++ case -ESTALE: ++ return "stamped otherwise than the create read it"; ++ default: ++ return spdk_strerror(-rc); ++ } ++} ++ ++struct raid_bdev_survey { ++ uint32_t remaining; ++ raid_bdev_survey_cb cb; ++ void *cb_ctx; ++}; ++ ++struct raid_bdev_survey_leg { ++ struct raid_bdev_survey *survey; ++ struct raid_bdev_stamp *stamp; ++ struct spdk_bdev_desc *desc; ++ struct spdk_io_channel *ch; ++}; ++ ++static void ++raid_bdev_survey_leg_done(struct raid_bdev_survey *survey) ++{ ++ assert(survey->remaining > 0); ++ if (--survey->remaining == 0) { ++ survey->cb(survey->cb_ctx); ++ free(survey); ++ } ++} ++ ++static void ++raid_bdev_survey_event_cb(enum spdk_bdev_event_type type, struct spdk_bdev *bdev, void *event_ctx) ++{ ++} ++ ++static void ++raid_bdev_survey_load_cb(const struct raid_bdev_superblock *sb, int status, void *ctx) ++{ ++ struct raid_bdev_survey_leg *leg = ctx; ++ struct raid_bdev_survey *survey = leg->survey; ++ ++ /* -EINVAL: the leg carries no valid superblock, a read like any other */ ++ leg->stamp->status = status == -EINVAL ? 0 : status; ++ if (status == 0) { ++ raid_bdev_read_stamp(sb, &leg->stamp->uuid, leg->stamp); ++ } ++ ++ spdk_put_io_channel(leg->ch); ++ spdk_bdev_close(leg->desc); ++ free(leg); ++ ++ raid_bdev_survey_leg_done(survey); ++} ++ ++static int ++raid_bdev_survey_leg(struct raid_bdev_survey *survey, const char *name, ++ struct raid_bdev_stamp *stamp) ++{ ++ struct raid_bdev_survey_leg *leg; ++ int rc; ++ ++ leg = calloc(1, sizeof(*leg)); ++ if (leg == NULL) { ++ return -ENOMEM; ++ } ++ leg->survey = survey; ++ leg->stamp = stamp; ++ ++ rc = spdk_bdev_open_ext(name, false, raid_bdev_survey_event_cb, NULL, &leg->desc); ++ if (rc != 0) { ++ free(leg); ++ return rc; ++ } ++ spdk_uuid_copy(&stamp->uuid, spdk_bdev_get_uuid(spdk_bdev_desc_get_bdev(leg->desc))); ++ ++ leg->ch = spdk_bdev_get_io_channel(leg->desc); ++ if (leg->ch == NULL) { ++ spdk_bdev_close(leg->desc); ++ free(leg); ++ return -ENOMEM; ++ } ++ ++ rc = raid_bdev_load_base_bdev_superblock(leg->desc, leg->ch, raid_bdev_survey_load_cb, leg); ++ if (rc != 0) { ++ spdk_put_io_channel(leg->ch); ++ spdk_bdev_close(leg->desc); ++ free(leg); ++ } ++ ++ return rc; ++} ++ ++int ++raid_bdev_survey(char *const *names, uint8_t num_legs, struct raid_bdev_stamp *stamps, ++ raid_bdev_survey_cb cb, void *cb_ctx) ++{ ++ struct raid_bdev_survey *survey; ++ uint8_t i; ++ int rc; ++ ++ assert(spdk_get_thread() == spdk_thread_get_app_thread()); ++ ++ survey = calloc(1, sizeof(*survey)); ++ if (survey == NULL) { ++ return -ENOMEM; ++ } ++ survey->cb = cb; ++ survey->cb_ctx = cb_ctx; ++ /* One per leg, and one the launch holds until every read has started: a ++ * read completing inline must not end the survey early. */ ++ survey->remaining = (uint32_t)num_legs + 1; ++ ++ for (i = 0; i < num_legs; i++) { ++ memset(&stamps[i], 0, sizeof(stamps[i])); ++ rc = raid_bdev_survey_leg(survey, names[i], &stamps[i]); ++ if (rc != 0) { ++ /* -EINVAL would read as a leg with no superblock */ ++ stamps[i].status = rc == -EINVAL ? -EIO : rc; ++ raid_bdev_survey_leg_done(survey); ++ } ++ } ++ ++ raid_bdev_survey_leg_done(survey); ++ ++ return 0; ++} ++ ++/* Evariops 0053: a create reads its legs before it takes any. The stamps of a ++ * lineage carry comparable generations: an ejection advances the generation of ++ * the members that stay, in the superblock transaction that records it, and a ++ * create configures its members at the generation they carry. So only the legs ++ * at the highest generation among those configured in their own superblock ++ * hold every acknowledged write: they are configured, and every other leg, one ++ * without a stamp of the lineage included, stays out to be rebuilt. When no leg ++ * carries a stamp of the lineage (blank, or stamped before minor 2, whose ++ * generations restarted at 0 with each create), there is nothing to compare: ++ * every listed leg is configured, as before, on the control plane's word. ++ * refused names the leg a refusal is about, num_legs for none. */ ++int ++raid_bdev_arbitrate(const struct raid_bdev_view *view, const struct raid_bdev_stamp *stamps, ++ uint8_t num_legs, bool *configure, uint64_t *generation, bool *migrates, ++ uint8_t *refused) ++{ ++ bool stamped = false, configured = false; ++ uint64_t top = 0; ++ uint8_t i; ++ int rc; ++ ++ *refused = num_legs; ++ ++ for (i = 0; i < num_legs; i++) { ++ if (stamps[i].status != 0) { ++ *refused = i; ++ return stamps[i].status; ++ } ++ } ++ ++ /* another volume's leg first: no epoch makes it this volume's */ ++ for (i = 0; i < num_legs; i++) { ++ if (stamps[i].stamped && spdk_uuid_compare(&stamps[i].lineage, &view->lineage) != 0) { ++ *refused = i; ++ return -EEXIST; ++ } ++ } ++ ++ for (i = 0; i < num_legs; i++) { ++ rc = raid_bdev_stamp_admits(view, &stamps[i]); ++ if (rc != 0) { ++ *refused = i; ++ return rc; ++ } ++ } ++ ++ for (i = 0; i < num_legs; i++) { ++ if (!stamps[i].stamped) { ++ continue; ++ } ++ stamped = true; ++ if (stamps[i].own_configured && (!configured || stamps[i].own_generation > top)) { ++ top = stamps[i].own_generation; ++ configured = true; ++ } ++ } ++ ++ if (!stamped) { ++ for (i = 0; i < num_legs; i++) { ++ configure[i] = true; ++ } ++ *generation = 0; ++ *migrates = true; ++ return 0; ++ } ++ ++ if (!configured) { ++ return -ENODATA; ++ } ++ ++ for (i = 0; i < num_legs; i++) { ++ configure[i] = stamps[i].stamped && stamps[i].own_configured && ++ stamps[i].own_generation == top; ++ } ++ *generation = top; ++ *migrates = false; ++ ++ return 0; ++} ++ ++int ++raid_bdev_lay_out(struct raid_bdev *raid_bdev, uint8_t num_members) ++{ ++ assert(raid_bdev->state == RAID_BDEV_STATE_CONFIGURING); ++ assert(raid_bdev->num_base_bdevs_discovered == 0); ++ ++ if (num_members == 0 || num_members > raid_bdev->num_base_bdevs || ++ num_members < raid_bdev->min_base_bdevs_operational) { ++ return -EINVAL; ++ } ++ ++ raid_bdev->num_base_bdevs_operational = num_members; ++ ++ return 0; ++} ++ ++/* Evariops 0053: the members take the first slots, in the order the caller adds ++ * them, at the arbitrated generation and the view's epoch. Each leg left out ++ * takes the next slot, empty, where the superblock records it MISSING at its ++ * own generation (0 without a stamp of the lineage), as it records an ejected ++ * member. The declared slots past them stay empty. */ ++int ++raid_bdev_lay_out_view(struct raid_bdev *raid_bdev, const struct raid_bdev_view *view, ++ const struct raid_bdev_stamp *stamps, const bool *configure, ++ uint8_t num_legs, uint64_t generation, bool migrates) ++{ ++ struct raid_base_bdev_info *base_info; ++ uint8_t i, members = 0, slot; ++ int rc; ++ ++ if (num_legs > raid_bdev->num_base_bdevs) { ++ return -EINVAL; ++ } ++ ++ for (i = 0; i < num_legs; i++) { ++ members += configure[i] ? 1 : 0; ++ } ++ ++ rc = raid_bdev_lay_out(raid_bdev, members); ++ if (rc != 0) { ++ return rc; ++ } ++ ++ spdk_uuid_copy(&raid_bdev->lineage, &view->lineage); ++ raid_bdev->view_epoch = view->epoch; ++ raid_bdev->arbitrating = true; ++ raid_bdev->arbitration_migrates = migrates; ++ raid_bdev->arbitrated_generation = generation; ++ raid_bdev->restamp_epoch = view->restamp_epoch; ++ ++ for (slot = 0; slot < members; slot++) { ++ base_info = &raid_bdev->base_bdev_info[slot]; ++ base_info->content_generation = generation; ++ base_info->view_epoch = view->epoch; ++ } ++ ++ for (i = 0; i < num_legs; i++) { ++ if (configure[i]) { ++ continue; ++ } ++ base_info = &raid_bdev->base_bdev_info[slot++]; ++ spdk_uuid_copy(&base_info->excluded_uuid, &stamps[i].uuid); ++ if (stamps[i].stamped) { ++ base_info->content_generation = stamps[i].own_generation; ++ base_info->view_epoch = stamps[i].own_view_epoch; ++ } ++ } ++ ++ return 0; ++} ++ + static void raid_bdev_examine_sb(const struct raid_bdev_superblock *sb, struct spdk_bdev *bdev, + raid_base_bdev_cb cb_fn, void *cb_ctx); + +@@ -4706,6 +5157,25 @@ raid_bdev_configure_base_bdev_check_sb_cb(const struct raid_bdev_superblock *sb, + struct raid_base_bdev_info *base_info = ctx; + raid_base_bdev_cb configure_cb = base_info->configure_cb; + ++ /* Evariops 0053: a raid of a lineage checks the leg by its stamp, whatever ++ * raid the superblock names. */ ++ if ((status == 0 || status == -EINVAL) && raid_bdev_checks_stamps(base_info->raid_bdev)) { ++ status = raid_bdev_check_stamp(base_info->raid_bdev, status == 0 ? sb : NULL, ++ &base_info->uuid); ++ if (status == 0) { ++ raid_bdev_configure_base_bdev_cont(base_info); ++ return; ++ } ++ SPDK_ERRLOG("raid bdev %s refuses bdev %s, %s\n", base_info->raid_bdev->bdev.name, ++ base_info->name, raid_bdev_stamp_refusal(status)); ++ base_info->configure_cb = NULL; ++ raid_bdev_free_base_bdev_resource(base_info); ++ if (configure_cb != NULL) { ++ configure_cb(base_info->configure_cb_ctx, status); ++ } ++ return; ++ } ++ + switch (status) { + case 0: + /* valid superblock found */ +@@ -5026,6 +5496,12 @@ raid_bdev_create_from_sb(const struct raid_bdev_superblock *sb, struct raid_bdev + assert(sb->length <= RAID_BDEV_SB_MAX_LENGTH); + memcpy(raid_bdev->sb, sb, sb->length); + ++ /* Evariops 0053: a raid reassembled from its superblock keeps its stamp. */ ++ if (raid_bdev_sb_stamped(sb)) { ++ spdk_uuid_copy(&raid_bdev->lineage, &sb->lineage); ++ raid_bdev->view_epoch = sb->view_epoch; ++ } ++ + for (i = 0; i < sb->base_bdevs_size; i++) { + const struct raid_bdev_sb_base_bdev *sb_base_bdev = &sb->base_bdevs[i]; + struct raid_base_bdev_info *base_info = &raid_bdev->base_bdev_info[sb_base_bdev->slot]; +diff --git a/module/bdev/raid/bdev_raid.h b/module/bdev/raid/bdev_raid.h +index 8f5d692..2594564 100644 +--- a/module/bdev/raid/bdev_raid.h ++++ b/module/bdev/raid/bdev_raid.h +@@ -151,6 +151,12 @@ struct raid_base_bdev_info { + uint64_t content_generation; + uint64_t view_epoch; + ++ /* Evariops 0053: the leg a view create left out of this empty slot, behind ++ * the generation of the members it configured. The superblock records it ++ * MISSING at its own generation, as it records an ejected member; the slot ++ * stays free for the next add. */ ++ struct spdk_uuid excluded_uuid; ++ + /* Evariops 0017: unix seconds of the last observable member-state change + * (configured/failed/write_only flips), for callers applying hysteresis. + * 0 = never transitioned. */ +@@ -301,6 +307,25 @@ struct raid_bdev { + * until bdev_raid_claim; the create RPC requires it. */ + char *incarnation; + ++ /* Evariops 0053: the view this raid belongs to: the volume's lineage, and ++ * the last view epoch the raid stamped on its members. A null lineage for a ++ * raid created outside the view protocol. A raid reassembled from its ++ * superblock keeps the stamp it reads there. */ ++ struct spdk_uuid lineage; ++ uint64_t view_epoch; ++ ++ /* Evariops 0053: set while the create that arbitrated the members adds ++ * them, until the raid configures. Each member is re-read as it is claimed ++ * and must still be what the arbitration read: a stamp of the lineage, ++ * configured in its own superblock at arbitrated_generation, or, when no leg ++ * carried one (arbitration_migrates), no stamp of the lineage at all. ++ * restamp_epoch is the epoch of another owner's stamps the create may ++ * overwrite (force_restamp), 0 for none. */ ++ bool arbitrating; ++ bool arbitration_migrates; ++ uint64_t arbitrated_generation; ++ uint64_t restamp_epoch; ++ + /* Superblock */ + bool superblock_enabled; + struct raid_bdev_superblock *sb; +@@ -360,6 +385,61 @@ int raid_bdev_add_base_bdev(struct raid_bdev *raid_bdev, const char *name, + bool skip_rebuild, bool write_only, + raid_base_bdev_cb cb_fn, void *cb_ctx); + ++/* Evariops 0053: a leg's superblock as a view create reads it. A leg with no ++ * superblock, or with one below minor 2 or without a lineage, is unstamped: it ++ * carries no generation of any lineage. */ ++struct raid_bdev_stamp { ++ /* 0 once the leg is read, whether it carries a superblock or not; else the ++ * error its open or its read met */ ++ int status; ++ /* the leg's bdev */ ++ struct spdk_uuid uuid; ++ bool stamped; ++ struct spdk_uuid lineage; ++ char owner[RAID_INCARNATION_MAX]; ++ uint64_t epoch; ++ /* the leg's own slot in its own superblock */ ++ bool own_configured; ++ uint64_t own_generation; ++ uint64_t own_view_epoch; ++}; ++ ++/* Evariops 0053: the view a create assembles its raid in: the volume's lineage, ++ * the view epoch the control plane read at its gate, the creating incarnation, ++ * and the epoch of another owner's stamps the create may overwrite, 0 for none. */ ++struct raid_bdev_view { ++ struct spdk_uuid lineage; ++ uint64_t epoch; ++ const char *incarnation; ++ uint64_t restamp_epoch; ++}; ++ ++typedef void (*raid_bdev_survey_cb)(void *ctx); ++ ++/* Evariops 0053: reads each leg's superblock, opened read-only, into stamps[i], ++ * and calls cb once every read has completed; a leg that does not open carries ++ * the open's error as its status. */ ++int raid_bdev_survey(char *const *names, uint8_t num_legs, struct raid_bdev_stamp *stamps, ++ raid_bdev_survey_cb cb, void *cb_ctx); ++/* Evariops 0053: which legs a view create configures. 0 with configure[i], the ++ * generation they carry, and whether no leg carried a stamp of the lineage. ++ * Otherwise the refusal, with the leg it is about in *refused (num_legs for ++ * none): -EEXIST for a leg of another lineage, -EBUSY for a leg another owner ++ * stamped in the view's epoch or a newer one, -ESTALE for such a stamp at an ++ * epoch other than the one the view restamps, -ENODATA when no stamped leg ++ * holds a configured copy, or a leg's read error. */ ++int raid_bdev_arbitrate(const struct raid_bdev_view *view, const struct raid_bdev_stamp *stamps, ++ uint8_t num_legs, bool *configure, uint64_t *generation, bool *migrates, ++ uint8_t *refused); ++/* Evariops 0053: a raid just created comes online with num_members members, ++ * every slot past them left empty. -EINVAL below what its level operates on. */ ++int raid_bdev_lay_out(struct raid_bdev *raid_bdev, uint8_t num_members); ++/* Evariops 0053: lays a raid just created out for the verdict of ++ * raid_bdev_arbitrate. The caller then adds the configured legs, in order. */ ++int raid_bdev_lay_out_view(struct raid_bdev *raid_bdev, const struct raid_bdev_view *view, ++ const struct raid_bdev_stamp *stamps, const bool *configure, ++ uint8_t num_legs, uint64_t generation, bool migrates); ++ + /* Evariops 0013: one dirty range, in blocks, seeding a partial rebuild. The + * array must be sorted by offset and non-overlapping — the RPC validates it and + * the rebuild engine's skip cursor relies on it. */ +@@ -596,8 +676,11 @@ raid_bdev_flush_blocks(struct raid_base_bdev_info *base_info, struct spdk_io_cha + #define RAID_BDEV_SB_VERSION_MAJOR 1 + /* Evariops 0016: minor 1 adds the per-member content_generation and view_epoch, + * carved out of reserved bytes, so a minor-0 superblock reads back as +- * generation 0 / epoch 0. */ +-#define RAID_BDEV_SB_VERSION_MINOR 1 ++ * generation 0 / epoch 0. ++ * Evariops 0053: minor 2 adds the raid's lineage, owner and view epoch, carved ++ * out of the header's reserved bytes, so a minor-1 superblock reads back as an ++ * unstamped one. */ ++#define RAID_BDEV_SB_VERSION_MINOR 2 + + #define RAID_BDEV_SB_NAME_SIZE 64 + +@@ -671,7 +754,16 @@ struct raid_bdev_superblock { + /* number of raid base devices */ + uint8_t num_base_bdevs; + +- uint8_t reserved[118]; ++ uint8_t reserved0[7]; ++ ++ /* Evariops 0053: the stamp. The volume the raid serves, null when it was ++ * created outside the view protocol; the control-plane incarnation that ++ * created it, NUL-terminated; and the last view epoch it stamped. */ ++ struct spdk_uuid lineage; ++ uint8_t owner[RAID_INCARNATION_MAX]; ++ uint64_t view_epoch; ++ ++ uint8_t reserved[23]; + + /* size of the base bdevs array */ + uint8_t base_bdevs_size; +@@ -679,6 +771,11 @@ struct raid_bdev_superblock { + struct raid_bdev_sb_base_bdev base_bdevs[]; + }; + SPDK_STATIC_ASSERT(sizeof(struct raid_bdev_superblock) == 256, "incorrect size"); ++/* Evariops 0053: the stamp sits where minor 1 kept reserved bytes. */ ++SPDK_STATIC_ASSERT(offsetof(struct raid_bdev_superblock, lineage) == 144, "incorrect offset"); ++SPDK_STATIC_ASSERT(offsetof(struct raid_bdev_superblock, owner) == 160, "incorrect offset"); ++SPDK_STATIC_ASSERT(offsetof(struct raid_bdev_superblock, view_epoch) == 224, "incorrect offset"); ++SPDK_STATIC_ASSERT(offsetof(struct raid_bdev_superblock, base_bdevs_size) == 255, "incorrect offset"); + + #define RAID_BDEV_SB_MAX_LENGTH (sizeof(struct raid_bdev_superblock) + UINT8_MAX * sizeof(struct raid_bdev_sb_base_bdev)) + +diff --git a/module/bdev/raid/bdev_raid_rpc.c b/module/bdev/raid/bdev_raid_rpc.c +index d4127d2..9119230 100644 +--- a/module/bdev/raid/bdev_raid_rpc.c ++++ b/module/bdev/raid/bdev_raid_rpc.c +@@ -145,6 +145,20 @@ struct rpc_bdev_raid_create { + /* Evariops 0014: identity of the creating control-plane incarnation. + * Mandatory — there is no creation without an identity. */ + char *incarnation; ++ ++ /* Evariops 0053: the view the raid is created in. A view_epoch arbitrates ++ * the members from their superblocks: the volume's lineage, the view epoch ++ * the control plane read at its gate, and, to overwrite another owner's ++ * stamps of that epoch or a newer one, force_restamp with the epoch read ++ * from them. */ ++ struct spdk_uuid lineage; ++ uint64_t view_epoch; ++ bool force_restamp; ++ uint64_t expected_view_epoch; ++ ++ /* Evariops 0053: the raid's width when wider than the listed legs; the ++ * slots past them stay empty, each for a member added later. */ ++ uint32_t declared_slots; + }; + + /* +@@ -193,6 +207,11 @@ static const struct spdk_json_object_decoder rpc_bdev_raid_create_decoders[] = { + {"uuid", offsetof(struct rpc_bdev_raid_create, uuid), spdk_json_decode_uuid, true}, + {"superblock", offsetof(struct rpc_bdev_raid_create, superblock_enabled), spdk_json_decode_bool, true}, + {"incarnation", offsetof(struct rpc_bdev_raid_create, incarnation), spdk_json_decode_string}, ++ {"lineage", offsetof(struct rpc_bdev_raid_create, lineage), spdk_json_decode_uuid, true}, ++ {"view_epoch", offsetof(struct rpc_bdev_raid_create, view_epoch), spdk_json_decode_uint64, true}, ++ {"force_restamp", offsetof(struct rpc_bdev_raid_create, force_restamp), spdk_json_decode_bool, true}, ++ {"expected_view_epoch", offsetof(struct rpc_bdev_raid_create, expected_view_epoch), spdk_json_decode_uint64, true}, ++ {"declared_slots", offsetof(struct rpc_bdev_raid_create, declared_slots), spdk_json_decode_uint32, true}, + }; + + struct rpc_bdev_raid_create_ctx { +@@ -201,6 +220,10 @@ struct rpc_bdev_raid_create_ctx { + struct spdk_jsonrpc_request *request; + uint8_t remaining; + int status; ++ /* Evariops 0053: a view create's read of each listed leg, and the legs its ++ * arbitration configures */ ++ struct raid_bdev_stamp *stamps; ++ bool *configure; + }; + + static void +@@ -221,6 +244,8 @@ free_rpc_bdev_raid_create_ctx(struct rpc_bdev_raid_create_ctx *ctx) + free(req->base_bdevs.base_bdevs[i]); + } + ++ free(ctx->stamps); ++ free(ctx->configure); + free(ctx); + } + +@@ -261,16 +286,186 @@ rpc_bdev_raid_create_add_base_bdev_cb(void *_ctx, int status) + * returns: + * none + */ ++/* Evariops 0053: creates the raid and adds its members: every listed leg, or, ++ * for a create in a view, the legs its arbitration configured. A create in a ++ * view fails on a leg gone since it read it, where a plain create waits for it. */ ++static void ++rpc_bdev_raid_create_members(struct rpc_bdev_raid_create_ctx *ctx, const struct raid_bdev_view *view, ++ uint64_t generation, bool migrates) ++{ ++ struct rpc_bdev_raid_create *req = &ctx->req; ++ uint8_t num_legs = req->base_bdevs.num_base_bdevs; ++ uint8_t num_slots = spdk_max(req->declared_slots, num_legs); ++ uint8_t num_members = 0, added = 0; ++ struct raid_bdev *raid_bdev; ++ int rc = 0; ++ size_t i; ++ ++ rc = raid_bdev_create(req->name, req->strip_size_kb, num_slots, ++ req->level, req->superblock_enabled, &req->uuid, ++ req->incarnation, &raid_bdev); ++ if (rc != 0) { ++ spdk_jsonrpc_send_error_response_fmt(ctx->request, rc, ++ "Failed to create RAID bdev %s: %s", ++ req->name, spdk_strerror(-rc)); ++ free_rpc_bdev_raid_create_ctx(ctx); ++ return; ++ } ++ ++ if (view != NULL) { ++ rc = raid_bdev_lay_out_view(raid_bdev, view, ctx->stamps, ctx->configure, num_legs, ++ generation, migrates); ++ } else if (num_slots > num_legs) { ++ rc = raid_bdev_lay_out(raid_bdev, num_legs); ++ } ++ if (rc != 0) { ++ raid_bdev_delete(raid_bdev, NULL, NULL); ++ spdk_jsonrpc_send_error_response_fmt(ctx->request, rc, ++ "Failed to create RAID bdev %s: %u members of %u slots: %s", ++ req->name, num_legs, num_slots, spdk_strerror(-rc)); ++ free_rpc_bdev_raid_create_ctx(ctx); ++ return; ++ } ++ ++ for (i = 0; i < num_legs; i++) { ++ if (view == NULL || ctx->configure[i]) { ++ num_members++; ++ } else { ++ SPDK_NOTICELOG("raid bdev %s leaves bdev %s out, behind the generation %" PRIu64 ++ " of its members (Evariops 0053)\n", req->name, ++ req->base_bdevs.base_bdevs[i], generation); ++ } ++ } ++ ++ ctx->raid_bdev = raid_bdev; ++ ctx->remaining = num_members; ++ ++ assert(num_members > 0); ++ ++ /* The add of the last member can complete the request, and free ctx, before ++ * it returns: the loop stops there. */ ++ for (i = 0; i < num_legs && added < num_members; i++) { ++ const char *base_bdev_name = req->base_bdevs.base_bdevs[i]; ++ ++ if (view != NULL && !ctx->configure[i]) { ++ continue; ++ } ++ ++ rc = raid_bdev_add_base_bdev(raid_bdev, base_bdev_name, false, false, ++ rpc_bdev_raid_create_add_base_bdev_cb, ctx); ++ if (rc == -ENODEV && view == NULL) { ++ SPDK_DEBUGLOG(bdev_raid, "base bdev %s doesn't exist now\n", base_bdev_name); ++ assert(ctx->remaining > 1 || added + 1 == num_members); ++ rpc_bdev_raid_create_add_base_bdev_cb(ctx, 0); ++ } else if (rc != 0) { ++ SPDK_DEBUGLOG(bdev_raid, "Failed to add base bdev %s to RAID bdev %s: %s", ++ base_bdev_name, req->name, spdk_strerror(-rc)); ++ ctx->remaining -= (num_members - added - 1); ++ rpc_bdev_raid_create_add_base_bdev_cb(ctx, rc); ++ break; ++ } ++ added++; ++ } ++} ++ ++/* Evariops 0053: why a create in a view refuses. */ ++static const char * ++rpc_bdev_raid_create_refusal(int rc) ++{ ++ switch (rc) { ++ case -EEXIST: ++ return "stamped for another volume"; ++ case -EBUSY: ++ return "stamped by another owner in the view's epoch or a newer one"; ++ case -ESTALE: ++ return "stamped by another owner at an epoch other than the one to restamp"; ++ case -ENODATA: ++ return "no leg stamped for the volume holds a configured copy"; ++ default: ++ return spdk_strerror(-rc); ++ } ++} ++ ++static void ++rpc_bdev_raid_create_surveyed(void *_ctx) ++{ ++ struct rpc_bdev_raid_create_ctx *ctx = _ctx; ++ struct rpc_bdev_raid_create *req = &ctx->req; ++ struct raid_bdev_view view = { ++ .epoch = req->view_epoch, ++ .incarnation = req->incarnation, ++ .restamp_epoch = req->force_restamp ? req->expected_view_epoch : 0, ++ }; ++ uint8_t num_legs = req->base_bdevs.num_base_bdevs; ++ uint64_t generation; ++ uint8_t refused; ++ bool migrates; ++ int rc; ++ ++ spdk_uuid_copy(&view.lineage, &req->lineage); ++ ++ rc = raid_bdev_arbitrate(&view, ctx->stamps, num_legs, ctx->configure, &generation, &migrates, ++ &refused); ++ if (rc == 0) { ++ rpc_bdev_raid_create_members(ctx, &view, generation, migrates); ++ return; ++ } ++ ++ if (refused < num_legs && ctx->stamps[refused].status == 0) { ++ const struct raid_bdev_stamp *stamp = &ctx->stamps[refused]; ++ ++ spdk_jsonrpc_send_error_response_fmt(ctx->request, rc, ++ "Failed to create RAID bdev %s: base bdev %s is %s " ++ "(owner '%s', view epoch %" PRIu64 ")", ++ req->name, req->base_bdevs.base_bdevs[refused], ++ rpc_bdev_raid_create_refusal(rc), stamp->owner, stamp->epoch); ++ } else if (refused < num_legs) { ++ spdk_jsonrpc_send_error_response_fmt(ctx->request, rc, ++ "Failed to create RAID bdev %s: base bdev %s: %s", ++ req->name, req->base_bdevs.base_bdevs[refused], ++ spdk_strerror(-rc)); ++ } else { ++ spdk_jsonrpc_send_error_response_fmt(ctx->request, rc, ++ "Failed to create RAID bdev %s: %s", ++ req->name, rpc_bdev_raid_create_refusal(rc)); ++ } ++ free_rpc_bdev_raid_create_ctx(ctx); ++} ++ ++/* Evariops 0053: what a create in a view needs: a raid1 with a superblock, the ++ * volume's lineage, and force_restamp named with the epoch it overwrites. */ ++static int ++rpc_bdev_raid_create_check_view(struct rpc_bdev_raid_create *req, const char **why) ++{ ++ if (req->view_epoch == 0) { ++ *why = "lineage, force_restamp and expected_view_epoch need a view_epoch"; ++ return spdk_uuid_is_null(&req->lineage) && !req->force_restamp && ++ req->expected_view_epoch == 0 ? 0 : -EINVAL; ++ } ++ ++ if (req->level != RAID1 || !req->superblock_enabled || spdk_uuid_is_null(&req->lineage)) { ++ *why = "a create in a view needs a raid1 with a superblock and a lineage"; ++ return -EINVAL; ++ } ++ ++ if (req->force_restamp != (req->expected_view_epoch != 0)) { ++ *why = "force_restamp needs the expected_view_epoch it overwrites, and only it"; ++ return -EINVAL; ++ } ++ ++ return 0; ++} ++ + static void + rpc_bdev_raid_create(struct spdk_jsonrpc_request *request, + const struct spdk_json_val *params) + { + struct rpc_bdev_raid_create *req; +- struct raid_bdev *raid_bdev; + int rc; + size_t i; + struct rpc_bdev_raid_create_ctx *ctx; + uint8_t num_base_bdevs; ++ const char *why = NULL; + + ctx = calloc(1, sizeof(*ctx)); + if (ctx == NULL) { +@@ -278,6 +473,7 @@ rpc_bdev_raid_create(struct spdk_jsonrpc_request *request, + goto cleanup; + } + req = &ctx->req; ++ ctx->request = request; + + if (spdk_json_decode_object(params, rpc_bdev_raid_create_decoders, + SPDK_COUNTOF(rpc_bdev_raid_create_decoders), +@@ -297,38 +493,38 @@ rpc_bdev_raid_create(struct spdk_jsonrpc_request *request, + } + } + +- rc = raid_bdev_create(req->name, req->strip_size_kb, num_base_bdevs, +- req->level, req->superblock_enabled, &req->uuid, +- req->incarnation, &raid_bdev); +- if (rc != 0) { +- spdk_jsonrpc_send_error_response_fmt(request, rc, +- "Failed to create RAID bdev %s: %s", +- req->name, spdk_strerror(-rc)); ++ if (req->declared_slots > RPC_MAX_BASE_BDEVS) { ++ spdk_jsonrpc_send_error_response_fmt(request, -EINVAL, ++ "declared_slots %u exceeds %d", ++ req->declared_slots, RPC_MAX_BASE_BDEVS); + goto cleanup; + } + +- ctx->raid_bdev = raid_bdev; +- ctx->request = request; +- ctx->remaining = num_base_bdevs; ++ rc = rpc_bdev_raid_create_check_view(req, &why); ++ if (rc != 0) { ++ spdk_jsonrpc_send_error_response_fmt(request, rc, "Failed to create RAID bdev %s: %s", ++ req->name, why); ++ goto cleanup; ++ } + +- assert(num_base_bdevs > 0); ++ if (req->view_epoch == 0) { ++ rpc_bdev_raid_create_members(ctx, NULL, 0, false); ++ return; ++ } + +- for (i = 0; i < num_base_bdevs; i++) { +- const char *base_bdev_name = req->base_bdevs.base_bdevs[i]; ++ ctx->stamps = calloc(num_base_bdevs, sizeof(*ctx->stamps)); ++ ctx->configure = calloc(num_base_bdevs, sizeof(*ctx->configure)); ++ if (ctx->stamps == NULL || ctx->configure == NULL) { ++ spdk_jsonrpc_send_error_response(request, -ENOMEM, spdk_strerror(ENOMEM)); ++ goto cleanup; ++ } + +- rc = raid_bdev_add_base_bdev(raid_bdev, base_bdev_name, false, false, +- rpc_bdev_raid_create_add_base_bdev_cb, ctx); +- if (rc == -ENODEV) { +- SPDK_DEBUGLOG(bdev_raid, "base bdev %s doesn't exist now\n", base_bdev_name); +- assert(ctx->remaining > 1 || i + 1 == num_base_bdevs); +- rpc_bdev_raid_create_add_base_bdev_cb(ctx, 0); +- } else if (rc != 0) { +- SPDK_DEBUGLOG(bdev_raid, "Failed to add base bdev %s to RAID bdev %s: %s", +- base_bdev_name, req->name, spdk_strerror(-rc)); +- ctx->remaining -= (num_base_bdevs - i - 1); +- rpc_bdev_raid_create_add_base_bdev_cb(ctx, rc); +- break; +- } ++ rc = raid_bdev_survey(req->base_bdevs.base_bdevs, num_base_bdevs, ctx->stamps, ++ rpc_bdev_raid_create_surveyed, ctx); ++ if (rc != 0) { ++ spdk_jsonrpc_send_error_response_fmt(request, rc, "Failed to create RAID bdev %s: %s", ++ req->name, spdk_strerror(-rc)); ++ goto cleanup; + } + return; + cleanup: +diff --git a/module/bdev/raid/bdev_raid_sb.c b/module/bdev/raid/bdev_raid_sb.c +index bb0fcfc..eb092d8 100644 +--- a/module/bdev/raid/bdev_raid_sb.c ++++ b/module/bdev/raid/bdev_raid_sb.c +@@ -81,13 +81,37 @@ raid_bdev_init_superblock(struct raid_bdev *raid_bdev) + sb->num_base_bdevs = sb->base_bdevs_size = raid_bdev->num_base_bdevs; + sb->length = sizeof(*sb) + sizeof(*sb_base_bdev) * sb->base_bdevs_size; + ++ /* Evariops 0053: a raid created in a view stamps it. */ ++ if (!spdk_uuid_is_null(&raid_bdev->lineage)) { ++ spdk_uuid_copy(&sb->lineage, &raid_bdev->lineage); ++ snprintf((char *)sb->owner, sizeof(sb->owner), "%s", ++ raid_bdev->incarnation != NULL ? raid_bdev->incarnation : ""); ++ sb->view_epoch = raid_bdev->view_epoch; ++ } ++ ++ /* Evariops 0053: a slot its create left empty is recorded MISSING, under ++ * the leg it left out or, for a slot no leg ever held, a UUID of its own: ++ * every slot needs an entry, which a member added there later takes over, ++ * and examine expects each entry to name one. Upstream creates a raid with ++ * every slot filled. The generations are the ones the create set. */ + sb_base_bdev = &sb->base_bdevs[0]; + RAID_FOR_EACH_BASE_BDEV(raid_bdev, base_info) { +- spdk_uuid_copy(&sb_base_bdev->uuid, &base_info->uuid); ++ if (base_info->is_configured) { ++ spdk_uuid_copy(&sb_base_bdev->uuid, &base_info->uuid); ++ sb_base_bdev->state = RAID_SB_BASE_BDEV_CONFIGURED; ++ } else { ++ if (!spdk_uuid_is_null(&base_info->excluded_uuid)) { ++ spdk_uuid_copy(&sb_base_bdev->uuid, &base_info->excluded_uuid); ++ } else { ++ spdk_uuid_generate(&sb_base_bdev->uuid); ++ } ++ sb_base_bdev->state = RAID_SB_BASE_BDEV_MISSING; ++ } + sb_base_bdev->data_offset = base_info->data_offset; + sb_base_bdev->data_size = base_info->data_size; +- sb_base_bdev->state = RAID_SB_BASE_BDEV_CONFIGURED; + sb_base_bdev->slot = raid_bdev_base_bdev_slot(base_info); ++ sb_base_bdev->content_generation = base_info->content_generation; ++ sb_base_bdev->view_epoch = base_info->view_epoch; + sb_base_bdev++; + } + } +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index c359160..0f21d2c 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -51,6 +51,8 @@ uint32_t g_max_io_size; + uint8_t g_max_base_drives; + uint8_t g_max_raids; + uint8_t g_rpc_err; ++/* Evariops 0053: the code of the last error response */ ++int g_rpc_err_code; + char *g_get_raids_output[MAX_RAIDS]; + uint32_t g_get_raids_count; + uint8_t g_json_decode_obj_err; +@@ -142,6 +144,19 @@ static struct raid_bdev_module g_ut_raid_module = { + }; + RAID_MODULE_REGISTER(&g_ut_raid_module) + ++/* Evariops 0053: a create in a view is raid1 only. This module carries raid1's ++ * constraint and the suite's I/O. */ ++static struct raid_bdev_module g_ut_raid1_module = { ++ .level = RAID1, ++ .base_bdevs_min = 2, ++ .base_bdevs_constraint = {CONSTRAINT_MIN_BASE_BDEVS_OPERATIONAL, 1}, ++ .start = ut_raid_start, ++ .submit_rw_request = ut_raid_submit_rw_request, ++ .submit_null_payload_request = ut_raid_submit_null_payload_request, ++ .submit_process_request = ut_raid_submit_process_request, ++}; ++RAID_MODULE_REGISTER(&g_ut_raid1_module) ++ + DEFINE_STUB_V(spdk_bdev_module_examine_done, (struct spdk_bdev_module *module)); + DEFINE_STUB_V(spdk_bdev_module_list_add, (struct spdk_bdev_module *bdev_module)); + DEFINE_STUB(spdk_bdev_io_type_supported, bool, (struct spdk_bdev *bdev, +@@ -203,8 +218,29 @@ DEFINE_STUB(spdk_bdev_notify_blockcnt_change, int, (struct spdk_bdev *bdev, uint + DEFINE_STUB(spdk_json_write_named_uuid, int, (struct spdk_json_write_ctx *w, const char *name, + const struct spdk_uuid *val), 0); + DEFINE_STUB_V(raid_bdev_init_superblock, (struct raid_bdev *raid_bdev)); +-DEFINE_STUB(raid_bdev_alloc_superblock, int, (struct raid_bdev *raid_bdev, uint32_t block_size), 0); +-DEFINE_STUB_V(raid_bdev_free_superblock, (struct raid_bdev *raid_bdev)); ++ ++/* Evariops 0053: the suite's raids have no superblock buffer, except while a test ++ * reassembles one from a superblock. */ ++static bool g_ut_alloc_superblock; ++ ++int ++raid_bdev_alloc_superblock(struct raid_bdev *raid_bdev, uint32_t block_size) ++{ ++ if (g_ut_alloc_superblock) { ++ raid_bdev->sb = calloc(1, RAID_BDEV_SB_MAX_LENGTH); ++ if (raid_bdev->sb == NULL) { ++ return -ENOMEM; ++ } ++ } ++ return 0; ++} ++ ++void ++raid_bdev_free_superblock(struct raid_bdev *raid_bdev) ++{ ++ free(raid_bdev->sb); ++ raid_bdev->sb = NULL; ++} + DEFINE_STUB(spdk_bdev_readv_blocks_ext, int, (struct spdk_bdev_desc *desc, + struct spdk_io_channel *ch, struct iovec *iov, int iovcnt, uint64_t offset_blocks, + uint64_t num_blocks, spdk_bdev_io_completion_cb cb, void *cb_arg, +@@ -312,11 +348,104 @@ spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) + return g_block_len; + } + ++/* Evariops 0053: what reading each base bdev's superblock returns, by bdev name. ++ * A bdev without an entry, or whose entry has no superblock, carries none ++ * (-EINVAL). Otherwise the entry returns its read error (on its first read only ++ * when the error is transient), or its superblock, and from its second read on ++ * its later superblock when it has one: a leg that changed between a create's ++ * survey and its claim. */ ++#define UT_LEGS_MAX 4 ++#define UT_SBS_MAX 8 ++ ++struct ut_leg { ++ const char *name; ++ struct raid_bdev_superblock *sb; ++ struct raid_bdev_superblock *later_sb; ++ int read_status; ++ bool read_error_transient; ++ uint32_t reads; ++}; ++ ++static struct ut_leg g_ut_legs[UT_LEGS_MAX]; ++static struct raid_bdev_superblock *g_ut_sbs[UT_SBS_MAX]; ++ ++static struct ut_leg * ++ut_leg(const char *name) ++{ ++ size_t i; ++ ++ for (i = 0; i < SPDK_COUNTOF(g_ut_legs); i++) { ++ if (g_ut_legs[i].name != NULL && strcmp(g_ut_legs[i].name, name) == 0) { ++ return &g_ut_legs[i]; ++ } ++ } ++ for (i = 0; i < SPDK_COUNTOF(g_ut_legs); i++) { ++ if (g_ut_legs[i].name == NULL) { ++ g_ut_legs[i].name = name; ++ return &g_ut_legs[i]; ++ } ++ } ++ SPDK_CU_ASSERT_FATAL(false); ++ return NULL; ++} ++ ++static uint32_t ++ut_leg_reads(void) ++{ ++ uint32_t reads = 0; ++ size_t i; ++ ++ for (i = 0; i < SPDK_COUNTOF(g_ut_legs); i++) { ++ reads += g_ut_legs[i].reads; ++ } ++ return reads; ++} ++ ++static void ++ut_legs_reset(void) ++{ ++ size_t i; ++ ++ memset(g_ut_legs, 0, sizeof(g_ut_legs)); ++ for (i = 0; i < SPDK_COUNTOF(g_ut_sbs); i++) { ++ free(g_ut_sbs[i]); ++ g_ut_sbs[i] = NULL; ++ } ++} ++ + int + raid_bdev_load_base_bdev_superblock(struct spdk_bdev_desc *desc, struct spdk_io_channel *ch, + raid_bdev_load_sb_cb cb, void *cb_ctx) + { +- cb(NULL, -EINVAL, cb_ctx); ++ const char *name = spdk_bdev_desc_get_bdev(desc)->name; ++ struct raid_bdev_superblock *sb; ++ struct ut_leg *leg = NULL; ++ size_t i; ++ ++ for (i = 0; i < SPDK_COUNTOF(g_ut_legs); i++) { ++ if (g_ut_legs[i].name != NULL && strcmp(g_ut_legs[i].name, name) == 0) { ++ leg = &g_ut_legs[i]; ++ break; ++ } ++ } ++ ++ if (leg == NULL) { ++ cb(NULL, -EINVAL, cb_ctx); ++ return 0; ++ } ++ ++ leg->reads++; ++ if (leg->read_status != 0 && (!leg->read_error_transient || leg->reads == 1)) { ++ cb(NULL, leg->read_status, cb_ctx); ++ return 0; ++ } ++ ++ sb = leg->reads > 1 && leg->later_sb != NULL ? leg->later_sb : leg->sb; ++ if (sb == NULL) { ++ cb(NULL, -EINVAL, cb_ctx); ++ } else { ++ cb(sb, 0, cb_ctx); ++ } + + return 0; + } +@@ -643,6 +772,11 @@ spdk_json_decode_object(const struct spdk_json_val *values, + _out->incarnation = strdup(req->incarnation); + SPDK_CU_ASSERT_FATAL(_out->incarnation != NULL); + } ++ spdk_uuid_copy(&_out->lineage, &req->lineage); ++ _out->view_epoch = req->view_epoch; ++ _out->force_restamp = req->force_restamp; ++ _out->expected_view_epoch = req->expected_view_epoch; ++ _out->declared_slots = req->declared_slots; + _out->base_bdevs.num_base_bdevs = req->base_bdevs.num_base_bdevs; + for (i = 0; i < req->base_bdevs.num_base_bdevs; i++) { + _out->base_bdevs.base_bdevs[i] = strdup(req->base_bdevs.base_bdevs[i]); +@@ -666,6 +800,7 @@ spdk_jsonrpc_send_error_response(struct spdk_jsonrpc_request *request, + int error_code, const char *msg) + { + g_rpc_err = 1; ++ g_rpc_err_code = error_code; + } + + void +@@ -673,6 +808,7 @@ spdk_jsonrpc_send_error_response_fmt(struct spdk_jsonrpc_request *request, + int error_code, const char *fmt, ...) + { + g_rpc_err = 1; ++ g_rpc_err_code = error_code; + } + + struct spdk_bdev * +@@ -957,6 +1093,7 @@ create_test_req(struct rpc_bdev_raid_create *r, const char *raid_name, + char name[16]; + uint8_t bbdev_idx = bbdev_start_idx; + ++ memset(r, 0, sizeof(*r)); + r->name = strdup(raid_name); + SPDK_CU_ASSERT_FATAL(r->name != NULL); + r->strip_size_kb = (g_strip_size * g_block_len) / 1024; +@@ -2327,6 +2464,723 @@ test_raid1_ejection_of_a_member_behind_opens_no_epoch(void) + member, sizeof(member)) == 0); + } + ++/* Evariops 0053: the volume the view tests create a raid for, and another one. */ ++static struct spdk_uuid g_ut_lineage; ++static struct spdk_uuid g_ut_other_lineage; ++static uint8_t g_ut_saved_max_base_drives; ++ ++static void ++ut_view_setup(void) ++{ ++ g_ut_saved_max_base_drives = g_max_base_drives; ++ g_max_base_drives = 3; ++ set_globals(); ++ CU_ASSERT(raid_bdev_init() == 0); ++ ut_legs_reset(); ++ spdk_uuid_generate(&g_ut_lineage); ++ spdk_uuid_generate(&g_ut_other_lineage); ++ g_rpc_err_code = 0; ++} ++ ++static void ++ut_view_teardown(void) ++{ ++ struct rpc_bdev_raid_delete delete_req; ++ ++ if (raid_bdev_find_by_name("raid1") != NULL) { ++ create_raid_bdev_delete_req(&delete_req, "raid1", 0); ++ rpc_bdev_raid_delete(NULL, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ } ++ poll_app_thread(); ++ raid_bdev_exit(); ++ base_bdevs_cleanup(); ++ reset_globals(); ++ ut_legs_reset(); ++ g_ut_alloc_superblock = false; ++ g_max_base_drives = g_ut_saved_max_base_drives; ++} ++ ++/* Evariops 0053: a create over Nvme0n1, Nvme1n1 and Nvme2n1, which exist once it ++ * returns, in the view of the test volume at epoch, under the suite's ++ * incarnation. */ ++static void ++ut_view_create_req(struct rpc_bdev_raid_create *r, uint64_t epoch) ++{ ++ create_raid_bdev_create_req(r, "raid1", 0, true, 0, true); ++ r->level = RAID1; ++ r->strip_size_kb = 0; ++ spdk_uuid_copy(&r->lineage, &g_ut_lineage); ++ r->view_epoch = epoch; ++} ++ ++/* Evariops 0053: a superblock as a raid of lineage writes it, owned by owner at ++ * epoch; with no lineage, an unstamped one. Freed with the legs. */ ++static struct raid_bdev_superblock * ++ut_sb(const struct spdk_uuid *lineage, const char *owner, uint64_t epoch) ++{ ++ struct raid_bdev_superblock *sb; ++ size_t i; ++ ++ for (i = 0; i < SPDK_COUNTOF(g_ut_sbs) && g_ut_sbs[i] != NULL; i++) { ++ } ++ SPDK_CU_ASSERT_FATAL(i < SPDK_COUNTOF(g_ut_sbs)); ++ sb = calloc(1, RAID_BDEV_SB_MAX_LENGTH); ++ SPDK_CU_ASSERT_FATAL(sb != NULL); ++ g_ut_sbs[i] = sb; ++ ++ memcpy(sb->signature, RAID_BDEV_SB_SIG, sizeof(sb->signature)); ++ sb->version.major = RAID_BDEV_SB_VERSION_MAJOR; ++ sb->version.minor = RAID_BDEV_SB_VERSION_MINOR; ++ sb->length = sizeof(*sb); ++ spdk_uuid_generate(&sb->uuid); ++ snprintf((char *)sb->name, sizeof(sb->name), "%s", "raid-before"); ++ sb->block_size = g_block_len; ++ sb->level = RAID1; ++ if (lineage != NULL) { ++ spdk_uuid_copy(&sb->lineage, lineage); ++ snprintf((char *)sb->owner, sizeof(sb->owner), "%s", owner); ++ sb->view_epoch = epoch; ++ } ++ ++ return sb; ++} ++ ++/* Evariops 0053: records bdev name in the next slot of sb, in state at generation. */ ++static void ++ut_sb_slot(struct raid_bdev_superblock *sb, const char *name, uint32_t state, uint64_t generation) ++{ ++ struct spdk_bdev *bdev = spdk_bdev_get_by_name(name); ++ struct raid_bdev_sb_base_bdev *slot = &sb->base_bdevs[sb->base_bdevs_size]; ++ ++ SPDK_CU_ASSERT_FATAL(bdev != NULL); ++ spdk_uuid_copy(&slot->uuid, &bdev->uuid); ++ slot->state = state; ++ slot->content_generation = generation; ++ slot->slot = sb->base_bdevs_size; ++ sb->base_bdevs_size++; ++ sb->num_base_bdevs = sb->base_bdevs_size; ++ sb->length = sizeof(*sb) + sb->base_bdevs_size * sizeof(*slot); ++} ++ ++/* Evariops 0053: the superblock every leg of a raid of the test volume carries, ++ * owned by owner at epoch, its three legs configured at generation. */ ++static struct raid_bdev_superblock * ++ut_sb_all_configured(const char *owner, uint64_t epoch, uint64_t generation) ++{ ++ struct raid_bdev_superblock *sb = ut_sb(&g_ut_lineage, owner, epoch); ++ ++ ut_sb_slot(sb, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, generation); ++ ut_sb_slot(sb, "Nvme1n1", RAID_SB_BASE_BDEV_CONFIGURED, generation); ++ ut_sb_slot(sb, "Nvme2n1", RAID_SB_BASE_BDEV_CONFIGURED, generation); ++ ut_leg("Nvme0n1")->sb = sb; ++ ut_leg("Nvme1n1")->sb = sb; ++ ut_leg("Nvme2n1")->sb = sb; ++ ++ return sb; ++} ++ ++/* Evariops 0053: Nvme0n1 was ejected at generation 6 from a raid of the test ++ * volume that incarnation-before created at epoch 3: the two members that ++ * stayed advanced to 7 in their own superblocks, and its own still has it ++ * configured at 6. */ ++static struct raid_bdev_superblock * ++ut_sb_after_an_ejection(void) ++{ ++ struct raid_bdev_superblock *survivors, *ejected; ++ ++ survivors = ut_sb(&g_ut_lineage, "incarnation-before", 3); ++ ut_sb_slot(survivors, "Nvme0n1", RAID_SB_BASE_BDEV_MISSING, 6); ++ ut_sb_slot(survivors, "Nvme1n1", RAID_SB_BASE_BDEV_CONFIGURED, 7); ++ ut_sb_slot(survivors, "Nvme2n1", RAID_SB_BASE_BDEV_CONFIGURED, 7); ++ ejected = ut_sb(&g_ut_lineage, "incarnation-before", 3); ++ ut_sb_slot(ejected, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 6); ++ ut_sb_slot(ejected, "Nvme1n1", RAID_SB_BASE_BDEV_CONFIGURED, 6); ++ ut_sb_slot(ejected, "Nvme2n1", RAID_SB_BASE_BDEV_CONFIGURED, 6); ++ ut_leg("Nvme0n1")->sb = ejected; ++ ut_leg("Nvme1n1")->sb = survivors; ++ ut_leg("Nvme2n1")->sb = survivors; ++ ++ return ejected; ++} ++ ++/* Evariops 0053: the raid the view tests created, online with num_members. */ ++static struct raid_bdev * ++ut_view_raid_online(uint8_t num_members) ++{ ++ struct raid_bdev *raid_bdev = raid_bdev_find_by_name("raid1"); ++ ++ SPDK_CU_ASSERT_FATAL(raid_bdev != NULL); ++ CU_ASSERT(raid_bdev->state == RAID_BDEV_STATE_ONLINE); ++ CU_ASSERT(raid_bdev->num_base_bdevs_operational == num_members); ++ CU_ASSERT(raid_bdev->num_base_bdevs_discovered == num_members); ++ CU_ASSERT(!raid_bdev->arbitrating); ++ /* sized by its members, not by the slots they left empty */ ++ CU_ASSERT(raid_bdev->bdev.blockcnt == raid_bdev->base_bdev_info[0].data_size); ++ CU_ASSERT(raid_bdev->bdev.blockcnt != 0); ++ ++ return raid_bdev; ++} ++ ++/* Evariops 0053: the member of slot, bdev name, at generation and epoch. */ ++static void ++ut_view_member(struct raid_bdev *raid_bdev, uint8_t slot, const char *name, uint64_t generation, ++ uint64_t epoch) ++{ ++ struct raid_base_bdev_info *base_info = &raid_bdev->base_bdev_info[slot]; ++ struct spdk_bdev *bdev = spdk_bdev_get_by_name(name); ++ ++ SPDK_CU_ASSERT_FATAL(bdev != NULL); ++ CU_ASSERT(base_info->is_configured); ++ CU_ASSERT(base_info->name != NULL && strcmp(base_info->name, name) == 0); ++ CU_ASSERT(base_info->content_generation == generation); ++ CU_ASSERT(base_info->view_epoch == epoch); ++ CU_ASSERT(bdev->internal.claim_type == SPDK_BDEV_CLAIM_EXCL_WRITE); ++} ++ ++/* Evariops 0053: slot is empty, free for the next add, and records the leg ++ * name left out at its own generation; the leg is not claimed. */ ++static void ++ut_view_left_out(struct raid_bdev *raid_bdev, uint8_t slot, const char *name, uint64_t generation) ++{ ++ struct raid_base_bdev_info *base_info = &raid_bdev->base_bdev_info[slot]; ++ struct spdk_bdev *bdev = spdk_bdev_get_by_name(name); ++ ++ SPDK_CU_ASSERT_FATAL(bdev != NULL); ++ CU_ASSERT(!base_info->is_configured); ++ CU_ASSERT(base_info->name == NULL); ++ CU_ASSERT(spdk_uuid_is_null(&base_info->uuid)); ++ CU_ASSERT(spdk_uuid_compare(&base_info->excluded_uuid, &bdev->uuid) == 0); ++ CU_ASSERT(base_info->content_generation == generation); ++ CU_ASSERT(base_info->data_size == raid_bdev->bdev.blockcnt); ++ CU_ASSERT(bdev->internal.claim_type == SPDK_BDEV_CLAIM_NONE); ++} ++ ++/* Evariops 0053: a refused create leaves no raid and takes no leg. */ ++static void ++ut_view_refused(int code) ++{ ++ const char *names[] = { "Nvme0n1", "Nvme1n1", "Nvme2n1" }; ++ size_t i; ++ ++ CU_ASSERT(g_rpc_err == 1); ++ CU_ASSERT(g_rpc_err_code == code); ++ CU_ASSERT(raid_bdev_find_by_name("raid1") == NULL); ++ for (i = 0; i < SPDK_COUNTOF(names); i++) { ++ struct spdk_bdev *bdev = spdk_bdev_get_by_name(names[i]); ++ ++ SPDK_CU_ASSERT_FATAL(bdev != NULL); ++ CU_ASSERT(bdev->internal.claim_type == SPDK_BDEV_CLAIM_NONE); ++ } ++} ++ ++/* Evariops 0053: legs that carry no superblock, a new volume's, have nothing to ++ * compare: every leg is configured, at generation 0, and the raid carries the ++ * view it was created in. */ ++static void ++test_raid_view_create_over_blank_legs_configures_every_leg(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev *raid_bdev; ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(3); ++ CU_ASSERT(spdk_uuid_compare(&raid_bdev->lineage, &g_ut_lineage) == 0); ++ CU_ASSERT(raid_bdev->view_epoch == 5); ++ ut_view_member(raid_bdev, 0, "Nvme0n1", 0, 5); ++ ut_view_member(raid_bdev, 1, "Nvme1n1", 0, 5); ++ ut_view_member(raid_bdev, 2, "Nvme2n1", 0, 5); ++ ++ ut_view_teardown(); ++} ++ ++/* Evariops 0053: only the legs at the highest generation hold every write the ++ * raid before acknowledged. They are configured, in the first slots, at that ++ * generation; the ejected leg stays out, in the next slot, at its own, never ++ * claimed: the control plane rebuilds it. */ ++static void ++test_raid_view_create_configures_the_legs_at_the_highest_generation(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev *raid_bdev; ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 4); ++ ut_sb_after_an_ejection(); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(2); ++ CU_ASSERT(raid_bdev->num_base_bdevs == 3); ++ CU_ASSERT(raid_bdev->view_epoch == 4); ++ ut_view_member(raid_bdev, 0, "Nvme1n1", 7, 4); ++ ut_view_member(raid_bdev, 1, "Nvme2n1", 7, 4); ++ ut_view_left_out(raid_bdev, 2, "Nvme0n1", 6); ++ /* read once, by the survey, and never claimed */ ++ CU_ASSERT(ut_leg("Nvme0n1")->reads == 1); ++ ++ ut_view_teardown(); ++} ++ ++/* Evariops 0053: a leg with no stamp of the volume carries no generation of it ++ * once another leg does: a blank one, and one stamped before minor 2, whatever ++ * generation it claims, stay out. A copy its leg never completed (its own slot ++ * not configured) does not count either, however high its generation. */ ++static void ++test_raid_view_create_leaves_out_a_leg_without_a_configured_stamp(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev_superblock *stamped, *before_minor_2, *never_completed; ++ struct raid_bdev *raid_bdev; ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 4); ++ stamped = ut_sb(&g_ut_lineage, "incarnation-before", 2); ++ ut_sb_slot(stamped, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 4); ++ ut_sb_slot(stamped, "Nvme1n1", RAID_SB_BASE_BDEV_MISSING, 4); ++ ut_leg("Nvme0n1")->sb = stamped; ++ before_minor_2 = ut_sb(&g_ut_lineage, "incarnation-before", 2); ++ ut_sb_slot(before_minor_2, "Nvme2n1", RAID_SB_BASE_BDEV_CONFIGURED, 9); ++ before_minor_2->version.minor = 1; ++ ut_leg("Nvme2n1")->sb = before_minor_2; ++ /* Nvme1n1 blank */ ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(1); ++ ut_view_member(raid_bdev, 0, "Nvme0n1", 4, 4); ++ ut_view_left_out(raid_bdev, 1, "Nvme1n1", 0); ++ ut_view_left_out(raid_bdev, 2, "Nvme2n1", 0); ++ ut_view_teardown(); ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 4); ++ stamped = ut_sb(&g_ut_lineage, "incarnation-before", 2); ++ ut_sb_slot(stamped, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 5); ++ ut_sb_slot(stamped, "Nvme1n1", RAID_SB_BASE_BDEV_CONFIGURED, 5); ++ ut_sb_slot(stamped, "Nvme2n1", RAID_SB_BASE_BDEV_MISSING, 5); ++ ut_leg("Nvme0n1")->sb = stamped; ++ ut_leg("Nvme1n1")->sb = stamped; ++ never_completed = ut_sb(&g_ut_lineage, "incarnation-before", 2); ++ ut_sb_slot(never_completed, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 5); ++ ut_sb_slot(never_completed, "Nvme2n1", RAID_SB_BASE_BDEV_MISSING, 8); ++ ut_leg("Nvme2n1")->sb = never_completed; ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(2); ++ ut_view_member(raid_bdev, 0, "Nvme0n1", 5, 4); ++ ut_view_member(raid_bdev, 1, "Nvme1n1", 5, 4); ++ ut_view_left_out(raid_bdev, 2, "Nvme2n1", 8); ++ ut_view_teardown(); ++} ++ ++/* Evariops 0053: legs another owner stamped in the view's epoch or a newer one ++ * belong to a raid that a view at least as recent still counts: the create is ++ * refused before it claims any leg (-EBUSY). */ ++static void ++test_raid_view_create_refuses_legs_another_owner_holds(void) ++{ ++ const uint64_t stamp_epochs[] = { 5, 6 }; ++ struct rpc_bdev_raid_create req; ++ size_t i; ++ ++ for (i = 0; i < SPDK_COUNTOF(stamp_epochs); i++) { ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ ut_sb_all_configured("incarnation-elsewhere", stamp_epochs[i], 2); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ ut_view_refused(-EBUSY); ++ ut_view_teardown(); ++ } ++} ++ ++/* Evariops 0053: legs stamped at an older epoch belong to an owner that left ++ * the view, and legs stamped by this create's own incarnation to a replay of ++ * it: both are taken. */ ++static void ++test_raid_view_create_takes_legs_of_an_older_epoch_or_of_its_own(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev *raid_bdev; ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ ut_sb_all_configured("incarnation-elsewhere", 4, 2); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(3); ++ ut_view_member(raid_bdev, 0, "Nvme0n1", 2, 5); ++ ut_view_teardown(); ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ ut_sb_all_configured(UT_INCARNATION, 7, 3); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(3); ++ ut_view_member(raid_bdev, 2, "Nvme2n1", 3, 5); ++ ut_view_teardown(); ++} ++ ++/* Evariops 0053: a leg stamped for another volume is refused whatever its ++ * epoch (-EEXIST), before any other refusal. */ ++static void ++test_raid_view_create_refuses_a_leg_of_another_volume(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev_superblock *held, *other_volume; ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ held = ut_sb(&g_ut_lineage, "incarnation-elsewhere", 9); ++ ut_sb_slot(held, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 1); ++ ut_leg("Nvme0n1")->sb = held; ++ other_volume = ut_sb(&g_ut_other_lineage, "incarnation-elsewhere", 1); ++ ut_sb_slot(other_volume, "Nvme1n1", RAID_SB_BASE_BDEV_CONFIGURED, 1); ++ ut_leg("Nvme1n1")->sb = other_volume; ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ ut_view_refused(-EEXIST); ++ ut_view_teardown(); ++} ++ ++/* Evariops 0053: force_restamp overwrites another owner's stamps only at the ++ * epoch the control plane read from them; at any other, the create is refused ++ * (-ESTALE). */ ++static void ++test_raid_view_create_restamps_only_the_epoch_it_names(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev *raid_bdev; ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ req.force_restamp = true; ++ req.expected_view_epoch = 6; ++ ut_sb_all_configured("incarnation-elsewhere", 6, 3); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(3); ++ ut_view_member(raid_bdev, 1, "Nvme1n1", 3, 5); ++ ut_view_teardown(); ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ req.force_restamp = true; ++ req.expected_view_epoch = 5; ++ ut_sb_all_configured("incarnation-elsewhere", 6, 3); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ ut_view_refused(-ESTALE); ++ ut_view_teardown(); ++} ++ ++/* Evariops 0053: when no stamped leg holds a configured copy, nothing tells ++ * which one holds the volume's writes: refused (-ENODATA). */ ++static void ++test_raid_view_create_refuses_without_a_configured_copy(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev_superblock *sb; ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ sb = ut_sb(&g_ut_lineage, "incarnation-before", 2); ++ ut_sb_slot(sb, "Nvme0n1", RAID_SB_BASE_BDEV_MISSING, 3); ++ ut_sb_slot(sb, "Nvme1n1", RAID_SB_BASE_BDEV_FAILED, 3); ++ ut_leg("Nvme0n1")->sb = sb; ++ ut_leg("Nvme1n1")->sb = sb; ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ ut_view_refused(-ENODATA); ++ ut_view_teardown(); ++} ++ ++/* Evariops 0053: a leg the survey cannot read refuses the create with the ++ * error: an unread leg is not a blank one. Read as blank, the one leg at the ++ * highest generation would be left out, never to be read again, and the raid ++ * would come up on the legs behind it, even if the next read succeeds. */ ++static void ++test_raid_view_create_refuses_a_leg_it_cannot_read(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev_superblock *ahead, *behind; ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ ahead = ut_sb(&g_ut_lineage, "incarnation-before", 3); ++ ut_sb_slot(ahead, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 7); ++ ut_leg("Nvme0n1")->sb = ahead; ++ ut_leg("Nvme0n1")->read_status = -EIO; ++ ut_leg("Nvme0n1")->read_error_transient = true; ++ behind = ut_sb(&g_ut_lineage, "incarnation-before", 3); ++ ut_sb_slot(behind, "Nvme1n1", RAID_SB_BASE_BDEV_CONFIGURED, 6); ++ ut_sb_slot(behind, "Nvme2n1", RAID_SB_BASE_BDEV_CONFIGURED, 6); ++ ut_leg("Nvme1n1")->sb = behind; ++ ut_leg("Nvme2n1")->sb = behind; ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ ut_view_refused(-EIO); ++ ut_view_teardown(); ++} ++ ++/* Evariops 0053: each member is read again as the raid claims it, and must still ++ * be what the arbitration read: a leg whose generation moved, or a leg the ++ * survey found blank that now carries a stamp of the volume, makes the create ++ * stale (-ESTALE), and the raid goes with every leg it had claimed. */ ++static void ++test_raid_view_create_fails_when_a_leg_changed_since_it_was_read(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev_superblock *advanced, *stamped; ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ ut_sb_all_configured("incarnation-before", 4, 4); ++ advanced = ut_sb(&g_ut_lineage, "incarnation-before", 4); ++ ut_sb_slot(advanced, "Nvme1n1", RAID_SB_BASE_BDEV_CONFIGURED, 5); ++ ut_leg("Nvme1n1")->later_sb = advanced; ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ ut_view_refused(-ESTALE); ++ ut_view_teardown(); ++ ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ stamped = ut_sb(&g_ut_lineage, "incarnation-elsewhere", 1); ++ ut_sb_slot(stamped, "Nvme2n1", RAID_SB_BASE_BDEV_CONFIGURED, 1); ++ ut_leg("Nvme2n1")->later_sb = stamped; ++ ut_leg("Nvme2n1")->sb = NULL; ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ ut_view_refused(-ESTALE); ++ ut_view_teardown(); ++} ++ ++/* Evariops 0053: a create in a view is a raid1 with a superblock and the ++ * volume's lineage, and force_restamp names the epoch it overwrites; view ++ * parameters without a view_epoch are refused too. Each is refused before any ++ * leg is read. */ ++static void ++test_raid_view_create_refuses_what_it_cannot_arbitrate(void) ++{ ++ struct rpc_bdev_raid_create req; ++ int c; ++ ++ for (c = 0; c < 7; c++) { ++ ut_view_setup(); ++ ut_view_create_req(&req, 5); ++ ut_leg("Nvme0n1"); ++ ut_leg("Nvme1n1"); ++ ut_leg("Nvme2n1"); ++ switch (c) { ++ case 0: ++ req.superblock_enabled = false; ++ break; ++ case 1: ++ spdk_uuid_set_null(&req.lineage); ++ break; ++ case 2: ++ req.level = 123; ++ req.strip_size_kb = (g_strip_size * g_block_len) / 1024; ++ break; ++ case 3: ++ req.force_restamp = true; ++ break; ++ case 4: ++ req.expected_view_epoch = 5; ++ break; ++ case 5: ++ req.view_epoch = 0; ++ break; ++ case 6: ++ req.declared_slots = 256; ++ break; ++ } ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ ut_view_refused(-EINVAL); ++ CU_ASSERT(ut_leg_reads() == 0); ++ ut_view_teardown(); ++ } ++} ++ ++/* Evariops 0053: declared_slots makes a raid wider than its legs, the slots ++ * past them empty and sized like its members, for members added later; a ++ * level that operates only whole refuses them, and takes no leg. */ ++static void ++test_raid_create_with_declared_slots_leaves_the_slots_past_its_legs_empty(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev *raid_bdev; ++ struct raid_base_bdev_info *empty; ++ ++ ut_view_setup(); ++ create_raid_bdev_create_req(&req, "raid1", 0, true, 0, false); ++ req.level = RAID1; ++ req.strip_size_kb = 0; ++ req.declared_slots = 4; ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(3); ++ CU_ASSERT(raid_bdev->num_base_bdevs == 4); ++ CU_ASSERT(spdk_uuid_is_null(&raid_bdev->lineage)); ++ empty = &raid_bdev->base_bdev_info[3]; ++ CU_ASSERT(!empty->is_configured); ++ CU_ASSERT(empty->name == NULL); ++ CU_ASSERT(empty->data_size == raid_bdev->bdev.blockcnt); ++ CU_ASSERT(raid_bdev->bdev.blockcnt == raid_bdev->base_bdev_info[0].data_size); ++ ut_view_teardown(); ++ ++ ut_view_setup(); ++ create_raid_bdev_create_req(&req, "raid1", 0, true, 0, false); ++ req.declared_slots = 4; ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ++ ut_view_refused(-EINVAL); ++ ut_view_teardown(); ++} ++ ++static void ++ut_add_cb(void *ctx, int status) ++{ ++ *(int *)ctx = status; ++} ++ ++/* Evariops 0053: an add into an online raid of a lineage checks the leg by its ++ * stamp, where upstream refuses any leg another raid stamped: refused when ++ * another owner stamped it in the raid's epoch or a newer one (-EBUSY), or for ++ * another volume (-EEXIST), the slot staying free; taken, as a new member, when ++ * an owner that left the view stamped it. */ ++static void ++test_raid_add_into_a_raid_of_a_lineage_checks_the_stamp(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev_superblock *ejected, *held, *other_volume; ++ struct raid_bdev *raid_bdev; ++ struct spdk_bdev *leg; ++ int status; ++ ++ ut_view_setup(); ++ g_ut_alloc_superblock = true; ++ ut_view_create_req(&req, 4); ++ ejected = ut_sb_after_an_ejection(); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(2); ++ leg = spdk_bdev_get_by_name("Nvme0n1"); ++ SPDK_CU_ASSERT_FATAL(leg != NULL); ++ ++ held = ut_sb(&g_ut_lineage, "incarnation-elsewhere", 4); ++ ut_sb_slot(held, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 9); ++ ut_leg("Nvme0n1")->sb = held; ++ status = 1; ++ CU_ASSERT(raid_bdev_add_base_bdev(raid_bdev, "Nvme0n1", true, false, ut_add_cb, &status) == 0); ++ poll_app_thread(); ++ CU_ASSERT(status == -EBUSY); ++ ut_view_left_out(raid_bdev, 2, "Nvme0n1", 6); ++ ++ other_volume = ut_sb(&g_ut_other_lineage, "incarnation-elsewhere", 1); ++ ut_sb_slot(other_volume, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 1); ++ ut_leg("Nvme0n1")->sb = other_volume; ++ status = 1; ++ CU_ASSERT(raid_bdev_add_base_bdev(raid_bdev, "Nvme0n1", true, false, ut_add_cb, &status) == 0); ++ poll_app_thread(); ++ CU_ASSERT(status == -EEXIST); ++ CU_ASSERT(leg->internal.claim_type == SPDK_BDEV_CLAIM_NONE); ++ ++ ut_leg("Nvme0n1")->sb = ejected; ++ status = 1; ++ CU_ASSERT(raid_bdev_add_base_bdev(raid_bdev, "Nvme0n1", true, false, ut_add_cb, &status) == 0); ++ poll_app_thread(); ++ CU_ASSERT(status == 0); ++ CU_ASSERT(raid_bdev->base_bdev_info[2].is_configured); ++ CU_ASSERT(raid_bdev->num_base_bdevs_discovered == 3); ++ CU_ASSERT(leg->internal.claim_type == SPDK_BDEV_CLAIM_EXCL_WRITE); ++ ++ ut_view_teardown(); ++} ++ ++/* Evariops 0053: a raid reassembled from a stamped superblock keeps the view it ++ * was stamped in, so an add into it checks stamps; one reassembled from a ++ * superblock below minor 2 has none, whatever its reserved bytes held. */ ++static void ++test_raid_reassembled_from_a_stamped_superblock_keeps_its_view(void) ++{ ++ struct raid_bdev_superblock *sb; ++ struct raid_bdev *raid_bdev; ++ ++ ut_view_setup(); ++ g_ut_alloc_superblock = true; ++ create_base_bdevs(0); ++ ++ sb = ut_sb(&g_ut_lineage, "incarnation-before", 4); ++ ut_sb_slot(sb, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 2); ++ ut_sb_slot(sb, "Nvme1n1", RAID_SB_BASE_BDEV_CONFIGURED, 2); ++ CU_ASSERT(raid_bdev_create_from_sb(sb, &raid_bdev) == 0); ++ CU_ASSERT(spdk_uuid_compare(&raid_bdev->lineage, &g_ut_lineage) == 0); ++ CU_ASSERT(raid_bdev->view_epoch == 4); ++ CU_ASSERT(raid_bdev->incarnation == NULL); ++ raid_bdev_delete(raid_bdev, NULL, NULL); ++ ++ sb = ut_sb(&g_ut_lineage, "incarnation-before", 4); ++ ut_sb_slot(sb, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 2); ++ ut_sb_slot(sb, "Nvme1n1", RAID_SB_BASE_BDEV_CONFIGURED, 2); ++ sb->version.minor = 1; ++ CU_ASSERT(raid_bdev_create_from_sb(sb, &raid_bdev) == 0); ++ CU_ASSERT(spdk_uuid_is_null(&raid_bdev->lineage)); ++ CU_ASSERT(raid_bdev->view_epoch == 0); ++ raid_bdev_delete(raid_bdev, NULL, NULL); ++ ++ ut_view_teardown(); ++} ++ + static int + test_new_thread_fn(struct spdk_thread *thread) + { +@@ -2380,6 +3234,20 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_raid_process_with_qos); + CU_ADD_TEST(suite, test_raid1_ejection_opens_an_epoch_on_the_raids_cbt); + CU_ADD_TEST(suite, test_raid1_ejection_of_a_member_behind_opens_no_epoch); ++ CU_ADD_TEST(suite, test_raid_view_create_over_blank_legs_configures_every_leg); ++ CU_ADD_TEST(suite, test_raid_view_create_configures_the_legs_at_the_highest_generation); ++ CU_ADD_TEST(suite, test_raid_view_create_leaves_out_a_leg_without_a_configured_stamp); ++ CU_ADD_TEST(suite, test_raid_view_create_refuses_legs_another_owner_holds); ++ CU_ADD_TEST(suite, test_raid_view_create_takes_legs_of_an_older_epoch_or_of_its_own); ++ CU_ADD_TEST(suite, test_raid_view_create_refuses_a_leg_of_another_volume); ++ CU_ADD_TEST(suite, test_raid_view_create_restamps_only_the_epoch_it_names); ++ CU_ADD_TEST(suite, test_raid_view_create_refuses_without_a_configured_copy); ++ CU_ADD_TEST(suite, test_raid_view_create_refuses_a_leg_it_cannot_read); ++ CU_ADD_TEST(suite, test_raid_view_create_fails_when_a_leg_changed_since_it_was_read); ++ CU_ADD_TEST(suite, test_raid_view_create_refuses_what_it_cannot_arbitrate); ++ CU_ADD_TEST(suite, test_raid_create_with_declared_slots_leaves_the_slots_past_its_legs_empty); ++ CU_ADD_TEST(suite, test_raid_add_into_a_raid_of_a_lineage_checks_the_stamp); ++ CU_ADD_TEST(suite, test_raid_reassembled_from_a_stamped_superblock_keeps_its_view); + + spdk_thread_lib_init(test_new_thread_fn, 0); + g_app_thread = spdk_thread_create("app_thread", NULL); +diff --git a/test/unit/lib/bdev/raid/bdev_raid_sb.c/bdev_raid_sb_ut.c b/test/unit/lib/bdev/raid/bdev_raid_sb.c/bdev_raid_sb_ut.c +index f5ce7ec..14125a4 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid_sb.c/bdev_raid_sb_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid_sb.c/bdev_raid_sb_ut.c +@@ -377,6 +377,121 @@ test_raid_bdev_parse_superblock(void) + CU_ASSERT(raid_bdev_parse_superblock(&ctx) == -EINVAL); + } + ++/* Evariops 0053: a raid created in a view stamps its lineage, its owner and its ++ * epoch, at minor 2, and records each slot as its create laid it out: a member ++ * CONFIGURED, a leg left out MISSING under its UUID, a slot no leg held MISSING ++ * under a UUID of its own, each at the generation and epoch the create set. The ++ * stamp survives the write and the read back. */ ++static void ++test_raid_bdev_init_superblock_stamps_the_view(void) ++{ ++ struct raid_base_bdev_info base_info[3] = {}; ++ char incarnation[] = "incarnation-ut"; ++ struct raid_bdev raid_bdev = { ++ .num_base_bdevs = SPDK_COUNTOF(base_info), ++ .base_bdev_info = base_info, ++ .bdev = g_bdev, ++ .incarnation = incarnation, ++ .view_epoch = 4, ++ }; ++ struct raid_bdev_superblock *sb; ++ struct spdk_uuid member, left_out; ++ int status; ++ uint8_t i; ++ ++ spdk_uuid_generate(&raid_bdev.lineage); ++ spdk_uuid_generate(&member); ++ spdk_uuid_generate(&left_out); ++ for (i = 0; i < SPDK_COUNTOF(base_info); i++) { ++ base_info[i].raid_bdev = &raid_bdev; ++ base_info[i].data_size = 1024; ++ } ++ base_info[0].is_configured = true; ++ spdk_uuid_copy(&base_info[0].uuid, &member); ++ base_info[0].data_offset = 2048; ++ base_info[0].content_generation = 7; ++ base_info[0].view_epoch = 4; ++ spdk_uuid_copy(&base_info[1].excluded_uuid, &left_out); ++ base_info[1].content_generation = 6; ++ base_info[1].view_epoch = 3; ++ ++ CU_ASSERT(raid_bdev_alloc_superblock(&raid_bdev, ++ spdk_bdev_get_data_block_size(&raid_bdev.bdev)) == 0); ++ raid_bdev_init_superblock(&raid_bdev); ++ sb = raid_bdev.sb; ++ ++ CU_ASSERT(sb->version.minor == 2); ++ CU_ASSERT(spdk_uuid_compare(&sb->lineage, &raid_bdev.lineage) == 0); ++ CU_ASSERT(strcmp((const char *)sb->owner, incarnation) == 0); ++ CU_ASSERT(sb->view_epoch == 4); ++ CU_ASSERT(sb->base_bdevs_size == 3); ++ ++ CU_ASSERT(sb->base_bdevs[0].state == RAID_SB_BASE_BDEV_CONFIGURED); ++ CU_ASSERT(spdk_uuid_compare(&sb->base_bdevs[0].uuid, &member) == 0); ++ CU_ASSERT(sb->base_bdevs[0].data_offset == 2048); ++ CU_ASSERT(sb->base_bdevs[0].content_generation == 7); ++ CU_ASSERT(sb->base_bdevs[0].view_epoch == 4); ++ ++ CU_ASSERT(sb->base_bdevs[1].state == RAID_SB_BASE_BDEV_MISSING); ++ CU_ASSERT(spdk_uuid_compare(&sb->base_bdevs[1].uuid, &left_out) == 0); ++ CU_ASSERT(sb->base_bdevs[1].content_generation == 6); ++ CU_ASSERT(sb->base_bdevs[1].view_epoch == 3); ++ ++ CU_ASSERT(sb->base_bdevs[2].state == RAID_SB_BASE_BDEV_MISSING); ++ CU_ASSERT(!spdk_uuid_is_null(&sb->base_bdevs[2].uuid)); ++ CU_ASSERT(spdk_uuid_compare(&sb->base_bdevs[2].uuid, &member) != 0); ++ CU_ASSERT(spdk_uuid_compare(&sb->base_bdevs[2].uuid, &left_out) != 0); ++ CU_ASSERT(sb->base_bdevs[2].content_generation == 0); ++ CU_ASSERT(sb->base_bdevs[2].data_size == 1024); ++ ++ status = INT_MAX; ++ raid_bdev_write_superblock(&raid_bdev, write_sb_cb, &status); ++ process_io_completions(); ++ CU_ASSERT(status == 0); ++ status = INT_MAX; ++ CU_ASSERT(raid_bdev_load_base_bdev_superblock(NULL, NULL, load_sb_cb, &status) == 0); ++ CU_ASSERT(status == 0); ++ ++ raid_bdev_free_superblock(&raid_bdev); ++} ++ ++/* Evariops 0053: a raid created outside the view protocol writes minor 2 with no ++ * stamp: a null lineage, no owner, epoch 0. */ ++static void ++test_raid_bdev_init_superblock_outside_a_view_leaves_the_stamp_empty(void) ++{ ++ struct raid_base_bdev_info base_info[2] = {}; ++ char incarnation[] = "incarnation-ut"; ++ struct raid_bdev raid_bdev = { ++ .num_base_bdevs = SPDK_COUNTOF(base_info), ++ .base_bdev_info = base_info, ++ .bdev = g_bdev, ++ .incarnation = incarnation, ++ }; ++ struct raid_bdev_superblock *sb; ++ uint8_t i; ++ ++ for (i = 0; i < SPDK_COUNTOF(base_info); i++) { ++ base_info[i].raid_bdev = &raid_bdev; ++ base_info[i].is_configured = true; ++ spdk_uuid_generate(&base_info[i].uuid); ++ } ++ ++ CU_ASSERT(raid_bdev_alloc_superblock(&raid_bdev, ++ spdk_bdev_get_data_block_size(&raid_bdev.bdev)) == 0); ++ raid_bdev_init_superblock(&raid_bdev); ++ sb = raid_bdev.sb; ++ ++ CU_ASSERT(sb->version.minor == 2); ++ CU_ASSERT(spdk_uuid_is_null(&sb->lineage)); ++ CU_ASSERT(sb->owner[0] == '\0'); ++ CU_ASSERT(sb->view_epoch == 0); ++ CU_ASSERT(sb->base_bdevs[0].state == RAID_SB_BASE_BDEV_CONFIGURED); ++ CU_ASSERT(sb->base_bdevs[1].state == RAID_SB_BASE_BDEV_CONFIGURED); ++ ++ raid_bdev_free_superblock(&raid_bdev); ++} ++ + int + main(int argc, char **argv) + { +@@ -385,6 +500,8 @@ main(int argc, char **argv) + { "test_raid_bdev_write_superblock", test_raid_bdev_write_superblock }, + { "test_raid_bdev_load_base_bdev_superblock", test_raid_bdev_load_base_bdev_superblock }, + { "test_raid_bdev_parse_superblock", test_raid_bdev_parse_superblock }, ++ { "test_raid_bdev_init_superblock_stamps_the_view", test_raid_bdev_init_superblock_stamps_the_view }, ++ { "test_raid_bdev_init_superblock_outside_a_view_leaves_the_stamp_empty", test_raid_bdev_init_superblock_outside_a_view_leaves_the_stamp_empty }, + CU_TEST_INFO_NULL, + }; + CU_SuiteInfo suites[] = { diff --git a/patches/README.md b/patches/README.md index 41421bd..d7bec22 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0052`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0053`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -58,6 +58,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0052`) | 0050 | a reservation held by all registrants outlives the registrant that acquired it — when that registrant leaves (preempted by another, or unregistering), the reservation, which every registrant holds, now stands under the key of the registrant that holds it next. It kept the key of the one that left, and the persisted state (`ptpl_file`) recorded a reservation key no registrant held: the restore refused it, so after a writer handover, the next restart of a target failed to add the namespace back, and the target stayed unexported. The restore also takes a state persisted before this patch, under the key of its first registrant; a single holder's key that no registrant holds is still refused. `subsystem_ut` pins the preempt, the unregister and the restore | lib/nvmf (subsystem.c), test/unit/lib/nvmf/subsystem.c | — | | 0051 | an I/O passthrough needs a connected qpair — `bdev_nvme_send_cmd` with an I/O command (NVMe reservations are I/O commands) took the controller channel's qpair without a check and handed it to the NVMe library. bdev_nvme frees that qpair from a disconnect until the reconnect (reset, failed controller, delayed reconnect), so a reservation sent to a member whose path had just gone down dereferenced NULL and killed the process, and every raid it carried with it. The command is now refused (`-ENXIO`, the channel given back), as is one for which no channel can be had. `bdev_nvme_ut` pins the disconnected and the connected case | module/bdev/nvme (nvme_rpc.c), test/unit/lib/bdev/nvme | — | | 0052 | a raid1 ejection opens its epoch on the raid's cbt — when a current member leaves a raid1 (removed, hot-removed or failed), the cbt stacked on the raid opens an epoch named after the member's UUID, the identity the control plane derives from the replica it placed there, and records from then on every write the member misses, so its return is a delta instead of a full rebuild. 0019 asked each survivor's own name for it, a bdev no cbt sits on, and no epoch ever opened. Only for a member that left current (its superblock slot configured at its peers' generation, not a write-only joiner, not a failed member an earlier removal failed to take), and only while the cbt carries no live epoch: a freeze empties the live bitmap, and an epoch opened after one may start short of writes the member missed just before it left. `bdev_raid_ut` pins the removed and the failed member, and the five ejections that open none | bdev_raid, cbt module (`vbdev_cbt_auto_epoch_open`), test/unit/lib/bdev/raid | 0016, 0019, 0033 | +| 0053 | a raid create in a view arbitrates its members from their superblocks — `bdev_raid_create` with a `view_epoch` (a raid1 with a superblock, the volume's `lineage`, the creating `incarnation`) reads every listed leg's superblock, opened read-only, before it claims any, and takes only what the stamps allow. A superblock now carries a stamp (minor 2, carved out of reserved bytes): the volume's `lineage`, the `owner` incarnation and the last `view_epoch` it stamped. A leg stamped for another volume is refused (`-EEXIST`), and so is one another owner stamped in the view's epoch or a newer one (`-EBUSY`), with nothing claimed or written, unless `force_restamp` names that exact epoch (`expected_view_epoch`, `-ESTALE` otherwise). Among the legs stamped for the volume, only those at the highest content generation, configured in their own superblock, are configured, at that generation; every other leg (behind, never completed, unstamped, or stamped before minor 2) stays out, recorded MISSING at its own generation, for the control plane to rebuild, and no stamped leg holding a configured copy is `-ENODATA`. When no leg carries a stamp of the volume, every leg is configured, as before. Each member is read again as it is claimed and must not have moved (`-ESTALE`); a leg the survey cannot read refuses the create, since read as blank the freshest leg would be left out and never read again. An add into an online raid of a lineage checks the leg the same way, where upstream refuses any leg another raid stamped. `declared_slots` creates a raid wider than its legs, the slots past them empty and sized like the members. `bdev_raid_ut` and `bdev_raid_sb_ut` pin each rule, and each was seen to fail with its rule taken out | bdev_raid (raid, rpc, superblock), test/unit/lib/bdev/raid | 0014, 0016 | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From 7382e98f8288c260c7328f969d0d8ce0f6b04aeb Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Sat, 3 Oct 2026 14:28:48 +0200 Subject: [PATCH 23/26] feat(raid): the raid's owner journals each view epoch on its members (0054) 0053 refuses a create on legs another owner stamped in the create's view epoch or a newer one. A raid stamps its legs once, at its create, so an owner that stays in the view while the view moves on held them with an old epoch, and a create from a view the owner already left behind took them. bdev_raid_journal_view_epoch {name, view_epoch, expected_incarnation, lineage} stamps a new epoch on an online raid: the superblock's view_epoch, its owner (the raid's incarnation, which takes over the stamp of a raid it claimed after a reassembly), and the view_epoch of each member it serves, then writes the superblock to them. An owner that stays in the view journals each epoch, so a create from a view no more recent finds the legs held for as long as the owner serves. - Only the raid's incarnation journals (-ESTALE for another). - The epoch never goes back (-ESTALE). The same epoch is written again, so a journal whose write failed is retried as it was. - A raid created outside the view protocol joins its volume's view when the journal names the lineage (-EINVAL without one, -EEXIST for another lineage). - A superblock an older image wrote is raised to minor 2, where the stamp is read. bdev_raid_ut pins the stamp, each refusal, the join and a failed write. With superblock writes landing on the legs, a create by another incarnation is refused at the journaled epoch and at the one before, and takes the legs at the next. Taken out in turn, the stamp (3 tests, the end-to-end one included), the monotonic epoch (1), the minor 2 upgrade (2) and the join (1) each turned their tests red. Every raid suite passes: bdev_raid 44, bdev_raid_sb 15, raid0 8, raid1 7, raid5f 8, concat 3. --- ...rnals-each-view-epoch-on-its-members.patch | 596 ++++++++++++++++++ patches/README.md | 3 +- 2 files changed, 598 insertions(+), 1 deletion(-) create mode 100644 patches/0054-raid-owner-journals-each-view-epoch-on-its-members.patch diff --git a/patches/0054-raid-owner-journals-each-view-epoch-on-its-members.patch b/patches/0054-raid-owner-journals-each-view-epoch-on-its-members.patch new file mode 100644 index 0000000..95aa9f6 --- /dev/null +++ b/patches/0054-raid-owner-journals-each-view-epoch-on-its-members.patch @@ -0,0 +1,596 @@ +diff --git a/module/bdev/raid/bdev_raid.c b/module/bdev/raid/bdev_raid.c +index c18bde7..0bd65e1 100644 +--- a/module/bdev/raid/bdev_raid.c ++++ b/module/bdev/raid/bdev_raid.c +@@ -5147,6 +5147,77 @@ raid_bdev_lay_out_view(struct raid_bdev *raid_bdev, const struct raid_bdev_view + return 0; + } + ++/* Evariops 0054: an owner that stays in the view stamps each new epoch on the ++ * members it serves, so a create from a view no more recent finds the legs held ++ * and is refused for as long as the owner serves. The stamp names the raid's ++ * incarnation, which takes over the stamp of a raid it claimed after a ++ * reassembly, and a raid created outside the view protocol joins its volume's ++ * view through it. The same epoch is written again, so a journal whose write ++ * failed is retried as it was. */ ++int ++raid_bdev_journal_view_epoch(struct raid_bdev *raid_bdev, uint64_t epoch, ++ const struct spdk_uuid *lineage, raid_bdev_write_sb_cb cb, void *cb_ctx) ++{ ++ struct raid_bdev_superblock *sb = raid_bdev->sb; ++ struct raid_base_bdev_info *base_info; ++ uint8_t i; ++ ++ assert(spdk_get_thread() == spdk_thread_get_app_thread()); ++ ++ if (raid_bdev->state != RAID_BDEV_STATE_ONLINE || sb == NULL || ++ raid_bdev->incarnation == NULL || epoch == 0) { ++ return -EINVAL; ++ } ++ ++ if (lineage != NULL && !spdk_uuid_is_null(&raid_bdev->lineage) && ++ spdk_uuid_compare(lineage, &raid_bdev->lineage) != 0) { ++ return -EEXIST; ++ } ++ ++ if (lineage == NULL && spdk_uuid_is_null(&raid_bdev->lineage)) { ++ return -EINVAL; ++ } ++ ++ if (epoch < raid_bdev->view_epoch) { ++ return -ESTALE; ++ } ++ ++ if (spdk_uuid_is_null(&raid_bdev->lineage)) { ++ spdk_uuid_copy(&raid_bdev->lineage, lineage); ++ } ++ raid_bdev->view_epoch = epoch; ++ ++ /* a superblock an older image wrote has no stamp to read before minor 2 */ ++ if (sb->version.minor < 2) { ++ sb->version.minor = 2; ++ } ++ spdk_uuid_copy(&sb->lineage, &raid_bdev->lineage); ++ memset(sb->owner, 0, sizeof(sb->owner)); ++ snprintf((char *)sb->owner, sizeof(sb->owner), "%s", raid_bdev->incarnation); ++ sb->view_epoch = epoch; ++ ++ for (i = 0; i < sb->base_bdevs_size; i++) { ++ struct raid_bdev_sb_base_bdev *sb_base_bdev = &sb->base_bdevs[i]; ++ ++ if (sb_base_bdev->state != RAID_SB_BASE_BDEV_CONFIGURED || ++ sb_base_bdev->slot >= raid_bdev->num_base_bdevs) { ++ continue; ++ } ++ ++ base_info = &raid_bdev->base_bdev_info[sb_base_bdev->slot]; ++ if (!base_info->is_configured) { ++ continue; ++ } ++ ++ sb_base_bdev->view_epoch = epoch; ++ base_info->view_epoch = epoch; ++ } ++ ++ raid_bdev_write_superblock(raid_bdev, cb, cb_ctx); ++ ++ return 0; ++} ++ + static void raid_bdev_examine_sb(const struct raid_bdev_superblock *sb, struct spdk_bdev *bdev, + raid_base_bdev_cb cb_fn, void *cb_ctx); + +diff --git a/module/bdev/raid/bdev_raid.h b/module/bdev/raid/bdev_raid.h +index 2594564..5c30917 100644 +--- a/module/bdev/raid/bdev_raid.h ++++ b/module/bdev/raid/bdev_raid.h +@@ -793,6 +793,14 @@ void raid_bdev_write_superblock(struct raid_bdev *raid_bdev, raid_bdev_write_sb_ + int raid_bdev_load_base_bdev_superblock(struct spdk_bdev_desc *desc, struct spdk_io_channel *ch, + raid_bdev_load_sb_cb cb, void *cb_ctx); + ++/* Evariops 0054: stamps view epoch on an online raid's members, under the raid's ++ * incarnation; lineage joins a raid created outside the view protocol to its ++ * volume's view. cb reports the superblock write. -ESTALE for an epoch older ++ * than the raid's, -EEXIST for another lineage, -EINVAL for a raid offline, ++ * without a superblock or an owner, or outside the view without a lineage. */ ++int raid_bdev_journal_view_epoch(struct raid_bdev *raid_bdev, uint64_t epoch, ++ const struct spdk_uuid *lineage, raid_bdev_write_sb_cb cb, void *cb_ctx); ++ + struct spdk_raid_bdev_opts { + /* Size of the background process window in KiB */ + uint32_t process_window_size_kb; +diff --git a/module/bdev/raid/bdev_raid_rpc.c b/module/bdev/raid/bdev_raid_rpc.c +index 9119230..aada046 100644 +--- a/module/bdev/raid/bdev_raid_rpc.c ++++ b/module/bdev/raid/bdev_raid_rpc.c +@@ -597,6 +597,83 @@ rpc_bdev_raid_claim(struct spdk_jsonrpc_request *request, const struct spdk_json + } + SPDK_RPC_REGISTER("bdev_raid_claim", rpc_bdev_raid_claim, SPDK_RPC_RUNTIME) + ++/* Evariops 0054: the raid's owner stamps a new view epoch on its members. Only ++ * the raid's incarnation journals (-ESTALE for another), and lineage joins a ++ * raid created outside the view protocol to its volume's view. */ ++struct rpc_bdev_raid_journal_view_epoch { ++ char *name; ++ uint64_t view_epoch; ++ char *expected_incarnation; ++ struct spdk_uuid lineage; ++}; ++ ++static const struct spdk_json_object_decoder rpc_bdev_raid_journal_view_epoch_decoders[] = { ++ {"name", offsetof(struct rpc_bdev_raid_journal_view_epoch, name), spdk_json_decode_string}, ++ {"view_epoch", offsetof(struct rpc_bdev_raid_journal_view_epoch, view_epoch), spdk_json_decode_uint64}, ++ {"expected_incarnation", offsetof(struct rpc_bdev_raid_journal_view_epoch, expected_incarnation), spdk_json_decode_string}, ++ {"lineage", offsetof(struct rpc_bdev_raid_journal_view_epoch, lineage), spdk_json_decode_uuid, true}, ++}; ++ ++static void ++rpc_bdev_raid_journal_view_epoch_written(int status, struct raid_bdev *raid_bdev, void *ctx) ++{ ++ struct spdk_jsonrpc_request *request = ctx; ++ ++ if (status != 0) { ++ spdk_jsonrpc_send_error_response_fmt(request, status, ++ "raid bdev %s: view epoch %" PRIu64 " not written: %s", ++ raid_bdev->bdev.name, raid_bdev->view_epoch, ++ spdk_strerror(-status)); ++ return; ++ } ++ ++ spdk_jsonrpc_send_bool_response(request, true); ++} ++ ++static void ++rpc_bdev_raid_journal_view_epoch(struct spdk_jsonrpc_request *request, ++ const struct spdk_json_val *params) ++{ ++ struct rpc_bdev_raid_journal_view_epoch req = {}; ++ struct raid_bdev *raid_bdev; ++ int rc; ++ ++ if (spdk_json_decode_object(params, rpc_bdev_raid_journal_view_epoch_decoders, ++ SPDK_COUNTOF(rpc_bdev_raid_journal_view_epoch_decoders), &req)) { ++ spdk_jsonrpc_send_error_response(request, SPDK_JSONRPC_ERROR_INVALID_PARAMS, ++ "Invalid parameters"); ++ goto cleanup; ++ } ++ ++ raid_bdev = raid_bdev_find_by_name(req.name); ++ if (raid_bdev == NULL) { ++ spdk_jsonrpc_send_error_response_fmt(request, -ENODEV, "No such raid bdev: %s", req.name); ++ goto cleanup; ++ } ++ ++ rc = raid_bdev_check_incarnation(raid_bdev, req.expected_incarnation); ++ if (rc != 0) { ++ spdk_jsonrpc_send_error_response_fmt(request, rc, ++ "raid bdev %s belongs to another incarnation than '%s'", ++ req.name, req.expected_incarnation); ++ goto cleanup; ++ } ++ ++ rc = raid_bdev_journal_view_epoch(raid_bdev, req.view_epoch, ++ spdk_uuid_is_null(&req.lineage) ? NULL : &req.lineage, ++ rpc_bdev_raid_journal_view_epoch_written, request); ++ if (rc != 0) { ++ spdk_jsonrpc_send_error_response_fmt(request, rc, ++ "raid bdev %s: view epoch %" PRIu64 " not journaled " ++ "(the raid's: %" PRIu64 "): %s", req.name, req.view_epoch, ++ raid_bdev->view_epoch, spdk_strerror(-rc)); ++ } ++cleanup: ++ free(req.name); ++ free(req.expected_incarnation); ++} ++SPDK_RPC_REGISTER("bdev_raid_journal_view_epoch", rpc_bdev_raid_journal_view_epoch, SPDK_RPC_RUNTIME) ++ + /* + * Input structure for RPC deleting a raid bdev + */ +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index 0f21d2c..beaec9b 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -355,7 +355,7 @@ spdk_bdev_get_data_block_size(const struct spdk_bdev *bdev) + * its later superblock when it has one: a leg that changed between a create's + * survey and its claim. */ + #define UT_LEGS_MAX 4 +-#define UT_SBS_MAX 8 ++#define UT_SBS_MAX 16 + + struct ut_leg { + const char *name; +@@ -413,6 +413,23 @@ ut_legs_reset(void) + } + } + ++/* Evariops 0054: a superblock buffer, freed with the legs. */ ++static struct raid_bdev_superblock * ++ut_sb_alloc(void) ++{ ++ struct raid_bdev_superblock *sb; ++ size_t i; ++ ++ for (i = 0; i < SPDK_COUNTOF(g_ut_sbs) && g_ut_sbs[i] != NULL; i++) { ++ } ++ SPDK_CU_ASSERT_FATAL(i < SPDK_COUNTOF(g_ut_sbs)); ++ sb = calloc(1, RAID_BDEV_SB_MAX_LENGTH); ++ SPDK_CU_ASSERT_FATAL(sb != NULL); ++ g_ut_sbs[i] = sb; ++ ++ return sb; ++} ++ + int + raid_bdev_load_base_bdev_superblock(struct spdk_bdev_desc *desc, struct spdk_io_channel *ch, + raid_bdev_load_sb_cb cb, void *cb_ctx) +@@ -450,10 +467,33 @@ raid_bdev_load_base_bdev_superblock(struct spdk_bdev_desc *desc, struct spdk_io_ + return 0; + } + ++/* Evariops 0054: the superblock writes the raids ask for, the status they ++ * complete with, and, while g_ut_sb_writes_land, each lands on the raid's ++ * configured members as the superblock their next read returns. */ ++static uint32_t g_ut_sb_writes; ++static int g_ut_sb_write_status; ++static bool g_ut_sb_writes_land; ++ + void + raid_bdev_write_superblock(struct raid_bdev *raid_bdev, raid_bdev_write_sb_cb cb, void *cb_ctx) + { +- cb(0, raid_bdev, cb_ctx); ++ struct raid_base_bdev_info *base_info; ++ struct raid_bdev_superblock *landed; ++ ++ g_ut_sb_writes++; ++ if (g_ut_sb_write_status == 0 && g_ut_sb_writes_land && raid_bdev->sb != NULL) { ++ landed = ut_sb_alloc(); ++ memcpy(landed, raid_bdev->sb, RAID_BDEV_SB_MAX_LENGTH); ++ RAID_FOR_EACH_BASE_BDEV(raid_bdev, base_info) { ++ if (base_info->is_configured && base_info->desc != NULL) { ++ struct ut_leg *leg = ut_leg(spdk_bdev_desc_get_bdev(base_info->desc)->name); ++ ++ leg->sb = landed; ++ leg->later_sb = NULL; ++ } ++ } ++ } ++ cb(g_ut_sb_write_status, raid_bdev, cb_ctx); + } + + const struct spdk_uuid * +@@ -2498,6 +2538,8 @@ ut_view_teardown(void) + reset_globals(); + ut_legs_reset(); + g_ut_alloc_superblock = false; ++ g_ut_sb_write_status = 0; ++ g_ut_sb_writes_land = false; + g_max_base_drives = g_ut_saved_max_base_drives; + } + +@@ -2519,15 +2561,7 @@ ut_view_create_req(struct rpc_bdev_raid_create *r, uint64_t epoch) + static struct raid_bdev_superblock * + ut_sb(const struct spdk_uuid *lineage, const char *owner, uint64_t epoch) + { +- struct raid_bdev_superblock *sb; +- size_t i; +- +- for (i = 0; i < SPDK_COUNTOF(g_ut_sbs) && g_ut_sbs[i] != NULL; i++) { +- } +- SPDK_CU_ASSERT_FATAL(i < SPDK_COUNTOF(g_ut_sbs)); +- sb = calloc(1, RAID_BDEV_SB_MAX_LENGTH); +- SPDK_CU_ASSERT_FATAL(sb != NULL); +- g_ut_sbs[i] = sb; ++ struct raid_bdev_superblock *sb = ut_sb_alloc(); + + memcpy(sb->signature, RAID_BDEV_SB_SIG, sizeof(sb->signature)); + sb->version.major = RAID_BDEV_SB_VERSION_MAJOR; +@@ -3181,6 +3215,296 @@ test_raid_reassembled_from_a_stamped_superblock_keeps_its_view(void) + ut_view_teardown(); + } + ++/* Evariops 0054: fills the superblock the suite's stubbed init leaves empty, as ++ * init writes it, at minor: each slot under its member, CONFIGURED or MISSING, ++ * at the generation and epoch the raid holds for it. */ ++static void ++ut_fill_sb(struct raid_bdev *raid_bdev, uint16_t minor) ++{ ++ struct raid_bdev_superblock *sb = raid_bdev->sb; ++ struct raid_base_bdev_info *base_info; ++ uint8_t i = 0; ++ ++ SPDK_CU_ASSERT_FATAL(sb != NULL); ++ memcpy(sb->signature, RAID_BDEV_SB_SIG, sizeof(sb->signature)); ++ sb->version.major = RAID_BDEV_SB_VERSION_MAJOR; ++ sb->version.minor = minor; ++ spdk_uuid_copy(&sb->uuid, &raid_bdev->bdev.uuid); ++ sb->block_size = g_block_len; ++ sb->level = raid_bdev->level; ++ sb->num_base_bdevs = sb->base_bdevs_size = raid_bdev->num_base_bdevs; ++ sb->length = sizeof(*sb) + sb->base_bdevs_size * sizeof(sb->base_bdevs[0]); ++ RAID_FOR_EACH_BASE_BDEV(raid_bdev, base_info) { ++ struct raid_bdev_sb_base_bdev *slot = &sb->base_bdevs[i]; ++ ++ spdk_uuid_copy(&slot->uuid, base_info->is_configured ? &base_info->uuid : ++ &base_info->excluded_uuid); ++ slot->state = base_info->is_configured ? RAID_SB_BASE_BDEV_CONFIGURED : ++ RAID_SB_BASE_BDEV_MISSING; ++ slot->slot = i++; ++ slot->content_generation = base_info->content_generation; ++ slot->view_epoch = base_info->view_epoch; ++ } ++} ++ ++/* Evariops 0054: journals epoch on raid name as incarnation, with lineage ++ * unless NULL. */ ++static void ++ut_journal(const char *name, uint64_t epoch, const char *incarnation, const struct spdk_uuid *lineage) ++{ ++ struct rpc_bdev_raid_journal_view_epoch req = {}; ++ ++ req.name = strdup(name); ++ req.expected_incarnation = strdup(incarnation); ++ SPDK_CU_ASSERT_FATAL(req.name != NULL && req.expected_incarnation != NULL); ++ req.view_epoch = epoch; ++ if (lineage != NULL) { ++ spdk_uuid_copy(&req.lineage, lineage); ++ } ++ g_rpc_req = &req; ++ g_rpc_req_size = sizeof(req); ++ g_json_decode_obj_create = 0; ++ g_json_decode_obj_err = 0; ++ g_rpc_err = 0; ++ g_rpc_err_code = 0; ++ rpc_bdev_raid_journal_view_epoch(NULL, NULL); ++ poll_app_thread(); ++ g_rpc_req = NULL; ++} ++ ++/* Evariops 0054: the raid of the view tests after an ejection, created at epoch ++ * 4: Nvme1n1 and Nvme2n1 configured at generation 7, Nvme0n1 left out. Its ++ * superblock is filled at minor. */ ++static struct raid_bdev * ++ut_view_raid_after_an_ejection(uint16_t minor) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev *raid_bdev; ++ ++ g_ut_alloc_superblock = true; ++ ut_view_create_req(&req, 4); ++ ut_sb_after_an_ejection(); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(2); ++ ut_fill_sb(raid_bdev, minor); ++ ++ return raid_bdev; ++} ++ ++/* Evariops 0054: the owner journals a new epoch: the raid, its superblock and ++ * the members it serves carry it, under the owner's incarnation, and a ++ * superblock an older image wrote is raised to minor 2, where the stamp is ++ * read. The slot of a member not serving keeps its epoch. The same epoch is ++ * written again. */ ++static void ++test_raid_journal_stamps_the_view_epoch_on_the_members(void) ++{ ++ struct raid_bdev_superblock *sb; ++ struct raid_bdev *raid_bdev; ++ uint32_t writes; ++ ++ ut_view_setup(); ++ raid_bdev = ut_view_raid_after_an_ejection(1); ++ sb = raid_bdev->sb; ++ writes = g_ut_sb_writes; ++ ++ ut_journal("raid1", 6, UT_INCARNATION, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ CU_ASSERT(g_ut_sb_writes == writes + 1); ++ CU_ASSERT(raid_bdev->view_epoch == 6); ++ CU_ASSERT(sb->version.minor == 2); ++ CU_ASSERT(spdk_uuid_compare(&sb->lineage, &g_ut_lineage) == 0); ++ CU_ASSERT(strcmp((const char *)sb->owner, UT_INCARNATION) == 0); ++ CU_ASSERT(sb->view_epoch == 6); ++ CU_ASSERT(sb->base_bdevs[0].view_epoch == 6); ++ CU_ASSERT(sb->base_bdevs[1].view_epoch == 6); ++ CU_ASSERT(sb->base_bdevs[2].view_epoch == 0); ++ CU_ASSERT(raid_bdev->base_bdev_info[0].view_epoch == 6); ++ CU_ASSERT(raid_bdev->base_bdev_info[1].view_epoch == 6); ++ CU_ASSERT(raid_bdev->base_bdev_info[2].view_epoch == 0); ++ ++ ut_journal("raid1", 6, UT_INCARNATION, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ CU_ASSERT(g_ut_sb_writes == writes + 2); ++ ++ ut_view_teardown(); ++} ++ ++/* Evariops 0054: no journal goes back in time, none comes from another ++ * incarnation or names another volume, and none journals epoch 0 or a raid that ++ * is not there. Each leaves the raid's epoch and writes nothing. */ ++static void ++test_raid_journal_refuses_what_its_owner_did_not_send(void) ++{ ++ struct raid_bdev *raid_bdev; ++ uint32_t writes; ++ ++ ut_view_setup(); ++ raid_bdev = ut_view_raid_after_an_ejection(2); ++ writes = g_ut_sb_writes; ++ ++ ut_journal("raid1", 3, UT_INCARNATION, NULL); ++ CU_ASSERT(g_rpc_err == 1); ++ CU_ASSERT(g_rpc_err_code == -ESTALE); ++ ++ ut_journal("raid1", 5, "incarnation-elsewhere", NULL); ++ CU_ASSERT(g_rpc_err == 1); ++ CU_ASSERT(g_rpc_err_code == -ESTALE); ++ ++ ut_journal("raid1", 5, UT_INCARNATION, &g_ut_other_lineage); ++ CU_ASSERT(g_rpc_err == 1); ++ CU_ASSERT(g_rpc_err_code == -EEXIST); ++ ++ ut_journal("raid1", 0, UT_INCARNATION, NULL); ++ CU_ASSERT(g_rpc_err == 1); ++ CU_ASSERT(g_rpc_err_code == -EINVAL); ++ ++ ut_journal("raid2", 5, UT_INCARNATION, NULL); ++ CU_ASSERT(g_rpc_err == 1); ++ CU_ASSERT(g_rpc_err_code == -ENODEV); ++ ++ CU_ASSERT(raid_bdev->view_epoch == 4); ++ CU_ASSERT(g_ut_sb_writes == writes); ++ ++ ut_view_teardown(); ++} ++ ++/* Evariops 0054: a raid created outside the view protocol joins its volume's ++ * view through the journal, which names the lineage; without one it is ++ * refused. */ ++static void ++test_raid_journal_joins_a_raid_created_outside_the_view(void) ++{ ++ struct rpc_bdev_raid_create req; ++ struct raid_bdev *raid_bdev; ++ ++ ut_view_setup(); ++ g_ut_alloc_superblock = true; ++ create_raid_bdev_create_req(&req, "raid1", 0, true, 0, true); ++ req.level = RAID1; ++ req.strip_size_kb = 0; ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(3); ++ ut_fill_sb(raid_bdev, 1); ++ ++ ut_journal("raid1", 2, UT_INCARNATION, NULL); ++ CU_ASSERT(g_rpc_err == 1); ++ CU_ASSERT(g_rpc_err_code == -EINVAL); ++ CU_ASSERT(spdk_uuid_is_null(&raid_bdev->lineage)); ++ ++ ut_journal("raid1", 2, UT_INCARNATION, &g_ut_lineage); ++ CU_ASSERT(g_rpc_err == 0); ++ CU_ASSERT(spdk_uuid_compare(&raid_bdev->lineage, &g_ut_lineage) == 0); ++ CU_ASSERT(spdk_uuid_compare(&raid_bdev->sb->lineage, &g_ut_lineage) == 0); ++ CU_ASSERT(raid_bdev->sb->version.minor == 2); ++ CU_ASSERT(raid_bdev->sb->view_epoch == 2); ++ ++ ut_journal("raid1", 3, UT_INCARNATION, &g_ut_lineage); ++ CU_ASSERT(g_rpc_err == 0); ++ CU_ASSERT(raid_bdev->view_epoch == 3); ++ ++ ut_view_teardown(); ++} ++ ++/* Evariops 0054: a journal whose write fails reports it, and is retried at the ++ * same epoch. */ ++static void ++test_raid_journal_reports_a_failed_write(void) ++{ ++ uint32_t writes; ++ ++ ut_view_setup(); ++ ut_view_raid_after_an_ejection(2); ++ writes = g_ut_sb_writes; ++ ++ g_ut_sb_write_status = -EIO; ++ ut_journal("raid1", 5, UT_INCARNATION, NULL); ++ CU_ASSERT(g_rpc_err == 1); ++ CU_ASSERT(g_rpc_err_code == -EIO); ++ g_ut_sb_write_status = 0; ++ ++ ut_journal("raid1", 5, UT_INCARNATION, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ CU_ASSERT(g_ut_sb_writes == writes + 2); ++ ++ ut_view_teardown(); ++} ++ ++/* Evariops 0054: once the owner journals epoch 6 on its members, a create by ++ * another incarnation from a view no more recent is refused while those stamps ++ * stand, and one from a newer view takes the legs. */ ++static void ++test_raid_journaled_epoch_refuses_a_create_from_an_older_view(void) ++{ ++ const uint64_t refused_epochs[] = { 5, 6 }; ++ struct rpc_bdev_raid_create req; ++ struct rpc_bdev_raid_delete delete_req; ++ struct raid_bdev *raid_bdev; ++ size_t i; ++ ++ ut_view_setup(); ++ g_ut_alloc_superblock = true; ++ g_ut_sb_writes_land = true; ++ ut_view_create_req(&req, 4); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ CU_ASSERT(g_rpc_err == 0); ++ raid_bdev = ut_view_raid_online(3); ++ ut_fill_sb(raid_bdev, 2); ++ ut_journal("raid1", 6, UT_INCARNATION, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ ++ create_raid_bdev_delete_req(&delete_req, "raid1", 0); ++ rpc_bdev_raid_delete(NULL, NULL); ++ CU_ASSERT(g_rpc_err == 0); ++ ++ for (i = 0; i < SPDK_COUNTOF(refused_epochs); i++) { ++ create_test_req(&req, "raid1", 0, false, true); ++ g_rpc_err = 0; ++ g_json_decode_obj_create = 1; ++ g_json_decode_obj_err = 0; ++ g_test_multi_raids = 0; ++ req.level = RAID1; ++ req.strip_size_kb = 0; ++ spdk_uuid_copy(&req.lineage, &g_ut_lineage); ++ req.view_epoch = refused_epochs[i]; ++ free(req.incarnation); ++ req.incarnation = strdup("incarnation-elsewhere"); ++ SPDK_CU_ASSERT_FATAL(req.incarnation != NULL); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ ut_view_refused(-EBUSY); ++ } ++ ++ create_test_req(&req, "raid1", 0, false, true); ++ g_rpc_err = 0; ++ g_json_decode_obj_create = 1; ++ g_json_decode_obj_err = 0; ++ g_test_multi_raids = 0; ++ req.level = RAID1; ++ req.strip_size_kb = 0; ++ spdk_uuid_copy(&req.lineage, &g_ut_lineage); ++ req.view_epoch = 7; ++ free(req.incarnation); ++ req.incarnation = strdup("incarnation-elsewhere"); ++ SPDK_CU_ASSERT_FATAL(req.incarnation != NULL); ++ rpc_bdev_raid_create(NULL, NULL); ++ poll_app_thread(); ++ free_test_req(&req); ++ CU_ASSERT(g_rpc_err == 0); ++ ut_view_raid_online(3); ++ ++ ut_view_teardown(); ++} ++ + static int + test_new_thread_fn(struct spdk_thread *thread) + { +@@ -3248,6 +3572,11 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_raid_create_with_declared_slots_leaves_the_slots_past_its_legs_empty); + CU_ADD_TEST(suite, test_raid_add_into_a_raid_of_a_lineage_checks_the_stamp); + CU_ADD_TEST(suite, test_raid_reassembled_from_a_stamped_superblock_keeps_its_view); ++ CU_ADD_TEST(suite, test_raid_journal_stamps_the_view_epoch_on_the_members); ++ CU_ADD_TEST(suite, test_raid_journal_refuses_what_its_owner_did_not_send); ++ CU_ADD_TEST(suite, test_raid_journal_joins_a_raid_created_outside_the_view); ++ CU_ADD_TEST(suite, test_raid_journal_reports_a_failed_write); ++ CU_ADD_TEST(suite, test_raid_journaled_epoch_refuses_a_create_from_an_older_view); + + spdk_thread_lib_init(test_new_thread_fn, 0); + g_app_thread = spdk_thread_create("app_thread", NULL); diff --git a/patches/README.md b/patches/README.md index d7bec22..bf4788c 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0053`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0054`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -59,6 +59,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0053`) | 0051 | an I/O passthrough needs a connected qpair — `bdev_nvme_send_cmd` with an I/O command (NVMe reservations are I/O commands) took the controller channel's qpair without a check and handed it to the NVMe library. bdev_nvme frees that qpair from a disconnect until the reconnect (reset, failed controller, delayed reconnect), so a reservation sent to a member whose path had just gone down dereferenced NULL and killed the process, and every raid it carried with it. The command is now refused (`-ENXIO`, the channel given back), as is one for which no channel can be had. `bdev_nvme_ut` pins the disconnected and the connected case | module/bdev/nvme (nvme_rpc.c), test/unit/lib/bdev/nvme | — | | 0052 | a raid1 ejection opens its epoch on the raid's cbt — when a current member leaves a raid1 (removed, hot-removed or failed), the cbt stacked on the raid opens an epoch named after the member's UUID, the identity the control plane derives from the replica it placed there, and records from then on every write the member misses, so its return is a delta instead of a full rebuild. 0019 asked each survivor's own name for it, a bdev no cbt sits on, and no epoch ever opened. Only for a member that left current (its superblock slot configured at its peers' generation, not a write-only joiner, not a failed member an earlier removal failed to take), and only while the cbt carries no live epoch: a freeze empties the live bitmap, and an epoch opened after one may start short of writes the member missed just before it left. `bdev_raid_ut` pins the removed and the failed member, and the five ejections that open none | bdev_raid, cbt module (`vbdev_cbt_auto_epoch_open`), test/unit/lib/bdev/raid | 0016, 0019, 0033 | | 0053 | a raid create in a view arbitrates its members from their superblocks — `bdev_raid_create` with a `view_epoch` (a raid1 with a superblock, the volume's `lineage`, the creating `incarnation`) reads every listed leg's superblock, opened read-only, before it claims any, and takes only what the stamps allow. A superblock now carries a stamp (minor 2, carved out of reserved bytes): the volume's `lineage`, the `owner` incarnation and the last `view_epoch` it stamped. A leg stamped for another volume is refused (`-EEXIST`), and so is one another owner stamped in the view's epoch or a newer one (`-EBUSY`), with nothing claimed or written, unless `force_restamp` names that exact epoch (`expected_view_epoch`, `-ESTALE` otherwise). Among the legs stamped for the volume, only those at the highest content generation, configured in their own superblock, are configured, at that generation; every other leg (behind, never completed, unstamped, or stamped before minor 2) stays out, recorded MISSING at its own generation, for the control plane to rebuild, and no stamped leg holding a configured copy is `-ENODATA`. When no leg carries a stamp of the volume, every leg is configured, as before. Each member is read again as it is claimed and must not have moved (`-ESTALE`); a leg the survey cannot read refuses the create, since read as blank the freshest leg would be left out and never read again. An add into an online raid of a lineage checks the leg the same way, where upstream refuses any leg another raid stamped. `declared_slots` creates a raid wider than its legs, the slots past them empty and sized like the members. `bdev_raid_ut` and `bdev_raid_sb_ut` pin each rule, and each was seen to fail with its rule taken out | bdev_raid (raid, rpc, superblock), test/unit/lib/bdev/raid | 0014, 0016 | +| 0054 | the raid's owner journals each view epoch on its members — `bdev_raid_journal_view_epoch {name, view_epoch, expected_incarnation, lineage?}` stamps a new epoch on an online raid: the superblock's `view_epoch`, its `owner` (the raid's incarnation, which takes over the stamp of a raid it claimed after a reassembly), and the `view_epoch` of each member it serves, written to them. An owner that stays in the view journals each epoch, so a create from a view no more recent finds the legs held (0053's `-EBUSY`) for as long as the owner serves. Only the raid's incarnation journals (`-ESTALE` for another), the epoch never goes back (`-ESTALE`), and the same epoch is written again, so a journal whose write failed is retried as it was. A raid created outside the view protocol joins its volume's view when the journal names the `lineage` (`-EINVAL` without one, `-EEXIST` for another); a superblock an older image wrote is raised to minor 2, where the stamp is read. `bdev_raid_ut` pins the stamp, each refusal, the join, a failed write, and, with superblock writes landing on the legs, a create by another incarnation refused at the journaled epoch and taken at the next | bdev_raid (raid, rpc), test/unit/lib/bdev/raid | 0037, 0053 | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From 619d23f88543fb6f7f1eea6bfbdb26135fdff094 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Sat, 3 Oct 2026 14:33:49 +0200 Subject: [PATCH 24/26] fix(tier): a superblock read that fails is an error, not a disk without a superblock bdev_tier_read_sb answered valid:false both for a disk that carries no valid superblock (-EILSEQ) and for a read that failed (-EIO, -ENOMEM). The superblock module tells the two apart; the RPC folded them. A caller that reads valid:false as "no superblock" then lets an older generation win the reassembly when the read of the disk holding the newest one fails, and, when the reads fail on every band, lays a fresh layout over disks that hold data. The RPC now answers an error for a failed read, with its errno, and valid:false only for a disk without a valid superblock. The verdict is tier_sb_read_answer() in vbdev_tier.h, which test_tier_units pins: a failed read, and a read with neither a superblock nor an error, is FAILED. With a failed read mapped to "none" the test goes red (2 assertions). The tier and lvol modules build against the series. --- docs/RPC-CONTRACT.md | 2 +- module/bdev/tier/test/test_tier_units.c | 18 ++++++++++++++++++ module/bdev/tier/vbdev_tier.h | 20 ++++++++++++++++++++ module/bdev/tier/vbdev_tier_rpc.c | 12 +++++++++++- 4 files changed, 50 insertions(+), 2 deletions(-) diff --git a/docs/RPC-CONTRACT.md b/docs/RPC-CONTRACT.md index 3d48877..486e765 100644 --- a/docs/RPC-CONTRACT.md +++ b/docs/RPC-CONTRACT.md @@ -27,7 +27,7 @@ No params, read-only, idempotent. Returns `boot_id` (per-process uuid), `tier_sb | `bdev_tier_retire_band` | not an md-mirror band, else `-EBUSY` | idempotent: a re-run re-persists and re-closes | async; acks only once the superblock is durable. `rc ≠ 0` means retry | | `bdev_tier_resync_md` | target is a DEGRADED md leg; a healthy source leg exists | re-runnable; the leg stays DEGRADED on failure | copies under an md-range quiesce; acks after activate + persist | | `bdev_tier_delete` | — | `-ENODEV` if absent | unregister, then destruct | -| `bdev_tier_get_bands`, `bdev_tier_read_sb` | — | read-only | `read_sb` returns the highest-seq valid slot plus `generation_uuid` | +| `bdev_tier_get_bands`, `bdev_tier_read_sb` | — | read-only | `read_sb` returns the highest-seq valid slot plus `generation_uuid`; `valid: false` for a disk with no valid superblock, and an error (its errno) for a read that failed, which says nothing about the disk | **Assembly rules.** `bdev_tier_read_sb` exposes `version`, `seq`, `generation_uuid`, `created_epoch_sec`. Read every candidate disk's superblock, group by `generation_uuid` (this fences stale disks from a previous instance), take the highest `seq` per band, and when the two md legs disagree on `seq`, assemble the higher one ACTIVE and the other DEGRADED, then `bdev_tier_resync_md`. The fork persists DEGRADED but cannot arbitrate a split-brain across disks; that is the control-plane's. diff --git a/module/bdev/tier/test/test_tier_units.c b/module/bdev/tier/test/test_tier_units.c index c5eb121..87b9e68 100644 --- a/module/bdev/tier/test/test_tier_units.c +++ b/module/bdev/tier/test/test_tier_units.c @@ -183,6 +183,23 @@ test_identity_tuple_validate(void) CHECK(tier_unit_identity_validate(uuid, 0, 0) == -EINVAL); } +/* ---- tier_sb_read_answer: a failed read is not a disk without a superblock -- */ + +static void +test_sb_read_answer(void) +{ + struct tier_superblock sb = {0}; + + CHECK(tier_sb_read_answer(&sb, 0) == TIER_SB_READ_FOUND); + /* The disk carries no valid superblock: a fact about it. */ + CHECK(tier_sb_read_answer(NULL, -EILSEQ) == TIER_SB_READ_NONE); + /* A failed read says nothing about the disk. */ + CHECK(tier_sb_read_answer(NULL, -EIO) == TIER_SB_READ_FAILED); + CHECK(tier_sb_read_answer(NULL, -ENOMEM) == TIER_SB_READ_FAILED); + /* No superblock and no error is no answer either. */ + CHECK(tier_sb_read_answer(NULL, 0) == TIER_SB_READ_FAILED); +} + int main(void) { @@ -191,6 +208,7 @@ main(void) test_ranges_overlap(); test_identity_conflict(); test_identity_tuple_validate(); + test_sb_read_answer(); if (g_failures != 0) { fprintf(stderr, "test_tier_units: %d FAILURE(S)\n", g_failures); diff --git a/module/bdev/tier/vbdev_tier.h b/module/bdev/tier/vbdev_tier.h index 4a2f457..5bd3639 100644 --- a/module/bdev/tier/vbdev_tier.h +++ b/module/bdev/tier/vbdev_tier.h @@ -134,6 +134,26 @@ tier_sb_slot_for_seq(uint64_t seq) return (uint32_t)(seq & 1); } +/* What a superblock read says of a disk. A disk with no valid superblock + * (-EILSEQ) is a fact about the disk. A read that failed, -EIO or any other + * error, says nothing about it: taken for "no superblock", the disk carrying + * the newest one lets an older generation win the reassembly, and reads failing + * on every band lay a fresh layout over disks that hold data. */ +enum tier_sb_read_answer { + TIER_SB_READ_FOUND, + TIER_SB_READ_NONE, + TIER_SB_READ_FAILED, +}; + +static inline enum tier_sb_read_answer +tier_sb_read_answer(const struct tier_superblock *sb, int rc) +{ + if (rc == 0 && sb != NULL) { + return TIER_SB_READ_FOUND; + } + return rc == -EILSEQ ? TIER_SB_READ_NONE : TIER_SB_READ_FAILED; +} + /* Render a part_uuid as 32 lowercase hex chars into out (>= 33 bytes). An * all-zero uuid renders as the EMPTY string — the wire convention for "no * partition identity" on every RPC that emits the field. Returns out. */ diff --git a/module/bdev/tier/vbdev_tier_rpc.c b/module/bdev/tier/vbdev_tier_rpc.c index 69e3cbf..1fd6ab1 100644 --- a/module/bdev/tier/vbdev_tier_rpc.c +++ b/module/bdev/tier/vbdev_tier_rpc.c @@ -522,12 +522,22 @@ static void rpc_read_sb_done(void *cb_arg, const struct tier_superblock *sb, int rc) { struct rpc_read_sb_ctx *c = cb_arg; + enum tier_sb_read_answer answer = tier_sb_read_answer(sb, rc); struct spdk_json_write_ctx *w; uint32_t i; + if (answer == TIER_SB_READ_FAILED) { + rc = rc != 0 ? rc : -EIO; + spdk_jsonrpc_send_error_response_fmt(c->request, rc, "read_sb failed: %s", + spdk_strerror(-rc)); + spdk_bdev_close(c->desc); + free(c); + return; + } + w = spdk_jsonrpc_begin_result(c->request); spdk_json_write_object_begin(w); - if (rc != 0 || sb == NULL) { + if (answer == TIER_SB_READ_NONE) { spdk_json_write_named_bool(w, "valid", false); } else { char uuid_str[SPDK_UUID_STRING_LEN]; From 327e9f74ffe77686bc9eaf7428da85823ed0d945 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Sun, 4 Oct 2026 14:28:30 +0200 Subject: [PATCH 25/26] feat(raid): a copied member takes its source's standing (0055) A member copied from another member of its raid while no raid was assembled carries the source's superblock, where its own entry says what the source knew of it: MISSING when it was behind. The next create in a view leaves it out and rebuilds it in full, and with the source gone refuses the volume although the copy holds the content. bdev_raid_restamp_copied_member gives its own entry the source's standing: CONFIGURED at the source's content generation and view epoch. --- ...ed-member-takes-its-sources-standing.patch | 451 ++++++++++++++++++ patches/README.md | 3 +- 2 files changed, 453 insertions(+), 1 deletion(-) create mode 100644 patches/0055-raid-a-copied-member-takes-its-sources-standing.patch diff --git a/patches/0055-raid-a-copied-member-takes-its-sources-standing.patch b/patches/0055-raid-a-copied-member-takes-its-sources-standing.patch new file mode 100644 index 0000000..627b2a0 --- /dev/null +++ b/patches/0055-raid-a-copied-member-takes-its-sources-standing.patch @@ -0,0 +1,451 @@ +diff --git a/module/bdev/raid/bdev_raid.h b/module/bdev/raid/bdev_raid.h +index 5c30917..427addc 100644 +--- a/module/bdev/raid/bdev_raid.h ++++ b/module/bdev/raid/bdev_raid.h +@@ -793,6 +793,22 @@ void raid_bdev_write_superblock(struct raid_bdev *raid_bdev, raid_bdev_write_sb_ + int raid_bdev_load_base_bdev_superblock(struct spdk_bdev_desc *desc, struct spdk_io_channel *ch, + raid_bdev_load_sb_cb cb, void *cb_ctx); + ++/* Evariops 0055: a member copied from another member of its raid while no raid ++ * was assembled (a cold repair) carries the source's superblock, where its own ++ * entry says what the source knew of it: MISSING, when it was behind. It holds ++ * the source's content, so it takes the source's standing: its own entry ++ * CONFIGURED at the source's content generation and view epoch, seq_number and ++ * crc renewed. -EINVAL for a superblock without a stamp, -EEXIST for another ++ * lineage, -ENOENT when either entry is absent, -ESTALE for a source the ++ * superblock does not record CONFIGURED. 1 when the copy already has that ++ * standing (nothing changed), 0 once restamped; either way content_generation ++ * receives the copy's generation. */ ++int raid_bdev_sb_restamp_copied_member(struct raid_bdev_superblock *sb, ++ const struct spdk_uuid *lineage, ++ const struct spdk_uuid *member, ++ const struct spdk_uuid *source, ++ uint64_t *content_generation); ++ + /* Evariops 0054: stamps view epoch on an online raid's members, under the raid's + * incarnation; lineage joins a raid created outside the view protocol to its + * volume's view. cb reports the superblock write. -ESTALE for an epoch older +diff --git a/module/bdev/raid/bdev_raid_rpc.c b/module/bdev/raid/bdev_raid_rpc.c +index aada046..b67facc 100644 +--- a/module/bdev/raid/bdev_raid_rpc.c ++++ b/module/bdev/raid/bdev_raid_rpc.c +@@ -1763,3 +1763,199 @@ rpc_bdev_raid_clear_superblock(struct spdk_jsonrpc_request *request, + } + } + SPDK_RPC_REGISTER("bdev_raid_clear_superblock", rpc_bdev_raid_clear_superblock, SPDK_RPC_RUNTIME) ++ ++/* Evariops 0055: a member copied from another member of its raid, while no raid ++ * was assembled, takes the source's standing in its own superblock (see ++ * raid_bdev_sb_restamp_copied_member): without it, the next create in a view ++ * leaves the copy out and rebuilds it, and with its source gone refuses the ++ * volume (-ENODATA) although the copy holds the content. As for the clear, ++ * opening for WRITE is the gate: a member of an assembled raid, or a namespace ++ * a subsystem exposes, is claimed and cannot be restamped. */ ++struct rpc_bdev_raid_restamp_copied_member { ++ char *base_bdev; ++ struct spdk_uuid member_uuid; ++ struct spdk_uuid source_uuid; ++ struct spdk_uuid lineage; ++}; ++ ++static void ++free_rpc_bdev_raid_restamp_copied_member(struct rpc_bdev_raid_restamp_copied_member *req) ++{ ++ free(req->base_bdev); ++} ++ ++static const struct spdk_json_object_decoder rpc_bdev_raid_restamp_copied_member_decoders[] = { ++ {"base_bdev", offsetof(struct rpc_bdev_raid_restamp_copied_member, base_bdev), spdk_json_decode_string}, ++ {"member_uuid", offsetof(struct rpc_bdev_raid_restamp_copied_member, member_uuid), spdk_json_decode_uuid}, ++ {"source_uuid", offsetof(struct rpc_bdev_raid_restamp_copied_member, source_uuid), spdk_json_decode_uuid}, ++ {"lineage", offsetof(struct rpc_bdev_raid_restamp_copied_member, lineage), spdk_json_decode_uuid}, ++}; ++ ++struct rpc_restamp_ctx { ++ struct rpc_bdev_raid_restamp_copied_member req; ++ struct spdk_jsonrpc_request *request; ++ struct spdk_bdev_desc *desc; ++ struct spdk_io_channel *ch; ++ void *buf; ++ bool restamped; ++ uint64_t content_generation; ++}; ++ ++static void ++rpc_restamp_ctx_free(struct rpc_restamp_ctx *ctx) ++{ ++ spdk_dma_free(ctx->buf); ++ if (ctx->ch != NULL) { ++ spdk_put_io_channel(ctx->ch); ++ } ++ if (ctx->desc != NULL) { ++ spdk_bdev_close(ctx->desc); ++ } ++ free_rpc_bdev_raid_restamp_copied_member(&ctx->req); ++ free(ctx); ++} ++ ++static void ++rpc_restamp_complete(struct rpc_restamp_ctx *ctx, int status) ++{ ++ struct spdk_jsonrpc_request *request = ctx->request; ++ struct spdk_json_write_ctx *w; ++ ++ if (status != 0) { ++ spdk_jsonrpc_send_error_response(request, status, spdk_strerror(-status)); ++ } else { ++ w = spdk_jsonrpc_begin_result(request); ++ spdk_json_write_object_begin(w); ++ spdk_json_write_named_bool(w, "restamped", ctx->restamped); ++ spdk_json_write_named_uint64(w, "content_generation", ctx->content_generation); ++ spdk_json_write_object_end(w); ++ spdk_jsonrpc_end_result(request, w); ++ } ++ rpc_restamp_ctx_free(ctx); ++} ++ ++static void ++rpc_restamp_write_done(struct spdk_bdev_io *bdev_io, bool success, void *cb_arg) ++{ ++ struct rpc_restamp_ctx *ctx = cb_arg; ++ ++ spdk_bdev_free_io(bdev_io); ++ ctx->restamped = success; ++ rpc_restamp_complete(ctx, success ? 0 : -EIO); ++} ++ ++static void ++rpc_restamp_load_cb(const struct raid_bdev_superblock *sb, int status, void *cb_ctx) ++{ ++ struct rpc_restamp_ctx *ctx = cb_ctx; ++ struct spdk_bdev *bdev = spdk_bdev_desc_get_bdev(ctx->desc); ++ struct raid_bdev_superblock *copy; ++ uint32_t length; ++ int rc; ++ ++ if (status != 0 || sb == NULL) { ++ SPDK_ERRLOG("bdev '%s': no valid raid superblock to restamp: %s\n", ++ ctx->req.base_bdev, spdk_strerror(-status)); ++ rpc_restamp_complete(ctx, status != 0 ? status : -ENODATA); ++ return; ++ } ++ ++ /* The superblock as read is the loader's: the restamp works on a copy sized ++ * for the write, whole blocks of a bdev without interleaved metadata. */ ++ if (spdk_bdev_is_md_interleaved(bdev)) { ++ rpc_restamp_complete(ctx, -ENOTSUP); ++ return; ++ } ++ ++ length = SPDK_ALIGN_CEIL(sb->length, spdk_bdev_get_block_size(bdev)); ++ ctx->buf = spdk_dma_zmalloc(length, spdk_bdev_get_buf_align(bdev), NULL); ++ if (ctx->buf == NULL) { ++ rpc_restamp_complete(ctx, -ENOMEM); ++ return; ++ } ++ memcpy(ctx->buf, sb, sb->length); ++ copy = ctx->buf; ++ ++ rc = raid_bdev_sb_restamp_copied_member(copy, &ctx->req.lineage, &ctx->req.member_uuid, ++ &ctx->req.source_uuid, &ctx->content_generation); ++ if (rc < 0) { ++ SPDK_ERRLOG("bdev '%s': superblock of raid '%.*s' not restamped: %s\n", ++ ctx->req.base_bdev, (int)sizeof(copy->name), copy->name, spdk_strerror(-rc)); ++ rpc_restamp_complete(ctx, rc); ++ return; ++ } ++ ++ if (rc == 1) { ++ /* Already the source's standing: nothing to write. */ ++ rpc_restamp_complete(ctx, 0); ++ return; ++ } ++ ++ SPDK_NOTICELOG("bdev '%s': copied member takes its source's standing in raid '%.*s' " ++ "(content generation %" PRIu64 ", seq %" PRIu64 ")\n", ctx->req.base_bdev, ++ (int)sizeof(copy->name), copy->name, ctx->content_generation, copy->seq_number); ++ ++ rc = spdk_bdev_write(ctx->desc, ctx->ch, ctx->buf, 0, length, rpc_restamp_write_done, ctx); ++ if (rc != 0) { ++ rpc_restamp_complete(ctx, rc); ++ } ++} ++ ++static void ++rpc_restamp_event_cb(enum spdk_bdev_event_type type, struct spdk_bdev *bdev, void *event_ctx) ++{ ++ SPDK_DEBUGLOG(bdev_raid, "restamp_copied_member: unhandled bdev event %d\n", type); ++} ++ ++static void ++rpc_bdev_raid_restamp_copied_member(struct spdk_jsonrpc_request *request, ++ const struct spdk_json_val *params) ++{ ++ struct rpc_restamp_ctx *ctx; ++ char detail[256]; ++ int rc; ++ ++ ctx = calloc(1, sizeof(*ctx)); ++ if (ctx == NULL) { ++ spdk_jsonrpc_send_error_response(request, SPDK_JSONRPC_ERROR_INTERNAL_ERROR, ++ spdk_strerror(ENOMEM)); ++ return; ++ } ++ ctx->request = request; ++ ++ if (spdk_json_decode_object(params, rpc_bdev_raid_restamp_copied_member_decoders, ++ SPDK_COUNTOF(rpc_bdev_raid_restamp_copied_member_decoders), ++ &ctx->req)) { ++ spdk_jsonrpc_send_error_response(request, SPDK_JSONRPC_ERROR_INVALID_PARAMS, ++ "spdk_json_decode_object failed"); ++ free_rpc_bdev_raid_restamp_copied_member(&ctx->req); ++ free(ctx); ++ return; ++ } ++ ++ snprintf(detail, sizeof(detail), "base_bdev=%s", ctx->req.base_bdev); ++ spdk_jsonrpc_request_audit(request, "bdev_raid_restamp_copied_member", detail); ++ ++ rc = spdk_bdev_open_ext(ctx->req.base_bdev, true, rpc_restamp_event_cb, NULL, &ctx->desc); ++ if (rc != 0) { ++ SPDK_ERRLOG("bdev '%s': open for restamp failed: %s\n", ++ ctx->req.base_bdev, spdk_strerror(-rc)); ++ spdk_jsonrpc_send_error_response(request, rc, spdk_strerror(-rc)); ++ free_rpc_bdev_raid_restamp_copied_member(&ctx->req); ++ free(ctx); ++ return; ++ } ++ ++ ctx->ch = spdk_bdev_get_io_channel(ctx->desc); ++ if (ctx->ch == NULL) { ++ rpc_restamp_complete(ctx, -ENOMEM); ++ return; ++ } ++ ++ rc = raid_bdev_load_base_bdev_superblock(ctx->desc, ctx->ch, rpc_restamp_load_cb, ctx); ++ if (rc != 0) { ++ rpc_restamp_complete(ctx, rc); ++ } ++} ++SPDK_RPC_REGISTER("bdev_raid_restamp_copied_member", rpc_bdev_raid_restamp_copied_member, ++ SPDK_RPC_RUNTIME) +diff --git a/module/bdev/raid/bdev_raid_sb.c b/module/bdev/raid/bdev_raid_sb.c +index eb092d8..6465fbe 100644 +--- a/module/bdev/raid/bdev_raid_sb.c ++++ b/module/bdev/raid/bdev_raid_sb.c +@@ -156,6 +156,62 @@ raid_bdev_sb_check_crc(struct raid_bdev_superblock *sb) + return crc == prev; + } + ++/* Evariops 0055: the entry a raid of this superblock knows a member by — the ++ * first one naming it, as examine reads it. */ ++static struct raid_bdev_sb_base_bdev * ++raid_bdev_sb_entry_of(struct raid_bdev_superblock *sb, const struct spdk_uuid *uuid) ++{ ++ uint8_t i; ++ ++ for (i = 0; i < sb->base_bdevs_size; i++) { ++ if (spdk_uuid_compare(&sb->base_bdevs[i].uuid, uuid) == 0) { ++ return &sb->base_bdevs[i]; ++ } ++ } ++ ++ return NULL; ++} ++ ++int ++raid_bdev_sb_restamp_copied_member(struct raid_bdev_superblock *sb, const struct spdk_uuid *lineage, ++ const struct spdk_uuid *member, const struct spdk_uuid *source, ++ uint64_t *content_generation) ++{ ++ struct raid_bdev_sb_base_bdev *own, *from; ++ ++ if (sb->version.minor < 2 || spdk_uuid_is_null(&sb->lineage)) { ++ return -EINVAL; ++ } ++ ++ if (spdk_uuid_compare(&sb->lineage, lineage) != 0) { ++ return -EEXIST; ++ } ++ ++ own = raid_bdev_sb_entry_of(sb, member); ++ from = raid_bdev_sb_entry_of(sb, source); ++ if (own == NULL || from == NULL || own == from) { ++ return -ENOENT; ++ } ++ ++ if (from->state != RAID_SB_BASE_BDEV_CONFIGURED) { ++ return -ESTALE; ++ } ++ ++ *content_generation = from->content_generation; ++ if (own->state == RAID_SB_BASE_BDEV_CONFIGURED && ++ own->content_generation == from->content_generation) { ++ return 1; ++ } ++ ++ own->state = RAID_SB_BASE_BDEV_CONFIGURED; ++ own->content_generation = from->content_generation; ++ own->view_epoch = from->view_epoch; ++ sb->seq_number++; ++ raid_bdev_sb_update_crc(sb); ++ ++ return 0; ++} ++ + static int + raid_bdev_parse_superblock(struct raid_bdev_read_sb_ctx *ctx) + { +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index beaec9b..fc08614 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -218,6 +218,11 @@ DEFINE_STUB(spdk_bdev_notify_blockcnt_change, int, (struct spdk_bdev *bdev, uint + DEFINE_STUB(spdk_json_write_named_uuid, int, (struct spdk_json_write_ctx *w, const char *name, + const struct spdk_uuid *val), 0); + DEFINE_STUB_V(raid_bdev_init_superblock, (struct raid_bdev *raid_bdev)); ++/* Evariops 0055: the restamp of a copied member is bdev_raid_sb.c's, tested by its ++ * own suite; the RPC that calls it is not exercised here. */ ++DEFINE_STUB(raid_bdev_sb_restamp_copied_member, int, (struct raid_bdev_superblock *sb, ++ const struct spdk_uuid *lineage, const struct spdk_uuid *member, ++ const struct spdk_uuid *source, uint64_t *content_generation), -ENOTSUP); + + /* Evariops 0053: the suite's raids have no superblock buffer, except while a test + * reassembles one from a superblock. */ +diff --git a/test/unit/lib/bdev/raid/bdev_raid_sb.c/bdev_raid_sb_ut.c b/test/unit/lib/bdev/raid/bdev_raid_sb.c/bdev_raid_sb_ut.c +index 14125a4..8310a4a 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid_sb.c/bdev_raid_sb_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid_sb.c/bdev_raid_sb_ut.c +@@ -492,6 +492,123 @@ test_raid_bdev_init_superblock_outside_a_view_leaves_the_stamp_empty(void) + raid_bdev_free_superblock(&raid_bdev); + } + ++/* Evariops 0055: the superblock a cold repair copies onto a member: the ++ * source's own, stamped for the lineage at epoch 4, where the source is ++ * CONFIGURED at generation 7, the copy MISSING at 6 (it was behind) and a ++ * third member MISSING at 5. */ ++static struct raid_bdev_superblock * ++ut_copied_sb(const struct spdk_uuid *lineage, const struct spdk_uuid *source, ++ const struct spdk_uuid *copy, const struct spdk_uuid *third) ++{ ++ struct raid_bdev_superblock *sb; ++ ++ sb = calloc(1, sizeof(*sb) + 3 * sizeof(struct raid_bdev_sb_base_bdev)); ++ SPDK_CU_ASSERT_FATAL(sb != NULL); ++ memcpy(&sb->signature, RAID_BDEV_SB_SIG, sizeof(sb->signature)); ++ sb->version.major = RAID_BDEV_SB_VERSION_MAJOR; ++ sb->version.minor = 2; ++ sb->length = sizeof(*sb) + 3 * sizeof(struct raid_bdev_sb_base_bdev); ++ sb->num_base_bdevs = sb->base_bdevs_size = 3; ++ sb->seq_number = 40; ++ spdk_uuid_copy(&sb->lineage, lineage); ++ sb->view_epoch = 4; ++ ++ spdk_uuid_copy(&sb->base_bdevs[0].uuid, source); ++ sb->base_bdevs[0].state = RAID_SB_BASE_BDEV_CONFIGURED; ++ sb->base_bdevs[0].content_generation = 7; ++ sb->base_bdevs[0].view_epoch = 4; ++ sb->base_bdevs[0].slot = 0; ++ ++ spdk_uuid_copy(&sb->base_bdevs[1].uuid, copy); ++ sb->base_bdevs[1].state = RAID_SB_BASE_BDEV_MISSING; ++ sb->base_bdevs[1].content_generation = 6; ++ sb->base_bdevs[1].view_epoch = 3; ++ sb->base_bdevs[1].slot = 1; ++ ++ spdk_uuid_copy(&sb->base_bdevs[2].uuid, third); ++ sb->base_bdevs[2].state = RAID_SB_BASE_BDEV_MISSING; ++ sb->base_bdevs[2].content_generation = 5; ++ sb->base_bdevs[2].view_epoch = 2; ++ sb->base_bdevs[2].slot = 2; ++ ++ raid_bdev_sb_update_crc(sb); ++ return sb; ++} ++ ++/* Evariops 0055: the copy takes its source's standing — its own entry CONFIGURED ++ * at the source's generation and epoch, the sequence number moved on and the ++ * crc renewed — and only its own: the third member stays as it was. Done once, ++ * it is a no-op that writes nothing. */ ++static void ++test_raid_bdev_sb_restamp_copied_member_takes_the_sources_standing(void) ++{ ++ struct spdk_uuid lineage, source, copy, third; ++ struct raid_bdev_superblock *sb; ++ uint64_t generation = 0; ++ ++ spdk_uuid_generate(&lineage); ++ spdk_uuid_generate(&source); ++ spdk_uuid_generate(©); ++ spdk_uuid_generate(&third); ++ sb = ut_copied_sb(&lineage, &source, ©, &third); ++ ++ CU_ASSERT(raid_bdev_sb_restamp_copied_member(sb, &lineage, ©, &source, &generation) == 0); ++ CU_ASSERT(generation == 7); ++ CU_ASSERT(sb->base_bdevs[1].state == RAID_SB_BASE_BDEV_CONFIGURED); ++ CU_ASSERT(sb->base_bdevs[1].content_generation == 7); ++ CU_ASSERT(sb->base_bdevs[1].view_epoch == 4); ++ CU_ASSERT(sb->base_bdevs[2].state == RAID_SB_BASE_BDEV_MISSING); ++ CU_ASSERT(sb->base_bdevs[2].content_generation == 5); ++ CU_ASSERT(sb->seq_number == 41); ++ CU_ASSERT(raid_bdev_sb_check_crc(sb)); ++ ++ generation = 0; ++ CU_ASSERT(raid_bdev_sb_restamp_copied_member(sb, &lineage, ©, &source, &generation) == 1); ++ CU_ASSERT(generation == 7); ++ CU_ASSERT(sb->seq_number == 41); ++ ++ free(sb); ++} ++ ++/* Evariops 0055: each refusal leaves the superblock as it was: no stamp, another ++ * volume's, an entry missing or the same one twice, a source not CONFIGURED. */ ++static void ++test_raid_bdev_sb_restamp_copied_member_refusals(void) ++{ ++ struct spdk_uuid lineage, other, source, copy, third, stranger; ++ struct raid_bdev_superblock *sb; ++ uint64_t generation = 0; ++ uint32_t crc; ++ ++ spdk_uuid_generate(&lineage); ++ spdk_uuid_generate(&other); ++ spdk_uuid_generate(&source); ++ spdk_uuid_generate(©); ++ spdk_uuid_generate(&third); ++ spdk_uuid_generate(&stranger); ++ sb = ut_copied_sb(&lineage, &source, ©, &third); ++ crc = sb->crc; ++ ++ CU_ASSERT(raid_bdev_sb_restamp_copied_member(sb, &other, ©, &source, &generation) == -EEXIST); ++ CU_ASSERT(raid_bdev_sb_restamp_copied_member(sb, &lineage, &stranger, &source, &generation) == -ENOENT); ++ CU_ASSERT(raid_bdev_sb_restamp_copied_member(sb, &lineage, ©, &stranger, &generation) == -ENOENT); ++ CU_ASSERT(raid_bdev_sb_restamp_copied_member(sb, &lineage, ©, ©, &generation) == -ENOENT); ++ /* The third member was behind too: copied from it, the copy would be behind. */ ++ CU_ASSERT(raid_bdev_sb_restamp_copied_member(sb, &lineage, ©, &third, &generation) == -ESTALE); ++ CU_ASSERT(sb->base_bdevs[1].state == RAID_SB_BASE_BDEV_MISSING); ++ CU_ASSERT(sb->seq_number == 40); ++ CU_ASSERT(sb->crc == crc); ++ ++ sb->version.minor = 1; ++ CU_ASSERT(raid_bdev_sb_restamp_copied_member(sb, &lineage, ©, &source, &generation) == -EINVAL); ++ sb->version.minor = 2; ++ spdk_uuid_set_null(&sb->lineage); ++ CU_ASSERT(raid_bdev_sb_restamp_copied_member(sb, &lineage, ©, &source, &generation) == -EINVAL); ++ CU_ASSERT(sb->base_bdevs[1].state == RAID_SB_BASE_BDEV_MISSING); ++ ++ free(sb); ++} ++ + int + main(int argc, char **argv) + { +@@ -502,6 +619,8 @@ main(int argc, char **argv) + { "test_raid_bdev_parse_superblock", test_raid_bdev_parse_superblock }, + { "test_raid_bdev_init_superblock_stamps_the_view", test_raid_bdev_init_superblock_stamps_the_view }, + { "test_raid_bdev_init_superblock_outside_a_view_leaves_the_stamp_empty", test_raid_bdev_init_superblock_outside_a_view_leaves_the_stamp_empty }, ++ { "test_raid_bdev_sb_restamp_copied_member_takes_the_sources_standing", test_raid_bdev_sb_restamp_copied_member_takes_the_sources_standing }, ++ { "test_raid_bdev_sb_restamp_copied_member_refusals", test_raid_bdev_sb_restamp_copied_member_refusals }, + CU_TEST_INFO_NULL, + }; + CU_SuiteInfo suites[] = { diff --git a/patches/README.md b/patches/README.md index bf4788c..414e964 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0054`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0055`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -60,6 +60,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0054`) | 0052 | a raid1 ejection opens its epoch on the raid's cbt — when a current member leaves a raid1 (removed, hot-removed or failed), the cbt stacked on the raid opens an epoch named after the member's UUID, the identity the control plane derives from the replica it placed there, and records from then on every write the member misses, so its return is a delta instead of a full rebuild. 0019 asked each survivor's own name for it, a bdev no cbt sits on, and no epoch ever opened. Only for a member that left current (its superblock slot configured at its peers' generation, not a write-only joiner, not a failed member an earlier removal failed to take), and only while the cbt carries no live epoch: a freeze empties the live bitmap, and an epoch opened after one may start short of writes the member missed just before it left. `bdev_raid_ut` pins the removed and the failed member, and the five ejections that open none | bdev_raid, cbt module (`vbdev_cbt_auto_epoch_open`), test/unit/lib/bdev/raid | 0016, 0019, 0033 | | 0053 | a raid create in a view arbitrates its members from their superblocks — `bdev_raid_create` with a `view_epoch` (a raid1 with a superblock, the volume's `lineage`, the creating `incarnation`) reads every listed leg's superblock, opened read-only, before it claims any, and takes only what the stamps allow. A superblock now carries a stamp (minor 2, carved out of reserved bytes): the volume's `lineage`, the `owner` incarnation and the last `view_epoch` it stamped. A leg stamped for another volume is refused (`-EEXIST`), and so is one another owner stamped in the view's epoch or a newer one (`-EBUSY`), with nothing claimed or written, unless `force_restamp` names that exact epoch (`expected_view_epoch`, `-ESTALE` otherwise). Among the legs stamped for the volume, only those at the highest content generation, configured in their own superblock, are configured, at that generation; every other leg (behind, never completed, unstamped, or stamped before minor 2) stays out, recorded MISSING at its own generation, for the control plane to rebuild, and no stamped leg holding a configured copy is `-ENODATA`. When no leg carries a stamp of the volume, every leg is configured, as before. Each member is read again as it is claimed and must not have moved (`-ESTALE`); a leg the survey cannot read refuses the create, since read as blank the freshest leg would be left out and never read again. An add into an online raid of a lineage checks the leg the same way, where upstream refuses any leg another raid stamped. `declared_slots` creates a raid wider than its legs, the slots past them empty and sized like the members. `bdev_raid_ut` and `bdev_raid_sb_ut` pin each rule, and each was seen to fail with its rule taken out | bdev_raid (raid, rpc, superblock), test/unit/lib/bdev/raid | 0014, 0016 | | 0054 | the raid's owner journals each view epoch on its members — `bdev_raid_journal_view_epoch {name, view_epoch, expected_incarnation, lineage?}` stamps a new epoch on an online raid: the superblock's `view_epoch`, its `owner` (the raid's incarnation, which takes over the stamp of a raid it claimed after a reassembly), and the `view_epoch` of each member it serves, written to them. An owner that stays in the view journals each epoch, so a create from a view no more recent finds the legs held (0053's `-EBUSY`) for as long as the owner serves. Only the raid's incarnation journals (`-ESTALE` for another), the epoch never goes back (`-ESTALE`), and the same epoch is written again, so a journal whose write failed is retried as it was. A raid created outside the view protocol joins its volume's view when the journal names the `lineage` (`-EINVAL` without one, `-EEXIST` for another); a superblock an older image wrote is raised to minor 2, where the stamp is read. `bdev_raid_ut` pins the stamp, each refusal, the join, a failed write, and, with superblock writes landing on the legs, a create by another incarnation refused at the journaled epoch and taken at the next | bdev_raid (raid, rpc), test/unit/lib/bdev/raid | 0037, 0053 | +| 0055 | a copied member takes its source's standing — `bdev_raid_restamp_copied_member {base_bdev, member_uuid, source_uuid, lineage}` restamps the superblock of a member copied from another member of its raid while no raid was assembled (a cold repair). The copy carries the source's superblock, where its own entry says what the source knew of it: MISSING, when it was behind. A create in a view (0053) would then leave it out and rebuild it in full and, with the source gone, refuse the volume (`-ENODATA`) although the copy holds the content. Its own entry takes the source's standing instead: CONFIGURED at the source's content generation and view epoch, with `seq_number` and crc renewed, written back at the head of the bdev; no other entry and no stamp field changes. Refused with nothing written for a superblock without a stamp (`-EINVAL`), of another lineage (`-EEXIST`), without an entry for either member or naming the same one twice (`-ENOENT`), whose source entry is not CONFIGURED (`-ESTALE`), or on a bdev with interleaved metadata (`-ENOTSUP`). A copy that already has that standing is not written again (`restamped: false`); the response carries the copy's `content_generation`. Opening for write is the gate, as for `bdev_raid_clear_superblock`: a member of an assembled raid, or a namespace a subsystem exposes, is claimed and cannot be restamped. `bdev_raid_sb_ut` pins the restamp and each refusal, and was seen to fail with the restamp taken out | bdev_raid (rpc, superblock), test/unit/lib/bdev/raid | 0016, 0053 | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters). From dc27c8e0e89c17004a326883cbd65d4cb08aa4a9 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Rapha=C3=ABl=20DUCOM?= Date: Sun, 4 Oct 2026 15:42:59 +0200 Subject: [PATCH 26/26] feat(raid): discovery leaves a stamped member to a create in view (0056) Examine assembled a raid of a lineage from whichever member it read first: the raid came online on that member's record of the others, wrote its superblock and rebuilt the others in full before the control plane's fence, and the members' own standings were never arbitrated. A member stamped for a lineage now stays unclaimed until a create in its view assembles it; an unstamped one is reassembled as upstream does. --- ...a-stamped-member-to-a-create-in-view.patch | 89 +++++++++++++++++++ patches/README.md | 3 +- 2 files changed, 91 insertions(+), 1 deletion(-) create mode 100644 patches/0056-raid-discovery-leaves-a-stamped-member-to-a-create-in-view.patch diff --git a/patches/0056-raid-discovery-leaves-a-stamped-member-to-a-create-in-view.patch b/patches/0056-raid-discovery-leaves-a-stamped-member-to-a-create-in-view.patch new file mode 100644 index 0000000..dcc3276 --- /dev/null +++ b/patches/0056-raid-discovery-leaves-a-stamped-member-to-a-create-in-view.patch @@ -0,0 +1,89 @@ +diff --git a/module/bdev/raid/bdev_raid.c b/module/bdev/raid/bdev_raid.c +index 0bd65e1..fbb86cb 100644 +--- a/module/bdev/raid/bdev_raid.c ++++ b/module/bdev/raid/bdev_raid.c +@@ -5982,6 +5982,19 @@ raid_bdev_examine_cont(struct spdk_bdev *bdev, const struct raid_bdev_superblock + case 0: + /* valid superblock found */ + SPDK_DEBUGLOG(bdev_raid, "raid superblock found on bdev %s\n", bdev->name); ++ /* Evariops 0056: a member stamped for a lineage is assembled by a create in its ++ * view (0053), which the control plane issues once its fence has landed, and ++ * which arbitrates the members from their own standings. Discovery assembled ++ * it from whichever member it read first: the raid came online on that ++ * member's view of the others, wrote its superblock and rebuilt the others in ++ * full before any fence, and a member whose own standing was current was ++ * rebuilt all the same. Discovery leaves it alone, unclaimed. */ ++ if (raid_bdev_sb_stamped(sb)) { ++ SPDK_NOTICELOG("bdev %s: member of raid %.*s stamped for a lineage, left to a " ++ "create in its view (Evariops 0056)\n", bdev->name, ++ (int)sizeof(sb->name), sb->name); ++ break; ++ } + raid_bdev_examine_sb(sb, bdev, raid_bdev_examine_done, bdev); + return; + case -EINVAL: +diff --git a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +index fc08614..e7b9d33 100644 +--- a/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c ++++ b/test/unit/lib/bdev/raid/bdev_raid.c/bdev_raid_ut.c +@@ -3220,6 +3220,52 @@ test_raid_reassembled_from_a_stamped_superblock_keeps_its_view(void) + ut_view_teardown(); + } + ++/* Evariops 0056: discovery leaves a member stamped for a lineage alone, unclaimed: ++ * a create in its view assembles it, after the control plane's fence. A member ++ * whose superblock carries no stamp is reassembled as upstream does. */ ++static void ++test_raid_discovery_leaves_a_stamped_member_to_a_create_in_view(void) ++{ ++ struct raid_bdev_superblock *sb; ++ struct spdk_bdev *leg; ++ struct raid_bdev *raid_bdev; ++ ++ ut_view_setup(); ++ g_ut_alloc_superblock = true; ++ create_base_bdevs(0); ++ leg = spdk_bdev_get_by_name("Nvme0n1"); ++ SPDK_CU_ASSERT_FATAL(leg != NULL); ++ ++ sb = ut_sb(&g_ut_lineage, "incarnation-before", 4); ++ ut_sb_slot(sb, "Nvme0n1", RAID_SB_BASE_BDEV_CONFIGURED, 2); ++ ut_sb_slot(sb, "Nvme1n1", RAID_SB_BASE_BDEV_MISSING, 1); ++ ut_leg("Nvme0n1")->sb = sb; ++ ++ raid_bdev_examine(leg); ++ poll_app_thread(); ++ ++ CU_ASSERT(ut_leg("Nvme0n1")->reads == 1); ++ CU_ASSERT(raid_bdev_find_by_uuid(&sb->uuid) == NULL); ++ CU_ASSERT(leg->internal.claim_type == SPDK_BDEV_CLAIM_NONE); ++ ++ /* The control: the same member, its superblock below minor 2, has discovery create ++ * the raid it belongs to (the suite's slots carry no data range, so the member ++ * itself is not configured in it). */ ++ sb->version.minor = 1; ++ raid_bdev_examine(leg); ++ poll_app_thread(); ++ ++ CU_ASSERT(ut_leg("Nvme0n1")->reads == 2); ++ raid_bdev = raid_bdev_find_by_uuid(&sb->uuid); ++ CU_ASSERT(raid_bdev != NULL); ++ if (raid_bdev != NULL) { ++ raid_bdev_delete(raid_bdev, NULL, NULL); ++ poll_app_thread(); ++ } ++ ++ ut_view_teardown(); ++} ++ + /* Evariops 0054: fills the superblock the suite's stubbed init leaves empty, as + * init writes it, at minor: each slot under its member, CONFIGURED or MISSING, + * at the generation and epoch the raid holds for it. */ +@@ -3577,6 +3623,7 @@ main(int argc, char **argv) + CU_ADD_TEST(suite, test_raid_create_with_declared_slots_leaves_the_slots_past_its_legs_empty); + CU_ADD_TEST(suite, test_raid_add_into_a_raid_of_a_lineage_checks_the_stamp); + CU_ADD_TEST(suite, test_raid_reassembled_from_a_stamped_superblock_keeps_its_view); ++ CU_ADD_TEST(suite, test_raid_discovery_leaves_a_stamped_member_to_a_create_in_view); + CU_ADD_TEST(suite, test_raid_journal_stamps_the_view_epoch_on_the_members); + CU_ADD_TEST(suite, test_raid_journal_refuses_what_its_owner_did_not_send); + CU_ADD_TEST(suite, test_raid_journal_joins_a_raid_created_outside_the_view); diff --git a/patches/README.md b/patches/README.md index 414e964..2a44ec7 100644 --- a/patches/README.md +++ b/patches/README.md @@ -4,7 +4,7 @@ Out-of-tree patches applied on top of upstream SPDK during the container build ( ## Application order -Patches are applied in **lexicographic order of filename** (`0001` … `0055`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: +Patches are applied in **lexicographic order of filename** (`0001` … `0056`) — the Dockerfile globs `patches/*.patch` and `git apply`s each. The numeric prefix IS the contract; do not rely on any other ordering. Order matters: | # | Patch | Touches | Depends on | |--:|:------|:--------|:-----------| @@ -61,6 +61,7 @@ Patches are applied in **lexicographic order of filename** (`0001` … `0055`) | 0053 | a raid create in a view arbitrates its members from their superblocks — `bdev_raid_create` with a `view_epoch` (a raid1 with a superblock, the volume's `lineage`, the creating `incarnation`) reads every listed leg's superblock, opened read-only, before it claims any, and takes only what the stamps allow. A superblock now carries a stamp (minor 2, carved out of reserved bytes): the volume's `lineage`, the `owner` incarnation and the last `view_epoch` it stamped. A leg stamped for another volume is refused (`-EEXIST`), and so is one another owner stamped in the view's epoch or a newer one (`-EBUSY`), with nothing claimed or written, unless `force_restamp` names that exact epoch (`expected_view_epoch`, `-ESTALE` otherwise). Among the legs stamped for the volume, only those at the highest content generation, configured in their own superblock, are configured, at that generation; every other leg (behind, never completed, unstamped, or stamped before minor 2) stays out, recorded MISSING at its own generation, for the control plane to rebuild, and no stamped leg holding a configured copy is `-ENODATA`. When no leg carries a stamp of the volume, every leg is configured, as before. Each member is read again as it is claimed and must not have moved (`-ESTALE`); a leg the survey cannot read refuses the create, since read as blank the freshest leg would be left out and never read again. An add into an online raid of a lineage checks the leg the same way, where upstream refuses any leg another raid stamped. `declared_slots` creates a raid wider than its legs, the slots past them empty and sized like the members. `bdev_raid_ut` and `bdev_raid_sb_ut` pin each rule, and each was seen to fail with its rule taken out | bdev_raid (raid, rpc, superblock), test/unit/lib/bdev/raid | 0014, 0016 | | 0054 | the raid's owner journals each view epoch on its members — `bdev_raid_journal_view_epoch {name, view_epoch, expected_incarnation, lineage?}` stamps a new epoch on an online raid: the superblock's `view_epoch`, its `owner` (the raid's incarnation, which takes over the stamp of a raid it claimed after a reassembly), and the `view_epoch` of each member it serves, written to them. An owner that stays in the view journals each epoch, so a create from a view no more recent finds the legs held (0053's `-EBUSY`) for as long as the owner serves. Only the raid's incarnation journals (`-ESTALE` for another), the epoch never goes back (`-ESTALE`), and the same epoch is written again, so a journal whose write failed is retried as it was. A raid created outside the view protocol joins its volume's view when the journal names the `lineage` (`-EINVAL` without one, `-EEXIST` for another); a superblock an older image wrote is raised to minor 2, where the stamp is read. `bdev_raid_ut` pins the stamp, each refusal, the join, a failed write, and, with superblock writes landing on the legs, a create by another incarnation refused at the journaled epoch and taken at the next | bdev_raid (raid, rpc), test/unit/lib/bdev/raid | 0037, 0053 | | 0055 | a copied member takes its source's standing — `bdev_raid_restamp_copied_member {base_bdev, member_uuid, source_uuid, lineage}` restamps the superblock of a member copied from another member of its raid while no raid was assembled (a cold repair). The copy carries the source's superblock, where its own entry says what the source knew of it: MISSING, when it was behind. A create in a view (0053) would then leave it out and rebuild it in full and, with the source gone, refuse the volume (`-ENODATA`) although the copy holds the content. Its own entry takes the source's standing instead: CONFIGURED at the source's content generation and view epoch, with `seq_number` and crc renewed, written back at the head of the bdev; no other entry and no stamp field changes. Refused with nothing written for a superblock without a stamp (`-EINVAL`), of another lineage (`-EEXIST`), without an entry for either member or naming the same one twice (`-ENOENT`), whose source entry is not CONFIGURED (`-ESTALE`), or on a bdev with interleaved metadata (`-ENOTSUP`). A copy that already has that standing is not written again (`restamped: false`); the response carries the copy's `content_generation`. Opening for write is the gate, as for `bdev_raid_clear_superblock`: a member of an assembled raid, or a namespace a subsystem exposes, is claimed and cannot be restamped. `bdev_raid_sb_ut` pins the restamp and each refusal, and was seen to fail with the restamp taken out | bdev_raid (rpc, superblock), test/unit/lib/bdev/raid | 0016, 0053 | +| 0056 | discovery leaves a stamped member to a create in view — examine no longer assembles a raid from a member whose superblock is stamped for a lineage (minor 2 with a lineage, 0053): the member stays unclaimed, and a create in its view assembles it. Discovery assembled it from whichever member it read first: the raid came online on that member's record of the others, wrote its superblock and rebuilt the others in full, all before the control plane's fence, and the arbitration of the members' own standings (0053, 0055) never ran — a member whose own standing was current was rebuilt all the same. A member without a stamp is still reassembled as upstream does, and an add the control plane issues still reads the member's superblock (0053). `bdev_raid_ut` pins a stamped member left unclaimed with no raid created, and, as the control, the same member below minor 2 reassembled; it was seen to fail with the check taken out | bdev_raid | 0053 | 0005 `#include`s `vbdev_tier.h` and adds `-I module/bdev/tier` to the lvol module CFLAGS via its own Makefile hunk; the Dockerfile injects the module dirs before applying patches (copy-before-apply ordering matters).