Skip to content

test: collection-upgrade harness - #250

Merged
gene-redpanda merged 9 commits into
mainfrom
devprod-4109-upgrade-test
Jul 24, 2026
Merged

gene-redpanda merged 9 commits into
mainfrom
devprod-4109-upgrade-test

Conversation

@gene-redpanda

Copy link
Copy Markdown
Member

Turns every Redpanda CI lane into a collection upgrade test: build a cluster with the last released redpanda.cluster collection at a pinned RP version, seed a known dataset, then re-converge the same cluster with the candidate collection at the SAME version and assert a safe, lossless re-converge. Redpanda is pinned across both phases, so the collection is the only variable — a red lane means "this collection change breaks re-convergence of an existing cluster."

What each lane does

prereqs → override→BASELINE → provision @ pinned RP (present) → capture + seed
        → install CANDIDATE → re-provision @ same RP → deploy monitor/console
        → profile tests → upgrade:assert (repo/key cutover, version unchanged,
          service active, N/N records survived) → destroy

Knobs

  • BASELINE_COLLECTION_REF — the "last released" collection (Galaxy latest default; override to a prior release).
  • CANDIDATE_COLLECTION_REF — the collection under test: a git branch (git+<url>,<branch>), a released Galaxy version (redpanda.cluster:0.12.0), or unset → requirements.yml.
  • REDPANDA_VERSION — pinned across both phases (single pin knob).

Replaces the earlier standalone ci:*:rp:upgrade lanes; unstable lanes stay single-phase. Full model in docs/COLLECTION_UPGRADE_TEST.md.

Why this exists / what it caught

Built to validate the Redpanda Artifact Registry migration. It found real upgrade/fresh-install NO_PUBKEY failures in the collection's apt key handling (fixed in redpanda-ansible-collection#150) and an invalid-JSON tiered config template (redpanda-ansible-collection#149).

Validation status

  • Basic lane (ci:gcp:rp): green live on GCP — full re-converge + 500-record survival + version-stable assertions.
  • Tiered / connect lanes: code-complete + dry-validated. Tiered end-to-end is blocked until a released collection carries the tiered_storage.j2 JSON fix (Fixed a couple of errors in readme comparing with vars.tf #149); connect needs CONNECT_RPM_TOKEN.

🤖 Generated with Claude Code

…> candidate

Reworks the Redpanda CI lanes into collection-upgrade tests: each lane builds a
cluster with the last released redpanda.cluster collection at a pinned RP version,
seeds a known dataset, then re-converges the same cluster with the candidate
collection at the SAME version and asserts a safe re-converge (repo/key cutover,
version unchanged, service active) + full data survival. Pinning RP isolates the
collection as the only variable.

- cluster:provision / cluster:tiered take RP_VERSION + RP_INSTALL_STATUS and no
  longer install collections themselves (the workflow controls the active one).
- test:upgrade:{capture:baseline,seed,assert} + scripts/upgrade/*.sh do the
  seed/verify + per-node repo/key/version/service assertions (TLS via RPK_EXTRA).
- baseline via BASELINE_COLLECTION_REF (Galaxy latest default); candidate via
  CANDIDATE_COLLECTION_REF (git branch or Galaxy version; default requirements.yml).
- Removes the standalone ci:*:rp:upgrade composites; docs/COLLECTION_UPGRADE_TEST.md.

Validated live on GCP (basic lane). Tiered/connect lanes are code-complete +
dry-validated; tiered end-to-end needs a released collection with the
tiered_storage.j2 JSON fix.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@gene-redpanda gene-redpanda changed the title test: collection-upgrade harness (every lane upgrades last-release → candidate) test: collection-upgrade harness Jul 1, 2026
@gene-redpanda gene-redpanda added the ci-ready indicates that a PR is ready for builds to run label Jul 1, 2026
Wiring — the knobs must actually reach the containers:
- forward BASELINE_COLLECTION_REF / CANDIDATE_COLLECTION_REF /
  REDPANDA_VERSION through the docker plugin env allowlist (build 425's
  candidate phase silently fell back to Galaxy latest without this)
- hoist the shared lane vars to taskfile level (GA pin bump = one edit)

Unstable lanes — single-phase ci:aws:rp:tiered:unstable at latest: the GA
pin only exists in the stable repo (build 426), and the lane owns its
IS_USING_UNSTABLE identity so local runs test the unstable repo too.

Fail loud, not silent:
- non-clobbering :ansible:prereqs (marker vs requirements.yml mtime), so
  phase-2 deploys can't revert the candidate collection; provision/tiered
  regain the dep and work standalone again
- timeout on the rpk consume data-survival checks ('-n N' blocks forever
  on partial loss — the exact case they exist to catch)
- hard-fail on a missing baseline version file; baseline moved to /var/tmp
  (Fedora /tmp is tmpfs)
- sentinel topic create failures surface (describe fallback replaces
  '|| true'); version/status extra-vars only passed when explicitly set

Fail fast:
- ConnectionAttempts 100 -> 30 (a dead host said nothing for ~50 min)
- timeout_in_minutes on every step
- ts-connect builds the extra cluster at point of use, not 30-60 idle
  minutes early; deferred extra destroy is a no-op when nothing was built

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@gene-redpanda
gene-redpanda force-pushed the devprod-4109-upgrade-test branch from 82f9813 to b0c02cb Compare July 13, 2026 21:57
Comment thread .buildkite/pipeline.yml Outdated
gene-redpanda and others added 2 commits July 17, 2026 13:59
Build 430: three builds queued from successive pushes interleaved their
lanes (per-lane concurrency groups serialize lanes, not builds). An earlier
build's cleanup-aws job then mass-terminated every ci-* instance in the
account at 22:09:49Z — including 23 instances belonging to build 430's
still-running lanes. Established SSH sessions froze ('privilege escalation
timeout') and freshly-launched nodes died before answering SSH
(UNREACHABLE). CloudTrail confirms: one aws-sdk-go-v2 TerminateInstances
call with all 5 of aws-fedora-tiered's instance IDs in it.

The taskfile already plumbed a MIN_AGE knob, but neither Go program
implemented --min-age (passing it would have crashed on an unknown flag).

- AWS cleanup: -min-age duration flag; age-gate instances (LaunchTime),
  volumes (CreateTime), key pairs (CreateTime), IAM roles/profiles
  (CreateDate), S3 buckets (CreationDate). Security groups carry no
  timestamp, so gate them on in-use instead: an SG attached to any ENI is
  skipped entirely — previously the code revoked all rules from a live
  lane's SG even when the delete itself failed.
- GCP cleanup: same flag; age-gate instances, instance groups, firewalls,
  addresses, subnets, networks (CreationTimestamp) and buckets
  (TimeCreated). Service accounts expose no creation time and are skipped
  under min-age rather than risking a live lane's credentials.
- cleanup:{aws,gcp}:ci default MIN_AGE=3h (> the 2h lane timeout, so
  nothing a live lane owns can be reaped). MIN_AGE=0 restores a full sweep;
  manual cleanup:aws / cleanup:gcp are unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…anda minimum)

us-west2-b has been out of n2-standard-2 stock for 2+ weeks (builds
425-431). Zone discovery filters on zone status (UP), not capacity, and the
module round-robins nodes across all discovered zones, so every 3+ node
cluster placed at least one node in the dead zone and failed at terraform
apply. Quota is not the issue (N2_CPUS 0/750 used) — it is Google-side
zonal stock exhaustion.

n2d-standard-4 because the replacement family must satisfy all three:
- local-SSD capable: the module attaches local-ssd scratch disks, which
  rules out E2 (400: '[e2-standard-2, local-ssd] features are not
  compatible', build 432)
- separate stock pool: N2D is AMD silicon, drawing from a different pool
  than the exhausted Intel N2; offered in all three us-west2 zones
- Redpanda-compliant shape: the production requirements specify a minimum
  of two physical cores per node; standard-2 is 1 physical core with SMT,
  so standard-4 is the smallest compliant shape

Override GCP_INSTANCE_TYPE to pin a different type per build.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@gene-redpanda
gene-redpanda force-pushed the devprod-4109-upgrade-test branch 2 times, most recently from c52ded7 to 3c9e21f Compare July 17, 2026 20:48
gene-redpanda and others added 4 commits July 23, 2026 12:26
…otation

Collection-upgrade validation: assert-node.sh signed-metadata/key checks (the RPM
metadata-verification assert gated on the dnf backend), the sabotage-deb-keyring
fixture (retaining the baseline key across rotation), and the ci/testing task wiring.
…E_COLLECTION_REF

Previously CANDIDATE_COLLECTION_REF only took effect in the upgrade/unstable CI
lanes (which call :ansible:collection:candidate explicitly). Every other provision
workflow installed redpanda.cluster from requirements.yml (Galaxy) only, so testing
a branch there meant hand-editing requirements.yml.

Add :ansible:collection:pin, run as part of :ansible:prereqs, which force-installs
redpanda.cluster from CANDIDATE_COLLECTION_REF on top of the requirements.yml install
whenever the env var is set (git branch or Galaxy version); no-op when unset. Now any
workflow honors the one env var with no file fiddling.

run: once keeps the repeated prereqs deps from re-installing and clobbering the
upgrade test's phase-1 baseline; the upgrade lanes' explicit override/candidate steps
run afterward and still win. The pin is intentionally outside the requirements mtime
status marker so a warm collections cache does not skip it. The env var is already
forwarded through the Buildkite docker steps, so no pipeline plumbing changes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Add a terse, developer-oriented CLAUDE.md (Task interface, ci:* lanes,
  CANDIDATE_COLLECTION_REF, upgrade-test model, lint).
- Fix README: broken deploy-client.yml ref, Basic Usage shell bugs, add a short
  Task section, point branch-testing at CANDIDATE_COLLECTION_REF, link docs/.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…aform pin

- Remove the unused, drifted Makefile (CI uses Task; its ci-* targets diverged
  from the task ci:* lanes).
- Remove Dockerfile_FEDORA (EOL base) and .pre-commit-config.yaml (CI enforces
  ansible-lint + terraform fmt via GitHub Actions); update README/CLAUDE.md.
- Bump Dockerfile_UBUNTU terraform 1.4.5 -> 1.7.4 to match Taskfile.yml.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@gene-redpanda
gene-redpanda force-pushed the devprod-4109-upgrade-test branch 2 times, most recently from 1f00b7c to a6ddc78 Compare July 23, 2026 20:36
Add terraform validate (dflook/terraform-validate) alongside the existing fmt
check, scoped to aws/, so it runs on every PR. Remove the now-unused .tflint.hcl
(pre-commit, its only consumer, is gone).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@gene-redpanda
gene-redpanda force-pushed the devprod-4109-upgrade-test branch from a6ddc78 to b8c3e2b Compare July 24, 2026 01:52
@gene-redpanda
gene-redpanda merged commit 63b71dc into main Jul 24, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-ready indicates that a PR is ready for builds to run

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants