Repository navigation
test: collection-upgrade harness - #250
Merged
Merged
Conversation
…> candidate
Reworks the Redpanda CI lanes into collection-upgrade tests: each lane builds a
cluster with the last released redpanda.cluster collection at a pinned RP version,
seeds a known dataset, then re-converges the same cluster with the candidate
collection at the SAME version and asserts a safe re-converge (repo/key cutover,
version unchanged, service active) + full data survival. Pinning RP isolates the
collection as the only variable.
- cluster:provision / cluster:tiered take RP_VERSION + RP_INSTALL_STATUS and no
longer install collections themselves (the workflow controls the active one).
- test:upgrade:{capture:baseline,seed,assert} + scripts/upgrade/*.sh do the
seed/verify + per-node repo/key/version/service assertions (TLS via RPK_EXTRA).
- baseline via BASELINE_COLLECTION_REF (Galaxy latest default); candidate via
CANDIDATE_COLLECTION_REF (git branch or Galaxy version; default requirements.yml).
- Removes the standalone ci:*:rp:upgrade composites; docs/COLLECTION_UPGRADE_TEST.md.
Validated live on GCP (basic lane). Tiered/connect lanes are code-complete +
dry-validated; tiered end-to-end needs a released collection with the
tiered_storage.j2 JSON fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
RafalKorepta
approved these changes
Jul 2, 2026
Wiring — the knobs must actually reach the containers:
- forward BASELINE_COLLECTION_REF / CANDIDATE_COLLECTION_REF /
REDPANDA_VERSION through the docker plugin env allowlist (build 425's
candidate phase silently fell back to Galaxy latest without this)
- hoist the shared lane vars to taskfile level (GA pin bump = one edit)
Unstable lanes — single-phase ci:aws:rp:tiered:unstable at latest: the GA
pin only exists in the stable repo (build 426), and the lane owns its
IS_USING_UNSTABLE identity so local runs test the unstable repo too.
Fail loud, not silent:
- non-clobbering :ansible:prereqs (marker vs requirements.yml mtime), so
phase-2 deploys can't revert the candidate collection; provision/tiered
regain the dep and work standalone again
- timeout on the rpk consume data-survival checks ('-n N' blocks forever
on partial loss — the exact case they exist to catch)
- hard-fail on a missing baseline version file; baseline moved to /var/tmp
(Fedora /tmp is tmpfs)
- sentinel topic create failures surface (describe fallback replaces
'|| true'); version/status extra-vars only passed when explicitly set
Fail fast:
- ConnectionAttempts 100 -> 30 (a dead host said nothing for ~50 min)
- timeout_in_minutes on every step
- ts-connect builds the extra cluster at point of use, not 30-60 idle
minutes early; deferred extra destroy is a no-op when nothing was built
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
gene-redpanda
force-pushed
the
devprod-4109-upgrade-test
branch
from
July 13, 2026 21:57
82f9813 to
b0c02cb
Compare
PrzemekZglinicki
requested changes
Jul 15, 2026
Build 430: three builds queued from successive pushes interleaved their
lanes (per-lane concurrency groups serialize lanes, not builds). An earlier
build's cleanup-aws job then mass-terminated every ci-* instance in the
account at 22:09:49Z — including 23 instances belonging to build 430's
still-running lanes. Established SSH sessions froze ('privilege escalation
timeout') and freshly-launched nodes died before answering SSH
(UNREACHABLE). CloudTrail confirms: one aws-sdk-go-v2 TerminateInstances
call with all 5 of aws-fedora-tiered's instance IDs in it.
The taskfile already plumbed a MIN_AGE knob, but neither Go program
implemented --min-age (passing it would have crashed on an unknown flag).
- AWS cleanup: -min-age duration flag; age-gate instances (LaunchTime),
volumes (CreateTime), key pairs (CreateTime), IAM roles/profiles
(CreateDate), S3 buckets (CreationDate). Security groups carry no
timestamp, so gate them on in-use instead: an SG attached to any ENI is
skipped entirely — previously the code revoked all rules from a live
lane's SG even when the delete itself failed.
- GCP cleanup: same flag; age-gate instances, instance groups, firewalls,
addresses, subnets, networks (CreationTimestamp) and buckets
(TimeCreated). Service accounts expose no creation time and are skipped
under min-age rather than risking a live lane's credentials.
- cleanup:{aws,gcp}:ci default MIN_AGE=3h (> the 2h lane timeout, so
nothing a live lane owns can be reaped). MIN_AGE=0 restores a full sweep;
manual cleanup:aws / cleanup:gcp are unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…anda minimum) us-west2-b has been out of n2-standard-2 stock for 2+ weeks (builds 425-431). Zone discovery filters on zone status (UP), not capacity, and the module round-robins nodes across all discovered zones, so every 3+ node cluster placed at least one node in the dead zone and failed at terraform apply. Quota is not the issue (N2_CPUS 0/750 used) — it is Google-side zonal stock exhaustion. n2d-standard-4 because the replacement family must satisfy all three: - local-SSD capable: the module attaches local-ssd scratch disks, which rules out E2 (400: '[e2-standard-2, local-ssd] features are not compatible', build 432) - separate stock pool: N2D is AMD silicon, drawing from a different pool than the exhausted Intel N2; offered in all three us-west2 zones - Redpanda-compliant shape: the production requirements specify a minimum of two physical cores per node; standard-2 is 1 physical core with SMT, so standard-4 is the smallest compliant shape Override GCP_INSTANCE_TYPE to pin a different type per build. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
gene-redpanda
force-pushed
the
devprod-4109-upgrade-test
branch
2 times, most recently
from
July 17, 2026 20:48
c52ded7 to
3c9e21f
Compare
…otation Collection-upgrade validation: assert-node.sh signed-metadata/key checks (the RPM metadata-verification assert gated on the dnf backend), the sabotage-deb-keyring fixture (retaining the baseline key across rotation), and the ci/testing task wiring.
…E_COLLECTION_REF Previously CANDIDATE_COLLECTION_REF only took effect in the upgrade/unstable CI lanes (which call :ansible:collection:candidate explicitly). Every other provision workflow installed redpanda.cluster from requirements.yml (Galaxy) only, so testing a branch there meant hand-editing requirements.yml. Add :ansible:collection:pin, run as part of :ansible:prereqs, which force-installs redpanda.cluster from CANDIDATE_COLLECTION_REF on top of the requirements.yml install whenever the env var is set (git branch or Galaxy version); no-op when unset. Now any workflow honors the one env var with no file fiddling. run: once keeps the repeated prereqs deps from re-installing and clobbering the upgrade test's phase-1 baseline; the upgrade lanes' explicit override/candidate steps run afterward and still win. The pin is intentionally outside the requirements mtime status marker so a warm collections cache does not skip it. The env var is already forwarded through the Buildkite docker steps, so no pipeline plumbing changes. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Add a terse, developer-oriented CLAUDE.md (Task interface, ci:* lanes, CANDIDATE_COLLECTION_REF, upgrade-test model, lint). - Fix README: broken deploy-client.yml ref, Basic Usage shell bugs, add a short Task section, point branch-testing at CANDIDATE_COLLECTION_REF, link docs/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…aform pin - Remove the unused, drifted Makefile (CI uses Task; its ci-* targets diverged from the task ci:* lanes). - Remove Dockerfile_FEDORA (EOL base) and .pre-commit-config.yaml (CI enforces ansible-lint + terraform fmt via GitHub Actions); update README/CLAUDE.md. - Bump Dockerfile_UBUNTU terraform 1.4.5 -> 1.7.4 to match Taskfile.yml. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
gene-redpanda
force-pushed
the
devprod-4109-upgrade-test
branch
2 times, most recently
from
July 23, 2026 20:36
1f00b7c to
a6ddc78
Compare
Add terraform validate (dflook/terraform-validate) alongside the existing fmt check, scoped to aws/, so it runs on every PR. Remove the now-unused .tflint.hcl (pre-commit, its only consumer, is gone). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
gene-redpanda
force-pushed
the
devprod-4109-upgrade-test
branch
from
July 24, 2026 01:52
a6ddc78 to
b8c3e2b
Compare
PrzemekZglinicki
approved these changes
Jul 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Turns every Redpanda CI lane into a collection upgrade test: build a cluster with the last released
redpanda.clustercollection at a pinned RP version, seed a known dataset, then re-converge the same cluster with the candidate collection at the SAME version and assert a safe, lossless re-converge. Redpanda is pinned across both phases, so the collection is the only variable — a red lane means "this collection change breaks re-convergence of an existing cluster."What each lane does
Knobs
BASELINE_COLLECTION_REF— the "last released" collection (Galaxy latest default; override to a prior release).CANDIDATE_COLLECTION_REF— the collection under test: a git branch (git+<url>,<branch>), a released Galaxy version (redpanda.cluster:0.12.0), or unset →requirements.yml.REDPANDA_VERSION— pinned across both phases (single pin knob).Replaces the earlier standalone
ci:*:rp:upgradelanes; unstable lanes stay single-phase. Full model indocs/COLLECTION_UPGRADE_TEST.md.Why this exists / what it caught
Built to validate the Redpanda Artifact Registry migration. It found real upgrade/fresh-install
NO_PUBKEYfailures in the collection's apt key handling (fixed in redpanda-ansible-collection#150) and an invalid-JSON tiered config template (redpanda-ansible-collection#149).Validation status
ci:gcp:rp): green live on GCP — full re-converge + 500-record survival + version-stable assertions.tiered_storage.j2JSON fix (Fixed a couple of errors in readme comparing with vars.tf #149); connect needsCONNECT_RPM_TOKEN.🤖 Generated with Claude Code