Repository navigation
SASL E2E lane, rolling-restart coverage, collection pin, and credential-handling fixes - #251
Draft
gene-redpanda wants to merge 13 commits into
Draft
gene-redpanda wants to merge 13 commits into
gene-redpanda wants to merge 13 commits into
Conversation
The harness had zero coverage of SASL, user_management, or the broker rolling restart -- exactly the areas where the collection's worst bugs lived. The new ci:aws:rp:sasl lane fresh-installs the candidate collection with kafka_enable_authorization, reconciles a test user and topic ACL via the user_management role, and asserts real enforcement: superuser access works, anonymous access is denied, and the ACL-scoped user can use exactly its topic and nothing else. The lane then performs a SASL-authenticated rolling restart via operation-rolling-restart.yml (previously exercised by no lane) and re-verifies enforcement and data access afterwards. The restart playbook gains conditional RPK_USER/RPK_PASS environment for SASL clusters -- populated only when authorization is on, since rpk rejects empty credential env vars. requirements.yml pins redpanda.cluster to the released 0.12.0: floating on Galaxy latest meant any bad collection release broke every lane immediately; bumps are now deliberate. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Taskfile auto-loads .env (dotenv) for local credentials; it must never be committable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The env block re-exported only the static access key pair, silently stripping the session token -- any developer running lanes locally with aws sso credentials got InvalidClientTokenId from terraform. Static CI keys leave the token empty, which AWS ignores. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…block
Global env: templates evaluate before dotenv merges, so the aliasing
entries ('{{.AWS_ACCESS_KEY_ID | default .DA_AWS_ACCESS_KEY_ID}}' etc.)
resolved empty for anyone supplying credentials via .env and overrode
them -- terraform then fell back to stale ~/.aws keys and failed with
InvalidClientTokenId. Verified against go-task 3.39 and 3.52; the
ordering is by design. CI passes real process env, which flows through
untouched without these lines; the legacy DA_AWS_* aliases are dropped.
This also supersedes the earlier session-token passthrough attempt,
which could never work from dotenv for the same ordering reason.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
go-task merges included-taskfile env into the whole run, so cleanup.yml's
env block applied far beyond the reaper tasks: its credential templates
resolve empty before dotenv merges (blanking .env-supplied keys
everywhere) and it exported the comma-joined reaper region list
('us-east-2,us-west-2') as the global AWS_DEFAULT_REGION, which the aws
CLI rejects outright. The reaper already receives its regions via
--region arguments and credentials flow from process env or .env
untouched, so the block is deleted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The post-restart re-verification re-created the app topic and failed on TOPIC_ALREADY_EXISTS; creation now falls back to describe, the same idiom the upgrade seed task uses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
DO NOT MERGE this commit -- it exists so CI validates this harness against the 0.13.0 candidate collection. Revert to the pinned Galaxy release before merging. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
DO NOT MERGE. The agent environment supplies CANDIDATE_COLLECTION_REF pointing at collection main via the pipeline secret, which force-installs main over the requirements.yml branch pin -- build 445 validated main instead of the candidate (its sasl-lane auth failure is main's broken SASL bootstrap, fixed in the candidate). Pipeline-level env pins the candidate for every lane, same mechanism the devprod-4109 branch used. Drop together with the requirements TEMP commit before merging. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The node_exporter tarball download from GitHub fails intermittently under concurrent-lane rate limiting (multiple lanes each fetching to several hosts from one egress IP -- builds 445 and 446 both lost lanes to 'Remote end closed connection without response'). The deploy plays are idempotent, so a single bounded retry with a 30s backoff absorbs the flake without masking real failures. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
DO NOT MERGE. Build 446 proved the pipeline-level env pin is dead on arrival: the aws-sm secret hook exports CANDIDATE_COLLECTION_REF (git+...,main) into the job environment after yaml env applies, so every lane still validated collection main. The docker plugin's KEY=VALUE environment entries are set inside the container and cannot be clobbered by the job env; each step now pins the collection-hardening branch directly. Drop with the other TEMP commits before merging. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Lint for the redpanda.cluster collection now runs in the collection repo's own CI (production profile, pinned tooling); linting it again from the consuming harness double-reports the same findings against whichever collection version happens to be pinned here. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ling restart With serial: 1 the per-host entry health check runs immediately after the previous node rejoined, while partitions are still healing -- the one-shot probe read 'Healthy: false' and failed the play (build 447's sasl lane). The entry gate now uses the same --watch --exit-when-healthy + retry treatment as the post-restart checks. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The geerlingguy.node_exporter role streams the tarball from GitHub on every host with no retry, and resolves 'latest' via a GitHub API call per host -- under concurrent-lane rate limiting GitHub resets those connections, which has now killed five separate CI lanes across builds 445-448 (including one that outlived the lane-level retry). Both monitor playbooks now prefetch the tarball via get_url with 10 retries and point the role at the local file, pinning 1.12.1 (which also skips the per-host latest lookup). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds the harness's first SASL/user-management/rolling-restart E2E coverage, pins the collection dependency, and fixes three credential-handling bugs that broke local (non-CI) runs. Validates redpanda-data/redpanda-ansible-collection#152.
New:
ci:aws:rp:sasllane (+ Buildkite stepaws-up-ubuntu-sasl)The harness previously had zero coverage of SASL, the user_management role, or the broker rolling restart — exactly where the collection's worst historical bugs lived. The new single-phase lane:
kafka_enable_authorization: true(newprovision-cluster-sasl.yml)operation-user-management.yml,ops:users:apply)operation-rolling-restart.yml(previously exercised by no lane) — the playbook gains conditionalRPK_USER/RPK_PASSfor SASL clustersRun green end to end against the collection-hardening candidate on AWS (jammy, 3 brokers). The lane's first live runs surfaced two real collection bugs (node-id lookup vs advertised addresses; filters not loading from packaged collections), both fixed in the companion PR — which is the point of this coverage.
Collection dependency pinned
requirements.ymlpreviously floated on Galaxylatest, so any bad collection release broke every lane immediately. Now pinned; bumps become deliberate.Credential-handling fixes (hit by anyone running lanes locally)
env:aliasing entries evaluate before dotenv merges, so.env-supplied AWS credentials were silently clobbered with empty strings (terraform then failed on stale~/.awskeys). Verified against go-task 3.39 and 3.52. The aliases are removed; process env (CI) and.envboth flow through untouched. LegacyDA_AWS_*fallbacks are dropped.cleanup.yml's taskfile-level env block leaked into the whole run (included-taskfile env merges globally), blanking credentials and exporting the comma-joined reaper region list as everyone'sAWS_DEFAULT_REGION. The reaper already takes--regionargs; block removed..envis now gitignored (the Taskfile auto-loads it via dotenv; it was committable).Merge checklist
TEMP: point redpanda.cluster at the collection-hardening branchcommit (restore the Galaxy pin) once the collection PR merges and 0.13.0 is published — then bump the pin to 0.13.0.🤖 Generated with Claude Code