Skip to content

SASL E2E lane, rolling-restart coverage, collection pin, and credential-handling fixes - #251

Draft
gene-redpanda wants to merge 13 commits into
mainfrom
collection-hardening-e2e
Draft

gene-redpanda wants to merge 13 commits into
mainfrom
collection-hardening-e2e

Conversation

@gene-redpanda

Copy link
Copy Markdown
Member

Adds the harness's first SASL/user-management/rolling-restart E2E coverage, pins the collection dependency, and fixes three credential-handling bugs that broke local (non-CI) runs. Validates redpanda-data/redpanda-ansible-collection#152.

New: ci:aws:rp:sasl lane (+ Buildkite step aws-up-ubuntu-sasl)

The harness previously had zero coverage of SASL, the user_management role, or the broker rolling restart — exactly where the collection's worst historical bugs lived. The new single-phase lane:

  1. Fresh-installs the candidate collection with kafka_enable_authorization: true (new provision-cluster-sasl.yml)
  2. Reconciles a test user + topic ACL via the user_management role (new operation-user-management.yml, ops:users:apply)
  3. Asserts real enforcement: superuser works, anonymous access denied, ACL-scoped user can use exactly its topic and is refused others
  4. Performs a SASL-authenticated rolling restart via operation-rolling-restart.yml (previously exercised by no lane) — the playbook gains conditional RPK_USER/RPK_PASS for SASL clusters
  5. Re-verifies enforcement and topic data survival post-restart

Run green end to end against the collection-hardening candidate on AWS (jammy, 3 brokers). The lane's first live runs surfaced two real collection bugs (node-id lookup vs advertised addresses; filters not loading from packaged collections), both fixed in the companion PR — which is the point of this coverage.

Collection dependency pinned

requirements.yml previously floated on Galaxy latest, so any bad collection release broke every lane immediately. Now pinned; bumps become deliberate.

Credential-handling fixes (hit by anyone running lanes locally)

  • Global env: aliasing entries evaluate before dotenv merges, so .env-supplied AWS credentials were silently clobbered with empty strings (terraform then failed on stale ~/.aws keys). Verified against go-task 3.39 and 3.52. The aliases are removed; process env (CI) and .env both flow through untouched. Legacy DA_AWS_* fallbacks are dropped.
  • cleanup.yml's taskfile-level env block leaked into the whole run (included-taskfile env merges globally), blanking credentials and exporting the comma-joined reaper region list as everyone's AWS_DEFAULT_REGION. The reaper already takes --region args; block removed.
  • .env is now gitignored (the Taskfile auto-loads it via dotenv; it was committable).

Merge checklist

  • Drop the final TEMP: point redpanda.cluster at the collection-hardening branch commit (restore the Galaxy pin) once the collection PR merges and 0.13.0 is published — then bump the pin to 0.13.0.

🤖 Generated with Claude Code

gene-redpanda and others added 7 commits August 11, 2026 14:37
The harness had zero coverage of SASL, user_management, or the broker
rolling restart -- exactly the areas where the collection's worst bugs
lived. The new ci:aws:rp:sasl lane fresh-installs the candidate
collection with kafka_enable_authorization, reconciles a test user and
topic ACL via the user_management role, and asserts real enforcement:
superuser access works, anonymous access is denied, and the ACL-scoped
user can use exactly its topic and nothing else. The lane then performs
a SASL-authenticated rolling restart via operation-rolling-restart.yml
(previously exercised by no lane) and re-verifies enforcement and data
access afterwards. The restart playbook gains conditional
RPK_USER/RPK_PASS environment for SASL clusters -- populated only when
authorization is on, since rpk rejects empty credential env vars.

requirements.yml pins redpanda.cluster to the released 0.12.0: floating
on Galaxy latest meant any bad collection release broke every lane
immediately; bumps are now deliberate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Taskfile auto-loads .env (dotenv) for local credentials; it must
never be committable.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The env block re-exported only the static access key pair, silently
stripping the session token -- any developer running lanes locally with
aws sso credentials got InvalidClientTokenId from terraform. Static CI
keys leave the token empty, which AWS ignores.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…block

Global env: templates evaluate before dotenv merges, so the aliasing
entries ('{{.AWS_ACCESS_KEY_ID | default .DA_AWS_ACCESS_KEY_ID}}' etc.)
resolved empty for anyone supplying credentials via .env and overrode
them -- terraform then fell back to stale ~/.aws keys and failed with
InvalidClientTokenId. Verified against go-task 3.39 and 3.52; the
ordering is by design. CI passes real process env, which flows through
untouched without these lines; the legacy DA_AWS_* aliases are dropped.
This also supersedes the earlier session-token passthrough attempt,
which could never work from dotenv for the same ordering reason.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
go-task merges included-taskfile env into the whole run, so cleanup.yml's
env block applied far beyond the reaper tasks: its credential templates
resolve empty before dotenv merges (blanking .env-supplied keys
everywhere) and it exported the comma-joined reaper region list
('us-east-2,us-west-2') as the global AWS_DEFAULT_REGION, which the aws
CLI rejects outright. The reaper already receives its regions via
--region arguments and credentials flow from process env or .env
untouched, so the block is deleted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The post-restart re-verification re-created the app topic and failed on
TOPIC_ALREADY_EXISTS; creation now falls back to describe, the same
idiom the upgrade seed task uses.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
DO NOT MERGE this commit -- it exists so CI validates this harness
against the 0.13.0 candidate collection. Revert to the pinned Galaxy
release before merging.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@gene-redpanda gene-redpanda added the ci-ready indicates that a PR is ready for builds to run label Aug 12, 2026
gene-redpanda and others added 6 commits August 12, 2026 16:43
DO NOT MERGE. The agent environment supplies CANDIDATE_COLLECTION_REF
pointing at collection main via the pipeline secret, which force-installs
main over the requirements.yml branch pin -- build 445 validated main
instead of the candidate (its sasl-lane auth failure is main's broken
SASL bootstrap, fixed in the candidate). Pipeline-level env pins the
candidate for every lane, same mechanism the devprod-4109 branch used.
Drop together with the requirements TEMP commit before merging.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The node_exporter tarball download from GitHub fails intermittently under
concurrent-lane rate limiting (multiple lanes each fetching to several
hosts from one egress IP -- builds 445 and 446 both lost lanes to
'Remote end closed connection without response'). The deploy plays are
idempotent, so a single bounded retry with a 30s backoff absorbs the
flake without masking real failures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
DO NOT MERGE. Build 446 proved the pipeline-level env pin is dead on
arrival: the aws-sm secret hook exports CANDIDATE_COLLECTION_REF
(git+...,main) into the job environment after yaml env applies, so every
lane still validated collection main. The docker plugin's KEY=VALUE
environment entries are set inside the container and cannot be clobbered
by the job env; each step now pins the collection-hardening branch
directly. Drop with the other TEMP commits before merging.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Lint for the redpanda.cluster collection now runs in the collection
repo's own CI (production profile, pinned tooling); linting it again
from the consuming harness double-reports the same findings against
whichever collection version happens to be pinned here.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ling restart

With serial: 1 the per-host entry health check runs immediately after the
previous node rejoined, while partitions are still healing -- the one-shot
probe read 'Healthy: false' and failed the play (build 447's sasl lane).
The entry gate now uses the same --watch --exit-when-healthy + retry
treatment as the post-restart checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The geerlingguy.node_exporter role streams the tarball from GitHub on
every host with no retry, and resolves 'latest' via a GitHub API call per
host -- under concurrent-lane rate limiting GitHub resets those
connections, which has now killed five separate CI lanes across builds
445-448 (including one that outlived the lane-level retry). Both monitor
playbooks now prefetch the tarball via get_url with 10 retries and point
the role at the local file, pinning 1.12.1 (which also skips the per-host
latest lookup).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-ready indicates that a PR is ready for builds to run

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant