Skip to content

Orphaned aces-operation-record-prune Deployment crashlooped for weeks after the ACES to RAES rename #1480

Description

@Brad-Edwards-SecOps

Summary

The gcp-dev cluster carried an orphaned aces-operation-record-prune Deployment
left behind by the ACES -> RAES rename. It invoked a management command that no
longer exists and had been in CrashLoopBackOff for approximately 3 days 21 hours
with 1095 restarts.

Running: python manage.py run_aces_operation_record_prune
Unknown command: 'run_aces_operation_record_prune'. Did you mean run_raes_operation_record_prune?

What it actually was

Not a broken command reference to be corrected -- a duplicate resource that should
have been removed when the rename landed:

aces-operation-record-prune raes-operation-record-prune
created 2026-07-18T22:52:31Z 2026-07-29T00:20:33Z
image ...@sha256:743081f2... (stale) ...@sha256:01bc2475... (current)
command python manage.py run_aces_operation_record_prune image default
state CrashLoopBackOff, 1095 restarts Running, healthy
defined in platform/ no yes

The replacement had been running healthily since 2026-07-29. The management command
itself is fine and correctly named:
shifter/shifter_platform/shared/management/commands/run_raes_operation_record_prune.py.

Action taken

Deleted the orphaned Deployment from the live gcp-dev cluster:

kubectl -n shifter-platform delete deploy aces-operation-record-prune

Safe to delete outright: it is not defined anywhere under platform/, so no chart
or manifest recreates it, and its function is already covered by
raes-operation-record-prune. Correcting its command instead would have been wrong
-- that would have run the prune twice on the same records.

No crashlooping pods remain in the namespace.

Why this needs an issue anyway

The cluster is not the only place this can exist, and nothing detected it for
nearly four weeks:

  1. Other environments. Any tenant deployed before the rename may still carry the
    same orphan. aws-dev, aws-proof, and aws-prod should be checked for a
    aces-operation-record-prune Deployment (and for any other aces--prefixed
    workload) and cleaned up the same way.

  2. Renames leave orphans. Applying a renamed manifest creates the new resource;
    it does not remove the old one. Whatever applies these manifests does not prune
    resources that disappear from source, so every rename risks leaving a duplicate
    running against production data. Worth deciding whether the GCP apply path should
    prune resources absent from source, or whether renames need an explicit cleanup
    step in their change checklist.

  3. A pod crashlooped 1095 times without anyone noticing. There is no alert on
    sustained CrashLoopBackOff in shifter-platform. That is the more valuable fix
    here: this instance was harmless, but the same silence would cover a workload that
    actually matters.

Acceptance criteria

  • No aces--prefixed workloads remain in any environment.
  • Renames have a defined path that removes the superseded resource, either by
    pruning on apply or by an explicit step.
  • Sustained CrashLoopBackOff in shifter-platform raises an alert rather than
    persisting silently for weeks.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions