Summary
The gcp-dev cluster carried an orphaned aces-operation-record-prune Deployment
left behind by the ACES -> RAES rename. It invoked a management command that no
longer exists and had been in CrashLoopBackOff for approximately 3 days 21 hours
with 1095 restarts.
Running: python manage.py run_aces_operation_record_prune
Unknown command: 'run_aces_operation_record_prune'. Did you mean run_raes_operation_record_prune?
What it actually was
Not a broken command reference to be corrected -- a duplicate resource that should
have been removed when the rename landed:
|
aces-operation-record-prune |
raes-operation-record-prune |
| created |
2026-07-18T22:52:31Z |
2026-07-29T00:20:33Z |
| image |
...@sha256:743081f2... (stale) |
...@sha256:01bc2475... (current) |
| command |
python manage.py run_aces_operation_record_prune |
image default |
| state |
CrashLoopBackOff, 1095 restarts |
Running, healthy |
defined in platform/ |
no |
yes |
The replacement had been running healthily since 2026-07-29. The management command
itself is fine and correctly named:
shifter/shifter_platform/shared/management/commands/run_raes_operation_record_prune.py.
Action taken
Deleted the orphaned Deployment from the live gcp-dev cluster:
kubectl -n shifter-platform delete deploy aces-operation-record-prune
Safe to delete outright: it is not defined anywhere under platform/, so no chart
or manifest recreates it, and its function is already covered by
raes-operation-record-prune. Correcting its command instead would have been wrong
-- that would have run the prune twice on the same records.
No crashlooping pods remain in the namespace.
Why this needs an issue anyway
The cluster is not the only place this can exist, and nothing detected it for
nearly four weeks:
-
Other environments. Any tenant deployed before the rename may still carry the
same orphan. aws-dev, aws-proof, and aws-prod should be checked for a
aces-operation-record-prune Deployment (and for any other aces--prefixed
workload) and cleaned up the same way.
-
Renames leave orphans. Applying a renamed manifest creates the new resource;
it does not remove the old one. Whatever applies these manifests does not prune
resources that disappear from source, so every rename risks leaving a duplicate
running against production data. Worth deciding whether the GCP apply path should
prune resources absent from source, or whether renames need an explicit cleanup
step in their change checklist.
-
A pod crashlooped 1095 times without anyone noticing. There is no alert on
sustained CrashLoopBackOff in shifter-platform. That is the more valuable fix
here: this instance was harmless, but the same silence would cover a workload that
actually matters.
Acceptance criteria
- No
aces--prefixed workloads remain in any environment.
- Renames have a defined path that removes the superseded resource, either by
pruning on apply or by an explicit step.
- Sustained
CrashLoopBackOff in shifter-platform raises an alert rather than
persisting silently for weeks.
Summary
The
gcp-devcluster carried an orphanedaces-operation-record-pruneDeploymentleft behind by the ACES -> RAES rename. It invoked a management command that no
longer exists and had been in
CrashLoopBackOfffor approximately 3 days 21 hourswith 1095 restarts.
What it actually was
Not a broken command reference to be corrected -- a duplicate resource that should
have been removed when the rename landed:
aces-operation-record-pruneraes-operation-record-prune...@sha256:743081f2...(stale)...@sha256:01bc2475...(current)python manage.py run_aces_operation_record_pruneplatform/The replacement had been running healthily since 2026-07-29. The management command
itself is fine and correctly named:
shifter/shifter_platform/shared/management/commands/run_raes_operation_record_prune.py.Action taken
Deleted the orphaned Deployment from the live
gcp-devcluster:Safe to delete outright: it is not defined anywhere under
platform/, so no chartor manifest recreates it, and its function is already covered by
raes-operation-record-prune. Correcting its command instead would have been wrong-- that would have run the prune twice on the same records.
No crashlooping pods remain in the namespace.
Why this needs an issue anyway
The cluster is not the only place this can exist, and nothing detected it for
nearly four weeks:
Other environments. Any tenant deployed before the rename may still carry the
same orphan.
aws-dev,aws-proof, andaws-prodshould be checked for aaces-operation-record-pruneDeployment (and for any otheraces--prefixedworkload) and cleaned up the same way.
Renames leave orphans. Applying a renamed manifest creates the new resource;
it does not remove the old one. Whatever applies these manifests does not prune
resources that disappear from source, so every rename risks leaving a duplicate
running against production data. Worth deciding whether the GCP apply path should
prune resources absent from source, or whether renames need an explicit cleanup
step in their change checklist.
A pod crashlooped 1095 times without anyone noticing. There is no alert on
sustained
CrashLoopBackOffinshifter-platform. That is the more valuable fixhere: this instance was harmless, but the same silence would cover a workload that
actually matters.
Acceptance criteria
aces--prefixed workloads remain in any environment.pruning on apply or by an explicit step.
CrashLoopBackOffinshifter-platformraises an alert rather thanpersisting silently for weeks.