Summary
The portal web deployment uses the dependency-aware /health/ endpoint for its
liveness probe. /health/ is a CoarseHealthCheckView that checks Postgres,
both Redis caches, the channel layer, and default file storage. When any of those
dependencies is unavailable the endpoint returns HTTP 500, liveness fails, and the
kubelet restarts the container.
The effect is that a dependency outage becomes a portal restart storm: the pods are
killed precisely while they are waiting for the dependency to come back, so they
never reach a state where they can serve.
Observed
During a Cloud SQL maintenance window on the gcp-dev tenant (a tier change plus
ZONAL -> REGIONAL conversion, ~6 minutes), every portal-web pod entered a
restart loop:
Warning Unhealthy kubelet Readiness probe failed: HTTP probe failed with statuscode: 500
Warning Unhealthy kubelet Liveness probe failed: HTTP probe failed with statuscode: 500
Normal Killing kubelet Container portal failed liveness probe, will be restarted
Pod logs show the application reaching Application startup complete and being sent
SIGTERM roughly one second later. The portal returned 502 externally for the whole
window and did not recover on its own until the database returned; requests in flight
failed with HTTP 500.
Once the database was healthy, /health/ returned 200 with every subsystem working,
confirming the application itself was fine — only the probe wiring was wrong.
Root cause
Liveness and readiness answer different questions:
- liveness — "is this process wedged and in need of a restart?"
- readiness — "can this pod serve traffic right now?"
A dependency check belongs in readiness, where failure removes the pod from the
Service and traffic stops, but the process is left alone to recover. Wiring it into
liveness means an outage in any dependency restarts every replica.
timeoutSeconds: 1 on both probes is a secondary problem: a check that queries
Postgres, two Redis caches, and object storage will not reliably answer inside one
second under load, so the probes will also flap during traffic spikes independent
of any outage.
Applied hotfix (live cluster only, not in source)
Patched directly on the gcp-dev cluster to get through the current event:
kubectl -n shifter-platform patch deploy portal-web --type=json -p '[
{"op":"replace","path":"/spec/template/spec/containers/0/livenessProbe","value":{
"tcpSocket":{"port":"http"},
"initialDelaySeconds":60,"periodSeconds":30,"timeoutSeconds":5,
"failureThreshold":3,"successThreshold":1}},
{"op":"replace","path":"/spec/template/spec/containers/0/readinessProbe/timeoutSeconds","value":5}
]'
This is not in the chart, so the next Helm apply reverts it.
Permanent fix
Make the equivalent change in source. Both copies must stay consistent:
platform/charts/shifter/templates/web-deployment.yaml (liveness at line ~62)
platform/k8s/gcp/base/web-deployment.yaml
Suggested shape:
- liveness — process-level check that does not touch dependencies, either
tcpSocket on the http port or a dedicated shallow HTTP endpoint that returns
200 whenever the WSGI/ASGI app is serving. A shallow endpoint is preferable to
tcpSocket because it proves the application, not just the listening socket, is
responsive; it requires a small view alongside CoarseHealthCheckView.
- readiness — keep
/health/, with timeoutSeconds raised to a value that
tolerates dependency latency under load (5s is what the hotfix uses).
Audit the other deployments carrying livenessProbe for the same pattern —
worker-engine, worker-reconciler, worker-provisioner-launcher,
ctf-scheduler, worker-operation-result-applier, guacamole-client,
guacamole-bootstrap-prune — and their platform/k8s/gcp/base/ counterparts.
Acceptance criteria
- Portal liveness no longer fails when Postgres or Redis is unavailable.
- Readiness still fails in that situation, so unready pods are removed from the Service.
- A dependency outage of several minutes leaves portal pods running with restart
count unchanged, and they serve traffic again once the dependency returns without
manual intervention.
- Chart and static base manifest remain byte-consistent.
Summary
The portal
webdeployment uses the dependency-aware/health/endpoint for itsliveness probe.
/health/is aCoarseHealthCheckViewthat checks Postgres,both Redis caches, the channel layer, and default file storage. When any of those
dependencies is unavailable the endpoint returns HTTP 500, liveness fails, and the
kubelet restarts the container.
The effect is that a dependency outage becomes a portal restart storm: the pods are
killed precisely while they are waiting for the dependency to come back, so they
never reach a state where they can serve.
Observed
During a Cloud SQL maintenance window on the
gcp-devtenant (a tier change plusZONAL -> REGIONALconversion, ~6 minutes), everyportal-webpod entered arestart loop:
Pod logs show the application reaching
Application startup completeand being sentSIGTERMroughly one second later. The portal returned 502 externally for the wholewindow and did not recover on its own until the database returned; requests in flight
failed with HTTP 500.
Once the database was healthy,
/health/returned 200 with every subsystemworking,confirming the application itself was fine — only the probe wiring was wrong.
Root cause
Liveness and readiness answer different questions:
A dependency check belongs in readiness, where failure removes the pod from the
Service and traffic stops, but the process is left alone to recover. Wiring it into
liveness means an outage in any dependency restarts every replica.
timeoutSeconds: 1on both probes is a secondary problem: a check that queriesPostgres, two Redis caches, and object storage will not reliably answer inside one
second under load, so the probes will also flap during traffic spikes independent
of any outage.
Applied hotfix (live cluster only, not in source)
Patched directly on the
gcp-devcluster to get through the current event:This is not in the chart, so the next Helm apply reverts it.
Permanent fix
Make the equivalent change in source. Both copies must stay consistent:
platform/charts/shifter/templates/web-deployment.yaml(liveness at line ~62)platform/k8s/gcp/base/web-deployment.yamlSuggested shape:
tcpSocketon thehttpport or a dedicated shallow HTTP endpoint that returns200 whenever the WSGI/ASGI app is serving. A shallow endpoint is preferable to
tcpSocketbecause it proves the application, not just the listening socket, isresponsive; it requires a small view alongside
CoarseHealthCheckView./health/, withtimeoutSecondsraised to a value thattolerates dependency latency under load (5s is what the hotfix uses).
Audit the other deployments carrying
livenessProbefor the same pattern —worker-engine,worker-reconciler,worker-provisioner-launcher,ctf-scheduler,worker-operation-result-applier,guacamole-client,guacamole-bootstrap-prune— and theirplatform/k8s/gcp/base/counterparts.Acceptance criteria
count unchanged, and they serve traffic again once the dependency returns without
manual intervention.