Skip to content

Portal liveness probe uses dependency-aware health check, turning dependency outages into restart storms #1479

Description

@Brad-Edwards-SecOps

Summary

The portal web deployment uses the dependency-aware /health/ endpoint for its
liveness probe. /health/ is a CoarseHealthCheckView that checks Postgres,
both Redis caches, the channel layer, and default file storage. When any of those
dependencies is unavailable the endpoint returns HTTP 500, liveness fails, and the
kubelet restarts the container.

The effect is that a dependency outage becomes a portal restart storm: the pods are
killed precisely while they are waiting for the dependency to come back, so they
never reach a state where they can serve.

Observed

During a Cloud SQL maintenance window on the gcp-dev tenant (a tier change plus
ZONAL -> REGIONAL conversion, ~6 minutes), every portal-web pod entered a
restart loop:

Warning  Unhealthy  kubelet  Readiness probe failed: HTTP probe failed with statuscode: 500
Warning  Unhealthy  kubelet  Liveness probe failed: HTTP probe failed with statuscode: 500
Normal   Killing    kubelet  Container portal failed liveness probe, will be restarted

Pod logs show the application reaching Application startup complete and being sent
SIGTERM roughly one second later. The portal returned 502 externally for the whole
window and did not recover on its own until the database returned; requests in flight
failed with HTTP 500.

Once the database was healthy, /health/ returned 200 with every subsystem working,
confirming the application itself was fine — only the probe wiring was wrong.

Root cause

Liveness and readiness answer different questions:

  • liveness — "is this process wedged and in need of a restart?"
  • readiness — "can this pod serve traffic right now?"

A dependency check belongs in readiness, where failure removes the pod from the
Service and traffic stops, but the process is left alone to recover. Wiring it into
liveness means an outage in any dependency restarts every replica.

timeoutSeconds: 1 on both probes is a secondary problem: a check that queries
Postgres, two Redis caches, and object storage will not reliably answer inside one
second under load, so the probes will also flap during traffic spikes independent
of any outage.

Applied hotfix (live cluster only, not in source)

Patched directly on the gcp-dev cluster to get through the current event:

kubectl -n shifter-platform patch deploy portal-web --type=json -p '[
 {"op":"replace","path":"/spec/template/spec/containers/0/livenessProbe","value":{
   "tcpSocket":{"port":"http"},
   "initialDelaySeconds":60,"periodSeconds":30,"timeoutSeconds":5,
   "failureThreshold":3,"successThreshold":1}},
 {"op":"replace","path":"/spec/template/spec/containers/0/readinessProbe/timeoutSeconds","value":5}
]'

This is not in the chart, so the next Helm apply reverts it.

Permanent fix

Make the equivalent change in source. Both copies must stay consistent:

  • platform/charts/shifter/templates/web-deployment.yaml (liveness at line ~62)
  • platform/k8s/gcp/base/web-deployment.yaml

Suggested shape:

  • liveness — process-level check that does not touch dependencies, either
    tcpSocket on the http port or a dedicated shallow HTTP endpoint that returns
    200 whenever the WSGI/ASGI app is serving. A shallow endpoint is preferable to
    tcpSocket because it proves the application, not just the listening socket, is
    responsive; it requires a small view alongside CoarseHealthCheckView.
  • readiness — keep /health/, with timeoutSeconds raised to a value that
    tolerates dependency latency under load (5s is what the hotfix uses).

Audit the other deployments carrying livenessProbe for the same pattern —
worker-engine, worker-reconciler, worker-provisioner-launcher,
ctf-scheduler, worker-operation-result-applier, guacamole-client,
guacamole-bootstrap-prune — and their platform/k8s/gcp/base/ counterparts.

Acceptance criteria

  • Portal liveness no longer fails when Postgres or Redis is unavailable.
  • Readiness still fails in that situation, so unready pods are removed from the Service.
  • A dependency outage of several minutes leaves portal pods running with restart
    count unchanged, and they serve traffic again once the dependency returns without
    manual intervention.
  • Chart and static base manifest remain byte-consistent.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions