Skip to content

gcp-dev event staging: platform faults found and corrected scaling to 300 concurrent ranges #1481

Description

@Brad-Edwards-SecOps

Summary

Record of the platform faults found and corrected while staging the gcp-dev
tenant for the KeplerOps CTF (200-300 concurrent ranges). Several were latent
defects that only surfaced under the attempt to scale, and several fixes are
currently live on the cluster but not in source, so they revert on the next
Helm apply. This is a description of what was found and what state it is in.

Deploy runs

run commit outcome
31083273037 9688b3c50 failure
31084576511 a84555d71 failure -- Redis resize, zonal capacity
31086472277 a84555d71 cancelled -- would have replaced node pools
31088808310 caa1f9517 in flight

Faults found

1. GKE pod IP range caps the cluster at 16 nodes

The blocking fault. All three node pools (web, workers, provisioner) are
pinned to the gke-provisioner-pods secondary range (10.46.0.0/20) at
maxPodsPerNode: 110. GKE allocates a /24 per node, so that range holds
4096/256 = 16 nodes total across all three pools combined. It reported
utilization: 1.0.

The cluster's default pod range gke-pods (10.44.0.0/16, 256 nodes) was
entirely unused.

Every attempt to add a node failed with:

IP space of 'projects/prod-ksqdkj/regions/us-central1/subnetworks/
shifter-gcp-dev-gke' is exhausted. Insufficient free IP addresses in the
IP range '10.46.0.0/20'.

This does not fail fast. Three SET_NODE_POOL_SIZE operations hung for ~20
minutes retrying, holding the cluster operation lock so no other cluster
operation could proceed. 83 pods sat Pending.

Resolved by creating node pool shifter-gcp-dev-web2 (36 nodes,
e2-standard-4) on the gke-pods range. Workloads carry no nodeSelectors,
tolerations, or hard affinity -- only soft pod anti-affinity -- so pending pods
scheduled onto it without any drain or migration.

The durable fix is to move the original pools onto gke-pods. Pod range is
immutable on an existing pool, so that means replacing them: a planned
migration, not a deploy-time change.

2. tfvars declared a machine type that forces node pool replacement

terraform.tfvars declared e2-standard-8 for web/workers while the live pools
are e2-standard-4. GKE cannot change machine type in place, so the next
terraform apply would have replaced all three node pools -- during event
staging. Run 31086472277 was cancelled for this reason.

Corrected in c5255d3cc and caa1f9517: machine types match the live pools,
node counts total 15 of the 16 the /20 can hold, with the constraint documented
in the file so the counts are not "optimised" back upward.

3. Portal liveness probe turns dependency outages into restart storms

portal-web used the dependency-aware /health/ endpoint (Postgres, both Redis
caches, channel layer, file storage) as its liveness probe. During a Cloud
SQL maintenance window the endpoint returned 500, liveness failed, and the
kubelet restarted every replica while they were waiting for the database to come
back. Pods reached Application startup complete and were sent SIGTERM about
one second later. The portal served 502 for the duration.

timeoutSeconds: 1 on both probes is a secondary problem for a check that
queries four dependencies.

Filed as #1479. Hotfixed live (liveness -> tcpSocket, readiness timeout -> 5s);
not in the chart, so Helm reverts it.

4. Orphaned aces-operation-record-prune Deployment

Left behind by the ACES -> RAES rename, invoking a management command that no
longer exists. In CrashLoopBackOff for 3d21h with 1095 restarts. Superseded
by a healthy raes-operation-record-prune running since 2026-07-29, and not
defined anywhere under platform/.

Deleted from the cluster. Filed as #1480, which also covers checking other
environments and the absence of any alert on sustained CrashLoopBackOff.

5. Range access authorized against the wrong instance registry

cms.services._range_access resolved instance ownership from
cms.models.Instance, but realized range instances are rows of
engine.models.Instance -- the CMS table is written only by NGFW provisioning.
Every terminal/RDP/SSH open failed with "Instance not found" before the workspace
check was reached, on every provisioned range.

Fixed in c17b34795, joining on the shared request_id UUID rather than on
primary key (the two request tables have independent pks). 252b6ba90 guards the
lookup against null user/request. Both are in 31088808310 and were not in
the previously running image.

6. Per-range OpenVPN gateway provisioned unconditionally

Every range on a capable backend whose scenario exposes a single
participant-visible Kali attacker provisions a dedicated gateway VM, an external
address, and per-range VPN secrets, and runs an extra guest-setup step that can
fail. Four of the 11 range VMs currently alive are vpn-gateway instances. At
300 ranges that is 300 additional VMs and external IPs for an access path this
event does not use.

348ba69fb puts this behind RANGE_OPENVPN_ENABLED. Note the setting defaults
to true
, so the code alone changes nothing -- the environment must set it
false.

7. Range placement pinned to a single zone

Range cells were placed in RANGE_NETWORK_ZONE only. us-central1-a has already
returned ZONE_RESOURCE_POOL_EXHAUSTED_WITH_DETAILS for a range host
(shifter-r-157, 2026-07-29), and the live ranges are concentrated in
us-central1-a/-c.

d2614557f adds a RANGE_NETWORK_ZONES pool with per-slot round-robin
placement; a84555d71 allows the key in the provisioner-Job admission policy,
without which admission denies every Job carrying it.

Infrastructure applied

Applied directly rather than through a deploy, after the deploy carrying them
failed on the Redis resize and the event timeline did not permit 25-minute
round trips per attempt:

  • Cloud SQL: db-custom-8-30720, REGIONAL, 100GB (this one did apply via
    terraform in 31084576511 before that run failed)
  • Redis: 1GB -> 16GB. Terraform failed with "Not enough zonal resources are
    available"; the same API call succeeded directly once capacity freed
  • Node pools off 1 node each; new web2 pool on the unused /16
  • Replica counts: portal 2 -> 24, guacd 1 -> 120, guacamole-client 1 -> 24,
    workers 1 -> 3 each
  • Resource requests/limits raised to 1 CPU / 2Gi requests, 2 CPU / 4Gi limits.
    guacd had been capped at a 500m CPU limit and portal at a 1Gi memory limit

State that is NOT durable

These are live on the cluster and absent from source. The next Helm apply
reverts them:

terraform.tfvars is the one piece reconciled to reality, so a terraform apply
converges rather than fighting the cluster.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions