Summary
Record of the platform faults found and corrected while staging the gcp-dev
tenant for the KeplerOps CTF (200-300 concurrent ranges). Several were latent
defects that only surfaced under the attempt to scale, and several fixes are
currently live on the cluster but not in source, so they revert on the next
Helm apply. This is a description of what was found and what state it is in.
Deploy runs
| run |
commit |
outcome |
31083273037 |
9688b3c50 |
failure |
31084576511 |
a84555d71 |
failure -- Redis resize, zonal capacity |
31086472277 |
a84555d71 |
cancelled -- would have replaced node pools |
31088808310 |
caa1f9517 |
in flight |
Faults found
1. GKE pod IP range caps the cluster at 16 nodes
The blocking fault. All three node pools (web, workers, provisioner) are
pinned to the gke-provisioner-pods secondary range (10.46.0.0/20) at
maxPodsPerNode: 110. GKE allocates a /24 per node, so that range holds
4096/256 = 16 nodes total across all three pools combined. It reported
utilization: 1.0.
The cluster's default pod range gke-pods (10.44.0.0/16, 256 nodes) was
entirely unused.
Every attempt to add a node failed with:
IP space of 'projects/prod-ksqdkj/regions/us-central1/subnetworks/
shifter-gcp-dev-gke' is exhausted. Insufficient free IP addresses in the
IP range '10.46.0.0/20'.
This does not fail fast. Three SET_NODE_POOL_SIZE operations hung for ~20
minutes retrying, holding the cluster operation lock so no other cluster
operation could proceed. 83 pods sat Pending.
Resolved by creating node pool shifter-gcp-dev-web2 (36 nodes,
e2-standard-4) on the gke-pods range. Workloads carry no nodeSelectors,
tolerations, or hard affinity -- only soft pod anti-affinity -- so pending pods
scheduled onto it without any drain or migration.
The durable fix is to move the original pools onto gke-pods. Pod range is
immutable on an existing pool, so that means replacing them: a planned
migration, not a deploy-time change.
2. tfvars declared a machine type that forces node pool replacement
terraform.tfvars declared e2-standard-8 for web/workers while the live pools
are e2-standard-4. GKE cannot change machine type in place, so the next
terraform apply would have replaced all three node pools -- during event
staging. Run 31086472277 was cancelled for this reason.
Corrected in c5255d3cc and caa1f9517: machine types match the live pools,
node counts total 15 of the 16 the /20 can hold, with the constraint documented
in the file so the counts are not "optimised" back upward.
3. Portal liveness probe turns dependency outages into restart storms
portal-web used the dependency-aware /health/ endpoint (Postgres, both Redis
caches, channel layer, file storage) as its liveness probe. During a Cloud
SQL maintenance window the endpoint returned 500, liveness failed, and the
kubelet restarted every replica while they were waiting for the database to come
back. Pods reached Application startup complete and were sent SIGTERM about
one second later. The portal served 502 for the duration.
timeoutSeconds: 1 on both probes is a secondary problem for a check that
queries four dependencies.
Filed as #1479. Hotfixed live (liveness -> tcpSocket, readiness timeout -> 5s);
not in the chart, so Helm reverts it.
4. Orphaned aces-operation-record-prune Deployment
Left behind by the ACES -> RAES rename, invoking a management command that no
longer exists. In CrashLoopBackOff for 3d21h with 1095 restarts. Superseded
by a healthy raes-operation-record-prune running since 2026-07-29, and not
defined anywhere under platform/.
Deleted from the cluster. Filed as #1480, which also covers checking other
environments and the absence of any alert on sustained CrashLoopBackOff.
5. Range access authorized against the wrong instance registry
cms.services._range_access resolved instance ownership from
cms.models.Instance, but realized range instances are rows of
engine.models.Instance -- the CMS table is written only by NGFW provisioning.
Every terminal/RDP/SSH open failed with "Instance not found" before the workspace
check was reached, on every provisioned range.
Fixed in c17b34795, joining on the shared request_id UUID rather than on
primary key (the two request tables have independent pks). 252b6ba90 guards the
lookup against null user/request. Both are in 31088808310 and were not in
the previously running image.
6. Per-range OpenVPN gateway provisioned unconditionally
Every range on a capable backend whose scenario exposes a single
participant-visible Kali attacker provisions a dedicated gateway VM, an external
address, and per-range VPN secrets, and runs an extra guest-setup step that can
fail. Four of the 11 range VMs currently alive are vpn-gateway instances. At
300 ranges that is 300 additional VMs and external IPs for an access path this
event does not use.
348ba69fb puts this behind RANGE_OPENVPN_ENABLED. Note the setting defaults
to true, so the code alone changes nothing -- the environment must set it
false.
7. Range placement pinned to a single zone
Range cells were placed in RANGE_NETWORK_ZONE only. us-central1-a has already
returned ZONE_RESOURCE_POOL_EXHAUSTED_WITH_DETAILS for a range host
(shifter-r-157, 2026-07-29), and the live ranges are concentrated in
us-central1-a/-c.
d2614557f adds a RANGE_NETWORK_ZONES pool with per-slot round-robin
placement; a84555d71 allows the key in the provisioner-Job admission policy,
without which admission denies every Job carrying it.
Infrastructure applied
Applied directly rather than through a deploy, after the deploy carrying them
failed on the Redis resize and the event timeline did not permit 25-minute
round trips per attempt:
- Cloud SQL:
db-custom-8-30720, REGIONAL, 100GB (this one did apply via
terraform in 31084576511 before that run failed)
- Redis: 1GB -> 16GB. Terraform failed with "Not enough zonal resources are
available"; the same API call succeeded directly once capacity freed
- Node pools off 1 node each; new
web2 pool on the unused /16
- Replica counts: portal 2 -> 24, guacd 1 -> 120, guacamole-client 1 -> 24,
workers 1 -> 3 each
- Resource requests/limits raised to 1 CPU / 2Gi requests, 2 CPU / 4Gi limits.
guacd had been capped at a 500m CPU limit and portal at a 1Gi memory limit
State that is NOT durable
These are live on the cluster and absent from source. The next Helm apply
reverts them:
terraform.tfvars is the one piece reconciled to reality, so a terraform apply
converges rather than fighting the cluster.
Summary
Record of the platform faults found and corrected while staging the
gcp-devtenant for the KeplerOps CTF (200-300 concurrent ranges). Several were latent
defects that only surfaced under the attempt to scale, and several fixes are
currently live on the cluster but not in source, so they revert on the next
Helm apply. This is a description of what was found and what state it is in.
Deploy runs
310832730379688b3c5031084576511a84555d7131086472277a84555d7131088808310caa1f9517Faults found
1. GKE pod IP range caps the cluster at 16 nodes
The blocking fault. All three node pools (
web,workers,provisioner) arepinned to the
gke-provisioner-podssecondary range (10.46.0.0/20) atmaxPodsPerNode: 110. GKE allocates a /24 per node, so that range holds4096/256 = 16 nodes total across all three pools combined. It reported
utilization: 1.0.The cluster's default pod range
gke-pods(10.44.0.0/16, 256 nodes) wasentirely unused.
Every attempt to add a node failed with:
This does not fail fast. Three
SET_NODE_POOL_SIZEoperations hung for ~20minutes retrying, holding the cluster operation lock so no other cluster
operation could proceed. 83 pods sat Pending.
Resolved by creating node pool
shifter-gcp-dev-web2(36 nodes,e2-standard-4) on thegke-podsrange. Workloads carry no nodeSelectors,tolerations, or hard affinity -- only soft pod anti-affinity -- so pending pods
scheduled onto it without any drain or migration.
The durable fix is to move the original pools onto
gke-pods. Pod range isimmutable on an existing pool, so that means replacing them: a planned
migration, not a deploy-time change.
2. tfvars declared a machine type that forces node pool replacement
terraform.tfvarsdeclarede2-standard-8for web/workers while the live poolsare
e2-standard-4. GKE cannot change machine type in place, so the nextterraform apply would have replaced all three node pools -- during event
staging. Run
31086472277was cancelled for this reason.Corrected in
c5255d3ccandcaa1f9517: machine types match the live pools,node counts total 15 of the 16 the /20 can hold, with the constraint documented
in the file so the counts are not "optimised" back upward.
3. Portal liveness probe turns dependency outages into restart storms
portal-webused the dependency-aware/health/endpoint (Postgres, both Rediscaches, channel layer, file storage) as its liveness probe. During a Cloud
SQL maintenance window the endpoint returned 500, liveness failed, and the
kubelet restarted every replica while they were waiting for the database to come
back. Pods reached
Application startup completeand were sentSIGTERMaboutone second later. The portal served 502 for the duration.
timeoutSeconds: 1on both probes is a secondary problem for a check thatqueries four dependencies.
Filed as #1479. Hotfixed live (liveness ->
tcpSocket, readiness timeout -> 5s);not in the chart, so Helm reverts it.
4. Orphaned
aces-operation-record-pruneDeploymentLeft behind by the ACES -> RAES rename, invoking a management command that no
longer exists. In
CrashLoopBackOfffor 3d21h with 1095 restarts. Supersededby a healthy
raes-operation-record-prunerunning since 2026-07-29, and notdefined anywhere under
platform/.Deleted from the cluster. Filed as #1480, which also covers checking other
environments and the absence of any alert on sustained
CrashLoopBackOff.5. Range access authorized against the wrong instance registry
cms.services._range_accessresolved instance ownership fromcms.models.Instance, but realized range instances are rows ofengine.models.Instance-- the CMS table is written only by NGFW provisioning.Every terminal/RDP/SSH open failed with "Instance not found" before the workspace
check was reached, on every provisioned range.
Fixed in
c17b34795, joining on the sharedrequest_idUUID rather than onprimary key (the two request tables have independent pks).
252b6ba90guards thelookup against null user/request. Both are in
31088808310and were not inthe previously running image.
6. Per-range OpenVPN gateway provisioned unconditionally
Every range on a capable backend whose scenario exposes a single
participant-visible Kali attacker provisions a dedicated gateway VM, an external
address, and per-range VPN secrets, and runs an extra guest-setup step that can
fail. Four of the 11 range VMs currently alive are
vpn-gatewayinstances. At300 ranges that is 300 additional VMs and external IPs for an access path this
event does not use.
348ba69fbputs this behindRANGE_OPENVPN_ENABLED. Note the setting defaultsto true, so the code alone changes nothing -- the environment must set it
false.
7. Range placement pinned to a single zone
Range cells were placed in
RANGE_NETWORK_ZONEonly.us-central1-ahas alreadyreturned
ZONE_RESOURCE_POOL_EXHAUSTED_WITH_DETAILSfor a range host(
shifter-r-157, 2026-07-29), and the live ranges are concentrated inus-central1-a/-c.d2614557fadds aRANGE_NETWORK_ZONESpool with per-slot round-robinplacement;
a84555d71allows the key in the provisioner-Job admission policy,without which admission denies every Job carrying it.
Infrastructure applied
Applied directly rather than through a deploy, after the deploy carrying them
failed on the Redis resize and the event timeline did not permit 25-minute
round trips per attempt:
db-custom-8-30720,REGIONAL, 100GB (this one did apply viaterraform in
31084576511before that run failed)available"; the same API call succeeded directly once capacity freed
web2pool on the unused /16workers 1 -> 3 each
guacd had been capped at a 500m CPU limit and portal at a 1Gi memory limit
State that is NOT durable
These are live on the cluster and absent from source. The next Helm apply
reverts them:
RANGE_NETWORK_ZONESandRANGE_OPENVPN_ENABLED=falsein theplatform-runtimeConfigMap. Neither key is invalues-gcp-dev.yamlruntimeEnvvalues-gcp-dev.yamlbutwere applied by hand because Helm never ran
shifter-gcp-dev-web2, which is not in terraform state at allterraform.tfvarsis the one piece reconciled to reality, so a terraform applyconverges rather than fighting the cluster.