From b887a49e7024ba36a59bac7ee085be3b4cf67d1d Mon Sep 17 00:00:00 2001 From: Matthias Linhuber Date: Fri, 28 Aug 2026 03:14:03 +0200 Subject: [PATCH 1/3] docs: two runbooks and a release trap, all from real incidents Everything here was hit during the 2.1.x rollout and the production cutover. **403 on session launch after a deploy.** A restarted operator recreates the eager-start warm-pool instances and their oauth2-proxy ConfigMaps, rebuilding `authenticated-emails-list` empty. Sessions already assigned to one keep pointing at it, and oauth2-proxy denies everybody when the list is empty. Login succeeds, the session 403s, and nothing is logged. The runbook gives the check that identifies it and the one-command fix. **Certificate stuck with no challenges outstanding.** Distinguishes a transient ACME finalize failure - which cert-manager retries only after an hour of backoff, and which is cleared by patching the status - from a name whose listener or DNS is genuinely wrong. Both look the same from outside. **The release tag trap.** The shared build workflow strips a leading `v`, but an `image-tag` override is used verbatim, so a caller passing the raw tag publishes `v1.2.0`. The spelling that passes the tag-format check is the one that breaks the images, which is what kept it hidden. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_019qeiQRFu8xAMRYWPdZewjG --- docs/admins/maintenance/release-policy.md | 12 +++++ docs/admins/operations/incident-response.md | 59 +++++++++++++++++++++ 2 files changed, 71 insertions(+) diff --git a/docs/admins/maintenance/release-policy.md b/docs/admins/maintenance/release-policy.md index 9ee3c38..81e77b5 100644 --- a/docs/admins/maintenance/release-policy.md +++ b/docs/admins/maintenance/release-policy.md @@ -28,6 +28,18 @@ Three spellings of a git tag existed historically — `1.1.0`, `v1.1.0` and this image". A shared CI check now rejects anything that is not `vX.Y.Z`. Old tags are not retagged. +:::warning A caller can defeat the `v` stripping +The shared build workflow derives the image tag with `${RELEASE_TAG#v}`, but an +`image-tag` **override is used verbatim**. A repository that passes +`github.event.release.tag_name` as that override therefore publishes `v1.2.0`, +which no chart can consume, while a tag spelled `1.2.0` produces the right image +and merely fails the tag-format check. + +That combination hides the fault: the spelling that passes CI is the one that +breaks the images. If a release publishes `v`-prefixed images, this is why. +Callers should not pass `image-tag` on a release event at all. +::: + ## Cutting a release The release train in EduIDE-Helm does it, and defaults to `dry_run: true`. diff --git a/docs/admins/operations/incident-response.md b/docs/admins/operations/incident-response.md index 00503fd..16b2565 100644 --- a/docs/admins/operations/incident-response.md +++ b/docs/admins/operations/incident-response.md @@ -97,6 +97,34 @@ This indicates the session image is unavailable. Verify the image tag in the App --- +## Runbook: 403 on session launch after a deploy + +**Symptoms:** Login at Keycloak succeeds, then the session URL returns **403 Forbidden**. Only some users or some sessions are affected. Nothing appears in the operator or service logs, and `status.operatorMessage` on the Session is empty. + +**Cause:** the session was assigned a **warm-pool instance**, and a deploy restarted the operator, which recreated those instances and their oauth2-proxy ConfigMaps. Each instance's `authenticated-emails-list` was rebuilt empty. The Session object still points at the instance, but oauth2-proxy denies every authenticated user when that list is empty. + +**Confirm it:** + +```bash +kubectl -n get cm -o name | grep -- '-email' | \ + xargs -n1 kubectl -n get -o \ + jsonpath='{.metadata.name}={.data.authenticated-emails-list}{"\n"}' +``` + +An affected instance shows an empty list. A session with its own dedicated pod shows the user's address and works, which is the contrast that identifies this. + +**Fix:** delete the orphaned Session so a fresh one is issued. + +```bash +kubectl -n delete sessions.theia.cloud +``` + +The user then launches again normally. Their workspace volume is untouched. + +:::note Affects every installation +`eagerStart: true` is the default, so any deploy landing while a user holds a session on a warm instance can cause this. Tracked in EduIDE-Cloud issue 135. +::: + ## Runbook: Authentication outage **Symptoms:** All users are redirected to Keycloak but cannot log in, or receive "Access Denied" after successful login. @@ -335,6 +363,37 @@ kubectl describe certificate -n eduide-system --- +## Runbook: Certificate stuck, no challenges outstanding + +**Symptoms:** a `Certificate` sits `Ready=False`, its listener never programs, and there are no `Challenge` resources left to look at. + +**Check the order:** + +```bash +kubectl -n eduide-system get certificate \ + -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.message}{"\n"}{end}' +``` + +Two different situations look the same from outside. + +**Finalize failed.** The challenges validated and the CA then failed the final step: + +``` +Failed to finalize Order: 404 urn:ietf:params:acme:error:malformed: +Certificate not found +``` + +This is transient. cert-manager will retry, but only after an exponential backoff starting at an hour. Clear the backoff to retry now: + +```bash +kubectl -n eduide-system patch certificate --type=merge --subresource=status \ + -p '{"status":{"lastFailureTime":null,"failedIssuanceAttempts":null}}' +``` + +It normally issues within a minute. Nothing is wrong with the configuration. + +**Challenges pending and staying pending.** That is a real fault: the hostname has no listener on the Gateway, or DNS for it does not reach the Gateway's address. A name on a certificate with no listener answers 404 to its HTTP-01 challenge and blocks the whole certificate, including every other name on it. + ## Runbook: Storage exhaustion **Symptoms:** New workspace creation fails with storage errors. Existing sessions are unaffected. From 324a4111b2834efbb49aadd8e81c3ed9ca4af762 Mon Sep 17 00:00:00 2001 From: Matthias Linhuber Date: Fri, 28 Aug 2026 14:30:03 +0200 Subject: [PATCH 2/3] docs: correct the 403 runbook - a generation change wipes it, not a restart Verifying the fix on a test environment showed the mechanism I had written down was wrong. `ensureCapacity` only creates ConfigMaps that are missing, so restarting the operator with everything present is a no-op - confirmed empirically against the old build. The wipe comes from `reconcile`, where the instance ConfigMaps are treated as outdated when the AppDefinition's metadata.generation changes, which is what a helm upgrade does. The user-facing symptom is unchanged, and 'it broke after a deploy' still holds. The distinction matters for anyone reading the operator, and for deciding whether a plain restart is a safe thing to try. Also records that the operator fix is merged, and that a broken installation heals within about twenty seconds of the fixed operator starting - verified against a real reproduction rather than a simulated one. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_019qeiQRFu8xAMRYWPdZewjG --- docs/admins/operations/incident-response.md | 10 +++++++--- 1 file changed, 7 insertions(+), 3 deletions(-) diff --git a/docs/admins/operations/incident-response.md b/docs/admins/operations/incident-response.md index 16b2565..bee74cb 100644 --- a/docs/admins/operations/incident-response.md +++ b/docs/admins/operations/incident-response.md @@ -101,7 +101,9 @@ This indicates the session image is unavailable. Verify the image tag in the App **Symptoms:** Login at Keycloak succeeds, then the session URL returns **403 Forbidden**. Only some users or some sessions are affected. Nothing appears in the operator or service logs, and `status.operatorMessage` on the Session is empty. -**Cause:** the session was assigned a **warm-pool instance**, and a deploy restarted the operator, which recreated those instances and their oauth2-proxy ConfigMaps. Each instance's `authenticated-emails-list` was rebuilt empty. The Session object still points at the instance, but oauth2-proxy denies every authenticated user when that list is empty. +**Cause:** the session was assigned a **warm-pool instance**, and a deploy changed the AppDefinition's `metadata.generation`. The operator treats the instance's oauth2-proxy ConfigMaps as outdated and recreates them, so each instance's `authenticated-emails-list` is rebuilt empty. The Session object still points at the instance, but oauth2-proxy denies every authenticated user when that list is empty. + +A restart of the operator on its own does **not** cause this - it only creates ConfigMaps that are missing. The trigger is the generation change that a helm upgrade produces, which is why it presents as "broken by the last deploy". **Confirm it:** @@ -121,8 +123,10 @@ kubectl -n delete sessions.theia.cloud The user then launches again normally. Their workspace volume is untouched. -:::note Affects every installation -`eagerStart: true` is the default, so any deploy landing while a user holds a session on a warm instance can cause this. Tracked in EduIDE-Cloud issue 135. +:::note Fixed in the operator from 2026-08-28 +`eagerStart: true` is the default, so any deploy landing while a user holds a session on a warm instance could cause this. EduIDE-Cloud issue 135 fixed it: the operator now writes the claiming session's user back into the instance's ConfigMap during the same reconcile that recreates it. + +Verified on a test environment against a real reproduction: an installation broken by the old operator was **healed within about twenty seconds** of the fixed one starting, with no manual intervention. If you are running an operator from before that date, this runbook still applies and upgrading is the fix. ::: ## Runbook: Authentication outage From 5b860e6cecd04bd60c0d33eb4055ee41e37b6273 Mon Sep 17 00:00:00 2001 From: Matthias Linhuber Date: Fri, 28 Aug 2026 18:54:07 +0200 Subject: [PATCH 3/3] docs: address review feedback on the certificate runbook Three findings, all valid. The runbook went straight to patching away cert-manager's backoff. It now says to confirm the finalize 404 by reading down the chain the Certificate owns - CertificateRequest, Order, Challenge - since the Certificate's own conditions rarely say enough, and to apply the patch only for that confirmed failure. Notes that --subresource needs kubectl v1.24 or newer, and says what to do when the cleared retry fails too, rather than asserting the cause is always transient. Pending challenges are now split by what the Challenge actually reports. An HTTP 404 means something answered but did not route the solver path; a DNS failure, refused connection or timeout means nothing answered at all, and no amount of listener configuration helps until the challenge reaches the cluster. The previous text ran the two together. Also tags the error fence, which markdownlint flagged as MD040. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_019qeiQRFu8xAMRYWPdZewjG --- docs/admins/operations/incident-response.md | 20 +++++++++++++------- 1 file changed, 13 insertions(+), 7 deletions(-) diff --git a/docs/admins/operations/incident-response.md b/docs/admins/operations/incident-response.md index bee74cb..3ae8b8e 100644 --- a/docs/admins/operations/incident-response.md +++ b/docs/admins/operations/incident-response.md @@ -371,32 +371,38 @@ kubectl describe certificate -n eduide-system **Symptoms:** a `Certificate` sits `Ready=False`, its listener never programs, and there are no `Challenge` resources left to look at. -**Check the order:** +**Check the order.** The `Certificate`'s own conditions rarely say enough - a `Certificate` owns a `CertificateRequest`, which owns an `Order`, which owns the `Challenge`s, and the real error is usually further down that chain than you started: ```bash -kubectl -n eduide-system get certificate \ - -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.message}{"\n"}{end}' +kubectl -n eduide-system describe certificate +kubectl -n eduide-system get certificaterequest,order,challenge +kubectl -n eduide-system describe order ``` Two different situations look the same from outside. **Finalize failed.** The challenges validated and the CA then failed the final step: -``` +```text Failed to finalize Order: 404 urn:ietf:params:acme:error:malformed: Certificate not found ``` -This is transient. cert-manager will retry, but only after an exponential backoff starting at an hour. Clear the backoff to retry now: +This is transient. cert-manager will retry, but only after an exponential backoff starting at an hour. Clear the backoff to retry now - only once you have read that exact message off the `Order`, since the patch below does nothing for any other cause and just hides how long the certificate has been failing: ```bash kubectl -n eduide-system patch certificate --type=merge --subresource=status \ -p '{"status":{"lastFailureTime":null,"failedIssuanceAttempts":null}}' ``` -It normally issues within a minute. Nothing is wrong with the configuration. +`--subresource` requires kubectl v1.24 or newer; on older clients the flag is silently unavailable and the patch will not apply to status. + +It normally issues within a minute. If a cleared retry fails again, it was not transient after all - work back down the chain above, and check the `ClusterIssuer` and the cert-manager controller logs before changing any configuration. + +**Challenges pending and staying pending.** That is a real fault, and the error recorded on the `Challenge` distinguishes the two causes: -**Challenges pending and staying pending.** That is a real fault: the hostname has no listener on the Gateway, or DNS for it does not reach the Gateway's address. A name on a certificate with no listener answers 404 to its HTTP-01 challenge and blocks the whole certificate, including every other name on it. +- **HTTP 404.** Something answered on port 80 but did not route the `/.well-known/acme-challenge/` path. The hostname has no listener on the Gateway, or its listener is on a Gateway the solver route does not attach to. A name on a certificate with no listener blocks the whole certificate, including every other name on it. +- **DNS failure, connection refused or timeout.** Nothing answered at all: the hostname does not resolve, or it resolves to an address that is not this Gateway. Fix DNS first - no amount of listener configuration helps until the challenge reaches the cluster. ## Runbook: Storage exhaustion