diff --git a/docs/admins/install/certificates.md b/docs/admins/install/certificates.md new file mode 100644 index 0000000..5a7daa4 --- /dev/null +++ b/docs/admins/install/certificates.md @@ -0,0 +1,214 @@ +--- +title: Certificates and DNS +description: The four hostnames every installation needs, why one of them must be a wildcard, and the two ways to get a certificate for it. +--- + +# Certificates and DNS + +Read this before you install. The wildcard requirement below is the single thing +most likely to delay a new installation, because it often needs a request to a +DNS or PKI team that takes days. + +## Four hostnames per installation + +For an installation whose landing page is at `eduide.example.edu`: + +| Hostname | Serves | +|---|---| +| `eduide.example.edu` | the landing page | +| `service.eduide.example.edu` | the REST service the landing page calls | +| `instance.eduide.example.edu` | session ingress | +| `*.webview.instance.eduide.example.edu` | **per-session webviews** | + +All four need DNS pointing at your Gateway's external address. The fourth is a +**wildcard record**, because every session gets its own subdomain under it. + +:::caution Check your DNS policy first +Some institutions do not hand out wildcard records, or require a specific +approval. Find out before you plan the rest of the install — there is no +configuration that avoids the wildcard if you want webviews. +::: + +## Why the wildcard exists + +Inside the IDE, anything rendered in a panel — a Markdown preview, a notebook, a +rendered PDF, embedded documentation — is served from its own origin so that it +cannot script against the IDE itself. Those origins are +`.webview.instance.`. + +Without the wildcard, sessions still start and the IDE still loads. **Only the +previews break**, with an opaque failure inside the IDE. That means this is +usually discovered by a student weeks after go-live rather than by you during +installation. + +## Why ACME cannot issue it over HTTP-01 + +This is the part that catches people out. + +cert-manager will happily issue certificates for the first three hostnames using +an **HTTP-01** challenge: it serves a token over plain HTTP on port 80 at that +exact hostname, and the CA fetches it. + +That cannot work for a wildcard. To prove control of `*.webview.instance.` +you would have to serve a token at every possible name under it, which is +infinite. **The ACME specification therefore does not permit HTTP-01 for +wildcards at all** — it is not a cert-manager limitation and no configuration +changes it. + +You have two options. + +## Option A — DNS-01 (recommended) + +With a **DNS-01** challenge, cert-manager proves control by writing a TXT record +into your zone, which works for a wildcard because it proves control of the +whole zone. This is the better answer: cert-manager then issues *and renews* the +wildcard automatically, and you never think about it again. + +It needs an API credential for your DNS zone. cert-manager has built-in support +for Route53, CloudDNS, AzureDNS, Cloudflare, DigitalOcean, ACME-DNS and RFC-2136 +dynamic updates, plus community webhook providers for many university DNS +systems. + +**Ask your DNS team for a scoped API credential before assuming this is +impossible.** RFC-2136 works with BIND, which many universities run. + +Store the credential as a Secret, then set `gatewayAcmeIssuer.solvers` in the +cluster chart values: + +```yaml +gatewayAcmeIssuer: + enabled: true + email: platform@example.edu + solvers: + # DNS-01 for the wildcard + - dns01: + cloudflare: + apiTokenSecretRef: { name: cloudflare-api-token, key: token } + selector: + dnsNames: ["*.webview.instance.eduide.example.edu"] + # HTTP-01 for everything else + - http01: + gatewayHTTPRoute: + parentRefs: + - group: gateway.networking.k8s.io + kind: Gateway + name: theia-shared-gateway + namespace: eduide-system +``` + +Then simply list the wildcard alongside the other names: + +```yaml +managedCertificates: + enabled: true + certificates: + - name: eduide-tls + secretName: eduide-tls + dnsNames: + - eduide.example.edu + - service.eduide.example.edu + - instance.eduide.example.edu + - "*.webview.instance.eduide.example.edu" +``` + +## Option B — bring your own wildcard certificate + +If you cannot get a DNS-01 credential, obtain the wildcard some other way — a +commercial CA, or your institution's certificate service — and hand it to the +cluster chart. + +```bash +cat wildcard.crt | base64 | tr -d '\n' # -> certificate +cat wildcard.key | base64 | tr -d '\n' # -> key +``` + +```yaml +wildcardTLSSecret: + create: true + name: eduide-webview-tls + certificate: "" + key: "" +``` + +and point the webview listener at it: + +```yaml +gateway: + listeners: + - name: prod-webview + hostname: "*.webview.instance.eduide.example.edu" + tlsSecretName: eduide-webview-tls +``` + +The trade-off is that **this does not renew itself.** Put the expiry date in +whatever your team uses to track such things. A wildcard that expires takes out +every webview at once. + +## Nothing tells you when a certificate is wrong + +This is worth internalising, because it has already cost one installation six +months. + +Gateway API **never compares a certificate's names against the listener's +hostname.** A listener whose Secret holds a certificate for an entirely +different host reports: + +``` +Programmed=True Accepted=True ResolvedRefs=True +``` + +Everything looks healthy. The first symptom is a browser certificate warning. +The second is subtler and much more confusing: the landing page loads (the user +clicks through the warning), then its JavaScript calls `service.` — a +different origin — and the browser **silently blocks that request** because that +certificate is invalid too. The launch never reaches the server, so there is +nothing in any log to find. + +**Verify explicitly, and never with `curl -k`:** + +```bash +for h in eduide.example.edu service.eduide.example.edu instance.eduide.example.edu; do + echo | openssl s_client -connect "$h:443" -servername "$h" 2>/dev/null \ + | openssl x509 -noout -checkhost "$h" +done +``` + +Each should print `Host matches certificate`. Then check the wildcard by +testing a name under it: + +```bash +h=probe.webview.instance.eduide.example.edu +echo | openssl s_client -connect "$h:443" -servername "$h" 2>/dev/null \ + | openssl x509 -noout -checkhost "$h" +``` + +And confirm a browser would accept it — `%{ssl_verify_result}` must be `0`: + +```bash +curl -s -o /dev/null -w '%{http_code} %{ssl_verify_result}\n' https://eduide.example.edu +``` + +## Adding an installation later + +When you add a second installation to a cluster, its three non-wildcard names +must be added to the certificate, and it needs its own webview wildcard. + +A certificate that covers your existing installations but not the new one will +not fail anything visibly — see above. Re-check with the commands in this page +after every change. + +:::tip Only put names with a listener on a certificate +cert-manager proves each name with its own challenge. A name on the certificate +that has no matching Gateway listener will fail its HTTP-01 challenge with a +404, and that one pending challenge **blocks the whole certificate** — including +every name that would otherwise have worked. +::: + +## Checklist + +- [ ] Four DNS records per installation, one of them a wildcard +- [ ] DNS policy permits wildcards, or Option B is agreed +- [ ] A DNS-01 credential obtained, or a wildcard certificate obtained and its + expiry tracked +- [ ] Every hostname verified with `openssl ... -checkhost`, not `curl -k` +- [ ] `ssl_verify_result` is `0` for the landing page **and** the service host diff --git a/docs/admins/install/prerequisites.md b/docs/admins/install/prerequisites.md new file mode 100644 index 0000000..3f77503 --- /dev/null +++ b/docs/admins/install/prerequisites.md @@ -0,0 +1,169 @@ +--- +title: Cluster Prerequisites +description: Everything that must exist on the cluster before EduIDE is installed, with the commands to put it there. +--- + +# Cluster Prerequisites + +EduIDE does not install its own platform layer. Work through this page first; +`helm install` will otherwise fail, or — worse — succeed and not work. + +Everything here is installed **once per cluster**, not once per installation. + +## What you need before you start + +| | | +|---|---| +| **Kubernetes** | 1.26 or later. Gateway API `v1` and the CRD conversion webhooks both need it | +| **Cluster-admin** | You will install CRDs, ClusterRoles and ClusterIssuers | +| **A LoadBalancer** | Something must give the Gateway an external address — a cloud provider's controller, MetalLB, or equivalent | +| **DNS you can change** | Four records per installation, **one of them a wildcard**. See [Certificates and DNS](certificates.md) — check early that your DNS team permits wildcards | +| **An OIDC provider** | Keycloak is what EduIDE is tested against | + +## 1. Gateway API CRDs + +EduIDE routes every session through Gateway API. The CRDs are not part of +Kubernetes and must be installed separately. + +```bash +kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.2.1/standard-install.yaml +``` + +Pin the version deliberately rather than tracking latest — a Gateway API major +version bump is a breaking change. + +```bash +kubectl get crd gateways.gateway.networking.k8s.io # should exist +``` + +## 2. A Gateway controller + +The CRDs above are just types; something has to implement them. EduIDE is +developed and tested against **Envoy Gateway**. + +```bash +helm install eg oci://docker.io/envoyproxy/gateway-helm \ + --version v1.2.1 -n envoy-gateway-system --create-namespace +kubectl -n envoy-gateway-system rollout status deploy/envoy-gateway +``` + +Any conformant controller should work, but nothing else has been tried. If you +use another, expect to adjust `gateway.className` in the cluster chart. + +```bash +kubectl get gatewayclass # you should see one, and it should be Accepted +``` + +## 3. cert-manager — with Gateway API support turned on + +```bash +helm repo add jetstack https://charts.jetstack.io && helm repo update +helm install cert-manager jetstack/cert-manager \ + --namespace cert-manager --create-namespace \ + --version v1.16.2 \ + --set crds.enabled=true \ + --set config.apiVersion="controller.config.cert-manager.io/v1alpha1" \ + --set config.kind="ControllerConfiguration" \ + --set config.enableGatewayAPI=true +``` + +:::warning `enableGatewayAPI` is not optional +Without it, cert-manager ignores Gateway resources entirely and your HTTP-01 +challenges never complete. If you installed cert-manager earlier without this +flag, upgrade with it and restart the controller: + +```bash +kubectl -n cert-manager rollout restart deploy/cert-manager +``` +::: + +cert-manager also provides the CA injection that EduIDE's CRD conversion webhook +depends on, so it is required even if you bring your own certificates. + +```bash +kubectl -n cert-manager get pods # controller, webhook and cainjector all Running +``` + +## 4. A storage class + +Every session that is not ephemeral gets a PersistentVolumeClaim. + +```bash +kubectl get storageclass +``` + +Requirements: + +- **ReadWriteOnce** is sufficient; sessions do not share volumes. +- **Dynamic provisioning** is required — volumes are created on demand as + students start sessions. +- Volume expansion is not required but is useful. + +Note the name; you will set it as `operator.storageClassName`. Do not rely on +the cluster's default class being the one you want. + +:::caution More than one default class +If two storage classes are both marked default, a PVC that omits the class gets +an arbitrary one, and you will get inconsistent behaviour that is hard to +attribute. Check with: + +```bash +kubectl get storageclass -o custom-columns=NAME:.metadata.name,DEFAULT:.metadata.annotations.'storageclass\.kubernetes\.io/is-default-class' +``` +::: + +## 5. Node disk for preloaded images + +EduIDE preloads every IDE image onto **every node**, so a session starts in +seconds rather than waiting on a multi-gigabyte pull. + +Budget accordingly: each IDE image is roughly 0.5–1 GB, and the default set is +eight. Allow **at least 20 GB** of image storage per node, and more if you add +languages. Trim the app list in the tenant values if that is too much — see +[Applications](../platform/app-definitions.md). + +## 6. Prometheus Operator — only if you want metrics + +EduIDE ships PodMonitors and Grafana dashboards, and they are **off by default** +because they are not portable: they need the Prometheus Operator CRDs, and the +namespaces they target differ per monitoring stack. + +```bash +kubectl get crd podmonitors.monitoring.coreos.com +``` + +If that is absent, leave `monitoring.enabled: false`. Everything else works +without it. + +## 7. An OIDC provider + +EduIDE authenticates through OIDC, tested against Keycloak. You need: + +- A realm your students can authenticate to. +- A **public** client — EduIDE's landing page is a browser application and holds + no secret. +- Redirect URIs for **all four** of the installation's hostnames. Three is a + common mistake and breaks webviews specifically. + +See [Access Control](../platform/access-control.md) for the full client setup. + +You can defer this: an installation can be brought up with +`keycloak.allowUnauthenticated: true` to prove the platform works before wiring +identity. **That installation has no authentication and must not be exposed.** + +## Checklist + +Before installing EduIDE: + +- [ ] Kubernetes 1.26+, cluster-admin +- [ ] Gateway API CRDs installed, pinned +- [ ] A Gateway controller running, GatewayClass `Accepted` +- [ ] cert-manager running, **with `enableGatewayAPI=true`** +- [ ] A storage class chosen, with dynamic provisioning +- [ ] ~20 GB of image space per node +- [ ] DNS you can change, **wildcards permitted** +- [ ] A certificate plan — read [Certificates and DNS](certificates.md) before + going further, because the wildcard requirement surprises people +- [ ] An OIDC realm and public client, or a deliberate decision to defer + +Then continue to [Installing EduIDE](installing.md). diff --git a/docs/admins/intro.md b/docs/admins/intro.md index fa3d70e..31e61d7 100644 --- a/docs/admins/intro.md +++ b/docs/admins/intro.md @@ -1,70 +1,88 @@ --- title: Admin Overview -description: Entry point for platform operators and administrators of EduIDE. +description: Entry point for anyone deploying and operating EduIDE, at TUM or anywhere else. --- # Admin Overview -This section is for platform operators and administrators responsible for deploying, configuring, and maintaining EduIDE. +This section is for whoever runs EduIDE: installing it, keeping it up, and +handing out access. It assumes cluster-admin on a Kubernetes cluster and admin +rights in an identity provider. -Admin concerns are kept separate from instructor, student, and developer sections because they involve cluster-level access, service secrets, identity provider configuration, and infrastructure lifecycle decisions that are irrelevant to end users and should not be mixed with usage or development guidance. +EduIDE is not TUM-specific software. It is installed the same way anywhere, with +two Helm charts and no dependency on TUM's automation. Where a page uses a TUM +hostname or realm name it is an example, and is marked as one. -:::info TUM-specific deployment details -This page sometimes uses TUM-specific environment names, domains, and operational examples to describe the current EduIDE rollout. EduIDE itself is an independent product and can be deployed and operated for any institution with its own infrastructure and policies. -::: - -## Who this section is for - -This section assumes you have: +If you are building on or extending EduIDE rather than running it, see the +[Developer](/developer/intro) section. -- Kubernetes cluster access (`kubectl` and Helm) -- Admin permissions in the Keycloak realm -- Access to the deployment repository and GitHub environment secrets -- Familiarity with the three EduIDE deployment environments (production, staging, test) +## Installing it for the first time -If you are an engineer building on or extending EduIDE, see the [Developer](/developer/intro) section instead. +In this order. The first two are the ones that surprise people. -## What lives here +1. **[Cluster Prerequisites](install/prerequisites.md)** — Gateway API, + a Gateway controller, cert-manager with Gateway support, storage, node disk. +2. **[Certificates and DNS](install/certificates.md)** — four hostnames per + installation, one of them a wildcard that ACME cannot issue over HTTP-01. + **Read this early**; it often needs a request to another team. +3. **[Installing EduIDE](install/installing.md)** — the two charts. +4. **[Access Control](platform/access-control.md)** — the identity provider. +5. **[Adding an Installation](install/adding-an-installation.md)** — for the + second and subsequent ones on the same cluster. -### Platform +## How it is put together -Core setup and ongoing configuration of the platform. +Two charts, because the two halves have different cardinality. -- [Provisioning](/admins/platform/provisioning) — bootstrapping a new environment from cluster prerequisites through first launch -- [Access Control](/admins/platform/access-control) — Keycloak client setup, admin group assignment, and access reviews -- [App Definitions](/admins/platform/app-definitions) — managing launchable IDE session types and their scaling parameters -- [Storage and Quotas](/admins/platform/storage-and-quotas) — persistent volume sizing, storage classes, and per-namespace resource limits +| Chart | Installed | Owns | +|---|---|---| +| `eduide-cluster` | once per **cluster**, into `eduide-system` | CRDs, the conversion webhook, ClusterRoles, cert-manager issuers, the shared Gateway, PodMonitors and dashboards | +| `eduide` | once per **installation**, into its own namespace | the operator, the REST service, the landing page, routes, app definitions, image preloading | -### Operations +An **installation** is one namespace on one cluster with its own hostnames, +branding and identity provider. There is no fixed set of environments: you might +run one, or a test and a production one, or one per department. Namespaces are +conventionally `eduide-`. -Keeping the platform running and responding when it does not. +Sessions are routed with **Gateway API** — a shared Gateway per cluster, and an +HTTPRoute per installation attaching to it by listener name. -- [Monitoring Basics](/admins/operations/monitoring-basics) — signals, dashboards, alert thresholds, and health check cadence -- [Incident Response](/admins/operations/incident-response) — runbooks for the most common incident classes -- [Session Management](/admins/operations/session-management) — admin-level oversight of active and stuck sessions -- [Garbage Collection](/admins/operations/garbage-collection) — workspace TTL configuration and cleanup operations - -### Security - -Admin API protection and compliance practices. +:::note The `theia.cloud` API group +The custom resources are `appdefinitions.theia.cloud`, `sessions.theia.cloud` +and `workspaces.theia.cloud`. The `theia.cloud` group name is historical and has +not been renamed, because renaming an API group is a migration. Do not grep for +`eduide` and conclude the CRDs are missing. +::: -- [Admin API Tokens](/admins/security/admin-api-tokens) — token issuance, rotation, and request authentication -- [Audit and Compliance](/admins/security/audit-and-compliance) — what is logged, retention policy, and access review checklist +## Running it -### Maintenance +| | | +|---|---| +| [Session Management](operations/session-management.md) | Sessions, limits, and what times them out | +| [Storage and Quotas](platform/storage-and-quotas.md) | Volumes, sizing, namespace quotas | +| [Garbage Collection](operations/garbage-collection.md) | How old workspaces are reclaimed | +| [App Definitions](platform/app-definitions.md) | Which languages and templates are offered | +| [Monitoring](operations/monitoring-basics.md) | Metrics, and what EduIDE actually exposes | +| [Incident Response](operations/incident-response.md) | Runbooks, including routing and certificate failures | -Planned operational procedures. +## Changing versions -- [Upgrades](/admins/maintenance/upgrades) — upgrading the EduIDE Cloud service, operator, and supporting charts +| | | +|---|---| +| [Release and Version Policy](maintenance/release-policy.md) | What a version number means here | +| [Upgrades](maintenance/upgrades.md) | Moving an installation to a new version | +| [Rollback](maintenance/rollback.md) | Undoing a deploy that succeeded and behaves badly | -## Environments +## Security -EduIDE runs three deployment environments: +| | | +|---|---| +| [Access Control](platform/access-control.md) | The identity provider and the session proxy | +| [Admin API Tokens](security/admin-api-tokens.md) | The bearer token for the scaling API | +| [Audit and Compliance](security/audit-and-compliance.md) | What is logged, retention, decommissioning | -| Environment | Namespace | Domain | Deploy trigger | -|---|---|---|---| -| Production | `theia-prod` | `theia.artemis.cit.tum.de` | Manual with approval | -| Staging | `theia-staging` | `theia-staging.artemis.cit.tum.de` | Push to main | -| Test | `test1` | `test1.theia-test.artemis.cit.tum.de` | PR push with approval | +## Deployment automation -Most admin procedures apply to all three environments. Where behavior differs, it is called out explicitly. +[Provisioning](platform/provisioning.md) describes the GitHub Actions pipeline +TUM uses. You do not need it — the charts install identically by hand — but it +is a working example if you want to automate the same way. diff --git a/docs/admins/maintenance/upgrades.md b/docs/admins/maintenance/upgrades.md index 9e035b2..d5e3a52 100644 --- a/docs/admins/maintenance/upgrades.md +++ b/docs/admins/maintenance/upgrades.md @@ -1,127 +1,107 @@ --- title: Upgrades -description: Procedures for upgrading EduIDE Cloud charts, the operator, and supporting components. +description: Moving an installation to a new EduIDE version. --- # Upgrades -EduIDE releases new versions approximately every three months. This page covers how to upgrade the platform components, what to check before and after an upgrade, and how to roll back if needed. +An upgrade is a version change and a `helm upgrade`. There is no multi-chart +ordering to remember any more — there are two charts, and only one of them is +usually involved. -## Release cadence and versioning +## What you are upgrading -EduIDE uses a three-month release cycle. Between releases, chart versions carry a `-next.X` suffix (e.g., `1.2.0-next.2`). At release time: -- The `-next` suffix is dropped. -- All image tags are pinned to specific SHAs. -- After release, the version is bumped and `-next.0` is added for the next cycle. +| Chart | When it changes | Blast radius | +|---|---|---| +| `eduide` | Most upgrades | One installation | +| `eduide-cluster` | CRDs, the conversion webhook, Gateway or issuer changes | **Every installation on the cluster** | -In production, always deploy pinned release versions, never floating tags like `latest`. +Both carry the same version. Read the release notes to see whether the cluster +chart changed; if it did, upgrade it first and expect it to affect everyone. -## How upgrades are deployed +## Before -Upgrades go through the same GitHub Actions pipelines as all other deployments. To upgrade an environment: +- Read the release notes, particularly for CRD changes. +- Check what you are running: `helm list -n `. +- Know how to go back: [Rollback](rollback.md). +- On a cluster with more than one installation, upgrade a test one first. -1. Update the chart versions and image tags in the environment's values files in the deployment repository. -2. Commit and push (staging) or trigger the workflow manually (production). -3. The pipeline runs all chart upgrades in the correct order and waits for each rollout to complete. - -**Always upgrade staging before production.** The staging pipeline runs automatically on push to main, making it the natural first target. - -The underlying commands the pipeline executes are documented in the steps below — they are useful for understanding what happens and for emergency manual intervention when the pipeline is unavailable. - -## Pre-upgrade checklist - -Before merging an upgrade to production: - -- [ ] The upgrade has been deployed to staging and verified stable. -- [ ] Release notes have been checked for breaking changes, especially CRD schema changes. -- [ ] No large cohort exercise is in progress (an upgrade restarts pods, briefly queuing launches). -- [ ] Current chart versions are noted: `helm list -n theia-prod` -- [ ] Values files are committed and reviewed. - -## Upgrade order - -The pipeline respects this order. CRDs and cluster-scoped resources must come before environment-scoped resources. - -1. `theia-cloud-crds` -2. `theia-cloud-base` -3. `theia-cloud-combined` (service + operator + landing page) -4. `theia-appdefinitions` -5. `theia-certificates` -6. `theia-monitoring` -7. `theia-workspace-garbage-collector` - -## CRD upgrades - -CRDs must be upgraded before the operator because the operator validates resources against the CRD schema at startup. - -After any CRD upgrade, verify that existing custom resources are still valid before proceeding: +## Upgrading an installation ```bash -kubectl get workspaces -n theia-prod -kubectl get sessions -n theia-prod +helm upgrade eduide oci://ghcr.io/eduide/charts/eduide \ + --version -n \ + -f values.yaml -f secrets.yaml \ + --wait --timeout 15m ``` -If resources show validation errors, do not proceed until they are resolved. - -## App Definition image upgrades +Keep the same release name and the same values files. Changing the release name +is not an upgrade — Helm treats it as a new install and will refuse to adopt the +existing resources. -When new IDE images are released, update the image tags in the environment's `appdefinitions.yaml` values file and deploy via the pipeline. +:::tip Preview it first +`helm diff upgrade` (from the helm-diff plugin) shows what will change before +anything is applied. On a live installation this is worth the extra step. +::: -After the upgrade, the operator replaces pre-warmed sessions with the new image. **Existing running sessions are not affected** — they continue on the old image until they are stopped and restarted by the user. +### `--wait` and image preloading -## Checking rollout status +The preloading DaemonSet pulls every IDE image onto every node. On an upgrade +that changes the image set, that can take longer than a sensible Helm timeout. -After the pipeline completes, verify the rollout: +Either allow a generous `--timeout`, or drop `--wait` and watch the rollout +yourself: ```bash -kubectl rollout status deployment/operator -n theia-prod -kubectl rollout status deployment/service -n theia-prod -kubectl get pods -n theia-prod +kubectl -n rollout status deploy/operator-deployment +kubectl -n rollout status deploy/service-deployment +kubectl -n rollout status deploy/landing-page-deployment ``` -## Post-upgrade validation - -- [ ] All pods in `theia-prod` are `Running`: `kubectl get pods -n theia-prod` -- [ ] Service admin ping responds: `GET /service/admin/{appId}` with `X-Admin-Api-Token` -- [ ] A test session can be launched and connects to the IDE -- [ ] Grafana dashboards show data for the namespace -- [ ] App Definition scaling values are intact: `GET /service/admin/appdefinition` - -## Rolling back +Note the `-deployment` suffix; the Deployments are not named `operator` or +`service`. -If an upgrade introduces a regression, roll back via Helm: +## Upgrading the cluster chart ```bash -# Check release history -helm history theia-cloud -n theia-prod - -# Roll back to a previous revision -helm rollback theia-cloud -n theia-prod +helm upgrade eduide-cluster oci://ghcr.io/eduide/charts/eduide-cluster \ + --version -n eduide-system \ + -f cluster-values.yaml --wait ``` -For CRD rollbacks, Helm does not manage this automatically. If a CRD schema change is incompatible with running resources, you may need to manually apply the previous CRD version from the deployment repository: - -```bash -kubectl apply -f charts/theia-cloud-crds/templates/ -``` +CRDs are ordinary templates in this chart rather than files in `crds/`, which +means `helm upgrade` genuinely updates them — unlike many charts, where CRD +changes need a manual `kubectl apply`. -This is rare and only necessary when the CRD schema changed in a breaking way. +That cuts both ways: a CRD change lands the moment you upgrade, for every +installation on the cluster at once. **A release that changes a stored CRD +version cannot be undone by rolling back the tenant chart** — see +[Rollback](rollback.md). -## Upgrading Keycloak +## After -Keycloak is managed separately from the EduIDE charts. If a Keycloak upgrade is required: +- All four Deployments Ready, no pod in `ImagePullBackOff`. +- The landing page returns 200 **and its TLS validates without `-k`**. +- Start a real session and open something in a webview panel. Everything above + can be healthy while sessions do not start. +- If the AppDefinition set changed, check the landing page offers what you + expect. -1. Coordinate with the identity provider admin team. -2. Test authentication in the staging environment after the upgrade. -3. Verify that all three token claim mappers (username, audience, groups) still work correctly. -4. Confirm OAuth2 proxy compatibility with the new Keycloak version. +## Two things that do not behave as you would guess -## Image tag pinning +**`minInstances` and `maxInstances` stop taking effect after the first install.** +The chart deliberately reads the live values back so that Helm does not reset +scaling that the admin API has changed. The consequence is that changing them in +your values file on an upgrade does nothing. Change them through the admin API +instead — see [App Definitions](../platform/app-definitions.md). -Production values files pin all images to specific SHA-based tags (e.g., `latest-e431a13`). When updating images: +**`hosts.allWildcardInstances` does not update cleanly on upgrade.** If you +change it, reinstall rather than upgrade. This is called out in the chart's own +values file. -1. Identify the new image SHA from the release or container registry. -2. Update the tag in the relevant values file in the deployment repository. -3. Deploy via the pipeline. +## Coordinating with your identity provider -Never use `latest` or branch-based tags in production — they change without a deployment, causing invisible configuration drift. +An EduIDE upgrade does not change your Keycloak configuration. But if an upgrade +adds a hostname — a new installation, or a renamed one — the client's redirect +URIs need updating in the same change window, or login breaks for that host +only. See [Access Control](../platform/access-control.md). diff --git a/docs/admins/operations/garbage-collection.md b/docs/admins/operations/garbage-collection.md index e71ca88..ee05c0d 100644 --- a/docs/admins/operations/garbage-collection.md +++ b/docs/admins/operations/garbage-collection.md @@ -5,7 +5,47 @@ description: Workspace TTL configuration, cleanup scheduling, and manual reclama # Garbage Collection -The workspace garbage collector is a dedicated Kubernetes operator (`theia-workspace-garbage-collector`) that periodically deletes workspaces that have exceeded their configured time-to-live. Without it, workspaces accumulate indefinitely, consuming PVC quota and storage capacity. +The workspace garbage collector periodically deletes workspaces that have exceeded their configured time-to-live. Without it, workspaces accumulate indefinitely, consuming PVC quota and storage capacity. + +:::note Placeholders on this page + +Each EduIDE installation lives in its own namespace, named `eduide-` by convention (`eduide-staging`, `eduide-cs101`, and so on). Commands below use `-n ` - substitute your installation's namespace. `eduide` is the conventional Helm release name for an installation. + +::: + +## Where it lives + +The garbage collector is a **conditional subchart of the `eduide` chart**, not a separately installed component. It is declared as a chart dependency named `theia-workspace-garbage-collector`, gated on `theia-workspace-garbage-collector.enabled`, which the `eduide` chart defaults to `true`. + +This has two consequences that drive everything else on this page: + +1. **You do not install it.** It comes and goes with the installation's `helm upgrade`. There is no separate release to manage. +2. **Its values must be nested** under a top-level `theia-workspace-garbage-collector:` key in the installation values file. That is how Helm passes values to a subchart, and the key must match the dependency name exactly. A bare top-level `env:` block is not an error - Helm accepts it, the subchart never sees it, and the garbage collector runs with its defaults. This fails silently, so check the running configuration after any change (see [Verifying the running configuration](#verifying-the-running-configuration)). + +The Kubernetes objects it creates in the installation namespace: + +| Object | Name | +|---|---| +| Deployment | `garbage-collector` | +| Pod label | `app: theia-workspace-garbage-collector` | +| Container | `garbage-collector` | +| ServiceAccount | `garbage-collector-sa` | + +The Deployment name and the pod label are different strings. Log selectors use the label; `rollout` commands use the Deployment name. + +### About the pinned image + +The image is pinned to a **commit SHA** rather than a version tag: + +```yaml +theia-workspace-garbage-collector: + image: + tag: "599557839e5c5893eb0c20785dac671ae70f7e8a" +``` + +The garbage collector's own repository has never cut a release, so its registry holds only `latest`, `main` and per-commit SHAs. Its own chart defaults to `latest`, which would mean two installs of the same EduIDE version could get different garbage collector builds, and `helm upgrade` would see no diff when the image changed underneath it. The pin exists to make an install reproducible. + +If you fork or rebuild the garbage collector, override the tag in your installation values rather than editing the chart. Expect this pin to be replaced by a semver tag once that repository publishes releases. ## How it works @@ -16,131 +56,192 @@ The garbage collector runs as a loop inside the cluster. On each tick: 3. If the workspace age exceeds `WORKSPACE_TTL`, the workspace resource is deleted. 4. Deletion of the workspace resource triggers the operator to clean up associated Kubernetes resources, including the PVC (depending on the storage class reclaim policy). -The garbage collector operates on creation time, not last-activity time. A workspace created 14 days ago will be deleted regardless of whether a session was active recently. This is a known limitation of the current implementation. +The garbage collector operates on **creation time, not last-use time**. A workspace created 14 days ago will be deleted regardless of whether a session was active in it five minutes ago. This is a known limitation of the current implementation and the single most important thing to understand before choosing a TTL: the TTL is a hard lifespan for the workspace, not an idle timeout. At startup, the garbage collector prints its configuration: ``` Starting garbage collector... -- Namespace: theia-prod +- Namespace: - Check interval: 30m0s - Workspace TTL: 336h0m0s ``` ## Configuration -The garbage collector is configured via environment variables. These are set in the Helm chart values file. +The garbage collector is configured through environment variables, which the subchart builds from a **map** under `env`: + +| Variable | Source | Default | Description | +|---|---|---|---| +| `K8S_NAMESPACE` | the Helm release namespace | - | The namespace to watch. **Not configurable** - the subchart always sets it to the release namespace, which is the installation's own namespace | +| `WORKSPACE_TTL` | `theia-workspace-garbage-collector.env.WORKSPACE_TTL` | `1209600s` (14 days) | Maximum age of a workspace before deletion | +| `CHECK_INTERVAL` | `theia-workspace-garbage-collector.env.CHECK_INTERVAL` | `1800s` (30 minutes) | How often the GC loop runs | -| Variable | Default | Description | -|---|---|---| -| `K8S_NAMESPACE` | `theia-prod` | The Kubernetes namespace to watch | -| `WORKSPACE_TTL` | `336h` (14 days) | Maximum age of a workspace before deletion. Accepts Go duration strings | -| `CHECK_INTERVAL` | `30m` | How often the GC loop runs. Accepts Go duration strings | +Two things differ from what you might expect: -Go duration string format: `336h` (hours), `72h30m` (hours and minutes), `1440m` (minutes). Do not use days — Go's duration parser does not support `d`. +- `K8S_NAMESPACE` has no values key. Because the garbage collector ships with each installation and watches only that installation, there is nothing to point it at. If you want a different namespace watched, that is a different installation with its own garbage collector. +- `env` is a **map of variable name to value**, not a list of `{name, value}` objects. Writing it as a list produces a rendering error or a broken Deployment, depending on where Helm gives up. + +Durations are Go duration strings. The chart ships plain-seconds values (`1209600s`, `1800s`); hour and minute suffixes (`168h`, `72h30m`, `1440m`) are the same format. Do not use days - Go's duration parser has no `d` unit. If you change notation, confirm the result in the startup log rather than assuming it parsed. ### Example values For a short course where workspaces should be cleaned up after 7 days: ```yaml -env: - - name: WORKSPACE_TTL - value: "168h" - - name: CHECK_INTERVAL - value: "1h" - - name: K8S_NAMESPACE - value: "theia-prod" +theia-workspace-garbage-collector: + enabled: true + env: + WORKSPACE_TTL: "604800s" # 7 days + CHECK_INTERVAL: "3600s" # 1 hour ``` For a long-running research environment where workspaces should persist for 90 days: ```yaml -env: - - name: WORKSPACE_TTL - value: "2160h" - - name: CHECK_INTERVAL - value: "6h" +theia-workspace-garbage-collector: + enabled: true + env: + WORKSPACE_TTL: "7776000s" # 90 days + CHECK_INTERVAL: "21600s" # 6 hours +``` + +To turn the garbage collector off entirely - appropriate where workspace lifetime is managed by some other process, or where losing a student's work to a TTL is unacceptable: + +```yaml +theia-workspace-garbage-collector: + enabled: false ``` +With it disabled, nothing reclaims workspace storage automatically. Plan a manual cleanup cadence and watch your PVC quota. + ## Deploying and updating -The garbage collector is deployed via its own Helm chart: +Because it is a subchart, you change the garbage collector by upgrading the installation: ```bash -helm upgrade --install theia-workspace-garbage-collector \ - ./helm -f ./helm/values.yaml +helm upgrade --install eduide oci://ghcr.io/eduide/charts/eduide \ + --version \ + -n \ + -f .yaml ``` -To change the TTL or interval, update the values file and run the upgrade. The deployment restarts and picks up the new values immediately. +There is no standalone chart to install. A command of the form `helm upgrade --install theia-workspace-garbage-collector ./helm` does not work: there is no such chart to install on its own, and even if there were, a separate release would not receive the installation's namespace or share its lifecycle. -To verify the running configuration: +The Deployment restarts and picks up the new environment variables as part of the upgrade. + +### Verifying the running configuration + +Given that misplaced values fail silently, verify rather than assume. Check the environment the container actually received: + +```bash +kubectl get deployment garbage-collector -n \ + -o jsonpath='{range .spec.template.spec.containers[0].env[*]}{.name}={.value}{"\n"}{end}' +``` + +And confirm against the startup log: ```bash -kubectl logs -n theia-prod -l app=theia-workspace-garbage-collector | head -10 +kubectl logs -n -l app=theia-workspace-garbage-collector --tail=10 ``` -The startup log lines show the effective configuration. +If the log shows the defaults after you set something else, your values are nested wrongly. Check that the top-level key is exactly `theia-workspace-garbage-collector` and that `env` is a map. + +You can also ask Helm what it resolved, without touching the cluster: + +```bash +helm get values eduide -n --all \ + | grep -A5 'theia-workspace-garbage-collector' +``` ## Manual cleanup -If you need to reclaim space immediately — for example, when the namespace is approaching quota limits before the next scheduled GC run — you can delete workspaces manually. +If you need to reclaim space immediately - when the namespace is approaching quota limits before the next scheduled GC run, for instance - you can delete workspaces manually. ```bash # List all workspaces with their age -kubectl get workspaces -n theia-prod \ +kubectl get workspaces.theia.cloud -n \ --sort-by=.metadata.creationTimestamp \ -o custom-columns='NAME:.metadata.name,CREATED:.metadata.creationTimestamp,USER:.spec.user' # Delete a specific workspace -kubectl delete workspace -n theia-prod +kubectl delete workspace -n -# Delete all workspaces older than a specific date (use with caution) -kubectl get workspaces -n theia-prod -o json \ - | jq -r '.items[] | select(.metadata.creationTimestamp < "2025-01-01T00:00:00Z") | .metadata.name' \ - | xargs -I {} kubectl delete workspace {} -n theia-prod +# Delete all workspaces created before a specific date (use with caution) +kubectl get workspaces.theia.cloud -n -o json \ + | jq -r '.items[] | select(.metadata.creationTimestamp < "T00:00:00Z") | .metadata.name' \ + | xargs -I {} kubectl delete workspace {} -n ``` +Run the `jq` pipeline without the final `xargs` first, and read the list before deleting anything. + Always confirm the workspace does not have an active session before deleting it. Deleting a workspace with a live session attached will leave the session broken. +The custom resources are in the `theia.cloud` API group even though the charts are named `eduide*` - a historical name that was never migrated. `kubectl get workspaces` resolves to the same resource as `kubectl get workspaces.theia.cloud`; the long form is used here so the group is unambiguous. + ## PVC cleanup after workspace deletion Deleting a workspace resource removes the Kubernetes custom resource and triggers the operator to delete associated pod resources. The PVC lifecycle depends on the storage class reclaim policy: - **Delete** policy: The PVC and the underlying volume are deleted automatically. -- **Retain** policy: The PVC is released but the underlying PersistentVolume remains. You must delete the PV manually to reclaim the storage. +- **Retain** policy: The PVC is removed but the underlying PersistentVolume remains, holding the data and the capacity. You must delete the PV manually to reclaim the storage. Check the current policy: ```bash -kubectl get storageclass csi-rbd-sc \ - -o jsonpath='{.reclaimPolicy}' +kubectl get storageclass \ + -o jsonpath='{.reclaimPolicy}{"\n"}' + +# Or check every storage class at once +kubectl get storageclass \ + -o custom-columns='NAME:.metadata.name,RECLAIM:.reclaimPolicy,DEFAULT:.metadata.annotations.storageclass\.kubernetes\.io/is-default-class' +``` + +With a `Retain` policy, orphaned volumes show up as PersistentVolumes in the `Released` phase: + +```bash +kubectl get pv --field-selector=status.phase=Released ``` -If orphaned PVCs accumulate: +**`Released` is a PersistentVolume phase, not a PersistentVolumeClaim phase.** A PVC is only ever `Pending`, `Bound` or `Lost`, so `kubectl get pvc --field-selector=status.phase=Released` matches nothing whether or not there is storage to reclaim - it reports an empty list and reads as "nothing to clean up". PVs are also cluster-scoped, so there is no `-n` flag. + +To find the released volumes that belonged to one installation: ```bash -# Find Released PVCs (no longer bound to any claim) -kubectl get pvc -n theia-prod --field-selector=status.phase=Released +kubectl get pv -o json \ + | jq -r '.items[] | select(.status.phase=="Released") + | select(.spec.claimRef.namespace=="") + | "\(.metadata.name)\t\(.spec.capacity.storage)\t\(.spec.claimRef.name)"' +``` + +Once you are certain the data is not needed: -# Delete them -kubectl delete pvc -n theia-prod +```bash +kubectl delete pv ``` +Deleting a `Released` PV with a `Retain` policy removes the Kubernetes object. Whether it removes the data depends on your storage backend - some drivers leave the underlying volume behind for a separate cleanup. Confirm with whoever operates your storage before assuming the capacity has come back. + ## Temporary TTL reduction -During storage pressure situations, lower the TTL temporarily to accelerate cleanup without manual intervention: +During storage pressure, lower the TTL temporarily to accelerate cleanup without deleting workspaces by hand: -1. Update the values file: lower `WORKSPACE_TTL` to e.g. `72h`. -2. Run `helm upgrade` to apply. -3. Wait for the next GC run (or trigger one by restarting the pod). -4. Once storage pressure is resolved, restore the original TTL value. +1. Lower `theia-workspace-garbage-collector.env.WORKSPACE_TTL` in the installation values file - to `259200s` (3 days), for example. +2. Run `helm upgrade` on the installation to apply. +3. Confirm the new value landed with the verification commands above. +4. Wait for the next GC run, or restart the pod to trigger one immediately. +5. Once storage pressure is resolved, restore the original TTL. ```bash -# Restart the GC pod to trigger an immediate run with new config -kubectl rollout restart deployment/theia-workspace-garbage-collector -n theia-prod +# Restart the GC pod to trigger an immediate run with the new config +kubectl rollout restart deployment/garbage-collector -n +kubectl rollout status deployment/garbage-collector -n ``` +The Deployment is named `garbage-collector`. `kubectl rollout restart deployment/theia-workspace-garbage-collector` fails with `not found` - that string is the pod label and the chart dependency name, not the Deployment name. + +Be deliberate about step 5. A TTL lowered during an incident and never restored quietly deletes work that users expected to keep. + ## What is not covered The garbage collector currently only deletes workspaces based on age since creation. It does not: @@ -148,5 +249,7 @@ The garbage collector currently only deletes workspaces based on age since creat - Consider last session activity time - Handle partial deletion failures gracefully (it stops on the first error) - Send notifications before deleting a workspace +- Clean up `Released` PersistentVolumes left behind by a `Retain` storage class +- Expose any metrics or a health endpoint - it has no Service, so its only observable output is its log If a deletion fails mid-run, the error is logged and the run halts. The next scheduled run will retry. diff --git a/docs/admins/operations/incident-response.md b/docs/admins/operations/incident-response.md index f6d07c2..00503fd 100644 --- a/docs/admins/operations/incident-response.md +++ b/docs/admins/operations/incident-response.md @@ -7,17 +7,30 @@ description: Runbooks for common EduIDE incident classes, triage sequence, and p This page contains runbooks for the most common incident classes in EduIDE, a triage sequence for unknown incidents, and the post-incident procedure. +:::note Placeholders on this page + +Each EduIDE installation lives in its own namespace, named `eduide-` by convention (`eduide-staging`, `eduide-cs101`, and so on). Commands use `-n ` - substitute the affected installation's namespace. The cluster-level chart installs into `eduide-system`, which is spelled out literally where it applies. + +Hostnames appear as ``, `` and `` - the DNS names you configured for the REST service, the landing page and session instances. + +The Deployments are named `operator-deployment`, `service-deployment` and `landing-page-deployment`. Their pods carry the shorter labels `app=operator`, `app=service` and `app=landing-page`, so log selectors and rollout commands use different names. Both forms appear below; they are not typos. + +The CRDs are in the `theia.cloud` API group even though the charts are named `eduide*` - a historical name that was never migrated. Searching the CRD list for `eduide` finds nothing and does not mean the CRDs are missing. + +::: + ## Triage sequence for unknown incidents When an incident is reported and the root cause is not immediately clear, work through this sequence before jumping to a specific runbook: -1. **Confirm impact scope** — Is it one user, one session type, one environment, or all environments? -2. **Check pod health** — `kubectl get pods -n theia-prod | grep -v Running` -3. **Check operator logs** — `kubectl logs -n theia-prod -l app=operator --tail=100` -4. **Check service logs** — `kubectl logs -n theia-prod -l app=service --tail=100` -5. **Check recent deployments** — Did a deployment run in the last 30 minutes? -6. **Check Keycloak** — Are authentication failures spiking? -7. **Apply the relevant runbook below.** +1. **Confirm impact scope** - Is it one user, one session type, one installation, or every installation on the cluster? A cluster-wide symptom points at `eduide-system` (Gateway, certificates, CRDs) rather than any one namespace. +2. **Check pod health** - `kubectl get pods -n | grep -v Running` +3. **Check operator logs** - `kubectl logs -n -l app=operator --tail=100` +4. **Check service logs** - `kubectl logs -n -l app=service --tail=100` +5. **Check routing** - `kubectl get gateway -n eduide-system` and `kubectl get httproute -n ` +6. **Check recent deployments** - Did a `helm upgrade` run in the last 30 minutes? `helm history eduide -n ` +7. **Check Keycloak** - Are authentication failures spiking? +8. **Apply the relevant runbook below.** --- @@ -28,24 +41,24 @@ When an incident is reported and the root cause is not immediately clear, work t **Step 1: Check for pending session pods** ```bash -kubectl get pods -n theia-prod | grep -E 'Pending|ContainerCreating' +kubectl get pods -n | grep -E 'Pending|ContainerCreating' ``` Pending pods indicate a scheduling or resource problem. Investigate: ```bash -kubectl describe pod -n theia-prod +kubectl describe pod -n ``` Common causes in the `Events` section: -- `Insufficient memory` or `Insufficient cpu` — cluster is at capacity -- `no nodes available` — all nodes are unschedulable -- `PodToleratesNodeTaints` — taints misconfiguration +- `Insufficient memory` or `Insufficient cpu` - cluster is at capacity +- `no nodes available` - all nodes are unschedulable +- `PodToleratesNodeTaints` - taints misconfiguration **Step 2: Check resource quota** ```bash -kubectl describe resourcequota -n theia-prod +kubectl describe resourcequota -n ``` If `persistentvolumeclaims` or `requests.memory` are at the hard limit, no new sessions can start. @@ -53,17 +66,26 @@ If `persistentvolumeclaims` or `requests.memory` are at the hard limit, no new s **Step 3: Check the operator** ```bash -kubectl logs -n theia-prod -l app=operator --tail=200 +kubectl logs -n -l app=operator --tail=200 ``` Look for reconciliation errors or repeated error messages on the same resource. -**Step 4: Check image availability** +**Step 4: Check the per-user session limit** + +If only some users are affected and they all already have a session running, they may be hitting the per-user cap. `operator.sessionsPerUser` defaults to **`1`**, so on a default install a user with one live session cannot start a second one. This is configuration working as intended, not an outage. + +```bash +kubectl get sessions.theia.cloud -n \ + -o custom-columns='NAME:.metadata.name,USER:.spec.user,CREATED:.metadata.creationTimestamp' +``` + +**Step 5: Check image availability** If pods are in `ErrImagePull` or `ImagePullBackOff`: ```bash -kubectl describe pod -n theia-prod | grep -A 5 Events +kubectl describe pod -n | grep -A 5 Events ``` This indicates the session image is unavailable. Verify the image tag in the App Definition still exists in the container registry. @@ -82,35 +104,237 @@ This indicates the session image is unavailable. Verify the image tag in the App **Step 1: Identify the failure point** - If the Keycloak login page itself fails to load: the problem is upstream of EduIDE. Contact the Keycloak instance admin. -- If login succeeds but users are rejected by EduIDE: the problem is in the OAuth2 proxy or token claim configuration. +- If login succeeds but users are rejected by EduIDE: the problem is in oauth2-proxy or the token claim configuration. + +**Step 2: Check oauth2-proxy logs** -**Step 2: Check OAuth2 proxy logs** +oauth2-proxy runs as a **sidecar container inside each session pod**. There is no oauth2-proxy Deployment, no oauth2-proxy Service, and no pod labelled `app=oauth2-proxy` - a selector like `kubectl logs -l app=oauth2-proxy` matches nothing and returns silently, which reads like "no errors" and is not. + +You must pick a session pod and name the container: ```bash -kubectl logs -n theia-prod -l app=oauth2-proxy --tail=100 +# Find a session pod: everything that is not a platform component +kubectl get pods -n \ + -l 'app notin (operator,service,landing-page,image-preloading)' + +# Read that pod's oauth2-proxy sidecar +kubectl logs -n -c oauth2-proxy --tail=100 ``` +Because each session has its own proxy, a configuration fault shows up identically in every session pod. Checking two or three is enough to tell a global misconfiguration from one broken pod. + Look for: -- `invalid cookie` — cookie secret mismatch, likely after a redeployment with a changed secret -- `failed to verify token` — audience claim missing or wrong -- `upstream response 403` — service is rejecting the proxied request +- `invalid cookie` - cookie secret mismatch, likely after an upgrade with a changed secret +- `failed to verify token` - audience claim missing or wrong +- `upstream response 403` - the service is rejecting the proxied request + +**Step 3: Check the oauth2-proxy configuration** -**Step 3: Verify Keycloak client scope** +The sidecar's configuration comes from ConfigMaps that the `eduide` chart renders into the installation namespace and the operator mounts by name: + +```bash +kubectl get configmap oauth2-proxy-config -n -o yaml +``` + +A change here only reaches sessions started **after** the change. Existing sessions keep the configuration they were launched with. + +**Step 4: Verify the Keycloak client scope** If token claims are missing (username, groups, audience), the client scope mappers may have been removed or the scope unassigned from the client. Check in the Keycloak admin console: 1. Open the client → **Client scopes**. -2. Confirm `theia-cloud-dedicated` is listed as a Default scope. +2. Confirm the EduIDE client scope is listed as a Default scope. 3. Open the scope → **Mappers** and verify all three mappers exist. -**Step 4: Check cookie secret** +**Step 5: Check the cookie secret** -If the cookie secret was rotated (new deployment with a different `THEIA_KEYCLOAK_COOKIE_SECRET`), all existing sessions are invalidated. Users need to clear cookies and log in again. This is expected behaviour, not a bug. +If the cookie secret was rotated, all existing sessions are invalidated. Users need to clear cookies and log in again. This is expected behaviour, not a bug. **Mitigation:** Communicate to affected users that they need to clear browser cookies for the domain and log in again. --- +## Runbook: Gateway or HTTPRoute failure + +**Symptoms:** Users get a connection error, a 404, or a 503 from the load balancer rather than an EduIDE page. Pods are healthy and logs show no inbound requests at all - the traffic never reaches them. + +Routing is Gateway API, served by Envoy Gateway. A **shared Gateway** lives in `eduide-system` and carries one set of listeners per installation; each installation namespace contributes **HTTPRoutes** that attach to specific listeners by name. There is no Ingress resource anywhere in the system, so checking for one is a dead end. + +The failure mode that makes this worth its own runbook: a route that does not attach produces **no error anywhere in EduIDE's own logs**. Nothing is broken from the operator's or the service's point of view; the request simply never arrives. + +**Step 1: Check the Gateway itself** + +```bash +kubectl get gateway -n eduide-system +``` + +The default name is `theia-shared-gateway`. The `PROGRAMMED` column is the one that matters - `True` means Envoy Gateway has translated the Gateway into a live listener configuration. + +```bash +kubectl describe gateway theia-shared-gateway -n eduide-system +``` + +Under `Status`, the Gateway carries top-level `Accepted` and `Programmed` conditions. `Accepted=False` means the spec was rejected (a bad `gatewayClassName`, most often). `Accepted=True, Programmed=False` means the spec is valid but the controller could not realise it - typically no address is available, or the controller is not running: + +```bash +kubectl get pods -n envoy-gateway-system +``` + +Envoy Gateway is **not installed by the EduIDE charts**. It is assumed to be present. If that namespace is empty, that is your incident. + +**Step 2: Check individual listener conditions** + +Each listener reports its own conditions, and one broken listener does not stop the others. This is why a single installation can be unreachable while every other installation on the cluster is fine. + +```bash +kubectl get gateway theia-shared-gateway -n eduide-system \ + -o jsonpath='{range .status.listeners[*]}{.name}{"\t"}{range .conditions[*]}{.type}={.status}({.reason}) {end}{"\n"}{end}' +``` + +What to look for per listener: + +- `Accepted=False` - the listener spec is invalid. A duplicate hostname across listeners or an unsupported protocol will do this. +- `Programmed=False` - accepted but not serving. +- `ResolvedRefs=False` with reason `InvalidCertificateRef` - the listener names a TLS Secret that does not exist, is not of type `kubernetes.io/tls`, or is in a namespace the Gateway may not read. Check the Secret exists: + + ```bash + kubectl get secret -n eduide-system + ``` + +- `Conflicted=True` - two listeners claim the same port and hostname combination. + +Also check `attachedRoutes` per listener in the same status block. A listener with `attachedRoutes: 0` that should have routes is the direct link to step 3. + +**Step 3: Check the HTTPRoutes** + +The `eduide` chart creates three HTTPRoutes in the installation namespace: + +| Route | Serves | Backend | +|---|---|---| +| `landing-route` | the landing page host | `landing-page-service` | +| `service-route` | the REST service host, path prefix `/service` | `service-service` | +| the instances route (default name `theia-cloud-demo-ws-route`) | session instance hosts | patched by the operator, one rule per live session | + +```bash +kubectl get httproute -n +``` + +Then read the attachment status, which is the part `kubectl get` does not show: + +```bash +kubectl get httproute -n \ + -o jsonpath='{range .status.parents[*]}{.parentRef.name}/{.parentRef.sectionName}{"\t"}{range .conditions[*]}{.type}={.status}({.reason}) {end}{"\n"}{end}' +``` + +The conditions to read: + +- **`Accepted=False`, reason `NoMatchingListenerHostname`** - the most common routing fault in this system. The route declares a hostname that does not intersect the hostname of the listener it names. Gateway API requires the two to overlap; if they do not, the route is silently dropped. This happens when an installation's hostnames are changed in the tenant values without re-running the cluster bootstrap that generates the matching listeners, so the listener still carries the old hostname. +- `Accepted=False`, reason `NoMatchingParent` - the `sectionName` names a listener that does not exist on that Gateway. Same root cause: listener names are generated per installation, and the tenant values must reference them exactly. +- `Accepted=False`, reason `NotAllowedByListeners` - the listener's `allowedRoutes` does not permit routes from this namespace. +- `ResolvedRefs=False`, reason `BackendNotFound` - the backend Service does not exist. Check `kubectl get svc -n ` for `service-service` and `landing-page-service`. + +Compare the two hostnames directly when you suspect a mismatch: + +```bash +kubectl get httproute -n -o jsonpath='{.spec.hostnames}{"\n"}' +kubectl get gateway theia-shared-gateway -n eduide-system \ + -o jsonpath='{range .spec.listeners[*]}{.name}{"\t"}{.hostname}{"\n"}{end}' +``` + +**Step 4: If only session URLs are broken** + +The instances route ships with an empty rule list; the **operator** patches one rule into it per live session. If the landing page and the service work but session URLs 404, check that the operator is running and reconciling (see the operator runbook below), then inspect the route's rules: + +```bash +kubectl get httproute theia-cloud-demo-ws-route -n -o yaml +``` + +An empty `rules` list with sessions running means the operator is not patching. An operator restart is the usual fix. + +**Mitigation:** Re-run the cluster-level `helm upgrade` for `eduide-cluster` with listener definitions that match the installation's current hostnames, then re-check the route conditions. Fixing the tenant side alone does not help if the listener is what is stale. + +--- + +## Runbook: Certificate does not cover a hostname + +**Symptoms:** Browsers show a certificate warning on one hostname. More confusingly, the landing page loads but stays empty or reports that it cannot reach the service, with no error in any EduIDE log. + +This one deserves care because **the cluster reports itself healthy throughout**. Gateway API never compares a certificate's Subject Alternative Names against the listener's hostname. If a listener for `` references a certificate valid only for ``, the listener still reports `Programmed=True` and `ResolvedRefs=True`: from Envoy's point of view the Secret exists, parses, and was loaded. The mismatch is a client-side judgement, so the only place it appears is in the browser. + +The second-order effect is the one that generates support tickets. The landing page and the REST service are on **different hostnames**, so the landing page's calls to the service are cross-origin. The browser refuses to complete a TLS handshake it does not trust, and because the request is a background `fetch` rather than a top-level navigation, the user is never offered the "proceed anyway" interstitial. They see a landing page that simply does not work, with no warning explaining why. Meanwhile the service is running perfectly and logs nothing, because no request ever reached it. + +So: if the landing page is blank or cannot list workspaces, and the service pod looks healthy and idle, check the service host's certificate before anything else. + +**Step 1: Check what the certificate actually covers** + +Ask the live endpoint, which is what the browser sees: + +```bash +echo | openssl s_client -connect :443 -servername 2>/dev/null \ + | openssl x509 -noout -checkhost +``` + +This prints either `Host does match certificate` or `Host does NOT match certificate`. The `-servername` flag matters: without it, SNI is not sent and you may be handed a different listener's certificate than the browser gets. + +Run it for every hostname the installation uses - landing, service, and at least one instance host: + +```bash +for h in ; do + printf '%s: ' "$h" + echo | openssl s_client -connect "$h":443 -servername "$h" 2>/dev/null \ + | openssl x509 -noout -checkhost "$h" +done +``` + +**Step 2: List the SANs to see what is missing** + +```bash +echo | openssl s_client -connect :443 -servername 2>/dev/null \ + | openssl x509 -noout -text \ + | grep -A1 'Subject Alternative Name' +``` + +Also check the validity window while you have the certificate in hand - an expired certificate produces the same silent cross-origin failure: + +```bash +echo | openssl s_client -connect :443 -servername 2>/dev/null \ + | openssl x509 -noout -dates -subject -issuer +``` + +**Step 3: Check the certificate in the cluster** + +Verify the Secret the listener references, without going over the network: + +```bash +kubectl get secret -n eduide-system \ + -o jsonpath='{.data.tls\.crt}' | base64 -d \ + | openssl x509 -noout -checkhost -dates +``` + +If this passes but step 1 fails, the listener is referencing a **different** Secret than you think. Confirm which one: + +```bash +kubectl get gateway theia-shared-gateway -n eduide-system \ + -o jsonpath='{range .spec.listeners[*]}{.name}{"\t"}{.hostname}{"\t"}{range .tls.certificateRefs[*]}{.name}{" "}{end}{"\n"}{end}' +``` + +That command prints listener name, hostname and certificate Secret side by side, which makes a mismatched pairing obvious. + +**Step 4: If cert-manager manages the certificate** + +The `eduide-cluster` chart can create cert-manager `Certificate` resources, off by default. When enabled, check the resource rather than the Secret: + +```bash +kubectl get certificate -n eduide-system +kubectl describe certificate -n eduide-system +``` + +`Ready=False` with a `CertificateRequest` stuck in the events usually means the ACME challenge is not completing. Note that an HTTP-01 challenge **cannot** issue a wildcard certificate, so a listener that needs a wildcard hostname requires a DNS-01 solver or a certificate you supply yourself. + +**Mitigation:** Re-issue the certificate with every hostname the installation serves in its SAN list, then let the listener reload. Envoy Gateway picks up a changed Secret without a Gateway restart, but confirm with step 1 rather than assuming. Adding a hostname to an installation always means widening the certificate as well as adding a listener - the two are separate changes and forgetting the second produces exactly this incident. + +--- + ## Runbook: Storage exhaustion **Symptoms:** New workspace creation fails with storage errors. Existing sessions are unaffected. @@ -118,7 +342,7 @@ If the cookie secret was rotated (new deployment with a different `THEIA_KEYCLOA **Step 1: Check PVC quota** ```bash -kubectl describe resourcequota -n theia-prod | grep persistentvolumeclaims +kubectl describe resourcequota -n | grep persistentvolumeclaims ``` If at the hard limit, no new PVCs can be created. @@ -126,18 +350,18 @@ If at the hard limit, no new PVCs can be created. **Step 2: Check storage capacity** ```bash -kubectl describe resourcequota -n theia-prod | grep requests.storage +kubectl describe resourcequota -n | grep requests.storage ``` **Step 3: Identify stale workspaces** ```bash # List workspaces sorted by age -kubectl get workspaces -n theia-prod \ +kubectl get workspaces.theia.cloud -n \ --sort-by=.metadata.creationTimestamp # Count total workspaces -kubectl get workspaces -n theia-prod --no-headers | wc -l +kubectl get workspaces.theia.cloud -n --no-headers | wc -l ``` **Step 4: Delete old workspaces** @@ -145,18 +369,41 @@ kubectl get workspaces -n theia-prod --no-headers | wc -l If the garbage collector has not yet run, manually delete workspaces older than the TTL: ```bash -kubectl delete workspace -n theia-prod +kubectl delete workspace -n ``` The PVC is released according to the storage class reclaim policy. See [Storage and Quotas](/admins/platform/storage-and-quotas) for PVC cleanup. +**Step 5: Look for orphaned PersistentVolumes** + +With a `Retain` reclaim policy, deleting a workspace leaves the underlying PV behind in the `Released` phase, still counting against your storage backend. + +`Released` is a **PersistentVolume** phase, not a PersistentVolumeClaim phase. PVCs are only ever `Pending`, `Bound` or `Lost`, so a field selector for `Released` PVCs matches nothing and will mislead you into concluding there is nothing to clean up. PVs are also cluster-scoped, so no `-n` applies: + +```bash +# Released PVs: correct +kubectl get pv --field-selector=status.phase=Released + +# Narrow to the ones that belonged to this installation +kubectl get pv -o json \ + | jq -r '.items[] | select(.status.phase=="Released") + | select(.spec.claimRef.namespace=="") + | .metadata.name' +``` + +Delete them once you are certain the data is not needed: + +```bash +kubectl delete pv +``` + **Mitigation:** Lower the garbage collection TTL temporarily to accelerate cleanup. See [Garbage Collection](/admins/operations/garbage-collection). --- ## Runbook: Node pressure / cluster capacity -**Symptoms:** Many pods are in `Pending` state. Session launches are slow or failing. Grafana shows high node CPU or memory utilisation. +**Symptoms:** Many pods are in `Pending` state. Session launches are slow or failing. Dashboards show high node CPU or memory utilisation. **Step 1: Identify the bottleneck** @@ -168,16 +415,22 @@ kubectl describe nodes | grep -A 5 "Allocated resources" **Step 2: Reduce pre-warmed instances temporarily** ```bash -# Drop minInstances to 0 for all App Definitions to stop warming new sessions curl -X PATCH \ -H "X-Admin-Api-Token: $ADMIN_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{"minInstances": 0}' \ - https://service.theia.artemis.cit.tum.de/service/admin/appdefinition/java-17-latest + https:///service/admin/appdefinition/ ``` Repeat for each affected App Definition. This frees scheduling space for active user sessions. +This endpoint needs the admin API token, which is not configured by default. See [Admin API Tokens](/admins/security/admin-api-tokens). If you have not set one up, edit the App Definition resource directly instead: + +```bash +kubectl patch appdefinitions.theia.cloud -n \ + --type=merge -p '{"spec":{"minInstances":0}}' +``` + **Step 3: Communicate status** If user-visible impact is ongoing, post a status update to the relevant channel. Include: @@ -187,25 +440,33 @@ If user-visible impact is ongoing, post a status update to the relevant channel. **Step 4: Scale cluster if needed** -If the pressure is sustained and expected (e.g., large cohort exercise), coordinate with the infrastructure team to add nodes. +If the pressure is sustained and expected (a large cohort exercise, for instance), coordinate with whoever manages the cluster's node pool to add capacity. --- ## Runbook: Operator not reconciling -**Symptoms:** App Definitions are updated via the API but no new pre-warmed sessions appear. Sessions that should be cleaned up remain running. +**Symptoms:** App Definitions are updated but no new pre-warmed sessions appear. Sessions that should be cleaned up remain running. Session URLs stop being added to the instances HTTPRoute. **Step 1: Check operator pod status** ```bash -kubectl get pods -n theia-prod -l app=operator +kubectl get pods -n -l app=operator ``` -In production, 3 replicas should be `Running`. If fewer, check for crash loops: +The chart default is **`operator.replicas: 1`**. One `Running` operator pod is a healthy operator, not a degraded one - do not escalate on a single replica unless you deliberately configured more. Check what you actually asked for: ```bash -kubectl describe pod -n theia-prod -kubectl logs -n theia-prod +kubectl get deployment operator-deployment -n \ + -o jsonpath='{.spec.replicas} desired / {.status.readyReplicas} ready{"\n"}' +``` + +If a pod is missing or restarting, check for crash loops: + +```bash +kubectl describe pod -n +kubectl logs -n +kubectl logs -n --previous ``` **Step 2: Check for CRD version mismatches** @@ -213,29 +474,39 @@ kubectl logs -n theia-prod After a CRD upgrade, the operator may fail to process resources if it is running an older version: ```bash -kubectl get crd | grep theia +kubectl get crd | grep theia.cloud +``` + +Expect `workspaces.theia.cloud`, `sessions.theia.cloud` and `appdefinitions.theia.cloud`. To see the stored version of each: + +```bash +kubectl get crd appdefinitions.theia.cloud \ + -o jsonpath='{.status.storedVersions}{"\n"}' ``` -Confirm the CRD versions match what the currently running operator expects. +Confirm the CRD versions match what the currently running operator expects. The CRDs are installed by the **cluster-level** `eduide-cluster` chart and carry a `helm.sh/resource-policy: keep` annotation, so they survive a tenant uninstall and are upgraded independently of any one installation. **Step 3: Restart the operator** -If logs show the operator is running but not reconciling (e.g., stuck in a watch loop): +If logs show the operator is running but not reconciling (stuck in a watch loop, for instance): ```bash -kubectl rollout restart deployment/operator -n theia-prod +kubectl rollout restart deployment/operator-deployment -n +kubectl rollout status deployment/operator-deployment -n ``` +The Deployment is `operator-deployment`. `kubectl rollout restart deployment/operator` fails with `deployments.apps "operator" not found` - the short name is the pod **label**, not the Deployment name. + --- ## Post-incident procedure -After every incident affecting production: +After every incident affecting a production installation: -1. **Confirm resolution** — Verify the platform is fully operational with a smoke-test session launch. -2. **Write up the timeline** — Document what happened, when it was detected, what was done, and when it was resolved. -3. **Identify root cause** — Was it a deployment change, a capacity event, a configuration drift, or an external dependency? -4. **Record follow-up actions** — Create tasks for any changes needed to prevent recurrence or improve detection speed. -5. **Update runbooks** — If this incident class was not covered, add it here. +1. **Confirm resolution** - Verify the platform is fully operational with a smoke-test session launch. +2. **Write up the timeline** - Document what happened, when it was detected, what was done, and when it was resolved. +3. **Identify root cause** - Was it a chart upgrade, a capacity event, a configuration drift, a certificate change, or an external dependency? +4. **Record follow-up actions** - Create tasks for any changes needed to prevent recurrence or improve detection speed. +5. **Update runbooks** - If this incident class was not covered, add it here. Keep incident records even for minor events. Patterns across small incidents often predict larger ones. diff --git a/docs/admins/operations/monitoring-basics.md b/docs/admins/operations/monitoring-basics.md index 97a5623..8a40682 100644 --- a/docs/admins/operations/monitoring-basics.md +++ b/docs/admins/operations/monitoring-basics.md @@ -5,97 +5,232 @@ description: Key signals, Prometheus metrics, alert thresholds, and health check # Monitoring Basics -EduIDE's monitoring stack is built on Prometheus and Grafana, managed through the `theia-monitoring` chart and the Rancher monitoring system (`cattle-monitoring-system`). This page describes the key signals to watch, what normal and degraded states look like, and how to verify platform health proactively. +EduIDE ships optional Prometheus integration as part of the `eduide-cluster` chart. It is off by default, and it assumes you already run a Prometheus and a Grafana somewhere in the cluster. This page describes what the chart actually creates, what it does and does not give you, the key signals to watch, and how to verify platform health by hand. + +:::note Placeholders on this page + +Every EduIDE installation lives in its own namespace, named `eduide-` by convention (`eduide-staging`, `eduide-cs101`, and so on). Commands below use `-n ` - substitute your installation's namespace. + +Hostnames appear as ``, `` and ``. These are the DNS names you configured for the REST service, the landing page and the optional shared build cache. + +::: + +:::warning `monitoring` and `monitor` are two different settings + +`monitoring.*` (this page) is Prometheus scrape configuration. `monitor.*` in the `eduide` chart is the operator's **session activity tracker** - it pings running session pods so idle sessions can be shut down, and has nothing to do with metrics. Turning off `monitor.enable` stops idle-session cleanup; turning off `monitoring.enabled` stops Prometheus discovery. They are easy to confuse and the failure modes look nothing alike. + +::: ## Monitoring infrastructure -The `theia-monitoring` chart deploys Prometheus `ServiceMonitor` resources that configure scrape targets for each environment namespace. Grafana dashboards are discovered from the `cattle-dashboards` namespace. +Monitoring is configured under `monitoring.*` in the **`eduide-cluster`** chart values, not in a per-installation values file, and not in a separate chart. An older standalone `theia-monitoring` chart existed; it has been removed. + +The chart creates: + +- Two `PodMonitor` resources (`monitoring.coreos.com/v1`) - **not** `ServiceMonitor`s. +- Two Grafana dashboard `ConfigMap`s, labelled `grafana_dashboard: "1"`. + +### Prerequisites -Namespaces in scope: -- `theia-prod` -- `theia-staging` -- `test1`, `test2`, `test3` +`monitoring.enabled` defaults to **`false`**, deliberately. Before turning it on: -If a new environment namespace is added, update the `targetNamespaces` and `sessionNamespaces` lists in the monitoring chart values before deploying. +1. **Prometheus Operator CRDs must exist.** `PodMonitor` is not a core Kubernetes kind. Verify: -### Shared build cache metrics + ```bash + kubectl get crd podmonitors.monitoring.coreos.com + ``` -The `theia-shared-cache` (Gradle and Bazel build cache) exposes Prometheus metrics at `/metrics` via a Redis Exporter sidecar. Enable the ServiceMonitor for it in the chart values: + If this returns `NotFound`, you have Prometheus but not the Prometheus Operator, and the chart render will fail. A standard `kube-prometheus-stack` install provides these CRDs. + +2. **Both target namespaces must already exist.** Helm does not create namespaces it was not told to create, so if `monitoring.namespace` or `monitoring.dashboardNamespace` is absent the install fails outright. + +3. **Your Prometheus must actually select these PodMonitors.** Prometheus Operator only scrapes `PodMonitor`s that match its `podMonitorSelector` and `podMonitorNamespaceSelector`. On a default `kube-prometheus-stack` install the selector is scoped to the release's own label, so a `PodMonitor` dropped into an arbitrary namespace is ignored silently. Either place the `PodMonitor`s where your Prometheus looks, or relax the selector (`prometheus.prometheusSpec.podMonitorSelectorNilUsesHelmValues: false` on `kube-prometheus-stack`). + +### Values ```yaml -metrics: - serviceMonitor: - enabled: true +monitoring: + enabled: false + # Namespace the PodMonitors go in. Must be a namespace your Prometheus + # discovers, and must already exist. + namespace: + # Namespace the Grafana dashboard ConfigMaps go in. Must already exist and + # be watched by Grafana's dashboard sidecar. + dashboardNamespace: + # Namespaces to scrape the REST service in. + targetNamespaces: [] + # Namespaces to scrape session pods in. + sessionNamespaces: [] +``` + +The chart's defaults for these two namespaces are `cattle-monitoring-system` and `cattle-dashboards`. Those are **Rancher's** namespaces - they are defaults inherited from where EduIDE was first deployed, not requirements. On a vanilla `kube-prometheus-stack` install both are typically the namespace you installed the stack into (often `monitoring`), because that stack's Grafana sidecar watches all namespaces for the `grafana_dashboard` label by default. + +`targetNamespaces` and `sessionNamespaces` are flat lists of namespace names. They are derived per cluster: you list every EduIDE installation namespace on that cluster. If either list is empty, the corresponding `PodMonitor` is not rendered at all. + +```yaml +monitoring: + enabled: true + namespace: monitoring + dashboardNamespace: monitoring + targetNamespaces: + - eduide-staging + - eduide-cs101 + sessionNamespaces: + - eduide-staging + - eduide-cs101 +``` + +Because the lists are cluster-level, adding a new installation means re-running the `eduide-cluster` upgrade with the namespace added. There is no way for a tenant to add itself. + +### Opting one installation out + +The **`eduide`** (per-installation) chart has its own `monitoring.enabled`, defaulting to `true`: + +```yaml +# in the tenant values file +monitoring: + enabled: false +``` + +This value renders nothing on its own - no template in the `eduide` chart reads it. It is a declaration that whoever assembles the cluster-level `targetNamespaces` and `sessionNamespaces` lists is expected to honour by leaving that namespace out. If you assemble those lists by hand, you have to read this flag yourself. + +### Verifying the wiring + +```bash +kubectl get podmonitor -n +# expect: theia-cloud-service, theia-cloud-sessions + +kubectl get configmap -n -l grafana_dashboard=1 +# expect: theia-cloud-dashboard-overview, theia-cloud-dashboard-session-startup ``` +Then confirm Prometheus picked them up: open the Prometheus UI under **Status → Targets** and look for the two job names, or port-forward and check the service discovery page. If the `PodMonitor`s exist but no targets appear, the cause is almost always the `podMonitorSelector` problem described above. + +## What is actually scraped + +Two scrape configurations, and it is worth knowing exactly what each one covers. + +### `theia-cloud-service` + +Selects pods labelled `app: service` in every namespace listed in `targetNamespaces`, scrapes the container port named `http` at path `/q/metrics` every 15s. + +The REST service is a Quarkus application using SmallRye Metrics. What you get is the MicroProfile base and vendor metric set: JVM heap and non-heap usage, thread counts, GC statistics, CPU, and generic REST request counters. This is genuinely useful for spotting a service that is leaking memory or wedged on GC. + +### `theia-cloud-sessions` + +Selects pods in `sessionNamespaces` whose `app` label is **not** one of `conversion-webhook`, `landing-page`, `operator`, `service` - in other words, everything left over, which in an installation namespace is the session pods. It scrapes the container port named `application` at the default path `/metrics` every 15s. + +Session pods are labelled `app: -`, so there is no stable label value to select on positively. That is why the selector is a negative match. + +Whether anything answers on `/metrics` depends entirely on the IDE image you run. A stock Theia or Code image does not serve Prometheus metrics on its application port, in which case these targets appear in Prometheus as `DOWN`. That is expected, not a fault to chase. + +### What is not scraped + +- **The operator exposes no Prometheus metrics.** There is no metrics dependency in the operator build and no `PodMonitor` for it. Everything you learn about the operator comes from its logs and its pod status. +- **The landing page exposes no metrics.** +- There are **no EduIDE-specific metric names** - nothing like `eduide_sessions_active`. The signals below are therefore derived from Kubernetes-level exporters (kube-state-metrics, cAdvisor) rather than read off an EduIDE counter. Where a signal has no metric behind it at all, this page says so. + ## Key signals +These assume `kube-state-metrics` and cAdvisor metrics are available, which is the case on any `kube-prometheus-stack` install. Thresholds are starting points; tune them to your cohort size. + ### Session launch latency -The most user-visible signal. A session launch that takes longer than ~10 seconds indicates either: -- All pre-warmed instances are consumed (increase `minInstances`) -- The cluster is under node pressure (check CPU/memory on nodes) -- The operator is backlogged in reconciliation +The most user-visible signal, and the one with no metric behind it. + +Nothing in EduIDE times a session launch and exports it. The `theia-cloud-dashboard-session-startup` dashboard shipped by the chart approximates it from pod lifecycle timestamps. To measure it directly, run a synthetic check: launch a session through the API on a schedule and time it from request to a reachable session URL. -Monitor: time from `POST /service` to a reachable session URL. +A launch that takes longer than roughly 10 seconds usually means one of: -**Alert threshold:** p95 > 15 seconds sustained for 5 minutes. +- All pre-warmed instances are consumed (increase `minInstances` on the App Definition). +- The cluster is under node pressure (check CPU and memory on nodes). +- The operator is backlogged in reconciliation. + +**Suggested alert:** p95 of your synthetic launch check > 15 seconds sustained for 5 minutes. ### Session availability -The fraction of launch requests that succeed. A drop indicates cluster instability, image pull failures, or storage attachment problems. +The fraction of launch attempts that succeed. A drop indicates cluster instability, image pull failures, or storage attachment problems. -Monitor: ratio of successful session starts to total start attempts. +Also has no first-class metric. The nearest proxies from kube-state-metrics are session pods stuck outside `Running`: -**Alert threshold:** success rate < 95% over a 10-minute window. +``` +kube_pod_status_phase{namespace="", phase="Pending"} +kube_pod_container_status_waiting_reason{namespace="", reason=~"ErrImagePull|ImagePullBackOff|CrashLoopBackOff"} +``` + +**Suggested alert:** any session pod `Pending` for more than 5 minutes. ### Pod memory utilisation -Individual session pods have a memory limit (e.g., `3000M` for `java-17-latest`). Pods consistently near their limit will OOMKill, which surfaces as unexpected session terminations. +Session pods have a memory limit - `2400M` by default for App Definitions, raised per app where needed. Pods that sit near their limit will be OOMKilled, which users experience as a session dying without warning. + +``` +container_memory_working_set_bytes{namespace=""} + / on(pod, container) kube_pod_container_resource_limits{resource="memory"} +``` -Monitor: `container_memory_working_set_bytes` for session pods vs. `limits.memory`. +Also watch `kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}` - a rising count is unambiguous. -**Alert threshold:** > 85% of memory limit for > 5 minutes. +**Suggested alert:** > 85% of the memory limit for > 5 minutes. ### Authentication error rate -Failed authentication attempts spike during misconfiguration (e.g., after a Keycloak change) or during an attack. Normal rate should be near zero for legitimate users. +Failed authentication spikes during misconfiguration (after a Keycloak change, for instance) or during an attack. The normal rate for legitimate users is near zero. + +oauth2-proxy runs as a **sidecar inside each session pod**, not as a central Deployment, so there is no single proxy to scrape and no aggregate authentication metric. What you can observe: -Monitor: HTTP 401 and 403 response rates on the OAuth2 proxy and service. +- Rejected admin API token requests, in the REST service logs (see [Audit and Compliance](/admins/security/audit-and-compliance)). +- Keycloak's own `LOGIN_ERROR` event rate, from Keycloak's metrics or event log. +- Per-pod oauth2-proxy logs, one pod at a time: -**Alert threshold:** > 5% of requests returning 401/403 over a 5-minute window. + ```bash + kubectl logs -n -c oauth2-proxy --tail=100 + ``` + +If authentication failures matter to you as an alertable signal, collect them at Keycloak. EduIDE does not aggregate them. ### Workspace storage usage -PVC count and total storage consumption against the namespace quota. Approaching the quota hard limit will prevent new workspace creation. +PVC count and total storage against the namespace quota. Hitting the quota hard limit prevents new workspace creation, while existing sessions carry on working - which makes it a confusing failure to diagnose from user reports alone. + +``` +kube_resourcequota{namespace="", resource="persistentvolumeclaims"} +kube_resourcequota{namespace="", resource="requests.storage"} +``` -Monitor: `kube_resourcequota` for `persistentvolumeclaims` and `requests.storage`. +Compare the `used` and `hard` values of each. -**Alert threshold:** > 80% of quota consumed. +**Suggested alert:** > 80% of quota consumed. ### Build cache hit rate -A low cache hit rate for the shared cache degrades CI build times but does not affect user sessions directly. Gradle and Bazel are tracked with separate counters. +The shared build cache is the optional `eduide-shared-cache` subchart of `eduide`, disabled by default (`eduide-shared-cache.enabled: false`). If you do not run it, skip this section. -Monitor (Gradle): `gradle_cache_cache_hits` / (`gradle_cache_cache_hits` + `gradle_cache_cache_misses`). +The subchart ships its own `ServiceMonitor` resources, gated on its own flag, which is on by default within the subchart: -Monitor (Bazel): `bazel_cache_cache_hits` / (`bazel_cache_cache_hits` + `bazel_cache_cache_misses`). +```yaml +eduide-shared-cache: + enabled: true + monitoring: + enabled: true +``` -Also watch `bazel_cache_hash_mismatches` — a non-zero rate indicates Bazel CAS uploads failing content-hash verification and should be investigated. +Note these are `ServiceMonitor`s, unlike the `PodMonitor`s above, so the same "does my Prometheus select them" question applies with `serviceMonitorSelector`. -**Informational threshold:** < 50% over a 24-hour window warrants investigation. +A low hit rate degrades build times but does not affect user sessions. The exact metric names the cache exports are not documented in the charts, and this page does not claim to know them. Check the `/metrics` output of a running cache pod to see what is actually there before writing alert rules against it. ## Health check procedures ### Service health ```bash -# Public ping (requires service auth token, not admin token) -curl https://service.theia.artemis.cit.tum.de/service/{appId} +# Public ping (requires the service auth token, not the admin token) +curl https:///service/ -# Admin ping (requires admin API token) -curl -H "X-Admin-Api-Token: $ADMIN_API_TOKEN" \ - https://service.theia.artemis.cit.tum.de/service/admin/{appId} +# Admin ping (requires an OAuth token with the admin group claim) +curl -H "Authorization: Bearer $OAUTH_TOKEN" \ + https:///service/admin/ ``` Both return `true` when the service is healthy. @@ -103,51 +238,88 @@ Both return `true` when the service is healthy. ### Operator health ```bash -kubectl get pods -n theia-prod -l app=operator -kubectl logs -n theia-prod -l app=operator --tail=50 +kubectl get pods -n -l app=operator +kubectl logs -n -l app=operator --tail=50 ``` -The operator runs 3 replicas in production. If fewer than 3 are `Running`, investigate immediately. +The operator runs **1 replica** by default (`operator.replicas: 1`). One `Running` pod is a healthy operator. Raise the replica count only if you have deliberately configured it; do not treat a single operator pod as a degraded state. + +The Deployment is named `operator-deployment`, not `operator`: + +```bash +kubectl rollout status deployment/operator-deployment -n +``` ### Session pod health ```bash -# Count running session pods -kubectl get pods -n theia-prod --field-selector=status.phase=Running | grep -c session +# Count running pods that are not platform components +kubectl get pods -n --field-selector=status.phase=Running \ + -l 'app notin (operator,service,landing-page,image-preloading)' --no-headers | wc -l # Find stuck or crash-looping pods -kubectl get pods -n theia-prod | grep -E 'CrashLoopBackOff|Error|Pending' +kubectl get pods -n | grep -E 'CrashLoopBackOff|Error|Pending' ``` -### Build cache health +### Custom resource health + +The CRDs are still in the `theia.cloud` API group for historical reasons, even though the charts are named `eduide*`. Grepping for `eduide` in the CRD list finds nothing and does not mean the CRDs are missing. + +```bash +kubectl get crd | grep theia.cloud +kubectl get workspaces.theia.cloud -n +kubectl get sessions.theia.cloud -n +kubectl get appdefinitions.theia.cloud -n +``` + +### Gateway health + +Routing is Gateway API (Envoy Gateway) with a shared Gateway in `eduide-system` and HTTPRoutes in each installation namespace. There is no Ingress resource to check. ```bash -# Readiness check -curl https://cache.theia.artemis.cit.tum.de/health +kubectl get gateway -n eduide-system +kubectl get httproute -n +``` + +See [Incident Response](/admins/operations/incident-response) for what to do when either reports a problem. + +### Build cache health -# Liveness check -curl https://cache.theia.artemis.cit.tum.de/ping +Only if you run the optional shared cache: + +```bash +curl https:///health +curl https:///ping ``` Returns `200 OK` when healthy. ## Grafana dashboards -Dashboards are deployed to the `cattle-dashboards` namespace. To access them: +The chart installs two dashboards as `ConfigMap`s labelled `grafana_dashboard: "1"` into `monitoring.dashboardNamespace`: + +- `theia-cloud-dashboard-overview` +- `theia-cloud-dashboard-session-startup` + +That label is the convention the `kube-prometheus-stack` Grafana sidecar watches for, so on a standard install the dashboards appear automatically once the ConfigMaps land in a namespace the sidecar covers. On Rancher, the equivalent namespace is `cattle-dashboards` and the dashboards show up in the Rancher monitoring UI. + +If the dashboards do not appear: -1. Open the Rancher UI and navigate to the monitoring section. -2. Look for dashboards prefixed with `theia-`. -3. The main session dashboard shows launch latency, active sessions, and pod resource usage per namespace. +1. Confirm the ConfigMaps exist in the namespace you configured. +2. Confirm your Grafana sidecar watches that namespace and looks for that label. +3. Confirm the panels have data - a dashboard renders empty rather than disappearing when Prometheus has no matching series, which is what you will see if the `PodMonitor`s are not being selected. -If dashboards are missing after a new environment is added, verify that the monitoring chart has been redeployed with the updated namespace list. +If dashboards render but a newly added installation is missing from them, the cause is the cluster-level namespace lists: re-run the `eduide-cluster` upgrade with the new namespace in `targetNamespaces` and `sessionNamespaces`. ## Routine health check cadence | Check | Frequency | Method | |---|---|---| | Session launch smoke test | Daily | Launch a session manually and verify it starts | -| Pod status overview | Daily | `kubectl get pods -n theia-prod` | -| Resource quota utilisation | Weekly | `kubectl describe resourcequota -n theia-prod` | +| Pod status overview | Daily | `kubectl get pods -n ` | +| Gateway and route status | Daily | `kubectl get gateway -n eduide-system` and `kubectl get httproute -n ` | +| Resource quota utilisation | Weekly | `kubectl describe resourcequota -n ` | | PVC growth rate | Weekly | Compare PVC count to previous week | -| Alert rule review | Monthly | Confirm alert thresholds are still appropriate for current cohort size | -| Dashboard coverage | On namespace addition | Verify new namespaces appear in Grafana | +| Certificate expiry and coverage | Monthly | See [Incident Response](/admins/operations/incident-response) | +| Alert rule review | Monthly | Confirm thresholds still suit the current cohort size | +| Namespace list coverage | On installation addition | Verify the new namespace is in `targetNamespaces` and `sessionNamespaces` | diff --git a/docs/admins/operations/session-management.md b/docs/admins/operations/session-management.md index 799f09f..deb3967 100644 --- a/docs/admins/operations/session-management.md +++ b/docs/admins/operations/session-management.md @@ -5,7 +5,15 @@ description: Admin-level oversight of active sessions, workspaces, and handling # Session Management -As an admin, you can inspect and manage all sessions and workspaces across the platform. This is distinct from what individual users can do — users can only see their own workspaces and sessions. Admin operations cover the entire namespace and are intended for operational oversight, not routine use. +As an admin, you can inspect and manage all sessions and workspaces across the platform. This is distinct from what individual users can do - users can only see their own workspaces and sessions. Admin operations cover the entire namespace and are intended for operational oversight, not routine use. + +:::note Placeholders on this page + +Each EduIDE installation lives in its own namespace, named `eduide-` by convention (`eduide-staging`, `eduide-cs101`, and so on). Commands below use `-n ` - substitute your installation's namespace. `` stands for the DNS name you configured for the REST service. + +The custom resources are in the `theia.cloud` API group even though the charts are named `eduide*`. This is a historical name that was never migrated, so `kubectl get crd | grep eduide` returns nothing. The full resource names are `workspaces.theia.cloud`, `sessions.theia.cloud` and `appdefinitions.theia.cloud`; the short forms `workspaces`, `sessions` and `appdefinitions` work too, and `ws`, `appdef` and `ad` are registered short names. + +::: ## Understanding the session and workspace model @@ -13,25 +21,41 @@ A **workspace** is the durable context for a user. It owns the PVC and persists A **session** is the live IDE runtime. It is bound to a user and a workspace, exposes the IDE over a URL, and is destroyed when it ends. -A user can have multiple workspaces, but the platform enforces a per-user session limit (`sessionsPerUser`, default: 10). If a user hits this limit, they cannot start new sessions until existing ones are stopped. +A user can have multiple workspaces, but the platform enforces a per-user session limit (`operator.sessionsPerUser`, chart default: **`1`**). If a user hits this limit, they cannot start new sessions until existing ones are stopped. + +The default of one session per user catches people out: a user who leaves a session running in another browser tab and then tries to start a second one is refused, and the symptom looks like a launch failure. Raise the limit in the installation values if your teaching model needs concurrent sessions: + +```yaml +operator: + sessionsPerUser: "3" +``` + +The value is a string in the chart values. ## Listing sessions and workspaces via kubectl The operator manages workspaces and sessions as custom resources. You can inspect them directly: ```bash -# List all workspaces in the production namespace -kubectl get workspaces -n theia-prod +# List all workspaces in an installation namespace +kubectl get workspaces.theia.cloud -n # List all sessions -kubectl get sessions -n theia-prod +kubectl get sessions.theia.cloud -n # List sessions for a specific user -kubectl get sessions -n theia-prod \ +kubectl get sessions.theia.cloud -n \ -o jsonpath='{range .items[?(@.spec.user=="")]}{.metadata.name}{"\n"}{end}' # Get full details on a session -kubectl describe session -n theia-prod +kubectl describe session -n +``` + +To see who owns what at a glance: + +```bash +kubectl get sessions.theia.cloud -n \ + -o custom-columns='NAME:.metadata.name,USER:.spec.user,APP:.spec.appDefinition,CREATED:.metadata.creationTimestamp' ``` ## Listing sessions via the service API @@ -59,31 +83,39 @@ If a session is stuck, has become unresponsive, or needs to be terminated to fre curl -X DELETE \ -H "Content-Type: application/json" \ -d '{"appId": "", "user": "", "sessionName": ""}' \ - https://service.theia.artemis.cit.tum.de/service/session + https:///service/session # Or delete the session resource directly via kubectl (immediate, bypasses service) -kubectl delete session -n theia-prod +kubectl delete session -n ``` Direct `kubectl delete` bypasses the service layer and removes the resource immediately. Use this when the session pod is not responding to normal termination or when the service itself is unavailable. ## Force-deleting a workspace -Deleting a workspace removes the custom resource and — depending on the storage class reclaim policy — may also remove the associated PVC. +Deleting a workspace removes the custom resource and - depending on the storage class reclaim policy - may also remove the associated PVC. ```bash # Via the service API curl -X DELETE \ -H "Content-Type: application/json" \ -d '{"appId": "", "user": "", "workspaceName": ""}' \ - https://service.theia.artemis.cit.tum.de/service/workspace + https:///service/workspace # Via kubectl -kubectl delete workspace -n theia-prod +kubectl delete workspace -n ``` Before deleting a workspace, confirm the user does not have an active session attached to it. Deleting an active workspace while a session is running may leave the session in a broken state. +With a `Retain` reclaim policy the underlying PersistentVolume survives the deletion and moves to the `Released` phase. `Released` is a **PersistentVolume** phase; PVCs only ever reach `Pending`, `Bound` or `Lost`, so looking for released PVCs finds nothing whether or not there is anything to clean up. PVs are cluster-scoped, so no namespace flag applies: + +```bash +kubectl get pv --field-selector=status.phase=Released +``` + +See [Storage and Quotas](/admins/platform/storage-and-quotas) for the full cleanup procedure. + ## Identifying stuck or runaway sessions Sessions that remain running far beyond normal usage patterns (hours or days) may indicate: @@ -93,14 +125,27 @@ Sessions that remain running far beyond normal usage patterns (hours or days) ma ```bash # Find sessions older than expected (sort by creation time) -kubectl get sessions -n theia-prod \ +kubectl get sessions.theia.cloud -n \ --sort-by=.metadata.creationTimestamp # Find pods with very high memory consumption -kubectl top pods -n theia-prod --sort-by=memory | head -20 +kubectl top pods -n --sort-by=memory | head -20 # Find pods that are running but not in Ready state -kubectl get pods -n theia-prod | grep 'Running' | grep '0/' +kubectl get pods -n | grep 'Running' | grep '0/' +``` + +Session pods are labelled `app: -`, so there is no single label that selects all of them. To list session pods and nothing else, exclude the platform components instead: + +```bash +kubectl get pods -n \ + -l 'app notin (operator,service,landing-page,image-preloading)' +``` + +Each session pod runs at least two containers: the IDE itself, named after the App Definition, and the `oauth2-proxy` sidecar. When reading logs from a session pod you must name the container: + +```bash +kubectl logs -n -c oauth2-proxy --tail=50 ``` ## Session activity reporting @@ -121,13 +166,27 @@ Request body: This endpoint is called by the IDE client to prevent inactivity shutdown. If you need to verify whether a session has been reporting activity, check the session resource's status in the operator logs. +The operator side of this is controlled by `monitor.*` in the installation values, which is on by default: + +```yaml +monitor: + enable: true + activityTracker: + enable: true + # minutes between re-pinging the pods + interval: 1 +``` + +If idle sessions are never being shut down, confirm these are still enabled before investigating further. Note that `monitor` here is unrelated to `monitoring`, which is the Prometheus scrape configuration described in [Monitoring Basics](/admins/operations/monitoring-basics). The names are one letter apart and the settings do entirely different things. + ## Managing sessions before a large exercise Before a scheduled exercise where many users will launch sessions simultaneously: -1. **Increase `minInstances`** for the relevant App Definition to pre-warm sessions (see [App Definitions](/admins/platform/app-definitions)). +1. **Increase `minInstances`** for the relevant App Definition to pre-warm sessions (see [App Definitions](/admins/platform/app-definitions)). The chart default is `0`, meaning nothing is pre-warmed at all. 2. **Check current session count** to confirm enough capacity exists. -3. **Verify quota headroom** with `kubectl describe resourcequota -n theia-prod`. +3. **Verify quota headroom** with `kubectl describe resourcequota -n `. +4. **Check the per-user session limit** is high enough for what you are asking students to do, given the default of `1`. After the exercise peak passes, reduce `minInstances` back to the baseline to free cluster resources. @@ -137,7 +196,7 @@ For an active session, the service exposes a performance endpoint: ```bash curl \ - https://service.theia.artemis.cit.tum.de/service/session/performance/{appId}/{sessionName} + https:///service/session/performance// ``` This returns metrics from inside the running session container (if the session type supports it). Use this to investigate reports of a specific session being slow. diff --git a/docs/admins/platform/access-control.md b/docs/admins/platform/access-control.md index 3aae268..94e6451 100644 --- a/docs/admins/platform/access-control.md +++ b/docs/admins/platform/access-control.md @@ -53,16 +53,31 @@ Click **Next**. Set URLs based on the target environment domain. For production: ``` -Root URL: https://theia.artemis.cit.tum.de -Home URL: https://theia.artemis.cit.tum.de +Root URL: https:// +Home URL: https:// Valid redirect URIs: - https://theia.artemis.cit.tum.de/* - https://instance.theia.artemis.cit.tum.de/* + https:///* + https://service./* + https://instance./* + https://*.webview.instance./* Valid post-logout redirect URIs: + Web origins: + ``` -For a test environment, replace the domain with the test domain (e.g., `test1.theia-test.artemis.cit.tum.de`). +:::warning All four, not two +An installation serves four hostnames and every one of them takes part in the +login flow. Listing only the landing and instance hosts is a common mistake with +a confusing result: login appears to work, and then **webviews inside the IDE +fail to authenticate** — a failure users hit days later and report as "previews +are broken". + +Substitute your own landing host; for an installation at `eduide.example.edu` +the four are `eduide.example.edu`, `service.eduide.example.edu`, +`instance.eduide.example.edu` and `*.webview.instance.eduide.example.edu`. +::: + +Repeat this for each installation — one client per installation, because each +has its own hostnames. Click **Save**. diff --git a/docs/admins/platform/app-definitions.md b/docs/admins/platform/app-definitions.md index e9db5e2..aecc465 100644 --- a/docs/admins/platform/app-definitions.md +++ b/docs/admins/platform/app-definitions.md @@ -23,50 +23,71 @@ Each App Definition specifies: | `maxInstances` | Maximum number of concurrent sessions allowed for this type | | `options` | Additional per-definition configuration, such as data bridge settings | -In production, the current App Definitions are: +The chart ships eight, and they are not all visible: -| Name | Image | Memory request | Min instances | Max instances | -|---|---|---|---|---| -| `java-17-latest` | `ghcr.io/eduide/eduide/java-17` | 2000M | 3 | 1000 | -| `python-latest` | `ghcr.io/eduide/eduide/python` | 2000M | configurable | configurable | -| `c-latest` | `ghcr.io/eduide/eduide/c` | 2000M | configurable | configurable | -| `javascript-latest` | `ghcr.io/eduide/eduide/javascript` | 2000M | configurable | configurable | -| `ocaml-latest` | `ghcr.io/eduide/eduide/ocaml` | 2000M | configurable | configurable | -| `rust-latest` | `ghcr.io/eduide/eduide/rust` | 2000M | configurable | configurable | +| Name | Image | Offered on the landing page | +|---|---|---| +| `java-17-templates-latest` | `eduide/java-17-templates` | yes, with a Maven/Gradle choice | +| `java-17-latest` | `eduide/java-17` | no — hidden behind the templates variant | +| `c-templates-latest` | `eduide/c-templates` | yes, with a Make/Bazel choice | +| `c-latest` | `eduide/c` | no | +| `javascript-latest` | `eduide/javascript` | yes | +| `ocaml-latest` | `eduide/ocaml` | yes | +| `python-latest` | `eduide/python` | yes | +| `rust-latest` | `eduide/rust` | yes | -## Managing App Definitions via Helm +The `-templates` variants ship a starter project and a build-system picker; the +plain ones give an empty workspace. Both are deployable — a hidden definition can +still be launched by name — but only the visible ones appear in the drop-down. -App Definitions are defined in the `theia-appdefinitions` Helm chart. The values file (`appdefinitions.yaml`) for each environment controls the deployed set. +Defaults are `requestsMemory: 500M`, `requestsCpu: 200m`, `limitsMemory: 2400M`, +`limitsCpu: "2"`, `minInstances: 0`, `maxInstances: 1000`, with Java overriding +the CPU and memory upward. -To add a new App Definition: +## Changing which applications are offered -1. Add an entry under `apps` in the environment's appdefinitions values file: +App definitions live in your installation's values for the `eduide` chart, under +`appDefinitions.apps`. It is a **map keyed by definition name**, with a +`defaults` block that every entry inherits. - ```yaml - apps: - - name: haskell-latest - image: ghcr.io/eduide/eduide/haskell:latest - requestsMemory: 2000M - requestsCpu: 500m - limitsMemory: 3000M - minInstances: 0 - maxInstances: 200 - ``` +```yaml +appDefinitions: + apps: + haskell-latest: + image: eduide/haskell # tag comes from versions.ide + requestsMemory: 800M # only what differs from defaults + landingPage: + label: Haskell # omit this key to deploy it but hide it +``` + +Then upgrade the installation as usual: + +```bash +helm upgrade eduide oci://ghcr.io/eduide/charts/eduide \ + --version -n -f values.yaml -f secrets.yaml +``` + +### One entry, three effects + +That single map drives the AppDefinition custom resource, the landing page's +app list, **and** the set of images preloaded onto every node. You do not +maintain a preload list — an earlier design did, and production ended up +offering an application whose image was never preloaded, so every student who +picked it waited for a cold multi-gigabyte pull. -2. Ensure the image is available and the tag is pinned for production deployments. +### Two things that are easy to get wrong -3. Deploy the updated chart: +**Removing `buildSystems` removes the picker.** A `-templates` application whose +`landingPage` block has no `buildSystems` list offers no build-system choice, +which is the entire reason those images exist. - ```bash - helm upgrade theia-appdefinitions \ - oci://ghcr.io/eduide/charts/theia-appdefinitions \ - -n theia-prod \ - -f deployments/theia.artemis.cit.tum.de/appdefinitions.yaml - ``` +**Adding an application costs disk on every node.** Each image is preloaded +cluster-wide. Trimming the list to what your courses actually use is a +legitimate and effective way to reclaim node storage. -4. Add the new name to `additionalApps` in the main values file so it appears in the landing page. +Removing an entry stops new sessions using it. Sessions already running are not +disturbed. -To remove an App Definition, remove its entry from the values file and run the same upgrade. Sessions already running on that definition are not affected immediately; the operator will no longer reconcile new ones. ## Adjusting scaling at runtime @@ -83,11 +104,11 @@ This is the recommended approach for live capacity adjustments before a schedule ```bash # List all app definitions and their scaling config curl -H "X-Admin-Api-Token: $ADMIN_API_TOKEN" \ - https://service.theia.artemis.cit.tum.de/service/admin/appdefinition + https://service./service/admin/appdefinition # Get a specific app definition curl -H "X-Admin-Api-Token: $ADMIN_API_TOKEN" \ - https://service.theia.artemis.cit.tum.de/service/admin/appdefinition/java-17-latest + https://service./service/admin/appdefinition/java-17-latest ``` Response shape: @@ -106,7 +127,7 @@ curl -X PATCH \ -H "X-Admin-Api-Token: $ADMIN_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{"minInstances": 10, "maxInstances": 500}' \ - https://service.theia.artemis.cit.tum.de/service/admin/appdefinition/java-17-latest + https://service./service/admin/appdefinition/java-17-latest ``` Validation rules enforced by the service: diff --git a/docs/admins/platform/provisioning.md b/docs/admins/platform/provisioning.md index ce87268..a02708a 100644 --- a/docs/admins/platform/provisioning.md +++ b/docs/admins/platform/provisioning.md @@ -7,6 +7,13 @@ description: Bootstrap guide for deploying a new EduIDE environment from cluster This page covers the steps required to bootstrap a new EduIDE environment. It is intended for operators setting up production, staging, or test deployments. The steps must be followed in order because later charts depend on resources installed by earlier ones. +:::note This page describes deployment automation, not installation +If you are installing EduIDE at your own institution, you do not need any of +this — start at [Cluster Prerequisites](../install/prerequisites.md). The +pipeline below is how TUM drives its own installations; the charts are installed +the same way either way. +::: + ## How deployments work All EduIDE environments are deployed through GitHub Actions pipelines defined in @@ -33,170 +40,21 @@ the workflow — which is why `e2e-test` deliberately has no required reviewers. The pre-restructure workflows `deploy-production.yml`, `deploy-pr.yml` and `deploy-theia.yml` no longer exist. -## Prerequisites - -Before a first deployment to a new cluster, confirm the following are in place: - -- A Kubernetes cluster is available and reachable via `kubectl` -- Helm 3 is installed on the runner (handled automatically in CI) -- You have cluster-admin permissions -- **`cert-manager`** is installed on the cluster — manages TLS certificates via the `cert-manager.io` CRDs -- A Keycloak instance is running and you have admin access to the relevant realm -- DNS entries for the environment domain are configured and propagating -- GitHub environment secrets are set (see [GitHub environment secrets](#github-environment-secrets)) - -### Installing cert-manager - -This is a cluster-level prerequisite installed once. If it is already present, skip this. - -```bash -# cert-manager (check https://cert-manager.io for the latest version) -helm upgrade --install cert-manager jetstack/cert-manager \ - --namespace cert-manager --create-namespace \ - --set crds.enabled=true \ - --set config.enableGatewayAPI=true -``` - -Verify it is running: -```bash -kubectl get pods -n cert-manager -``` - -## Helm chart overview - -EduIDE is deployed through a set of layered Helm charts. Cluster-scoped charts are installed once and shared across all environments. Environment-scoped charts are deployed independently per environment. - -| Chart | Scope | Purpose | -|---|---|---| -| `theia-cloud-crds` | Cluster | Custom resource definitions for workspaces, sessions, app definitions | -| `theia-cloud-base` | Cluster | Cluster-wide shared resources | -| `theia-shared-gateway` | Cluster | Shared Envoy Gateway API entry point | -| `theia-cloud-combined` | Environment | Main application: service, operator, landing page, OAuth2 proxy | -| `theia-appdefinitions` | Environment | IDE session type definitions | -| `theia-certificates` | Environment | TLS certificate management for the environment domain | -| `theia-monitoring` | Environment | Prometheus ServiceMonitors and Grafana dashboards | +## Where the rest of this page went -## Step 1: Install cluster-scoped charts +Everything that used to follow — a seven-chart install order, `deployments//` +values paths, `deploy-production.yml`, and a secrets table — described the +platform before the 2.0.0 restructure. Every chart it named has been deleted and +every command it gave would fail. -These are installed once per cluster. The pipeline handles this via the `deploy_shared_gateway` input flag. If setting up a brand new cluster manually: +It has been replaced by pages that are kept in step with the charts: -```bash -# CRDs first — operator and service depend on these -helm upgrade --install theia-cloud-crds ./charts/theia-cloud-crds - -# Cluster base resources -helm upgrade --install theia-cloud-base ./charts/theia-cloud-base - -# Shared Envoy Gateway (namespace: gateway-system) -helm upgrade --install theia-shared-gateway ./charts/theia-shared-gateway \ - -n gateway-system --create-namespace \ - -f -``` - -## Step 2: Prepare Keycloak - -Before running the first environment deployment, the Keycloak client must be configured. See [Access Control](/admins/platform/access-control) for the full procedure. - -At minimum, you need: - -- A Keycloak client with the correct redirect URIs for the target environment domain -- A dedicated client scope with username, audience, and groups mappers -- Users who need admin access added to the `theia-cloud/admin` Keycloak group -- The client ID and cookie secret ready for the GitHub environment secrets - -## GitHub environment secrets - -The deployment pipeline reads credentials from GitHub environment secrets. Set the following in the GitHub environment before triggering a deployment: - -| Secret | Description | +| For | See | |---|---| -| `KUBECONFIG` | Kubernetes cluster configuration for the target cluster | -| `THEIA_ADMIN_API_TOKEN` | Token protecting the admin scaling endpoints. Must be a strong random string | -| `THEIA_KEYCLOAK_REALM` | Keycloak realm name, e.g. `tum` | -| `THEIA_KEYCLOAK_CLIENT_ID` | Client ID configured in Keycloak | -| `THEIA_KEYCLOAK_CLIENT_SECRET` | Keycloak client secret | -| `THEIA_KEYCLOAK_COOKIE_SECRET` | Base64-encoded 32-byte key for OAuth2 proxy cookie encryption | -| `THEIA_WILDCARD_CERTIFICATE_CERT` | Wildcard TLS certificate (base64 encoded) | -| `THEIA_WILDCARD_CERTIFICATE_KEY` | Wildcard TLS certificate key (base64 encoded) | - -Generate the cookie secret: -```bash -openssl rand -base64 32 -``` - -GitHub environment variables (not secrets) also required: - -| Variable | Example | -|---|---| -| `NAMESPACE` | `theia-prod` | -| `HELM_VALUES_PATH` | `deployments/theia.artemis.cit.tum.de` | - -## Step 3: Configure the environment values - -Each environment has its own `values.yaml` in the deployment repository under `deployments//`. The key sections to review before a first deployment: - -**Hosts:** -```yaml -hosts: - configuration: - baseHost: artemis.cit.tum.de - service: service.theia - landing: theia - instance: instance.theia -``` - -**Operator:** -```yaml -theia-cloud: - operator: - replicas: 3 - sessionsPerUser: 10 - storageClassName: csi-rbd-sc - requestedStorage: 250Mi - eagerStart: false -``` - -`sessionsPerUser` sets the hard upper limit on concurrent active sessions per user. Set `eagerStart: true` only if you intend to use session pre-warming. - -**Landing page:** -```yaml - landingPage: - appDefinition: java-17-latest - ephemeralStorage: true - additionalApps: - java-17-latest: { label: Java 17 } - python-latest: { label: Python } -``` - -`additionalApps` controls which IDE session types appear in the UI. Only include app definitions that are also deployed in `theia-appdefinitions`. - -## Step 4: Trigger the deployment pipeline - -Once secrets, variables, and values files are in place, trigger the deployment: - -- **Production:** Go to Actions → `deploy-production.yml` → Run workflow. Requires manual approval. -- **Staging:** Push to the main branch. The pipeline runs automatically. -- **Test/PR:** Open or push to a PR. Requires approval to run. - -The pipeline runs all chart installs in sequence, injects secrets, and waits for the rollout to complete before returning success. - -## Step 5: Validate the installation - -After the pipeline completes successfully: - -- [ ] All pods in the environment namespace are `Running` or `Completed`: `kubectl get pods -n theia-prod` -- [ ] Landing page is reachable at the environment domain -- [ ] Keycloak login redirect works correctly -- [ ] A test session can be launched and connects to the IDE -- [ ] Admin ping responds: `GET /service/admin/{appId}` with `X-Admin-Api-Token` -- [ ] Grafana dashboards show the environment namespace in scope - -## Required secrets summary - -| Secret (k8s) | Content | Created by | -|---|---|---| -| `service-admin-api-token` | `ADMIN_API_TOKEN` | Deployment pipeline | -| Keycloak client secret | OAuth2 proxy client credential | Deployment pipeline | -| OAuth2 proxy cookie secret | Cookie encryption key | Deployment pipeline | -| Wildcard TLS secret | Certificate and key | Deployment pipeline | -| Redis password (shared cache) | Redis auth credential | Helm chart (auto-generated) | +| What must exist on the cluster first, with commands | [Cluster Prerequisites](../install/prerequisites.md) | +| DNS, the wildcard requirement, and getting certificates | [Certificates and DNS](../install/certificates.md) | +| Installing both charts | [Installing EduIDE](../install/installing.md) | +| Adding a second or third installation | [Adding an Installation](../install/adding-an-installation.md) | +| Identity provider setup | [Access Control](access-control.md) | +| Moving to a new version | [Release and Version Policy](../maintenance/release-policy.md) | +| Undoing a bad deploy | [Rollback](../maintenance/rollback.md) | diff --git a/docs/admins/platform/storage-and-quotas.md b/docs/admins/platform/storage-and-quotas.md index a83cbec..c3ef9b5 100644 --- a/docs/admins/platform/storage-and-quotas.md +++ b/docs/admins/platform/storage-and-quotas.md @@ -37,16 +37,33 @@ theia-cloud: To change the default, update `requestedStorage` in the values file and redeploy the operator. Changes apply to newly created workspaces only. Existing PVCs are not resized automatically. -## Ephemeral storage mode +## Ephemeral storage mode — and it is the default -If `ephemeralStorage: true` is set on the landing page configuration, workspaces launched from that entry point do not get a persistent PVC. Sessions are fully ephemeral — all data is lost when the session ends. +:::danger Sessions do not persist unless you turn persistence on +`landingPage.ephemeralStorage` defaults to **`true`**. A session started from the +landing page gets no PersistentVolumeClaim, and **everything in it is destroyed +when the session ends**. + +If your students are expected to keep work between sessions, you must set: ```yaml landingPage: - ephemeralStorage: true + ephemeralStorage: false +``` + +Check what you actually deployed before a cohort relies on it: + +```bash +kubectl -n get cm landing-page-config -o jsonpath='{.data.config\.js}' | grep useEphemeralStorage ``` +::: + +Ephemeral mode is the right default for evaluation and demos — nothing to clean +up, no storage cost. It is the wrong setting for a course. -This is useful for demo or evaluation environments where persistence is not needed and storage overhead should be minimised. +Note that sessions launched from Artemis carry their own workspace regardless, +so this setting governs the landing-page path specifically. Test the path your +students will actually use. ## Namespace resource quotas @@ -58,8 +75,8 @@ Recommended quota structure for a production namespace: apiVersion: v1 kind: ResourceQuota metadata: - name: theia-prod-quota - namespace: theia-prod + name: -quota + namespace: spec: hard: requests.cpu: "200" # total CPU requests across all pods @@ -115,14 +132,14 @@ See [Garbage Collection](/admins/operations/garbage-collection) for TTL configur ```bash # List all PVCs in the production namespace -kubectl get pvc -n theia-prod +kubectl get pvc -n # Check total PVC count and storage consumption -kubectl get pvc -n theia-prod \ +kubectl get pvc -n \ -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.resources.requests.storage}{"\n"}{end}' # Check resource quota utilisation -kubectl describe resourcequota -n theia-prod +kubectl describe resourcequota -n ``` ## Cleaning up stale PVCs @@ -133,10 +150,13 @@ If PVCs are not automatically released: ```bash # List unbound PVCs -kubectl get pvc -n theia-prod --field-selector=status.phase=Released +# NOTE: `Released` is a PersistentVolume phase, not a PVC phase, so filtering +# PVCs by it always returns nothing. PVC phases are Pending, Bound and Lost. +# To find volumes left behind by deleted workspaces, look at the PVs: +kubectl get pv --field-selector=status.phase=Released # Delete a specific unbound PVC -kubectl delete pvc -n theia-prod +kubectl delete pvc -n ``` Confirm the storage class reclaim policy with: diff --git a/docs/admins/security/admin-api-tokens.md b/docs/admins/security/admin-api-tokens.md index b570bcd..b1a07e1 100644 --- a/docs/admins/security/admin-api-tokens.md +++ b/docs/admins/security/admin-api-tokens.md @@ -5,11 +5,17 @@ description: How the admin API token protects scaling endpoints, how to issue an # Admin API Tokens -The EduIDE service exposes a set of scaling endpoints that allow changing App Definition instance counts at runtime. These endpoints are not protected by Keycloak group membership — they use a dedicated static token passed via a request header. This page explains how the token works, how it is provisioned, and how to rotate it. +The EduIDE service exposes a set of scaling endpoints that allow changing App Definition instance counts at runtime. These endpoints are not protected by Keycloak group membership - they use a dedicated static token passed via a request header. This page explains how the token works, how it is provisioned, and how to rotate it. + +:::note Placeholders on this page + +Each EduIDE installation lives in its own namespace, named `eduide-` by convention (`eduide-staging`, `eduide-cs101`, and so on). Commands use `-n ` - substitute your installation's namespace. `eduide` is the conventional Helm release name for an installation, and `` stands for the DNS name you configured for the REST service. + +::: ## Why a separate token -The admin API token exists because the scaling endpoints are designed to be called by automation (deployment workflows, scripts, CI pipelines) rather than by human users logged in through a browser. OAuth flows are not well-suited to machine-to-machine calls. A static bearer token is simpler to use in that context. +The admin API token exists because the scaling endpoints are designed to be called by automation (deployment pipelines, scripts, cron jobs) rather than by human users logged in through a browser. OAuth flows are not well-suited to machine-to-machine calls. A static bearer token is simpler to use in that context. The trade-off is that token compromise grants direct access to the scaling endpoints. This is why token rotation and secure storage matter. @@ -23,34 +29,121 @@ The following endpoints require a valid `X-Admin-Api-Token` header: | `GET` | `/service/admin/appdefinition/{name}` | Get scaling config for a specific App Definition | | `PATCH` | `/service/admin/appdefinition/{name}` | Update `minInstances` or `maxInstances` | -Requests without the header receive `401 Unauthorized`. Requests with a wrong token receive `403 Forbidden`. +The Keycloak-protected admin ping endpoint (`GET /service/admin/{appId}`) is a **different** mechanism. It requires a valid OAuth token carrying the admin group claim, and the admin API token has no effect on it. The two live under the same `/service/admin` path prefix but are guarded by different filters. -The Keycloak-protected admin endpoints (`/service/admin/{appId}` ping endpoint, annotated with `@AdminOnly`) are separate from these and require a valid OAuth token with the admin group claim instead. +## Is the admin API even enabled? -## Token provisioning +By default it is not, and this is a deliberate choice rather than an oversight. -The token is injected into the service as a Kubernetes secret. The deployment workflow creates the secret from the GitHub environment secret `THEIA_ADMIN_API_TOKEN`. +The relevant values in the `eduide` chart: -The service reads it via: -- **Service property:** `theia.cloud.admin.api.token` -- **Container env var:** `ADMIN_API_TOKEN` -- **Secret reference in values:** - ```yaml - service: - adminApiTokenSecret: - name: service-admin-api-token - key: ADMIN_API_TOKEN - ``` +```yaml +service: + adminApiTokenSecret: + # Have the chart create the Secret from `adminApiToken` below. + create: false + # Set true if you created the Secret yourself, outside the chart. + external: false + name: service-admin-api-token + key: ADMIN_API_TOKEN + # Base64-encoded admin API token. Only read when adminApiTokenSecret.create is true. + adminApiToken: "" +``` + +With both `create` and `external` left at `false`, the service is deployed **without** the `ADMIN_API_TOKEN` environment variable at all, and the scaling endpoints are effectively closed. That is a valid way to run an installation: if you never call the scaling API, not issuing a token is the safest configuration. + +You must therefore choose one of the two provisioning paths below before the admin API works. + +## Provisioning: path A, let the chart create the Secret + +Use this when your installation values file is itself managed as a secret (a sealed values file, a CI secret store, a vault-rendered file). + +**Step 1: Generate the token and its base64 encoding** + +The value you put in `service.adminApiToken` goes into the Secret's `data:` field, which Kubernetes requires to be **base64-encoded**. The chart passes the value straight through without encoding it for you. + +```bash +# The token itself - this is what callers send in the header +TOKEN=$(openssl rand -hex 32) +echo "$TOKEN" + +# The value to put in service.adminApiToken +printf '%s' "$TOKEN" | base64 | tr -d '\n' +``` + +Two details that cause real failures: + +- **`openssl rand -hex 32` produces hex, not base64.** Pasting its output directly into `service.adminApiToken` produces a Secret whose decoded content is binary garbage, and every request is then rejected as an invalid token with no clue as to why. The value must be base64 of the token, which is what the second command produces. +- **`tr -d '\n'` is not optional.** A 64-character token base64-encodes to 88 characters, and GNU `base64` wraps its output at 76 columns. The embedded newline is preserved through Helm's quoting and corrupts the token. + +**Step 2: Set the values** + +```yaml +service: + adminApiTokenSecret: + create: true + adminApiToken: "" +``` + +**Step 3: Apply** + +```bash +helm upgrade --install eduide oci://ghcr.io/eduide/charts/eduide \ + --version \ + -n \ + -f .yaml +``` + +If `create: true` and `adminApiToken` is empty, the render fails with an explicit message rather than producing a broken Secret. + +## Provisioning: path B, bring your own Secret + +This is the better path for a hand-installed cluster, because the token never has to appear in a values file at all. Create the Secret with `kubectl` and point the chart at it. + +**Step 1: Create the Secret** + +```bash +TOKEN=$(openssl rand -hex 32) + +kubectl create secret generic service-admin-api-token \ + -n \ + --from-literal=ADMIN_API_TOKEN="$TOKEN" -## Generating a token +echo "Store this token in your secret manager: $TOKEN" +``` + +Note the contrast with path A: `kubectl create secret --from-literal` base64-encodes the value **for** you. Do not pre-encode here, or you will end up with a doubly-encoded token. + +**Step 2: Tell the chart the Secret exists** + +```yaml +service: + adminApiTokenSecret: + create: false + external: true + name: service-admin-api-token + key: ADMIN_API_TOKEN +``` + +`external: true` is what makes the chart mount the environment variable from a Secret it does not own. Leaving it at `false` while the Secret exists means the service is still deployed without the variable, and the admin API stays closed - a confusing state, because the Secret is right there in the namespace. -Use a cryptographically random value of at least 32 bytes: +**Step 3: Apply** ```bash -openssl rand -hex 32 +helm upgrade --install eduide oci://ghcr.io/eduide/charts/eduide \ + --version \ + -n \ + -f .yaml ``` -Store the output in the GitHub environment secret `THEIA_ADMIN_API_TOKEN` before running a deployment. +If you change the Secret's contents later, the service does not notice on its own. See [Rotating the token](#rotating-the-token). + +## How the service reads it + +- **Container env var:** `ADMIN_API_TOKEN`, mounted via `secretKeyRef` from the Secret named in `service.adminApiTokenSecret.name`, key `service.adminApiTokenSecret.key` +- **Service property:** `theia.cloud.admin.api.token` + +The value is read at startup. It is not re-read while the service is running. ## Authenticating requests @@ -63,50 +156,110 @@ X-Admin-Api-Token: Example: ```bash -export ADMIN_API_TOKEN="your-token-value" +export ADMIN_API_TOKEN="" curl -H "X-Admin-Api-Token: $ADMIN_API_TOKEN" \ - https://service.theia.artemis.cit.tum.de/service/admin/appdefinition + https:///service/admin/appdefinition ``` +The header carries the **raw** token, not its base64 encoding. The base64 step exists only because Kubernetes Secrets store data that way. + Never pass the token as a query parameter or in a URL. It will appear in server access logs. -## Rotating the token +## Response codes -Token rotation requires a redeployment because the service reads the token at startup from the mounted secret. +Knowing which code you got narrows the problem considerably: -1. Generate a new token: `openssl rand -hex 32` -2. Update the GitHub environment secret `THEIA_ADMIN_API_TOKEN` with the new value. -3. Trigger a deployment for the target environment. The deployment workflow recreates the Kubernetes secret and redeploys the service. -4. Update any automation or scripts that use the old token. +| Response | Meaning | +|---|---| +| `401 Unauthorized` - "Admin API token required in X-Admin-Api-Token header." | No header was sent. The token itself may be fine. | +| `403 Forbidden` - "Valid admin API token required." | A header was sent but does not match the configured token. | +| `403 Forbidden` - "Admin API token authentication is not configured." | **No token is configured on the service at all.** Every request gets this, with or without a header. Go back to the provisioning section. | -There is no grace period — after the redeployment, the old token is immediately invalid. +That third case is the one to watch for: it is a `403`, which reads like a rejected credential, but the cause is on the server side and no token you send will ever work. The service also logs a warning each time it happens, so the service log distinguishes the two `403`s clearly: -## Token scope +```bash +kubectl logs -n -l app=service --tail=100 | grep -i "admin API token" +``` -Each environment (production, staging, test) has its own independent admin API token. The token for production should not be shared with staging or test environments. This limits the impact of a token leak to one environment. +## Verifying the configuration -## What happens if the token is not configured +Confirm the Secret exists and holds a plausible value: -If `ADMIN_API_TOKEN` is empty or the secret is absent, the service starts but the admin API token check behaviour changes: +```bash +kubectl get secret service-admin-api-token -n \ + -o jsonpath='{.data.ADMIN_API_TOKEN}' | base64 --decode | wc -c +``` -- The `AppDefinitionAdminApiTokenFilter` returns an empty string for the configured token. -- Any request with any non-empty `X-Admin-Api-Token` header will be rejected as invalid. -- In effect, the scaling endpoints become inaccessible via this mechanism. +A count of 64 matches a token from `openssl rand -hex 32`. A count of zero, or a wildly different number, means the encoding went wrong somewhere. -Verify the token is correctly configured after any deployment: +Confirm the service actually received it: ```bash -kubectl get secret service-admin-api-token -n theia-prod \ - -o jsonpath='{.data.ADMIN_API_TOKEN}' | base64 --decode | wc -c +kubectl get deployment service-deployment -n \ + -o jsonpath='{range .spec.template.spec.containers[0].env[*]}{.name}{"\n"}{end}' \ + | grep ADMIN_API_TOKEN +``` + +No output means the env var was not mounted, which means neither `create` nor `external` is `true`. This is worth checking explicitly after every upgrade, because the failure is a `403` rather than a crash - the service starts and serves everything else normally. + +Finally, an end-to-end check: + +```bash +curl -s -o /dev/null -w '%{http_code}\n' \ + -H "X-Admin-Api-Token: $ADMIN_API_TOKEN" \ + https:///service/admin/appdefinition ``` -A non-zero character count confirms the secret is present. +`200` confirms the whole chain. + +## Rotating the token + +Rotation requires a service restart, because the token is read at startup from the mounted Secret. + +**If the chart owns the Secret (path A):** + +1. Generate a new token and its base64 encoding, as in path A step 1. +2. Update `service.adminApiToken` in your values file. +3. Run `helm upgrade` on the installation. +4. Restart the service if the upgrade did not roll the pods: + ```bash + kubectl rollout restart deployment/service-deployment -n + kubectl rollout status deployment/service-deployment -n + ``` +5. Update any automation that uses the old token. + +**If you own the Secret (path B):** + +1. Generate a new token. +2. Replace the Secret: + ```bash + kubectl create secret generic service-admin-api-token \ + -n \ + --from-literal=ADMIN_API_TOKEN="$NEW_TOKEN" \ + --dry-run=client -o yaml | kubectl apply -f - + ``` +3. Restart the service - this step is mandatory here, since nothing in the Helm release changed and no rollout is triggered on its own: + ```bash + kubectl rollout restart deployment/service-deployment -n + kubectl rollout status deployment/service-deployment -n + ``` +4. Update any automation that uses the old token. + +The Deployment is named `service-deployment`. `kubectl rollout restart deployment/service` fails with `not found` - `service` is the pod label, not the Deployment name. + +There is no grace period. Once the new pod is serving, the old token is immediately invalid, and there is a brief window during the rollout where both old and new pods are running and either token may be accepted depending on which pod answers. + +## Token scope + +Each installation has its own independent admin API token, because each installation has its own Secret in its own namespace. A production installation's token should not be reused for staging or test installations. This limits the impact of a leak to one installation. + +If you run several installations from one automation account, store the tokens separately rather than sharing one across namespaces. Sharing removes the only isolation this design gives you. ## Rotation schedule Rotate the admin API token: - Every 6 months as routine hygiene -- Immediately if you suspect the token has been exposed (logs, error reports, rotation of a developer who had access) +- Immediately if you suspect the token has been exposed (logs, error reports, a screen share) - After any off-boarding of a person or system that had access to it diff --git a/docs/admins/security/audit-and-compliance.md b/docs/admins/security/audit-and-compliance.md index 0a8e753..1761585 100644 --- a/docs/admins/security/audit-and-compliance.md +++ b/docs/admins/security/audit-and-compliance.md @@ -7,17 +7,27 @@ description: What EduIDE logs, recommended retention policies, and the access re This page describes what EduIDE logs across its components, how long to retain those logs, which operations are considered sensitive, and the recommended cadence for access reviews. +:::note Placeholders on this page + +Each EduIDE installation lives in its own namespace, named `eduide-` by convention (`eduide-staging`, `eduide-cs101`, and so on). Commands use `-n ` - substitute your installation's namespace. The cluster-level chart installs into `eduide-system`, spelled out literally where it applies. `eduide` is the conventional Helm release name for an installation. + +::: + ## What is logged +Everything below is written to container stdout. EduIDE ships no log shipper, no log storage and no log retention mechanism of its own. Whatever you can query after a pod restarts is entirely a function of the log aggregation you run alongside it (Loki, Elasticsearch, a cloud provider's logging service). Without one, `kubectl logs` gives you the current container's output and nothing more, and a rolled pod takes its history with it. + ### EduIDE Cloud service The service logs all inbound requests including: - Request method, path, and response status - The authenticated user identity (from the JWT `username` claim) where applicable -- Admin endpoint access, including whether the `X-Admin-Api-Token` was accepted or rejected +- Admin endpoint access, including whether the `X-Admin-Api-Token` was accepted or rejected, and whether a token was configured at all - Workspace and session lifecycle events (create, delete, stop) -Logs are emitted to stdout and collected by the Kubernetes log aggregation pipeline (Rancher / Prometheus stack). +```bash +kubectl logs -n -l app=service --tail=200 +``` ### EduIDE Cloud operator @@ -27,87 +37,154 @@ The operator logs every reconciliation action: - Errors during reconciliation - Scaling changes applied to App Definitions -### OAuth2 proxy +```bash +kubectl logs -n -l app=operator --tail=200 +``` + +The operator is your primary audit trail for session and workspace lifecycle, because it is the component that actually creates and destroys them. -The OAuth2 proxy logs every authentication event: +### oauth2-proxy + +oauth2-proxy logs every authentication event it handles: - Successful logins with the authenticated username - Failed authentication attempts with the failure reason - Token validation results - Logout events +**There is no single login log stream.** oauth2-proxy runs as a **sidecar container inside each session pod**, not as a shared Deployment in front of the platform. Every session gets its own proxy instance, writing its own log, which lives and dies with that pod. + +The consequences for auditing are significant, and worth being explicit about: + +- There is no pod labelled `app=oauth2-proxy`. A selector like `kubectl logs -l app=oauth2-proxy` matches nothing and exits successfully, which reads like a clean authentication log and is not. +- Reading these logs means naming a pod and a container: + + ```bash + # Find session pods: everything that is not a platform component + kubectl get pods -n \ + -l 'app notin (operator,service,landing-page,image-preloading)' + + kubectl logs -n -c oauth2-proxy --tail=200 + ``` + +- **The logs are as ephemeral as the session.** When a session ends, its pod is deleted and its authentication log goes with it. There is no post-hoc way to ask "who logged into this platform last Tuesday" from EduIDE itself. +- Failed logins that never produced a session produce no oauth2-proxy log anywhere, because there was no session pod to run a proxy in. + +If authentication events matter for your compliance obligations, you have two options, and you need at least one of them: + +1. **Treat Keycloak as the authoritative authentication log.** Keycloak sees every login attempt for every session, records them centrally, and retains them independently of pod lifecycle. This is the right answer for almost everyone. +2. **Ship session pod logs off-cluster as they are produced**, so the oauth2-proxy sidecar output survives the pod. This gives you the proxy's view but requires log collection that captures short-lived pods reliably. + +Do not plan an audit process around retrieving oauth2-proxy logs with `kubectl` after the fact. That data is gone. + ### Keycloak Keycloak maintains its own audit event log. Relevant event types: -- `LOGIN` — successful user authentication -- `LOGIN_ERROR` — failed login attempt -- `LOGOUT` — user logout -- `TOKEN_EXCHANGE` — token operations -- `CLIENT_LOGIN` — service-to-Keycloak authentication +- `LOGIN` - successful user authentication +- `LOGIN_ERROR` - failed login attempt +- `LOGOUT` - user logout +- `TOKEN_EXCHANGE` - token operations +- `CLIENT_LOGIN` - service-to-Keycloak authentication - Admin events: user creation, group membership changes, client configuration changes -Keycloak audit events are accessed via the Keycloak admin console under **Realm Settings → Events**. +Keycloak audit events are accessed via the Keycloak admin console under **Realm Settings → Events**. Event logging and admin event logging are configured separately there, and both are off by default in a fresh Keycloak. Confirm both are enabled and that the expiration is set to at least your intended retention before relying on this. ### Workspace garbage collector The garbage collector logs each workspace deletion it performs, including the workspace name and its creation timestamp. If a deletion fails, the error is logged with the workspace name. +```bash +kubectl logs -n -l app=theia-workspace-garbage-collector --tail=200 +``` + +This is the only record that a user's workspace was reclaimed. It has no Service and no metrics endpoint, so the log is the whole story. + +### Helm + +`helm history` is an audit record of configuration change, and one that survives longer than most container logs: + +```bash +helm history eduide -n +helm get values eduide -n --revision +``` + +Note that `helm get values` will show secret material if you provisioned the admin API token through the values file. Treat its output accordingly. + ## Sensitive operations inventory The following operations should be treated as high-sensitivity events and reviewed if unusual patterns appear: | Operation | Where logged | Sensitivity | |---|---|---| -| Admin API token request (success) | Service logs | High — means the token is in use | -| Admin API token request (failure) | Service logs | High — may indicate brute force or stale token | -| App Definition scaling change | Service + operator logs | Medium — affects platform capacity | -| Workspace force-deletion by admin | Service/kubectl audit | Medium — user data deletion | -| Keycloak group membership change | Keycloak admin events | High — grants or removes elevated access | -| OAuth2 proxy cookie secret rotation | Deployment logs | Medium — invalidates all active sessions | -| New Keycloak client creation | Keycloak admin events | High — creates new authentication entry point | +| Admin API token request (success) | Service logs | High - means the token is in use | +| Admin API token request (failure) | Service logs | High - may indicate brute force or a stale token | +| App Definition scaling change | Service + operator logs | Medium - affects platform capacity | +| Workspace force-deletion by admin | Operator logs, plus the Kubernetes API audit log if enabled | Medium - user data deletion | +| Workspace deletion by TTL | Garbage collector logs | Medium - user data deletion, unattended | +| Keycloak group membership change | Keycloak admin events | High - grants or removes elevated access | +| oauth2-proxy cookie secret rotation | Helm history | Medium - invalidates all active sessions | +| New Keycloak client creation | Keycloak admin events | High - creates a new authentication entry point | +| Gateway listener or certificate change | Helm history for `eduide-cluster` | High - changes what hostnames the cluster serves | +| Chart upgrade of an installation | `helm history` | Medium - the record of what configuration was live when | + +A note on `kubectl` actions: deleting a workspace or session by hand leaves a trace in the operator's logs (it observes the deletion), but the record of **who** did it lives only in the Kubernetes API server audit log. That is a cluster-level feature you enable on the API server; it is not part of EduIDE. If attribution of admin actions matters to you, enable it. ## Log retention recommendations | Log source | Recommended retention | Rationale | |---|---|---| -| Service and operator logs (k8s) | 30 days | Sufficient for most incident investigations | -| OAuth2 proxy logs | 30 days | Covers typical authentication incident lookback | +| Service and operator logs | 30 days | Sufficient for most incident investigations | +| Session pod logs, including oauth2-proxy | 30 days, if collected at all | Only survives with off-cluster log shipping | | Keycloak event log | 90 days | Covers quarterly access reviews | -| Deployment workflow logs (GitHub) | 90 days | Covers audit of configuration changes | -| Garbage collector logs | 30 days | Low sensitivity, high volume | +| Kubernetes API audit log | 90 days | Attribution of admin actions | +| Garbage collector logs | 30 days | Low volume, but the only record of automated data deletion | +| Helm release history | Indefinite | Small, and it is your configuration change record | -Kubernetes default log retention depends on your cluster configuration. If you are using a centralised log aggregation system (e.g., Loki, Elasticsearch), set the retention policies there. +Kubernetes default log retention depends on your cluster configuration - typically the kubelet's per-container rotation, which is measured in megabytes rather than days. If you use a centralised log aggregation system, set the retention policies there. ## Access review checklist -Run through this checklist on the indicated cadence. +Run through this checklist on the indicated cadence. Substitute your own group and secret-store names where the checklist refers to them generically. ### Monthly -- [ ] Review the `theia-cloud/admin` group membership in Keycloak. Remove any users who no longer require admin access. -- [ ] Check service logs for any unexplained spikes in admin API token requests (successful or failed). -- [ ] Verify that all running environments correspond to active deployments. Decommission any orphaned environment namespaces. +- [ ] Review membership of the EduIDE admin group in Keycloak (the group named in your installation's Keycloak configuration). Remove any users who no longer require admin access. +- [ ] Check service logs for unexplained spikes in admin API token requests, successful or failed. +- [ ] Verify that every EduIDE namespace on the cluster corresponds to an installation you still intend to run: + ```bash + kubectl get namespace -o name | grep eduide- + helm list --all-namespaces --filter '^eduide$' + ``` + Decommission anything orphaned, using the procedure below. ### Quarterly - [ ] Review all Keycloak user accounts. Disable accounts for users who have left the organisation or completed their course. -- [ ] Review Keycloak client configurations. Remove stale redirect URIs for decommissioned environments. -- [ ] Review who has access to GitHub environment secrets for each deployment environment. Revoke access for anyone no longer actively deploying. -- [ ] Check GitHub Actions deployment logs for any unexpected deployments or configuration changes. +- [ ] Review Keycloak client configurations. Remove stale redirect URIs for decommissioned installations - these are the most commonly forgotten artefact of a decommissioning. +- [ ] Review who can read the secret store holding your admin API tokens and installation values files. Revoke access for anyone no longer operating the platform. +- [ ] Review `helm history` for each installation for upgrades you cannot account for. +- [ ] Check Gateway listeners in `eduide-system` against the installations that actually exist: + ```bash + kubectl get gateway -n eduide-system \ + -o jsonpath='{range .spec.listeners[*]}{.name}{"\t"}{.hostname}{"\n"}{end}' + ``` + A listener for a hostname no longer served is a loose end, not a fault. ### Every 6 months -- [ ] Rotate the admin API token for all environments. See [Admin API Tokens](/admins/security/admin-api-tokens). -- [ ] Rotate the OAuth2 proxy cookie secret for all environments. -- [ ] Review and update alert thresholds in the monitoring chart to reflect current cohort sizes. +- [ ] Rotate the admin API token for every installation. See [Admin API Tokens](/admins/security/admin-api-tokens). +- [ ] Rotate the oauth2-proxy cookie secret for every installation. This invalidates all active sessions, so schedule it. +- [ ] Review your alert thresholds against current cohort sizes. Note that the charts ship Grafana dashboards but **no alerting rules** - any alerts you have are ones you wrote, and they live wherever you defined them, not in the EduIDE charts. See [Monitoring Basics](/admins/operations/monitoring-basics). +- [ ] Verify certificate coverage for every hostname each installation serves. A certificate that has stopped covering a hostname produces no cluster-side error; see [Incident Response](/admins/operations/incident-response). ### On personnel changes When a person with platform access leaves or changes role: -- [ ] Remove them from the `theia-cloud/admin` Keycloak group immediately. -- [ ] Revoke their access to GitHub environment secrets. -- [ ] Rotate the admin API token if they had access to it. -- [ ] Rotate the OAuth2 proxy cookie secret if they had access to it. +- [ ] Remove them from the EduIDE admin group in Keycloak immediately. +- [ ] Revoke their access to the secret store and to installation values files. +- [ ] Revoke their `kubectl` credentials for the cluster. +- [ ] Rotate the admin API token for every installation they had access to. +- [ ] Rotate the oauth2-proxy cookie secret if they had access to it. - [ ] Disable their Keycloak account. Do not defer these steps. Credentials are not invalidated automatically when a person's role changes. @@ -115,9 +192,122 @@ Do not defer these steps. Credentials are not invalidated automatically when a p ## Checking for failed admin token attempts ```bash -# Search service logs for rejected admin token requests -kubectl logs -n theia-prod -l app=service --since=24h \ - | grep -i "admin.*token\|401\|403" +kubectl logs -n -l app=service --since=24h \ + | grep -i "admin API token" +``` + +The service logs a distinct message for each failure mode: a missing header, an invalid token, and no token configured at all. Reading which one you are getting matters. A steady stream of *invalid token* rejections without a matching change to your automation suggests either a stale script or an attempt to guess the token - rotate immediately if the source cannot be explained. A stream of *not configured* messages means the admin API was never provisioned on that installation, which is a configuration gap rather than a security event. + +## Decommissioning an installation + +Removing an installation is not just `helm uninstall`. Several artefacts live outside the installation's namespace and outside its Helm release, and every one of them is a loose end if left behind - a Gateway listener for a hostname that no longer resolves, a certificate SAN for a service that no longer exists, a Keycloak client that still accepts redirects to a domain you no longer control. + +Work through these in order. The ordering matters in two places, noted below. + +**Step 0: Deal with the data first** + +Uninstalling destroys student work. Confirm what exists and that anyone who needs it has been told: + +```bash +kubectl get workspaces.theia.cloud -n \ + -o custom-columns='NAME:.metadata.name,USER:.spec.user,CREATED:.metadata.creationTimestamp' +kubectl get pvc -n +``` + +Export anything that must be kept before going further. There is no undo after step 2. + +**Step 1: Delete the custom resources while the operator is still running** + +This ordering matters. Sessions and workspaces are reconciled by the operator, and the operator is part of the release you are about to uninstall. Removing the resources first lets the operator tear down pods, PVCs and route entries cleanly. Uninstalling first leaves the operator gone and any resources with finalizers stuck, which then need manual finalizer removal. + +```bash +kubectl -n delete sessions.theia.cloud --all +kubectl -n delete workspaces.theia.cloud --all +kubectl -n delete appdefinitions.theia.cloud --all ``` -A steady stream of 401 responses on `/service/admin/` endpoints without a matching successful deployment suggests either a misconfigured automation script or an attempted brute force. Rotate the token immediately if the source cannot be explained. +Wait for the session pods to disappear before continuing: + +```bash +kubectl get pods -n -w +``` + +**Step 2: Uninstall the installation release** + +```bash +helm uninstall eduide -n +``` + +This removes the operator, the service, the landing page, the garbage collector subchart, the HTTPRoutes and the ConfigMaps. It does **not** remove the CRDs - those belong to the cluster-level chart and are annotated to survive - and it does not remove anything in `eduide-system`. + +**Step 3: Re-run the cluster bootstrap so the Gateway and certificates shrink** + +This is the step most easily forgotten, because nothing breaks if you skip it. The shared Gateway in `eduide-system` still carries this installation's listeners, and the certificates still carry its hostnames in their SANs. + +Remove the installation from your cluster-level configuration: + +- Drop its listeners from `gateway.listeners`. There are typically four per installation (landing, service, instances, webview), plus any ACME challenge listeners. +- Drop its hostnames from the relevant entry in `managedCertificates.certificates`, so the certificate is re-issued with a smaller SAN list. +- Drop its namespace from `monitoring.targetNamespaces` and `monitoring.sessionNamespaces`. + +Then upgrade the cluster chart: + +```bash +helm upgrade eduide-cluster oci://ghcr.io/eduide/charts/eduide-cluster \ + --version \ + -n eduide-system \ + -f .yaml +``` + +Verify the listeners are gone and that the remaining ones are still healthy - a botched listener list breaks every other installation on the cluster, so do not skip the check: + +```bash +kubectl get gateway theia-shared-gateway -n eduide-system \ + -o jsonpath='{range .status.listeners[*]}{.name}{"\t"}{range .conditions[*]}{.type}={.status} {end}{"\n"}{end}' +``` + +If a certificate was re-issued with fewer names, confirm the surviving hostnames are still covered before you consider this step done. See the certificate runbook in [Incident Response](/admins/operations/incident-response). + +**Step 4: Remove the Keycloak configuration** + +In the Keycloak admin console, for the client this installation used: + +- Remove its redirect URIs and web origins. A redirect URI pointing at a domain you no longer control is a genuine security problem, not just untidiness: whoever acquires that hostname next can receive authorization codes issued for your realm. +- If the client existed solely for this installation, delete the client. +- Remove any groups or roles created solely for this installation's users. + +**Step 5: Delete the namespace** + +```bash +kubectl delete namespace +``` + +If this hangs, a resource with a finalizer is still present - usually a custom resource that survived step 1 because the operator was already gone. Find it: + +```bash +kubectl api-resources --verbs=list --namespaced -o name \ + | xargs -n1 kubectl get --show-kind --ignore-not-found -n +``` + +**Step 6: Reclaim the storage** + +With a `Retain` reclaim policy, the PersistentVolumes survive the namespace and still occupy capacity on your storage backend. They show up in the `Released` phase - a **PersistentVolume** phase, not a PVC phase, and cluster-scoped, so no namespace flag applies: + +```bash +kubectl get pv -o json \ + | jq -r '.items[] | select(.status.phase=="Released") + | select(.spec.claimRef.namespace=="") + | "\(.metadata.name)\t\(.spec.capacity.storage)"' +``` + +Delete them once you are certain the data is not needed: + +```bash +kubectl delete pv +``` + +Whether this frees the underlying storage depends on your CSI driver. Confirm with whoever operates your storage rather than assuming the capacity has returned. + +**Step 7: Record it** + +Note the decommissioning date, who authorised it, what data was exported and what was destroyed. This is the record you will want at the next quarterly review when someone asks why a namespace disappeared. diff --git a/docs/developer/projects/eduide-deployment.md b/docs/developer/projects/eduide-deployment.md index 96f0a07..2376a9d 100644 --- a/docs/developer/projects/eduide-deployment.md +++ b/docs/developer/projects/eduide-deployment.md @@ -12,12 +12,14 @@ The **EduIDE Deployment** repository is the central hub for the infrastructure-a ## Key Features - **Automated CI/CD**: Seamless deployment pipelines using GitHub Actions for various branches and environments. -- **Environment Management**: Specialized configurations for `production`, `staging`, and multiple `test` environments. -- **Custom Helm Charts**: - - `theia-cloud-combined`: A master chart that bundles all necessary components. - - `theia-appdefinitions`: Configures the custom IDE environments (images and resources). - - `theia-certificates`: Manages SSL/TLS certificates, including TUM-specific processes. - - `theia-monitoring`: Sets up Prometheus and Grafana dashboards for observability. +- **Environment Management**: One directory per installation under + `environments/`, split into `env.yaml` (how it is deployed) and `values.yaml` + (how the chart is configured). +- **No chart source.** The charts live in + [EduIDE-Helm](https://github.com/EduIDE/EduIDE-Helm) and are pulled from + `ghcr.io/eduide/charts`. There are two: `eduide-cluster` once per cluster, and + `eduide` once per installation. The five charts this repository used to carry + were consolidated into those two at 2.0.0. - **GitOps Workflow**: Deployments are managed through Git with approval gates and automated rollouts to staging. - **Authentication Integration**: Detailed configuration for Keycloak to manage user access and session security. diff --git a/sidebarsAdmins.ts b/sidebarsAdmins.ts index ca9d4c6..d749a58 100644 --- a/sidebarsAdmins.ts +++ b/sidebarsAdmins.ts @@ -6,7 +6,12 @@ const sidebars: SidebarsConfig = { { type: 'category', label: 'Install', - items: ['install/installing', 'install/adding-an-installation'], + items: [ + 'install/prerequisites', + 'install/certificates', + 'install/installing', + 'install/adding-an-installation', + ], }, { type: 'category',