Skip to content

Latest commit

 

History

History
408 lines (324 loc) · 17.5 KB

File metadata and controls

408 lines (324 loc) · 17.5 KB

Preparing a cluster

What has to exist on a cluster before EduIDE can be deployed to it, which parts Bootstrap cluster does for you, and which parts it does not. Written around the TUM clusters, but the eduide cluster is not one of them and the places that differ say so.

Read this end to end before bootstrapping a cluster for the first time. Two of the manual steps below cannot be discovered by trying: they produce a cluster that reports itself healthy and serves nothing.

The short version

Who does it
Gateway API CRDs, Envoy Gateway, cert-manager, storage you, once per cluster
A data plane answering on the address DNS publishes - a load balancer's, or the node's you
The ACME ClusterIssuer Bootstrap cluster, from spec.acmeEmail
The GatewayClass and its EnvoyProxy Bootstrap cluster, from spec.gatewayClass and spec.envoyProxy
Redirects from hostnames you used to serve Bootstrap cluster, from spec.redirects
The webview wildcard certificate you, see tum-certificates.md
DNS records you (RBG)
Keycloak client and redirect URIs you, see keycloak-setup.md
The KUBECONFIG and certificate secrets you, see github-environments.md
CRDs, conversion webhook, ClusterRoles Bootstrap cluster
The shared Gateway and all its listeners Bootstrap cluster
The certificate covering every environment's hosts Bootstrap cluster
PodMonitors, Grafana dashboards, watched namespaces Bootstrap cluster
The cluster identity ConfigMap Bootstrap cluster
The operator, service, landing page, routes, apps Deploy

Bootstrap cluster installs the eduide-cluster chart and nothing else. It does not install a Gateway controller, cert-manager or a CNI, and it assumes the platform underneath already works. It does create the ACME ClusterIssuer the certificates it derives point at - see step 3.

Step 1: the platform layer

Install these once per cluster. On TUM's clusters they are already present and owned by whoever runs the cluster, so in practice this step is a check rather than an install.

kubectl get crd gateways.gateway.networking.k8s.io      # Gateway API
kubectl get gatewayclass                                # a controller, Accepted
kubectl -n cert-manager get deploy                      # cert-manager
kubectl get storageclass                                # dynamic provisioning
kubectl get crd podmonitors.monitoring.coreos.com       # only if you want metrics

cert-manager must have been installed with config.enableGatewayAPI=true. Without it, cert-manager ignores Gateway resources entirely and every HTTP-01 challenge stays pending forever. If it was installed without the flag, upgrade and restart it:

kubectl -n cert-manager rollout restart deploy/cert-manager

Verified as present on both TUM clusters at the time of writing:

tum-student (stud-cp) the current test cluster (theia-prod)
Gateway API CRDs yes yes
Envoy Gateway yes yes
cert-manager yes yes
Prometheus Operator CRDs yes yes
cattle-monitoring-system, cattle-dashboards yes yes
Storage classes csi-rbd-sc (default), longhorn, longhorn-static csi-rbd-sc and longhorn, both default
Nodes 28

The test cluster has two default StorageClasses. A PVC that names no class gets an arbitrary one. Nothing in this repository can fix that; it is worth raising with whoever owns the cluster. The deploy sidesteps it by always naming spec.storageClassName from clusters/<name>.yaml explicitly.

28 nodes is a real number here. Image preloading pulls every IDE image onto every node, so the default set of eight is roughly 20 GB per node, once. Budget it before the first bootstrap rather than discovering it as disk pressure.

Step 2: decide which address serves EduIDE

This is the step that produces a healthy-looking cluster that serves nothing, so do it before anything else.

Envoy Gateway can be configured to merge gateways: every Gateway using a GatewayClass shares one Envoy deployment and therefore one external address. tum-student is configured that way today:

GatewayClass envoy
  -> parametersRef: EnvoyProxy envoy-gateway-system/artemis-envoy-proxy
       mergeGateways: true
       metallb.universe.tf/address-pool: lb2   ->  131.159.88.15

and the EduIDE hostnames point somewhere else:

eduide.student.k8s.aet.cit.tum.de  ->  k8s-stud-lb3  ->  131.159.88.14   (pool lb3)

So a Gateway created with gatewayClassName: envoy on that cluster comes up Programmed=True, is served on 131.159.88.15, and is unreachable at every name DNS actually publishes. Nothing reports an error.

There are three ways out, and it is a decision, not a default:

(a) Join the existing merged gateway. Ask for the EduIDE DNS names to point at 131.159.88.15 instead. Nothing in this repository changes. EduIDE then shares an Envoy data plane with Artemis, so a configuration mistake in either can affect the other.

(b) Give EduIDE its own GatewayClass. Create a second GatewayClass with its own EnvoyProxy pinned to pool lb3, and set gatewayClassName in clusters/tum-student.yaml to match. EduIDE then has its own data plane on 131.159.88.14, which is what DNS already says.

(c) There is no load balancer at all. A single node cluster outside TUM may have neither MetalLB nor servicelb, and then a Service of type LoadBalancer sits Pending for ever while DNS publishes the node's own address. Bind the two public ports on the node instead: envoyService.type: ClusterIP, and a StrategicMerge patch on the Envoy Deployment giving its container hostPort: 80 and hostPort: 443. The eduide cluster does this, and clusters/eduide.yaml carries the whole story - including why hostNetwork plus useListenerPortAsContainerPort is the wrong answer: Envoy runs as non-root, Kubernetes cannot grant NET_BIND_SERVICE effectively, and every listener then fails with cannot bind '0.0.0.0:80': Permission denied while the pod reports Running.

bootstrap-cluster.yml passes both from the cluster manifest, so options (b) and (c) are spec.gatewayClass.create: true plus a spec.envoyProxy block that sets create: true itself. The chart defaults envoyProxy.create to false, so a block without it renders no EnvoyProxy at all while the GatewayClass still gets a parametersRef naming one - a dangling reference, and no data plane. envoyProxy.name is the name of the EnvoyProxy resource itself; for (b) the MetalLB pool goes inside it, under spec.provider.kubernetes.envoyService.annotations. That is exactly what tum-production does. See envoy-gateway-setup.md for the MetalLB details.

On (b) and (c), state the data plane's replica count. Envoy Gateway reconciles the Envoy Deployment's replicas only while the EnvoyProxy names them; left unset it writes the field once and never looks again, so one kubectl scale --replicas=0 takes every Gateway on the class down until somebody scales it back by hand. That is what took Bonn and Mannheim off the air on 2026-09-23. Put replicas in the envoyDeployment block, as clusters/eduide.yaml does.

Option (a) has no such block here, because the EnvoyProxy belongs to whoever owns the merged gateway. The exposure does not go away - it moves. Ask them whether their EnvoyProxy states its replicas, and remember that scaling that data plane to zero takes EduIDE down with everything else sharing it.

clusters/tum-production.yaml carries spec.loadBalancerIP: 131.159.88.82. Nothing reads it. It records the intent; it does not enforce it.

Whichever option is chosen, verify it after bootstrap:

kubectl -n eduide-system get gateway theia-shared-gateway \
  -o jsonpath='{.status.addresses[*].value}{"\n"}'
dig +short <the landing host of an environment on this cluster> A

Those two must end at the same address. On a cluster serving option (c) the Gateway's address is the Service's ClusterIP, which DNS never publishes - check the node's own address instead, and that something answers on :443 there.

Step 3: the ACME issuer

Bootstrap cluster derives a cert-manager Certificate covering every environment's landing, service and instance hostnames, and creates the ClusterIssuer it points at rather than assuming a suitable one exists. That happens whenever the cluster sets spec.tls.acmeHttp: true, and it needs one thing from the cluster manifest:

spec:
  acmeEmail: admin.aet@xcit.tum.de        # required when acmeHttp is true
  acmeIssuerName: letsencrypt-prod-gateway   # optional, this is the default

cert-manager will not register an ACME account without a contact address, so bootstrap fails early if it is missing, and test-deploy-logic.sh catches it before that.

:::warning Do not reuse the existing letsencrypt-prod Both TUM clusters carry a ClusterIssuer of that name whose only solver is an nginx Ingress solver:

solvers:
  - http01:
      ingress:
        class: nginx

cert-manager cannot answer a Gateway API challenge with that. Certificates pointed at it stay pending indefinitely, and the Gateway keeps reporting Programmed=True while serving whatever the secret already held. The issuer bootstrap creates uses http01.gatewayHTTPRoute against the shared Gateway. :::

The wildcard webview certificate is not issued this way and never can be. ACME does not permit HTTP-01 for wildcards. Those hosts keep a long-lived certificate supplied through wildcardTLSSecret; see tum-certificates.md.

Step 4: DNS

Four records per environment, all pointing at the address from step 2:

<landing>                             the landing page
service.<landing>                     the REST service
instance.<landing>                    session ingress
*.webview.instance.<landing>          per-session webviews

The fourth is a wildcard. RBG issues these; ask early.

Current state:

Host Resolves
*.eduide.student.k8s.aet.cit.tum.de 131.159.88.14
eduide.artemis.cit.tum.de 131.159.88.82
bonn.eduide.aet.cit.tum.de, mannheim.… 131.159.88.106 (parma itself)

Step 5: the GitHub Environment

Bootstrap reads its KUBECONFIG and the wildcard certificate from a GitHub Environment named in clusters/<name>.yaml as spec.bootstrapEnvironment. Deploying reads its own KUBECONFIG from a GitHub Environment named after the environment. They are different environments and hold different secrets.

Full instructions, including what is missing today, are in github-environments.md.

Step 6: bootstrap

Actions -> Bootstrap cluster
  cluster:       tum-student
  chart_version: 2.1.0
  dry_run:       true

Read the rendered values and the helm diff in the job summary. The workflow prints the listeners, watched namespaces and certificate names it derived, so this is where a wrong sectionName or a missing environment shows up.

What it derives, from the environments that claim the cluster:

  • Gateway listeners, from each environment's gateway.parentRefs[].sectionName plus its hosts. Four per environment, plus a plain :80 listener per non-wildcard host when the cluster sets spec.tls.acmeHttp: true.
  • The TLS secret per listener, from spec.tls in the cluster manifest. A listener with no secret renders an empty certificateRefs, is accepted, and never programs TLS, so the chart fails the render instead.
  • The certificate's dnsNames, from the same pass, so a listener and its certificate cannot disagree. Only landing, service and instance hosts go on it: a name without a listener answers 404 to its HTTP-01 challenge, stays pending, and blocks the certificate for every other name on it.
  • The namespaces the PodMonitors watch, from each environment's spec.namespace, skipping any that set monitoring.enabled: false.

It refuses to run when no environment claims the cluster, and it refuses to bootstrap a cluster that already carries a different name in eduide-system/eduide-cluster-identity - that almost always means the KUBECONFIG on the GitHub Environment points somewhere unexpected.

Then run it again with dry_run: false.

Step 7: verify

kubectl -n eduide-system get cm eduide-cluster-identity \
  -o jsonpath='{.data.clusterName}{"\n"}'
kubectl get crd | grep theia.cloud                    # 3 CRDs
kubectl -n eduide-system get gateway -o wide          # Programmed, with an address
kubectl -n eduide-system get certificate              # Ready=True
kubectl -n cattle-monitoring-system get podmonitor    # 2

Every listener should report Programmed. A listener that reports ResolvedRefs=False names a Secret that does not exist yet, which is normal until the certificate issues for the first time.

Do not check TLS with curl -k. It suppresses exactly the failure that is worth finding. Check each hostname against the certificate the server actually presents:

for h in test1.eduide.student.k8s.aet.cit.tum.de \
         service.test1.eduide.student.k8s.aet.cit.tum.de \
         instance.test1.eduide.student.k8s.aet.cit.tum.de; do
  echo | openssl s_client -connect "$h:443" -servername "$h" 2>/dev/null \
    | openssl x509 -noout -checkhost "$h"
done

Each must print Host <name> matches certificate. Gateway API never compares a certificate's names against its listener's hostname, so a listener holding a certificate for entirely different hosts reports Programmed=True ResolvedRefs=True and looks perfect. test3 ran that way for 184 days: the landing page loaded past the browser warning, then silently could not call its own REST service, because the browser blocks a cross-origin request over an invalid certificate. Nothing appeared in any log.

Step 8: deploy an environment

Actions -> Deploy (dispatch) -> environment: test1, dry_run: true

The deploy asserts the cluster identity before touching anything, so a KUBECONFIG pointing at the wrong cluster stops here rather than installing EduIDE somewhere unexpected.

Adding a cluster

  1. clusters/<name>.yaml. It is validated against schemas/cluster.schema.json, which rejects unknown keys.

    apiVersion: eduide.dev/v1
    kind: Cluster
    metadata:
      name: eduide
      displayName: EduIDE cluster
    spec:
      storageClassName: local        # must exist on the cluster
      gatewayClassName: envoy        # see step 2 before accepting this default
      tls:                           # the Secret each listener role terminates with
        landing: shared-theia-cert
        service: shared-theia-cert
        instances: shared-theia-cert
        webview: static-theia-cert   # always separate - two labels lower
        acmeHttp: false              # true adds :80 listeners for HTTP-01
      acmeEmail: you@example.edu     # required when acmeHttp is true
      sharedGateway: { namespace: eduide-system, name: theia-shared-gateway }
      runner: ubuntu-latest          # must be able to reach the API server
      bootstrapEnvironment: cluster-eduide
  2. Add the name to the cluster choice list in .github/workflows/bootstrap-cluster.yml. Workflow inputs cannot be derived from a file.

  3. Create the GitHub Environment named in bootstrapEnvironment.

  4. Work through steps 1 to 7 above.

runner is per cluster because deploying is not building: a deploy has to reach the API server, and the clusters may differ in how they are reachable.

When a certificate will not issue

Two failures look identical from outside - the listener never programs - and have different causes.

No challenges outstanding, order errored. Let's Encrypt occasionally fails to finalize an order that has already validated:

Failed to finalize Order: 404 urn:ietf:params:acme:error:malformed:
Certificate not found

Confirm it is that failure before doing anything, by reading down the chain the Certificate owns - its conditions alone do not tell you:

kubectl -n eduide-system describe certificate <name>
kubectl -n eduide-system get certificaterequest,order,challenge

cert-manager retries on its own, but only after an exponential backoff that starts at an hour. Clear the backoff to retry immediately (--subresource needs kubectl v1.24 or newer):

kubectl -n eduide-system patch certificate <name> --type=merge --subresource=status \
  -p '{"status":{"lastFailureTime":null,"failedIssuanceAttempts":null}}'

It issued on the first retry when this happened during the tum-production bootstrap, so one clean retry is the expected outcome and no configuration change is called for. If the retry fails too, the transient explanation is wrong: go back over the CertificateRequest, Order, Challenge and ClusterIssuer, and the cert-manager controller logs, before changing anything.

Challenges pending, staying pending. That is a real problem, and the error on the Challenge says which:

  • HTTP 404 - something answered, but nothing routes the solver path. The hostname has no listener on the Gateway, or the listener exists on a Gateway the solver's HTTPRoute does not attach to.
  • DNS, connection refused or timeout - nothing answered at all. The name does not resolve, or it resolves somewhere that is not this Gateway.

See step 3.

Known gaps

Gap Consequence
spec.loadBalancerIP is in the schema and in tum-production.yaml Nothing reads it. spec.envoyProxy is what actually pins the address (step 2)