What has to exist on a cluster before EduIDE can be deployed to it, which parts
Bootstrap cluster does for you, and which parts it does not. Written around
the TUM clusters, but the eduide cluster is not one of them and the places
that differ say so.
Read this end to end before bootstrapping a cluster for the first time. Two of the manual steps below cannot be discovered by trying: they produce a cluster that reports itself healthy and serves nothing.
| Who does it | |
|---|---|
| Gateway API CRDs, Envoy Gateway, cert-manager, storage | you, once per cluster |
| A data plane answering on the address DNS publishes - a load balancer's, or the node's | you |
The ACME ClusterIssuer |
Bootstrap cluster, from spec.acmeEmail |
| The GatewayClass and its EnvoyProxy | Bootstrap cluster, from spec.gatewayClass and spec.envoyProxy |
| Redirects from hostnames you used to serve | Bootstrap cluster, from spec.redirects |
| The webview wildcard certificate | you, see tum-certificates.md |
| DNS records | you (RBG) |
| Keycloak client and redirect URIs | you, see keycloak-setup.md |
The KUBECONFIG and certificate secrets |
you, see github-environments.md |
| CRDs, conversion webhook, ClusterRoles | Bootstrap cluster |
| The shared Gateway and all its listeners | Bootstrap cluster |
| The certificate covering every environment's hosts | Bootstrap cluster |
| PodMonitors, Grafana dashboards, watched namespaces | Bootstrap cluster |
| The cluster identity ConfigMap | Bootstrap cluster |
| The operator, service, landing page, routes, apps | Deploy |
Bootstrap cluster installs the eduide-cluster chart and nothing else. It
does not install a Gateway controller, cert-manager or a CNI, and it assumes the
platform underneath already works. It does create the ACME ClusterIssuer the
certificates it derives point at - see step 3.
Install these once per cluster. On TUM's clusters they are already present and owned by whoever runs the cluster, so in practice this step is a check rather than an install.
kubectl get crd gateways.gateway.networking.k8s.io # Gateway API
kubectl get gatewayclass # a controller, Accepted
kubectl -n cert-manager get deploy # cert-manager
kubectl get storageclass # dynamic provisioning
kubectl get crd podmonitors.monitoring.coreos.com # only if you want metricscert-manager must have been installed with config.enableGatewayAPI=true.
Without it, cert-manager ignores Gateway resources entirely and every HTTP-01
challenge stays pending forever. If it was installed without the flag, upgrade
and restart it:
kubectl -n cert-manager rollout restart deploy/cert-managerVerified as present on both TUM clusters at the time of writing:
tum-student (stud-cp) |
the current test cluster (theia-prod) |
|
|---|---|---|
| Gateway API CRDs | yes | yes |
| Envoy Gateway | yes | yes |
| cert-manager | yes | yes |
| Prometheus Operator CRDs | yes | yes |
cattle-monitoring-system, cattle-dashboards |
yes | yes |
| Storage classes | csi-rbd-sc (default), longhorn, longhorn-static |
csi-rbd-sc and longhorn, both default |
| Nodes | 28 |
The test cluster has two default StorageClasses. A PVC that names no class
gets an arbitrary one. Nothing in this repository can fix that; it is worth
raising with whoever owns the cluster. The deploy sidesteps it by always naming
spec.storageClassName from clusters/<name>.yaml explicitly.
28 nodes is a real number here. Image preloading pulls every IDE image onto every node, so the default set of eight is roughly 20 GB per node, once. Budget it before the first bootstrap rather than discovering it as disk pressure.
This is the step that produces a healthy-looking cluster that serves nothing, so do it before anything else.
Envoy Gateway can be configured to merge gateways: every Gateway using a
GatewayClass shares one Envoy deployment and therefore one external address.
tum-student is configured that way today:
GatewayClass envoy
-> parametersRef: EnvoyProxy envoy-gateway-system/artemis-envoy-proxy
mergeGateways: true
metallb.universe.tf/address-pool: lb2 -> 131.159.88.15
and the EduIDE hostnames point somewhere else:
eduide.student.k8s.aet.cit.tum.de -> k8s-stud-lb3 -> 131.159.88.14 (pool lb3)
So a Gateway created with gatewayClassName: envoy on that cluster comes up
Programmed=True, is served on 131.159.88.15, and is unreachable at every
name DNS actually publishes. Nothing reports an error.
There are three ways out, and it is a decision, not a default:
(a) Join the existing merged gateway. Ask for the EduIDE DNS names to point
at 131.159.88.15 instead. Nothing in this repository changes. EduIDE then
shares an Envoy data plane with Artemis, so a configuration mistake in either
can affect the other.
(b) Give EduIDE its own GatewayClass. Create a second GatewayClass with its
own EnvoyProxy pinned to pool lb3, and set gatewayClassName in
clusters/tum-student.yaml to match. EduIDE then has its own data plane on
131.159.88.14, which is what DNS already says.
(c) There is no load balancer at all. A single node cluster outside TUM may
have neither MetalLB nor servicelb, and then a Service of type LoadBalancer
sits Pending for ever while DNS publishes the node's own address. Bind the two
public ports on the node instead: envoyService.type: ClusterIP, and a
StrategicMerge patch on the Envoy Deployment giving its container hostPort: 80
and hostPort: 443. The eduide cluster does this, and
clusters/eduide.yaml carries the whole story - including why hostNetwork plus
useListenerPortAsContainerPort is the wrong answer: Envoy runs as non-root,
Kubernetes cannot grant NET_BIND_SERVICE effectively, and every listener then
fails with cannot bind '0.0.0.0:80': Permission denied while the pod reports
Running.
bootstrap-cluster.yml passes both from the cluster manifest, so options (b) and
(c) are spec.gatewayClass.create: true plus a spec.envoyProxy block that
sets create: true itself. The chart defaults envoyProxy.create to false,
so a block without it renders no EnvoyProxy at all while the GatewayClass still
gets a parametersRef naming one - a dangling reference, and no data plane.
envoyProxy.name is the name of the EnvoyProxy resource itself; for (b) the
MetalLB pool goes inside it, under
spec.provider.kubernetes.envoyService.annotations. That is exactly what
tum-production does. See
envoy-gateway-setup.md for the MetalLB details.
On (b) and (c), state the data plane's replica count. Envoy Gateway
reconciles the Envoy Deployment's replicas only while the EnvoyProxy names them;
left unset it writes the field once and never looks again, so one
kubectl scale --replicas=0 takes every Gateway on the class down until
somebody scales it back by hand. That is what took Bonn and Mannheim off the air
on 2026-09-23. Put replicas in the envoyDeployment block, as
clusters/eduide.yaml does.
Option (a) has no such block here, because the EnvoyProxy belongs to whoever owns the merged gateway. The exposure does not go away - it moves. Ask them whether their EnvoyProxy states its replicas, and remember that scaling that data plane to zero takes EduIDE down with everything else sharing it.
clusters/tum-production.yaml carries spec.loadBalancerIP: 131.159.88.82.
Nothing reads it. It records the intent; it does not enforce it.
Whichever option is chosen, verify it after bootstrap:
kubectl -n eduide-system get gateway theia-shared-gateway \
-o jsonpath='{.status.addresses[*].value}{"\n"}'
dig +short <the landing host of an environment on this cluster> AThose two must end at the same address. On a cluster serving option (c) the
Gateway's address is the Service's ClusterIP, which DNS never publishes - check
the node's own address instead, and that something answers on :443 there.
Bootstrap cluster derives a cert-manager Certificate covering every
environment's landing, service and instance hostnames, and creates the
ClusterIssuer it points at rather than assuming a suitable one exists. That
happens whenever the cluster sets spec.tls.acmeHttp: true, and it needs one
thing from the cluster manifest:
spec:
acmeEmail: admin.aet@xcit.tum.de # required when acmeHttp is true
acmeIssuerName: letsencrypt-prod-gateway # optional, this is the defaultcert-manager will not register an ACME account without a contact address, so
bootstrap fails early if it is missing, and test-deploy-logic.sh catches it
before that.
:::warning Do not reuse the existing letsencrypt-prod
Both TUM clusters carry a ClusterIssuer of that name whose only solver is an
nginx Ingress solver:
solvers:
- http01:
ingress:
class: nginxcert-manager cannot answer a Gateway API challenge with that. Certificates
pointed at it stay pending indefinitely, and the Gateway keeps reporting
Programmed=True while serving whatever the secret already held. The issuer
bootstrap creates uses http01.gatewayHTTPRoute against the shared Gateway.
:::
The wildcard webview certificate is not issued this way and never can be.
ACME does not permit HTTP-01 for wildcards. Those hosts keep a long-lived
certificate supplied through wildcardTLSSecret; see
tum-certificates.md.
Four records per environment, all pointing at the address from step 2:
<landing> the landing page
service.<landing> the REST service
instance.<landing> session ingress
*.webview.instance.<landing> per-session webviews
The fourth is a wildcard. RBG issues these; ask early.
Current state:
| Host | Resolves |
|---|---|
*.eduide.student.k8s.aet.cit.tum.de |
131.159.88.14 |
eduide.artemis.cit.tum.de |
131.159.88.82 |
bonn.eduide.aet.cit.tum.de, mannheim.… |
131.159.88.106 (parma itself) |
Bootstrap reads its KUBECONFIG and the wildcard certificate from a GitHub
Environment named in clusters/<name>.yaml as spec.bootstrapEnvironment.
Deploying reads its own KUBECONFIG from a GitHub Environment named after the
environment. They are different environments and hold different secrets.
Full instructions, including what is missing today, are in github-environments.md.
Actions -> Bootstrap cluster
cluster: tum-student
chart_version: 2.1.0
dry_run: true
Read the rendered values and the helm diff in the job summary. The workflow
prints the listeners, watched namespaces and certificate names it derived, so
this is where a wrong sectionName or a missing environment shows up.
What it derives, from the environments that claim the cluster:
- Gateway listeners, from each environment's
gateway.parentRefs[].sectionNameplus its hosts. Four per environment, plus a plain:80listener per non-wildcard host when the cluster setsspec.tls.acmeHttp: true. - The TLS secret per listener, from
spec.tlsin the cluster manifest. A listener with no secret renders an emptycertificateRefs, is accepted, and never programs TLS, so the chart fails the render instead. - The certificate's
dnsNames, from the same pass, so a listener and its certificate cannot disagree. Only landing, service and instance hosts go on it: a name without a listener answers 404 to its HTTP-01 challenge, stays pending, and blocks the certificate for every other name on it. - The namespaces the PodMonitors watch, from each environment's
spec.namespace, skipping any that setmonitoring.enabled: false.
It refuses to run when no environment claims the cluster, and it refuses to
bootstrap a cluster that already carries a different name in
eduide-system/eduide-cluster-identity - that almost always means the
KUBECONFIG on the GitHub Environment points somewhere unexpected.
Then run it again with dry_run: false.
kubectl -n eduide-system get cm eduide-cluster-identity \
-o jsonpath='{.data.clusterName}{"\n"}'
kubectl get crd | grep theia.cloud # 3 CRDs
kubectl -n eduide-system get gateway -o wide # Programmed, with an address
kubectl -n eduide-system get certificate # Ready=True
kubectl -n cattle-monitoring-system get podmonitor # 2Every listener should report Programmed. A listener that reports
ResolvedRefs=False names a Secret that does not exist yet, which is normal
until the certificate issues for the first time.
Do not check TLS with curl -k. It suppresses exactly the failure that is
worth finding. Check each hostname against the certificate the server actually
presents:
for h in test1.eduide.student.k8s.aet.cit.tum.de \
service.test1.eduide.student.k8s.aet.cit.tum.de \
instance.test1.eduide.student.k8s.aet.cit.tum.de; do
echo | openssl s_client -connect "$h:443" -servername "$h" 2>/dev/null \
| openssl x509 -noout -checkhost "$h"
doneEach must print Host <name> matches certificate. Gateway API never compares a
certificate's names against its listener's hostname, so a listener holding a
certificate for entirely different hosts reports Programmed=True ResolvedRefs=True and looks perfect. test3 ran that way for 184 days: the
landing page loaded past the browser warning, then silently could not call its
own REST service, because the browser blocks a cross-origin request over an
invalid certificate. Nothing appeared in any log.
Actions -> Deploy (dispatch) -> environment: test1, dry_run: true
The deploy asserts the cluster identity before touching anything, so a
KUBECONFIG pointing at the wrong cluster stops here rather than installing
EduIDE somewhere unexpected.
-
clusters/<name>.yaml. It is validated againstschemas/cluster.schema.json, which rejects unknown keys.apiVersion: eduide.dev/v1 kind: Cluster metadata: name: eduide displayName: EduIDE cluster spec: storageClassName: local # must exist on the cluster gatewayClassName: envoy # see step 2 before accepting this default tls: # the Secret each listener role terminates with landing: shared-theia-cert service: shared-theia-cert instances: shared-theia-cert webview: static-theia-cert # always separate - two labels lower acmeHttp: false # true adds :80 listeners for HTTP-01 acmeEmail: you@example.edu # required when acmeHttp is true sharedGateway: { namespace: eduide-system, name: theia-shared-gateway } runner: ubuntu-latest # must be able to reach the API server bootstrapEnvironment: cluster-eduide
-
Add the name to the
clusterchoice list in.github/workflows/bootstrap-cluster.yml. Workflow inputs cannot be derived from a file. -
Create the GitHub Environment named in
bootstrapEnvironment. -
Work through steps 1 to 7 above.
runner is per cluster because deploying is not building: a deploy has to reach
the API server, and the clusters may differ in how they are reachable.
Two failures look identical from outside - the listener never programs - and have different causes.
No challenges outstanding, order errored. Let's Encrypt occasionally fails
to finalize an order that has already validated:
Failed to finalize Order: 404 urn:ietf:params:acme:error:malformed:
Certificate not found
Confirm it is that failure before doing anything, by reading down the chain the
Certificate owns - its conditions alone do not tell you:
kubectl -n eduide-system describe certificate <name>
kubectl -n eduide-system get certificaterequest,order,challengecert-manager retries on its own, but only after an exponential backoff that
starts at an hour. Clear the backoff to retry immediately (--subresource
needs kubectl v1.24 or newer):
kubectl -n eduide-system patch certificate <name> --type=merge --subresource=status \
-p '{"status":{"lastFailureTime":null,"failedIssuanceAttempts":null}}'It issued on the first retry when this happened during the tum-production
bootstrap, so one clean retry is the expected outcome and no configuration
change is called for. If the retry fails too, the transient explanation is
wrong: go back over the CertificateRequest, Order, Challenge and
ClusterIssuer, and the cert-manager controller logs, before changing anything.
Challenges pending, staying pending. That is a real problem, and the error
on the Challenge says which:
- HTTP 404 - something answered, but nothing routes the solver path. The hostname has no listener on the Gateway, or the listener exists on a Gateway the solver's HTTPRoute does not attach to.
- DNS, connection refused or timeout - nothing answered at all. The name does not resolve, or it resolves somewhere that is not this Gateway.
See step 3.
| Gap | Consequence |
|---|---|
spec.loadBalancerIP is in the schema and in tum-production.yaml |
Nothing reads it. spec.envoyProxy is what actually pins the address (step 2) |