Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion charts/eduide-cluster/Chart.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -4,5 +4,5 @@ description: |
Cluster-scoped half of EduIDE: CRDs, the conversion webhook, ClusterRoles and
cert-manager issuers. Install once per cluster, before any eduide release.
type: application
version: 2.1.5
version: 2.2.0
appVersion: "1.2.0"
27 changes: 25 additions & 2 deletions charts/eduide-cluster/README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# eduide-cluster

![Version: 2.1.5](https://img.shields.io/badge/Version-2.1.5-informational?style=flat-square) ![Type: application](https://img.shields.io/badge/Type-application-informational?style=flat-square) ![AppVersion: 1.2.0](https://img.shields.io/badge/AppVersion-1.2.0-informational?style=flat-square)
![Version: 2.2.0](https://img.shields.io/badge/Version-2.2.0-informational?style=flat-square) ![Type: application](https://img.shields.io/badge/Type-application-informational?style=flat-square) ![AppVersion: 1.2.0](https://img.shields.io/badge/AppVersion-1.2.0-informational?style=flat-square)

Cluster-scoped half of EduIDE: CRDs, the conversion webhook, ClusterRoles and
cert-manager issuers. Install once per cluster, before any eduide release.
Expand Down Expand Up @@ -50,10 +50,33 @@ cert-manager issuers. Install once per cluster, before any eduide release.
| managedCertificates.enabled | bool | `false` | |
| managedCertificates.issuerRef.kind | string | `"ClusterIssuer"` | |
| managedCertificates.issuerRef.name | string | `"letsencrypt-prod"` | |
| monitoring | object | `{"dashboardNamespace":"cattle-dashboards","enabled":false,"namespace":"cattle-monitoring-system","sessionNamespaces":[],"targetNamespaces":[]}` | ------------------------------------------------------------------------ |
| monitoring | object | `{"alerting":{"channels":[],"enabled":false,"grafanaUrl":"","groupInterval":"5m","groupWait":"30s","minSeverity":"warning","namespace":"eduide-system","repeatInterval":"4h","runbookUrl":"https://eduide.github.io/Docs/admins/operations","thresholds":{"certExpiryDays":21,"componentRestarts":3,"sessionCrashPerHour":5,"sessionOOMPerHour":3,"startupSeconds":90,"volumePercent":85},"webhookSecret":{"create":false,"data":{},"name":"eduide-alert-webhooks"}},"certManager":{"enabled":false,"namespace":"cert-manager","portName":"http-metrics","selectorLabels":{"app.kubernetes.io/component":"controller","app.kubernetes.io/name":"cert-manager"}},"dashboardNamespace":"cattle-dashboards","enabled":false,"namespace":"cattle-monitoring-system","targetNamespaces":[]}` | ------------------------------------------------------------------------ |
| monitoring.alerting.channels | list | `[]` | Where to send alerts. Each entry needs `name`, `type` (`slack` or `discord`) and `secretKey`; Slack entries may also set `channel`. An empty list with alerting enabled means alerts fire but notify nobody, so the chart fails the render instead. A channel may set `environments: [ns, ...]` to receive only that installation's alerts. One cluster can host installations belonging to different people - Bonn and Mannheim share a cluster and each has its own Discord - and without scoping, both would see the other's incidents. Anything no scoped channel claims goes to **every** channel. That is how cluster-scoped alerts still get out: a certificate expiring or the conversion webhook failing is not about any one tenant's namespace, and dropping it for failing to match a tenant route would lose exactly the alerts that matter most. channels: - name: mannheim type: discord secretKey: discord-mannheim environments: [eduide-mannheim] - name: platform type: slack secretKey: slack-platform channel: "#eduide-alerts" |
| monitoring.alerting.enabled | bool | `false` | Create the PrometheusRule and the AlertmanagerConfig. |
| monitoring.alerting.grafanaUrl | string | `""` | Base URL of the Grafana that serves the EduIDE dashboards, with no trailing slash. Used to deep-link a notification straight to the affected session. Left empty, notifications carry no dashboard link. |
| monitoring.alerting.groupInterval | string | `"5m"` | How long to wait before notifying about new alerts added to a group. |
| monitoring.alerting.groupWait | string | `"30s"` | How long to wait for more alerts in a group before the first notification. |
| monitoring.alerting.minSeverity | string | `"warning"` | Lowest severity that reaches the notification channels: `warning` or `critical`. Anything below stays in Alertmanager and on the dashboards. This is what keeps the channels worth reading. |
| monitoring.alerting.namespace | string | `"eduide-system"` | Namespace both resources go in. This is also the `namespace` label every EduIDE alert carries, and it is a routing artifact rather than the environment the alert is about. The Prometheus Operator defaults `alertmanagerConfigMatcherStrategy` to `OnNamespace`, which prepends `namespace = <the AlertmanagerConfig's own namespace>` to every route generated from it, so an alert must claim this namespace to reach a receiver at all. The environment an alert is about is in `eduide_namespace`. Silence on that label, never on `namespace` - `namespace` is identical across every EduIDE alert and silencing it silences all of them. |
| monitoring.alerting.repeatInterval | string | `"4h"` | How long before an unresolved alert is sent again. |
| monitoring.alerting.runbookUrl | string | `"https://eduide.github.io/Docs/admins/operations"` | Base URL of the operations runbooks, with no trailing slash. |
| monitoring.alerting.thresholds.certExpiryDays | int | `21` | Days before certificate expiry to start warning. |
| monitoring.alerting.thresholds.componentRestarts | int | `3` | Restarts of a platform container within 15m before it counts as crash looping. |
| monitoring.alerting.thresholds.sessionCrashPerHour | int | `5` | Session crashes per hour in one environment before alerting. |
| monitoring.alerting.thresholds.sessionOOMPerHour | int | `3` | Session OOM kills per hour in one environment before alerting. |
| monitoring.alerting.thresholds.startupSeconds | int | `90` | p95 session startup, in seconds, that counts as too slow. |
| monitoring.alerting.thresholds.volumePercent | int | `85` | Workspace volume fill percentage that counts as nearly full. |
| monitoring.alerting.webhookSecret.create | bool | `false` | Create the Secret from `data` below. Off means the Secret already exists and is referenced by name only. |
| monitoring.alerting.webhookSecret.data | object | `{}` | Base64-encoded webhook URLs, keyed by the `secretKey` a channel names. Supplied by the workflow, not committed. |
| monitoring.alerting.webhookSecret.name | string | `"eduide-alert-webhooks"` | Name of the Secret channels read their URL from. |
| monitoring.certManager.enabled | bool | `false` | Create a ServiceMonitor for cert-manager's controller metrics. |
| monitoring.certManager.namespace | string | `"cert-manager"` | Namespace cert-manager runs in. |
| monitoring.certManager.portName | string | `"http-metrics"` | Name of the metrics port on that Service. |
| monitoring.certManager.selectorLabels | object | `{"app.kubernetes.io/component":"controller","app.kubernetes.io/name":"cert-manager"}` | Labels identifying cert-manager's controller Service. A ServiceMonitor selects on Service **labels**; there is no matcher for `metadata.name`. These are the labels cert-manager's own chart sets, and they are values because a release installed under a different name keeps the labels and changes their values. Get them wrong and the ServiceMonitor is created with no targets and reports no error. Check with: kubectl -n cert-manager get svc cert-manager --show-labels |
| monitoring.dashboardNamespace | string | `"cattle-dashboards"` | Namespace the Grafana dashboard ConfigMaps go in. Must already exist and be watched by Grafana's sidecar. `cattle-dashboards` is Rancher's. |
| monitoring.enabled | bool | `false` | Create the PodMonitors and Grafana dashboards. Off by default because it is not portable: it needs the Prometheus Operator CRDs (`monitoring.coreos.com/v1`) to exist, and the two namespaces below are Rancher's. Helm does not create namespaces it was not told to, so on a cluster without them the install fails outright on the dashboards. Turn it on once you know where your Prometheus and Grafana look for these. |
| monitoring.namespace | string | `"cattle-monitoring-system"` | Namespace the PodMonitors go in. Must be somewhere your Prometheus discovers. `cattle-monitoring-system` is Rancher's; change it for any other monitoring stack. |
| monitoring.targetNamespaces | list | `[]` | Namespaces to watch. Bootstrap derives this from the environments on the cluster; it is also the namespace list every dashboard's picker offers. |
| operatorrole.name | string | `"operator-api-access"` | name for the operator's cluster role |
| servicerole.name | string | `"service-api-access"` | name for the services' cluster role |
| wildcardTLSSecret.certificate | string | `""` | |
Expand Down
101 changes: 101 additions & 0 deletions charts/eduide-cluster/templates/_helpers.tpl
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
{{/*
A RE2 alternation of the namespaces this cluster monitors, anchored, for use in
a PromQL label matcher and in Grafana variable queries.

The dashboards used to carry a hand-written list of namespaces. It went stale
the moment the environments were renamed - the picker still offered `theia`,
`theia-staging` and `test1`, none of which exist any more, so every panel was
empty whatever you selected. Deriving it from the same list the PodMonitors use
means it cannot drift again.

An empty list yields `^$`, which matches nothing. That is deliberate: a cluster
where every environment opted out of monitoring should show an empty picker,
not silently fall back to `.*` and start graphing other tenants' namespaces.
*/}}
{{- define "eduide.monitoredNamespaceRegex" -}}
{{- $ns := .Values.monitoring.targetNamespaces | default list -}}
{{- if $ns -}}
^({{ join "|" $ns }})$
{{- else -}}
^$
{{- end -}}
{{- end -}}

{{/*
Alerting depends on monitoring, and saying so beats rendering nothing.

Both the PrometheusRule and the AlertmanagerConfig are gated on
`monitoring.enabled` as well as `monitoring.alerting.enabled`, because alerts
built on metrics nobody scrapes are decoration. Asking for alerting with
monitoring off used to render neither resource and succeed, which is the worst
outcome: the install reports fine and the feature is absent.
*/}}
{{- define "eduide.checkAlertingPrerequisites" -}}
{{- if and .Values.monitoring.alerting.enabled (not .Values.monitoring.enabled) -}}
{{- fail "monitoring.alerting.enabled is true but monitoring.enabled is false - alerting needs the PodMonitors and the metrics they collect, so this would install nothing at all" -}}
{{- end -}}
{{- end -}}

{{/*
The labels every EduIDE alert carries.

`namespace` is a routing label and is NOT the environment the alert is about.
The Prometheus Operator defaults `alertmanagerConfigMatcherStrategy` to
`OnNamespace`, which prepends `namespace = <the AlertmanagerConfig's own
namespace>` to every route it generates. Our AlertmanagerConfig lives in one
namespace, so alerts have to claim that namespace or they reach no receiver at
all.

The environment an alert is actually about is `eduide_namespace`, templated
from the series. Silences and inhibition rules must match on that one;
`namespace` is the same string on every EduIDE alert and matching it would
silence all of them at once.
*/}}
{{- define "eduide.alertRoutingLabels" -}}
namespace: {{ .Values.monitoring.alerting.namespace }}
eduide_namespace: '{{ "{{" }} $labels.namespace {{ "}}" }}'
{{- end -}}

{{/*
The receiver bodies for a set of channels.

Shared between the catch-all receiver and each environment-scoped one, so a
change to the message format cannot apply to one and not the other. Lives here
rather than in the template because a `define` may not sit inside an `if`, and
the whole AlertmanagerConfig is guarded by one.
*/}}
{{- define "eduide.receiverConfigs" }}
{{- $ctx := .ctx }}
{{- $a := $ctx.Values.monitoring.alerting }}
{{- $slack := list }}
{{- $discord := list }}
{{- range .channels }}
{{- if eq (.type | toString) "slack" }}{{ $slack = append $slack . }}{{ else }}{{ $discord = append $discord . }}{{ end }}
{{- end }}
{{- if $slack }}
slackConfigs:
{{- range $slack }}
- apiURL:
name: {{ $a.webhookSecret.name }}
key: {{ .secretKey }}
{{- if .channel }}
channel: {{ .channel | quote }}
{{- end }}
sendResolved: true
username: EduIDE
title: {{ $.title | quote }}
text: {{ $.text | quote }}
{{- end }}
{{- end }}
{{- if $discord }}
discordConfigs:
{{- range $discord }}
- apiURL:
name: {{ $a.webhookSecret.name }}
key: {{ .secretKey }}
sendResolved: true
title: {{ $.title | quote }}
message: {{ $.text | quote }}
{{- end }}
{{- end }}
{{- end }}
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
{{- if and .Values.monitoring.enabled .Values.monitoring.alerting.enabled .Values.monitoring.alerting.webhookSecret.create }}
{{- /*
The webhook URLs the channels post to.

A Slack or Discord webhook URL is a credential: anyone holding it can post into
the channel. It never goes in a values file in git and never on a `--set`, which
would put it in the process list and in Actions debug logs. Bootstrap reads it
from a GitHub Environment secret and passes it here already base64-encoded, the
same way the wildcard TLS certificate is handled.

Chart-managed rather than applied alongside, so it is removed when the release
is, and so a channel referring to a key that was never supplied fails the deploy
instead of producing an AlertmanagerConfig that silently notifies nobody.
*/}}
{{- $a := .Values.monitoring.alerting }}
{{- range $i, $c := $a.channels }}
{{- if not (hasKey $a.webhookSecret.data $c.secretKey) }}
{{- fail (printf "channel %q reads key %q from secret %q, but that key was not supplied in monitoring.alerting.webhookSecret.data" $c.name $c.secretKey $a.webhookSecret.name) }}
{{- end }}
{{- end }}
apiVersion: v1
kind: Secret
type: Opaque
metadata:
name: {{ $a.webhookSecret.name }}
namespace: {{ $a.namespace }}
labels:
app.kubernetes.io/part-of: eduide
data:
{{- range $k, $v := $a.webhookSecret.data }}
{{ $k }}: {{ $v | quote }}
{{- end }}
{{- end }}
118 changes: 118 additions & 0 deletions charts/eduide-cluster/templates/monitoring/alertmanagerconfig.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,118 @@
{{- include "eduide.checkAlertingPrerequisites" . }}
{{- if and .Values.monitoring.enabled .Values.monitoring.alerting.enabled }}
{{- $a := .Values.monitoring.alerting }}
{{- /*
Where EduIDE's alerts go.

Three things about this file are not obvious.

**The route matches on `namespace`, and that label is a routing artifact.** The
Prometheus Operator defaults `alertmanagerConfigMatcherStrategy` to
`OnNamespace`, which silently prepends `namespace = {{ $a.namespace }}` to the
route generated from this resource. Every EduIDE alert therefore sets that
label, and carries the environment it is really about in `eduide_namespace`.
Grouping and any inhibition rule must use `eduide_namespace`; grouping on
`namespace` would merge an outage in production with one in staging into a
single notification, because the label is identical across all of them.

**A channel may be scoped to environments.** Several installations can share one
cluster and belong to different people: Bonn and Mannheim both live on the
`eduide` cluster and each has its own Discord. A channel naming `environments`
gets a sub-route matching those namespaces, and first match wins.

**Anything not claimed by a scoped channel goes to every channel.** The
top-level receiver holds all of them deliberately. Cluster-scoped alerts - the
conversion webhook being down, a certificate about to expire - are not about any
one tenant's namespace, and dropping them because they failed to match a tenant
route would silently lose exactly the alerts that matter most.

**Only warning and above is notified.** Everything fires and is visible in
Alertmanager and on the dashboards, but the channels get a filtered subset. A
channel that receives every capacity blip stops being read, and then the
platform is unmonitored no matter how many rules exist.
*/}}
{{- if not $a.channels }}
{{- fail "monitoring.alerting.enabled is true but monitoring.alerting.channels is empty - alerts would fire and notify nobody" }}
{{- end }}
{{- $severities := list "critical" }}
{{- if eq ($a.minSeverity | toString) "warning" }}
{{- $severities = list "critical" "warning" }}
{{- else if ne ($a.minSeverity | toString) "critical" }}
{{- fail (printf "monitoring.alerting.minSeverity must be \"warning\" or \"critical\", got %q" ($a.minSeverity | toString)) }}
{{- end }}
{{- $slack := list }}
{{- $discord := list }}
{{- $scoped := list }}
{{- $names := dict }}
{{- range $i, $c := $a.channels }}
{{- if not $c.name }}{{ fail (printf "monitoring.alerting.channels[%d] has no name" $i) }}{{ end }}
{{- if hasKey $names $c.name }}{{ fail (printf "two channels are both called %q; receiver names must be unique" $c.name) }}{{ end }}
{{- $_ := set $names $c.name true }}
{{- if not $c.secretKey }}{{ fail (printf "channel %q names no secretKey - webhook URLs are only ever read from a Secret, never from values" $c.name) }}{{ end }}
{{- $type := $c.type | default "" | toString }}
{{- if eq $type "slack" }}{{ $slack = append $slack $c }}
{{- else if eq $type "discord" }}{{ $discord = append $discord $c }}
{{- else }}{{ fail (printf "channel %q has type %q; supported types are \"slack\" and \"discord\"" $c.name $type) }}{{ end }}
{{- if $c.environments }}{{ $scoped = append $scoped $c }}{{ end }}
{{- end }}
{{- /*
The message body. Alertmanager templating, not Helm's - the doubled braces are
escapes. Written to stand on its own in a chat client: what broke, what it means
for students, and where to look, without needing the dashboard open. `.Alerts`
is grouped by alertname and environment, so one message covers one problem in
one environment however many pods it touched.
*/}}
{{- $text := `{{ range .Alerts }}*{{ .Annotations.summary }}*
{{ .Annotations.description }}
{{ if .Labels.eduide_namespace }}Environment: {{ .Labels.eduide_namespace }}{{ end }}{{ if .Labels.pod }} | Pod: {{ .Labels.pod }}{{ end }}{{ if .Labels.node }} | Node: {{ .Labels.node }}{{ end }}
Severity: {{ .Labels.severity }} | Since: {{ .StartsAt.Format "15:04 MST" }}
{{ if .Annotations.runbook_url }}Runbook: {{ .Annotations.runbook_url }}
{{ end }}{{ if .Annotations.dashboard_url }}Dashboard: {{ .Annotations.dashboard_url }}
{{ end }}{{ end }}` }}
{{- $title := `{{ if eq .Status "firing" }}[FIRING {{ .Alerts.Firing | len }}]{{ else }}[RESOLVED]{{ end }} {{ .GroupLabels.alertname }}{{ if .GroupLabels.eduide_namespace }} - {{ .GroupLabels.eduide_namespace }}{{ end }}` }}
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata:
name: eduide-alerts
namespace: {{ $a.namespace }}
labels:
app.kubernetes.io/part-of: eduide
spec:
route:
receiver: eduide-all
groupBy:
- alertname
- eduide_namespace
groupWait: {{ $a.groupWait }}
groupInterval: {{ $a.groupInterval }}
repeatInterval: {{ $a.repeatInterval }}
matchers:
- name: severity
matchType: =~
value: {{ join "|" $severities | quote }}
{{- if $scoped }}
routes:
{{- range $scoped }}
- receiver: {{ printf "eduide-%s" .name }}
groupBy:
- alertname
- eduide_namespace
groupWait: {{ $a.groupWait }}
groupInterval: {{ $a.groupInterval }}
repeatInterval: {{ $a.repeatInterval }}
matchers:
- name: eduide_namespace
matchType: =~
value: {{ printf "^(%s)$" (join "|" .environments) | quote }}
{{- end }}
{{- end }}
receivers:
{{- /* Every channel. Reached by anything no scoped route claimed, which is
how cluster-scoped alerts still get out. */}}
- name: eduide-all
{{- include "eduide.receiverConfigs" (dict "ctx" $ "channels" $a.channels "title" $title "text" $text) }}
{{- range $scoped }}
- name: {{ printf "eduide-%s" .name }}
{{- include "eduide.receiverConfigs" (dict "ctx" $ "channels" (list .) "title" $title "text" $text) }}
{{- end }}
{{- end }}
Loading
Loading