From b6c01badf8b08e9720e7efc28cb13f13de3786e6 Mon Sep 17 00:00:00 2001 From: Matthias Linhuber Date: Fri, 28 Aug 2026 20:02:31 +0200 Subject: [PATCH] fix: stop EduIDEPVCPending firing for every idle workspace Found by the alerting itself, within minutes of the first cluster switching it on: EduIDEPVCPending was firing for Bonn and Mannheim, and neither had a problem. Most storage classes bind WaitForFirstConsumer - parma's local-path does - so a workspace volume with no running session sits Pending indefinitely and that is correct rather than a fault. It binds when a pod first mounts it. Both claims showed `Used By: `. The alert now requires some pod to actually reference the claim, which is what makes a stuck Pending a real problem. Checked against the live cluster: the old expression returns 2 series there, the new one returns none. This is the failure mode the rest of these rules were written to avoid - an alert nobody can act on trains people to ignore the channel - and it got in anyway, on the one rule where a healthy zero at review time looked like proof it was quiet. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_019qeiQRFu8xAMRYWPdZewjG --- charts/eduide-cluster/Chart.yaml | 2 +- charts/eduide-cluster/README.md | 2 +- .../templates/monitoring/prometheusrule.yaml | 21 ++++++++++++++++--- 3 files changed, 20 insertions(+), 5 deletions(-) diff --git a/charts/eduide-cluster/Chart.yaml b/charts/eduide-cluster/Chart.yaml index 91fa447..3eb5336 100644 --- a/charts/eduide-cluster/Chart.yaml +++ b/charts/eduide-cluster/Chart.yaml @@ -4,5 +4,5 @@ description: | Cluster-scoped half of EduIDE: CRDs, the conversion webhook, ClusterRoles and cert-manager issuers. Install once per cluster, before any eduide release. type: application -version: 2.2.0 +version: 2.2.1 appVersion: "1.2.0" diff --git a/charts/eduide-cluster/README.md b/charts/eduide-cluster/README.md index 72da256..6e4b6fd 100644 --- a/charts/eduide-cluster/README.md +++ b/charts/eduide-cluster/README.md @@ -1,6 +1,6 @@ # eduide-cluster -![Version: 2.2.0](https://img.shields.io/badge/Version-2.2.0-informational?style=flat-square) ![Type: application](https://img.shields.io/badge/Type-application-informational?style=flat-square) ![AppVersion: 1.2.0](https://img.shields.io/badge/AppVersion-1.2.0-informational?style=flat-square) +![Version: 2.2.1](https://img.shields.io/badge/Version-2.2.1-informational?style=flat-square) ![Type: application](https://img.shields.io/badge/Type-application-informational?style=flat-square) ![AppVersion: 1.2.0](https://img.shields.io/badge/AppVersion-1.2.0-informational?style=flat-square) Cluster-scoped half of EduIDE: CRDs, the conversion webhook, ClusterRoles and cert-manager issuers. Install once per cluster, before any eduide release. diff --git a/charts/eduide-cluster/templates/monitoring/prometheusrule.yaml b/charts/eduide-cluster/templates/monitoring/prometheusrule.yaml index 0e2ed11..8f4b736 100644 --- a/charts/eduide-cluster/templates/monitoring/prometheusrule.yaml +++ b/charts/eduide-cluster/templates/monitoring/prometheusrule.yaml @@ -360,9 +360,21 @@ spec: device". Checking free space will not show the problem. runbook_url: {{ $runbook }} + {{- /* + Narrowed to claims a pod is actually waiting on. + + Most storage classes bind `WaitForFirstConsumer`, so a workspace volume + with no running session sits Pending indefinitely and that is correct, + not a fault - it binds when a pod first mounts it. Alerting on Pending + alone fired permanently for every idle workspace on the first cluster + this was deployed to. The join requires some pod to reference the claim, + which is what makes a stuck Pending a real problem. + */}} - alert: EduIDEPVCPending expr: |- - kube_persistentvolumeclaim_status_phase{namespace=~"{{ $ns }}", phase="Pending"} == 1 + (kube_persistentvolumeclaim_status_phase{namespace=~"{{ $ns }}", phase="Pending"} == 1) + and on (namespace, persistentvolumeclaim) + kube_pod_spec_volumes_persistentvolumeclaims_info{namespace=~"{{ $ns }}"} for: 10m labels: severity: warning @@ -371,8 +383,11 @@ spec: summary: 'A workspace volume will not provision in {{ "{{" }} $labels.namespace {{ "}}" }}' description: >- PersistentVolumeClaim {{ "{{" }} $labels.persistentvolumeclaim {{ "}}" }} in - {{ "{{" }} $labels.namespace {{ "}}" }} has been Pending for 10 minutes, so the session - that needs it cannot start. + {{ "{{" }} $labels.namespace {{ "}}" }} has been Pending for 10 minutes while a pod is + waiting on it, so the session that needs it cannot start. + A workspace volume with no pod is a different thing and does not alert: + most storage classes bind on first consumer, so an idle workspace is + Pending by design. The storage class is the first thing to check - it is a cluster property set once in clusters/.yaml, and a claim naming a class the cluster does not offer stays Pending forever without any other symptom.