The 'instance' and 'pod' labels represent the KSM instance and need to be filtered out to avoid mismatch errors ("many-to-many matching not allowed: matching labels must be unique on one side") , in the situation where KSM restarts (or runs >1 replica).
This works for the one side of the CPU query:
max without (instance,pod) (max_over_time(kube_pod_container_resource_requests_cpu_cores{node != ""}[48h]))
But no matter what I tried I could not get a vector result from the other side; this is where the mismatch error is coming from:
max_over_time(kube_pod_completion_time[48h]) - on (exported_pod) max_over_time(kube_pod_start_time[48h])
I don't understand that because 'on (exported_pod)' should mean only the exported_pod label is used for matching.
Particularly odd: both of these queries work, which each ignore just one label:
max_over_time(kube_pod_completion_time[48h]) - ignoring(pod) max_over_time(kube_pod_start_time[48h])
max_over_time(kube_pod_completion_time[48h]) - ignoring(instance) max_over_time(kube_pod_start_time[48h])
But ignoring both labels causes a many-to-many matching error, just like using on(exported_pod):
max_over_time(kube_pod_completion_time[48h]) - ignoring (pod, instance) max_over_time(kube_pod_start_time[48h])
And this works but it has problematic duplicate entries:
max_over_time(kube_pod_completion_time[48h]) - max_over_time(kube_pod_start_time[48h])
so that is how it is working currently, and it relies on the rearrange function ignoring duplicates.
It might be preferable to avoid duplicates at the PromQL level instead of in the python code - however the prometheus queries are subject to complex vagaries and occasional syntax changes so maybe deduplicating in python is safer.
The 'instance' and 'pod' labels represent the KSM instance and need to be filtered out to avoid mismatch errors ("many-to-many matching not allowed: matching labels must be unique on one side") , in the situation where KSM restarts (or runs >1 replica).
This works for the one side of the CPU query:
But no matter what I tried I could not get a vector result from the other side; this is where the mismatch error is coming from:
I don't understand that because 'on (exported_pod)' should mean only the exported_pod label is used for matching.
Particularly odd: both of these queries work, which each ignore just one label:
But ignoring both labels causes a many-to-many matching error, just like using on(exported_pod):
And this works but it has problematic duplicate entries:
so that is how it is working currently, and it relies on the rearrange function ignoring duplicates.
It might be preferable to avoid duplicates at the PromQL level instead of in the python code - however the prometheus queries are subject to complex vagaries and occasional syntax changes so maybe deduplicating in python is safer.