From 40eb369d48f0d9a5491e6f33336aa2efbdcb45b0 Mon Sep 17 00:00:00 2001 From: Arve Knudsen Date: Tue, 2 Jun 2026 11:25:56 +0200 Subject: [PATCH 1/3] proposal: /api/v1/info_labels endpoint for info() autocomplete Signed-off-by: Arve Knudsen --- proposals/0085-info-labels-endpoint.md | 278 +++++++++++++++++++++++++ 1 file changed, 278 insertions(+) create mode 100644 proposals/0085-info-labels-endpoint.md diff --git a/proposals/0085-info-labels-endpoint.md b/proposals/0085-info-labels-endpoint.md new file mode 100644 index 00000000..3dcc92df --- /dev/null +++ b/proposals/0085-info-labels-endpoint.md @@ -0,0 +1,278 @@ +## Info-metric label discovery API for `info()` autocomplete + +* **Owners:** + * Arve Knudsen [@aknuds1](https://github.com/aknuds1) arve.knudsen@gmail.com + * Ismail Simsek [@itsmylife](https://github.com/itsmylife) ismail.simsek@grafana.com + +* **Implementation Status:** Not implemented (working branch: https://github.com/aknuds1/prometheus/tree/arve/info-autocomplete) + +* **Related Issues and PRs:** + * [PROM-74 — V2 API for labels and values discovery](https://github.com/prometheus/proposals/pull/74). This proposal builds on PROM-74's NDJSON contract, feature gate, and shared parameter set. + * [PROM-37 — Simplify joins with info metrics in PromQL](./0037-native-support-for-info-metrics-metadata.md). Introduces the `info()` PromQL function; the proposed endpoint supports its autocomplete. + +* **Other docs or links:** + * PromQL `info()` function documentation in `docs/querying/functions.md`. + +> TL;DR: Add `GET|POST /api/v1/info_labels` as the info-metric companion to the experimental `/api/v1/search/*` family. The endpoint streams the non-identifying data labels of info metrics (e.g. `target_info`) as NDJSON. Queries can be scoped either by an explicit info-metric specifier (`metric_match`), or by evaluating a PromQL expression (`expr`) and harvesting its identifying labels. The endpoint reuses PROM-74's NDJSON contract, fuzzy/sort/score/limit/batch parameters, and `--enable-feature=search-api` gate. It additionally requires `--enable-feature=promql-experimental-functions`, since its only purpose is to support autocomplete for the experimental `info()` function. + +## Why + +The PromQL `info()` function (added under `--enable-feature=promql-experimental-functions`) enriches base metrics with the data labels carried on info metrics like `target_info`. A PromQL editor offering autocomplete inside `info(, {…})` needs to know **which data labels are available for the in-scope info metrics**, where the in-scope set is determined by the user's in-progress expression. + +This is a new gap created by the introduction of `info()` itself — it is not specific to any one editor or vendor. Anyone building a PromQL editor that supports `info()` will hit the same problem. + +The in-tree client implementation lives in `web/ui/module/lezer-promql/src/client/`; any PromQL editor that links it (Prometheus' own React UI via `codemirror-promql`, Grafana, others) inherits `/info_labels` autocomplete. + +### Pitfalls of the current solution + +Constructing the same response with the endpoints that exist today requires either fetching far more than needed or multiple round-trips per autocomplete suggestion. + +* **`/api/v1/labels` + per-name `/api/v1/label/{name}/values`** returns *all* labels matching a selector, including labels carried by the base metric, not just data labels on info metrics. Filtering client-side requires downloading the full universe of labels and then making one `/label/{name}/values` call per label of interest (N+1). +* **`/api/v1/series` + client aggregation** returns one row per series, then the client must dedupe labels and prune identifying labels. The wire size is O(series), and there is no server-side scoring for relevance ranking. +* **`/api/v1/search/label_names` + `/api/v1/search/label_values`** (from PROM-74) is closer but still costs **1 + N round-trips**, where N is the number of returned label names: one call to discover the names, then one per name to fetch its values. Neither endpoint understands the info-metric scoping pattern (find label NAMES on info metrics whose identifying labels intersect with a PromQL expression's result set). +* **`/api/v1/search/*` with `match[]` alone** cannot scope autocomplete to be context-aware against arbitrary PromQL like `rate(http_requests_total[5m])`, which is what users actually type. The endpoint needs to evaluate an expression and harvest its identifying labels, not merely accept a series selector. + +For interactive autocomplete this manifests as visible latency on every keystroke and unnecessary load on the server. + +## Goals + +* One round-trip per autocomplete suggestion, regardless of how many data labels match. +* Server-side scoping by either an explicit info-metric specifier (`metric_match`) or by evaluating a PromQL expression (`expr`) and using its identifying labels. +* Reuse PROM-74's NDJSON contract, search/fuzzy/sort/score/limit/batch parameters, and `--enable-feature=search-api` gate so operators see one coherent experimental autocomplete API surface. +* Couple the endpoint to the function it serves: also require `--enable-feature=promql-experimental-functions`, so an operator cannot reach a state where the endpoint is up but `info()` queries fail to parse. +* Bound response size and provide predictable degradation under pathological input (e.g. `target_info` with very high `version`/`env`/`region` cardinality, or a `metric_match` that matches far more than expected). + +## Non-Goals + +* Replace or modify `/api/v1/labels` or `/api/v1/label/{name}/values`. +* Add or modify the `/api/v1/search/*` family. This proposal *builds on* PROM-74; it does not extend it. +* Server-side evaluation of `info()`. This is a metadata endpoint; `info()` queries continue to use `/api/v1/query`. +* Cursor-based pagination. Same stance as PROM-74; `has_more` is informational only. +* Special handling of `__type__` / `__unit__` labels. They are not in `infohelper.DefaultIdentifyingLabels` (which is `{"instance", "job"}`), so they pass through the extractor as ordinary data labels and may appear in the response if set on info metrics. + +## How + +A new endpoint, gated behind two feature flags, that streams NDJSON using the same wire framing (batches + trailer) as `/api/v1/search/*`. The per-record shape is `/info_labels`-specific (`{name, values, score?}`). + +### Implementation notes + +* The parameter set extends PROM-74's shared parameters (`search[]`, `fuzz_*`, `sort_*`, `include_score`, `start`, `end`, `limit`, `batch_size`) with three /info_labels-specific parameters: `expr`, `metric_match`, and `values_limit`. +* `match[]` is not accepted. Scoping happens via `metric_match` (which info metrics to consider) and `expr` (which identifying-label values to restrict to). Accepting `match[]` as well would create two equivalent matcher mechanisms with unclear precedence. +* The response unit is one record per data label NAME, carrying that label's values inline. Search-api emits names and values via separate endpoints; collapsing them here matches the autocomplete UI pattern: each label name is a top-level suggestion, and selecting one reveals its values. +* The response is NDJSON (`application/x-ndjson`) with the same batch + trailer contract as PROM-74. +* The endpoint is dual-gated: `--enable-feature=search-api` covers the NDJSON + parsing infrastructure it reuses; `--enable-feature=promql-experimental-functions` covers the only consumer (`info()`). Either missing flag returns the standard Prometheus JSON error with `errorType: unavailable` and a flag-specific message. + +### `GET|POST /api/v1/info_labels` + +#### Request + +**Method:** `GET` or `POST` + +`POST` is recommended when the `expr` parameter exceeds URL length limits. + +**Path:** `/api/v1/info_labels` + +**Query parameters:** + +| Name | Type | Required | Default | Description | +|------------------|--------------------------|----------|-----------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `expr` | string (PromQL) | No | | Optional PromQL expression evaluated server-side. Identifying labels (`job`, `instance`) extracted from the result restrict the info-metric query so autocomplete sees only the labels relevant to the user's expression. | +| `metric_match` | string | No | `target_info` (exact) | Info-metric specifier. Bare value (no prefix) is exact match; `=~` / `!=` / `!~` prefixes give the corresponding matcher type (e.g. `target_info`, `=~.*_info`, `!=target_info`, `!~build_info`). | +| `search[]` | []string | No | | One or more search terms matched against data label NAMES (OR semantics). As per PROM-74. | +| `fuzz_threshold` | int [0..100] | No | 0 | Fuzzy threshold. As per PROM-74. | +| `fuzz_alg` | string | No | `subsequence` | Fuzzy algorithm. As per PROM-74. | +| `case_sensitive` | bool | No | true | As per PROM-74. | +| `sort_by` | `alpha` / `score` | No | | As per PROM-74. | +| `sort_dir` | `asc` / `dsc` | No | `asc` | As per PROM-74. | +| `include_score` | bool | No | false | As per PROM-74. | +| `start`, `end` | rfc3339 / unix_timestamp | No | last 1h | As per PROM-74. | +| `limit` | int ≥ 0 | No | 100, capped by `--web.search.max-limit` | Maximum number of label NAMES returned. Same `has_more` semantics as PROM-74. | +| `values_limit` | int ≥ 0 | No | 0 (no cap) | Maximum number of values returned per label. New to this endpoint. | +| `batch_size` | int > 0 | No | 100 | As per PROM-74. | + +**Notes:** + +***Feature gating*** + +The endpoint requires *both* `--enable-feature=search-api` (NDJSON + parsing infrastructure) and `--enable-feature=promql-experimental-functions` (the only consumer is `info()` autocomplete; without `info()`, no client can use the response). The handler checks both flags upfront and, if either is missing, responds with the standard non-streaming JSON error (`errorType: unavailable`) carrying a single message that names every missing flag, so an operator sees the whole gap in one round-trip: + +* `"info_labels requires --enable-feature=search-api"` — only search-api missing. +* `"info_labels requires --enable-feature=promql-experimental-functions"` — only experimental-functions missing. +* `"info_labels requires --enable-feature=search-api,promql-experimental-functions"` — both missing. + +***expr*** + +If `expr` is provided, the server evaluates it as an instant query at the request's `end` timestamp. The result must be a Vector or Matrix; Scalar/String returns `errorBadData`. The server harvests the identifying labels (`job`, `instance`) from each sample/series and uses their values as additional matchers against the info-metric query. If the eval produces no identifying-label values, the stream consists of a single empty batch (carrying any warnings) followed by the success trailer. + +This is what makes the endpoint context-aware: `expr=rate(http_requests_total[5m])` gives an autocomplete experience scoped to the info metrics actually relevant to that PromQL expression, without requiring the client to interpret PromQL itself. + +When `metric_match` is also set, the two filters apply jointly: an info-metric series must match `metric_match`'s name pattern AND have identifying labels intersecting the `expr` eval result. This lets callers narrow both the info-metric family (e.g. `metric_match==~.*_info`) and the relevant series (e.g. `expr=rate(http_requests_total[5m])`). + +***metric_match*** + +Specifies which info metrics to consider. Prefix syntax mirrors PromQL matcher operators: + +* `metric_match=target_info` → `__name__="target_info"` (default) +* `metric_match==~.*_info` → `__name__=~".*_info"` +* `metric_match=!=target_info` → `__name__!="target_info"` +* `metric_match=!~build_info` → `__name__!~"build_info"` + +The double-equals in `metric_match==~.*_info` is not a typo — the first `=` is the URL form-value separator and the second `=` is the matcher-type prefix. + +***match[]*** + +Not accepted. Requests carrying `match[]` are rejected with `errorBadData`: `"match[] is not supported by info_labels; use metric_match or expr"`. + +***values_limit*** + +Caps the number of values per label. Distinct from `limit`, which caps the number of records (per PROM-74's semantics). Values are sorted alphabetically and truncated after the sort, so the truncation is deterministic. + +***start/end*** + +The default window is the last hour (inherited from PROM-74's `parseSearchParams`). This is conservative — info metrics like `target_info` typically only emit one sample per scrape and may be scraped infrequently, so a deployment with a long scrape interval and an idle window can land outside a 1-hour lookback and return no data. Operators with infrequent scrapes should pass an explicit `start` to widen the window; clients that know the user's session age can do the same. See Known unknowns for the open question of whether `/info_labels` should ship with a longer default than `/search/*`. + +***Storage*** + +The endpoint uses the existing `storage.Querier` interface, not PROM-74's `Searcher`. The unit of work — enumerate info-metric series, collect their non-identifying labels deduped across series, with values — does not map onto `SearchLabelNames` (which returns names without values and would force one `SearchLabelValues` call per name). The search-api `storage.Filter` produced by `buildSearchFilter` is applied to label names in-memory inside the extractor loop, so fuzzy/score behaviour is identical to `/api/v1/search/label_names` for the same input names (iteration order differs because `/info_labels` is series-driven rather than index-driven, but the difference is invisible once results pass through the post-sort). + +***Security*** + +* `expr` is evaluated through the standard `QueryEngine`, so the existing per-query timeout, max-samples limit, and any deployed query authorisation policy apply unchanged. The endpoint adds no new evaluation path. +* The `namesLimit` cap described in the *Memory bound* note bounds extractor memory against pathological `metric_match` values, so a malicious or careless caller cannot exhaust server memory by matching the entire metric universe. +* The endpoint exposes no metadata that a caller cannot already obtain by combining `/api/v1/labels`, `/api/v1/label/{name}/values`, and client-side post-filtering. It packages the same information more efficiently; it does not widen the read surface. +* Expensive `expr` is bounded by the existing QueryEngine timeout and max-samples (`--query.timeout`, `--query.max-samples`); `/info_labels` adds no further protection beyond what `/api/v1/query` already enforces. + +***Memory bound*** + +The extractor drains all matching info-metric series before streaming the first batch. To prevent unbounded growth on pathological `metric_match` values (e.g. one that matches the entire metric universe), the extractor caps the number of distinct label NAMES it collects at `--web.search.max-limit` (default 10000; setting the flag to 0 disables the cap and makes the extractor unbounded). Once the cap is hit, the extractor keeps draining the series set to gather values for already-collected names but rejects any not-yet-seen names, and a warning is surfaced in the first batch. + +`sort_by=score` under a hit cap may miss the true top-N because some high-scoring names were never seen — documented as a known limitation; operators who want full coverage can raise the cap. Streaming gives clients incremental rendering after the storage scan completes; it does not eliminate the scan itself. + +***Performance*** + +* Time: O(series matching `metric_match`) × O(labels per series). Per-series work is one label-set iteration plus a memoised filter-accept lookup per distinct label name, so the filter cost is amortised to O(1) per name across all series. +* Memory: bounded by `namesLimit` (distinct names) × `values_limit` (values per name) × average value length. With defaults (`--web.search.max-limit=10000`, no `values_limit`), the in-practice ceiling is set by the cardinality of the real info-metric data; the cap exists to bound the worst case, not the typical one. When `values_limit` is unset, the per-name values count is bounded only by the data's natural cardinality, which is typically small (tens) for info metrics. +* No additional storage round-trips beyond the one info-metric `Select`. The optional `expr` adds one instant-query evaluation through the standard QueryEngine path. + +#### Response + +* `Content-Type: application/x-ndjson; charset=utf-8` + +Zero or more `{results, warnings?}` batch lines, followed by a `{status, has_more, warnings?}` trailer, or — on mid-stream failure after the first batch has been written — a `{status, errorType, error}` line in place of the trailer. Errors that occur before the first batch is written are returned as the usual non-streaming Prometheus JSON error with the matching 4xx/5xx status code. + +This is the same contract as the `/api/v1/search/*` endpoints. + +URLs in the examples below are shown unescaped for readability — single-quote them on the shell or let curl perform the percent-encoding. + +##### Example: no scoping + +```ndjson +{"results":[{"name":"cluster","values":["us-east","us-west"]},{"name":"env","values":["prod","staging"]},{"name":"region","values":["us-east"]},{"name":"version","values":["v1.0","v2.0","v2.1"]}]} +{"status":"success","has_more":false} +``` + +##### Example: scoped by expression + +```bash +curl -N 'http://localhost:9090/api/v1/info_labels?expr=rate(http_requests_total{job="api-gateway"}[5m])' +``` + +```ndjson +{"results":[{"name":"cluster","values":["us-east","us-west"]},{"name":"env","values":["prod","staging"]},{"name":"version","values":["v1.0","v2.0"]}]} +{"status":"success","has_more":false} +``` + +##### Example: scoped by search with relevance scoring + +```bash +curl -N 'http://localhost:9090/api/v1/info_labels?search[]=ver&sort_by=score&include_score=true' +``` + +```ndjson +{"results":[{"name":"version","values":["v1.0","v2.0","v2.1"],"score":1},{"name":"server","values":["nginx","envoy"],"score":0.83}]} +{"status":"success","has_more":false} +``` + +##### Example: `match[]` rejected + +```bash +curl -N 'http://localhost:9090/api/v1/info_labels?match[]={job="prometheus"}' +``` + +```json +{"status":"error","errorType":"bad_data","error":"match[] is not supported by info_labels; use metric_match or expr"} +``` + +##### Example: names truncated at namesLimit + +When the extractor hits the `--web.search.max-limit` cap, the first batch carries a warning naming the cap and what to do about it. The trailer's `status` stays `success` because the result set is still well-formed — just incomplete. + +```ndjson +{"results":[{"name":"cluster","values":["us-east","us-west"]},...],"warnings":["info-labels names truncated at 10000; narrow metric_match or raise --web.search.max-limit"]} +{"status":"success","has_more":false} +``` + +### Interface changes + +Smaller than PROM-74's contribution. This proposal does not introduce new storage interfaces: + +* PROM-74's `Searcher.SearchLabelNames` is not used; the endpoint stays on `storage.Querier`. +* The search-api request preamble (CORS, feature-gate, form parsing, `parseSearchParams`, querier acquisition, `buildSearchFilter`, sort/order plumbing) is factored out of `newSearchRequest` into a reusable `newAutocompleteRequest` so both `/api/v1/search/*` and `/api/v1/info_labels` share it. This is a small refactor of code that landed with PROM-74: `newSearchRequest` is preserved as a thin wrapper that adds the `storage.Searcher` type assertion the search-api endpoints require. +* A small helper package (`promql/infohelper`) exposes `ExtractDataLabels(ctx, querier, infoMetricMatcher, identifyingLabelValues, hints, filter, namesLimit, valuesLimit) → []InfoLabelRecord` where `InfoLabelRecord = {Name, Values, Score}`. The filter is applied to label names; the score carries through to the wire when `include_score=true`. + +### Extensibility for Mimir, Thanos, Cortex + +The per-record JSON shape inherits PROM-74's extensibility story: downstream implementations can add an `extensions` field per record (and an `extensions` map keyed by provider name) without changing the core contract. No new mechanism is required. + +```ndjson +{"results":[{"name":"cluster","values":["us-east","us-west"],"extensions":{"mimir":{"cardinality":42}}}]} +{"status":"success","has_more":false} +``` + +The pure Prometheus implementation does not include an `extensions` field. Downstream implementations adding it should tag it `json:",omitempty"` so records without provider-specific data stay clean on the wire. + +### Testing and verification + +* **Recommended manual verification (not automated today):** a client can enumerate `target_info` series via `/api/v1/series`, dedupe label names client-side, and compare against `/api/v1/info_labels` output for the same time window. +* The implementation's unit tests cover the parameter surface, the dual gate, `match[]` rejection, sort/score/fuzzy combinations, `values_limit` truncation, multi-batch streaming, and error paths (invalid expr, scalar result, queryable error). + +### Migration + +New endpoint. Both feature flags it depends on are experimental and opt-in. No migration is required. + +### Known unknowns + +* **Default start/end lookback.** Aligned with PROM-74's 1h default via the shared `parseSearchParams` helper. Reasonable, but `target_info` can be sparse on short windows — reviewers may want a longer default specific to `/info_labels`. +* **Whether to accept `match[]` in addition to `expr`.** The current design rejects it on cohesion grounds; reviewers may push back. +* **Whether identifying labels should be configurable per request.** Today the set is `{"job", "instance"}` from `infohelper.DefaultIdentifyingLabels`. Per-request override is plausible if downstream implementations have different identifier conventions. +* **One combined `info-labels-api` flag.** Rejected here on grounds of flag proliferation, but worth re-litigating if reviewers prefer a single-flag operator UX. + +## Alternatives + +### 1. Extend `/api/v1/search/label_names` to optionally return values per name + +Would collapse two endpoints into one. Rejected on cohesion grounds: PROM-74 keeps names and values in separate endpoints precisely so that each endpoint's response shape stays simple. The values payload is only useful when the names are scoped to info metrics, so the coupling does not belong on the general search endpoint. + +### 2. Add per-name expansion via a follow-up `/api/v1/search/label_values` call + +Works in principle but reintroduces the N+1 round-trip that PROM-74 is trying to fix. For autocomplete this re-creates the latency problem and increases backend load. + +### 3. Pure client-side composition (`/api/v1/labels` + `/api/v1/label/{name}/values` + filter) + +Feasible today but slow at scale and does not address the info-metric-specific scoping. In particular it cannot evaluate a PromQL expression server-side to harvest identifying labels, so the autocomplete cannot be context-aware against arbitrary PromQL. + +### 4. Subsume into a hypothetical `/api/v1/info()` evaluation endpoint + +Out of scope. `info()` is a PromQL function that already uses the existing `/api/v1/query` flow. `/api/v1/info_labels` is a metadata endpoint that supports building the queries; it is not a query endpoint itself. + +### 5. Use `match[]` instead of `expr` + +`match[]` can only express series selectors, not arbitrary PromQL like `rate(http_requests_total[5m])`. Autocomplete must work against expressions users actually type, not the subset that fits inside a series selector. + +## Action Plan + +* [X] Build a working implementation (working branch `arve/info-autocomplete` in `aknuds1/prometheus`). +* [X] Add the dual feature-gate check to the implementation (search-api + promql-experimental-functions). +* [ ] Finalize this proposal based on community feedback. +* [ ] Open the upstream implementation PR against `prometheus/prometheus`. +* [ ] Update `docs/querying/api.md` (the implementation already includes this; the action item is post-merge alignment). From c0f386407f7d630c19f5a766d8375e0c1b151baf Mon Sep 17 00:00:00 2001 From: Arve Knudsen Date: Thu, 16 Jul 2026 14:41:04 +0200 Subject: [PATCH 2/3] proposal: incorporate Grafana client feedback Document the request profiles and client requirements learned from the Grafana info() autocomplete proof of concept. Clarify expression interpolation, response bounds, and terminal NDJSON handling. Signed-off-by: Arve Knudsen --- proposals/0085-info-labels-endpoint.md | 56 +++++++++++++++++++------- 1 file changed, 42 insertions(+), 14 deletions(-) diff --git a/proposals/0085-info-labels-endpoint.md b/proposals/0085-info-labels-endpoint.md index 3dcc92df..0310b189 100644 --- a/proposals/0085-info-labels-endpoint.md +++ b/proposals/0085-info-labels-endpoint.md @@ -4,11 +4,12 @@ * Arve Knudsen [@aknuds1](https://github.com/aknuds1) arve.knudsen@gmail.com * Ismail Simsek [@itsmylife](https://github.com/itsmylife) ismail.simsek@grafana.com -* **Implementation Status:** Not implemented (working branch: https://github.com/aknuds1/prometheus/tree/arve/info-autocomplete) +* **Implementation Status:** Not implemented upstream ([WIP implementation PR](https://github.com/prometheus/prometheus/pull/17930)) * **Related Issues and PRs:** * [PROM-74 — V2 API for labels and values discovery](https://github.com/prometheus/proposals/pull/74). This proposal builds on PROM-74's NDJSON contract, feature gate, and shared parameter set. * [PROM-37 — Simplify joins with info metrics in PromQL](./0037-native-support-for-info-metrics-metadata.md). Introduces the `info()` PromQL function; the proposed endpoint supports its autocomplete. + * [Grafana Prometheus datasource PR #244](https://github.com/grafana/grafana-prometheus-datasource/pull/244). Independent client PoC for `info()` data-label autocomplete using this endpoint. * **Other docs or links:** * PromQL `info()` function documentation in `docs/querying/functions.md`. @@ -21,26 +22,26 @@ The PromQL `info()` function (added under `--enable-feature=promql-experimental- This is a new gap created by the introduction of `info()` itself — it is not specific to any one editor or vendor. Anyone building a PromQL editor that supports `info()` will hit the same problem. -The in-tree client implementation lives in `web/ui/module/lezer-promql/src/client/`; any PromQL editor that links it (Prometheus' own React UI via `codemirror-promql`, Grafana, others) inherits `/info_labels` autocomplete. +The Prometheus UI PoC integrates the endpoint through `web/ui/module/lezer-promql/src/client/`. Grafana has a separate Monaco-based client PoC, demonstrating that the endpoint contract is usable without sharing Prometheus' editor implementation. ### Pitfalls of the current solution -Constructing the same response with the endpoints that exist today requires either fetching far more than needed or multiple round-trips per autocomplete suggestion. +Constructing an equivalently scoped interaction with the endpoints that exist today requires either fetching far more than needed or combining endpoints that cannot express arbitrary-PromQL info-metric scoping consistently. * **`/api/v1/labels` + per-name `/api/v1/label/{name}/values`** returns *all* labels matching a selector, including labels carried by the base metric, not just data labels on info metrics. Filtering client-side requires downloading the full universe of labels and then making one `/label/{name}/values` call per label of interest (N+1). * **`/api/v1/series` + client aggregation** returns one row per series, then the client must dedupe labels and prune identifying labels. The wire size is O(series), and there is no server-side scoring for relevance ranking. -* **`/api/v1/search/label_names` + `/api/v1/search/label_values`** (from PROM-74) is closer but still costs **1 + N round-trips**, where N is the number of returned label names: one call to discover the names, then one per name to fetch its values. Neither endpoint understands the info-metric scoping pattern (find label NAMES on info metrics whose identifying labels intersect with a PromQL expression's result set). +* **`/api/v1/search/label_names` + `/api/v1/search/label_values`** (from PROM-74) is closer, but neither endpoint understands the info-metric scoping pattern (find label names on info metrics whose identifying labels intersect with a PromQL expression's result set). A client can defer the value lookup until the user selects a name, avoiding eager 1 + N calls, but it still has to reproduce the missing info-metric scoping across two endpoint families. * **`/api/v1/search/*` with `match[]` alone** cannot scope autocomplete to be context-aware against arbitrary PromQL like `rate(http_requests_total[5m])`, which is what users actually type. The endpoint needs to evaluate an expression and harvest its identifying labels, not merely accept a series selector. For interactive autocomplete this manifests as visible latency on every keystroke and unnecessary load on the server. ## Goals -* One round-trip per autocomplete suggestion, regardless of how many data labels match. +* One bounded round-trip for each autocomplete interaction: discover label names, then fetch values only for the label the user selects. * Server-side scoping by either an explicit info-metric specifier (`metric_match`) or by evaluating a PromQL expression (`expr`) and using its identifying labels. * Reuse PROM-74's NDJSON contract, search/fuzzy/sort/score/limit/batch parameters, and `--enable-feature=search-api` gate so operators see one coherent experimental autocomplete API surface. * Couple the endpoint to the function it serves: also require `--enable-feature=promql-experimental-functions`, so an operator cannot reach a state where the endpoint is up but `info()` queries fail to parse. -* Bound response size and provide predictable degradation under pathological input (e.g. `target_info` with very high `version`/`env`/`region` cardinality, or a `metric_match` that matches far more than expected). +* Give callers independent caps for label names and values, and provide predictable degradation under pathological input (e.g. `target_info` with very high `version`/`env`/`region` cardinality, or a `metric_match` that matches far more than expected). ## Non-Goals @@ -58,10 +59,21 @@ A new endpoint, gated behind two feature flags, that streams NDJSON using the sa * The parameter set extends PROM-74's shared parameters (`search[]`, `fuzz_*`, `sort_*`, `include_score`, `start`, `end`, `limit`, `batch_size`) with three /info_labels-specific parameters: `expr`, `metric_match`, and `values_limit`. * `match[]` is not accepted. Scoping happens via `metric_match` (which info metrics to consider) and `expr` (which identifying-label values to restrict to). Accepting `match[]` as well would create two equivalent matcher mechanisms with unclear precedence. -* The response unit is one record per data label NAME, carrying that label's values inline. Search-api emits names and values via separate endpoints; collapsing them here matches the autocomplete UI pattern: each label name is a top-level suggestion, and selecting one reveals its values. +* The response unit is one record per data label name, carrying that label's values inline. Clients should use different limits for the two autocomplete phases instead of eagerly retrieving all values for every candidate name. * The response is NDJSON (`application/x-ndjson`) with the same batch + trailer contract as PROM-74. * The endpoint is dual-gated: `--enable-feature=search-api` covers the NDJSON + parsing infrastructure it reuses; `--enable-feature=promql-experimental-functions` covers the only consumer (`info()`). Either missing flag returns the standard Prometheus JSON error with `errorType: unavailable` and a flag-specific message. +### Client integration feedback + +The independent Grafana PoC validates the request and response shape while exposing several requirements that apply to PromQL editors generally: + +* `expr` must be valid PromQL. Clients with editor-specific variables or macros must resolve them with the same semantics as normal query execution before sending the request. +* Each request must use the current query time range. A client cache should be bounded and freshness-limited, keyed by the effective resolved request, retain only complete successful responses, and allow retries after failures. +* A label-name completion request should send the current `start`/`end`, resolved `expr`, optional `metric_match`, and typed `search[]`. Omitting `limit` uses the server's at-most-100-name default, reduced further when the operator configures a lower maximum; `values_limit=1` prevents the name-completion response from carrying every value for every name. +* A label-value completion request can use the existing search contract with `search[]=`, `fuzz_alg=subsequence`, `fuzz_threshold=100`, `case_sensitive=true`, `sort_by=alpha`, `sort_dir=asc`, `limit=1`, and an explicit client-selected `values_limit`. Subsequence prefix matches can tie with the exact name, so the client must still verify that the returned record name equals the requested name before using its values. + +The last request profile is a bounded workaround, not an exact-name API primitive. A first-class name-only mode and exact-name filter remain open design questions. + ### `GET|POST /api/v1/info_labels` #### Request @@ -104,6 +116,8 @@ The endpoint requires *both* `--enable-feature=search-api` (NDJSON + parsing inf If `expr` is provided, the server evaluates it as an instant query at the request's `end` timestamp. The result must be a Vector or Matrix; Scalar/String returns `errorBadData`. The server harvests the identifying labels (`job`, `instance`) from each sample/series and uses their values as additional matchers against the info-metric query. If the eval produces no identifying-label values, the stream consists of a single empty batch (carrying any warnings) followed by the success trailer. +The endpoint accepts PromQL, not editor-specific template syntax. A client must resolve local variables and macros before sending `expr`; using the same interpolation path as normal query execution avoids autocomplete evaluating a different expression than the query itself. + This is what makes the endpoint context-aware: `expr=rate(http_requests_total[5m])` gives an autocomplete experience scoped to the info metrics actually relevant to that PromQL expression, without requiring the client to interpret PromQL itself. When `metric_match` is also set, the two filters apply jointly: an info-metric series must match `metric_match`'s name pattern AND have identifying labels intersecting the `expr` eval result. This lets callers narrow both the info-metric family (e.g. `metric_match==~.*_info`) and the relevant series (e.g. `expr=rate(http_requests_total[5m])`). @@ -119,6 +133,8 @@ Specifies which info metrics to consider. Prefix syntax mirrors PromQL matcher o The double-equals in `metric_match==~.*_info` is not a typo — the first `=` is the URL form-value separator and the second `=` is the matcher-type prefix. +Clients deriving `metric_match` from a PromQL `__name__` matcher must preserve the full operator. In particular, a regex matcher is encoded as the form value `=~.*_info`; dropping the leading `=` changes it into an exact metric name. + ***match[]*** Not accepted. Requests carrying `match[]` are rejected with `errorBadData`: `"match[] is not supported by info_labels; use metric_match or expr"`. @@ -127,9 +143,13 @@ Not accepted. Requests carrying `match[]` are rejected with `errorBadData`: `"ma Caps the number of values per label. Distinct from `limit`, which caps the number of records (per PROM-74's semantics). Values are sorted alphabetically and truncated after the sort, so the truncation is deterministic. +The two dimensions multiply: a response can contain up to `limit * values_limit` values when both are non-zero. Omitting `values_limit`, or setting it to zero, leaves value cardinality uncapped. Interactive clients should therefore always set it explicitly. + +To make `values_limit` a server-memory bound as well as a wire bound, implementations must enforce it while collecting values, retaining only the lexicographically smallest values needed for the deterministic result. Collecting every distinct value and truncating only after the scan bounds the response but not extractor memory. + ***start/end*** -The default window is the last hour (inherited from PROM-74's `parseSearchParams`). This is conservative — info metrics like `target_info` typically only emit one sample per scrape and may be scraped infrequently, so a deployment with a long scrape interval and an idle window can land outside a 1-hour lookback and return no data. Operators with infrequent scrapes should pass an explicit `start` to widen the window; clients that know the user's session age can do the same. See Known unknowns for the open question of whether `/info_labels` should ship with a longer default than `/search/*`. +The default window is the last hour (inherited from PROM-74's `parseSearchParams`). This is conservative — info metrics like `target_info` typically only emit one sample per scrape and may be scraped infrequently, so a deployment with a long scrape interval and an idle window can land outside a 1-hour lookback and return no data. Operators with infrequent scrapes should pass an explicit `start` to widen the window; interactive clients should pass the current editor time range rather than retaining the range from editor initialization. See Known unknowns for the open question of whether `/info_labels` should ship with a longer default than `/search/*`. ***Storage*** @@ -138,20 +158,22 @@ The endpoint uses the existing `storage.Querier` interface, not PROM-74's `Searc ***Security*** * `expr` is evaluated through the standard `QueryEngine`, so the existing per-query timeout, max-samples limit, and any deployed query authorisation policy apply unchanged. The endpoint adds no new evaluation path. -* The `namesLimit` cap described in the *Memory bound* note bounds extractor memory against pathological `metric_match` values, so a malicious or careless caller cannot exhaust server memory by matching the entire metric universe. +* The `namesLimit` cap described in the *Memory bound* note bounds the number of distinct label names, but not the number of values retained for each name. An implementation must apply the name cap to all retained per-name state and enforce a positive `values_limit` during extraction to bound both dimensions; requests with an omitted or zero `values_limit` remain unbounded in the value dimension. * The endpoint exposes no metadata that a caller cannot already obtain by combining `/api/v1/labels`, `/api/v1/label/{name}/values`, and client-side post-filtering. It packages the same information more efficiently; it does not widen the read surface. * Expensive `expr` is bounded by the existing QueryEngine timeout and max-samples (`--query.timeout`, `--query.max-samples`); `/info_labels` adds no further protection beyond what `/api/v1/query` already enforces. ***Memory bound*** -The extractor drains all matching info-metric series before streaming the first batch. To prevent unbounded growth on pathological `metric_match` values (e.g. one that matches the entire metric universe), the extractor caps the number of distinct label NAMES it collects at `--web.search.max-limit` (default 10000; setting the flag to 0 disables the cap and makes the extractor unbounded). Once the cap is hit, the extractor keeps draining the series set to gather values for already-collected names but rejects any not-yet-seen names, and a warning is surfaced in the first batch. +The extractor drains all matching info-metric series before streaming the first batch. To prevent unbounded growth in distinct names on pathological `metric_match` values (e.g. one that matches the entire metric universe), the extractor caps the number of names it collects at `--web.search.max-limit` (default 10000; setting the flag to 0 disables the name cap). The cap applies to all retained per-name state, including memoized filter decisions. Once the cap is hit, the extractor keeps draining the series set to gather values for already-collected names but does not retain state for not-yet-seen names, and a warning is surfaced in the first batch. + +Value cardinality is an independent dimension. With a positive `values_limit`, the extractor must retain at most that many values per collected name while preserving the lexicographically smallest values for deterministic output. With `values_limit` unset or zero, values remain unbounded by this endpoint. The extractor is fully memory-bounded by these dimensions only when both the operator name cap and a positive per-request value cap are active. `sort_by=score` under a hit cap may miss the true top-N because some high-scoring names were never seen — documented as a known limitation; operators who want full coverage can raise the cap. Streaming gives clients incremental rendering after the storage scan completes; it does not eliminate the scan itself. ***Performance*** -* Time: O(series matching `metric_match`) × O(labels per series). Per-series work is one label-set iteration plus a memoised filter-accept lookup per distinct label name, so the filter cost is amortised to O(1) per name across all series. -* Memory: bounded by `namesLimit` (distinct names) × `values_limit` (values per name) × average value length. With defaults (`--web.search.max-limit=10000`, no `values_limit`), the in-practice ceiling is set by the cardinality of the real info-metric data; the cap exists to bound the worst case, not the typical one. When `values_limit` is unset, the per-name values count is bounded only by the data's natural cardinality, which is typically small (tens) for info metrics. +* Time: O(series matching `metric_match`) × O(labels per series). Per-series work is one label-set iteration plus filter evaluation. Decisions for retained names can be memoized; names encountered after the cap may be re-evaluated rather than retained, trading bounded memory for repeated filter work. +* Memory: with a positive `values_limit` enforced during extraction, bounded by `namesLimit` (distinct names) × `values_limit` (values per name) × average value length. With the defaults (`--web.search.max-limit=10000`, no `values_limit`), the values dimension is bounded only by the data's natural cardinality. A post-scan truncation does not provide this memory bound. * No additional storage round-trips beyond the one info-metric `Select`. The optional `expr` adds one instant-query evaluation through the standard QueryEngine path. #### Response @@ -162,6 +184,8 @@ Zero or more `{results, warnings?}` batch lines, followed by a `{status, has_mor This is the same contract as the `/api/v1/search/*` endpoints. +The terminal record defines stream completeness. A client must reject the response if EOF occurs before a terminal success or error record, if a line is malformed, if more than one terminal record appears, or if content follows the terminal record. Partial batches must not be treated or cached as a successful response. + URLs in the examples below are shown unescaped for readability — single-quote them on the shell or let curl perform the percent-encoding. ##### Example: no scoping @@ -235,6 +259,8 @@ The pure Prometheus implementation does not include an `extensions` field. Downs * **Recommended manual verification (not automated today):** a client can enumerate `target_info` series via `/api/v1/series`, dedupe label names client-side, and compare against `/api/v1/info_labels` output for the same time window. * The implementation's unit tests cover the parameter surface, the dual gate, `match[]` rejection, sort/score/fuzzy combinations, `values_limit` truncation, multi-batch streaming, and error paths (invalid expr, scalar result, queryable error). +* Extraction tests must also verify that the name and value limits constrain retained state during the scan, not only the final response. +* The Grafana client PoC validates independent editor integration, including expression scoping, metric matching, label-name and label-value completion, and buffered NDJSON consumption. Its production-readiness review motivated the bounded request profiles and completeness requirements documented above. ### Migration @@ -246,6 +272,8 @@ New endpoint. Both feature flags it depends on are experimental and opt-in. No m * **Whether to accept `match[]` in addition to `expr`.** The current design rejects it on cohesion grounds; reviewers may push back. * **Whether identifying labels should be configurable per request.** Today the set is `{"job", "instance"}` from `infohelper.DefaultIdentifyingLabels`. Per-request override is plausible if downstream implementations have different identifier conventions. * **One combined `info-labels-api` flag.** Rejected here on grounds of flag proliferation, but worth re-litigating if reviewers prefer a single-flag operator UX. +* **First-class name-only and exact-name requests.** The current contract can bound name discovery with `values_limit=1` and emulate exact lookup with strict search parameters plus client-side verification, but dedicated semantics would be clearer and avoid returning an unused value during name completion. +* **A bounded default for values.** Leaving `values_limit` unset preserves all values but leaves response size and extractor memory unbounded in that dimension. A future revision may prefer a server default or operator cap. ## Alternatives @@ -255,7 +283,7 @@ Would collapse two endpoints into one. Rejected on cohesion grounds: PROM-74 kee ### 2. Add per-name expansion via a follow-up `/api/v1/search/label_values` call -Works in principle but reintroduces the N+1 round-trip that PROM-74 is trying to fix. For autocomplete this re-creates the latency problem and increases backend load. +An interactive client can defer this call until the user selects a label, making the normal path two requests rather than an eager 1 + N requests. This is viable, but the search endpoints cannot apply the `expr`-derived info-metric scope, so the client cannot keep name and value results consistently scoped for arbitrary PromQL. The dedicated endpoint uses the same scoping contract for both phases. ### 3. Pure client-side composition (`/api/v1/labels` + `/api/v1/label/{name}/values` + filter) @@ -273,6 +301,6 @@ Out of scope. `info()` is a PromQL function that already uses the existing `/api * [X] Build a working implementation (working branch `arve/info-autocomplete` in `aknuds1/prometheus`). * [X] Add the dual feature-gate check to the implementation (search-api + promql-experimental-functions). +* [X] Open the [WIP upstream implementation PR](https://github.com/prometheus/prometheus/pull/17930). * [ ] Finalize this proposal based on community feedback. -* [ ] Open the upstream implementation PR against `prometheus/prometheus`. * [ ] Update `docs/querying/api.md` (the implementation already includes this; the action item is post-merge alignment). From 477232b7207d6eb429b8b2542b34af07e35a0c52 Mon Sep 17 00:00:00 2001 From: Arve Knudsen Date: Fri, 17 Jul 2026 08:31:57 +0200 Subject: [PATCH 3/3] proposal: split info label discovery endpoints Signed-off-by: Arve Knudsen --- proposals/0085-info-labels-endpoint.md | 349 +++++++++++-------------- 1 file changed, 152 insertions(+), 197 deletions(-) diff --git a/proposals/0085-info-labels-endpoint.md b/proposals/0085-info-labels-endpoint.md index 0310b189..ca2b7d89 100644 --- a/proposals/0085-info-labels-endpoint.md +++ b/proposals/0085-info-labels-endpoint.md @@ -1,306 +1,261 @@ -## Info-metric label discovery API for `info()` autocomplete +## Info-metric label discovery APIs for `info()` autocomplete * **Owners:** * Arve Knudsen [@aknuds1](https://github.com/aknuds1) arve.knudsen@gmail.com * Ismail Simsek [@itsmylife](https://github.com/itsmylife) ismail.simsek@grafana.com -* **Implementation Status:** Not implemented upstream ([WIP implementation PR](https://github.com/prometheus/prometheus/pull/17930)) +* **Implementation Status:** Not implemented upstream ([WIP implementation PR](https://github.com/prometheus/prometheus/pull/17930)). * **Related Issues and PRs:** - * [PROM-74 — V2 API for labels and values discovery](https://github.com/prometheus/proposals/pull/74). This proposal builds on PROM-74's NDJSON contract, feature gate, and shared parameter set. - * [PROM-37 — Simplify joins with info metrics in PromQL](./0037-native-support-for-info-metrics-metadata.md). Introduces the `info()` PromQL function; the proposed endpoint supports its autocomplete. - * [Grafana Prometheus datasource PR #244](https://github.com/grafana/grafana-prometheus-datasource/pull/244). Independent client PoC for `info()` data-label autocomplete using this endpoint. + * [PROM-74 — V2 API for labels and values discovery](https://github.com/prometheus/proposals/pull/74). This proposal builds on PROM-74's NDJSON model, feature gate, search behavior, and storage interfaces. The request bounds below follow the current search implementation and are authoritative where PROM-74's original text differs. + * [PROM-37 — Simplify joins with info metrics in PromQL](./0037-native-support-for-info-metrics-metadata.md). Introduces the `info()` PromQL function supported by this autocomplete API. + * [Prometheus PR #19557 — fix `info()` with mixed identifying-label presence](https://github.com/prometheus/prometheus/pull/19557). Independent correctness fix for the existing experimental evaluator; it is not part of this proposal. + * [Grafana Prometheus datasource PR #244](https://github.com/grafana/grafana-prometheus-datasource/pull/244). Independent client PoC and source of the client-integration feedback incorporated here. -* **Other docs or links:** - * PromQL `info()` function documentation in `docs/querying/functions.md`. - -> TL;DR: Add `GET|POST /api/v1/info_labels` as the info-metric companion to the experimental `/api/v1/search/*` family. The endpoint streams the non-identifying data labels of info metrics (e.g. `target_info`) as NDJSON. Queries can be scoped either by an explicit info-metric specifier (`metric_match`), or by evaluating a PromQL expression (`expr`) and harvesting its identifying labels. The endpoint reuses PROM-74's NDJSON contract, fuzzy/sort/score/limit/batch parameters, and `--enable-feature=search-api` gate. It additionally requires `--enable-feature=promql-experimental-functions`, since its only purpose is to support autocomplete for the experimental `info()` function. +> TL;DR: Add `GET|POST /api/v1/info_labels` and `GET|POST /api/v1/info_label_values` as info-metric companions to PROM-74's experimental search API. Both endpoints apply the same `expr` and repeated `data_match[]` scope. The first searches data-label names; the second searches values for one exact data-label name. They reuse the current search API's NDJSON, search, sort, score, limit, batch, and storage behavior, apply one timeout across expression evaluation and storage search, and require both `search-api` and `promql-experimental-functions`. ## Why -The PromQL `info()` function (added under `--enable-feature=promql-experimental-functions`) enriches base metrics with the data labels carried on info metrics like `target_info`. A PromQL editor offering autocomplete inside `info(, {…})` needs to know **which data labels are available for the in-scope info metrics**, where the in-scope set is determined by the user's in-progress expression. - -This is a new gap created by the introduction of `info()` itself — it is not specific to any one editor or vendor. Anyone building a PromQL editor that supports `info()` will hit the same problem. - -The Prometheus UI PoC integrates the endpoint through `web/ui/module/lezer-promql/src/client/`. Grafana has a separate Monaco-based client PoC, demonstrating that the endpoint contract is usable without sharing Prometheus' editor implementation. - -### Pitfalls of the current solution +The PromQL `info()` function enriches base series with data labels from info metrics such as `target_info`. An editor completing `info(, {…})` must first discover which data-label names exist for the in-scope info series, then discover values only for the name the user selected. -Constructing an equivalently scoped interaction with the endpoints that exist today requires either fetching far more than needed or combining endpoints that cannot express arbitrary-PromQL info-metric scoping consistently. +This is not editor-specific. The Prometheus UI and the Grafana Prometheus datasource use different editor implementations, but both require the same server-side operations. -* **`/api/v1/labels` + per-name `/api/v1/label/{name}/values`** returns *all* labels matching a selector, including labels carried by the base metric, not just data labels on info metrics. Filtering client-side requires downloading the full universe of labels and then making one `/label/{name}/values` call per label of interest (N+1). -* **`/api/v1/series` + client aggregation** returns one row per series, then the client must dedupe labels and prune identifying labels. The wire size is O(series), and there is no server-side scoring for relevance ranking. -* **`/api/v1/search/label_names` + `/api/v1/search/label_values`** (from PROM-74) is closer, but neither endpoint understands the info-metric scoping pattern (find label names on info metrics whose identifying labels intersect with a PromQL expression's result set). A client can defer the value lookup until the user selects a name, avoiding eager 1 + N calls, but it still has to reproduce the missing info-metric scoping across two endpoint families. -* **`/api/v1/search/*` with `match[]` alone** cannot scope autocomplete to be context-aware against arbitrary PromQL like `rate(http_requests_total[5m])`, which is what users actually type. The endpoint needs to evaluate an expression and harvest its identifying labels, not merely accept a series selector. +Existing APIs cannot express those operations efficiently and consistently: -For interactive autocomplete this manifests as visible latency on every keystroke and unnecessary load on the server. +* `/api/v1/labels` and `/api/v1/label/{name}/values` include labels from base metrics as well as info metrics and cannot derive the info scope from arbitrary PromQL. +* `/api/v1/series` transfers one row per series and requires the client to deduplicate labels and remove identifying labels. +* PROM-74's general label search endpoints support filtering and ranking, but do not evaluate an arbitrary expression to derive the identifying-label values that scope the relevant info series. +* `match[]` can express selectors, not expressions such as `rate(http_requests_total[5m])` that users actually type. ## Goals -* One bounded round-trip for each autocomplete interaction: discover label names, then fetch values only for the label the user selects. -* Server-side scoping by either an explicit info-metric specifier (`metric_match`) or by evaluating a PromQL expression (`expr`) and using its identifying labels. -* Reuse PROM-74's NDJSON contract, search/fuzzy/sort/score/limit/batch parameters, and `--enable-feature=search-api` gate so operators see one coherent experimental autocomplete API surface. -* Couple the endpoint to the function it serves: also require `--enable-feature=promql-experimental-functions`, so an operator cannot reach a state where the endpoint is up but `info()` queries fail to parse. -* Give callers independent caps for label names and values, and provide predictable degradation under pathological input (e.g. `target_info` with very high `version`/`env`/`region` cardinality, or a `metric_match` that matches far more than expected). +* One bounded request for each autocomplete phase: names first, values for one selected name second. +* Identical server-side scoping for both phases, using full PromQL name and data-label matchers and optionally the identifying labels harvested from a PromQL expression. +* Exact label-name semantics for value lookup. Fuzzy name search must not be used to emulate selecting one label. +* Reuse the current search API's request, streaming, storage, and feature-gate behavior, with the explicit bounds below. +* Keep the response units simple and independently bounded: one name per name record, one value per value record. ## Non-Goals -* Replace or modify `/api/v1/labels` or `/api/v1/label/{name}/values`. -* Add or modify the `/api/v1/search/*` family. This proposal *builds on* PROM-74; it does not extend it. -* Server-side evaluation of `info()`. This is a metadata endpoint; `info()` queries continue to use `/api/v1/query`. -* Cursor-based pagination. Same stance as PROM-74; `has_more` is informational only. -* Special handling of `__type__` / `__unit__` labels. They are not in `infohelper.DefaultIdentifyingLabels` (which is `{"instance", "job"}`), so they pass through the extractor as ordinary data labels and may appear in the response if set on info metrics. +* Replace or modify the existing label or series endpoints. +* Add server-side evaluation of `info()`; query evaluation remains on `/api/v1/query`. +* Modify the `info()` evaluator; mixed identifying-label presence correctness is tracked independently. +* Add cursor-based pagination. As in PROM-74, `has_more` reports truncation but is not a cursor. +* Make identifying labels configurable per request. The initial set is `job` and `instance`. ## How -A new endpoint, gated behind two feature flags, that streams NDJSON using the same wire framing (batches + trailer) as `/api/v1/search/*`. The per-record shape is `/info_labels`-specific (`{name, values, score?}`). - -### Implementation notes - -* The parameter set extends PROM-74's shared parameters (`search[]`, `fuzz_*`, `sort_*`, `include_score`, `start`, `end`, `limit`, `batch_size`) with three /info_labels-specific parameters: `expr`, `metric_match`, and `values_limit`. -* `match[]` is not accepted. Scoping happens via `metric_match` (which info metrics to consider) and `expr` (which identifying-label values to restrict to). Accepting `match[]` as well would create two equivalent matcher mechanisms with unclear precedence. -* The response unit is one record per data label name, carrying that label's values inline. Clients should use different limits for the two autocomplete phases instead of eagerly retrieving all values for every candidate name. -* The response is NDJSON (`application/x-ndjson`) with the same batch + trailer contract as PROM-74. -* The endpoint is dual-gated: `--enable-feature=search-api` covers the NDJSON + parsing infrastructure it reuses; `--enable-feature=promql-experimental-functions` covers the only consumer (`info()`). Either missing flag returns the standard Prometheus JSON error with `errorType: unavailable` and a flag-specific message. - -### Client integration feedback - -The independent Grafana PoC validates the request and response shape while exposing several requirements that apply to PromQL editors generally: - -* `expr` must be valid PromQL. Clients with editor-specific variables or macros must resolve them with the same semantics as normal query execution before sending the request. -* Each request must use the current query time range. A client cache should be bounded and freshness-limited, keyed by the effective resolved request, retain only complete successful responses, and allow retries after failures. -* A label-name completion request should send the current `start`/`end`, resolved `expr`, optional `metric_match`, and typed `search[]`. Omitting `limit` uses the server's at-most-100-name default, reduced further when the operator configures a lower maximum; `values_limit=1` prevents the name-completion response from carrying every value for every name. -* A label-value completion request can use the existing search contract with `search[]=`, `fuzz_alg=subsequence`, `fuzz_threshold=100`, `case_sensitive=true`, `sort_by=alpha`, `sort_dir=asc`, `limit=1`, and an explicit client-selected `values_limit`. Subsequence prefix matches can tie with the exact name, so the client must still verify that the returned record name equals the requested name before using its values. - -The last request profile is a bounded workaround, not an exact-name API primitive. A first-class name-only mode and exact-name filter remain open design questions. - -### `GET|POST /api/v1/info_labels` - -#### Request +Add two dual-gated endpoints with a shared scope and distinct result types: -**Method:** `GET` or `POST` +| Endpoint | Search target | Result record | +|-----------------------------|---------------------------------------|-----------------------------------------| +| `/api/v1/info_labels` | Non-identifying label names | `{ "name": string, "score"?: number }` | +| `/api/v1/info_label_values` | Values of the exact `label` parameter | `{ "value": string, "score"?: number }` | -`POST` is recommended when the `expr` parameter exceeds URL length limits. +These are top-level `info_*` endpoints rather than `/api/v1/search/*` resources because `expr`, the info-specific matcher split, and the dual `info()` feature gate are function-specific semantics. Reusing PROM-74's `Searcher`, request parameters, and NDJSON contract does not make the operations general label search. -**Path:** `/api/v1/info_labels` +The separation follows the two editor interactions and PROM-74's label-name and label-value split. It avoids a combined `{name, values[]}` response whose two independent cardinality dimensions require `values_limit`, encourages eager value retrieval, and cannot give the selected label first-class exact semantics. -**Query parameters:** +Requests that reach streaming return `200 OK` with `application/x-ndjson` using PROM-74's zero-or-more batch lines followed by exactly one terminal success or error record. Validation and setup failures, plus first-iteration failures that produce no batch results, return the usual non-2xx Prometheus JSON error. Both endpoints use the `storage.Searcher` interface: `SearchLabelNames` for names and `SearchLabelValues` for values. -| Name | Type | Required | Default | Description | -|------------------|--------------------------|----------|-----------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| `expr` | string (PromQL) | No | | Optional PromQL expression evaluated server-side. Identifying labels (`job`, `instance`) extracted from the result restrict the info-metric query so autocomplete sees only the labels relevant to the user's expression. | -| `metric_match` | string | No | `target_info` (exact) | Info-metric specifier. Bare value (no prefix) is exact match; `=~` / `!=` / `!~` prefixes give the corresponding matcher type (e.g. `target_info`, `=~.*_info`, `!=target_info`, `!~build_info`). | -| `search[]` | []string | No | | One or more search terms matched against data label NAMES (OR semantics). As per PROM-74. | -| `fuzz_threshold` | int [0..100] | No | 0 | Fuzzy threshold. As per PROM-74. | -| `fuzz_alg` | string | No | `subsequence` | Fuzzy algorithm. As per PROM-74. | -| `case_sensitive` | bool | No | true | As per PROM-74. | -| `sort_by` | `alpha` / `score` | No | | As per PROM-74. | -| `sort_dir` | `asc` / `dsc` | No | `asc` | As per PROM-74. | -| `include_score` | bool | No | false | As per PROM-74. | -| `start`, `end` | rfc3339 / unix_timestamp | No | last 1h | As per PROM-74. | -| `limit` | int ≥ 0 | No | 100, capped by `--web.search.max-limit` | Maximum number of label NAMES returned. Same `has_more` semantics as PROM-74. | -| `values_limit` | int ≥ 0 | No | 0 (no cap) | Maximum number of values returned per label. New to this endpoint. | -| `batch_size` | int > 0 | No | 100 | As per PROM-74. | +Search capability is fail-closed across the full active query path. Before streaming begins, every participating non-noop storage querier must implement `storage.Searcher`; otherwise the request returns the standard non-streaming Prometheus JSON error with `errorType: unavailable`. A federated or fanout implementation must not silently omit results from a storage backend that lacks search support. Runtime errors from search-capable secondary storage retain PROM-74's best-effort warning behavior. -**Notes:** +### Shared request parameters -***Feature gating*** +| Name | Type | Required | Default | Description | +|------------------|-------------------------------|----------|--------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `expr` | string (PromQL) | No | | Instant-vector expression evaluated to derive `job` and `instance` values that restrict the info-series scope. | +| `data_match[]` | []string | No | `__name__="target_info"` | Repeated full PromQL matchers from `info()`'s data-label selector, including its special `__name__` matcher. Multiple matchers have AND semantics. | +| `time` | rfc3339 / unix timestamp | No | `end` | Instant evaluation time for `expr`. | +| `lookback_delta` | duration / float seconds | No | server default | Lookback used both to evaluate `expr` and to select matching info series. | +| `timeout` | duration / float seconds | No | `--query.timeout` | Shared timeout for expression evaluation and the subsequent info-series search. Explicit values are capped by `--query.timeout`. | +| `start`, `end` | rfc3339 / unix timestamp | No | last 1h | Storage search window when `expr` is absent. With `expr`, `start` must parse but is ignored for range selection and ordering validation; `end` supplies the default `time`. | +| `search[]` | []string | No | | At most 32 search terms, matched against names or values according to the endpoint. Multiple terms have OR semantics. | +| `fuzz_threshold` | int [0..100] | No | 0 | Fuzzy threshold, as in PROM-74. | +| `fuzz_alg` | `jarowinkler` / `subsequence` | No | `jarowinkler` | Fuzzy algorithm, as in PROM-74. | +| `case_sensitive` | bool | No | true | Case sensitivity, as in PROM-74. | +| `sort_by` | `alpha` / `score` | No | value ascending | Ordering, as in PROM-74. `sort_by=score` requires `search[]`. | +| `sort_dir` | `asc` / `dsc` | No | `asc` | Direction for alphabetical ordering. Accepted only with explicit `sort_by=alpha`. | +| `include_score` | bool | No | false | Include the deterministic relevance score in each record. | +| `limit` | positive int | No | 100 | Maximum records returned after filtering and ordering. A positive `--web.search.max-limit` caps explicit values and may reduce the default; 0 disables that operator cap, not the request limit. | +| `batch_size` | positive int, at most 10000 | No | 100 | Requested records per NDJSON batch line. Its effective value is no greater than `limit`. | -The endpoint requires *both* `--enable-feature=search-api` (NDJSON + parsing infrastructure) and `--enable-feature=promql-experimental-functions` (the only consumer is `info()` autocomplete; without `info()`, no client can use the response). The handler checks both flags upfront and, if either is missing, responds with the standard non-streaming JSON error (`errorType: unavailable`) carrying a single message that names every missing flag, so an operator sees the whole gap in one round-trip: +`limit` bounds the endpoint's retained result state and response size. It is passed to the storage search as a hint, but the API does not promise that an implementation can avoid scanning or enumerating additional index data to determine the requested ordered results or `has_more`. Unlike PROM-74's original request table, `limit=0` and `batch_size=0` are rejected. -* `"info_labels requires --enable-feature=search-api"` — only search-api missing. -* `"info_labels requires --enable-feature=promql-experimental-functions"` — only experimental-functions missing. -* `"info_labels requires --enable-feature=search-api,promql-experimental-functions"` — both missing. +At most 32 `data_match[]` values may be supplied. Each value must parse as exactly one full PromQL matcher. Clients can therefore preserve every completed matcher from the data-label selector as its original source without classifying `__name__` separately or inventing an operator-prefix encoding. -***expr*** +`match[]` is rejected on both endpoints. It would introduce a second series-scoping mechanism with unclear precedence relative to `data_match[]` and `expr`. -If `expr` is provided, the server evaluates it as an instant query at the request's `end` timestamp. The result must be a Vector or Matrix; Scalar/String returns `errorBadData`. The server harvests the identifying labels (`job`, `instance`) from each sample/series and uses their values as additional matchers against the info-metric query. If the eval produces no identifying-label values, the stream consists of a single empty batch (carrying any warnings) followed by the success trailer. +### Scope semantics -The endpoint accepts PromQL, not editor-specific template syntax. A client must resolve local variables and macros before sending `expr`; using the same interpolation path as normal query execution avoids autocomplete evaluating a different expression than the query itself. +`data_match[]` carries the completed matchers from `info()`'s data-label selector. The server partitions them by label name: `__name__` matchers select which info metric families to search, while all remaining matchers filter those info series. -This is what makes the endpoint context-aware: `expr=rate(http_requests_total[5m])` gives an autocomplete experience scoped to the info metrics actually relevant to that PromQL expression, without requiring the client to interpret PromQL itself. +* No name matcher means `__name__="target_info"`. +* `data_match[]=__name__=~".+_info"` selects matching info metric names. +* Repeated name matchers are ANDed, for example `__name__=~".+_info"` and `__name__!="custom_info"`. +* A negative-only request is implicitly restricted by `__name__=~".+_info"`, so `__name__!="target_info"` cannot broaden the search to every non-target metric. -When `metric_match` is also set, the two filters apply jointly: an info-metric series must match `metric_match`'s name pattern AND have identifying labels intersecting the `expr` eval result. This lets callers narrow both the info-metric family (e.g. `metric_match==~.*_info`) and the relevant series (e.g. `expr=rate(http_requests_total[5m])`). +The remaining matchers filter the selected info series before either endpoint discovers names or values. They are ANDed with one another, the effective name matchers, and the expression-derived scope. -***metric_match*** +When `expr` is present, Prometheus evaluates it as an instant query at `time`, which defaults to `end`. The expression must have instant-vector type; matrix, scalar, and string expressions are rejected before execution. Series already matching the effective info-metric name scope are ignored, consistent with `info()` not enriching info series. An expression that produces no remaining series with a non-empty identifying label returns a successful empty stream rather than falling back to an unscoped search. -Specifies which info metrics to consider. Prefix syntax mirrors PromQL matcher operators: +The subsequent info-series search uses the same temporal selection hints as an instant `info(expr, ...)` evaluation: it applies the effective lookback delta and honors selector `offset` and `@` modifiers. The general search `start` parameter does not broaden or narrow this expression-derived range. -* `metric_match=target_info` → `__name__="target_info"` (default) -* `metric_match==~.*_info` → `__name__=~".*_info"` -* `metric_match=!=target_info` → `__name__!="target_info"` -* `metric_match=!~build_info` → `__name__!~"build_info"` +When every vector-producing path has the same effective reference, including enclosing subquery offsets and `@` anchors, the search uses that shared reference. If the paths disagree or any vector-producing path is selector-free, selection falls back to the query evaluation time rather than choosing an arbitrary selector. -The double-equals in `metric_match==~.*_info` is not a typo — the first `=` is the URL form-value separator and the second `=` is the matcher-type prefix. +The handler groups expression results by identifying-label presence: `job` only, `instance` only, and both. A missing identifying label becomes an exact empty matcher instead of a wildcard. Within a presence group, observed values become escaped regular-expression alternatives. This avoids false negatives and keeps the number of storage selections bounded by the number of presence patterns rather than the number of expression series. A group containing both labels may conservatively inspect cross-pairs of independently observed `job` and `instance` values. That can make autocomplete suggestions broader, but it cannot change query results: `info()` performs the exact identifying-label join when the completed query executes. -Clients deriving `metric_match` from a PromQL `__name__` matcher must preserve the full operator. In particular, a regex matcher is encoded as the form value `=~.*_info`; dropping the leading `=` changes it into an exact metric name. +Matcher construction is also bounded before the info-series `Searcher` is opened. At most 10,000 unique `(presence group, identifying label, value)` entries and 1,048,576 total escaped regular-expression bytes, including alternation separators, are accepted. Duplicate values do not consume the value budget. Exceeding either bound returns a pre-stream `bad_data` error asking the caller to narrow `expr`; values are never truncated because truncation could silently omit valid completions. -***match[]*** +The name matchers, data matchers, and `expr` apply jointly. Expression warnings are merged with storage warnings and exposed on the stream. -Not accepted. Requests carrying `match[]` are rejected with `errorBadData`: `"match[] is not supported by info_labels; use metric_match or expr"`. +Clients must resolve editor-specific variables and macros before sending `expr`, using the same semantics as normal query execution. Otherwise autocomplete and query execution can observe different expressions. -***values_limit*** +### Security and admission control -Caps the number of values per label. Distinct from `limit`, which caps the number of records (per PROM-74's semantics). Values are sorted alphabetically and truncated after the sort, so the truncation is deterministic. +An `expr` request is a query, not metadata-only discovery. Response presence can depend on sample values, for example `expr=up == 0`, even though the response contains only label names or values. Expression evaluation and the subsequent info-series search must therefore use the same request and tenant context, and callers must have the same query authorization required by `/api/v1/query`. -The two dimensions multiply: a response can contain up to `limit * values_limit` values when both are non-zero. Omitting `values_limit`, or setting it to zero, leaves value cardinality uncapped. Interactive clients should therefore always set it explicitly. +Deployments whose authorization layer cannot distinguish requests by form or query parameter must protect both endpoint routes as query surfaces, including requests that omit `expr`. The endpoints consolidate existing discovery operations, but `expr` makes it incorrect to claim that their observable result is limited to metadata already available through label APIs. -To make `values_limit` a server-memory bound as well as a wire bound, implementations must enforce it while collecting values, retaining only the lexicographically smallest values needed for the deterministic result. Collecting every distinct value and truncating only after the scan bounds the response but not extractor memory. - -***start/end*** - -The default window is the last hour (inherited from PROM-74's `parseSearchParams`). This is conservative — info metrics like `target_info` typically only emit one sample per scrape and may be scraped infrequently, so a deployment with a long scrape interval and an idle window can land outside a 1-hour lookback and return no data. Operators with infrequent scrapes should pass an explicit `start` to widen the window; interactive clients should pass the current editor time range rather than retaining the range from editor initialization. See Known unknowns for the open question of whether `/info_labels` should ship with a longer default than `/search/*`. - -***Storage*** - -The endpoint uses the existing `storage.Querier` interface, not PROM-74's `Searcher`. The unit of work — enumerate info-metric series, collect their non-identifying labels deduped across series, with values — does not map onto `SearchLabelNames` (which returns names without values and would force one `SearchLabelValues` call per name). The search-api `storage.Filter` produced by `buildSearchFilter` is applied to label names in-memory inside the extractor loop, so fuzzy/score behaviour is identical to `/api/v1/search/label_names` for the same input names (iteration order differs because `/info_labels` is series-driven rather than index-driven, but the difference is invisible once results pass through the post-sort). +### `GET|POST /api/v1/info_labels` -***Security*** +This endpoint searches non-identifying label names on the scoped info series. `__name__`, `job`, and `instance` are filtered before the result limit is applied, so they cannot consume autocomplete slots. -* `expr` is evaluated through the standard `QueryEngine`, so the existing per-query timeout, max-samples limit, and any deployed query authorisation policy apply unchanged. The endpoint adds no new evaluation path. -* The `namesLimit` cap described in the *Memory bound* note bounds the number of distinct label names, but not the number of values retained for each name. An implementation must apply the name cap to all retained per-name state and enforce a positive `values_limit` during extraction to bound both dimensions; requests with an omitted or zero `values_limit` remain unbounded in the value dimension. -* The endpoint exposes no metadata that a caller cannot already obtain by combining `/api/v1/labels`, `/api/v1/label/{name}/values`, and client-side post-filtering. It packages the same information more efficiently; it does not widen the read surface. -* Expensive `expr` is bounded by the existing QueryEngine timeout and max-samples (`--query.timeout`, `--query.max-samples`); `/info_labels` adds no further protection beyond what `/api/v1/query` already enforces. +The `label` parameter is rejected; callers seeking values must use `/api/v1/info_label_values`. -***Memory bound*** +Example: -The extractor drains all matching info-metric series before streaming the first batch. To prevent unbounded growth in distinct names on pathological `metric_match` values (e.g. one that matches the entire metric universe), the extractor caps the number of names it collects at `--web.search.max-limit` (default 10000; setting the flag to 0 disables the name cap). The cap applies to all retained per-name state, including memoized filter decisions. Once the cap is hit, the extractor keeps draining the series set to gather values for already-collected names but does not retain state for not-yet-seen names, and a warning is surfaced in the first batch. +```bash +curl -N -g 'http://localhost:9090/api/v1/info_labels?expr=rate(http_requests_total{job="api"}[5m])&data_match[]=__name__=~".+_info"&data_match[]=env="prod"&search[]=ver&sort_by=score&include_score=true' +``` -Value cardinality is an independent dimension. With a positive `values_limit`, the extractor must retain at most that many values per collected name while preserving the lexicographically smallest values for deterministic output. With `values_limit` unset or zero, values remain unbounded by this endpoint. The extractor is fully memory-bounded by these dimensions only when both the operator name cap and a positive per-request value cap are active. +```ndjson +{"results":[{"name":"version","score":1},{"name":"server","score":0.83}]} +{"status":"success","has_more":false} +``` -`sort_by=score` under a hit cap may miss the true top-N because some high-scoring names were never seen — documented as a known limitation; operators who want full coverage can raise the cap. Streaming gives clients incremental rendering after the storage scan completes; it does not eliminate the scan itself. +### `GET|POST /api/v1/info_label_values` -***Performance*** +This endpoint requires `label`, interpreted as one exact decoded label name. It is a query parameter rather than a path component so UTF-8 label names do not require a second path-specific quoting contract. -* Time: O(series matching `metric_match`) × O(labels per series). Per-series work is one label-set iteration plus filter evaluation. Decisions for retained names can be memoized; names encountered after the cap may be re-evaluated rather than retained, trading bounded memory for repeated filter work. -* Memory: with a positive `values_limit` enforced during extraction, bounded by `namesLimit` (distinct names) × `values_limit` (values per name) × average value length. With the defaults (`--web.search.max-limit=10000`, no `values_limit`), the values dimension is bounded only by the data's natural cardinality. A post-scan truncation does not provide this memory bound. -* No additional storage round-trips beyond the one info-metric `Select`. The optional `expr` adds one instant-query evaluation through the standard QueryEngine path. +Empty `label`, `__name__`, `job`, and `instance` are rejected because they are not info data labels. `search[]`, if present, filters and ranks values of the selected label; it never changes which label is selected. -#### Response +Example: -* `Content-Type: application/x-ndjson; charset=utf-8` +```bash +curl -N -g 'http://localhost:9090/api/v1/info_label_values?label=version&expr=rate(http_requests_total{job="api"}[5m])&search[]=v2&sort_by=score' +``` -Zero or more `{results, warnings?}` batch lines, followed by a `{status, has_more, warnings?}` trailer, or — on mid-stream failure after the first batch has been written — a `{status, errorType, error}` line in place of the trailer. Errors that occur before the first batch is written are returned as the usual non-streaming Prometheus JSON error with the matching 4xx/5xx status code. +```ndjson +{"results":[{"value":"v2.1"},{"value":"v2.0"}]} +{"status":"success","has_more":false} +``` -This is the same contract as the `/api/v1/search/*` endpoints. +### Feature gating -The terminal record defines stream completeness. A client must reject the response if EOF occurs before a terminal success or error record, if a line is malformed, if more than one terminal record appears, or if content follows the terminal record. Partial batches must not be treated or cached as a successful response. +Both endpoints require: -URLs in the examples below are shown unescaped for readability — single-quote them on the shell or let curl perform the percent-encoding. +* `--enable-feature=search-api`, for the experimental search and NDJSON infrastructure. +* `--enable-feature=promql-experimental-functions`, because this API serves the experimental `info()` function. -##### Example: no scoping +Missing gates return the standard non-streaming Prometheus JSON error with `errorType: unavailable`; the message names every missing feature in one response. -```ndjson -{"results":[{"name":"cluster","values":["us-east","us-west"]},{"name":"env","values":["prod","staging"]},{"name":"region","values":["us-east"]},{"name":"version","values":["v1.0","v2.0","v2.1"]}]} -{"status":"success","has_more":false} -``` +`GET /api/v1/features` advertises the pair through `data.api.info_label_search`. It is `true` only when both flags are enabled and Prometheus is running in server rather than Agent mode. This advertises that the routes are enabled, not that every storage backend participating in a particular request supports search; clients must still handle a per-request `unavailable` response. Clients must treat an absent or false value as unsupported and ignore unknown feature keys. Development PoCs targeting servers that predate this capability may retain an explicit, off-by-default operator override. -##### Example: scoped by expression +### Timeout and cancellation -```bash -curl -N 'http://localhost:9090/api/v1/info_labels?expr=rate(http_requests_total{job="api-gateway"}[5m])' -``` +`timeout` uses the same duration and float-seconds syntax as `/api/v1/query`. The effective duration is the shorter of the requested value and `--query.timeout`; when omitted, the server flag supplies the duration. One deadline covers expression evaluation and the subsequent context-aware storage search, so time spent evaluating `expr` reduces the time remaining for discovery. -```ndjson -{"results":[{"name":"cluster","values":["us-east","us-west"]},{"name":"env","values":["prod","staging"]},{"name":"version","values":["v1.0","v2.0"]}]} -{"status":"success","has_more":false} -``` +Expiry before streaming returns the normal Prometheus JSON error with HTTP 503 and `errorType: timeout`. Expiry after a batch has been written terminates the NDJSON stream with an error record carrying `errorType: timeout`; partial results are not successful or cacheable. Caller cancellation remains an abrupt EOF without a terminal record. -##### Example: scoped by search with relevance scoring +### Stream completeness -```bash -curl -N 'http://localhost:9090/api/v1/info_labels?search[]=ver&sort_by=score&include_score=true' -``` +A non-2xx response uses the standard Prometheus JSON error format and is not an NDJSON stream. A successful NDJSON stream ends with: ```ndjson -{"results":[{"name":"version","values":["v1.0","v2.0","v2.1"],"score":1},{"name":"server","values":["nginx","envoy"],"score":0.83}]} {"status":"success","has_more":false} ``` -##### Example: `match[]` rejected - -```bash -curl -N 'http://localhost:9090/api/v1/info_labels?match[]={job="prometheus"}' -``` - -```json -{"status":"error","errorType":"bad_data","error":"match[] is not supported by info_labels; use metric_match or expr"} -``` +A client must reject a response when EOF arrives before a terminal record, a line is malformed, more than one terminal record appears, or content follows the terminal record. A mid-stream error record is terminal. Partial batches must not be returned or cached as a successful response. -##### Example: names truncated at namesLimit +Clients should cache only complete successful responses, bound cache size and freshness, key entries by the effective resolved request, deduplicate identical in-flight requests, and evict failures so a later request can retry. Clients should normally omit `limit` and accept the operator-controlled default. -When the extractor hits the `--web.search.max-limit` cap, the first batch carries a warning naming the cap and what to do about it. The trailer's `status` stays `success` because the result set is still well-formed — just incomplete. +Editor integrations should forward every completed matcher in the second `info()` argument as its original full PromQL source. The matcher currently being edited must be omitted, while other matchers on the same label remain in scope. The second argument has no metric-name prefix syntax; clients encountering an unquoted identifier or quoted metric-name prefix must use generic completion rather than synthesizing a `__name__` matcher. Users select another info metric explicitly with a matcher such as `{__name__="build_info", ...}`. Variables and macros must be resolved before the request, and quoted UTF-8 label names must be decoded before use as the exact `label`. Completion for `__name__`, `job`, and `instance` remains the general metric or label completion path because those are not discoverable info data labels. -```ndjson -{"results":[{"name":"cluster","values":["us-east","us-west"]},...],"warnings":["info-labels names truncated at 10000; narrow metric_match or raise --web.search.max-limit"]} -{"status":"success","has_more":false} -``` +### Storage and performance -### Interface changes +The name endpoint calls `Searcher.SearchLabelNames` with the common scope and a filter that excludes identifying labels before applying search and limit. The value endpoint calls `Searcher.SearchLabelValues` with the exact label and the same scope. Composite storage advertises `Searcher` only when every participating non-noop querier does, including secondary and out-of-order readers. This keeps search, scoring, ordering, limiting, and downstream storage optimizations aligned with PROM-74 rather than reimplementing them in an info-specific series extractor. -Smaller than PROM-74's contribution. This proposal does not introduce new storage interfaces: +The optional `expr` adds one standard instant-query evaluation. Existing max-samples and lookback behavior applies, while the endpoint-level deadline continues through the storage search. Query origin metadata is recorded as for `/api/v1/query`, and both phases retain the same request context. The storage search uses the derived matchers over the exact range that `info()` would select. -* PROM-74's `Searcher.SearchLabelNames` is not used; the endpoint stays on `storage.Querier`. -* The search-api request preamble (CORS, feature-gate, form parsing, `parseSearchParams`, querier acquisition, `buildSearchFilter`, sort/order plumbing) is factored out of `newSearchRequest` into a reusable `newAutocompleteRequest` so both `/api/v1/search/*` and `/api/v1/info_labels` share it. This is a small refactor of code that landed with PROM-74: `newSearchRequest` is preserved as a thin wrapper that adds the `storage.Searcher` type assertion the search-api endpoints require. -* A small helper package (`promql/infohelper`) exposes `ExtractDataLabels(ctx, querier, infoMetricMatcher, identifyingLabelValues, hints, filter, namesLimit, valuesLimit) → []InfoLabelRecord` where `InfoLabelRecord = {Name, Values, Score}`. The filter is applied to label names; the score carries through to the wire when `include_score=true`. +The `limit` contract bounds result retention and wire output, not the cardinality of the underlying index or the worst-case storage work. The actual work depends on the `Searcher` implementation, matcher selectivity, requested ordering, and whether it can stop early while still determining `has_more`. -### Extensibility for Mimir, Thanos, Cortex +### Extensibility for Mimir, Thanos, and Cortex -The per-record JSON shape inherits PROM-74's extensibility story: downstream implementations can add an `extensions` field per record (and an `extensions` map keyed by provider name) without changing the core contract. No new mechanism is required. +As in PROM-74, downstream implementations may add optional per-record extensions without changing the core record shapes: ```ndjson -{"results":[{"name":"cluster","values":["us-east","us-west"],"extensions":{"mimir":{"cardinality":42}}}]} +{"results":[{"name":"cluster","extensions":{"mimir":{"cardinality":42}}}]} {"status":"success","has_more":false} ``` -The pure Prometheus implementation does not include an `extensions` field. Downstream implementations adding it should tag it `json:",omitempty"` so records without provider-specific data stay clean on the wire. +The Prometheus implementation does not emit extensions. ### Testing and verification -* **Recommended manual verification (not automated today):** a client can enumerate `target_info` series via `/api/v1/series`, dedupe label names client-side, and compare against `/api/v1/info_labels` output for the same time window. -* The implementation's unit tests cover the parameter surface, the dual gate, `match[]` rejection, sort/score/fuzzy combinations, `values_limit` truncation, multi-batch streaming, and error paths (invalid expr, scalar result, queryable error). -* Extraction tests must also verify that the name and value limits constrain retained state during the scan, not only the final response. -* The Grafana client PoC validates independent editor integration, including expression scoping, metric matching, label-name and label-value completion, and buffered NDJSON consumption. Its production-readiness review motivated the bounded request profiles and completeness requirements documented above. +Implementation tests cover: + +* both endpoints over GET and POST; +* the dual feature gate; +* `expr`, repeated `data_match[]` values including `__name__`, negative-only name matching, and their joint scope; +* mixed identifying-label presence and conservative cross-pair scoping without false negatives; +* instant-vector type validation, historical `end`, ignored `start`, and temporal equivalence with `info()` for lookback, offsets, `@` modifiers, equivalent and mixed references, selector-free vector paths, and enclosing subqueries; +* exact matcher-construction bounds, duplicate accounting, and rejection before the info `Searcher` opens; +* exact and UTF-8 label names for value lookup; +* identifying-label filtering and rejection; +* matcher count and syntax validation, and rejection of `match[]` and `label` on the name endpoint; +* search, Jaro-Winkler and subsequence fuzzy matching, scoring, ordering, limit, `has_more`, and batching; +* omitted, shorter, capped, invalid, pre-stream, and mid-stream timeout behavior, plus caller cancellation; +* route-capability advertisement for every feature-gate and Agent-mode combination; +* fail-closed behavior for mixed search-capable and incapable primary or secondary storage, while ignoring noop storage; +* searches whose matching labels exist only in out-of-order head data, including ordering and limiting; +* strict client parsing, bounded caches, in-flight deduplication, and retry after failure. + +Manual verification can compare names and values against client-side aggregation of `/api/v1/series` for the same info-metric scope and time window. ### Migration -New endpoint. Both feature flags it depends on are experimental and opt-in. No migration is required. - -### Known unknowns - -* **Default start/end lookback.** Aligned with PROM-74's 1h default via the shared `parseSearchParams` helper. Reasonable, but `target_info` can be sparse on short windows — reviewers may want a longer default specific to `/info_labels`. -* **Whether to accept `match[]` in addition to `expr`.** The current design rejects it on cohesion grounds; reviewers may push back. -* **Whether identifying labels should be configurable per request.** Today the set is `{"job", "instance"}` from `infohelper.DefaultIdentifyingLabels`. Per-request override is plausible if downstream implementations have different identifier conventions. -* **One combined `info-labels-api` flag.** Rejected here on grounds of flag proliferation, but worth re-litigating if reviewers prefer a single-flag operator UX. -* **First-class name-only and exact-name requests.** The current contract can bound name discovery with `values_limit=1` and emulate exact lookup with strict search parameters plus client-side verification, but dedicated semantics would be clearer and avoid returning an unused value during name completion. -* **A bounded default for values.** Leaving `values_limit` unset preserves all values but leaves response size and extractor memory unbounded in that dimension. A future revision may prefer a server default or operator cap. +These are new, experimental, opt-in endpoints. No stored-data migration is required. Implementations and PoC clients should use the two-endpoint, repeated-full-matcher contract. ## Alternatives -### 1. Extend `/api/v1/search/label_names` to optionally return values per name +### 1. One combined `{name, values[]}` endpoint + +Rejected. Name and value completion are separate user interactions with different search targets and cardinalities. A combined response either eagerly downloads unused values or requires a second `values_limit` dimension. It also turns exact label selection into a fuzzy-name-search workaround and does not map cleanly to PROM-74's `SearchLabelNames` and `SearchLabelValues` interfaces. -Would collapse two endpoints into one. Rejected on cohesion grounds: PROM-74 keeps names and values in separate endpoints precisely so that each endpoint's response shape stays simple. The values payload is only useful when the names are scoped to info metrics, so the coupling does not belong on the general search endpoint. +If measured client latency later justifies removing the second round trip, an optional `include_values` extension on the name endpoint could be proposed separately. The split endpoints remain the canonical bounded operations, and such an extension would need an explicit independent value bound. -### 2. Add per-name expansion via a follow-up `/api/v1/search/label_values` call +### 2. Reuse `/api/v1/search/label_names` and `/api/v1/search/label_values` -An interactive client can defer this call until the user selects a label, making the normal path two requests rather than an eager 1 + N requests. This is viable, but the search endpoints cannot apply the `expr`-derived info-metric scope, so the client cannot keep name and value results consistently scoped for arbitrary PromQL. The dedicated endpoint uses the same scoping contract for both phases. +These endpoints have the right split but cannot derive the info-series scope from an arbitrary PromQL expression. Adding `expr` and info-specific identifying-label behavior to the general endpoints would mix function-specific semantics into otherwise general label search. -### 3. Pure client-side composition (`/api/v1/labels` + `/api/v1/label/{name}/values` + filter) +### 3. Pure client-side composition -Feasible today but slow at scale and does not address the info-metric-specific scoping. In particular it cannot evaluate a PromQL expression server-side to harvest identifying labels, so the autocomplete cannot be context-aware against arbitrary PromQL. +Combining labels, per-label values, or series endpoints is possible but transfers substantially more data, may require N+1 requests, and cannot evaluate arbitrary PromQL server-side to keep both completion phases consistently scoped. -### 4. Subsume into a hypothetical `/api/v1/info()` evaluation endpoint +### 4. Use `match[]` instead of `expr` -Out of scope. `info()` is a PromQL function that already uses the existing `/api/v1/query` flow. `/api/v1/info_labels` is a metadata endpoint that supports building the queries; it is not a query endpoint itself. +`match[]` handles selectors only. Autocomplete must work for arbitrary expressions users type, including functions and operators. -### 5. Use `match[]` instead of `expr` +### 5. Add an `/api/v1/info()` query endpoint -`match[]` can only express series selectors, not arbitrary PromQL like `rate(http_requests_total[5m])`. Autocomplete must work against expressions users actually type, not the subset that fits inside a series selector. +Out of scope. `info()` already uses the normal PromQL query endpoints. This proposal is metadata discovery for building that query. ## Action Plan -* [X] Build a working implementation (working branch `arve/info-autocomplete` in `aknuds1/prometheus`). -* [X] Add the dual feature-gate check to the implementation (search-api + promql-experimental-functions). -* [X] Open the [WIP upstream implementation PR](https://github.com/prometheus/prometheus/pull/17930). +* [X] Build a working Prometheus implementation. +* [X] Validate the contract in the Prometheus UI client. +* [X] Validate independent integration in the Grafana Prometheus datasource. +* [X] Split name and exact-label value discovery into separate endpoints based on client feedback. * [ ] Finalize this proposal based on community feedback. -* [ ] Update `docs/querying/api.md` (the implementation already includes this; the action item is post-merge alignment). +* [ ] Merge the implementation after proposal acceptance.