You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The api can accept a single GraphQL request it cannot survive. On 2026-08-19 one client sent ~19 requests/min of a query that allocates ~1.5 GB each against a 2 GiB pod limit, killing 5 pods before it was tracked down and fixed in the client's repo. That resolution depended on having write access to another team's code, which doesn't generalize. The API should refuse what it can't serve.
What the measurements rule out
Canary runs (isolated pod, app: api-canary, one request each, cgroup sampled every 200 ms):
root first × nested relations
response
duration
peak RSS delta
50 × 1000
3.05 MB
6.6 s
+152 Mi
100 × 100
5.89 MB
5.8 s
+260 Mi
500 × 100
29.0 MB
18.1 s
+724 Mi
1000 × 1000
62.9 MB
33.3 s
+1,569 Mi
Not a static budget over declared limits.50 × 1000 and 500 × 100 are the same 50k product but cost 152 Mi vs 724 Mi, because declared limits are upper bounds and only realized rows are paid for. A walker over the query document over-rejects cheap queries and still misses expensive ones.
Not the cost ceiling.GRAPHQL_COST_REJECT_THRESHOLD is armed at 500; the killer scored 274, and the 08-06 incident scored 228. A ceiling low enough to catch 274 also catches ~2–3% of ordinary traffic (≤260 is 96.8% of requests). Cost counts nodes; the OOM is driven by bytes hydrated per node.
Not concurrency. See #897/#900 — the cap moved 10→6→10; 6 doubled 503s to 1414/h with no OOM reduction, because this is a single-request kill.
Not a memory-limit bump. 4 of 8 pods share one node with 6.5 Gi allocatable; 3 Gi limits would risk node-level eviction, which is a worse failure mode.
So the only axis the data supports is realized bytes/rows, counted at runtime, with an abort.
Proposed shape
Count as the response materializes and abort past a ceiling, returning a structured error instead of dying. A ~10–15 MB ceiling would have fired on ~78 requests/18 h (~4/h) while bounding peak RSS to roughly 300–750 Mi.
Open design question — where to count:
Per-connection realized rows, accumulated per request, checked as each connection resolves. Aborts before the whole response exists. Note paginationCapPlugin's docstring warns makeWrapResolversPlugin only reliably influences SQL at root level — but counting rows doesn't need to influence SQL, so wrapping should be viable here.
During serialization, which is where the 25–50× amplification actually happens. Catches everything, but by then the object graph is already in memory, so it bounds the string not the graph.
(1) prevents more of the allocation; (2) is closer to the real cost. Possibly both, with (1) as the guard and (2) as the backstop.
Follow the existing Phase 1 / Phase 2 convention from costLoggerPlugin: land the accounting observational with an env-gated ceiling (0 = off), confirm against real traffic what a ceiling would have rejected, then enforce.
Related
perf(api): stop double-serializing the responses that OOM pods #901 removes an amplifier on this path — the response-size metric was calling JSON.stringify a second time on exactly these responses. Its estimateJsonBytes walker is a possible basis for (2), though as written it runs post-hoc in onExecuteDone.
The deeper fix, separate and larger, is cutting the 25–50× factor itself: streaming/chunked serialization rather than building one giant JSON string.
geobrowser/geo-skills#5 fixes the docs that taught clients to write this shape.
Root cause writeup: geo-explorers/postgres_to_geo#36, and their follow-up #37.
The api can accept a single GraphQL request it cannot survive. On 2026-08-19 one client sent ~19 requests/min of a query that allocates ~1.5 GB each against a 2 GiB pod limit, killing 5 pods before it was tracked down and fixed in the client's repo. That resolution depended on having write access to another team's code, which doesn't generalize. The API should refuse what it can't serve.
What the measurements rule out
Canary runs (isolated pod,
app: api-canary, one request each, cgroup sampled every 200 ms):first× nestedrelationsNot a static budget over declared limits.
50 × 1000and500 × 100are the same 50k product but cost 152 Mi vs 724 Mi, because declared limits are upper bounds and only realized rows are paid for. A walker over the query document over-rejects cheap queries and still misses expensive ones.Not the cost ceiling.
GRAPHQL_COST_REJECT_THRESHOLDis armed at 500; the killer scored 274, and the 08-06 incident scored 228. A ceiling low enough to catch 274 also catches ~2–3% of ordinary traffic (≤260is 96.8% of requests). Cost counts nodes; the OOM is driven by bytes hydrated per node.Not concurrency. See #897/#900 — the cap moved 10→6→10; 6 doubled 503s to 1414/h with no OOM reduction, because this is a single-request kill.
Not a memory-limit bump. 4 of 8 pods share one node with 6.5 Gi allocatable; 3 Gi limits would risk node-level eviction, which is a worse failure mode.
So the only axis the data supports is realized bytes/rows, counted at runtime, with an abort.
Proposed shape
Count as the response materializes and abort past a ceiling, returning a structured error instead of dying. A ~10–15 MB ceiling would have fired on ~78 requests/18 h (~4/h) while bounding peak RSS to roughly 300–750 Mi.
Open design question — where to count:
paginationCapPlugin's docstring warnsmakeWrapResolversPluginonly reliably influences SQL at root level — but counting rows doesn't need to influence SQL, so wrapping should be viable here.(1) prevents more of the allocation; (2) is closer to the real cost. Possibly both, with (1) as the guard and (2) as the backstop.
Follow the existing Phase 1 / Phase 2 convention from
costLoggerPlugin: land the accounting observational with an env-gated ceiling (0= off), confirm against real traffic what a ceiling would have rejected, then enforce.Related
JSON.stringifya second time on exactly these responses. ItsestimateJsonByteswalker is a possible basis for (2), though as written it runs post-hoc inonExecuteDone.geobrowser/geo-skills#5fixes the docs that taught clients to write this shape.geo-explorers/postgres_to_geo#36, and their follow-up#37.