You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Active revision:ca-grubify-45ne7zvrqsjxa--0000013 (100% traffic)
Summary
Alert alert-slow-response-sre-lab (Sev3) fired at 2026-06-25T05:52:47Z on App Insights resource appi-45ne7zvrqsjxa, indicating Grubify API response times exceeded 3 seconds. Investigation confirmed this is a false positive — the Grubify API was healthy throughout (30–61ms response times). The alert was triggered by two ARM management API calls that timed out at ~100,000ms, pushing the average requests/duration above the 3000ms threshold.
This is the 3rd occurrence of this pattern, following Issue #3 (2026-06-11) and Issue #4 (2026-06-24). The root cause remains unchanged: the .NET application does not include the Application Insights SDK, so no application telemetry flows — only ARM management API calls are tracked.
Impact
User impact: None — the Grubify API was functioning normally with response times of 30–61ms
Alert noise: One false positive Sev3 alert, consuming SRE investigation time (~5 minutes)
Monitoring gap: Application telemetry (requests, exceptions, dependencies) is still not flowing to App Insights — real application issues would not be detected by App Insights-based alerts
Timeline (UTC)
Time
Event
2026-06-16T01:15:39Z
Current revision 0000013 created (last deployment)
2026-06-25T05:50:22Z
ARM AlertsManagement API call timed out at 100,005ms (~100s) — tracked in App Insights
2026-06-25T05:52:47Z
Alert fired: alert-slow-response-sre-lab — requests/duration avg > 3000ms over 5 min
2026-06-25T05:53:03Z
Second ARM AlertsManagement API call timed out at 99,999ms (~100s)
2026-06-25T05:54:00Z
SRE Agent investigation started
2026-06-25T05:56:00Z
Root cause confirmed: Zero Grubify app telemetry in App Insights; all requests are ARM API calls; API healthy at 30–61ms via direct curl
Evidence
App Insights Requests — ALL ARM API Calls, Zero Grubify Telemetry
-- Confirm no Grubify app requests in App Insights (last 1 hour)
requests
| where timestamp > ago(1h)
| where cloud_RoleName has"grubify"or name has"/api/"
| summarizecount() by name, resultCode
-- Result: ZERO_ROWS_RETURNED
ARM API Calls That Triggered the Alert
Timestamp (UTC)
Request
Duration (ms)
cloud_RoleName
2026-06-25T05:50:22Z
GET AlertsManagement/alerts
100,005
(empty)
2026-06-25T05:53:03Z
GET AlertsManagement/alerts
99,999
(empty)
These are ARM management API calls from the SRE agent's monitoring loop — NOT Grubify application requests.
Grubify API Health (Direct Endpoint Testing)
Request 1: HTTP 200 in 0.062s
Request 2: HTTP 200 in 0.037s
Request 3: HTTP 200 in 0.031s
Request 4: HTTP 200 in 0.037s
All endpoints healthy, well under the 500ms baseline.
App Insights SDK in code: ❌ Missing — GrubifyApi.csproj does not reference Microsoft.ApplicationInsights.AspNetCore, and Program.cs does not call builder.Services.AddApplicationInsightsTelemetry()
Alert Rule Configuration
Rule: alert-slow-response-sre-lab
Metric: requests/duration on microsoft.insights/components
Condition: Average > 3000ms over PT5M window, evaluated every PT1M
Severity: 3
No dimension filter — captures ALL requests including ARM management API calls
Root Cause
Missing Application Insights SDK in the Grubify API application code (unchanged from Issues #3 and #4).
For .NET App Insights telemetry to flow, both conditions must be met:
❌ Microsoft.ApplicationInsights.AspNetCore NuGet package + AddApplicationInsightsTelemetry() in Program.cs — still missing
With zero application telemetry, the only requests data in App Insights comes from ARM management API calls. When two of these calls timed out at ~100s each (05:50Z and 05:53Z), they pushed the average requests/duration above the 3000ms threshold, causing a false positive.
Remediation
No immediate remediation needed — the API is healthy. This is an alert fidelity issue, not a service issue.
Required Code Changes (P1 — OVERDUE from Issue #4)
Incident Report: False Positive Slow Response Alert (3rd Recurrence)
8da7bba0-48e2-46f4-9db3-c77990edf000ca-grubify-45ne7zvrqsjxa(rg:rg-sre-lab)573c01a0-ef63-4535-a78e-0bc7f79c87c9ca-grubify-45ne7zvrqsjxa.graycoast-6fc58ae6.eastus2.azurecontainerapps.ioca-grubify-45ne7zvrqsjxa--0000013(100% traffic)Summary
Alert
alert-slow-response-sre-lab(Sev3) fired at 2026-06-25T05:52:47Z on App Insights resourceappi-45ne7zvrqsjxa, indicating Grubify API response times exceeded 3 seconds. Investigation confirmed this is a false positive — the Grubify API was healthy throughout (30–61ms response times). The alert was triggered by two ARM management API calls that timed out at ~100,000ms, pushing the averagerequests/durationabove the 3000ms threshold.This is the 3rd occurrence of this pattern, following Issue #3 (2026-06-11) and Issue #4 (2026-06-24). The root cause remains unchanged: the .NET application does not include the Application Insights SDK, so no application telemetry flows — only ARM management API calls are tracked.
Impact
Timeline (UTC)
0000013created (last deployment)alert-slow-response-sre-lab—requests/durationavg > 3000ms over 5 minEvidence
App Insights Requests — ALL ARM API Calls, Zero Grubify Telemetry
ARM API Calls That Triggered the Alert
These are ARM management API calls from the SRE agent's monitoring loop — NOT Grubify application requests.
Grubify API Health (Direct Endpoint Testing)
All endpoints healthy, well under the 500ms baseline.
Resource Metrics (Azure Monitor, 05:25–05:55Z)
No CPU or memory pressure whatsoever.
Container App Configuration
0000013(created 2026-06-16T01:15:39Z)acrcagrubify45ne7zvrqsjxa.azurecr.io/grubify-api:latestAPPLICATIONINSIGHTS_CONNECTION_STRING: ✅ PresentSIMULATE_SLOW: Not present (no artificial delay)GrubifyApi.csprojdoes not referenceMicrosoft.ApplicationInsights.AspNetCore, andProgram.csdoes not callbuilder.Services.AddApplicationInsightsTelemetry()Alert Rule Configuration
alert-slow-response-sre-labrequests/durationonmicrosoft.insights/componentsRoot Cause
Missing Application Insights SDK in the Grubify API application code (unchanged from Issues #3 and #4).
For .NET App Insights telemetry to flow, both conditions must be met:
APPLICATIONINSIGHTS_CONNECTION_STRINGenv var — present (fixed after Issue [Incident] False Positive Slow Response Alert — Missing App Insights Instrumentation #3)Microsoft.ApplicationInsights.AspNetCoreNuGet package +AddApplicationInsightsTelemetry()inProgram.cs— still missingWith zero application telemetry, the only
requestsdata in App Insights comes from ARM management API calls. When two of these calls timed out at ~100s each (05:50Z and 05:53Z), they pushed the averagerequests/durationabove the 3000ms threshold, causing a false positive.Remediation
No immediate remediation needed — the API is healthy. This is an alert fidelity issue, not a service issue.
Required Code Changes (P1 — OVERDUE from Issue #4)
GrubifyApi.csproj:Program.cs:Alert Rule Improvement (P2 — OVERDUE from Issue #3)
Add a dimension filter to
alert-slow-response-sre-lab:cloud/roleNamecontainsgrubifynamestarts with/api/Action Items
Microsoft.ApplicationInsights.AspNetCoreNuGet package toGrubifyApi.csprojbuilder.Services.AddApplicationInsightsTelemetry()toProgram.csalert-slow-response-sre-labRecurrence History
References
/subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourcegroups/rg-sre-lab/providers/microsoft.insights/components/appi-45ne7zvrqsjxa/providers/Microsoft.AlertsManagement/alerts/8da7bba0-48e2-46f4-9db3-c77990edf000/subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourcegroups/rg-sre-lab/providers/Microsoft.Insights/metricAlerts/alert-slow-response-sre-lab/subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourceGroups/rg-sre-lab/providers/Microsoft.App/containerapps/ca-grubify-45ne7zvrqsjxa/subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourceGroups/rg-sre-lab/providers/Microsoft.Insights/components/appi-45ne7zvrqsjxa24fa3c3f-c6cb-4be9-ba43-5f09562e869dCreated by Azure SRE Agent: Open thread on Saw | Open thread on MSFT
This issue was created by sre-agent-45ne7zvrqsjxa--fcd9117a
Tracked by the SRE agent here