Skip to content

[Incident] False Positive Slow Response Alert — Missing App Insights Instrumentation #3

Description

@DarrenJohns

Summary

Alert alert-slow-response-sre-lab (Sev3) fired at 2026-06-11T08:44:26Z on App Insights resource appi-45ne7zvrqsjxa, indicating Grubify API response times exceeded 3 seconds. Investigation determined this was a false positive — the Grubify API was healthy throughout (33–55ms response times). The alert was triggered by slow Azure Resource Manager (ARM) management API calls tracked in the same App Insights instance, due to missing application-level instrumentation.

Impact

  • User impact: None — the Grubify API was functioning normally with response times of 33–55ms
  • Alert noise: One false positive Sev3 alert, consuming SRE investigation time (~10 minutes)
  • Monitoring gap: Application telemetry (requests, exceptions, dependencies) was not flowing to App Insights, meaning real application issues would not have been detected by this alert

Timeline (UTC)

Time Event
2026-06-11T02:23:23Z Revision 0000010 created (last deployment before incident)
2026-06-11T08:42:22Z ARM AlertsManagement API call timed out at 99,999ms — tracked in App Insights
2026-06-11T08:44:26Z Alert fired: alert-slow-response-sre-lab — requests/duration avg > 3000ms over 5 min
2026-06-11T08:46:34Z Second slow ARM API call at 5,939ms
2026-06-11T08:47:00Z SRE Agent investigation started
2026-06-11T08:49:44Z Remediation applied: Added APPLICATIONINSIGHTS_CONNECTION_STRING env var, revision 0000011 deployed
2026-06-11T08:50:00Z Verified API healthy (33–109ms), new revision running

Evidence

Alert Rule Configuration

  • Rule: alert-slow-response-sre-lab
  • Metric: requests/duration on microsoft.insights/components
  • Condition: Average > 3000ms over 5-minute window, evaluated every 1 minute
  • Scope: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourceGroups/rg-sre-lab/providers/Microsoft.Insights/components/appi-45ne7zvrqsjxa

Requests That Triggered the Alert (App Insights)

| Timestamp                  | Request (ARM API)           | Duration (ms) |
|----------------------------|-----------------------------|---------------|
| 2026-06-11T08:42:22Z       | GET AlertsManagement/alerts | 99,999        |
| 2026-06-11T08:46:34Z       | GET alert details           | 5,939         |

These are ARM management API calls from the SRE agent's own monitoring loop, NOT Grubify application requests.

Grubify API Health (Direct Testing)

Request 1: HTTP 200 in 0.055s
Request 2: HTTP 200 in 0.039s
Request 3: HTTP 200 in 0.033s
Request 4: HTTP 200 in 0.034s
Request 5: HTTP 200 in 0.033s

Container App Configuration (at time of alert)

  • Revision: 0000010 (created 2026-06-11T02:23:23Z)
  • CPU: 0.5 cores, Memory: 1Gi — normal
  • Scale: 1–5 replicas (1 running)
  • Env vars: ASPNETCORE_URLS, ASPNETCORE_ENVIRONMENT, AllowedOrigins__0
  • Missing: APPLICATIONINSIGHTS_CONNECTION_STRING — no application telemetry flowing

KQL Queries Used

-- Confirm no Grubify app requests in App Insights
requests
| where timestamp > ago(1h)
| where cloud_RoleName has "grubify" or name has "/api/"
| summarize count() by name, resultCode
-- Result: ZERO_ROWS_RETURNED

-- Find requests that exceeded threshold
requests
| where timestamp > ago(1h)
| where duration > 3000
| project timestamp, name, duration, resultCode
-- Result: 2 ARM API calls (99999ms, 5939ms)

Root Cause

Missing Application Insights instrumentation on the Grubify container app.

The container app (ca-grubify-45ne7zvrqsjxa) did not have the APPLICATIONINSIGHTS_CONNECTION_STRING environment variable configured. As a result:

  1. No application-level telemetry (HTTP requests, exceptions, dependencies) was being sent to App Insights
  2. The only requests data in App Insights came from ARM management API calls (made by the SRE agent's monitoring infrastructure)
  3. When an ARM API call timed out at ~100 seconds, it pushed the average requests/duration metric above the 3000ms threshold
  4. The alert fired as a false positive

Remediation

Action taken: Added APPLICATIONINSIGHTS_CONNECTION_STRING environment variable to ca-grubify-45ne7zvrqsjxa with the connection string for appi-45ne7zvrqsjxa.

az containerapp update \
  --name ca-grubify-45ne7zvrqsjxa \
  --resource-group rg-sre-lab \
  --subscription 573c01a0-ef63-4535-a78e-0bc7f79c87c9 \
  --set-env-vars "APPLICATIONINSIGHTS_CONNECTION_STRING=InstrumentationKey=953e936a-..."

New revision 0000011 created and verified running successfully.

Rollback command (if needed):

az containerapp update \
  --name ca-grubify-45ne7zvrqsjxa \
  --resource-group rg-sre-lab \
  --subscription 573c01a0-ef63-4535-a78e-0bc7f79c87c9 \
  --remove-env-vars APPLICATIONINSIGHTS_CONNECTION_STRING

Action Items

  • P1: Verify application telemetry is flowing to App Insights after revision 0000011 deploys (check requests table for Grubify API endpoints within 30 minutes)
  • P2: Update Bicep IaC templates (infra/core/host/container-app.bicep) to include APPLICATIONINSIGHTS_CONNECTION_STRING so future azd up deployments don't regress this fix
  • P3: Consider adding a dimension filter to the alert rule (e.g., filter on cloud/roleName = grubify-api) to prevent ARM API calls from triggering application-level alerts
  • P3: Audit the frontend container app (ca-grubify-fe-45ne7zvrqsjxa) for the same missing instrumentation

References

  • Alert ID: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourcegroups/rg-sre-lab/providers/microsoft.insights/components/appi-45ne7zvrqsjxa/providers/Microsoft.AlertsManagement/alerts/a9532934-b998-4941-a3ac-5ef85618f000
  • Alert Rule: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourcegroups/rg-sre-lab/providers/Microsoft.Insights/metricAlerts/alert-slow-response-sre-lab
  • Container App: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourceGroups/rg-sre-lab/providers/Microsoft.App/containerapps/ca-grubify-45ne7zvrqsjxa
  • App Insights: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourceGroups/rg-sre-lab/providers/Microsoft.Insights/components/appi-45ne7zvrqsjxa
  • Log Analytics Workspace ID: 24fa3c3f-c6cb-4be9-ba43-5f09562e869d
  • Azure Portal Alert: View Alert

Created by Azure SRE Agent: Open thread on Saw | Open thread on MSFT

This issue was created by sre-agent-45ne7zvrqsjxa--fcd9117a
Tracked by the SRE agent here

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions