Skip to content

[Incident] Recurring False Positive Slow Response Alert — App Insights SDK Missing from Application Code #4

Description

@DarrenJohns

Summary

Alert alert-slow-response-sre-lab (Sev3) fired at 2026-06-24T00:49:33Z on App Insights resource appi-45ne7zvrqsjxa, indicating Grubify API response times exceeded 3 seconds. Investigation determined this was a false positive — the Grubify API was healthy throughout (34–44ms response times). The alert was triggered by slow Azure Resource Manager (ARM) management API calls tracked in the same App Insights instance.

This is a recurrence of Issue #3 (2026-06-11), but with an important difference: the APPLICATIONINSIGHTS_CONNECTION_STRING env var is present on the container app this time. The root cause is that the .NET application itself does not include the Application Insights SDK NuGet package (Microsoft.ApplicationInsights.AspNetCore), so the connection string is ignored and no application telemetry flows.

Impact

  • User impact: None — the Grubify API was functioning normally with response times of 34–44ms
  • Alert noise: One false positive Sev3 alert, consuming SRE investigation time (~5 minutes)
  • Monitoring gap: Application telemetry (requests, exceptions, dependencies) is not flowing to App Insights — real application issues would not be detected by App Insights-based alerts

Timeline (UTC)

Time Event
2026-06-16T01:15:39Z Current revision 0000013 created (last deployment)
2026-06-24T00:47:00Z ARM AlertsManagement API call timed out at 99,997ms (~100s) — tracked in App Insights
2026-06-24T00:49:33Z Alert fired: alert-slow-response-sre-lab — requests/duration avg > 3000ms over 5 min
2026-06-24T00:50:00Z Second ARM API call at 44,301ms (~44s)
2026-06-24T00:50:30Z SRE Agent investigation started
2026-06-24T00:52:00Z Root cause confirmed: Zero Grubify app telemetry in App Insights; all requests are ARM API calls; API healthy at 34–44ms via direct curl

Evidence

App Insights Requests — ALL ARM API Calls, Zero Grubify Telemetry

-- Confirm no Grubify app requests in App Insights (last 1 hour)
requests
| where timestamp > ago(1h)
| where cloud_RoleName has "grubify" or name has "/api/"
| summarize count() by name, resultCode
-- Result: ZERO_ROWS_RETURNED

ARM API Calls That Triggered the Alert

Timestamp (UTC) Request Duration (ms)
2026-06-24T00:47:00Z GET AlertsManagement/alerts 99,997
2026-06-24T00:50:00Z GET AlertsManagement/alerts 44,301

These are ARM management API calls from the SRE agent's monitoring loop — NOT Grubify application requests. A single ~100s timeout at 00:47 pushed the per-minute average to 5,744ms, breaching the 3000ms threshold.

Per-Minute Average Duration (App Insights requests — all ARM calls)

Time (UTC) Avg Duration (ms) Requests Slow (>3s)
00:44 308 6 0
00:45 338 3 0
00:46 538 1 0
00:47 5,744 18 1
00:48 226 6 0
00:49 386 4 0
00:50 11,386 4 1
00:51 206 17 0

Grubify API Health (Direct Endpoint Testing)

/api/restaurants:  HTTP 200 in 0.037s
/api/fooditems:    HTTP 200 in 0.034s
/weatherforecast:  HTTP 200 in 0.041s

All endpoints healthy, well under the 500ms baseline.

Container App Configuration (at time of alert)

  • Revision: 0000013 (created 2026-06-16T01:15:39Z)
  • CPU: 0.5 cores, Memory: 1Gi — normal
  • Scale: 1 replica running
  • Image: acrcagrubify45ne7zvrqsjxa.azurecr.io/grubify-api:latest
  • APPLICATIONINSIGHTS_CONNECTION_STRING: ✅ Present (set after Issue [Incident] False Positive Slow Response Alert — Missing App Insights Instrumentation #3)
  • App Insights SDK in code: ❌ Missing — GrubifyApi.csproj does not reference Microsoft.ApplicationInsights.AspNetCore, and Program.cs does not call builder.Services.AddApplicationInsightsTelemetry()

Why Telemetry Isn't Flowing Despite Env Var Being Set

The .NET app needs two things for App Insights to work:

  1. ✅ APPLICATIONINSIGHTS_CONNECTION_STRING env var — present
  2. ❌ Microsoft.ApplicationInsights.AspNetCore NuGet package + SDK initialization in Program.cs — missing

Current GrubifyApi.csproj only has:

<PackageReference Include="Microsoft.AspNetCore.OpenApi" Version="9.0.8" />

Current Program.cs has no App Insights initialization:

var builder = WebApplication.CreateBuilder(args);
builder.Services.AddControllers();
builder.Services.AddOpenApi();
// ... no AddApplicationInsightsTelemetry() call

Root Cause

Missing Application Insights SDK in the Grubify API application code.

The APPLICATIONINSIGHTS_CONNECTION_STRING env var was added after Issue #3, but the application code itself does not include the App Insights NuGet package or SDK initialization. The .NET runtime does not auto-instrument without the SDK, so the connection string is silently ignored.

With zero application telemetry, the only requests data in App Insights comes from ARM management API calls. When one of these calls timed out at ~100s, it pushed the average requests/duration above the 3000ms alert threshold, causing a false positive.

Remediation

No immediate remediation needed — the API is healthy. This is an alert fidelity issue, not a service issue.

Required Code Changes (P1)

  1. Add NuGet package to GrubifyApi.csproj:
<PackageReference Include="Microsoft.ApplicationInsights.AspNetCore" Version="2.22.0" />
  1. Add SDK initialization to Program.cs:
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddApplicationInsightsTelemetry(); // Add this line
builder.Services.AddControllers();
  1. Rebuild and deploy the container image.

Alert Rule Improvement (P2)

Add a dimension filter to alert-slow-response-sre-lab to exclude ARM API calls:

  • Filter on cloud/roleName contains grubify
  • OR filter name starts with /api/
  • This prevents ARM management API latency from triggering application-level alerts

Action Items

  • P1: Add Microsoft.ApplicationInsights.AspNetCore NuGet package to GrubifyApi.csproj
  • P1: Add builder.Services.AddApplicationInsightsTelemetry() to Program.cs
  • P1: Rebuild and deploy container image with App Insights SDK
  • P2: Add dimension filter to alert-slow-response-sre-lab to exclude ARM API calls (filter on cloud/roleName or request name)
  • P3: Verify telemetry is flowing after code fix (check requests table for /api/* entries)
  • P3: Close Issue [Incident] False Positive Slow Response Alert — Missing App Insights Instrumentation #3 after code fix is deployed (env var alone is insufficient)

References

  • Alert ID: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourcegroups/rg-sre-lab/providers/microsoft.insights/components/appi-45ne7zvrqsjxa/providers/Microsoft.AlertsManagement/alerts/7414993e-96ff-483b-8761-fafcf99ff000
  • Alert Rule: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourcegroups/rg-sre-lab/providers/Microsoft.Insights/metricAlerts/alert-slow-response-sre-lab
  • Container App: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourceGroups/rg-sre-lab/providers/Microsoft.App/containerapps/ca-grubify-45ne7zvrqsjxa
  • App Insights: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourceGroups/rg-sre-lab/providers/Microsoft.Insights/components/appi-45ne7zvrqsjxa
  • Log Analytics Workspace ID: 24fa3c3f-c6cb-4be9-ba43-5f09562e869d
  • Azure Portal Alert: View Alert
  • Previous Incident: Issue #3 — False Positive Slow Response Alert (2026-06-11)

Created by Azure SRE Agent: Open thread on Saw | Open thread on MSFT

This issue was created by sre-agent-45ne7zvrqsjxa--fcd9117a
Tracked by the SRE agent here

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions