Skip to content

[Incident] False Positive Slow Response Alert — 5th Recurrence (2026-07-15) #7

Description

@DarrenJohns

Summary

The alert-slow-response-sre-lab alert fired at 2026-07-15T02:52:47Z on App Insights resource appi-45ne7zvrqsjxa. Investigation confirmed this is a false positive — the 5th recurrence of the same known issue. The Grubify API is healthy (HTTP 200, 93ms response time). The alert was triggered by ARM management API calls timing out at ~100 seconds, which inflated the App Insights requests/duration average above the 3000ms threshold.

Root Cause: The .NET Grubify API is missing the Application Insights SDK (Microsoft.ApplicationInsights.AspNetCore package and builder.Services.AddApplicationInsightsTelemetry() call in Program.cs). As a result, no application-level requests telemetry flows to App Insights. Only ARM management API calls (from the Azure platform polling management.azure.com) are recorded. When these ARM calls time out, they push the average duration above the alert threshold.

Impact

  • User Impact: None. The API was healthy throughout. No Grubify users were affected.
  • Service Impact: None. This is a monitoring false positive only.
  • Alert Noise: This is the 5th false positive from this alert rule, contributing to alert fatigue.

Timeline

Time (UTC) Event
02:50:17Z ARM AlertsManagement GET timed out at 99,998ms
02:52:03Z ARM AlertsManagement GET slow at 46,179ms
02:52:47Z alert-slow-response-sre-lab alert fired (Sev3)
02:53:43Z ARM AlertsManagement GET timed out at 100,000ms
02:55:00Z SRE Agent investigation started
02:56:00Z False positive confirmed — API healthy (HTTP 200, 93ms)

Evidence

App Insights Request Analysis (last 30 min)

  • Total requests: 126 — ALL are ARM management API calls (management.azure.com)
  • Grubify application requests (/api/*): ZERO (ZERO_ROWS_RETURNED)
  • Requests > 3000ms: 3 ARM AlertsManagement GET calls:
    • 100,000.38ms at 02:53:43Z (timeout, resultCode: unknown)
    • 99,998.07ms at 02:50:17Z (timeout, resultCode: unknown)
    • 46,179.85ms at 02:52:03Z (slow, resultCode: 200)

Container App Resource Metrics (02:49–02:53 UTC)

Metric Values Status
CPU (UsageNanoCores) 244K–351K (~27–35% of 1 vCPU) ✅ Healthy
Memory (WorkingSetBytes) ~53MB ✅ Healthy (well under 400MB)

Container App Configuration

Setting Value Status
CPU 0.5 cores ✅ Normal
Memory 1Gi ✅ Normal
Image acrcagrubify45ne7zvrqsjxa.azurecr.io/grubify-api:latest ✅ Correct
SIMULATE_SLOW Not present ✅ Not injected
APPLICATIONINSIGHTS_CONNECTION_STRING Present ⚠️ SDK not integrated

Live Endpoint Test

curl /api/restaurants → HTTP 200 in 0.093s

KQL Queries Used

// Confirm zero application requests
requests | where timestamp > ago(1h)
| where cloud_RoleName has "grubify" or name has "/api/"
→ ZERO_ROWS_RETURNED

// Find requests > 3000ms (alert threshold)
requests | where timestamp > ago(1h) | where duration > 3000
→ 3 ARM AlertsManagement GET calls (100,000ms, 99,998ms, 46,179ms)

Root Cause

Confirmed (recurring): The .NET Grubify API has APPLICATIONINSIGHTS_CONNECTION_STRING set as an environment variable but does not have the Application Insights SDK integrated:

  1. GrubifyApi.csproj is missing the Microsoft.ApplicationInsights.AspNetCore NuGet package
  2. Program.cs does not call builder.Services.AddApplicationInsightsTelemetry()

Without the SDK, no application-level telemetry (requests, dependencies, traces, exceptions) flows to App Insights. The only requests recorded are ARM management API calls made by the Azure platform. When these ARM calls time out (~100 seconds), they push the average requests/duration metric above the 3000ms alert threshold, triggering a false positive.

Remediation

No immediate remediation needed — the API is healthy and no users are affected.

Permanent Fixes Required (same as issues #3, #4, #5, #6):

  1. Integrate App Insights SDK (Priority: High)

    • Add Microsoft.ApplicationInsights.AspNetCore to GrubifyApi.csproj
    • Add builder.Services.AddApplicationInsightsTelemetry() to Program.cs
    • Redeploy with azd up
  2. Fix alert rule to exclude ARM calls (Priority: High)

    • Add a dimension filter to alert-slow-response-sre-lab:
      • Filter cloud/roleName contains grubify, OR
      • Filter request name starts with /api/
    • This would eliminate false positives even without the SDK fix

Action Items

References


This issue was created by sre-agent-45ne7zvrqsjxa--fcd9117a
Tracked by the SRE agent here

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions