Skip to content

[Incident Report] Container Restart Loop — Port Mismatch on Revision --etqi6yl (2026-06-09) #1

Description

@DarrenJohns

Summary

On 2026-06-09, the Grubify API container app (ca-grubify-45ne7zvrqsjxa) entered a crash-loop restart cycle. The container accumulated 11 restarts between 02:22–03:17 UTC. Azure Monitor alert alert-restarts-sre-lab fired at 02:39:45 UTC (Sev2). The root cause was a port mismatch: the deployed container image on revision --etqi6yl was listening on port 80, while the Container App ingress targetPort was configured for port 8080. This caused health probe failures, leading to repeated container kills and restarts.

The incident was resolved at 04:05 UTC when Darren Johns deployed a new revision (--0000002) with the correct .NET 9 application image listening on port 8080.


Impact

Metric Value
Duration ~1 hour 43 minutes (02:22 – 04:05 UTC)
Severity Sev2
User-facing impact API unavailable/degraded during restart loop; frontend unable to reach backend
Error rate 0% captured in App Insights (requests failed at infrastructure level before reaching app)
Total restarts 11

Timeline

Time (UTC) Event
02:16 Revision --etqi6yl starts, logs show Listening on :80... (incorrect port)
02:22 First restart detected (RestartCount = 1)
02:25 SRE Agent begins automated resource discovery (74 management API calls)
02:27 RestartCount = 2
02:32 RestartCount = 3
02:37 RestartCount = 4
02:39:45 Alert alert-restarts-sre-lab FIRED (threshold: RestartCount > 3 in 5 min)
02:42 RestartCount = 6
02:57 RestartCount = 8
03:12 RestartCount = 11 (peak)
03:17 Last restart on revision --etqi6yl
03:20 Gap in traffic — app mostly unreachable
03:56 Brief attempt with revision --omathd2, still listening on :80
04:04 Revision --0000002 deployed — .NET 9 app starts on port 8080
04:05 Last modification by Darren.Johns@microsoft.com
04:30+ Traffic normalizes, 0% error rate, stable operation

Evidence

RestartCount Metrics (02:22–03:17 UTC)

02:22 → 1 restart
02:27 → 2 restarts
02:32 → 3 restarts
02:37 → 4 restarts
02:42 → 6 restarts
02:47 → 7 restarts
02:52 → 7 restarts
02:57 → 8 restarts
03:02 → 9 restarts
03:07 → 9 restarts
03:12 → 11 restarts (peak)
03:17 → 11 restarts
03:57 → 1 restart (counter reset — new revision)

Container Console Logs (Log Analytics)

  • Revision --etqi6yl logged Listening on :80... every ~4 minutes (restart cycle)
  • Revision --0000002 logged successful ASP.NET Core startup on http://[::]:8080

Current Health (post-recovery)

  • Provisioning State: Succeeded
  • Running Status: Running
  • Replicas: 1/1 ready
  • CPU: ~0% (idle)
  • Memory: ~5%
  • Error rate: 0%
  • Image: acrcagrubify45ne7zvrqsjxa.azurecr.io/grubify-api:latest

Root Cause

Port mismatch between container application and Container App ingress configuration.

The container image deployed on revision --etqi6yl was running an application that bound to port 80. However, the Container App ingress was configured with targetPort: 8080. This mismatch caused:

  1. Container starts and binds to port 80
  2. Container Apps health probes target port 8080 — no listener found
  3. Health probes fail repeatedly
  4. Container is killed after probe failure threshold
  5. Container restarts, cycle repeats (~4-minute intervals matching probe timeout + back-off)

The log format (2026/06/09 02:16:24 Listening on :80...) indicates the revision was running a different application (likely a Node.js or Go image) rather than the intended .NET 9 Grubify API.


Remediation

  • Immediate fix: Darren Johns deployed revision --0000002 at 04:05 UTC with the correct .NET 9 container image (grubify-api:latest) listening on port 8080.
  • Alert closed: Azure Monitor alert 539ca273-1d38-4938-85ef-844b956bf000 closed by SRE Agent after recovery verification.

Action Items

  • CI/CD guard: Add a deployment validation step that verifies the container's EXPOSE port matches the Container App targetPort before promoting to production.
  • Health probe monitoring: Add a startup probe with a shorter failure threshold to detect port mismatches faster (currently takes ~4 min per cycle).
  • Image tag governance: Investigate why revision --etqi6yl received an image listening on port 80. Consider pinning image tags instead of using latest to prevent unexpected image changes.
  • Runbook update: Document port mismatch as a root cause pattern in the crash-recovery runbook — check Listening on :XX in logs vs targetPort in ingress config.
  • Rollback automation: Configure Container Apps revision rollback policy to automatically revert to last healthy revision when restart count exceeds threshold.

References

  • Container App Resource ID: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourceGroups/rg-sre-lab/providers/Microsoft.App/containerapps/ca-grubify-45ne7zvrqsjxa
  • App Insights Resource ID: /subscriptions/573c01a0-ef63-4535-a78e-0bc7f79c87c9/resourceGroups/rg-sre-lab/providers/Microsoft.Insights/components/appi-45ne7zvrqsjxa
  • Log Analytics Workspace ID: 24fa3c3f-c6cb-4be9-ba43-5f09562e869d
  • Alert ID: 539ca273-1d38-4938-85ef-844b956bf000
  • Alert Rule: alert-restarts-sre-lab
  • Azure Portal: View Alert

This issue was created by sre-agent-45ne7zvrqsjxa--fcd9117a
Tracked by the SRE agent here

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions