Skip to content

[Incident] Sev1: PostgreSQL Server Manually Stopped — zava-api Total Database Outage (2026-06-22 18:56–19:03 UTC) #3

Description

@wilkinshum

Summary

PostgreSQL Flexible Server zava-pg-auf6gw was manually stopped by admin@MngEnvMCAP602713.onmicrosoft.com at 18:56:39 UTC on 2026-06-22, causing a complete database outage for the Zava food delivery API (zava-api). All PostgreSQL connection attempts from AKS pods timed out at the 5-second threshold, resulting in health check failures, product query failures, and total service unavailability for approximately 6 minutes.

The server was restarted by a service principal (48f51566-903d-4fc5-b28b-13d8b3340c7c) at 19:01:51 UTC and became ready at 19:02:58 UTC. Service fully recovered at 19:03 UTC.

Alert: postgres-network-blocked (Sev1) fired at 18:59:38 UTC.

Impact

Metric Value
Duration ~6 minutes (18:58-19:03 UTC)
Failed PostgreSQL connections ~1,466 (343+353+341+337+92)
Container log timeout errors ~1,529
Affected endpoints /api/products, /api/products/category/*, health checks
Affected containers 2 pods
Service availability 0% during outage window
User impact Complete inability to browse products or place orders

Pre-incident baseline (18:33-18:57 UTC)

  • ~660 PostgreSQL calls/min, 0 failures
  • Average connection duration: ~1ms
  • All health checks passing

During incident (18:58-19:02 UTC)

  • 100% PostgreSQL connection failure rate
  • All connections timing out at exactly 5000ms
  • 360 timeout errors/min in container logs
  • Health checks: "DB unreachable"

Timeline

Time (UTC) Event
18:56:39 admin@MngEnvMCAP602713.onmicrosoft.com initiates PostgreSQL stop action
18:58:00 First connection timeouts appear (92/580 calls fail)
18:59:00 100% failure rate - all 343 PostgreSQL calls timeout at 5000ms
18:59:38 Alert postgres-network-blocked (Sev1) fires
18:59:43 PostgreSQL stop action completes (status: Succeeded)
19:01:51 Service principal 48f51566-903d-4fc5-b28b-13d8b3340c7c initiates server start
19:02:58 PostgreSQL logs: "database system is ready to accept connections"
19:03:00 Service fully recovered - 532 calls, 0 failures, ~1.3ms avg latency

Evidence

1. Activity Log - Stop Action

Caller: admin@MngEnvMCAP602713.onmicrosoft.com
Operation: Microsoft.DBforPostgreSQL/flexibleServers/stop/action
Status: Succeeded
Time: 2026-06-22T18:59:43Z (initiated at 18:56:39Z)

2. Activity Log - Start Action

Caller: 48f51566-903d-4fc5-b28b-13d8b3340c7c (service principal)
Operation: Microsoft.DBforPostgreSQL/flexibleServers/start/action
Status: Accepted
Time: 2026-06-22T19:01:51Z

3. Container Logs (sample)

{"level":"error","message":"Health check: DB unreachable","timestamp":"2026-06-22T19:01:35.147Z","service":"zava-api","error":"timeout exceeded when trying to connect"}
{"level":"error","message":"Failed to fetch products","timestamp":"2026-06-22T19:01:35.148Z","service":"zava-api","error":"timeout exceeded when trying to connect","endpoint":"/api/products"}
{"level":"error","message":"Failed to fetch products by category","timestamp":"2026-06-22T19:01:35.149Z","service":"zava-api","error":"timeout exceeded when trying to connect","category":"__probe"}

4. AppDependencies - Connection Duration Spike

  • Pre-incident: avg ~1ms, 0 failures
  • During incident: avg ~5001ms (5s timeout), 100% failure
  • Post-recovery: avg ~1.3ms, 0 failures

5. PostgreSQL Diagnostics Log

2026-06-22 19:02:58 UTC - "database system is ready to accept connections"

6. NSG State (at time of investigation)

  • nsg-aks-auf6gw on AKS subnet: allow-all-outbound (Allow, priority 1000) - outbound path was clear
  • allow-web-inbound rule changed from Allow to Deny (unrelated security fix)
  • NSG was NOT the cause - outbound VNet traffic was permitted throughout

7. Resource State at Investigation Time

Resource State
PostgreSQL zava-pg-auf6gw Ready (recovered)
AKS aks-Zava-auf6gw Succeeded / Running
Network path (AKS to PG) Open via VNet integration

Root Cause

Manual operator action. User admin@MngEnvMCAP602713.onmicrosoft.com executed a stop action on the PostgreSQL Flexible Server zava-pg-auf6gw at 18:56:39 UTC via the Azure control plane. This immediately began shutting down the database, causing all application connections to time out. The alert name postgres-network-blocked is misleading - the issue was not a network block but a stopped server.

The server was subsequently restarted by service principal 48f51566-903d-4fc5-b28b-13d8b3340c7c at 19:01:51 UTC, restoring service at 19:02:58 UTC.

Remediation

Immediate (Completed)

  • PostgreSQL server restarted at 19:01:51 UTC by service principal
  • Server ready and accepting connections at 19:02:58 UTC
  • Service fully recovered at 19:03 UTC - 0 failures confirmed

Short-term

  • Investigate why admin@MngEnvMCAP602713.onmicrosoft.com stopped the server - determine if intentional or accidental
  • Rename alert postgres-network-blocked to better reflect all connectivity failure modes (e.g., postgres-connection-failure)
  • Add Azure Policy or RBAC constraint to prevent unauthorized stop/start of production PostgreSQL servers

Long-term

  • Implement Azure Resource Lock (CanNotDelete/ReadOnly) on critical database resources
  • Create postgres-server-stopped activity log alert that fires immediately on stop actions
  • Add circuit breaker pattern in zava-api to fail fast instead of accumulating 5s timeouts
  • Implement connection retry with exponential backoff in the application layer

Action Items

# Action Owner Priority
1 Determine if the stop action was intentional or accidental Ops team P0
2 Add Azure Resource Lock on zava-pg-auf6gw Infra team P1
3 Create activity log alert for PostgreSQL stop/start actions SRE team P1
4 Restrict RBAC for server stop/start to authorized principals only Security team P1
5 Rename alert to postgres-connection-failure SRE team P2
6 Implement circuit breaker in zava-api DB connection layer Dev team P2

References

  • Alert ID: /subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourcegroups/rg-zava-aks-postgres/providers/microsoft.operationalinsights/workspaces/law-zava-auf6gw/providers/Microsoft.AlertsManagement/alerts/570973c0-9944-793b-c390-d2ee03ca0033
  • Alert Rule: /subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/microsoft.insights/scheduledqueryrules/postgres-network-blocked
  • PostgreSQL Server: /subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/Microsoft.DBforPostgreSQL/flexibleServers/zava-pg-auf6gw
  • AKS Cluster: /subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/Microsoft.ContainerService/managedClusters/aks-Zava-auf6gw
  • Log Analytics Workspace: law-Zava-auf6gw (ID: 69f6cc9c-c59f-4624-8b95-e12b93220ece)
  • App Insights: ai-Zava-auf6gw (ARM: /subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/Microsoft.Insights/components/ai-Zava-auf6gw)
  • NSG: /subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/Microsoft.Network/networkSecurityGroups/nsg-aks-auf6gw
  • Subscription: f627598e-05c5-4093-8667-5730c4026ea3
  • Resource Group: rg-zava-aks-postgres
  • Related incidents: Incident: postgres-server-down Alert Fired During Infrastructure Deployment (Zava AKS+PostgreSQL) dm-chelupati/grubify#309 (false positive deployment alert), Incident: Slow Products Category Query — Zava API (PostgreSQL Cold Cache) dm-chelupati/grubify#310 (cold cache slow query)
  • Portal Link: Alert in Azure Portal

This issue was created by sre-agent-jkazytu5kl5py--70975bf6
Tracked by the SRE agent here

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions