You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[Incident] Sev1: PostgreSQL Server Manually Stopped — zava-api Total Database Outage (2026-06-22 18:56–19:03 UTC) #3
PostgreSQL Flexible Server zava-pg-auf6gw was manually stopped by admin@MngEnvMCAP602713.onmicrosoft.com at 18:56:39 UTC on 2026-06-22, causing a complete database outage for the Zava food delivery API (zava-api). All PostgreSQL connection attempts from AKS pods timed out at the 5-second threshold, resulting in health check failures, product query failures, and total service unavailability for approximately 6 minutes.
The server was restarted by a service principal (48f51566-903d-4fc5-b28b-13d8b3340c7c) at 19:01:51 UTC and became ready at 19:02:58 UTC. Service fully recovered at 19:03 UTC.
Alert:postgres-network-blocked (Sev1) fired at 18:59:38 UTC.
Impact
Metric
Value
Duration
~6 minutes (18:58-19:03 UTC)
Failed PostgreSQL connections
~1,466 (343+353+341+337+92)
Container log timeout errors
~1,529
Affected endpoints
/api/products, /api/products/category/*, health checks
Affected containers
2 pods
Service availability
0% during outage window
User impact
Complete inability to browse products or place orders
{"level":"error","message":"Health check: DB unreachable","timestamp":"2026-06-22T19:01:35.147Z","service":"zava-api","error":"timeout exceeded when trying to connect"}
{"level":"error","message":"Failed to fetch products","timestamp":"2026-06-22T19:01:35.148Z","service":"zava-api","error":"timeout exceeded when trying to connect","endpoint":"/api/products"}
{"level":"error","message":"Failed to fetch products by category","timestamp":"2026-06-22T19:01:35.149Z","service":"zava-api","error":"timeout exceeded when trying to connect","category":"__probe"}
4. AppDependencies - Connection Duration Spike
Pre-incident: avg ~1ms, 0 failures
During incident: avg ~5001ms (5s timeout), 100% failure
Post-recovery: avg ~1.3ms, 0 failures
5. PostgreSQL Diagnostics Log
2026-06-22 19:02:58 UTC - "database system is ready to accept connections"
6. NSG State (at time of investigation)
nsg-aks-auf6gw on AKS subnet: allow-all-outbound (Allow, priority 1000) - outbound path was clear
allow-web-inbound rule changed from Allow to Deny (unrelated security fix)
NSG was NOT the cause - outbound VNet traffic was permitted throughout
7. Resource State at Investigation Time
Resource
State
PostgreSQL zava-pg-auf6gw
Ready (recovered)
AKS aks-Zava-auf6gw
Succeeded / Running
Network path (AKS to PG)
Open via VNet integration
Root Cause
Manual operator action. User admin@MngEnvMCAP602713.onmicrosoft.com executed a stop action on the PostgreSQL Flexible Server zava-pg-auf6gw at 18:56:39 UTC via the Azure control plane. This immediately began shutting down the database, causing all application connections to time out. The alert name postgres-network-blocked is misleading - the issue was not a network block but a stopped server.
The server was subsequently restarted by service principal 48f51566-903d-4fc5-b28b-13d8b3340c7c at 19:01:51 UTC, restoring service at 19:02:58 UTC.
Remediation
Immediate (Completed)
PostgreSQL server restarted at 19:01:51 UTC by service principal
Server ready and accepting connections at 19:02:58 UTC
Service fully recovered at 19:03 UTC - 0 failures confirmed
Short-term
Investigate why admin@MngEnvMCAP602713.onmicrosoft.com stopped the server - determine if intentional or accidental
Rename alert postgres-network-blocked to better reflect all connectivity failure modes (e.g., postgres-connection-failure)
Add Azure Policy or RBAC constraint to prevent unauthorized stop/start of production PostgreSQL servers
Long-term
Implement Azure Resource Lock (CanNotDelete/ReadOnly) on critical database resources
Create postgres-server-stopped activity log alert that fires immediately on stop actions
Add circuit breaker pattern in zava-api to fail fast instead of accumulating 5s timeouts
Implement connection retry with exponential backoff in the application layer
Action Items
#
Action
Owner
Priority
1
Determine if the stop action was intentional or accidental
Ops team
P0
2
Add Azure Resource Lock on zava-pg-auf6gw
Infra team
P1
3
Create activity log alert for PostgreSQL stop/start actions
SRE team
P1
4
Restrict RBAC for server stop/start to authorized principals only
Security team
P1
5
Rename alert to postgres-connection-failure
SRE team
P2
6
Implement circuit breaker in zava-api DB connection layer
AKS Cluster:/subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/Microsoft.ContainerService/managedClusters/aks-Zava-auf6gw
Summary
PostgreSQL Flexible Server
zava-pg-auf6gwwas manually stopped byadmin@MngEnvMCAP602713.onmicrosoft.comat 18:56:39 UTC on 2026-06-22, causing a complete database outage for the Zava food delivery API (zava-api). All PostgreSQL connection attempts from AKS pods timed out at the 5-second threshold, resulting in health check failures, product query failures, and total service unavailability for approximately 6 minutes.The server was restarted by a service principal (
48f51566-903d-4fc5-b28b-13d8b3340c7c) at 19:01:51 UTC and became ready at 19:02:58 UTC. Service fully recovered at 19:03 UTC.Alert:
postgres-network-blocked(Sev1) fired at 18:59:38 UTC.Impact
/api/products,/api/products/category/*, health checksPre-incident baseline (18:33-18:57 UTC)
During incident (18:58-19:02 UTC)
Timeline
admin@MngEnvMCAP602713.onmicrosoft.cominitiates PostgreSQL stop actionpostgres-network-blocked(Sev1) fires48f51566-903d-4fc5-b28b-13d8b3340c7cinitiates server startEvidence
1. Activity Log - Stop Action
2. Activity Log - Start Action
3. Container Logs (sample)
{"level":"error","message":"Health check: DB unreachable","timestamp":"2026-06-22T19:01:35.147Z","service":"zava-api","error":"timeout exceeded when trying to connect"} {"level":"error","message":"Failed to fetch products","timestamp":"2026-06-22T19:01:35.148Z","service":"zava-api","error":"timeout exceeded when trying to connect","endpoint":"/api/products"} {"level":"error","message":"Failed to fetch products by category","timestamp":"2026-06-22T19:01:35.149Z","service":"zava-api","error":"timeout exceeded when trying to connect","category":"__probe"}4. AppDependencies - Connection Duration Spike
5. PostgreSQL Diagnostics Log
6. NSG State (at time of investigation)
nsg-aks-auf6gwon AKS subnet:allow-all-outbound(Allow, priority 1000) - outbound path was clearallow-web-inboundrule changed from Allow to Deny (unrelated security fix)7. Resource State at Investigation Time
zava-pg-auf6gwaks-Zava-auf6gwRoot Cause
Manual operator action. User
admin@MngEnvMCAP602713.onmicrosoft.comexecuted astopaction on the PostgreSQL Flexible Serverzava-pg-auf6gwat 18:56:39 UTC via the Azure control plane. This immediately began shutting down the database, causing all application connections to time out. The alert namepostgres-network-blockedis misleading - the issue was not a network block but a stopped server.The server was subsequently restarted by service principal
48f51566-903d-4fc5-b28b-13d8b3340c7cat 19:01:51 UTC, restoring service at 19:02:58 UTC.Remediation
Immediate (Completed)
Short-term
admin@MngEnvMCAP602713.onmicrosoft.comstopped the server - determine if intentional or accidentalpostgres-network-blockedto better reflect all connectivity failure modes (e.g.,postgres-connection-failure)Long-term
postgres-server-stoppedactivity log alert that fires immediately on stop actionszava-apito fail fast instead of accumulating 5s timeoutsAction Items
zava-pg-auf6gwpostgres-connection-failureReferences
/subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourcegroups/rg-zava-aks-postgres/providers/microsoft.operationalinsights/workspaces/law-zava-auf6gw/providers/Microsoft.AlertsManagement/alerts/570973c0-9944-793b-c390-d2ee03ca0033/subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/microsoft.insights/scheduledqueryrules/postgres-network-blocked/subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/Microsoft.DBforPostgreSQL/flexibleServers/zava-pg-auf6gw/subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/Microsoft.ContainerService/managedClusters/aks-Zava-auf6gwlaw-Zava-auf6gw(ID:69f6cc9c-c59f-4624-8b95-e12b93220ece)ai-Zava-auf6gw(ARM:/subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/Microsoft.Insights/components/ai-Zava-auf6gw)/subscriptions/f627598e-05c5-4093-8667-5730c4026ea3/resourceGroups/rg-zava-aks-postgres/providers/Microsoft.Network/networkSecurityGroups/nsg-aks-auf6gwf627598e-05c5-4093-8667-5730c4026ea3rg-zava-aks-postgresThis issue was created by sre-agent-jkazytu5kl5py--70975bf6
Tracked by the SRE agent here