Skip to content

fix(deploy): prune unused tagged images before pull — ENOSPC on EC2 - #206

Merged
Anuraj-dev merged 1 commit into
mainfrom
fix/deploy-disk-prune
Jul 30, 2026
Merged

Anuraj-dev merged 1 commit into
mainfrom
fix/deploy-disk-prune

Conversation

@Anuraj-dev

Copy link
Copy Markdown
Owner

Summary

Today's backend deploy (run 30553109392, after #199 merged) failed on the EC2 side with no space left on device while extracting the new image layer. Prod stayed up — the pull died before the container restart, so the previous release kept serving — but no backend release can land until the disk is cleared.

Root cause: releases are immutable :<sha> tags, and the pipeline's only cleanup is docker image prune -f after up -d, which removes dangling images only. Tagged release images are never dangling, so one full image per deploy accumulated since the 2026-07-01 cutover until the disk filled.

Fix: docker image prune -af immediately before the pull, in both the deploy and rollback scripts. At that point the running last-good container pins its own image (the recorded rollback target), so -a removes only genuinely unused older releases. Disk usage steadies at ~2 release images, and the prune doubles as self-healing for the current full-disk state — the first deploy after this merges frees the space it needs.

Testing

  • Workflow-file change only; no app code. docker image prune semantics (-f = dangling-only vs -a = unused): per Docker docs.
  • The real gate is the first deploy run after merge — this PR's merge itself triggers one (the workflow file is in its own path filter), which validates the fix against the currently-full disk.

Every release is an immutable :<sha> tag, and the post-deploy
'docker image prune -f' removes only dangling images — tagged release
images were never cleaned up, one accumulating per deploy until the
instance disk filled and today's deploy died mid-pull with
'no space left on device' (run 30553109392; the old container kept
serving, so prod stayed up on the previous release).

Prune with -a before the pull in both the deploy and rollback scripts:
at that point the running last-good container pins its image (the
recorded rollback target), so only genuinely unused older releases are
removed, and the disk steadies at roughly two release images.
@vercel

vercel Bot commented Jul 30, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
a-meet Ready Ready Preview Jul 30, 2026 3:12pm

@Anuraj-dev
Anuraj-dev merged commit c442c9a into main Jul 30, 2026
15 checks passed
@Anuraj-dev
Anuraj-dev deleted the fix/deploy-disk-prune branch July 30, 2026 15:16

This branch was successfully deployed

1 active deployment
Preview — bcf243d5 Deployed Jul 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant