Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
4 changes: 2 additions & 2 deletions .dockerignore
Original file line number Diff line number Diff line change
Expand Up @@ -26,8 +26,8 @@
**/Dockerfile
**/.dockerignore

# Deploy/infra lives in the civfix-infra repo now; the only thing left here is the email-worker
# (a standalone Cloudflare Worker) which must never be baked into the api/worker image.
# infra/ holds only the email-worker (a standalone Cloudflare Worker), which must never be baked into
# the api/worker image.
infra

# Tests + fixtures are not needed for a production image build (build stage only runs tsup, not tests)
Expand Down
6 changes: 3 additions & 3 deletions .github/workflows/build-images.yml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
name: Build release images

# Builds the two artifacts a civfix-backend release consists of — the api and the media-worker — and
# Builds the two artifacts a civfix-backend release consists of (the api and the media-worker) and
# pushes them to GHCR. This is the ONLY place either image is built: the VPS deploy pulls what this
# workflow produced (civfix-infra ops/deploy.sh) instead of rebuilding from source on the box, so
# staging and production run byte-identical bytes rather than two builds of one commit.
Expand All @@ -12,13 +12,13 @@ name: Build release images
#
# NATIVE ARM64, NOT QEMU. Both VPSes are Oracle Ampere (arm64) boxes. `ubuntu-24.04-arm` is a GitHub
# larger runner (available on the Team plan, billed per-minute rather than from the free allotment) and
# builds ~5-10x faster than cross-building an image with this much native compilation in it — the
# builds ~5-10x faster than cross-building an image with this much native compilation in it: the
# media-worker alone pulls a static ffmpeg and builds sharp/libvips bindings. `platform` is an input so
# a future amd64 box is a one-line change at the call site rather than a rewrite.
#
# TAGGING. The immutable tag is `sha-<full 40-char commit sha>`, and that shape is a CROSS-REPO CONTRACT:
# civfix-infra's ops/deploy.sh reconstructs exactly this tag from the commit it checked out. Do not
# shorten it, do not switch to a `metadata-action` tag scheme, and do not remove it — a deploy that
# shorten it, do not switch to a `metadata-action` tag scheme, and do not remove it: a deploy that
# cannot find its own release image fails closed and refuses to deploy. `main` is a moving convenience
# tag for humans; nothing reads it.
#
Expand Down
12 changes: 4 additions & 8 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,6 @@ on:
pull_request:
branches: [main]

# Cancel superseded runs on the same ref.
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
Expand All @@ -33,13 +32,12 @@ jobs:
- name: Install dependencies
run: pnpm install --frozen-lockfile

# Stage 0.7 (UI-unification) backend-isolation guard: fail CI if the backend ever pulls in
# The backend depends ONLY on the framework-free @civfix/shared contract; fail if it ever pulls in
# frontend runtime deps (react / react-dom / react-native / react-native-web / @civfix/ui).
# The backend depends ONLY on the framework-free @civfix/shared contract; this keeps it that way.
- name: Assert backend stays React/RN-free
run: node scripts/assert-no-frontend.mjs

# H13 gate: fail the PR on a HIGH/CRITICAL advisory in the PRODUCTION dependency tree. `--prod`
# Fail the PR on a HIGH/CRITICAL advisory in the PRODUCTION dependency tree. `--prod`
# deliberately ignores dev-only tooling (a DoS in a test runner is not a production exposure).
# When a finding has no upstream fix yet, add a `pnpm.overrides` entry in the root package.json
# (with a comment explaining why the forced version is safe) rather than removing this step.
Expand All @@ -61,8 +59,8 @@ jobs:
integration:
name: integration (postgis + redis)
runs-on: ubuntu-latest
# Integration tests that require Docker-backed services run here. They are populated by later
# steps; for now this job installs, builds, and is ready to run them against live containers.
# Placeholder: the Docker-backed integration suite runs inside the build job's `pnpm test`
# (testcontainers). This job stays because its name is a required status check on main.
services:
postgres:
image: postgis/postgis:16-3.4
Expand Down Expand Up @@ -110,8 +108,6 @@ jobs:
- name: Build
run: pnpm build

# Later steps add integration test scripts; this is the hook for them. Kept as a no-op-safe
# placeholder so the job is green until those tests exist.
- name: Integration tests
run: echo "integration tests are added by later steps; services are up at DATABASE_URL / REDIS_URL"

Expand Down
19 changes: 10 additions & 9 deletions .github/workflows/deploy-staging.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,10 +3,10 @@ name: Deploy (staging)
# Deploys the civfix backend to the STAGING VPS (193.122.142.107). TWO jobs, in order:
#
# 1. `images` builds the api + media-worker images for this commit and pushes them to GHCR
# (build-images.yml). This is the ONLY build of this release — production is later
# (build-images.yml). This is the ONLY build of this release: production is later
# promoted onto these exact digests rather than rebuilding.
# 2. `deploy` SSHes the box, whose forced command runs /opt/civfix/infra/ops/deploy.sh. That script
# git-pulls the infra repo (single `main` branch — the same tree prod runs; the box's env
# git-pulls the infra repo (single `main` branch, the same tree prod runs; the box's env
# marker at /etc/civfix/env selects env/staging.sh and secrets/staging/) + this backend
# (at CIVFIX_REF, which env/staging.sh defaults to `main`), PULLS the images job 1 just
# published, SOPS-decrypts the secrets, then blue/green swaps the compose stack.
Expand All @@ -16,8 +16,8 @@ name: Deploy (staging)
# Mirrors deploy.yml (prod). Issue: https://github.com/civfix/issue-tracker/issues/108
#
# Triggers:
# - Every push to main. `main` IS the staging lane — there is no `dev` branch any more; feature
# branches PR into main, and prod is promoted from a `v*` release cut off main.
# - Every push to main. `main` IS the staging lane; feature branches PR into main, and prod is
# promoted from a `v*` release cut off main.
# - Manually via the Actions "Run workflow" button (workflow_dispatch).
# Both deploy the staging env's CIVFIX_REF (from env/staging.sh, default `main`); the SSH forced command
# ignores any client-sent command, so a deploy can never target an arbitrary ref.
Expand All @@ -28,14 +28,15 @@ name: Deploy (staging)
# Against the old stop-then-start deploy the API genuinely goes down for tens of seconds and this job
# fails every time. This backend branch must merge AFTER the civfix-infra change, never alone.
#
# THE AVAILABILITY CHECK IS AN ASSERTION, NOT A "did it come back" WAIT — see the deploy.yml header for
# THE AVAILABILITY CHECK IS AN ASSERTION, NOT A "did it come back" WAIT; see the deploy.yml header for
# the reasoning. GET /healthz is polled once a second for the whole deploy (pure and
# rate-limit-allowlisted, so a 1s poll cannot trip the limiter); the job fails when the longest gap
# between serving samples exceeds MAX_OUTAGE_S seconds, with a 503 carrying the x-civfix-draining
# response header (the retiring color, by design) not counted as an outage. /readyz (db + redis, not
# allowlisted) stays a separate gate afterwards, and the availability + deploy-rc assertion runs last:
# both steps are `if: always()`, so a red availability result cannot skip the readiness gate and a
# failed ssh still fails the job. It all runs after the flip, so it reports rather than prevents.
# both steps are `if: success() || failure()`, so a red availability result cannot skip the readiness
# gate and a failed ssh still fails the job. It all runs after the flip, so it reports rather than
# prevents.
#
# Required repo secret:
# STAGING_DEPLOY_SSH_KEY Private half of the dedicated staging CI deploy key (ed25519). Its public half
Expand All @@ -56,7 +57,7 @@ permissions:
# NOTE THE SCOPE: the serialization is on the `deploy` JOB, not on the workflow.
#
# It has to be. Workflow-level concurrency queues the whole RUN, and a queued run starts none of its
# jobs — including its image build. So while one box was waiting for the image of a newer commit (it
# jobs, including its image build. So while one box was waiting for the image of a newer commit (it
# deploys the branch tip, not the commit that triggered the run), the run that would have built that
# image could not start until the waiting deploy finished. The wait could never succeed, and every
# back-to-back merge burned the full retry budget and then failed.
Expand Down Expand Up @@ -157,7 +158,7 @@ jobs:

# No remote command is passed: the deploy key's forced command runs deploy.sh (CIVFIX_REF=main).
# Its stdout/stderr stream here and its exit code becomes `rc`, which the assertion step
# turns into a job failure. Keepalives keep the connection up during the (quiet) build.
# turns into a job failure. Keepalives keep the connection up while deploy.sh is quiet.
if ssh -i ~/.ssh/deploy_key \
-o BatchMode=yes -o IdentitiesOnly=yes -o StrictHostKeyChecking=yes \
-o ConnectTimeout=20 -o ServerAliveInterval=15 -o ServerAliveCountMax=8 \
Expand Down
39 changes: 20 additions & 19 deletions .github/workflows/deploy.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,9 +3,9 @@ name: Deploy (production)
# Deploys the civfix backend to the production VPS by SSHing in and running the on-box deploy script
# (/opt/civfix/ops/deploy.sh, a shim that execs /opt/civfix/infra/ops/deploy.sh). That script git-pulls
# the infra repo (single `main` branch; the box's env marker at /etc/civfix/env selects env/prod.sh and
# secrets/prod/) + this backend (origin/main), SOPS-decrypts the secrets, then rebuilds and blue/green
# swaps the docker compose stack. The build happens ON the VPS (4-vCPU arm64 box); this workflow only
# triggers it and then proves the public API stayed up.
# secrets/prod/) + this backend, SOPS-decrypts the secrets, PULLS the promoted release images and
# blue/green swaps the docker compose stack. This workflow only triggers it and then proves the public
# API stayed up.
#
# PRODUCTION IS PROMOTED, NOT BUILT. A merge to main deploys STAGING (deploy-staging.yml) and publishes
# the release images; this workflow takes those exact digests and puts them in front of production.
Expand All @@ -14,15 +14,15 @@ name: Deploy (production)
# Triggers:
# - A PUBLISHED GitHub release whose tag matches `v*`. Cutting a release IS the production deploy.
# - Manually via the Actions "Run workflow" button (workflow_dispatch), which re-deploys the current
# newest release — the recovery path when a deploy failed after the release was already published.
# newest release: the recovery path when a deploy failed after the release was already published.
# Neither carries a ref: the SSH forced command discards any client-sent command, and the on-box
# deploy.sh resolves env/prod.sh's CIVFIX_REF=@latest-release to the highest `v*` tag MERGED INTO
# origin/main. Production therefore cannot run code that did not go through the staging lane.
#
# BECAUSE THE BOX RESOLVES "NEWEST RELEASE" ITSELF, the promote job below refuses to run when the tag
# being released is not the newest one. Publishing a release for an older tag would otherwise deploy the
# NEWER code while reporting success against the older tag — a silent wrong-deploy, and exactly the shape
# a panicked "re-release the last good version" rollback takes. Roll back on the box instead:
# NEWER code while reporting success against the older tag: a silent wrong-deploy, and exactly the
# shape a panicked "re-release the last good version" rollback takes. Roll back on the box instead:
# sudo -u civfix CIVFIX_REF=<tag|sha> /opt/civfix/infra/ops/deploy.sh
# or, for a bad-but-healthy release, ops/rollback.sh (~10s, flips back to the retired color).
#
Expand All @@ -47,18 +47,19 @@ name: Deploy (production)
# memoized 5s server-side.
#
# STEP ORDER IS LOAD-BEARING: the probe step only measures and records; the readiness gate runs next
# with `if: always()`; the availability + deploy-rc assertion runs LAST, also `if: always()`. A red
# availability result therefore can never skip /readyz, and a failed ssh still fails the job.
# with `if: success() || failure()`; the availability + deploy-rc assertion runs LAST, with the same
# condition. A red availability result therefore can never skip /readyz, and a failed ssh still
# fails the job.
#
# It all runs AFTER the flip, so it cannot prevent a bad deploy — it fails the job loudly. Rolling back
# is still `CIVFIX_REF=<sha>` on the box (or `docker compose start api-<previous color>`).
# It all runs AFTER the flip, so it cannot prevent a bad deploy; it fails the job loudly. Roll back on
# the box as described above.
#
# Required repo secret:
# DEPLOY_SSH_KEY Private half of the dedicated CI deploy key (ed25519). Its public half is the
# `deployer` user on the VPS, whose authorized_keys forces this key to run ONLY
# `sudo -n -u civfix -H /opt/civfix/ops/deploy.sh` (sudoers-scoped to that one
# command). A leaked key can at most trigger a deploy of origin/main — never a
# shell, other commands, or root.
# command). A leaked key can at most trigger a deploy of the newest release, never
# a shell, other commands, or root.
#
# The VPS origin IP and its SSH host key are pinned in the job below (neither is secret; SSH 22 is
# firewalled to key-only/no-root and ingress 80/443 is Cloudflare-only).
Expand Down Expand Up @@ -108,19 +109,19 @@ jobs:
# compare against.
git fetch --quiet --force --tags origin '+refs/heads/main:refs/remotes/origin/main'

# THE SAME SELECTION THE BOX MAKES — civfix-infra env/prod.sh CIVFIX_RELEASE_TAG_REGEX and
# THE SAME SELECTION THE BOX MAKES: civfix-infra env/prod.sh CIVFIX_RELEASE_TAG_REGEX and
# ops/deploy.sh. Keep this regex byte-identical to that one: the box resolves the release
# independently, so if the two rules disagree this job blesses one commit and the box deploys
# another.
#
# git sorts the whole `v*` set and the grep filters afterwards; that is sound because
# filtering preserves relative order, so the first surviving line is still the highest. The
# regex is what excludes prereleases — git's -v:refname ranks `v2.0.0-beta.1` ABOVE `v2.0.0`,
# regex is what excludes prereleases: git's -v:refname ranks `v2.0.0-beta.1` ABOVE `v2.0.0`,
# the opposite of semver. The glob alone is not a release rule either: `v*` also matches
# `vtest` and `v1.2`, which would outrank every real release and wedge this job permanently.
RELEASE_RE='^v[0-9]+\.[0-9]+\.[0-9]+$'
# `|| true`: under `set -euo pipefail` a grep that matches nothing exits 1 and errexit kills
# the step AT THIS ASSIGNMENT, making the ::error:: annotation below unreachable — the job
# the step AT THIS ASSIGNMENT, making the ::error:: annotation below unreachable, so the job
# would fail with a bare exit code and no explanation. The guard reports it instead.
newest="$(git tag --list 'v*' --sort=-v:refname --merged origin/main | grep -E "$RELEASE_RE" | head -n 1)" || true
[ -n "$newest" ] || { echo "::error::no release tag matching $RELEASE_RE is merged into origin/main." >&2; exit 1; }
Expand All @@ -129,7 +130,7 @@ jobs:

if ! printf '%s' "$tag" | grep -qE "$RELEASE_RE"; then
echo "::error::'$tag' is not a vMAJOR.MINOR.PATCH release tag; refusing to deploy production." >&2
echo "::error::Prereleases are excluded by design — they never reach prod." >&2
echo "::error::Prereleases are excluded by design; they never reach prod." >&2
exit 1
fi

Expand Down Expand Up @@ -161,7 +162,7 @@ jobs:

- name: Promote the staging images to this release
# NO REBUILD, BY DESIGN. `imagetools create` adds a tag to a manifest that already exists in the
# registry — a registry-side operation, no pull, no push of layers — so the bytes production runs
# registry (a registry-side operation, no pull, no push of layers), so the bytes production runs
# are provably the bytes staging exercised. A missing image means this commit never completed a
# staging build, and the deploy is refused here rather than half-way through on the box.
env:
Expand All @@ -186,7 +187,7 @@ jobs:
fi
echo "$image: $digest"
# `:$TAG` is the durable, immutable record of this release. `:prod` is a MOVING tag and is
# for humans reading the package list only — nothing consumes it: the box pulls
# for humans reading the package list only; nothing consumes it: the box pulls
# `sha-<40>` and immediately re-pins to the digest (civfix-infra/ops/deploy.sh). Never make
# anything depend on `:prod`.
docker buildx imagetools create \
Expand Down Expand Up @@ -277,7 +278,7 @@ jobs:

# No remote command is passed: the deploy key's forced command runs /opt/civfix/ops/deploy.sh.
# Its stdout/stderr stream here and its exit code becomes `rc`, which the assertion step
# turns into a job failure. Keepalives keep the connection up during the (quiet) build.
# turns into a job failure. Keepalives keep the connection up while deploy.sh is quiet.
if ssh -i ~/.ssh/deploy_key \
-o BatchMode=yes -o IdentitiesOnly=yes -o StrictHostKeyChecking=yes \
-o ConnectTimeout=20 -o ServerAliveInterval=15 -o ServerAliveCountMax=8 \
Expand Down
2 changes: 1 addition & 1 deletion NOTICE
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
civfix — the API and media worker
civfix: the API and media worker
Copyright (c) 2026 Reach Out Los Angeles Inc. and contributors

This program is free software: you can redistribute it and/or modify it under
Expand Down
Loading
Loading