Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
51 commits
Select commit Hold shift + click to select a range
a7379fa
host comms: refused sends are not charged, failed enqueues roll back,…
theobong Sep 23, 2026
4432d15
reports and media: scratch failures retry instead of rejecting upload…
theobong Sep 23, 2026
e416169
adapters and env: storage timeouts, per-batch push pruning, strict en…
theobong Sep 23, 2026
c75aefb
auth: suspension during login, handle rename race, trimmed redirects,…
theobong Sep 23, 2026
4ee8351
events and hours: refused guest sign-ups roll back, joined means atte…
theobong Sep 23, 2026
764b62b
social and notifications: fan-out failures are reported, coalesced ro…
theobong Sep 23, 2026
3ff37a1
chat: report fan-out keeps its cap on every path, late members can be…
theobong Sep 23, 2026
7116f02
host registration: slots cannot oversell, waitlists wake on transfers…
theobong Sep 23, 2026
1d831b2
admin: keyset cursors keep microseconds, operator removals reach the …
theobong Sep 23, 2026
32ca1f2
db scripts: microsecond backfill cursors, migrate never pools a broke…
theobong Sep 23, 2026
ece56ab
admin mail: in-flight sends block resends, deadlines keep the outreac…
theobong Sep 23, 2026
9febd58
outreach jobs log a failed claim release
theobong Sep 23, 2026
f0078ab
jurisdiction health ignores bounced contacts
theobong Sep 23, 2026
06df83c
cleanup duplicate test reads the clock once
theobong Sep 23, 2026
222e4a7
docs: apns gateway flag, in-flight sends, bounce marker, export reaper
theobong Sep 23, 2026
e526d63
stuck media requeues keep the upload etag
theobong Sep 23, 2026
02756bb
city forwarding and worker abuse checks log through the injected logger
theobong Sep 23, 2026
84c069c
media byte quota uses the shared counter window
theobong Sep 23, 2026
71782f2
in-memory owner resolve mirrors the repository; takedown doc matches …
theobong Sep 23, 2026
5a92671
one microsecond-exact keyset cursor for admin and consumer lists
theobong Sep 23, 2026
e9f6fcc
merge the security fixes into the correctness branch
theobong Sep 23, 2026
01ceda3
event edits apply scalar fields, links and slots in one transaction
theobong Sep 23, 2026
1849993
certificate verify prompt names the deployment's own page
theobong Sep 23, 2026
6fc6a37
guest upserts always report whether they created the row
theobong Sep 23, 2026
a196d2d
roster pages sorted by check-in keep every row
theobong Sep 23, 2026
388cc4e
websocket buffers a full frame burst before auth
theobong Sep 23, 2026
11bb492
notification dedupe serializes writers; broadcasts re-pend only the r…
theobong Sep 23, 2026
bdbbe0c
a cancellation notice whose plan enqueue failed is re-planned on retry
theobong Sep 23, 2026
762e180
refused sends give back their counter charges
theobong Sep 23, 2026
de4ec35
each export run writes its own object
theobong Sep 23, 2026
26791c1
host broadcast and admin page lists page on microsecond keysets; susp…
theobong Sep 23, 2026
dca9aab
demo seeders share a guard-free seat module and the api's ticket secret
theobong Sep 23, 2026
d59ace4
container close lets draining jobs finish before refusing handles
theobong Sep 23, 2026
819e63f
storage transfers get the transfer timeout per request
theobong Sep 23, 2026
ded9d9c
env: lenient drain delay, single-letter apns gateway flags
theobong Sep 23, 2026
7d68aec
a forged cursor with an impossible instant starts from the first page
theobong Sep 23, 2026
610052a
guest roster pages on the shared keyset
theobong Sep 23, 2026
a622a03
media worker refuses an image timeout that outruns the job budget
theobong Sep 23, 2026
eeb2806
export runs fence before deleting and never lose track of an object
theobong Sep 23, 2026
5e68410
recipient budget gives a refused charge back with a decrement
theobong Sep 23, 2026
be92edc
registration twins serialize on their key, slots lock before members,…
theobong Sep 23, 2026
03e51ad
chat fan-out retries a total bell failure and logs once per fan-out
theobong Sep 23, 2026
d15754d
guest refusals carry a reason and doomed code requests send no code
theobong Sep 23, 2026
3783cd2
unreposting an unreadable post answers a withdrawn card, never its co…
theobong Sep 23, 2026
0bc7392
report routing and the routable count skip bounced contacts
theobong Sep 23, 2026
59ed2ba
owner takedowns skip a strike only on items the owner alone opened; r…
theobong Sep 23, 2026
d5b7dec
admin row locks leave foreign-key checks unblocked
theobong Sep 23, 2026
87d92e9
data exports back off across mail outages and are always on record
theobong Sep 23, 2026
58a49fe
bounces record before discovery, failing bounces park after six tries…
theobong Sep 23, 2026
894045c
report chat and event mail skip bounced city contacts
theobong Sep 23, 2026
1891b0e
href edge trim runs in linear time
theobong Sep 23, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
45 changes: 42 additions & 3 deletions docs/inbound-mail-effects.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# Mail effects and the outbound send triad

**Audience:** internal (engineering). Not served publicly.
**Last updated:** 2026-09-22 (sender authentication policy; unauthenticated token replies file as unaffiliated).
**Last updated:** 2026-09-23 (an attempt with no outcome yet counts as in flight; bounce bookkeeping and its completion marker).

An inbound message that correlates to a mail thread can drive **public** effects: a report status
transition, a public `report_timeline` row, a report-chat system message, and a push to the reporter.
Expand Down Expand Up @@ -175,6 +175,34 @@ The exposure is bounded and one-sided:
If `report_timeline` later gains a `meta` jsonb for another reason, the fix is to stamp the mail message
id into the row and make the timeline/chat/notify steps no-op when that marker is already present.

## Bounces: bookkeeping and its completion marker

A delivery status notification (`detectBounce` in `services/api/src/services/admin/inbound-bounce.ts`:
a `mailer-daemon@` or `postmaster@` sender, a `report-type=delivery-status` content type, or an
`X-Failed-Recipients` header) is stored in the Inbox and then runs `handleBounce`. It acts only when the
DSN names both a failed recipient and an original Message-ID, the sender passes
`isPlausibleBounceSender`, and the Message-ID maps to a thread that actually sent to that recipient. It
then, in order:

1. sets the thread status to `bounced`;
2. stamps `bounced_at` on the jurisdiction's `jurisdiction_contacts` rows for that address and enqueues
`jurisdiction.discovery` for the geoid (when a geoid is known, from the thread or the contact);
3. writes the `bounced` mail event, with meta `{ failedRecipient, originalMessageId }`.

The `bounced` event is written **last** because it is the completion marker. When a step throws, the
object stays in `inbound/pending/` and the next sweep replays it; a replay runs `handleBounce` again,
which first asks `hasBounceEvent` (same thread, `type = 'bounced'`, same `originalMessageId`, same
`failedRecipient` ignoring case). With no marker it repeats the steps, all of which are safe to repeat.
With the marker it does nothing, so a duplicate delivery of a finished DSN cannot force a thread back to
`bounced` after an operator has changed its status.

A legacy `contact_emails` address has no `bounced_at` column, so the event's `failedRecipient` is what
marks it unusable: `legacyContactEmailUsable` skips an address with a `bounced` event on that
jurisdiction's threads newer than `contact_updated_at`, or with a bounced per-category row for the same
address. Discovery, the jurisdiction health probe behind `resolveForPoint` and the outreach digest all
apply it, and all ignore a per-category contact whose `bounced_at` is set. Saving the contact again moves
`contact_updated_at` past the old event, which makes a corrected address usable again.


---

Expand All @@ -188,7 +216,7 @@ service and the admin report repository import.
|---|---|
| **Send deadline** (`outboundSendDeadlineMs`) | Total wall clock for one delivery: a phase budget (`OCI_EMAIL_SMTP_TIMEOUT_MS × 3`, covering connect + greeting + socket, all of which are INACTIVITY timeouts and so bound nothing on a trickling relay) plus the payload's time at a floor throughput (`OUTBOUND_SEND_MIN_THROUGHPUT_BPS`, default 256 KiB/s). Clamped to `2^31 - 1` so it can never overflow `setTimeout`, which Node silently clamps to 1 ms. |
| **In-flight window** (`ROUTE_DEADLINE_INFLIGHT_SECONDS`, 900 s) | How long a `failed` event whose meta says `reason: "deadline"` counts as *still in flight* rather than as a delivery failure. |
| **Stale claim** (`ROUTE_CLAIM_STALE_SECONDS`, 900 s) | How long an outbound row with NO event at all counts as in flight before it is treated as a crashed claim and becomes re-routable. |
| **Stale claim** (`ROUTE_CLAIM_STALE_SECONDS`, 900 s) | How long an outbound row with no `sent` and no `failed` event counts as in flight (reply, resend and re-route are refused) before it is treated as a crashed claim and becomes re-routable. |

## Why the deadline is not an abort

Expand All @@ -213,7 +241,18 @@ non-delivery**:
`mail_events.message_id`), never thread-wide. Thread-wide, any earlier reason-less `failed` (attempt 1
connect timeout, say) satisfied the predicate and killed the in-flight guard for every later attempt.
The verdict is: a hard `failed` on the newest attempt → failed; a `deadline` failure inside the window →
in flight; any other `failed` → failed; no event and older than the stale window → crashed claim.
in flight; any other `failed` → failed; no event and younger than the stale window → in flight; no event
and older than the stale window → crashed claim.

`sendInFlightExpr` (`services/api/src/services/admin/outbound-send-sql.ts`) is the in-flight half of that
verdict, read from the same newest attempt: no `sent` event, and either a `deadline` failure inside the
window or no `failed` event at all while the outbound row is younger than the stale window. The outbound
row is inserted before transmission starts, so an attempt with no outcome yet is a send still on the
wire, or one whose process died mid-send; the two cannot be told apart until the stale window passes.
Until then reply and resend (`MailService`, through `hasSendInFlight`) and re-route (`assertRoutable`,
through the report's `send_in_flight`) answer 409 `SEND_IN_FLIGHT_CONFLICT`. Once a `sent` or `failed`
event lands, or the row passes the stale window, the guard lifts and `sendFailedExpr` decides whether the
attempt reads as failed.

## Why misconfiguration fails closed

Expand Down
5 changes: 3 additions & 2 deletions docs/media-pipeline-hardening.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,8 +31,9 @@ serving an object the uploader can still overwrite with unchecked bytes after
the asset is already published on a report.

The worker also binds the row to the exact object version it inspected: the API
records the upload's ETag at finalize and passes it in the job payload, the
worker compares it to what it downloaded, and it re-HEADs the upload key
records the upload's ETag at finalize (on the row as `media_assets.upload_etag`,
migration 0184, so a stuck-sweep requeue carries it too) and passes it in the job
payload, the worker compares it to what it downloaded, and it re-HEADs the upload key
immediately before publishing. A mismatch (or a vanished object) is a
`rejected` outcome — never a throw, per the pipeline's never-throw invariant.

Expand Down
11 changes: 6 additions & 5 deletions docs/report-takedown.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,11 +42,12 @@ spam the queue.

## Offline / no-DB behavior

`reportOwnedBy` is DB-gated and fail-safe: with no `DATABASE_URL` (all-fakes boot)
or on any query error it returns `false`, so the route degrades to the ordinary
user-report path (the request is still filed, just not flagged as an owner
takedown). The audit write only runs when ownership was confirmed, which implies
a real DB is present.
`reportOwnedBy` is DB-gated: with no `DATABASE_URL` (all-fakes boot) it returns
`false`, so the route degrades to the ordinary user-report path (the request is
still filed, just not flagged as an owner takedown). A query error is not caught:
it propagates and the request fails with 500 (not filed), so a transient DB error
never downgrades an owner takedown to a third-party report. The audit write only
runs when ownership was confirmed, which implies a real DB is present.

## What this does NOT do (product/counsel DECISIONS)

Expand Down
24 changes: 21 additions & 3 deletions docs/retention-cleanup.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# Retention cleanup jobs (civfix-backend)

**Audience:** internal (engineering + ops). Not served publicly.
**Last updated:** 2026-09-02 (audit-fix pass: inbound_emails TTL + doc-drift corrections).
**Last updated:** 2026-09-23 (host export reaper: abandoned runs, failed-run objects, row retention waits for the object).

Backs the "written retention schedule + scheduled cleanup jobs" item in
`documents/21-privacy-compliance.md` §7.1 for the TTL-able auth artifacts that
Expand Down Expand Up @@ -386,11 +386,29 @@ a failed lane is logged and the next lane still runs.
|---|---|---|
| `broadcast_deliveries` rows | **180 d** from `created_at` | `host.retention.sweep` → `broadcast_deliveries` |
| `broadcasts` subject + body + CTA | scrubbed **180 d** after `finished_at`; the counts and the row are kept indefinitely as the audit record that a message was sent | `host.retention.sweep` → `broadcast_content` |
| `host_exports` rows | **90 d** from `requested_at` | `host.retention.sweep` → `host_exports` |
| host export OBJECTS | **24 h** (`HOST_EXPORT_TTL_HOURS`); objects are deleted BEFORE their rows | `host.export.reap` |
| `host_exports` rows | **90 d** from `requested_at`, and only once the row no longer names an object (`r2_key IS NULL`) | `host.retention.sweep` → `host_exports` |
| host export OBJECTS | **24 h** (`HOST_EXPORT_TTL_HOURS`) for a ready export; a failed run's object at the next reaper pass; objects are deleted BEFORE their rows | `host.export.reap` |
| `event_metrics_daily` | **never** — aggregates with no identifier of any kind, and the only long-run record a host has | — |
| `broadcast_unsubscribes`, `email_suppressions` | **indefinite, deliberately** — a suppression list that expires re-enables mailing someone who said stop (same reasoning as `sms_opt_outs`) | — |

`host.export.reap` (`HOST_EXPORT_REAP_CRON`, default hourly at :40; up to 200 rows per
lane per pass) runs two lanes in `services/api/src/services/host/export-service.ts`:

- **Expired.** A `ready` row whose `expires_at` has passed has its object deleted, then
becomes `expired` with `r2_key = NULL`.
- **Orphaned.** A `failed` row that still names an object, a `running` row whose
`started_at` is more than 1 hour old, and a `queued` row whose `requested_at` is more
than 1 hour old (`EXPORT_ABANDON_AFTER_MS`). Any object is deleted first; the row then
becomes `failed` with `r2_key = NULL`, keeps an existing `error_code` or else takes
`not_started` (was queued) or `build_failed` (was running), and gets `completed_at` if
it had none. The update applies only while the row's status and run token are unchanged,
so a run that claimed the row in between is not overwritten.

When an object delete fails, the row is left as it is and the next pass retries it. A
failed run also deletes its own object straight away; the reaper finishes that cleanup
when it could not. `host.retention.sweep` deletes `host_exports` rows past 90 days only
when `r2_key IS NULL`, so it never drops the last pointer to an object still in storage.

### Donation links, retired payments tables, and legal

civfix no longer processes donations. A donation link is now a single nullable
Expand Down
4 changes: 3 additions & 1 deletion docs/security/2026-07-24-full-backend-security-review.md
Original file line number Diff line number Diff line change
Expand Up @@ -322,7 +322,7 @@ Everything an operator has to do by hand, in order, plus the two infra facts and
node dist/db/migrate.js # == pnpm --filter @civfix/api db:migrate
```

`drizzle/` holds **165 files**, `0000_extensions.sql` … `0182_organization_invites_invited_by_idx.sql`. The nine rows
`drizzle/` holds **167 files**, `0000_extensions.sql` … `0185_inbound_bounce_attempts.sql`. The nine rows
below are exactly what this change set adds — `0052`–`0060`, contiguous, no gaps — and everything from
`0000` through `0051_social_posts.sql` predates it. (`0060` arrived later than the rest, with the feed
redesign; it is listed here because this table is the single operator runbook. `0061`–`0064` arrived
Expand Down Expand Up @@ -572,6 +572,8 @@ foreign key into `organizations` from `0105`.
| `0180_civfix_official_account.sql` | DATA only, no DDL (the Drizzle mirror is unchanged): creates the official `@civfix` account (`users.id` `00000000-0000-4000-8000-00000000c1f1`, display name `CivFix`, role `citizen`, no email, DMs closed) that admin-panel chat posts are authored by. First it frees the handle: any OTHER row holding `civfix` (citext, so any case) gets its auto-placeholder handle `user` + the first 12 hex digits of its id, and nothing else about that row changes - not its display name, and not `handle_changed_at`, so its owner may pick a new handle at once. The INSERT conflicts on the id only, so a placeholder or handle clash fails the file instead of skipping it. At most two `users` rows are touched: milliseconds, row locks only. `SessionService` refuses to mint or resolve a session for the id, and with no email and no OAuth identity no sign-in path can reach it | Admin-panel report and event chat posts fail on the `chat_messages.sender_id` foreign key once the code that authors them as the official account is deployed; nothing else reads the row |
| `0181_media_uploader.sql` | adds the nullable `media_assets.uploader` column (`u:<user id>`, `a:<anon token id>` or `anon`), written by upload create, so a report, post, chat or avatar claim by uploadId binds only the caller's own upload. Nullable with no default and no backfill: a catalog-only change on a hot table, lock held for milliseconds, no index (claims find the row through the unique `upload_id`). Rows from before the deploy stay NULL and are claimable only inside the claim window | Upload create fails writing a column that does not exist; claims fail on the missing column |
| `0182_organization_invites_invited_by_idx.sql` | adds the partial index `organization_invites_inviter_pending_idx (invited_by) WHERE status = 'pending'`, so the account-erasure statement that revokes the pending organization invites a user sent or received (`invited_by = $1 OR user_id = $1`) can BitmapOr this index with `organization_invites_invitee_pending_idx` instead of scanning the table. `organization_invites` is not on the hot-table list, so it builds inline under the migration transaction (milliseconds at current size); if the table has grown, build it `CONCURRENTLY` by hand first and the `IF NOT EXISTS` guard makes the file a no-op | Account deletion still works, with a sequential scan of `organization_invites` inside the erasure transaction |
| `0184_media_assets_upload_etag.sql` | adds `media_assets.upload_etag`, a nullable `text` with no default, no index and no backfill, so a catalog-only `ADD COLUMN IF NOT EXISTS` on a hot table (a brief ACCESS EXCLUSIVE lock, no rewrite, no scan). Finalize now writes the HEAD etag of the uploaded object in the same UPDATE that claims `finalized_at`, and the media-worker stuck sweep reads it back so a requeued `media.checks` job still rejects bytes re-PUT after finalize. Rows finalized before this file keep NULL and requeue with no etag, which skips the comparison exactly as before. An opaque object-version string, no PII: it lives and dies with its row | Every finalize and every stuck-sweep pick fails on the missing column once the code that writes and reads it is deployed |
| `0185_inbound_bounce_attempts.sql` | creates `inbound_bounce_attempts (object_key text PRIMARY KEY, attempts int, last_attempt_at timestamptz)`, a counter of failed bounce-bookkeeping runs per pending inbound object. The inbound processor increments it when a DSN's bookkeeping fails and, at `INBOUND_BOUNCE_MAX_ATTEMPTS`, parks the object under `inbound/failed/` and deletes the row, so a DSN that can never be recorded stops taking a slot in every sweep batch. New empty table, no index beyond the key, no backfill; a row holds only the pending object's own key and is deleted on success or park | Every DSN, including one whose bookkeeping succeeds, throws on the missing table at the counter write and stays under `inbound/pending/`, so every sweep replays it |


**Deferred to the NEXT release** (expand/contract, `docs/migrations-expand-contract.md`): 0.43.0
Expand Down
8 changes: 7 additions & 1 deletion services/api/.env.example
Original file line number Diff line number Diff line change
Expand Up @@ -218,7 +218,13 @@ APNS_KEY_ID= # [OPT] APNs auth key id
APNS_TEAM_ID= # [OPT] Apple team id for APNs
APNS_PRIVATE_KEY= # [OPT] APNs .p8 private key contents
APNS_BUNDLE_ID= # [OPT] iOS app bundle id (apns topic)
APNS_PRODUCTION= # [OPT] 1/true to use the production APNs gateway
APNS_PRODUCTION= # [BOOT] in production once all four APNS_* credentials above are set,
# USE_FAKE_PUSH or not (else [OPT]). true = production gateway
# (App Store and TestFlight builds), false = sandbox gateway.
# Parsed strictly: true/false, t/f, yes/no, y/n, on/off, 1/0 (any case);
# anything else fails boot. Unset outside production = the
# production gateway. A gateway mismatch makes APNs reject every
# token as BadDeviceToken.

# ===== push: FCM (bypassed by USE_FAKE_PUSH) [OPT] =====
FCM_SERVICE_ACCOUNT_JSON= # [OPT] Firebase service account JSON (string)
Expand Down
28 changes: 28 additions & 0 deletions services/api/drizzle/0184_media_assets_upload_etag.sql
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
-- =============================================================================
-- 0184_media_assets_upload_etag.sql
-- -----------------------------------------------------------------------------
-- WHY. Finalize HEADs the uploaded object and hands its ETag to the media.checks
-- job, and the worker rejects the upload when the bytes it downloads carry a
-- different ETag: the client's presigned PUT stays valid after finalize, so a
-- re-PUT in that window must not be processed as the bytes that were finalized.
-- The ETag only lived in that first job's payload. When the job is lost and the
-- stuck sweep requeues it from the row, the payload had no ETag and the check
-- was skipped. upload_etag keeps it on the row, written in the same UPDATE that
-- claims finalized_at, so the sweep's requeue carries it too.
--
-- BACK-COMPAT: rows finalized before this file keep NULL, and a NULL ETag means
-- the worker skips the comparison exactly as it did before. The column is an
-- opaque object-version string for the row's own bytes (no PII, no location),
-- so it needs no retention rule of its own: it lives and dies with the row.
--
-- HOT TABLE: media_assets. ADD COLUMN IF NOT EXISTS of a NULLable column with
-- NO default and NO index is a catalog-only change on Postgres 11+ (a brief
-- ACCESS EXCLUSIVE lock, no rewrite, no scan). No index: upload_etag is only
-- read off rows already selected by id through the stuck-sweep index.
--
-- CANONICAL DDL: hand-authored source of truth. Mirror: schema/media.ts.
-- Forward-only, no down. Requires 0001_core.sql (media_assets).
-- =============================================================================

ALTER TABLE media_assets
ADD COLUMN IF NOT EXISTS upload_etag text;
26 changes: 26 additions & 0 deletions services/api/drizzle/0185_inbound_bounce_attempts.sql
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
-- =============================================================================
-- 0185_inbound_bounce_attempts.sql
-- -----------------------------------------------------------------------------
-- WHY. A DSN stays under inbound/pending/ until its bounce bookkeeping succeeds,
-- so the next sweep can finish what a failed run left undone. One whose
-- bookkeeping fails every time used to stay there forever and take a slot in
-- every sweep's 200-object batch. This table counts the failed runs per pending
-- object; at the cap the processor parks the object under inbound/failed/ and
-- deletes the row. The storage seam has no custom object metadata to hold it.
--
-- BACK-COMPAT: new table, nothing reads it until the code that writes it ships.
-- A missing row means no failed run yet.
--
-- RETENTION: a row holds only the pending object's key (a Message-ID slug plus
-- a content digest, the same name the object already carries) and lives only
-- while that object is pending: success and parking both delete the row.
--
-- CANONICAL DDL: hand-authored source of truth. Mirror:
-- schema/inbound_bounce_attempts.ts. Forward-only, no down.
-- =============================================================================

CREATE TABLE IF NOT EXISTS inbound_bounce_attempts (
object_key text PRIMARY KEY,
attempts integer NOT NULL DEFAULT 0,
last_attempt_at timestamptz NOT NULL DEFAULT now()
);
Loading
Loading