Skip to content

fix(db): batch route reconfirmation to avoid ingest stalls - #180

Merged
MrAlders0n merged 1 commit into
devfrom
fix/route-reconfirm-batches
Sep 28, 2026
Merged

MrAlders0n merged 1 commit into
devfrom
fix/route-reconfirm-batches

Conversation

@MrAlders0n

Copy link
Copy Markdown
Member

What this PR does

Route reconfirmation held locks across a transaction of up to 750,000 routes. Production wait-state captures showed it blocking ingest route upserts, followed by timeouts and MQTT queue overflows.

Process 1,000 routes per transaction with a five-second batch timeout, retaining the 750,000-route run limit. Skip busy rows and use one cutoff per run so completed rows are not selected again, including when application time is ahead of database time. Existing stale/ambiguous route validation stays intact; busy rows and cancelled batches remain eligible for a later run.

Type of change

  • Bug fix
  • Tests

Checklist

  • go build ./... passes
  • gofmt -l . is empty
  • go vet ./... passes
  • go test ./... passes with local PostgreSQL integration tests enabled
  • Regression tests cover batching, deadlines, lock release, rollback, route validation and clock skew
  • sqlc generate and mock regeneration completed; no schema migration needed
  • No API, Swagger or dependency changes
  • I have read CONTRIBUTING.md

Testing notes

  • Race tests passed for db, internal/background and internal/ingest with PostgreSQL enabled.
  • Local replay with deliberately slow reconfirmation preserved all 99 packets, 4,950 observations and 14,850 live hearings with zero MQTT drops or ingest/enrichment errors. Stable stored and live enrichment fields matched the clean control, including decrypted messages and node capabilities. The old implementation dropped 27,603 MQTT messages in the same scenario; these include duplicates.
  • A local 750,000-route / 30,000-node benchmark completed in 17.82 seconds versus 10.98 seconds before. New batches had 28.1 ms p99 and 75.5 ms maximum duration. This trades additional total maintenance work for shorter lock lifetimes.

Local tests demonstrate preservation for the exercised cases, not a universal production losslessness guarantee. Busy routes may have validation deferred. Production has not been changed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant