This runbook covers the minimal production operating path for AgentBridge SaaS deployments: required environment, database backups, migrations, application startup, rollback caveats, smoke checks, and log inspection.
- Use a managed or externally operated PostgreSQL database for production.
- Treat
docker-compose.ymland its bundleddbservice as a local development or small self-host reference only. It is not a complete production database operations plan unless you add external backups, monitoring, restore drills, disk management, and access controls. - The production Docker image starts the Next.js web app and currently runs
prisma migrate deployindocker-entrypoint.shbefore starting the server. - The app listens on container port
3000and exposes an unauthenticated readiness endpoint atGET /api/health.
Set these runtime variables for the web app or container:
| Variable | Required | Notes |
|---|---|---|
DATABASE_URL |
Yes | PostgreSQL connection string. Include ?schema=public unless using a different Prisma schema intentionally. Use a production database user with only the privileges the app needs. |
AUTH_SECRET |
Yes | Long random secret for session signing. Keep stable across app restarts; rotating it invalidates existing sessions. |
NEXT_TELEMETRY_DISABLED |
Recommended | Set to 1 in production images and runtime. |
PORT |
Optional | Host-side Compose port. The container listens on 3000. |
Do not print, commit, or paste real database credentials, company bearer tokens, token hashes, AUTH_SECRET, .env files, or production logs containing secrets into issues or agent tasks.
- Confirm the release artifact or image tag you are deploying.
- Confirm the target database and environment are correct.
- Confirm
DATABASE_URLandAUTH_SECRETare present in the runtime secret store. - Take or verify a fresh database backup before running migrations.
- Review pending Prisma migrations for destructive operations or long-running locks.
- Ensure only one migration runner will execute for the release.
- Have rollback steps ready: previous app image/tag, backup location, and operator access.
Before every production migration:
- Trigger a provider snapshot or logical backup for the production PostgreSQL database.
- Record the backup identifier, timestamp, environment, and release/image tag being deployed.
- Confirm the backup completed successfully before migration starts.
- Prefer periodic restore drills in a separate database so backup validity is tested before an incident.
Example logical backup shape for self-managed PostgreSQL:
pg_dump "$DATABASE_URL" --format=custom --file="agentbridge-$(date -u +%Y%m%dT%H%M%SZ).dump"Store backups outside the app container and outside the primary database volume. Protect them with the same or stronger access controls as production data.
Prisma migrations are applied with:
corepack pnpm prisma migrate deployThe Docker image runs the equivalent startup step through docker-entrypoint.sh:
pnpm prisma migrate deployRun migrations as a single one-off job before starting or rolling out multiple app replicas:
docker run --rm \
-e DATABASE_URL="$DATABASE_URL" \
-e AUTH_SECRET="$AUTH_SECRET" \
ghcr.io/Vann-Dev/AgentBridge:<tag> \
pnpm prisma migrate deployThen start or roll the web app containers:
docker run --rm \
-p 3000:3000 \
-e DATABASE_URL="$DATABASE_URL" \
-e AUTH_SECRET="$AUTH_SECRET" \
-e NEXT_TELEMETRY_DISABLED="1" \
ghcr.io/Vann-Dev/AgentBridge:<tag>Because the current image entrypoint always runs prisma migrate deploy, every app container attempts the migration step at startup. This is acceptable for a single-container deployment and usually harmless when no migrations are pending, but it is not the safest multi-replica rollout pattern.
For multi-replica production deployments, avoid migration races by ensuring only one instance starts with the migration-capable entrypoint at a time, or by using platform controls/entrypoint overrides to run a dedicated migration job before scaling the web app. If your platform cannot separate migrations from app startup, document that limitation in the release notes and roll out one replica at a time.
The migration chain has a compatibility path for older non-empty AgentBridge databases that existed before company-scoped bearer tokens and stable AgentId values were introduced. Historical token migrations backfill deterministic placeholder values so prisma migrate deploy can complete instead of failing on required non-null columns.
After upgrading a legacy pre-token database:
- Sign in to the dashboard.
- Review agents and rename any generated
legacy-<uuid>AgentIds to the intended stable API identifiers before configuring external agents. - Rotate the company bearer token in Settings and update external agent configs with the newly shown token. The migration placeholders are not usable plaintext bearer tokens.
- Run the Agent API smoke check with the intended
AgentIdand new company token.
For Docker Compose with an external database:
DATABASE_URL="postgresql://USER:PASSWORD@HOST:5432/agentbridge?schema=public" \
AUTH_SECRET="replace-with-a-long-random-string" \
docker compose up --build appFor the local bundled PostgreSQL reference only:
docker compose --profile local-db up --buildDo not use the bundled Compose database as production unless you have added and tested backups, monitoring, restore, upgrades, and storage operations.
Run these checks after migrations and app startup.
curl -i https://<host>/api/healthUse the same unauthenticated command for Vercel preview/production URLs, Docker hosts, or local Compose by replacing the host, for example curl -fsS http://localhost:3000/api/health after docker compose up --build.
Expected ready response:
- HTTP
200 - JSON
statusCode: 200 - JSON
status: "healthy" - JSON
checks.app: "ok" - JSON
checks.database: "ok"
Expected degraded response when the app can answer but the database check fails:
- HTTP
503 - JSON
statusCode: 503 - JSON
status: "degraded" - JSON
checks.database: "unavailable"
The health response intentionally omits secrets, environment values, user/company data, stack traces, and raw database errors.
- Visit
/login. - Sign in with an existing operator account or complete the approved first-owner setup flow for a fresh database if enabled in the deployed version.
- Confirm
/dashboardloads. - Open Projects and one project board.
- Confirm task cards, Notes, Audit Logs, Docs, Agents, and Settings pages load as expected for the selected company.
Use a test company token and AgentId from the target company. Never paste real tokens into shared logs.
curl "$AGENTBRIDGE_BASE_URL/api/agent" \
-H "Authorization: Bearer $AGENTBRIDGE_COMPANY_TOKEN" \
-H "AgentId: $AGENTBRIDGE_AGENT_ID" \
-H "Accept: application/json"Expected response:
- HTTP
200 - JSON
statusCode: 200 agent.AgentIdmatches the requested agentagent.companymatches the intended company
Then list assigned tasks:
curl "$AGENTBRIDGE_BASE_URL/api/agent/tasks" \
-H "Authorization: Bearer $AGENTBRIDGE_COMPANY_TOKEN" \
-H "AgentId: $AGENTBRIDGE_AGENT_ID" \
-H "Accept: application/json"Expected response:
- HTTP
200 - JSON
statusCode: 200 tasksis an array scoped to the authenticated company and requesting agent.
For repeatable local, disposable, or production-like validation, run the scripted smoke against an AgentBridge instance that is safe to mutate. The script resolves the authenticated agent profile from /api/agent, uses that profile id as the task assignedAgentId UUID, creates a disposable task in the configured project, lists tasks, marks that task done, writes a note summary, and marks it read by the smoke agent.
Required environment:
| Variable | Notes |
|---|---|
AGENTBRIDGE_BASE_URL |
Base URL for the instance, for example http://localhost:3000 or a disposable preview URL. |
AGENTBRIDGE_COMPANY_TOKEN |
Company bearer token for the target company. Use a test/disposable token and never paste it into logs. |
AGENTBRIDGE_AGENT_ID |
API-facing AgentId sent in the AgentId header. The script resolves its database UUID from /api/agent before task creation. |
AGENTBRIDGE_PROJECT_ID |
Project ID where the disposable smoke task can be created and updated. |
AGENTBRIDGE_BASE_URL="http://localhost:3000" \
AGENTBRIDGE_COMPANY_TOKEN="<test-company-token>" \
AGENTBRIDGE_AGENT_ID="<test-agent-id>" \
AGENTBRIDGE_PROJECT_ID="<test-project-id>" \
corepack pnpm smoke:saas-productionExpected result:
GET /api/healthreturns HTTP200andchecks.database: "ok".GET /api/agentreturns the configured agent profile and provides the database UUID used asassignedAgentIdfor task creation.- Agent API task create/list/update all return success.
- Console output ends with
PASS smoke-saas-production taskId=<created-task-id>.
Only run this script against production when the configured project is explicitly approved for smoke data. Do not run it with customer data projects unless the resulting disposable task is acceptable and auditable.
Inspect platform/container logs for:
- Migration start/completion from
prisma migrate deploy. - Next.js app startup on
0.0.0.0:3000. - Repeated
GET /api/healthfailures or HTTP503responses. - Authentication failures that may indicate wrong company token or AgentId.
- Database connection timeouts or pool exhaustion.
Examples:
docker compose logs app
docker logs <container-id>For managed platforms, use the provider's deployment logs and runtime logs. Do not copy logs containing secrets or raw credentials into public channels.
Application rollback and database rollback are separate operations.
- Rolling back the app image to a previous tag does not roll back database migrations.
- Prisma does not provide automatic production migration rollback.
- If a migration causes bad schema/data state, choose between:
- restoring the database from the pre-deploy backup, or
- applying a forward-fix migration/code patch.
- Restoring a database reverts data to the backup timestamp and can lose writes made after that backup.
- Coordinate downtime or write freezes before restore when data loss risk matters.
Minimum rollback procedure:
- Stop or pause new app rollout.
- Decide whether app-only rollback is safe with the current schema.
- If app-only rollback is safe, deploy the previous known-good image/tag and run smoke checks.
- If database restore is required, stop writers, restore the chosen backup, deploy a compatible app version, and rerun smoke checks.
- Record the incident, root cause, release tag, migration names, and follow-up fixes.
Release:
Image/tag:
Commit:
Database backup id:
Migration command/result:
Deployed by:
Started at:
Completed at:
Health check: 200/503 details
Login/dashboard smoke: pass/fail
Agent API smoke: pass/fail
Rollback plan:
Notes/follow-ups: