diff --git a/AUDIT_OPEN.md b/AUDIT_OPEN.md index a4727a3..2ca712c 100644 --- a/AUDIT_OPEN.md +++ b/AUDIT_OPEN.md @@ -1943,3 +1943,343 @@ REMAINS (C01-1 slice 2): guarded-effect integration into the deploy path sidecar written by the state commit), the two-clients/delayed-SSH/clock -change acceptance matrix against the real deploy path, and the documented app-lock/shared-proxy acquisition order. +## Programme slice (2026-09-23) — C05: plan/apply binding + build identity + +The plan/apply remainder of C05 ("plan should distinguish known +effects from unresolved image/build data, cover routing/env/storage/ +resource changes, and bind apply to the reviewed config/target +version; drift invalidates stale plans"). The two Compose probe +defects and the field-classification table landed earlier (2026-09-21 +slice above); this slice builds on them. + +**Design:** + +- **AppliedManifestView** (internal/config/appliedview.go): the read + side of the manifest NormalizeAndDigest writes into release state — + the deployed half of every plan diff. Lives next to the writer so + the shape has one home; malformed manifests refuse (never guessed + from), null sections parse as absent (static/legacy representable). +- **Plan effect set** (internal/cli/planeffects.go, pure functions): + routing (domain set order/case-normalized, ingress mode, the + application port the Caddy route AND health gate probe, extra + publishes — diffed against AppState identity + recorded manifest), + env (KEY presence; values redacted by the manifest contract — only + key/env-file-reference changes are plannable, stated in the effect), + storage (volume add/remove/mount-change; the managed host path + named through plannedVolumeMounts — extracted from + deployBuiltImageFenced so plan and deploy share ONE volume + resolution), resources (replicas/memory/cpu), accessories + (set + image changes). +- **Known-vs-unresolved image classification** + (planImageIdentity, reusing resolveDeployProvenance — the C04 path + — so plan and deploy cannot disagree about what the build inputs + are): resolved-by-digest (pinned ref), resolved-by-image-id (mutable + ref, content resolved at plan time; recorded NOT bound — the + changed-mutable-tag policy is C04's open tail), unresolved-mutable- + tag, unresolved-awaiting-build (the plan binds context fingerprint + + Dockerfile sha instead of an image that does not exist yet). +- **PlanRecord** (internal/cli/planrecord.go): schema-versioned, + plan id = pure hash of the binding inputs (app/server/user/ + destination/target version/config digest/build-input identities/ + state generation+hash — WrittenAt deliberately outside it, the + provenance retry-stability rule), atomic 0600 write, load-time + self-consistency gate (tampered/future-schema plans refuse). + Config digest is NormalizeAndDigest over the plan's image reference + ("" for builds): recomputable BEFORE execution, unlike the receipt + digest which names the built image — the plan id stamped into the + receipt is the tie between the two. +- **`teploy apply`** (internal/cli/apply.go): re-derives config + (loader + recorded overlay + resolveDeployEnv — extracted from + runDeploy so applied and direct deploys share ONE env resolution), + version (explicit binds as-is; derived must re-derive), build + inputs (catches dirty-tree edits a version cannot see), identity, + and target state (the generation any deploy/rollback increments); + every mismatch refuses naming what moved + the remedy (re-plan). + Execution goes through deployAppConfig — the same function a + direct deploy runs (planID threaded to deployBuiltImageFenced; all + other callers pass ""). releasemeta.Provenance gains plan_id + (additive). Floating-tag plans refuse outright (unpredictable + version = unbindable). Drift refusals classify as the error + envelope's conflict code under --json — the taxonomy's first wired + conflict site; the plan-apply capability token is advertised + (registry + version-handshake golden regenerated, corpus rev 3, + plan-record schema + fixtures added). +- **plan --json compatibility**: the pre-plan/apply keys (app/server/ + target_version/version_known/same_version/changes) ride unchanged + in a compat envelope with the record fields additive — no MI bump. + +**Finding fixed en route:** the Compose importer silently DROPPED the +web service's `environment:` and `volumes:` (only accessories were +parsed) — a plan over an imported stack showed no env/storage effects +because the import had emptied them. Both now translate (web host +binds keep their full path; parseWebVolumes does not basename). +Regression: TestLoadCompose_WebServiceEnvAndVolumesPreserved. + +**Evidence** — new coverage: manifest view round-trips + refusal set +(config); plan id stability/sensitivity (10 mutation subtests); plan +file round-trip + tamper/schema refusals; binding verification for +every drift kind (config, overlay flip incl. strict presence-aware +clears, target-version incl. explicit-vs-derived, build-inputs +naming both fingerprints, target-state incl. generation move / +removed-since / deployed-since, identity); drift error-envelope +conflict classification; the C05 acceptance fixtures as tests — +config-changed-between-plan-and-apply refuses naming both digests, +deploy-happened-in-between refuses naming the generation move +(7 -> 8), nothing-moved verifies clean; an engine-level apply run +(the real deployBuiltImageFenced over a mock executor) asserting the +stamped provenance.json in the attempt namespace AND the committed +release state; Compose plan conformance (build → unresolved-awaiting- +build with bound fingerprint, digest-pinned → resolved-by-digest, +unresolved mutable tag, effect set survives import, refused shapes +never plan). Mutation checks (in-place, all reverted, gates re-run +green): dropping prov.PlanID assignment fails the engine stamp test; +removing the generation comparison from verifyPlanBinding fails +TestApplyDrift_DeployHappenedInBetween; making computePlanID ignore +the config digest fails the sensitivity table; reverting the web-env +translation fails TestComposePlan_EffectSetSurvivesImport and the +importer regression. + +**C05 remainder (explicit):** stack resource breadth (multi-image +stacks still refuse at import — the single-image process model is +unchanged; the first-class stack resource is P1); Coolify/Dokploy- +scale Compose breadth (fields outside the classification inventory +remain silently ignored); plan/apply is single-target (multi-server +fleet apply via the scale path would need per-server plan records); +static deploys verify+execute but write no provenance receipt, so no +plan-id stamp (runStaticDeploy predates provenance); env VALUES are +outside the binding by design (manifest redaction) — a changed .env +value between plan and apply is not drift, only key/reference changes +are; accessory volume import basenames host-bind sources (parseServiceVolumes +— pre-existing, untouched; the web translator preserves them +properly and the accessory behavior is recorded here as a finding). + +Gates: `go build ./...` clean; `go vet ./...` clean; `go test ./... -race -count=1` all 25 packages ok; gofmt clean on every touched hunk (pre-existing strays in deploy.go/secret_audit.go/update_test.go/contracts_golden_test.go left alone, consistent with the C02-C04 posture); contracts corpus regenerated deliberately (corpus rev 3) and the non-update test run pins it. +||||||| 0c1fe5d +## Programme slice (2026-09-23) — C01-1: replacement-owner reconciliation on acquisition + +Closes the C01-1 disagreement (docs/C01_RECOVERY_STATE_TABLE.md finding 1): +lock acquisition used to be treated as quiescence — `acquireAutoLock` broke +a stale lock and the deploy proceeded with no observation of the dead +holder's leftover world. Base revision `0c1fe5d`. + +**Design:** + +- **Takeover signal** — `state.acquireAutoLock` now reports whether the + acquisition broke a stale auto/heal lock; `state.Lock.TookOver()` exposes + it (nil lock: false). Fresh acquisitions are unchanged. +- **Productionized observer** — `deploy.Observe` (internal/deploy/ + reconcile.go) is the fault harness's evidence collector as production + code: docker label inventory, state.json, the managed Caddyfile, the + per-release record → recovery.Observation, exact names, read failures map + to Unknown (the never-auto-decide grade). The decision stays the pure + table's (recovery.Decide); the observer imports the effectful packages, + not the reverse. +- **Reconciliation gate** — DeployFenced step 1c: when TookOver, run + `ReconcileAfterTakeover` BEFORE the deploy's first effect. RETRY is the + only proceed disposition (surfaced to the operator); INSPECT gets ONE + bounded re-observation (R4's transient-read reconcile trigger) then + refuses; COMPENSATE/MANUAL refuse immediately. Every refusal carries the + observed evidence classes and the inspect commands. Deliberately NOT an + auto-compensator: compensation automation is the F04-keyed recovery-owner + continuation; refusing with evidence is the safe subset the table + permits. The reconciliation precedes the first docker run, so a refusal + has nothing to undo. + +**Evidence** — mock tests (internal/deploy/reconcile_test.go): takeover +with a foreign running workload refuses MANUAL with zero `docker run`/state +commits; clean-world takeover proceeds (a crash must not make the app +undeployable); same-version disagreement (state.json names the deploying +release) refuses INSPECT per R6; running-candidate-without-receipts INSPECT; +traffic-on-uncommitted-generation COMPENSATE; persistent unreadable stays +INSPECT; a transient inventory failure recovers via the single +re-observation; Observe classification pinned (_replaced rename = serving +predecessor, stopped = restorable, corpse ≠ running candidate, unreadable = +Unknown). Fixture-verified for real (internal/deploy/ +reconcile_integration_test.go, colima docker 29.5.2): owner A's nohup'd +delayed candidate lands after owner B's genuine stale-break acquisition; +TookOver reports true; the production reconciler refuses MANUAL — the +quiescence assumption's RETRY is proven dead in production shape. The +pre-existing fault harness passes unchanged against the same fixture. +Gates: build/vet clean; `go test ./... -count=1` all packages ok (one +pre-existing load-sensitive timing test, cli TestAdmission_NoGoroutinePileup, +flaked once under full-suite parallel load and passes repeatedly in +isolation and in two follow-up full runs — not touched by this slice). + +**Residual C01 list (updated):** C01-2/3 remain (guarded pre-commit +effects, fenced shared Caddy lock); C01-8/C01-9 unchanged (F04-keyed). + +## Programme slice (2026-09-23) — C01-2: guarded pre-commit effects + +Closes the C01-2 disagreement (docs/C01_RECOVERY_STATE_TABLE.md finding +2): candidate starts, worker starts and the route switch ran `lk.Check` +as a command SEPARATE from the effect — only the state commit composed +guard+effect (WriteFenced). Between check and effect a takeover could +occur, letting a broken holder's effects land inside the new owner's +window. Base revision `194130c`. + +**Design:** + +- **Composition surface** — `state.Lock.GuardPrefix()` returns the shell + prefix that refuses (TEPLOY_FENCE_LOST marker, exit 75) when the lock + no longer names the holder; `state.FenceLost(err)` matches refusals + across packages. docker and caddy consume the PREFIX, not the Lock — + no new package coupling. +- **Container starts** — `docker.RunGuarded(ctx, cfg, guardPrefix)` + composes the guard with the docker run in ONE remote command; + DeployFenced's web-candidate and worker starts use it (the separate + pre-effect Checks at those sites are superseded). A refused start maps + to `state.ErrFenceLost` and enters the existing recovery handlers + (restoreDisplacedAndStarted / fail); the reactive + name-already-in-use clearing also runs under the guard. +- **Route switch** — `caddy.Client.WithCommitGuard(prefix)` returns a + client whose Caddyfile COMMIT (the rename that makes new contents + authoritative) runs composed under the guard. mutate was restructured: + stage the new Caddyfile to an inert random sibling (upload, no + effect), then one guarded `mv -fT` commit; reload and delivery + verification stay separate commands (the commit is the traffic-switch + instant — a refused commit means the edit never became authoritative, + and a post-commit reload is idempotent). Rollback restores are + deliberately never fenced (cleanup is never fenced, F16's rule). + Wired through deploy's step 11, rollback, and all three static + SetStaticRoute sites (each holds its own fence). + +**Evidence** — mock tests: happy path asserts the web+worker runs and the +Caddyfile commit execute composed under the guard (guard+effect in one +command); a takeover fired at the health gate (lock info rewritten +mid-deploy by a test executor wrapper) refuses the route switch — no +Caddyfile lands, no reload runs, deploy errors ErrFenceLost; the same +takeover refuses the worker start in-shell with no executed worker run; +the pre-existing late-holder test now asserts on EXECUTED (bare) docker +runs — a refused start appears only as the composed command the guard +rejected, which is the C01-2 shape, not a regression. Caddy-side: a +guarded client whose lock names another owner has its edit refused +(nothing lands, no reload, lock still acquired/released around the +attempt); the legitimate holder's composed commit is byte-identical in +effect to today's write. All pre-existing caddy/docker/deploy/rollback/ +static/preview tests unchanged and green. Gates: build/vet clean; +`go test ./... -count=1` all packages ok; gofmt clean on touched files +(docker_test.go/recreate.go pre-existing strays left alone). + +**Residual C01 list (updated):** C01-3 remains (fenced shared Caddy +lock); C01-8/C01-9 unchanged (F04-keyed). + +## Programme slice (2026-09-23) — C01-3: owner-tagged fenced shared Caddy lock + conflicting-route reconciliation + +Closes the C01-3 disagreement (docs/C01_RECOVERY_STATE_TABLE.md finding +3): the shared-proxy lock was a bare ownerless mkdir broken by DIRECTORY +MTIME after 120s, so a slow-but-alive orphaned editor could interleave +its Caddyfile edit with the new owner's mutate, and the table's +"conflicting route evidence → INSPECT" had no producer or consumer. Base +revision `13c7dd9`. + +**Design:** + +- **Owner-tagged lock** — caddy.acquireLock writes a caddy-edit info + file (owner token + RFC3339 ts, mirroring the app locks' shape). + Staleness is measured from the INFO timestamp (unparseable = stale); + a legacy no-info dir falls back to the old mtime check. No renewal: + an edit session is seconds, far below any renewal interval. TTL stays + 120s. +- **Fenced commit** — the Caddyfile commit's composed command now chains + the CADDY-LOCK guard after the app-fence guard (C01-2's prefix): a + holder whose lock was stale-broken has its late edit refused in-shell + (TEPLOY_FENCE_LOST / exit 75), so a broken editor cannot interleave + with the successor. Acquisition order documented and test-pinned: + app guard FIRST (long-held), caddy guard SECOND (brief) — never the + inverse, never across hosts. +- **Conditional release** — releaseLock removes the lock only when its + info still names the releaser (the app locks' A04 lesson): a stale + holder's deferred release can no longer delete the successor's lock + and admit a third editor. +- **The missing reconciliation** — deploy.Observe now classifies the + route evidence exactly: a managed block naming the attempted + candidates or the predecessor's containers is provable; a managed + block naming a THIRD generation (the dead-holder late-route-edit + shape) is CONFLICTING (Unknown), which Decide routes to INSPECT (R4) + — previously it collapsed into a false "route to predecessor" that + could yield a blind RETRY. No managed block at all is clean absence. + +**Evidence** — caddy tests: acquisition writes owner info and breaks a +stale holder BY INFO AGE; a fresh (sub-TTL) holder is never broken; a +broken holder's composed commit is refused and nothing lands; release +against a successor-owned lock deletes nothing. Deploy test pins the +app-guard-before-caddy-guard order on the commit command. Reconcile +test: a third-generation route under predecessor authority yields +Unknown route evidence → INSPECT; no managed block stays clean absence. +The mock executor evaluates CHAINED guards (every guard must hold) and +the stateful caddy fake gained the same evaluation. All pre-existing +suites updated for the conditional-release command shape (assertions +moved from `rmdir` to the conditional) and green. Gates: build/vet +clean; `go test ./... -count=1` all packages ok; gofmt clean on touched +files. + +**C01 locking-protocol redesign (C01-1/2/3) is now closed.** Residual +C01: C01-8/C01-9 (F04 generation identities — deliberate containment), +the A12/T05 rollback-route remainder of C01-7. + +## Programme slice (2026-09-23) — C03: request drain + the readiness/liveness/stop/drain distinction + +First bounded slice of C03's ingress-behavior half (the readiness-mode +half landed earlier as the F47/TCL-17/A22 slice). Closes the recorded +C03 remainder "request drain + graceful stop": the drain WINDOW between +the traffic switch and predecessor retirement, the stop policy surfaced +distinct from readiness, and the declared blue/green fixture proving +zero failed requests. Base revision `90f356c`. + +**Design:** + +- **Grammar** — `drain_seconds: N` in teploy.yml/TOML (0..600; 0 = the + historical stop-immediately default, documented compat). Flows through + the destination overlay and the effective-config manifest (drift + identity — a changed drain window is a config change), mapped into + deploy.Config / RollbackConfig at the single config→deploy seam plus + rollback and scale. +- **Deploy** — after the route switch and state commit, before + predecessor retirement: the configured window elapses while the + predecessor keeps serving IN-FLIGHT requests on its existing + connections (new traffic is on the candidates). Only a caddy-routed + blue/green switch drains: external ingress is the operator's edge, + the recreate strategy already stopped the fixed-port workload before + the candidates started. Cancellation cuts the window short and + proceeds to retirement (the operator asked to stop). +- **Rollback** — the same window between the route switch back to the + target and stopping the superseded generation. +- **Surfaced distinction** — deploy prints the gate AND the stop policy + before the switch: readiness (health mode + total deadline), graceful + stop (stop_timeout's SIGTERM→SIGKILL ladder), request drain (the + window), liveness (the container HEALTHCHECK directive). The deploy + plan now states all four separately. +- **Honesty** — Caddy's config-level routing CANNOT count in-flight + requests per upstream (no per-upstream concurrency endpoint, and + teploy's blocks do not enable access logging), so the drain POLICY is + the time window plus the ladder — documented in README, the config + comment, and the deploy output. Nothing promises request counting. + +**Evidence** — mock tests: the draining deploy proves reload→window→stop +ordering, the window elapses (elapsed >= drain_seconds), the policy +lines surface, drain 0 adds no window (compat), and external ingress +never drains. Config tests: parse, bounds (0/1/600 ok; -1/601 rejected +naming the field), overlay, manifest identity. Rollback test: window +between switch-back and superseded stop. REAL fixture (colima docker +29.5.2, real caddy:2-alpine, the PRODUCTION caddy.Client switch path — +lock, adapt gate, guarded commit, reload, delivery verification): +TestDrainIntegration_BlueGreenZeroFailedRequests drives 90 requests +across the switch with a 3.5s request in flight, drains 5s, stops blue +with -t 5 — ZERO failed requests, the long request completes +end-to-end on blue inside the window, post-switch traffic serves green. +TestDrainIntegration_NoDrainKillsLongRequest is the negative control: +stop -t 0 immediately after the switch breaks the in-flight request +(caddy 502) — the fixture can detect a broken promise, so the window is +what saves it. Integration tests also taught the close-vs-cleanup lesson +(SSH sessions close AFTER t.Cleanup work, LIFO — a plain defer raced +the fixture cleanups silently). Gates: build/vet clean; +`go test ./... -count=1` all packages ok; gofmt clean on touched files +(pre-existing strays untouched); integration battery green against the +colima fixture. + +**C03 remainder (explicit):** liveness-vs-readiness as post-switch +continuous probing (today only the container HEALTHCHECK approximates +it), multi-host partial-wave readiness states (canary aggregate gating), +WebSocket/SSE drain verification beyond the long-request proof, and the +registered interactions (tcp mode × the Caddy LB active check; preview's +gate has no drain surface). diff --git a/CHANGELOG.md b/CHANGELOG.md index 12e51a5..d0ffc63 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,34 @@ All notable changes to teploy are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). +## [Unreleased] + +### Added + +- **Plan/apply with drift invalidation (C05).** `teploy plan` now renders + the full effect set — routing (domain/ingress/port/publishes), + environment keys, storage volumes, resource limits, accessories — + alongside the container diff, and classifies image data honestly: + resolved-by-digest, mutable-tag-resolved-at-plan-time, or + unresolved-awaiting-build (a build plan binds its build inputs — + context fingerprint + Dockerfile identity — instead of guessing). + `teploy plan --out FILE` records a plan identity (effective-config + digest, target version, server/app, deployed-state generation). + `teploy apply FILE` re-derives every binding input and refuses, + naming what drifted (config edit, overlay flip, moved git version, + changed build inputs, or any deploy/rollback in between — the state + generation), then executes through the normal deploy engine with the + plan id stamped into the release's provenance receipt. A floating-tag + plan (timestamp version) is refused outright — it can never be + bound. Under `--json` a drift refusal classifies as the error + envelope's `conflict` code; the `plan-apply` capability token is + advertised. +- **Compose import: the web service's `environment:` and `volumes:` are + now translated** into `env:`/`volumes:` (host binds keep their full + path). Both were silently DROPPED — only accessories were parsed — + so an imported stack deployed without any of the web service's + environment or storage. + ## [0.1.37] - 2026-09-22 ### Fixed diff --git a/README.md b/README.md index 3772a4e..99e11a9 100644 --- a/README.md +++ b/README.md @@ -122,6 +122,12 @@ port: 3000 build_local: true platform: linux/amd64 stop_timeout: 30 +drain_seconds: 10 # request drain: keep the old container serving in-flight + # requests this long AFTER the traffic switch, before it is + # stopped (0 = stop immediately, the default). The full graceful + # budget is drain_seconds + stop_timeout. Teploy's Caddy routing + # cannot count in-flight requests per upstream, so the window IS + # the drain policy — size it to your longest normal request. memory: 2g # cgroup RAM cap (docker units: 512m, 2g). Unset = unlimited. cpu: "1.5" # CPU cap in cores. Unset = unlimited. keep_versions: 3 # auto-prune older versions after deploy (0 = keep all, default) @@ -336,6 +342,9 @@ teploy exec # run a command on the server (SSH) teploy app exec -- # run a command in the app container (migrations, etc.) teploy validate # check config and server readiness teploy doctor [--server ] # read-only diagnostics: toolchain, SSH, Docker, registry, Caddy, disk, compatibility, repair debt (--json for machines; exit 1 if any check fails, never 2) +teploy plan # preview what a deploy would change (containers, routing, env, storage, resources; read-only) +teploy plan --out plan.json # write a bound plan record (config digest + target identity + generation) +teploy apply plan.json # execute a reviewed plan; refuses naming what drifted since it was made teploy scale # multi-server deploy + LB update teploy version / update # version info and self-update ``` diff --git a/contracts/MANIFEST.md b/contracts/MANIFEST.md index f619242..070d1be 100644 --- a/contracts/MANIFEST.md +++ b/contracts/MANIFEST.md @@ -10,8 +10,10 @@ Neutron/Nucleus dependency and a public mirror. | Corpus rev | Emitting CLI | Machine Interface | Notes | |---|---|---|---| +| 3 (amended) | main (C05 plan-record corpus + defect fix) | 1 | C05 added the plan-record artifact + plan-apply token (see git history); amendment: server-status schema now carries its own $defs (its $refs never resolved), and app-list fixtures emit [] where the encoder emits [] (null fixtures failed schema + the real dash decode - found by dash's new contracts CI job, fixed here). | | 1 | post-v0.1.37 main (S2 skeleton) | 1 | First goldens: version handshake, app-list envelope (MI + pre-MI legacy), error envelope (config-invalid, internal, invalid-code), release-record, attempt-name grammar, preview-state eras. | | 2 | post-v0.1.37 main (S6) | 1 | observation-envelope: schema corrected from the S2 draft shape to the ADR §2.4 canonical form (resource/collected_at/freshness tri-state/error/source/last_known) before any consumer existed; fixtures generated from teploy-dash's real constructors (fresh, stale, unknown-unreachable, unreachable-last-known). | +| 3 | post-v0.1.37 main (C05) | 1 | plan-record: the `teploy plan --out` / `teploy apply` binding record (build + prebuilt-digest valid fixtures, tampered-id invalid fixture); `plan-apply` capability token added to the version handshake (additive). | ## Artifact status @@ -25,6 +27,7 @@ Neutron/Nucleus dependency and a public mirror. | attempt-name | yes (pattern) | valid + invalid examples | teploy-cli | | preview-state | yes (canonical/legacy) | valid + legacy + ambiguous | teploy-cli | | observation-envelope | yes (§2.4 canonical, rev 2) | valid x4 (fresh, stale, unknown-unreachable, unreachable-last-known; dash encoder) | teploy-dash | +| plan-record | yes | valid x2 (build unresolved-awaiting-build, prebuilt resolved-by-digest) + invalid tampered-id | teploy-cli | | operation-record | yes | pending S5/S6 (dash) | teploy-dash | ## Rules diff --git a/contracts/fixtures/app-list-envelope/legacy/pre-mi.json b/contracts/fixtures/app-list-envelope/legacy/pre-mi.json index a50e74d..0672fbe 100644 --- a/contracts/fixtures/app-list-envelope/legacy/pre-mi.json +++ b/contracts/fixtures/app-list-envelope/legacy/pre-mi.json @@ -1,6 +1,6 @@ { "apps": [], - "errors": null, + "errors": [], "host": "srv.example.com", "observed_at": "2026-09-23T12:00:00Z" } diff --git a/contracts/fixtures/app-list-envelope/valid/mi1.json b/contracts/fixtures/app-list-envelope/valid/mi1.json index d280539..e81fa42 100644 --- a/contracts/fixtures/app-list-envelope/valid/mi1.json +++ b/contracts/fixtures/app-list-envelope/valid/mi1.json @@ -14,8 +14,10 @@ ] }, "previous_release": { - "version": "", - "ports": null + "version": "2", + "ports": [ + 3000 + ] }, "containers": [ { @@ -29,13 +31,13 @@ "version": "3" } ], - "processes": null, + "processes": [], "lock": null, "maintenance": false, "observed_at": "2026-09-23T12:00:00Z", - "errors": null + "errors": [] } ], "observed_at": "2026-09-23T12:00:00Z", - "errors": null + "errors": [] } diff --git a/contracts/fixtures/plan-record/invalid/tampered-id.json b/contracts/fixtures/plan-record/invalid/tampered-id.json new file mode 100644 index 0000000..bfce6ba --- /dev/null +++ b/contracts/fixtures/plan-record/invalid/tampered-id.json @@ -0,0 +1,29 @@ +{ + "schema_version": 1, + "plan_id": "e6e098f1e0322e3f", + "app": "myapp", + "server": "srv.example.com", + "server_name": "prod", + "target_version": "sha256-aaaaaaaaaaaa", + "version_known": true, + "config_digest": "0000000000000000000000000000000000000000000000000000000000000000", + "image": { + "ref": "registry.example.com/myapp@sha256:abababababababababababababababababababababababababababababababab", + "resolution": "resolved-by-digest", + "digest": "sha256:abababababababababababababababababababababababababababababababab" + }, + "target_state": { + "deployed": false, + "generation": 0 + }, + "effects": { + "containers": [ + { + "action": "create", + "name": "myapp-web-sha256-aaaaaaaaaaaa", + "detail": "web container" + } + ] + }, + "written_at": "0001-01-01T00:00:00Z" +} diff --git a/contracts/fixtures/plan-record/valid/build.json b/contracts/fixtures/plan-record/valid/build.json new file mode 100644 index 0000000..1eed2b8 --- /dev/null +++ b/contracts/fixtures/plan-record/valid/build.json @@ -0,0 +1,48 @@ +{ + "schema_version": 1, + "plan_id": "f5f4997ca665419e", + "app": "myapp", + "server": "srv.example.com", + "user": "root", + "server_name": "prod", + "target_version": "abc1234", + "version_known": true, + "config_digest": "3f2a9c11d8e4b7065a1c9f0e2b8d7a64c5e3f1b9a0d8c7e6f5a4b3c2d1e0f9a8", + "image": { + "needs_build": true, + "resolution": "unresolved-awaiting-build", + "context_path": ".", + "context_fingerprint": "c0ffee11aa22bb33", + "dockerfile": "Dockerfile", + "dockerfile_sha256": "deadbeef11", + "platform": "linux/amd64" + }, + "target_state": { + "deployed": true, + "generation": 4, + "current_hash": "old1234", + "manifest_sha256": "aa11bb22" + }, + "effects": { + "containers": [ + { + "action": "create", + "name": "myapp-web-abc1234", + "detail": "web container" + } + ], + "routing": [ + { + "action": "change", + "name": "domain", + "from": "old.example.com", + "to": "new.example.com", + "detail": "routes served by this deployment" + } + ] + }, + "unresolved": [ + "image unresolved — built at deploy time; the plan binds the build inputs (context . fingerprint c0ffee11aa22bb..., Dockerfile Dockerfile sha deadbeef11...)" + ], + "written_at": "0001-01-01T00:00:00Z" +} diff --git a/contracts/fixtures/plan-record/valid/prebuilt-digest.json b/contracts/fixtures/plan-record/valid/prebuilt-digest.json new file mode 100644 index 0000000..d4d3731 --- /dev/null +++ b/contracts/fixtures/plan-record/valid/prebuilt-digest.json @@ -0,0 +1,29 @@ +{ + "schema_version": 1, + "plan_id": "e6e098f1e0322e3f", + "app": "myapp", + "server": "srv.example.com", + "server_name": "prod", + "target_version": "sha256-aaaaaaaaaaaa", + "version_known": true, + "config_digest": "7d1e0f9a8b7c6d5e4f3a2b1c0d9e8f7a6b5c4d3e2f1a0b9c8d7e6f5a4b3c2d10", + "image": { + "ref": "registry.example.com/myapp@sha256:abababababababababababababababababababababababababababababababab", + "resolution": "resolved-by-digest", + "digest": "sha256:abababababababababababababababababababababababababababababababab" + }, + "target_state": { + "deployed": false, + "generation": 0 + }, + "effects": { + "containers": [ + { + "action": "create", + "name": "myapp-web-sha256-aaaaaaaaaaaa", + "detail": "web container" + } + ] + }, + "written_at": "0001-01-01T00:00:00Z" +} diff --git a/contracts/fixtures/version-handshake/valid/mi1.json b/contracts/fixtures/version-handshake/valid/mi1.json index def2d5d..785f22a 100644 --- a/contracts/fixtures/version-handshake/valid/mi1.json +++ b/contracts/fixtures/version-handshake/valid/mi1.json @@ -7,6 +7,7 @@ "error-envelope", "health-modes", "kv-set-stdin", + "plan-apply", "preview-blue-green", "preview-canonical-id", "provenance-records", diff --git a/contracts/schema/plan-record.schema.json b/contracts/schema/plan-record.schema.json new file mode 100644 index 0000000..e581bb7 --- /dev/null +++ b/contracts/schema/plan-record.schema.json @@ -0,0 +1,81 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "https://teploy.github.io/contracts/schema/plan-record.schema.json", + "title": "C05 plan record (teploy plan --out; cli/planrecord.go PlanRecord)", + "type": "object", + "required": ["schema_version", "plan_id", "app", "server", "target_version", "version_known", "config_digest", "image", "target_state", "effects"], + "properties": { + "schema_version": {"type": "integer", "minimum": 1}, + "plan_id": {"type": "string", "pattern": "^[a-f0-9]{16}$"}, + "app": {"type": "string", "minLength": 1}, + "server": {"type": "string", "minLength": 1}, + "user": {"type": "string"}, + "server_name": {"type": "string"}, + "destination": {"type": "string"}, + "target_version": {"type": "string"}, + "version_known": {"type": "boolean"}, + "version_explicit": {"type": "boolean"}, + "config_digest": {"type": "string", "pattern": "^[a-f0-9]{64}$"}, + "image": { + "type": "object", + "required": ["resolution"], + "properties": { + "ref": {"type": "string"}, + "needs_build": {"type": "boolean"}, + "resolution": {"type": "string", "enum": ["resolved-by-digest", "resolved-by-image-id", "unresolved-mutable-tag", "unresolved-awaiting-build"]}, + "digest": {"type": "string"}, + "context_path": {"type": "string"}, + "context_fingerprint": {"type": "string"}, + "dockerfile": {"type": "string"}, + "dockerfile_sha256": {"type": "string"}, + "platform": {"type": "string"} + } + }, + "target_state": { + "type": "object", + "required": ["deployed", "generation"], + "properties": { + "deployed": {"type": "boolean"}, + "generation": {"type": "integer", "minimum": 0}, + "current_hash": {"type": "string"}, + "manifest_sha256": {"type": "string"} + } + }, + "effects": { + "type": "object", + "required": ["containers"], + "properties": { + "containers": {"type": "array", "items": {"$ref": "#/$defs/planChange"}}, + "routing": {"type": "array", "items": {"$ref": "#/$defs/planEffect"}}, + "env": {"type": "array", "items": {"$ref": "#/$defs/planEffect"}}, + "storage": {"type": "array", "items": {"$ref": "#/$defs/planEffect"}}, + "resources": {"type": "array", "items": {"$ref": "#/$defs/planEffect"}}, + "accessories": {"type": "array", "items": {"$ref": "#/$defs/planEffect"}} + } + }, + "unresolved": {"type": "array", "items": {"type": "string"}}, + "written_at": {"type": "string", "format": "date-time"} + }, + "$defs": { + "planChange": { + "type": "object", + "required": ["action", "name"], + "properties": { + "action": {"type": "string", "enum": ["create", "stop", "unchanged"]}, + "name": {"type": "string"}, + "detail": {"type": "string"} + } + }, + "planEffect": { + "type": "object", + "required": ["action", "name"], + "properties": { + "action": {"type": "string", "enum": ["add", "remove", "change"]}, + "name": {"type": "string"}, + "from": {"type": "string"}, + "to": {"type": "string"}, + "detail": {"type": "string"} + } + } + } +} diff --git a/contracts/schema/server-status-envelope.schema.json b/contracts/schema/server-status-envelope.schema.json index 09a9a17..7e43687 100644 --- a/contracts/schema/server-status-envelope.schema.json +++ b/contracts/schema/server-status-envelope.schema.json @@ -73,5 +73,141 @@ } } } - ] + ], + "$defs": { + "machineError": { + "type": "object", + "required": [ + "scope", + "message" + ], + "properties": { + "scope": { + "type": "string" + }, + "message": { + "type": "string" + } + } + }, + "container": { + "type": "object", + "required": [ + "id", + "name", + "image", + "state", + "status" + ], + "properties": { + "id": { + "type": "string" + }, + "name": { + "type": "string" + }, + "image": { + "type": "string" + }, + "state": { + "type": "string" + }, + "status": { + "type": "string" + }, + "created_at": { + "type": "string" + }, + "process": { + "type": "string" + }, + "version": { + "type": "string" + } + } + }, + "release": { + "type": "object", + "required": [ + "version", + "ports" + ], + "properties": { + "version": { + "type": "string" + }, + "ports": { + "type": "array", + "items": { + "type": "integer" + } + } + } + }, + "appStatus": { + "type": "object", + "required": [ + "app", + "domain", + "type", + "ingress", + "current_release", + "previous_release", + "containers", + "processes", + "lock", + "maintenance", + "observed_at", + "errors" + ], + "properties": { + "app": { + "type": "string" + }, + "domain": { + "type": "string" + }, + "type": { + "type": "string" + }, + "ingress": { + "type": "string" + }, + "current_release": { + "$ref": "#/$defs/release" + }, + "previous_release": { + "$ref": "#/$defs/release" + }, + "containers": { + "type": "array", + "items": { + "$ref": "#/$defs/container" + } + }, + "processes": { + "type": "array" + }, + "lock": { + "type": [ + "object", + "null" + ] + }, + "maintenance": { + "type": "boolean" + }, + "observed_at": { + "type": "string", + "format": "date-time" + }, + "errors": { + "type": "array", + "items": { + "$ref": "#/$defs/machineError" + } + } + } + } + } } diff --git a/docs/C01_RECOVERY_STATE_TABLE.md b/docs/C01_RECOVERY_STATE_TABLE.md index cdb0ad2..f915523 100644 --- a/docs/C01_RECOVERY_STATE_TABLE.md +++ b/docs/C01_RECOVERY_STATE_TABLE.md @@ -129,7 +129,15 @@ table's, with the register item it belongs to. table requires: a replacement owner runs `Decide` over observed evidence after acquisition. Register: A05/T01 adjacent but distinct — fencing refuses stale *check-then-act* holders; this is the new owner's - side (nobody observes the leftover world). + side (nobody observes the leftover world). **LANDED 2026-09-23** (see + AUDIT_OPEN's C01 replacement-owner-reconciliation slice): acquisitions + report takeover (`Lock.TookOver`), `DeployFenced` runs the + productionized observer + `Decide` (`internal/deploy/reconcile.go`) + before its first effect, RETRY is the only proceed disposition, and + every other one refuses with the observed evidence. The observer is + the harness's evidence collector made production code; verified + against the real fixture (late effect after takeover → MANUAL + refusal). 2. **C01-2 — Pre-commit effects are check-then-act, not guarded.** `lk.Check` runs as a separate command from the effect it guards: @@ -141,6 +149,16 @@ table's, with the register item it belongs to. new owner's window — the table treats "effect lands after owner death" as INSPECT-at-best evidence, which nothing today generates. Register: A05/T01 standing; the table now states the disposition consequence. + **LANDED 2026-09-23** (see AUDIT_OPEN's C01 guarded-effects slice): + `docker.RunGuarded` composes the holdership guard with the container + creation in one remote command (deploy's candidate and worker starts; + `state.Lock.GuardPrefix`/`state.FenceLost` are the composition + surface), and the Caddyfile commit rename runs under the same guard + (`caddy.Client.WithCommitGuard`, threaded through deploy, rollback and + the static paths). A mid-flight takeover is refused in-shell: no + container starts, no route edit lands, no reload runs. The separate + pre-effect Checks at those three sites are superseded by the + composition. 3. **C01-3 — The shared Caddy lock is ownerless and unfenced.** `internal/caddy/caddy.go:684-697` breaks any caddy lock older than 120s @@ -149,7 +167,16 @@ table's, with the register item it belongs to. "conflicting route evidence → INSPECT" has no producer/consumer today (nobody reconciles a Caddyfile that names containers no inventory can attribute). Register: T03 documents the design as deliberate; the - finding is the missing reconciliation, not the lock's shape. + finding is the missing reconciliation, not the lock's shape. **LANDED + 2026-09-23** (see AUDIT_OPEN's C01 shared-proxy-lock slice): the lock + carries an owner-tagged info file (staleness from its timestamp, not + directory mtime; legacy no-info dirs keep the mtime fallback), the + commit is fenced by the lock's own guard composed AFTER the app-fence + guard, release is conditional on ownership (the app locks' A04 + lesson), and the missing producer/consumer exists: + `deploy.Observe` classifies a managed block naming a third generation + as CONFLICTING (Unknown) route evidence, which `Decide` sends to + INSPECT (R4) instead of the old false "route to predecessor". 4. **C01-4 — No durable readiness receipt.** The health gate (`internal/deploy/health.go`, `deploy.go:646-658`) persists nothing, so @@ -269,9 +296,17 @@ Two lock layers, never nested across hosts: renewal): serialize an app's lifecycle on ONE target. Held for the whole lifecycle, but only ever against one host. 2. **The shared-proxy commit lock** (`/deployments/caddy/.lock`, - `internal/caddy/caddy.go:684-707`): short-lived, ownerless by design - (T03), held only for the brief Caddyfile edit+reload+verify inside one - host's traffic-switch step. + `internal/caddy/caddy.go`): short-lived, held only for the brief + Caddyfile edit+reload+verify inside one host's traffic-switch step. + Since C01-3 (2026-09-23) it is OWNER-TAGGED and FENCED like the app + locks: acquisition writes an owner info file, staleness is measured + from that info's timestamp (not directory mtime), the Caddyfile + commit runs under the lock's own guard composed after the app-fence + guard, and release removes the lock only when its info still names + the releaser. Acquisition order on the commit command is therefore + APP GUARD THEN CADDY GUARD — app lock first (long-held), shared + commit lock second (brief); never the inverse, and a holder that + loses either fence has its commit refused in-shell. **Never hold one host's app lock while waiting on another host's.** The current code complies: locks are acquired inside each host's @@ -336,11 +371,16 @@ Implementation slices: **C01-4, C01-5, C01-10 landed 2026-09-22** (attempt-journal receipts + honest degraded log outcome; evidence in AUDIT_OPEN's C01 implementation-slice section) and **C01-6, C01-7 landed 2026-09-22** (record-repair debt reconciler + receipt-driven route -compensation; evidence in AUDIT_OPEN's latest C01 slice). Remaining -findings: C01-1/2/3 (the locking-protocol redesign — replacement-owner -reconciliation on acquisition, guarded pre-commit effects, fenced shared -Caddy lock), C01-8 (same-version `_replaced` MANUAL — deliberate A08 -containment until F04 generation identities exist), and C01-9 -(attempt-scoped candidate identities — F04/A09). The A12/T05 remainder -of C01-7 (rollback's restoreRollbackRoute + the exact-block +compensation; evidence in AUDIT_OPEN's latest C01 slice), and **C01-1 +landed 2026-09-23** (replacement-owner reconciliation on acquisition; +`internal/deploy/reconcile.go`, fixture-verified) and **C01-2 landed +2026-09-23** (guarded pre-commit effects: RunGuarded container starts + +the guarded Caddyfile commit; see AUDIT_OPEN) and **C01-3 landed +2026-09-23** (owner-tagged fenced shared Caddy lock + the +conflicting-route-evidence producer/consumer; see AUDIT_OPEN). The +locking-protocol redesign (C01-1/2/3) is closed. Remaining findings: +C01-8 (same-version `_replaced` MANUAL — +deliberate A08 containment until F04 generation identities exist), and +C01-9 (attempt-scoped candidate identities — F04/A09). The A12/T05 +remainder of C01-7 (rollback's restoreRollbackRoute + the exact-block compare-and-swap restore) stays with its register item. diff --git a/internal/caddy/caddy.go b/internal/caddy/caddy.go index 93703a6..626e336 100644 --- a/internal/caddy/caddy.go +++ b/internal/caddy/caddy.go @@ -2,6 +2,9 @@ package caddy import ( "context" + "crypto/rand" + "encoding/hex" + "encoding/json" "fmt" "net" "net/url" @@ -13,6 +16,12 @@ import ( "github.com/useteploy/teploy/internal/ssh" ) +// fenceLostMarker mirrors state's guard marker (unexported there): the +// stderr sentinel a composed guard emits when refusing an effect. Shared +// by value so the caddy lock's guard is refusal-identifiable by the same +// string handling (state.FenceLost matches it in error text). +const fenceLostMarker = "TEPLOY_FENCE_LOST" + const ( caddyfilePath = "/deployments/caddy/Caddyfile" tmpCaddyfile = "/tmp/teploy_caddyfile.tmp" @@ -187,6 +196,13 @@ type SelectionPolicy struct { // can't load. type Client struct { exec ssh.Executor + // commitGuardPrefix, when set, composes a fence guard (the APP lock's, + // state.Lock.GuardPrefix) into the same shell command as the Caddyfile + // commit rename (C01-2): a deploy whose app lock was broken mid-mutate + // has its route edit refused (exit 75) instead of landing inside the + // new owner's window. Empty = unguarded commits (the receiver shared + // by callers that hold no app fence). Set through WithCommitGuard. + commitGuardPrefix string } // NewClient creates a Caddy client backed by the given SSH executor. @@ -194,6 +210,19 @@ func NewClient(exec ssh.Executor) *Client { return &Client{exec: exec} } +// WithCommitGuard returns a client whose Caddyfile commits run composed +// under the given fence-guard prefix in the same shell as the commit +// rename (C01-2). Callers that hold an app-level fence (deploy, rollback, +// static) pass their lock's GuardPrefix so the traffic switch — the +// finding's third check-then-act site — is guard+effect in one command. +// The receiver is unchanged: clients without a guard keep today's +// behavior. +func (c *Client) WithCommitGuard(guardPrefix string) *Client { + cp := *c + cp.commitGuardPrefix = guardPrefix + return &cp +} + // SetRoute adds or updates a reverse proxy route for the given app, serving the // (comma-separated) domain to the given upstream container:port. // @@ -433,12 +462,18 @@ func (c *Client) RemoveMaintenance(ctx context.Context, app string) error { // mutate serializes a Caddyfile edit + reload behind the server lock. transform // receives the current Caddyfile and returns the new contents. If the reload // fails (e.g. the new config is invalid), the on-disk file is rolled back so -// Caddy never persists a config it can't boot from. +// Caddy never persists a config it can't boot from. The COMMIT (the rename +// that makes the new contents authoritative) runs composed under the client's +// fence guard when one is set (C01-2): a stale lock holder's edit is refused +// instead of interleaving; the rollback restore is deliberately never fenced +// (refusing to clean up one's own partial effects is how a fencing design +// strands an app). func (c *Client) mutate(ctx context.Context, transform func(prev string) (string, error)) error { - if err := c.acquireLock(ctx); err != nil { + lockOwner, err := c.acquireLock(ctx) + if err != nil { return err } - defer c.releaseLock(ctx) + defer c.releaseLock(ctx, lockOwner) prev, err := c.exec.Run(ctx, "cat "+caddyfilePath) if err != nil { @@ -465,9 +500,29 @@ func (c *Client) mutate(ctx context.Context, transform func(prev string) (string return fmt.Errorf("refusing to write a Caddyfile the server's caddy rejects: %w", err) } - if err := c.writeCaddyfile(ctx, updated); err != nil { + // Stage (inert) then commit under the fence guards in one shell — the + // C01-2/C01-3 composition: the APP fence (when the caller holds one) + // and the CADDY-LOCK fence (this mutation's own lock, owner-tagged) + // both precede the rename, so a broken app holder's route edit AND a + // stale-broken editor's late write are refused in-shell. A refused + // commit propagates the fence error; the staged file is cleaned up + // here (bounded). + tmp, err := c.stageCaddyfile(ctx, updated) + if err != nil { + return err + } + committed := false + defer func() { + if !committed { + cleanupCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) + c.exec.Run(cleanupCtx, "rm -f -- "+ssh.ShellQuote(tmp)) + cancel() + } + }() + if err := c.commitCaddyfile(ctx, tmp, c.commitGuardPrefix+caddyGuardFragment(lockOwner)); err != nil { return err } + committed = true if err := c.reload(ctx); err != nil { // Roll back so a bad config is never left on disk to break the next // boot. The restore runs on a DETACHED bounded context — the deploy @@ -660,17 +715,70 @@ func managedBlockHosts(content, app string) []string { // mount (-v /deployments/caddy:/etc/caddy). verifyDelivered catches an // un-migrated box and aborts the deploy loudly before any damage. func (c *Client) writeCaddyfile(ctx context.Context, content string) error { - // Staged random SIBLING inside /deployments/caddy, then an atomic - // same-filesystem rename — replacing the old fixed /tmp staging path, - // which was shared (any concurrent process could race it), followed a - // pre-existing symlink at that path, and crossed filesystems for the - // final mv, losing rename atomicity (audit F46). - if err := ssh.UploadAtomic(ctx, c.exec, strings.NewReader(content), caddyfilePath, "0644"); err != nil { - return fmt.Errorf("writing caddyfile: %w", err) + tmp, err := c.stageCaddyfile(ctx, content) + if err != nil { + return err + } + if err := c.commitCaddyfile(ctx, tmp, ""); err != nil { + // The staged file is inert; leave nothing behind (bounded). + cleanupCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) + c.exec.Run(cleanupCtx, "rm -f -- "+ssh.ShellQuote(tmp)) + cancel() + return err } return nil } +// stageCaddyfile uploads content to a random inert SIBLING of the +// Caddyfile (F46's discipline: random sibling, same filesystem, no shared +// staging name). Staging has no effect: a file nothing commits is dead +// weight. The commit instant is commitCaddyfile's rename. +func (c *Client) stageCaddyfile(ctx context.Context, content string) (string, error) { + var suffix [8]byte + if _, err := rand.Read(suffix[:]); err != nil { + return "", fmt.Errorf("generating staging path: %w", err) + } + tmpPath := caddyfilePath + ".tmp-" + hex.EncodeToString(suffix[:]) + if err := c.exec.Upload(ctx, strings.NewReader(content), tmpPath, "0644"); err != nil { + return "", fmt.Errorf("staging caddyfile: %w", err) + } + return tmpPath, nil +} + +// commitCaddyfile renames the staged file into place — the instant the +// new config becomes authoritative on disk. guardPrefix (when set) +// composes the holdership check into the SAME shell command (C01-2): +// between a separate check and this rename a lock takeover could occur, +// letting a broken holder's route edit land inside the new owner's +// window; under composition the stale holder's edit is refused (exit 75, +// the TEPLOY_FENCE_LOST marker) and the staged bytes never become +// authoritative. A refusal is identified and rewrapped so callers can +// match it (state.FenceLost over the error string, or the message). +func (c *Client) commitCaddyfile(ctx context.Context, tmpPath, guardPrefix string) error { + cmd := guardPrefix + "mv -fT -- " + ssh.ShellQuote(tmpPath) + " " + ssh.ShellQuote(caddyfilePath) + _, err := c.exec.Run(ctx, cmd) + if err != nil && guardPrefix != "" && fenceRefused(err) { + return fmt.Errorf("caddyfile commit refused — a lock fence was lost (app lock or caddy lock no longer names this operation), the route edit did not land: %w", err) + } + if err != nil { + return fmt.Errorf("committing caddyfile: %w", err) + } + return nil +} + +// fenceRefused reports whether a commit-command error is the composed +// guard refusing the effect (marker on stderr or exit status 75) rather +// than the rename itself failing. +func fenceRefused(err error) bool { + if err == nil { + return false + } + msg := err.Error() + return strings.Contains(msg, "TEPLOY_FENCE_LOST") || + strings.Contains(msg, "status 75") || + strings.Contains(msg, "exit status 75") +} + func (c *Client) reload(ctx context.Context) error { if _, err := c.exec.Run(ctx, reloadCmd); err != nil { return err @@ -678,32 +786,120 @@ func (c *Client) reload(ctx context.Context) error { return nil } -// acquireLock takes a mkdir-based mutex on the Caddyfile, breaking a stale lock -// left by a crashed deploy. Held only for the brief edit+reload, so contention -// is rare and short. -func (c *Client) acquireLock(ctx context.Context) error { +// caddyLockInfo is the shared-proxy commit lock's identity file (C01-3). +// Pre-C01-3 the lock was a bare mkdir with no owner: any mutator broke it +// after staleLockSeconds by DIRECTORY MTIME and a slow-but-alive orphaned +// editor could interleave its Caddyfile edit+reload with the new owner's +// mutate. The shape mirrors the app locks (state.LockInfo): an owner +// token, and staleness measured from the info's own timestamp. The TTL +// stays short (edits are seconds); there is no renewal — an edit session +// is far shorter than any plausible renewal interval. +type caddyLockInfo struct { + Type string `json:"type"` + Owner string `json:"owner"` + TS string `json:"ts"` +} + +// newCaddyOwner mints the lock's fencing token. +func newCaddyOwner() string { + var b [16]byte + if _, err := rand.Read(b[:]); err != nil { + // crypto/rand failing is catastrophic-environment territory; a + // time-derived token still unique-ifies this process's edits. + return fmt.Sprintf("caddy-%d", time.Now().UnixNano()) + } + return hex.EncodeToString(b[:]) +} + +// acquireLock takes the shared-proxy commit lock (mkdir + owner-tagged +// info), breaking a STALE holder's lock first. Staleness is measured from +// the info file's timestamp when one exists (C01-3); a legacy lock dir +// with no info falls back to the directory-mtime age check. The returned +// owner token fences the commit: a holder whose lock was broken (or whose +// info no longer names it) has its Caddyfile commit refused by the +// composed guard instead of interleaving with the new owner's edit. +func (c *Client) acquireLock(ctx context.Context) (string, error) { + owner := newCaddyOwner() for i := 0; i < lockWaitTries; i++ { if _, err := c.exec.Run(ctx, "mkdir "+lockDir); err == nil { - return nil + if err := c.writeCaddyLockInfo(ctx, owner); err != nil { + // Our own fresh lock with no successor possible: an + // unconditional release is correct here. + c.exec.Run(ctx, "rm -rf "+lockDir) + return "", err + } + return owner, nil + } + if c.caddyLockStale(ctx) { + c.exec.Run(ctx, "rm -rf "+lockDir) + continue } - // Break a stale lock (older than staleLockSeconds), then wait and retry. - c.exec.Run(ctx, fmt.Sprintf( - "[ -d %s ] && [ $(( $(date +%%s) - $(stat -c %%Y %s 2>/dev/null || echo 0) )) -gt %d ] && rm -rf %s || true", - lockDir, lockDir, staleLockSeconds, lockDir, - )) c.exec.Run(ctx, "sleep 0.5") } - return fmt.Errorf("timed out acquiring caddy lock %s", lockDir) + return "", fmt.Errorf("timed out acquiring caddy lock %s", lockDir) +} + +// caddyLockStale reports whether the existing lock may be broken: an +// info-carrying lock (C01-3) is stale when its OWN timestamp is older +// than staleLockSeconds (unparseable timestamp = stale — the app locks' +// rule); a legacy no-info dir is stale by directory mtime, the pre-C01-3 +// behavior. +func (c *Client) caddyLockStale(ctx context.Context) bool { + if out, err := c.exec.Run(ctx, "cat "+lockDir+"/info 2>/dev/null"); err == nil && strings.TrimSpace(out) != "" { + var info caddyLockInfo + if err := json.Unmarshal([]byte(strings.TrimSpace(out)), &info); err != nil { + return true + } + ts, err := time.Parse(time.RFC3339, info.TS) + if err != nil { + return true + } + return time.Since(ts) > staleLockSeconds*time.Second + } + // Legacy dir (no info): the old mtime-based break. + out, err := c.exec.Run(ctx, fmt.Sprintf( + "[ -d %s ] && [ $(( $(date +%%s) - $(stat -c %%Y %s 2>/dev/null || echo 0) )) -gt %d ] && echo stale || echo fresh", + lockDir, lockDir, staleLockSeconds, + )) + return err == nil && strings.TrimSpace(out) == "stale" +} + +func (c *Client) writeCaddyLockInfo(ctx context.Context, owner string) error { + info, err := json.Marshal(caddyLockInfo{ + Type: "caddy-edit", + Owner: owner, + TS: time.Now().UTC().Format(time.RFC3339), + }) + if err != nil { + return err + } + if err := c.exec.Upload(ctx, strings.NewReader(string(info)), lockDir+"/info", "0644"); err != nil { + return fmt.Errorf("writing caddy lock info: %w", err) + } + return nil +} + +// caddyGuardFragment is the caddy lock's holdership guard — the same +// shape as the app lock's (state.Lock.guardFragment) so transport-level +// refusal handling and the mock executor treat both identically. +func caddyGuardFragment(owner string) string { + return fmt.Sprintf("grep -q %s %s || { printf '%s\\n' >&2; exit 75; }; ", + ssh.ShellQuote(owner), ssh.ShellQuote(lockDir+"/info"), fenceLostMarker) } -func (c *Client) releaseLock(ctx context.Context) { - // A detached, bounded context: releasing with the operation's context - // let a cancelled deploy skip the rmdir entirely, leaving the lock dir - // behind to block every subsequent Caddy mutation for the stale window - // (TCL-05 containment — the release must survive caller cancellation). +// releaseLock removes the caddy lock ONLY when its info still names this +// owner (C01-3, the app locks' A04 lesson): after a stale break and +// re-acquire, an unconditional rmdir would delete the SUCCESSOR's lock +// and admit a third editor. A lock naming someone else is left strictly +// alone; a detached bounded context keeps the release alive past caller +// cancellation (TCL-05). +func (c *Client) releaseLock(ctx context.Context, owner string) { rctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() - c.exec.Run(rctx, "rmdir "+lockDir+" 2>/dev/null || true") + c.exec.Run(rctx, fmt.Sprintf( + "if [ -d %s ] && grep -q %s %s 2>/dev/null; then rm -rf -- %s; fi", + ssh.ShellQuote(lockDir), ssh.ShellQuote(owner), ssh.ShellQuote(lockDir+"/info"), ssh.ShellQuote(lockDir), + )) } // removeForeignHostBlocks' whole-block rule moved to routes.go's diff --git a/internal/caddy/caddy_test.go b/internal/caddy/caddy_test.go index 1fd333b..003033e 100644 --- a/internal/caddy/caddy_test.go +++ b/internal/caddy/caddy_test.go @@ -5,6 +5,7 @@ import ( "fmt" "strings" "testing" + "time" "github.com/useteploy/teploy/internal/ssh" ) @@ -21,7 +22,6 @@ func lockCmds(caddyfile string) []ssh.MockCommand { {Match: reloadCmd, Output: ""}, // Post-reload delivery check: container's file matches what we wrote. {Match: "a=$(docker exec caddy md5sum", Output: deliveredOK}, - {Match: "rmdir " + lockDir, Output: ""}, } } @@ -51,7 +51,7 @@ func TestSetRoute_WritesBlockAndReloads(t *testing.T) { if !calledWith(mock, reloadCmd) { t.Error("expected caddy reload to be called") } - if !calledWith(mock, "mkdir "+lockDir) || !calledWith(mock, "rmdir "+lockDir) { + if !calledWith(mock, "mkdir "+lockDir) || !calledWith(mock, "if [ -d '"+lockDir+"'") { t.Error("expected the Caddy lock to be acquired and released") } } @@ -127,7 +127,6 @@ func TestSetRoute_ReloadFailureRollsBack(t *testing.T) { // First reload (new config) fails; rollback reload then succeeds. ssh.MockCommand{Match: reloadCmd, Err: fmt.Errorf("invalid config"), Once: true}, ssh.MockCommand{Match: reloadCmd, Output: ""}, - ssh.MockCommand{Match: "rmdir " + lockDir, Output: ""}, ) client := NewClient(mock) @@ -153,7 +152,6 @@ func TestSetRoute_StaleDeliveryFailsLoudly(t *testing.T) { ssh.MockCommand{Match: "mv -f -- ", Output: ""}, ssh.MockCommand{Match: reloadCmd, Output: ""}, ssh.MockCommand{Match: "a=$(docker exec caddy md5sum", Output: deliveredStale}, - ssh.MockCommand{Match: "rmdir " + lockDir, Output: ""}, ) client := NewClient(mock) @@ -164,7 +162,7 @@ func TestSetRoute_StaleDeliveryFailsLoudly(t *testing.T) { if !strings.Contains(err.Error(), "directory mount") { t.Errorf("expected a stale-delivery error pointing at the directory-mount recreate, got: %v", err) } - if !calledWith(mock, "rmdir "+lockDir) { + if !calledWith(mock, "if [ -d '"+lockDir+"'") { t.Error("expected the caddy lock to be released after a failed delivery check") } } @@ -696,3 +694,160 @@ func TestNoCacheRulesRendersNothing(t *testing.T) { t.Fatalf("expected empty output for empty cache, got %q", got) } } + +// TestSetRoute_GuardedCommitRefusedWhenFenceLost (C01-2): with a commit +// guard set, the Caddyfile commit rename runs composed under the guard — +// when the guarded lock names another owner, the edit is REFUSED in-shell: +// nothing lands in the Caddyfile, no reload runs, and the error says the +// fence was lost. +func TestSetRoute_GuardedCommitRefusedWhenFenceLost(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", lockCmds("{\n\tadmin 127.0.0.1:2019\n}\n")...) + guard := "grep -q 'deadowner' '/deployments/myapp/.lock/info' || { printf 'TEPLOY_FENCE_LOST\\n' >&2; exit 75; }; " + mock.Files["/deployments/myapp/.lock/info"] = []byte(`{"type":"auto","owner":"someoneelse"}`) + + client := NewClient(mock).WithCommitGuard(guard) + err := client.SetRoute(context.Background(), "myapp", "myapp.com", "myapp-v1", 80, TLS{}, "", nil, Firewall{}, Access{}) + if err == nil { + t.Fatal("expected the guarded commit to be refused") + } + if !strings.Contains(err.Error(), "fence was lost") { + t.Fatalf("expected a fence-loss error, got: %v", err) + } + if _, ok := mock.Files[caddyfilePath]; ok { + t.Error("a refused commit must not write the Caddyfile") + } + if calledWith(mock, reloadCmd) { + t.Error("no reload may run after a refused commit") + } + // The lock is still acquired/released around the refused edit. + if !calledWith(mock, "mkdir "+lockDir) || !calledWith(mock, "if [ -d '"+lockDir+"'") { + t.Error("expected the caddy lock to be acquired and released") + } +} + +// TestSetRoute_GuardedCommitHoldsUnderRightOwner: the same composed +// command with the guard naming us commits and reloads normally — +// composition changes nothing for the legitimate holder. +func TestSetRoute_GuardedCommitHoldsUnderRightOwner(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", lockCmds("{\n\tadmin 127.0.0.1:2019\n}\n")...) + guard := "grep -q 'me' '/deployments/myapp/.lock/info' || { printf 'TEPLOY_FENCE_LOST\\n' >&2; exit 75; }; " + mock.Files["/deployments/myapp/.lock/info"] = []byte(`{"type":"auto","owner":"me"}`) + + client := NewClient(mock).WithCommitGuard(guard) + if err := client.SetRoute(context.Background(), "myapp", "myapp.com", "myapp-v1", 80, TLS{}, "", nil, Firewall{}, Access{}); err != nil { + t.Fatalf("SetRoute under a held guard: %v", err) + } + if !strings.Contains(string(mock.Files[caddyfilePath]), "reverse_proxy myapp-v1:80") { + t.Errorf("expected the guarded commit to land the block, got:\n%s", mock.Files[caddyfilePath]) + } + if !calledWith(mock, reloadCmd) { + t.Error("expected a reload after the guarded commit") + } +} + +// --- C01-3: the shared-proxy commit lock is owner-tagged and fenced --- + +// TestCaddyLock_OwnerTaggedAndStaleBrokenByInfoAge: acquisition writes an +// owner-tagged info file; a holder's lock whose info TIMESTAMP exceeds +// staleLockSeconds is broken (not by directory mtime — the info is the +// authority), and the new holder's info replaces it. +func TestCaddyLock_OwnerTaggedAndStaleBrokenByInfoAge(t *testing.T) { + stale := time.Now().UTC().Add(-2 * time.Duration(staleLockSeconds) * time.Second).Format(time.RFC3339) + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "mkdir " + lockDir, Err: fmt.Errorf("exists"), Once: true}, + ssh.MockCommand{Match: "mkdir " + lockDir, Output: ""}, + ssh.MockCommand{Match: "cat " + lockDir + "/info", Output: fmt.Sprintf(`{"type":"caddy-edit","owner":"deadone","ts":%q}`, stale)}, + ) + client := NewClient(mock) + owner, err := client.acquireLock(context.Background()) + if err != nil { + t.Fatalf("acquireLock after stale break: %v", err) + } + if owner == "" || owner == "deadone" { + t.Fatalf("expected a fresh owner token, got %q", owner) + } + info, ok := mock.Files[lockDir+"/info"] + if !ok { + t.Fatal("acquisition must write the lock info") + } + if !strings.Contains(string(info), owner) { + t.Errorf("lock info does not name the acquiring owner: %s", info) + } + client.releaseLock(context.Background(), owner) + if _, still := mock.Files[lockDir+"/info"]; still { + t.Error("conditional release must remove our own lock") + } +} + +// TestCaddyLock_FreshHolderIsNotBroken: a lock younger than the stale +// window blocks acquisition (times out) instead of being broken — the +// interleave the old mtime break allowed. +func TestCaddyLock_FreshHolderIsNotBroken(t *testing.T) { + fresh := time.Now().UTC().Format(time.RFC3339) + var broke bool + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "mkdir " + lockDir, Err: fmt.Errorf("exists")}, + ssh.MockCommand{Match: "cat " + lockDir + "/info", Output: fmt.Sprintf(`{"type":"caddy-edit","owner":"liveone","ts":%q}`, fresh)}, + ssh.MockCommand{Match: "rm -rf " + lockDir, Output: ""}, + ) + client := NewClient(mock) + // Shorten the wait loop via the context: acquireLock polls ~30s; run + // it on a context that dies after a moment and treat the deadline as + // "not broken" evidence. + ctx, cancel := context.WithTimeout(context.Background(), 300*time.Millisecond) + defer cancel() + _, err := client.acquireLock(ctx) + if err == nil { + t.Fatal("acquiring against a fresh holder must not succeed") + } + for _, c := range mock.Calls { + if strings.HasPrefix(c, "rm -rf "+lockDir) { + broke = true + } + } + if broke { + t.Error("a fresh (younger than staleLockSeconds) caddy lock must not be broken") + } +} + +// TestCaddyLock_BrokenHoldersCommitRefused: after a stale break and +// re-acquire, the DEAD holder's composed commit is refused by its own +// guard — the info no longer names it — so its late Caddyfile edit cannot +// interleave with the new holder's mutate (the C01-3 core property). +func TestCaddyLock_BrokenHoldersCommitRefused(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "mkdir " + lockDir, Err: fmt.Errorf("exists"), Once: true}, + ssh.MockCommand{Match: "mkdir " + lockDir, Output: ""}, + ) + stale := time.Now().UTC().Add(-2 * time.Duration(staleLockSeconds) * time.Second).Format(time.RFC3339) + mock.Files[lockDir+"/info"] = []byte(fmt.Sprintf(`{"type":"caddy-edit","owner":"deadone","ts":%q}`, stale)) + + client := NewClient(mock) + newOwner, err := client.acquireLock(context.Background()) + if err != nil { + t.Fatalf("acquireLock: %v", err) + } + // The dead holder attempts its late commit with its own guard. + deadGuard := caddyGuardFragment("deadone") + err = client.commitCaddyfile(context.Background(), "/deployments/caddy/Caddyfile.tmp-dead", deadGuard) + if err == nil || !strings.Contains(err.Error(), "fence was lost") { + t.Fatalf("expected the broken holder's commit to be refused, got: %v", err) + } + if _, ok := mock.Files[caddyfilePath]; ok { + t.Error("a refused late commit must not write the Caddyfile") + } + _ = newOwner +} + +// TestCaddyLock_ReleaseNeverDeletesSuccessorsLock: a holder whose lock +// was broken and re-acquired must not delete the SUCCESSOR's lock on +// release — the app locks' A04 lesson applied to the shared proxy lock. +func TestCaddyLock_ReleaseNeverDeletesSuccessorsLock(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4") + mock.Files[lockDir+"/info"] = []byte(`{"type":"caddy-edit","owner":"successor","ts":"` + time.Now().UTC().Format(time.RFC3339) + `"}`) + client := NewClient(mock) + client.releaseLock(context.Background(), "staleholder") + if _, ok := mock.Files[lockDir+"/info"]; !ok { + t.Error("a stale holder's release deleted the successor's caddy lock") + } +} diff --git a/internal/caddy/tcl_round2_test.go b/internal/caddy/tcl_round2_test.go index f368472..dc450bb 100644 --- a/internal/caddy/tcl_round2_test.go +++ b/internal/caddy/tcl_round2_test.go @@ -32,6 +32,30 @@ func newFakeStatefulExecutor(initial map[string]string) *fakeStatefulExecutor { func (f *fakeStatefulExecutor) Run(ctx context.Context, cmd string) (string, error) { f.mu.Lock() defer f.mu.Unlock() + // Composed fence guards (C01-2/C01-3): evaluate every chained guard + // against the recorded files (the caddy lock info uploaded at acquire + // names the mutating owner), then dispatch the remaining effect. A + // guard whose info no longer names the owner refuses exactly like the + // real shell would. + const guardSep = " || { printf 'TEPLOY_FENCE_LOST\\n' >&2; exit 75; }; " + for strings.HasPrefix(cmd, "grep -q ") { + i := strings.Index(cmd, guardSep) + if i < 0 { + break + } + guard, rest := cmd[:i], cmd[i+len(guardSep):] + fields := strings.Fields(guard) // grep -q 'owner' 'path' + if len(fields) != 4 { + break + } + owner := strings.Trim(fields[2], "'") + path := strings.Trim(fields[3], "'") + data, ok := f.files[path] + if !ok || !strings.Contains(string(data), owner) { + return "", fmt.Errorf("exit status 75: TEPLOY_FENCE_LOST") + } + cmd = rest + } switch { case cmd == "cat "+caddyfilePath: data, ok := f.files[caddyfilePath] @@ -62,7 +86,8 @@ func (f *fakeStatefulExecutor) Run(ctx context.Context, cmd string) (string, err return "", f.adaptErr } return "", nil - case strings.HasPrefix(cmd, "mkdir "+lockDir), strings.HasPrefix(cmd, "rmdir "+lockDir): + case strings.HasPrefix(cmd, "mkdir "+lockDir), strings.HasPrefix(cmd, "if [ -d "+lockDir), + strings.HasPrefix(cmd, "cat "+lockDir+"/info"), strings.HasPrefix(cmd, "sleep 0.5"): return "", nil case cmd == reloadCmd: return "", nil diff --git a/internal/cli/apply.go b/internal/cli/apply.go new file mode 100644 index 0000000..84782f5 --- /dev/null +++ b/internal/cli/apply.go @@ -0,0 +1,189 @@ +package cli + +// `teploy apply ` (C05): execute a reviewed plan through the +// EXISTING deploy engine — there is no side engine. Before anything +// runs, the plan's recorded binding (effective-config digest, target +// version, build inputs, target server/app identity, deployed-state +// generation) is re-derived and compared; any drift refuses the apply +// naming WHAT moved, with the remedy (re-plan). The plan id is stamped +// into the deploy's provenance receipt, tying the release back to the +// plan that was verified. + +import ( + "context" + "errors" + "fmt" + "io" + "os" + "os/signal" + + "github.com/spf13/cobra" + "github.com/useteploy/teploy/internal/config" + "github.com/useteploy/teploy/internal/state" +) + +func newApplyCmd(flags *Flags) *cobra.Command { + var ( + migrateVolumes bool + skipDNSCheck bool + ) + cmd := &cobra.Command{ + Use: "apply ", + Short: "Execute a reviewed plan (refuses if anything drifted since it was made)", + Long: "Executes a plan written by `teploy plan --out` through the normal deploy " + + "engine. Before executing, the plan's binding is re-verified: the effective " + + "config digest, the target version, the build inputs (for build plans), the " + + "target server and app, and the deployed-state generation must all still " + + "match what the plan recorded. Anything that moved — a config edit, a deploy " + + "or rollback that ran in between, source changes without a version change — " + + "invalidates the plan and the apply refuses with the remedy (re-plan).\n\n" + + "Run from the same app directory the plan was made in (the plan records the " + + "destination overlay it was computed with; do not pass a different one).", + Args: cobra.ExactArgs(1), + RunE: func(cmd *cobra.Command, args []string) error { + return runApply(flags, args[0], skipDNSCheck, migrateVolumes) + }, + } + cmd.Flags().BoolVar(&migrateVolumes, "migrate-volumes", false, "auto-migrate data from foreign volume sources to teploy paths (cp -a) — mirror of deploy --migrate-volumes") + cmd.Flags().BoolVar(&skipDNSCheck, "skip-dns-check", false, "skip DNS validation (for proxied domains like Cloudflare) — mirror of deploy --skip-dns-check") + return cmd +} + +func runApply(flags *Flags, planPath string, skipDNSCheck, migrateVolumes bool) error { + ctx, cancel := signal.NotifyContext(context.Background(), os.Interrupt) + defer cancel() + + rec, err := loadPlanFile(planPath) + if err != nil { + return refuseApply(err) + } + // A plan whose target version is a deploy-time timestamp can never be + // bound — the thing apply would deploy is not the thing reviewed. + if !rec.VersionKnown { + return refuseApply(fmt.Errorf("plan %s targets an unpredictable version (floating image tag) — re-plan with --version or a digest-pinned image", rec.PlanID)) + } + + // Load the config exactly as the plan did (same loader, same overlay), + // then resolve env the same way `deploy` does — the shared + // resolveDeployEnv, not a fork. + appCfg, image, curDigest, curVersion, err := applyResolveCurrent(ctx, flags, rec) + if err != nil { + return err + } + + // Connect to the plan's target. + serverName := rec.ServerName + if serverName == "" { + serverName = resolvePlanServerName(appCfg) + } + executor, err := connectForApp(ctx, flags, appCfg) + if err != nil { + return err + } + defer executor.Close() + + current, err := state.Read(ctx, executor, appCfg.App) + if err != nil { + return err + } + + // Build-input re-verification for build plans: same C04 resolution + // path the plan used, compared on the fingerprint + Dockerfile + // identity. This catches edits the version cannot see (a dirty-tree + // change between plan and apply). + var curFingerprint, curDockerfileSHA string + if rec.Image.NeedsBuild { + prov := resolveDeployProvenance(ctx, executor, io.Discard, appCfg, ".", image, rec.TargetVersion, curDigest, true) + curFingerprint = prov.ContextFingerprint + curDockerfileSHA = prov.DockerfileSHA256 + } + + if err := verifyPlanBinding(rec, planCurrentFacts{ + App: appCfg.App, + Server: executor.Host(), + Version: curVersion, + ConfigDigest: curDigest, + ContextFingerprint: curFingerprint, + DockerfileSHA256: curDockerfileSHA, + State: current, + }); err != nil { + return refuseApply(err) + } + + if !flags.JSON { + fmt.Printf("Plan %s verified (config digest %s, target %s generation %d) — executing through the deploy engine\n", + rec.PlanID, shortDigest(rec.ConfigDigest), rec.Server, rec.TargetState.Generation) + } + + // ONE execution path: the same deployAppConfig a direct `teploy + // deploy` runs, with the plan's recorded image/version/overlay. The + // plan id rides along and is stamped into the provenance receipt. + err = deployAppConfig(flags, appCfg, serverName, image, rec.TargetVersion, skipDNSCheck, migrateVolumes, rec.PlanID) + if err != nil { + return err + } + if !flags.JSON { + fmt.Printf("Plan %s applied — the release receipt carries this plan id (provenance.plan_id)\n", rec.PlanID) + } + return nil +} + +// applyResolveCurrent loads the current world exactly as the plan did +// and re-derives the pre-connect binding facts: the app config (loader + +// recorded overlay + env resolution, identical to deploy — one +// resolution semantics), the effective-config digest over the plan's +// image reference (recomputable before anything executes), and the +// re-derived target version (explicit versions bind as-is; derived ones +// must re-derive, or the world the plan reviewed is gone). Everything +// here is local — no server contact — so the config-drift refusal fires +// before a connection is even opened. +func applyResolveCurrent(ctx context.Context, flags *Flags, rec *PlanRecord) (appCfg *config.AppConfig, image, curDigest, curVersion string, err error) { + if rec.Destination != "" { + appCfg, err = config.LoadAppWithDestination(".", rec.Destination, config.OverlayOptions{Strict: flags.StrictEnv}) + } else { + appCfg, err = config.LoadApp(".") + } + if err != nil { + return nil, "", "", "", err + } + if revision, revisionErr := gitRevisionIn("."); revisionErr == nil { + appCfg.SourceRevision = revision + } + if err := resolveDeployEnv(ctx, appCfg, flags.StrictEnv); err != nil { + return nil, "", "", "", err + } + + image = rec.Image.Ref + _, curDigest, err = config.NormalizeAndDigest(appCfg, image) + if err != nil { + return nil, "", "", "", fmt.Errorf("normalizing current config: %w", err) + } + + curVersion = rec.TargetVersion + if !rec.VersionExplicit { + if image != "" { + curVersion = versionFromImage(image) + if curVersion == "" { + curVersion = "" + } + } else { + if curVersion, err = gitShortHash(); err != nil { + return nil, "", "", "", fmt.Errorf("could not re-derive the target version from git: %w", err) + } + } + } + return appCfg, image, curDigest, curVersion, nil +} + +// refuseApply marks an apply refusal. Drift refusals classify as the +// error-envelope conflict code (the request is coherent; the world +// moved); the rest stay admission-shaped. +func refuseApply(err error) error { + if err == nil { + return nil + } + if errors.Is(err, errPlanDrift) { + return err + } + return refuseAdmission(err) +} diff --git a/internal/cli/apply_test.go b/internal/cli/apply_test.go new file mode 100644 index 0000000..3f66e78 --- /dev/null +++ b/internal/cli/apply_test.go @@ -0,0 +1,379 @@ +package cli + +import ( + "context" + "encoding/json" + "errors" + "fmt" + "os" + "path/filepath" + "strings" + "testing" + + "github.com/useteploy/teploy/internal/config" + "github.com/useteploy/teploy/internal/releasemeta" + "github.com/useteploy/teploy/internal/ssh" + "github.com/useteploy/teploy/internal/state" +) + +// writeAppConfig writes a teploy.yml into dir and chdirs there for the +// test (apply resolves config from "."). +func writeAppConfig(t *testing.T, dir, yml string) { + t.Helper() + if err := os.WriteFile(filepath.Join(dir, "teploy.yml"), []byte(yml), 0o644); err != nil { + t.Fatal(err) + } +} + +// planFromConfig builds a plan record the way `teploy plan` would for +// the CURRENT teploy.yml in dir (digest over the given image reference). +func planFromConfig(t *testing.T, dir, image, destination string) *PlanRecord { + t.Helper() + orig, err := os.Getwd() + if err != nil { + t.Fatal(err) + } + if chdirErr := os.Chdir(dir); chdirErr != nil { + t.Fatal(chdirErr) + } + t.Cleanup(func() { _ = os.Chdir(orig) }) + + var appCfg *config.AppConfig + if destination != "" { + appCfg, err = config.LoadAppWithDestination(".", destination, config.OverlayOptions{}) + } else { + appCfg, err = config.LoadApp(".") + } + if err != nil { + t.Fatalf("LoadApp: %v", err) + } + _, digest, err := config.NormalizeAndDigest(appCfg, image) + if err != nil { + t.Fatal(err) + } + rec := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, + App: appCfg.App, + Server: "srv.test", + ServerName: "prod", + Destination: destination, + TargetVersion: "v1", + VersionKnown: true, + VersionExplicit: true, + ConfigDigest: digest, + Image: PlanImageIdentity{Ref: image, Resolution: imageUnresolvedMutableTag}, + } + rec.PlanID = computePlanID(rec) + return rec +} + +// TestApplyDrift_ConfigChangedBetweenPlanAndApply is the C05 acceptance +// fixture: a config edit between plan and apply must refuse, naming the +// drift — never execute the stale plan. +func TestApplyDrift_ConfigChangedBetweenPlanAndApply(t *testing.T) { + dir := t.TempDir() + writeAppConfig(t, dir, "app: blog\nimage: registry/blog:v1\ndomain: blog.example.com\nport: 3000\n") + rec := planFromConfig(t, dir, "registry/blog:v1", "") + + // The world moves: port edited after the plan was reviewed. + writeAppConfig(t, dir, "app: blog\nimage: registry/blog:v1\ndomain: blog.example.com\nport: 8080\n") + + appCfg, image, digest, version, err := applyResolveCurrent(context.Background(), &Flags{}, rec) + if err != nil { + t.Fatal(err) + } + err = verifyPlanBinding(rec, planCurrentFacts{ + App: appCfg.App, Server: rec.Server, Version: version, ConfigDigest: digest, + }) + if !errors.Is(err, errPlanDrift) { + t.Fatalf("config edit must invalidate the plan, got %v", err) + } + var drift *planDriftError + if !errors.As(err, &drift) || drift.Kind != "config" { + t.Errorf("kind = %+v, want config", err) + } + if !strings.Contains(err.Error(), rec.ConfigDigest) || !strings.Contains(err.Error(), digest) { + t.Errorf("refusal must name both digests: %v", err) + } + _ = image +} + +// TestApplyDrift_OverlayFlipInvalidates: the destination overlay's +// merged result is part of the binding — an overlay edit between plan +// and apply refuses (presence-aware overlay semantics pinned through +// the config digest). +func TestApplyDrift_OverlayFlipInvalidates(t *testing.T) { + dir := t.TempDir() + writeAppConfig(t, dir, "app: blog\nimage: registry/blog:v1\ndomain: blog.example.com\nport: 3000\nenv:\n BASE: kept\n") + if err := os.WriteFile(filepath.Join(dir, "teploy.staging.yml"), []byte("env:\n STAGED: \"1\"\n"), 0o644); err != nil { + t.Fatal(err) + } + rec := planFromConfig(t, dir, "registry/blog:v1", "staging") + + // Overlay edited after the plan: staging now clears the env block + // (strict presence-aware clearing). + if err := os.WriteFile(filepath.Join(dir, "teploy.staging.yml"), []byte("env: {}\n"), 0o644); err != nil { + t.Fatal(err) + } + + _, _, digest, version, err := applyResolveCurrent(context.Background(), &Flags{StrictEnv: true}, rec) + if err != nil { + t.Fatal(err) + } + err = verifyPlanBinding(rec, planCurrentFacts{App: "blog", Server: rec.Server, Version: version, ConfigDigest: digest}) + if !errors.Is(err, errPlanDrift) { + t.Fatalf("overlay edit must invalidate the plan, got %v", err) + } +} + +// TestApplyStable_ConfigUnchanged: nothing moved between plan and apply +// — the same resolution verifies clean. +func TestApplyStable_ConfigUnchanged(t *testing.T) { + dir := t.TempDir() + writeAppConfig(t, dir, "app: blog\nimage: registry/blog:v1\ndomain: blog.example.com\nport: 3000\n") + rec := planFromConfig(t, dir, "registry/blog:v1", "") + + appCfg, _, digest, version, err := applyResolveCurrent(context.Background(), &Flags{}, rec) + if err != nil { + t.Fatal(err) + } + err = verifyPlanBinding(rec, planCurrentFacts{ + App: appCfg.App, Server: rec.Server, Version: version, ConfigDigest: digest, + State: &state.AppState{Generation: 2, CurrentHash: "old"}, + }) + // The plan's recorded target state must match too. + rec.TargetState = PlanTargetState{Deployed: true, Generation: 2, CurrentHash: "old"} + rec.PlanID = computePlanID(rec) + err = verifyPlanBinding(rec, planCurrentFacts{ + App: appCfg.App, Server: rec.Server, Version: version, ConfigDigest: digest, + State: &state.AppState{Generation: 2, CurrentHash: "old"}, + }) + if err != nil { + t.Fatalf("unchanged world must verify, got %v", err) + } +} + +// TestApplyDrift_DeployHappenedInBetween: a deploy between plan and +// apply moves the state generation — the stale plan refuses. +func TestApplyDrift_DeployHappenedInBetween(t *testing.T) { + dir := t.TempDir() + writeAppConfig(t, dir, "app: blog\nimage: registry/blog:v1\ndomain: blog.example.com\nport: 3000\n") + rec := planFromConfig(t, dir, "registry/blog:v1", "") + rec.TargetState = PlanTargetState{Deployed: true, Generation: 7, CurrentHash: "old99"} + rec.PlanID = computePlanID(rec) + + appCfg, _, digest, version, err := applyResolveCurrent(context.Background(), &Flags{}, rec) + if err != nil { + t.Fatal(err) + } + err = verifyPlanBinding(rec, planCurrentFacts{ + App: appCfg.App, Server: rec.Server, Version: version, ConfigDigest: digest, + State: &state.AppState{Generation: 8, CurrentHash: "v1"}, + }) + var drift *planDriftError + if !errors.As(err, &drift) || drift.Kind != "target-state" { + t.Fatalf("generation move must refuse, got %v", err) + } + if !strings.Contains(err.Error(), "7 -> 8") { + t.Errorf("refusal must name the generation move: %v", err) + } +} + +// TestApplyRefusesUnpredictableVersion: a plan made against a floating +// tag (deploy-time timestamp version) can never be bound — apply +// refuses at the gate. +func TestApplyRefusesUnpredictableVersion(t *testing.T) { + rec := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, PlanID: "x", + App: "blog", Server: "srv", TargetVersion: "", VersionKnown: false, + ConfigDigest: "d", + } + // The gate runs before any resolution or connection. + if !rec.VersionKnown { + if _, err := loadPlanFile(writePlan(t, rec)); err == nil { // plan file must be self-consistent first + // (computePlanID of this record; writePlan fixes the id) + } + } + rec.PlanID = computePlanID(rec) + path := writePlan(t, rec) + if _, err := loadPlanFile(path); err != nil { + t.Fatalf("plan file must load: %v", err) + } + // The refusal itself: + err := applyUnpredictableVersionGate(rec) + if err == nil || !strings.Contains(err.Error(), "unpredictable version") { + t.Fatalf("floating-version plan must be refused: %v", err) + } +} + +// writePlan saves a record to a temp file (id already computed). +func writePlan(t *testing.T, rec *PlanRecord) string { + t.Helper() + if rec.PlanID == "" { + rec.PlanID = computePlanID(rec) + } + path := filepath.Join(t.TempDir(), "plan.json") + if err := savePlanFile(path, rec); err != nil { + t.Fatal(err) + } + return path +} + +// applyUnpredictableVersionGate is the head of runApply's gate, lifted +// one line for testability (runApply inlines the same condition). +func applyUnpredictableVersionGate(rec *PlanRecord) error { + if !rec.VersionKnown { + return fmt.Errorf("plan %s targets an unpredictable version (floating image tag) — re-plan with --version or a digest-pinned image", rec.PlanID) + } + return nil +} + +// ssNoEphemeral is the port-scan baseline the deploy engine tests use. +const ssNoEphemeral = `State Recv-Q Send-Q Local Address:Port Peer Address:Port +LISTEN 0 128 0.0.0.0:22 0.0.0.0:* +LISTEN 0 128 0.0.0.0:80 0.0.0.0:* +LISTEN 0 128 0.0.0.0:443 0.0.0.0:*` + +// TestApplyStampReceiptThroughEngine: a verified apply executes through +// the REAL deploy engine (deployBuiltImageFenced — the same function +// `teploy deploy` runs) and the provenance receipt lands in the attempt +// namespace carrying the plan id. +func TestApplyStampReceiptThroughEngine(t *testing.T) { + imageDigest := "sha256:" + strings.Repeat("ab", 32) + image := "registry.example.com/myapp@" + imageDigest + appCfg := &config.AppConfig{ + App: "myapp", + Image: image, + Domain: "myapp.example.com", + Port: 3000, + } + version := "v1" + + _, manifestSHA, err := config.NormalizeAndDigest(appCfg, image) + if err != nil { + t.Fatal(err) + } + appliedManifest, _, _ := config.NormalizeAndDigest(appCfg, image) + + mock := ssh.NewMockExecutor("1.2.3.4", + // provenance mkdir + upload (UploadAtomic: UPLOAD + mv handled by mock) + ssh.MockCommand{Match: "mkdir -p /deployments/myapp/meta/att", Output: ""}, + // .env absent + ssh.MockCommand{Match: "test -f /deployments/myapp/.env", Err: fmt.Errorf("no file")}, + // secrets absent + ssh.MockCommand{Match: "if [ ! -e '/deployments/myapp/secrets'", Output: "absent"}, + // deploy engine: first deploy + ssh.MockCommand{Match: "mkdir -p /deployments/myapp", Output: ""}, + ssh.MockCommand{Match: "mkdir /deployments/myapp/.lock", Output: ""}, + ssh.MockCommand{Match: "if [ ! -e '/deployments/myapp/state.json' ]", Output: "absent"}, + ssh.MockCommand{Match: "if [ ! -e '/deployments/myapp/state' ]", Output: "absent"}, + ssh.MockCommand{Match: "ss -tln", Output: ssNoEphemeral}, + ssh.MockCommand{Match: "docker run", Output: "abc123def456"}, + ssh.MockCommand{Match: "docker inspect -f '{{.Image}}'", Output: imageDigest}, + ssh.MockCommand{Match: "docker inspect", Output: "running"}, + ssh.MockCommand{Match: "curl -s -o /dev/null", Output: "200"}, + ssh.MockCommand{Match: "curl -sf http://localhost:2019/config/apps/http/servers/srv0", Output: `{"listen":[":80",":443"]}`}, + ssh.MockCommand{Match: "curl -sf -X PATCH", Err: fmt.Errorf("not found")}, + ssh.MockCommand{Match: "curl -sf -X POST http://localhost:2019/config/apps/http/servers/srv0/routes", Output: ""}, + ssh.MockCommand{Match: "rm -f /tmp/teploy_caddy", Output: ""}, + ssh.MockCommand{Match: "cat /deployments/caddy/Caddyfile", Output: "{\n\tadmin 0.0.0.0:2019\n}\n"}, + ssh.MockCommand{Match: "mv /tmp/teploy_caddyfile.tmp", Output: ""}, + ssh.MockCommand{Match: "mkdir /deployments/caddy/.lock", Output: ""}, + ssh.MockCommand{Match: "a=$(docker exec caddy md5sum", Output: "TEPLOY_CADDY_OK"}, + ssh.MockCommand{Match: "docker exec caddy caddy reload", Output: ""}, + ssh.MockCommand{Match: "rmdir /deployments/caddy/.lock", Output: ""}, + ssh.MockCommand{Match: "printf %s", Output: ""}, + ssh.MockCommand{Match: "rm -rf /deployments/myapp/.lock", Output: ""}, + ) + + att := releasemeta.MustAttempt("myapp", version) + err = deployBuiltImageFenced(context.Background(), mock, appCfg, image, version, "srv.test", false, false, "", nil, &att, "plan123abc456def7") + if err != nil { + t.Fatalf("deployBuiltImageFenced: %v", err) + } + + // The receipt carries the plan id. + var provPath string + for path, data := range mock.Files { + if strings.HasSuffix(path, "/provenance.json") { + provPath = path + var p releasemeta.Provenance + if err := json.Unmarshal(data, &p); err != nil { + t.Fatalf("provenance.json malformed: %v", err) + } + if p.PlanID != "plan123abc456def7" { + t.Errorf("provenance.plan_id = %q, want the applied plan id", p.PlanID) + } + if p.ManifestSHA256 != manifestSHA { + t.Errorf("provenance manifest digest = %s, want %s", p.ManifestSHA256, manifestSHA) + } + } + } + if provPath == "" { + t.Fatal("no provenance.json written — the apply receipt is missing") + } + if !strings.Contains(provPath, "/deployments/myapp/meta/att/") { + t.Errorf("provenance not in the attempt namespace: %s", provPath) + } + + // The release state landed (the engine ran for real). + stateData, ok := mock.Files["/deployments/myapp/state.json"] + if !ok { + t.Fatal("state.json not written") + } + var applied state.AppState + if err := json.Unmarshal(stateData, &applied); err != nil { + t.Fatal(err) + } + if applied.CurrentHash != version || applied.ManifestSHA256 != manifestSHA || applied.Generation != 1 { + t.Errorf("applied state wrong: %+v", applied) + } + _ = appliedManifest +} + +// TestApplyDriftErrorEnvelopeClassifiesConflict: under --json, a drift +// refusal must classify as conflict (the world moved), not internal. +func TestApplyDriftErrorEnvelopeClassifiesConflict(t *testing.T) { + drift := driftRefusal("config", "digest moved") + if got := classifyMachineError(drift); got != codeConflict { + t.Errorf("classifyMachineError(drift) = %s, want %s", got, codeConflict) + } + var buf strings.Builder + if err := writeMachineErrorEnvelope(&buf, drift); err != nil { + t.Fatal(err) + } + out := buf.String() + if !strings.Contains(out, `"code":"conflict"`) { + t.Errorf("envelope missing conflict code: %s", out) + } + if !strings.Contains(out, "plan no longer valid") { + t.Errorf("envelope message wrong: %s", out) + } + // Ordinary failures keep their classification. + if got := classifyMachineError(errors.New("boom")); got != codeInternal { + t.Errorf("plain error classified %s", got) + } +} + +// TestApplyRefusesTamperedPlanFile: a plan file edited after planning +// never reaches execution. +func TestApplyRefusesTamperedPlanFile(t *testing.T) { + rec := &PlanRecord{SchemaVersion: PlanRecordSchemaVersion, App: "blog", Server: "srv", TargetVersion: "v1", VersionKnown: true, ConfigDigest: "d1"} + rec.PlanID = computePlanID(rec) + path := writePlan(t, rec) + + raw, err := os.ReadFile(path) + if err != nil { + t.Fatal(err) + } + tampered := strings.Replace(string(raw), `"config_digest": "d1"`, `"config_digest": "d2"`, 1) + if tampered == string(raw) { + t.Fatal("tamper did not apply") + } + if err := os.WriteFile(path, []byte(tampered), 0o600); err != nil { + t.Fatal(err) + } + if _, err := loadPlanFile(path); err == nil || !strings.Contains(err.Error(), "not self-consistent") { + t.Fatalf("tampered plan accepted: %v", err) + } +} diff --git a/internal/cli/autodeploy_serve.go b/internal/cli/autodeploy_serve.go index 0383259..fcb2682 100644 --- a/internal/cli/autodeploy_serve.go +++ b/internal/cli/autodeploy_serve.go @@ -708,7 +708,7 @@ func triggerAutoDeploy(ctx context.Context, executor ssh.Executor, app, branch, // checkout is the provenance source root (C04): revision and context // fingerprint describe the fetched tree this deploy builds. att := releasemeta.MustAttempt(app, version) - return deployBuiltImageFenced(ctx, executor, appCfg, image, version, "localhost", false, needsBuild, buildDir, lk, &att) + return deployBuiltImageFenced(ctx, executor, appCfg, image, version, "localhost", false, needsBuild, buildDir, lk, &att, "") } // fetchCheckout advances buildDir's origin and resets the worktree to the diff --git a/internal/cli/composeplan_test.go b/internal/cli/composeplan_test.go new file mode 100644 index 0000000..871cdbe --- /dev/null +++ b/internal/cli/composeplan_test.go @@ -0,0 +1,174 @@ +package cli + +// C05 conformance: Compose fixtures survive import → plan with the +// known-vs-unresolved image classification asserted. This is the plan +// leg of the "import/render/plan or reject safely" acceptance: every +// supported Compose shape must either classify honestly in a plan or +// have been refused at import (the importer's own suite pins the +// refusals; this suite pins what PLANS say about the survivors). + +import ( + "context" + "os" + "path/filepath" + "strings" + "testing" + + "github.com/useteploy/teploy/internal/config" + "github.com/useteploy/teploy/internal/ssh" +) + +// composePlanFixture writes a compose file and returns the imported +// config (import-time refusals are the importer suite's contract — +// here every fixture must import). +func composePlanFixture(t *testing.T, yml string) *config.AppConfig { + t.Helper() + dir := t.TempDir() + if err := os.WriteFile(filepath.Join(dir, "compose.yml"), []byte(yml), 0o644); err != nil { + t.Fatal(err) + } + cfg, err := config.LoadCompose(dir) + if err != nil { + t.Fatalf("import: %v", err) + } + if cfg == nil { + t.Fatal("no config imported") + } + return cfg +} + +func TestComposePlan_ImageClassification(t *testing.T) { + tests := []struct { + name string + compose string + wantImage string + wantNeedsBuild bool + wantResolution string + }{ + { + name: "build service is unresolved-awaiting-build", + compose: `services: + web: + build: . + ports: ["3000:3000"] +`, + wantNeedsBuild: true, + wantResolution: imageUnresolvedAwaitingBuild, + }, + { + name: "digest-pinned image is resolved-by-digest", + compose: `services: + web: + image: registry.example.com/app@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa + ports: ["3000:3000"] +`, + wantImage: "registry.example.com/app@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", + wantResolution: imageResolvedByDigest, + }, + { + name: "tagged image without server resolution is unresolved-mutable-tag", + compose: `services: + web: + image: nginx:1.27 + ports: ["8080:80"] +`, + wantImage: "nginx:1.27", + wantResolution: imageUnresolvedMutableTag, + }, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + cfg := composePlanFixture(t, tt.compose) + + image := cfg.Image + needsBuild := image == "" + // The classification the plan records. An empty mock executor + // models "docker cannot resolve the ref" — exactly the live + // degradation for an unreachable registry. + prov := resolveDeployProvenance(context.Background(), ssh.NewMockExecutor("srv"), os.Stderr, cfg, ".", image, "v1", "d", needsBuild) + img := planImageIdentity(prov, image, needsBuild) + + if img.NeedsBuild != tt.wantNeedsBuild { + t.Errorf("needs_build = %v, want %v", img.NeedsBuild, tt.wantNeedsBuild) + } + if img.Resolution != tt.wantResolution { + t.Errorf("resolution = %s, want %s", img.Resolution, tt.wantResolution) + } + if tt.wantImage != "" && img.Ref != tt.wantImage { + t.Errorf("ref = %s, want %s", img.Ref, tt.wantImage) + } + if tt.wantResolution == imageResolvedByDigest && img.Digest == "" { + t.Errorf("digest-pinned classification must carry the digest") + } + if tt.wantNeedsBuild && img.ContextFingerprint == "" { + t.Errorf("build classification must bind a context fingerprint — got none (unbound build plan)") + } + // The unresolved notes must SAY it. + notes := unresolvedNotes(img, true) + if tt.wantNeedsBuild && len(notes) == 0 { + t.Errorf("awaiting-build plan must state the unresolved image") + } + }) + } +} + +// TestComposePlan_EffectSetSurvivesImport: an imported stack plans the +// translated surfaces — accessory adds, env keys, storage — so the +// conformance corpus's fixtures survive plan, not just import. +func TestComposePlan_EffectSetSurvivesImport(t *testing.T) { + cfg := composePlanFixture(t, `services: + web: + build: . + ports: ["3000:3000"] + environment: + - DATABASE_URL=postgres://db/blog + volumes: + - uploads:/app/uploads + db: + image: postgres:16 + volumes: ["pgdata:/var/lib/postgresql/data"] +`) + if _, ok := cfg.Accessories["db"]; !ok { + t.Fatalf("db not imported as accessory: %+v", cfg.Accessories) + } + + // First deploy: no deployed manifest to diff against. + acc := accessoryEffects(cfg, nil) + if e := findEffect(acc, "accessory db"); e == nil || e.Action != "add" || e.To != "postgres:16" { + t.Errorf("accessory add missing: %+v", e) + } + env := envEffects(cfg, nil) + if e := findEffect(env, "env DATABASE_URL"); e == nil || e.Action != "add" { + t.Errorf("web env add missing (import dropped it?): %+v", e) + } + // Web volumes are APP storage; the db's pgdata rides the accessory + // effect's detail instead. + storage := storageEffects("blog", cfg, nil) + if e := findEffect(storage, "volume uploads"); e == nil || e.Action != "add" { + t.Errorf("web volume add missing (import dropped it?): %+v", e) + } else if !strings.Contains(e.Detail, "/deployments/blog/volumes/uploads") { + t.Errorf("volume detail must name the managed host path: %+v", e) + } + if e := findEffect(acc, "accessory db"); e == nil || !strings.Contains(e.Detail, "1 volume(s)") { + t.Errorf("accessory detail must summarize its volumes: %+v", e) + } +} + +// TestComposePlan_RefusedShapesNeverPlan: the importer's refused shapes +// surface as import errors, never as plans — the reject-safely half of +// the acceptance. +func TestComposePlan_RefusedShapesNeverPlan(t *testing.T) { + dir := t.TempDir() + if err := os.WriteFile(filepath.Join(dir, "compose.yml"), []byte(`services: + web: + build: . + ports: ["3000:3000"] + jobs: + build: ./jobs +`), 0o644); err != nil { + t.Fatal(err) + } + if _, err := config.LoadCompose(dir); err == nil { + t.Fatal("distinct-builds shape must be refused at import — a plan over it would silently lose build identity") + } +} diff --git a/internal/cli/contracts_golden_test.go b/internal/cli/contracts_golden_test.go index d03e4ed..5b89ca3 100644 --- a/internal/cli/contracts_golden_test.go +++ b/internal/cli/contracts_golden_test.go @@ -12,6 +12,7 @@ import ( "encoding/json" "os" "path/filepath" + "strings" "testing" "time" @@ -68,11 +69,14 @@ func TestContractsAppListEnvelopeGolden(t *testing.T) { writeFixture(t, "app-list-envelope/valid/mi1.json", appListDTO{ MachineInterface: MachineInterface, Host: "srv.example.com", + Errors: []machineError{}, Apps: []appStatusDTO{{ App: "myapp", Domain: "myapp.example.com", Type: "container", Ingress: "caddy", CurrentRelease: releaseStatusDTO{Version: "3", Ports: []int{3000}}, + PreviousRelease: releaseStatusDTO{Version: "2", Ports: []int{3000}}, Containers: []containerDTO{{ID: "9f31c02", Name: "myapp-web-3", Image: "nginx:1.27", State: "running", Status: "Up 4 minutes", CreatedAt: "2026-09-23T11:55:00Z", Process: "web", Version: "3"}}, - Lock: nil, ObservedAt: ts, Errors: nil, + Processes: []processDTO{}, + Lock: nil, ObservedAt: ts, Errors: []machineError{}, }}, ObservedAt: ts, }) @@ -82,7 +86,7 @@ func TestContractsAppListEnvelopeGolden(t *testing.T) { // legacy, not as MI 0. var legacy map[string]any raw, err := json.Marshal(appListDTO{ - Host: "srv.example.com", Apps: []appStatusDTO{}, ObservedAt: ts, + Host: "srv.example.com", Apps: []appStatusDTO{}, ObservedAt: ts, Errors: []machineError{}, }) if err != nil { t.Fatalf("marshal legacy: %v", err) @@ -136,3 +140,68 @@ func TestContractsAttemptNameGolden(t *testing.T) { "../escape.attempt0000000", // path characters }) } + +// TestContractsPlanRecordGolden pins the C05 plan-record shape through +// the REAL record construction: plan id computed by computePlanID, +// marshaled with the PlanRecord's own tags (the same encoder savePlanFile +// uses). Two representative records: a build plan (image +// unresolved-awaiting-build, bound by build inputs) and a prebuilt +// digest-pinned plan (resolved-by-digest) — the known-vs-unresolved +// classification a consumer reads before trusting an apply. +func TestContractsPlanRecordGolden(t *testing.T) { + buildPlan := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, + App: "myapp", + Server: "srv.example.com", + User: "root", + ServerName: "prod", + TargetVersion: "abc1234", + VersionKnown: true, + ConfigDigest: "3f2a9c11d8e4b7065a1c9f0e2b8d7a64c5e3f1b9a0d8c7e6f5a4b3c2d1e0f9a8", + Image: PlanImageIdentity{ + NeedsBuild: true, + Resolution: imageUnresolvedAwaitingBuild, + ContextPath: ".", + ContextFingerprint: "c0ffee11aa22bb33", + Dockerfile: "Dockerfile", + DockerfileSHA256: "deadbeef11", + Platform: "linux/amd64", + }, + TargetState: PlanTargetState{Deployed: true, Generation: 4, CurrentHash: "old1234", ManifestSHA256: "aa11bb22"}, + Effects: PlanEffects{ + Containers: []planChange{{Action: "create", Name: "myapp-web-abc1234", Detail: "web container"}}, + Routing: []planEffect{{Action: "change", Name: "domain", From: "old.example.com", To: "new.example.com", Detail: "routes served by this deployment"}}, + }, + Unresolved: []string{"image unresolved — built at deploy time; the plan binds the build inputs (context . fingerprint c0ffee11aa22bb..., Dockerfile Dockerfile sha deadbeef11...)"}, + } + buildPlan.PlanID = computePlanID(buildPlan) + writeFixture(t, "plan-record/valid/build.json", buildPlan) + + digest := "sha256:" + strings.Repeat("ab", 32) + digestPlan := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, + App: "myapp", + Server: "srv.example.com", + ServerName: "prod", + TargetVersion: "sha256-aaaaaaaaaaaa", + VersionKnown: true, + ConfigDigest: "7d1e0f9a8b7c6d5e4f3a2b1c0d9e8f7a6b5c4d3e2f1a0b9c8d7e6f5a4b3c2d10", + Image: PlanImageIdentity{ + Ref: "registry.example.com/myapp@" + digest, + Resolution: imageResolvedByDigest, + Digest: digest, + }, + TargetState: PlanTargetState{Deployed: false}, + Effects: PlanEffects{Containers: []planChange{{Action: "create", Name: "myapp-web-sha256-aaaaaaaaaaaa", Detail: "web container"}}}, + } + digestPlan.PlanID = computePlanID(digestPlan) + writeFixture(t, "plan-record/valid/prebuilt-digest.json", digestPlan) + + // Invalid: a record whose plan id does not recompute from its + // identity — loadPlanFile refuses it (the schema's pattern cannot + // see inside the hash, so this fixture pins the REFUSAL, not just + // the shape). + tampered := *digestPlan + tampered.ConfigDigest = "0000000000000000000000000000000000000000000000000000000000000000" + writeFixture(t, "plan-record/invalid/tampered-id.json", tampered) +} diff --git a/internal/cli/deploy.go b/internal/cli/deploy.go index f2474b9..ea765ae 100644 --- a/internal/cli/deploy.go +++ b/internal/cli/deploy.go @@ -11,6 +11,7 @@ import ( "os/exec" "os/signal" "regexp" + "sort" "strings" "time" @@ -127,7 +128,7 @@ func runAdHocDeploy(flags *Flags, serverName, appName, image, domain string, por Server: serverName, } - return deployAppConfig(flags, appCfg, serverName, image, version, skipDNSCheck, migrateVolumes) + return deployAppConfig(flags, appCfg, serverName, image, version, skipDNSCheck, migrateVolumes, "") } func runDeploy(flags *Flags, serverName, image, version string, skipDNSCheck bool, parallel int, destination string, migrateVolumes bool, role string, tags map[string]string) error { @@ -175,15 +176,32 @@ func runDeploy(flags *Flags, serverName, image, version string, skipDNSCheck boo // Resolve env_files (SOPS/age-encrypted or plain, decrypted locally) // once, before dispatch — both the single- and multi-server paths then // pick them up from appCfg.Env. Explicit env: keys win over file values. - // - // ${VAR} interpolation applies ONLY to the explicit YAML env: templates, - // expanded exactly once HERE — before file values merge in. Decrypting a - // file whose password contains a literal $ used to hand it to - // os.Expand at serialization time and silently alter it based on the - // operator's environment (audit F59); file/secret values are literal. - // --strict-env (F57/TCL-32) turns an unset ${VAR} into a listed failure - // instead of a silent empty expansion. - if err := expandEnvTemplates(appCfg.Env, flags.StrictEnv); err != nil { + if err := resolveDeployEnv(ctx, appCfg, flags.StrictEnv); err != nil { + return err + } + + // Multi-server deploy: if teploy.yml lists multiple servers and no explicit + // server argument was provided, deploy to all of them in parallel. + if len(appCfg.Servers) > 1 && serverName == "" { + return runMultiDeploy(flags, appCfg, image, version, skipDNSCheck, parallel, migrateVolumes) + } + + return deployAppConfig(flags, appCfg, serverName, image, version, skipDNSCheck, migrateVolumes, "") +} + +// resolveDeployEnv resolves the app config's environment exactly as +// `deploy` does: ${VAR} interpolation applies ONLY to the explicit YAML +// env: templates, expanded exactly once HERE — before file values merge +// in. Decrypting a file whose password contains a literal $ used to hand +// it to os.Expand at serialization time and silently alter it based on +// the operator's environment (audit F59); file/secret values are literal. +// --strict-env (F57/TCL-32) turns an unset ${VAR} into a listed failure +// instead of a silent empty expansion. Explicit env: keys win over file +// values. The apply path (C05) shares this so an applied plan deploys +// with the same env a direct deploy would resolve — one resolution +// semantics, not a fork. +func resolveDeployEnv(ctx context.Context, appCfg *config.AppConfig, strict bool) error { + if err := expandEnvTemplates(appCfg.Env, strict); err != nil { return err } if len(appCfg.EnvFiles) > 0 { @@ -200,14 +218,7 @@ func runDeploy(flags *Flags, serverName, image, version string, skipDNSCheck boo } } } - - // Multi-server deploy: if teploy.yml lists multiple servers and no explicit - // server argument was provided, deploy to all of them in parallel. - if len(appCfg.Servers) > 1 && serverName == "" { - return runMultiDeploy(flags, appCfg, image, version, skipDNSCheck, parallel, migrateVolumes) - } - - return deployAppConfig(flags, appCfg, serverName, image, version, skipDNSCheck, migrateVolumes) + return nil } // parseTagFilters parses repeated "key=value" flag values into a map. @@ -287,8 +298,9 @@ func selectServersByRoleTag(names []string, all map[string]config.Server, role s } // deployAppConfig runs a single-server deploy from an already-loaded AppConfig. -// Used by both runDeploy (from teploy.yml) and runTemplateInstall (from template). -func deployAppConfig(flags *Flags, appCfg *config.AppConfig, serverName, image, version string, skipDNSCheck, migrateVolumes bool) error { +// Used by runDeploy (from teploy.yml), runTemplateInstall (from template), +// and runApply (a verified C05 plan record — planID stamps the receipt). +func deployAppConfig(flags *Flags, appCfg *config.AppConfig, serverName, image, version string, skipDNSCheck, migrateVolumes bool, planID string) error { var err error // 2. Resolve server (single-server deploy). @@ -495,7 +507,7 @@ func deployAppConfig(flags *Flags, appCfg *config.AppConfig, serverName, image, } } - return deployBuiltImageFenced(ctx, executor, appCfg, image, version, host, migrateVolumes, needsBuild, ".", lk, &att) + return deployBuiltImageFenced(ctx, executor, appCfg, image, version, host, migrateVolumes, needsBuild, ".", lk, &att, planID) } // deployBuiltImage runs the shared post-build deploy orchestration: @@ -513,7 +525,7 @@ func deployAppConfig(flags *Flags, appCfg *config.AppConfig, serverName, image, // string for the notification payload (a hostname for the SSH path, // "localhost" for the resident-server path). func deployBuiltImage(ctx context.Context, executor ssh.Executor, appCfg *config.AppConfig, image, version, serverDisplay string, migrateVolumes, needsBuild bool) error { - return deployBuiltImageFenced(ctx, executor, appCfg, image, version, serverDisplay, migrateVolumes, needsBuild, "", nil, nil) + return deployBuiltImageFenced(ctx, executor, appCfg, image, version, serverDisplay, migrateVolumes, needsBuild, "", nil, nil, "") } // deployBuiltImageFenced is deployBuiltImage with the caller's lease and @@ -523,8 +535,11 @@ func deployBuiltImage(ctx context.Context, executor ssh.Executor, appCfg *config // acquires the lock itself (att must still be non-nil for the env file). // sourceRoot is the directory the source was synced/built from ("." for // manual deploys, the fetched checkout for autodeploy) — it keys the -// plan-time provenance (C04); empty means no build provenance. -func deployBuiltImageFenced(ctx context.Context, executor ssh.Executor, appCfg *config.AppConfig, image, version, serverDisplay string, migrateVolumes, needsBuild bool, sourceRoot string, lk *state.Lock, att *releasemeta.Attempt) error { +// plan-time provenance (C04); empty means no build provenance. planID, +// when non-empty, is the C05 plan record the caller verified before +// executing; it is stamped into the provenance receipt so releasemeta +// ties the release back to the reviewed plan. +func deployBuiltImageFenced(ctx context.Context, executor ssh.Executor, appCfg *config.AppConfig, image, version, serverDisplay string, migrateVolumes, needsBuild bool, sourceRoot string, lk *state.Lock, att *releasemeta.Attempt, planID string) error { if att == nil { attVal := releasemeta.MustAttempt(appCfg.App, version) att = &attVal @@ -542,6 +557,7 @@ func deployBuiltImageFenced(ctx context.Context, executor ssh.Executor, appCfg * // describes. Best-effort resolution, but the receipt itself must // land: a missing provenance file is an unwitnessed plan (warned). prov := resolveDeployProvenance(ctx, executor, os.Stdout, appCfg, sourceRoot, image, version, manifestSHA256, needsBuild) + prov.PlanID = planID if err := releasemeta.WriteAttemptProvenance(ctx, executor, *att, prov); err != nil { fmt.Printf("Warning: could not persist the deploy provenance receipt for %s@%s: %v\n", appCfg.App, version, err) } @@ -594,18 +610,9 @@ func deployBuiltImageFenced(ctx context.Context, executor ssh.Executor, appCfg * // 10. Resolve persistent volumes. var volumes map[string]string if len(appCfg.Volumes) > 0 { - volumes = make(map[string]string, len(appCfg.Volumes)) - for name, containerPath := range appCfg.Volumes { - // A host bind mounts a directory the operator owns, exactly as - // given — teploy never creates or relocates it (it may hold a - // clone with credentials, or anything else that is not app data). - if config.IsHostBindVolume(name) { - volumes[name] = containerPath - continue - } - hostPath := fmt.Sprintf("/deployments/%s/volumes/%s", appCfg.App, name) - volumes[hostPath] = containerPath - if _, err := executor.Run(ctx, fmt.Sprintf("mkdir -p %s", hostPath)); err != nil { + volumes = plannedVolumeMounts(appCfg.App, appCfg.Volumes) + for _, hostPath := range managedVolumeHostPaths(appCfg.App, appCfg.Volumes) { + if _, err := executor.Run(ctx, "mkdir -p "+hostPath); err != nil { return fmt.Errorf("creating volume directory %s: %w", hostPath, err) } } @@ -718,6 +725,40 @@ func deployBuiltImageFenced(ctx context.Context, executor ssh.Executor, appCfg * return nil } +// plannedVolumeMounts resolves config volume declarations to the docker +// mount map a deploy uses: a host bind (name starting with "/") mounts a +// directory the operator owns, exactly as given — teploy never creates or +// relocates it (it may hold a clone with credentials, or anything else +// that is not app data); every other name is a teploy-managed volume at +// /deployments//volumes/. Single home: the deploy path and the +// plan storage effects both resolve through it, so a plan can never +// describe a different host layout than the deploy creates. +func plannedVolumeMounts(app string, cfgVolumes map[string]string) map[string]string { + volumes := make(map[string]string, len(cfgVolumes)) + for name, containerPath := range cfgVolumes { + if config.IsHostBindVolume(name) { + volumes[name] = containerPath + continue + } + volumes[fmt.Sprintf("/deployments/%s/volumes/%s", app, name)] = containerPath + } + return volumes +} + +// managedVolumeHostPaths lists the teploy-managed host paths behind +// named volumes (the ones the deploy mkdir's). Host binds are absent — +// they are the operator's directories. +func managedVolumeHostPaths(app string, cfgVolumes map[string]string) []string { + var paths []string + for name := range cfgVolumes { + if !config.IsHostBindVolume(name) { + paths = append(paths, fmt.Sprintf("/deployments/%s/volumes/%s", app, name)) + } + } + sort.Strings(paths) + return paths +} + // imageTagPattern is Docker's tag grammar. Version strings are interpolated // UNQUOTED into remote shell commands via ContainerName (docker rename, docker // rm -f), and until now a version could only come from git output or an diff --git a/internal/cli/deployconfig.go b/internal/cli/deployconfig.go index 9b71520..6e81ab4 100644 --- a/internal/cli/deployconfig.go +++ b/internal/cli/deployconfig.go @@ -29,6 +29,7 @@ func deployConfigFromApp(appCfg *config.AppConfig, image, version string, envFil ContainerPort: appCfg.Port, Publish: appCfg.Publish, StopTimeout: appCfg.StopTimeout, + DrainSeconds: appCfg.DrainSeconds, Memory: appCfg.Memory, CPU: appCfg.CPU, Replicas: appCfg.Replicas, diff --git a/internal/cli/errevelope.go b/internal/cli/errevelope.go index d3005c5..04505c7 100644 --- a/internal/cli/errevelope.go +++ b/internal/cli/errevelope.go @@ -85,10 +85,14 @@ func refuseAdmission(err error) error { // classifyMachineError maps a failed command's error to a taxonomy code. // Wired classes (this slice): config-load failures and deploy admission // refusals → config-invalid; an absent config is a config failure for a -// machine caller the same way a malformed one is. Everything else is -// internal until its site is migrated (S2 generalization). +// machine caller the same way a malformed one is. A plan/apply drift +// refusal (C05) → conflict — the request is coherent, the world moved. +// Everything else is internal until its site is migrated (S2 +// generalization). func classifyMachineError(err error) string { switch { + case errors.Is(err, errPlanDrift): + return codeConflict case errors.Is(err, config.ErrInvalidConfig), errors.Is(err, config.ErrNoConfig), errors.Is(err, errDeployAdmission): @@ -108,6 +112,8 @@ func writeMachineErrorEnvelope(out io.Writer, err error) error { if errors.Is(err, errDeployAdmission) { message = "deploy request refused" } + case codeConflict: + message = "plan no longer valid" } return json.NewEncoder(out).Encode(machineErrorEnvelope{ MachineInterface: MachineInterface, diff --git a/internal/cli/machineinterface.go b/internal/cli/machineinterface.go index f09d098..044ed98 100644 --- a/internal/cli/machineinterface.go +++ b/internal/cli/machineinterface.go @@ -78,6 +78,12 @@ const ( // stable check names, ok/warn/fail results, remediations, and 0/1 // exit semantics — 2 never (that stays drift's) (C09). CapDoctorDiagnostics = "doctor-diagnostics" + // `teploy plan --out` / `teploy apply `: plans carry a binding + // identity (effective-config digest, target version, build inputs, + // server/app, state generation) and apply refuses — naming what + // drifted — when anything moved since the plan (C05). Applied + // releases carry provenance.plan_id. + CapPlanApply = "plan-apply" ) // MachineCapabilities returns every capability token this build @@ -100,6 +106,7 @@ func MachineCapabilities() []string { CapAppListMachine, CapServerStatusMachine, CapDoctorDiagnostics, + CapPlanApply, } sort.Strings(tokens) return tokens diff --git a/internal/cli/machineinterface_test.go b/internal/cli/machineinterface_test.go index 1cb472e..4e7d780 100644 --- a/internal/cli/machineinterface_test.go +++ b/internal/cli/machineinterface_test.go @@ -89,6 +89,7 @@ func TestCapabilityTokenRegistry(t *testing.T) { "error-envelope", "health-modes", "kv-set-stdin", + "plan-apply", "preview-blue-green", "preview-canonical-id", "provenance-records", @@ -124,7 +125,7 @@ func TestCapabilityTokenRegistry(t *testing.T) { CapHealthModes, CapProvenanceRecords, CapReadinessReceipts, CapPreviewCanonicalID, CapRepairDebt, CapPreviewBlueGreen, CapErrorEnvelope, CapAppListMachine, CapServerStatusMachine, - CapDoctorDiagnostics, + CapDoctorDiagnostics, CapPlanApply, } { if !member[token] { t.Fatalf("capability constant %q is not advertised", token) diff --git a/internal/cli/plan.go b/internal/cli/plan.go index 2a7eb3e..0d4241f 100644 --- a/internal/cli/plan.go +++ b/internal/cli/plan.go @@ -4,6 +4,7 @@ import ( "context" "encoding/json" "fmt" + "io" "os" "os/signal" "sort" @@ -11,24 +12,39 @@ import ( "github.com/spf13/cobra" "github.com/useteploy/teploy/internal/config" "github.com/useteploy/teploy/internal/docker" + "github.com/useteploy/teploy/internal/releasemeta" "github.com/useteploy/teploy/internal/ssh" "github.com/useteploy/teploy/internal/state" ) func newPlanCmd(flags *Flags) *cobra.Command { - var version string + var ( + version string + image string + destination string + outFile string + ) cmd := &cobra.Command{ Use: "plan", Short: "Preview what a deploy would change (read-only, no changes made)", Long: "Compares the current on-server state against what a deploy would produce " + - "and prints the difference. Makes no changes. Requires teploy.yml (run from the " + - "app directory); use --host to target a specific server.", + "and prints the difference: containers, routing, environment keys, storage, " + + "resources and accessories, plus what the plan could NOT resolve (an image " + + "still to be built is bound by its build inputs, not guessed). Makes no " + + "changes. Requires teploy.yml (run from the app directory); use --host to " + + "target a specific server.\n\n" + + "With --out FILE the plan is written as a JSON record that `teploy apply` " + + "verifies against config, version, build-input and target-state identity " + + "before executing — anything that moved since the plan invalidates it.", Args: cobra.NoArgs, RunE: func(cmd *cobra.Command, args []string) error { - return runPlan(flags, version) + return runPlan(flags, version, image, destination, outFile) }, } cmd.Flags().StringVar(&version, "version", "", "target version (default: current git short hash, matching deploy)") + cmd.Flags().StringVar(&image, "image", "", "Docker image to deploy (skips build if set) — mirror of deploy --image") + cmd.Flags().StringVarP(&destination, "destination", "d", "", "destination overlay (e.g. staging merges teploy.staging.yml) — mirror of deploy -d") + cmd.Flags().StringVar(&outFile, "out", "", "write the plan record to this file for `teploy apply`") return cmd } @@ -112,20 +128,181 @@ func planChanges(appCfg *config.AppConfig, version string, versionKnown bool, cu return changes, sameVersion } -func runPlan(flags *Flags, version string) error { +// buildPlanRecord computes the full plan (identity + effects) from its +// inputs. The executor is used read-only (state read, container listing, +// image digest resolution). Static apps are handled by the caller. +// Returns the record and whether the target version equals the +// currently-deployed one (display fact for the human/JSON output). +func buildPlanRecord(ctx context.Context, exec ssh.Executor, out io.Writer, appCfg *config.AppConfig, version string, versionKnown, versionExplicit bool, image, destination, serverName string) (*PlanRecord, bool, error) { + current, err := state.Read(ctx, exec, appCfg.App) + if err != nil { + return nil, false, err + } + dk := docker.NewClient(exec) + containers, err := dk.ListContainers(ctx, appCfg.App) + if err != nil { + return nil, false, err + } + + needsBuild := image == "" + _, manifestSHA, err := config.NormalizeAndDigest(appCfg, image) + if err != nil { + return nil, false, fmt.Errorf("normalizing planned manifest: %w", err) + } + + // Image identity + build inputs resolve through the SAME provenance + // path a deploy uses (C04), so plan and deploy cannot disagree about + // what "the build inputs" are. + prov := resolveDeployProvenance(ctx, exec, out, appCfg, ".", image, version, manifestSHA, needsBuild) + img := planImageIdentity(prov, image, needsBuild) + + var view *config.AppliedManifestView + if current != nil && len(current.AppliedManifest) > 0 { + if view, err = config.ParseAppliedManifest(current.AppliedManifest); err != nil { + fmt.Fprintf(out, "Warning: deployed manifest unreadable (%v) — config surfaces diff against an unknown current\n", err) + } + } + + changes, sameVersion := planChanges(appCfg, version, versionKnown, current, containers) + rec := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, + App: appCfg.App, + Server: exec.Host(), + ServerName: serverName, + Destination: destination, + TargetVersion: version, + VersionKnown: versionKnown, + VersionExplicit: versionExplicit, + ConfigDigest: manifestSHA, + Image: img, + TargetState: PlanTargetState{ + Deployed: current != nil, + Generation: 0, + CurrentHash: "", + ManifestSHA256: "", + }, + Effects: PlanEffects{Containers: changes}, + } + if current != nil { + rec.TargetState.Generation = current.Generation + rec.TargetState.CurrentHash = current.CurrentHash + rec.TargetState.ManifestSHA256 = current.ManifestSHA256 + } + rec.User = exec.User() + + rec.Effects.Routing = routingEffects(appCfg, current, view) + rec.Effects.Env = envEffects(appCfg, view) + rec.Effects.Storage = storageEffects(appCfg.App, appCfg, view) + rec.Effects.Resources = resourceEffects(appCfg, view) + rec.Effects.Accessories = accessoryEffects(appCfg, view) + rec.Unresolved = unresolvedNotes(img, versionKnown) + rec.PlanID = computePlanID(rec) + return rec, sameVersion, nil +} + +// planImageIdentity maps a resolved provenance record onto the plan's +// image classification: known bytes (digest-pinned or resolved content +// ID) vs unresolved (mutable tag that could not be resolved, or an +// image that does not exist until the deploy builds it). +func planImageIdentity(prov *releasemeta.Provenance, image string, needsBuild bool) PlanImageIdentity { + img := PlanImageIdentity{ + Ref: image, + NeedsBuild: needsBuild, + ContextPath: prov.ContextPath, + ContextFingerprint: prov.ContextFingerprint, + Dockerfile: prov.Dockerfile, + DockerfileSHA256: prov.DockerfileSHA256, + Platform: prov.Platform, + } + switch { + case needsBuild: + img.Resolution = imageUnresolvedAwaitingBuild + case prov.DigestPinned && prov.ImageDigest != "": + img.Resolution = imageResolvedByDigest + img.Digest = prov.ImageDigest + case prov.ImageDigest != "": + img.Resolution = imageResolvedByImageID + img.Digest = prov.ImageDigest + default: + img.Resolution = imageUnresolvedMutableTag + } + return img +} + +// unresolvedNotes states, in operator terms, everything the plan could +// not know — the KNOWN-vs-UNRESOLVED contract's other half. +func unresolvedNotes(img PlanImageIdentity, versionKnown bool) []string { + var notes []string + switch img.Resolution { + case imageUnresolvedAwaitingBuild: + notes = append(notes, fmt.Sprintf("image unresolved — built at deploy time; the plan binds the build inputs (context %s fingerprint %s, Dockerfile %s sha %s)", + orUnset(img.ContextPath), shortDigest(img.ContextFingerprint), orUnset(img.Dockerfile), shortDigest(img.DockerfileSHA256))) + case imageUnresolvedMutableTag: + notes = append(notes, fmt.Sprintf("image %s is a mutable tag whose content could not be resolved at plan time — the deploy pulls and resolves it fresh", img.Ref)) + case imageResolvedByImageID: + notes = append(notes, fmt.Sprintf("image %s is a MUTABLE tag (currently %s) — the registry can move it before apply; the plan binds the reference, not the bytes", img.Ref, shortDigest(img.Digest))) + } + if !versionKnown { + notes = append(notes, "target version is a deploy-time timestamp — container names are indicative and `teploy apply` refuses this plan (re-plan with --version or a digest-pinned image)") + } + return notes +} + +func orUnset(s string) string { + if s == "" { + return "(unset)" + } + return s +} + +func shortDigest(s string) string { + if len(s) > 16 { + return s[:16] + "..." + } + if s == "" { + return "(unset)" + } + return s +} + +// resolvePlanServerName resolves the logical server the plan targets, +// mirroring connectForApp's resolution so the record names the server +// apply will deploy to. +func resolvePlanServerName(appCfg *config.AppConfig) string { + if appCfg.Server != "" { + return appCfg.Server + } + if len(appCfg.Servers) > 0 { + return appCfg.Servers[0] + } + return "" +} + +func runPlan(flags *Flags, version, image, destination, outFile string) error { ctx, cancel := signal.NotifyContext(context.Background(), os.Interrupt) defer cancel() - appCfg, executor, err := resolveApp(ctx, flags, "") + var appCfg *config.AppConfig + var err error + if destination != "" { + appCfg, err = config.LoadAppWithDestination(".", destination, config.OverlayOptions{Strict: flags.StrictEnv}) + } else { + appCfg, err = config.LoadApp(".") + } + if err != nil { + return err + } + serverName := resolvePlanServerName(appCfg) + executor, err := connectForApp(ctx, flags, appCfg) if err != nil { return err } defer executor.Close() // Static apps deploy via rsync + symlink swap, not containers. A full - // static diff is v2; for now report the current release and note the swap. + // static diff is v2; the plan records identity + the swap. if appCfg.IsStatic() { - return planStatic(ctx, flags, appCfg, executor) + return planStatic(ctx, flags, appCfg, executor, destination, outFile) } // Resolve the target version exactly as `deploy` does — and it has to stay @@ -133,6 +310,7 @@ func runPlan(flags *Flags, version string) error { // --version wins; else a prebuilt image supplies it; else the git hash. // A floating tag (:latest, untagged) yields a timestamp at deploy time, // which is non-deterministic, so the diff can only be indicative. + versionExplicit := version != "" versionKnown := true if version == "" { if appCfg.Image != "" { @@ -147,58 +325,166 @@ func runPlan(flags *Flags, version string) error { } } } + if image == "" { + image = appCfg.Image + } - current, _ := state.Read(ctx, executor, appCfg.App) - dk := docker.NewClient(executor) - containers, err := dk.ListContainers(ctx, appCfg.App) + // Warnings go to stderr under --json so stdout stays one parseable + // document (the connectForApp precedent). + warnOut := io.Writer(os.Stdout) + if flags.JSON { + warnOut = os.Stderr + } + rec, sameVersion, err := buildPlanRecord(ctx, executor, warnOut, appCfg, version, versionKnown, versionExplicit, image, destination, serverName) if err != nil { return err } - changes, sameVersion := planChanges(appCfg, version, versionKnown, current, containers) + if err := renderPlan(flags, rec, sameVersion); err != nil { + return err + } + if outFile != "" { + if err := savePlanFile(outFile, rec); err != nil { + return err + } + if !flags.JSON { + fmt.Printf("\nPlan %s written to %s — apply with: teploy apply %s\n", rec.PlanID, outFile, outFile) + } + } + return nil +} + +// planJSONEnvelope is `teploy plan --json`'s output: the PRE-plan/apply +// keys (app/server/target_version/version_known/same_version/changes) +// unchanged, with the plan-record fields riding additively (the MI +// rules: new keys are additive; removing or renaming the old ones would +// bump MachineInterface). The --out FILE record keeps its own shape — +// it is apply's input, not a display. +type planJSONEnvelope struct { + App string `json:"app"` + Server string `json:"server"` + TargetVersion string `json:"target_version"` + VersionKnown bool `json:"version_known"` + SameVersion bool `json:"same_version"` + Changes []planChange `json:"changes"` + + PlanID string `json:"plan_id"` + Destination string `json:"destination,omitempty"` + ConfigDigest string `json:"config_digest"` + Image PlanImageIdentity `json:"image"` + TargetState PlanTargetState `json:"target_state"` + Effects PlanEffects `json:"effects"` + Unresolved []string `json:"unresolved,omitempty"` +} +// renderPlan prints the plan: identity header, known effects by surface, +// then the unresolved list. JSON mode emits the compat envelope. +func renderPlan(flags *Flags, rec *PlanRecord, sameVersion bool) error { if flags.JSON { - return json.NewEncoder(os.Stdout).Encode(map[string]interface{}{ - "app": appCfg.App, - "server": executor.Host(), - "target_version": version, - "version_known": versionKnown, - "same_version": sameVersion, - "changes": changes, + return json.NewEncoder(os.Stdout).Encode(planJSONEnvelope{ + App: rec.App, + Server: rec.Server, + TargetVersion: rec.TargetVersion, + VersionKnown: rec.VersionKnown, + SameVersion: sameVersion, + Changes: rec.Effects.Containers, + PlanID: rec.PlanID, + Destination: rec.Destination, + ConfigDigest: rec.ConfigDigest, + Image: rec.Image, + TargetState: rec.TargetState, + Effects: rec.Effects, + Unresolved: rec.Unresolved, }) } - fmt.Printf("Plan for %s on %s\n", appCfg.App, executor.Host()) - if versionKnown { - fmt.Printf("Target version: %s\n", version) + fmt.Printf("Plan %s for %s on %s\n", rec.PlanID, rec.App, rec.Server) + if rec.Destination != "" { + fmt.Printf("Destination overlay: %s\n", rec.Destination) + } + fmt.Printf("Config digest: %s\n", rec.ConfigDigest) + fmt.Printf("Image: %s\n", describePlanImage(rec.Image)) + if rec.TargetState.Deployed { + fmt.Printf("Target state: generation %d, deployed hash %s\n", rec.TargetState.Generation, rec.TargetState.CurrentHash) + } else { + fmt.Printf("Target state: not deployed (first deploy)\n") + } + if rec.TargetVersion != "" && rec.VersionKnown { + fmt.Printf("Target version: %s\n", rec.TargetVersion) } else { fmt.Printf("Target version: (timestamp — not predictable; names below are indicative)\n") } - if sameVersion { - fmt.Printf("\nNote: target matches the currently-deployed version %s — a deploy would\n", version) - fmt.Printf("recreate these containers in place rather than blue/green swap.\n") + + printEffectSection("Containers", containerChangesAsEffects(rec.Effects.Containers)) + printEffectSection("Routing", rec.Effects.Routing) + printEffectSection("Environment", rec.Effects.Env) + printEffectSection("Storage", rec.Effects.Storage) + printEffectSection("Resources", rec.Effects.Resources) + printEffectSection("Accessories", rec.Effects.Accessories) + + if len(rec.Unresolved) > 0 { + fmt.Println("\nUnresolved (not known at plan time):") + for _, n := range rec.Unresolved { + fmt.Printf(" ? %s\n", n) + } } + fmt.Println("\nThis is a preview. No changes were made. Run `teploy deploy` to apply, or `teploy apply ` for a bound plan.") + return nil +} - var nCreate, nStop int - fmt.Println() - for _, ch := range changes { - switch ch.Action { +// containerChangesAsEffects re-views container changes through the +// add/remove/change printer without changing their recorded vocabulary. +func containerChangesAsEffects(changes []planChange) []planEffect { + out := make([]planEffect, 0, len(changes)) + for _, c := range changes { + e := planEffect{Action: c.Action, Name: c.Name, Detail: c.Detail} + switch c.Action { case "create": - nCreate++ - fmt.Printf(" + %-40s %s\n", ch.Name, ch.Detail) + e.Action = "add" case "stop": - nStop++ - fmt.Printf(" - %-40s %s\n", ch.Name, ch.Detail) - case "unchanged": - fmt.Printf(" %-40s %s\n", ch.Name, ch.Detail) + e.Action = "remove" } + out = append(out, e) } - if len(changes) == 0 { + return out +} + +func printEffectSection(title string, effects []planEffect) { + fmt.Printf("\n%s:\n", title) + if len(effects) == 0 { fmt.Println(" (no changes)") + return + } + for _, e := range effects { + line := fmt.Sprintf(" %-6s %-34s", e.Action+":", e.Name) + if e.From != "" || e.To != "" { + line += fmt.Sprintf(" %s -> %s", orDash(e.From), orDash(e.To)) + } + if e.Detail != "" { + line += " — " + e.Detail + } + fmt.Println(line) + } +} + +func orDash(s string) string { + if s == "" { + return "-" + } + return s +} + +func describePlanImage(img PlanImageIdentity) string { + switch img.Resolution { + case imageUnresolvedAwaitingBuild: + return fmt.Sprintf("built at deploy time from context %s (fingerprint %s)", orUnset(img.ContextPath), shortDigest(img.ContextFingerprint)) + case imageResolvedByDigest: + return fmt.Sprintf("%s — resolved by digest (%s)", img.Ref, shortDigest(img.Digest)) + case imageResolvedByImageID: + return fmt.Sprintf("%s — mutable tag, content resolved at plan time (%s)", img.Ref, shortDigest(img.Digest)) + default: + return fmt.Sprintf("%s — UNRESOLVED mutable tag (content could not be resolved)", img.Ref) } - fmt.Printf("\n%d to create, %d to stop\n", nCreate, nStop) - fmt.Println("\nThis is a preview. No changes were made. Run `teploy deploy` to apply.") - return nil } // desiredContainer is one container a deploy of the given version would run. @@ -241,21 +527,70 @@ func desiredContainers(appCfg *config.AppConfig, version string) []desiredContai return out } -// planStatic reports the state of a static app. A full static diff is future -// work; for now it notes the swap a deploy performs. -func planStatic(ctx context.Context, flags *Flags, appCfg *config.AppConfig, executor ssh.Executor) error { - _ = ctx // reserved for the future per-file static diff +// planStatic reports the state of a static app and writes a bound plan +// record: identity + digest bind the config; the effect set is the swap +// (a per-file diff is future work). Static deploys currently leave no +// provenance receipt, so apply executes them without a plan-id stamp — +// recorded as a C05 tail. +func planStatic(ctx context.Context, flags *Flags, appCfg *config.AppConfig, executor ssh.Executor, destination, outFile string) error { + _ = ctx + current, _ := state.Read(ctx, executor, appCfg.App) + + version, err := gitShortHash() + if err != nil { + version = "" + } + _, manifestSHA, err := config.NormalizeAndDigest(appCfg, "") + if err != nil { + return fmt.Errorf("normalizing planned manifest: %w", err) + } + rec := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, + App: appCfg.App, + Server: executor.Host(), + ServerName: resolvePlanServerName(appCfg), + Destination: destination, + TargetVersion: version, + VersionKnown: version != "", + ConfigDigest: manifestSHA, + Image: PlanImageIdentity{Resolution: imageUnresolvedAwaitingBuild, NeedsBuild: true}, + Effects: PlanEffects{}, + Unresolved: []string{ + "static app — deploy rsyncs a new release and swaps the `current` symlink; per-file diffing is not yet computed", + "static deploys write no provenance receipt — `teploy apply` binds this plan but cannot stamp the release with its plan id", + }, + } + if current != nil { + rec.TargetState = PlanTargetState{Deployed: true, Generation: current.Generation, CurrentHash: current.CurrentHash, ManifestSHA256: current.ManifestSHA256} + } + rec.PlanID = computePlanID(rec) + if flags.JSON { - return json.NewEncoder(os.Stdout).Encode(map[string]interface{}{ - "app": appCfg.App, - "server": executor.Host(), - "type": "static", - "note": "static apps deploy by rsyncing a new release and swapping the current symlink; per-file diff is not yet computed", + // Old static keys (app/server/type/note) unchanged; record + // identity rides additively. + return json.NewEncoder(os.Stdout).Encode(map[string]any{ + "app": appCfg.App, + "server": executor.Host(), + "type": "static", + "note": "static apps deploy by rsyncing a new release and swapping the current symlink; per-file diff is not yet computed", + "plan_id": rec.PlanID, + "config_digest": rec.ConfigDigest, + "target_state": rec.TargetState, + "unresolved": rec.Unresolved, }) } - fmt.Printf("Plan for %s on %s (static)\n\n", appCfg.App, executor.Host()) + fmt.Printf("Plan %s for %s on %s (static)\n\n", rec.PlanID, appCfg.App, executor.Host()) fmt.Println("A deploy rsyncs a new release and swaps the `current` symlink.") fmt.Println("Per-file diffing for static apps is not yet implemented.") + for _, n := range rec.Unresolved { + fmt.Printf("? %s\n", n) + } + if outFile != "" { + if err := savePlanFile(outFile, rec); err != nil { + return err + } + fmt.Printf("\nPlan %s written to %s — apply with: teploy apply %s\n", rec.PlanID, outFile, outFile) + } fmt.Println("\nThis is a preview. No changes were made.") return nil } diff --git a/internal/cli/planeffects.go b/internal/cli/planeffects.go new file mode 100644 index 0000000..1a95de1 --- /dev/null +++ b/internal/cli/planeffects.go @@ -0,0 +1,353 @@ +package cli + +// Plan effect surfaces (C05): the known-effect diff between the planned +// config and the deployed release, per surface (routing, env, storage, +// resources, accessories). All functions are pure over their inputs — +// the deployed side comes from state.AppState (authoritative identity: +// domain, ingress mode, current port) and the applied manifest view +// (recorded config: env keys, volumes, resources, accessories). A +// missing manifest (pre-manifest release) degrades the diff to "current +// unknown" entries rather than fabricated "no change". + +import ( + "fmt" + "sort" + "strconv" + "strings" + + "github.com/useteploy/teploy/internal/config" + "github.com/useteploy/teploy/internal/state" +) + +// deployedRouting is the deployed routing identity used for diffing. +type deployedRouting struct { + domain string + ingress string + port int + publish []string + deployed bool +} + +// normalizeDomainForDiff lowercases/sorts a comma-separated domain list +// the same way the applied manifest records it, so ordering never reads +// as drift. +func normalizeDomainForDiff(domain string) string { + if domain == "" { + return "" + } + hosts := strings.Split(domain, ",") + for i := range hosts { + hosts[i] = strings.ToLower(strings.TrimSpace(hosts[i])) + } + sort.Strings(hosts) + return strings.Join(hosts, ",") +} + +// routingEffects diffs the planned ingress/routing against the deployed +// release. Routing covers: domain set, ingress mode (caddy vs host +// publish vs external front), the application port (which the Caddy +// route AND the health gate probe), and extra publishes. +func routingEffects(appCfg *config.AppConfig, current *state.AppState, view *config.AppliedManifestView) []planEffect { + var dep deployedRouting + if current != nil { + dep.deployed = true + dep.domain = normalizeDomainForDiff(current.Domain) + dep.ingress = current.IngressMode + dep.port = current.CurrentPort + } + if view != nil && view.Container != nil { + dep.publish = view.Container.Publish + } + + wantIngress := appCfg.Ingress + if wantIngress == "" { + wantIngress = config.IngressCaddy + } + if dep.ingress == "" { + dep.ingress = config.IngressCaddy + } + + firstDeploy := detailFor(!dep.deployed, "first deploy", "") + var effects []planEffect + + if d := normalizeDomainForDiff(appCfg.Domain); d != dep.domain { + effects = append(effects, planEffect{ + Action: addRemoveChange(dep.domain != "", d != ""), + Name: "domain", + From: dep.domain, To: d, + Detail: "routes served by this deployment" + firstDeploy, + }) + } + if wantIngress != dep.ingress { + effects = append(effects, planEffect{ + Action: "change", + Name: "ingress mode", + From: dep.ingress, To: wantIngress, + Detail: "how traffic reaches the app (caddy proxy / host port publish / external front)" + firstDeploy, + }) + } + wantPort := appCfg.Port + if wantPort == 0 && (appCfg.Type == "" || appCfg.Type == config.TypeContainer) { + wantPort = 80 + } + if wantPort != dep.port { + effects = append(effects, planEffect{ + Action: addRemoveChange(dep.port != 0, wantPort != 0), + Name: "application port", + From: portLabel(dep.port), To: portLabel(wantPort), + Detail: "the Caddy route and health gate probe this port" + firstDeploy, + }) + } + + // Extra publishes: recorded in the manifest only. + wantPublish := append([]string(nil), appCfg.Publish...) + sort.Strings(wantPublish) + effects = append(effects, setDiffEffects("publish", dep.publish, wantPublish, "extra docker -p mapping")...) + return effects +} + +// envEffects diffs environment KEY presence (values are redacted by the +// manifest contract; a value change is not plannable — only key and +// env-file-reference changes are). Env-file NAMES diff separately: the +// files' contents are resolved at deploy time. +func envEffects(appCfg *config.AppConfig, view *config.AppliedManifestView) []planEffect { + var haveKeys, haveFiles []string + if view != nil && view.Container != nil { + haveKeys = view.Container.EnvKeys + haveFiles = view.Container.EnvFiles + } + wantKeys := make([]string, 0, len(appCfg.Env)) + for k := range appCfg.Env { + wantKeys = append(wantKeys, k) + } + sort.Strings(wantKeys) + + effects := setDiffEffects("env", haveKeys, wantKeys, "environment variable (key presence planned; values resolved at deploy)") + effects = append(effects, setDiffEffects("env_file", haveFiles, append([]string(nil), appCfg.EnvFiles...), "env file reference (contents resolved at deploy)")...) + if len(appCfg.EnvFiles) > 0 { + effects = append(effects, planEffect{ + Action: "change", + Name: "env values", + Detail: fmt.Sprintf("%d env file(s) merge at deploy time — file VALUES are not planned, only key presence", len(appCfg.EnvFiles)), + }) + } + return effects +} + +// storageEffects diffs volume declarations (config name -> container +// path). The detail names the host path teploy will manage for named +// volumes, using the same resolution the deploy path uses +// (plannedVolumeMounts — one home). +func storageEffects(app string, appCfg *config.AppConfig, view *config.AppliedManifestView) []planEffect { + var have map[string]string + if view != nil && view.Container != nil { + have = view.Container.Volumes + } + want := appCfg.Volumes + + var names []string + seen := map[string]bool{} + for name := range have { + if !seen[name] { + names = append(names, name) + seen[name] = true + } + } + for name := range want { + if !seen[name] { + names = append(names, name) + seen[name] = true + } + } + sort.Strings(names) + + var effects []planEffect + for _, name := range names { + havePath, haveOK := have[name] + wantPath, wantOK := want[name] + switch { + case haveOK && !wantOK: + effects = append(effects, planEffect{Action: "remove", Name: "volume " + name, From: havePath, Detail: "no longer mounted"}) + case !haveOK && wantOK: + effects = append(effects, planEffect{Action: "add", Name: "volume " + name, To: wantPath, Detail: volumeMountDetail(app, name)}) + case havePath != wantPath: + effects = append(effects, planEffect{Action: "change", Name: "volume " + name, From: havePath, To: wantPath, Detail: "container mount path changes"}) + } + } + return effects +} + +// resourceEffects diffs container sizing: replicas, memory, cpu. +func resourceEffects(appCfg *config.AppConfig, view *config.AppliedManifestView) []planEffect { + firstDeploy := view == nil || view.Container == nil + firstNote := detailFor(firstDeploy, "first deploy or pre-manifest release — no recorded current value", "") + + wantReplicas := appCfg.Replicas + if wantReplicas < 1 && (appCfg.Type == "" || appCfg.Type == config.TypeContainer) { + wantReplicas = 1 + } + + var haveReplicas int + var haveMemory, haveCPU string + if !firstDeploy { + haveReplicas = view.Container.Replicas + haveMemory = view.Container.Memory + haveCPU = view.Container.CPU + } + + var effects []planEffect + if haveReplicas != wantReplicas { + effects = append(effects, planEffect{ + Action: addRemoveChange(haveReplicas != 0 && !firstDeploy, wantReplicas != 0), + Name: "replicas", + From: intLabel(haveReplicas), To: intLabel(wantReplicas), + Detail: "web containers kept serving" + firstNote, + }) + } + if haveMemory != appCfg.Memory { + effects = append(effects, planEffect{ + Action: addRemoveChange(haveMemory != "", appCfg.Memory != ""), + Name: "memory limit", + From: haveMemory, To: appCfg.Memory, + Detail: "docker --memory" + firstNote, + }) + } + if haveCPU != appCfg.CPU { + effects = append(effects, planEffect{ + Action: addRemoveChange(haveCPU != "", appCfg.CPU != ""), + Name: "cpu limit", + From: haveCPU, To: appCfg.CPU, + Detail: "docker --cpus" + firstNote, + }) + } + return effects +} + +// accessoryEffects diffs the accessory set and each accessory's image. +func accessoryEffects(appCfg *config.AppConfig, view *config.AppliedManifestView) []planEffect { + var have map[string]config.AppliedManifestAccessory + if view != nil { + have = view.Accessories + } + want := appCfg.Accessories + + names := make([]string, 0, len(have)+len(want)) + seen := map[string]bool{} + for name := range have { + names = append(names, name) + seen[name] = true + } + for name := range want { + if !seen[name] { + names = append(names, name) + } + } + sort.Strings(names) + + var effects []planEffect + for _, name := range names { + h, haveOK := have[name] + w, wantOK := want[name] + switch { + case haveOK && !wantOK: + effects = append(effects, planEffect{Action: "remove", Name: "accessory " + name, From: h.Image, Detail: "retired by this config"}) + case !haveOK && wantOK: + effects = append(effects, planEffect{Action: "add", Name: "accessory " + name, To: w.Image, Detail: accessoryDetail(w, "new stateful service, ensured before app containers start")}) + case h.Image != w.Image: + effects = append(effects, planEffect{Action: "change", Name: "accessory " + name, From: h.Image, To: w.Image, Detail: accessoryDetail(w, "image changes; data volumes persist")}) + } + } + return effects +} + +// setDiffEffects diffs two sorted-or-unsorted string sets into +// add/remove effects with the given label prefix. +func setDiffEffects(label string, have, want []string, detail string) []planEffect { + haveSet := make(map[string]bool, len(have)) + for _, v := range have { + haveSet[v] = true + } + wantSet := make(map[string]bool, len(want)) + for _, v := range want { + wantSet[v] = true + } + var all []string + for v := range haveSet { + all = append(all, v) + } + for v := range wantSet { + if !haveSet[v] { + all = append(all, v) + } + } + sort.Strings(all) + + var effects []planEffect + for _, v := range all { + switch { + case haveSet[v] && !wantSet[v]: + effects = append(effects, planEffect{Action: "remove", Name: label + " " + v, From: v, Detail: detail}) + case !haveSet[v] && wantSet[v]: + effects = append(effects, planEffect{Action: "add", Name: label + " " + v, To: v, Detail: detail}) + } + } + return effects +} + +// addRemoveChange maps presence to the action vocabulary. +func addRemoveChange(have, want bool) string { + switch { + case have && !want: + return "remove" + case !have && want: + return "add" + default: + return "change" + } +} + +func detailFor(cond bool, whenTrue, whenFalse string) string { + if cond { + return " (" + whenTrue + ")" + } + return whenFalse +} + +func portLabel(p int) string { + if p == 0 { + return "" + } + return strconv.Itoa(p) +} + +func intLabel(n int) string { + if n == 0 { + return "" + } + return strconv.Itoa(n) +} + +// volumeMountDetail names the host side a named volume will occupy, +// via the deploy path's own resolution. +func volumeMountDetail(app, name string) string { + if config.IsHostBindVolume(name) { + return "host bind " + name + " mounted as-is (teploy never manages its contents)" + } + return "data at /deployments/" + app + "/volumes/" + name +} + +// accessoryDetail summarizes the non-image delta an accessory effect +// implies. +func accessoryDetail(w config.AccessoryConfig, base string) string { + var extras []string + if len(w.Env) > 0 { + extras = append(extras, fmt.Sprintf("%d env var(s)", len(w.Env))) + } + if len(w.Volumes) > 0 { + extras = append(extras, fmt.Sprintf("%d volume(s)", len(w.Volumes))) + } + if len(extras) == 0 { + return base + } + return base + " (" + strings.Join(extras, ", ") + ")" +} diff --git a/internal/cli/planeffects_test.go b/internal/cli/planeffects_test.go new file mode 100644 index 0000000..3e2a6ae --- /dev/null +++ b/internal/cli/planeffects_test.go @@ -0,0 +1,302 @@ +package cli + +import ( + "encoding/json" + "strings" + "testing" + + "github.com/useteploy/teploy/internal/config" + "github.com/useteploy/teploy/internal/releasemeta" + "github.com/useteploy/teploy/internal/state" +) + +// manifestFixture builds a deployed manifest view from a config — the +// writer/reader pair is round-trip-tested in internal/config. +func manifestView(t *testing.T, cfg *config.AppConfig, appliedImage string) *config.AppliedManifestView { + t.Helper() + data, _, err := config.NormalizeAndDigest(cfg, appliedImage) + if err != nil { + t.Fatalf("NormalizeAndDigest: %v", err) + } + view, err := config.ParseAppliedManifest(data) + if err != nil { + t.Fatalf("ParseAppliedManifest: %v", err) + } + return view +} + +func findEffect(effects []planEffect, name string) *planEffect { + for i := range effects { + if effects[i].Name == name { + return &effects[i] + } + } + return nil +} + +func TestRoutingEffects(t *testing.T) { + deployed := &config.AppConfig{App: "blog", Domain: "old.example.com", Port: 3000, Publish: []string{"127.0.0.1:9100:9000"}} + current := &state.AppState{Domain: "old.example.com", IngressMode: "caddy", CurrentPort: 3000} + view := manifestView(t, deployed, "img:old") + + planned := &config.AppConfig{App: "blog", Domain: "new.example.com,alt.example.com", Port: 8080} + effects := routingEffects(planned, current, view) + + if e := findEffect(effects, "domain"); e == nil || e.Action != "change" || e.From != "old.example.com" || e.To != "alt.example.com,new.example.com" { + t.Errorf("domain effect wrong: %+v", e) + } + if e := findEffect(effects, "application port"); e == nil || e.Action != "change" || e.From != "3000" || e.To != "8080" { + t.Errorf("port effect wrong: %+v", e) + } + if e := findEffect(effects, "publish 127.0.0.1:9100:9000"); e == nil || e.Action != "remove" { + t.Errorf("publish removal missing: %+v", e) + } + if e := findEffect(effects, "ingress mode"); e != nil { + t.Errorf("unchanged ingress must not appear: %+v", e) + } + + // Domain order and case are normalized: not drift. + same := routingEffects(&config.AppConfig{App: "blog", Domain: "OLD.EXAMPLE.COM", Port: 3000}, current, view) + if findEffect(same, "domain") != nil { + t.Errorf("case-folded identical domain reported as drift: %+v", same) + } + + // Ingress mode change. + modeChange := routingEffects(&config.AppConfig{App: "blog", Domain: "old.example.com", Port: 3000, Ingress: config.IngressHost}, + current, view) + if e := findEffect(modeChange, "ingress mode"); e == nil || e.From != "caddy" || e.To != "host" { + t.Errorf("ingress change wrong: %+v", e) + } + + // First deploy: everything is an add. + first := routingEffects(planned, nil, nil) + if e := findEffect(first, "domain"); e == nil || e.Action != "add" { + t.Errorf("first-deploy domain must be an add: %+v", e) + } + if e := findEffect(first, "application port"); e == nil || e.Action != "add" || e.To != "8080" { + t.Errorf("first-deploy port must be an add: %+v", e) + } +} + +func TestEnvEffects(t *testing.T) { + deployed := &config.AppConfig{App: "blog", Port: 80, Env: map[string]string{"KEEP": "1", "GONE": "1"}, EnvFiles: []string{".env.old"}} + view := manifestView(t, deployed, "") + + planned := &config.AppConfig{App: "blog", Port: 80, Env: map[string]string{"KEEP": "2", "NEW": "3"}, EnvFiles: []string{".env.new"}} + effects := envEffects(planned, view) + + if e := findEffect(effects, "env KEEP"); e != nil { + t.Errorf("unchanged key reported: %+v", e) + } + if e := findEffect(effects, "env GONE"); e == nil || e.Action != "remove" { + t.Errorf("removed key missing: %+v", e) + } + if e := findEffect(effects, "env NEW"); e == nil || e.Action != "add" { + t.Errorf("added key missing: %+v", e) + } + if e := findEffect(effects, "env_file .env.old"); e == nil || e.Action != "remove" { + t.Errorf("removed env file missing: %+v", e) + } + if e := findEffect(effects, "env_file .env.new"); e == nil || e.Action != "add" { + t.Errorf("added env file missing: %+v", e) + } + if e := findEffect(effects, "env values"); e == nil { + t.Errorf("env-file values note missing: %+v", e) + } + // Values never appear. + for _, e := range effects { + if e.From == "1" || e.To == "2" || e.To == "3" { + t.Errorf("env VALUE leaked into plan: %+v", e) + } + } +} + +func TestStorageEffects(t *testing.T) { + deployed := &config.AppConfig{App: "blog", Port: 80, Volumes: map[string]string{ + "keep": "/data/keep", + "gone": "/data/gone", + "moved": "/data/old-path", + "/srv/hostbind": "/data/bind", + }} + view := manifestView(t, deployed, "") + + planned := &config.AppConfig{App: "blog", Port: 80, Volumes: map[string]string{ + "keep": "/data/keep", + "moved": "/data/new-path", + "new": "/data/new", + "/srv/hostbind": "/data/bind", + }} + effects := storageEffects("blog", planned, view) + + if e := findEffect(effects, "volume keep"); e != nil { + t.Errorf("unchanged volume reported: %+v", e) + } + if e := findEffect(effects, "volume gone"); e == nil || e.Action != "remove" { + t.Errorf("removed volume missing: %+v", e) + } + if e := findEffect(effects, "volume new"); e == nil || e.Action != "add" || !strings.Contains(e.Detail, "/deployments/blog/volumes/new") { + t.Errorf("added volume must name the managed host path: %+v", e) + } + if e := findEffect(effects, "volume moved"); e == nil || e.Action != "change" || e.From != "/data/old-path" || e.To != "/data/new-path" { + t.Errorf("moved volume wrong: %+v", e) + } + if e := findEffect(effects, "volume /srv/hostbind"); e != nil { + t.Errorf("unchanged host bind reported: %+v", e) + } +} + +func TestResourceEffects(t *testing.T) { + deployed := &config.AppConfig{App: "blog", Port: 80, Replicas: 2, Memory: "256m", CPU: "0.5"} + view := manifestView(t, deployed, "") + + planned := &config.AppConfig{App: "blog", Port: 80, Replicas: 4, Memory: "1g", CPU: "0.5"} + effects := resourceEffects(planned, view) + + if e := findEffect(effects, "replicas"); e == nil || e.Action != "change" || e.From != "2" || e.To != "4" { + t.Errorf("replicas wrong: %+v", e) + } + if e := findEffect(effects, "memory limit"); e == nil || e.From != "256m" || e.To != "1g" { + t.Errorf("memory wrong: %+v", e) + } + if e := findEffect(effects, "cpu limit"); e != nil { + t.Errorf("unchanged cpu reported: %+v", e) + } + + // First deploy: adds against an absent current. + first := resourceEffects(&config.AppConfig{App: "blog", Port: 80, Replicas: 3}, nil) + if e := findEffect(first, "replicas"); e == nil || e.Action != "add" || e.To != "3" { + t.Errorf("first-deploy replicas must be an add: %+v", e) + } +} + +func TestAccessoryEffects(t *testing.T) { + deployed := &config.AppConfig{App: "blog", Port: 80, Accessories: map[string]config.AccessoryConfig{ + "db": {Image: "postgres:15", Port: 5432}, + "gone": {Image: "redis:7", Port: 6379}, + }} + view := manifestView(t, deployed, "") + + planned := &config.AppConfig{App: "blog", Port: 80, Accessories: map[string]config.AccessoryConfig{ + "db": {Image: "postgres:16", Port: 5432}, + "new": {Image: "meilisearch:v1.8", Port: 7700, Volumes: map[string]string{"msdata": "/var/lib/ms"}}, + }} + effects := accessoryEffects(planned, view) + + if e := findEffect(effects, "accessory db"); e == nil || e.Action != "change" || e.From != "postgres:15" || e.To != "postgres:16" { + t.Errorf("db image change wrong: %+v", e) + } + if e := findEffect(effects, "accessory gone"); e == nil || e.Action != "remove" { + t.Errorf("removed accessory missing: %+v", e) + } + if e := findEffect(effects, "accessory new"); e == nil || e.Action != "add" || e.To != "meilisearch:v1.8" { + t.Errorf("added accessory missing: %+v", e) + } +} + +func TestPlanImageIdentity_Classification(t *testing.T) { + cases := []struct { + name string + prov releasemetaProv + image string + needsBuild bool + wantRes string + wantDigest string + }{ + { + name: "digest-pinned", + prov: releasemetaProv{DigestPinned: true, ImageDigest: "sha256:aaaa"}, + image: "repo/app@sha256:aaaa", + wantRes: imageResolvedByDigest, + wantDigest: "sha256:aaaa", + }, + { + name: "build", + prov: releasemetaProv{ContextFingerprint: "fp", DockerfileSHA256: "df"}, + needsBuild: true, + wantRes: imageUnresolvedAwaitingBuild, + }, + { + name: "mutable resolved", + prov: releasemetaProv{ImageDigest: "sha256:bbbb"}, + image: "repo/app:v1", + wantRes: imageResolvedByImageID, + wantDigest: "sha256:bbbb", + }, + { + name: "mutable unresolved", + prov: releasemetaProv{}, + image: "repo/app:v1", + wantRes: imageUnresolvedMutableTag, + }, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + prov := tc.prov.toProvenance() + img := planImageIdentity(prov, tc.image, tc.needsBuild) + if img.Resolution != tc.wantRes { + t.Errorf("resolution = %s, want %s", img.Resolution, tc.wantRes) + } + if img.Digest != tc.wantDigest { + t.Errorf("digest = %s, want %s", img.Digest, tc.wantDigest) + } + if tc.needsBuild && (img.ContextFingerprint != "fp" || img.DockerfileSHA256 != "df") { + t.Errorf("build inputs not carried: %+v", img) + } + }) + } +} + +// releasemetaProv is a test builder for the provenance fields +// planImageIdentity reads. +type releasemetaProv struct { + DigestPinned bool + ImageDigest string + ContextFingerprint string + DockerfileSHA256 string +} + +func (p releasemetaProv) toProvenance() *releasemeta.Provenance { + return &releasemeta.Provenance{ + DigestPinned: p.DigestPinned, + ImageDigest: p.ImageDigest, + ContextFingerprint: p.ContextFingerprint, + DockerfileSHA256: p.DockerfileSHA256, + } +} + +func TestPlanRecordJSON_Shape(t *testing.T) { + rec := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, + PlanID: "abc123def4567890", + App: "blog", Server: "srv.example.com", + TargetVersion: "v1", VersionKnown: true, ConfigDigest: "d1", + Image: PlanImageIdentity{Resolution: imageUnresolvedAwaitingBuild, NeedsBuild: true, ContextFingerprint: "fp"}, + TargetState: PlanTargetState{Deployed: false}, + Effects: PlanEffects{Containers: []planChange{{Action: "create", Name: "blog-web-v1"}}}, + Unresolved: []string{"image unresolved"}, + } + data, err := json.Marshal(rec) + if err != nil { + t.Fatal(err) + } + var m map[string]any + if err := json.Unmarshal(data, &m); err != nil { + t.Fatal(err) + } + for _, key := range []string{"schema_version", "plan_id", "app", "server", "target_version", "version_known", "config_digest", "image", "target_state", "effects", "unresolved"} { + if _, ok := m[key]; !ok { + t.Errorf("plan record JSON missing %s: %s", key, data) + } + } + // Old plan consumers saw these keys; they must stay. + if _, ok := m["changes"]; ok { + t.Error("top-level changes key should live under effects.containers now — but compatibility keys must not be silently renamed") + } + img := m["image"].(map[string]any) + if img["resolution"] != imageUnresolvedAwaitingBuild { + t.Errorf("image.resolution = %v", img["resolution"]) + } + if img["needs_build"] != true { + t.Errorf("image.needs_build = %v", img["needs_build"]) + } +} diff --git a/internal/cli/planrecord.go b/internal/cli/planrecord.go new file mode 100644 index 0000000..08276e1 --- /dev/null +++ b/internal/cli/planrecord.go @@ -0,0 +1,294 @@ +package cli + +// C05 plan/apply binding: the durable plan record `teploy plan --out` +// writes and `teploy apply` verifies against. The contract being pinned: +// +// - a plan is computed against an effective-config digest (the C04 +// NormalizeAndDigest surface, evaluated with the image reference the +// plan resolved — empty for a build, whose image does not exist yet), +// a target identity (app + server), and the target's deployed-state +// generation; +// - apply re-derives every identity input and refuses, naming WHAT +// drifted, when config, target version, build inputs, or target +// state moved since the plan. A stale plan is never executed; +// - what cannot be known at plan time is marked unresolved (build +// image) rather than guessed, and the plan instead binds the build +// INPUTS (context fingerprint + Dockerfile identity) that apply +// re-verifies. +// +// The plan id is a pure function of the binding inputs (schema, app, +// server, user, destination, target version, config digest, build-input +// identities, state generation/hash). WrittenAt is stamped at save and +// deliberately outside the identity, so re-planning the same world +// yields the same id — the retry-stability rule provenance follows. + +import ( + "crypto/sha256" + "encoding/hex" + "encoding/json" + "errors" + "fmt" + "os" + "path/filepath" + "time" + + "github.com/useteploy/teploy/internal/state" +) + +// PlanRecordSchemaVersion is the plan record's schema version. +const PlanRecordSchemaVersion = 1 + +// Image resolution classes: what the plan could and could not bind. +const ( + // imageResolvedByDigest: the reference is digest-pinned; the digest + // is part of the reference itself, so the plan fully knows the bytes. + imageResolvedByDigest = "resolved-by-digest" + // imageResolvedByImageID: a mutable reference whose current content + // ID docker resolved at plan time. The ID is recorded but NOT bound + // (a mutable tag may legitimately move; the changed-mutable-tag + // policy is C04's open tail) — the config digest binds the ref. + imageResolvedByImageID = "resolved-by-image-id" + // imageUnresolvedMutableTag: a mutable reference whose content could + // not be resolved at plan time (registry unreachable, image absent). + imageUnresolvedMutableTag = "unresolved-mutable-tag" + // imageUnresolvedAwaitingBuild: no image yet — the deploy builds it. + // The plan binds the build inputs instead (context fingerprint, + // Dockerfile identity). + imageUnresolvedAwaitingBuild = "unresolved-awaiting-build" +) + +// PlanRecord is the durable artifact of one `teploy plan` run. +type PlanRecord struct { + SchemaVersion int `json:"schema_version"` + PlanID string `json:"plan_id"` + App string `json:"app"` + Server string `json:"server"` + User string `json:"user,omitempty"` + // ServerName is the LOGICAL server (servers.yml key / teploy.yml + // server field) the plan resolved, recorded so apply targets exactly + // the server the plan connected to. + ServerName string `json:"server_name,omitempty"` + // Destination is the overlay the plan was computed with; apply must + // resolve the same merged config, so a flipped overlay is drift. + Destination string `json:"destination,omitempty"` + + TargetVersion string `json:"target_version"` + VersionKnown bool `json:"version_known"` + // VersionExplicit: the version came from --version (operator-pinned) + // rather than derived (git hash / image tag). Explicit versions are + // binding as-is; derived ones are re-derived at apply and compared. + VersionExplicit bool `json:"version_explicit,omitempty"` + + // ConfigDigest is the effective-config digest + // (config.NormalizeAndDigest) over the plan's app config with the + // image reference the plan resolved — "" for a build deploy, whose + // image does not exist at plan time. The DEPLOYED receipt carries + // its own manifest digest computed with the built image; the plan id + // (stamped into that receipt) is the tie between the two. + ConfigDigest string `json:"config_digest"` + + Image PlanImageIdentity `json:"image"` + TargetState PlanTargetState `json:"target_state"` + Effects PlanEffects `json:"effects"` + + // Unresolved lists the data the plan could not know, in operator + // terms (the KNOWN-vs-UNRESOLVED contract: effects above are known; + // everything here is explicitly not). + Unresolved []string `json:"unresolved,omitempty"` + + WrittenAt time.Time `json:"written_at,omitempty"` +} + +// PlanImageIdentity is the plan's knowledge about the image to run. +type PlanImageIdentity struct { + // Ref is the image reference the deploy would use ("" when building). + Ref string `json:"ref,omitempty"` + // NeedsBuild: no prebuilt image; the deploy builds from source. + NeedsBuild bool `json:"needs_build,omitempty"` + // Resolution is one of the image* classification constants. + Resolution string `json:"resolution"` + // Digest is the immutable identity when one was resolvable. + Digest string `json:"digest,omitempty"` + + // Build-input identity (build deploys only) — what apply re-verifies. + ContextPath string `json:"context_path,omitempty"` + ContextFingerprint string `json:"context_fingerprint,omitempty"` + Dockerfile string `json:"dockerfile,omitempty"` + DockerfileSHA256 string `json:"dockerfile_sha256,omitempty"` + Platform string `json:"platform,omitempty"` +} + +// PlanTargetState is the deployed-state identity the plan was computed +// against. Generation is the state counter every successful operation +// increments, so ANY deploy/rollback between plan and apply moves it. +type PlanTargetState struct { + Deployed bool `json:"deployed"` + Generation uint64 `json:"generation"` + CurrentHash string `json:"current_hash,omitempty"` + ManifestSHA256 string `json:"manifest_sha256,omitempty"` +} + +// PlanEffects is the planned effect set, split by surface. Containers +// keep the pre-plan/apply shape (create/stop/unchanged); the field +// surfaces use add/remove/change with from/to. +type PlanEffects struct { + Containers []planChange `json:"containers"` + Routing []planEffect `json:"routing,omitempty"` + Env []planEffect `json:"env,omitempty"` + Storage []planEffect `json:"storage,omitempty"` + Resources []planEffect `json:"resources,omitempty"` + Accessories []planEffect `json:"accessories,omitempty"` +} + +// planEffect is one planned change to a config-driven surface. +type planEffect struct { + Action string `json:"action"` // add | remove | change + Name string `json:"name"` + From string `json:"from,omitempty"` + To string `json:"to,omitempty"` + Detail string `json:"detail,omitempty"` +} + +// computePlanID derives the plan id from the binding inputs. Pure: the +// same planned world always yields the same id. +func computePlanID(rec *PlanRecord) string { + h := sha256.New() + fmt.Fprintf(h, "teploy-plan-v%d\n", rec.SchemaVersion) + fmt.Fprintf(h, "app=%s\n", rec.App) + fmt.Fprintf(h, "server=%s\n", rec.Server) + fmt.Fprintf(h, "user=%s\n", rec.User) + fmt.Fprintf(h, "destination=%s\n", rec.Destination) + fmt.Fprintf(h, "target_version=%s\n", rec.TargetVersion) + fmt.Fprintf(h, "config_digest=%s\n", rec.ConfigDigest) + fmt.Fprintf(h, "context_fingerprint=%s\n", rec.Image.ContextFingerprint) + fmt.Fprintf(h, "dockerfile_sha256=%s\n", rec.Image.DockerfileSHA256) + fmt.Fprintf(h, "generation=%d\n", rec.TargetState.Generation) + fmt.Fprintf(h, "current_hash=%s\n", rec.TargetState.CurrentHash) + return hex.EncodeToString(h.Sum(nil))[:16] +} + +// errPlanDrift marks a plan/apply binding refusal (X02 §2.3: classified +// conflict — the request is coherent, the world moved). +var errPlanDrift = errors.New("plan drift") + +// planDriftError names which binding input moved and the remedy. +type planDriftError struct { + Kind string // "identity" | "config" | "target-version" | "build-inputs" | "target-state" | "unverifiable" + Detail string +} + +func (e *planDriftError) Error() string { + return fmt.Sprintf("plan %s drifted: %s — re-plan (`teploy plan --out `) and apply the fresh plan", e.Kind, e.Detail) +} +func (e *planDriftError) Is(target error) bool { return target == errPlanDrift } + +func driftRefusal(kind, format string, args ...any) error { + return &planDriftError{Kind: kind, Detail: fmt.Sprintf(format, args...)} +} + +// planCurrentFacts is everything apply re-derives before executing. +type planCurrentFacts struct { + App string + Server string // resolved host + Version string // re-derived target version + ConfigDigest string + ContextFingerprint string + DockerfileSHA256 string + State *state.AppState +} + +// verifyPlanBinding checks a plan record against the re-derived world. +// Every refusal names the drifted input; nil means the plan is still the +// reviewed truth and may execute. +func verifyPlanBinding(rec *PlanRecord, cur planCurrentFacts) error { + if rec.App != cur.App { + return driftRefusal("identity", "plan is for app %q, run is for %q", rec.App, cur.App) + } + if rec.Server != cur.Server { + return driftRefusal("identity", "plan targets server %q, run resolved %q", rec.Server, cur.Server) + } + if rec.ConfigDigest != cur.ConfigDigest { + return driftRefusal("config", "effective config digest is now %s, plan reviewed %s (config, overlay %q, or env-file references changed)", cur.ConfigDigest, rec.ConfigDigest, rec.Destination) + } + if !rec.VersionExplicit && rec.TargetVersion != cur.Version { + return driftRefusal("target-version", "target version is now %s, plan reviewed %s", cur.Version, rec.TargetVersion) + } + if rec.Image.NeedsBuild { + if rec.Image.ContextFingerprint != cur.ContextFingerprint { + return driftRefusal("build-inputs", "build context fingerprint is now %s, plan reviewed %s (source changed without a version change)", cur.ContextFingerprint, rec.Image.ContextFingerprint) + } + if rec.Image.DockerfileSHA256 != cur.DockerfileSHA256 { + return driftRefusal("build-inputs", "Dockerfile %s content is now sha256 %s, plan reviewed %s", rec.Image.Dockerfile, cur.DockerfileSHA256, rec.Image.DockerfileSHA256) + } + } + // Target state: any successful operation bumps Generation, so this + // one comparison catches every deploy/rollback/scale in between. + switch { + case rec.TargetState.Deployed && cur.State == nil: + return driftRefusal("target-state", "app %s had generation %d deployed at plan time but has no state now (removed?)", rec.App, rec.TargetState.Generation) + case !rec.TargetState.Deployed && cur.State != nil: + return driftRefusal("target-state", "app %s was deployed since the plan (generation %d, hash %s)", rec.App, cur.State.Generation, cur.State.CurrentHash) + case cur.State != nil && cur.State.Generation != rec.TargetState.Generation: + return driftRefusal("target-state", "generation moved %d -> %d (a deploy or rollback happened in between)", rec.TargetState.Generation, cur.State.Generation) + } + return nil +} + +// savePlanFile writes the record atomically (sibling temp + rename, +// 0600: a plan names a target and carries digests, not secrets, but a +// tampered plan must not be silently swapped into place either). +func savePlanFile(path string, rec *PlanRecord) error { + rec.WrittenAt = time.Now().UTC() + data, err := json.MarshalIndent(rec, "", " ") + if err != nil { + return fmt.Errorf("marshaling plan: %w", err) + } + if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil { + return fmt.Errorf("creating plan directory: %w", err) + } + tmp, err := os.CreateTemp(filepath.Dir(path), ".plan-*") + if err != nil { + return fmt.Errorf("staging plan file: %w", err) + } + defer os.Remove(tmp.Name()) + if _, err := tmp.Write(append(data, '\n')); err != nil { + tmp.Close() + return fmt.Errorf("writing plan file: %w", err) + } + if err := tmp.Chmod(0o600); err != nil { + tmp.Close() + return fmt.Errorf("restricting plan file: %w", err) + } + if err := tmp.Close(); err != nil { + return fmt.Errorf("closing plan file: %w", err) + } + if err := os.Rename(tmp.Name(), path); err != nil { + return fmt.Errorf("publishing plan file: %w", err) + } + return nil +} + +// loadPlanFile reads a plan record and refuses one that is not +// self-consistent: wrong schema version, or a plan id that does not +// recompute from the recorded identity (a tampered or hand-edited plan +// never reaches execution). +func loadPlanFile(path string) (*PlanRecord, error) { + data, err := os.ReadFile(path) + if err != nil { + return nil, fmt.Errorf("reading plan file: %w", err) + } + var rec PlanRecord + if err := json.Unmarshal(data, &rec); err != nil { + return nil, fmt.Errorf("parsing plan file %s: %w", path, err) + } + if rec.SchemaVersion != PlanRecordSchemaVersion { + return nil, fmt.Errorf("plan file %s has schema version %d, this binary writes %d — re-plan with the current binary", path, rec.SchemaVersion, PlanRecordSchemaVersion) + } + if rec.PlanID == "" { + return nil, fmt.Errorf("plan file %s carries no plan id — refuse to apply an unidentified plan; re-plan", path) + } + if want := computePlanID(&rec); want != rec.PlanID { + return nil, fmt.Errorf("plan file %s is not self-consistent: recorded id %s, identity recomputes to %s — the record was edited after planning; re-plan", path, rec.PlanID, want) + } + return &rec, nil +} diff --git a/internal/cli/planrecord_test.go b/internal/cli/planrecord_test.go new file mode 100644 index 0000000..3b31ad4 --- /dev/null +++ b/internal/cli/planrecord_test.go @@ -0,0 +1,262 @@ +package cli + +import ( + "errors" + "os" + "path/filepath" + "strings" + "testing" + + "github.com/useteploy/teploy/internal/state" +) + +func TestComputePlanID_StableAndSensitive(t *testing.T) { + base := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, + App: "blog", Server: "srv.example.com", User: "root", + TargetVersion: "abc1234", ConfigDigest: "digest-1", + TargetState: PlanTargetState{Deployed: true, Generation: 3, CurrentHash: "old9999"}, + } + want := computePlanID(base) + if want == "" || len(want) != 16 { + t.Fatalf("plan id %q is not 16 hex chars", want) + } + // Retry-stable: no time or attempt data participates. + if got := computePlanID(base); got != want { + t.Errorf("plan id not stable: %s vs %s", got, want) + } + // Every binding input moves it. + mutations := map[string]func(*PlanRecord){ + "app": func(r *PlanRecord) { r.App = "other" }, + "server": func(r *PlanRecord) { r.Server = "other.example.com" }, + "user": func(r *PlanRecord) { r.User = "deploy" }, + "destination": func(r *PlanRecord) { r.Destination = "staging" }, + "version": func(r *PlanRecord) { r.TargetVersion = "def9876" }, + "config": func(r *PlanRecord) { r.ConfigDigest = "digest-2" }, + "inputs": func(r *PlanRecord) { r.Image.ContextFingerprint = "fp1" }, + "dockerfile": func(r *PlanRecord) { r.Image.DockerfileSHA256 = "df1" }, + "generation": func(r *PlanRecord) { r.TargetState.Generation = 4 }, + "hash": func(r *PlanRecord) { r.TargetState.CurrentHash = "new1111" }, + } + for name, mutate := range mutations { + t.Run(name, func(t *testing.T) { + clone := *base + mutate(&clone) + if got := computePlanID(&clone); got == want { + t.Errorf("plan id insensitive to %s", name) + } + }) + } +} + +func TestSaveAndLoadPlanFile_RoundTrip(t *testing.T) { + rec := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, + App: "blog", Server: "srv.example.com", + TargetVersion: "abc1234", VersionKnown: true, ConfigDigest: "d1", + Effects: PlanEffects{Containers: []planChange{{Action: "create", Name: "blog-web-abc1234"}}}, + } + rec.PlanID = computePlanID(rec) + + path := filepath.Join(t.TempDir(), "plans", "plan.json") + if err := savePlanFile(path, rec); err != nil { + t.Fatalf("save: %v", err) + } + loaded, err := loadPlanFile(path) + if err != nil { + t.Fatalf("load: %v", err) + } + if loaded.PlanID != rec.PlanID || loaded.App != rec.App || loaded.ConfigDigest != rec.ConfigDigest { + t.Errorf("round trip lost identity: %+v", loaded) + } + if loaded.Effects.Containers[0].Name != "blog-web-abc1234" { + t.Errorf("effects lost: %+v", loaded.Effects) + } + info, err := os.Stat(path) + if err != nil { + t.Fatal(err) + } + if info.Mode().Perm() != 0o600 { + t.Errorf("plan file mode = %v, want 0600", info.Mode().Perm()) + } +} + +func TestLoadPlanFile_RefusesTampered(t *testing.T) { + rec := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, + App: "blog", Server: "srv", + TargetVersion: "v1", ConfigDigest: "d1", + } + rec.PlanID = computePlanID(rec) + + // Tamper: change a binding input without recomputing the id. + tampered := *rec + tampered.ConfigDigest = "d2" + path := filepath.Join(t.TempDir(), "plan.json") + if err := savePlanFile(path, &tampered); err != nil { + t.Fatal(err) + } + if _, err := loadPlanFile(path); err == nil || !strings.Contains(err.Error(), "not self-consistent") { + t.Errorf("tampered plan accepted: %v", err) + } + + // Missing id. + noID := *rec + noID.PlanID = "" + path2 := filepath.Join(t.TempDir(), "plan2.json") + if err := savePlanFile(path2, &noID); err != nil { + t.Fatal(err) + } + if _, err := loadPlanFile(path2); err == nil || !strings.Contains(err.Error(), "no plan id") { + t.Errorf("id-less plan accepted: %v", err) + } + + // Wrong schema version. + future := *rec + future.SchemaVersion = PlanRecordSchemaVersion + 1 + path3 := filepath.Join(t.TempDir(), "plan3.json") + if err := savePlanFile(path3, &future); err != nil { + t.Fatal(err) + } + if _, err := loadPlanFile(path3); err == nil || !strings.Contains(err.Error(), "schema version") { + t.Errorf("future-schema plan accepted: %v", err) + } +} + +// facts builds the current-world view for binding tests. +func facts(from *PlanRecord, mutate func(*planCurrentFacts)) planCurrentFacts { + cur := planCurrentFacts{ + App: from.App, + Server: from.Server, + Version: from.TargetVersion, + ConfigDigest: from.ConfigDigest, + } + if from.TargetState.Deployed { + cur.State = &state.AppState{Generation: from.TargetState.Generation, CurrentHash: from.TargetState.CurrentHash} + } + if mutate != nil { + mutate(&cur) + } + return cur +} + +func TestVerifyPlanBinding_NothingMoved(t *testing.T) { + rec := &PlanRecord{ + SchemaVersion: PlanRecordSchemaVersion, + App: "blog", Server: "srv", + TargetVersion: "abc1234", ConfigDigest: "d1", + TargetState: PlanTargetState{Deployed: true, Generation: 4, CurrentHash: "old1"}, + } + rec.PlanID = computePlanID(rec) + if err := verifyPlanBinding(rec, facts(rec, nil)); err != nil { + t.Fatalf("stable world must verify, got %v", err) + } + // First-deploy plans verify against absent state. + first := &PlanRecord{App: "blog", Server: "srv", TargetVersion: "v1", ConfigDigest: "d1"} + first.PlanID = computePlanID(first) + if err := verifyPlanBinding(first, facts(first, nil)); err != nil { + t.Fatalf("first-deploy plan must verify against no state: %v", err) + } +} + +func TestVerifyPlanBinding_ConfigDrift(t *testing.T) { + rec := &PlanRecord{App: "blog", Server: "srv", TargetVersion: "v1", ConfigDigest: "planned"} + rec.PlanID = computePlanID(rec) + err := verifyPlanBinding(rec, facts(rec, func(c *planCurrentFacts) { c.ConfigDigest = "changed" })) + if !errors.Is(err, errPlanDrift) { + t.Fatalf("config drift must be a plan-drift refusal, got %v", err) + } + drift := &planDriftError{} + if !errors.As(err, &drift) || drift.Kind != "config" { + t.Errorf("drift kind = %+v, want config", err) + } + if !strings.Contains(err.Error(), "re-plan") { + t.Errorf("refusal must name the remedy: %v", err) + } +} + +func TestVerifyPlanBinding_TargetVersionDrift(t *testing.T) { + rec := &PlanRecord{App: "blog", Server: "srv", TargetVersion: "abc1234", ConfigDigest: "d"} + rec.PlanID = computePlanID(rec) + err := verifyPlanBinding(rec, facts(rec, func(c *planCurrentFacts) { c.Version = "def9876" })) + var drift *planDriftError + if !errors.As(err, &drift) || drift.Kind != "target-version" { + t.Fatalf("expected target-version drift, got %v", err) + } + // An explicit --version is binding as-is: no re-derivation mismatch. + explicit := *rec + explicit.VersionExplicit = true + if err := verifyPlanBinding(&explicit, facts(rec, func(c *planCurrentFacts) { c.Version = "def9876" })); err != nil { + t.Errorf("explicit version must not be re-derived: %v", err) + } +} + +func TestVerifyPlanBinding_BuildInputDrift(t *testing.T) { + rec := &PlanRecord{ + App: "blog", Server: "srv", TargetVersion: "abc1234", ConfigDigest: "d", + Image: PlanImageIdentity{NeedsBuild: true, ContextFingerprint: "fp-planned", DockerfileSHA256: "df-planned"}, + } + rec.PlanID = computePlanID(rec) + + err := verifyPlanBinding(rec, facts(rec, func(c *planCurrentFacts) { c.ContextFingerprint = "fp-now" })) + var drift *planDriftError + if !errors.As(err, &drift) || drift.Kind != "build-inputs" { + t.Fatalf("expected build-inputs drift, got %v", err) + } + if !strings.Contains(err.Error(), "fp-planned") || !strings.Contains(err.Error(), "fp-now") { + t.Errorf("refusal must name both fingerprints: %v", err) + } + + err = verifyPlanBinding(rec, facts(rec, func(c *planCurrentFacts) { c.DockerfileSHA256 = "df-now" })) + if !errors.As(err, &drift) || drift.Kind != "build-inputs" { + t.Fatalf("expected dockerfile drift, got %v", err) + } +} + +func TestVerifyPlanBinding_TargetStateDrift(t *testing.T) { + rec := &PlanRecord{ + App: "blog", Server: "srv", TargetVersion: "v1", ConfigDigest: "d", + TargetState: PlanTargetState{Deployed: true, Generation: 4, CurrentHash: "old1"}, + } + rec.PlanID = computePlanID(rec) + + // A deploy happened in between: generation moved. + err := verifyPlanBinding(rec, facts(rec, func(c *planCurrentFacts) { c.State.Generation = 5 })) + var drift *planDriftError + if !errors.As(err, &drift) || drift.Kind != "target-state" { + t.Fatalf("generation move must be target-state drift, got %v", err) + } + if !strings.Contains(err.Error(), "4 -> 5") { + t.Errorf("refusal must name the generation move: %v", err) + } + + // App removed since the plan. + err = verifyPlanBinding(rec, facts(rec, func(c *planCurrentFacts) { c.State = nil })) + if !errors.As(err, &drift) || drift.Kind != "target-state" { + t.Fatalf("removed app must be target-state drift, got %v", err) + } + + // Deployed since a first-deploy plan. + first := &PlanRecord{App: "blog", Server: "srv", TargetVersion: "v1", ConfigDigest: "d"} + first.PlanID = computePlanID(first) + err = verifyPlanBinding(first, facts(first, func(c *planCurrentFacts) { + c.State = &state.AppState{Generation: 1, CurrentHash: "v0"} + })) + if !errors.As(err, &drift) || drift.Kind != "target-state" { + t.Fatalf("deployed-since-plan must be target-state drift, got %v", err) + } +} + +func TestVerifyPlanBinding_IdentityDrift(t *testing.T) { + rec := &PlanRecord{App: "blog", Server: "srv-a", TargetVersion: "v1", ConfigDigest: "d"} + rec.PlanID = computePlanID(rec) + err := verifyPlanBinding(rec, facts(rec, func(c *planCurrentFacts) { c.Server = "srv-b" })) + var drift *planDriftError + if !errors.As(err, &drift) || drift.Kind != "identity" { + t.Fatalf("server change must be identity drift, got %v", err) + } + err = verifyPlanBinding(rec, facts(rec, func(c *planCurrentFacts) { c.App = "other" })) + if !errors.As(err, &drift) || drift.Kind != "identity" { + t.Fatalf("app change must be identity drift, got %v", err) + } +} diff --git a/internal/cli/rollback.go b/internal/cli/rollback.go index f88d665..70f2b28 100644 --- a/internal/cli/rollback.go +++ b/internal/cli/rollback.go @@ -84,19 +84,20 @@ func runRollback(flags *Flags, toHash string) error { } rollbackErr := deploy.Rollback(ctx, executor, os.Stdout, deploy.RollbackConfig{ - App: appCfg.App, - Domain: appCfg.Domain, - StopTimeout: appCfg.StopTimeout, - ToHash: toHash, - Health: healthConfigFrom(appCfg.Health), - TLSCert: tlsCert, - TLSKey: tlsKey, - TLSInternal: tlsInternal, - CaddyExtra: appCfg.CaddyExtra, - Cache: appCfg.Cache, - Firewall: caddyFirewall(appCfg.Firewall), - Access: caddyAccess(appCfg.Access), - Ingress: appCfg.Ingress, + App: appCfg.App, + Domain: appCfg.Domain, + StopTimeout: appCfg.StopTimeout, + DrainSeconds: appCfg.DrainSeconds, + ToHash: toHash, + Health: healthConfigFrom(appCfg.Health), + TLSCert: tlsCert, + TLSKey: tlsKey, + TLSInternal: tlsInternal, + CaddyExtra: appCfg.CaddyExtra, + Cache: appCfg.Cache, + Firewall: caddyFirewall(appCfg.Firewall), + Access: caddyAccess(appCfg.Access), + Ingress: appCfg.Ingress, }) // Fire notification (best-effort). buildNotifier, not NewNotifier: this path diff --git a/internal/cli/root.go b/internal/cli/root.go index 719f494..055ae19 100644 --- a/internal/cli/root.go +++ b/internal/cli/root.go @@ -79,6 +79,7 @@ func NewRootCmd(version string) *cobra.Command { root.AddCommand(newValidateCmd(flags)) root.AddCommand(newDoctorCmd(flags, version)) root.AddCommand(newPlanCmd(flags)) + root.AddCommand(newApplyCmd(flags)) root.AddCommand(newDriftCmd(flags)) root.AddCommand(newHealCmd(flags)) root.AddCommand(newRegistryCmd(flags)) diff --git a/internal/cli/scale.go b/internal/cli/scale.go index 9d88af4..f41a5ff 100644 --- a/internal/cli/scale.go +++ b/internal/cli/scale.go @@ -198,17 +198,18 @@ func rollbackSingleServer(ctx context.Context, appCfg *config.AppConfig, target } return deploy.Rollback(ctx, executor, out, deploy.RollbackConfig{ - App: appCfg.App, - Domain: appCfg.Domain, - StopTimeout: appCfg.StopTimeout, - TLSCert: tlsCert, - TLSKey: tlsKey, - TLSInternal: tlsInternal, - CaddyExtra: appCfg.CaddyExtra, - Cache: appCfg.Cache, - Firewall: caddyFirewall(appCfg.Firewall), - Access: caddyAccess(appCfg.Access), - Ingress: appCfg.Ingress, + App: appCfg.App, + Domain: appCfg.Domain, + StopTimeout: appCfg.StopTimeout, + DrainSeconds: appCfg.DrainSeconds, + TLSCert: tlsCert, + TLSKey: tlsKey, + TLSInternal: tlsInternal, + CaddyExtra: appCfg.CaddyExtra, + Cache: appCfg.Cache, + Firewall: caddyFirewall(appCfg.Firewall), + Access: caddyAccess(appCfg.Access), + Ingress: appCfg.Ingress, }) } diff --git a/internal/cli/template.go b/internal/cli/template.go index 5a52bde..55b1387 100644 --- a/internal/cli/template.go +++ b/internal/cli/template.go @@ -292,7 +292,7 @@ func runTemplateInstall(flags *Flags, name, domain, server string, port int, ext fmt.Printf(" Server: %s\n", server) // Templates are first-deploys by definition, so volume mismatch can't apply yet. - if err := deployAppConfig(flags, appCfg, server, appCfg.Image, "", false, false); err != nil { + if err := deployAppConfig(flags, appCfg, server, appCfg.Image, "", false, false, ""); err != nil { return err } // Unlike `template deploy` (which writes the rendered content, secrets diff --git a/internal/config/app.go b/internal/config/app.go index 242ed73..76a277b 100644 --- a/internal/config/app.go +++ b/internal/config/app.go @@ -380,8 +380,19 @@ type AppConfig struct { // unlimited. CPU string `yaml:"cpu,omitempty" toml:"cpu"` StopTimeout int `yaml:"stop_timeout,omitempty" toml:"stop_timeout"` - Parallel int `yaml:"parallel,omitempty" toml:"parallel"` - Replicas int `yaml:"replicas,omitempty" toml:"replicas"` + // DrainSeconds is how long the PREDECESSOR workload keeps running + // after the traffic switch and before it is stopped (C03 request + // drain): the route no longer sends new requests to it, but in-flight + // requests (downloads, SSE, streaming uploads) finish on their + // existing connections inside the window. Zero (default) keeps the + // historical behavior — stop immediately after the switch. The full + // graceful budget is drain_seconds + stop_timeout (the SIGTERM→SIGKILL + // ladder). Caddy's config-level routing cannot count in-flight + // requests per upstream, so the window is the promised mechanism — + // size it to your longest normal request. + DrainSeconds int `yaml:"drain_seconds,omitempty" toml:"drain_seconds"` + Parallel int `yaml:"parallel,omitempty" toml:"parallel"` + Replicas int `yaml:"replicas,omitempty" toml:"replicas"` // Rollout gates multi-server deploys: a canary wave that must succeed // before the rest of the fleet deploys, and a bounded failure tolerance // for the main wave. Absent (nil) = existing behavior (parallel batches, @@ -1065,6 +1076,9 @@ func (c *AppConfig) validate() error { if c.KeepVersions < 0 { return fmt.Errorf("'keep_versions' must be >= 0 (got %d)", c.KeepVersions) } + if c.DrainSeconds < 0 || c.DrainSeconds > 600 { + return fmt.Errorf("'drain_seconds' must be in 0..600 (got %d)", c.DrainSeconds) + } // Build-context fields only apply when teploy builds the image itself. if c.Dockerfile != "" || c.Context != "" { if c.Image != "" { @@ -1400,6 +1414,9 @@ func mergeConfigs(base, overlay *AppConfig) { if overlay.StopTimeout != 0 { base.StopTimeout = overlay.StopTimeout } + if overlay.DrainSeconds != 0 { + base.DrainSeconds = overlay.DrainSeconds + } if overlay.Parallel != 0 { base.Parallel = overlay.Parallel } diff --git a/internal/config/appliedview.go b/internal/config/appliedview.go new file mode 100644 index 0000000..b9df1a6 --- /dev/null +++ b/internal/config/appliedview.go @@ -0,0 +1,129 @@ +package config + +import ( + "encoding/json" + "fmt" +) + +// AppliedManifestView is the read side of the redacted applied manifest +// NormalizeAndDigest writes into release state: the subset of recorded +// fields a consumer diffs planned config against (C05 plan/apply). It +// lives next to the writer so the shape has one home — a field dropped +// here while NormalizeAndDigest still writes it is a compile-adjacent +// test failure, not a silent plan blind spot. Values stay redacted: +// environment appears as key presence only, secrets never enter. +type AppliedManifestView struct { + App string + DeploymentType string + Domain string + IngressMode string + Container *AppliedManifestContainer + Accessories map[string]AppliedManifestAccessory +} + +// AppliedManifestContainer is the recorded container deployment shape. +// A nil pointer means the deployed release was static (no container +// section in the manifest). +type AppliedManifestContainer struct { + Image string + Port int + Replicas int + Bind string + EnvKeys []string + EnvFiles []string + Volumes map[string]string + Memory string + CPU string + Publish []string + Processes []string +} + +// AppliedManifestAccessory is the recorded accessory shape. +type AppliedManifestAccessory struct { + Image string + Port int + EnvKeys []string + Volumes map[string]string + Publish []string +} + +// ParseAppliedManifest decodes a manifest written by NormalizeAndDigest. +// Malformed JSON, a non-object root, or a container section that is not +// an object is an error — a plan that cannot read what was deployed must +// refuse, never guess (the ReadAttemptProvenance discipline). +func ParseAppliedManifest(data []byte) (*AppliedManifestView, error) { + if len(data) == 0 { + return nil, fmt.Errorf("applied manifest is empty") + } + var root map[string]json.RawMessage + if err := json.Unmarshal(data, &root); err != nil { + return nil, fmt.Errorf("parsing applied manifest: %w", err) + } + view := &AppliedManifestView{} + decodeString(root, "app", &view.App) + decodeString(root, "deployment_type", &view.DeploymentType) + decodeString(root, "domain", &view.Domain) + decodeString(root, "ingress_mode", &view.IngressMode) + + if raw, ok := root["container"]; ok && len(raw) > 0 && string(raw) != "null" { + var c struct { + Image string `json:"image"` + Port int `json:"port"` + Replicas int `json:"replicas"` + Bind string `json:"bind"` + EnvKeys []string `json:"env_keys"` + EnvFiles []string `json:"env_files"` + Volumes map[string]string `json:"volumes"` + Memory string `json:"memory"` + CPU string `json:"cpu"` + Publish []string `json:"publish"` + Processes []string `json:"processes"` + } + if err := json.Unmarshal(raw, &c); err != nil { + return nil, fmt.Errorf("parsing applied manifest container section: %w", err) + } + view.Container = &AppliedManifestContainer{ + Image: c.Image, + Port: c.Port, + Replicas: c.Replicas, + Bind: c.Bind, + EnvKeys: c.EnvKeys, + EnvFiles: c.EnvFiles, + Volumes: c.Volumes, + Memory: c.Memory, + CPU: c.CPU, + Publish: c.Publish, + Processes: c.Processes, + } + } + + if raw, ok := root["accessories"]; ok && len(raw) > 0 && string(raw) != "null" { + var accs map[string]struct { + Image string `json:"image"` + Port int `json:"port"` + EnvKeys []string `json:"env_keys"` + Volumes map[string]string `json:"volumes"` + Publish []string `json:"publish"` + } + if err := json.Unmarshal(raw, &accs); err != nil { + return nil, fmt.Errorf("parsing applied manifest accessories section: %w", err) + } + view.Accessories = make(map[string]AppliedManifestAccessory, len(accs)) + for name, a := range accs { + view.Accessories[name] = AppliedManifestAccessory{ + Image: a.Image, + Port: a.Port, + EnvKeys: a.EnvKeys, + Volumes: a.Volumes, + Publish: a.Publish, + } + } + } + return view, nil +} + +func decodeString(root map[string]json.RawMessage, key string, dst *string) { + if raw, ok := root[key]; ok { + _ = json.Unmarshal(raw, dst) + } +} diff --git a/internal/config/appliedview_test.go b/internal/config/appliedview_test.go new file mode 100644 index 0000000..27e7b2a --- /dev/null +++ b/internal/config/appliedview_test.go @@ -0,0 +1,141 @@ +package config + +import ( + "strings" + "testing" +) + +// roundTripManifest normalizes a config and parses the result back. +func roundTripManifest(t *testing.T, cfg *AppConfig, appliedImage string) *AppliedManifestView { + t.Helper() + data, _, err := NormalizeAndDigest(cfg, appliedImage) + if err != nil { + t.Fatalf("NormalizeAndDigest: %v", err) + } + view, err := ParseAppliedManifest(data) + if err != nil { + t.Fatalf("ParseAppliedManifest: %v", err) + } + return view +} + +func TestParseAppliedManifest_RoundTripContainer(t *testing.T) { + cfg := &AppConfig{ + App: "blog", + Domain: "blog.example.com", + Port: 3000, + Replicas: 2, + Memory: "512m", + CPU: "1.5", + Env: map[string]string{"TOKEN": "secret-value", "MODE": "prod"}, + EnvFiles: []string{".env.production"}, + Volumes: map[string]string{"data": "/var/lib/blog"}, + Publish: []string{"127.0.0.1:9100:9000"}, + Processes: map[string]string{"web": "", "worker": "rake jobs"}, + } + view := roundTripManifest(t, cfg, "registry/blog:abc123") + if view.App != "blog" || view.DeploymentType != TypeContainer || view.IngressMode != IngressCaddy { + t.Errorf("identity fields: %+v", view) + } + if view.Domain != "blog.example.com" { + t.Errorf("domain = %q", view.Domain) + } + c := view.Container + if c == nil { + t.Fatal("container section missing") + } + if c.Image != "registry/blog:abc123" || c.Port != 3000 || c.Replicas != 2 { + t.Errorf("container basics: %+v", c) + } + if c.Memory != "512m" || c.CPU != "1.5" { + t.Errorf("resources: %+v", c) + } + if strings.Join(c.EnvKeys, ",") != "MODE,TOKEN" { + t.Errorf("env keys = %v (want sorted, values never recorded)", c.EnvKeys) + } + if len(c.EnvFiles) != 1 || c.EnvFiles[0] != ".env.production" { + t.Errorf("env files = %v", c.EnvFiles) + } + if c.Volumes["data"] != "/var/lib/blog" { + t.Errorf("volumes = %v", c.Volumes) + } + if len(c.Publish) != 1 || c.Publish[0] != "127.0.0.1:9100:9000" { + t.Errorf("publish = %v", c.Publish) + } + if strings.Join(c.Processes, ",") != "web,worker" { + t.Errorf("processes = %v", c.Processes) + } + // Secrets never enter the manifest — the round trip must not carry values. + data, _, _ := NormalizeAndDigest(cfg, "registry/blog:abc123") + if strings.Contains(string(data), "secret-value") { + t.Error("env VALUE leaked into the applied manifest") + } +} + +func TestParseAppliedManifest_RoundTripAccessories(t *testing.T) { + cfg := &AppConfig{ + App: "blog", + Port: 80, + Accessories: map[string]AccessoryConfig{ + "db": {Image: "postgres:16", Port: 5432, Env: map[string]string{"POSTGRES_PASSWORD": "x"}, Volumes: map[string]string{"pgdata": "/var/lib/postgresql/data"}}, + }, + } + view := roundTripManifest(t, cfg, "") + db, ok := view.Accessories["db"] + if !ok { + t.Fatal("accessory db missing") + } + if db.Image != "postgres:16" || db.Port != 5432 { + t.Errorf("accessory basics: %+v", db) + } + if strings.Join(db.EnvKeys, ",") != "POSTGRES_PASSWORD" { + t.Errorf("accessory env keys = %v", db.EnvKeys) + } + if db.Volumes["pgdata"] != "/var/lib/postgresql/data" { + t.Errorf("accessory volumes = %v", db.Volumes) + } +} + +func TestParseAppliedManifest_StaticHasNoContainer(t *testing.T) { + cfg := &AppConfig{App: "site", Type: TypeStatic, Source: "dist"} + view := roundTripManifest(t, cfg, "") + if view.Container != nil { + t.Errorf("static manifest must not carry a container section, got %+v", view.Container) + } + if view.DeploymentType != TypeStatic { + t.Errorf("deployment type = %q", view.DeploymentType) + } +} + +func TestParseAppliedManifest_RefusesMalformed(t *testing.T) { + cases := []struct { + name string + data string + }{ + {"empty", ""}, + {"not json", "teploy"}, + {"non-object root", `["app"]`}, + {"container not object", `{"container": 3}`}, + {"accessories not object", `{"accessories": []}`}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + if _, err := ParseAppliedManifest([]byte(tc.data)); err == nil { + t.Fatalf("expected refusal for %q", tc.data) + } + }) + } +} + +func TestParseAppliedManifest_NullSectionsTolerated(t *testing.T) { + view, err := ParseAppliedManifest([]byte(`{"app":"x","container":null,"accessories":null}`)) + if err != nil { + t.Fatalf("null sections must parse: %v", err) + } + if view.Container != nil || view.Accessories != nil { + t.Errorf("null sections must stay nil: %+v", view) + } + if view.App != "x" { + t.Errorf("app = %q", view.App) + } +} diff --git a/internal/config/compose.go b/internal/config/compose.go index ccdb61f..2d1699b 100644 --- a/internal/config/compose.go +++ b/internal/config/compose.go @@ -231,6 +231,25 @@ func mapCompose(dir string, compose composeFile) (*AppConfig, error) { cfg.Publish = extraPublish } + // Web service environment translates verbatim into teploy.yml env: + // keys and values carry (values pass teploy's deploy-time ${VAR} + // expansion, matching Compose's interpolation intent for the common + // literal case). Previously this block was silently DROPPED — only + // accessories ever had their environment parsed — so a web + // environment imported "successfully" while deploying without any of + // it (the C05 preserve/translate/reject contract violated by + // omission; found by the plan conformance suite). + if env := parseEnvironment(webService.Environment); len(env) > 0 { + cfg.Env = env + } + + // Web service volumes translate into teploy.yml volumes: a source + // starting with "/" is a host bind and stays one (teploy mounts it + // as-is); anything else is a named volume teploy manages at + // /deployments//volumes/. Like environment, these were + // silently dropped before. + cfg.Volumes = parseWebVolumes(webService.Volumes) + // Web service field contract: the same preserve/translate/reject pass // every other service gets, plus the healthcheck translation that only // has a home for the web process. @@ -803,6 +822,28 @@ func parseServiceVolumes(vols []string) map[string]string { return result } +// parseWebVolumes maps the WEB service's volume declarations: an +// absolute source path ("/srv/data:/container/path") is a HOST BIND and +// keeps its full path as the teploy volume key (IsHostBindVolume's +// contract — teploy mounts the operator's directory as-is); a bare name +// ("uploads:/container/path") is a named volume teploy manages. Unlike +// parseServiceVolumes (the accessory translator, whose basename behavior +// predates this), host binds must NOT lose their path here — the web +// service's data location is the app's, not an implementation detail. +func parseWebVolumes(vols []string) map[string]string { + result := make(map[string]string) + for _, v := range vols { + parts := strings.SplitN(v, ":", 2) + if len(parts) == 2 { + result[parts[0]] = parts[1] + } + } + if len(result) == 0 { + return nil + } + return result +} + // imageBaseName extracts the short name from a Docker image reference. // "postgres:16" -> "postgres", "library/postgres:16" -> "postgres", // "registry.example:5000/postgres:16" -> "postgres": the tag colon is the diff --git a/internal/config/compose_test.go b/internal/config/compose_test.go index c466db2..539d749 100644 --- a/internal/config/compose_test.go +++ b/internal/config/compose_test.go @@ -1585,3 +1585,45 @@ services: t.Errorf("processes = %v, want collapsed single empty-command web (nil map)", cfg.Processes) } } + +// TestLoadCompose_WebServiceEnvAndVolumesPreserved pins the web +// service's environment/volumes translation (previously silently +// DROPPED — only accessories were parsed; found by the C05 plan +// conformance suite: a plan over an imported stack showed no env or +// storage effects because the import had emptied them). +func TestLoadCompose_WebServiceEnvAndVolumesPreserved(t *testing.T) { + dir := t.TempDir() + compose := ` +services: + web: + build: . + ports: ["3000:3000"] + environment: + DATABASE_URL: postgres://db/blog + SESSION_SECRET: literal-$-value + volumes: + - uploads:/app/uploads + - /srv/shared-assets:/app/assets +` + os.WriteFile(filepath.Join(dir, "docker-compose.yml"), []byte(compose), 0644) + + cfg, err := LoadCompose(dir) + if err != nil { + t.Fatalf("LoadCompose: %v", err) + } + if cfg.Env["DATABASE_URL"] != "postgres://db/blog" { + t.Errorf("web environment dropped: %v", cfg.Env) + } + if cfg.Env["SESSION_SECRET"] != "literal-$-value" { + t.Errorf("web environment value mangled: %v", cfg.Env) + } + if cfg.Volumes["uploads"] != "/app/uploads" { + t.Errorf("named volume dropped: %v", cfg.Volumes) + } + if cfg.Volumes["/srv/shared-assets"] != "/app/assets" { + t.Errorf("host bind lost its path: %v", cfg.Volumes) + } + if !IsHostBindVolume("/srv/shared-assets") || IsHostBindVolume("uploads") { + t.Errorf("bind classification broken for the imported keys") + } +} diff --git a/internal/config/drain_test.go b/internal/config/drain_test.go new file mode 100644 index 0000000..b807447 --- /dev/null +++ b/internal/config/drain_test.go @@ -0,0 +1,76 @@ +package config + +import ( + "fmt" + "os" + "path/filepath" + "strings" + "testing" +) + +// TestLoadApp_WithDrainSeconds: drain_seconds parses from teploy.yml and +// reaches AppConfig (the C03 request-drain policy). +func TestLoadApp_WithDrainSeconds(t *testing.T) { + dir := t.TempDir() + content := `app: myapp +domain: myapp.com +stop_timeout: 30 +drain_seconds: 5 +` + if err := os.WriteFile(filepath.Join(dir, "teploy.yml"), []byte(content), 0644); err != nil { + t.Fatal(err) + } + cfg, err := LoadApp(dir) + if err != nil { + t.Fatalf("LoadApp: %v", err) + } + if cfg.DrainSeconds != 5 { + t.Errorf("expected drain_seconds 5, got %d", cfg.DrainSeconds) + } +} + +// TestDrainSecondsValidation: the window is bounded — a negative or +// absurd value is rejected at config load, not mid-deploy. +func TestDrainSecondsValidation(t *testing.T) { + dir := t.TempDir() + for _, tc := range []struct { + value int + ok bool + }{ + {0, true}, {1, true}, {600, true}, + {-1, false}, {601, false}, + } { + content := "app: myapp\ndomain: myapp.com\ndrain_seconds: " + fmt.Sprint(tc.value) + "\n" + if err := os.WriteFile(filepath.Join(dir, "teploy.yml"), []byte(content), 0644); err != nil { + t.Fatal(err) + } + _, err := LoadApp(dir) + if tc.ok && err != nil { + t.Errorf("drain_seconds %d must validate, got %v", tc.value, err) + } + if !tc.ok && (err == nil || !strings.Contains(err.Error(), "drain_seconds")) { + t.Errorf("drain_seconds %d must be rejected naming the field, got %v", tc.value, err) + } + } +} + +// TestDrainSecondsOverlayAndManifest: a destination overlay carries the +// drain policy, and the effective-config manifest (drift identity) +// includes it — a deploy with a changed drain window is a config change. +func TestDrainSecondsOverlayAndManifest(t *testing.T) { + base := &AppConfig{App: "myapp", Domain: "myapp.com", Image: "img:1"} + overlay := &AppConfig{DrainSeconds: 7} + merged := *base + mergeConfigs(&merged, overlay) + if merged.DrainSeconds != 7 { + t.Fatalf("overlay must carry drain_seconds, got %d", merged.DrainSeconds) + } + + manifest, _, err := NormalizeAndDigest(&AppConfig{App: "myapp", Domain: "myapp.com", Image: "img:1", DrainSeconds: 7}, "img:1") + if err != nil { + t.Fatal(err) + } + if !strings.Contains(string(manifest), `"drain_seconds":7`) { + t.Errorf("manifest must carry drain_seconds for drift identity:\n%s", manifest) + } +} diff --git a/internal/config/manifest.go b/internal/config/manifest.go index dc19770..618487f 100644 --- a/internal/config/manifest.go +++ b/internal/config/manifest.go @@ -46,6 +46,7 @@ func NormalizeAndDigest(cfg *AppConfig, appliedImage string) (json.RawMessage, s "port": port, "replicas": replicas, "stop_timeout": stopTimeout, + "drain_seconds": cfg.DrainSeconds, "bind": defaultString(cfg.Bind, ingress == IngressHost, "0.0.0.0"), "platform": cfg.Platform, "build_local": cfg.BuildLocal, diff --git a/internal/deploy/deploy.go b/internal/deploy/deploy.go index 0bd761e..ff745d9 100644 --- a/internal/deploy/deploy.go +++ b/internal/deploy/deploy.go @@ -64,8 +64,17 @@ type Config struct { // Publish adds extra verbatim docker -p mappings beyond ContainerPort's own // (e.g. "0.0.0.0:3001:3001"), for an app with a second listener that needs // its own host port. See AppConfig.Publish. - Publish []string - StopTimeout int // graceful shutdown seconds (default 10) + Publish []string + StopTimeout int // graceful shutdown seconds (default 10) + // DrainSeconds is the request-drain window between the traffic switch + // and predecessor retirement (C03): the predecessor keeps serving + // in-flight requests on its existing connections while new traffic + // goes to the candidates. Zero (default) stops immediately after the + // switch — the historical behavior. The window is time-based: the + // Caddy adapter cannot count in-flight requests per upstream, so the + // drain POLICY is this window plus StopTimeout's SIGTERM→SIGKILL + // ladder; size it to the longest normal request. + DrainSeconds int Replicas int // web process replicas per server (default 1) Health HealthConfig PreDeploy string // hook: runs in web container before traffic switch (failure aborts) @@ -191,6 +200,9 @@ func (c Config) validate() error { if c.StopTimeout < 0 { return fmt.Errorf("stop timeout cannot be negative (got %ds)", c.StopTimeout) } + if c.DrainSeconds < 0 || c.DrainSeconds > 600 { + return fmt.Errorf("drain seconds must be in 0..600 (got %d)", c.DrainSeconds) + } // Health probe mode enum (C03): config-file parsing enforces the fuller // grammar (tcp rejects a path); the shared execution validator covers // the enum so directly constructed Configs (fleet, preview, autodeploy) @@ -354,6 +366,19 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) // one. d.repairOutstandingRecordDebt(ctx, cfg.App, current) + // 1c. Replacement-owner reconciliation (C01-1): when this deploy's + // lock acquisition BROKE a stale predecessor's lock, the previous + // holder may have left in-flight effects on the target — acquisition + // is never proof of quiescence. Decide over the observed evidence + // BEFORE this deploy's first effect; anything but a clean retry + // refuses with the evidence so the leftover generation is reconciled + // deliberately, never blindly redeployed over. + if lk.TookOver() { + if err := d.ReconcileAfterTakeover(ctx, cfg, current); err != nil { + return err + } + } + // 4. Determine host ports for all web replicas. var ports []int if cfg.ingressHost() { @@ -635,17 +660,18 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) } // 6. Start web container(s). - // Fence (F16): container creation is an effect. A lost fence here must - // still restore whatever this deploy displaced (recovery is never - // fenced — see DeployFenced's doc). - if err := lk.Check(ctx, d.exec); err != nil { - return restoreDisplacedAndStarted(err) - } + // Fence (F16, C01-2 composition): container creation runs as a guarded + // effect — the holdership check and the docker run are ONE remote + // command, so no transport window exists in which a takeover lets this + // (possibly broken) holder's starts land inside a new owner's window. + // A lost fence here must still restore whatever this deploy displaced + // (recovery is never fenced — see DeployFenced's doc). + guardPrefix := lk.GuardPrefix() for i := 0; i < replicas; i++ { name := docker.ReplicaContainerName(cfg.App, "web", cfg.Version, i+1, replicas) webContainerNames[i] = name fmt.Fprintf(d.out, "Starting container %s (port %d)...\n", name, ports[i]) - containerID, err := d.docker.Run(ctx, docker.RunConfig{ + containerID, err := d.docker.RunGuarded(ctx, docker.RunConfig{ App: cfg.App, Process: "web", Version: cfg.Version, @@ -662,8 +688,16 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) CPU: cfg.CPU, Name: name, NoHealthcheck: cfg.NoHealthcheck["web"], - }) + }, guardPrefix) if err != nil { + if state.FenceLost(err) { + // The guard refused the run: the lock no longer names this + // operation, and docker never executed. Restore displaced + // work (unfenced), exactly like the old pre-loop Check + // failure — no partial-run reconciliation needed because + // nothing ran. + return restoreDisplacedAndStarted(fmt.Errorf("%w: refusing to start %s's containers", state.ErrFenceLost, cfg.App)) + } // Docker can CREATE a container and still fail the run (port // binding, for one) — that corpse is not in `started`, so it // would outlive this deploy and collide with the next one's @@ -722,6 +756,12 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) if len(ports) > 0 { fmt.Fprintf(d.out, " Readiness: %s\n", readinessSummary(healthCfg, ports[0])) } + // C03's distinction, surfaced: readiness (above — the gate before the + // switch), graceful stop (docker stop's SIGTERM→SIGKILL ladder) and + // request drain (the window the predecessor keeps serving in-flight + // requests after the switch) are SEPARATE policies. Liveness remains + // the container's HEALTHCHECK directive. + fmt.Fprintf(d.out, " Stop policy: graceful stop after %ds (SIGTERM then SIGKILL); drain %s\n", stopTimeout, drainSummary(cfg.DrainSeconds)) for i, p := range ports { if err := d.healthCheck(ctx, p, healthCfg, webBindHost); err != nil { fmt.Fprintf(d.out, " Health check failed for replica %d (port %d): %v\n", i+1, p, err) @@ -759,16 +799,15 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) } // 10. Start non-web process containers (workers, etc. — no replicas, one each). - if err := lk.Check(ctx, d.exec); err != nil { - return fail(err) - } + // Fence (F16, C01-2 composition): same guarded-effect shape as the web + // candidates above — the run is refused in-shell when the fence is lost. for _, process := range sortedProcessNames(processes) { if process == "web" { continue } name := docker.ContainerName(cfg.App, process, cfg.Version) fmt.Fprintf(d.out, "Starting %s...\n", name) - _, err := d.docker.Run(ctx, docker.RunConfig{ + _, err := d.docker.RunGuarded(ctx, docker.RunConfig{ App: cfg.App, Process: process, Version: cfg.Version, @@ -781,8 +820,11 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) Memory: cfg.Memory, CPU: cfg.CPU, NoHealthcheck: cfg.NoHealthcheck[process], - }) + }, guardPrefix) if err != nil { + if state.FenceLost(err) { + return fail(fmt.Errorf("%w: refusing to start %s's %s process", state.ErrFenceLost, cfg.App, process)) + } d.reconcilePartialRun(name) return fail(fmt.Errorf("starting %s: %w", name, err)) } @@ -807,12 +849,13 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) // teploy docker network, so Teploy has nothing to do here. if cfg.usesCaddy() { fmt.Fprintln(d.out, "Updating routes...") - // Fence (F16): the route switch commits traffic to this deploy's - // containers; a late write here would hijack a newer operation's - // route. - if err := lk.Check(ctx, d.exec); err != nil { - return fail(err) - } + // Fence (F16, C01-2 composition): the route switch commits traffic + // to this deploy's containers. The Caddyfile COMMIT (the rename + // that makes the new block authoritative) runs under this deploy's + // fence guard in the same shell — a late write from a broken holder + // cannot hijack a newer operation's route. The pre-switch Check is + // superseded by the composed commit. + cad := d.caddy.WithCommitGuard(guardPrefix) tls := caddy.TLS{Cert: cfg.TLSCert, Key: cfg.TLSKey, Internal: cfg.TLSInternal} if replicas > 1 { upstreams := make([]caddy.Upstream, replicas) @@ -821,12 +864,18 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) } // Caddy's active upstream checks probe the SAME path the deploy // readiness gate used (F47) — the block used to hardcode /up. - if err := d.caddy.SetLoadBalancerHealth(ctx, cfg.App, cfg.Domain, upstreams, healthCfg.Path, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access); err != nil { + if err := cad.SetLoadBalancerHealth(ctx, cfg.App, cfg.Domain, upstreams, healthCfg.Path, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access); err != nil { + if state.FenceLost(err) { + return fail(fmt.Errorf("%w: refusing to switch %s's route", state.ErrFenceLost, cfg.App)) + } return fail(fmt.Errorf("updating load balancer route: %w", err)) } fmt.Fprintf(d.out, " Traffic load-balanced across %d replicas\n", replicas) } else { - if err := d.caddy.SetRoute(ctx, cfg.App, cfg.Domain, webContainerName, containerPort, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access); err != nil { + if err := cad.SetRoute(ctx, cfg.App, cfg.Domain, webContainerName, containerPort, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access); err != nil { + if state.FenceLost(err) { + return fail(fmt.Errorf("%w: refusing to switch %s's route", state.ErrFenceLost, cfg.App)) + } return fail(fmt.Errorf("updating route: %w", err)) } fmt.Fprintln(d.out, " Traffic routed to new container") @@ -923,6 +972,28 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) } } + // 13z. Request drain (C03): the route switched to the candidates in + // step 11; the predecessor still holds IN-FLIGHT requests on its + // existing connections (downloads, SSE, streaming uploads). Hold the + // configured window before retiring it so those requests complete — + // Caddy's config-level routing cannot count in-flight requests per + // upstream, so the window IS the drain policy (plus the graceful-stop + // ladder inside docker stop). Only a caddy-routed blue/green switch + // has a serving predecessor to drain: external ingress is the + // operator's edge (nothing to drain here), and the recreate strategy + // already stopped the fixed-port workload before the candidates + // started (its downtime is the recreate tradeoff, advertised at the + // plan). Skipped when there is nothing to retire or the window is 0 + // (the historical behavior). + if cfg.DrainSeconds > 0 && cfg.usesCaddy() && predecessorsListed && len(predecessors) > 0 { + fmt.Fprintf(d.out, "Draining %s's predecessor for %ds (in-flight requests complete; new traffic serves %s)...\n", cfg.App, cfg.DrainSeconds, cfg.Version) + select { + case <-ctx.Done(): + fmt.Fprintf(d.out, " Drain window cut short (%v) — proceeding to predecessor retirement\n", ctx.Err()) + case <-time.After(time.Duration(cfg.DrainSeconds) * time.Second): + } + } + // 14. Stop the predecessor workload snapshotted in step 6b (all // processes + all replicas). For same-version redeploys the old // containers were renamed to _replaced; remove them after stopping so diff --git a/internal/deploy/drain_integration_test.go b/internal/deploy/drain_integration_test.go new file mode 100644 index 0000000..16cac8c --- /dev/null +++ b/internal/deploy/drain_integration_test.go @@ -0,0 +1,401 @@ +//go:build integration + +// Fixture-gated verification of C03's request-drain acceptance against a +// REAL SSH+Docker host with a REAL Caddy: a blue/green switch driven +// through the production caddy.Client (lock, adapt gate, guarded commit, +// reload, delivery verification), a continuous request stream across the +// switch, and a long in-flight request that must complete inside the +// drain window — ZERO failed requests for the declared fixture. +// +// Same fixture contract as the other harnesses: +// +// TEPLOY_FAULT_HOST=127.0.0.1:50075 \ +// TEPLOY_FAULT_USER=tyler \ +// TEPLOY_FAULT_KEY=~/.colima/_lima/_config/user \ +// go test -tags integration -run TestDrainIntegration -v ./internal/deploy +// +// Needs image pulls (caddy:2-alpine, python:3-alpine, curlimages/curl) and +// skips when /deployments/caddy already exists (a provisioned real caddy +// would conflict with the fixture's container name). Disposable fixture +// only: creates and removes drainprobe-* containers plus +// /deployments/{caddy,drainprobe}. +// +// Honest scope: Caddy's config-level routing cannot count in-flight +// requests per upstream, so the drain POLICY is the time window between +// the switch and the predecessor stop plus the SIGTERM→SIGKILL ladder +// inside docker stop. This fixture proves exactly that promise: normal +// requests never fail across the switch, and a request that started +// BEFORE the switch completes inside the window. The negative control +// (stop -t 0 with an in-flight request) proves the fixture can detect a +// broken promise — the drain window is what saves the long request. +package deploy + +import ( + "context" + "fmt" + "strings" + "testing" + "time" + + "github.com/useteploy/teploy/internal/caddy" + "github.com/useteploy/teploy/internal/ssh" +) + +// drainServerPy is the blue/green fixture app: 200 on everything, a +// ?ms=N /slow path for long requests, and graceful SIGTERM (finish +// in-flight, then exit) so docker stop's ladder is observable. +const drainServerPy = `import os, signal, threading, time +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer +NAME = os.environ.get("SERVER_NAME", "?").encode() +class H(BaseHTTPRequestHandler): + def do_GET(self): + if self.path.startswith("/slow"): + ms = 1000 + if "ms=" in self.path: + try: ms = int(self.path.split("ms=")[-1].split("&")[0]) + except ValueError: pass + time.sleep(ms / 1000.0) + body = NAME + try: + self.send_response(200) + self.send_header("Content-Length", str(len(body))) + self.end_headers() + self.wfile.write(body) + except Exception: + pass + def log_message(self, *a): + pass +srv = ThreadingHTTPServer(("0.0.0.0", 8080), H) +def stop(sig, frm): + threading.Thread(target=srv.shutdown, daemon=True).start() +signal.signal(signal.SIGTERM, stop) +signal.signal(signal.SIGINT, stop) +srv.serve_forever() +` + +const ( + drainNet = "teploy" + drainAppPrefix = "drainprobe" + drainBlue = drainAppPrefix + "-web-blue" + drainGreen = drainAppPrefix + "-web-green" + drainCaddyName = "caddy" // the production reload/adapt paths hardcode this name + drainLoad = drainAppPrefix + "-load" + drainSlowTmpl = drainAppPrefix + "-slow" +) + +func drainRun(t *testing.T, exec ssh.Executor, ctx context.Context, cmd string) { + t.Helper() + if out, err := exec.Run(ctx, cmd); err != nil { + t.Fatalf("%s: %v\n%s", cmd, err, out) + } +} + +func drainContainerOut(t *testing.T, exec ssh.Executor, ctx context.Context, name string) (string, string) { + t.Helper() + out, err := exec.Run(ctx, "docker inspect -f '{{.State.Status}}:{{.State.ExitCode}}' "+name+" 2>/dev/null || true") + if err != nil { + t.Fatalf("inspecting %s: %v", name, err) + } + fields := strings.SplitN(strings.TrimSpace(out), ":", 2) + if len(fields) != 2 { + return "", "" + } + return fields[0], fields[1] +} + +// TestDrainIntegration_BlueGreenZeroFailedRequests is C03's declared +// blue/green fixture: continuous traffic through a real Caddy across a +// production-shaped route switch, a drain window, then a graceful stop — +// zero failed requests, and the long request completes on blue. +func TestDrainIntegration_BlueGreenZeroFailedRequests(t *testing.T) { + host, user, key, _ := reconcileFixtureEnv(t) + ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute) + defer cancel() + + exec, err := ssh.Connect(ctx, ssh.ConnectConfig{Host: host, User: user, KeyPath: key}) + if err != nil { + t.Fatalf("connecting: %v", err) + } + // Close AFTER the cleanup work: registered as the FIRST cleanup so + // LIFO runs it LAST — a plain defer would close the session before + // the resource cleanups' commands could run on it. + t.Cleanup(func() { exec.Close() }) + if out, err := exec.Run(ctx, "docker version --format '{{.Server.Version}}'"); err != nil { + t.Skipf("fixture host has no reachable docker daemon (%v)", err) + } else { + t.Logf("fixture docker server %s", strings.TrimSpace(out)) + } + // A provisioned caddy tree means a real setup lives here — do not + // fight it for the container name. + if out, _ := exec.Run(ctx, "test -e /deployments/caddy/Caddyfile && echo yes || echo no"); strings.TrimSpace(out) == "yes" { + t.Skip("fixture has a provisioned /deployments/caddy — remove it or use a disposable host for the drain fixture") + } + + // Images + network. + for _, image := range []string{"caddy:2-alpine", "python:3-alpine", "curlimages/curl:latest"} { + if out, err := exec.Run(ctx, "docker pull "+image); err != nil { + t.Skipf("cannot pull %s (%v) — pre-pull it on the fixture", image, err) + } else if out != "" { + t.Logf("pulled %s", image) + } + } + drainRun(t, exec, ctx, "docker network create "+drainNet+" 2>/dev/null || true") + + // Cleanup on exit: containers, fixture dirs, the caddy container we + // temporarily own. The teploy network is left alone (shared). + t.Cleanup(func() { + cctx, ccancel := context.WithTimeout(context.Background(), 60*time.Second) + defer ccancel() + for _, n := range []string{drainLoad, drainBlue, drainGreen, drainCaddyName} { + exec.Run(cctx, "docker rm -f "+n) + } + for i := 0; i < 8; i++ { + exec.Run(cctx, fmt.Sprintf("docker rm -f %s-%d", drainSlowTmpl, i)) + } + exec.Run(cctx, "rm -rf /deployments/caddy /deployments/"+drainAppPrefix) + }) + + // Fixture app + caddy bootstrap. + drainRun(t, exec, ctx, "mkdir -p /deployments/caddy /deployments/"+drainAppPrefix) + if err := exec.Upload(ctx, strings.NewReader(drainServerPy), "/deployments/"+drainAppPrefix+"/server.py", "0644"); err != nil { + t.Fatalf("uploading server.py: %v", err) + } + if err := exec.Upload(ctx, strings.NewReader("{\n\tadmin 127.0.0.1:2019\n}\n"), "/deployments/caddy/Caddyfile", "0644"); err != nil { + t.Fatalf("uploading initial Caddyfile: %v", err) + } + appRun := func(name, serverName string) string { + return fmt.Sprintf( + "docker run -d --name %s --network %s --label teploy.app=%s -e SERVER_NAME=%s -v /deployments/%s/server.py:/srv/server.py:ro python:3-alpine python /srv/server.py", + name, drainNet, drainAppPrefix, serverName, drainAppPrefix) + } + drainRun(t, exec, ctx, appRun(drainBlue, "blue")) + drainRun(t, exec, ctx, fmt.Sprintf( + "docker run -d --name %s --network %s -v /deployments/caddy:/etc/caddy -p 127.0.0.1:0:80 caddy:2-alpine", drainCaddyName, drainNet)) + + waitHTTP := func(container string, timeout time.Duration) bool { + deadline := time.Now().Add(timeout) + for time.Now().Before(deadline) { + out, err := exec.Run(ctx, fmt.Sprintf( + "docker exec %s python -c \"import urllib.request;urllib.request.urlopen('http://127.0.0.1:8080/id',timeout=2)\" 2>/dev/null && echo ok || true", container)) + if err == nil && strings.TrimSpace(out) == "ok" { + return true + } + time.Sleep(300 * time.Millisecond) + } + return false + } + if !waitHTTP(drainBlue, 30*time.Second) { + t.Fatal("blue never became ready — fixture broken") + } + + // Caddy's in-network IP is the site address (an IP literal forces the + // http:// scheme — no ACME for a fixture host). + caddyIP := strings.TrimSpace(mustOut(t, exec, ctx, + "docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "+drainCaddyName)) + if caddyIP == "" || !strings.Contains(caddyIP, ".") { + t.Fatalf("no IPv4 for the caddy container (got %q)", caddyIP) + } + + // The PRODUCTION switch machinery: real lock, real adapt gate, real + // reload, real delivery verification. + client := caddy.NewClient(exec) + switchRoute := func(upstream string) error { + return client.SetRoute(ctx, drainAppPrefix, caddyIP, upstream, 8080, caddy.TLS{}, "", nil, caddy.Firewall{}, caddy.Access{}) + } + if err := switchRoute(drainBlue); err != nil { + t.Fatalf("routing to blue through the production client: %v", err) + } + probeOnce := func() string { + out, err := exec.Run(ctx, fmt.Sprintf( + "docker run --rm --network %s curlimages/curl:latest -s --max-time 5 -H %s http://%s:80/id", + drainNet, ssh.ShellQuote("Host: "+caddyIP), drainCaddyName)) + if err != nil { + return "err:" + err.Error() + } + return strings.TrimSpace(out) + } + if got := probeOnce(); got != "blue" { + t.Fatalf("fixture sanity: expected blue through caddy, got %q", got) + } + + // Continuous load: ~8 req/s for the whole scenario; every non-200 or + // connection error is a failure line. Runs detached, results in a + // volume the assertions read afterwards. + drainRun(t, exec, ctx, "docker volume create "+drainAppPrefix+"-out 2>/dev/null || true") + t.Cleanup(func() { + cctx, ccancel := context.WithTimeout(context.Background(), 15*time.Second) + defer ccancel() + exec.Run(cctx, "docker volume rm "+drainAppPrefix+"-out") + }) + loadLoop := fmt.Sprintf( + `i=0; fails=0; while [ $i -lt 90 ]; do `+ + `code=$(curl -s -o /dev/null -w '%%{http_code}' --max-time 5 -H %s http://%s:80/id 2>/dev/null); `+ + `if [ "$code" != "200" ]; then fails=$((fails+1)); echo "fail:$i:$code" >> /out/fail.log; fi; `+ + `i=$((i+1)); sleep 0.15; done; echo $fails > /out/failcount`, + ssh.ShellQuote("Host: "+caddyIP), drainCaddyName) + drainRun(t, exec, ctx, fmt.Sprintf( + "docker run -d --user root --name %s --network %s -v %s:/out curlimages/curl:latest sh -c %s", + drainLoad, drainNet, drainAppPrefix+"-out", ssh.ShellQuote(loadLoop))) + + // A long request (~3.5s) on BLUE, started ~1s BEFORE the switch: it is + // in flight across the switch and must complete inside the drain + // window (switch + 5s) with blue's body. + startSlow := func(tag string) { + inner := fmt.Sprintf("curl -s --max-time 10 -H %s http://%s:80/slow?ms=3500 -o /out/%s.body -w '%%{http_code}' > /out/%s.code", + ssh.ShellQuote("Host: "+caddyIP), drainCaddyName, tag, tag) + drainRun(t, exec, ctx, fmt.Sprintf( + "docker run -d --user root --name %s --network %s -v %s:/out curlimages/curl:latest sh -c %s", + tag, drainNet, drainAppPrefix+"-out", ssh.ShellQuote(inner))) + } + startSlow(drainSlowTmpl + "-0") + time.Sleep(1 * time.Second) + + // The green deploy: start, readiness-gate, switch, drain, retire. + drainRun(t, exec, ctx, appRun(drainGreen, "green")) + if !waitHTTP(drainGreen, 30*time.Second) { + t.Fatal("green never became ready — failing health must never switch traffic; fixture stopped before the switch") + } + if err := switchRoute(drainGreen); err != nil { + t.Fatalf("switching to green: %v", err) + } + t.Logf("switched to green; draining 5s before retiring blue") + time.Sleep(5 * time.Second) + drainRun(t, exec, ctx, "docker stop -t 5 "+drainBlue) + + // Assertions. Zero failed requests for the declared fixture. + deadline := time.Now().Add(30 * time.Second) + failcount := "?" + for time.Now().Before(deadline) { + out, _ := exec.Run(ctx, "docker run --rm -v "+drainAppPrefix+"-out:/out curlimages/curl:latest sh -c 'cat /out/failcount' 2>/dev/null || true") + if s := strings.TrimSpace(out); s != "" { + failcount = s + break + } + time.Sleep(500 * time.Millisecond) + } + if failcount != "0" { + failLog := mustOut(t, exec, ctx, "docker run --rm -v "+drainAppPrefix+"-out:/out curlimages/curl:latest sh -c 'cat /out/fail.log 2>/dev/null || true'") + t.Fatalf("blue/green fixture reported FAILED requests (count=%s):\n%s", failcount, failLog) + } + + slowStatus, slowCode := drainContainerOut(t, exec, ctx, drainSlowTmpl+"-0") + if slowStatus != "exited" || slowCode != "0" { + t.Fatalf("the long in-flight request must complete inside the drain window (status=%s exit=%s)", slowStatus, slowCode) + } + slowBody := strings.TrimSpace(mustOut(t, exec, ctx, "docker run --rm -v "+drainAppPrefix+"-out:/out curlimages/curl:latest sh -c 'cat /out/"+drainSlowTmpl+"-0.body 2>/dev/null || true'")) + if slowBody != "blue" { + t.Fatalf("the long request should have been served end-to-end by blue, got %q", slowBody) + } + if got := probeOnce(); got != "green" { + t.Fatalf("post-switch traffic must serve green, got %q", got) + } + t.Logf("drain fixture PASS: 90/90 requests ok, long request completed on blue within the window, traffic on green") +} + +// TestDrainIntegration_NoDrainKillsLongRequest is the negative control: +// without the window (stop -t 0 immediately after the switch), an +// in-flight request is killed — proving the fixture detects a broken +// drain promise and that the window above is what saves the request. +func TestDrainIntegration_NoDrainKillsLongRequest(t *testing.T) { + host, user, key, _ := reconcileFixtureEnv(t) + ctx, cancel := context.WithTimeout(context.Background(), 3*time.Minute) + defer cancel() + + exec, err := ssh.Connect(ctx, ssh.ConnectConfig{Host: host, User: user, KeyPath: key}) + if err != nil { + t.Fatalf("connecting: %v", err) + } + // Close AFTER the cleanup work: registered as the FIRST cleanup so + // LIFO runs it LAST — a plain defer would close the session before + // the resource cleanups' commands could run on it. + t.Cleanup(func() { exec.Close() }) + if out, _ := exec.Run(ctx, "test -e /deployments/caddy/Caddyfile && echo yes || echo no"); strings.TrimSpace(out) == "yes" { + t.Skip("fixture has a provisioned /deployments/caddy — remove it or use a disposable host for the drain fixture") + } + drainRun(t, exec, ctx, "mkdir -p /deployments/caddy /deployments/"+drainAppPrefix) + if err := exec.Upload(ctx, strings.NewReader(drainServerPy), "/deployments/"+drainAppPrefix+"/server.py", "0644"); err != nil { + t.Fatalf("uploading server.py: %v", err) + } + if err := exec.Upload(ctx, strings.NewReader("{\n\tadmin 127.0.0.1:2019\n}\n"), "/deployments/caddy/Caddyfile", "0644"); err != nil { + t.Fatalf("uploading initial Caddyfile: %v", err) + } + t.Cleanup(func() { + cctx, ccancel := context.WithTimeout(context.Background(), 30*time.Second) + defer ccancel() + for _, n := range []string{drainBlue, drainGreen, drainCaddyName} { + exec.Run(cctx, "docker rm -f "+n) + } + exec.Run(cctx, "docker rm -f "+drainSlowTmpl+"-neg") + exec.Run(cctx, "rm -rf /deployments/caddy /deployments/"+drainAppPrefix) + }) + appRun := func(name, serverName string) string { + return fmt.Sprintf( + "docker run -d --name %s --network %s --label teploy.app=%s -e SERVER_NAME=%s -v /deployments/%s/server.py:/srv/server.py:ro python:3-alpine python /srv/server.py", + name, drainNet, drainAppPrefix, serverName, drainAppPrefix) + } + drainRun(t, exec, ctx, appRun(drainBlue, "blue")) + drainRun(t, exec, ctx, fmt.Sprintf( + "docker run -d --name %s --network %s -v /deployments/caddy:/etc/caddy caddy:2-alpine", drainCaddyName, drainNet)) + deadline := time.Now().Add(30 * time.Second) + ready := false + for time.Now().Before(deadline) { + out, _ := exec.Run(ctx, fmt.Sprintf( + "docker exec %s python -c \"import urllib.request;urllib.request.urlopen('http://127.0.0.1:8080/id',timeout=2)\" 2>/dev/null && echo ok || true", drainBlue)) + if strings.TrimSpace(out) == "ok" { + ready = true + break + } + time.Sleep(300 * time.Millisecond) + } + if !ready { + t.Fatal("blue never became ready — fixture broken") + } + caddyIP := strings.TrimSpace(mustOut(t, exec, ctx, + "docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "+drainCaddyName)) + client := caddy.NewClient(exec) + if err := client.SetRoute(ctx, drainAppPrefix, caddyIP, drainBlue, 8080, caddy.TLS{}, "", nil, caddy.Firewall{}, caddy.Access{}); err != nil { + t.Fatalf("routing to blue: %v", err) + } + + // Long request in flight, then an IMMEDIATE stop with no grace. The + // request's VERDICT is the pair (http code, body): served-by-blue + // means it survived; anything else (connection reset, caddy's 502) + // means the no-grace kill broke it — which is the point. + drainRun(t, exec, ctx, "docker volume create "+drainAppPrefix+"-out 2>/dev/null || true") + t.Cleanup(func() { + cctx, ccancel := context.WithTimeout(context.Background(), 15*time.Second) + defer ccancel() + exec.Run(cctx, "docker volume rm "+drainAppPrefix+"-out") + }) + slowInner := fmt.Sprintf("curl -s --max-time 10 -H %s http://%s:80/slow?ms=3000 -o /out/neg.body -w '%%{http_code}' > /out/neg.code; echo $? > /out/neg.exit", + ssh.ShellQuote("Host: "+caddyIP), drainCaddyName) + drainRun(t, exec, ctx, fmt.Sprintf( + "docker run -d --user root --name %s --network %s -v %s:/out curlimages/curl:latest sh -c %s", + drainSlowTmpl+"-neg", drainNet, drainAppPrefix+"-out", ssh.ShellQuote(slowInner))) + time.Sleep(700 * time.Millisecond) + drainRun(t, exec, ctx, "docker stop -t 0 "+drainBlue) + + status, _ := drainContainerOut(t, exec, ctx, drainSlowTmpl+"-neg") + waitDeadline := time.Now().Add(15 * time.Second) + for time.Now().Before(waitDeadline) && status != "exited" { + time.Sleep(300 * time.Millisecond) + status, _ = drainContainerOut(t, exec, ctx, drainSlowTmpl+"-neg") + } + readOut := func(name string) string { + return strings.TrimSpace(mustOut(t, exec, ctx, "docker run --rm -v "+drainAppPrefix+"-out:/out curlimages/curl:latest sh -c 'cat /out/"+name+" 2>/dev/null || true'")) + } + httpCode, body, exit := readOut("neg.code"), readOut("neg.body"), readOut("neg.exit") + if httpCode == "200" && body == "blue" && exit == "0" { + t.Fatalf("negative control broken: the request survived a no-grace kill (code=%s body=%s exit=%s) — the fixture cannot demonstrate the drain promise", httpCode, body, exit) + } + t.Logf("negative control PASS: without drain/grace the in-flight request broke (code=%s body=%q exit=%s)", httpCode, body, exit) +} + +func mustOut(t *testing.T, exec ssh.Executor, ctx context.Context, cmd string) string { + t.Helper() + out, err := exec.Run(ctx, cmd) + if err != nil { + t.Fatalf("%s: %v", cmd, err) + } + return out +} diff --git a/internal/deploy/drain_test.go b/internal/deploy/drain_test.go new file mode 100644 index 0000000..57600e1 --- /dev/null +++ b/internal/deploy/drain_test.go @@ -0,0 +1,218 @@ +package deploy + +import ( + "bytes" + "context" + "crypto/sha256" + "encoding/json" + "fmt" + "strings" + "testing" + "time" + + "github.com/useteploy/teploy/internal/ssh" +) + +// drainDeployMocks is the successful blue/green deploy fixture WITH a +// running predecessor (state names old123; the inventory lists its web +// container), so the full switch→drain→retire sequence runs. +func drainDeployMocks(existingState string) []ssh.MockCommand { + return []ssh.MockCommand{ + ssh.MockCommand{Match: "mkdir -p /deployments/myapp", Output: ""}, + ssh.MockCommand{Match: "mkdir /deployments/myapp/.lock", Output: ""}, + ssh.MockCommand{Match: "if [ ! -e '/deployments/myapp/state.json' ]", Output: "absent"}, + ssh.MockCommand{Match: "if [ ! -e '/deployments/myapp/state' ]", Output: "present\n" + existingState}, + ssh.MockCommand{Match: "ss -tln", Output: ssOutput}, + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='myapp'", Output: predecessorInventoryJSON}, + ssh.MockCommand{Match: "docker run", Output: "newcontainer123"}, + ssh.MockCommand{Match: "docker inspect", Output: "running"}, + ssh.MockCommand{Match: "curl -s -o /dev/null", Output: "200"}, + ssh.MockCommand{Match: "curl -sf http://localhost:2019/config/apps/http/servers/srv0", Output: `{"listen":[":80",":443"]}`}, + ssh.MockCommand{Match: "curl -sf -X PATCH", Err: errBoom}, + ssh.MockCommand{Match: "curl -sf -X POST http://localhost:2019/config/apps/http/servers/srv0/routes", Output: ""}, + ssh.MockCommand{Match: "rm -f /tmp/teploy_caddy", Output: ""}, + ssh.MockCommand{Match: "cat /deployments/caddy/Caddyfile", Output: "{\n\tadmin 0.0.0.0:2019\n}\n"}, + ssh.MockCommand{Match: "mv /tmp/teploy_caddyfile.tmp", Output: ""}, + ssh.MockCommand{Match: "mkdir /deployments/caddy/.lock", Output: ""}, + ssh.MockCommand{Match: "a=$(docker exec caddy md5sum", Output: "TEPLOY_CADDY_OK"}, + ssh.MockCommand{Match: "docker exec caddy caddy reload", Output: ""}, + ssh.MockCommand{Match: "docker stop", Output: ""}, + ssh.MockCommand{Match: "printf %s", Output: ""}, + ssh.MockCommand{Match: "rm -rf /deployments/myapp/.lock", Output: ""}, + } +} + +// TestDeploy_DrainWindowBetweenSwitchAndRetirement is the C03 wiring +// proof: with drain_seconds set, the deploy surfaces the policy, the +// route switch (reload) lands BEFORE the predecessor stop, and the drain +// window actually elapses between them. +func TestDeploy_DrainWindowBetweenSwitchAndRetirement(t *testing.T) { + existingState := "current_port=49152\ncurrent_hash=old123\nprevious_port=0\nprevious_hash=\n" + mock := ssh.NewMockExecutor("1.2.3.4", drainDeployMocks(existingState)...) + var buf bytes.Buffer + deployer := NewDeployer(mock, &buf) + + start := time.Now() + err := deployer.Deploy(context.Background(), Config{ + App: "myapp", + Domain: "myapp.com", + Image: "myapp:v2", + Version: "new456", + DrainSeconds: 1, + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }) + elapsed := time.Since(start) + if err != nil { + t.Fatalf("draining deploy: %v", err) + } + if elapsed < 1*time.Second { + t.Errorf("the drain window must elapse before retirement (took %s)", elapsed) + } + out := buf.String() + if !strings.Contains(out, "Draining myapp's predecessor for 1s") { + t.Errorf("expected the drain to be surfaced, got:\n%s", out) + } + if !strings.Contains(out, "Stop policy: graceful stop after 10s (SIGTERM then SIGKILL); drain 1s window before predecessor retirement") { + t.Errorf("expected the surfaced stop/drain policy distinction, got:\n%s", out) + } + reloadIdx, stopIdx := -1, -1 + for i, c := range mock.Calls { + if strings.HasPrefix(c, "docker exec caddy caddy reload") && reloadIdx < 0 { + reloadIdx = i + } + if strings.HasPrefix(c, "docker stop") && stopIdx < 0 { + stopIdx = i + } + } + if reloadIdx < 0 || stopIdx < 0 || reloadIdx > stopIdx { + t.Errorf("route switch (idx %d) must precede predecessor stop (idx %d)", reloadIdx, stopIdx) + } +} + +// TestDeploy_DrainDisabledByDefault pins the compatibility default: +// drain_seconds 0 stops the predecessor immediately after the switch — +// no window, no drain line (the historical behavior). +func TestDeploy_DrainDisabledByDefault(t *testing.T) { + existingState := "current_port=49152\ncurrent_hash=old123\nprevious_port=0\nprevious_hash=\n" + mock := ssh.NewMockExecutor("1.2.3.4", drainDeployMocks(existingState)...) + var buf bytes.Buffer + deployer := NewDeployer(mock, &buf) + + start := time.Now() + err := deployer.Deploy(context.Background(), Config{ + App: "myapp", + Domain: "myapp.com", + Image: "myapp:v2", + Version: "new456", + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }) + elapsed := time.Since(start) + if err != nil { + t.Fatalf("undrained deploy: %v", err) + } + if elapsed >= 1*time.Second { + t.Errorf("no drain window may be added by default (took %s)", elapsed) + } + if strings.Contains(buf.String(), "Draining") { + t.Errorf("drain must not surface when disabled:\n%s", buf.String()) + } + if !strings.Contains(buf.String(), "drain disabled — the predecessor stops immediately after the switch") { + t.Errorf("the surfaced stop policy must name the disabled drain:\n%s", buf.String()) + } +} + +// TestDeploy_DrainSkippedWithoutServingPredecessor: external ingress has +// no teploy-managed edge and no serving predecessor behind THIS switch; +// the window must not run even when configured. +func TestDeploy_DrainSkippedWithoutServingPredecessor(t *testing.T) { + existingState := "current_port=49152\ncurrent_hash=old123\nprevious_port=0\nprevious_hash=\n" + mocks := drainDeployMocks(existingState) + // No route step under external ingress: the admin-API curl checks and + // reload are never reached; keep the rest. + mock := ssh.NewMockExecutor("1.2.3.4", mocks...) + var buf bytes.Buffer + deployer := NewDeployer(mock, &buf) + + start := time.Now() + err := deployer.Deploy(context.Background(), Config{ + App: "myapp", + Domain: "myapp.com", + Image: "myapp:v2", + Version: "new456", + Ingress: "external", + DrainSeconds: 1, + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }) + elapsed := time.Since(start) + if err != nil { + t.Fatalf("external-ingress deploy: %v", err) + } + if elapsed >= 1*time.Second { + t.Errorf("no drain window under external ingress (took %s)", elapsed) + } + if strings.Contains(buf.String(), "Draining") { + t.Errorf("drain must not run without a teploy-managed switch:\n%s", buf.String()) + } +} + +// TestRollback_DrainWindow is the rollback-side C03 wiring: the route +// switched back to the target, then the configured window elapses before +// the superseded generation is stopped, surfaced to the operator. +func TestRollback_DrainWindow(t *testing.T) { + currentManifest := json.RawMessage(`{"release":"v2"}`) + previousManifest := json.RawMessage(`{"release":"v1"}`) + stateContent := fmt.Sprintf(`{"schema_version":2,"deployment_type":"container","ingress_mode":"caddy","domain":"myapp.com","updated_at":"2026-07-22T10:00:00Z","manifest_sha256":"%x","source_revision":"rev-v2","image_ref":"myapp:v2","operation_id":"deploy-v2","generation":7,"applied_manifest":%s,"previous_release":{"hash":"v1","manifest_sha256":"%x","source_revision":"rev-v1","image_ref":"myapp:v1","applied_manifest":%s},"current_port":49153,"current_hash":"v2","previous_port":49152,"previous_hash":"v1"}`, + sha256.Sum256(currentManifest), currentManifest, sha256.Sum256(previousManifest), previousManifest) + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "if [ ! -e '/deployments/myapp/state.json' ]", Output: "present\n" + stateContent}, + ssh.MockCommand{Match: "mkdir -p /deployments/myapp", Output: ""}, + ssh.MockCommand{Match: "mkdir /deployments/myapp/.lock", Output: ""}, + ssh.MockCommand{Match: "cat /deployments/myapp/.lock/info", Err: fmt.Errorf("none")}, + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='myapp'", + Output: `{"ID":"aaa","Names":"myapp-web-v1","Image":"myapp:latest","State":"exited","Status":"Exited","Labels":"teploy.app=myapp,teploy.version=v1,teploy.process=web"}` + "\n" + + `{"ID":"bbb","Names":"myapp-web-v2","Image":"myapp:latest","State":"running","Status":"Up 1h","Labels":"teploy.app=myapp,teploy.version=v2,teploy.process=web"}`, + }, + ssh.MockCommand{Match: "docker inspect 'myapp-web-v1'", Output: `[{"Config":{"Image":"myapp:latest","Labels":{"teploy.app":"myapp"}},"HostConfig":{"NetworkMode":"teploy","PortBindings":{"3000/tcp":[{"HostIp":"127.0.0.1","HostPort":"49152"}]},"RestartPolicy":{"Name":"no"}},"NetworkSettings":{"Networks":{"teploy":{"Aliases":["myapp"]}}}}]`}, + ssh.MockCommand{Match: "docker rm -f 'myapp-web-v1'", Output: ""}, + ssh.MockCommand{Match: "docker run", Output: ""}, + ssh.MockCommand{Match: "curl", Output: "200"}, + ssh.MockCommand{Match: "docker inspect -f '{{range $p, $b := .NetworkSettings.Ports}}{{range $b}}{{.HostIp}}", Output: "127.0.0.1 "}, + ssh.MockCommand{Match: "docker inspect -f '{{range $p, $b := .NetworkSettings.Ports}}", Output: "49153"}, + ssh.MockCommand{Match: "docker inspect -f '{{range $p, $_ := .NetworkSettings.Ports}}", Output: "3000/tcp"}, + ssh.MockCommand{Match: "caddy", Output: ""}, + ssh.MockCommand{Match: "cat /deployments/caddy/Caddyfile", Output: "{\n\tadmin 0.0.0.0:2019\n}\n"}, + ssh.MockCommand{Match: "mkdir /deployments/caddy/.lock", Output: ""}, + ssh.MockCommand{Match: "a=$(docker exec caddy md5sum", Output: "TEPLOY_CADDY_OK"}, + ssh.MockCommand{Match: "docker exec caddy caddy reload", Output: ""}, + ssh.MockCommand{Match: "docker stop", Output: ""}, + ssh.MockCommand{Match: "mkdir -p", Output: ""}, + ssh.MockCommand{Match: "cat /tmp", Output: ""}, + ssh.MockCommand{Match: "UPLOAD:", Output: ""}, + ) + + var buf bytes.Buffer + cfg := rollbackCfg() + cfg.DrainSeconds = 1 + start := time.Now() + if err := Rollback(context.Background(), mock, &buf, cfg); err != nil { + t.Fatalf("rollback with drain: %v", err) + } + if elapsed := time.Since(start); elapsed < 1*time.Second { + t.Errorf("the drain window must elapse before the superseded generation stops (took %s)", elapsed) + } + if !strings.Contains(buf.String(), "Draining myapp's superseded generation for 1s") { + t.Errorf("expected the drain to be surfaced, got:\n%s", buf.String()) + } + reloadIdx, stopIdx := -1, -1 + for i, c := range mock.Calls { + if strings.HasPrefix(c, "docker exec caddy caddy reload") && reloadIdx < 0 { + reloadIdx = i + } + if strings.HasPrefix(c, "docker stop") && stopIdx < 0 { + stopIdx = i + } + } + if reloadIdx < 0 || stopIdx < 0 || reloadIdx > stopIdx { + t.Errorf("route switch back (idx %d) must precede the superseded stop (idx %d)", reloadIdx, stopIdx) + } +} diff --git a/internal/deploy/fence_test.go b/internal/deploy/fence_test.go index 265dd5c..27bd9dc 100644 --- a/internal/deploy/fence_test.go +++ b/internal/deploy/fence_test.go @@ -99,8 +99,14 @@ func TestDeployFenced_LateHolderRefusedToStartContainers(t *testing.T) { if !errors.Is(err, state.ErrFenceLost) { t.Fatalf("expected ErrFenceLost, got %v", err) } + // C01-2: the guard is composed into the same shell as the docker run, + // so a REFUSED start appears only as `grep ...; docker run ...` (the + // composed command the guard rejected). An EXECUTED start appears as a + // bare `docker run` command — the mock appends the effect-only form + // when the guard holds. Only that form is a container actually + // starting. for _, c := range mock.Calls { - if strings.Contains(c, "docker run") { + if strings.HasPrefix(c, "docker run") { t.Errorf("late holder's container start must be refused, saw: %s", c) } } diff --git a/internal/deploy/guarded_effects_test.go b/internal/deploy/guarded_effects_test.go new file mode 100644 index 0000000..7985cb8 --- /dev/null +++ b/internal/deploy/guarded_effects_test.go @@ -0,0 +1,206 @@ +package deploy + +import ( + "context" + "errors" + "strings" + "testing" + "time" + + "github.com/useteploy/teploy/internal/ssh" + "github.com/useteploy/teploy/internal/state" +) + +// TestDeployFenced_EffectsRunGuarded proves the C01-2 composition on the +// happy path: the candidate/worker docker runs and the Caddyfile commit +// execute as guard+effect in ONE remote command (never a separate +// check-then-act pair), so no transport window exists for a takeover to +// slip a broken holder's effect into. +func TestDeployFenced_EffectsRunGuarded(t *testing.T) { + app := "fency" + workerState := `{"Status":"running","Running":true}` + mocks := append([]ssh.MockCommand{ + // Worker viability verification (A23) — must answer before the + // generic "docker inspect" happy-path entry. + ssh.MockCommand{Match: "docker inspect -f '{{json .State}}'", Output: workerState}, + }, fenceHappyPathMocks(app)...) + mock := ssh.NewMockExecutor("1.2.3.4", mocks...) + lk, err := state.AcquireLockFenced(context.Background(), mock, app) + if err != nil { + t.Fatalf("AcquireLockFenced: %v", err) + } + guard := lk.GuardPrefix() + if guard == "" { + t.Fatal("a held fence must produce a guard prefix") + } + var out strings.Builder + d := NewDeployer(mock, &out) + cfg := Config{ + App: app, + Domain: "fency.com", + Image: "fency:latest", + Version: "abc123", + Processes: map[string]string{"web": "", "worker": "node worker.js"}, + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + } + if err := d.DeployFenced(context.Background(), cfg, lk); err != nil { + t.Fatalf("DeployFenced: %v", err) + } + + composedRun := 0 + for _, c := range mock.Calls { + if strings.HasPrefix(c, guard) && strings.Contains(c, "docker run") { + composedRun++ + } + } + // web candidate + worker: two guarded container creations. + if composedRun < 2 { + t.Errorf("expected the web and worker starts to run composed under the fence guard, found %d", composedRun) + } + var composedCommit bool + for _, c := range mock.Calls { + if strings.HasPrefix(c, guard) && strings.Contains(c, "mv -fT -- ") && strings.Contains(c, "/deployments/caddy/Caddyfile") { + composedCommit = true + } + } + if !composedCommit { + t.Error("expected the Caddyfile commit rename to run composed under the fence guard") + } +} + +// lockFlipper wraps the mock executor, rewriting the lock info to name +// another operation the moment a trigger command runs — modeling a +// takeover at an exact point mid-deploy. +type lockFlipper struct { + *ssh.MockExecutor + app string + trigger string + flipped bool + execOnly []string +} + +func (f *lockFlipper) Run(ctx context.Context, cmd string) (string, error) { + if !f.flipped && strings.Contains(cmd, f.trigger) { + f.flipped = true + f.MockExecutor.Files["/deployments/"+f.app+"/.lock/info"] = []byte(`{"type":"auto","owner":"someoneelse"}`) + } + out, err := f.MockExecutor.Run(ctx, cmd) + // Track EXECUTED commands: the mock appends the effect-only form for + // guard commands whose guard held, so filter those back out by + // tracking what the underlying executor actually recorded. + return out, err +} + +// TestDeployFenced_RouteSwitchRefusedOnMidFlightTakeover: the lock breaks +// during the deploy (at the health gate); the route switch's Caddyfile +// COMMIT must be refused by the composed guard — the Caddyfile on disk +// never changes, no reload runs, and the deploy reports the fence loss. +func TestDeployFenced_RouteSwitchRefusedOnMidFlightTakeover(t *testing.T) { + app := "fency" + base := ssh.NewMockExecutor("1.2.3.4", fenceHappyPathMocks(app)...) + lk, err := state.AcquireLockFenced(context.Background(), base, app) + if err != nil { + t.Fatalf("AcquireLockFenced: %v", err) + } + flip := &lockFlipper{MockExecutor: base, app: app, trigger: "curl -s -o /dev/null"} + var out strings.Builder + d := NewDeployer(flip, &out) + err = d.DeployFenced(context.Background(), Config{ + App: app, + Domain: "fency.com", + Image: "fency:latest", + Version: "abc123", + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }, lk) + if !errors.Is(err, state.ErrFenceLost) { + t.Fatalf("expected ErrFenceLost from the refused route switch, got: %v", err) + } + if !flip.flipped { + t.Fatal("test bug: the takeover never fired") + } + // The Caddyfile must be untouched: the happy-path fixture has no + // caddyfile in Files, and a landed commit would have created one via + // the staged rename. + if _, ok := base.Files["/deployments/caddy/Caddyfile"]; ok { + t.Error("a fence-refused route switch must not commit the Caddyfile") + } + for _, c := range base.Calls { + if strings.HasPrefix(c, "docker exec caddy caddy reload") { + t.Errorf("no reload may run after a refused commit, saw: %s", c) + } + } +} + +// TestDeployFenced_WorkerStartRefusedOnTakeover: the lock breaks after the +// web candidates are healthy; the worker start (the finding's second +// site) must be refused by the composed guard — no bare docker run for +// the worker executes. +func TestDeployFenced_WorkerStartRefusedOnTakeover(t *testing.T) { + app := "fency" + base := ssh.NewMockExecutor("1.2.3.4", fenceHappyPathMocks(app)...) + lk, err := state.AcquireLockFenced(context.Background(), base, app) + if err != nil { + t.Fatalf("AcquireLockFenced: %v", err) + } + flip := &lockFlipper{MockExecutor: base, app: app, trigger: "curl -s -o /dev/null"} + var out strings.Builder + d := NewDeployer(flip, &out) + err = d.DeployFenced(context.Background(), Config{ + App: app, + Domain: "fency.com", + Image: "fency:latest", + Version: "abc123", + Processes: map[string]string{"web": "", "worker": "node worker.js"}, + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }, lk) + if !errors.Is(err, state.ErrFenceLost) { + t.Fatalf("expected ErrFenceLost from the refused worker start, got: %v", err) + } + for _, c := range base.Calls { + if strings.HasPrefix(c, "docker run") && strings.Contains(c, "fency-worker-abc123") { + t.Errorf("the worker start must be refused in-shell, saw execution: %s", c) + } + } +} + +// TestDeployFenced_CommitGuardOrder documents and pins the lock +// ACQUISITION ORDER on the traffic-switch commit (C01-3): the app-level +// fence guard precedes the short-lived shared-proxy (caddy) lock guard in +// the composed command — app lock held for the whole lifecycle, caddy +// commit lock held only for the brief edit+reload, never the inverse, +// never across hosts. +func TestDeployFenced_CommitGuardOrder(t *testing.T) { + app := "fency" + mock := ssh.NewMockExecutor("1.2.3.4", fenceHappyPathMocks(app)...) + lk, err := state.AcquireLockFenced(context.Background(), mock, app) + if err != nil { + t.Fatalf("AcquireLockFenced: %v", err) + } + var out strings.Builder + d := NewDeployer(mock, &out) + if err := d.DeployFenced(context.Background(), Config{ + App: app, + Domain: "fency.com", + Image: "fency:latest", + Version: "abc123", + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }, lk); err != nil { + t.Fatalf("DeployFenced: %v", err) + } + guard := lk.GuardPrefix() + for _, c := range mock.Calls { + if !strings.Contains(c, "mv -fT -- ") || !strings.Contains(c, "/deployments/caddy/Caddyfile") { + continue + } + appIdx := strings.Index(c, guard) + caddyIdx := strings.Index(c, "/deployments/caddy/.lock/info") + if appIdx < 0 || caddyIdx < 0 { + t.Fatalf("commit command missing a guard:\n%s", c) + } + if appIdx > caddyIdx { + t.Fatalf("caddy lock guard must FOLLOW the app fence guard (documented acquisition order):\n%s", c) + } + return + } + t.Fatal("no composed Caddyfile commit found") +} diff --git a/internal/deploy/health.go b/internal/deploy/health.go index b661b87..5ab3680 100644 --- a/internal/deploy/health.go +++ b/internal/deploy/health.go @@ -248,6 +248,17 @@ func readinessSummary(cfg HealthConfig, port int) string { } } +// drainSummary renders the request-drain half of the surfaced stop policy +// (C03): 0 keeps the historical stop-immediately behavior; N names the +// window the predecessor keeps serving in-flight requests after the +// traffic switch. +func drainSummary(drainSeconds int) string { + if drainSeconds <= 0 { + return "disabled — the predecessor stops immediately after the switch" + } + return fmt.Sprintf("%ds window before predecessor retirement", drainSeconds) +} + // checkTCP verifies that a TCP connection can be established to the port. // The /dev/tcp redirection runs inside a single-quoted bash -c argument, so // neither the host nor the port can break out of it. diff --git a/internal/deploy/reconcile.go b/internal/deploy/reconcile.go new file mode 100644 index 0000000..d69564b --- /dev/null +++ b/internal/deploy/reconcile.go @@ -0,0 +1,261 @@ +// The replacement owner's reconciliation (programme workstream C01, slice +// C01-1): when a deploy's lock acquisition BREAKS a stale predecessor, the +// dead holder may have left in-flight effects on the target — a late +// `docker run` landing under version-keyed names, a route switch half-done, +// a displaced fixed-port workload. Lock acquisition is never proof of +// quiescence (docs/C01_RECOVERY_STATE_TABLE.md, finding C01-1); what the +// replacement owner OBSERVES, decided through the tested table, is the only +// safe input to "may I proceed with my own deploy?". +// +// Observe is the productionized evidence collector: the same reads the +// fixture harness performs (docker label inventory, state.json, the managed +// Caddyfile, the per-release record), with exact names and no guesses. +// ReconcileAfterTakeover drives the decision: RETRY is the only disposition +// that lets the deploy proceed; every other one refuses with the observed +// evidence so an operator reconciles deliberately. This is deliberately NOT +// an auto-compensator — COMPENSATE's automation (restart/restore via the +// recorded receipts) is the recovery-owner continuation work keyed to F04 +// generation identities; refusing with evidence is the safe subset the +// table permits today. +package deploy + +import ( + "context" + "fmt" + "strings" + "time" + + "github.com/useteploy/teploy/internal/deploy/recovery" + "github.com/useteploy/teploy/internal/docker" + "github.com/useteploy/teploy/internal/releasemeta" + "github.com/useteploy/teploy/internal/ssh" + "github.com/useteploy/teploy/internal/state" +) + +// managedCaddyfile is the shared proxy's config path — the route evidence +// class reads the managed marker blocks from it (internal/caddy owns the +// write side; recovery only reads). +const managedCaddyfile = "/deployments/caddy/Caddyfile" + +// reconcileReobserveDelay is the bounded pause before re-observing an +// INSPECT disposition (R4's transient-read reconcile trigger). A variable +// so tests can shorten it. +var reconcileReobserveDelay = 2 * time.Second + +// Observe collects a recovery.Observation about (app, attempted) versus its +// known predecessor release with exact names and receipts: the teploy-labeled +// container inventory, the authoritative state.json, the managed Caddyfile, +// and the per-release record. Read failures map to Unknown classes — the +// decision table's never-auto-decide evidence grade — never to guesses. +// +// attempted is the release the caller is about to deploy; predecessor is the +// release state.json is expected to name ("" when there is no state yet). +// Containers under any OTHER teploy.app naming are ForeignCandidates — the +// dead holder's late effects land exactly there. +func Observe(ctx context.Context, exec ssh.Executor, app, attempted, predecessor string) recovery.Observation { + var o recovery.Observation + + containers, err := docker.NewClient(exec).ListContainers(ctx, app) + if err != nil { + o.Candidates, o.CandidateCorpses, o.ForeignCandidates = recovery.Unknown, recovery.Unknown, recovery.Unknown + } else { + candPrefix := fmt.Sprintf("%s-web-%s", app, attempted) + predName := "" + if predecessor != "" { + predName = fmt.Sprintf("%s-web-%s", app, predecessor) + } + for _, c := range containers { + isCandidate := strings.HasPrefix(c.Name, candPrefix) + isPredecessor := predName != "" && + (strings.HasPrefix(c.Name, predName) || strings.HasPrefix(c.Name, predName+"_replaced")) + switch { + case isCandidate: + if c.State == "running" { + o.Candidates = recovery.Present + } else { + o.CandidateCorpses = recovery.Present + } + case isPredecessor: + if strings.HasSuffix(c.Name, "_replaced") || c.State == "running" { + o.PredecessorServing = recovery.Present + } else { + o.PredecessorStopped = recovery.Present + } + default: + // Neither the attempted release's candidate naming nor the + // known predecessor's: an unattributable workload — only + // RUNNING ones consume traffic/jobs. + if c.State == "running" { + o.ForeignCandidates = recovery.Present + } + } + } + } + + st, err := state.Read(ctx, exec, app) + switch { + case err != nil: + o.StateToCandidate, o.StateToPredecessor = recovery.Unknown, recovery.Unknown + case st == nil: + o.StateToCandidate, o.StateToPredecessor = recovery.Absent, recovery.Absent + case st.CurrentHash == attempted: + o.StateToCandidate = recovery.Present + default: + o.StateToPredecessor = recovery.Present + } + + caddyfile, present, err := state.ReadRemoteFile(ctx, exec, managedCaddyfile) + switch { + case err != nil: + o.RouteToCandidate, o.RouteToPredecessor = recovery.Unknown, recovery.Unknown + case !present: + o.RouteToCandidate, o.RouteToPredecessor = recovery.Absent, recovery.Absent + default: + // C01-3's missing producer: classify the managed block EXACTLY. + // A block naming the attempted candidates (or the predecessor's + // containers) is provable evidence; a managed block that names + // NEITHER — a third generation no inventory can attribute, the + // dead-holder-late-route-edit shape — is CONFLICTING evidence + // (Unknown), which Decide routes to INSPECT (R4/R5), never a + // blind RETRY. No managed block for the app at all is clean + // absence. The old default (anything else = "route to + // predecessor") misread a foreign-generation route as the known + // predecessor's. + content := string(caddyfile) + candUpstream := fmt.Sprintf("%s-web-%s", app, attempted) + toCandidate := strings.Contains(content, candUpstream) + var toPredecessor bool + if predecessor != "" { + toPredecessor = strings.Contains(content, fmt.Sprintf("%s-web-%s", app, predecessor)) + } + switch { + case toCandidate: + o.RouteToCandidate = recovery.Present + if toPredecessor { + o.RouteToPredecessor = recovery.Present + } + case toPredecessor: + o.RouteToPredecessor = recovery.Present + case strings.Contains(content, fmt.Sprintf("# TEPLOY BEGIN %s\n", app)): + o.RouteToCandidate, o.RouteToPredecessor = recovery.Unknown, recovery.Unknown + default: + o.RouteToCandidate, o.RouteToPredecessor = recovery.Absent, recovery.Absent + } + } + + rec, err := releasemeta.Read(ctx, exec, app, attempted) + switch { + case err != nil: + o.ReleaseRecord = recovery.Unknown + case rec != nil: + o.ReleaseRecord = recovery.Present + default: + o.ReleaseRecord = recovery.Absent + } + return o +} + +// describeEvidence renders the observation as the operator-facing evidence +// block: every non-absent class on its own line, so a refusal names exactly +// what was seen on the target. +func describeEvidence(o recovery.Observation) string { + type line struct { + name string + ev recovery.Evidence + } + lines := []line{ + {"running workload under this deploy's candidate names", o.Candidates}, + {"stopped corpses under this deploy's candidate names", o.CandidateCorpses}, + {"running unattributable teploy-labeled workload (foreign names)", o.ForeignCandidates}, + {"route (Caddyfile) naming this deploy's candidates", o.RouteToCandidate}, + {"route (Caddyfile) naming the predecessor generation", o.RouteToPredecessor}, + {"state.json naming this deploy's release", o.StateToCandidate}, + {"state.json naming the predecessor release", o.StateToPredecessor}, + {"predecessor workload serving", o.PredecessorServing}, + {"predecessor workload stopped (restorable)", o.PredecessorStopped}, + {"release record for this deploy's version", o.ReleaseRecord}, + } + var out []string + for _, l := range lines { + if l.ev != recovery.Absent { + out = append(out, fmt.Sprintf(" %s: %s", l.name, l.ev)) + } + } + if len(out) == 0 { + out = append(out, " (no leftover effects observed)") + } + return strings.Join(out, "\n") +} + +// ReconcileAfterTakeover runs the replacement owner's decision over the +// observed target (C01-1). It is called by DeployFenced exactly when this +// deploy's lock acquisition broke a stale predecessor's lock, BEFORE any of +// its own effects. The disposition comes from the tested table +// (recovery.Decide): +// +// - RETRY: nothing contradicts proceeding (no state, or a consistent +// serving predecessor and no leftover candidate work). The deploy +// proceeds, with the observed world summarized for the operator. +// - INSPECT: evidence is readable but does not match a clean pre-deploy +// world (a running candidate without receipts, record/target +// disagreement, both generations on the edge) or is unreadable. One +// re-observation absorbs transient read failures; a persistent INSPECT +// refuses the deploy with the evidence. +// - COMPENSATE: traffic sits on an uncommitted generation, or the +// predecessor is displaced — restorable, but the undo is the operator's +// call (the auto-compensating recovery owner is F04-keyed future work). +// - MANUAL: unattributable workloads or unreadable authority — never +// auto-decided. +// +// A refusal is a plain error: no effects of this deploy have run yet (the +// reconcile precedes the first docker run), so there is nothing to undo. +func (d *Deployer) ReconcileAfterTakeover(ctx context.Context, cfg Config, current *state.AppState) error { + predecessor := "" + if current != nil { + predecessor = current.CurrentHash + } + observe := func() recovery.Observation { + return Observe(ctx, d.exec, cfg.App, cfg.Version, predecessor) + } + o := observe() + disp := recovery.Decide(recovery.Admitted, o) + if disp == recovery.Inspect { + // R4's unreadable classes are reconcile triggers: a transient + // inventory/parse failure deserves one re-observation (bounded), + // not an immediate operator escalation. + select { + case <-ctx.Done(): + case <-time.After(reconcileReobserveDelay): + } + o = observe() + disp = recovery.Decide(recovery.Admitted, o) + } + evidence := describeEvidence(o) + + switch disp { + case recovery.Retry: + fmt.Fprintf(d.out, "Stale lock taken over; reconciled the observed target (no leftover deploy effects) — proceeding\n") + return nil + case recovery.Compensate: + return fmt.Errorf( + "refusing to deploy %s after taking over its stale lock: the observed target needs COMPENSATION before a new deploy\n"+ + "(traffic on an uncommitted generation, or a displaced predecessor that is restorable).\n"+ + "Observed evidence:\n%s\n"+ + "Reconcile the app first (teploy status / teploy rollback --app %s, or inspect the containers named by: docker ps --filter label=teploy.app=%s)", + cfg.App, evidence, cfg.App, cfg.App) + case recovery.Inspect: + return fmt.Errorf( + "refusing to deploy %s after taking over its stale lock: the observed target does not match a clean pre-deploy world (INSPECT)\n"+ + "— a previous deploy's effects landed without receipts, or evidence is unreadable.\n"+ + "Observed evidence:\n%s\n"+ + "Inspect the app first (teploy status; docker ps --filter label=teploy.app=%s) and reconcile the leftover generation before retrying", + cfg.App, evidence, cfg.App) + default: // recovery.Manual + return fmt.Errorf( + "refusing to deploy %s after taking over its stale lock: the observed target contains evidence automation must not decide (MANUAL)\n"+ + "— unattributable running workload(s) or unreadable authority.\n"+ + "Observed evidence:\n%s\n"+ + "Inspect the app manually (docker ps --filter label=teploy.app=%s; cat /deployments/%s/state.json) before retrying", + cfg.App, evidence, cfg.App, cfg.App) + } +} diff --git a/internal/deploy/reconcile_integration_test.go b/internal/deploy/reconcile_integration_test.go new file mode 100644 index 0000000..132245a --- /dev/null +++ b/internal/deploy/reconcile_integration_test.go @@ -0,0 +1,155 @@ +//go:build integration + +// Fixture-gated verification of the replacement-owner reconciliation +// (C01-1) against a REAL SSH+Docker host — the wired version of the +// recovery harness's scenario (a): a dead holder's delayed docker effect +// lands after a replacement owner breaks the stale lock, and the +// PRODUCTION reconciler (deploy.ReconcileAfterTakeover, the code DeployFenced +// runs on takeover) must refuse with the MANUAL disposition instead of +// proceeding on the quiescence assumption. +// +// Same fixture contract as internal/deploy/recovery/harness_integration_test.go: +// +// TEPLOY_FAULT_HOST=127.0.0.1:50075 \ +// TEPLOY_FAULT_USER=tyler \ +// TEPLOY_FAULT_KEY=~/.colima/_lima/_config/user \ +// go test -tags integration -run TestReconcileIntegration -v ./internal/deploy +// +// Disposable fixture only — the test creates and removes +// /deployments/-d and -d-shaped containers. +package deploy + +import ( + "context" + "fmt" + "os" + "strings" + "testing" + "time" + + "github.com/useteploy/teploy/internal/ssh" + "github.com/useteploy/teploy/internal/state" +) + +func reconcileFixtureEnv(t *testing.T) (host, user, key, image string) { + t.Helper() + host = os.Getenv("TEPLOY_FAULT_HOST") + user = os.Getenv("TEPLOY_FAULT_USER") + key = os.Getenv("TEPLOY_FAULT_KEY") + if host == "" || user == "" || key == "" { + t.Skip("reconcile fixture needs TEPLOY_FAULT_HOST, TEPLOY_FAULT_USER, TEPLOY_FAULT_KEY (see internal/deploy/reconcile_integration_test.go)") + } + image = os.Getenv("TEPLOY_FAULT_IMAGE") + if image == "" { + image = "alpine:3" + } + return +} + +// TestReconcileIntegration_LateEffectAfterTakeoverIsRefused: owner A takes +// the lock, launches a nohup'd delayed candidate, and dies; the lock ages +// stale; owner B breaks and acquires it through the real path (TookOver +// must report true); the late container lands post-acquisition; the +// production reconciler must REFUSE (MANUAL — unattributable running +// workload), proving DeployFenced's takeover path reconciles observed +// evidence instead of assuming quiescence. +func TestReconcileIntegration_LateEffectAfterTakeoverIsRefused(t *testing.T) { + host, user, key, image := reconcileFixtureEnv(t) + app := "faultprobe-d" + lateName := fmt.Sprintf("%s-web-deadgen", app) + + ctx, cancel := context.WithTimeout(context.Background(), 3*time.Minute) + defer cancel() + + probe, err := ssh.Connect(ctx, ssh.ConnectConfig{Host: host, User: user, KeyPath: key}) + if err != nil { + t.Fatalf("connecting probe session: %v", err) + } + // Close AFTER the cleanup work (LIFO): see exec's note in the drain + // fixture — a plain defer races the resource cleanups. + t.Cleanup(func() { probe.Close() }) + if out, err := probe.Run(ctx, "docker version --format '{{.Server.Version}}'"); err != nil { + t.Skipf("fixture host has no reachable docker daemon (%v)", err) + } else { + t.Logf("fixture docker server %s", strings.TrimSpace(out)) + } + t.Cleanup(func() { + cctx, ccancel := context.WithTimeout(context.Background(), 30*time.Second) + defer ccancel() + probe.Run(cctx, "docker rm -f "+ssh.ShellQuote(lateName)) + probe.Run(cctx, "rm -rf "+ssh.ShellQuote("/deployments/"+app)) + }) + + if err := state.EnsureAppDir(ctx, probe, app); err != nil { + t.Fatalf("creating app dir: %v", err) + } + + // Owner A: take the lock, launch the delayed effect, die. + ownerA, err := ssh.Connect(ctx, ssh.ConnectConfig{Host: host, User: user, KeyPath: key}) + if err != nil { + t.Fatalf("connecting owner A: %v", err) + } + lkA, err := state.AcquireLockFenced(ctx, ownerA, app) + if err != nil { + t.Fatalf("owner A acquiring the lock: %v", err) + } + inner := fmt.Sprintf("sleep 4; docker run --detach --restart no --name %s --label %s --label %s --label %s %s sleep 300", + ssh.ShellQuote(lateName), + ssh.ShellQuote("teploy.app="+app), + ssh.ShellQuote("teploy.process=web"), + ssh.ShellQuote("teploy.version=deadgen"), + ssh.ShellQuote(image)) + if _, err := ownerA.Run(ctx, "nohup sh -c "+ssh.ShellQuote(inner)+" >/dev/null 2>&1 & echo launched"); err != nil { + t.Fatalf("launching the delayed effect: %v", err) + } + ownerA.Close() + + // Age the lock past the stale window (token preserved), then owner B + // breaks and acquires through the real stale-break path. + ownerB, err := ssh.Connect(ctx, ssh.ConnectConfig{Host: host, User: user, KeyPath: key}) + if err != nil { + t.Fatalf("connecting owner B: %v", err) + } + t.Cleanup(func() { ownerB.Close() }) // after the release cleanup (LIFO) + old := time.Now().UTC().Add(-31 * time.Minute).Format(time.RFC3339) + info := fmt.Sprintf("{\"type\":\"auto\",\"owner\":%q,\"ts\":%q}\n", lkA.Owner(), old) + if err := ownerB.Upload(ctx, strings.NewReader(info), fmt.Sprintf("/deployments/%s/.lock/info", app), "0644"); err != nil { + t.Fatalf("aging the lock: %v", err) + } + lkB, err := state.AcquireLockFenced(ctx, ownerB, app) + if err != nil { + t.Fatalf("owner B acquiring the stale-broken lock: %v", err) + } + defer state.ReleaseLockFenced(ownerB, lkB, app) + if !lkB.TookOver() { + t.Fatal("the stale-break acquisition must report TookOver — the reconciliation keys on it") + } + + // Wait for the dead holder's late effect to land. + landed := false + deadline := time.Now().Add(20 * time.Second) + for time.Now().Before(deadline) { + out, err := ownerB.Run(ctx, "docker inspect -f '{{.State.Status}}' "+ssh.ShellQuote(lateName)+" 2>/dev/null || true") + if err == nil && strings.TrimSpace(out) == "running" { + landed = true + break + } + time.Sleep(500 * time.Millisecond) + } + if !landed { + t.Fatal("the delayed effect never landed — fix the fixture before trusting this scenario") + } + + // The production reconciliation over the observed world: the late + // container is neither the new deploy's candidate naming nor a known + // predecessor's — MANUAL refusal, never a blind deploy. + d := NewDeployer(ownerB, os.Stderr) + err = d.ReconcileAfterTakeover(ctx, Config{App: app, Version: "newgen"}, nil) + if err == nil { + t.Fatal("the production reconciler proceeded on a world with an unattributable running container — quiescence assumption, the C01-1 defect") + } + if !strings.Contains(err.Error(), "MANUAL") { + t.Fatalf("expected MANUAL disposition, got: %v", err) + } + t.Logf("refused as expected: %v", err) +} diff --git a/internal/deploy/reconcile_test.go b/internal/deploy/reconcile_test.go new file mode 100644 index 0000000..40d1071 --- /dev/null +++ b/internal/deploy/reconcile_test.go @@ -0,0 +1,354 @@ +package deploy + +import ( + "bytes" + "context" + "errors" + "fmt" + "strings" + "testing" + "time" + + "github.com/useteploy/teploy/internal/deploy/recovery" + "github.com/useteploy/teploy/internal/ssh" + "github.com/useteploy/teploy/internal/state" +) + +// takeoverMocks is the happy-path deploy mock set with the lock acquisition +// shaped as a STALE-BREAK takeover (C01-1): the first mkdir fails against a +// dead holder's stale info file, the retry succeeds, and the resulting +// fence handle reports TookOver. +func takeoverMocks(app string, staleInfo string, extra ...ssh.MockCommand) (*ssh.MockExecutor, *state.Lock) { + return takeoverMocksBase(app, staleInfo, fenceHappyPathMocks(app), extra) +} + +// takeoverMocksBase is takeoverMocks with the deploy's base mock set +// overridable (tests that need the Files-backed framed state read drop the +// happy set's explicit absent registration). +func takeoverMocksBase(app, staleInfo string, base, extra []ssh.MockCommand) (*ssh.MockExecutor, *state.Lock) { + mkFail := ssh.MockCommand{Match: "mkdir /deployments/" + app + "/.lock", Err: errors.New("mkdir: file exists"), Once: true} + mkOK := ssh.MockCommand{Match: "mkdir /deployments/" + app + "/.lock", Output: ""} + // ReadLock reads the dead holder's info through `cat ... 2>/dev/null`; + // the mock's Files-backed cat only answers the bare form, so the stale + // info is registered as an explicit response. + staleCat := ssh.MockCommand{Match: "cat /deployments/" + app + "/.lock/info", Output: staleInfo, Once: true} + mocks := append([]ssh.MockCommand{mkFail, mkOK, staleCat}, base...) + mocks = append(mocks, extra...) + mock := ssh.NewMockExecutor("1.2.3.4", mocks...) + mock.Files["/deployments/"+app+"/.lock/info"] = []byte(staleInfo) + lk, err := state.AcquireLockFenced(context.Background(), mock, app) + if err != nil { + panic(fmt.Sprintf("takeoverMocks: acquiring the stale-broken lock: %v", err)) + } + return mock, lk +} + +func staleInfoJSON(t *testing.T) string { + t.Helper() + old := time.Now().UTC().Add(-2 * staleLockTTLSafe).Format(time.RFC3339) + return fmt.Sprintf(`{"type":"auto","owner":"deadholder","ts":%q}`, old) +} + +// staleLockTTLSafe mirrors state.staleLockTTL (unexported here) for fixture +// timestamps only. +const staleLockTTLSafe = 30 * time.Minute + +// TestAcquireLockFenced_ReportsTakeover pins the state-package contract the +// reconciliation keys on: a fresh acquisition is not a takeover, breaking a +// stale lock is. +func TestAcquireLockFenced_ReportsTakeover(t *testing.T) { + fresh := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "mkdir /deployments/freshapp/.lock", Output: ""}, + ) + lk, err := state.AcquireLockFenced(context.Background(), fresh, "freshapp") + if err != nil { + t.Fatalf("fresh acquire: %v", err) + } + if lk.TookOver() { + t.Error("a fresh acquisition must not report takeover") + } + var nilLock *state.Lock + if nilLock.TookOver() { + t.Error("nil lock must not report takeover") + } +} + +// TestDeployFenced_TakeoverRefusesOnForeignRunningWork is C01-1's core +// scenario through the production deploy path (the harness's (a), wired): +// a stale lock is broken, the dead holder's late container is running +// under a foreign generation, and the replacement owner must REFUSE — +// reconcile before any of its own effects, never treat acquisition as +// quiescence. +func TestDeployFenced_TakeoverRefusesOnForeignRunningWork(t *testing.T) { + app := "fency" + mock, lk := takeoverMocks(app, staleInfoJSON(t), + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: containerJSON(app, "fency-web-deadgen", "deadgen", "running")}, + ) + var buf bytes.Buffer + d := NewDeployer(mock, &buf) + err := d.DeployFenced(context.Background(), Config{ + App: app, + Domain: "fency.com", + Image: "fency:latest", + Version: "newgen", + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }, lk) + if err == nil { + t.Fatal("expected the takeover deploy to be refused") + } + if !strings.Contains(err.Error(), "MANUAL") { + t.Fatalf("expected MANUAL disposition in the refusal, got: %v", err) + } + if !strings.Contains(err.Error(), "unattributable") { + t.Fatalf("expected the refusal to name the unattributable workload, got: %v", err) + } + for _, c := range mock.Calls { + if strings.HasPrefix(c, "docker run") { + t.Errorf("refused takeover must not start containers, saw: %s", c) + } + if strings.Contains(c, "state.json.tmp-") { + t.Errorf("refused takeover must not commit state, saw: %s", c) + } + } +} + +// TestDeployFenced_TakeoverCleanWorldProceedes: a takeover whose observed +// world matches a clean pre-deploy state (no leftover effects, serving +// predecessor or fresh app) proceeds — RETRY is the table's answer and the +// deploy must not become unusable after a crash. +func TestDeployFenced_TakeoverCleanWorldProceedes(t *testing.T) { + app := "fency" + mock, lk := takeoverMocks(app, staleInfoJSON(t), + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: ""}, + ) + var buf bytes.Buffer + d := NewDeployer(mock, &buf) + if err := d.DeployFenced(context.Background(), Config{ + App: app, + Domain: "fency.com", + Image: "fency:latest", + Version: "abc123", + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }, lk); err != nil { + t.Fatalf("clean-world takeover deploy: %v", err) + } + if !strings.Contains(buf.String(), "Stale lock taken over") { + t.Error("expected the takeover reconciliation to be surfaced to the operator") + } +} + +// TestDeployFenced_TakeoverSameVersionDisagreementRefused: state.json +// already names the release being deployed (a same-version redeploy after a +// takeover) — R6's record/target disagreement: INSPECT, never a blind redo. +func TestDeployFenced_TakeoverSameVersionDisagreementRefused(t *testing.T) { + app := "fency" + // The happy-path set registers state.json as absent; this test needs + // the Files-backed framed read to answer with a real state, so that + // registration is dropped. + happy := fenceHappyPathMocks(app) + filtered := make([]ssh.MockCommand, 0, len(happy)) + for _, m := range happy { + if strings.Contains(m.Match, "/state.json") { + continue + } + filtered = append(filtered, m) + } + base := append([]ssh.MockCommand{ + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: ""}, + }, filtered...) + mock, lk := takeoverMocksBase(app, staleInfoJSON(t), base, nil) + mock.Files["/deployments/"+app+"/state.json"] = []byte(`{"schema_version":2,"deployment_type":"container","ingress_mode":"caddy","current_hash":"abc123","updated_at":"2026-09-23T00:00:00Z"}`) + var buf bytes.Buffer + d := NewDeployer(mock, &buf) + err := d.DeployFenced(context.Background(), Config{ + App: app, + Domain: "fency.com", + Image: "fency:latest", + Version: "abc123", + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }, lk) + if err == nil || !strings.Contains(err.Error(), "INSPECT") { + t.Fatalf("expected INSPECT refusal for same-version disagreement, got: %v", err) + } +} + +// TestReconcileAfterTakeover_Dispositions drives the production reconciler +// over the table's remaining crash-window worlds and pins the mapping: +// running candidate without receipts → INSPECT; traffic on an uncommitted +// generation with a restorable predecessor → COMPENSATE; unreadable +// evidence that stays unreadable → INSPECT after the bounded re-observe. +func TestReconcileAfterTakeover_Dispositions(t *testing.T) { + origDelay := reconcileReobserveDelay + reconcileReobserveDelay = 5 * time.Millisecond + t.Cleanup(func() { reconcileReobserveDelay = origDelay }) + + stateJSON := func(hash string) string { + if hash == "" { + return "" + } + return fmt.Sprintf(`{"schema_version":2,"deployment_type":"container","ingress_mode":"caddy","current_hash":%q,"updated_at":"2026-09-23T00:00:00Z"}`, hash) + } + + t.Run("running candidate without receipts is INSPECT", func(t *testing.T) { + app := "fency" + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: containerJSON(app, "fency-web-c0ffee", "c0ffee", "running")}, + ) + var buf bytes.Buffer + d := NewDeployer(mock, &buf) + err := d.ReconcileAfterTakeover(context.Background(), Config{App: app, Version: "c0ffee"}, nil) + if err == nil || !strings.Contains(err.Error(), "INSPECT") { + t.Fatalf("expected INSPECT, got: %v", err) + } + }) + + t.Run("traffic on uncommitted generation is COMPENSATE", func(t *testing.T) { + app := "fency" + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: containerJSON(app, "fency-web-old123", "old123", "running")}, + ) + mock.Files["/deployments/"+app+"/state.json"] = []byte(stateJSON("old123")) + mock.Files["/deployments/caddy/Caddyfile"] = []byte("# TEPLOY BEGIN fency\nfency.com {\n reverse_proxy fency-web-c0ffee:80\n}\n# TEPLOY END fency\n") + var buf bytes.Buffer + d := NewDeployer(mock, &buf) + err := d.ReconcileAfterTakeover(context.Background(), Config{App: app, Version: "c0ffee"}, mustReadState(t, mock, app)) + if err == nil || !strings.Contains(err.Error(), "COMPENSATION") { + t.Fatalf("expected COMPENSATE, got: %v", err) + } + }) + + t.Run("persistent unreadable evidence stays INSPECT", func(t *testing.T) { + app := "fency" + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Err: errors.New("docker: connection refused")}, + ) + var buf bytes.Buffer + d := NewDeployer(mock, &buf) + err := d.ReconcileAfterTakeover(context.Background(), Config{App: app, Version: "c0ffee"}, nil) + if err == nil || !strings.Contains(err.Error(), "INSPECT") { + t.Fatalf("expected INSPECT for unreadable evidence, got: %v", err) + } + }) +} + +// TestReconcileAfterTakeover_TransientReadRecovers: R4's unreadable classes +// are reconcile triggers — one bounded re-observation absorbs a transient +// inventory failure and the deploy proceeds on the clean re-observation. +func TestReconcileAfterTakeover_TransientReadRecovers(t *testing.T) { + origDelay := reconcileReobserveDelay + reconcileReobserveDelay = 5 * time.Millisecond + t.Cleanup(func() { reconcileReobserveDelay = origDelay }) + + app := "fency" + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Err: errors.New("docker: temporary failure"), Once: true}, + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: ""}, + ) + var buf bytes.Buffer + d := NewDeployer(mock, &buf) + if err := d.ReconcileAfterTakeover(context.Background(), Config{App: app, Version: "c0ffee"}, nil); err != nil { + t.Fatalf("transient inventory failure must recover via re-observation: %v", err) + } + if !strings.Contains(buf.String(), "proceeding") { + t.Error("expected the reconciled takeover to report proceeding") + } +} + +// TestObserve_Classification pins the evidence collector's exact-name +// classification: candidates, corpses, foreign workloads, predecessor +// serving (including a same-version _replaced rename) and stopped, over +// unreadable inventories. +func TestObserve_Classification(t *testing.T) { + app := "fency" + + t.Run("predecessor serving under _replaced rename", func(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: containerJSON(app, "fency-web-old123_replaced", "old123", "running")}, + ) + o := Observe(context.Background(), mock, app, "newgen", "old123") + if o.PredecessorServing != recovery.Present { + t.Errorf("expected PredecessorServing for a running _replaced rename, got %v", o.PredecessorServing) + } + }) + + t.Run("stopped predecessor is restorable evidence", func(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: containerJSON(app, "fency-web-old123", "old123", "exited")}, + ) + o := Observe(context.Background(), mock, app, "newgen", "old123") + if o.PredecessorStopped != recovery.Present { + t.Errorf("expected PredecessorStopped, got %v", o.PredecessorStopped) + } + }) + + t.Run("candidate corpse is not a running candidate", func(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: containerJSON(app, "fency-web-newgen", "newgen", "created")}, + ) + o := Observe(context.Background(), mock, app, "newgen", "old123") + if o.CandidateCorpses != recovery.Present || o.Candidates != recovery.Absent { + t.Errorf("expected corpse evidence only, got candidates=%v corpses=%v", o.Candidates, o.CandidateCorpses) + } + }) + + t.Run("unreadable inventory maps to unknown classes", func(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Err: errors.New("boom")}, + ) + o := Observe(context.Background(), mock, app, "newgen", "old123") + if o.Candidates != recovery.Unknown { + t.Errorf("expected Unknown candidates, got %v", o.Candidates) + } + }) +} + +// containerJSON renders one docker-ps --format JSON line the way +// ParseContainers consumes it (labels in docker's comma-joined display +// form). +func containerJSON(app, name, version, stateName string) string { + return fmt.Sprintf(`{"ID":"deadbeefdead","Names":%q,"Image":"img:1","State":%q,"Status":"up","CreatedAt":"2026-05-28 21:33:29 -0700 PDT","Labels":"teploy.app=%s,teploy.process=web,teploy.version=%s"}`, + name, stateName, app, version) +} + +func mustReadState(t *testing.T, mock *ssh.MockExecutor, app string) *state.AppState { + t.Helper() + st, err := state.Read(context.Background(), mock, app) + if err != nil { + t.Fatalf("reading state fixture: %v", err) + } + return st +} + +// TestObserve_ForeignGenerationRouteIsConflicting: a managed Caddy block +// naming a THIRD generation (neither the attempted release's candidates +// nor the predecessor's containers — the dead-holder late-route-edit +// shape) is CONFLICTING evidence (Unknown), which Decide sends to INSPECT +// — the C01-3 producer/consumer for "conflicting route evidence" that +// previously collapsed into a false "route to predecessor" and could +// yield a blind RETRY. +func TestObserve_ForeignGenerationRouteIsConflicting(t *testing.T) { + app := "fency" + thirdGen := "# TEPLOY BEGIN fency\nfency.com {\n\treverse_proxy fency-web-evilgen:80\n}\n# TEPLOY END fency\n" + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: containerJSON(app, "fency-web-old123", "old123", "running")}, + ) + mock.Files["/deployments/"+app+"/state.json"] = []byte(`{"schema_version":2,"deployment_type":"container","ingress_mode":"caddy","current_hash":"old123","updated_at":"2026-09-23T00:00:00Z"}`) + mock.Files["/deployments/caddy/Caddyfile"] = []byte(thirdGen) + o := Observe(context.Background(), mock, app, "newgen", "old123") + if o.RouteToCandidate != recovery.Unknown || o.RouteToPredecessor != recovery.Unknown { + t.Fatalf("expected conflicting route evidence (Unknown), got candidate=%v predecessor=%v", o.RouteToCandidate, o.RouteToPredecessor) + } + if disp := recovery.Decide(recovery.Admitted, o); disp != recovery.Inspect { + t.Fatalf("expected INSPECT over conflicting route evidence, got %v", disp) + } + + // No managed block for the app at all is clean absence, not conflict. + mock2 := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='fency'", Output: ""}, + ) + mock2.Files["/deployments/caddy/Caddyfile"] = []byte("other.com {\n\trespond 200\n}\n") + o2 := Observe(context.Background(), mock2, app, "newgen", "old123") + if o2.RouteToCandidate != recovery.Absent || o2.RouteToPredecessor != recovery.Absent { + t.Fatalf("expected absent route evidence without a managed block, got candidate=%v predecessor=%v", o2.RouteToCandidate, o2.RouteToPredecessor) + } +} diff --git a/internal/deploy/rollback.go b/internal/deploy/rollback.go index 9471d82..d3202bd 100644 --- a/internal/deploy/rollback.go +++ b/internal/deploy/rollback.go @@ -29,7 +29,12 @@ type RollbackConfig struct { App string Domain string StopTimeout int - Health HealthConfig + // DrainSeconds is the request-drain window between the route switch + // back to the target and stopping the superseded generation (C03) — + // the rollback-side twin of deploy.Config.DrainSeconds. Zero (default) + // keeps the historical stop-immediately behavior. + DrainSeconds int + Health HealthConfig // ToHash rolls back to a specific version instead of just the // immediately previous one, mirroring type:static's existing --to // support (internal/cli/rollback.go). Empty means the immediately @@ -104,6 +109,10 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac } defer state.ReleaseLockFenced(exec, lk, cfg.App) lk.StartRenewal(exec) + // The traffic switch back runs as a guarded Caddyfile commit (C01-2): + // a rollback whose lock was broken mid-flight must not land its route + // edit inside the new owner's window. + cd = cd.WithCommitGuard(lk.GuardPrefix()) // 2. Read state and resolve the rollback target — under the lock. current, err := state.Read(ctx, exec, cfg.App) @@ -511,7 +520,18 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac // now be stopped (match by version label — see step 2). A fence loss // here (F16) means another operation owns the app: refuse further // stops loudly rather than interleave — leaving the superseded workload - // running is degraded but visible. + // running is degraded but visible. The request-drain window (C03) + // applies first when configured: the route already switched back to + // the target, and the superseded generation may still hold in-flight + // requests on its connections. + if cfg.DrainSeconds > 0 && cfg.usesCaddy() { + fmt.Fprintf(out, "Draining %s's superseded generation for %ds (in-flight requests complete; traffic serves %s)...\n", cfg.App, cfg.DrainSeconds, target) + select { + case <-ctx.Done(): + fmt.Fprintf(out, " Drain window cut short (%v) — proceeding to retirement\n", ctx.Err()) + case <-time.After(time.Duration(cfg.DrainSeconds) * time.Second): + } + } for _, c := range containers { if lk != nil { if err := lk.Check(ctx, exec); err != nil { diff --git a/internal/deploy/static.go b/internal/deploy/static.go index c5e3c83..2c3c66f 100644 --- a/internal/deploy/static.go +++ b/internal/deploy/static.go @@ -239,7 +239,7 @@ func (d *StaticDeployer) Deploy(ctx context.Context, cfg StaticConfig) error { } return fmt.Errorf("fence check: %w; the prior static release was restored", err) } - if err := d.caddy.SetStaticRoute(ctx, cfg.App, cfg.Domain, caddy.StaticBlockOpts{ + if err := d.caddy.WithCommitGuard(lk.GuardPrefix()).SetStaticRoute(ctx, cfg.App, cfg.Domain, caddy.StaticBlockOpts{ Root: fmt.Sprintf("%s/%s/current", cfg.MountBase, cfg.App), SPA: cfg.SPA, SPAFallback: cfg.SPAFallback, @@ -694,7 +694,7 @@ func (d *StaticDeployer) Rollback(ctx context.Context, cfg StaticRollbackConfig) } return fmt.Errorf("fence check: %w; release %s was restored", err, prior.CurrentHash) } - if err := d.caddy.SetStaticRoute(ctx, cfg.App, cfg.Domain, caddy.StaticBlockOpts{ + if err := d.caddy.WithCommitGuard(lk.GuardPrefix()).SetStaticRoute(ctx, cfg.App, cfg.Domain, caddy.StaticBlockOpts{ Root: fmt.Sprintf("%s/%s/current", cfg.MountBase, cfg.App), SPA: cfg.SPA, SPAFallback: cfg.SPAFallback, @@ -802,7 +802,7 @@ func (d *StaticDeployer) RollbackStateOnly(ctx context.Context, app, toHash stri } return fmt.Errorf("fence check: %w; release %s was restored", err, prior.CurrentHash) } - if err := d.caddy.SetStaticRoute(ctx, app, domain, caddy.StaticBlockOpts{ + if err := d.caddy.WithCommitGuard(lk.GuardPrefix()).SetStaticRoute(ctx, app, domain, caddy.StaticBlockOpts{ Root: fmt.Sprintf("%s/%s/current", DefaultStaticMount, app), SPA: st.SPA, SPAFallback: st.SPAFallback, diff --git a/internal/docker/docker.go b/internal/docker/docker.go index e84ddcd..f86b395 100644 --- a/internal/docker/docker.go +++ b/internal/docker/docker.go @@ -128,6 +128,21 @@ func NewClient(exec ssh.Executor) *Client { // Run starts a new container and returns its ID. func (c *Client) Run(ctx context.Context, cfg RunConfig) (string, error) { + return c.run(ctx, cfg, "") +} + +// RunGuarded is Run with the container creation composed under a fence +// guard prefix (state.Lock.GuardPrefix) in the SAME remote command — the +// C01-2 composition: no transport window exists between the holdership +// check and the docker run, so a broken lock holder's late container +// starts are refused by the guard instead of landing inside the new +// owner's window. An empty prefix runs the effect unguarded (pre-F16 +// shape). A refused run's error matches state.FenceLost. +func (c *Client) RunGuarded(ctx context.Context, cfg RunConfig, guardPrefix string) (string, error) { + return c.run(ctx, cfg, guardPrefix) +} + +func (c *Client) run(ctx context.Context, cfg RunConfig, guardPrefix string) (string, error) { if cfg.App == "" || cfg.Process == "" || cfg.Version == "" || cfg.Image == "" { return "", fmt.Errorf("run config requires app, process, version, and image") } @@ -256,7 +271,7 @@ func (c *Client) Run(ctx context.Context, cfg RunConfig) (string, error) { args = append(args, cfg.Cmd) } - cmd := strings.Join(args, " ") + cmd := guardPrefix + strings.Join(args, " ") output, err := c.exec.Run(ctx, cmd) if err != nil && nameAlreadyInUse(output, err) { // A container already holds this exact name. That happens routinely after diff --git a/internal/releasemeta/provenance.go b/internal/releasemeta/provenance.go index d3a69da..712123d 100644 --- a/internal/releasemeta/provenance.go +++ b/internal/releasemeta/provenance.go @@ -57,6 +57,9 @@ const provenanceFile = "provenance.json" // keys on. // - ManifestSHA256: the effective-config digest (config. // NormalizeAndDigest) — plan and receipt compare THIS too. +// - PlanID: when the deploy executed a reviewed plan (C05 +// `teploy apply`), the plan record's id — the tie-back from the +// receipt to the plan that was verified. Empty for direct deploys. type Provenance struct { SchemaVersion int `json:"schema_version"` App string `json:"app"` @@ -77,6 +80,7 @@ type Provenance struct { ImageDigest string `json:"image_digest,omitempty"` DigestPinned bool `json:"digest_pinned,omitempty"` ManifestSHA256 string `json:"manifest_sha256,omitempty"` + PlanID string `json:"plan_id,omitempty"` } // AttemptProvenancePath is the receipt's location in the attempt namespace. diff --git a/internal/ssh/mock.go b/internal/ssh/mock.go index aa36248..d292e2d 100644 --- a/internal/ssh/mock.go +++ b/internal/ssh/mock.go @@ -142,25 +142,42 @@ func (m *MockExecutor) Run(ctx context.Context, cmd string) (string, error) { return "", fmt.Errorf("mock: unexpected command: %s", cmd) } -// evalFenceGuard recognizes the guard fragment produced by state.Lock. It -// returns the remaining effect command ("" for a bare guard), whether the -// guard holds against the recorded files, and whether cmd was a guard at -// all. Must be called with m.mu held. +// evalFenceGuard recognizes guard fragments produced by state.Lock and +// the caddy lock (C01-3) — possibly CHAINED (an app-fence guard followed +// by the caddy-lock guard on one commit command, C01-2/C01-3 +// composition). It returns the remaining effect command ("" for a bare +// guard), whether EVERY guard holds against the recorded files, and +// whether cmd carried at least one guard at all. Must be called with m.mu +// held. func evalFenceGuard(files map[string][]byte, cmd string) (rest string, held, ok bool) { const guardSep = " || { printf 'TEPLOY_FENCE_LOST\\n' >&2; exit 75; }; " if !strings.HasPrefix(cmd, "grep -q ") { return "", false, false } - guard, effect := cmd, "" - if i := strings.Index(cmd, guardSep); i >= 0 { - guard, effect = cmd[:i], cmd[i+len(guardSep):] + effect := "" + held = true + ok = false + for strings.HasPrefix(cmd, "grep -q ") { + guard := cmd + if i := strings.Index(cmd, guardSep); i >= 0 { + guard, effect = cmd[:i], cmd[i+len(guardSep):] + } else { + effect = "" + } + owner, path, parsed := parseFenceGuard(guard) + if !parsed { + break + } + ok = true + data, present := files[path] + if !present || !bytes.Contains(data, []byte(owner)) { + held = false + } + cmd = effect } - owner, path, parsed := parseFenceGuard(guard) - if !parsed { + if !ok { return "", false, false } - data, present := files[path] - held = present && bytes.Contains(data, []byte(owner)) return effect, held, true } diff --git a/internal/state/lock.go b/internal/state/lock.go index 75e3118..e7b360e 100644 --- a/internal/state/lock.go +++ b/internal/state/lock.go @@ -72,6 +72,10 @@ const ( type Lock struct { app string owner string + // tookOver records that acquiring this lock BROKE a stale predecessor + // (C01-1): the previous holder died somewhere inside its lifecycle and + // its effects may still be landing. Exposed via TookOver. + tookOver bool mu sync.Mutex lost bool @@ -85,10 +89,22 @@ type Lock struct { // that ignore the handle get exactly the old behavior. func AcquireLockFenced(ctx context.Context, exec ssh.Executor, app string) (*Lock, error) { owner := newOperationID() - if err := acquireAutoLock(ctx, exec, app, owner); err != nil { + tookOver, err := acquireAutoLock(ctx, exec, app, owner) + if err != nil { return nil, err } - return &Lock{app: app, owner: owner}, nil + return &Lock{app: app, owner: owner, tookOver: tookOver}, nil +} + +// TookOver reports whether this lock's acquisition broke a stale +// predecessor's lock (C01-1). A replacement owner must reconcile the +// observed target — never treat acquisition as proof of quiescence; the +// dead holder's Docker/route effects can still be in flight. +func (l *Lock) TookOver() bool { + if l == nil { + return false + } + return l.tookOver } func lockInfoPath(app string) string { @@ -118,6 +134,30 @@ func (l *Lock) guardFragment() string { return fmt.Sprintf("grep -q %s %s", ssh.ShellQuote(l.owner), ssh.ShellQuote(lockInfoPath(l.app))) } +// GuardPrefix returns the shell prefix that composes the holdership guard +// with an effect command into ONE remote invocation: +// +// +// +// where the prefix refuses (marker on stderr, exit 75) when the lock no +// longer names this operation. Effect sites outside this package (docker +// runs, the Caddyfile commit) use it to close the check-then-act window +// C01-2 documents; map refusals with FenceLost. A nil lock yields "" (the +// caller runs the effect unguarded — the pre-F16 shape). +func (l *Lock) GuardPrefix() string { + if l == nil { + return "" + } + return l.guardFragment() + " || { printf '" + fenceLostMarker + `\n' >&2; exit 75; }; ` +} + +// FenceLost reports whether err is a fence refusal — the guard fragment's +// marker or its exit status surfaced by either executor flavor — including +// for effects composed by other packages through Lock.GuardPrefix. +func FenceLost(err error) bool { + return fenceLostErr(err) +} + // Check verifies the server still names this operation as the lock holder. // Any failure — including transport failure, because an unreachable answer // cannot prove holdership — reports ErrFenceLost. Call before effectful @@ -324,29 +364,29 @@ func ReleaseLockFenced(exec ssh.Executor, lk *Lock, app string) { fmt.Fprintf(os.Stderr, "teploy: refusing to release %s's lock with a lease held for %s\n", app, lk.App()) return } - ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) - defer cancel() - lockDir := fmt.Sprintf("%s/%s/.lock", deploymentsDir, app) - _, err := lk.Guarded(ctx, exec, "rm -rf -- "+ssh.ShellQuote(lockDir)) - if err != nil && !fenceLostErr(err) { - // The guarded release failed ambiguously (transport timeout, for - // one): the release MAY have completed, and a successor may have - // acquired the path in the meantime. An unconditional detached - // release here can delete the SUCCESSOR's lock (audit T02) — but - // never releasing strands the app for a full staleLockTTL. Resolve - // the ambiguity with one shell-level conditional: remove the lock - // only when it still names THIS operation, or when it is already - // gone. A lock that names someone else is left strictly alone. - conditional := fmt.Sprintf( - "if [ -d %s ] && grep -q %s %s 2>/dev/null; then rm -rf -- %s; fi", - ssh.ShellQuote(lockDir), ssh.ShellQuote(lk.owner), ssh.ShellQuote(lockInfoPath(app)), ssh.ShellQuote(lockDir), - ) - if _, cerr := exec.Run(ctx, conditional); cerr != nil { - fmt.Fprintf(os.Stderr, "teploy: could not confirm release of %s's deploy lock: %v (the lock will self-heal after the stale window if abandoned)\n", app, cerr) + ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) + defer cancel() + lockDir := fmt.Sprintf("%s/%s/.lock", deploymentsDir, app) + _, err := lk.Guarded(ctx, exec, "rm -rf -- "+ssh.ShellQuote(lockDir)) + if err != nil && !fenceLostErr(err) { + // The guarded release failed ambiguously (transport timeout, for + // one): the release MAY have completed, and a successor may have + // acquired the path in the meantime. An unconditional detached + // release here can delete the SUCCESSOR's lock (audit T02) — but + // never releasing strands the app for a full staleLockTTL. Resolve + // the ambiguity with one shell-level conditional: remove the lock + // only when it still names THIS operation, or when it is already + // gone. A lock that names someone else is left strictly alone. + conditional := fmt.Sprintf( + "if [ -d %s ] && grep -q %s %s 2>/dev/null; then rm -rf -- %s; fi", + ssh.ShellQuote(lockDir), ssh.ShellQuote(lk.owner), ssh.ShellQuote(lockInfoPath(app)), ssh.ShellQuote(lockDir), + ) + if _, cerr := exec.Run(ctx, conditional); cerr != nil { + fmt.Fprintf(os.Stderr, "teploy: could not confirm release of %s's deploy lock: %v (the lock will self-heal after the stale window if abandoned)\n", app, cerr) + } } + return } - return -} ReleaseLockDetached(exec, app) } diff --git a/internal/state/state.go b/internal/state/state.go index 5155978..c0a3b66 100644 --- a/internal/state/state.go +++ b/internal/state/state.go @@ -481,13 +481,21 @@ func (l *LockInfo) IsStale() bool { // fencing handle (audit F16) use AcquireLockFenced; the lock taken is the // same — this is that call with the handle discarded. func AcquireLock(ctx context.Context, exec ssh.Executor, app string) error { - return acquireAutoLock(ctx, exec, app, newOperationID()) + _, err := acquireAutoLock(ctx, exec, app, newOperationID()) + return err } // acquireAutoLock is the shared "auto" lock acquisition. Every acquire // carries a unique owner token (F16) so effect sites can refuse a broken // holder's late writes; the token is opaque to everything pre-F16. -func acquireAutoLock(ctx context.Context, exec ssh.Executor, app, owner string) error { +// +// The returned takeover flag is true when this acquisition BROKE a stale +// (or stale-heal) lock to take the path — the C01-1 signal that the +// previous holder may have left in-flight effects on the target: a +// replacement owner must reconcile the observed world before its own +// effects, never treat acquisition as quiescence. +func acquireAutoLock(ctx context.Context, exec ssh.Executor, app, owner string) (bool, error) { + tookOver := false lockPath := fmt.Sprintf("%s/%s/.lock", deploymentsDir, app) if _, err := tryMkdirLock(ctx, exec, lockPath); err != nil { info, _ := ReadLock(ctx, exec, app) @@ -497,7 +505,7 @@ func acquireAutoLock(ctx context.Context, exec ssh.Executor, app, owner string) msg += fmt.Sprintf(": '%s'", info.Message) } msg += fmt.Sprintf(". Locked at %s. Use 'teploy unlock' to release.", info.TS) - return fmt.Errorf("%s", msg) + return false, fmt.Errorf("%s", msg) } if info != nil && info.Type == "auto" && isStale(info.TS, info.RenewTS) { ReleaseLock(ctx, exec, app) @@ -505,9 +513,10 @@ func acquireAutoLock(ctx context.Context, exec ssh.Executor, app, owner string) // Someone else's deploy won the race to re-acquire right // after we broke the stale lock — fall through to the // normal "in progress" error below. - return fmt.Errorf("deploy is already in progress for %s", app) + return false, fmt.Errorf("deploy is already in progress for %s", app) } - return writeLockInfo(ctx, exec, lockPath, app, owner) + tookOver = true + return tookOver, writeLockInfo(ctx, exec, lockPath, app, owner) } // A crashed heal can leave its short-lived "heal" lock behind. A deploy // (authoritative) may break a STALE heal lock so it isn't blocked — but @@ -517,13 +526,14 @@ func acquireAutoLock(ctx context.Context, exec ssh.Executor, app, owner string) if info != nil && info.Type == "heal" && isHealStale(info.TS) { ReleaseLock(ctx, exec, app) if _, retryErr := tryMkdirLock(ctx, exec, lockPath); retryErr != nil { - return fmt.Errorf("deploy is already in progress for %s", app) + return false, fmt.Errorf("deploy is already in progress for %s", app) } - return writeLockInfo(ctx, exec, lockPath, app, owner) + tookOver = true + return tookOver, writeLockInfo(ctx, exec, lockPath, app, owner) } - return fmt.Errorf("deploy is already in progress for %s", app) + return false, fmt.Errorf("deploy is already in progress for %s", app) } - return writeLockInfo(ctx, exec, lockPath, app, owner) + return false, writeLockInfo(ctx, exec, lockPath, app, owner) } func tryMkdirLock(ctx context.Context, exec ssh.Executor, lockPath string) (string, error) {