diff --git a/.gitignore b/.gitignore index d9682c7..cca8ff8 100644 --- a/.gitignore +++ b/.gitignore @@ -2,6 +2,9 @@ /teploy /cmd/teploy/teploy +# Release-verify local build matrix (scripts/release-verify.sh) +/dist-verify/ + # Go *.exe *.test diff --git a/Dockerfile b/Dockerfile new file mode 100644 index 0000000..eabb7c3 --- /dev/null +++ b/Dockerfile @@ -0,0 +1,13 @@ +# Smoke-test vehicle for the R01 release-verify built-image check +# (scripts/release-verify.sh smoke). NOT a published release artifact — +# goreleaser ships binaries/archives; this image exists so the exact +# verified binary proves it runs in a minimal (scratch) container. +# The binary is COPY'd from the release-verify dist directory, so the +# image provenance is always the checksummed matrix build. +# +# Build (as the smoke script does): +# docker build --build-arg BINARY=dist-verify/teploy_linux_ -t teploy:release-smoke . +FROM scratch +ARG BINARY=dist-verify/teploy_linux_amd64 +COPY ${BINARY} /teploy +ENTRYPOINT ["/teploy"] diff --git a/Makefile b/Makefile index 0fe770a..1bbc269 100644 --- a/Makefile +++ b/Makefile @@ -1,4 +1,4 @@ -.PHONY: build test lint vet clean +.PHONY: build test lint vet clean quickstart release-verify release-record release-smoke build: go build -o teploy ./cmd/teploy @@ -14,3 +14,22 @@ vet: clean: rm -f teploy + +# Executable quickstart (C09): deploy the maintained fixture app to the +# local colima VM and verify it answers. Skips honestly (exit 0) when no +# local docker target exists. +quickstart: + ./examples/quickstart/run.sh + +# R01 release receipts: build the goreleaser matrix locally, checksum, +# and diff against recorded expectations (see release/RELEASE_RECEIPT.md). +release-verify: + ./scripts/release-verify.sh verify + +release-record: + ./scripts/release-verify.sh record + +# R01 built-image smoke: build the container image from the verified +# matrix binary and run version + doctor in it. Skips without docker. +release-smoke: + ./scripts/release-verify.sh smoke diff --git a/README.md b/README.md index 99e11a9..18d411f 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@

teploy

-

Zero-downtime Docker deploys to any server via SSH.
Single binary. No management server. No dependencies.

+

Docker deploys to any Linux server you can SSH into, with blue/green zero-downtime under Caddy.
Single binary. No management server.

@@ -13,7 +13,7 @@ ## Why teploy? -Most deploy tools require either a management server (Coolify, Dokploy) or complex configuration (Kamal). Teploy is a single binary that deploys Docker containers to any server you can SSH into. Three lines of config, one command to deploy. +Most deploy tools require either a management server (Coolify, Dokploy) or longer configuration (Kamal). Teploy is a single binary that deploys Docker containers to any server you can SSH into. Three lines of config, one command to deploy. ```yaml # teploy.yml @@ -26,7 +26,11 @@ server: 1.2.3.4 teploy deploy ``` -Your app is live with HTTPS, zero-downtime deploys, and automatic rollback on failure. +With the default Caddy ingress your app is live with automatic HTTPS, +blue/green zero-downtime deploys, and rollback to the previous version if +the readiness gate fails. (`ingress: host` publishes a raw port instead: +recreate-style deploys with seconds of downtime — see +[docs/supported-workloads.md](docs/supported-workloads.md).) ## Install @@ -81,16 +85,16 @@ standalone. 2. **Starts** a new container alongside the old one 3. **Health checks** the new container 4. **Routes traffic** via Caddy (automatic HTTPS) -5. **Stops** the old container — zero downtime -6. **Rolls back** automatically if anything fails +5. **Stops** the old container — no downtime during the switch under Caddy +6. **Rolls back** to the previous version if the readiness gate fails ## Features | Feature | Description | |---|---| -| **Zero-downtime deploys** | New container starts and passes health checks before old one stops | +| **Zero-downtime deploys** | New container starts and passes health checks before old one stops (Caddy blue/green; `ingress: host` deploys by recreate — brief downtime, documented) | | **Automatic HTTPS** | Caddy provisions and renews TLS certificates | -| **Rollback** | `teploy rollback` reverts to the previous version instantly | +| **Rollback** | `teploy rollback` reverts to the previous version and health-gates it before answering | | **Multi-process** | Run web, worker, and scheduler from the same image | | **Accessories** | Manage Postgres, Redis, etc. alongside your app | | **Environment variables** | `teploy env set KEY=value` — stored securely on server | @@ -631,8 +635,20 @@ the server), and a five-command human-confirmed rebuild runbook. Read ## Requirements -- A server with SSH access (any Linux VPS — Hetzner, DigitalOcean, Linode, etc.) -- That's it. `teploy setup` handles the rest. +- A Linux server with SSH key access (any VPS — Hetzner, DigitalOcean, Linode, etc.) +- That's it for the standard path. `teploy setup` installs Docker, Caddy, and rsync. Already-provisioned Docker hosts work too (an `ingress: host` deploy needs nothing else). + +One app runs one image; multi-image stacks are refused, not approximated — see the full matrix and declared limits in [docs/supported-workloads.md](docs/supported-workloads.md). + +## Docs + +- [First success](docs/first-success.md) — install to verified deploy and rollback, with failure modes and remedies +- [Supported workloads](docs/supported-workloads.md) — what deploys today, what is refused, ingress guarantees, operational limits +- [Failure and recovery](docs/failure-and-recovery.md) — error envelope and exit codes, `teploy doctor`, interrupted deploys and repair debt, DR bundles (`teploy dr`) +- [Migration](docs/migration.md) — Dokploy/Coolify/Compose import: what converts, what refuses, concept mapping +- [CI/CD](docs/ci-deploy.md) — push-to-deploy with Forgejo/GitHub Actions +- [Secrets scanning](docs/secrets-scanning.md) — Gitleaks recipe +- [Resilience](docs/resilience.md) — surviving server loss (topology + runbook) ## Comparison diff --git a/contracts/MANIFEST.md b/contracts/MANIFEST.md index b3375e6..54e8618 100644 --- a/contracts/MANIFEST.md +++ b/contracts/MANIFEST.md @@ -10,6 +10,7 @@ Neutron/Nucleus dependency and a public mirror. | Corpus rev | Emitting CLI | Machine Interface | Notes | |---|---|---|---| +| 5 | main (X02 S2 tail: server-status fixtures + schema correction) | 2 | server-status-envelope fixtures landed (was "pending live capture"): valid x2 (full healthy observation, partial-caddy-unavailable — the class a target without a caddy container produces) + legacy pre-MI (machine_interface absent, the 42243e2-era shape). Encoder-derived: generated from the REAL `collectServerStatus` via a mock SSH executor (`contracts_golden_test.go`, TEPLOY_UPDATE_CONTRACTS) — synthetic values, real encoder and parse stages; the wire shape was verified against a live `server status --json` run before pinning. Defect fixed in the same commit: the schema had copied the appStatus root since its S2 draft (its own defect-fix commit 08cfb1b said so) and never described the actual serverStatusDTO wire format (server/host/uptime/load/memory/disks/docker/caddy) — rewritten to the real root with strict required-key coverage of the DTO's no-omitempty fields. Additive to consumers (a schema that matched nothing before now matches the wire); no MI bump. | | 4 | main (X02 S2 tail: server-list reshape) | 2 | **The MI 2 bump** (D8 non-additive): `server list --json` now emits the envelope `{machine_interface, servers[], observed_at}` carrying the per-server fields unchanged (name + id/host/user/role/tags/vpn_ip); the pre-reshape bare map-of-servers root is GONE on the wire and is pinned as the artifact's legacy class. New artifact server-list-envelope (schema + valid + legacy fixtures); version-handshake schema maximum 1→2 and its valid fixture renamed mi1→mi2 (app-list valid likewise — both envelopes now report MI 2). Capability tokens unchanged. Coordinated consumer: teploy-dash decodes both shapes during the transition (MaxSupportedMachineInterface 2). | | 3 (amended) | main (C05 plan-record corpus + defect fix) | 1 | C05 added the plan-record artifact + plan-apply token (see git history); amendment: server-status schema now carries its own $defs (its $refs never resolved), and app-list fixtures emit [] where the encoder emits [] (null fixtures failed schema + the real dash decode - found by dash's new contracts CI job, fixed here). | | 1 | post-v0.1.37 main (S2 skeleton) | 1 | First goldens: version handshake, app-list envelope (MI + pre-MI legacy), error envelope (config-invalid, internal, invalid-code), release-record, attempt-name grammar, preview-state eras. | @@ -23,7 +24,7 @@ Neutron/Nucleus dependency and a public mirror. | version-handshake | yes | valid (real `writeVersion` encoder) | teploy-cli | | app-list-envelope | yes | valid (real DTO tags) + legacy pre-MI | teploy-cli | | server-list-envelope | yes (MI 2 reshape) | valid (real `writeServerList` encoder) + legacy bare-map | teploy-cli | -| server-status-envelope | yes (appStatus root) | pending S2 tail (live `server status` capture) | teploy-cli | +| server-status-envelope | yes (serverStatusDTO root, corrected rev 5) | valid x2 (full, partial-caddy-unavailable; real `collectServerStatus` encoder over mock executor) + legacy pre-MI | teploy-cli | | error-envelope | yes | valid x2 + invalid code | teploy-cli | | release-record | yes | valid container | teploy-cli | | attempt-name | yes (pattern) | valid + invalid examples | teploy-cli | diff --git a/contracts/fixtures/server-status-envelope/legacy/pre-mi.json b/contracts/fixtures/server-status-envelope/legacy/pre-mi.json new file mode 100644 index 0000000..14fcada --- /dev/null +++ b/contracts/fixtures/server-status-envelope/legacy/pre-mi.json @@ -0,0 +1,93 @@ +{ + "caddy": { + "available": true, + "routes": [ + { + "handlers": [ + "reverse_proxy" + ], + "hosts": [ + "myapp.example.com" + ], + "id": "", + "server": "srv0", + "status_code": "", + "upstreams": [ + "myapp-web-3:3000" + ] + }, + { + "handlers": [ + "subroute" + ], + "hosts": [ + "myapp.example.com" + ], + "id": "myapp", + "server": "srv0", + "status_code": "", + "upstreams": [] + } + ] + }, + "disks": [ + { + "available_bytes": 750, + "filesystem": "/dev/vda1", + "mountpoint": "/", + "total_bytes": 1000, + "used_bytes": 250, + "used_percent": "25%" + }, + { + "available_bytes": 1500, + "filesystem": "/dev/vdb1", + "mountpoint": "/srv", + "total_bytes": 2000, + "used_bytes": 500, + "used_percent": "26%" + } + ], + "docker": { + "containers": [ + { + "created_at": "2026-09-23 11:55:00 +0000 UTC", + "id": "9f31c02", + "image": "example/myapp:3", + "name": "myapp-web-3", + "process": "web", + "state": "running", + "status": "Up 4 minutes", + "version": "3" + } + ], + "images": [ + { + "created_at": "2026-09-23 11:50:00 +0000 UTC", + "id": "sha256:1a2b3c4d5e6f", + "repository": "example/myapp", + "size": "25MB", + "tag": "3" + } + ], + "installed": true, + "version": "29.0.0" + }, + "errors": [], + "host": "192.0.2.10", + "load": { + "fifteen": 0.3, + "five": 0.2, + "one": 0.1 + }, + "memory": { + "available_bytes": 409600, + "total_bytes": 1024000, + "used_bytes": 614400 + }, + "observed_at": "2026-09-23T12:00:00Z", + "server": "prod", + "uptime": { + "seconds": 3600.5 + } +} diff --git a/contracts/fixtures/server-status-envelope/valid/full.json b/contracts/fixtures/server-status-envelope/valid/full.json new file mode 100644 index 0000000..9271c4f --- /dev/null +++ b/contracts/fixtures/server-status-envelope/valid/full.json @@ -0,0 +1,94 @@ +{ + "machine_interface": 2, + "server": "prod", + "host": "192.0.2.10", + "uptime": { + "seconds": 3600.5 + }, + "load": { + "one": 0.1, + "five": 0.2, + "fifteen": 0.3 + }, + "memory": { + "total_bytes": 1024000, + "used_bytes": 614400, + "available_bytes": 409600 + }, + "disks": [ + { + "filesystem": "/dev/vda1", + "mountpoint": "/", + "total_bytes": 1000, + "used_bytes": 250, + "available_bytes": 750, + "used_percent": "25%" + }, + { + "filesystem": "/dev/vdb1", + "mountpoint": "/srv", + "total_bytes": 2000, + "used_bytes": 500, + "available_bytes": 1500, + "used_percent": "26%" + } + ], + "docker": { + "installed": true, + "version": "29.0.0", + "containers": [ + { + "id": "9f31c02", + "name": "myapp-web-3", + "image": "example/myapp:3", + "state": "running", + "status": "Up 4 minutes", + "created_at": "2026-09-23 11:55:00 +0000 UTC", + "process": "web", + "version": "3" + } + ], + "images": [ + { + "id": "sha256:1a2b3c4d5e6f", + "repository": "example/myapp", + "tag": "3", + "size": "25MB", + "created_at": "2026-09-23 11:50:00 +0000 UTC" + } + ] + }, + "caddy": { + "available": true, + "routes": [ + { + "server": "srv0", + "id": "", + "hosts": [ + "myapp.example.com" + ], + "handlers": [ + "reverse_proxy" + ], + "upstreams": [ + "myapp-web-3:3000" + ], + "status_code": "" + }, + { + "server": "srv0", + "id": "myapp", + "hosts": [ + "myapp.example.com" + ], + "handlers": [ + "subroute" + ], + "upstreams": [], + "status_code": "" + } + ] + }, + "observed_at": "2026-09-23T12:00:00Z", + "errors": [] +} diff --git a/contracts/fixtures/server-status-envelope/valid/partial-caddy-unavailable.json b/contracts/fixtures/server-status-envelope/valid/partial-caddy-unavailable.json new file mode 100644 index 0000000..f75bfbe --- /dev/null +++ b/contracts/fixtures/server-status-envelope/valid/partial-caddy-unavailable.json @@ -0,0 +1,45 @@ +{ + "machine_interface": 2, + "server": "staging", + "host": "192.0.2.20", + "uptime": { + "seconds": 86400 + }, + "load": { + "one": 0, + "five": 0.01, + "fifteen": 0.05 + }, + "memory": { + "total_bytes": 512000, + "used_bytes": 256000, + "available_bytes": 256000 + }, + "disks": [ + { + "filesystem": "/dev/vda1", + "mountpoint": "/", + "total_bytes": 500, + "used_bytes": 100, + "available_bytes": 400, + "used_percent": "20%" + } + ], + "docker": { + "installed": true, + "version": "29.0.0", + "containers": [], + "images": [] + }, + "caddy": { + "available": false, + "routes": [] + }, + "observed_at": "2026-09-23T12:00:00Z", + "errors": [ + { + "scope": "caddy.routes", + "message": "Error response from daemon: No such container: caddy" + } + ] +} diff --git a/contracts/schema/server-status-envelope.schema.json b/contracts/schema/server-status-envelope.schema.json index 7e43687..24f33d9 100644 --- a/contracts/schema/server-status-envelope.schema.json +++ b/contracts/schema/server-status-envelope.schema.json @@ -1,91 +1,166 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", "$id": "https://teploy.github.io/contracts/schema/server-status-envelope.schema.json", - "title": "server status --json envelope (MI 1, appStatusDTO root)", - "allOf": [ - { + "title": "server status --json envelope (serverStatusDTO root)", + "type": "object", + "required": [ + "machine_interface", + "server", + "host", + "uptime", + "load", + "memory", + "disks", + "docker", + "caddy", + "observed_at", + "errors" + ], + "properties": { + "machine_interface": { + "type": "integer", + "minimum": 1 + }, + "server": { + "type": "string" + }, + "host": { + "type": "string" + }, + "uptime": { + "$ref": "#/$defs/uptime" + }, + "load": { + "$ref": "#/$defs/load" + }, + "memory": { + "$ref": "#/$defs/memory" + }, + "disks": { + "type": "array", + "items": { + "$ref": "#/$defs/disk" + } + }, + "docker": { + "$ref": "#/$defs/dockerInventory" + }, + "caddy": { + "$ref": "#/$defs/caddyObservation" + }, + "observed_at": { + "type": "string", + "format": "date-time" + }, + "errors": { + "type": "array", + "items": { + "$ref": "#/$defs/machineError" + } + } + }, + "$defs": { + "machineError": { "type": "object", "required": [ - "app", - "domain", - "type", - "ingress", - "current_release", - "previous_release", - "containers", - "processes", - "lock", - "maintenance", - "observed_at", - "errors", - "machine_interface" + "scope", + "message" ], "properties": { - "app": { - "type": "string" - }, - "domain": { - "type": "string" - }, - "type": { + "scope": { "type": "string" }, - "ingress": { + "message": { "type": "string" + } + } + }, + "uptime": { + "type": "object", + "required": [ + "seconds" + ], + "properties": { + "seconds": { + "type": "number", + "minimum": 0 + } + } + }, + "load": { + "type": "object", + "required": [ + "one", + "five", + "fifteen" + ], + "properties": { + "one": { + "type": "number", + "minimum": 0 }, - "current_release": { - "$ref": "#/$defs/release" - }, - "previous_release": { - "$ref": "#/$defs/release" - }, - "containers": { - "type": "array", - "items": { - "$ref": "#/$defs/container" - } - }, - "processes": { - "type": "array" - }, - "lock": { - "type": [ - "object", - "null" - ] - }, - "maintenance": { - "type": "boolean" + "five": { + "type": "number", + "minimum": 0 }, - "observed_at": { - "type": "string", - "format": "date-time" + "fifteen": { + "type": "number", + "minimum": 0 + } + } + }, + "memory": { + "type": "object", + "required": [ + "total_bytes", + "used_bytes", + "available_bytes" + ], + "properties": { + "total_bytes": { + "type": "integer", + "minimum": 0 }, - "errors": { - "type": "array", - "items": { - "$ref": "#/$defs/machineError" - } + "used_bytes": { + "type": "integer", + "minimum": 0 }, - "machine_interface": { + "available_bytes": { "type": "integer", - "minimum": 1 + "minimum": 0 } } - } - ], - "$defs": { - "machineError": { + }, + "disk": { "type": "object", "required": [ - "scope", - "message" + "filesystem", + "mountpoint", + "total_bytes", + "used_bytes", + "available_bytes", + "used_percent" ], "properties": { - "scope": { + "filesystem": { "type": "string" }, - "message": { + "mountpoint": { + "type": "string" + }, + "total_bytes": { + "type": "integer", + "minimum": 0 + }, + "used_bytes": { + "type": "integer", + "minimum": 0 + }, + "available_bytes": { + "type": "integer", + "minimum": 0 + }, + "used_percent": { "type": "string" } } @@ -97,7 +172,10 @@ "name", "image", "state", - "status" + "status", + "created_at", + "process", + "version" ], "properties": { "id": { @@ -126,85 +204,116 @@ } } }, - "release": { + "image": { "type": "object", "required": [ + "id", + "repository", + "tag", + "size", + "created_at" + ], + "properties": { + "id": { + "type": "string" + }, + "repository": { + "type": "string" + }, + "tag": { + "type": "string" + }, + "size": { + "type": "string" + }, + "created_at": { + "type": "string" + } + } + }, + "dockerInventory": { + "type": "object", + "required": [ + "installed", "version", - "ports" + "containers", + "images" ], "properties": { + "installed": { + "type": "boolean" + }, "version": { "type": "string" }, - "ports": { + "containers": { "type": "array", "items": { - "type": "integer" + "$ref": "#/$defs/container" + } + }, + "images": { + "type": "array", + "items": { + "$ref": "#/$defs/image" } } } }, - "appStatus": { + "caddyRoute": { "type": "object", "required": [ - "app", - "domain", - "type", - "ingress", - "current_release", - "previous_release", - "containers", - "processes", - "lock", - "maintenance", - "observed_at", - "errors" + "server", + "id", + "hosts", + "handlers", + "upstreams", + "status_code" ], "properties": { - "app": { + "server": { "type": "string" }, - "domain": { - "type": "string" - }, - "type": { - "type": "string" - }, - "ingress": { + "id": { "type": "string" }, - "current_release": { - "$ref": "#/$defs/release" - }, - "previous_release": { - "$ref": "#/$defs/release" - }, - "containers": { + "hosts": { "type": "array", "items": { - "$ref": "#/$defs/container" + "type": "string" } }, - "processes": { - "type": "array" + "handlers": { + "type": "array", + "items": { + "type": "string" + } }, - "lock": { - "type": [ - "object", - "null" - ] + "upstreams": { + "type": "array", + "items": { + "type": "string" + } }, - "maintenance": { + "status_code": { + "type": "string" + } + } + }, + "caddyObservation": { + "type": "object", + "required": [ + "available", + "routes" + ], + "properties": { + "available": { "type": "boolean" }, - "observed_at": { - "type": "string", - "format": "date-time" - }, - "errors": { + "routes": { "type": "array", "items": { - "$ref": "#/$defs/machineError" + "$ref": "#/$defs/caddyRoute" } } } diff --git a/docs/C01_RECOVERY_STATE_TABLE.md b/docs/C01_RECOVERY_STATE_TABLE.md index f915523..6cdfc34 100644 --- a/docs/C01_RECOVERY_STATE_TABLE.md +++ b/docs/C01_RECOVERY_STATE_TABLE.md @@ -242,15 +242,48 @@ table's, with the register item it belongs to. INSPECT with an adopt-as-predecessor continuation; automating it requires generation-scoped identities (register F04/A09). The current code's MANUAL is the safe subset — recorded as a disagreement, not a - defect. + defect. **The generation-identity sub-slice LANDED 2026-09-24** + (F04/A09's contained core, see AUDIT_OPEN's C01 generation-identity + slice): every teploy container now carries an immutable + `teploy.generation` label (RunConfig.Generation; preserved across + Recreate), the state commit publishes a `.generation` sidecar + (`/deployments//.generation`, the targetguard contract) under the + same guarded command as state.json, and rollback — the operation the + finding is about — resolves an explicit `fromGeneration` identity and + fences every destructive effect on it: retirement stops and fixed-port + displacements stop by exact container ID under a composed + holdership+label check (`docker.StopGenerationFenced`), target + restarts guard their force-remove the same way + (`docker.RestartFenced`), and the state commit CASes on the sidecar + (`state.WriteFencedGeneration` refuses ErrGenerationFenced naming both + generations when a successor committed in between). A stale rollback + can no longer stop, remove, or overwrite a newer generation's + workloads — proven against the live fixture with a genuinely DELAYED + stop command (nohup) refused at execution time by the live label. The + INSPECT→adopt continuation for a running `_replaced` stays MANUAL + (A08) — the identity surface it needs now exists; the automation + remains the F04-keyed recovery-owner decision. 9. **C01-9 — Candidate identities are version-keyed, not - attempt/generation-keyed.** `{app}-{process}-{version}[-{index}]` - (`internal/docker/docker.go:82-117`) means two attempts of the same - release hash share candidate names; evidence attribution between them - relies on the `_replaced` convention alone. The table's "exact - container IDs" evidence requirement points at attempt-scoped names — - F08's attempt ids are the existing keying surface (register F04/A09). + attempt/generation-keyed.** `{app}-{process}-{version}[-{index}]` + (`internal/docker/docker.go:82-117`) means two attempts of the same + release hash share candidate names; evidence attribution between them + relies on the `_replaced` convention alone. The table's "exact + container IDs" evidence requirement points at attempt-scoped names — + F08's attempt ids are the existing keying surface (register F04/A09). + **The generation-keyed sub-slice LANDED 2026-09-24**: candidates are + labeled `teploy.generation=` (deploy stamps the generation it + creates; rollback's recreates preserve the label), every destructive + effect addresses containers by the exact ID from the operation's own + under-lock inventory, and the ID-bearing command reads the label AT + EXECUTION — the immutable identity that survives name reuse, which is + what closes the delayed-SSH-effect window holdership guards cannot. + Managed Caddy blocks carry the same identity inside the markers + (`# TEPLOY GENERATION `), so both evidence surfaces (containers, + routes) are generation-attributable. Attempt-scoped NAMES remain + F04/A09 (the `_replaced` convention still bridges same-hash attempts; + renaming containers is the breaking change this slice deliberately + does not make). 10. **C01-10 — The predecessor snapshot is in-memory only.** `deploy.go:464-473` snapshots predecessors before candidates start, @@ -305,8 +338,12 @@ Two lock layers, never nested across hosts: guard, and release removes the lock only when its info still names the releaser. Acquisition order on the commit command is therefore APP GUARD THEN CADDY GUARD — app lock first (long-held), shared - commit lock second (brief); never the inverse, and a holder that - loses either fence has its commit refused in-shell. + commit lock second (brief); never the inverse, and a holder that + loses either fence has its commit refused in-shell. Since the + 2026-09-24 generation-identity slice, the exact-block route CAS + (and, on the state commit, the `.generation` sidecar CAS) chains + AFTER both guards in the same command — identity fences last, so + they are evaluated against the world the guards just proved current. **Never hold one host's app lock while waiting on another host's.** The current code complies: locks are acquired inside each host's @@ -378,9 +415,17 @@ landed 2026-09-23** (replacement-owner reconciliation on acquisition; the guarded Caddyfile commit; see AUDIT_OPEN) and **C01-3 landed 2026-09-23** (owner-tagged fenced shared Caddy lock + the conflicting-route-evidence producer/consumer; see AUDIT_OPEN). The -locking-protocol redesign (C01-1/2/3) is closed. Remaining findings: -C01-8 (same-version `_replaced` MANUAL — -deliberate A08 containment until F04 generation identities exist), and -C01-9 (attempt-scoped candidate identities — F04/A09). The A12/T05 -remainder of C01-7 (rollback's restoreRollbackRoute + the exact-block -compare-and-swap restore) stays with its register item. +locking-protocol redesign (C01-1/2/3) is closed. **The generation-identity +sub-slices of C01-8/C01-9 landed 2026-09-24** (teploy.generation labels + +the `.generation` sidecar + rollback's explicit fromGeneration fencing + +exact-block route CAS; see AUDIT_OPEN's C01 generation-identity slice) — +closing the A12/T05 remainder of C01-7 with it: rollback's +`restoreRollbackRoute` now renders from the F14 record (deploy's +`restoreRouteFromReceipt` shape) with reconstruct-from-inspection only as +the announced legacy fallback, and every route switch/restore runs under +the exact-block compare-and-swap (`caddy.Client.WithRouteCAS`), whose +refusal names both generations. What remains under the register: +C01-8's INSPECT→adopt continuation for a running `_replaced` (deliberate +A08 containment until the F04-keyed recovery-owner decision), C01-9's +attempt-scoped container NAMES (F04/A09 — the breaking rename), and F04's +RouteSwitch handoff boundary (T22/A16). diff --git a/docs/failure-and-recovery.md b/docs/failure-and-recovery.md new file mode 100644 index 0000000..6f1409d --- /dev/null +++ b/docs/failure-and-recovery.md @@ -0,0 +1,160 @@ +# Failure and recovery + +What failure looks like from the outside (exit codes, the structured error +envelope), what `teploy doctor` diagnoses, what happens when a deploy is +interrupted mid-flight, and the disaster-recovery bundle family. The +behavioral contract behind the crash-recovery machinery is +[C01_RECOVERY_STATE_TABLE.md](C01_RECOVERY_STATE_TABLE.md); this page is +the operator-facing version. + +## Exit codes and the error envelope + +Exit codes are stable and minimal: `0` success, `1` command failure. `2` +is reserved as `teploy drift --exit-code`'s CI signal and is never used +by other commands (`doctor` fails with 1, never 2). + +Under `--json`, a failed command additionally writes one JSON document to +**stderr** (stdout keeps only successful output): + +```json +{"machine_interface":2,"code":"config-invalid","message":"invalid teploy configuration","detail":"invalid teploy.yml: 'domain' is required"} +``` + +The closed code taxonomy: + +| Code | Meaning | Wired today | +|---|---|---| +| `config-invalid` | Config (teploy.yml/TOML/destination/Compose) failed to load or validate, or a deploy request was refused before any effect (admission) | Yes — config errors and deploy admission refusals | +| `conflict` | The request is coherent but the world moved under it (plan/apply drift refusal) | Yes | +| `internal` | Everything else, including not-yet-migrated failure sites | Yes (default) | +| `target-unreachable` | SSH/transport failure | Reserved — currently classifies as `internal` | +| `unsupported` | The verb needs a capability this binary lacks | Reserved | +| `uncertain-outcome` | The effect's fate is unknown pending reconciliation | Reserved | +| `degraded` | Success with a flag — traffic switched but the outcome is not clean | Reserved | + +Reserved means the code exists in the schema and consumers must decode it, +but no error site classifies into it yet; treat an unknown code as +`internal` (forward-safe). Machine consumers: stderr is the error channel, +stdout is data, and the exit code stays the pass/fail signal. + +A note on honest outcomes: a deploy whose traffic switched but whose +predecessor retirement partially failed is recorded in `teploy log` as +`DEGRADED` (with the reason), not as a clean `ok` — filter on status, not +just success. + +## teploy doctor + +Read-only diagnostics; a run that fails checks deploys nothing. Checks: + +| Check | What it verifies | +|---|---| +| `git` | Local git present (warn only — prebuilt-image deploys do not need it) | +| `config` | teploy.yml/TOML or the Compose importer accepts the project (grammar errors arrive verbatim) | +| `ssh` | Key-auth connectivity to the app's server (or `--server`) | +| `docker` | Remote Docker reachable | +| `disk` | Root filesystem headroom (fail < 2 GiB free; warn < 10 GiB or > 85% used — deploys write layers, backups and attempt artifacts to `/`) | +| `registry` | The configured `image:` is reachable; auth failures distinguished from unreachable | +| `caddy` | Caddy admin API (Caddy ingress only; skipped-ok for host/external ingress) | +| `compatibility` | Local teploy version vs the server-side teploy binary when one exists (autodeploy installs one at `/deployments/.bin/teploy`) | +| `repair-debt` | Outstanding release-record repair debt (below) | + +`--json` emits `{machine_interface, checks[{name,result,detail,remediation}], summary}`; +result is `ok|warn|fail` and every check always carries all four keys. +Failing checks each print a `fix:` line in human mode. + +## Interrupted deploys + +A deploy is a sequence of durable steps (attempt artifacts, containers, +readiness receipt, traffic switch, fenced state commit, predecessor +retirement, terminal record). The machinery around an owner that dies +mid-sequence: + +- **Fenced per-app locks.** A new owner that takes over a stale lock does + not assume quiescence: it observes the target's real evidence (running + containers by exact name/label, receipts, authority state) and runs the + recovery decision before its first effect. Safe-to-retry is the only + disposition that proceeds automatically; anything ambiguous refuses + with the observed evidence rather than guessing. +- **Dispositions** (what the decision function can order): RETRY + (idempotent tail — record writes, predecessor retirement), INSPECT + (effects landed without receipts — re-observe, never invent success), + COMPENSATE (uncommitted traffic — undo via the recorded predecessor), + MANUAL (unattributable workloads, unreadable authority, traffic on an + uncommitted generation with the predecessor gone). MANUAL means the + operator reconciles by hand; the refusal message names what was seen. +- **Repair debt.** If the post-commit release-record write fails, a + `/deployments//repair-debt.json` marker is persisted; the next + deploy repairs the record from live containers before its own work, + then clears the marker. `teploy status` and `teploy doctor` report + outstanding debt. +- **Predecessor snapshot persistence.** The predecessor container + identities are journaled with the attempt + (`meta/att/./predecessors.json`) before any new container + starts, so compensation after a crash targets exactly the recorded + containers rather than an inference. + +What this means operationally: after a crashed or canceled deploy, the +command to run first is `teploy doctor`, then `teploy status`. If the +deploy refused with a recovery-evidence error, that is the MANUAL +disposition — read the named evidence, `teploy rollback` if a predecessor +is recorded, and only then redeploy. Never respond to an uncertain deploy +by blind re-deploying; the CLI will refuse where it cannot attribute, and +that refusal is the safety working. + +Fleet note: a multi-server rollout is a sequence of per-host recorded +outcomes, not a global transaction — see the staged-rollout section of the +README for canary and failure-budget behavior. + +## Self-heal (steady state) + +`teploy heal enable` installs a systemd-timer probe that restarts an +unhealthy **web** container in place (bounded attempts/backoff) — for +"container up but failing", not for deploys. `teploy heal status` / +`teploy heal disable` manage it. + +## DR bundles (teploy dr) + +`teploy backup` is the data-only family (volume archives, single +accessories). `teploy dr` bundles the whole application: state, release +records, the applied manifest, secret references (or explicitly opted-in +encrypted material), routing identity, and consistency-labeled data +snapshots. + +```bash +# create (S3 or a plain directory on the server) +teploy dr create --bucket my-bucket # or --dir /srv/dr-bundles +teploy dr create --include-secrets # opt in: age ciphertexts, resolved .env, + # accessory credentials (never default) +teploy dr create --include-age-key # + the key itself, so a fresh host can + # decrypt (requires --include-secrets) +teploy dr create --stop-app # quiesced volume snapshots (app restarted after) + +teploy dr list --dir ... # bundle ids +teploy dr show --dir ... # manifest: snapshots + consistency, + # secrets mode, routing, recovery plan + +teploy dr restore --dir ... # isolated restore into /var/tmp/teploy-dr, + # boots scratch engines + app container, + # validates, writes an RPO/RTO receipt. + # Nothing under /deployments is touched. +teploy dr cutover # the explicit mutation: promote a + # validated staged restore over the live app + # (originals kept for two-phase recovery) +``` + +Verified against a scratch host (this branch): `create --dir`, `list`, +`show`, and `restore` — the restore receipt reported staging path, RPO/RTO, +per-check results (`app ... pass image=... running`), and the exact next +command (`teploy dr cutover `). A restore that fails validation +exits non-zero with "nothing live was touched"; missing secret keys fail +before any mutation. After a cutover, run `teploy deploy` to bring the +app container and routing live from the restored state. + +Snapshot consistency is labeled per snapshot: engine dumps are +engine-consistent; raw volume copies are crash-consistent unless the app +was stopped (`--stop-app`) or you assert a volume quiesced by hand +(`--quiesced-volume NAME`). The labels are recorded in the manifest — +read them before trusting a volume snapshot of a writing database. + +For the topology-level version (N+1, state off-box, the dead-server +runbook), see [resilience.md](resilience.md). diff --git a/docs/first-success.md b/docs/first-success.md new file mode 100644 index 0000000..63b7c4e --- /dev/null +++ b/docs/first-success.md @@ -0,0 +1,146 @@ +# First success: install to rollback + +The shortest honest path from nothing to a deployed app you have verified +and can roll back. Every command below was executed against a scratch +Docker host over SSH (a colima VM); where a step needs an environment this +walk-through cannot assume (a public domain, a cloud VPS), it says so. + +## 0. Install the CLI + +```bash +brew install useteploy/tap/teploy # macOS/Linux +# or download a release binary / Scoop on Windows / go install — see README +teploy version +``` + +Verified here with a worktree build (`teploy version` prints the embedded +version; a source build prints `dev`). + +## 1. Have a server + +Any Linux server you can SSH into with key auth, with Docker installable. +`teploy setup ` provisions it: Docker, Caddy, firewall, and by +default host audit hardening (auditd + sudo session recording; skip with +`--no-harden`). + +```bash +teploy setup 203.0.113.10 +``` + +Gated here: `setup` against a fresh public VPS was not re-run for this +walk-through (no spare public host); the command's behavior is covered by +its own tests, and the fixture below used an already-provisioned Docker +host. If your host already runs Docker, a deploy works without `setup` +for `ingress: host` — Caddy is only required for domain routing. + +Register the server so commands can name it (writes +`~/.teploy/servers.yml`): + +```bash +teploy server add box1 203.0.113.10 --user root +``` + +Non-standard SSH port? Use `host:port` — `teploy server add box1 +203.0.113.10:2222 --user root` and `server: box1` in `teploy.yml`. + +## 2. Create the app + +A project directory with a `Dockerfile` listening on one port: + +```dockerfile +FROM nginx:1.27-alpine +COPY index.html /usr/share/nginx/html/index.html +``` + +Either run `teploy init` (interactive: app name, domain or raw port, +server) or write the three-line config yourself: + +```yaml +# teploy.yml +app: demo +domain: demo.example.com # or: ingress: host + port: 8080 for a raw port +server: box1 +``` + +Check yourself before deploying: + +```bash +teploy validate # config grammar + server reference +teploy doctor # read-only: git, config, SSH, Docker, disk, registry, + # Caddy, version compatibility, repair debt +``` + +`doctor` never mutates server state and exits 1 (never 2) when any check +fails; every failing check prints a `fix:` line. A `warn` (e.g. missing +local git) does not fail the run. + +## 3. Deploy + +```bash +teploy deploy +``` + +What you should see (abridged, from the verified run): + +``` +Built image: demo-build- +Deploying demo (version )... +Publishing on 0.0.0.0:80 (host ingress)... +Starting container demo-web- (port 80)... + Readiness: auto — HTTP then TCP fallback (compat, 30s deadline) + Health check passed +Deployed demo version in 1.237s +Receipt: image sha256:..., revision , config manifest ... +``` + +The deploy output names the target, the version, the readiness mode and +deadline, and ends with a receipt (image digest, revision, config +manifest) — the identity of exactly what landed. + +Failure modes at this step, with the product's remedy: + +| Symptom | Meaning | Remedy | +|---|---|---| +| `could not determine version from git ... (use --version flag)` | Project is not a git repository | `git init && git commit`, or `teploy deploy --version v1` | +| `rsync failed: exit status 127` | Server lacks `rsync` | Install it (`apt-get install rsync`); `teploy setup` includes it | +| `authentication failed for root@...; try --user ...` | SSH user/key mismatch | `--user`, `--key`, or `TEPLOY_USER`/`TEPLOY_SSH_KEY` | +| `'domain' is required` | No domain and no `ingress: host` | Set `domain:`, or `ingress: host` + `port:` | +| Health check never passes | Readiness gate refuses to switch traffic | The deploy fails and the previous version keeps serving. Fix the app's health endpoint or set `health: {mode: tcp}` — see README "Config" | +| Doctor fails on `ssh`/`docker`/`disk` | Target not ready | Follow the per-check `fix:` line; disk fails below 2 GiB free | + +## 4. Verify + +```bash +teploy status # containers + version +teploy health # run the readiness probe on the live app +teploy logs --tail 20 # stream logs (Ctrl-C to exit — it follows) +teploy log # deploy history: deploys, rollbacks, failures +``` + +And from any machine that can reach the app: `curl http(s):///`. + +## 5. Change something, deploy again, roll back + +Commit a change and deploy; `teploy status` now shows current and previous +hashes. To revert: + +```bash +teploy rollback +``` + +On Caddy ingress this switches traffic back to the still-known predecessor +version and health-gates it. On `ingress: host` the prior container was +removed at deploy, so rollback redeploys the previous version and +health-gates it (seconds, not an instant switch) — verified on the +fixture: `Rolled back demo to version in 369ms`. + +## 6. Next steps + +- Secrets and env: `teploy env set`, `teploy secret set` (encrypted at + rest), SOPS/age `env_files:`. +- Backups: `teploy backup create --bucket ...`, verified restores via + `teploy accessory verify-backup`. +- Whole-app disaster recovery bundles: [failure-and-recovery.md](failure-and-recovery.md). +- CI: [ci-deploy.md](ci-deploy.md). +- Fleet and surviving server loss: [resilience.md](resilience.md). diff --git a/docs/migration.md b/docs/migration.md new file mode 100644 index 0000000..5e32b49 --- /dev/null +++ b/docs/migration.md @@ -0,0 +1,96 @@ +# Migrating to teploy (from Dokploy / Coolify / raw Compose) + +Both platforms center on Compose files, and teploy can import that subset +of Compose it can preserve exactly. The contract (programme workstream +C05): **every supplied field is preserved, explicitly translated, or +rejected with a named, actionable error — never silently dropped.** The +classification table below is a summary; the executable authority is the +importer and its conformance tests (`internal/config/compose.go`, +`TestLoadCompose_FieldClassificationInventory` in `compose_test.go`). + +## What a migration looks like + +1. `teploy setup ` on a fresh host (Docker + Caddy + firewall). +2. Put your `docker-compose.yml` in the project directory (or run + `teploy init`, which offers to import it). +3. `teploy validate` — this either imports or refuses, naming every + problem field. +4. Fix refusals (below), set `server:` and `domain:`, deploy. +5. Cut traffic over (DNS or proxy) when the app is verified — both stacks + can run side by side; nothing forces a destructive cutover. + +## What converts + +| Compose | Becomes | +|---|---| +| The single non-accessory service with ports | The app (`web` process); its container port from short-form `"host:container"` (host side deliberately not preserved — teploy allocates host ports and routes via Caddy) | +| `build:` on the web service | Build context for the image | +| `image:` on the web service | `image:` (no build) | +| Same-`build:` second service | A `processes:` worker running the same image with the service's `command:` | +| `postgres/redis/mysql/mariadb/mongo/clickhouse/meilisearch/elasticsearch/memcached/rabbitmq/nats` images | Accessories (known default ports) | +| Any other standalone-image service | An accessory | +| `environment:` (map or list) | `env:` verbatim (deploy-time `${VAR}` expansion applies) | +| `volumes:` on web or services | Named volumes (managed under `/deployments//volumes/`) or host binds (source starting `/`) | +| Web `healthcheck.test` exec-list HTTP probe | `health:` path/interval | +| `healthcheck: {disable: true}` / `test: ["NONE"]` | `healthcheck..disable: true` | +| `restart: always` / `unless-stopped` | Tolerated (teploy's own policies match) | +| `deploy:` with only `replicas: 1` / `mode: replicated` | Tolerated as the no-op default | +| `networks: [default]` (or absent) | Tolerated as the implicit default | +| Services under non-default `profiles:` | Skipped, deliberately — `docker compose up` without `--profile` would not deploy them either | +| `depends_on` | Parsed, not translated: teploy already starts every accessory before any app container. Readiness conditions (`service_healthy`) are **not** waited for | + +## What refuses (named errors, before any effect) + +- A second service with a **different build context**: "unsupported + independent build in compose import: `` (build `""`) while + `"web"` builds from ... — teploy runs one image per app and cannot + preserve a separately built service; use the same build context as the + app, a prebuilt image, or write teploy.yml". +- **Multiple port-publishing non-accessory services**: "ambiguous compose + import: multiple non-accessory services publish ports (...)". +- Per-field refusals, one named error each: non-default `networks:`, + `env_file`, `secrets`, `configs`, `extends`, non-default `deploy:`, + `container_name`, `hostname`, `working_dir`, `entrypoint`, + `privileged: true`, non-empty `cap_add`, other `restart:` policies, + long-form/ranged/multi-port publishes, non-web healthchecks without a + home. Each error names the field, why it cannot be preserved, and the + alternative ("remove it or write teploy.yml"). + +What multi-image stacks should do instead: model each independently built +service as its own teploy app on the same server (they share the teploy +network and can address each other by app name), or prebuild images and +run them as accessories. + +## Concept mapping from Dokploy/Coolify + +| There | Here | +|---|---| +| Project / Application | One directory with `teploy.yml` (or an imported compose file) | +| Environment variables UI | `teploy env set KEY=value` (stored server-side) | +| Secrets | `teploy secret set` (age-encrypted at rest) or SOPS/age `env_files:` | +| Traefik + Let's Encrypt | Caddy, written and reloaded by teploy (ACME default; custom `tls:` supported) | +| Domains / routes | `domain:` per app; preview subdomains via `teploy preview` | +| Databases | `accessories:` (managed containers with `--restart always`) | +| Webhook auto-deploy | `teploy autodeploy setup` (+ `autodeploy.paths:` filters for monorepos) | +| Dashboard | [teploy-dash](https://github.com/useteploy/teploy-dash) (optional, read-state + delegate; the CLI stays the source of truth) | +| Backups | `teploy backup` (data-only) and `teploy dr` (whole-app bundles) | + +## Reversible adoption + +Nothing in a teploy migration touches the origin platform: state lives in +`/deployments//` on the host you point at, containers are plain +Docker containers with `teploy.*` labels, and +`teploy remove` retires an app's containers, proxy route, and deploy +state when you want it gone. Rolling back to Dokploy/Coolify means +pointing DNS at the old deployment — run both in parallel until the new +one has proven itself, then decommission the old. + +## After importing + +- `teploy doctor` — the `config` check re-runs the importer; `disk`, + `docker`, `caddy` validate the target. +- `teploy plan` — read-only preview of the first deploy's container and + routing changes. +- Data: migrate database contents with a dump/restore into the new + accessory (`teploy accessory backup/restore` on the source platform's + volume export), then cut over. diff --git a/docs/supported-workloads.md b/docs/supported-workloads.md new file mode 100644 index 0000000..d3bb1c0 --- /dev/null +++ b/docs/supported-workloads.md @@ -0,0 +1,61 @@ +# Supported workloads + +What `teploy deploy` accepts today, and what it refuses — with the refusal +behavior named. If something you need is in the refused column, the answer +is "not yet", not "silently degraded": every refusal below is a named, +actionable error issued before any container effect unless stated +otherwise. + +## Matrix + +| Workload | Status | Notes | +|---|---|---| +| Single-image container app (Dockerfile) | Supported | Default path. Build on the server (`rsync` context + `docker build`), or `build_local: true`. Nixpacks used when no Dockerfile is present (requires Nixpacks on the server or locally). | +| Single-image app from a registry (`image:`) | Supported | Private registries via `teploy registry login`. | +| Multi-process from one image (`processes:` web/worker/cron) | Supported | All processes run from the same image; `healthcheck:` per-process overrides for inherited probes. | +| Static site (`type: static`) | Supported | rsync to the server, served by the managed Caddy. Requires Caddy ingress; `ingress: host`/`external` and `tls:` are rejected for static. | +| Compose file as an importer (subset) | Supported (subset) | `docker-compose.yml` in the project dir imports when no `teploy.yml` exists. Every supplied field is preserved, translated, or rejected with a named error — see [migration.md](migration.md) for the classification summary. | +| Templates (`teploy template install`) | Supported | One-command deploys of reviewed community apps (Postgres+Adminer, WordPress, Immich, ...). Catalog: `teploy template list`. | +| Preview environments (`teploy preview`) | Supported | Branch slugs on `preview-.` against a pre-built image (`teploy build`). Requires Teploy-managed Caddy. | +| Accessories (Postgres, Redis, MySQL, Mariaadb, Mongo, ClickHouse, Meilisearch, Elasticsearch, Memcached, RabbitMQ, NATS, or any standalone image) | Supported | Managed alongside the app with `--restart always`, volumes, ports, env. | +| Multi-image stacks (several independently built services) | Refused | One image per app is the model. A Compose file whose service builds from a different context than the web service refuses at config load: `unsupported independent build in compose import: ... — teploy runs one image per app and cannot preserve a separately built service`. Model as separate teploy apps, or prebuilt images. | +| Multiple web candidates in one Compose file | Refused | `ambiguous compose import: multiple non-accessory services publish ports (...)`. Remove ports from non-app services or write `teploy.yml`. | +| Host port ranges / long-form Compose ports | Refused | Named error from the port grammar; short-form `"host:container"` and non-TCP publishes are what convert. | +| Compose `networks:` (non-default), `env_file`, `secrets`, `configs`, `extends`, `deploy:` (non-default), `container_name`, `hostname`, `working_dir`, `entrypoint`, `privileged`, `cap_add`, `restart:` (other than always/unless-stopped) | Refused per field | One named, actionable error per supplied field whose meaning would be lost (`unsupported compose fields: ...`). Full table: [migration.md](migration.md). | +| Kubernetes-style scheduling, cross-host replica scheduling | Not supported, by design | No scheduler. Fleet semantics are per-host deploys + Caddy LB. See [resilience.md](resilience.md). | + +## Ingress modes and their guarantees + +| Mode | Deploy strategy | Downtime | Rollback | +|---|---|---|---| +| `ingress: caddy` (default) | Blue/green: new container starts, passes the readiness gate, then traffic switches; predecessor stops | None during the switch (drain via `drain_seconds`) | `teploy rollback` switches back to the predecessor version | +| `ingress: external` | Same container lifecycle; Teploy never touches the proxy (the container joins the teploy network with its app-name alias) | Whatever your proxy's cutover does | `teploy rollback` (container-level) | +| `ingress: host` | Recreate: stop old, start new, on a fixed host port | Seconds per deploy (a fixed port cannot be blue/green) | `teploy rollback` redeploys the previous version and health-gates it (there is no instant container switch — the prior container was removed) | + +Static apps (`type: static`) are Caddy-served and use release directories, +not containers; `keep_releases` prunes them. + +## Operational limits + +Declared, not aspirational: + +- **One app = one image.** Every process runs from the same image; there is + no per-service build. This is the boundary the Compose importer enforces. +- **Single-writer deploys per app per host.** Deploys serialize behind a + fenced per-app lock; a crashed holder's effects are reconciled by the next + owner (see [failure-and-recovery.md](failure-and-recovery.md)). +- **Fixed host ports are single-replica.** `ingress: host` and `publish:` + entries cannot be load-balanced across containers on one host; + replicas require Caddy ingress. +- **Readiness is a gate, not a liveness system.** `health:` defines what + "healthy" means before traffic switches; steady-state restart-in-place is + `teploy heal enable` (bounded, systemd-timer driven). +- **The server needs `rsync` on PATH** for source sync (present on normal + Debian/Ubuntu images; `teploy setup` installs it). Missing `rsync` fails + the sync step with `rsync failed: exit status 127` before anything lands. +- **No scheduler.** Surviving server loss is a topology concern — + [resilience.md](resilience.md) is the supported pattern and runbook. +- **Version identity comes from git** (short hash) or `--version`. A project + that is not a git repository must deploy with `--version` — the failure + message says so: `could not determine version from git ... (use --version + flag)`. diff --git a/examples/quickstart/README.md b/examples/quickstart/README.md new file mode 100644 index 0000000..7a624c0 --- /dev/null +++ b/examples/quickstart/README.md @@ -0,0 +1,32 @@ +# teploy quickstart (executable) + +The C09 acceptance line: *a new user deploys the maintained fixture from +documentation without undocumented repair.* This directory is that +documentation, and it runs. + +``` +make quickstart # from the repo root +``` + +`run.sh` deploys `app/` (the maintained fixture: a busybox httpd app with +a `teploy.yml`) to the **local colima VM's own SSH endpoint**, then: + +1. builds the CLI from this checkout, +2. bootstraps the target once (`/deployments` directory; the only + target-side setup, via the VM's passwordless sudo), +3. deploys version `qs1` — build-on-target, health-gated start, host + ingress on `127.0.0.1:18080`, +4. verifies the app answers with the v1 content, +5. redeploys as `qs2` with changed content and verifies the switch, +6. shows `teploy status`, then removes everything it created (app, + containers, images, its own known_hosts lines). + +Requirements: docker CLI, a **running** colima VM (the script never +starts one — `colima start` yourself), ssh/ssh-keyscan/curl/python3. +When any is missing the script prints `SKIP: ...` and exits 0 — a +skipped quickstart is not a failed one. + +No remote server, no domain, no DNS: host ingress publishes a plain port +on the target, which is exactly what makes the loop runnable on a +laptop. The deploy path exercised is the real one — same config loader, +same build/health/commit machinery as production. diff --git a/examples/quickstart/app/Dockerfile b/examples/quickstart/app/Dockerfile new file mode 100644 index 0000000..53ee879 --- /dev/null +++ b/examples/quickstart/app/Dockerfile @@ -0,0 +1,5 @@ +FROM busybox:1.37 +COPY index.html /www/index.html +# teploy injects PORT (the published port) as an env var; listen there so +# the health gate and the published port see the same listener. +CMD ["sh", "-c", "httpd -f -p ${PORT:-80} -h /www"] diff --git a/examples/quickstart/app/index.html b/examples/quickstart/app/index.html new file mode 100644 index 0000000..ea153cc --- /dev/null +++ b/examples/quickstart/app/index.html @@ -0,0 +1,3 @@ + +teploy quickstart +

teploy quickstart v1

diff --git a/examples/quickstart/app/teploy.yml b/examples/quickstart/app/teploy.yml new file mode 100644 index 0000000..107da89 --- /dev/null +++ b/examples/quickstart/app/teploy.yml @@ -0,0 +1,9 @@ +# The maintained quickstart fixture app (C09). Deployed by +# examples/quickstart/run.sh against a local docker target; everything +# teploy needs travels in this directory — no undocumented repair. +app: quickstart +server: colima-vm +ingress: host +port: 18080 +health: + mode: tcp diff --git a/examples/quickstart/run.sh b/examples/quickstart/run.sh new file mode 100755 index 0000000..3c52cb1 --- /dev/null +++ b/examples/quickstart/run.sh @@ -0,0 +1,125 @@ +#!/usr/bin/env bash +# Executable quickstart (C09): deploy the maintained fixture app from +# examples/quickstart/app against a LOCAL docker target — the colima VM's +# own SSH endpoint — with no undocumented repair, then verify the app +# answers, redeploy a second version, and clean up after itself. +# +# Honest gating: this script needs (a) the docker CLI, (b) a running +# colima VM (it will NOT start one), and (c) ssh/ssh-keyscan/curl on +# PATH. When any is missing it prints SKIP and exits 0 — a skipped +# quickstart must not read as a failed one. +# +# What it proves: a new user path from `git clean` checkout to a +# responding application — config in teploy.yml, build on the target, +# health-gated start, published port, idempotent redeploy, status, +# removal. Exit 0 only if every step held. +set -euo pipefail + +REPO_ROOT="$(cd "$(dirname "$0")/../.." && pwd)" +APP_DIR="$REPO_ROOT/examples/quickstart/app" +WORK_DIR="$(mktemp -d "${TMPDIR:-/tmp}/teploy-quickstart.XXXXXX")" +TEPLOY_BIN="" +SSH_HOST=""; SSH_PORT=""; SSH_USER=""; SSH_KEY="" +CONTAINER_KEY_FILE="" # known_hosts lines added by this run +APP_PORT=18080 + +log() { printf '==> %s\n' "$*"; } +skip() { printf 'SKIP: %s\n' "$*"; exit 0; } + +cleanup() { + local code=$? + set +e + if [ -n "$TEPLOY_BIN" ] && [ -n "$SSH_HOST" ]; then + "$TEPLOY_BIN" remove --purge --yes --app quickstart --host "$SSH_HOST" \ + --user "$SSH_USER" --key "$SSH_KEY" >/dev/null 2>&1 + ssh -i "$SSH_KEY" -p "$SSH_PORT" -o BatchMode=yes "$SSH_USER@$SSH_HOST" \ + 'docker rm -f quickstart-web-qs1 quickstart-web-qs2 >/dev/null 2>&1; docker rmi quickstart-build-qs1 quickstart-build-qs2 >/dev/null 2>&1; sudo rm -rf /deployments/quickstart' 2>/dev/null + fi + if [ -n "$CONTAINER_KEY_FILE" ] && [ -f "$HOME/.ssh/known_hosts" ]; then + # Remove only the lines this run appended. + python3 - "$CONTAINER_KEY_FILE" <<'PY' +import sys +added = set(open(sys.argv[1]).read().splitlines()) +path = __import__("os").path.expanduser("~/.ssh/known_hosts") +lines = open(path).read().splitlines() +kept = [l for l in lines if l not in added] +open(path, "w").write("\n".join(kept) + ("\n" if kept else "")) +PY + fi + rm -rf "$WORK_DIR" + exit $code +} +trap cleanup EXIT + +# --- gates ----------------------------------------------------------------- +for bin in docker ssh ssh-keyscan ssh-keygen curl python3; do + command -v "$bin" >/dev/null 2>&1 || skip "$bin not found on PATH" +done +command -v colima >/dev/null 2>&1 || skip "colima not found (this quickstart targets a local colima VM)" +colima status >/dev/null 2>&1 || skip "colima VM not running (start it with: colima start), then re-run" + +# --- resolve the colima VM's SSH endpoint from colima's own config --------- +COLIMA_HOME_DIR="${COLIMA_HOME:-$HOME/.colima}" +SSH_CFG="" +for candidate in "$COLIMA_HOME_DIR/ssh_config" "$COLIMA_HOME_DIR/default/ssh_config" "$COLIMA_HOME_DIR/_lima/colima/ssh_config"; do + [ -f "$candidate" ] && SSH_CFG="$candidate" && break +done +[ -n "$SSH_CFG" ] || skip "no colima ssh_config under $COLIMA_HOME_DIR" +SSH_HOST="$(awk '$1=="Hostname"{print $2}' "$SSH_CFG")" +SSH_PORT="$(awk '$1=="Port"{print $2}' "$SSH_CFG")" +SSH_USER="$(awk '$1=="User"{print $2}' "$SSH_CFG")" +SSH_KEY="$(awk '$1=="IdentityFile"{print $2}' "$SSH_CFG" | tr -d '"')" +for v in "$SSH_HOST" "$SSH_PORT" "$SSH_USER" "$SSH_KEY"; do + [ -n "$v" ] || skip "could not parse endpoint from $SSH_CFG" +done +[ -f "$SSH_KEY" ] || skip "colima SSH key missing ($SSH_KEY) — restart colima to regenerate" +log "target: $SSH_USER@$SSH_HOST:$SSH_PORT (colima VM)" + +ssh -i "$SSH_KEY" -p "$SSH_PORT" -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null \ + -o BatchMode=yes "$SSH_USER@$SSH_HOST" true 2>/dev/null \ + || skip "cannot SSH to the colima VM ($SSH_USER@$SSH_HOST:$SSH_PORT)" + +# --- host key + /deployments bootstrap (the only target-side setup) -------- +mkdir -p "$HOME/.ssh" +touch "$HOME/.ssh/known_hosts" +CONTAINER_KEY_FILE="$WORK_DIR/added-host-keys" +ssh-keyscan -p "$SSH_PORT" "$SSH_HOST" >"$CONTAINER_KEY_FILE" 2>/dev/null +grep -q . "$CONTAINER_KEY_FILE" || skip "ssh-keyscan produced no keys for $SSH_HOST:$SSH_PORT" +cat "$CONTAINER_KEY_FILE" >>"$HOME/.ssh/known_hosts" + +ssh -i "$SSH_KEY" -p "$SSH_PORT" -o BatchMode=yes "$SSH_USER@$SSH_HOST" \ + 'sudo mkdir -p /deployments && sudo chown "$(id -un)" /deployments' 2>/dev/null \ + || skip "cannot bootstrap /deployments on the VM (needs passwordless sudo)" + +# --- build this checkout's CLI ---------------------------------------------- +log "building teploy from this checkout" +TEPLOY_BIN="$WORK_DIR/teploy" +(cd "$REPO_ROOT" && go build -o "$TEPLOY_BIN" ./cmd/teploy) + +TEPLOY=("$TEPLOY_BIN" --host "$SSH_HOST:$SSH_PORT" --user "$SSH_USER" --key "$SSH_KEY") + +# --- deploy v1 -------------------------------------------------------------- +cp -R "$APP_DIR/." "$WORK_DIR/app/" +log "deploying quickstart v1 (build-on-target, health-gated, host ingress :$APP_PORT)" +(cd "$WORK_DIR/app" && "${TEPLOY[@]}" deploy --version qs1) 2>&1 | sed 's/^/ /' + +BODY="$(curl -fsS -m 10 "http://127.0.0.1:$APP_PORT/")" +grep -q "quickstart v1" <<<"$BODY" || { echo "FAIL: v1 content not served: $BODY" >&2; exit 1; } +log "verified: http://127.0.0.1:$APP_PORT/ serves the v1 fixture" + +# --- redeploy v2 (recreate path) -------------------------------------------- +log "redeploying as qs2 with changed content" +sed 's/quickstart v1/quickstart v2/' "$WORK_DIR/app/index.html" >"$WORK_DIR/app/index.html.tmp" +mv "$WORK_DIR/app/index.html.tmp" "$WORK_DIR/app/index.html" +(cd "$WORK_DIR/app" && "${TEPLOY[@]}" deploy --version qs2) 2>&1 | sed 's/^/ /' + +BODY="$(curl -fsS -m 10 "http://127.0.0.1:$APP_PORT/")" +grep -q "quickstart v2" <<<"$BODY" || { echo "FAIL: v2 content not served after redeploy: $BODY" >&2; exit 1; } +log "verified: redeploy switched the served content to v2" + +# --- status ------------------------------------------------------------------ +log "teploy status sees the deployment" +"${TEPLOY[@]}" status --app quickstart 2>&1 | sed 's/^/ /' + +log "quickstart complete: deployed, verified, redeployed, verified again" +log "cleanup follows (teploy remove --purge, known_hosts lines, temp dir)" diff --git a/internal/caddy/caddy.go b/internal/caddy/caddy.go index 626e336..98a1e92 100644 --- a/internal/caddy/caddy.go +++ b/internal/caddy/caddy.go @@ -3,6 +3,7 @@ package caddy import ( "context" "crypto/rand" + "crypto/sha256" "encoding/hex" "encoding/json" "fmt" @@ -10,6 +11,7 @@ import ( "net/url" "regexp" "sort" + "strconv" "strings" "time" @@ -203,6 +205,12 @@ type Client struct { // new owner's window. Empty = unguarded commits (the receiver shared // by callers that hold no app fence). Set through WithCommitGuard. commitGuardPrefix string + // stampGeneration, when > 0, stamps the next managed block this client + // upserts with the deployment generation it serves (WithGeneration). + stampGeneration uint64 + // routeCAS, when set, applies an exact-block compare-and-swap to the + // next managed-block commit (WithRouteCAS) — A12/T05. + routeCAS *routeExpectation } // NewClient creates a Caddy client backed by the given SSH executor. @@ -223,6 +231,57 @@ func (c *Client) WithCommitGuard(guardPrefix string) *Client { return &cp } +// generationStampFmt is the managed block's generation identity line +// (C01-8/9): the FIRST line inside the marker region, naming the +// deployment generation the block serves. A comment, so caddy ignores it; +// inside the markers, so every teploy rewrite replaces it and it can never +// leak outside the managed region. Blocks written before this convention +// carry no stamp (RegionGeneration reports ok=false) — the exact-block CAS +// compares whole-region hashes, so unstamped blocks are still protected +// byte-exactly; the stamp is what lets a refusal NAME both generations. +const generationStampFmt = "# TEPLOY GENERATION %d" + +// WithGeneration returns a client whose next managed-block upsert stamps +// the block with the given generation (A12/T05 + C01-8/9): the block the +// operation writes names the generation it is creating, making the route +// attributable by evidence the same way containers are (teploy.generation +// labels) and giving the exact-block CAS its refusal evidence. +func (c *Client) WithGeneration(generation uint64) *Client { + cp := *c + cp.stampGeneration = generation + return &cp +} + +// routeExpectation is an exact-block compare-and-swap (A12/T05): the set of +// region hashes the app's managed block may hold at commit time for the +// edit to proceed, plus the generation this operation prepared against +// (refusal evidence). +type routeExpectation struct { + app string + acceptableRegionHashes []string + expectedGeneration uint64 +} + +// WithRouteCAS returns a client whose next managed-block upsert commits +// under an exact-block compare-and-swap: the app's managed region on the +// live Caddyfile must hash (ManagedRegionHash) to one of the given +// acceptable hashes — the precise predecessor block this operation +// resolved — or the commit is refused (ErrRouteCAS) with evidence naming +// both generations. Concurrent writers cannot interleave: a successor's +// block swap between resolution and commit changes the region hash and +// fences the stale edit out, independently of lock holdership (the +// delayed-SSH-effect window guards cannot close). The expectation applies +// to the NEXT applyManagedBlock-derived edit on this client instance. +func (c *Client) WithRouteCAS(app string, acceptableRegionHashes []string, expectedGeneration uint64) *Client { + cp := *c + cp.routeCAS = &routeExpectation{ + app: app, + acceptableRegionHashes: acceptableRegionHashes, + expectedGeneration: expectedGeneration, + } + return &cp +} + // SetRoute adds or updates a reverse proxy route for the given app, serving the // (comma-separated) domain to the given upstream container:port. // @@ -469,6 +528,12 @@ func (c *Client) RemoveMaintenance(ctx context.Context, app string) error { // (refusing to clean up one's own partial effects is how a fencing design // strands an app). func (c *Client) mutate(ctx context.Context, transform func(prev string) (string, error)) error { + return c.mutateWithCAS(ctx, nil, transform) +} + +// mutateWithCAS is mutate with an optional exact-block compare-and-swap on +// the app's managed region, evaluated inside the COMMIT command (A12/T05). +func (c *Client) mutateWithCAS(ctx context.Context, cas *routeExpectation, transform func(prev string) (string, error)) error { lockOwner, err := c.acquireLock(ctx) if err != nil { return err @@ -504,9 +569,23 @@ func (c *Client) mutate(ctx context.Context, transform func(prev string) (string // C01-2/C01-3 composition: the APP fence (when the caller holds one) // and the CADDY-LOCK fence (this mutation's own lock, owner-tagged) // both precede the rename, so a broken app holder's route edit AND a - // stale-broken editor's late write are refused in-shell. A refused - // commit propagates the fence error; the staged file is cleaned up - // here (bounded). + // stale-broken editor's late write are refused in-shell. The exact- + // block CAS (when the caller resolved one) chains AFTER the guards and + // BEFORE the rename: the region being replaced must still be the one + // this operation prepared against, or the edit is refused with both + // generations named (ErrRouteCAS). A refused commit propagates the + // fence/CAS error; the staged file is cleaned up here (bounded). + casPrefix := "" + if cas != nil { + hashes := make([]string, len(cas.acceptableRegionHashes)) + for i, h := range cas.acceptableRegionHashes { + hashes[i] = h + } + if len(hashes) == 0 { + hashes = []string{ManagedRegionHash("")} + } + casPrefix = routeCASFragment(cas.app, hashes, cas.expectedGeneration) + } tmp, err := c.stageCaddyfile(ctx, updated) if err != nil { return err @@ -519,7 +598,10 @@ func (c *Client) mutate(ctx context.Context, transform func(prev string) (string cancel() } }() - if err := c.commitCaddyfile(ctx, tmp, c.commitGuardPrefix+caddyGuardFragment(lockOwner)); err != nil { + if err := c.commitCaddyfile(ctx, tmp, c.commitGuardPrefix+caddyGuardFragment(lockOwner)+casPrefix); err != nil { + if cas != nil && routeCASRefused(err) { + return c.routeCASError(ctx, cas, err) + } return err } committed = true @@ -620,8 +702,22 @@ func (c *Client) adaptCheck(ctx context.Context, content string) error { // The app's persisted webhook fragment (webhook.go) is re-applied to the // rendered block, so deploys and rollbacks can no longer erase the webhook // route the way they erased the old runtime-API injection (audit T26). +// +// C01-8/9 + A12/T05: when the client carries a generation (WithGeneration) +// the block is stamped with it, and when it carries an exact-block +// expectation (WithRouteCAS) the COMMIT runs under the compare-and-swap — +// both consumed by THIS edit. An expectation naming a different app is a +// caller bug and refuses loudly rather than silently skipping the CAS. func (c *Client) applyManagedBlock(ctx context.Context, app string, hosts []string, block string) error { - return c.mutate(ctx, func(prev string) (string, error) { + if c.routeCAS != nil && c.routeCAS.app != app { + return fmt.Errorf("route CAS expectation names app %s but the edit targets %s — refusing to edit without the compare-and-swap", c.routeCAS.app, app) + } + if block != "" && c.stampGeneration > 0 { + block = fmt.Sprintf(generationStampFmt, c.stampGeneration) + "\n" + block + } + cas := c.routeCAS + c.routeCAS = nil // consume-once: one expectation guards one edit + return c.mutateWithCAS(ctx, cas, func(prev string) (string, error) { updated, err := renderUpdated(prev, app, hosts, block) if err != nil { return "", err @@ -682,12 +778,14 @@ func renderUpdated(prev, app string, hosts []string, block string) (string, erro // managedBlockHosts returns the site-address hosts of the app's managed // block (its first non-empty line), or nil when the app has no managed // block. Used to prove whether a legacy lb- block belongs to this app -// rather than to a distinct application of the same name (TCL-23). +// rather than to a distinct application of the same name (TCL-23). The +// generation stamp line (C01-8/9) is a comment inside the region and is +// skipped — the address line is the first non-comment content line. func managedBlockHosts(content, app string) []string { block := extractCaddyfileBlock(content, fmt.Sprintf(markerBeginFmt, app), fmt.Sprintf(markerEndFmt, app)) for _, line := range strings.Split(block, "\n") { line = strings.TrimSpace(line) - if line == "" { + if line == "" || strings.HasPrefix(line, "# TEPLOY GENERATION ") { continue } addr := strings.TrimSpace(strings.TrimSuffix(line, "{")) @@ -779,6 +877,37 @@ func fenceRefused(err error) bool { strings.Contains(msg, "exit status 75") } +// routeCASRefused reports whether a commit-command error is the exact-block +// CAS refusing the edit (marker on stderr or exit status 76). +func routeCASRefused(err error) bool { + if err == nil { + return false + } + msg := err.Error() + return strings.Contains(msg, routeCASMarker) || + strings.Contains(msg, "status 76") || + strings.Contains(msg, "exit status 76") +} + +// routeCASError turns a raw CAS refusal into the evidence-carrying +// ErrRouteCAS: the live region is re-read (safe — nothing landed) to name +// the found generation and its upstreams alongside the expected one. +func (c *Client) routeCASError(ctx context.Context, cas *routeExpectation, cause error) error { + e := &ErrRouteCAS{ + App: cas.app, + ExpectedGeneration: cas.expectedGeneration, + } + if len(cas.acceptableRegionHashes) > 0 { + e.ExpectedRegionHash = cas.acceptableRegionHashes[0] + } + if region, present, rerr := c.ReadManagedBlock(ctx, cas.app); rerr == nil && present { + e.FoundRegionHash = ManagedRegionHash(region) + e.FoundGeneration, e.FoundStamped = RegionGeneration(region) + e.FoundUpstreams = RegionUpstreams(region) + } + return fmt.Errorf("%w: %v", e, cause) +} + func (c *Client) reload(ctx context.Context) error { if _, err := c.exec.Run(ctx, reloadCmd); err != nil { return err @@ -977,6 +1106,173 @@ const ( markerEndPrefix = "# TEPLOY END " ) +// extractManagedRegion returns the app's managed region MARKER-INCLUSIVE — +// the exact bytes between and including the `# TEPLOY BEGIN ` and +// `# TEPLOY END ` lines — or "" when the app has no region. This is +// the exact-block CAS's identity: what ManagedRegionHash hashes and the +// commit-time sed extraction prints, so a Go-computed expectation and the +// server-side comparison are over identical bytes. An UNTERMINATED region +// (begin marker with no end) yields "" here while the server-side sed +// would print to end-of-file — the hash mismatch refuses the edit, the +// fail-closed direction for a corrupted Caddyfile. +func extractManagedRegion(content, app string) string { + var out []string + active := false + for _, line := range strings.Split(content, "\n") { + if name, ok := beginMarkerApp(line); ok { + if name == app { + active = true + out = append(out, line) + } + continue + } + if name, ok := endMarkerApp(line); ok { + if active && name == app { + out = append(out, line) + return strings.Join(out, "\n") + } + continue + } + if active { + out = append(out, line) + } + } + return "" +} + +// ManagedRegionHash normalizes a managed region for the exact-block CAS. +// The server-side comparison pipes sed's extraction through sha256sum; sed +// terminates every printed line, so a non-empty region hashes as +// region+"\n" while an absent region hashes as the empty input. Both sides +// (Go expectations, shell fragments) MUST use this normalization. +func ManagedRegionHash(region string) string { + payload := []byte(region) + if region != "" { + payload = append(append([]byte(nil), region...), '\n') + } + sum := sha256.Sum256(payload) + return hex.EncodeToString(sum[:]) +} + +// RegionGeneration parses a managed region's generation stamp line +// (generationStampFmt). ok=false means the region carries no stamp — a +// legacy block written before C01-8/9 or a foreign/unstamped region. +func RegionGeneration(region string) (gen uint64, ok bool) { + for _, line := range strings.Split(region, "\n") { + t := strings.TrimSpace(line) + if !strings.HasPrefix(t, "# TEPLOY GENERATION ") { + continue + } + n, err := strconv.ParseUint(strings.TrimSpace(strings.TrimPrefix(t, "# TEPLOY GENERATION ")), 10, 64) + if err != nil { + return 0, false + } + return n, true + } + return 0, false +} + +// RegionUpstreams lists the upstream dials a managed region routes to +// (`name:port` tokens on reverse_proxy lines) — refusal evidence so a CAS +// mismatch can name WHAT the conflicting generation serves, not just its +// number. +func RegionUpstreams(region string) []string { + var upstreams []string + for _, line := range strings.Split(region, "\n") { + code := line + if i := strings.IndexByte(code, '#'); i >= 0 { + code = code[:i] + } + if !strings.Contains(code, "reverse_proxy") { + continue + } + for _, tok := range strings.Fields(code) { + if tok == "reverse_proxy" || tok == "{" { + continue + } + if strings.HasPrefix(tok, "#") { + break + } + if i := strings.IndexByte(tok, ':'); i > 0 { + if _, err := strconv.Atoi(tok[i+1:]); err == nil { + upstreams = append(upstreams, tok) + } + } + } + } + return upstreams +} + +// ReadManagedBlock reads the app's managed region from the live Caddyfile +// (marker-inclusive; "" when absent). Callers resolve the region under +// their app lock at effect-preparation time and hand ManagedRegionHash of +// it to WithRouteCAS — the exact predecessor block identity. +func (c *Client) ReadManagedBlock(ctx context.Context, app string) (region string, present bool, err error) { + data, present, err := readServerFile(ctx, c.exec, caddyfilePath) + if err != nil { + return "", false, fmt.Errorf("reading the Caddyfile for %s's managed block: %w", app, err) + } + if !present { + return "", false, nil + } + region = extractManagedRegion(string(data), app) + if region == "" { + return "", false, nil + } + return region, true, nil +} + +// ErrRouteCAS is the exact-block compare-and-swap refusal (A12/T05): the +// app's managed region on the live Caddyfile is NOT one this operation +// resolved, so a newer writer switched the route in between and this edit +// must not land over it. The error names BOTH generations — the one this +// operation prepared against and the one the live block serves — plus the +// live block's upstreams, the operator's reconciliation evidence. +type ErrRouteCAS struct { + App string + ExpectedGeneration uint64 + FoundGeneration uint64 + FoundStamped bool + FoundUpstreams []string + ExpectedRegionHash string + FoundRegionHash string +} + +func (e *ErrRouteCAS) Error() string { + found := "unstamped (legacy or foreign)" + if e.FoundStamped { + found = fmt.Sprintf("generation %d", e.FoundGeneration) + } + upstreams := "none parsed" + if len(e.FoundUpstreams) > 0 { + upstreams = strings.Join(e.FoundUpstreams, ", ") + } + return fmt.Sprintf( + "route compare-and-swap refused for %s: the managed block changed since this operation resolved it — prepared against generation %d, the live block serves %s (upstreams: %s); another operation switched the route, reconcile before retrying", + e.App, e.ExpectedGeneration, found, upstreams) +} + +// routeCASMarker is the stderr sentinel of the commit-time CAS fragment; +// exit 76 distinguishes it from the fence guards' 75. +const routeCASMarker = "TEPLOY_ROUTE_CAS_MISMATCH" + +// routeCASFragment renders the commit-time exact-block check: extract the +// app's managed region, hash it, and refuse (marker + both generations on +// stderr, exit 76) unless it matches one of the acceptable hashes. Composed +// into the SAME shell command as the commit rename — after the holdership +// guards — so the comparison happens on the target against the bytes being +// replaced, with no client-side read in between. +func routeCASFragment(app string, acceptable []string, expectedGeneration uint64) string { + begin := fmt.Sprintf(markerBeginFmt, app) + end := fmt.Sprintf(markerEndFmt, app) + extract := fmt.Sprintf("sed -n '/^%s$/,/^%s$/p' %s 2>/dev/null", begin, end, ssh.ShellQuote(caddyfilePath)) + cases := strings.Join(acceptable, "|") + return fmt.Sprintf( + `cur=$(%s | sha256sum | cut -d' ' -f1); case "$cur" in %s) ;; *) fg=$(%s | grep -m1 '^# TEPLOY GENERATION ' | cut -d' ' -f4); printf '%s found_generation=%%s expected_generation=%d\n' "${fg:--}" %d >&2; exit 76;; esac; `, + extract, cases, extract, routeCASMarker, expectedGeneration, expectedGeneration, + ) +} + // extractCaddyfileBlock returns the content between the app's exact begin // and end marker lines (markers excluded), or "" if not found. Marker lines // must match COMPLETELY — app "web" no longer matches inside diff --git a/internal/caddy/generation_test.go b/internal/caddy/generation_test.go new file mode 100644 index 0000000..2258e00 --- /dev/null +++ b/internal/caddy/generation_test.go @@ -0,0 +1,221 @@ +package caddy + +import ( + "context" + "crypto/sha256" + "encoding/hex" + "errors" + "strings" + "testing" + + "github.com/useteploy/teploy/internal/ssh" +) + +// caddyMutateMocks satisfies one full mutate transaction (lock, read, +// adapt, commit, reload, verify, release). +func caddyMutateMocks(caddyfile string) (*ssh.MockExecutor, []ssh.MockCommand) { + extra := []ssh.MockCommand{ + ssh.MockCommand{Match: "mkdir /deployments/caddy/.lock", Output: ""}, + ssh.MockCommand{Match: "docker exec caddy caddy reload", Output: ""}, + ssh.MockCommand{Match: "a=$(docker exec caddy md5sum", Output: "TEPLOY_CADDY_OK"}, + } + mock := ssh.NewMockExecutor("1.2.3.4", extra...) + if caddyfile != "" { + mock.Files["/deployments/caddy/Caddyfile"] = []byte(caddyfile) + } + return mock, extra +} + +// TestWithGeneration_StampAndRegionHash pins the C01-8/9 identity: the +// managed block carries a generation stamp line inside the markers, and +// ManagedRegionHash(ReadManagedBlock) round-trips the exact region bytes. +func TestWithGeneration_StampAndRegionHash(t *testing.T) { + mock, _ := caddyMutateMocks("{\n\tadmin 0.0.0.0:2019\n}\n") + cd := NewClient(mock).WithGeneration(7) + if err := cd.SetRoute(context.Background(), "myapp", "myapp.com", "myapp-web-v1", 3000, TLS{}, "", nil, Firewall{}, Access{}); err != nil { + t.Fatalf("SetRoute: %v", err) + } + content := string(mock.Files["/deployments/caddy/Caddyfile"]) + if !strings.Contains(content, "# TEPLOY GENERATION 7\n") { + t.Errorf("the managed block must be stamped with its generation:\n%s", content) + } + + region, present, err := NewClient(mock).ReadManagedBlock(context.Background(), "myapp") + if err != nil || !present { + t.Fatalf("ReadManagedBlock: %q %v", region, err) + } + for _, want := range []string{ + "# TEPLOY BEGIN myapp", + "# TEPLOY GENERATION 7", + "myapp-web-v1:3000", + "# TEPLOY END myapp", + } { + if !strings.Contains(region, want) { + t.Errorf("region missing %q:\n%s", want, region) + } + } + // The hash normalization is the CAS's identity contract: region + "\n" + // (what sed prints), the empty input for an absent region. + want := sha256.Sum256([]byte(region + "\n")) + if got := ManagedRegionHash(region); got != hex.EncodeToString(want[:]) { + t.Errorf("ManagedRegionHash mismatch: %s", got) + } + empty := sha256.Sum256(nil) + if ManagedRegionHash("") != hex.EncodeToString(empty[:]) { + t.Error("absent region must hash the empty input") + } +} + +// TestRegionGenerationAndUpstreams pins the refusal-evidence parsers. +func TestRegionGenerationAndUpstreams(t *testing.T) { + region := "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 9\nmyapp.com {\n\treverse_proxy myapp-web-v9:3000 myapp-web-v9-2:3000 {\n\t}\n}\n# TEPLOY END myapp" + gen, ok := RegionGeneration(region) + if !ok || gen != 9 { + t.Fatalf("expected stamp 9, got %d %v", gen, ok) + } + if gen, ok := RegionGeneration(strings.ReplaceAll(region, "# TEPLOY GENERATION 9\n", "")); ok { + t.Fatalf("an unstamped region must report ok=false, got %d", gen) + } + ups := RegionUpstreams(region) + if len(ups) != 2 || ups[0] != "myapp-web-v9:3000" || ups[1] != "myapp-web-v9-2:3000" { + t.Fatalf("upstreams must parse for the evidence, got %v", ups) + } +} + +// TestRouteCAS_RefusesWhenPredecessorBlockChanged is the A12/T05 core +// acceptance: the block changed since this operation resolved it (a newer +// writer switched the route) — the commit is refused, nothing lands, no +// reload runs, and the refusal names BOTH generations plus the live +// upstreams. +func TestRouteCAS_RefusesWhenPredecessorBlockChanged(t *testing.T) { + predecessor := "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 7\nmyapp.com {\n\treverse_proxy myapp-web-v7:3000\n}\n# TEPLOY END myapp\n" + successor := "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 8\nmyapp.com {\n\treverse_proxy myapp-web-v8:3000\n}\n# TEPLOY END myapp\n" + mock, _ := caddyMutateMocks(predecessor) + + // The operation resolved the PREDECESSOR's region… + resolved, _, err := NewClient(mock).ReadManagedBlock(context.Background(), "myapp") + if err != nil { + t.Fatalf("ReadManagedBlock: %v", err) + } + // …and a successor switched the route before its commit. + mock.Files["/deployments/caddy/Caddyfile"] = []byte(successor) + + cd := NewClient(mock).WithRouteCAS("myapp", []string{ManagedRegionHash(resolved)}, 7) + err = cd.SetRoute(context.Background(), "myapp", "myapp.com", "myapp-web-v7", 3000, TLS{}, "", nil, Firewall{}, Access{}) + var cas *ErrRouteCAS + if !errors.As(err, &cas) { + t.Fatalf("expected *ErrRouteCAS, got %v", err) + } + if cas.ExpectedGeneration != 7 || !cas.FoundStamped || cas.FoundGeneration != 8 { + t.Errorf("the refusal must name both generations: %+v", cas) + } + if len(cas.FoundUpstreams) != 1 || cas.FoundUpstreams[0] != "myapp-web-v8:3000" { + t.Errorf("the refusal must name the live upstreams: %+v", cas) + } + for _, want := range []string{"generation 7", "generation 8", "myapp-web-v8:3000"} { + if !strings.Contains(err.Error(), want) { + t.Errorf("error must carry %q: %v", want, err) + } + } + // Nothing landed: the live Caddyfile is still the successor's block. + if string(mock.Files["/deployments/caddy/Caddyfile"]) != successor { + t.Errorf("a refused CAS must not modify the Caddyfile:\n%s", mock.Files["/deployments/caddy/Caddyfile"]) + } + for _, c := range mock.Calls { + if strings.HasPrefix(c, "docker exec caddy caddy reload") { + t.Error("no reload may run after a refused CAS") + } + } +} + +// TestRouteCAS_PassesOnExactPredecessor: the same region resolved is the +// region live at commit — the switch lands. +func TestRouteCAS_PassesOnExactPredecessor(t *testing.T) { + predecessor := "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 7\nmyapp.com {\n\treverse_proxy myapp-web-v7:3000\n}\n# TEPLOY END myapp\n" + mock, _ := caddyMutateMocks(predecessor) + resolved, _, err := NewClient(mock).ReadManagedBlock(context.Background(), "myapp") + if err != nil { + t.Fatal(err) + } + cd := NewClient(mock).WithGeneration(8).WithRouteCAS("myapp", []string{ManagedRegionHash(resolved)}, 7) + if err := cd.SetRoute(context.Background(), "myapp", "myapp.com", "myapp-web-v8", 3000, TLS{}, "", nil, Firewall{}, Access{}); err != nil { + t.Fatalf("SetRoute over the exact predecessor must land: %v", err) + } + content := string(mock.Files["/deployments/caddy/Caddyfile"]) + if !strings.Contains(content, "# TEPLOY GENERATION 8") || !strings.Contains(content, "myapp-web-v8:3000") { + t.Errorf("the switch must land stamped with its generation:\n%s", content) + } +} + +// TestRouteCAS_AcceptsRestoreSet: the compensation restore accepts either +// the region the operation resolved or the one it switched to, and refuses +// a successor's block — an UNFENCED restore cannot clobber the newer +// generation's route (the acceptance's delayed-effect race). +func TestRouteCAS_AcceptsRestoreSet(t *testing.T) { + // Region consts carry NO trailing newline (ReadManagedBlock's shape); + // seeded files add it. + resolvedBlock := "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 7\nmyapp.com {\n\treverse_proxy myapp-web-v7:3000\n}\n# TEPLOY END myapp" + switchedBlock := "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 8\nmyapp.com {\n\treverse_proxy myapp-web-v8:3000\n}\n# TEPLOY END myapp" + successorBlock := "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 9\nmyapp.com {\n\treverse_proxy myapp-web-v9:3000\n}\n# TEPLOY END myapp" + + // Restore over our own uncommitted switch: live == switched — lands. + mock, _ := caddyMutateMocks(switchedBlock + "\n") + cd := NewClient(mock).WithGeneration(7).WithRouteCAS("myapp", + []string{ManagedRegionHash(resolvedBlock), ManagedRegionHash(switchedBlock)}, 7) + if err := cd.SetRoute(context.Background(), "myapp", "myapp.com", "myapp-web-v7", 3000, TLS{}, "", nil, Firewall{}, Access{}); err != nil { + t.Fatalf("restore over the operation's own switch must land: %v", err) + } + if !strings.Contains(string(mock.Files["/deployments/caddy/Caddyfile"]), "# TEPLOY GENERATION 7") { + t.Error("the restored block must carry the restored generation's stamp") + } + + // Restore against a successor's block: refused, nothing lands. + mock2, _ := caddyMutateMocks(successorBlock + "\n") + cd2 := NewClient(mock2).WithGeneration(7).WithRouteCAS("myapp", + []string{ManagedRegionHash(resolvedBlock), ManagedRegionHash(switchedBlock)}, 7) + err := cd2.SetRoute(context.Background(), "myapp", "myapp.com", "myapp-web-v7", 3000, TLS{}, "", nil, Firewall{}, Access{}) + var cas *ErrRouteCAS + if !errors.As(err, &cas) || cas.FoundGeneration != 9 { + t.Fatalf("expected a CAS refusal naming the successor's generation 9, got %v", err) + } + if string(mock2.Files["/deployments/caddy/Caddyfile"]) != successorBlock+"\n" { + t.Error("a refused restore must not modify the Caddyfile") + } +} + +// TestRouteCAS_FirstDeployAbsentRegion: no managed region at resolution and +// none at commit — the CAS passes (hash of the empty input on both sides). +func TestRouteCAS_FirstDeployAbsentRegion(t *testing.T) { + mock, _ := caddyMutateMocks("{\n\tadmin 0.0.0.0:2019\n}\n") + cd := NewClient(mock).WithGeneration(1).WithRouteCAS("myapp", []string{ManagedRegionHash("")}, 0) + if err := cd.SetRoute(context.Background(), "myapp", "myapp.com", "myapp-web-v1", 3000, TLS{}, "", nil, Firewall{}, Access{}); err != nil { + t.Fatalf("first deploy over an absent region must land: %v", err) + } +} + +// TestManagedBlockHostsSkipsStamp keeps the legacy lb- detection honest +// with stamped blocks: the stamp comment is not the address line. +func TestManagedBlockHostsSkipsStamp(t *testing.T) { + content := "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 3\nmyapp.com, www.myapp.com {\n\treverse_proxy x:3000\n}\n# TEPLOY END myapp\n" + hosts := managedBlockHosts(content, "myapp") + if len(hosts) != 2 || hosts[0] != "myapp.com" || hosts[1] != "www.myapp.com" { + t.Fatalf("stamp line must be skipped, got %v", hosts) + } +} + +// TestRouteCASFragment_HashParityWithSed proves the Go normalization and +// the shell extraction agree: the fragment's acceptable hash is exactly +// sha256sum of sed's marker-inclusive output. +func TestRouteCASFragment_HashParityWithSed(t *testing.T) { + region := "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 7\nmyapp.com {\n\treverse_proxy myapp-web-v7:3000\n}\n# TEPLOY END myapp" + content := "{\n\tadmin 0.0.0.0:2019\n}\n\n" + region + "\n" + mock, _ := caddyMutateMocks(content) + goHash := ManagedRegionHash(region) + + // Run the actual fragment shape through the mock: the commit's CAS + // evaluates against Files and must PASS with the Go-computed hash. + cd := NewClient(mock).WithGeneration(8).WithRouteCAS("myapp", []string{goHash}, 7) + if err := cd.SetRoute(context.Background(), "myapp", "myapp.com", "myapp-web-v8", 3000, TLS{}, "", nil, Firewall{}, Access{}); err != nil { + t.Fatalf("the Go-computed region hash must satisfy the shell-side sha256sum comparison: %v", err) + } +} diff --git a/internal/cli/contracts_golden_test.go b/internal/cli/contracts_golden_test.go index 1a2d21a..4ac7f9c 100644 --- a/internal/cli/contracts_golden_test.go +++ b/internal/cli/contracts_golden_test.go @@ -9,7 +9,9 @@ package cli import ( "bytes" + "context" "encoding/json" + "errors" "os" "path/filepath" "strings" @@ -18,6 +20,7 @@ import ( "github.com/useteploy/teploy/internal/config" "github.com/useteploy/teploy/internal/releasemeta" + "github.com/useteploy/teploy/internal/ssh" ) const contractsDir = "../../contracts" @@ -131,6 +134,70 @@ func TestContractsAppListEnvelopeGolden(t *testing.T) { writeFixture(t, "app-list-envelope/legacy/pre-mi.json", legacy) } +// TestContractsServerStatusEnvelopeGolden drives the REAL +// collectServerStatus encoder (the same collection path `server status +// --json` runs) through a mock SSH executor — the DTO values are +// synthetic, the encoder and every parse stage (memory, disks, docker +// inventory, Caddy routes) are the real ones. The wire shape was +// verified against a live `server status --json` capture before this +// fixture was pinned (see MANIFEST rev 5). Two valid classes: a full +// healthy observation, and a partial one with the Caddy probe failing — +// the class a real deployment without a caddy container produces. The +// legacy fixture is the pre-MI shape (machine_interface absent), which a +// 42243e2-era CLI emitted. +func TestContractsServerStatusEnvelopeGolden(t *testing.T) { + observedAt := time.Date(2026, 9, 23, 12, 0, 0, 0, time.UTC) + container := `{"ID":"9f31c02","Names":"myapp-web-3","Image":"example/myapp:3","State":"running","Status":"Up 4 minutes","CreatedAt":"2026-09-23 11:55:00 +0000 UTC","Labels":"teploy.app=myapp,teploy.process=web,teploy.version=3"}` + image := `{"ID":"sha256:1a2b3c4d5e6f","Repository":"example/myapp","Tag":"3","Size":"25MB","CreatedAt":"2026-09-23 11:50:00 +0000 UTC"}` + caddy := `{"servers":{"srv0":{"routes":[{"@id":"myapp","match":[{"host":["myapp.example.com"]}],"handle":[{"handler":"subroute","routes":[{"handle":[{"handler":"reverse_proxy","upstreams":[{"dial":"myapp-web-3:3000"}]}]}]}]}]}}}` + full := ssh.NewMockExecutor("192.0.2.10", + ssh.MockCommand{Match: "cat /proc/uptime", Output: "3600.50 1200.00"}, + ssh.MockCommand{Match: "cat /proc/loadavg", Output: "0.10 0.20 0.30 1/100 1"}, + ssh.MockCommand{Match: "cat /proc/meminfo", Output: "MemTotal: 1000 kB\nMemAvailable: 400 kB\n"}, + ssh.MockCommand{Match: "df -B1 -P", Output: "Filesystem 1-blocks Used Available Capacity Mounted on\n/dev/vda1 1000 250 750 25% /\n/dev/vdb1 2000 500 1500 26% /srv\n"}, + ssh.MockCommand{Match: "docker version", Output: "29.0.0"}, + ssh.MockCommand{Match: "docker ps --all", Output: container}, + ssh.MockCommand{Match: "docker image ls", Output: image}, + ssh.MockCommand{Match: "docker exec caddy", Output: caddy}, + ) + got := collectServerStatus(context.Background(), full, "prod", observedAt) + if len(got.Errors) != 0 { + t.Fatalf("full observation reported errors: %#v", got.Errors) + } + fullStatus := got + writeFixture(t, "server-status-envelope/valid/full.json", got) + + partial := ssh.NewMockExecutor("192.0.2.20", + ssh.MockCommand{Match: "cat /proc/uptime", Output: "86400.00 86400.00"}, + ssh.MockCommand{Match: "cat /proc/loadavg", Output: "0.00 0.01 0.05 1/100 1"}, + ssh.MockCommand{Match: "cat /proc/meminfo", Output: "MemTotal: 500 kB\nMemAvailable: 250 kB\n"}, + ssh.MockCommand{Match: "df -B1 -P", Output: "Filesystem 1-blocks Used Available Capacity Mounted on\n/dev/vda1 500 100 400 20% /\n"}, + ssh.MockCommand{Match: "docker version", Output: "29.0.0"}, + ssh.MockCommand{Match: "docker ps --all", Output: ""}, + ssh.MockCommand{Match: "docker image ls", Output: ""}, + ssh.MockCommand{Match: "docker exec caddy", Err: errors.New("Error response from daemon: No such container: caddy")}, + ) + got = collectServerStatus(context.Background(), partial, "staging", observedAt) + if len(got.Errors) != 1 || got.Errors[0].Scope != "caddy.routes" { + t.Fatalf("partial observation missing its caddy error: %#v", got.Errors) + } + writeFixture(t, "server-status-envelope/valid/partial-caddy-unavailable.json", got) + + // Legacy: the pre-MI envelope (no machine_interface field) a CLI + // between 42243e2 and dda4911 emitted. Same delete-from-map approach + // as the app-list legacy class. + raw, err := json.Marshal(fullStatus) + if err != nil { + t.Fatalf("marshal full status: %v", err) + } + var legacy map[string]any + if err := json.Unmarshal(raw, &legacy); err != nil { + t.Fatalf("unmarshal legacy: %v", err) + } + delete(legacy, "machine_interface") + writeFixture(t, "server-status-envelope/legacy/pre-mi.json", legacy) +} + // TestContractsErrorEnvelopeGolden pins the two wired error classes. func TestContractsErrorEnvelopeGolden(t *testing.T) { writeFixture(t, "error-envelope/valid/config-invalid.json", machineErrorEnvelope{ diff --git a/internal/cli/deploy.go b/internal/cli/deploy.go index ea765ae..f4cb0ef 100644 --- a/internal/cli/deploy.go +++ b/internal/cli/deploy.go @@ -61,7 +61,15 @@ If a previously-deployed container mounted volumes from a different host path (common when migrating from Dokploy or hand-rolled docker run setups), the deploy aborts safely rather than orphaning data. Pass --migrate-volumes to copy data from the existing source into the teploy-expected path before -swapping traffic.`, +swapping traffic. + +Outcomes and exit codes: 0 = deployed (warnings possible — read the output); +1 = refused before any effect (config/admission problems), failed, or +INTERRUPTED — an interrupted deploy (Ctrl-C, timeout) has an unknown outcome +until reconciled; re-running is safe (the next attempt reconciles partial +state), or inspect with teploy status. Under --json a failure is reported as +a machine error envelope on stderr whose code distinguishes the classes +(config-invalid, conflict, uncertain-outcome, internal).`, Args: cobra.MaximumNArgs(1), RunE: func(cmd *cobra.Command, args []string) error { var serverName string @@ -713,7 +721,7 @@ func deployBuiltImageFenced(ctx context.Context, executor ssh.Executor, appCfg * emitDeployAudit(ctx, appCfg, "deploy.run", version, serverDisplay, deployErr) if deployErr != nil { - return deployErr + return wrapDeployOutcomeError(deployErr) } // 13. Prune old build images (best-effort). @@ -725,6 +733,19 @@ func deployBuiltImageFenced(ctx context.Context, executor ssh.Executor, appCfg * return nil } +// wrapDeployOutcomeError renders a failed deploy's error for the human +// path. An interrupted deploy (Ctrl-C, automation timeout) leaves the +// outcome UNKNOWN — the C01 crash-window contract — so it names the +// recovery action instead of letting a bare "context canceled" read as a +// clean failure. Under --json the machine envelope classifies the same +// error as uncertain-outcome (errevelope.go). +func wrapDeployOutcomeError(err error) error { + if errors.Is(err, context.Canceled) || errors.Is(err, context.DeadlineExceeded) { + return fmt.Errorf("deploy interrupted: %w\nThe outcome is unknown until reconciled. Re-running this deploy is safe (the next attempt reconciles any partial state); inspect first with: teploy status", err) + } + return err +} + // plannedVolumeMounts resolves config volume declarations to the docker // mount map a deploy uses: a host bind (name starting with "/") mounts a // directory the operator owns, exactly as given — teploy never creates or @@ -969,18 +990,18 @@ func runMultiDeploy(flags *Flags, appCfg *config.AppConfig, image, version strin fmt.Printf("\n%d of %d servers failed — rolling back the %d server(s) that succeeded...\n", failCount, len(targets), len(successTargets)) - // Best-effort: attempt to roll back EVERY succeeded server even if one - // rollback fails — otherwise a single rollback failure would fail-fast - // and strand the remaining servers on the new version (M1). The wave - // runs on a bounded DETACHED recovery context (audit T58): the deploy - // context is signal-cancelled exactly when the operator interrupts, - // and recovery work that skips itself because the cancelled context - // disappeared is how a Ctrl-C strands half a fleet on the new version. - rollbackCtx, rollbackCancel := context.WithTimeout(context.WithoutCancel(ctx), 10*time.Minute) - defer rollbackCancel() - rollbackResults := multideploy.ParallelDeployAll(rollbackCtx, successTargets, parallel, func(ctx context.Context, target multideploy.ServerTarget, out io.Writer) error { - return rollbackSingleServer(ctx, appCfg, target, out) - }, os.Stdout) + // Best-effort: attempt to roll back EVERY succeeded server even if one + // rollback fails — otherwise a single rollback failure would fail-fast + // and strand the remaining servers on the new version (M1). The wave + // runs on a bounded DETACHED recovery context (audit T58): the deploy + // context is signal-cancelled exactly when the operator interrupts, + // and recovery work that skips itself because the cancelled context + // disappeared is how a Ctrl-C strands half a fleet on the new version. + rollbackCtx, rollbackCancel := context.WithTimeout(context.WithoutCancel(ctx), 10*time.Minute) + defer rollbackCancel() + rollbackResults := multideploy.ParallelDeployAll(rollbackCtx, successTargets, parallel, func(ctx context.Context, target multideploy.ServerTarget, out io.Writer) error { + return rollbackSingleServer(ctx, appCfg, target, out) + }, os.Stdout) var rolledBack, firstDeploys, rollbackFailed []string for _, r := range rollbackResults { diff --git a/internal/cli/errevelope.go b/internal/cli/errevelope.go index 04505c7..b313f6f 100644 --- a/internal/cli/errevelope.go +++ b/internal/cli/errevelope.go @@ -1,6 +1,7 @@ package cli import ( + "context" "encoding/json" "errors" "fmt" @@ -39,7 +40,9 @@ const ( // (UNMIGRATED: currently internal). codeConflict = "conflict" // The effect's fate is unknown pending reconciliation; never - // rendered or recorded as failure (UNMIGRATED: currently internal). + // rendered or recorded as failure. Wired for interrupted commands + // (context canceled / deadline exceeded — SIGINT, automation + // timeouts); deeper per-phase migration stays on the S2 list. codeUncertainOutcome = "uncertain-outcome" // Success with a flag — traffic switched but the outcome is not // clean (UNMIGRATED: currently internal). @@ -87,12 +90,19 @@ func refuseAdmission(err error) error { // refusals → config-invalid; an absent config is a config failure for a // machine caller the same way a malformed one is. A plan/apply drift // refusal (C05) → conflict — the request is coherent, the world moved. -// Everything else is internal until its site is migrated (S2 +// An interrupted command (context canceled or deadline exceeded — the +// SIGINT path, an automation timeout) → uncertain-outcome: whatever +// effect was in flight has an unknown fate until reconciled (the C01 +// crash-window machinery treats every interrupted window as INSPECT), +// which is precisely the class automation must not render as a plain +// failure. Everything else is internal until its site is migrated (S2 // generalization). func classifyMachineError(err error) string { switch { case errors.Is(err, errPlanDrift): return codeConflict + case errors.Is(err, context.Canceled), errors.Is(err, context.DeadlineExceeded): + return codeUncertainOutcome case errors.Is(err, config.ErrInvalidConfig), errors.Is(err, config.ErrNoConfig), errors.Is(err, errDeployAdmission): @@ -114,6 +124,8 @@ func writeMachineErrorEnvelope(out io.Writer, err error) error { } case codeConflict: message = "plan no longer valid" + case codeUncertainOutcome: + message = "interrupted — outcome unknown until reconciled" } return json.NewEncoder(out).Encode(machineErrorEnvelope{ MachineInterface: MachineInterface, diff --git a/internal/cli/errevelope_test.go b/internal/cli/errevelope_test.go index 079fde1..8381caf 100644 --- a/internal/cli/errevelope_test.go +++ b/internal/cli/errevelope_test.go @@ -2,8 +2,10 @@ package cli import ( "bytes" + "context" "encoding/json" "errors" + "fmt" "strings" "testing" @@ -150,6 +152,63 @@ func TestMachineErrorEnvelopeUnclassifiedIsInternal(t *testing.T) { } } +// TestMachineErrorEnvelopeInterruptedIsUncertain: an interrupted command +// (context canceled — the SIGINT path — or a deadline exceeded) is NOT a +// plain failure: its outcome is unknown until reconciled (C09's +// uncertain/canceled distinction for automation). Wrapped and chained +// errors must classify the same way. +func TestMachineErrorEnvelopeInterruptedIsUncertain(t *testing.T) { + for name, err := range map[string]error{ + "canceled": context.Canceled, + "deadline": context.DeadlineExceeded, + "wrapped": fmt.Errorf("deploying myapp: %w", context.Canceled), + "double-wrapped": fmt.Errorf("running docker run: %w", fmt.Errorf("build: %w", context.DeadlineExceeded)), + } { + var out bytes.Buffer + if writeErr := writeMachineErrorEnvelope(&out, err); writeErr != nil { + t.Fatal(writeErr) + } + var decoded map[string]any + if jsonErr := json.Unmarshal(out.Bytes(), &decoded); jsonErr != nil { + t.Fatalf("%s: envelope not JSON: %q", name, out.String()) + } + if decoded["code"] != "uncertain-outcome" { + t.Fatalf("%s: code = %v, want uncertain-outcome", name, decoded["code"]) + } + if decoded["message"] != "interrupted — outcome unknown until reconciled" { + t.Fatalf("%s: message = %v", name, decoded["message"]) + } + } +} + +// TestDeployInterruptedErrorNamesRecovery: the human-path error for an +// interrupted deploy names the recovery action instead of a bare +// "context canceled" (which reads as a clean failure), while a plain +// deploy failure passes through unchanged. +func TestDeployInterruptedErrorNamesRecovery(t *testing.T) { + interrupted := wrapDeployOutcomeError(fmt.Errorf("deploying myapp: %w", context.Canceled)) + msg := interrupted.Error() + if !strings.Contains(msg, "deploy interrupted") || !strings.Contains(msg, "reconcil") || !strings.Contains(msg, "teploy status") { + t.Fatalf("interrupted deploy error must name the recovery action: %q", msg) + } + var out bytes.Buffer + if err := writeMachineErrorEnvelope(&out, interrupted); err != nil { + t.Fatal(err) + } + var decoded map[string]any + if err := json.Unmarshal(out.Bytes(), &decoded); err != nil { + t.Fatalf("envelope not JSON: %q", out.String()) + } + if decoded["code"] != "uncertain-outcome" { + t.Fatalf("interrupted deploy code = %v, want uncertain-outcome", decoded["code"]) + } + + plain := errors.New("image pull failed") + if got := wrapDeployOutcomeError(plain); got != plain { + t.Fatalf("plain failure must pass through unchanged: %v", got) + } +} + // TestExitCodesPinned: 0/1/2 semantics are unchanged by the envelope — // drift's 2 remains a signal (not a failure) gated on --exit-code, and // every other failure exits 1 after the error is reported. diff --git a/internal/cli/version.go b/internal/cli/version.go index c97b39b..9bf60e5 100644 --- a/internal/cli/version.go +++ b/internal/cli/version.go @@ -21,6 +21,12 @@ func newVersionCmd(flags *Flags, version string) *cobra.Command { return &cobra.Command{ Use: "version", Short: "Show teploy version", + Long: `Show teploy version. + +With --json, emits the machine-interface handshake instead: the version, +the machine_interface number (fail closed if a consumer supports less), +and the capability tokens this build advertises. One call replaces +help-text scraping as the compatibility check.`, RunE: func(cmd *cobra.Command, args []string) error { return writeVersion(cmd.OutOrStdout(), version, flags.JSON) }, diff --git a/internal/deploy/audit_fixes_test.go b/internal/deploy/audit_fixes_test.go index ea6fa4d..2314acb 100644 --- a/internal/deploy/audit_fixes_test.go +++ b/internal/deploy/audit_fixes_test.go @@ -44,7 +44,7 @@ func TestDeploy_SameVersion_DedupesReplicaAndPlainNames(t *testing.T) { ssh.MockCommand{Match: "a=$(docker exec caddy md5sum", Output: "TEPLOY_CADDY_OK"}, ssh.MockCommand{Match: "UPLOAD:", Output: ""}, ssh.MockCommand{Match: "mv -f -- ", Output: ""}, - ssh.MockCommand{Match: "if [ ! -e '/deployments/caddy/Caddyfile'", Err: fmt.Errorf("none")}, + ssh.MockCommand{Match: "if [ ! -e '/deployments/caddy/Caddyfile'", Output: "present\n{\n\tadmin 0.0.0.0:2019\n}\n"}, ) var buf bytes.Buffer @@ -111,7 +111,7 @@ func TestDeploy_RemovedWorkerIsStoppedByInventory(t *testing.T) { ssh.MockCommand{Match: "a=$(docker exec caddy md5sum", Output: "TEPLOY_CADDY_OK"}, ssh.MockCommand{Match: "UPLOAD:", Output: ""}, ssh.MockCommand{Match: "mv -f -- ", Output: ""}, - ssh.MockCommand{Match: "if [ ! -e '/deployments/caddy/Caddyfile'", Err: fmt.Errorf("none")}, + ssh.MockCommand{Match: "if [ ! -e '/deployments/caddy/Caddyfile'", Output: "present\n{\n\tadmin 0.0.0.0:2019\n}\n"}, ) var buf bytes.Buffer @@ -130,7 +130,10 @@ func TestDeploy_RemovedWorkerIsStoppedByInventory(t *testing.T) { stoppedWorker := false for _, call := range mock.Calls { - if strings.Contains(call, "myapp-worker-old123") && strings.HasPrefix(call, "docker stop") { + // The retirement stop addresses the container by its exact ID from + // the deploy's under-lock snapshot (C01-9); the fixture's inventory + // gives the worker ID b. + if (strings.Contains(call, "myapp-worker-old123") || strings.Contains(call, "'b'")) && strings.HasPrefix(call, "docker stop") { stoppedWorker = true } } diff --git a/internal/deploy/deploy.go b/internal/deploy/deploy.go index ff745d9..554177c 100644 --- a/internal/deploy/deploy.go +++ b/internal/deploy/deploy.go @@ -356,6 +356,40 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) return fmt.Errorf("refusing to deploy with unreadable state for %s: %w", cfg.App, err) } + // 1z. Resolve the generation identities this deploy prepares against + // (C01-8/9). Every successful operation bumps the state generation by + // one; the candidates this deploy starts are labeled with the NEXT + // generation, and its destructive effects (commit, route switch, + // retirement) fence on the CURRENT one — a stale operation whose + // commands land after a takeover can neither commit over nor stop a + // newer generation, no matter when its SSH effects arrive. + fromGeneration := uint64(0) + if current != nil { + fromGeneration = current.Generation + } + nextGeneration := fromGeneration + 1 + + // The exact-block route CAS expectation (A12/T05): resolve the app's + // managed region NOW, under the lock, before any effect — the commit + // of the route switch compares it and refuses if a newer writer + // swapped the block in between (evidence names both generations). A + // Caddyfile read failure is a deploy refusal exactly like the state + // read above: an unreadable route authority cannot be compared, and an + // uncomparable switch is not made. + resolvedRouteHash := "" + if cfg.usesCaddy() { + region, _, rerr := d.caddy.ReadManagedBlock(ctx, cfg.App) + if rerr != nil { + return fmt.Errorf("refusing to deploy %s with an unreadable route authority: %w", cfg.App, rerr) + } + resolvedRouteHash = caddy.ManagedRegionHash(region) + } + // switchedRouteHash is re-resolved after a SUCCESSFUL switch (below): + // the block this deploy wrote but has not yet committed — the restore + // path's CAS accepts exactly {resolved, switched} and refuses anything + // else (a successor's block). + switchedRouteHash := resolvedRouteHash + // 1b. Converge outstanding record repair debt (C01-6): a previous // deploy whose releasemeta record write failed after the live commit // left a repair-debt marker. Rebuild that record from the live @@ -599,6 +633,12 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) // returned without either, orphaning the first replica and leaving a // host-ingress app down (audit F05). var started []string + // startedIDs carries the exact container IDs docker run returned, + // paired with `started`'s names — cleanup stops by ID under the + // generation label check (C01-9), so a takeover + same-hash redeploy + // cannot have this (possibly stale) holder's cleanup stop the + // successor's same-named containers. + var startedIDs []string // candidateIDs collects the container IDs docker run returned for the // web candidates — the exact identities the readiness receipt records // (C01-4) and the strongest attribution evidence a recovery owner has. @@ -614,12 +654,28 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) // (or one of its replicas) down — "at least one thing worked" must // never read as "recovered". var cleanupFailures []string - for _, n := range started { - if err := d.docker.Stop(recoveryCtx, n, 5); err != nil { + for i, n := range started { + // Stop by the EXACT identity this deploy created, fenced on + // the generation label (C01-9): a successor's same-hash + // containers carry a newer generation and are refused here — + // cleanup of our own effects must never destroy a newer + // generation's workload. Recovery is deliberately NOT + // holdership-fenced (A07); the generation check is identity, + // not holdership, and applies in recovery too. + ref := n + if i < len(startedIDs) && startedIDs[i] != "" { + ref = startedIDs[i] + } + expected := nextGeneration + if err := d.docker.StopGenerationFenced(recoveryCtx, ref, 5, expected, ""); err != nil { + if state.GenerationFenced(err) { + cleanupFailures = append(cleanupFailures, fmt.Sprintf("stop %s: refused — a newer generation owns the name", n)) + continue + } cleanupFailures = append(cleanupFailures, fmt.Sprintf("stop %s: %v", n, err)) continue } - if err := d.docker.Remove(recoveryCtx, n); err != nil { + if err := d.docker.Remove(recoveryCtx, ref); err != nil { cleanupFailures = append(cleanupFailures, fmt.Sprintf("remove %s: %v", n, err)) } } @@ -634,7 +690,16 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) } restored := len(displaced) == 0 for _, old := range displaced { - if err := d.docker.Restart(recoveryCtx, old, nil); err != nil { + // The generation check guards the recreate's force-remove + // (C01-9): a successor may have re-created the name at a newer + // generation — restoring over it would destroy the new owner's + // workload, exactly what the check refuses. + if err := d.docker.RestartFenced(recoveryCtx, old, nil, fromGeneration, ""); err != nil { + if state.GenerationFenced(err) { + cleanupFailures = append(cleanupFailures, fmt.Sprintf("restore %s: refused — a newer generation owns the name", old)) + fmt.Fprintf(d.out, " WARNING: could not restore displaced container %s — a newer generation now owns the name\n", old) + continue + } cleanupFailures = append(cleanupFailures, fmt.Sprintf("restore %s: %v", old, err)) fmt.Fprintf(d.out, " WARNING: could not restore displaced container %s: %v\n", old, err) } else { @@ -675,6 +740,7 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) App: cfg.App, Process: "web", Version: cfg.Version, + Generation: nextGeneration, Image: runImage, Port: ports[i], BindHost: webBindHost, @@ -708,6 +774,7 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) return restoreDisplacedAndStarted(fmt.Errorf("starting container %s: %w", name, err)) } started = append(started, name) + startedIDs = append(startedIDs, containerID) candidateIDs[i] = containerID fmt.Fprintf(d.out, " Container %s started\n", containerID[:min(12, len(containerID))]) } @@ -807,10 +874,11 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) } name := docker.ContainerName(cfg.App, process, cfg.Version) fmt.Fprintf(d.out, "Starting %s...\n", name) - _, err := d.docker.RunGuarded(ctx, docker.RunConfig{ + workerID, err := d.docker.RunGuarded(ctx, docker.RunConfig{ App: cfg.App, Process: process, Version: cfg.Version, + Generation: nextGeneration, Image: runImage, Port: 0, // non-web processes don't get a port EnvFiles: cfg.EnvFiles, @@ -829,6 +897,7 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) return fail(fmt.Errorf("starting %s: %w", name, err)) } started = append(started, name) + startedIDs = append(startedIDs, workerID) // A detached `docker run` proves nothing about the worker's // viability — a bad command or an instantly-crashing process used // to be recorded as a successful deploy while no jobs were consumed @@ -855,7 +924,16 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) // fence guard in the same shell — a late write from a broken holder // cannot hijack a newer operation's route. The pre-switch Check is // superseded by the composed commit. - cad := d.caddy.WithCommitGuard(guardPrefix) + // + // C01-8/9 + A12/T05: the new block is STAMPED with the generation + // this deploy creates, and the commit runs under the exact-block + // compare-and-swap on the region resolved at step 1z — if a newer + // writer swapped the block in between, the edit is refused with + // both generations named (caddy.ErrRouteCAS) instead of landing + // over the successor's route. + cad := d.caddy.WithCommitGuard(guardPrefix). + WithGeneration(nextGeneration). + WithRouteCAS(cfg.App, []string{resolvedRouteHash}, fromGeneration) tls := caddy.TLS{Cert: cfg.TLSCert, Key: cfg.TLSKey, Internal: cfg.TLSInternal} if replicas > 1 { upstreams := make([]caddy.Upstream, replicas) @@ -868,6 +946,10 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) if state.FenceLost(err) { return fail(fmt.Errorf("%w: refusing to switch %s's route", state.ErrFenceLost, cfg.App)) } + var routeCAS *caddy.ErrRouteCAS + if errors.As(err, &routeCAS) { + return fail(fmt.Errorf("%w", routeCAS)) + } return fail(fmt.Errorf("updating load balancer route: %w", err)) } fmt.Fprintf(d.out, " Traffic load-balanced across %d replicas\n", replicas) @@ -876,10 +958,18 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) if state.FenceLost(err) { return fail(fmt.Errorf("%w: refusing to switch %s's route", state.ErrFenceLost, cfg.App)) } + var routeCAS *caddy.ErrRouteCAS + if errors.As(err, &routeCAS) { + return fail(fmt.Errorf("%w", routeCAS)) + } return fail(fmt.Errorf("updating route: %w", err)) } fmt.Fprintln(d.out, " Traffic routed to new container") } + // The switched-to region is this deploy's own uncommitted block. + if region, _, rerr := d.caddy.ReadManagedBlock(ctx, cfg.App); rerr == nil { + switchedRouteHash = caddy.ManagedRegionHash(region) + } } else { fmt.Fprintf(d.out, "Skipping Caddy route update (ingress: %s)\n", cfg.Ingress) } @@ -917,9 +1007,13 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) } // The commit runs under the fence (F16): the atomic rename that makes // this deploy authoritative is a guarded effect, so a broken holder - // commits nothing. - if err := state.WriteFenced(ctx, d.exec, cfg.App, newState, lk); err != nil { - return d.abortStateCommit(ctx, cfg, current, started, displacedHostWeb, assetAttempt, start, err) + // commits nothing. C01-8/9: the guard chains a compare-and-swap on the + // committed-generation sidecar — a deploy that resolved the world at + // generation G cannot commit over a successor's G+N even if its + // commands land inside the successor's window (WriteFencedGeneration + // refuses with ErrGenerationFenced naming both generations). + if err := state.WriteFencedGeneration(ctx, d.exec, cfg.App, newState, lk, fromGeneration); err != nil { + return d.abortStateCommit(ctx, cfg, current, started, startedIDs, displacedHostWeb, assetAttempt, restoreRouteCAS(resolvedRouteHash, switchedRouteHash), start, err) } // 13b. Record the release metadata (F14). The containers are live and @@ -1018,7 +1112,7 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) // running part of the superseded generation. var retireIncomplete []string if predecessorsListed { - retireIncomplete = d.stopPredecessorSnapshot(ctx, predecessors, sameVersion, stopTimeout, lk) + retireIncomplete = d.stopPredecessorSnapshot(ctx, predecessors, sameVersion, stopTimeout, lk, fromGeneration, guardPrefix) } else if current != nil && current.CurrentHash != "" { // Fence (F16): post-commit cleanup never interleaves with a new // holder. @@ -1032,7 +1126,7 @@ func (d *Deployer) DeployFenced(ctx context.Context, cfg Config, lk *state.Lock) } if fenceOK { if inv, invErr := d.docker.ListContainers(ctx, cfg.App); invErr == nil { - retireIncomplete = append(retireIncomplete, d.stopPredecessorSnapshot(ctx, selectPredecessors(inv, current, sameVersion), sameVersion, stopTimeout, lk)...) + retireIncomplete = append(retireIncomplete, d.stopPredecessorSnapshot(ctx, selectPredecessors(inv, current, sameVersion), sameVersion, stopTimeout, lk, fromGeneration, guardPrefix)...) } else { fmt.Fprintf(d.out, "Warning: container inventory still unreadable (%v) — cleaning up by derived names; a removed worker process may escape retirement\n", invErr) if err := stopOldWorkloadsByName(ctx, d.docker, d.out, cfg, current, processes, stopTimeout); err != nil { @@ -1160,31 +1254,43 @@ func selectPredecessors(inv []docker.Container, current *state.AppState, sameVer // stopPredecessorSnapshot retires exactly the snapshotted predecessor set // and returns what escaped retirement (C01-5: the caller records it as a // degraded outcome — never clean success, never a failed deploy). -// Fence checks precede each stop: the deploy is already committed, and a -// fence loss mid-cleanup means another operation owns the app — refuse -// further stops (loudly) rather than interleaving. -func (d *Deployer) stopPredecessorSnapshot(ctx context.Context, predecessors []docker.Container, sameVersion bool, stopTimeout int, lk *state.Lock) []string { +// C01-8/9: each stop is a composed guarded effect — holdership guard, +// generation label check, and docker stop in ONE command, addressing the +// container by its exact ID from this deploy's own under-lock snapshot. A +// fence loss refuses further stops loudly; a generation refusal means the +// name now belongs to a newer generation's container (a takeover with a +// same-hash redeploy) and is skipped with the evidence, never executed — +// the retirement of OUR predecessor may not stop the successor's workload. +func (d *Deployer) stopPredecessorSnapshot(ctx context.Context, predecessors []docker.Container, sameVersion bool, stopTimeout int, lk *state.Lock, expectedGeneration uint64, guardPrefix string) []string { var incomplete []string for _, ct := range predecessors { - if lk != nil { - if err := lk.Check(ctx, d.exec); err != nil { + ref := ct.Name + if ct.ID != "" { + ref = ct.ID + } + if err := d.docker.StopGenerationFenced(ctx, ref, stopTimeout, expectedGeneration, guardPrefix); err != nil { + switch { + case state.GenerationFenced(err): + fmt.Fprintf(d.out, "Warning: retirement of %s refused — %v\n", ct.Name, err) + incomplete = append(incomplete, fmt.Sprintf("stop %s refused: a newer generation owns the container", ct.Name)) + continue + case state.FenceLost(err): fmt.Fprintf(d.out, "Warning: predecessor cleanup stopped — %v\n", err) incomplete = append(incomplete, fmt.Sprintf("cleanup interrupted before %s (fence lost)", ct.Name)) - break + return incomplete + default: + // Traffic is already committed to the new generation; a + // failed predecessor stop is degraded cleanup, not a failed + // deploy — but it must be reported, never silent (TCL-19): + // a leftover old worker keeps consuming jobs. + fmt.Fprintf(d.out, "Warning: could not stop old container %s: %v\n", ct.Name, err) + incomplete = append(incomplete, fmt.Sprintf("stop %s: %v", ct.Name, err)) + continue } } fmt.Fprintf(d.out, "Stopping old container %s...\n", ct.Name) - if err := d.docker.Stop(ctx, ct.Name, stopTimeout); err != nil { - // Traffic is already committed to the new generation; a - // failed predecessor stop is degraded cleanup, not a failed - // deploy — but it must be reported, never silent (TCL-19): - // a leftover old worker keeps consuming jobs. - fmt.Fprintf(d.out, "Warning: could not stop old container %s: %v\n", ct.Name, err) - incomplete = append(incomplete, fmt.Sprintf("stop %s: %v", ct.Name, err)) - continue - } if sameVersion { - if err := d.docker.Remove(ctx, ct.Name); err != nil { + if err := d.docker.Remove(ctx, ref); err != nil { fmt.Fprintf(d.out, "Warning: could not remove old container %s: %v\n", ct.Name, err) incomplete = append(incomplete, fmt.Sprintf("remove %s: %v", ct.Name, err)) } @@ -1237,7 +1343,21 @@ func stopOldWorkloadsByName(ctx context.Context, dk *docker.Client, out io.Write return errors.Join(failures...) } -func (d *Deployer) abortStateCommit(ctx context.Context, cfg Config, current *state.AppState, started, displacedHostWeb []string, att releasemeta.Attempt, start time.Time, commitErr error) error { +// restoreRouteCAS builds the restore path's exact-block CAS set: the +// region this operation resolved (A), the region it switched to but has +// not committed (B). A restore may only overwrite one of those; anything +// else is a successor's block and the restore refuses (A12/T05 — a +// deliberately UNFENCED compensation must not clobber a newer owner's +// route). +func restoreRouteCAS(resolvedHash, switchedHash string) []string { + set := []string{resolvedHash} + if switchedHash != resolvedHash { + set = append(set, switchedHash) + } + return set +} + +func (d *Deployer) abortStateCommit(ctx context.Context, cfg Config, current *state.AppState, started, startedIDs []string, displacedHostWeb []string, att releasemeta.Attempt, routeCAS []string, start time.Time, commitErr error) error { // Compensation runs on a DETACHED bounded context (A11): if the commit // failed because the deploy context was cancelled, reusing that context // would skip the very stops/restarts/route restores that undo the @@ -1246,6 +1366,13 @@ func (d *Deployer) abortStateCommit(ctx context.Context, cfg Config, current *st defer cancel() d.logDeploy(recoveryCtx, cfg, false, "", start) + fromGeneration := uint64(0) + nextGeneration := uint64(1) + if current != nil { + fromGeneration = current.Generation + nextGeneration = fromGeneration + 1 + } + // The displaced fixed-port workload: the in-memory list when this // process did the displacing; a process recovering a crashed attempt // reads it from the durable predecessor snapshot (C01-10) — exactly @@ -1256,15 +1383,28 @@ func (d *Deployer) abortStateCommit(ctx context.Context, cfg Config, current *st } if cfg.ingressHost() || len(cfg.Publish) > 0 { - for _, name := range started { - d.docker.Stop(recoveryCtx, name, 5) + for i, name := range started { + ref := name + if i < len(startedIDs) && startedIDs[i] != "" { + ref = startedIDs[i] + } + d.docker.StopGenerationFenced(recoveryCtx, ref, 5, nextGeneration, "") } for _, old := range displaced { - if err := d.docker.Restart(recoveryCtx, old, nil); err != nil { - for _, name := range started { - d.docker.Start(recoveryCtx, name) + if err := d.docker.RestartFenced(recoveryCtx, old, nil, fromGeneration, ""); err != nil { + if !state.GenerationFenced(err) { + for i, name := range started { + ref := name + if i < len(startedIDs) && startedIDs[i] != "" { + ref = startedIDs[i] + } + d.docker.Start(recoveryCtx, ref) + } + return fmt.Errorf("committing authoritative applied state after replacing the fixed-port host workload: %w; restoring the original workload failed: %v; Teploy attempted to restart the new workload to avoid an outage", commitErr, err) } - return fmt.Errorf("committing authoritative applied state after replacing the fixed-port host workload: %w; restoring the original workload failed: %v; Teploy attempted to restart the new workload to avoid an outage", commitErr, err) + // A newer generation owns the name: the successor is + // responsible for it now — leave it strictly alone. + fmt.Fprintf(d.out, " WARNING: displaced container %s was not restored — a newer generation owns the name\n", old) } } // A caddy-ingress app with publish entries entered this branch too @@ -1275,15 +1415,23 @@ func (d *Deployer) abortStateCommit(ctx context.Context, cfg Config, current *st // if that fails, keep the candidates running rather than routing to // nothing. if cfg.usesCaddy() { - if err := d.restorePreviousRoute(recoveryCtx, cfg, current); err != nil { - for _, name := range started { - d.docker.Start(recoveryCtx, name) + if err := d.restorePreviousRoute(recoveryCtx, cfg, current, routeCAS, fromGeneration); err != nil { + for i, name := range started { + ref := name + if i < len(startedIDs) && startedIDs[i] != "" { + ref = startedIDs[i] + } + d.docker.Start(recoveryCtx, ref) } return fmt.Errorf("committing authoritative applied state after replacing the fixed-port host workload: %w; the original workload restarted but its route could not be restored: %v; the uncommitted workload was restarted to avoid an outage", commitErr, err) } } - for _, name := range started { - d.docker.Remove(recoveryCtx, name) + for i, name := range started { + ref := name + if i < len(startedIDs) && startedIDs[i] != "" { + ref = startedIDs[i] + } + d.docker.Remove(recoveryCtx, ref) } if len(displaced) == 0 { return fmt.Errorf("committing authoritative applied state after starting the first host-ingress workload: %w; the uncommitted workload was stopped and removed", commitErr) @@ -1292,14 +1440,18 @@ func (d *Deployer) abortStateCommit(ctx context.Context, cfg Config, current *st } if cfg.usesCaddy() { - if err := d.restorePreviousRoute(recoveryCtx, cfg, current); err != nil { + if err := d.restorePreviousRoute(recoveryCtx, cfg, current, routeCAS, fromGeneration); err != nil { return fmt.Errorf("committing authoritative applied state after route switch: %w; restoring the previous route failed: %v; old and new workloads were left running to avoid routing to a stopped container", commitErr, err) } } - for _, name := range started { - d.docker.Stop(recoveryCtx, name, 5) - d.docker.Remove(recoveryCtx, name) + for i, name := range started { + ref := name + if i < len(startedIDs) && startedIDs[i] != "" { + ref = startedIDs[i] + } + d.docker.StopGenerationFenced(recoveryCtx, ref, 5, nextGeneration, "") + d.docker.Remove(recoveryCtx, ref) } if current == nil { return fmt.Errorf("committing authoritative applied state after route switch: %w; the new route was removed and the uncommitted workload was stopped", commitErr) @@ -1317,12 +1469,25 @@ func (d *Deployer) abortStateCommit(ctx context.Context, cfg Config, current *st // config drifted (and compensating to a drifted block is compensating to // the wrong route). Reconstruct-from-inspection remains only as the // documented fallback for legacy installs without a record — and it says so. -func (d *Deployer) restorePreviousRoute(ctx context.Context, cfg Config, current *state.AppState) error { +// +// C01-8/9 + A12/T05: the restored block is stamped with the generation +// being restored TO, and the commit runs under the exact-block CAS over +// routeCAS — the region this deploy resolved plus the region it switched +// to. This compensation is deliberately UNFENCED (A07: refusing to clean up +// one's own partial effects strands an app), which is exactly why its +// writes must be identity-fenced instead: a successor that took over and +// switched the route has a block in NEITHER set, and the restore refuses +// (ErrRouteCAS) rather than clobber the newer generation's route. +func (d *Deployer) restorePreviousRoute(ctx context.Context, cfg Config, current *state.AppState, routeCAS []string, generation uint64) error { + restoreCaddy := d.caddy.WithGeneration(generation) + if len(routeCAS) > 0 { + restoreCaddy = restoreCaddy.WithRouteCAS(cfg.App, routeCAS, generation) + } if current == nil || current.CurrentHash == "" { - return d.caddy.RemoveRoute(ctx, cfg.App) + return restoreCaddy.RemoveRoute(ctx, cfg.App) } if current.IngressMode != "" && current.IngressMode != "caddy" { - return d.caddy.RemoveRoute(ctx, cfg.App) + return restoreCaddy.RemoveRoute(ctx, cfg.App) } rec, recErr := releasemeta.Read(ctx, d.exec, cfg.App, current.CurrentHash) @@ -1333,14 +1498,14 @@ func (d *Deployer) restorePreviousRoute(ctx context.Context, cfg Config, current fmt.Fprintf(d.out, "Warning: no release record for %s@%s (pre-F14 install) — restoring the previous route from live inspection instead of the recorded receipt\n", cfg.App, current.CurrentHash) default: if port, ok := releasemeta.PrimaryContainerPort(rec); ok { - return d.restoreRouteFromReceipt(ctx, cfg, current, rec, port) + return d.restoreRouteFromReceipt(ctx, cfg, current, rec, port, restoreCaddy) } // A record without a designated primary port (a backfilled record // whose bindings identified none) cannot render the receipt's // upstream port; that piece falls back to inspection, loudly. fmt.Fprintf(d.out, "Warning: the release record for %s@%s names no primary container port — restoring the previous route from live inspection instead of the recorded receipt\n", cfg.App, current.CurrentHash) } - return d.restoreRouteFromInspection(ctx, cfg, current) + return d.restoreRouteFromInspection(ctx, cfg, current, restoreCaddy) } // restoreRouteFromReceipt renders the predecessor route from the recorded @@ -1351,7 +1516,7 @@ func (d *Deployer) restorePreviousRoute(ctx context.Context, cfg Config, current // carries them; a backfilled record cannot (nothing recoverable from // containers), and the CLI-passed config stays the fallback for it exactly // like rollback's applyRecordToRollback. -func (d *Deployer) restoreRouteFromReceipt(ctx context.Context, cfg Config, current *state.AppState, rec *releasemeta.Record, containerPort int) error { +func (d *Deployer) restoreRouteFromReceipt(ctx context.Context, cfg Config, current *state.AppState, rec *releasemeta.Record, containerPort int, cd *caddy.Client) error { replicas := rec.Replicas if replicas <= 0 { replicas = 1 @@ -1397,15 +1562,15 @@ func (d *Deployer) restoreRouteFromReceipt(ctx context.Context, cfg Config, curr } if replicas > 1 { - return d.caddy.SetLoadBalancerHealth(ctx, cfg.App, domain, upstreams, healthPath, tls, caddyExtra, cache, fw, access) + return cd.SetLoadBalancerHealth(ctx, cfg.App, domain, upstreams, healthPath, tls, caddyExtra, cache, fw, access) } - return d.caddy.SetRoute(ctx, cfg.App, domain, names[0], containerPort, tls, caddyExtra, cache, fw, access) + return cd.SetRoute(ctx, cfg.App, domain, names[0], containerPort, tls, caddyExtra, cache, fw, access) } // restoreRouteFromInspection is the legacy fallback (pre-F14 installs, or a // record that cannot name its route): reconstruct the previous block from // the current config plus a live inspect of the predecessor containers. -func (d *Deployer) restoreRouteFromInspection(ctx context.Context, cfg Config, current *state.AppState) error { +func (d *Deployer) restoreRouteFromInspection(ctx context.Context, cfg Config, current *state.AppState, cd *caddy.Client) error { replicas := len(current.CurrentPorts) if replicas == 0 { replicas = 1 @@ -1435,9 +1600,9 @@ func (d *Deployer) restoreRouteFromInspection(ctx context.Context, cfg Config, c } tls := caddy.TLS{Cert: cfg.TLSCert, Key: cfg.TLSKey, Internal: cfg.TLSInternal} if replicas > 1 { - return d.caddy.SetLoadBalancerHealth(ctx, cfg.App, domain, upstreams, cfg.Health.Path, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access) + return cd.SetLoadBalancerHealth(ctx, cfg.App, domain, upstreams, cfg.Health.Path, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access) } - return d.caddy.SetRoute(ctx, cfg.App, domain, names[0], primaryPort, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access) + return cd.SetRoute(ctx, cfg.App, domain, names[0], primaryPort, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access) } // logDeploy appends the terminal receipt for a deploy attempt. success diff --git a/internal/deploy/deploy_tcl01_tcl02_test.go b/internal/deploy/deploy_tcl01_tcl02_test.go index ad057a9..0e4d25b 100644 --- a/internal/deploy/deploy_tcl01_tcl02_test.go +++ b/internal/deploy/deploy_tcl01_tcl02_test.go @@ -40,7 +40,7 @@ func TestDeploy_AcquiresLockExactlyOnce(t *testing.T) { ssh.MockCommand{Match: "a=$(docker exec caddy md5sum", Output: "TEPLOY_CADDY_OK"}, ssh.MockCommand{Match: "UPLOAD:", Output: ""}, ssh.MockCommand{Match: "mv -f -- ", Output: ""}, - ssh.MockCommand{Match: "if [ ! -e '/deployments/caddy/Caddyfile'", Err: fmt.Errorf("none")}, + ssh.MockCommand{Match: "if [ ! -e '/deployments/caddy/Caddyfile'", Output: "present\n{\n\tadmin 0.0.0.0:2019\n}\n"}, ) var buf bytes.Buffer @@ -103,7 +103,7 @@ func TestDeploy_SameVersionCleanupNeverTouchesReplacement(t *testing.T) { ssh.MockCommand{Match: "a=$(docker exec caddy md5sum", Output: "TEPLOY_CADDY_OK"}, ssh.MockCommand{Match: "UPLOAD:", Output: ""}, ssh.MockCommand{Match: "mv -f -- ", Output: ""}, - ssh.MockCommand{Match: "if [ ! -e '/deployments/caddy/Caddyfile'", Err: fmt.Errorf("none")}, + ssh.MockCommand{Match: "if [ ! -e '/deployments/caddy/Caddyfile'", Output: "present\n{\n\tadmin 0.0.0.0:2019\n}\n"}, ) var buf bytes.Buffer diff --git a/internal/deploy/f14_wiring_test.go b/internal/deploy/f14_wiring_test.go index 2eaefce..7b354f1 100644 --- a/internal/deploy/f14_wiring_test.go +++ b/internal/deploy/f14_wiring_test.go @@ -467,7 +467,8 @@ func TestRollback_FixedPortTargetDisplacesCurrent(t *testing.T) { } stopIdx, runIdx := -1, -1 for i, c := range mock.Calls { - if stopIdx < 0 && strings.HasPrefix(c, "docker stop") && strings.Contains(c, "myapp-web-v2") { + // The displacement stop addresses the exact container ID (C01-9). + if stopIdx < 0 && strings.HasPrefix(c, "docker stop") && (strings.Contains(c, "myapp-web-v2") || strings.Contains(c, "'bbb'")) { stopIdx = i } if runIdx < 0 && strings.HasPrefix(c, "docker run") && strings.Contains(c, "myapp-web-v1") { diff --git a/internal/deploy/generation_integration_test.go b/internal/deploy/generation_integration_test.go new file mode 100644 index 0000000..2429f6c --- /dev/null +++ b/internal/deploy/generation_integration_test.go @@ -0,0 +1,261 @@ +//go:build integration + +// Fixture-gated verification of the C01-8/9 + A12/T05 acceptance against a +// REAL SSH+Docker host: generation labels read at EXECUTION time (a delayed +// stop command landing after the world moved on refuses against the live +// label), the committed-generation sidecar CAS on the real FS, and the +// exact-block route CAS against a real Caddyfile through the production +// caddy.Client (lock, adapt gate, guarded commit, reload, delivery +// verification). +// +// Same fixture contract as the other harnesses: +// +// TEPLOY_FAULT_HOST=127.0.0.1:50075 \ +// TEPLOY_FAULT_USER=tyler \ +// TEPLOY_FAULT_KEY=~/.colima/_lima/_config/user \ +// go test -tags integration -run TestGenerationIntegration -v ./internal/deploy +// +// Needs caddy:2-alpine + alpine:3 pullable; skips when /deployments/caddy +// already exists (a provisioned real caddy would conflict). Disposable +// fixture only: creates and removes genprobe-* containers plus +// /deployments/{caddy,genprobe}. +package deploy + +import ( + "context" + "errors" + "fmt" + "strings" + "testing" + "time" + + "github.com/useteploy/teploy/internal/caddy" + "github.com/useteploy/teploy/internal/docker" + "github.com/useteploy/teploy/internal/ssh" + "github.com/useteploy/teploy/internal/state" +) + +const genProbeApp = "genprobe" + +func genRun(t *testing.T, exec ssh.Executor, ctx context.Context, cmd string) string { + t.Helper() + out, err := exec.Run(ctx, cmd) + if err != nil { + t.Fatalf("%s: %v\n%s", cmd, err, out) + } + return out +} + +// TestGenerationIntegration_DelayedStopRefusesNewerGeneration proves the +// delayed-SSH-effect acceptance on real docker: a container labeled +// generation 9 (a successor's same-hash redeploy) is NOT stoppable by an +// operation prepared against generation 7 — including a stop command whose +// EXECUTION is delayed past the label change (nohup), because the label +// check reads the container's identity at execution time, not at issue +// time. Legacy unlabeled containers keep stopping (compat). +func TestGenerationIntegration_DelayedStopRefusesNewerGeneration(t *testing.T) { + host, user, key, image := reconcileFixtureEnv(t) + ctx, cancel := context.WithTimeout(context.Background(), 3*time.Minute) + defer cancel() + + exec, err := ssh.Connect(ctx, ssh.ConnectConfig{Host: host, User: user, KeyPath: key}) + if err != nil { + t.Fatalf("connecting: %v", err) + } + t.Cleanup(func() { exec.Close() }) + if out, err := exec.Run(ctx, "docker version --format '{{.Server.Version}}'"); err != nil { + t.Skipf("fixture host has no reachable docker daemon (%v)", err) + } else { + t.Logf("fixture docker server %s", strings.TrimSpace(out)) + } + + name := genProbeApp + "-web-abc123" + legacy := genProbeApp + "-web-legacy" + t.Cleanup(func() { + cctx, ccancel := context.WithTimeout(context.Background(), 30*time.Second) + defer ccancel() + exec.Run(cctx, "docker rm -f "+name+" "+legacy) + }) + + dk := docker.NewClient(exec) + // A newer generation's container under the name a stale rollback + // resolved (generation 7), and a legacy unlabeled one. + genRun(t, exec, ctx, fmt.Sprintf( + "docker run -d --name %s --label teploy.app=%s --label teploy.version=abc123 --label teploy.generation=9 %s sleep 300", name, genProbeApp, image)) + genRun(t, exec, ctx, fmt.Sprintf( + "docker run -d --name %s --label teploy.app=%s --label teploy.version=old %s sleep 300", legacy, genProbeApp, image)) + + // The generation check reads the LIVE label: expected 7 refuses 9. + if err := dk.StopGenerationFenced(ctx, name, 2, 7, ""); err == nil { + t.Fatal("stopping a generation-9 container prepared against 7 must refuse") + } else if !state.GenerationFenced(err) || !strings.Contains(err.Error(), "TEPLOY_GENERATION_FENCED 9 7") { + t.Fatalf("refusal must name both generations, got %v", err) + } + if st := genRun(t, exec, ctx, "docker inspect -f '{{.State.Status}}' "+name); strings.TrimSpace(st) != "running" { + t.Fatalf("the newer generation's container must keep running, state %s", st) + } + + // The DELAYED effect: the composed stop is issued now, executes in 3s — + // the world it executes against still refuses it (label at execution). + if out, err := exec.Run(ctx, fmt.Sprintf( + "nohup sh -c 'sleep 3; %s' >/dev/null 2>&1 &", generationStopShell(name, 7))); err != nil { + t.Fatalf("launching the delayed stop: %v (%s)", err, out) + } + time.Sleep(5 * time.Second) + if st := genRun(t, exec, ctx, "docker inspect -f '{{.State.Status}}' "+name); strings.TrimSpace(st) != "running" { + t.Fatalf("the delayed stale stop must have been refused at execution, state %s", st) + } + + // An operation prepared against the CURRENT generation stops it; a + // legacy unlabeled container stops too (compat). + if err := dk.StopGenerationFenced(ctx, name, 2, 9, ""); err != nil { + t.Fatalf("stopping at the container's own generation must succeed: %v", err) + } + if err := dk.StopGenerationFenced(ctx, legacy, 2, 7, ""); err != nil { + t.Fatalf("a legacy unlabeled container must keep stopping: %v", err) + } +} + +// generationStopShell renders the composed stop exactly as +// docker.StopGenerationFenced issues it (unexported there), for the +// delayed-effect fixture. +func generationStopShell(ref string, expected uint64) string { + return fmt.Sprintf( + `g=$(docker inspect -f '{{index .Config.Labels "teploy.generation"}}' '%s' 2>/dev/null || printf '0'); case "$g" in ''|*[!0-9]*) g=0;; esac; [ "$g" -le %d ] || { printf 'TEPLOY_GENERATION_FENCED %%s %%s\n' "$g" %d >&2; exit 74; }; docker stop -t 2 '%s'`, + ref, expected, expected, ref) +} + +// TestGenerationIntegration_CommitCASAndSidecar proves the authority CAS +// on the real FS: WriteFencedGeneration commits state.json + the +// .generation sidecar together, and a stale prepared-against-7 commit +// refuses over a committed 8 with state.json left to the successor. +func TestGenerationIntegration_CommitCASAndSidecar(t *testing.T) { + host, user, key, _ := reconcileFixtureEnv(t) + ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute) + defer cancel() + + exec, err := ssh.Connect(ctx, ssh.ConnectConfig{Host: host, User: user, KeyPath: key}) + if err != nil { + t.Fatalf("connecting: %v", err) + } + t.Cleanup(func() { exec.Close() }) + t.Cleanup(func() { + cctx, ccancel := context.WithTimeout(context.Background(), 30*time.Second) + defer ccancel() + exec.Run(cctx, "rm -rf /deployments/"+genProbeApp) + }) + genRun(t, exec, ctx, "rm -rf /deployments/"+genProbeApp+" && mkdir -p /deployments/"+genProbeApp) + + lk, err := state.AcquireLockFenced(ctx, exec, genProbeApp) + if err != nil { + t.Fatalf("AcquireLockFenced: %v", err) + } + defer state.ReleaseLockFenced(exec, lk, genProbeApp) + + mk := func(hash string, gen uint64) *state.AppState { + return &state.AppState{ + SchemaVersion: state.SchemaVersionV2, DeploymentType: "container", IngressMode: "caddy", + Domain: "gen.example.com", UpdatedAt: time.Now().UTC(), CurrentHash: hash, Generation: gen, + } + } + if err := state.WriteFencedGeneration(ctx, exec, genProbeApp, mk("v3", 8), lk, 7); err != nil { + t.Fatalf("clean commit over generation 7 must land: %v", err) + } + gen := genRun(t, exec, ctx, "cat /deployments/"+genProbeApp+"/.generation") + if strings.TrimSpace(gen) != "8" { + t.Fatalf("the sidecar must carry the committed generation, got %q", gen) + } + + // The stale rollback: prepared against 7, commits after the successor. + err = state.WriteFencedGeneration(ctx, exec, genProbeApp, mk("v1", 8), lk, 7) + if !state.GenerationFenced(err) { + t.Fatalf("a commit prepared against 7 must refuse over committed 8, got %v", err) + } + cur := genRun(t, exec, ctx, "cat /deployments/"+genProbeApp+"/state.json") + if !strings.Contains(cur, `"current_hash":"v3"`) { + t.Fatalf("a refused commit must leave the successor's state, got %s", cur) + } +} + +// TestGenerationIntegration_RouteCASAgainstRealCaddy proves the exact-block +// CAS through the production caddy.Client against a real Caddyfile + real +// caddy reload: a switch over the exact predecessor lands stamped with its +// generation; the same expectation replayed after a successor switched +// refuses with BOTH generations named and nothing lands. +func TestGenerationIntegration_RouteCASAgainstRealCaddy(t *testing.T) { + host, user, key, _ := reconcileFixtureEnv(t) + ctx, cancel := context.WithTimeout(context.Background(), 4*time.Minute) + defer cancel() + + exec, err := ssh.Connect(ctx, ssh.ConnectConfig{Host: host, User: user, KeyPath: key}) + if err != nil { + t.Fatalf("connecting: %v", err) + } + t.Cleanup(func() { exec.Close() }) + if _, err := exec.Run(ctx, "docker version --format '{{.Server.Version}}'"); err != nil { + t.Skipf("fixture host has no reachable docker daemon (%v)", err) + } + if out, _ := exec.Run(ctx, "test -e /deployments/caddy/Caddyfile && echo yes || echo no"); strings.TrimSpace(out) == "yes" { + t.Skip("fixture has a provisioned /deployments/caddy — remove it or use a disposable host for this fixture") + } + for _, image := range []string{"caddy:2-alpine"} { + if _, err := exec.Run(ctx, "docker pull "+image); err != nil { + t.Skipf("cannot pull %s (%v) — pre-pull it on the fixture", image, err) + } + } + + t.Cleanup(func() { + cctx, ccancel := context.WithTimeout(context.Background(), 60*time.Second) + defer ccancel() + exec.Run(cctx, "docker rm -f caddy") + exec.Run(cctx, "rm -rf /deployments/caddy") + }) + genRun(t, exec, ctx, "docker network create teploy 2>/dev/null || true") + genRun(t, exec, ctx, "mkdir -p /deployments/caddy") + if err := exec.Upload(ctx, strings.NewReader("{\n\tadmin 127.0.0.1:2019\n}\n"), "/deployments/caddy/Caddyfile", "0644"); err != nil { + t.Fatalf("seeding Caddyfile: %v", err) + } + genRun(t, exec, ctx, "docker run -d --name caddy --network teploy -v /deployments/caddy:/etc/caddy -p 127.0.0.1:0:80 caddy:2-alpine") + + cd := caddy.NewClient(exec) + setRoute := func(gen uint64, upstream string) error { + return cd.WithGeneration(gen).SetRoute(ctx, genProbeApp, "gen.example.com", upstream, 3000, caddy.TLS{}, "", nil, caddy.Firewall{}, caddy.Access{}) + } + + // Generation 7's route lands. + if err := setRoute(7, genProbeApp+"-web-v7"); err != nil { + t.Fatalf("first route: %v", err) + } + resolved, present, err := cd.ReadManagedBlock(ctx, genProbeApp) + if err != nil || !present { + t.Fatalf("ReadManagedBlock after first route: %q %v", resolved, err) + } + if g, ok := caddy.RegionGeneration(resolved); !ok || g != 7 { + t.Fatalf("the live block must be stamped generation 7, got %d %v", g, ok) + } + + // The successor (generation 8) switches the route. + if err := setRoute(8, genProbeApp+"-web-v8"); err != nil { + t.Fatalf("successor route: %v", err) + } + + // The stale operation replays ITS switch, CAS'd on the generation-7 + // region: refused, both generations named, nothing lands. + err = cd.WithGeneration(7). + WithRouteCAS(genProbeApp, []string{caddy.ManagedRegionHash(resolved)}, 7). + SetRoute(ctx, genProbeApp, "gen.example.com", genProbeApp+"-web-v7", 3000, caddy.TLS{}, "", nil, caddy.Firewall{}, caddy.Access{}) + var cas *caddy.ErrRouteCAS + if !errors.As(err, &cas) { + t.Fatalf("expected *ErrRouteCAS, got %v", err) + } + if cas.ExpectedGeneration != 7 || !cas.FoundStamped || cas.FoundGeneration != 8 { + t.Errorf("the refusal must name both generations: %+v", cas) + } + if len(cas.FoundUpstreams) != 1 || !strings.Contains(cas.FoundUpstreams[0], "web-v8") { + t.Errorf("the refusal must name the successor's upstreams: %+v", cas) + } + live := genRun(t, exec, ctx, "cat /deployments/caddy/Caddyfile") + if !strings.Contains(live, "genprobe-web-v8") || strings.Contains(live, "TEPLOY GENERATION 7") { + t.Errorf("a refused CAS must leave the successor's block live:\n%s", live) + } +} diff --git a/internal/deploy/generation_test.go b/internal/deploy/generation_test.go new file mode 100644 index 0000000..54aea13 --- /dev/null +++ b/internal/deploy/generation_test.go @@ -0,0 +1,357 @@ +package deploy + +// C01-8/9 + A12/T05 acceptance coverage: generation identity for rollback +// and the exact-block CAS, proven through the REAL Rollback and +// DeployFenced entry points over the mock executor (the mock evaluates the +// composed guards, the generation sidecar CAS, the container label checks +// and the route CAS against recorded file state — the same evaluation the +// server shell performs). The racing shapes: +// +// - deploy vs rollback: the successor's route switch lands between the +// rollback's resolution and its own switch — the exact-block CAS +// refuses with BOTH generations named; nothing lands. +// - rollback vs rollback: same shape, the successor being another +// rollback (its block names its own target). +// - generation bump mid-flight: the successor commits (sidecar + state) +// between the rollback's switch and its commit — the commit refuses +// (ErrGenerationFenced) and the authority keeps the successor. +// - a delayed SSH effect landing post-takeover: the retirement stop +// addresses the container by ID and reads its immutable generation +// label at execution — a newer generation's container is refused, no +// matter when the command lands. + +import ( + "bytes" + "context" + "encoding/json" + "errors" + "strings" + "testing" + "time" + + "github.com/useteploy/teploy/internal/caddy" + "github.com/useteploy/teploy/internal/ssh" + "github.com/useteploy/teploy/internal/state" +) + +// genRollbackMocks is the rollback fixture: state at generation 7 (current +// v2, previous v1), v1 stopped (ID aaa), v2 running (ID bbb), a Caddyfile +// whose myapp block is the generation-7 route to v2. +func genRollbackMocks(t *testing.T, caddyfile string) *ssh.MockExecutor { + t.Helper() + stateContent := `{"schema_version":2,"deployment_type":"container","ingress_mode":"caddy","domain":"myapp.com","updated_at":"2026-09-24T00:00:00Z","operation_id":"op-v2","generation":7,"current_port":49153,"current_hash":"v2","previous_port":49152,"previous_hash":"v1"}` + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "mkdir -p /deployments/myapp", Output: ""}, + ssh.MockCommand{Match: "mkdir /deployments/myapp/.lock", Output: ""}, + ssh.MockCommand{Match: "if [ ! -e '/deployments/myapp/state.json' ]", Output: "present\n" + stateContent}, + ssh.MockCommand{Match: "docker ps --all --filter label=teploy.app='myapp'", + Output: `{"ID":"aaa","Names":"myapp-web-v1","Image":"myapp:latest","State":"exited","Status":"Exited","Labels":"teploy.app=myapp,teploy.version=v1,teploy.process=web,teploy.generation=5"}` + "\n" + + `{"ID":"bbb","Names":"myapp-web-v2","Image":"myapp:latest","State":"running","Status":"Up 1h","Labels":"teploy.app=myapp,teploy.version=v2,teploy.process=web,teploy.generation=7"}`, + }, + ssh.MockCommand{Match: "docker inspect 'myapp-web-v1'", Output: `[{"Config":{"Image":"myapp:latest","Labels":{"teploy.app":"myapp","teploy.version":"v1","teploy.generation":"5"}},"HostConfig":{"NetworkMode":"teploy","PortBindings":{"3000/tcp":[{"HostIp":"127.0.0.1","HostPort":"49152"}]},"RestartPolicy":{"Name":"no"}},"NetworkSettings":{"Networks":{"teploy":{"Aliases":["myapp"]}}}}]`}, + ssh.MockCommand{Match: "docker rm -f 'myapp-web-v1'", Output: ""}, + ssh.MockCommand{Match: "docker run", Output: "aaa"}, + ssh.MockCommand{Match: "curl -s -o /dev/null", Output: "200"}, + ssh.MockCommand{Match: "docker inspect -f '{{range $p, $b := .NetworkSettings.Ports}}{{range $b}}{{.HostIp}}", Output: "127.0.0.1 "}, + ssh.MockCommand{Match: "docker inspect -f '{{range $p, $b := .NetworkSettings.Ports}}", Output: "49152"}, + ssh.MockCommand{Match: "docker inspect -f '{{range $p, $_ := .NetworkSettings.Ports}}", Output: "3000/tcp"}, + ssh.MockCommand{Match: "docker exec caddy caddy reload", Output: ""}, + ssh.MockCommand{Match: "a=$(docker exec caddy md5sum", Output: "TEPLOY_CADDY_OK"}, + ssh.MockCommand{Match: "mkdir /deployments/caddy/.lock", Output: ""}, + ssh.MockCommand{Match: "docker stop", Output: ""}, + ssh.MockCommand{Match: "docker rm ", Output: ""}, + ssh.MockCommand{Match: "cat /tmp", Output: ""}, + ssh.MockCommand{Match: "printf %s", Output: ""}, + ssh.MockCommand{Match: "rm -rf /deployments/myapp/.lock", Output: ""}, + ) + // The Caddyfile lives in Files (not a cat registration) so the mock's + // shell-mode CAS evaluation reads the same bytes the framed reads do. + mock.Files["/deployments/caddy/Caddyfile"] = []byte(caddyfile) + mock.Files["/deployments/myapp/.generation"] = []byte("7\n") + return mock +} + +// genV2Block is the generation-7 managed block routing to v2. +const genV2Block = "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 7\nmyapp.com {\n\treverse_proxy myapp-web-v2:3000\n}\n# TEPLOY END myapp" + +// successorBlock models a successor generation's switch: stamped 8, naming +// ITS containers (a deploy's v3 candidates, or another rollback's target — +// the CAS refuses either way, which is the point). +func successorBlock(upstream string) string { + return "# TEPLOY BEGIN myapp\n# TEPLOY GENERATION 8\nmyapp.com {\n\treverse_proxy " + upstream + ":3000\n}\n# TEPLOY END myapp" +} + +// raceSwapper wraps the mock, landing a successor's committed world (route +// block + state.json + generation sidecar) the moment the trigger command +// runs — modeling the racing client's commit landing mid-flight. +type raceSwapper struct { + *ssh.MockExecutor + trigger string + fired bool + successorFile string + successorState string +} + +func (r *raceSwapper) Run(ctx context.Context, cmd string) (string, error) { + if !r.fired && strings.Contains(cmd, r.trigger) { + r.fired = true + r.MockExecutor.Files["/deployments/caddy/Caddyfile"] = []byte(r.successorFile) + if r.successorState != "" { + r.MockExecutor.Files["/deployments/myapp/state.json"] = []byte(r.successorState) + } + r.MockExecutor.Files["/deployments/myapp/.generation"] = []byte("8\n") + } + return r.MockExecutor.Run(ctx, cmd) +} + +// TestRollback_RouteSwitchCASRefusesAfterSuccessorSwitched is the +// deploy-vs-rollback race: the successor's route switch (generation 8, +// naming its own containers) lands between the rollback's resolution and +// its switch. The exact-block CAS refuses — BOTH generations named, the +// successor's upstreams as evidence — nothing lands, and the target the +// rollback started is cleaned up. +func TestRollback_RouteSwitchCASRefusesAfterSuccessorSwitched(t *testing.T) { + for _, tc := range []struct { + name string + upstream string + successor string + }{ + {"deploy vs rollback", "myapp-web-v3", "a deploy switched to its v3 candidates"}, + {"rollback vs rollback", "myapp-web-v1", "another rollback switched back to its target first"}, + } { + t.Run(tc.name, func(t *testing.T) { + base := genRollbackMocks(t, genV2Block+"\n") + swap := &raceSwapper{ + MockExecutor: base, + trigger: "docker inspect 'myapp-web-v1'", // after resolve, before the switch + successorFile: successorBlock(tc.upstream) + "\n", + } + var buf bytes.Buffer + err := Rollback(context.Background(), swap, &buf, rollbackCfg()) + var cas *caddy.ErrRouteCAS + if !errors.As(err, &cas) { + t.Fatalf("expected *ErrRouteCAS from the refused switch, got %v\noutput:\n%s", err, buf.String()) + } + if cas.ExpectedGeneration != 7 || !cas.FoundStamped || cas.FoundGeneration != 8 { + t.Errorf("the refusal must name both generations: %+v", cas) + } + if len(cas.FoundUpstreams) == 0 || !strings.Contains(cas.FoundUpstreams[0], tc.upstream) { + t.Errorf("the refusal must name the successor's upstreams: %+v", cas) + } + if !swap.fired { + t.Fatal("test bug: the racing switch never fired") + } + // Nothing landed: the live block is still the successor's. + if got := string(base.Files["/deployments/caddy/Caddyfile"]); got != successorBlock(tc.upstream)+"\n" { + t.Errorf("a refused switch must not modify the Caddyfile, got:\n%s", got) + } + for _, c := range base.Calls { + if strings.HasPrefix(c, "docker exec caddy caddy reload") { + t.Error("no reload may run after a refused switch") + } + } + }) + } +} + +// TestRollback_CommitRefusedOverMidFlightGenerationBump is the +// generation-bump-mid-flight race: the rollback's switch lands, then the +// successor COMMITS (state.json + sidecar at generation 8) before the +// rollback's own commit. The commit's generation CAS refuses; the +// authority keeps the successor's state; the rollback's own route is +// restored (the CAS set still owns it) and its target is stopped. +func TestRollback_CommitRefusedOverMidFlightGenerationBump(t *testing.T) { + base := genRollbackMocks(t, genV2Block+"\n") + successorState := `{"schema_version":2,"deployment_type":"container","ingress_mode":"caddy","updated_at":"2026-09-24T00:01:00Z","operation_id":"successor","generation":8,"current_hash":"v3","previous_hash":"v2"}` + swap := &raceSwapper{ + MockExecutor: base, + trigger: "docker exec caddy caddy reload", // the rollback's switch landed + successorFile: "", // the successor did NOT touch the route here + successorState: successorState, + } + var buf bytes.Buffer + err := Rollback(context.Background(), swap, &buf, rollbackCfg()) + if !state.GenerationFenced(err) { + t.Fatalf("expected a generation-fenced commit refusal, got %v\noutput:\n%s", err, buf.String()) + } + if !strings.Contains(err.Error(), "expected predecessor 7") || !strings.Contains(err.Error(), "generation 8") { + t.Errorf("the refusal must name both generations: %v", err) + } + // The successor's authority is intact. + var st state.AppState + if err := json.Unmarshal(base.Files["/deployments/myapp/state.json"], &st); err != nil { + t.Fatalf("successor state.json unreadable: %v", err) + } + if st.Generation != 8 || st.CurrentHash != "v3" { + t.Errorf("a refused commit must not touch the successor's state: %+v", st) + } + if string(base.Files["/deployments/myapp/.generation"]) != "8\n" { + t.Errorf("the sidecar must keep the successor's generation: %q", base.Files["/deployments/myapp/.generation"]) + } + // The rollback's own uncommitted switch was undone: the live block is + // the generation-7 predecessor route again (stamped 7). + live := string(base.Files["/deployments/caddy/Caddyfile"]) + if !strings.Contains(live, "# TEPLOY GENERATION 7") || !strings.Contains(live, "myapp-web-v2:3000") { + t.Errorf("the CAS-owned route must be restored to the predecessor block, got:\n%s", live) + } +} + +// TestRollback_RetirementRefusesNewerGenerationContainer is the +// delayed-SSH-effect acceptance: the retirement stop command executes LONG +// after the world moved on (takeover + same-hash redeploy put a generation +// 8 container under the same name/ID the rollback resolved). The stop +// reads the container's immutable generation label at execution time and +// refuses — the newer generation's container keeps running, and the +// rollback reports the degraded retirement instead of silently yo-yo'ing. +func TestRollback_RetirementRefusesNewerGenerationContainer(t *testing.T) { + base := genRollbackMocks(t, genV2Block+"\n") + // The v2 container was re-created by generation 8 under the same ID. + base.GenerationLabels = map[string]string{"bbb": "8"} + var buf bytes.Buffer + if err := Rollback(context.Background(), base, &buf, rollbackCfg()); err != nil { + t.Fatalf("a refused retirement is degraded cleanup, not a failed rollback: %v\n%s", err, buf.String()) + } + out := buf.String() + if !strings.Contains(out, "newer generation") { + t.Errorf("the retirement refusal must be reported, got:\n%s", out) + } + // The newer generation's container was never stopped. + for _, c := range base.Calls { + if strings.HasPrefix(c, "docker stop") { + t.Errorf("a generation-fenced stop must never execute, saw: %s", c) + } + } + if !strings.Contains(out, "Rolled back myapp to version v1") { + t.Errorf("the rollback itself must complete, got:\n%s", out) + } +} + +// TestRollback_HappyPathStampsAndSidecar pins the happy-path identity +// surface: the rollback's switched block is stamped with the generation it +// creates (8), and its commit publishes the sidecar — the world after a +// successful rollback is fully generation-attributable. +func TestRollback_HappyPathStampsAndSidecar(t *testing.T) { + base := genRollbackMocks(t, genV2Block+"\n") + var buf bytes.Buffer + if err := Rollback(context.Background(), base, &buf, rollbackCfg()); err != nil { + t.Fatalf("Rollback: %v\n%s", err, buf.String()) + } + live := string(base.Files["/deployments/caddy/Caddyfile"]) + if !strings.Contains(live, "# TEPLOY GENERATION 8") || !strings.Contains(live, "myapp-web-v1:3000") { + t.Errorf("the rolled-back route must be stamped with the new generation:\n%s", live) + } + if gen := strings.TrimSpace(string(base.Files["/deployments/myapp/.generation"])); gen != "8" { + t.Errorf("the commit must publish the sidecar at the new generation, got %q", gen) + } +} + +// TestDeploy_RouteSwitchCASRefusesAfterSuccessorSwitched is the deploy-side +// twin: a successor's block lands between the deploy's resolution and its +// switch — the switch refuses with both generations named and the +// candidates are cleaned up. +func TestDeploy_RouteSwitchCASRefusesAfterSuccessorSwitched(t *testing.T) { + app := "fency" + // A standing generation-7 world: state.json names it (drop the fixture's + // "absent" state registration so the Files-seeded state is read). + mocks := make([]ssh.MockCommand, 0, len(fenceHappyPathMocks(app))+2) + for _, m := range fenceHappyPathMocks(app) { + if m.Match == "if [ ! -e '/deployments/"+app+"/state.json' ]" { + continue + } + mocks = append(mocks, m) + } + mocks = append(mocks, + ssh.MockCommand{Match: "docker stop", Output: ""}, + ssh.MockCommand{Match: "docker rm -f", Output: ""}, + ) + base := ssh.NewMockExecutor("1.2.3.4", mocks...) + base.Files["/deployments/"+app+"/state.json"] = []byte(`{"schema_version":2,"deployment_type":"container","ingress_mode":"caddy","updated_at":"2026-09-24T00:00:00Z","operation_id":"op-7","generation":7,"current_hash":"old","current_port":49152}`) + predBlock := "# TEPLOY BEGIN " + app + "\n# TEPLOY GENERATION 7\nfency.com {\n\treverse_proxy " + app + "-web-old:3000\n}\n# TEPLOY END " + app + base.Files["/deployments/caddy/Caddyfile"] = []byte("{\n\tadmin 0.0.0.0:2019\n}\n\n" + predBlock + "\n") + base.Files["/deployments/"+app+"/.generation"] = []byte("7\n") + lk, err := state.AcquireLockFenced(context.Background(), base, app) + if err != nil { + t.Fatalf("AcquireLockFenced: %v", err) + } + swap := &raceSwapper{ + MockExecutor: base, + trigger: "curl -s -o /dev/null", // the health gate: after resolve, before the switch + successorFile: "{\n\tadmin 0.0.0.0:2019\n}\n\n" + successorBlockFency() + "\n", + } + var buf bytes.Buffer + d := NewDeployer(swap, &buf) + err = d.DeployFenced(context.Background(), Config{ + App: app, + Domain: "fency.com", + Image: "fency:latest", + Version: "abc123", + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }, lk) + var cas *caddy.ErrRouteCAS + if !errors.As(err, &cas) { + t.Fatalf("expected *ErrRouteCAS from the refused switch, got %v\noutput:\n%s", err, buf.String()) + } + if cas.ExpectedGeneration != 7 || !cas.FoundStamped || cas.FoundGeneration != 8 { + t.Errorf("the refusal must name both generations: %+v", cas) + } + if got := string(base.Files["/deployments/caddy/Caddyfile"]); !strings.Contains(got, successorUpstreamFency) { + t.Errorf("a refused switch must leave the successor's block live, got:\n%s", got) + } + if !strings.Contains(buf.String(), "cleanup incomplete") && !anyCall(base, "docker stop") { + t.Errorf("the deploy must clean up its started candidates:\n%s", buf.String()) + } +} + +func successorBlockFency() string { + return "# TEPLOY BEGIN fency\n# TEPLOY GENERATION 8\nfency.com {\n\treverse_proxy " + successorUpstreamFency + ":3000\n}\n# TEPLOY END fency" +} + +const successorUpstreamFency = "fency-web-newgen" + +func anyCall(mock *ssh.MockExecutor, prefix string) bool { + for _, c := range mock.Calls { + if strings.HasPrefix(c, prefix) { + return true + } + } + return false +} + +// TestDeploy_HappyPathStampsGeneration pins the deploy-side identity +// surface: candidates are labeled with the generation the deploy creates +// and the switched block carries its stamp. +func TestDeploy_HappyPathStampsGeneration(t *testing.T) { + app := "fency" + base := ssh.NewMockExecutor("1.2.3.4", fenceHappyPathMocks(app)...) + lk, err := state.AcquireLockFenced(context.Background(), base, app) + if err != nil { + t.Fatalf("AcquireLockFenced: %v", err) + } + var buf bytes.Buffer + d := NewDeployer(base, &buf) + if err := d.DeployFenced(context.Background(), Config{ + App: app, + Domain: "fency.com", + Image: "fency:latest", + Version: "abc123", + Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, + }, lk); err != nil { + t.Fatalf("DeployFenced: %v\n%s", err, buf.String()) + } + var labeledRun bool + for _, c := range base.Calls { + if strings.HasPrefix(c, "docker run") && strings.Contains(c, "--label 'teploy.generation=1'") { + labeledRun = true + } + } + if !labeledRun { + t.Error("first deploy must stamp its candidates with generation 1") + } + live := string(base.Files["/deployments/caddy/Caddyfile"]) + if !strings.Contains(live, "# TEPLOY GENERATION 1") { + t.Errorf("the switched block must be stamped with the deploy's generation:\n%s", live) + } + if gen := strings.TrimSpace(string(base.Files["/deployments/"+app+"/.generation"])); gen != "1" { + t.Errorf("the commit must publish the sidecar, got %q", gen) + } +} diff --git a/internal/deploy/journal_snapshot_test.go b/internal/deploy/journal_snapshot_test.go index 6600d60..73c84a9 100644 --- a/internal/deploy/journal_snapshot_test.go +++ b/internal/deploy/journal_snapshot_test.go @@ -177,7 +177,7 @@ func TestPredecessorSnapshot_CrashRecoveryCompensatesExactlyRecordedIDs(t *testi d := NewDeployer(recovered, &out) err := d.abortStateCommit(context.Background(), Config{ App: "web", Image: "web:new", Version: "new456", Ingress: "host", ContainerPort: 3000, - }, &state.AppState{SchemaVersion: 2, CurrentHash: "old123", IngressMode: "host"}, nil, nil, att, time.Now(), errBoom) + }, &state.AppState{SchemaVersion: 2, CurrentHash: "old123", IngressMode: "host"}, nil, nil, nil, att, nil, time.Now(), errBoom) if err == nil { t.Fatal("expected the commit error to surface") } diff --git a/internal/deploy/recovery_a_test.go b/internal/deploy/recovery_a_test.go index 276991e..45415d6 100644 --- a/internal/deploy/recovery_a_test.go +++ b/internal/deploy/recovery_a_test.go @@ -114,7 +114,7 @@ func TestAbortStateCommit_CancelledContextStillRunsCompensation(t *testing.T) { ContainerPort: 8080, } started := []string{"fency-web-abc123"} - err := d.abortStateCommit(ctx, cfg, nil, started, nil, releasemeta.MustAttempt(app, cfg.Version), time.Now(), errBoom) + err := d.abortStateCommit(ctx, cfg, nil, started, nil, nil, releasemeta.MustAttempt(app, cfg.Version), nil, time.Now(), errBoom) if err == nil { t.Fatal("expected the commit error to be returned") } @@ -160,7 +160,7 @@ func TestAbortStateCommit_CaddyPublishAppRestoresRoute(t *testing.T) { Health: HealthConfig{Timeout: 5 * time.Second, Interval: 10 * time.Millisecond}, } started := []string{"fency-web-abc123"} - err := d.abortStateCommit(context.Background(), cfg, current, started, nil, releasemeta.MustAttempt(app, cfg.Version), time.Now(), errBoom) + err := d.abortStateCommit(context.Background(), cfg, current, started, nil, nil, releasemeta.MustAttempt(app, cfg.Version), nil, time.Now(), errBoom) if err == nil { t.Fatal("expected the commit error to surface") } diff --git a/internal/deploy/rollback.go b/internal/deploy/rollback.go index d3202bd..3055cf0 100644 --- a/internal/deploy/rollback.go +++ b/internal/deploy/rollback.go @@ -133,6 +133,35 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac return fmt.Errorf("target version %s is already current", target) } + // 2b. The rollback's GENERATION identity (C01-8/9): everything this + // rollback does is prepared against `fromGeneration` — the generation + // state.json names right now — and creates `nextGeneration`. Container + // names stay version-keyed, so a takeover + same-hash redeploy can put + // a NEWER generation's container under a name this rollback resolved; + // every destructive effect below (route switch, state commit, target + // restarts, retirement stops) is fenced on the generation identity — + // label checks and sidecar/CAS compares composed into the same remote + // command as the effect — so a stale rollback can never stop, remove, + // or overwrite a newer generation, no matter when its SSH commands + // land. + fromGeneration := current.Generation + nextGeneration := fromGeneration + 1 + // The exact-block route CAS expectation (A12/T05): the app's managed + // region as resolved NOW, under the lock. The route switch commits + // only if this is still the live region; a successor's block refuses + // the switch with both generations named. + resolvedRouteHash := "" + if cfg.usesCaddy() { + region, _, rerr := cd.ReadManagedBlock(ctx, cfg.App) + if rerr != nil { + return fmt.Errorf("refusing to roll back %s with an unreadable route authority: %w", cfg.App, rerr) + } + resolvedRouteHash = caddy.ManagedRegionHash(region) + } + // switchedRouteHash is re-resolved after a SUCCESSFUL switch: the block + // this rollback wrote but has not committed. + switchedRouteHash := resolvedRouteHash + fmt.Fprintf(out, "Rolling back %s from %s to %s...\n", cfg.App, current.CurrentHash, target) // 3. Find the target version's containers. This happens BEFORE anything @@ -230,12 +259,20 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac // restoreDisplaced brings back the fixed-port workload displaced by host // ingress. Every failure path after the displacement runs it — an error // that returns while the displaced workload is still stopped leaves a - // host-ingress app down (audit F12). + // host-ingress app down (audit F12). The recreate's force-remove is + // generation-checked (C01-9): a successor that re-created the name at a + // newer generation is left strictly alone; recovery of OUR effects must + // not destroy the new owner's workload. Holdership is deliberately NOT + // checked here (A07: recovery is never fenced). restoreDisplaced := func() { recoveryCtx, cancel := context.WithTimeout(context.Background(), 60*time.Second) defer cancel() for _, name := range displacedHostWeb { - if rerr := dk.Restart(recoveryCtx, name, nil); rerr != nil { + if rerr := dk.RestartFenced(recoveryCtx, name, nil, fromGeneration, ""); rerr != nil { + if state.GenerationFenced(rerr) { + fmt.Fprintf(out, " WARNING: could not restore %s after the failed rollback — a newer generation owns the name\n", name) + continue + } fmt.Fprintf(out, " WARNING: could not restore %s after the failed rollback: %v\n", name, rerr) } else { fmt.Fprintf(out, " Restored %s\n", name) @@ -243,11 +280,12 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac } } if fixedPorts { - // Fence (F16): stopping the fixed-port workload is this rollback's - // first destructive effect; nothing of ours needs restoring yet. - if err := lk.Check(ctx, exec); err != nil { - return err - } + // Fence + generation check composed into the stop (F16 + C01-8/9): + // stopping the fixed-port workload is this rollback's first + // destructive effect — the composed command refuses when the lock + // was broken (holdership) or the name now belongs to a newer + // generation's container (identity). Nothing of ours needs + // restoring yet, so a refusal is a plain abort. for _, c := range containers { // Displace only the RUNNING web containers of the AUTHORITATIVE // current generation (TCL-07). The old filter (any non-target @@ -263,7 +301,11 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac if c.State != "running" { continue } - if err := dk.Stop(ctx, c.Name, 10); err != nil { + ref := c.Name + if c.ID != "" { + ref = c.ID + } + if err := dk.StopGenerationFenced(ctx, ref, 10, fromGeneration, lk.GuardPrefix()); err != nil { // Earlier containers may already be stopped — restore them // before bailing (F12). restoreDisplaced() @@ -277,26 +319,31 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac var started []string var targetWeb []docker.Container for _, c := range targetContainers { - // Fence (F16): each target restart is an effect; a lost fence - // unwinds what this rollback started and restores the displaced - // workload before bailing (recovery is never fenced). - if err := lk.Check(ctx, exec); err != nil { - for _, name := range started { - dk.Stop(ctx, name, 5) - } - restoreDisplaced() - return err - } - fmt.Fprintf(out, "Starting %s...\n", c.Name) + // Fence + generation identity composed into the recreate (F16 + + // C01-8/9): the holdership guard and the generation label check + // ride the same remote command as the force-remove — the + // destructive half of Restart. A lost fence unwinds what this + // rollback started and restores the displaced workload before + // bailing (recovery is never holdership-fenced); a generation + // refusal means a newer generation now owns the name (takeover + + // same-hash redeploy) and the recreate is refused instead of + // destroying the successor's container. The recreated container + // PRESERVES its labels (Recreate re-emits them), so the target + // workload keeps the generation identity it was created under. + // // Recreate rather than `docker start`: Docker 29 silently fails // to re-publish HostConfig.PortBindings on `docker start` when // another container has taken+released the host port in the - // interim — a common case if rolling back after deploying a - // neighboring app that reused the port. Restart() inspects the - // stopped container, force-removes, and `docker run`s fresh - // with the same config, reallocating any port binding that - // collides with avoidPorts. - if err := dk.Restart(ctx, c.Name, avoidPorts); err != nil { + // interim — see docker.Client.Restart's doc comment. + fmt.Fprintf(out, "Starting %s...\n", c.Name) + if err := dk.RestartFenced(ctx, c.Name, avoidPorts, fromGeneration, lk.GuardPrefix()); err != nil { + if state.FenceLost(err) || state.GenerationFenced(err) { + for _, name := range started { + dk.Stop(ctx, name, 5) + } + restoreDisplaced() + return err + } // Stop anything this rollback already (re)started, then put the // displaced fixed-port workload back (F12). for _, name := range started { @@ -392,15 +439,16 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac // container Teploy starts (in this case, the target version's). if cfg.usesCaddy() { fmt.Fprintln(out, "Updating routes...") - // Fence (F16): the route switch commits traffic to the target — - // a late write would hijack a newer operation's route. - if err := lk.Check(ctx, exec); err != nil { - for _, name := range started { - dk.Stop(ctx, name, 5) - } - restoreDisplaced() - return err - } + // Fence (F16, C01-2 composition) + the exact-block CAS (A12/T05): + // the Caddyfile commit runs under this rollback's holdership guard + // AND a compare-and-swap on the managed region resolved at step 2b + // — stamped with the generation this rollback creates. A late + // write from a broken holder cannot hijack a newer operation's + // route, and a newer writer's block (a successor that already + // switched) refuses this switch with BOTH generations named + // instead of being overwritten. + cad := cd.WithGeneration(nextGeneration). + WithRouteCAS(cfg.App, []string{resolvedRouteHash}, fromGeneration) // failRoutePhase unwinds a route-phase failure the same way a // health/start failure unwinds (audit T06): the upstream-port // inspections and SetRoute/SetLoadBalancerHealth used to return @@ -433,8 +481,8 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac } upstreams = append(upstreams, caddy.Upstream{Dial: fmt.Sprintf("%s:%d", c.Name, port)}) } - if err := cd.SetLoadBalancerHealth(ctx, cfg.App, cfg.Domain, upstreams, healthCfg.Path, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access); err != nil { - return failRoutePhase(fmt.Errorf("updating load balancer route: %w", err)) + if err := cad.SetLoadBalancerHealth(ctx, cfg.App, cfg.Domain, upstreams, healthCfg.Path, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access); err != nil { + return failRoutePhase(routePhaseErr(err)) } fmt.Fprintf(out, " Traffic load-balanced across %d replicas\n", len(targetWeb)) } else { @@ -442,11 +490,16 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac if err != nil { return failRoutePhase(fmt.Errorf("inspecting target container port: %w", err)) } - if err := cd.SetRoute(ctx, cfg.App, cfg.Domain, targetWeb[0].Name, port, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access); err != nil { - return failRoutePhase(fmt.Errorf("updating route: %w", err)) + if err := cad.SetRoute(ctx, cfg.App, cfg.Domain, targetWeb[0].Name, port, tls, cfg.CaddyExtra, cfg.Cache, cfg.Firewall, cfg.Access); err != nil { + return failRoutePhase(routePhaseErr(err)) } fmt.Fprintln(out, " Traffic routed to target version") } + // The switched-to region is this rollback's own uncommitted block — + // the restore path's CAS accepts exactly {resolved, this}. + if region, _, rerr := cd.ReadManagedBlock(ctx, cfg.App); rerr == nil { + switchedRouteHash = caddy.ManagedRegionHash(region) + } } else { fmt.Fprintf(out, "Skipping Caddy route restore (ingress: %s)\n", cfg.Ingress) } @@ -488,9 +541,13 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac if digest, digestErr := dk.ContainerImageDigest(ctx, targetWeb[0].Name); digestErr == nil { newState.ImageDigest = digest } - // The commit runs under the fence (F16): the atomic rename that makes - // the rollback authoritative is a guarded effect. - if err := state.WriteFenced(ctx, exec, cfg.App, newState, lk); err != nil { + // The commit runs under the fence (F16) and the generation CAS + // (C01-8/9): the atomic rename that makes the rollback authoritative + // is a guarded effect chained with a compare-and-swap on the committed + // generation — a rollback that resolved the world at fromGeneration + // cannot commit over a successor's newer generation (ErrGenerationFenced + // naming both), so a stale rollback can never become authority. + if err := state.WriteFencedGeneration(ctx, exec, cfg.App, newState, lk, fromGeneration); err != nil { // Fixed host ports: the target holds them. Stop it, restore the // displaced workload, then remove the uncommitted target — previously // this branch skipped the restore and still claimed "the original @@ -506,7 +563,19 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac return fmt.Errorf("committing authoritative applied state after rollback: %w; the fixed-port workload was restored and the uncommitted target was removed", err) } if cfg.usesCaddy() { - if restoreErr := restoreRollbackRoute(ctx, cd, dk, cfg, current, containers); restoreErr != nil { + // The restore is CAS-fenced (A12/T05): it may only overwrite + // the block this rollback resolved or the one it switched to. + // When the commit was refused because a NEWER generation + // committed (ErrGenerationFenced) and that successor also + // switched the route, the restore refuses too — the newer + // generation's route is never clobbered by this rollback's + // compensation. When the route is still ours to undo, it is + // undone. + if restoreErr := restoreRollbackRoute(ctx, exec, out, cd, dk, cfg, current, containers, restoreRouteCAS(resolvedRouteHash, switchedRouteHash), fromGeneration); restoreErr != nil { + var routeCAS *caddy.ErrRouteCAS + if errors.As(restoreErr, &routeCAS) { + return fmt.Errorf("committing authoritative applied state after rollback route switch: %w; the route was left to the newer generation that owns it (%v)", err, routeCAS) + } return fmt.Errorf("committing authoritative applied state after rollback route switch: %w; restoring the original route failed: %v; original and target workloads were left running to avoid routing to a stopped container", err, restoreErr) } } @@ -533,15 +602,33 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac } } for _, c := range containers { - if lk != nil { - if err := lk.Check(ctx, exec); err != nil { + if c.Labels["teploy.version"] != current.CurrentHash || c.State != "running" { + continue + } + // C01-8/9: the retirement stop is a composed guarded effect — + // holdership guard + generation label check + stop, addressing the + // container by its exact ID from this rollback's under-lock + // inventory. A delayed command landing after a takeover reads the + // live container's immutable generation label: a newer + // generation's container (same names, takeover + same-hash + // redeploy) is REFUSED, never stopped — the property this slice + // exists for. A fence loss stops the sweep loudly (degraded, + // visible), as before. + ref := c.Name + if c.ID != "" { + ref = c.ID + } + fmt.Fprintf(out, "Stopping %s...\n", c.Name) + if err := dk.StopGenerationFenced(ctx, ref, stopTimeout, fromGeneration, lk.GuardPrefix()); err != nil { + if state.FenceLost(err) { fmt.Fprintf(out, "Warning: current-workload cleanup stopped — %v\n", err) break } - } - if c.Labels["teploy.version"] == current.CurrentHash && c.State == "running" { - fmt.Fprintf(out, "Stopping %s...\n", c.Name) - dk.Stop(ctx, c.Name, stopTimeout) + if state.GenerationFenced(err) { + fmt.Fprintf(out, "Warning: retirement of %s refused — the container belongs to a newer generation (%v)\n", c.Name, err) + continue + } + fmt.Fprintf(out, "Warning: could not stop %s: %v\n", c.Name, err) } } @@ -560,7 +647,112 @@ func Rollback(ctx context.Context, exec ssh.Executor, out io.Writer, cfg Rollbac return nil } -func restoreRollbackRoute(ctx context.Context, cd *caddy.Client, dk *docker.Client, cfg RollbackConfig, current *state.AppState, containers []docker.Container) error { +// restoreRollbackRoute compensates a failed rollback's route switch by +// putting the CURRENT (rolled-back-from) release's route back — the +// failed-rollback twin of deploy's restorePreviousRoute (C01-7's A12/T05 +// remainder). The F14 record of the version being rolled back FROM is the +// receipt of what was serving before the switch and is AUTHORITATIVE for +// the restore: domain, replica upstream names, the recorded primary +// container port, TLS/extra/cache/firewall/access, and the LB health path. +// Reconstruct-from-inspection remains only as the announced legacy +// fallback (pre-F14 installs, unreadable records, no primary port). +// +// C01-8/9 + A12/T05: the restored block is stamped with fromGeneration and +// the commit runs under the exact-block CAS over routeCAS — {the region +// this rollback resolved, the region it switched to}. The compensation is +// deliberately UNFENCED (A07), which is why its write must be +// identity-fenced: a successor's block is in NEITHER set and the restore +// refuses (caddy.ErrRouteCAS) instead of clobbering the newer generation's +// route — the acceptance shape for the delayed-effect race. +func restoreRollbackRoute(ctx context.Context, exec ssh.Executor, out io.Writer, cd *caddy.Client, dk *docker.Client, cfg RollbackConfig, current *state.AppState, containers []docker.Container, routeCAS []string, generation uint64) error { + restoreCaddy := cd.WithGeneration(generation) + if len(routeCAS) > 0 { + restoreCaddy = restoreCaddy.WithRouteCAS(cfg.App, routeCAS, generation) + } + + rec, recErr := releasemeta.Read(ctx, exec, cfg.App, current.CurrentHash) + switch { + case recErr != nil: + fmt.Fprintf(out, "Warning: the release record for %s@%s could not be read (%v) — restoring the original route from live inspection instead of the recorded receipt\n", cfg.App, current.CurrentHash, recErr) + case rec == nil: + fmt.Fprintf(out, "Warning: no release record for %s@%s (pre-F14 install) — restoring the original route from live inspection instead of the recorded receipt\n", cfg.App, current.CurrentHash) + default: + if port, ok := releasemeta.PrimaryContainerPort(rec); ok { + return restoreRollbackRouteFromReceipt(ctx, restoreCaddy, cfg, current, rec, port) + } + fmt.Fprintf(out, "Warning: the release record for %s@%s names no primary container port — restoring the original route from live inspection instead of the recorded receipt\n", cfg.App, current.CurrentHash) + } + return restoreRollbackRouteFromInspection(ctx, restoreCaddy, dk, cfg, current, containers) +} + +// routePhaseErr keeps fence/CAS refusals classified instead of wrapping +// them into generic route failures the caller cannot distinguish. +func routePhaseErr(err error) error { + var routeCAS *caddy.ErrRouteCAS + if errors.As(err, &routeCAS) { + return routeCAS + } + if state.FenceLost(err) { + return fmt.Errorf("switching the route back: %w", err) + } + return fmt.Errorf("updating route: %w", err) +} + +// restoreRollbackRouteFromReceipt renders the rolled-back-from release's +// route from its F14 record — zero live inspect. The upstream NAMES are +// deterministic per release; rollback does not rename the from-version's +// containers, so no _replaced suffix applies here. +func restoreRollbackRouteFromReceipt(ctx context.Context, cd *caddy.Client, cfg RollbackConfig, current *state.AppState, rec *releasemeta.Record, containerPort int) error { + replicas := rec.Replicas + if replicas <= 0 { + replicas = 1 + } + names := make([]string, replicas) + upstreams := make([]caddy.Upstream, replicas) + for i := range replicas { + name := docker.ReplicaContainerName(cfg.App, "web", current.CurrentHash, i+1, replicas) + names[i] = name + upstreams[i] = caddy.Upstream{Dial: fmt.Sprintf("%s:%d", name, containerPort)} + } + + domain := rec.Domain + if domain == "" { + domain = current.Domain + } + if domain == "" { + domain = cfg.Domain + } + + tls := caddy.TLS{Cert: cfg.TLSCert, Key: cfg.TLSKey, Internal: cfg.TLSInternal} + caddyExtra := cfg.CaddyExtra + cache := cfg.Cache + fw := cfg.Firewall + access := cfg.Access + if rec.Caddy != nil { + tls = caddy.TLS{Cert: rec.Caddy.TLSCert, Key: rec.Caddy.TLSKey, Internal: rec.Caddy.TLSInternal} + caddyExtra = rec.Caddy.CaddyExtra + cache = rec.Caddy.Cache + if rec.Caddy.Firewall != nil { + fw = *rec.Caddy.Firewall + } + if rec.Caddy.Access != nil { + access = *rec.Caddy.Access + } + } + healthPath := cfg.Health.withDefaults().Path + if rec.Health != nil && rec.Health.Path != "" { + healthPath = rec.Health.Path + } + + if replicas > 1 { + return cd.SetLoadBalancerHealth(ctx, cfg.App, domain, upstreams, healthPath, tls, caddyExtra, cache, fw, access) + } + return cd.SetRoute(ctx, cfg.App, domain, names[0], containerPort, tls, caddyExtra, cache, fw, access) +} + +// restoreRollbackRouteFromInspection is the announced legacy fallback: +// reconstruct the original block from the running from-version containers. +func restoreRollbackRouteFromInspection(ctx context.Context, cd *caddy.Client, dk *docker.Client, cfg RollbackConfig, current *state.AppState, containers []docker.Container) error { var currentWeb []docker.Container for _, container := range containers { if container.Labels["teploy.version"] == current.CurrentHash && container.Labels["teploy.process"] == "web" && container.State == "running" { diff --git a/internal/deploy/rollback_test.go b/internal/deploy/rollback_test.go index 2e667a3..528dd3e 100644 --- a/internal/deploy/rollback_test.go +++ b/internal/deploy/rollback_test.go @@ -107,12 +107,15 @@ func TestRollback(t *testing.T) { t.Errorf("rollback route should not use host port 49152, got: %s", string(caddyfile)) } - // Verify current container was stopped. + // Verify the current container was stopped — by its exact ID from the + // under-lock inventory (C01-9), inside the composed guarded stop. Only + // the FINAL (fully guard-stripped) command form prefixes with + // "docker stop"; the composed and intermediate forms are recorded too. stopCalls := 0 for _, call := range mock.Calls { if strings.HasPrefix(call, "docker stop") { stopCalls++ - if !strings.Contains(call, "myapp-web-v2") { + if !strings.Contains(call, "'bbb'") && !strings.Contains(call, "myapp-web-v2") { t.Errorf("expected stop for v2 container, got: %s", call) } } @@ -432,7 +435,8 @@ func TestRollback_ToSpecificHash(t *testing.T) { // the target) must be left completely untouched. var stoppedV3, touchedV2 bool for _, c := range mock.Calls { - if strings.HasPrefix(c, "docker stop") && strings.Contains(c, "myapp-web-v3") { + // The retirement stop addresses the exact container ID (C01-9). + if strings.HasPrefix(c, "docker stop") && (strings.Contains(c, "myapp-web-v3") || strings.Contains(c, "'ccc'")) { stoppedV3 = true } if strings.Contains(c, "myapp-web-v2") { @@ -616,7 +620,8 @@ func TestRollback_HostIngressKeepsTheFixedPort(t *testing.T) { // would try to bind the same fixed port. stopIdx, runIdx := -1, -1 for i, c := range mock.Calls { - if stopIdx < 0 && strings.Contains(c, "docker stop") && strings.Contains(c, "myapp-web-v2") { + // The displacement stop addresses the exact container ID (C01-9). + if stopIdx < 0 && strings.HasPrefix(c, "docker stop") && (strings.Contains(c, "myapp-web-v2") || strings.Contains(c, "'bbb'")) { stopIdx = i } if runIdx < 0 && strings.Contains(c, "docker run") { diff --git a/internal/deploy/route_receipt_test.go b/internal/deploy/route_receipt_test.go index 2bfc44a..b3e9aab 100644 --- a/internal/deploy/route_receipt_test.go +++ b/internal/deploy/route_receipt_test.go @@ -51,7 +51,7 @@ func TestRestorePreviousRoute_RendersRouteFromRecordedReceipt(t *testing.T) { current := &state.AppState{SchemaVersion: 2, CurrentHash: "old123", Domain: "live.example.com", CurrentPorts: []int{49152}} cfg := Config{App: "myapp", Domain: "cfg.example.com", Version: "new456"} - if err := d.restorePreviousRoute(context.Background(), cfg, current); err != nil { + if err := d.restorePreviousRoute(context.Background(), cfg, current, nil, 0); err != nil { t.Fatalf("restorePreviousRoute: %v", err) } @@ -91,7 +91,7 @@ func TestRestorePreviousRoute_SameVersionUsesReplacedNaming(t *testing.T) { current := &state.AppState{SchemaVersion: 2, CurrentHash: "old123", CurrentPorts: []int{49152}} cfg := Config{App: "myapp", Domain: "myapp.com", Version: "old123"} // same version - if err := d.restorePreviousRoute(context.Background(), cfg, current); err != nil { + if err := d.restorePreviousRoute(context.Background(), cfg, current, nil, 0); err != nil { t.Fatalf("restorePreviousRoute: %v", err) } caddyfile := string(mock.Files["/deployments/caddy/Caddyfile"]) @@ -114,7 +114,7 @@ func TestRestorePreviousRoute_NoRecordFallsBackLoudly(t *testing.T) { current := &state.AppState{SchemaVersion: 2, CurrentHash: "old123", Domain: "live.example.com", CurrentPorts: []int{49152}} cfg := Config{App: "myapp", Domain: "cfg.example.com", Version: "new456"} - if err := d.restorePreviousRoute(context.Background(), cfg, current); err != nil { + if err := d.restorePreviousRoute(context.Background(), cfg, current, nil, 0); err != nil { t.Fatalf("restorePreviousRoute: %v", err) } @@ -149,7 +149,7 @@ func TestRestorePreviousRoute_RecordDisagreesWithLiveInspect_RecordWins(t *testi current := &state.AppState{SchemaVersion: 2, CurrentHash: "old123", Domain: "live.example.com", CurrentPorts: []int{49152, 49153}} cfg := Config{App: "myapp", Domain: "cfg.example.com", Version: "new456"} - if err := d.restorePreviousRoute(context.Background(), cfg, current); err != nil { + if err := d.restorePreviousRoute(context.Background(), cfg, current, nil, 0); err != nil { t.Fatalf("restorePreviousRoute: %v", err) } diff --git a/internal/docker/docker.go b/internal/docker/docker.go index 6d2ba65..a49369e 100644 --- a/internal/docker/docker.go +++ b/internal/docker/docker.go @@ -57,6 +57,16 @@ type RunConfig struct { CPU string // CPU limit, e.g. "1.0" Name string // explicit container name (overrides auto-generated) NoHealthcheck bool // pass --no-healthcheck so the container ignores the image HEALTHCHECK + // Generation stamps the teploy.generation label — the identity of the + // deployment generation this container is created for (C01-9's + // generation-keyed identity sub-slice). Because container NAMES remain + // version-keyed, two attempts of one release hash share names; the + // label is the immutable discriminator a stale operation can fence on + // (StopGenerationFenced): a container created by a NEWER generation is + // never removable by an operation prepared against an older one, no + // matter when its commands land. Zero means legacy/unlabeled (skipped) + // — pre-label containers keep working everywhere. + Generation uint64 } // publishBinding renders a docker -p binding "[ip:]host:container" with @@ -172,12 +182,17 @@ func (c *Client) run(ctx context.Context, cfg RunConfig, guardPrefix string) (st args = append(args, "--network-alias", q(cfg.App+"-"+cfg.Process)) } - // Labels for filtering containers by app, process, and version. + // Labels for filtering containers by app, process, and version. The + // generation label (when set) keys the container to the deployment + // generation that created it — the identity stale operations fence on. args = append(args, "--label", q("teploy.app="+cfg.App), "--label", q("teploy.process="+cfg.Process), "--label", q("teploy.version="+cfg.Version), ) + if cfg.Generation > 0 { + args = append(args, "--label", q("teploy.generation="+strconv.FormatUint(cfg.Generation, 10))) + } // Port publishing and PORT env var injection. if cfg.Port > 0 { @@ -336,6 +351,78 @@ func (c *Client) Stop(ctx context.Context, name string, timeout int) error { return nil } +// GenerationLabelName is the label that keys a container to the deployment +// generation that created it (RunConfig.Generation). +const GenerationLabelName = "teploy.generation" + +// generationLostMarker mirrors state's generation-fence markers (unexported +// there): the stderr sentinel a generation-checked effect emits on refusal, +// matched by state.GenerationFenced over the error text. +const generationLostMarker = "TEPLOY_GENERATION_FENCED" + +// generationCheckFragment renders the shell that refuses (marker on stderr, +// exit 74) when ref's teploy.generation label names a generation NEWER than +// expected — composed into the SAME command as the effect it guards, so an +// in-flight command landing after a takeover still sees the newer label and +// refuses: unlike holdership guards this check is over the target's OWN +// immutable identity (a container's generation label never changes after +// creation), which is what closes the delayed-SSH-effect window. A missing +// container or an unlabeled (legacy) container passes — generation 0. +func generationCheckFragment(ref string, expected uint64) string { + return fmt.Sprintf( + `g=$(docker inspect -f '{{index .Config.Labels %q}}' %s 2>/dev/null || printf '0'); case "$g" in ''|*[!0-9]*) g=0;; esac; [ "$g" -le %d ] || { printf '%s %%s %%s\n' "$g" %d >&2; exit 74; }; `, + GenerationLabelName, ssh.ShellQuote(ref), expected, generationLostMarker, expected, + ) +} + +// StopGenerationFenced stops the container named by ref (a name or — the +// exact form — a container ID from this operation's own under-lock +// inventory) only when its generation label does not name a NEWER +// generation than expectedGeneration (C01-8/9). The label check, the +// optional holdership guard prefix, and the stop are ONE remote command: +// a stale rollback's retirement stop cannot kill a newer generation's +// container even when the command lands long after a takeover. Refusal +// errors match state.GenerationFenced and name both generations. +func (c *Client) StopGenerationFenced(ctx context.Context, ref string, timeout int, expectedGeneration uint64, guardPrefix string) error { + cmd := guardPrefix + generationCheckFragment(ref, expectedGeneration) + + fmt.Sprintf("docker stop -t %d %s", timeout, ssh.ShellQuote(ref)) + _, err := c.exec.Run(ctx, cmd) + if err != nil { + if strings.Contains(err.Error(), generationLostMarker) || stateGenerationRefused(err) { + return fmt.Errorf("stopping %s refused: the container belongs to a generation newer than this operation prepared against (expected %d) — %w", ref, expectedGeneration, err) + } + return fmt.Errorf("stopping container %s: %w", ref, err) + } + return nil +} + +// stateGenerationRefused reports whether the executor surfaced the +// generation check's exit status (74) without the marker text. +func stateGenerationRefused(err error) bool { + msg := err.Error() + return strings.Contains(msg, "status 74") || strings.Contains(msg, "exit status 74") +} + +// ContainerGeneration reads ref's generation label. The second result is +// false when the container carries no label (legacy) or no longer exists. +func (c *Client) ContainerGeneration(ctx context.Context, ref string) (uint64, bool, error) { + out, err := c.exec.Run(ctx, fmt.Sprintf( + "docker inspect -f '{{index .Config.Labels %q}}' %s 2>/dev/null || true", + GenerationLabelName, ssh.ShellQuote(ref))) + if err != nil { + return 0, false, fmt.Errorf("inspecting %s's generation label: %w", ref, err) + } + v := strings.TrimSpace(out) + if v == "" { + return 0, false, nil + } + gen, perr := strconv.ParseUint(v, 10, 64) + if perr != nil { + return 0, false, fmt.Errorf("container %s carries an unreadable generation label %q", ref, v) + } + return gen, true, nil +} + // Exec runs a command inside a running container via docker exec. func (c *Client) Exec(ctx context.Context, name, command string) (string, error) { // Single-quote for the REMOTE shell so it doesn't expand $/backticks before diff --git a/internal/docker/generation_test.go b/internal/docker/generation_test.go new file mode 100644 index 0000000..f66b774 --- /dev/null +++ b/internal/docker/generation_test.go @@ -0,0 +1,159 @@ +package docker + +import ( + "context" + "strings" + "testing" + + "github.com/useteploy/teploy/internal/ssh" +) + +// TestRun_StampsGenerationLabel pins the C01-9 identity surface: a RunConfig +// with a generation labels the container teploy.generation=; zero keeps +// the legacy unlabeled shape. +func TestRun_StampsGenerationLabel(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker run", Output: "abc123"}, + ) + dk := NewClient(mock) + if _, err := dk.Run(context.Background(), RunConfig{ + App: "myapp", Process: "web", Version: "v1", Image: "img:1", + Generation: 8, + }); err != nil { + t.Fatalf("Run: %v", err) + } + var sawLabel bool + for _, c := range mock.Calls { + if strings.HasPrefix(c, "docker run") && strings.Contains(c, "--label 'teploy.generation=8'") { + sawLabel = true + } + } + if !sawLabel { + t.Errorf("the run must stamp the generation label:\n%s", mock.Calls[0]) + } + + mock2 := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker run", Output: "abc123"}, + ) + if _, err := NewClient(mock2).Run(context.Background(), RunConfig{ + App: "myapp", Process: "web", Version: "v1", Image: "img:1", + }); err != nil { + t.Fatalf("Run: %v", err) + } + if strings.Contains(mock2.Calls[0], "teploy.generation") { + t.Errorf("zero generation must keep the legacy unlabeled shape:\n%s", mock2.Calls[0]) + } +} + +// TestStopGenerationFenced_ComposedCommand pins the C01-8/9 shape: the +// holdership guard, the generation label check, and the stop are ONE remote +// command — no transport window between check and effect. +func TestStopGenerationFenced_ComposedCommand(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker stop", Output: ""}, + ) + mock.Files["/deployments/myapp/.lock/info"] = []byte(`{"type":"auto","owner":"own"}`) + guard := "grep -q 'own' '/deployments/myapp/.lock/info' || { printf 'TEPLOY_FENCE_LOST\\n' >&2; exit 75; }; " + dk := NewClient(mock) + if err := dk.StopGenerationFenced(context.Background(), "myapp-web-v1", 10, 7, guard); err != nil { + t.Fatalf("StopGenerationFenced: %v", err) + } + cmd := mock.Calls[0] + for _, want := range []string{ + "grep -q 'own'", + `docker inspect -f '{{index .Config.Labels "teploy.generation"}}'`, + `[ "$g" -le 7 ]`, + "docker stop -t 10 'myapp-web-v1'", + } { + if !strings.Contains(cmd, want) { + t.Errorf("composed stop missing %q:\n%s", want, cmd) + } + } + if !strings.HasPrefix(cmd, "grep -q ") { + t.Errorf("the holdership guard must prefix the whole command:\n%s", cmd) + } +} + +// TestStopGenerationFenced_RefusesNewerGeneration is the delayed-effect +// acceptance at the docker boundary: the name now belongs to a container +// created by a NEWER generation (takeover + same-hash redeploy) — the stop +// is refused by the label check and never executes, no matter when the +// command lands. Legacy unlabeled containers keep stopping (compat). +func TestStopGenerationFenced_RefusesNewerGeneration(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker stop", Output: ""}, + ) + mock.GenerationLabels = map[string]string{"myapp-web-v1": "9"} + dk := NewClient(mock) + err := dk.StopGenerationFenced(context.Background(), "myapp-web-v1", 10, 7, "") + if err == nil || !strings.Contains(err.Error(), "TEPLOY_GENERATION_FENCED 9 7") { + t.Fatalf("expected a generation refusal naming both generations, got %v", err) + } + for _, c := range mock.Calls { + if strings.HasPrefix(c, "docker stop") { + t.Fatalf("a refused stop must never execute, saw: %s", c) + } + } + + // Legacy container (no label) stops as before. + mock2 := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker stop", Output: ""}, + ) + if err := NewClient(mock2).StopGenerationFenced(context.Background(), "myapp-web-v1", 10, 7, ""); err != nil { + t.Fatalf("a legacy unlabeled container must keep stopping: %v", err) + } +} + +// TestRestartFenced_GuardsTheRemove pins the recreate composition: the +// force-remove — the destructive half of a rollback's target restart — +// carries the generation label check in the same shell, so a stale +// rollback cannot destroy a successor's same-named container. +func TestRestartFenced_GuardsTheRemove(t *testing.T) { + inspect := `[{"Config":{"Image":"myapp:latest","Labels":{"teploy.app":"myapp","teploy.generation":"3"}},"HostConfig":{"NetworkMode":"teploy","RestartPolicy":{"Name":"no"}},"NetworkSettings":{"Networks":{"teploy":{"Aliases":["myapp"]}}}}]` + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker inspect 'myapp-web-v1'", Output: inspect}, + ssh.MockCommand{Match: "docker rm -f 'myapp-web-v1'", Output: ""}, + ssh.MockCommand{Match: "docker run", Output: "newid"}, + ) + dk := NewClient(mock) + if err := dk.RestartFenced(context.Background(), "myapp-web-v1", nil, 7, ""); err != nil { + t.Fatalf("RestartFenced: %v", err) + } + var guardedRemove bool + for _, c := range mock.Calls { + if strings.HasPrefix(c, "g=$(docker inspect") && strings.Contains(c, "docker rm -f 'myapp-web-v1'") { + guardedRemove = true + } + } + if !guardedRemove { + t.Error("the recreate's force-remove must run composed with the generation check") + } + // The recreated container PRESERVES its labels (generation 3 rides along). + for _, c := range mock.Calls { + if strings.HasPrefix(c, "docker run") && !strings.Contains(c, "--label 'teploy.generation=3'") { + t.Errorf("recreate must preserve the generation label:\n%s", c) + } + } +} + +// TestRestartFenced_RefusesNewerGeneration: the name was re-created by a +// newer generation; the remove is refused and never executes. +func TestRestartFenced_RefusesNewerGeneration(t *testing.T) { + inspect := `[{"Config":{"Image":"myapp:latest","Labels":{"teploy.app":"myapp"}},"HostConfig":{"NetworkMode":"teploy","RestartPolicy":{"Name":"no"}},"NetworkSettings":{"Networks":{"teploy":{"Aliases":["myapp"]}}}}]` + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "docker inspect 'myapp-web-v1'", Output: inspect}, + ssh.MockCommand{Match: "docker rm -f 'myapp-web-v1'", Output: ""}, + ssh.MockCommand{Match: "docker run", Output: "newid"}, + ) + mock.GenerationLabels = map[string]string{"myapp-web-v1": "9"} + dk := NewClient(mock) + err := dk.RestartFenced(context.Background(), "myapp-web-v1", nil, 7, "") + if err == nil || !strings.Contains(err.Error(), "TEPLOY_GENERATION_FENCED 9 7") { + t.Fatalf("expected a generation refusal naming both generations, got %v", err) + } + for _, c := range mock.Calls { + if strings.HasPrefix(c, "docker rm -f") || strings.HasPrefix(c, "docker run") { + t.Fatalf("a refused recreate must never remove or run, saw: %s", c) + } + } +} diff --git a/internal/docker/recreate.go b/internal/docker/recreate.go index f0a3d9a..870c8e9 100644 --- a/internal/docker/recreate.go +++ b/internal/docker/recreate.go @@ -543,9 +543,38 @@ func recreatePublishBinding(b RecreateBinding) (string, error) { // Recreate force-removes the named container and runs a fresh one from the // spec. avoidPorts is the set of host ports currently held by containers // this recreation must not collide with; a binding whose original port is -// in the set gets a freshly allocated one (see Restart's doc comment for the -// rollback collision this exists for). +// in the set gets a freshly allocated one (see Restart's doc comment for +// the rollback collision this exists for). func (c *Client) Recreate(ctx context.Context, spec *RecreateSpec, avoidPorts map[int]bool) error { + return c.recreate(ctx, spec, avoidPorts, "") +} + +// RecreateGuarded is Recreate with the force-remove composed under a guard +// prefix (holdership and/or generation check): the rm — the destructive +// half of a recreate — is refused in-shell when the guard fails, so a +// stale holder cannot remove a container the successor now owns. +func (c *Client) RecreateGuarded(ctx context.Context, spec *RecreateSpec, avoidPorts map[int]bool, rmGuardPrefix string) error { + return c.recreate(ctx, spec, avoidPorts, rmGuardPrefix) +} + +// RestartFenced is Restart with the force-remove composed under BOTH the +// holdership guard prefix (state.Lock.GuardPrefix, when non-empty) and the +// generation label check (C01-8/9): recreating a stopped rollback target +// removes the container under its name first, and under a takeover with a +// same-hash redeploy that NAME may now belong to a NEWER generation's +// container — the label check refuses exactly that, in the same shell as +// the rm, so a stale rollback can never destroy the successor's workload +// even when its commands land post-takeover. Refusals match +// state.GenerationFenced (label check) and state.FenceLost (guard). +func (c *Client) RestartFenced(ctx context.Context, name string, avoidPorts map[int]bool, expectedGeneration uint64, guardPrefix string) error { + spec, err := c.InspectRecreate(ctx, name) + if err != nil { + return err + } + return c.recreate(ctx, spec, avoidPorts, guardPrefix+generationCheckFragment(name, expectedGeneration)) +} + +func (c *Client) recreate(ctx context.Context, spec *RecreateSpec, avoidPorts map[int]bool, rmGuardPrefix string) error { if spec == nil || spec.Name == "" { return fmt.Errorf("recreate requires a spec with a container name") } @@ -603,7 +632,10 @@ func (c *Client) Recreate(ctx context.Context, spec *RecreateSpec, avoidPorts ma return err } - if _, err := c.exec.Run(ctx, "docker rm -f "+ssh.ShellQuote(spec.Name)); err != nil { + if _, err := c.exec.Run(ctx, rmGuardPrefix+"docker rm -f "+ssh.ShellQuote(spec.Name)); err != nil { + if strings.Contains(err.Error(), generationLostMarker) || stateGenerationRefused(err) { + return fmt.Errorf("recreating %s refused: the name now belongs to a container of a newer generation than this operation prepared against — %w", spec.Name, err) + } return fmt.Errorf("removing old %s: %w", spec.Name, err) } if _, err := c.exec.Run(ctx, "docker run "+strings.Join(args, " ")); err != nil { diff --git a/internal/ssh/mock.go b/internal/ssh/mock.go index 70a6d2d..5837101 100644 --- a/internal/ssh/mock.go +++ b/internal/ssh/mock.go @@ -3,8 +3,11 @@ package ssh import ( "bytes" "context" + "crypto/sha256" + "encoding/hex" "fmt" "io" + "strconv" "strings" "sync" ) @@ -31,6 +34,12 @@ type MockExecutor struct { // instead of being evaluated against Files — modeling an SSH channel // dying mid-command, the ambiguous-release case of audit T02. GuardTransportFailures int + + // GenerationLabels models the docker generation labels (C01-8/9) for + // the composed generation-check fragment docker emits: container + // name → decimal generation. Absent names read as 0 (legacy/unlabeled + // — the compat rule), so tests that don't care keep working. + GenerationLabels map[string]string } // MockCommand maps a command prefix to a response. @@ -79,6 +88,47 @@ func (m *MockExecutor) Run(ctx context.Context, cmd string) (string, error) { m.Calls = append(m.Calls, cmd) } + // The committed-generation CAS prefix (internal/state, + // GenerationCASPrefix — C01-8/9): evaluate the sidecar read exactly as + // the server shell would, so a stale plan's commit is refused against + // the recorded file state and a current one proceeds. + if rest, refused, found, expected, ok := evalGenerationCAS(m.Files, cmd); ok { + if refused { + m.mu.Unlock() + if found == 0 && expected == 0 { + return "", fmt.Errorf("exit status 74: TEPLOY_GENERATION_BADGEN") + } + return "", fmt.Errorf("exit status 74: TEPLOY_GENERATION_FENCED %d %d", found, expected) + } + cmd = rest + m.Calls = append(m.Calls, cmd) + } + + // docker's composed generation label check (internal/docker, + // generationCheckFragment — C01-8/9): evaluate the container's + // generation label from GenerationLabels against the expected bound. + if rest, refused, found, expected, ok := m.evalGenerationLabelCheck(cmd); ok { + if refused { + m.mu.Unlock() + return "", fmt.Errorf("exit status 74: TEPLOY_GENERATION_FENCED %d %d", found, expected) + } + cmd = rest + m.Calls = append(m.Calls, cmd) + } + + // The exact-block route CAS on Caddyfile commits (internal/caddy, + // routeCASFragment — A12/T05): hash the app's managed region from the + // recorded Caddyfile with caddy's normalization and compare against + // the acceptable hashes embedded in the command. + if rest, refused, foundGen, expectedGen, ok := evalRouteCAS(m.Files, cmd); ok { + if refused { + m.mu.Unlock() + return "", fmt.Errorf("exit status 76: TEPLOY_ROUTE_CAS_MISMATCH found_generation=%d expected_generation=%d", foundGen, expectedGen) + } + cmd = rest + m.Calls = append(m.Calls, cmd) + } + // `cat ` answers from the recorded file state when the mock has // one (the real server re-reads whatever earlier writes left); an // explicit registration still wins for paths the mock has no file for. @@ -96,14 +146,26 @@ func (m *MockExecutor) Run(ctx context.Context, cmd string) (string, error) { } if c.Err == nil { m.applyFileCommand(cmd) + // A compound file op (`mv a b && mv c d`, C01-8/9's + // state+sidecar commit) matched a broad registration; + // apply each segment so the recorded file state still + // reflects what the server shell did. + if isFileOpCommand(cmd) { + for _, seg := range strings.Split(cmd, " && ") { + m.applyFileCommand(seg) + } + } } m.mu.Unlock() return c.Output, c.Err } } - if strings.HasPrefix(cmd, "mv -f -- ") || strings.HasPrefix(cmd, "mv -fT -- ") || - strings.HasPrefix(cmd, "rm -f -- ") || strings.HasPrefix(cmd, "rm -rf -- ") { - m.applyFileCommand(cmd) + if isFileOpCommand(cmd) { + for _, seg := range strings.Split(cmd, " && ") { + if isFileOpCommand(seg) { + m.applyFileCommand(seg) + } + } m.mu.Unlock() return "", nil } @@ -250,6 +312,235 @@ func parseConditionalLockRelease(cmd string) (dir, owner string, ok bool) { return dir, owner, true } +// isFileOpCommand reports whether cmd (or its first && -chained segment) +// is a plain file mutation the mock models against Files. Compound +// commands made ENTIRELY of such segments (state.WriteFenced's +// `mv state && mv sidecar`, C01-8/9) are applied segment by segment; a +// compound carrying anything else falls through to the registrations. +func isFileOpCommand(cmd string) bool { + for _, seg := range strings.Split(cmd, " && ") { + seg = strings.TrimSpace(seg) + if !strings.HasPrefix(seg, "mv -f -- ") && !strings.HasPrefix(seg, "mv -fT -- ") && + !strings.HasPrefix(seg, "rm -f -- ") && !strings.HasPrefix(seg, "rm -rf -- ") { + return false + } + } + return true +} + +// evalGenerationCAS recognizes the committed-generation CAS prefix emitted +// by state.GenerationCASPrefix: +// +// if [ -f '' ]; then tg=$(cat '' 2>/dev/null); case +// "$tg" in ''|*[!0-9]*) ... BADGEN ...;; esac; [ "$tg" -le ] || +// { ... TEPLOY_GENERATION_FENCED ...;; }; fi; +// +// It evaluates the sidecar against the recorded files: absent passes +// (generation 0 — the targetguard contract), a newer committed generation +// than expected refuses, and a present-but-non-numeric sidecar refuses +// fail-closed. Returns the remaining effect command, whether the CAS +// refused, the committed and expected generations (for the refusal error), +// and whether cmd carried the prefix at all. +func evalGenerationCAS(files map[string][]byte, cmd string) (rest string, refused bool, found, expected uint64, ok bool) { + const casHead = "if [ -f '" + if !strings.HasPrefix(cmd, casHead) { + return "", false, 0, 0, false + } + endQuote := strings.Index(cmd[len(casHead):], "'") + if endQuote < 0 { + return "", false, 0, 0, false + } + path := cmd[len(casHead) : len(casHead)+endQuote] + // The prefix ends at the first `fi; ` after the sidecar test. + fiAt := strings.Index(cmd, "fi; ") + if fiAt < 0 { + return "", false, 0, 0, false + } + prefix := cmd[:fiAt] + // Expected generation from the `[ "$tg" -le ]` comparison. + leAt := strings.Index(prefix, `[ "$tg" -le `) + if leAt < 0 { + return "", false, 0, 0, false + } + numRest := prefix[leAt+len(`[ "$tg" -le `):] + numEnd := strings.Index(numRest, " ]") + if numEnd < 0 { + return "", false, 0, 0, false + } + exp, err := strconv.ParseUint(numRest[:numEnd], 10, 64) + if err != nil { + return "", false, 0, 0, false + } + data, present := files[path] + if !present { + return cmd[fiAt+len("fi; "):], false, 0, exp, true + } + committed, perr := strconv.ParseUint(strings.TrimSpace(string(data)), 10, 64) + if perr != nil { + // BADGEN: found==expected==0 disambiguates from a numeric refusal. + return "", true, 0, 0, true + } + if committed > exp { + return "", true, committed, exp, true + } + return cmd[fiAt+len("fi; "):], false, committed, exp, true +} + +// evalGenerationLabelCheck recognizes docker's composed generation label +// check (internal/docker.generationCheckFragment): +// +// g=$(docker inspect -f '{{index .Config.Labels "teploy.generation"}}' +// '' 2>/dev/null || printf '0'); case "$g" in ''|*[!0-9]*) g=0;; +// esac; [ "$g" -le ] || { ...TEPLOY_GENERATION_FENCED...; exit 74; }; +// +// The label is answered from m.GenerationLabels (absent = 0, the legacy +// compat rule). Must be called with m.mu held. +func (m *MockExecutor) evalGenerationLabelCheck(cmd string) (rest string, refused bool, found, expected uint64, ok bool) { + const head = `g=$(docker inspect -f '{{index .Config.Labels "teploy.generation"}}' '` + if !strings.HasPrefix(cmd, head) { + return "", false, 0, 0, false + } + restQuote := cmd[len(head):] + endQuote := strings.Index(restQuote, "' 2>/dev/null") + if endQuote < 0 { + return "", false, 0, 0, false + } + ref := restQuote[:endQuote] + leAt := strings.Index(cmd, `[ "$g" -le `) + if leAt < 0 { + return "", false, 0, 0, false + } + numRest := cmd[leAt+len(`[ "$g" -le `):] + numEnd := strings.Index(numRest, " ]") + if numEnd < 0 { + return "", false, 0, 0, false + } + exp, err := strconv.ParseUint(numRest[:numEnd], 10, 64) + if err != nil { + return "", false, 0, 0, false + } + gen := uint64(0) + if v, present := m.GenerationLabels[ref]; present { + if parsed, perr := strconv.ParseUint(strings.TrimSpace(v), 10, 64); perr == nil { + gen = parsed + } + } + if gen > exp { + return "", true, gen, exp, true + } + // Strip through the refusal block's closing `; }; ` — the fragment's + // shape is `[ "$g" -le N ] || { ...; exit 74; }; `. + tail := numRest[numEnd:] + end := strings.Index(tail, "; }; ") + if end < 0 { + return "", false, 0, 0, false + } + return tail[end+len("; }; "):], false, gen, exp, true +} + +// evalRouteCAS recognizes caddy's exact-block compare-and-swap fragment +// (internal/caddy.routeCASFragment — A12/T05): +// +// cur=$(sed -n '/^# TEPLOY BEGIN $/,/^# TEPLOY END $/p' +// '' 2>/dev/null | sha256sum | cut -d' ' -f1); +// case "$cur" in

|

) ;; *) ...TEPLOY_ROUTE_CAS_MISMATCH...;; esac; +// +// It extracts the app's marker-inclusive region from the recorded +// Caddyfile, hashes it with caddy's normalization (region + "\n", empty +// input for an absent region) and compares against the acceptable hashes. +// foundGen reports the live region's stamp (0 when unstamped) for the +// refusal error, expectedGen the fragment's expectation. +func evalRouteCAS(files map[string][]byte, cmd string) (rest string, refused bool, foundGen, expectedGen uint64, ok bool) { + const head = `cur=$(sed -n '/^# TEPLOY BEGIN ` + if !strings.HasPrefix(cmd, head) { + return "", false, 0, 0, false + } + // App name: between "BEGIN " and the range separator "$/,/^# TEPLOY END". + beginIdx := strings.Index(cmd[len(head):], "$/,/^# TEPLOY END ") + if beginIdx < 0 { + return "", false, 0, 0, false + } + app := cmd[len(head) : len(head)+beginIdx] + // The case list: between `case "$cur" in ` and `) ;; *)`. + caseAt := strings.Index(cmd, `case "$cur" in `) + if caseAt < 0 { + return "", false, 0, 0, false + } + listRest := cmd[caseAt+len(`case "$cur" in `):] + listEnd := strings.Index(listRest, ") ;; *)") + if listEnd < 0 { + return "", false, 0, 0, false + } + acceptable := strings.Split(listRest[:listEnd], "|") + // The expectation rides the refusal printf: expected_generation=. + if at := strings.Index(cmd, "expected_generation="); at >= 0 { + num := cmd[at+len("expected_generation="):] + end := strings.IndexAny(num, " \n") + if end < 0 { + end = len(num) + } + if v, err := strconv.ParseUint(num[:end], 10, 64); err == nil { + expectedGen = v + } + } + // The fragment ends at `esac; `. + esacAt := strings.Index(cmd, "esac; ") + if esacAt < 0 { + return "", false, 0, 0, false + } + // Region extraction, mirroring the server sed: marker-inclusive lines. + fileData, present := files["/deployments/caddy/Caddyfile"] + if !present { + fileData = nil + } + regionText := extractMockManagedRegion(string(fileData), app) + var payload []byte + if regionText != "" { + payload = append([]byte(regionText), '\n') + } + sum := sha256.Sum256(payload) + actual := hex.EncodeToString(sum[:]) + for _, line := range strings.Split(regionText, "\n") { + if strings.HasPrefix(line, "# TEPLOY GENERATION ") { + if v, err := strconv.ParseUint(strings.TrimSpace(strings.TrimPrefix(line, "# TEPLOY GENERATION ")), 10, 64); err == nil { + foundGen = v + } + break + } + } + for _, h := range acceptable { + if h == actual { + return cmd[esacAt+len("esac; "):], false, foundGen, expectedGen, true + } + } + return "", true, foundGen, expectedGen, true +} + +// extractMockManagedRegion is the mock's mirror of caddy's +// extractManagedRegion (marker-inclusive region text, "" when absent) — +// duplicated rather than imported so the ssh package keeps no caddy +// dependency; the two MUST stay byte-compatible (exact-line marker +// matching, lines joined with \n, no trailing newline). +func extractMockManagedRegion(content, app string) string { + lines := strings.Split(content, "\n") + var out []string + active := false + for _, line := range lines { + if !active { + if line == "# TEPLOY BEGIN "+app { + active = true + out = append(out, line) + } + continue + } + out = append(out, line) + if line == "# TEPLOY END "+app { + return strings.Join(out, "\n") + } + } + return "" +} + func mockCommandMatches(cmd, match string) bool { if !strings.HasPrefix(cmd, match) { return false diff --git a/internal/state/generation_test.go b/internal/state/generation_test.go new file mode 100644 index 0000000..9deb12e --- /dev/null +++ b/internal/state/generation_test.go @@ -0,0 +1,154 @@ +package state + +import ( + "context" + "errors" + "strings" + "testing" + "time" + + "github.com/useteploy/teploy/internal/ssh" +) + +func TestGenerationCASPrefixSemantics(t *testing.T) { + ctx := context.Background() + + t.Run("absent sidecar reads as generation 0 and passes", func(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "printf ok", Output: "ok"}) + out, err := mock.Run(ctx, GenerationCASPrefix("myapp", 0)+"printf ok") + if err != nil || strings.TrimSpace(out) != "ok" { + t.Fatalf("expected the effect to run, got %q %v", out, err) + } + }) + + t.Run("equal committed generation passes", func(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "printf ok", Output: "ok"}) + mock.Files[GenerationSidecarPath("myapp")] = []byte("7\n") + out, err := mock.Run(ctx, GenerationCASPrefix("myapp", 7)+"printf ok") + if err != nil || strings.TrimSpace(out) != "ok" { + t.Fatalf("expected the effect to run, got %q %v", out, err) + } + }) + + t.Run("older committed generation passes", func(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4", + ssh.MockCommand{Match: "printf ok", Output: "ok"}) + mock.Files[GenerationSidecarPath("myapp")] = []byte("6\n") + if _, err := mock.Run(ctx, GenerationCASPrefix("myapp", 7)+"printf ok"); err != nil { + t.Fatalf("an older committed generation never fences: %v", err) + } + }) + + t.Run("newer committed generation refuses naming both", func(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4") + mock.Files[GenerationSidecarPath("myapp")] = []byte("8\n") + _, err := mock.Run(ctx, GenerationCASPrefix("myapp", 7)+"printf ok") + if err == nil || !strings.Contains(err.Error(), "TEPLOY_GENERATION_FENCED 8 7") { + t.Fatalf("refusal must name committed and expected generations, got %v", err) + } + if !GenerationFenced(err) { + t.Fatal("GenerationFenced must match the refusal") + } + }) + + t.Run("corrupt sidecar refuses fail-closed", func(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4") + mock.Files[GenerationSidecarPath("myapp")] = []byte("eight\n") + _, err := mock.Run(ctx, GenerationCASPrefix("myapp", 99)+"printf ok") + if err == nil || !strings.Contains(err.Error(), "TEPLOY_GENERATION_BADGEN") { + t.Fatalf("a corrupt sidecar must fail closed, got %v", err) + } + if !GenerationFenced(err) { + t.Fatal("GenerationFenced must match the BADGEN refusal") + } + }) +} + +func TestReadCommittedGeneration(t *testing.T) { + ctx := context.Background() + mock := ssh.NewMockExecutor("1.2.3.4") + if gen, err := ReadCommittedGeneration(ctx, mock, "myapp"); err != nil || gen != 0 { + t.Fatalf("absent sidecar must read as 0, got %d %v", gen, err) + } + mock.Files[GenerationSidecarPath("myapp")] = []byte("12\n") + if gen, err := ReadCommittedGeneration(ctx, mock, "myapp"); err != nil || gen != 12 { + t.Fatalf("expected 12, got %d %v", gen, err) + } + mock.Files[GenerationSidecarPath("myapp")] = []byte("junk\n") + if _, err := ReadCommittedGeneration(ctx, mock, "myapp"); err == nil { + t.Fatal("a corrupt sidecar must be an error, never a silent 0") + } +} + +// TestWriteFenced_PublishesGenerationSidecar pins the C01-8/9 commit shape: +// the guarded command renames state.json and the sidecar TOGETHER — +// authority first, fence second — so the generation a later CAS fences +// against is exactly the generation state.json names. +func TestWriteFenced_PublishesGenerationSidecar(t *testing.T) { + lk, mock := takeFencedLock(t, "myapp") + s := &AppState{SchemaVersion: SchemaVersionV2, CurrentHash: "abc", UpdatedAt: time.Now().UTC(), OperationID: "op", Generation: 8} + if err := WriteFenced(context.Background(), mock, "myapp", s, lk); err != nil { + t.Fatalf("WriteFenced: %v", err) + } + gen, ok := mock.Files[GenerationSidecarPath("myapp")] + if !ok || strings.TrimSpace(string(gen)) != "8" { + t.Fatalf("the sidecar must be published with the committed generation, got %q", string(gen)) + } + // Ordering: the state rename precedes the sidecar rename in the SAME + // guarded command (state.json is the authority; the sidecar trails + // in-shell, never leads). + var commit string + for _, c := range mock.Calls { + if strings.Contains(c, "mv -f -- ") && strings.Contains(c, "/state.json") && strings.Contains(c, ".generation") { + commit = c + } + } + if commit == "" { + t.Fatal("state and sidecar must commit in one guarded command") + } + stateIdx := strings.Index(commit, "state.json") + genIdx := strings.Index(commit, ".generation") + if stateIdx > genIdx { + t.Fatalf("state.json must be renamed before the sidecar:\n%s", commit) + } +} + +// TestWriteFencedGeneration_RefusesOverNewerCommittedGeneration is the +// stale-rollback acceptance at the authority boundary: an operation that +// resolved the world at generation 7 cannot commit when the target already +// committed generation 8 — the refusal names the expected generation and +// state.json keeps the successor's content. +func TestWriteFencedGeneration_RefusesOverNewerCommittedGeneration(t *testing.T) { + lk, mock := takeFencedLock(t, "myapp") + mock.Files[GenerationSidecarPath("myapp")] = []byte("8\n") + successorState := `{"schema_version":2,"deployment_type":"container","ingress_mode":"caddy","updated_at":"2026-09-24T00:00:00Z","operation_id":"successor","generation":8,"current_hash":"new"}` + mock.Files["/deployments/myapp/state.json"] = []byte(successorState) + + s := &AppState{SchemaVersion: SchemaVersionV2, CurrentHash: "old", UpdatedAt: time.Now().UTC(), OperationID: "stale", Generation: 8} + err := WriteFencedGeneration(context.Background(), mock, "myapp", s, lk, 7) + if !errors.Is(err, ErrGenerationFenced) { + t.Fatalf("expected ErrGenerationFenced, got %v", err) + } + if !strings.Contains(err.Error(), "expected predecessor 7") { + t.Errorf("refusal must name the expected generation: %v", err) + } + if string(mock.Files["/deployments/myapp/state.json"]) != successorState { + t.Fatal("a refused commit must not touch the successor's state.json") + } +} + +// TestWrite_PublishesSidecar pins the unfenced writer's sidecar +// maintenance (heal and static paths share it). +func TestWrite_PublishesSidecar(t *testing.T) { + mock := ssh.NewMockExecutor("1.2.3.4") + s := &AppState{SchemaVersion: SchemaVersionV2, CurrentHash: "abc", UpdatedAt: time.Now().UTC(), OperationID: "op", Generation: 3} + if err := Write(context.Background(), mock, "myapp", s); err != nil { + t.Fatalf("Write: %v", err) + } + gen, ok := mock.Files[GenerationSidecarPath("myapp")] + if !ok || strings.TrimSpace(string(gen)) != "3" { + t.Fatalf("Write must publish the sidecar, got %q", string(gen)) + } +} diff --git a/internal/state/lock.go b/internal/state/lock.go index e7b360e..1253808 100644 --- a/internal/state/lock.go +++ b/internal/state/lock.go @@ -39,6 +39,7 @@ import ( "errors" "fmt" "os" + "strconv" "strings" "sync" "time" @@ -394,8 +395,28 @@ func ReleaseLockFenced(exec ssh.Executor, lk *Lock, app string) { // staged to a unique sibling (no effect — a stale holder cannot clobber the // successor's staging, A06), and the atomic rename — the instant the new // state becomes authoritative — runs as a guarded effect. A holder that -// lost the lock commits nothing. +// lost the lock commits nothing. The committed-generation sidecar is +// published in the same guarded command, right after the state rename. func WriteFenced(ctx context.Context, exec ssh.Executor, app string, s *AppState, lk *Lock) error { + return writeFenced(ctx, exec, app, s, lk, nil) +} + +// WriteFencedGeneration is WriteFenced with the commit also fenced on the +// PREDECESSOR generation (C01-8/9): the guarded command chains the holdership +// guard, a compare-and-swap on the committed-generation sidecar (refuses when +// the target already committed a generation newer than expectedGeneration — +// ErrGenerationFenced), the state.json rename, and the sidecar update — one +// remote shell, so authority and its generation fence move together. A +// rollback or deploy that resolved the world at generation E cannot commit +// over a successor's E+N, no matter when its commands land. The sidecar CAS +// is defense-in-depth on top of the holdership guard: the guard refuses a +// BROKEN holder, the CAS refuses an effect that is stale on CONTENT even +// while the shell still holds (the in-flight/takeover-interleave window). +func WriteFencedGeneration(ctx context.Context, exec ssh.Executor, app string, s *AppState, lk *Lock, expectedGeneration uint64) error { + return writeFenced(ctx, exec, app, s, lk, &expectedGeneration) +} + +func writeFenced(ctx context.Context, exec ssh.Executor, app string, s *AppState, lk *Lock, expectedGeneration *uint64) error { if lk == nil { return Write(ctx, exec, app, s) } @@ -414,8 +435,31 @@ func WriteFenced(ctx context.Context, exec ssh.Executor, app string, s *AppState if err := exec.Upload(ctx, strings.NewReader(string(data)), tmpPath, "0644"); err != nil { return fmt.Errorf("uploading temporary state file: %w", err) } - if _, err := lk.Guarded(ctx, exec, "mv -f -- "+ssh.ShellQuote(tmpPath)+" "+ssh.ShellQuote(path)); err != nil { - // Leave the temp file for diagnosis; it is inert. + // The sidecar payload is staged like the state (unique owner-scoped + // sibling, A06's discipline) and renamed in the SAME guarded command, + // after state.json: authority first, fence second — a crash between the + // two renames can only leave the sidecar one generation BEHIND, the + // permissive direction (a stale effect may pass the CAS once; nothing + // can ever fence against a generation that never committed). + sidecarPath := GenerationSidecarPath(app) + sidecarTmp, err := ownerTempName(sidecarPath, lk.Owner()) + if err != nil { + return err + } + if err := exec.Upload(ctx, strings.NewReader(strconv.FormatUint(s.Generation, 10)+"\n"), sidecarTmp, "0644"); err != nil { + return fmt.Errorf("uploading temporary generation sidecar: %w", err) + } + genCAS := "" + if expectedGeneration != nil { + genCAS = GenerationCASPrefix(app, *expectedGeneration) + } + effect := genCAS + "mv -f -- " + ssh.ShellQuote(tmpPath) + " " + ssh.ShellQuote(path) + + " && mv -f -- " + ssh.ShellQuote(sidecarTmp) + " " + ssh.ShellQuote(sidecarPath) + if _, err := lk.Guarded(ctx, exec, effect); err != nil { + if GenerationFenced(err) { + return fmt.Errorf("%w: refusing to commit generation %d for %s over a newer committed generation (expected predecessor %d)", ErrGenerationFenced, s.Generation, app, *expectedGeneration) + } + // Leave the temp files for diagnosis; they are inert. return err } return nil diff --git a/internal/state/state.go b/internal/state/state.go index c0a3b66..bd12ad7 100644 --- a/internal/state/state.go +++ b/internal/state/state.go @@ -8,6 +8,7 @@ import ( "encoding/base64" "encoding/hex" "encoding/json" + "errors" "fmt" "strconv" "strings" @@ -21,6 +22,91 @@ const ( SchemaVersionV2 = 2 ) +// generationFencedMarker and generationBadgenMarker are the stderr sentinels +// of the generation CAS prefix (below): the first names the committed and the +// expected generation when a stale plan is fenced out; the second reports a +// present-but-unreadable .generation sidecar (fail closed, never pass a +// corrupt fence). Exit status 74 distinguishes the refusal class from the +// holdership fence's 75 in transport-wrapped errors. +const ( + generationFencedMarker = "TEPLOY_GENERATION_FENCED" + generationBadgenMarker = "TEPLOY_GENERATION_BADGEN" +) + +// ErrGenerationFenced is returned when an effect prepared against generation +// E discovers the target has committed a NEWER generation (sidecar > E) — +// the stale-rollback shape C01-8/9 exist to make impossible: reconcile +// against the observed target instead of retrying (the D11 honesty floor). +var ErrGenerationFenced = errors.New("generation fence refused the effect — the target committed a newer generation than this operation prepared against; reconcile before retrying") + +// GenerationFenced reports whether err is a generation-CAS refusal — the +// composed prefix's marker or its exit status (74) surfaced by either +// executor flavor — including effects composed by other packages through +// GenerationCASPrefix (docker's generation-checked stops use the same +// marker contract). +func GenerationFenced(err error) bool { + if err == nil { + return false + } + if errors.Is(err, ErrGenerationFenced) { + return true + } + msg := err.Error() + return strings.Contains(msg, generationFencedMarker) || + strings.Contains(msg, generationBadgenMarker) || + strings.Contains(msg, "status 74") || + strings.Contains(msg, "exit status 74") +} + +// GenerationSidecarPath returns the committed-generation sidecar's path: +// /deployments//.generation, a plain decimal integer written atomically +// by every state commit (Write/WriteFenced[Generation]). The path and the +// "absent = 0" rule are the targetguard helper's contract +// (internal/targetguard/guard.sh) — one sidecar both fences read. +func GenerationSidecarPath(app string) string { + return fmt.Sprintf("%s/%s/.generation", deploymentsDir, app) +} + +// ReadCommittedGeneration reads the sidecar: absent means 0 (pre-sidecar +// install), a parseable integer is returned as-is, anything else is an +// error — a corrupt fence must never silently read as generation 0. +func ReadCommittedGeneration(ctx context.Context, exec ssh.Executor, app string) (uint64, error) { + data, present, err := ReadRemoteFile(ctx, exec, GenerationSidecarPath(app)) + if err != nil { + return 0, fmt.Errorf("reading the committed generation for %s: %w", app, err) + } + if !present { + return 0, nil + } + gen, perr := strconv.ParseUint(strings.TrimSpace(string(data)), 10, 64) + if perr != nil { + return 0, fmt.Errorf("the committed-generation sidecar for %s is unreadable (%q) — inspect %s before retrying", app, strings.TrimSpace(string(data)), GenerationSidecarPath(app)) + } + return gen, nil +} + +// GenerationCASPrefix returns the shell prefix that composes the +// committed-generation compare-and-swap with an effect command into ONE +// remote invocation: +// +// +// +// The prefix refuses (marker on stderr, exit 74) when the app's .generation +// sidecar names a generation NEWER than expected, or when a present sidecar +// is unreadable (fail closed). An absent sidecar reads as generation 0 — +// the targetguard contract — so legacy installs keep working. Semantics are +// identical to internal/targetguard/guard.sh's fence (committed > expected +// fences; GUARD_BADGEN fails closed), implemented as an in-shell prefix so +// it composes with the holdership guard (state.Lock.GuardPrefix) and needs +// no helper upload. Effect sites map refusals with GenerationFenced. +func GenerationCASPrefix(app string, expected uint64) string { + sidecar := ssh.ShellQuote(GenerationSidecarPath(app)) + return fmt.Sprintf( + `if [ -f %s ]; then tg=$(cat %[1]s 2>/dev/null); case "$tg" in ''|*[!0-9]*) printf '%s\n' >&2; exit 74;; esac; [ "$tg" -le %d ] || { printf '%s %%s %%s\n' "$tg" %d >&2; exit 74; }; fi; `, + sidecar, generationBadgenMarker, expected, generationFencedMarker, expected, + ) +} + // staleLockTTL is how long an "auto" deploy lock (see LockInfo.Type) is // honored before AcquireLock treats it as abandoned and breaks it. Deploys // that crash or lose their SSH connection mid-flight — before the deferred @@ -374,14 +460,35 @@ func validateReleaseDigest(release *ReleaseMetadata) error { // Write atomically writes canonical v2 JSON. It never updates the legacy // key=value file, making that file an import-only migration source rather than -// a second writable authority. +// a second writable authority. The committed-generation sidecar (C01-8/9) is +// written right after the state rename — state.json is the authority and the +// sidecar trails it by one shell command, so the generation CAS +// (GenerationCASPrefix, targetguard) always fences against a generation that +// already committed. func Write(ctx context.Context, exec ssh.Executor, app string, s *AppState) error { data, err := prepareState(s) if err != nil { return err } path := fmt.Sprintf("%s/%s/state.json", deploymentsDir, app) - return ssh.UploadAtomic(ctx, exec, bytes.NewReader(data), path, "0644") + if err := ssh.UploadAtomic(ctx, exec, bytes.NewReader(data), path, "0644"); err != nil { + return err + } + return writeGenerationSidecar(ctx, exec, app, s.Generation) +} + +// writeGenerationSidecar publishes the committed generation atomically +// (stage + rename) — the targetguard contract's write side. +func writeGenerationSidecar(ctx context.Context, exec ssh.Executor, app string, generation uint64) error { + path := GenerationSidecarPath(app) + tmp := path + ".tmp-commit" + if err := exec.Upload(ctx, strings.NewReader(strconv.FormatUint(generation, 10)+"\n"), tmp, "0644"); err != nil { + return fmt.Errorf("staging the committed-generation sidecar for %s: %w", app, err) + } + if _, err := exec.Run(ctx, "mv -fT -- "+ssh.ShellQuote(tmp)+" "+ssh.ShellQuote(path)); err != nil { + return fmt.Errorf("publishing the committed-generation sidecar for %s: %w", app, err) + } + return nil } // prepareState validates and serializes s for writing (shared by Write and diff --git a/release/RELEASE_RECEIPT.md b/release/RELEASE_RECEIPT.md new file mode 100644 index 0000000..9735303 --- /dev/null +++ b/release/RELEASE_RECEIPT.md @@ -0,0 +1,80 @@ +# Release receipts (R01) + +Every release — and every local verification of one — leaves a receipt: +the recorded inputs and the checksums that prove a clean machine can +recreate the build. This file holds the PATTERN (template below) and +receipts from local verification runs. A receipt from an actual release +is added by the owner at tag time; nothing in this directory cuts a +release, tags, or publishes. + +## How to produce a receipt + +``` +make release-record # builds the goreleaser matrix, writes + # release/expectations/checksums--.txt +make release-verify # rebuilds from a clean checkout and diffs against + # the recorded expectation (inputs must match) +make release-smoke # builds the image (Dockerfile, scratch) from the + # verified binary and runs version + doctor in it +``` + +`release-record` prints the receipt block; paste it below under a new +heading. Commit the expectation file with it — the receipt and the +expectation are one artifact split across two files. + +## Template + +```markdown +### — — + +- commit : (git describe: ) +- go toolchain : (go.mod: ) +- build env : CGO_ENABLED=0, GOFLAGS unset +- ldflags : -s -w -X main.version= +- matrix : linux/darwin/windows × amd64/arm64 (goreleaser parity) +- checksums : release/expectations/checksums--.txt +- verify : make release-verify → REPRODUCIBLE (this machine, /) +- image smoke : make release-smoke → PASS (linux/, scratch image; + version exit 0; doctor emits the 9-check machine envelope, + exit 1 = documented failing-diagnosis semantics) +- goreleaser : +``` + +--- + +### 0.0.0-localverify-11d0709 — 2026-09-23 — LOCAL VERIFICATION, NOT A RELEASE + +First receipt, from the R01 lane's local run (worktree branch +`c09-x02f-r01-cli`, off `d9b652d`). Recorded at source commit `11d0709` +(the R01 tooling commit, clean tree — the expectation file itself lands +in the immediately following commit, which is why `release-verify` +compares via source identity, not strict HEAD equality). + +- commit : 11d0709fbe19b642739809c614d1ff64a0812e41 (git describe: v0.1.37-46-g11d0709) +- go toolchain : go1.26.6 (go.mod: 1.26.0) +- build env : CGO_ENABLED=0, GOFLAGS unset +- ldflags : -s -w -X main.version=0.0.0-localverify-11d0709 +- matrix : linux/darwin/windows × amd64/arm64 (goreleaser parity) +- checksums : release/expectations/checksums-0.0.0-localverify-11d0709-go1-26-6.txt + - bf185800ea9b731849cd0a4c61940586bb4fdfc5875636966abe01aad4a34d8e teploy_darwin_amd64 + - 553044b9f5d81fc99fb7ebd2369f0435a41daff3abbfe3f27e0fc192024653f8 teploy_darwin_arm64 + - bacf79c2f18585676e64ebbc06a409304fdb91fb72b9c386ae0b51b4adbcc259 teploy_linux_amd64 + - 33446a81515d1d7c4026b5de04aa56e7216389f6833ac56d13c35a25a41fa208 teploy_linux_arm64 + - 7d0255f55e8f6e74e9c85acff6d57b634b323dd5ea84caeb2ce86da966362fc9 teploy_windows_amd64 + - 13d24237797c3484a40388fa21af4834c1affa3ba16beb2a139960764c23de56 teploy_windows_arm64 +- verify : `./scripts/release-verify.sh verify 0.0.0-localverify-11d0709` + → REPRODUCIBLE (all 6 binaries bit-identical to the recording; clean + tree, Darwin/arm64 host) +- image smoke : `./scripts/release-verify.sh smoke` → PASS — + linux/arm64 scratch image from the verified binary; `teploy version` + exit 0; `teploy doctor --json` exit 1 with the full 9-check machine + envelope (machine_interface 2), which is the documented + failing-diagnosis semantics in a bare container, not a crash +- goreleaser : n/a (no tag, no publish — owner-controlled) + +Scope note: checksums cover the raw matrix binaries, the reproducible +unit; goreleaser's published `checksums.txt` covers archives (which +embed mtimes) and is recorded at release time. Bit-identical +reproduction is toolchain-scoped: the expectation file pins the go +version, and `release-verify` refuses to compare (exit 1, INPUTS +DIFFER) rather than report a meaningless mismatch across toolchains. diff --git a/release/expectations/checksums-0.0.0-localverify-11d0709-go1-26-6.txt b/release/expectations/checksums-0.0.0-localverify-11d0709-go1-26-6.txt new file mode 100644 index 0000000..6ffeccc --- /dev/null +++ b/release/expectations/checksums-0.0.0-localverify-11d0709-go1-26-6.txt @@ -0,0 +1,15 @@ +# teploy release-verify receipt +# version label : 0.0.0-localverify-11d0709 +# commit : 11d0709fbe19b642739809c614d1ff64a0812e41 +# git describe : v0.1.37-46-g11d0709 +# dirty files : 0 +# go toolchain : go1.26.6 (go.mod: 1.26.0) +# build env : CGO_ENABLED=0 GOFLAGS= +# ldflags : -s -w -X main.version=0.0.0-localverify-11d0709 +# recorded : 2026-09-24T01:26:56Z on Darwin/arm64 +9b41470e71554bb6a2899a5832ae20b67af06ec27aa6b3d31a99fa07a3706329 teploy_darwin_amd64 +fa51620f2abbbdfa8d9db71c18470c327f1ce7be8065cac7fb5045c01c5eef22 teploy_darwin_arm64 +2e37ab081802d45638dd5a4d0e50f685c9ef27dd438eafcb078653539b2dcd37 teploy_linux_amd64 +1e22ca85320fdef40039a430248ac5b46366ba41904d33e449e34742f0896c73 teploy_linux_arm64 +d0b2447f61520c97f4be25e9ef57cc9bc7ab8db253a3b5672a20413708e3cc5a teploy_windows_amd64 +7108ded33e71efc125601d7bbf93bb1beb57d51e360dbe2b763bdc8314f89bd5 teploy_windows_arm64 diff --git a/scripts/release-verify.sh b/scripts/release-verify.sh new file mode 100755 index 0000000..1c7b25a --- /dev/null +++ b/scripts/release-verify.sh @@ -0,0 +1,190 @@ +#!/usr/bin/env bash +# R01 release-reproducibility receipts for teploy-cli. +# +# Subcommands: +# verify build the exact goreleaser matrix locally, checksum every +# binary, and diff against the recorded expectation for this +# (version label, go toolchain). Exit 0 on match; exit 1 on +# checksum mismatch or recorded-input mismatch (a different +# go toolchain cannot reproduce the recorded bits — align or +# re-record); exit 1 with NO RECORDED EXPECTATION when none +# exists yet (run `record`). Honest results, never vacuous. +# record build the same matrix and write the expectation file +# (release/expectations/) plus a receipt block for +# release/RELEASE_RECEIPT.md. Recording is a deliberate act — +# expectations are committed and reviewed like any contract. +# smoke build the linux binary for the local docker architecture, +# build the container image (Dockerfile at repo root), and +# run `version` + `doctor` inside it as the built-image +# smoke. Prints SKIP (exit 0) when docker is unavailable. +# +# The matrix mirrors .goreleaser.yml exactly: CGO_ENABLED=0, +# -s -w -X main.version=