All notable changes to teploy are documented here. Format follows Keep a Changelog.
- Machine interface 2:
server list --jsonemits the machine envelope. The output is now{"machine_interface":2,"servers":[{"name","id","host","user","role","tags","vpn_ip"}],"observed_at"}— every per-server field the old bare map carried, with the map key promoted to anamefield — where it previously emitted a bare map-of-servers at the root. Removing the bare map from the wire is a non-additive machine-interface change: the whole binary now reportsmachine_interface: 2in every--jsonenvelope (MI 2 = MI 1 plus this reshape; no capability token changed). Machine consumers pinned to MI 1 must move to the envelope; teploy-dash ships a decoder for both shapes in the same release.
-
Tailnet preview mode.
teploy preview deploytakes--base-domain <domain>(hostname base instead of the app domain, e.g.100-64-1-2.sslip.io),--http-only(plain HTTP site, no ACME) and--allow-ip <ip|cidr>(repeatable; everything else gets 403). The mode is stored in the preview record (base_domain,http_only,allow_ips, all omitted for default previews) and inherited by later deploys of the branch unless overridden (--http-only=false,--allow-ip ""), so a blue/green update never silently re-enables HTTPS or drops the allowlist.preview list --jsonrows gainurl(http://for HTTP-only, elsehttps://);domainis unchanged. Thepreview-exposurecapability token is advertised. Preview identity (<app>-p-<hex8>) and blue/green are unchanged; records without the new fields behave exactly as before. -
Plan/apply with drift invalidation (C05).
teploy plannow renders the full effect set — routing (domain/ingress/port/publishes), environment keys, storage volumes, resource limits, accessories — alongside the container diff, and classifies image data honestly: resolved-by-digest, mutable-tag-resolved-at-plan-time, or unresolved-awaiting-build (a build plan binds its build inputs — context fingerprint + Dockerfile identity — instead of guessing).teploy plan --out FILErecords a plan identity (effective-config digest, target version, server/app, deployed-state generation).teploy apply FILEre-derives every binding input and refuses, naming what drifted (config edit, overlay flip, moved git version, changed build inputs, or any deploy/rollback in between — the state generation), then executes through the normal deploy engine with the plan id stamped into the release's provenance receipt. A floating-tag plan (timestamp version) is refused outright — it can never be bound. Under--jsona drift refusal classifies as the error envelope'sconflictcode; theplan-applycapability token is advertised. -
Compose import: the web service's
environment:andvolumes:are now translated intoenv:/volumes:(host binds keep their full path). Both were silently DROPPED — only accessories were parsed — so an imported stack deployed without any of the web service's environment or storage.
- Webhook deploys build the authenticated commit. A delivery now pins the build to the payload's exact commit (fetch-verify-reset); if that commit was force-pushed away the deploy fails loudly naming both commits instead of silently building the moving branch tip. The pin rides the durable admission ledger through supersede and crash-resume.
- Preview updates no longer take the preview down. A preview update
now runs blue/green: the candidate starts under a version-suffixed
name with its own network alias, passes a readiness gate, the route
switches, and only then is the predecessor retired — a failed
candidate leaves the old preview serving.
teploy preview pruneprunes expired previews across all apps (both record eras, idempotent, 72h default TTL) and is cron-able; the deploy-time prune uses the same core. - SSH host-key mismatches now name what was presented and what known_hosts holds (key algorithms included), instead of a bare mismatch error (found via ship's delivery provisioning).
- Crash-recovery evidence for deploys (C01 design obligations): readiness receipts (exact candidate IDs + probe outcomes) and predecessor snapshots (exact container IDs) persist per attempt at the moment they become true, so recovery can distinguish compensable states from inspect-only ones and restore exactly what was displaced; the deploy log records DEGRADED outcomes (traffic switched but retirement partially failed) instead of clean success.
- Preview environments no longer collide across branch names.
Branches that sanitize to the same slug (
feature/loginandfeature-login) used to share one preview record — the second deploy destroyed the first's state, container and route. Previews now carry canonical IDs (<app>-p-<hex8>, derived from the app and the FULL branch ref) through state files, container names, network aliases, Caddy routes and domains; the slug is only a display prefix. Existing slug-keyed records are adopted when their stored full branch matches; a genuine collision surfaces an explicit error naming both branches instead of guessing or deleting. - Webhook deliveries during a running deploy were acknowledged, then silently dropped. The deploy lock returns immediately when already held — it does not queue — so every push arriving mid-deploy got a 200 and then nothing. Admission is now durable before the acknowledgment (an fsync'd append-only ledger), with one bounded newest-wins pending slot per app (older pending deliveries are marked superseded, never piled up in goroutines), and a listener restart resumes admitted-but-unprocessed work exactly once. Persistence failure refuses with 503 + Retry-After instead of acknowledging something that is not durable.
- Scheduled redeploys run through the deploy engine. The cron
script no longer reconstructs the container from
docker inspectwith its own stop/rm/run (no lock, no health gate, no release record, no rollback, a stop-to-start downtime window). It keeps the cheap digest pre-check and, when the digest moved, invokes the new server-sideteploy autodeploy redeploy— the same fenced, health-gated deploy path as the webhook listener.schedulegains--branch, installs the server binary, and verifies it supports the command. Re-runteploy autodeploy scheduleon existing apps to upgrade an installed script; until then the old script keeps its previous behavior.
- Compose import no longer loses the application port.
ports: ['8080:3000']means container port 3000 bound to host 8080; the importer previously used ports only to pick the web service and discarded them, so the app deployed as:80, failed its health check and rolled back. The container port now imports as the application port (with bare ports, IP-prefixed bindings and IPv6 forms supported; non-TCP entries preserved intopublish), and unsupported grammar — ranges, long-form port objects, multiple distinct container ports, UDP-only services — is refused naming the service and the reason instead of being silently reinterpreted. - Compose import refuses services it cannot faithfully deploy.
Previously a service built from a different context than the app's
(
jobs: build ./jobsnext toweb: build ./web) was silently flattened into a process of the app's image: the wrong code ran under the right command and the import reported success. It is now refused naming the service, its build context and the remediation (shared context, prebuilt image, or teploy.yml). Same-build workers are unaffected. This is a deliberate behavior change: files that imported "successfully" while deploying something other than what they declared now fail fast at import time. - Compose fields are classified instead of silently ignored. The
importer's non-strict YAML parse accepted files using
healthcheck,networks,secrets,configs,profiles,deploy,env_file,entrypointand security options while dropping their semantics. Now: service healthchecks translate to teploy'shealth:block, no-op values are tolerated, non-default profiles skip the service (matchingdocker compose upsemantics), metadata is ignored with reasons, and everything with semantics teploy cannot preserve is rejected naming the service and field. Another deliberate behavior change in the same spirit as above.
- Crash-recovery state table for the deploy lifecycle (design spike
for the transaction work): an eight-state lattice with an exhaustive,
property-tested disposition function (retry / inspect / compensate /
manual) over crash evidence, an ADR mapping it onto the existing
fenced-lock and release-record machinery, and an integration-tagged
fault harness (
go test -tags integration ./internal/deploy/recoverywithTEPLOY_FAULT_*env) that drives a real SSH+Docker host through late-effect-after-owner-death, side-effect-without-receipt and stale-holder scenarios. No deploy behavior changed in this release by this table; it is the specification the recovery work implements against.
- Template rendering is YAML-safe with path-keyed generated secrets and a bounded, validated registry fetch; the templates corpus test renders the whole catalog.
- Homebrew formulas carry a version test block.
- Deterministic Compose import (sorted web-service selection) and
exec-form command quoting that survives the container's
sh -cre-parse. - Registry-port image references parse correctly in backup/restore, and
.envis archived/restored at its app-level location. - Scheduled backups: only the archive just created is uploaded,
--endpointapplies to every aws call,%escapes survive crond, and accessory restore uses a per-invocation temp directory; MySQL dump/restore passes the root password viaMYSQL_PWDinstead of argv.
-
deployserved a stale image whenever one was already on the server.ensureImageskippeddocker pullany time the image existed in the host's local cache, which treated a mutable tag exactly like an immutable one. Observed live: an app whoseteploy.ymlnamed a registry image with no tag (so:latest) had a:latestalready on the host; CI pushed a fresh:latestfor five days and teploy never pulled one of them. Every deploy started a container named after the newer commit, passed its health check and reported success while production served the build from five days earlier. Nothing could catch it from the outside — the container name is a label teploy writes, sodocker psagreed, andteploy driftcompares live state to deploy state by name, so it agreed too. A later manualdocker pullreported "Downloaded newer image".Now only a digest-pinned reference (
repo@sha256:...) is served from the local cache, because it is content-addressed and cannot move. Every tag, and a bare repo, is pulled on every deploy. A tag that looks like a git sha is still mutable by convention only, so it is not special-cased. The cost is a manifest check, not a re-download.The out-of-band case the skip was added for — an image built or
docker loaded on the server that exists in no registry — still works: when the pull fails and a copy is present, the deploy falls back to it and says so (WARNING: pull failed .../Falling back to the local copy ... it may be older than the registry), so "pulled fresh" can never be mistaken for "registry unreachable, using what was already here". A failed pull with nothing cached is still an error. OnlyensureImagehad this shape; the other pull sites (singledeploy,autodeploy serve,accessory upgrade) already pulled unconditionally. -
accessory backup --schedulegenerated a cron line calling a bareteploywhile installing the binary to/usr/local/bin, which is not on cron'sPATH=/usr/bin:/bin. The job was listed bycrontab -l, cron was running, and it had never executed once in 48 days. The generated line carries no output redirect, so the failure went nowhere. The cron entry now uses the absolute path. -
accessory verify-backuphad no--local, so it could not be pointed at the backups it exists to prove — the nightly jobs run as a local cron job writing to a MinIO on127.0.0.1, where there is no host for the SSH path to reach. -
accessory verify-backupstaged both the downloaded archive and its extract in/tmp, a tmpfs on most distributions and therefore sized by RAM. A 28 GB dataset on a host with a 15 GB/tmpfailed withtar: Wrote only 2560 of 10240 bytes, which reads as a corrupt archive and was not. Staging moved to the disk-backed/var/tmp;TEPLOY_VERIFY_TMPDIRoverrides it.
teploy secret rm KEY [KEY...]— delete secrets from the local age store.secrethad get/list/put/rotate/set/setup/status/audit/db and no way to remove anything, so a key that had to go away (a revoked vendor credential, a retired integration) could only be overwritten with a dummy value and stayed injected into every container. Reports each key as removed or as never set; a key that is not set is not an error, so re-running a revocation is safe. Local store only —--provider openbaois refused with a clear message rather than silently removing nothing, since versioned-KV delete/destroy is not the local store's unlink.
-
deploylabelled the deploy with the local git HEAD even when deploying a prebuilt image, so the container name,teploy logand the operator's screen all named a commit whose code was not running. Observed as "Deployed version fe82bf0" while running the image tagged462d7a7. The version now comes from the image when one is given;--versionstill wins over everything.This was never specific to the
--imageflag — ateploy.ymlpinning a prebuilt image hit the identical problem with no flag passed.Behaviour change: with an image that carries no usable version (untagged, or
:latest),deploynow labels with a timestamp rather than the git hash. A git hash there asserts a commit that may not be what:latestresolves to; a timestamp asserts nothing except uniqueness, which is the honest answer when the reference genuinely does not identify a build.:latestis refused outright as a version: reused, it gives every deploy the same container name and leavesCurrentHash == PreviousHash, disabling rollback. -
planresolved the target version by the old rule, so it predicted container namesdeploywould not create. It now uses the same resolution.
LogEntry.image, soteploy logrecords the image a deploy actually ran. The version alone could not be checked against anything, and the log is precisely where a rollback target gets chosen. Omitted for entries with no image (heal, lifecycle, static); old log lines parse unchanged.
-
teploy lock status— report the current deploy lock without touching it. Locks were writable from both ends (deploytakes an auto lock,locktakes a manual one,unlockreleases) but readable from neither: nothing could ask whether one was held. So "a deploy is running", "a human froze deploys" and "a deploy crashed and left its lock behind" were indistinguishable from the outside — the last of which self-heals on the next deploy and needs no intervention at all, while the middle one needs a person.Read-only: it never creates, refreshes or releases a lock, not even the app directory (
teploy lockcallsEnsureAppDirfirst; this doesn't). Staleness comes from the same TTLsAcquireLockbreaks locks by — 30 minutes for auto, 2 minutes for heal, never for manual — via a new exportedstate.LockInfo.IsStale, so the report can't drift from the behaviour it describes. Exits 0 whether or not a lock is held; a held lock is a state to report, not a command failure.--jsonemits{app, server, locked, type, user, message, since, stale}withlockedandstalealways present, which is what lets CI ask "is a deploy running?" without parsing prose or shelling out topgrep. Supports--app/--hostlikelockandunlock, so it works with noteploy.ymlon disk.
teploy remove— a crash-looping process (restarting) is now stopped before removal instead of aborting the whole remove with docker's "container is restarting" error.teploy remove --purge— accessory data owned by the engine's uid (Nucleus: 10001) could not be deleted by a non-root deploy user; the purge now falls back to an in-containerrmas root, matching how the directories were handed over on start.
- Accessories whose image runs as a non-root user (Nucleus runs as uid 10001)
crash-looped on their FIRST start: docker created the missing bind-mount
directory root-owned and the engine could not open it. The ownership
reconcile that already covered upgrades now runs on every start — teploy
creates the volume directories itself and chowns them to the image's user
from inside a throwaway container, so it works for a non-root deploy user
without sudo. Found by
teploy template install teploy-shipon a fresh server.
teploy template deploy|install—--domainis now optional foringress: hosttemplates, which publish onbind:portand have no domain to give (teploy-ship is the first such template). The pairing is validated after render against the template's actual ingress mode, so the error names the rule instead of demanding a flag the template cannot use. New--portflag overrides the template's host port; ontemplate deployit also patches the written teploy.yml, because the file is whatteploy deployre-reads — an override that lived only in memory was silently ignored on the next step.
-
teploy build— build the app's image and stop there. No container is replaced, no route changes, no state is written; the single postcondition is that the image named on the last line exists on the server.--jsonprints{image, version, built}and nothing else on stdout.It closes a real hole.
teploy preview deployruns an image that must ALREADY exist at the current git hash — it falls back to<app>-build-<hash>and never builds one. The only thing that produced such an image wasteploy deploy, which also replaces the running production containers, so "build this branch and look at it on a preview URL" was impossible without shipping the branch to production first. The command's own help said to do exactly that. -
teploy preview deploy --image— run a specific tag instead of re-deriving<app>-build-<git hash>. Without it, a build and the preview that consumes it agreed only because both read the same working directory, which fails silently the moment they run from different checkouts or at different commits.
-
memory:andcpu:limits for apps and each accessory, validated at parse time against docker's own syntax. The plumbing was complete and unreachable:docker.RunConfigcarried the values anddeploy.gopassed them, but no yaml key ever set them, so no user could cap any container. Validating at parse time matters because docker rejects a malformed limit when the container starts — during a deploy that is after the old container is gone, so a typo like8gbwas an outage. It is now a config error that names the offending accessory.This matters most for accessories. A storage engine's own memory budget is its accounting of its own allocations, so an engine that under-counts grows straight past it — one reached 30 GB RSS with a 16 GB budget set and the host OOM killer took an unrelated service down with it. Only the cgroup limit is enforced by the kernel. Limits are also fixed at container creation, so setting
memory:on an already-running accessory warns and names the command that applies it rather than silently doing nothing. -
teploy health --app <name> --host <server>— the last read-only command that still required ateploy.ymlin the working directory now reads server state likestatus,statsandlogs, which is what lets teploy-dash drive it.
-
Health checks probed
localhostregardless of thebind:address. Docker publishes a container on exactly the address it is given, so an app with a specific bind is not reachable at localhost — the probe never connected, retried to the timeout, and failed a deploy whose container was perfectly healthy. Becauseingress: hostdeploys by recreate, stopping the old container first, the failure was not a no-op: the deploy tore down the new container and left nothing running. Setting a bind address turned every deploy into a full outage.The same defect sat on the rollback and start/restart paths, where it is worse — a rollback that cannot health-check its target aborts, so the escape hatch failed exactly when it was needed. Fixed at all four call sites; the probe now resolves the address from the bind host, and the paths with no config in hand read it back from the running container.
-
App-level
memory:/cpu:never reached the container.deploy.Configis a separate struct the CLI populates field by field and nothing copied the two across, so both ends were correct, every unit test on either side passed, and a user settingmemory: 1ggot a container with no limit and no warning. Accessories were unaffected — they pass their config straight through, which is why verifying the accessory path live looked like proof the whole feature worked. The regression test guards the seam rather than either end: everydeploy.Configfield with an identically-namedAppConfigfield must actually be assigned. -
cache:was silently ignored for container apps. OnlyStaticBlockconsumedopts.Cache, so a static-site deploy honoured the rules while every reverse-proxied and load-balanced app dropped them —reverseProxyBlockandloadBalancerBlockdid not take a cache argument at all. The config parsed, validated and merged cleanly, and then produced nothing in the Caddyfile.The failure mode is the bad kind: no error, no warning, and a
teploy.ymlthat reads as if caching is configured. Found on a real deploy that declared"/assets/*": "public, max-age=31536000, immutable"and was serving its HTML, stylesheet and JS with noCache-Control, noETagand noLast-Modified— so every navigation refetched the entire page. It presented as a slow app, and the app was not slow: it rendered in ~2 ms and spent the rest on the wire.Cache rules now render for all three block types.
-
Cache rules are emitted in a stable order. The original loop ranged over a Go map, whose iteration order is randomised, so consecutive deploys that changed nothing still produced a different managed block. That made a Caddyfile diff useless for spotting genuine drift — which matters here, because a diff of that file is the tool for catching exactly this class of bug.
-
Deploy and backup shell commands were interpolated unquoted, SSH host-key verification silently fell back to
InsecureIgnoreHostKeyon a missing$HOMEor an unwritableknown_hosts, and volume/accessory tar restores extracted straight into a live path. Commands are now consistently shell-quoted, both host-key paths fail closed, and a restore extracts to a staging directory and only promotes on success.AcquireManualLockno longer leaves an orphaned non-expiring lock when the metadata upload fails, and the documented install commands mapaarch64->arm64and verify the release SHA-256 againstchecksums.txtbefore extracting.
teploy rollbackmoved aingress: hostapp to a random ephemeral port. The host port is fixed by config, so every version shares it and the current version's port always equals the target's — avoiding it made the recreate reallocate on every rollback rather than only on a genuine collision. A rollback that was supposed to restore service was what broke it. The fixed port is now preserved, and freed the waydeployalready does it: the current web containers stop before the target starts, because two containers cannot bind one fixed port. That ordering was the other half of the bug — rollback used blue/green order and only appeared to work because the port was being reallocated. A failed health check now restores the containers it displaced, so a failed rollback leaves the app where it started rather than down.- Redeploying a version that
rollbackleft stopped failed with docker's raw "name is already in use". Rollback deliberately keeps superseded containers, so the ordinary roll-back-then-fix-forward sequence hit this every time. A stopped container holding the name is now removed and the run retried once; a running one is refused with a clear message, since that means the exact version is already live.
- Webhook deliveries are now signed. Each carries
X-Teploy-TimestampandX-Teploy-Signature: sha256=hex(HMAC-SHA256(secret, timestamp + "." + body)), byte-identical to teploy-observe's and teploy-dash's scheme, so a receiver of all three writes one verifier. Signing the timestamp together with the body is what lets a receiver bound replay. Configure vianotifications.secretor, preferably,TEPLOY_WEBHOOK_SECRET—teploy.ymlis committed, and a signing secret in version control is not a secret. An unset secret sends unsigned, exactly as before. backup schedulereports whether its failure alert will be signed, so a verifying receiver rejecting an unsigned alert doesn't look like a wrong secret.
rollbackandstop/start/restartread only the legacynotifications.webhookkey whiledeploy/backup/previewwent through the multi-channel notifier. An install configured entirely withnotifications.channelstherefore received deploy events and silently never received a rollback — the one event you most want to hear about. All paths now use the same notifier, and event filters still apply, so this does not widen what existing channels receive.- Scheduled-backup failure alerts were the one delivery path that could not be
signed: the signature covers a timestamp,
date +%scontains a%, and crontab treats%as a newline escape. The alert now lives in a script at/deployments/<app>/backup-alert.sh(mode 0700) that cron invokes, where%has no special meaning. Falls back to an unsigned delivery ifopensslis absent, because a delivery that arrives and is rejected tells you something and silence tells you nothing.
teploy drift --app <name> --host <server>: drift detection without ateploy.yml. With no manifest to compare against, drift is derived from what the server itself records — containers of the deployed version that are no longer running, and containers of other versions still running. That covers manual stops and stale versions left behind, and it lets the dashboard report drift for apps it has no config for. JSON output carriesmode(manifestorstate) so consumers know which comparison ran; state mode also states its limitation, since replica-count drift is only visible to the manifest.
- Machine-readable observation commands for automation and the dashboard:
teploy app list(every app deployed on a server, with per-process container detail) andteploy server status <server-or-host>(resources, Docker inventory, and parsed Caddy routes). Both emit stable JSON under--json, contract-tested so consumers can pin to the shape. teploy deploy [server]accepts the server as a positional argument, so a delegating caller can target a server without ateploy.ymlin scope.--project-dirglobal flag: run as if teploy had started in that directory, so a caller never depends on its own working directory.- Canonical manifest model with revision semantics, the shared vocabulary for desired/applied/observed reconciliation.
- Server-side state writes are atomic (write-temp then rename), so concurrent operations can never observe or leave a half-written state file.
-
teploy remove(aliasdestroy) — deploy's inverse, completing the app lifecycle. Stops and removes the app's containers, removes its Caddy route, and deletes its deploy state. Volumes, accessory data, and running accessory containers are preserved unless--purge.--redirect <url>leaves a permanent redirect for the app's domains as a plain unmanaged Caddy block that later teploy operations never touch. Idempotent;--yesand--jsonfor automation; skips the route step oningress: hostservers. -
Zero-config first run:
teploy deploywith no config offers the init flow inline on a TTY, writes the resultingteploy.yml, and continues deploying. Non-TTY behavior is unchanged (hard error, now with ateploy inithint). -
Deploy failure diagnosis: failed deploys now print a rule-based
Likely cause+Try:under the container logs — OOM kills, missing env vars, database connection failures, wrongport:vs what the app actually listens on, 127.0.0.1-only binds, slow boots, disk-full, and bad entrypoints. Deterministic and local; no AI, no network calls. -
Release pinning:
teploy pin [version]/teploy unpin <version>/teploy pins. A pinned version is never removed bykeep_versionsauto-pruning, so a known-good rollback target survives outside the retention window. Pins are stored server-side, so the CLI, dashboard, and auto-deploy all honor the same set. -
Monorepo path filtering for auto-deploy: an
autodeploy.pathsblock inteploy.ymlrestricts which pushes redeploy an app — a push deploys only if it touched a matching file (trailing/**andpath.Matchglobs). Declarative, and fail-open when a push payload has no reliable file list (never silently skips a real change).
teploy initno longer pre-fills anapp.example.comdomain. An empty domain now generates a validingress: hostconfig (prompting for the port), and the server prompt lists knownservers.ymlnames.
teploy plan/teploy drift/teploy heal— read-only deploy dry-run diff, live-vs-declared drift detection (--exit-codefor CI), and bounded self-heal (host-side probe, restart-in-place with backoff, systemd timer, opt-inheal.conf). Web processes only; accessories excluded.teploy kv— shared Nucleus-backed KV store (get/set --ttl/del/exists/incr --by/list <glob>) for cross-app config and flags. Trusted-domain only: one global keyspace, prefixes are hygiene, not isolation.rollout:— staged multi-server deploys. Acanary: "N"or"P%"wave deploys first and rolls back on failure without touching the rest of the fleet;max_failuressets a tolerance budget for the main wave, with a named-straggler exit (never a silent mixed-version fleet) when the budget is exceeded.teploy secret— full OpenBao integration under the existing secret command (--provider local|openbao): setup with auto-unseal (static-env, KMS, or Transit seals), per-app AppRole least-privilege policy,put/get/list, deploy-timesecret:name#fieldenv injection, dynamic DB credentials with an Agent-sidecar for auto-rotation, static-role rotation of an existing DB user, multi-node Raft HA (--replicas), and continuous audit streaming into Observe's tamper-evident trail (secret audit enable/disable). Replaces the earlier standaloneteploy vaultcommand — OpenBao is not HashiCorp Vault, and the naming now matches every other provider-abstracted command (network --provider, etc.).teploy network grant/grants/revoke— just-in-time mesh access: ephemeral, tagged, TTL-bound pre-auth keys for Tailscale or Headscale, no permanent credentials to leak or revoke by hand.setup --hardennow also installs auditd (root/sudo execve plus writes under the Teploy-managed tree) and enables sudo I/O logging forsudoreplay-replayable sessions. The sudoers drop-in is visudo-validated before install so a malformed rule can't lock sudo on the box.env_files:— SOPS+age encrypted dotenv/YAML files merged into the container env at deploy time. File-based, at-rest secrets with no daemon required.scan: true— server-side Trivy vulnerability gate; blocks a deploy on fixable CRITICAL findings (--ignore-unfixedkeeps unpatchable base-image CVEs from wedging every release).firewall:— per-app IP allow/deny lists, user-agent blocking, and a request-body size cap, rendered as Caddy directives that execute before the reverse proxy.access:— self-hosted inbound gate:basic_auth(bcrypt-hashed) orforward_authdelegation to an external identity proxy (Authelia, oauth2-proxy, any OIDC gateway).- S3-compatible backup endpoints (MinIO, B2, R2) via
--endpoint; accessorycommand:override (the missing primitive for MinIO/ntfy accessories) andpublish:port mappings; backup retention (teploy backup prune,--keep-last/--max-age-days, also usable on the scheduled cron path); self-contained scheduled accessory backups (--local --app, noteploy.ymlneeded server-side);teploy accessory verify-backuprestores the latest backup into a throwaway scratch container and proves it's actually usable (real restore + row/key counts, not just "the archive exists"). type: ntfynotification channel; deploy and rollback events now emit to Observe's audit trail via a newaudit:config block.--role/--tagdeploy targeting; best-effort (non-fail-fast) fleet rollback so one server's rollback failure can't strand the rest of the fleet.
- The rsync channel used for static-site uploads now mirrors the control connection's host-key policy instead of unconditionally accepting new keys, closing a downgrade relative to the stricter default.
setup --harden's fail2ban no longer bans the operator's own SSH source — loopback, the Tailscale CGNAT range, and the live setup session's own IP are always exempted. Closes a real lockout seen in the wild (a management IP banned for 24h after two failed pubkey attempts).
- Image pull is now skipped when the image already exists on the server, unblocking locally-built or
docker load-ed images with no registry. - Backups and
verify-backupno longer leave partial archives in/tmpafter a failed dump or upload. - Generic volume backups no longer fail on a live-file tar warning from a WAL rotating mid-read — the archive is a crash-consistent snapshot, and
verify-backup's scratch-container boot remains the actual correctness gate.
- Resilience guide: N+1 topology, durable-state rules, and a human-confirmed dead-server rebuild runbook.
- CI/CD deploy recipe for Forgejo Actions and GitHub Actions, plus the no-secrets autodeploy-webhook alternative.
- OpenBao/secret feature docs (HA, seals, static-role rotation, audit streaming) and a Cloudflare caveat note for
firewall:'sremote_ipmatching.
- The embedded
teploy uiweb dashboard (internal/ui/) has been removed. It was an unauthenticated, localhost HTTP/WebSocket server that duplicated a strict subset of teploy-dash — the dedicated dashboard product (separate repo, real auth, monitoring, its own releases). Maintaining a second, weaker dashboard inside a CLI whose identity is "single binary, no management server" was a standing security surface (no auth/CSRF/Origin checks) and a maintenance/drift tax. The CLI is now a pure deploy engine; use teploy-dash for a dashboard (it runs as a single binary too, including locally).
teploy app exec [--process web] -- <cmd>— run a one-off command inside the app's running container (resolved by theteploy.versionlabel): database migrations, seeds, rake/manage.pytasks, etc. Runs in the existing container via its shell, streams output, and exits non-zero if the command fails. Works from an app directory or with--app+--host.teploy execremains server-level (raw SSH); this is the container-level counterpart.teploy accessory exec <name> -- <cmd>— the same for an accessory container, e.g.accessory exec db -- psql -U postgres -c 'SELECT 1'oraccessory exec cache -- redis-cli INFO. (An interactive REPL variant —console/db— was considered but deferred; it requires PTY handling that can't be unit-tested cleanly. Non-interactive queries are covered by this command.)--app <name>flag onstatus,stats,log,logs,lock,unlock,maintenance on/off,start,stop,restart,env set/get/list/unset, andaccessory list/stop/start/logs. With--app+--host, these commands act on a deployed app by reading server-side state instead of requiring ateploy.ymlin the working directory — the same modelrollback --appalready used. This is what lets teploy-dash (and any automation without an app checkout) drive these commands for arbitrary apps. A sharedresolveApphelper unifies the cwd-teploy.ymlpath and the--app+--hostpath.maintenancekeeps its Caddy-ingress guard only on the cwd path (server state doesn't record ingress mode).accessory upgrade/backup/restorestay cwd-bound — they need the accessory's image config fromteploy.yml, which server state doesn't carry.keep_versions: N— auto-prune older app versions on deploy. The current and immediately-previous versions are always protected; older versions beyondNare removed (containers + images). Container-deploy only — static deploys usekeep_releases. Default0= keep everything (legacy behavior).healthcheck.<proc>.disable— per-process override that passes--no-healthchecktodocker run. Useful for worker containers that share an image with a web container and would otherwise inherit a useless HTTP healthcheck.ingress: external— opt out of Caddy entirely. The container still joins theteployDocker network with its app-name alias, but Teploy doesn't write or reload the Caddyfile. For users fronting the app with Cloudflare Tunnel, Tailscale Funnel, nginx, AWS ALB, or any other external ingress.
secret setpreviously interpolated the secret value intoecho %q | age …, which uses double quotes — so the remote shell still expanded$, backticks and backslashes, silently corrupting values and executing a value like$(...)as the SSH user. Secrets are now passed viaprintf '%s'with single-quoting (no expansion). The same%q/double-quote pattern indocker.Execis fixed too, anddocker runnow single-quotes every interpolated value (name, image, env values, volumes, labels) so a value with a space or shell metacharacter can't break or inject into the command. (The command override is intentionally left raw — it's an operator-authored argv.)- SSH host-key verification no longer fails open. When
~/.ssh/known_hostsis absent (the common fresh-box / CI case) teploy previously accepted any host key with no record — no MITM protection ever. It now falls back to trust-on-first-use: the key is recorded on first connect and a mismatch errors thereafter (use--accept-newor clear the entry after a deliberate re-provision). - Auto-deploy webhooks are now authenticated. The listener never actually extracted the signature header, so with a secret every request was rejected and without one any POST to
0.0.0.0:9876triggered a deploy. The listener now captures and verifies theX-Hub-Signature-256HMAC over the raw body and binds127.0.0.1;autodeploy setuprequires a secret (the CLI generates and prints one when--secretis omitted).
- Multi-replica deploys (
replicas: N > 1) could never succeed — every replica was handed the same host port (the allocator only consulted currently-listening sockets, and no container was started yet), so the seconddocker run -pcollided. The deploy loop now excludes ports already claimed in the same pass. teploy rollbackdroppeddomainand the replica port arrays from the persisted state on swap, so a subsequent rollback failed with "no domain in state" and multi-replica apps orphaned replicas 2..N on the next deploy. The swap now carriesdomainthrough and mirrorscurrent_ports/previous_ports.rollback,start, andrestartmatched containers by a-<version>name suffix, which silently skipped every replica web container (named<app>-web-<version>-1). They now match on theteploy.versionlabel that every container carries.env setwrote values needing quoting with Go%q, butdocker run --env-filereads values literally — so the container received the surrounding quotes/escapes, andenv list(which unquotes) disagreed with the running container. Values are now written verbatim; newline-containing values are rejected (the env-file format can't represent them).maintenance offcould delete an app's Caddy route entirely. It ignored the error from reading the stashed pre-maintenance block, so any transient read failure rendered an empty block and removed the route, taking the domain offline. It now no-ops on a missing stash, aborts without touching the route on a read failure, and deletes the stash only after the reload succeeds.teploy updatewas permanently broken: it pointed at the wrong GitHub repo (teploy/teploy) and expected a bare per-platform binary, while releases shipteploy_{os}_{arch}.tar.gzarchives. It now usesuseteploy/teploy, downloads the archive +checksums.txt, verifies the SHA-256 before installing, and extracts the binary.- Preview environments routed Caddy to the host-published port instead of the container's internal port, so every preview 502'd. Rollback already inspected the internal port; preview now does the same.
- Redis accessory backups were unrestorable —
AccessoryRestorehad no redis case, so it looked for a.tar.gzwhile the backup is stored as.rdb.gz. Added a redis restore path (stop → copydump.rdb→ start so it loads the snapshot). env set/unsetcould wipe the env file on a transient SSH failure:readEnvrancat … 2>/dev/nulland treated any error as "empty", so a read that failed mid-operation returned an empty map and the subsequent write overwrote the real file with nothing. Reads now use atest -fguard so a genuine transport error propagates and the write is aborted; a missing file still yields an empty map.teploy server add(and the config layer) silently dropped a server'stagsand cleared itsvpn_ip/role/userwhen re-adding an existing entry —AddServerreplaced the whole record. Tags drive per-host env injection, so losing them broke deploys. AddServer now preservestagsand keeps any optional field not being changed.- Domains are now validated at the Caddy sink (
parseDomains), rejecting entries with characters that would break out of a Caddyfile site address (whitespace, newline,{}#"\). Defense-in-depth beyond config-time validation; a denylist, so legitimate wildcard /host:portaddresses still work. - Scheduled backups: the cron dedup matched the whole command as a
grepregex, but backup commands are full of regex metacharacters (. * $ ( ) /), so re-scheduling could leave duplicate cron entries. Managed lines now carry a stable# teploy-backup:<app>marker and dedup withgrep -vF. mergeConfigsdidn't carryreplicasfrom a destination overlay (teploy.<dest>.yml), so-d prodcouldn't change the replica count.PruneVersionscompared DockerCreatedAttimestamps as strings (lexicographic, wrong across timezones/DST); it now parses them totime.Timewith a deterministic tie-break and protects any container it can't date.RunStreamcould let the command goroutine write to the caller's stdout/stderr after returning on context cancellation (a data race); it now waits for the command to finish first.AppendLogand the autodeploy Caddy-route setup no longer stage through fixed/tmppaths that concurrent runs would clobber — the log line is appended in a single atomic command (base64 +>>, sub-PIPE_BUF) and the webhook route is piped tocurl -d @-over stdin..goreleaser.yml: migratedarchives.formatandformat_overrides.formatto the newformats: [tar.gz|zip]list syntax (goreleaser 2.x deprecation cleanup).teploy setuphung duringfail2baninstall on fresh Debian/Ubuntu VMs becauseapt-get's debconf step prompted for input. Allapt-getcalls (ufw, fail2ban, unattended-upgrades, sudo, age) now passDEBIAN_FRONTEND=noninteractiveto suppress prompts.- SSH auth-failure errors now suggest concrete next steps (
--user <name>,--key <path>,--password) instead of surfacing the rawcrypto/sshunable to authenticate, attempted methods [...]message. Root SSH is disabled on most modern distros, so the defaultrootuser fails silently for first-time users — the new message points to that fix. teploy rollbackpreviously calleddocker starton the stopped previous-version container. On Docker 29.5+ this silently fails to re-publishHostConfig.PortBindings(and detaches the container from custom networks) if another container had taken + released the host port in the interim — a common situation when rolling back after deploying a neighboring app that reused the port. Rollback now uses a newdocker.Client.Restartpath that inspects the stopped container, force-removes it, anddocker runs a fresh one with the same image, network mode + aliases, port bindings, env, bind mounts, named-volume + tmpfs mounts, command, working dir, user, labels, memory + CPU limits, restart policy, and--no-healthcheck NONEmarker. End-to-end verified on Docker 29.5.2 (single-process + multi-process apps, with and withoutingress: external).- Config:
tls:combined withingress: externalis now rejected at validation. With external ingress, the user's CF Tunnel / nginx / ALB handles TLS termination — the cert + key would be uploaded to the server but never wired into a Caddy block. Silent no-op caught upfront rather than silently wasting upload bandwidth.
- README Config section now documents
tls,keep_versions,healthcheck, andingress(previously undocumented). - CLAUDE.md updated for repo state (project structure reflects current packages; stale design-doc references removed).
- Per-app custom TLS certificate via
tls: { cert, key }. Cert + key are local file paths, uploaded to the server on deploy and referenced from the generated Caddy site block. Required when the public hostname is proxied (Cloudflare proxy, Cloudflare Tunnel) so ACME challenges can't reach the origin — terminate TLS with a Cloudflare Origin Certificate instead. The cert survives every deploy (unlike a hand-edited Caddyfile, which the authoritative model overwrites). teploy setupnow recreates legacy single-file-Caddyfile-mount containers, not just--resumeones.
- Stale Caddyfile bind-mount bug.
setup.gopreviously mounted the Caddyfile as a single file;caddy.gowrites it atomically (mv), which swaps the inode. Docker pins single-file mounts to the original inode, so the container never saw route updates andcaddy reloadreloaded stale config. Caddyfile is now mounted via its parent directory (/deployments/caddy:/etc/caddy), preserving updates after atomic rewrites. Affects every server using the prior mount model — recreate Caddy withteploy setupto pick up the fix.
- README URLs updated for
useteploy/teploy→useteploy/teploy-clirepo rename. GitHub auto-redirects, but explicit references now point at the canonical name.
- Caddyfile is now the single source of truth. Admin-API route changes are mirrored back to the Caddyfile so they survive
caddy reload.
- Published container ports bound to
127.0.0.1instead of0.0.0.0— closes a direct-from-internet exposure vector when the server's firewall is permissive. - Health checks tolerate HTTP 3xx responses (apps that redirect from
/to/loginare no longer marked unhealthy). - Asset bridging now runs as root and surfaces failures up the deploy pipeline. Fixes Next.js / Rails 404s on
/assets/*after deploys to non-root-USERimages.
type: staticdeploy mode — rsync + symlink + Caddyfile, no Docker. For static sites (Astro, Hugo, plain HTML).- Scheduled auto-deploy via cron (
autodeploy:block in teploy.yml). -d/--destinationflag for accessory subcommands (consistent with deploy).- Ad-hoc deploy flags:
--app,--image,--domainfor scripting and dashboard use without a teploy.yml. - Per-server replicas (different
replicas:per server in fleet). - Template install (deploy directly from a registry or git URL).
- Server hardening — firewall + SSH config + auto-updates wired into
teploy setup. - Password bootstrap for fresh servers.
- VPN integration in setup (Tailscale / Headscale / Netbird).
- Comma-separated
domain:field for apex + www served from one block.
- Encrypted secrets actually injected into container env (was silently no-op).
- Foreign volume mount detection prevents data loss on swap (deploy aborts with
--migrate-volumeshint when an existing container has a different host path). - Caddy route persistence survives reloads (precursor to the v0.1.4 source-of-truth fix).
- Rollback port bug — used the wrong port when the previous container's
ContainerPorthad changed. - Rollback missing domain in embedded UI.
- Scoop bucket for Windows installs (
scoop install teploy).
- README features table, tool comparison, and badges.
Patch release of initial drop.
Initial release.