-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathteploy.yml
More file actions
319 lines (308 loc) · 18.5 KB
/
Copy pathteploy.yml
File metadata and controls
319 lines (308 loc) · 18.5 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
# Self-hosted Teploy Ship — deployed with teploy.
#
# One image (deploy/build-image.sh), two processes:
# web the dashboard + webhook receivers (Forgejo + GitHub)
# worker executes durable runs (needs the model + git tokens)
# Nucleus runs as a teploy accessory (network alias `<app>-nucleus`, i.e.
# ship-nucleus) so any machine can list/approve runs against it.
#
# No domain: `ingress: host` publishes the web process directly at
# <server-ip>:7460 — the shape for a private/tailnet box. Gate the port
# with your firewall (e.g. `ufw allow from 100.64.0.0/10 to any port 7460`
# for Tailscale-only access).
#
# Required secrets — set ONCE via `teploy secret set` from this directory
# (never put values in this file):
# SHIP_WEB_TOKEN dashboard login + API bearer (web only)
# SHIP_WEBHOOK_SECRET Forgejo/GitHub webhook HMAC (web only)
# SHIP_SESSION_SECRET signs dashboard sessions (web only)
# SHIP_GIT_TOKEN deploy token for repo runs (worker only)
# SHIP_SANDBOX_TOKEN sandbox daemon credential (worker only)
# AI_GATEWAY_KEY project key from `teploy-gateway key create` (worker only)
# (or set ANTHROPIC_API_KEY instead and drop AI_GATEWAY_URL below)
#
# LEAST PRIVILEGE: teploy secrets are app-scoped — `processes:` takes a
# command string, so there is no per-process secret scoping to declare here —
# and they therefore land in BOTH containers. The web process has no use for
# the git deploy token or the model key: it renders a dashboard and verifies
# webhook signatures; it never clones, never pushes, and never calls a model.
#
# Ship handles that itself rather than waiting for the platform: `teploy-ship
# web` DROPS worker-only credentials from the environment before serving (see
# WORKER_ONLY_SECRETS in src/cli.ts), so a web-route bug cannot read a value
# that is not there. Nothing to configure; noted here because the deployed
# container will still show them in `teploy secret list`.
#
# Repository access (worker). Prefer the per-origin form — it names the
# credential AND the origins Ship may clone, so there is nothing else to set:
# SHIP_GIT_TOKENS={"https://github.com":"ghp_…"}
# The single-token SHIP_GIT_TOKEN still works, but because it has no host
# attached it also needs SHIP_REPO_ALLOWLIST to say where it may be sent —
# otherwise a repository URL arriving from a webhook, a Slack message, or an
# issue body is refused rather than being handed the token. Adding a repo on
# the dashboard's Projects page allows it too (the allowlist is the union);
# the env value is the floor.
app: ship
# infra-home. Ship's control plane lives beside the forge because its SANDBOXES
# do not: SHIP_SANDBOX_URL points at compute-1, and src/colocation.ts refuses to
# start if a sandbox daemon ever resolves to the forge's own machine.
#
# This was `smoke` (deploy-test) until 2026-08-27. Pointing the BASE file here
# rather than leaving it on a decommissioned box means a bare `teploy deploy`
# cannot quietly aim at a VM that no longer exists.
server: 100.108.123.49
user: root
ingress: host
port: 7460
platform: linux/amd64
# No image: teploy syncs this directory (see .teployignore) and builds
# from the root Dockerfile ON the server. Prerequisites, run locally
# first: `pnpm run build`, `cd web && pnpm run build`, and the SDK
# tarball pack from deploy/build-image.sh (the @neutron-build packages
# are ahead of their npm releases, so the image installs them from
# deploy/vendor/ tarballs).
processes:
web: web --store nucleus --port 7460
worker: worker --store nucleus --interval 5
# Ceiling per app container (web and worker each), not a reservation — they sit
# at ~140 MB and ~30 MB respectively, because durable-run work happens inside
# the sandbox container rather than in the worker process. This exists so a leak
# or a runaway run cannot take the whole 4 GB box down with it, which on this
# host would also mean the gateway and the Forgejo runner.
memory: 1g
# Nucleus sidecar. NUCLEUS_ALLOW_NO_AUTH: nucleus's pgwire SCRAM password
# auth panics upstream ("Salt required for SCRAM auth source"), so
# password auth is not usable yet; the accessory is only reachable on the
# app's private docker network. Same posture observe's teploy.yml ships.
accessories:
nucleus:
# PINNED, not :latest. Nucleus publishes :latest on every push to main, so a
# redeploy would silently adopt whatever landed since — and a storage engine
# is the last thing that should change underneath you unannounced.
#
# v0.1.8 (2026-08-21) supersedes v0.1.5. It keeps the 2026-08-03 KV/LSM
# memory rework that fixed the write-reject-under-memory-pressure outage,
# and adds the run of durability fixes released in 0.1.7/0.1.8 — several
# acked-but-unfsynced writes, three confirmed CRITICALs, and streaming
# routes that bypassed column masking and the SELECT grant.
#
# The concrete reason to move off v0.1.5: it cannot bootstrap a fresh
# Observe schema. Migration 28 fails with `events_pre028 not found in
# storage`, so a new install or a restore that replays the schema cannot
# come up. Verified on 2026-08-21 that Observe's suite passes against
# v0.1.8 (34 packages) and fails three of them against v0.1.5.
# Bump deliberately, after an instance has proven a release.
#
# v1.0.0 (2026-08-31), up from v0.1.8. Two wire-protocol fixes are why this
# release exists. The RowDescription a client received did not always
# describe the DataRows that followed it — 20,800 of 125,419 statements
# over a ten-second concurrent read were answered with a description that
# did not match their rows, and a client reading fields positionally throws
# while one reading by name silently decodes a row with fields missing.
# And Describe could EXECUTE a side-effecting statement, because whether to
# describe statically or by running the statement was decided by a text
# scan that a comment between a function name and its arguments defeats.
#
# ROLLING BACK TO v0.1.8 IS NOT A BINARY SWAP. DB_FORMAT_VERSION is still
# 2, so the engine will NOT refuse the older binary, and the rollback
# runbook's decision tree tells you to try exactly that first. S63 added
# transaction-tagged KV WAL records (0x06-0x09) that 0.1.8 has no case for;
# its replay fallthrough is RecordStep::Stop, so it treats the first tagged
# record as end-of-log and silently discards every KV record after it. The
# tags appear only where a KV mutation runs inside a coordinating SQL
# transaction. A v0.1.8 physical snapshot still restores into v1.0.0, so
# the pre-upgrade backup is the way back — fix forward, do not swap.
# v1.1.1: upgrade from the live v1.0.2 engine (the previous config pin
# incorrectly remained v1.0.0). First open migrates the WAL to v2.
# Stop writers and the engine, verify a full data-directory backup, and
# rehearse on a restored copy before upgrading. Keep the full archive:
# mvcc.wal.v1 can be retired on the next clean open. Rollback requires
# restoring pre-upgrade data with its matching image, not an image swap.
# Migration can take minutes; wait for health instead of restart-looping.
# The binary's startup banner still says v1.0.2; verify the image tag.
image: ghcr.io/neutron-build/nucleus:v1.1.1
port: 5432
# Hard cgroup cap. NUCLEUS_MAX_MEMORY_MB below is the engine's own
# accounting and is advisory — it was set to 16 GB on the day this engine
# reached 30 GB RSS and OOM-killed an unrelated service on the same host.
# Set above the engine budget so the engine's graceful eviction gets to run
# first, and the kernel cap is only the backstop.
memory: 1500m
env:
NUCLEUS_ALLOW_NO_AUTH: "1"
NUCLEUS_ALLOW_INSECURE_CLUSTER: "1"
NUCLEUS_ALLOW_INSECURE_REPLICATION: "1"
# Engine-wide memory budget (honored by nucleus >= 56a7f59; the engine
# rejects writes at 90% RSS of this, and its 512 MB default is too
# tight once real data is resident). Smoke box has 4 GB total.
NUCLEUS_MAX_MEMORY_MB: "1024"
volumes:
nucleus-data: /data
env:
NUCLEUS_URL: postgres://nucleus@ship-nucleus:5432/nucleus
# The trusted delivery copy (Package B): the clone mounted above. Unset
# means the whole approved-delivery execution path is OFF and approvals
# are held with the reason. Provisioned 2026-09-22 for the scratch
# acceptance target ship-delivery-proof (:7480 on this box).
SHIP_DELIVERY_DIR: /srv/ship-delivery/ship-journey-proof
SHIP_MODEL: anthropic/claude-sonnet-5
# Codebase indexing (A1): repo runs embed the clone into Nucleus vectors
# and the agent gets the ```search action. Embeddings ride the gateway's
# /v1/embeddings passthrough to its ollama accessory (nomic-embed-text,
# pulled once — see teploy-gateway/teploy.yml). Unset to disable indexing.
SHIP_EMBED_MODEL: ollama/nomic-embed-text
# Route model calls through teploy-gateway, deployed as its own internal
# teploy service (see teploy-gateway/teploy.yml) — reachable on the teploy
# network at its app-name alias. Drop this to call providers directly.
AI_GATEWAY_URL: http://ship-gateway:8089
# Auto-launch is OFF by default — a labeled issue only ever becomes a
# PROPOSED task you launch from the dashboard. To let the worker launch
# a source's tasks without a click, set (via `teploy secret set`, not
# here — earn it deliberately per source):
# SHIP_INTAKE_POLICIES={"forgejo":"auto"}
# The daily count cap, concurrency ceiling, and per-source spend budget
# still bound it (SHIP_DAILY_BUDGET_USD, SHIP_MAX_CONCURRENT_RUNS).
#
# SHIP_MAX_CONCURRENT_RUNS is now an OVERRIDE, not the setting (B1): the worker
# measures the box's cpu, memory and disk and derives its own ceiling, and the
# Fleet page says which limit is binding. Leave it unset unless you have a
# reason to disagree with the measurement. Measured on deploy-test 2026-08-26:
# the derived ceiling is 1, disk-bound — docs/capacity.md's recommended 4 was
# wrong for that box, and nothing was measuring disk at all.
#
# SHIP_MAX_STEPS caps how many model turns one durable run gets (default
# 40). Raise it for repos where tasks need more exploration; the spend
# caps above are what actually bound cost.
#
# SHIP_HARNESS picks what edits the tree: native (default, Ship's own loop,
# needs nothing in the sandbox image), claude-code or opencode (the vendor
# binary must be in SHIP_SANDBOX_IMAGE and the sandbox needs egress). It is
# recorded on each run at enqueue. See docs/adapters.md — including why a
# subscription-fed run is counted as unpriced rather than shown as $0.
#
# Run durable worker tasks in a teploy-sandbox container (instead of the
# slim worker image, which has no python/curl) so repo tasks whose tests
# need other tools can complete. The daemon is a host systemd service, so
# containers reach it at the teploy bridge gateway (172.18.0.1). Set
# SHIP_SANDBOX_TOKEN via `teploy secret set`.
# SHIP_SANDBOX_IMAGE is the worker-wide DEFAULT. A repo's project record
# (dashboard Projects page / `teploy-ship project set <repo> --image …`)
# overrides it for that repo's runs, so one worker serves a Go repo and a
# pnpm repo in their own images.
#
# SHIP_SANDBOX_URL takes a LIST (comma-separated) since B2: each run is placed
# on the least-loaded healthy daemon, and adding a box is adding it here.
#
# The image is built from `images/` in this repo (`images/build.sh`) — it used
# to be hand-built on the box and existed nowhere else. Go is pinned to 1.25:
# three repos need it, and on 2026-08-26 every Go pull request arrived marked
# `tests: failed` against 1.24. This line said `golang:1.24` while the live
# worker was already running the harness image, which is the drift that cost
# that day.
# compute-1. NOT a local address: nothing model-authored may run on the forge's
# box, and src/colocation.ts refuses to start the worker if a sandbox daemon
# ever resolves to the same machine as the forge. Comma-separate to add more
# compute; each run is placed on the least-loaded healthy daemon.
SHIP_SANDBOX_URL: http://100.107.192.39:7439
SHIP_SANDBOX_IMAGE: ship-sandbox-harness:dev
SHIP_SANDBOX_NETWORK: egress
# Pinned rather than derived, and this is a KNOWN GAP being worked around: B1
# senses the box the WORKER runs on — infra-home, 4 cpu / 30 GB — not
# compute-1 (4 cpu / 8 GB, shared with the Forgejo Actions runner), which is
# where sandboxes actually consume resources. The derived ceiling would be a
# number about the wrong machine. Raise it once the sensing understands
# remote pools.
SHIP_MAX_CONCURRENT_RUNS: "2"
# --- Evidence on the pull request (P1) ---------------------------------
# Each of these only ASKS. Whether the evidence can actually be produced is
# the worker's own configuration, and a leg that is not configured records
# itself as disabled rather than failing the run. A worker wired for none of
# them adds nothing to the PR at all, deliberately: printing "not deployed,
# not measured, not tested" on every pull request trains a reviewer to skip
# the section that sometimes carries the real thing.
#
# Ship runs the suite itself, after the agent stops and BEFORE the push, so
# "tests passed" describes the code that becomes the PR. The agent's own
# account of its testing is exactly the claim the finish gate exists because
# models get wrong.
#
# CAVEAT: SHIP_TEST_COMMAND is the worker-wide DEFAULT, used for repos with
# no per-repo evidence entry. Per-repo config exists — `teploy-ship evidence
# set <repo> --test-command "..." --observe-service <svc>` — and wins over
# this value, so one worker can serve repos with different suites. A command
# the sandbox image cannot run reports "not run", never "failed".
SHIP_TESTS: "1"
SHIP_TEST_COMMAND: "go test ./..."
# Read the affected service's error rate and latency either side of the
# change and put the numbers on the PR. Below OBSERVE_MIN_REQUESTS in either
# window it says "not enough data to compare" and shows raw counts instead —
# a preview environment serves almost nothing, so that is the common case and
# the refusal IS the feature. Where a comparison is reported it is labelled
# correlation; nothing claims causation.
#
# These three are the worker-wide DEFAULT (used for repos with no per-repo
# evidence entry). A repo's own service is configured per repo:
# teploy-ship evidence set tyler/fylun --observe-service fylun-web
# which also makes the run read THAT service regardless of OBSERVE_SERVICE.
#
# OBSERVE_READ_TOKEN is a share token and is set via `teploy secret set`.
SHIP_TELEMETRY: "1"
OBSERVE_URL: http://100.108.123.49:3000
OBSERVE_SERVICE: fylun-web
# Without a repo (here OR in per-repo evidence) the leg is OFF, by design:
# on 2026-08-21 a worker watching fylun-web put its RED metrics on a pull
# request that changed one line of Go in an unrelated repo, and the reviewer
# saw "p95 up 2653ms" under a change that could not have caused it.
OBSERVE_REPO: tyler/fylun
# --- Tailnet previews (the 2026-09-24 preview ruling) -------------------
# SHIP_PREVIEW_DIR is a preview ROOT: one clone per declared preview app at
# <dir>/<app> (resolvePreviewTarget in src/deploy.ts), origin = the forge's
# bare repo bind-mounted read-only below, the same shape as the delivery copy.
# SHIP_PREVIEW_TAILNET_IP is the preview TARGET's tailnet IPv4. It implies
# base `<ip>.sslip.io` + plain HTTP + a 100.64.0.0/10 allowlist, so a preview
# answers at http://preview-<slug>-<hex8>.100.107.192.39.sslip.io and the
# same Host header sent to any non-tailnet address gets 403. It is app-level
# env, so the web process gets it too — that is what lets the run page's CSP
# frame the preview (without it the panel only links out).
#
# The target is compute-1, NOT this box. A preview runs unreviewed
# model-authored code, and every container here shares the `teploy` docker
# network with ship-nucleus (NUCLEUS_ALLOW_NO_AUTH): a preview on infra-home
# could write this store directly, approvals included. compute-1 already runs
# model-authored code (the sandboxes), and nothing else answers :80 there.
# Its Caddy publishes :80 through iptables DNAT, so the allowlist sees the
# real client address (measured 2026-09-24: tailnet 200, LAN address 403).
SHIP_PREVIEW_DIR: /srv/ship-preview
SHIP_PREVIEW_TAILNET_IP: 100.107.192.39
# Main's URL per preview app, for the visual rung (tailnet mode cannot derive
# it from the preview host).
SHIP_PREVIEW_MAIN_URL: ship-preview-proof=http://100.107.192.39/
# Durable state for the worker's file paths (nucleus mode needs little,
# but the image declares VOLUME /data).
volumes:
ship-data: /data
# --- Trusted delivery copy (Package B, S14) — see src/delivery.ts -----
# The worker executes APPROVED deliveries against this clone of the
# repository being delivered: it fetches the merged SHA into a detached
# worktree here, builds and deploys from there, and reads the target back
# through the same copy's teploy.yml. Deploy credentials (the worker's key
# in /srv/ship-delivery/ssh) stay on the worker and never
# enter an agent sandbox. The clone's origin is the forge's own bare repo
# bind-mounted read-only — fetching needs no credential and cannot drift
# from what the forge actually holds. Both mounts land in web too (volumes
# are app-level); web has no code path that reads them.
/srv/ship-delivery/ship-journey-proof: /srv/ship-delivery/ship-journey-proof
/deployments/forgejo/volumes/data/git/repositories/tyler/ship-journey-proof-20260921.git: /srv/forge-bare/ship-journey-proof.git:ro
# --- Preview root (tailnet previews, see env above) -------------------
/srv/ship-preview: /srv/ship-preview
/deployments/forgejo/volumes/data/git/repositories/tyler/ship-preview-proof.git: /srv/forge-bare/ship-preview-proof.git:ro
# Operator-held config (read-only): SHIP_MODEL_ROUTING_POLICY points at a
# file under here (e.g. /srv/ship-config/routing-policy.json), so the
# routing shadow/on policy can change without an image rebuild.
/srv/ship-config: /srv/ship-config:ro
# The worker's ONE SSH identity (the CLI reads ~/.ssh/id_ed25519) plus
# known_hosts for every target it deploys to. Rotated 2026-09-24 (L8): the
# wave-9 delivery scratch key was retired (authorized nowhere); the current
# key is authorized on compute-1 for previews only. A delivery test against
# another target authorizes this key there and revokes it after.
/srv/ship-delivery/ssh: /home/node/.ssh:ro