Skip to content

feat: host-portable planning, a credential helper, and tofu-aware id resolution - #183

Merged
2000game merged 2 commits into
mainfrom
feat/portable-planning-tofu-bridge
Sep 21, 2026
Merged

2000game merged 2 commits into
mainfrom
feat/portable-planning-tofu-bridge

Conversation

@2000game

Copy link
Copy Markdown
Member

Closes #178, #179, #180, #181 — four changes the tier-0 OpenTofu cutover needs, kept in one PR because they land in the same files (the resolver, the plan operation, the auth stack) and are verified together.

#178 — a right this host does not have warns instead of aborting

A declared right absent from the active permission catalog aborted the whole plan, so one config could not serve an estate whose instances have different modules installed (13 of 113 rights do not exist on eqrm-dev). Now:

  • absent here, present in ct's bundled catalog → warning naming the right, the declaration and both catalog versions; that one grant is skipped for this host (never granted, never revoked) and everything else still plans;
  • absent from every catalog ct knows → still a hard error with the "did you mean" hint. That is a typo or a deleted right (churchreport:edit masterdata), which is what the error was written for.

The bundled catalog is the oracle that tells the two apart, so the lenient verdict requires a per-instance capture to be active — a repo that has not run ct permissions catalog --refresh sees exactly today's behaviour. --strict-catalog on plan/apply restores the old behaviour everywhere. Exit codes unchanged: a skip is not a pending change.

Same split for a preserveUnknown dimension — one no right on this host scopes by is reported and ignored, one no catalog knows is still rejected at config-eval time.

Worth stating explicitly: ct coverage never loads the config, so nothing there needed changing (contrary to my first read of the issue). Whatever a repo's coverage:check script wraps is ct plan, which is what this fixes.

#179 — ct auth token

Emits the ChurchTools session, not the personal login token. The token is permanent, unscopable and an admin credential on prod; the session it buys expires, is dropped by ct auth logout, and a copy that leaks into a tofu debug log dies within hours. The token never leaves the Keychain.

  • credential on stdout, nothing else there; every message on stderr, so $(ct auth token --raw) is safe;
  • failure → empty stdout, non-zero exit, remedy named;
  • a terminal is refused without --allow-tty;
  • expiresAt is ct's 12h reuse ceiling, not a promise — treat a 401 as "ask again".

Because this is meant to be called on every tofu run, it comes with a cross-process brake on login handshakes per host (3s spacing, waited out; 20/rolling hour, refused with the time the window frees up; CT_NO_LOGIN_THROTTLE=1 to disable). The counter is a cache file of timestamps — no credential in it.

Dependency: the provider's token attribute cannot consume a cookie. This half is inert until terraform-provider-churchtools learns a cookie+CSRF auth mode; the output shape is now concrete enough to file that against. CI is unaffected — it keeps passing the token from a GitHub secret.

#180 — a reference is not a declaration

Declaredness is now decided from the resources the config actually declares, matched on type and key, not from the key appearing somewhere. So the tier-0 cutover no longer needs --force for 49 of 50 entries, and the check survives for the case it is meant to catch.

A referenced-but-undeclared key is removed with a warning saying how many references stay behind and that they now resolve live by name. One class of reference still blocks removal, for the guard's original reason rather than a spelling one: a group is managed-only with no live catalog to fall back to, so dropping a group that a ct.groupRole domain or a group scope names makes the next plan fail to resolve it. That refusal now says so instead of claiming the key is declared — two existing tests changed expectations accordingly (behaviour unchanged, wording accurate).

#181 — resolving what OpenTofu now owns

Once tier-0 leaves ct's state, leftover references fall back to matching the key against the live name, which cannot work for keys that were never name-derived (status_unbekannt vs Unbekannt), and numeric ids are not portable (39 of 43 tier-0 ids differ between the two hosts).

ct now resolves through a committed .ct/ids.<host>.json, slotted after its own managed state (ct never stops trusting what it owns) and before the live catalog (an exact table beats a name guess). ct plan names the map in its header next to the permission catalog.

  • ct export tf writes it from the state it exports (--no-ids to skip);
  • ct ids sync --tofu-state <file|-> refreshes it from tofu's own state, so tofu state pull | ct ids sync -e prod --tofu-state - works on any backend and no S3 client enters ct-cli;
  • ct ids list shows what the resolver would use;
  • the map is host-checked on load (a foreign map would resolve every reference to a real, wrong resource), id: 0 round-trips, and a sync never replaces a populated map with an empty one.

I went with .ct/ids.<hostslug>.json rather than the ct-ids.<env>.json the issue names: it works without --env, sits beside permission-catalog.<host>.json, and cannot be mismatched to a host resolved some other way.

Verification

  • npm test — 1184 passed, 5 skipped (108 files); 43 new tests across permission-catalog-host-difference, auth-token-command, login-throttle, tofu-id-map, ids-sync, plus the state rm and export additions
  • npm run typecheck, npx eslint src tests, npm run format:check, npm run build — all clean
  • node .github/scripts/docs-staleness.mjs — all 6 pages current (4 re-read and re-signed)
  • node dist/index.js ids list against a real host — reports "no id map" cleanly, no network

Docs: new docs/opentofu-migration.md (the id map and ct auth token), a "A right this host does not have" section in docs/handbuch/permissions.md, and the state rm paragraph in the README rewritten.

…resolution

Four changes the tier-0 OpenTofu cutover needs, and one that makes a single
config usable across an estate whose instances differ.

**#178 — a right this host does not have warns instead of aborting.** A declared
right absent from the active permission catalog took down planning for
everything else, so one config could not serve two instances with different
modules installed (13 of 113 rights do not exist on eqrm-dev). It is now skipped
with a warning naming the right, the declaration and both catalog versions —
never granted, never revoked — while everything else still plans. A name no
catalog ct knows defines stays a hard error, which is the case that guard was
written for; ct's bundled catalog is the oracle that tells the two apart, so the
verdict needs a per-instance capture to be active. `--strict-catalog` on plan and
apply restores the old behaviour. Same split for a `preserveUnknown` dimension.

**#179 — `ct auth token`, a credential helper.** Emits the ChurchTools SESSION
rather than the personal login token: it expires, `ct auth logout` kills it, and
a leaked copy dies in hours instead of being a permanent admin credential. The
credential goes to stdout and nothing else does, a failure writes nothing there,
and a terminal is refused without `--allow-tty`. A cross-process brake bounds
login handshakes per host (3s spacing, 20/hour) so calling this on every tofu run
cannot burst into ChurchTools' login rate limit.

**#180 — a reference is not a declaration.** `ct state rm` refused a key the
config only referenced, which made `--force` mandatory for 49 of 50 tier-0
entries and suppressed the check for the case it is meant to catch. Declaredness
is now decided from the resources the config actually declares, matched on type
and key. References are reported, not refused — except a group, which is
managed-only with no live catalog to fall back to, and is refused in its own
words.

**#181 — resolving what OpenTofu now owns.** Once tier-0 leaves ct's state, the
references left behind fall back to matching keys against live names, which
cannot work for keys that were never name-derived (`status_unbekannt` vs
`Unbekannt`), and numeric ids are not portable (39 of 43 differ between hosts).
ct now resolves through a committed `.ct/ids.<host>.json` — after its own state,
before the live catalog — written by `ct export tf` and refreshed by
`ct ids sync --tofu-state -` from a `tofu state pull`, so no S3 client enters ct.
The map is host-checked on load, and a sync never replaces a populated map with
an empty one.

Closes #178, #179, #180, #181
Review findings on #183, each reproduced against the built binary before
being fixed and pinned by a test that fails without it.

The id map is the only thing resolving a leftover `campus:` or
`personStatus:` reference once tier-0 leaves ct's state, and its keys are
not name-derived, so a dropped entry is a hard plan error rather than a
degraded guess. Three ways it was dropped:

- a partial `ct export tf --only <type>` rewrote the whole file from that
  run's entries, so exporting one type at a time discarded every type it
  had not reached yet. Types the run covered are now rewritten wholesale,
  types it never selected are carried over.
- a full export against the POST-CUTOVER state — tier-0 gone from the
  config and the state file, which is the end state the map exists for —
  found nothing and wrote `entries: 0` over a good map. An export that
  maps nothing now keeps what is there and says so, matching the guard
  `ct ids sync` already had.
- `Number(attributes.id)` accepted null, "", [] and false as 0, and 0 is
  a real ChurchTools id, so a mid-create or tainted tofu resource became
  a reference silently resolving to the wrong live object. A `for_each`
  block likewise mapped its block label to one arbitrary instance id and
  reported every real key as removed; both are now refused and reported.

Also, from the same review:

- the login throttle's hourly cap gates every ct command, not just
  `ct auth token`, and there is no session cache off macOS — so 21 ct
  invocations in an hour on Linux CI hard-failed where they had always
  worked. Cap raised to 120: the 3s spacing is what protects the
  instance, this is the backstop for a loop that keeps going. The
  unlocked read-modify-write, and the fact that CI is *not* exempt, are
  now stated rather than implied away.
- `ct ids sync` refused to run while the map it exists to replace was
  malformed; it now regenerates and warns.
- `--dry-run` exited 0 on the empty-state condition the real run exits 1
  on, so a CI gate built on it passed exactly when it should fail.
- `--tofu-state` resolved against process.cwd() rather than the project
  cwd, unlike every other path in that operation.
- `--strict-catalog` set a process global nothing reset, which in the
  HTTP adapter contracts.ts anticipates would leave every later plan
  strict; now restored in a finally.
- `ct state rm` checked the key-only state-only-reference refusal before
  the type-and-key declaredness one, so a key that was both got the
  vaguer message.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A right absent from the host's catalog should warn, not hard-error — it makes a portable config unusable across instances

1 participant