Skip to content

fix(gcp): allow scalable custom-image preconfigured hosts #2382

Description

@Brad-Edwards

Bug

The preconfigured-machine-host capability currently requires a Compute Engine machine image. Google limits a source machine image to six VM creations in 60 minutes. A multi-participant event can therefore fail range provisioning even when project CPU, disk, and instance quotas are available. Retrying the same source within a short window cannot remove this hard limit.

Google documents the restriction at https://cloud.google.com/compute/docs/machine-images/create-instance-from-machine-image and documents a 20-instances-per-second image creation limit for custom images at https://cloud.google.com/compute/docs/instances/create-vm-from-custom-image.

Expected behavior

Allow a preconfigured host whose complete runtime is contained in one boot disk to use an exact administrator-bound custom image as well as the existing exact machine-image form. Preserve the same closed readiness, guest identity, nested virtualization, explicit service account, private networking, and deterministic image/profile binding. Existing machine-image hosts remain compatible.

Acceptance criteria

  • A preconfigured host with an exact custom-image reference validates and provisions through the existing generic image/profile seam.
  • The custom-image VM request explicitly enables nested virtualization and preserves the current security, network, identity, and guest-readiness contract.
  • Image-backed host retry/reconciliation cannot adopt a conflicting VM and leaves its boot disk owned by the VM lifecycle.
  • Existing machine-image and ordinary image-backed guest behavior remains unchanged.
  • Synthetic tests cover validation, request shape, RAES realization/reconcile, and regression cases, with no scenario-specific code or configuration.

Activity

  1. Brad-Edwards commented on Sep 24, 2026

    @Brad-Edwards
    OwnerAuthor

    🛠️ Picked up by /implement - driver codex, branch 2382-retry-image-throttle, 2026-09-24T04:16:19.408Z.

  2. Brad-Edwards commented on Sep 24, 2026

    @Brad-Edwards
    OwnerAuthor

    gc workflow phase recorded: preflight (issue #2382). Posted by the MCP server to enforce ordering between workflow steps (issue #794 MVP-2). Do not edit or delete — used by downstream tools to gate phase prerequisites.

  3. changed the title [-]fix(gcp): retry transient machine-image operation throttling[/-] [+]fix(gcp): allow scalable custom-image preconfigured hosts[/+] on Sep 24, 2026
  4. Brad-Edwards commented on Sep 24, 2026

    @Brad-Edwards
    OwnerAuthor

    The initial retry-only hypothesis was superseded by the provider-documented limit of six creates per machine image per 60 minutes. A bounded retry would still fail normal multi-range admission, so the acceptance contract now uses the existing custom-image provision path for single-boot-disk preconfigured hosts while retaining machine-image compatibility. This is the smallest generic correction that removes the hard source limit without introducing scenario-specific behavior.

  5. Brad-Edwards commented on Sep 24, 2026

    @Brad-Edwards
    OwnerAuthor

    gc workflow phase recorded: preflight (issue #2382). Posted by the MCP server to enforce ordering between workflow steps (issue #794 MVP-2). Do not edit or delete — used by downstream tools to gate phase prerequisites.

  6. Brad-Edwards commented on Sep 24, 2026

    @Brad-Edwards
    OwnerAuthor

    gc workflow phase recorded: plan (issue #2382). Posted by the MCP server to enforce ordering between workflow steps (issue #794 MVP-2). Do not edit or delete — used by downstream tools to gate phase prerequisites.

    Outcome

    Support a preconfigured GCE participant host from an administrator-pinned exact custom image as well as the existing machine-image source. The source type remains independent of the guest readiness capability. Keep legacy behavior and fail closed on stale deterministic VMs.

    Evidence and boundaries

    The existing machine-image source has a documented six-instance-per-60-minute create limit; a single-boot-disk custom image is the supported high-throughput source. The selected custom image is exact and project-qualified; no family, mutable alias, scenario-specific selector, executable adapter installation, or new IAM privilege. Preserve private network, host identity, Shielded VM settings, nested virtualization, explicit service account selection, boot-disk auto-delete, and the fixed participant-readiness contract.

    Test-first work

    1. Add synthetic failing tests for registry mappings, runtime plugin bindings, closed operation-input candidates, and config/profile validation: accept a complete host contract with an exact custom image; reject incomplete fields, families, malformed refs, and ordinary-image regressions.
    2. Add failing request/RAES tests: one auto-deleting custom-image boot disk, nested virtualization, Shielded settings, explicit no-service-account selection, output fields, conflicting existing-VM rejection, and unchanged machine-image behavior.
    3. Add failing tests for bounded participant liveness/readiness routing on both image kinds, including failure propagation before range readiness; investigate RAES apply/activation paths and make the canary boundary explicit.
    4. Implement minimal validator, request-renderer, reconcile, readiness, and UI changes. Update public architecture documentation with the source/capability distinction and portable operator prerequisites. No private scenario details enter this repository.

    Verification and delivery

    Run focused tests red then green; run affected frontend tests/typecheck, provisioner tests, platform ruff check/format, import-linter for Python import changes, and python3 scripts/adr_guard/adr_guard.py --all --level ci. Use the canonical Brad-Edwards/shifter repository. Publish a conventional-commit PR against origin/dev, complete pre-push review and CI/Sonar/readiness, and leave merging to the user. No protected-branch merge.

  7. Brad-Edwards commented on Sep 24, 2026

    @Brad-Edwards
    OwnerAuthor

    gc_codex_review — sanitized deferred publication for issue #2382, cycle 1 of 1
    Reviewed revision: e84202a69ee13aa7262c2fd050a5b1ef1d7d387eef899742ebcd5777b62961c9

    Architectural read

    The source-kind and guest-capability split fits the existing profile and readiness seams. Two resource-lifecycle defects remain before publication; no concrete security issue was identified in the reviewed diff.

    Verdict: ship-with-fixes

    Findings

    core-F1 — Late ownership conflicts can leave owned resources behind

    • Classification: one-off
    • Decision: fix
    • Location: shifter/engine/provisioner/raes_gcp_apply.py
    • Rationale: Add ownership-scoped partial cleanup that preserves any conflicting VM, with a regression test for a late conflict after earlier resources are created.

    core-F2 — Converge disk deletion policy for adopted custom-image hosts

    • Classification: class
    • Decision: fix
    • Location: shifter/engine/provisioner/raes_gcp_apply.py
    • Rationale: Apply attached-disk auto-delete convergence to either preconfigured source, and test adoption and teardown.
  8. Brad-Edwards commented on Sep 24, 2026

    @Brad-Edwards
    OwnerAuthor

    gc_codex_review pre-push cycle 1 of 1 complete for issue #2382 on branch '2382-retry-image-throttle'. Posted by the MCP server to enforce the pre-push hard-cap-1 contract (issues #796, #804, #906). Do not edit or delete — used by the next gc_codex_review (uncommitted) invocation to count cycles.

  9. Brad-Edwards commented on Sep 24, 2026

    @Brad-Edwards
    OwnerAuthor

    Review decision record — codex cycle 1 (issue #2382)

    Reviewer: codex
    Cycle: 1
    Verdict: ship-with-fixes

    Architectural read:

    The source-kind and guest-capability split fits the existing profile and readiness seams. Two resource-lifecycle defects remain before publication; no concrete security issue was identified in the reviewed diff.

    Blocking findings: 2

    Finding 1 — one-off

    • ID: core-F1
    • Title: Late ownership conflicts can leave owned resources behind
    • Location: shifter/engine/provisioner/raes_gcp_apply.py
    • Decision: fix
    • Rationale: Add ownership-scoped partial cleanup that preserves any conflicting VM, with a regression test for a late conflict after earlier resources are created.

    Finding 2 — class (3 instances)

    • ID: core-F2
    • Title: Converge disk deletion policy for adopted custom-image hosts
    • Location: shifter/engine/provisioner/raes_gcp_apply.py
    • Decision: fix
    • Rationale: Apply attached-disk auto-delete convergence to either preconfigured source, and test adoption and teardown.
    • Instances:
      • shifter/engine/provisioner/raes_gcp_apply.py
      • shifter/engine/provisioner/gcp_range_cells.py
      • shifter/engine/provisioner/raes_gcp_destroy.py
  10. Brad-Edwards commented on Sep 24, 2026

    @Brad-Edwards
    OwnerAuthor

    Pre-PR base synchronization

    • Source: refs/remotes/origin/dev at d406c68a2d98f0a5ff4056541412c16971e08893
    • Outcome: already_current
    • Published feature head: 430c9ea95163f4dd0b04a4abf46a53a3a97902b1
    • Synchronized tree: cd2ca5a7411348c28e652a7222011b98aebc1b4a
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingin-progressAn agent is actively working this issue via /implement

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions