Skip to content

feat: acquire scenario-declared feature artifacts and stream them to guests - #2501

Merged
Brad-Edwards merged 21 commits into
devfrom
feat/2463-feature-artifact-acquisition
Oct 5, 2026
Merged

Brad-Edwards merged 21 commits into
devfrom
feat/2463-feature-artifact-acquisition

Conversation

@Brad-Edwards

Copy link
Copy Markdown
Owner

Implements #2463 phases 1–5: Claude Code (and any recipe-backed artifact) is declared by scenarios, acquired once by the platform, and delivered to range guests. Nothing is redistributed in an image or pack.

Merge after #2495 and #2500. AWS delivery needs the provisioner's storage-key grant (#2495) and the storage bucket in its env (#2500); this PR makes smoke-linux-aws declare Claude Code, so without them the aws-dev post-deploy smoke fails in delivery.

What changes

Decisions (ADR-034-R4/R6/R7 amended, R11/R12 added; ADR-032-R3/R9 amended). Recipe-backed acquisition of pinned or open-scope upstream packages, tracked by Engine, integrity-verified, never redistributed; delivery streams with no fixed size cap.

Inventory and acquisition (Engine/shared).

  • AcquiredFeatureArtifact (migration 0088) records each artifact: source, resolved version, platform, content-addressed key, sha256, size, upstream ref/integrity, attempt state.
  • Platform recipes (first: claude-code, the npm @anthropic-ai/claude-code-linux-x64 binary, sha512-verified). Unknown sources fail; no generic URL fetch.
  • Single-flight claims under a row lock. The provisioner launcher runs one isolated acquisition Job per attempt and verifies the stored object independently before marking it ready. A storage error never reads as a missing object.

Isolation (chart, EKS). The Job runs in its own shifter-acquisition namespace (PSS restricted, default-deny plus HTTPS-only egress, quota), under a dedicated identity with PutObject on the delivery prefix only and no database access. A ValidatingAdmissionPolicy pins image, command, argument grammar, environment, filesystem and budget.

Triggers and launch gate (CMS). Pack registration, pack revision, and deploy bootstrap (acquire_feature_artifacts in the migrate Job) claim acquisitions before any range needs them. At launch, an unprojected artifact feature resolves to its ready artifact or fails that range only, honoring RAES explicitness and permitted routes.

Streaming delivery (provisioner and CMS). The provisioner downloads to a staging file bounded by the exact size, hashes in chunks, and streams raw bytes to the guest over SSH stdin. The guest checks free space, exact size and digest before publishing. CMS pack promotion streams through a hashing reader with a byte-identical deterministic tar. The 256 MiB caps are removed.

Scenario. smoke-linux-aws 0.3.0 declares claude-code on the Kali attacker with the version open, so the platform picks its reviewed default.

Observability. Content-delivery failures report which authored step failed, enforced by a source-scanning test.

Verification

  • Live on aws-dev: the acquisition Job fetched and stored Claude Code once (246,107,320 bytes). The post-deploy smoke then reached READY in under 5 minutes, which is gated on in-guest digest verification of the delivered binary, and its SSH probe passed.
  • Tests: targeted platform, provisioner, terraform, bootstrap and installation suites pass; ADR guard and lint pass.
  • Unchanged: the GKE chart render baseline. The acquisition resources are gated off on GKE; GCP support is feat(gcp): feature-artifact acquisition on GKE and Claude Code in smoke-linux #2479.

Remaining

Phase 6, removing the Claude binary from the AWS and GCE Kali/Ubuntu bakes, follows in a separate PR after this merges.

Refs #2463

Record the #2463 design so third-party software that Shifter must not
redistribute (first case: Claude Code, all rights reserved) reaches range
guests without any image or pack carrying its bytes.

- ADR-034-R11 (new): a source-backed feature artifact with no pack
  projection entry may be satisfied by a backend-owned artifact that
  Shifter acquires with a platform-owned recipe (RAES backend-owned-artifact
  or exact-artifact, acquisition pull, timing backend-preparation), with
  RAES exact/constrained/open version semantics, upstream integrity
  verification, content-addressed storage and CMS inventory tracking.
- ADR-034-R12 (new): acquire before ranges need it with no operator step;
  reuse ready artifacts; single flight via one isolated acquisition Job;
  bounded launch wait; per-range, non-sticky hard materialization failure.
- ADR-034-R4/R6/R7: credential-free public acquisition is not an
  entitlement system and is never republished; the R6 pack-input rule and
  the R7 no-parallel-store rule admit the R11 inventory explicitly.
- ADR-032-R3/R9: streaming delivery with constant memory and no fixed
  payload-size cap; the binding carries the object's exact byte count.

Design note: docs/architecture/raes-feature-artifact-acquisition-preflight-2463.md

Refs #2463
…-flight service

First slice of ADR-034-R11/R12 (#2463), not yet wired to launch or a Job:

- AcquiredFeatureArtifact: CMS inventory of backend-owned feature
  artifacts (source, resolved version, platform, content-addressed storage
  key, sha256, size, upstream ref/integrity) with an acquiring/ready/failed
  lifecycle, a fenced attempt lease, and failure backoff. Records where the
  bytes live, never the bytes.
- Platform recipe registry with the claude-code npm recipe (per-platform
  native binary package) and RAES version semantics: exact honored, open
  resolved to a reviewed default, anything else rejected; unknown sources
  fail closed.
- Streamed npm acquisition: pinned registry host, tarball must stay on it,
  sha512 dist.integrity verified, exactly one regular-file member extracted.
- Single-flight service: reuse a ready artifact whose object still matches,
  join an in-flight attempt, respect backoff, otherwise claim exactly one
  attempt; only the attempt holding the row may finish it; a launch gate
  that fails only the waiting range and refuses to wait inside a
  transaction.

Refs #2463
…es in shared

Layering allows cms -> engine.services but not engine -> cms, and the
provisioner launcher that will create and finalize acquisition Jobs is
Engine. Move the backend-owned feature-artifact inventory and single-flight
service into Engine (beside RaesImageMapping), and the Django-free recipe
registry and npm fetch into shared so the isolated acquisition Job can use
them without database access. Same platform database; the unpushed CMS
migration is replaced by engine 0088. ADR-034-R11 and the design note now
say Engine-owned inventory.

Refs #2463
Django-free entry point (python -m shared.feature_artifacts.job) for the
dedicated acquisition Job (ADR-034-R11/R12, #2463). It resolves the
platform recipe, fetches and integrity-verifies the artifact, stores it
under its content-addressed delivery key (never rewriting an existing
object), and prints one result line for the launcher to verify and
finalize. It holds no database access; acquisition failures are reported
with a bounded reason rather than raised.

Refs #2463
…rotocol

- KubernetesTaskProfile.container_command sets the container command
  (replacing the image entrypoint) only when a profile pins it, so admission
  can match it exactly. Existing profiles still emit args only.
- The acquisition Job exits 0 whenever it wrote a result line, including a
  reported failure: Kubernetes withholds a failed Job's output, and the
  launcher must read the failure reason. Only a crash fails the Job.

Refs #2463
Callers only claim an acquisition attempt; the provisioner launcher's
drain loop now reconciles in-flight attempts (ADR-034-R12, #2463):

- shared.cloud.feature_artifact_jobs: a fixed task profile in its own
  shifter-acquisition namespace and artifact-acquirer service account, a
  pinned Django-free command, read-only root with disk-backed scratch,
  resource limits and a 30-minute deadline; identity and Job name derive
  only from the attempt, so creation is create-or-observe safe.
- reconcile_feature_artifact_acquisitions launches one Job per attempt,
  reads its result line, and finalize_attempt verifies the stored object
  independently (content-addressed key for the digest, exact size) before
  marking the row ready; a crashed or unverifiable Job fails the attempt.
  Errors are isolated per attempt and only error types are logged. Without
  a configured FEATURE_ARTIFACT_JOB_IMAGE the reconcile is a no-op.
- The attempt lease now outlasts the Job deadline; models load lazily
  because engine.services is imported before the app registry is ready.

Refs #2463
…sition Jobs

Infrastructure for the feature-artifact acquisition Job (ADR-034-R12,
#2463), enabled only where the environment defines it:

- Chart (capabilities.featureArtifactAcquisition, default off):
  shifter-acquisition namespace (PSS restricted), artifact-acquirer SA with
  IRSA annotation and no token automount, default-deny plus DNS and HTTPS
  egress (provider-API CIDRs with the service CIDR carved out), a quota,
  a launcher-only Role, and restrict-feature-artifact-jobs: only the
  launcher may create Jobs there, which must run the pinned platform image
  and command with a bounded argument grammar, exactly seven literal
  (non-secret) env inputs, the restricted filesystem/process profile and the
  acquisition budget. A semantic CEL test admits the Job built by the real
  manifest builder with the launcher's env and denies tampering. GKE
  renders are unchanged.
- Terraform: opt-in feature_artifact_store_write grants get/put only under
  the content-addressed delivery prefix plus the storage key via S3; the
  dev root defines the artifactAcquirer identity.

Refs #2463
…tity

render_aws_values now treats artifactAcquirer as an optional workload
role (#2463). Where the environment's Terraform defines it, the renderer
projects its IRSA role, enables capabilities.featureArtifactAcquisition,
and emits FEATURE_ARTIFACT_JOB_IMAGE (the attested platform image) for the
launcher; elsewhere the capability stays off and the image is empty, so
the acquisition reconcile is a no-op. The effective-IRSA readiness probe
also verifies the acquirer identity when it is defined.
FEATURE_ARTIFACT_JOB_IMAGE is a renderer-owned runtime key; the published
backend-bundle contract is regenerated for it.

Refs #2463
…at launch

A scenario declares third-party software as an RAES artifact feature by
name and version only (ADR-034-R11, #2463). At materialization:

- prepare_content_delivery resolves each feature source the pack does not
  project through an injected acquire_feature resolver; a pack needs no
  projection document when it carries no such bytes, and an unresolved
  source still fails closed. Pack-projected content is unchanged.
- The CMS resolver honors the feature's RAES artifact requirement
  (constrained fails as unsupported; declared routes must permit pull),
  requires an artifact feature, derives the guest platform from the target
  node, waits at most RAES_FEATURE_ARTIFACT_WAIT_SECONDS (default 45s,
  inside the user's request) for the Engine artifact, and returns the same
  byte-free feature delivery binding a projected artifact produces.
  Unavailability fails only this range's materialization.

Refs #2463
Reuse verification swallowed every head_object failure as "object absent",
so an access-denied or throttled check reclaimed a ready artifact and
re-downloaded it. Only a definite absence (object_exists False) now counts
as missing; other storage errors propagate, and the launch gate turns them
into a per-range FeatureArtifactUnavailableError.

The acquisition Job no longer HEADs before writing: without s3:ListBucket a
missing key answers 403, not 404, so the check failed every first
acquisition. The key is the content digest, so an unconditional PutObject
rewrites identical bytes.

Refs #2463
…them

Registering a pack, admitting a new revision of it, and deploy content
bootstrap now claim acquisition for every unprojected artifact feature the
pack declares (ADR-034-R12). Realizability projects each source-backed
artifact feature with its node's OS family; a source the pack itself
projects is left to pack delivery. Claims never wait and never fail their
caller: the launcher completes them, and a launch that still finds an
artifact unavailable fails only that range.

The AWS EKS migrate Job runs acquire_feature_artifacts after the in-box
bootstrap, which covers packs registered before acquisition was enabled and
platform default-version changes. The migrator identity gains HEAD plus
ListBucket on the delivery prefix (so a missing key is 404) and no storage
KMS grant, so it cannot decrypt object bytes. The acquisition Job's grant
narrows to PutObject.

Refs #2463
The base AWS smoke scenario declares the claude-code artifact feature on its
Kali attacker with the version left open, so the platform resolves its
reviewed default and the post-deploy smoke exercises acquisition and
delivery end to end. The pack carries no Claude Code bytes.

Revision 0.3.0 re-binds the associated-artifact digest; the in-box entry
upgrades the prior 0.2.0 registration through expected_package_digest.

Refs #2463
Delivery held the whole payload in memory, base64-encoded it and rendered it
into the guest script, under a fixed 256 MiB cap, which a 246 MB feature
artifact cannot pass and an ISO never could. Per ADR-032-R9 the provisioner now
downloads to a staging file bounded by the binding's exact byte count (after a
staging free-space check), hashes it in chunks, and streams that file as raw
stdin to the guest with constant memory.

GuestSSHExecutor gains run_command_streaming: the script travels in argv as
for secret-bearing runs, stdin carries an optional header then the file, output
goes to temp files, and a watchdog enforces the deadline. SetupStep.stdin_path
routes a step through it; executors without the port fail the step, and the
range pod transport refuses explicitly (RAES delivery never uses it).

The Linux deliver scripts check destination free space, then receive stdin
into a private staging file and verify exact size and digest before publish.
The Windows scripts read four base64 header lines byte-wise from the raw
stdin stream and copy the binary remainder. The deliver budget scales with
size; the RAES_CONTENT_DELIVERY_MAX_BYTES setting is removed.

Refs #2463
Promotion materialized each source-backed payload into memory under the
SHIFTER_RAES_CONTENT_DELIVERY_MAX_PAYLOAD_BYTES cap (256 MiB), so a pack could
never carry a large file such as an ISO. Local staging is no substitute: the
portal and worker /tmp volumes are size-limited emptyDirs. Per ADR-032-R9
payloads now stream with constant memory and no fixed cap.

payload_chunks replaces materialize_payload: a file yields its bytes, and a
directory yields the same deterministic tar tarfile writes, byte for byte
(existing content addresses are unchanged), block by block. Every member is
read for exactly its recorded size, so a source that changes underneath fails
closed. Each pack input's read is bounded by its inventory record's exact
size instead of the cap.

Promotion knows the digest up front (the verified inventory digest for a
file, a streamed measuring pass for a directory), uploads through a hashing
reader, and deletes the object and fails if the stored bytes are not the
bytes its key names. The setting and DeliveryTarget.max_payload_bytes are
removed.

Refs #2463
… feature-artifact types

Sonar flagged the acquisition Job's HOME/TMPDIR of /tmp (S5443). The Job now
mounts its disk-backed scratch volume at a dedicated /work path, its only
writable location; the admission policy requires that mount, and a tamper
case proves a /tmp mount is denied. The suppression markers on those lines
are gone.

Code-smell cleanup on the feature-artifact code: concrete types in place of
Any (the inventory model, ObjectStorage, ContentRef, FeatureArtifactDemand,
KubernetesTaskRunner), docstrings, and fewer exits per function. The
precise types exposed that the launch resolver passed a nullable byte_count
into the binding; it now fails that range on an incomplete inventory row.
The version pattern uses \d with re.ASCII so non-ASCII digits stay
rejected. ContentRef and StorageTarget become public names because they
form the acquire_feature contract.

Refs #2463
…reaming

Sonar code-smell findings on the streaming delivery work, fixed in place:

- Provisioner: payload staging (download, chunked hashing, installed-tree
  digest) moves to raes_content_payload.py, taking raes_content_delivery.py
  under 500 lines. Its payload record becomes the public RaesContentPayload
  the delivery plan takes, replacing four constructor parameters. The
  streaming ssh seam is a staticmethod, and the result helper has a docstring.
- CMS/shared: the streamed materializer and promotion (hashing reader,
  measure, promote) move to shared/raes/content_payload.py, and pack-input
  verification against the inventory moves to
  shared/raes/pack_delivery_inputs.py, bringing content_delivery.py and
  content_delivery_prep.py under 500 lines. prepare_content_delivery delegates
  projection loading and the pack-or-acquire choice, cutting its complexity,
  and the tar emitter yields per member. readinto keeps RawIOBase's parameter
  kinds.

Behavior is unchanged; the streamed tar stays byte-identical to tarfile's.

Refs #2463
…ering

render_aws_values grew past Sonar's complexity and length limits when
feature-artifact acquisition was wired in. Workload IRSA role validation and
the acquisition Job image decision become helpers; the rendered values are
unchanged.

Refs #2463
- The acquisition-admission tamper case mounted scratch at a literal /tmp,
  which ruff S108 rejects in CI; any path other than /work proves the same
  denial.
- prepare_content_delivery delegates the bucket check and the
  pack-projection decision (complexity 11 -> under 10), keeping the same
  evaluation order.
- render_aws_values delegates the ALB edge values (103 -> under 100 lines).

Refs #2463
A failed delivery reached the result channel only as "(RaesContentDeliveryError)",
leaving roughly ten distinct delivery checks indistinguishable from CI. Every
RaesContentDeliveryError message is an authored constant naming the failing
step (no payload, key, path or guest output), so _classify_failure now reports
it, as it already does for RaesRealizationError, under the same closed
cloud_operation_failed code (ADR-043-R5 bars sensitive data, not authored
text).

The one raise that forwarded another exception's text now uses its own
literal, and a source-scanning test requires every RaesContentDeliveryError
call to pass a literal or named constant, so runtime data cannot leak in.

Refs #2463
Adding content-delivery errors gave _classify_failure four returns (Sonar
S1142). The authored-message cases become one table of (type, reason code,
template); behavior is unchanged.

Refs #2463
@Brad-Edwards
Brad-Edwards merged commit 835f8eb into dev Oct 5, 2026
7 checks passed
@Brad-Edwards
Brad-Edwards deleted the feat/2463-feature-artifact-acquisition branch October 5, 2026 17:09
@sonarqubecloud

sonarqubecloud Bot commented Oct 5, 2026

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant