Skip to content

feat(surface,sdk): f.memory via ai-hist — recall/why/learn (partial #307) #325

feat(surface,sdk): f.memory via ai-hist — recall/why/learn (partial #307)

feat(surface,sdk): f.memory via ai-hist — recall/why/learn (partial #307) #325

Workflow file for this run

name: Review swarm
on:
pull_request:
types: [opened, synchronize, reopened, ready_for_review]
permissions:
contents: read
pull-requests: write
concurrency:
group: review-swarm-${{ github.event.pull_request.number }}
cancel-in-progress: true
jobs:
review:
runs-on: ubuntu-latest
# Ordering invariant: swarm 60m < poll 65m < job 75m.
timeout-minutes: 75
# `agent-relay cloud run` authenticates using CLOUD_API_KEY.
# @agent-relay/cloud@11.10.3 adds WorkflowApiKeyClient.fromEnv, which
# prefers CLOUD_API_KEY over stored login, avoiding the interactive device
# flow. The credential is minted per AgentWorkforce/cloud →
# docs/runbooks/relay-ci-workflow-credential.md, profile workflow-invoke,
# scoped to workflow:invoke:read and workflow:invoke:write.
env:
CLOUD_API_URL: ${{ vars.CLOUD_API_URL || 'https://agentrelay.com/cloud' }}
CLOUD_API_KEY: ${{ secrets.CLOUD_API_KEY }}
RELAY_WORKSPACE_KEY: ${{ secrets.RELAY_WORKSPACE_KEY }}
RELAY_API_KEY: ${{ secrets.RELAY_WORKSPACE_KEY }}
steps:
- name: Check out PR head
uses: actions/checkout@v4
with:
ref: ${{ github.event.pull_request.head.sha }}
path: pr-head
fetch-depth: 0
- name: Check out immutable gate from main
uses: actions/checkout@v4
with:
# The base SHA is the immutable definition this PR is judged by. A
# moving `main` ref could change the judge while this run is live.
ref: ${{ github.event.pull_request.base.sha }}
path: gate-files
sparse-checkout: |
workflows/review-swarm.yaml
.github/workflows/scripts/swarm-post.sh
.github/workflows/scripts/swarm-prepare.sh
.github/workflows/scripts/swarm-verdict.sh
.github/workflows/scripts/swarm-gate.test.sh
.github/workflows/scripts/swarm-definition.sh
.github/workflows/scripts/swarm-definition.test.sh
# The gate decides whether code merges, so before it judges anything it
# proves it can still say no. `swarm-verdict.sh` is a handful of lines of
# shell; the failure that matters is not this gate going red -- a red gate
# announces itself -- but this gate quietly losing the ability to go red,
# which announces nothing and surfaces only after something broken has
# merged behind a green check. That is not hypothetical here: on
# 2026-09-09 a release shipped past a smoke test that could not fail.
#
# The suite asserts both directions. A genuine REVIEW_FAILED must fail the
# gate, and three clean passes must pass it -- without that second half
# every assertion would be satisfiable by an unconditional `exit 1`, and
# an always-red gate is as useless as an always-green one. It is hermetic
# (agent-relay and gh are stubbed) and runs offline in a few seconds, so
# it costs nothing to run ahead of a 20-minute cloud swarm.
#
# It runs the copy from main, alongside the scripts it tests, so a pull
# request cannot weaken the gate by editing the test that guards it. The
# guard is skipped only while the file has not yet reached main.
- name: Self-test the gate's verdict logic
env:
REVIEW_PR_NUMBER: ${{ github.event.pull_request.number }}
run: |
ruby --version
ruby -ryaml -e 'abort "Psych YAML parser unavailable" unless defined?(Psych)'
test_script=gate-files/.github/workflows/scripts/swarm-gate.test.sh
if [ ! -f "$test_script" ]; then
# Only the introducing PR may bootstrap before its test is on main.
# Later deletion, bad paths or incomplete checkouts must fail closed.
if [ "$REVIEW_PR_NUMBER" = 248 ]; then
echo "::notice::PR #248 bootstrap: self-test is not on main yet."
exit 0
fi
echo "::error::main-owned swarm-gate.test.sh is missing" >&2
exit 1
fi
bash "$test_script"
definition_test=gate-files/.github/workflows/scripts/swarm-definition.test.sh
if [ ! -f "$definition_test" ]; then
# The introducing PR cannot run a helper that is not on its base
# yet. Once this PR lands, absence is a deletion or checkout bug
# and must fail closed like the verdict self-test above.
if [ "$REVIEW_PR_NUMBER" = 265 ]; then
echo "::notice::PR #265 bootstrap: candidate validator is not on main yet."
else
echo "::error::main-owned candidate validator is missing" >&2
exit 1
fi
else
bash "$definition_test"
fi
# A PR that changes the relayflow definition cannot be judged by that
# definition without violating the immutable-gate rule. Parse and check
# the candidate's load-bearing policy against the trusted base, but keep
# the actual review run below on the base definition. This is a candidate
# contract check, not a candidate verdict.
- name: Validate candidate review definition
env:
REVIEW_PR_NUMBER: ${{ github.event.pull_request.number }}
REVIEW_BASE_SHA: ${{ github.event.pull_request.base.sha }}
REVIEW_HEAD_SHA: ${{ github.event.pull_request.head.sha }}
working-directory: pr-head
run: |
# Compare the PR patch (merge-base...head), not the two trees. A
# branch can be stale on this file without the PR changing it.
if git diff --quiet "$REVIEW_BASE_SHA...$REVIEW_HEAD_SHA" -- workflows/review-swarm.yaml; then
echo "Candidate review definition unchanged; trusted base definition is the effective file."
exit 0
fi
validator=../gate-files/.github/workflows/scripts/swarm-definition.sh
if [ ! -f "$validator" ]; then
if [ "$REVIEW_PR_NUMBER" = 265 ]; then
echo "::notice::PR #265 bootstrap: candidate validator is not on main yet."
else
echo "::error::main-owned candidate validator is missing" >&2
exit 1
fi
else
"$validator" workflows/review-swarm.yaml \
../gate-files/workflows/review-swarm.yaml
fi
# Fail here, in seconds, rather than in `Launch cloud swarm` ten minutes
# later. WorkflowApiKeyClient.fromEnv requires CLOUD_API_URL and
# CLOUD_API_KEY; if either is missing the CLI falls back to the device
# flow. Check both exactly as ops/NEXT.md specifies.
- name: Validate cloud authentication
run: |
test -n "$CLOUD_API_URL"
test -n "$CLOUD_API_KEY"
test -n "$RELAY_WORKSPACE_KEY"
# Presence is not validity. This step was named "Validate cloud
# authentication" while only asserting the variables were non-empty, so on
# 2026-09-07 it passed on every run while `agent-relay cloud run` failed
# immediately after with `Workflow prepare failed: 401 Unauthorized` --
# six PRs, repeatedly, behind a green check.
#
# Actually exercise the credential against the same host the CLI will use.
# /api/v1/workflows/runs requires a RESOLVED WORKSPACE and returns 401 for
# a fabricated or absent token (verified against production), so 200 here
# means the credential can genuinely act, not merely that a string was set.
#
# Also print a NON-REVERSIBLE fingerprint of the key. When this check
# passes and the launch still 401s, the fingerprint answers whether CI is
# even using the credential the mint installed -- otherwise unanswerable
# from outside, because the value is masked everywhere it appears.
fp="$(printf '%s' "$CLOUD_API_KEY" | shasum -a 256 | cut -c1-12)"
echo "CLOUD_API_KEY fingerprint (sha256, first 12): $fp"
# Bound the request and keep the three outcomes apart. Unbounded, an
# unreachable Cloud leaves curl waiting until the 75-minute job
# timeout; the `|| echo 000` then produced a non-200 and the one
# error message told a maintainer to re-mint a credential that was
# never the problem. A transport failure, an auth rejection and an
# unhealthy Cloud are three different diagnoses and must not share a
# sentence.
if ! status="$(curl --connect-timeout 10 --max-time 30 \
-s -o /dev/null -w '%{http_code}' \
-H "Authorization: Bearer $CLOUD_API_KEY" \
"${CLOUD_API_URL%/}/api/v1/workflows/runs")"; then
echo "::error::could not reach $CLOUD_API_URL to validate CLOUD_API_KEY (curl transport failure or timeout). This is not a credential verdict — re-run once Cloud is reachable." >&2
exit 1
fi
case "$status" in
200) ;;
401|403)
echo "::error::CLOUD_API_KEY is set but rejected by $CLOUD_API_URL (HTTP $status). Re-mint the credential; do not re-run this job." >&2
exit 1
;;
*)
echo "::error::$CLOUD_API_URL returned HTTP $status while validating CLOUD_API_KEY. That is not an authentication verdict — treat it as Cloud being unhealthy rather than the credential being bad." >&2
exit 1
;;
esac
echo "CLOUD_API_KEY authenticates against $CLOUD_API_URL; interactive login is unreachable from here."
# `agent-relay cloud run` launches the swarm, but nothing installed the
# CLI, so this job failed at `Launch cloud swarm` with
# `agent-relay: command not found` (exit 127) on every pull request.
# Pinned, not floating: the gate decides whether code merges, so it must
# not change behaviour because a new CLI was published overnight.
#
# 11.10.3, not 11.8.3. Both resolve an env-backed session, but they read
# different credentials: 11.8.3 accepts CLOUD_API_ACCESS_TOKEN plus its
# refresh token and expiry (a session, which ages out), while
# @agent-relay/cloud@11.10.3 adds CLOUD_API_KEY via
# `WorkflowApiKeyClient.fromEnv`, preferred over the stored login in
# workflows.js. The credential this gate is meant to carry is the one
# minted by cloud's docs/runbooks/relay-ci-workflow-credential.md, which
# issues an API key scoped to workflow:invoke:{read,write} -- so the
# runtime has to be a version that reads an API key. Verified: that
# symbol is absent from 11.8.3.
- name: Install the Agent Relay CLI
run: |
npm install -g agent-relay@11.10.3
agent-relay --version
- name: Prepare review input on GitHub runner
env:
GH_TOKEN: ${{ github.token }}
working-directory: pr-head
run: |
../gate-files/.github/workflows/scripts/swarm-prepare.sh \
"${{ github.event.pull_request.number }}"
mkdir -p .github/workflows/scripts
cp ../gate-files/.github/workflows/scripts/swarm-verdict.sh \
.github/workflows/scripts/swarm-verdict.sh
git add -f .github/workflows/scripts/swarm-verdict.sh
- name: Launch cloud swarm
id: launch
# Backstop only. The preflight above makes the device flow unreachable,
# but nothing in the CLI bounds an interactive login, so cap the step
# rather than let a future regression spend the job's 75 minutes on it.
timeout-minutes: 15
working-directory: pr-head
run: |
response=$(agent-relay cloud run \
../gate-files/workflows/review-swarm.yaml --sync-code --json)
run_id=$(jq -er '.runId // .id' <<<"$response")
echo "run_id=$run_id" >> "$GITHUB_OUTPUT"
- name: Wait for cloud swarm
id: wait
if: always() && steps.launch.outputs.run_id != ''
# `../gate-files/` below resolves relative to this step's cwd. Every
# other step that reaches into `gate-files/` runs from `pr-head/`;
# without a matching `working-directory` here, the diagnostic call
# resolves outside the workspace and `set +e` silently swallows the
# miss (Cursor Bugbot flagged as HIGH on #285).
working-directory: pr-head
run: |
set +e
# Ordering invariant: swarm 60m < this poll deadline 65m < job 75m.
deadline=$((SECONDS + 3900))
status=timed_out
while [ "$SECONDS" -lt "$deadline" ]; do
response=$(agent-relay cloud status "${{ steps.launch.outputs.run_id }}" --json) || {
status=status_error
break
}
parsed_status=$(jq -er '.status' <<<"$response") || {
status=status_error
break
}
status=$parsed_status
case "$status" in
completed|failed|cancelled) break ;;
esac
sleep 15
done
# `status=timed_out` above is set BEFORE the first poll and is
# overwritten by every successful one, so it can only survive if the
# very first `agent-relay cloud status` never returns a status. Any
# real timeout therefore reported the last thing it saw -- almost
# always `running` -- and "did not complete successfully: running"
# reads like a transient blip rather than "we waited the full 65
# minutes". Restore the distinction, and keep the last observed
# status, which is the more useful of the two facts.
if [ "$SECONDS" -ge "$deadline" ]; then
case "$status" in
completed|failed|cancelled|status_error) ;;
*)
echo "poll deadline reached after ${SECONDS}s;" \
"last observed swarm status: ${status}" >&2
status=timed_out
;;
esac
fi
echo "swarm_status=$status" >> "$GITHUB_OUTPUT"
if [ "$status" != completed ]; then
../gate-files/.github/workflows/scripts/swarm-status-diagnostic.sh "${response:-}"
fi
exit 0
- name: Post verdict and transcripts
if: always() && steps.launch.outputs.run_id != ''
env:
GH_TOKEN: ${{ github.token }}
working-directory: pr-head
run: ../gate-files/.github/workflows/scripts/swarm-post.sh "${{ steps.launch.outputs.run_id }}" "${{ github.event.pull_request.number }}"
- name: Enforce swarm result
if: always() && steps.wait.outputs.swarm_status != 'completed'
run: |
echo "Review swarm did not complete successfully: ${{ steps.wait.outputs.swarm_status }}" >&2
exit 1