Skip to content

sre-lab: add azd-first Azure SRE Agent event lab - #47

Merged
hellices merged 26 commits into
mainfrom
feature/sre-agent-azd-lab
Aug 15, 2026
Merged

hellices merged 26 commits into
mainfrom
feature/sre-agent-azd-lab

Conversation

@hellices

Copy link
Copy Markdown
Owner

Summary

  • reorganize the Azure SRE Agent event lab into an azd up-first guided workflow
  • deploy Container Apps, ACR, Storage, Application Insights, Log Analytics, VNet/private endpoint, and one-minute workspace alerts
  • add guarded HTTP 500, latency, and Storage RBAC scenarios with evidence capture, recovery, scoring, and safe cleanup
  • document the manual Agent/repository/knowledge/response-plan setup and the verified Azure deployment boundary

Safety

  • destructive scenarios require explicit acknowledgement and environment binding
  • any running or failed scenario blocks all new runs until repaired
  • recovery failures propagate and unrecovered runs always score 0
  • external role cleanup validates subscription, principal, role, and scope

Verification

  • python3 -m pytest scripts infra -q — 559 passed
  • app/.venv/bin/python -m pytest app -q — 10 passed
  • /bin/bash -n scripts/*.sh
  • python3 -m py_compile scripts/*.py
  • all infra/*.bicep templates built successfully
  • live azd provision, azd deploy, azd up, health/telemetry/alert checks, baseline, acknowledgement gate, and azd down --purge were previously completed

Reviewer notes

The live S1/S2/S3 Agent investigation sequence is intentionally left for the guided user run; temporary validation resources were deleted after infrastructure verification.

hellices and others added 25 commits August 14, 2026 13:13
Add azure.yaml and main.parameters.json so azd owns provisioning of
the SRE Agent event lab. Convert infra/main.bicep to a
subscription-scoped azd entry point (derives suffix, resource group,
and merged azd tags) and move the former resource-group module to
infra/lab.bicep, which now always deploys the Container App and
alert rules with a placeholder image. Remove the now-redundant
infra/subscription.bicep(param) in favor of the azd entry point.
Add scripts/azd-configure.sh (preprovision: provider registration and
env defaults) and scripts/azd-postprovision.sh (az acr build + update
the existing Container App to the immutable image, then health-check
it), keeping the lab Docker-free.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Address the Task 1 review findings on the azd migration:

- Add scripts/cleanup-external.sh so every azure.yaml hook points at an
  existing executable script. It removes only the recorded
  subscription-scoped role assignments and exits 0 when the Agent setup
  evidence is absent, so azd down never fails and never deletes broadly.
- Parameterize the Container App port and probes. The first provision
  runs the public placeholder on port 80 without /healthz probes, then
  postprovision moves ingress to 8000, swaps in the ACR-built image,
  waits for a healthy revision, and verifies /healthz. Probes now always
  match the port they check.
- Persist the built image in SRE_CONTAINER_IMAGE and bind it (plus
  ACTION_GROUP_RESOURCE_ID) in main.parameters.json, so a later provision
  keeps the lab image instead of reverting to the placeholder.
- Turn deploy.sh into a compatibility wrapper around azd up, refresh the
  README deployment/teardown commands, and update the tests, so no
  tracked file or documented command references the deleted
  subscription-scope templates.
- Pin every Azure CLI call in the azd hooks to AZURE_SUBSCRIPTION_ID and
  report a mismatched active account. Resource-group and app operations
  read the current azd values only.
- Delete infra/lab.bicepparam, which pinned a dead suffix and expiry.
- Restore the deployment outputs the lab scripts still read and ignore
  the project-local .azure/ directory.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… guard no-login, mark legacy scripts

Fixes the four remaining Task 1 review findings from
review-f960e97..a1cceb6.diff, under strict TDD (RED tests written and
confirmed failing first, then GREEN implementation):

1. README's `azd env new` command, prerequisites line, and the "실측
   환경" table no longer hardcode the original validation subscription
   ID. The command now omits --subscription (azd prompts interactively)
   with guidance to pass --subscription <YOUR_SUBSCRIPTION_ID> instead.
   New regression test
   test_azd_onboarding_docs_and_config_do_not_hardcode_a_subscription_id
   scans README/azure.yaml/Bicep/parameters/hook scripts for the literal
   ID (common.sh and validation-results.md are intentionally excluded,
   as historical/legacy records already covered by their own tests).

2. cleanup-external.sh (the predown hook) now requires azd and, when
   run with --yes (as azure.yaml's predown hook always does), clears
   SRE_CONTAINER_IMAGE and SRE_IMAGE_TAG via `azd env set KEY ""`
   before/independent of the evidence-gated role-assignment cleanup, so
   reusing the same azd environment falls back to the placeholder image
   on the next provision. A dry run (no --yes) only prints the planned
   action. New tests exercise this via a stubbed azd binary and confirm
   no az group/resource delete calls occur as a result.

3. README now carries an explicit transitional caveat at the top of
   "## 시나리오 실행" stating run-scenario.sh/query-evidence.sh still
   read common.sh's legacy pre-azd deployment lookup and are legacy-only
   until common.sh is rewritten to use azd env get-value (a separate,
   not-yet-done task); points readers to the Baseline section's
   azd env get-value-based checks in the meantime.

4. azd-configure.sh and azd-postprovision.sh now guard their
   `az account show --query id -o tsv` call and exit with a clear
   "Azure CLI is not signed in. Run 'az login'..." message instead of
   leaking raw Azure CLI stderr when the CLI is signed out.

Verified: azd env new --help / azd env set --help / a scratch azd env
new + azd env set KEY "" run locally confirm the empty-value clearing
semantics and that there is no `azd env unset`.

Tests: 124 passed (infra/tests + scripts/tests, up from 118 passed/6
failed RED baseline), app suite 10 passed unchanged, bash -n clean,
az bicep build clean for main/lab/workload.bicep, azure.yaml validated
against the official azd JSON schema.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replace common.sh's fixed subscription/resource-group/deployment-name
constants with load_lab_config, which resolves every setting as explicit
process environment > current `azd env get-value` > an allowed default via
a new setting()/require_setting() helper (no eval, no dynamic indirection).

- deployment_output() now reads the AZURE_-prefixed (and legacy-named)
  azd deployment outputs load_lab_config already resolved, instead of
  `az deployment sub show` against a hardcoded deployment name.
- verify_lab_resource_group() now requires both the purpose and
  azd-env-name tags to match the current azd environment, replacing the
  old single hardcoded resource-group-name safety boundary.
- run-scenario.sh, query-evidence.sh, capture-scenario.sh, and cleanup.sh
  all call require_lab_config (require_commands + load_lab_config) before
  reading SUBSCRIPTION_ID/RESOURCE_GROUP or any deployment_output() value.
- Added monitor/sre-agent-event-lab/.env.example documenting every
  setting name and its allowed default, with no secrets.
- Rewrote the common.sh tests that intentionally pinned the old hardcoded
  subscription/resource-group values to assert the new azd-backed
  resolution instead, and added test_azd_env.py (with a shared fake
  azd/az harness in azd_common_harness.py) covering precedence, missing
  required settings, and the tag-based resource-group safety check.
- Updated README's safety-boundary and scenario-execution sections: the
  scenario/query scripts are no longer legacy/transitional, they read
  the currently selected azd environment.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
`azd env get-value` writes its `ERROR: ...` diagnostics to stdout and
signals failure only through the exit status (verified against azd
1.29.0), so keeping stdout unconditionally turned "key not found" and
"no project exists" into configuration values: SUBSCRIPTION_ID became an
error sentence, documented defaults never applied, and the preprovision
hook skipped deriving AZURE_RESOURCE_GROUP. Every lookup now uses the
value only when azd exited 0 and pins `--cwd` to the lab's azd project so
the scripts work from the repository root or any other directory.

`load_lab_config` also declared `WORKSPACE_CUSTOMER_ID` /
`CONTAINER_APP_PRINCIPAL_ID` readonly, which aborted query-evidence.sh
under `set -e` on its own assignment; the resolved values moved to
LAB_-prefixed names.

Tests now run the four entry points as programs against fake
az/azd/python executables from a working directory outside the lab, the
fake azd reproduces azd 1.29's contract (both missing-value shapes), and
the text-order/eval-substring assertions were replaced by behaviour,
fail-closed and token-aware checks.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Add scripts/lab.sh as a single-command entry point dispatching to
  doctor|baseline|acknowledge agent-setup|run s1|s2|s3|capture s1|s2|s3|score.
  acknowledge/score depend on a later task's lab_state.py/score.py; until
  those land, lab.sh fails closed with a clear "not yet available"
  message (exit 3) instead of a raw file-not-found error.
- Add scripts/doctor.sh: a 15-row environment diagnostic printing the
  CHECK<TAB>STATUS<TAB>DETAIL contract (PASS/FAIL/MANUAL). Verifies
  required commands, Azure CLI login, azd configuration, subscription and
  resource-group pinning, Container App health, /healthz, Application
  Insights telemetry, alert rule enablement, the SRE Agent resource
  (when configured), and Reader role assignment -- all through official
  stable Azure CLI/REST reads. Fails closed: once any prerequisite check
  fails, every subsequent Azure-touching check is blocked rather than
  attempting further calls. Repository connection, knowledge source,
  incident platform, and response plan are always MANUAL since no stable
  API exposes them.
- Add scripts/baseline.sh: runs baseline load (30 orders + 10 documents
  requests) and verifies Application Insights shows telemetry for both
  endpoints via bounded polling (default 600s timeout / 20s interval,
  overridable via LAB_BASELINE_TELEMETRY_TIMEOUT_SECONDS/
  LAB_BASELINE_TELEMETRY_POLL_INTERVAL_SECONDS). Writes evidence under
  evidence/baseline-<UTC>/.
- Fix scripts/capture-scenario.sh executable bit (100644 -> 100755); it
  was the only lab script not marked executable, and lab.sh capture
  depends on invoking it directly.
- Add scripts/tests/doctor_harness.py: a fake-CLI harness (real az/azd/
  curl/python3 stubs on PATH) purpose-built for doctor.sh/baseline.sh/
  lab.sh's call surface (container app health, /healthz, per-rule alert
  reads, per-principal role assignment lookups), mirroring the existing
  lab_script_harness.py pattern.
- Add scripts/tests/test_doctor.py and scripts/tests/test_lab_cli.py:
  behavioural tests driving doctor.sh/lab.sh as real programs against the
  fake CLIs (not text-only assertions), covering the fully-healthy path,
  every individual FAIL condition, fail-closed gating once subscription/
  azd configuration checks fail, azd --cwd pinning, and lab.sh's
  dispatch/auto-discovery/not-yet-available behavior.
- Document the new lab.sh entry point and doctor check list in README.md.

173 tests pass (144 pre-existing + 29 new); bash -n and az bicep build
verified.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The doctor and baseline telemetry checks were written against inferred CLI
behaviour and could never report the truth against real Azure:

- `az monitor log-analytics query -o json` prints a flat JSON array of row
  objects (the log-analytics extension flattens the REST envelope), so the
  `.tables[0].rows` parse always yielded zero rows and telemetry could only
  ever FAIL. Parse the real shape through a shared `log_analytics_row_count`
  helper, and fake that shape in both harnesses.
- The doctor query used `| count`, whose single row exists even for an empty
  table, so row-count semantics could not distinguish data from no data.
  Query projected rows bounded by `take 1` instead.
- `az role assignment list` hides parent-scope grants without
  `--include-inherited`, reporting a subscription-scoped Reader as missing.
  Ask for inherited assignments and distinguish direct from inherited in the
  detail.
- A malformed `agent-setup.json` aborted the whole run through `jq` under
  `set -e`; it is now one FAIL row and exit 1 with the rest of the report
  intact.

Also add the two missing prerequisite rows -- the `log-analytics` CLI
extension (`az extension show`, with the exact `az extension add` remedy) and
azd's own login state (`azd auth login --check-status --output json`, whose
exit code is always 0, so only the printed status is trusted) -- give
baseline a poll that always attempts at least once and never sleeps past its
deadline, and correct the doctor comment that claimed it never reads the raw
process environment.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Scenario runs, captures and scoring now share one state file bound to the
azd environment they belong to, so a run can only start when the previous
scenario really recovered and was captured, and evidence is scored from
what the Agent actually produced.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ires

Address Task 4 review findings:

- LabState.mark_recovered/mark_failed now discard the scenario's
  previous capture_status (and its capture-side evidence_dir) before
  recording the new run's outcome. A conclusion captured against an
  earlier, superseded run could otherwise keep satisfying sX_captured
  -- and therefore the next scenario's gate and the scorer -- even
  though nothing has been captured for the current run yet. Both a
  recovered and a failed re-run now clear it; tested with the state
  reloaded from disk to prove it is actually persisted.

- run-scenario.sh now calls `lab_state mark-failed` with the evidence
  directory and a reason before exiting when the target alert never
  fires within the poll window, matching every other failure path.
  The wait is now bounded by LAB_ALERT_FIRE_TIMEOUT_SECONDS /
  LAB_ALERT_FIRE_POLL_INTERVAL_SECONDS (default 720s/20s, same
  pattern as the existing alert-resolve and health-check overrides)
  so this is exercised in the test suite instead of only in
  production. The pre-existing recovery trap is untouched and still
  reverts the injected fault on this path.

- LabState._load validates that `stages`, `scenarios`, and every entry
  inside them decode to JSON objects. A wrong type (list, string,
  null, number) now raises the same clean LabStateError a corrupt
  file already produces, instead of a raw TypeError/AttributeError
  traceback surfacing later from mark()/mark_recovered()/
  record_capture(); the file is never silently reset to defaults.

- Documented, in lab_state.py's module docstring, that state.json has
  no cross-process locking and concurrent operators racing the same
  environment can lose an update to a later writer -- a deliberate
  choice for a single-operator lab. No locking was added; nothing in
  the test suite showed a real concurrent-use need.

Tests (all written RED first): test_lab_state.py gains
test_rerunning_a_recovered_scenario_clears_the_stale_capture_status,
test_rerunning_a_failed_scenario_clears_the_stale_capture_status,
test_reloading_after_a_rerun_still_shows_the_cleared_capture_status,
a parametrized test_a_state_file_with_the_wrong_json_shape_is_refused_not_reset
(10 shapes), test_cli_reports_a_malformed_state_file_without_a_python_traceback,
and test_module_documents_that_concurrent_operators_are_unsupported.
test_lab_scripts.py gains test_run_scenario_marks_failed_when_the_alert_never_fires,
backed by a new alert_fires=False mode in lab_script_harness.py's fake az.

Verification:
  python3 -m pytest monitor/sre-agent-event-lab/scripts/tests -q
  307 passed (was 291 before this change)
  bash -n monitor/sre-agent-event-lab/scripts/*.sh
  python3 -c "import lab_state, score"   # Python 3.9.6, clean

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
`azd down` deletes the azd-owned resource group; the only lab artifacts it
cannot see are the subscription-scoped Monitoring Contributor assignments
the Azure SRE Agent setup recorded. cleanup-external.sh now removes exactly
those, after verifying every record, and cleanup.sh stops deleting a
resource group of its own.

- cleanup-external.sh verifies four things before deleting anything: the
  recorded assignment ID names a role assignment in the subscription this
  run resolved (AZURE_SUBSCRIPTION_ID > the current azd environment), the
  Azure CLI is signed in to exactly that subscription, the record carries
  the Agent principal it was created for, and the live assignment really
  holds that principal, the Monitoring Contributor role definition and
  subscription scope. Every record is verified before the first deletion,
  so one untrusted record leaves the subscription untouched. An empty,
  missing or already-absent record is a safe no-op (`RoleAssignmentNotFound`
  is what the CLI really answers -- recorded from azure-cli against a live
  subscription); malformed JSON, a missing principal, an unreadable
  assignment, a mismatch or a failed deletion all fail closed, before azd
  destroys anything.

- The hook-set SRE_CONTAINER_IMAGE/SRE_IMAGE_TAG reset moved out of
  `predown` into the `postdown` hook (`--reset-image-env`). `predown` runs
  before azd asks the operator to confirm the deletion, so clearing them
  there broke the environment of an operator who answered "no"; azd runs a
  post hook only after the action succeeded (HooksRunner.Invoke returns
  early on failure, cli/azd/pkg/ext/hooks_runner.go), which is exactly when
  the recorded image is really gone. `postdown` is an officially supported
  azd command hook.

- Both azd env writes are pinned with `--cwd "${LAB_ROOT}"`, so running the
  hook by hand from the repository root no longer resolves whatever azd
  project the working directory happens to hold.

- cleanup.sh is now a compatibility wrapper: it names `azd down --purge`
  and forwards to cleanup-external.sh. It never deletes a resource group
  unless `--legacy-delete-resource-group` is passed -- the documented
  recovery path for a lab whose azd environment was lost -- which still
  runs the subscription and `purpose`/`azd-env-name` tag checks first.

Tests: scripts/tests/test_cleanup_external.py drives the hook as a program
from a directory that is not the lab, on macOS's Bash 3.2, against a fake
`az` holding a staged subscription and the azd 1.29 fake; the cleanup tests
that lived in test_azd_hooks.py moved there. azd_fake now logs each argument
with %q so a value azd is asked to clear is visible instead of vanishing.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- cleanup-external.sh: capture az rest's stdout and stderr into separate
  files instead of merging them with 2>&1. A successful read that also
  prints an azure-cli warning to stderr no longer corrupts the ARM JSON,
  and a failing read's RoleAssignmentNotFound check now inspects stderr
  alone instead of a stream that could contain stray stdout content.
  Added --only-show-errors to ask azure-cli itself to drop most warnings
  first.
- Fixed the 'already verified' dedupe check: it matched any recorded ID
  that was merely a suffix of an already-verified one (case pattern
  *"${id}"$'\n'* has no left boundary), silently skipping validation
  of a second, distinct record. Now bounded by a leading newline as well,
  requiring a full \n<id>\n match, still a plain string (Bash 3.2 safe).
- README: documented that predown runs before azd's own delete
  confirmation prompt, so canceling that prompt does not restore already
  -removed role assignments; an operator must re-create the assignment
  and re-run 'lab.sh acknowledge agent-setup'. No longer implies canceling
  keeps a fully working environment.
- test_cleanup_runs_under_bash_32 now pins /bin/bash explicitly via
  run_cleanup(..., bash=system_bash) instead of relying on whatever
  'bash' resolves to on PATH.
- Added behaviour tests for all of the above plus fake az/az-fake harness
  support for staged stderr warnings and stdout noise (cleanup_harness.py
  Staged, lab_script_harness.py dispatch fix for --only-show-errors).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fixes six findings from the Task 6 review (review-f80be21..f14edd7.diff),
all with RED/GREEN tests:

1. The documented azd-first flow never prepared app/.venv: capture-scenario.sh
   hard-requires app/.venv/bin/python but nothing created it. Added a new,
   idempotent scripts/setup-venv.sh (uv venv --python ">=3.10" --allow-existing
   plus uv pip install -r requirements-dev.txt, verifying Pillow importability),
   invoked from the postprovision hook before any Azure CLI call. uv is
   mandatory, not merely preferred (this lab's proxy is configured for uv,
   not public PyPI), so there is no silent pip fallback -- missing uv is an
   actionable failure. Because postprovision runs after azd provision has
   already created cloud resources, every failure message states the exact
   rerun command (azd hooks run postprovision, or ./scripts/setup-venv.sh
   directly). doctor.sh gained a "Python environment" check (venv + Pillow
   readiness) and capture-scenario.sh gained its own Pillow-importability
   precondition, both pointing at the same rerun command.
2. Appended a correction section to the (gitignored) task-6-report.md
   documenting these inaccuracies, per its own "historical artifact" status,
   rather than rewriting the report in place.
3. Restored the verified https://azuresre.dev audience fact (deleted, not
   relocated, by the Task 6 README rewrite) into validation-results.md's
   historical bridge section, explicitly framed as legacy record for the
   non-default Logic App bridge.
4. Fixed README and guides/05-results.md cwd instructions: each document
   now uses exactly one cd, so every command block in it is runnable
   sequentially in one shell from that single working directory.
5. Reworded the README's scorecard row from "10 point max" (ambiguous as
   the overall max) to "10 per scenario, 30 overall", matching score.py's
   real MAX_POINTS (10 per scenario, 30 overall).
6. Fixed two accuracy issues in the numbered guides: guide 02's start
   conditions claimed an unenforced concurrency lock that lab_state.py's
   own docstring says does not exist; guide 05 ran the stdlib-only
   generate_notifications.py through app/.venv/bin/python, implying a
   Pillow dependency it does not have.

Tests: 414 passed (scripts/tests + infra/tests + app/tests). bash -n and
az bicep build both clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… wording

Fixes three remaining Task 6 review findings, all with RED/GREEN tests.

1. test_setup_venv.py ran the *real*, in-place setup-venv.sh: the script
   resolves SCRIPT_DIR/LAB_ROOT/VENV_DIR from its own ${BASH_SOURCE[0]},
   never from the caller's cwd, so pointing the process cwd at a tmp_path
   (as the old tests did) still targeted the real app/.venv -- and the
   tests' fake `uv` genuinely creates target/bin/python via `cat >`, which
   either overwrites the real app/.venv/bin/python or, since it is a
   symlink `uv` manages, writes straight through it into the real,
   shared-across-projects interpreter binary. A safe, non-mutating
   demonstration (fake `uv` that only logs argv, never touches disk) proved
   the exact target path was the real one, without risking it: confirmed in
   session, not committed. Every test now runs a `lab_copy` fixture that
   copies setup-venv.sh plus app/requirements.txt and
   app/requirements-dev.txt into tmp_path with the layout the script
   depends on, so its own path resolution lands entirely inside tmp_path.
   A new autouse fingerprint fixture (symlink type + target + content
   sha256) asserts the real app/.venv/bin/python is byte-identical before
   and after every test in the file, as a permanent regression tripwire.
   Verified real app/.venv/bin/python is unchanged and still imports
   Pillow after the full suite.

2. setup-venv.sh runs before azd-postprovision.sh's ACR build and
   Container App update, not after, so a failure here means the cloud app
   deployment for this run has not started yet -- only Bicep provisioning
   has. Re-running just this script would leave that deployment silently
   skipped. RERUN_HINT no longer offers "(or directly:
   ./scripts/setup-venv.sh)"; every failure message now points only at
   `cd <lab root> && azd hooks run postprovision`. doctor.sh and
   capture-scenario.sh keep recommending ./scripts/setup-venv.sh directly,
   since those run well after a deployment has already succeeded and are
   unaffected by this ordering. README's deploy section and guide 01's
   failure table are reworded to match: the README no longer implies "only
   a local step remains" when this step fails, and guide 01 now has a
   remedy row for doctor's Python environment check.

3. uv's real "no matching Python" failure (`error: No interpreter found for
   Python >=X ...`, verified against uv 0.10.9) now gets its own actionable
   line recommending `uv python install 3.12`, framed as reusing the same
   corporate proxy/mirror already configured for uv -- never a public pip
   install.

Tests: 421 passed (390 scripts/tests + 21 infra/tests + 10 app/tests).
bash -n scripts/*.sh clean.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… hooks

The remaining Task 6 blocking finding: test_azd_postprovision_reports_a
_clear_error_when_the_azure_cli_is_not_logged_in executed the real,
in-place azd-postprovision.sh via _run_hook_script(AZD_POSTPROVISION,
...), only pointing az stubs and AZURE_* env vars at tmp_path. Since
setup-venv.sh (called by azd-postprovision.sh before any Azure CLI call)
resolves its own SCRIPT_DIR/LAB_ROOT/VENV_DIR from its own
${BASH_SOURCE[0]}, never the caller's cwd, that always ran the real,
system uv against the real, developer-machine app/.venv -- a genuine
network/package-index call and filesystem mutation, unrelated to the
login-failure behaviour the test claims to check. The same
_run_hook_script helper, and _run_azd_configure, also ran
azd-configure.sh in place (lower risk, since it never touches
setup-venv.sh/uv, but still not isolated).

Every test in test_azd_hooks.py that executes a hook script now runs it
from a new lab_copy fixture: a throwaway copy of azd-configure.sh,
azd-postprovision.sh, and the real setup-venv.sh under scripts/, the
real app/requirements.txt / requirements-dev.txt under app/, and a
placeholder azure.yaml at the copied root -- laid out exactly as the
real lab does, so every script's own path resolution lands entirely
inside tmp_path. The postprovision login-failure test now puts a fake
uv (logging every call, never touching a real network or a real venv)
on PATH alongside the login-failing fake az, so the real, copied
setup-venv.sh genuinely runs its uv-venv/uv-pip-install/Pillow-import
logic against the fake before the login check fails -- the uv call log
assertion ("venv" in uv_calls, "pip install" in uv_calls) proves the
fake, not the real, uv did that work, and the created venv is asserted
to live under tmp_path, never under the real lab tree.

A new module-wide autouse fixture, real_lab_venv_tree_is_never_touched,
fingerprints the real app/.venv tree before and after every test in this
file: a manifest of every entry's relative path, kind
(file/dir/symlink), symlink target, size, and mtime, hashed into one
digest, plus a separate byte-content sha256 of bin/python's resolved
target. The manifest catches whole-tree structural/symlink/size/mtime
drift that a single canary file could miss; the interpreter content hash
independently catches a content-only mutation that happens to preserve
size and mtime, which the manifest alone would not (verified in
isolation: a same-size, same-mtime content change is still detected via
the interpreter hash). mtime is only ever one signal among several,
never checked alone.

test_no_execution_helper_runs_a_real_in_place_hook_script is a second,
purely static tripwire: a source-text regex scan of this file itself
(no subprocess, no filesystem access outside __file__) that fails if any
test again passes the real, in-place AZD_CONFIGURE/AZD_POSTPROVISION/
SETUP_VENV path constants straight to subprocess.run or
_run_hook_script. Confirmed RED against the pre-fix, git-committed
version of this file via this same safe regex scan (no execution): it
matched the AZD_CONFIGURE and AZD_POSTPROVISION call sites inside
_run_hook_script, and the AZD_CONFIGURE call site inside
_run_azd_configure. GREEN now: all patterns absent from the rewritten
file, and the meta-test itself passes.

Verified: pytest scripts/tests/test_azd_hooks.py -- 14 passed. Full
suite: pytest scripts/tests -- 391 passed; pytest infra/tests -- 21
passed; app/.venv/bin/python -m pytest app -- 10 passed (422 total).
Real app/.venv confirmed byte-for-byte and structurally unchanged
throughout: bin/python sha256 and file count (14735 entries) identical
before and after every run, and Pillow still imports from it afterward.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The lab is a Container Apps workload pulling from ACR with a
user-assigned managed identity, which cannot be provisioned and deployed
in one step: the lab image does not exist until the registry builds it,
and the AcrPull assignment the same ARM deployment creates is not
necessarily usable by a pull issued immediately afterwards. The single
postprovision hook did both halves -- setup-venv.sh, then az acr build,
az containerapp ingress update and az containerapp update --image --
so the very first pull raced the role assignment, with no check at all
that the identity could pull yet. That is the live-deployment blocker the
Azure-deploy preflight flags.

The two phases are now separate, and the gate sits between them:

* postprovision -> scripts/azd-postprovision-local.sh: setup-venv.sh and
  nothing else. It makes zero Azure calls, so `azd provision` ends with
  the public placeholder image (ingress 80, no probes) still serving, and
  says so on stdout ("Next: azd deploy").
* postdeploy -> scripts/azd-deploy-app.sh: guards the seven provision
  outputs it consumes (naming `azd provision` when one is missing),
  checks the CLI login, resolves the registry, then polls
  `az role assignment list` for the *exact* AcrPull assignment -- this
  principal, role definition 7f951dda-4ed3-4680-a7ca-43fe172d538d, the
  registry's own scope, compared case-insensitively -- for up to 300s
  (SRE_ACR_PULL_TIMEOUT_SECONDS, 10s interval) before its first deploy
  action. Only then: az acr build from app/ (never a local docker
  build), containerapp registry set --identity <UAMI>, ingress update
  --target-port 8000, containerapp update --image, wait for a new
  Healthy+active revision, wait for /healthz 200, and finally
  `azd env set SRE_IMAGE_TAG` / `SRE_CONTAINER_IMAGE` (--cwd-pinned).
  If the grant never appears, nothing is built or updated and the hook
  fails naming AcrPull and the scope it polled.

`azd up` runs the same postdeploy hook in its deploy phase, so the
documented single command still passes through the gate.

azd 1.29 behaviour was verified rather than assumed, because the project
declares no services (the only applicable service host, containerapp +
docker, would force a local Docker build): `azd deploy --no-prompt`
against a marker-hook copy of this exact azure.yaml shape ran postdeploy,
and a hook exiting 7 failed the command with "failed running post hooks:
'postdeploy' hook failed with exit code: '7'". Running deploy before any
provision fails fast with "infrastructure has not been provisioned" --
env.GetSubscriptionId() == "" in cli/azd/internal/cmd/deploy.go, not an
ARM query -- and cli/azd/internal/cmd/up_graph.go (tag
azure-dev-cli_1.29.0) adds cmdhook-predeploy/cmdhook-postdeploy
unconditionally with an explicit "Zero-service projects" branch, so
`azd up` reaches the same hook. No service definition was needed.

main.bicep now publishes AZURE_CONTAINER_APP_PRINCIPAL_ID,
AZURE_WORKLOAD_IDENTITY_RESOURCE_ID and AZURE_ACR_LOGIN_SERVER (azd
stores outputs under the template's declared names), with
workloadIdentityResourceId threaded through workload.bicep and lab.bicep.

Tests came first. A new deploy_app_harness.py runs the hook as a program
against a fake `az` that models the tenant instead of logging and
exiting 0: role assignments are records, `role assignment list` applies
azure-cli's assignee filter, exact-scope match (and refuses
--include-inherited) and the ends_with(roleDefinitionId, ...) projection,
available_after_attempts reproduces RBAC propagation delay, and
containerapp update rolls a new revision whose health revision list
reports. test_azd_deploy_app.py (25 tests) pins the ordering both in the
source and in the recorded call log -- no deploy call before the grant is
visible, three polls when the grant appears on the third, no build at all
when it never appears, wider scope / AcrPush / another principal all
rejected, a differently-cased scope accepted, and no image recorded when
/healthz never goes green. test_azd_hooks.py now proves the provision
phase makes no Azure call and builds nothing, and setup-venv.sh's rerun
hint no longer claims an image build follows it in the same hook.

README, doctor.sh, deploy.sh, cleanup.sh and cleanup-external.sh describe
the two phases and the new script names; .azure/deployment-plan.md
records the discovered gate and moves back to Status "Ready for
Validation" -- nothing was deployed to Azure by this change.

Validation: 455 tests, bash -n on every script, three Bicep builds,
azure.yaml against the official azd v1.0 schema, azd package --all, and
the azd deploy-hook probes above.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ph dependency, doctor placeholder state)

Follow-up review of the AcrPull poll (`acr_pull_is_visible` in
scripts/azd-deploy-app.sh, from dab05d6) found it calling
`az role assignment list --assignee-principal-type ServicePrincipal`.
That flag does not exist on `role assignment list` (only on `role
assignment create`); verified against the installed Azure CLI (2.89.1)
with a harmless real query -- it fails with `ERROR: unrecognized
arguments: --assignee-principal-type ServicePrincipal`, exit 2. Every
poll would have failed this way, and the deploy phase would never build
the image.

Fixed by removing that flag and adding the two flags the command does
support for exactly this purpose: `--fill-principal-name false
--fill-role-definition-name false`. `--assignee-object-id` already
bypasses Microsoft Graph for the assignee filter, but both `--fill-*`
flags default to true and would still query Graph to populate fields
this poll never reads -- setting them false keeps the poll working with
no Graph reachability at all (confirmed against the real CLI: exit 0,
no warning).

scripts/tests/deploy_app_harness.py's fake `az` was hardened to reject
any `role assignment list` flag the real parser does not recognise
(strict RED/GREEN): with the harness hardened but before the script fix,
11 tests failed with the fake's `ERROR: unrecognized arguments` (RED);
after the fix, all 30 tests in test_azd_deploy_app.py pass (GREEN). Added
a static contract test (no `--assignee-principal-type`, both `--fill-*`
flags present) and a behavioural test asserting the poll still succeeds
against the hardened fake.

Also, while reviewing the same two-phase gate:

* infra/main.bicep still attributed the image build/switch to
  "postprovision" in two comments; that work moved to the `postdeploy`
  hook (scripts/azd-deploy-app.sh) in dab05d6. Corrected the wording and
  added a regression test asserting the stale word is gone.
* doctor.sh's `/healthz` check reported the same generic "Investigate:
  curl -v" FAIL whether `azd deploy` simply had not run yet (the
  documented, expected placeholder state -- port 80, no `/healthz`) or
  the lab image was genuinely unhealthy after deployment, misclassifying
  a standard pre-deploy state as a broken one. It now reads the azd
  environment's SRE_CONTAINER_IMAGE (set only once the deploy phase
  succeeds, added to common.sh's load_lab_config): empty means the
  deploy phase has not run, and the FAIL detail now names the exact
  remedy (`azd deploy --no-prompt`) instead of the generic message; once
  set, a `/healthz` failure is still reported as a real regression. Added
  behavioural tests for both branches.

Full suite: 460 pytest tests green (app/infra/scripts), bash -n on every
script, `az bicep build` on all five templates, and azure.yaml validated
against the azd v1.0 schema -- all pass. Plan Status stays Ready for
Validation; recorded the fix and updated test counts in
.azure/deployment-plan.md.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Live `azd provision` failed ARM validation for all three
Microsoft.Insights/scheduledQueryRules@2023-12-01 alert rules with
`QueryNotContainKnownTable: One-minute frequency is not supported for
this query. Either switch to five-minute frequency or adapt the
query.`

Root cause: infra/alerts.bicep's requests/dependencies Application
Insights queries were paired with evaluationFrequency: 'PT1M', a
cadence those queries do not support (five minutes or coarser only).
`az bicep build` never calls ARM, so this was never caught before a
real deployment attempt.

TDD:
- RED: added test_evaluation_frequency_is_five_minutes_not_one_minute
  to infra/tests/test_alerts_bicep.py, plus two doc-contract tests in
  scripts/tests/test_lab_guides.py asserting README.md and
  dynamic-thresholds.md no longer claim one-minute static evaluation.
  3 failed against the unmodified template/docs.
- GREEN: changed evaluationFrequency to 'PT5M' for all three rules in
  alerts.bicep (windowSize stays 'PT5M', thresholds unchanged --
  neither is implicated by this failure); updated README.md's cost
  callout and dynamic-thresholds.md's Static Threshold section to say
  5 minutes. validation-results.md is left untouched: it is a
  historical record of a past run, not current design guidance.

Verified: full suite 463 passed (app/tests, infra/tests,
scripts/tests); `az bicep build` passes for main/lab/workload/alerts.bicep,
and the compiled alerts.bicep ARM JSON now shows
"evaluationFrequency": "PT5M". No Azure resources were deployed or
deleted.

.azure/deployment-plan.md keeps Status: Ready for Validation and now
records this failure and fix, since the live provision attempt that
surfaced it did not complete successfully.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The previous commit read `QueryNotContainKnownTable: One-minute frequency
is not supported for this query.` literally and moved all three
`Microsoft.Insights/scheduledQueryRules@2023-12-01` rules to PT5M. That
was the wrong root cause and is reverted here.

Real root cause: the rules were scoped to the Application Insights
component and queried the legacy resource-centric schema (`requests`,
`dependencies`, `timestamp`, `cloud_RoleName`, `duration`). On a
workspace-based component those legacy names are functions over the
workspace tables, and "the query calls a function that calls other
tables" is one of the documented one-minute-frequency limitations
(aka.ms/lsa_1m_limits -> alerts-create-log-alert-rule#frequency), which
is exactly what "does not contain a known table" reports.

Fix (TDD, RED first -- 11 failing tests):
- infra/alerts.bicep takes `workspaceResourceId`, queries the workspace
  known tables `AppRequests`/`AppDependencies` with the exact casing
  already used by scripts/query-evidence.sh (TimeGenerated, AppRoleName,
  Name, ResultCode, DurationMs, Target; percentile(DurationMs, 95) needs
  no timespan conversion), scopes each rule to the workspace, sets
  `targetResourceTypes: ['Microsoft.OperationalInsights/workspaces']`,
  and restores `evaluationFrequency: 'PT1M'` with `windowSize: 'PT5M'`.
- infra/lab.bicep passes `observability.outputs.workspaceId` to the
  alerts module; the workspaceId output chain (observability -> lab ->
  main -> AZURE_WORKSPACE_ID) and `appInsightsResourceId` are unchanged.
- Thresholds and the default fire/resolve timeouts (720s/900s) are
  untouched: the one-minute cadence they were sized for is back.
- README.md is back to "1분 주기 로그 검색 경고 규칙 3개",
  dynamic-thresholds.md to "evaluation: 1분(window 5분)" plus the
  workspace-scope precondition, and validation-results.md now annotates
  why its recorded one-minute run stays plausible.

Live evidence (nothing deployed, modified or deleted):
- `az deployment group validate` against the partially provisioned lab
  resource group: `provisioningState: Succeeded`, `error: null`, the
  three scheduledQueryRules validated.
- Honest limit, measured: the pre-fix template (component scope, legacy
  tables, PT1M) also passes preflight, so validate proves the template
  shape, not the query. The query side was proven directly with
  `az monitor log-analytics query` on the live workspace: all three
  alert queries resolve (Failures=0, P95DurationMs=None,
  DependencyFailures=0 with no traffic yet), while `requests | take 1`
  fails with SEM0100 "Failed to resolve table or column expression named
  'requests'".

Verified: 468 passed (app/tests, infra/tests, scripts/tests);
`az bicep build` passes for main/lab/workload/observability/alerts.bicep
and the compiled alerts.bicep emits PT1M/PT5M with the workspace scope
and target type. .azure/deployment-plan.md stays "Ready for Validation"
because the definitive proof is the still-pending live `azd provision`.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fixes discovered in review of the workspace-scope alert docs:

- Add doc-contract tests (RED first) requiring every scenario guide to
  name the AppRequests/AppDependencies workspace tables the deployed
  alert rules actually query, and forbidding the legacy Application
  Insights component-scope names (`requests`/`dependencies`) in that
  instructional content.
- Fix guides/02-04-scenario-*.md to name AppRequests/AppDependencies
  (with matching column casing) instead of the legacy schema.
- Document in runbooks/incident-response.md and guides/05-results.md
  that the alert rule's own scope/affected resource is the Log
  Analytics workspace, and that a telemetry row's _ResourceId column
  points to the Application Insights component -- neither is the
  workload. The Agent must still identify the actually affected
  Container App/service from AppRoleName/Name, and the impact_scope
  scoring criterion must not credit reporting the alert's own target as
  the answer.
- Correct .azure/deployment-plan.md's appInsightsResourceId wording:
  it remains an output for backward compatibility, but no module or
  script reads it -- the alerts module takes workspaceResourceId, and
  telemetry-querying scripts read workspaceCustomerId instead.
- Append the current, corrected workspace-scope alert record to the
  local task-7-report.md (Section 8), replacing the obsolete PT5M-only
  conclusion left by an earlier, reverted fix. (.superpowers/ is
  gitignored, so this file is not part of this commit.)

Plan status stays Ready for Validation; no Azure resources were
deployed, modified, or deleted.

Verified: app/.venv/bin/python -m pytest app/tests infra/tests
scripts/tests -> 472 passed (468 -> +4 new doc-contract tests).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ing them

`recover()` in run-scenario.sh is called as `if ! recover` from the EXIT
trap, and bash disables `set -e` inside a function invoked in a condition:
a rejected `az containerapp update`, a revision that never became ready or
a refused `az role assignment create` fell through to `RECOVERED=1` and
returned 0, so the run exited claiming a recovery that never happened while
the injected fault stayed live. `recover_on_exit`'s CRITICAL branch was
unreachable.

Recovery is now split into `restore_container_app_env` and
`restore_blob_role`, every az call, command substitution and wait is checked
explicitly, `RECOVERED=1` is reached only after a whole branch succeeded (so
a failed attempt is retried by the trap), and `recover_on_exit` prints what
failed plus the manual remedy before exiting non-zero. RED first: three
behaviour tests (S1 update rejected, S2 revision stalled, S3 role restore
refused on the exit-trap path) plus failure injection in the fake `az`.
`wait_for_new_revision_ready` gained an optional poll-interval argument so
those waits are bounded in tests; the production defaults are unchanged.

Also in this review pass:

- capture-scenario.sh takes the scenario only. The legacy optional evidence
  directory let an arbitrary directory's capture be recorded as the current
  environment's capture status; archived runs are re-rendered with
  capture_agent.py/render_capture.py, which record no state.
- .azure/deployment-plan.md is reconciled with what actually ran: the status
  is Validated for the base azd deployment only, and it says on the status
  line that the manual Agent connection and the live S1/S2/S3
  run/capture/score sequence have not been run. The stale "Ready for
  Validation", "no resources deployed", fresh-environment and unchecked
  preview notes are gone, the recorded test/Bicep counts match Section 7,
  and Live Deployment Proof gains a pending row for the manual sequence.
- dynamic-thresholds.md numbering no longer skips section 8.
- the lab .gitignore ignores the five compiled Bicep outputs by name;
  infra/main.parameters.json stays tracked, verified with git check-ignore.

Tests: 483 passed (472 before); bash -n, Python 3.9.6 imports and five
Bicep builds pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
run-scenario.sh only recorded an outcome at the end, so a re-run that died
early left the run it replaced in the state file. Successful S1, recovered
and captured; re-run S1; the injecting `az containerapp update` is rejected,
or recovery times out, before `mark-recovered`/`mark-failed` can run. The
file still said `s1.run_status: recovered` and `s1.capture_status:
conclusion`, so `require-run s2` admitted S2 and `score` scored a run that
no longer existed -- on evidence from an attempt whose fault may still be
live.

`LabState.begin_run(scenario, evidence_dir)` (CLI: `lab_state begin-run`)
now starts an attempt atomically: `require_run` first, so it cannot be
called out of order or against another environment's state file, then the
whole scenario entry is cleared -- run_status, capture_status,
failure_reason, alert_resolved_at, the previous evidence directory and any
terminal capture metadata -- and replaced with `run_status: running`,
`started_at` and the new evidence directory. The entry is cleared wholesale
rather than by a named list of fields so a field added later cannot silently
start surviving a re-run.

run-scenario.sh calls it after `require-run`, after the evidence directory
exists and before the EXIT trap is armed and the first destructive az call
is made. `running` satisfies no gate: `has('sX_recovered')` accepts only
`recovered` and `has('sX_captured')` only a recorded `conclusion`, so any
later early exit, trap failure or Ctrl-C leaves `running` or `failed`, the
next scenario stays blocked and `score` reports no captured evidence.
`mark_recovered`/`mark_failed` still complete a `running` attempt normally,
and still clear a stale `capture_status` themselves for state written
without a recorded start.

RED first: an end-to-end behaviour test replays old success -> re-run broken
at injection and at recovery -> S2 refused with no injection az call and
score blocked, plus tests that the attempt is recorded before the fault is
injected and completed by a healthy run; ten state unit tests and four CLI
tests cover the prerequisite and environment checks, the cleared fields,
persistence and the dropped evidence directory. The fake `az` gained an
injection-failure marker.

Also in this pass:

- validation-results.md opens by dating itself to the hand-built pre-azd lab
  (2026-08-12) so its resource group and Logic App bridge are not read as
  the current azd flow.
- dynamic-thresholds.md labels the two event paths: incident platform by
  default, Action Group -> Logic App as the legacy bridge that azd does not
  deploy.
- guides/01-agent-setup.md warns under both Learn screenshots that they show
  `Autonomous` while the lab must choose `Review`. No image changed.
- the scenario guides say a re-run clears the previous attempt's
  `sX_recovered`/`sX_captured` before injection.

Tests: 505 passed (483 before); bash -n under Bash 3.2, Python 3.9.6 imports
and five Bicep builds pass.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The ordered prerequisites look exactly one scenario back, so they cannot see
a run that broke somewhere else. Complete S1, S2 and S3 with real captures,
then re-run S1 and let the re-run die before it records an outcome: S1 is
`running` or `failed` -- its fault may still be live in the Container App all
three scenarios share -- but S2's entry is untouched, still `recovered` +
`conclusion`. `require_run("s3")` read only that and admitted S3, which
injected a third fault on top of an incident nobody had resolved and produced
two captures that can no longer be told apart.

`LabState.require_run` now applies a second, independent rule after the
ordered one: no run may start while any scenario is `running` or `failed`.
Re-running the *earliest* failed scenario is the single exception -- it is the
documented remedy, and making it unconditional is what guarantees the gate can
always be worked off from the top, so no state (including a hand-edited one
with two failures at once) can lock the lab. A scenario that is `running` is
never restartable, since two live injections of the same fault leave neither
capture readable. Every blocker is named in lab order with its status and the
command that clears it. The ordered rules still run first, so their more
specific message survives; `_remedy` became status-aware so it never points at
a command the gate would itself refuse -- while S1 is `running` the S2 refusal
now says `lab_state.py mark-failed s1`, not `lab.sh run s1`.

"Recovered but not captured" is deliberately *not* a blocker: a recovered run
is finished, so it keeps blocking only the scenario whose prerequisite names
it. The ordered rules are unchanged.

Defence in depth around the same state, since `state.json` is an editable file
and a truncated write looks like a valid one:

* `record_capture` refuses `conclusion` unless the run is `recovered`. The
  three missing markers stay recordable for any run status -- they measure
  what the Agent failed to produce, cannot unblock a scenario and cannot earn
  a point -- so diagnostic honesty is kept without a way to inflate a result.
  `capture-scenario.sh` reports that refusal instead of dying on a command
  substitution, and names the raw evidence still on disk.
* `score.py` takes `run_status` and fails every criterion, FAIL/0, when it is
  not `recovered`, whatever the capture says; the recorded capture status is
  still reported verbatim and the table gained a `run=` cell. Two independent
  checks must now be defeated before an unresolved incident can score.

The evidence directory is a name first and a directory second: `common.sh`
gained `evidence_dir_path`, which only builds the path, and `run-scenario.sh`
names it, calls `begin-run` (which re-checks the gate, closing the window
after `require-run`), then creates it. A refused run no longer litters
`evidence/` with an empty `sN-<timestamp>/` that reads like a real attempt.
`create_evidence_dir` stays for `baseline.sh`, implemented on top.

Negative tests no longer restate the real validation subscription ID to prove
it is absent: they forbid the UUID *shape*, minus Azure's tenant-independent
built-in role definition IDs, which also catches the next person's
subscription. `validation-results.md` remains the one file that records it,
and two new privacy guards derive it from there to prove no test source or
fixture repeats it.

RED first throughout: 17 failing state/CLI tests for the gate and the capture
rule, 17 failing scorer tests, then the shell and guide tests. New coverage
includes the end-to-end reproduction (broken S1 re-run refuses S3 with no
`role assignment delete` and no second S3 evidence directory), running/failed
S1 blocking S2 and S3, failed S2 blocking S1 and S3, normal re-runs once
everything recovered, the double-failure escape hatch, the gate reading disk
rather than process memory, a harness probe proving the evidence directory
does not exist when `begin-run` is called, and guide/README text for the new
precondition.

Verified: 557 passed (scripts/tests + infra/tests, was 495), 10 passed
(app/tests under the app venv), `bash -n` on all scripts under Bash 3.2.57,
`lab_state.py` and `score.py` import under Python 3.9.6, and `az bicep build`
clean on all five templates.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Keep privacy guards effective without requiring the historical validation report to retain a real subscription ID. Document that unrecovered runs always score zero.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Reworks the monitor/sre-agent-event-lab into an azd up-first guided workflow with a two-phase provision/deploy model (local venv prep in postprovision; ACR build + Container App switch gated on AcrPull propagation in postdeploy), plus a CLI-driven run/capture/score flow backed by explicit state and evidence.

Changes:

  • Introduces azd project structure (subscription-scope infra/main.bicep, azure.yaml hooks) and updates alerts to use Log Analytics workspace-schema tables at PT1M cadence.
  • Adds guided lab entrypoints (lab.sh, scenario/baseline/capture scripts) and a scoring tool (score.py) with extensive behavioral tests/harnesses.
  • Adds safer teardown for external subscription-scoped role assignments (cleanup-external.sh) and compatibility wrappers for legacy commands.

Reviewed changes

Copilot reviewed 62 out of 66 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
monitor/sre-agent-event-lab/validation-results.md Adds “legacy record” disclaimers and historical notes to validation results.
monitor/sre-agent-event-lab/scripts/tests/test_setup_venv.py New behavioral tests for setup-venv.sh with a fake uv and regression guard against touching real venv.
monitor/sre-agent-event-lab/scripts/tests/test_score.py New behavioral tests for score.py rubric/verdict logic and CLI behavior.
monitor/sre-agent-event-lab/scripts/tests/test_privacy.py Adds privacy guards around subscription-ID handling in docs/tests.
monitor/sre-agent-event-lab/scripts/tests/test_lab_cli.py New end-to-end tests for lab.sh dispatch and flow.
monitor/sre-agent-event-lab/scripts/tests/test_briefing_docs.py Adjusts briefing-doc tests to account for additional guide screenshots.
monitor/sre-agent-event-lab/scripts/tests/test_briefing_assets.py Strengthens asset privacy check to forbid GUID-shaped leaks.
monitor/sre-agent-event-lab/scripts/tests/test_baseline.py New behavioral tests for baseline.sh polling and parsing.
monitor/sre-agent-event-lab/scripts/tests/test_azd_env.py New tests for azd-backed config resolution and safety properties (no eval/indirection, correct precedence).
monitor/sre-agent-event-lab/scripts/tests/test_azd_deploy_app.py New tests for gated deploy hook ordering, subscription pinning, and AcrPull propagation wait.
monitor/sre-agent-event-lab/scripts/tests/cleanup_harness.py Harness to execute cleanup-external.sh against staged fake az/azd.
monitor/sre-agent-event-lab/scripts/tests/azd_fake.py Fake azd implementation capturing azd 1.29.0 observable contract for tests.
monitor/sre-agent-event-lab/scripts/tests/azd_common_harness.py Shared harness for driving common.sh with fake az/azd in tests.
monitor/sre-agent-event-lab/scripts/setup-venv.sh New uv-only idempotent venv setup with actionable failure guidance.
monitor/sre-agent-event-lab/scripts/score.py New rubric-based scorer producing scorecard.json and a TSV table, with failure/INCOMPLETE semantics.
monitor/sre-agent-event-lab/scripts/run-scenario.sh Tightens run gating via state, makes recovery explicit/validated, and records failed runs with reasons.
monitor/sre-agent-event-lab/scripts/query-evidence.sh Switches to require_lab_config preflight.
monitor/sre-agent-event-lab/scripts/lab.sh New single entrypoint CLI orchestrating doctor/baseline/ack/run/capture/score.
monitor/sre-agent-event-lab/scripts/deploy.sh Becomes a compatibility wrapper forwarding to azd up.
monitor/sre-agent-event-lab/scripts/cleanup.sh Becomes a compatibility wrapper forwarding to azd cleanup + optional legacy RG delete.
monitor/sre-agent-event-lab/scripts/cleanup-external.sh New teardown hook to remove recorded external subscription role assignments and clear hook-set image env values.
monitor/sre-agent-event-lab/scripts/capture-scenario.sh Simplifies capture interface (scenario-only), resolves evidence dir from state, records capture status, improves local-env errors.
monitor/sre-agent-event-lab/scripts/baseline.sh New baseline load + bounded polling of workspace telemetry, writes evidence and gates scenario execution.
monitor/sre-agent-event-lab/scripts/azd-postprovision-local.sh New azd postprovision hook to prepare local Python environment only.
monitor/sre-agent-event-lab/scripts/azd-deploy-app.sh New azd postdeploy hook implementing AcrPull gate + ACR build + Container App image/ingress update + health verification.
monitor/sre-agent-event-lab/scripts/azd-configure.sh New azd preprovision hook to register providers and ensure required env values exist.
monitor/sre-agent-event-lab/runbooks/incident-response.md Updates safety boundary and clarifies alert scope vs workload scope for incident reporting.
monitor/sre-agent-event-lab/infra/workload.bicep Parameterizes container target port and probe enablement for placeholder vs lab image.
monitor/sre-agent-event-lab/infra/tests/test_azd_project.py Adds infra/azd contract tests (hook wiring, placeholder behavior, ignore rules, GUID leak checks).
monitor/sre-agent-event-lab/infra/tests/test_alerts_bicep.py Updates alert tests to enforce workspace scope, known tables, casing, and PT1M evaluation.
monitor/sre-agent-event-lab/infra/subscription.bicepparam Deletes legacy hardcoded bicepparam (suffix/expiry/subscription-specific).
monitor/sre-agent-event-lab/infra/subscription.bicep Deletes legacy subscription-scope template superseded by azd-driven main.bicep.
monitor/sre-agent-event-lab/infra/observability.bicep Documents why alerts are scoped to the Log Analytics workspace resource ID.
monitor/sre-agent-event-lab/infra/main.parameters.json Adds azd parameter file mapping env vars into main.bicep.
monitor/sre-agent-event-lab/infra/main.bicepparam Deletes legacy hardcoded bicepparam.
monitor/sre-agent-event-lab/infra/main.bicep Moves to subscription scope, creates RG/tags, drives placeholder vs lab image behavior, exports azd outputs.
monitor/sre-agent-event-lab/infra/lab.bicep New RG-scope module composing observability/workload/alerts and exporting outputs.
monitor/sre-agent-event-lab/infra/alerts.bicep Switches queries from legacy AI schema to workspace schema, and scopes rules to workspace.
monitor/sre-agent-event-lab/guides/05-results.md New guide step for scoring and interpreting results.
monitor/sre-agent-event-lab/guides/04-scenario-s3.md New guide step for S3 (Storage RBAC) scenario and success criteria.
monitor/sre-agent-event-lab/guides/03-scenario-s2.md New guide step for S2 (latency) scenario and success criteria.
monitor/sre-agent-event-lab/guides/02-scenario-s1.md New guide step for S1 (HTTP 500) scenario and success criteria.
monitor/sre-agent-event-lab/guides/01-agent-setup.md New guide step for manual Agent setup + azd env settings + acknowledgement evidence.
monitor/sre-agent-event-lab/dynamic-thresholds.md Updates static-threshold notes and clarifies workspace-scope requirement for PT1M with known tables.
monitor/sre-agent-event-lab/azure.yaml Declares azd infra module and hooks for provision/deploy/down lifecycle.
monitor/sre-agent-event-lab/.gitignore Ignores azd local state and compiled bicep outputs (without hiding real parameter JSON).
monitor/sre-agent-event-lab/.env.example Adds documented env variables for azd + lab configuration defaults.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread monitor/sre-agent-event-lab/validation-results.md Outdated
Comment thread monitor/sre-agent-event-lab/scripts/cleanup-external.sh Outdated
Redact the historical subscription ID and make cancellation semantics explicit: predown can remove recorded external roles before the azd confirmation prompt, while postdown preserves image settings unless deletion succeeds.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@hellices
hellices requested a lite review from Copilot August 15, 2026 00:49

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 62 out of 66 changed files in this pull request and generated no new comments.

@hellices
hellices merged commit 3720a3e into main Aug 15, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants