sre-lab: add azd-first Azure SRE Agent event lab - #47
Merged
Merged
Conversation
Add azure.yaml and main.parameters.json so azd owns provisioning of the SRE Agent event lab. Convert infra/main.bicep to a subscription-scoped azd entry point (derives suffix, resource group, and merged azd tags) and move the former resource-group module to infra/lab.bicep, which now always deploys the Container App and alert rules with a placeholder image. Remove the now-redundant infra/subscription.bicep(param) in favor of the azd entry point. Add scripts/azd-configure.sh (preprovision: provider registration and env defaults) and scripts/azd-postprovision.sh (az acr build + update the existing Container App to the immutable image, then health-check it), keeping the lab Docker-free. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Address the Task 1 review findings on the azd migration: - Add scripts/cleanup-external.sh so every azure.yaml hook points at an existing executable script. It removes only the recorded subscription-scoped role assignments and exits 0 when the Agent setup evidence is absent, so azd down never fails and never deletes broadly. - Parameterize the Container App port and probes. The first provision runs the public placeholder on port 80 without /healthz probes, then postprovision moves ingress to 8000, swaps in the ACR-built image, waits for a healthy revision, and verifies /healthz. Probes now always match the port they check. - Persist the built image in SRE_CONTAINER_IMAGE and bind it (plus ACTION_GROUP_RESOURCE_ID) in main.parameters.json, so a later provision keeps the lab image instead of reverting to the placeholder. - Turn deploy.sh into a compatibility wrapper around azd up, refresh the README deployment/teardown commands, and update the tests, so no tracked file or documented command references the deleted subscription-scope templates. - Pin every Azure CLI call in the azd hooks to AZURE_SUBSCRIPTION_ID and report a mismatched active account. Resource-group and app operations read the current azd values only. - Delete infra/lab.bicepparam, which pinned a dead suffix and expiry. - Restore the deployment outputs the lab scripts still read and ignore the project-local .azure/ directory. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… guard no-login, mark legacy scripts Fixes the four remaining Task 1 review findings from review-f960e97..a1cceb6.diff, under strict TDD (RED tests written and confirmed failing first, then GREEN implementation): 1. README's `azd env new` command, prerequisites line, and the "실측 환경" table no longer hardcode the original validation subscription ID. The command now omits --subscription (azd prompts interactively) with guidance to pass --subscription <YOUR_SUBSCRIPTION_ID> instead. New regression test test_azd_onboarding_docs_and_config_do_not_hardcode_a_subscription_id scans README/azure.yaml/Bicep/parameters/hook scripts for the literal ID (common.sh and validation-results.md are intentionally excluded, as historical/legacy records already covered by their own tests). 2. cleanup-external.sh (the predown hook) now requires azd and, when run with --yes (as azure.yaml's predown hook always does), clears SRE_CONTAINER_IMAGE and SRE_IMAGE_TAG via `azd env set KEY ""` before/independent of the evidence-gated role-assignment cleanup, so reusing the same azd environment falls back to the placeholder image on the next provision. A dry run (no --yes) only prints the planned action. New tests exercise this via a stubbed azd binary and confirm no az group/resource delete calls occur as a result. 3. README now carries an explicit transitional caveat at the top of "## 시나리오 실행" stating run-scenario.sh/query-evidence.sh still read common.sh's legacy pre-azd deployment lookup and are legacy-only until common.sh is rewritten to use azd env get-value (a separate, not-yet-done task); points readers to the Baseline section's azd env get-value-based checks in the meantime. 4. azd-configure.sh and azd-postprovision.sh now guard their `az account show --query id -o tsv` call and exit with a clear "Azure CLI is not signed in. Run 'az login'..." message instead of leaking raw Azure CLI stderr when the CLI is signed out. Verified: azd env new --help / azd env set --help / a scratch azd env new + azd env set KEY "" run locally confirm the empty-value clearing semantics and that there is no `azd env unset`. Tests: 124 passed (infra/tests + scripts/tests, up from 118 passed/6 failed RED baseline), app suite 10 passed unchanged, bash -n clean, az bicep build clean for main/lab/workload.bicep, azure.yaml validated against the official azd JSON schema. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replace common.sh's fixed subscription/resource-group/deployment-name constants with load_lab_config, which resolves every setting as explicit process environment > current `azd env get-value` > an allowed default via a new setting()/require_setting() helper (no eval, no dynamic indirection). - deployment_output() now reads the AZURE_-prefixed (and legacy-named) azd deployment outputs load_lab_config already resolved, instead of `az deployment sub show` against a hardcoded deployment name. - verify_lab_resource_group() now requires both the purpose and azd-env-name tags to match the current azd environment, replacing the old single hardcoded resource-group-name safety boundary. - run-scenario.sh, query-evidence.sh, capture-scenario.sh, and cleanup.sh all call require_lab_config (require_commands + load_lab_config) before reading SUBSCRIPTION_ID/RESOURCE_GROUP or any deployment_output() value. - Added monitor/sre-agent-event-lab/.env.example documenting every setting name and its allowed default, with no secrets. - Rewrote the common.sh tests that intentionally pinned the old hardcoded subscription/resource-group values to assert the new azd-backed resolution instead, and added test_azd_env.py (with a shared fake azd/az harness in azd_common_harness.py) covering precedence, missing required settings, and the tag-based resource-group safety check. - Updated README's safety-boundary and scenario-execution sections: the scenario/query scripts are no longer legacy/transitional, they read the currently selected azd environment. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
`azd env get-value` writes its `ERROR: ...` diagnostics to stdout and signals failure only through the exit status (verified against azd 1.29.0), so keeping stdout unconditionally turned "key not found" and "no project exists" into configuration values: SUBSCRIPTION_ID became an error sentence, documented defaults never applied, and the preprovision hook skipped deriving AZURE_RESOURCE_GROUP. Every lookup now uses the value only when azd exited 0 and pins `--cwd` to the lab's azd project so the scripts work from the repository root or any other directory. `load_lab_config` also declared `WORKSPACE_CUSTOMER_ID` / `CONTAINER_APP_PRINCIPAL_ID` readonly, which aborted query-evidence.sh under `set -e` on its own assignment; the resolved values moved to LAB_-prefixed names. Tests now run the four entry points as programs against fake az/azd/python executables from a working directory outside the lab, the fake azd reproduces azd 1.29's contract (both missing-value shapes), and the text-order/eval-substring assertions were replaced by behaviour, fail-closed and token-aware checks. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Add scripts/lab.sh as a single-command entry point dispatching to doctor|baseline|acknowledge agent-setup|run s1|s2|s3|capture s1|s2|s3|score. acknowledge/score depend on a later task's lab_state.py/score.py; until those land, lab.sh fails closed with a clear "not yet available" message (exit 3) instead of a raw file-not-found error. - Add scripts/doctor.sh: a 15-row environment diagnostic printing the CHECK<TAB>STATUS<TAB>DETAIL contract (PASS/FAIL/MANUAL). Verifies required commands, Azure CLI login, azd configuration, subscription and resource-group pinning, Container App health, /healthz, Application Insights telemetry, alert rule enablement, the SRE Agent resource (when configured), and Reader role assignment -- all through official stable Azure CLI/REST reads. Fails closed: once any prerequisite check fails, every subsequent Azure-touching check is blocked rather than attempting further calls. Repository connection, knowledge source, incident platform, and response plan are always MANUAL since no stable API exposes them. - Add scripts/baseline.sh: runs baseline load (30 orders + 10 documents requests) and verifies Application Insights shows telemetry for both endpoints via bounded polling (default 600s timeout / 20s interval, overridable via LAB_BASELINE_TELEMETRY_TIMEOUT_SECONDS/ LAB_BASELINE_TELEMETRY_POLL_INTERVAL_SECONDS). Writes evidence under evidence/baseline-<UTC>/. - Fix scripts/capture-scenario.sh executable bit (100644 -> 100755); it was the only lab script not marked executable, and lab.sh capture depends on invoking it directly. - Add scripts/tests/doctor_harness.py: a fake-CLI harness (real az/azd/ curl/python3 stubs on PATH) purpose-built for doctor.sh/baseline.sh/ lab.sh's call surface (container app health, /healthz, per-rule alert reads, per-principal role assignment lookups), mirroring the existing lab_script_harness.py pattern. - Add scripts/tests/test_doctor.py and scripts/tests/test_lab_cli.py: behavioural tests driving doctor.sh/lab.sh as real programs against the fake CLIs (not text-only assertions), covering the fully-healthy path, every individual FAIL condition, fail-closed gating once subscription/ azd configuration checks fail, azd --cwd pinning, and lab.sh's dispatch/auto-discovery/not-yet-available behavior. - Document the new lab.sh entry point and doctor check list in README.md. 173 tests pass (144 pre-existing + 29 new); bash -n and az bicep build verified. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The doctor and baseline telemetry checks were written against inferred CLI behaviour and could never report the truth against real Azure: - `az monitor log-analytics query -o json` prints a flat JSON array of row objects (the log-analytics extension flattens the REST envelope), so the `.tables[0].rows` parse always yielded zero rows and telemetry could only ever FAIL. Parse the real shape through a shared `log_analytics_row_count` helper, and fake that shape in both harnesses. - The doctor query used `| count`, whose single row exists even for an empty table, so row-count semantics could not distinguish data from no data. Query projected rows bounded by `take 1` instead. - `az role assignment list` hides parent-scope grants without `--include-inherited`, reporting a subscription-scoped Reader as missing. Ask for inherited assignments and distinguish direct from inherited in the detail. - A malformed `agent-setup.json` aborted the whole run through `jq` under `set -e`; it is now one FAIL row and exit 1 with the rest of the report intact. Also add the two missing prerequisite rows -- the `log-analytics` CLI extension (`az extension show`, with the exact `az extension add` remedy) and azd's own login state (`azd auth login --check-status --output json`, whose exit code is always 0, so only the printed status is trusted) -- give baseline a poll that always attempts at least once and never sleeps past its deadline, and correct the doctor comment that claimed it never reads the raw process environment. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Scenario runs, captures and scoring now share one state file bound to the azd environment they belong to, so a run can only start when the previous scenario really recovered and was captured, and evidence is scored from what the Agent actually produced. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ires Address Task 4 review findings: - LabState.mark_recovered/mark_failed now discard the scenario's previous capture_status (and its capture-side evidence_dir) before recording the new run's outcome. A conclusion captured against an earlier, superseded run could otherwise keep satisfying sX_captured -- and therefore the next scenario's gate and the scorer -- even though nothing has been captured for the current run yet. Both a recovered and a failed re-run now clear it; tested with the state reloaded from disk to prove it is actually persisted. - run-scenario.sh now calls `lab_state mark-failed` with the evidence directory and a reason before exiting when the target alert never fires within the poll window, matching every other failure path. The wait is now bounded by LAB_ALERT_FIRE_TIMEOUT_SECONDS / LAB_ALERT_FIRE_POLL_INTERVAL_SECONDS (default 720s/20s, same pattern as the existing alert-resolve and health-check overrides) so this is exercised in the test suite instead of only in production. The pre-existing recovery trap is untouched and still reverts the injected fault on this path. - LabState._load validates that `stages`, `scenarios`, and every entry inside them decode to JSON objects. A wrong type (list, string, null, number) now raises the same clean LabStateError a corrupt file already produces, instead of a raw TypeError/AttributeError traceback surfacing later from mark()/mark_recovered()/ record_capture(); the file is never silently reset to defaults. - Documented, in lab_state.py's module docstring, that state.json has no cross-process locking and concurrent operators racing the same environment can lose an update to a later writer -- a deliberate choice for a single-operator lab. No locking was added; nothing in the test suite showed a real concurrent-use need. Tests (all written RED first): test_lab_state.py gains test_rerunning_a_recovered_scenario_clears_the_stale_capture_status, test_rerunning_a_failed_scenario_clears_the_stale_capture_status, test_reloading_after_a_rerun_still_shows_the_cleared_capture_status, a parametrized test_a_state_file_with_the_wrong_json_shape_is_refused_not_reset (10 shapes), test_cli_reports_a_malformed_state_file_without_a_python_traceback, and test_module_documents_that_concurrent_operators_are_unsupported. test_lab_scripts.py gains test_run_scenario_marks_failed_when_the_alert_never_fires, backed by a new alert_fires=False mode in lab_script_harness.py's fake az. Verification: python3 -m pytest monitor/sre-agent-event-lab/scripts/tests -q 307 passed (was 291 before this change) bash -n monitor/sre-agent-event-lab/scripts/*.sh python3 -c "import lab_state, score" # Python 3.9.6, clean Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
`azd down` deletes the azd-owned resource group; the only lab artifacts it
cannot see are the subscription-scoped Monitoring Contributor assignments
the Azure SRE Agent setup recorded. cleanup-external.sh now removes exactly
those, after verifying every record, and cleanup.sh stops deleting a
resource group of its own.
- cleanup-external.sh verifies four things before deleting anything: the
recorded assignment ID names a role assignment in the subscription this
run resolved (AZURE_SUBSCRIPTION_ID > the current azd environment), the
Azure CLI is signed in to exactly that subscription, the record carries
the Agent principal it was created for, and the live assignment really
holds that principal, the Monitoring Contributor role definition and
subscription scope. Every record is verified before the first deletion,
so one untrusted record leaves the subscription untouched. An empty,
missing or already-absent record is a safe no-op (`RoleAssignmentNotFound`
is what the CLI really answers -- recorded from azure-cli against a live
subscription); malformed JSON, a missing principal, an unreadable
assignment, a mismatch or a failed deletion all fail closed, before azd
destroys anything.
- The hook-set SRE_CONTAINER_IMAGE/SRE_IMAGE_TAG reset moved out of
`predown` into the `postdown` hook (`--reset-image-env`). `predown` runs
before azd asks the operator to confirm the deletion, so clearing them
there broke the environment of an operator who answered "no"; azd runs a
post hook only after the action succeeded (HooksRunner.Invoke returns
early on failure, cli/azd/pkg/ext/hooks_runner.go), which is exactly when
the recorded image is really gone. `postdown` is an officially supported
azd command hook.
- Both azd env writes are pinned with `--cwd "${LAB_ROOT}"`, so running the
hook by hand from the repository root no longer resolves whatever azd
project the working directory happens to hold.
- cleanup.sh is now a compatibility wrapper: it names `azd down --purge`
and forwards to cleanup-external.sh. It never deletes a resource group
unless `--legacy-delete-resource-group` is passed -- the documented
recovery path for a lab whose azd environment was lost -- which still
runs the subscription and `purpose`/`azd-env-name` tag checks first.
Tests: scripts/tests/test_cleanup_external.py drives the hook as a program
from a directory that is not the lab, on macOS's Bash 3.2, against a fake
`az` holding a staged subscription and the azd 1.29 fake; the cleanup tests
that lived in test_azd_hooks.py moved there. azd_fake now logs each argument
with %q so a value azd is asked to clear is visible instead of vanishing.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- cleanup-external.sh: capture az rest's stdout and stderr into separate
files instead of merging them with 2>&1. A successful read that also
prints an azure-cli warning to stderr no longer corrupts the ARM JSON,
and a failing read's RoleAssignmentNotFound check now inspects stderr
alone instead of a stream that could contain stray stdout content.
Added --only-show-errors to ask azure-cli itself to drop most warnings
first.
- Fixed the 'already verified' dedupe check: it matched any recorded ID
that was merely a suffix of an already-verified one (case pattern
*"${id}"$'\n'* has no left boundary), silently skipping validation
of a second, distinct record. Now bounded by a leading newline as well,
requiring a full \n<id>\n match, still a plain string (Bash 3.2 safe).
- README: documented that predown runs before azd's own delete
confirmation prompt, so canceling that prompt does not restore already
-removed role assignments; an operator must re-create the assignment
and re-run 'lab.sh acknowledge agent-setup'. No longer implies canceling
keeps a fully working environment.
- test_cleanup_runs_under_bash_32 now pins /bin/bash explicitly via
run_cleanup(..., bash=system_bash) instead of relying on whatever
'bash' resolves to on PATH.
- Added behaviour tests for all of the above plus fake az/az-fake harness
support for staged stderr warnings and stdout noise (cleanup_harness.py
Staged, lab_script_harness.py dispatch fix for --only-show-errors).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fixes six findings from the Task 6 review (review-f80be21..f14edd7.diff), all with RED/GREEN tests: 1. The documented azd-first flow never prepared app/.venv: capture-scenario.sh hard-requires app/.venv/bin/python but nothing created it. Added a new, idempotent scripts/setup-venv.sh (uv venv --python ">=3.10" --allow-existing plus uv pip install -r requirements-dev.txt, verifying Pillow importability), invoked from the postprovision hook before any Azure CLI call. uv is mandatory, not merely preferred (this lab's proxy is configured for uv, not public PyPI), so there is no silent pip fallback -- missing uv is an actionable failure. Because postprovision runs after azd provision has already created cloud resources, every failure message states the exact rerun command (azd hooks run postprovision, or ./scripts/setup-venv.sh directly). doctor.sh gained a "Python environment" check (venv + Pillow readiness) and capture-scenario.sh gained its own Pillow-importability precondition, both pointing at the same rerun command. 2. Appended a correction section to the (gitignored) task-6-report.md documenting these inaccuracies, per its own "historical artifact" status, rather than rewriting the report in place. 3. Restored the verified https://azuresre.dev audience fact (deleted, not relocated, by the Task 6 README rewrite) into validation-results.md's historical bridge section, explicitly framed as legacy record for the non-default Logic App bridge. 4. Fixed README and guides/05-results.md cwd instructions: each document now uses exactly one cd, so every command block in it is runnable sequentially in one shell from that single working directory. 5. Reworded the README's scorecard row from "10 point max" (ambiguous as the overall max) to "10 per scenario, 30 overall", matching score.py's real MAX_POINTS (10 per scenario, 30 overall). 6. Fixed two accuracy issues in the numbered guides: guide 02's start conditions claimed an unenforced concurrency lock that lab_state.py's own docstring says does not exist; guide 05 ran the stdlib-only generate_notifications.py through app/.venv/bin/python, implying a Pillow dependency it does not have. Tests: 414 passed (scripts/tests + infra/tests + app/tests). bash -n and az bicep build both clean. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… wording
Fixes three remaining Task 6 review findings, all with RED/GREEN tests.
1. test_setup_venv.py ran the *real*, in-place setup-venv.sh: the script
resolves SCRIPT_DIR/LAB_ROOT/VENV_DIR from its own ${BASH_SOURCE[0]},
never from the caller's cwd, so pointing the process cwd at a tmp_path
(as the old tests did) still targeted the real app/.venv -- and the
tests' fake `uv` genuinely creates target/bin/python via `cat >`, which
either overwrites the real app/.venv/bin/python or, since it is a
symlink `uv` manages, writes straight through it into the real,
shared-across-projects interpreter binary. A safe, non-mutating
demonstration (fake `uv` that only logs argv, never touches disk) proved
the exact target path was the real one, without risking it: confirmed in
session, not committed. Every test now runs a `lab_copy` fixture that
copies setup-venv.sh plus app/requirements.txt and
app/requirements-dev.txt into tmp_path with the layout the script
depends on, so its own path resolution lands entirely inside tmp_path.
A new autouse fingerprint fixture (symlink type + target + content
sha256) asserts the real app/.venv/bin/python is byte-identical before
and after every test in the file, as a permanent regression tripwire.
Verified real app/.venv/bin/python is unchanged and still imports
Pillow after the full suite.
2. setup-venv.sh runs before azd-postprovision.sh's ACR build and
Container App update, not after, so a failure here means the cloud app
deployment for this run has not started yet -- only Bicep provisioning
has. Re-running just this script would leave that deployment silently
skipped. RERUN_HINT no longer offers "(or directly:
./scripts/setup-venv.sh)"; every failure message now points only at
`cd <lab root> && azd hooks run postprovision`. doctor.sh and
capture-scenario.sh keep recommending ./scripts/setup-venv.sh directly,
since those run well after a deployment has already succeeded and are
unaffected by this ordering. README's deploy section and guide 01's
failure table are reworded to match: the README no longer implies "only
a local step remains" when this step fails, and guide 01 now has a
remedy row for doctor's Python environment check.
3. uv's real "no matching Python" failure (`error: No interpreter found for
Python >=X ...`, verified against uv 0.10.9) now gets its own actionable
line recommending `uv python install 3.12`, framed as reusing the same
corporate proxy/mirror already configured for uv -- never a public pip
install.
Tests: 421 passed (390 scripts/tests + 21 infra/tests + 10 app/tests).
bash -n scripts/*.sh clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… hooks
The remaining Task 6 blocking finding: test_azd_postprovision_reports_a
_clear_error_when_the_azure_cli_is_not_logged_in executed the real,
in-place azd-postprovision.sh via _run_hook_script(AZD_POSTPROVISION,
...), only pointing az stubs and AZURE_* env vars at tmp_path. Since
setup-venv.sh (called by azd-postprovision.sh before any Azure CLI call)
resolves its own SCRIPT_DIR/LAB_ROOT/VENV_DIR from its own
${BASH_SOURCE[0]}, never the caller's cwd, that always ran the real,
system uv against the real, developer-machine app/.venv -- a genuine
network/package-index call and filesystem mutation, unrelated to the
login-failure behaviour the test claims to check. The same
_run_hook_script helper, and _run_azd_configure, also ran
azd-configure.sh in place (lower risk, since it never touches
setup-venv.sh/uv, but still not isolated).
Every test in test_azd_hooks.py that executes a hook script now runs it
from a new lab_copy fixture: a throwaway copy of azd-configure.sh,
azd-postprovision.sh, and the real setup-venv.sh under scripts/, the
real app/requirements.txt / requirements-dev.txt under app/, and a
placeholder azure.yaml at the copied root -- laid out exactly as the
real lab does, so every script's own path resolution lands entirely
inside tmp_path. The postprovision login-failure test now puts a fake
uv (logging every call, never touching a real network or a real venv)
on PATH alongside the login-failing fake az, so the real, copied
setup-venv.sh genuinely runs its uv-venv/uv-pip-install/Pillow-import
logic against the fake before the login check fails -- the uv call log
assertion ("venv" in uv_calls, "pip install" in uv_calls) proves the
fake, not the real, uv did that work, and the created venv is asserted
to live under tmp_path, never under the real lab tree.
A new module-wide autouse fixture, real_lab_venv_tree_is_never_touched,
fingerprints the real app/.venv tree before and after every test in this
file: a manifest of every entry's relative path, kind
(file/dir/symlink), symlink target, size, and mtime, hashed into one
digest, plus a separate byte-content sha256 of bin/python's resolved
target. The manifest catches whole-tree structural/symlink/size/mtime
drift that a single canary file could miss; the interpreter content hash
independently catches a content-only mutation that happens to preserve
size and mtime, which the manifest alone would not (verified in
isolation: a same-size, same-mtime content change is still detected via
the interpreter hash). mtime is only ever one signal among several,
never checked alone.
test_no_execution_helper_runs_a_real_in_place_hook_script is a second,
purely static tripwire: a source-text regex scan of this file itself
(no subprocess, no filesystem access outside __file__) that fails if any
test again passes the real, in-place AZD_CONFIGURE/AZD_POSTPROVISION/
SETUP_VENV path constants straight to subprocess.run or
_run_hook_script. Confirmed RED against the pre-fix, git-committed
version of this file via this same safe regex scan (no execution): it
matched the AZD_CONFIGURE and AZD_POSTPROVISION call sites inside
_run_hook_script, and the AZD_CONFIGURE call site inside
_run_azd_configure. GREEN now: all patterns absent from the rewritten
file, and the meta-test itself passes.
Verified: pytest scripts/tests/test_azd_hooks.py -- 14 passed. Full
suite: pytest scripts/tests -- 391 passed; pytest infra/tests -- 21
passed; app/.venv/bin/python -m pytest app -- 10 passed (422 total).
Real app/.venv confirmed byte-for-byte and structurally unchanged
throughout: bin/python sha256 and file count (14735 entries) identical
before and after every run, and Pillow still imports from it afterward.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The lab is a Container Apps workload pulling from ACR with a
user-assigned managed identity, which cannot be provisioned and deployed
in one step: the lab image does not exist until the registry builds it,
and the AcrPull assignment the same ARM deployment creates is not
necessarily usable by a pull issued immediately afterwards. The single
postprovision hook did both halves -- setup-venv.sh, then az acr build,
az containerapp ingress update and az containerapp update --image --
so the very first pull raced the role assignment, with no check at all
that the identity could pull yet. That is the live-deployment blocker the
Azure-deploy preflight flags.
The two phases are now separate, and the gate sits between them:
* postprovision -> scripts/azd-postprovision-local.sh: setup-venv.sh and
nothing else. It makes zero Azure calls, so `azd provision` ends with
the public placeholder image (ingress 80, no probes) still serving, and
says so on stdout ("Next: azd deploy").
* postdeploy -> scripts/azd-deploy-app.sh: guards the seven provision
outputs it consumes (naming `azd provision` when one is missing),
checks the CLI login, resolves the registry, then polls
`az role assignment list` for the *exact* AcrPull assignment -- this
principal, role definition 7f951dda-4ed3-4680-a7ca-43fe172d538d, the
registry's own scope, compared case-insensitively -- for up to 300s
(SRE_ACR_PULL_TIMEOUT_SECONDS, 10s interval) before its first deploy
action. Only then: az acr build from app/ (never a local docker
build), containerapp registry set --identity <UAMI>, ingress update
--target-port 8000, containerapp update --image, wait for a new
Healthy+active revision, wait for /healthz 200, and finally
`azd env set SRE_IMAGE_TAG` / `SRE_CONTAINER_IMAGE` (--cwd-pinned).
If the grant never appears, nothing is built or updated and the hook
fails naming AcrPull and the scope it polled.
`azd up` runs the same postdeploy hook in its deploy phase, so the
documented single command still passes through the gate.
azd 1.29 behaviour was verified rather than assumed, because the project
declares no services (the only applicable service host, containerapp +
docker, would force a local Docker build): `azd deploy --no-prompt`
against a marker-hook copy of this exact azure.yaml shape ran postdeploy,
and a hook exiting 7 failed the command with "failed running post hooks:
'postdeploy' hook failed with exit code: '7'". Running deploy before any
provision fails fast with "infrastructure has not been provisioned" --
env.GetSubscriptionId() == "" in cli/azd/internal/cmd/deploy.go, not an
ARM query -- and cli/azd/internal/cmd/up_graph.go (tag
azure-dev-cli_1.29.0) adds cmdhook-predeploy/cmdhook-postdeploy
unconditionally with an explicit "Zero-service projects" branch, so
`azd up` reaches the same hook. No service definition was needed.
main.bicep now publishes AZURE_CONTAINER_APP_PRINCIPAL_ID,
AZURE_WORKLOAD_IDENTITY_RESOURCE_ID and AZURE_ACR_LOGIN_SERVER (azd
stores outputs under the template's declared names), with
workloadIdentityResourceId threaded through workload.bicep and lab.bicep.
Tests came first. A new deploy_app_harness.py runs the hook as a program
against a fake `az` that models the tenant instead of logging and
exiting 0: role assignments are records, `role assignment list` applies
azure-cli's assignee filter, exact-scope match (and refuses
--include-inherited) and the ends_with(roleDefinitionId, ...) projection,
available_after_attempts reproduces RBAC propagation delay, and
containerapp update rolls a new revision whose health revision list
reports. test_azd_deploy_app.py (25 tests) pins the ordering both in the
source and in the recorded call log -- no deploy call before the grant is
visible, three polls when the grant appears on the third, no build at all
when it never appears, wider scope / AcrPush / another principal all
rejected, a differently-cased scope accepted, and no image recorded when
/healthz never goes green. test_azd_hooks.py now proves the provision
phase makes no Azure call and builds nothing, and setup-venv.sh's rerun
hint no longer claims an image build follows it in the same hook.
README, doctor.sh, deploy.sh, cleanup.sh and cleanup-external.sh describe
the two phases and the new script names; .azure/deployment-plan.md
records the discovered gate and moves back to Status "Ready for
Validation" -- nothing was deployed to Azure by this change.
Validation: 455 tests, bash -n on every script, three Bicep builds,
azure.yaml against the official azd v1.0 schema, azd package --all, and
the azd deploy-hook probes above.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ph dependency, doctor placeholder state) Follow-up review of the AcrPull poll (`acr_pull_is_visible` in scripts/azd-deploy-app.sh, from dab05d6) found it calling `az role assignment list --assignee-principal-type ServicePrincipal`. That flag does not exist on `role assignment list` (only on `role assignment create`); verified against the installed Azure CLI (2.89.1) with a harmless real query -- it fails with `ERROR: unrecognized arguments: --assignee-principal-type ServicePrincipal`, exit 2. Every poll would have failed this way, and the deploy phase would never build the image. Fixed by removing that flag and adding the two flags the command does support for exactly this purpose: `--fill-principal-name false --fill-role-definition-name false`. `--assignee-object-id` already bypasses Microsoft Graph for the assignee filter, but both `--fill-*` flags default to true and would still query Graph to populate fields this poll never reads -- setting them false keeps the poll working with no Graph reachability at all (confirmed against the real CLI: exit 0, no warning). scripts/tests/deploy_app_harness.py's fake `az` was hardened to reject any `role assignment list` flag the real parser does not recognise (strict RED/GREEN): with the harness hardened but before the script fix, 11 tests failed with the fake's `ERROR: unrecognized arguments` (RED); after the fix, all 30 tests in test_azd_deploy_app.py pass (GREEN). Added a static contract test (no `--assignee-principal-type`, both `--fill-*` flags present) and a behavioural test asserting the poll still succeeds against the hardened fake. Also, while reviewing the same two-phase gate: * infra/main.bicep still attributed the image build/switch to "postprovision" in two comments; that work moved to the `postdeploy` hook (scripts/azd-deploy-app.sh) in dab05d6. Corrected the wording and added a regression test asserting the stale word is gone. * doctor.sh's `/healthz` check reported the same generic "Investigate: curl -v" FAIL whether `azd deploy` simply had not run yet (the documented, expected placeholder state -- port 80, no `/healthz`) or the lab image was genuinely unhealthy after deployment, misclassifying a standard pre-deploy state as a broken one. It now reads the azd environment's SRE_CONTAINER_IMAGE (set only once the deploy phase succeeds, added to common.sh's load_lab_config): empty means the deploy phase has not run, and the FAIL detail now names the exact remedy (`azd deploy --no-prompt`) instead of the generic message; once set, a `/healthz` failure is still reported as a real regression. Added behavioural tests for both branches. Full suite: 460 pytest tests green (app/infra/scripts), bash -n on every script, `az bicep build` on all five templates, and azure.yaml validated against the azd v1.0 schema -- all pass. Plan Status stays Ready for Validation; recorded the fix and updated test counts in .azure/deployment-plan.md. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Live `azd provision` failed ARM validation for all three Microsoft.Insights/scheduledQueryRules@2023-12-01 alert rules with `QueryNotContainKnownTable: One-minute frequency is not supported for this query. Either switch to five-minute frequency or adapt the query.` Root cause: infra/alerts.bicep's requests/dependencies Application Insights queries were paired with evaluationFrequency: 'PT1M', a cadence those queries do not support (five minutes or coarser only). `az bicep build` never calls ARM, so this was never caught before a real deployment attempt. TDD: - RED: added test_evaluation_frequency_is_five_minutes_not_one_minute to infra/tests/test_alerts_bicep.py, plus two doc-contract tests in scripts/tests/test_lab_guides.py asserting README.md and dynamic-thresholds.md no longer claim one-minute static evaluation. 3 failed against the unmodified template/docs. - GREEN: changed evaluationFrequency to 'PT5M' for all three rules in alerts.bicep (windowSize stays 'PT5M', thresholds unchanged -- neither is implicated by this failure); updated README.md's cost callout and dynamic-thresholds.md's Static Threshold section to say 5 minutes. validation-results.md is left untouched: it is a historical record of a past run, not current design guidance. Verified: full suite 463 passed (app/tests, infra/tests, scripts/tests); `az bicep build` passes for main/lab/workload/alerts.bicep, and the compiled alerts.bicep ARM JSON now shows "evaluationFrequency": "PT5M". No Azure resources were deployed or deleted. .azure/deployment-plan.md keeps Status: Ready for Validation and now records this failure and fix, since the live provision attempt that surfaced it did not complete successfully. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The previous commit read `QueryNotContainKnownTable: One-minute frequency is not supported for this query.` literally and moved all three `Microsoft.Insights/scheduledQueryRules@2023-12-01` rules to PT5M. That was the wrong root cause and is reverted here. Real root cause: the rules were scoped to the Application Insights component and queried the legacy resource-centric schema (`requests`, `dependencies`, `timestamp`, `cloud_RoleName`, `duration`). On a workspace-based component those legacy names are functions over the workspace tables, and "the query calls a function that calls other tables" is one of the documented one-minute-frequency limitations (aka.ms/lsa_1m_limits -> alerts-create-log-alert-rule#frequency), which is exactly what "does not contain a known table" reports. Fix (TDD, RED first -- 11 failing tests): - infra/alerts.bicep takes `workspaceResourceId`, queries the workspace known tables `AppRequests`/`AppDependencies` with the exact casing already used by scripts/query-evidence.sh (TimeGenerated, AppRoleName, Name, ResultCode, DurationMs, Target; percentile(DurationMs, 95) needs no timespan conversion), scopes each rule to the workspace, sets `targetResourceTypes: ['Microsoft.OperationalInsights/workspaces']`, and restores `evaluationFrequency: 'PT1M'` with `windowSize: 'PT5M'`. - infra/lab.bicep passes `observability.outputs.workspaceId` to the alerts module; the workspaceId output chain (observability -> lab -> main -> AZURE_WORKSPACE_ID) and `appInsightsResourceId` are unchanged. - Thresholds and the default fire/resolve timeouts (720s/900s) are untouched: the one-minute cadence they were sized for is back. - README.md is back to "1분 주기 로그 검색 경고 규칙 3개", dynamic-thresholds.md to "evaluation: 1분(window 5분)" plus the workspace-scope precondition, and validation-results.md now annotates why its recorded one-minute run stays plausible. Live evidence (nothing deployed, modified or deleted): - `az deployment group validate` against the partially provisioned lab resource group: `provisioningState: Succeeded`, `error: null`, the three scheduledQueryRules validated. - Honest limit, measured: the pre-fix template (component scope, legacy tables, PT1M) also passes preflight, so validate proves the template shape, not the query. The query side was proven directly with `az monitor log-analytics query` on the live workspace: all three alert queries resolve (Failures=0, P95DurationMs=None, DependencyFailures=0 with no traffic yet), while `requests | take 1` fails with SEM0100 "Failed to resolve table or column expression named 'requests'". Verified: 468 passed (app/tests, infra/tests, scripts/tests); `az bicep build` passes for main/lab/workload/observability/alerts.bicep and the compiled alerts.bicep emits PT1M/PT5M with the workspace scope and target type. .azure/deployment-plan.md stays "Ready for Validation" because the definitive proof is the still-pending live `azd provision`. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fixes discovered in review of the workspace-scope alert docs: - Add doc-contract tests (RED first) requiring every scenario guide to name the AppRequests/AppDependencies workspace tables the deployed alert rules actually query, and forbidding the legacy Application Insights component-scope names (`requests`/`dependencies`) in that instructional content. - Fix guides/02-04-scenario-*.md to name AppRequests/AppDependencies (with matching column casing) instead of the legacy schema. - Document in runbooks/incident-response.md and guides/05-results.md that the alert rule's own scope/affected resource is the Log Analytics workspace, and that a telemetry row's _ResourceId column points to the Application Insights component -- neither is the workload. The Agent must still identify the actually affected Container App/service from AppRoleName/Name, and the impact_scope scoring criterion must not credit reporting the alert's own target as the answer. - Correct .azure/deployment-plan.md's appInsightsResourceId wording: it remains an output for backward compatibility, but no module or script reads it -- the alerts module takes workspaceResourceId, and telemetry-querying scripts read workspaceCustomerId instead. - Append the current, corrected workspace-scope alert record to the local task-7-report.md (Section 8), replacing the obsolete PT5M-only conclusion left by an earlier, reverted fix. (.superpowers/ is gitignored, so this file is not part of this commit.) Plan status stays Ready for Validation; no Azure resources were deployed, modified, or deleted. Verified: app/.venv/bin/python -m pytest app/tests infra/tests scripts/tests -> 472 passed (468 -> +4 new doc-contract tests). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ing them `recover()` in run-scenario.sh is called as `if ! recover` from the EXIT trap, and bash disables `set -e` inside a function invoked in a condition: a rejected `az containerapp update`, a revision that never became ready or a refused `az role assignment create` fell through to `RECOVERED=1` and returned 0, so the run exited claiming a recovery that never happened while the injected fault stayed live. `recover_on_exit`'s CRITICAL branch was unreachable. Recovery is now split into `restore_container_app_env` and `restore_blob_role`, every az call, command substitution and wait is checked explicitly, `RECOVERED=1` is reached only after a whole branch succeeded (so a failed attempt is retried by the trap), and `recover_on_exit` prints what failed plus the manual remedy before exiting non-zero. RED first: three behaviour tests (S1 update rejected, S2 revision stalled, S3 role restore refused on the exit-trap path) plus failure injection in the fake `az`. `wait_for_new_revision_ready` gained an optional poll-interval argument so those waits are bounded in tests; the production defaults are unchanged. Also in this review pass: - capture-scenario.sh takes the scenario only. The legacy optional evidence directory let an arbitrary directory's capture be recorded as the current environment's capture status; archived runs are re-rendered with capture_agent.py/render_capture.py, which record no state. - .azure/deployment-plan.md is reconciled with what actually ran: the status is Validated for the base azd deployment only, and it says on the status line that the manual Agent connection and the live S1/S2/S3 run/capture/score sequence have not been run. The stale "Ready for Validation", "no resources deployed", fresh-environment and unchecked preview notes are gone, the recorded test/Bicep counts match Section 7, and Live Deployment Proof gains a pending row for the manual sequence. - dynamic-thresholds.md numbering no longer skips section 8. - the lab .gitignore ignores the five compiled Bicep outputs by name; infra/main.parameters.json stays tracked, verified with git check-ignore. Tests: 483 passed (472 before); bash -n, Python 3.9.6 imports and five Bicep builds pass. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
run-scenario.sh only recorded an outcome at the end, so a re-run that died
early left the run it replaced in the state file. Successful S1, recovered
and captured; re-run S1; the injecting `az containerapp update` is rejected,
or recovery times out, before `mark-recovered`/`mark-failed` can run. The
file still said `s1.run_status: recovered` and `s1.capture_status:
conclusion`, so `require-run s2` admitted S2 and `score` scored a run that
no longer existed -- on evidence from an attempt whose fault may still be
live.
`LabState.begin_run(scenario, evidence_dir)` (CLI: `lab_state begin-run`)
now starts an attempt atomically: `require_run` first, so it cannot be
called out of order or against another environment's state file, then the
whole scenario entry is cleared -- run_status, capture_status,
failure_reason, alert_resolved_at, the previous evidence directory and any
terminal capture metadata -- and replaced with `run_status: running`,
`started_at` and the new evidence directory. The entry is cleared wholesale
rather than by a named list of fields so a field added later cannot silently
start surviving a re-run.
run-scenario.sh calls it after `require-run`, after the evidence directory
exists and before the EXIT trap is armed and the first destructive az call
is made. `running` satisfies no gate: `has('sX_recovered')` accepts only
`recovered` and `has('sX_captured')` only a recorded `conclusion`, so any
later early exit, trap failure or Ctrl-C leaves `running` or `failed`, the
next scenario stays blocked and `score` reports no captured evidence.
`mark_recovered`/`mark_failed` still complete a `running` attempt normally,
and still clear a stale `capture_status` themselves for state written
without a recorded start.
RED first: an end-to-end behaviour test replays old success -> re-run broken
at injection and at recovery -> S2 refused with no injection az call and
score blocked, plus tests that the attempt is recorded before the fault is
injected and completed by a healthy run; ten state unit tests and four CLI
tests cover the prerequisite and environment checks, the cleared fields,
persistence and the dropped evidence directory. The fake `az` gained an
injection-failure marker.
Also in this pass:
- validation-results.md opens by dating itself to the hand-built pre-azd lab
(2026-08-12) so its resource group and Logic App bridge are not read as
the current azd flow.
- dynamic-thresholds.md labels the two event paths: incident platform by
default, Action Group -> Logic App as the legacy bridge that azd does not
deploy.
- guides/01-agent-setup.md warns under both Learn screenshots that they show
`Autonomous` while the lab must choose `Review`. No image changed.
- the scenario guides say a re-run clears the previous attempt's
`sX_recovered`/`sX_captured` before injection.
Tests: 505 passed (483 before); bash -n under Bash 3.2, Python 3.9.6 imports
and five Bicep builds pass.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The ordered prerequisites look exactly one scenario back, so they cannot see
a run that broke somewhere else. Complete S1, S2 and S3 with real captures,
then re-run S1 and let the re-run die before it records an outcome: S1 is
`running` or `failed` -- its fault may still be live in the Container App all
three scenarios share -- but S2's entry is untouched, still `recovered` +
`conclusion`. `require_run("s3")` read only that and admitted S3, which
injected a third fault on top of an incident nobody had resolved and produced
two captures that can no longer be told apart.
`LabState.require_run` now applies a second, independent rule after the
ordered one: no run may start while any scenario is `running` or `failed`.
Re-running the *earliest* failed scenario is the single exception -- it is the
documented remedy, and making it unconditional is what guarantees the gate can
always be worked off from the top, so no state (including a hand-edited one
with two failures at once) can lock the lab. A scenario that is `running` is
never restartable, since two live injections of the same fault leave neither
capture readable. Every blocker is named in lab order with its status and the
command that clears it. The ordered rules still run first, so their more
specific message survives; `_remedy` became status-aware so it never points at
a command the gate would itself refuse -- while S1 is `running` the S2 refusal
now says `lab_state.py mark-failed s1`, not `lab.sh run s1`.
"Recovered but not captured" is deliberately *not* a blocker: a recovered run
is finished, so it keeps blocking only the scenario whose prerequisite names
it. The ordered rules are unchanged.
Defence in depth around the same state, since `state.json` is an editable file
and a truncated write looks like a valid one:
* `record_capture` refuses `conclusion` unless the run is `recovered`. The
three missing markers stay recordable for any run status -- they measure
what the Agent failed to produce, cannot unblock a scenario and cannot earn
a point -- so diagnostic honesty is kept without a way to inflate a result.
`capture-scenario.sh` reports that refusal instead of dying on a command
substitution, and names the raw evidence still on disk.
* `score.py` takes `run_status` and fails every criterion, FAIL/0, when it is
not `recovered`, whatever the capture says; the recorded capture status is
still reported verbatim and the table gained a `run=` cell. Two independent
checks must now be defeated before an unresolved incident can score.
The evidence directory is a name first and a directory second: `common.sh`
gained `evidence_dir_path`, which only builds the path, and `run-scenario.sh`
names it, calls `begin-run` (which re-checks the gate, closing the window
after `require-run`), then creates it. A refused run no longer litters
`evidence/` with an empty `sN-<timestamp>/` that reads like a real attempt.
`create_evidence_dir` stays for `baseline.sh`, implemented on top.
Negative tests no longer restate the real validation subscription ID to prove
it is absent: they forbid the UUID *shape*, minus Azure's tenant-independent
built-in role definition IDs, which also catches the next person's
subscription. `validation-results.md` remains the one file that records it,
and two new privacy guards derive it from there to prove no test source or
fixture repeats it.
RED first throughout: 17 failing state/CLI tests for the gate and the capture
rule, 17 failing scorer tests, then the shell and guide tests. New coverage
includes the end-to-end reproduction (broken S1 re-run refuses S3 with no
`role assignment delete` and no second S3 evidence directory), running/failed
S1 blocking S2 and S3, failed S2 blocking S1 and S3, normal re-runs once
everything recovered, the double-failure escape hatch, the gate reading disk
rather than process memory, a harness probe proving the evidence directory
does not exist when `begin-run` is called, and guide/README text for the new
precondition.
Verified: 557 passed (scripts/tests + infra/tests, was 495), 10 passed
(app/tests under the app venv), `bash -n` on all scripts under Bash 3.2.57,
`lab_state.py` and `score.py` import under Python 3.9.6, and `az bicep build`
clean on all five templates.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Keep privacy guards effective without requiring the historical validation report to retain a real subscription ID. Document that unrecovered runs always score zero. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Contributor
There was a problem hiding this comment.
Pull request overview
Reworks the monitor/sre-agent-event-lab into an azd up-first guided workflow with a two-phase provision/deploy model (local venv prep in postprovision; ACR build + Container App switch gated on AcrPull propagation in postdeploy), plus a CLI-driven run/capture/score flow backed by explicit state and evidence.
Changes:
- Introduces azd project structure (subscription-scope
infra/main.bicep,azure.yamlhooks) and updates alerts to use Log Analytics workspace-schema tables at PT1M cadence. - Adds guided lab entrypoints (
lab.sh, scenario/baseline/capture scripts) and a scoring tool (score.py) with extensive behavioral tests/harnesses. - Adds safer teardown for external subscription-scoped role assignments (
cleanup-external.sh) and compatibility wrappers for legacy commands.
Reviewed changes
Copilot reviewed 62 out of 66 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
| monitor/sre-agent-event-lab/validation-results.md | Adds “legacy record” disclaimers and historical notes to validation results. |
| monitor/sre-agent-event-lab/scripts/tests/test_setup_venv.py | New behavioral tests for setup-venv.sh with a fake uv and regression guard against touching real venv. |
| monitor/sre-agent-event-lab/scripts/tests/test_score.py | New behavioral tests for score.py rubric/verdict logic and CLI behavior. |
| monitor/sre-agent-event-lab/scripts/tests/test_privacy.py | Adds privacy guards around subscription-ID handling in docs/tests. |
| monitor/sre-agent-event-lab/scripts/tests/test_lab_cli.py | New end-to-end tests for lab.sh dispatch and flow. |
| monitor/sre-agent-event-lab/scripts/tests/test_briefing_docs.py | Adjusts briefing-doc tests to account for additional guide screenshots. |
| monitor/sre-agent-event-lab/scripts/tests/test_briefing_assets.py | Strengthens asset privacy check to forbid GUID-shaped leaks. |
| monitor/sre-agent-event-lab/scripts/tests/test_baseline.py | New behavioral tests for baseline.sh polling and parsing. |
| monitor/sre-agent-event-lab/scripts/tests/test_azd_env.py | New tests for azd-backed config resolution and safety properties (no eval/indirection, correct precedence). |
| monitor/sre-agent-event-lab/scripts/tests/test_azd_deploy_app.py | New tests for gated deploy hook ordering, subscription pinning, and AcrPull propagation wait. |
| monitor/sre-agent-event-lab/scripts/tests/cleanup_harness.py | Harness to execute cleanup-external.sh against staged fake az/azd. |
| monitor/sre-agent-event-lab/scripts/tests/azd_fake.py | Fake azd implementation capturing azd 1.29.0 observable contract for tests. |
| monitor/sre-agent-event-lab/scripts/tests/azd_common_harness.py | Shared harness for driving common.sh with fake az/azd in tests. |
| monitor/sre-agent-event-lab/scripts/setup-venv.sh | New uv-only idempotent venv setup with actionable failure guidance. |
| monitor/sre-agent-event-lab/scripts/score.py | New rubric-based scorer producing scorecard.json and a TSV table, with failure/INCOMPLETE semantics. |
| monitor/sre-agent-event-lab/scripts/run-scenario.sh | Tightens run gating via state, makes recovery explicit/validated, and records failed runs with reasons. |
| monitor/sre-agent-event-lab/scripts/query-evidence.sh | Switches to require_lab_config preflight. |
| monitor/sre-agent-event-lab/scripts/lab.sh | New single entrypoint CLI orchestrating doctor/baseline/ack/run/capture/score. |
| monitor/sre-agent-event-lab/scripts/deploy.sh | Becomes a compatibility wrapper forwarding to azd up. |
| monitor/sre-agent-event-lab/scripts/cleanup.sh | Becomes a compatibility wrapper forwarding to azd cleanup + optional legacy RG delete. |
| monitor/sre-agent-event-lab/scripts/cleanup-external.sh | New teardown hook to remove recorded external subscription role assignments and clear hook-set image env values. |
| monitor/sre-agent-event-lab/scripts/capture-scenario.sh | Simplifies capture interface (scenario-only), resolves evidence dir from state, records capture status, improves local-env errors. |
| monitor/sre-agent-event-lab/scripts/baseline.sh | New baseline load + bounded polling of workspace telemetry, writes evidence and gates scenario execution. |
| monitor/sre-agent-event-lab/scripts/azd-postprovision-local.sh | New azd postprovision hook to prepare local Python environment only. |
| monitor/sre-agent-event-lab/scripts/azd-deploy-app.sh | New azd postdeploy hook implementing AcrPull gate + ACR build + Container App image/ingress update + health verification. |
| monitor/sre-agent-event-lab/scripts/azd-configure.sh | New azd preprovision hook to register providers and ensure required env values exist. |
| monitor/sre-agent-event-lab/runbooks/incident-response.md | Updates safety boundary and clarifies alert scope vs workload scope for incident reporting. |
| monitor/sre-agent-event-lab/infra/workload.bicep | Parameterizes container target port and probe enablement for placeholder vs lab image. |
| monitor/sre-agent-event-lab/infra/tests/test_azd_project.py | Adds infra/azd contract tests (hook wiring, placeholder behavior, ignore rules, GUID leak checks). |
| monitor/sre-agent-event-lab/infra/tests/test_alerts_bicep.py | Updates alert tests to enforce workspace scope, known tables, casing, and PT1M evaluation. |
| monitor/sre-agent-event-lab/infra/subscription.bicepparam | Deletes legacy hardcoded bicepparam (suffix/expiry/subscription-specific). |
| monitor/sre-agent-event-lab/infra/subscription.bicep | Deletes legacy subscription-scope template superseded by azd-driven main.bicep. |
| monitor/sre-agent-event-lab/infra/observability.bicep | Documents why alerts are scoped to the Log Analytics workspace resource ID. |
| monitor/sre-agent-event-lab/infra/main.parameters.json | Adds azd parameter file mapping env vars into main.bicep. |
| monitor/sre-agent-event-lab/infra/main.bicepparam | Deletes legacy hardcoded bicepparam. |
| monitor/sre-agent-event-lab/infra/main.bicep | Moves to subscription scope, creates RG/tags, drives placeholder vs lab image behavior, exports azd outputs. |
| monitor/sre-agent-event-lab/infra/lab.bicep | New RG-scope module composing observability/workload/alerts and exporting outputs. |
| monitor/sre-agent-event-lab/infra/alerts.bicep | Switches queries from legacy AI schema to workspace schema, and scopes rules to workspace. |
| monitor/sre-agent-event-lab/guides/05-results.md | New guide step for scoring and interpreting results. |
| monitor/sre-agent-event-lab/guides/04-scenario-s3.md | New guide step for S3 (Storage RBAC) scenario and success criteria. |
| monitor/sre-agent-event-lab/guides/03-scenario-s2.md | New guide step for S2 (latency) scenario and success criteria. |
| monitor/sre-agent-event-lab/guides/02-scenario-s1.md | New guide step for S1 (HTTP 500) scenario and success criteria. |
| monitor/sre-agent-event-lab/guides/01-agent-setup.md | New guide step for manual Agent setup + azd env settings + acknowledgement evidence. |
| monitor/sre-agent-event-lab/dynamic-thresholds.md | Updates static-threshold notes and clarifies workspace-scope requirement for PT1M with known tables. |
| monitor/sre-agent-event-lab/azure.yaml | Declares azd infra module and hooks for provision/deploy/down lifecycle. |
| monitor/sre-agent-event-lab/.gitignore | Ignores azd local state and compiled bicep outputs (without hiding real parameter JSON). |
| monitor/sre-agent-event-lab/.env.example | Adds documented env variables for azd + lab configuration defaults. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Redact the historical subscription ID and make cancellation semantics explicit: predown can remove recorded external roles before the azd confirmation prompt, while postdown preserves image settings unless deletion succeeds. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
azd up-first guided workflowSafety
Verification
python3 -m pytest scripts infra -q— 559 passedapp/.venv/bin/python -m pytest app -q— 10 passed/bin/bash -n scripts/*.shpython3 -m py_compile scripts/*.pyinfra/*.biceptemplates built successfullyazd provision,azd deploy,azd up, health/telemetry/alert checks, baseline, acknowledgement gate, andazd down --purgewere previously completedReviewer notes
The live S1/S2/S3 Agent investigation sequence is intentionally left for the guided user run; temporary validation resources were deleted after infrastructure verification.