diff --git a/.devcontainer/README.md b/.devcontainer/README.md new file mode 100644 index 0000000..95f2d06 --- /dev/null +++ b/.devcontainer/README.md @@ -0,0 +1,70 @@ +# 실습 도구 모음 + +이 저장소의 실습이 공통으로 쓰는 도구와 설치 방법입니다. 랩마다 목록을 복제하지 않고 이 문서 하나를 갱신합니다. + +## 컨테이너로 갖추기 (권장) + +`.devcontainer/devcontainer.json`이 저장소 전체의 기본 설정입니다. 별도로 고를 것이 없습니다. + +- **Codespaces**: 저장소에서 **Code > Codespaces > Create codespace** +- **로컬 VS Code**: **Dev Containers: Reopen in Container** + +컨테이너는 도구만 갖춰 줍니다. 로그인과 각 랩의 준비 명령은 실습 문서를 따라 직접 실행합니다. 무엇이 실행되는지 가려지지 않도록 컨테이너가 랩 명령을 대신 실행하지 않습니다. + +## 도구 목록 + +| 도구 | 쓰임 | 컨테이너에서 | +|---|---|---| +| `az` | Azure 리소스 조회·조작, 로그 쿼리 | `azure-cli` feature (`log-analytics`, `containerapp` 확장 포함) | +| `azd` | 랩 환경 프로비저닝(`azd up`)과 삭제(`azd down`) | `azure-dev/azd` feature | +| `gh` | GitHub 저장소·이슈 확인 | `github-cli` feature | +| `python3` | 랩의 Python 도구 실행 | `python` feature (3.12) | +| `uv` | Python 의존성 설치 | `python` feature의 `toolsToInstall` | +| `jq` | JSON 출력 파싱 | 베이스 이미지 | +| `curl` | 애플리케이션 엔드포인트 호출 | 베이스 이미지 | + +## 로컬에 직접 설치하기 + +컨테이너를 쓰지 않는다면 아래를 설치합니다. + +**macOS (Homebrew)** + +```bash +brew install azure-cli azure-dev gh python@3.12 uv jq +``` + +**Ubuntu / Debian** + +```bash +sudo apt-get update && sudo apt-get install -y jq curl python3 +curl -sSL https://aka.ms/InstallAzureCLIDeb | sudo bash +curl -fsSL https://aka.ms/install-azd.sh | bash +curl -LsSf https://astral.sh/uv/install.sh | sh +``` + +`gh`는 [GitHub CLI 설치 안내](https://github.com/cli/cli/blob/trunk/docs/install_linux.md)를 따릅니다. + +`az` 확장은 별도로 추가합니다. 랩에서 `az monitor log-analytics query`를 쓰려면 필요합니다. + +```bash +az extension add --name log-analytics +az extension add --name containerapp +``` + +사내 프록시 환경이라면 `uv`가 프록시로 구성된 인덱스를 그대로 씁니다. 랩의 Python 환경 준비 스크립트는 공개 PyPI로 폴백하지 않으므로, `uv`가 없으면 설치 안내와 함께 실패합니다. + +## 로그인 + +두 CLI는 자격 증명을 따로 관리하므로 각각 로그인합니다. + +```bash +az login --use-device-code +azd auth login +``` + +## 참고 + +- [Dev Containers 사양](https://containers.dev/implementors/json_reference/) +- [Codespaces 문서](https://docs.github.com/ko/codespaces) +- [azd 설치](https://learn.microsoft.com/azure/developer/azure-developer-cli/install-azd) +- [uv 설치](https://docs.astral.sh/uv/getting-started/installation/) diff --git a/.devcontainer/sre-agent-event-lab/devcontainer.json b/.devcontainer/devcontainer.json similarity index 72% rename from .devcontainer/sre-agent-event-lab/devcontainer.json rename to .devcontainer/devcontainer.json index afb0b51..9b3a0c6 100644 --- a/.devcontainer/sre-agent-event-lab/devcontainer.json +++ b/.devcontainer/devcontainer.json @@ -1,5 +1,5 @@ { - "name": "Azure SRE Agent Event Lab", + "name": "Azure DevGuide Sample", "image": "mcr.microsoft.com/devcontainers/base:ubuntu-24.04", "features": { "ghcr.io/devcontainers/features/azure-cli:1": { @@ -27,9 +27,5 @@ ] } }, - "postCreateCommand": "cd monitor/sre-agent-event-lab && ./scripts/setup-venv.sh", - "postAttachCommand": { - "next-steps": "echo 'Next: az login --use-device-code, then cd monitor/sre-agent-event-lab && source ./scripts/lab-env.sh'" - }, "remoteUser": "vscode" } diff --git a/monitor/sre-agent-event-lab/README.md b/monitor/sre-agent-event-lab/README.md index 3c4b6d8..d41a3a8 100644 --- a/monitor/sre-agent-event-lab/README.md +++ b/monitor/sre-agent-event-lab/README.md @@ -35,7 +35,7 @@ Azure SRE Agent는 이 실습이 만들지 않습니다. 미리 만들어 둔 Ag 이 실습은 **본인 fork에서** 진행합니다. Agent에 저장소를 연결하면 조사 결과 이슈가 그 저장소에 생성되므로, 원본을 연결하면 참가자 전원의 이슈가 한곳에 쌓이고 쓰기 권한도 없습니다. 1. 이 저장소를 본인 계정으로 fork합니다. -2. fork에서 **Code > Codespaces > New with options**를 열고 dev container로 `Azure SRE Agent Event Lab`을 고릅니다. `az`, `azd`, `gh`, Python, `uv`가 설치되고 `postCreateCommand`가 `setup-venv.sh`를 실행해 `app/.venv`까지 만듭니다. +2. fork에서 **Code > Codespaces > Create codespace**를 엽니다. 저장소 기본 dev container가 `az`, `azd`, `gh`, Python, `uv`, `jq`를 갖춰 줍니다. 도구 목록과 로컬 설치 방법은 [.devcontainer/README.md](../../.devcontainer/README.md)에 있습니다. 3. 터미널에서 로그인한 뒤 환경을 한 번 읽습니다. ```bash @@ -47,11 +47,11 @@ source ./monitor/sre-agent-event-lab/scripts/lab-env.sh `lab-env.sh`는 `azd`가 게시한 배포 출력만 읽어 리소스 그룹·구독·Container App·Storage 범위 등을 현재 셸에 export하고, 값이 하나라도 없으면 `LAB_READY=0`으로 알려 줍니다. 이후 가이드의 명령은 이 값들을 그대로 사용하므로 단계마다 다시 조회하지 않습니다. 비밀 값은 읽지도 출력하지도 않습니다. -로컬에서 진행한다면 아래 사전 조건을 직접 갖춘 뒤 같은 `source` 한 줄로 시작합니다. Codespaces에서는 dev container가 이미 갖춰 줍니다. +로컬에서 진행한다면 아래 사전 조건을 직접 갖춘 뒤 같은 `source` 한 줄로 시작합니다. ## 사전 조건 -- `az`, `azd`, `jq`, `curl`, `python3`, [`uv`](https://docs.astral.sh/uv/getting-started/installation/) — `uv`는 `app/.venv`를 만드는 `scripts/setup-venv.sh`가 쓰는 유일한 도구이며, 사내 프록시로 구성된 인덱스 설정을 그대로 씁니다(공개 PyPI `pip` 폴백 없음). +- `az`, `azd`, `jq`, `curl`, `python3`, `uv` — 설치 명령은 [.devcontainer/README.md](../../.devcontainer/README.md)가 관리합니다. `uv`는 `app/.venv`를 만드는 `scripts/setup-venv.sh`가 쓰는 유일한 도구이며, 사내 프록시로 구성된 인덱스 설정을 그대로 씁니다(공개 PyPI `pip` 폴백 없음). - `az extension add --name log-analytics` (`az monitor log-analytics query` 제공) - `az login`과 `azd auth login` — 두 CLI는 자격 증명을 따로 관리합니다. - 구독 Contributor, 역할 할당을 위한 Owner 또는 User Access Administrator @@ -65,7 +65,7 @@ source ./monitor/sre-agent-event-lab/scripts/lab-env.sh cd monitor/sre-agent-event-lab ``` -로컬 검증만 먼저 해 보려면 다음을 실행합니다. Codespaces에서는 `setup-venv.sh`가 이미 실행된 상태입니다. +로컬 검증만 먼저 해 보려면 다음을 실행합니다. `setup-venv.sh`는 `azd up`의 postprovision 단계에서도 실행되므로, 배포를 먼저 한 경우에는 이미 준비된 상태입니다. ```bash ./scripts/setup-venv.sh @@ -112,34 +112,24 @@ postprovision 단계가 실패하면 로컬 환경만 실패한 것입니다. `. 기본 실습에는 Logic App bridge를 배포하지 않습니다. 제품 표준 경로는 Azure Monitor를 incident platform으로 연결하는 것이고, 예전 실측에서 쓰던 Action Group + Logic App 인증 경로는 레거시 기록으로만 남아 있습니다([validation-results.md](validation-results.md)). -## 점검과 승인 +## 정상 상태 확인과 승인 + +장애를 주입하기 전에 정상 부하가 Application Insights까지 도달하는지 확인하고, 그 사실을 기록해야 S1이 열립니다. 텔레메트리가 없는 워크로드에 장애를 넣으면 주입한 장애와 원래부터 안 보이던 상태를 구별할 수 없습니다. 명령은 [guides/01-agent-setup.md](guides/01-agent-setup.md)에 있습니다. ```bash -./scripts/lab.sh doctor -./scripts/lab.sh baseline -./scripts/lab.sh acknowledge agent-setup +python3 scripts/lab_state.py mark baseline_passed --evidence-dir "${EVIDENCE_DIR}" +python3 scripts/lab_state.py acknowledge-agent ``` -`doctor`는 `CHECKSTATUSDETAIL` 한 줄씩 출력하고 `FAIL`이 하나라도 있으면 종료 코드 1을 반환합니다. 저장소 연결, 지식 원본, incident platform, 응답 계획은 공식 안정 API로 읽을 수 없어 항상 `MANUAL`입니다. `Python environment` 행은 `app/.venv`와 Pillow가 캡처(`capture-scenario.sh`)에 쓸 준비가 됐는지 확인하며, `FAIL`이면 `./scripts/setup-venv.sh`를 다시 실행하라고 안내합니다. - -`baseline`은 정상 부하를 넣고 Application Insights에 두 요청 종류가 모두 보일 때까지 최대 10분 기다립니다. `acknowledge agent-setup`은 대화형이며, 설정 값을 출력한 뒤 표준 입력으로 정확히 `acknowledge`를 입력해야 기록됩니다. +`acknowledge-agent`는 설정 값을 출력한 뒤 표준 입력으로 정확히 `acknowledge`를 입력해야 기록됩니다. ## 시나리오 실행 -각 시나리오 문서는 **수동 실행**을 먼저 설명합니다. `az containerapp update`, `az role assignment delete`처럼 실제로 Azure에 적용되는 명령을 그대로 실행하면서 무엇이 바뀌는지 확인하는 경로입니다. 처음 진행할 때는 이 경로를 권장합니다. - -같은 절차를 한 번에 실행하는 지름길도 각 문서 뒤쪽에 있습니다. +각 시나리오 문서는 실제로 Azure에 적용되는 명령을 그대로 실행하도록 안내합니다. `az containerapp update`, `az role assignment delete`처럼 무엇이 바뀌는지 보이는 명령만 씁니다. 시나리오를 대신 실행해 주는 스크립트는 없습니다. 장애를 넣고 되돌리는 일이 이 실습에서 배우는 내용이기 때문입니다. -```bash -./scripts/lab.sh run s1 -./scripts/lab.sh capture s1 -./scripts/lab.sh run s2 -./scripts/lab.sh capture s2 -./scripts/lab.sh run s3 -./scripts/lab.sh capture s3 -``` +진행 상태는 현재 azd 환경에 묶인 `evidence/state.json`에 기록되며, 순서를 어기면 첫 단계의 `lab_state.py begin-run`이 거부합니다. 순서와 별개로, 어떤 시나리오든 실행이 `running`이나 `failed`로 남아 있으면 세 시나리오 모두 새 실행이 거부됩니다. 세 시나리오는 Container App 하나를 공유하므로, 끝나지 않은 실행 하나가 남은 실습 전체를 막습니다. 복구 명령은 운영자가 직접 완료해야 합니다. -`run-scenario.sh`와 `capture-scenario.sh`는 `scripts/common.sh`의 `load_lab_config`로 "명시적 환경 변수 > 현재 `azd env get-value` > 허용된 기본값" 순서로 설정을 읽으므로, 고정된 구독이나 리소스 그룹이 스크립트 안에 없습니다. 진행 상태는 현재 azd 환경에 묶인 `evidence/state.json`에 기록되며 순서를 어기면 실행이 거부됩니다. 순서와 별개로, 어떤 시나리오든 실행이 `running`이나 `failed`로 남아 있으면 세 시나리오 모두 새 실행이 거부됩니다. 세 시나리오는 Container App 하나를 공유하므로, 끝나지 않은 실행 하나가 남은 실습 전체를 막습니다. 수동 실행도 첫 단계에서 `lab_state.py begin-run`을 호출해 같은 게이트를 적용받습니다. 차이는 복구입니다. 지름길은 종료 트랩이 장애를 자동으로 되돌리지만, 수동 실행에서는 복구 명령을 직접 완료해야 합니다. +증거 수집에 쓰는 `scripts/query-evidence.sh`는 `scripts/common.sh`의 `load_lab_config`로 "명시적 환경 변수 > 현재 `azd env get-value` > 허용된 기본값" 순서로 설정을 읽으므로, 고정된 구독이나 리소스 그룹이 스크립트 안에 없습니다. | 시나리오 | 주입하는 장애 | 안내 문서 | |---|---|---| @@ -150,7 +140,7 @@ postprovision 단계가 실패하면 로컬 환경만 실패한 것입니다. `. ## 결과 확인 ```bash -./scripts/lab.sh score +app/.venv/bin/python scripts/score.py --evidence-root evidence ``` 채점 기준, 사람이 채워야 하는 판정, 종합 판정 해석은 [guides/05-results.md](guides/05-results.md)에 있습니다. @@ -166,7 +156,7 @@ azd down --purge - predown hook `scripts/cleanup-external.sh --yes`: `evidence/agent-setup.json`에 기록된 구독 범위 Monitoring Contributor 할당만 제거합니다. 기록된 principal·역할·범위가 실제 할당과 모두 일치할 때만 삭제하고, 하나라도 어긋나면 아무것도 지우지 않습니다. - postdown hook `scripts/cleanup-external.sh --reset-image-env --yes`: 기록된 `SRE_CONTAINER_IMAGE`와 `SRE_IMAGE_TAG`를 비웁니다. 삭제가 실제로 성공한 뒤에만 실행되어야 하므로 predown이 아니라 postdown입니다. -중요: predown hook은 `azd down`이 삭제 **확인** 프롬프트를 띄우기 **전에** 실행됩니다. 그 프롬프트에서 **취소**해도 이미 제거된 Monitoring Contributor 할당은 돌아오지 않습니다. 리소스 그룹은 남지만 Agent의 구독 범위 권한은 사라진 상태이므로, 계속 쓰려면 역할 할당을 다시 만들고 `evidence/agent-setup.json`을 새 할당 ID로 직접 갱신한 뒤 `./scripts/lab.sh acknowledge agent-setup`을 실행해야 합니다. +중요: predown hook은 `azd down`이 삭제 **확인** 프롬프트를 띄우기 **전에** 실행됩니다. 그 프롬프트에서 **취소**해도 이미 제거된 Monitoring Contributor 할당은 돌아오지 않습니다. 리소스 그룹은 남지만 Agent의 구독 범위 권한은 사라진 상태이므로, 계속 쓰려면 역할 할당을 다시 만들고 `evidence/agent-setup.json`을 새 할당 ID로 직접 갱신한 뒤 `python3 scripts/lab_state.py acknowledge-agent`를 실행해야 합니다. hook이 실패해 손으로 다시 실행할 때는 아래를 직접 호출합니다. `--yes` 없이는 계획만 출력합니다. @@ -175,11 +165,19 @@ hook이 실패해 손으로 다시 실행할 때는 아래를 직접 호출합 ./scripts/cleanup-external.sh --reset-image-env --yes ``` -azd 환경을 잃어버린 실습을 정리할 때만 `./scripts/cleanup.sh --legacy-delete-resource-group`으로 예전 삭제 경로를 씁니다. 이 경로도 구독 일치와 태그 확인을 거치며, 첫 명령은 dry-run입니다. +azd 환경을 잃어버려 `azd down`을 쓸 수 없다면 리소스 그룹을 직접 지웁니다. 반드시 태그로 대상을 먼저 확인하세요. 이 실습이 만든 그룹만 두 태그를 모두 가집니다. + +```bash +az group list --subscription "${SUBSCRIPTION_ID}" \ + --query "[?tags.purpose=='sre-agent-event-lab'].{name:name, env:tags.\"azd-env-name\"}" \ + --output table + +az group delete --subscription "${SUBSCRIPTION_ID}" --name "${RESOURCE_GROUP}" --yes +``` ## 문제 해결 -먼저 `./scripts/lab.sh doctor`를 실행해 어떤 검사가 `FAIL`인지 확인하세요. 각 명령의 실패 처리와 복구 절차는 해당 단계 문서에 있습니다. +각 명령의 실패 처리와 복구 절차는 해당 단계 문서에 있습니다. | 증상 | 확인할 곳 | |---|---| diff --git a/monitor/sre-agent-event-lab/guides/01-agent-setup.md b/monitor/sre-agent-event-lab/guides/01-agent-setup.md index 4a7ab13..9515c96 100644 --- a/monitor/sre-agent-event-lab/guides/01-agent-setup.md +++ b/monitor/sre-agent-event-lab/guides/01-agent-setup.md @@ -182,26 +182,56 @@ jq -e \ } ``` -## 점검과 승인 +## 정상 상태 확인과 승인 + +장애를 주입하기 전에, 정상 상태의 요청이 Application Insights까지 도달하는지 확인합니다. 이 확인이 통과해야 S1을 시작할 수 있습니다. 텔레메트리가 도착하지 않는 워크로드에 장애를 넣으면, 주입한 장애와 원래부터 안 보이던 상태를 구별할 수 없기 때문입니다. + +먼저 정상 부하를 넣습니다. 두 엔드포인트 모두 200이어야 합니다. + +```bash +EVIDENCE_DIR="${PWD}/evidence/baseline-$(date -u +%Y%m%dT%H%M%SZ)" +mkdir -p "${EVIDENCE_DIR}" + +python3 scripts/loadgen.py "https://${APP_FQDN}/api/orders" \ + --requests 30 --concurrency 4 --expect-status 200 \ + --output "${EVIDENCE_DIR}/orders.json" + +python3 scripts/loadgen.py "https://${APP_FQDN}/api/documents" \ + --requests 10 --concurrency 2 --expect-status 200 \ + --output "${EVIDENCE_DIR}/documents.json" +``` + +두 요청 종류가 워크스페이스에 보이는지 확인합니다. 수집에는 보통 2~5분이 걸리므로, 결과가 비어 있으면 잠시 뒤 다시 실행합니다. ```bash -./scripts/lab.sh doctor -./scripts/lab.sh baseline -./scripts/lab.sh acknowledge agent-setup +az monitor log-analytics query \ + --workspace "${WORKSPACE_CUSTOMER_ID}" \ + --analytics-query "AppRequests | where AppRoleName == '${TELEMETRY_SERVICE_NAME}' | where TimeGenerated > ago(30m) | summarize count() by Name" \ + --output table ``` -`doctor`가 출력하는 네 줄은 언제나 `MANUAL`입니다. 저장소 연결, 지식 원본, incident platform, 응답 계획을 읽을 수 있는 공식 안정 API가 없기 때문입니다. 나머지 검사에 `FAIL`이 남아 있으면 먼저 해결합니다. +`/api/orders`와 `/api/documents`가 모두 보이면 통과입니다. 그 사실을 기록해야 S1이 열립니다. + +```bash +python3 scripts/lab_state.py mark baseline_passed --evidence-dir "${EVIDENCE_DIR}" +``` + +마지막으로 Agent 설정을 승인합니다. 이 명령은 설정 값을 출력한 뒤 표준 입력으로 정확히 `acknowledge`를 받아야 기록합니다. 어떤 환경 변수로도 대체할 수 없습니다. 값이 하나라도 다르면 그대로 중단하고 위 단계로 돌아가세요. + +```bash +python3 scripts/lab_state.py acknowledge-agent +``` -`acknowledge agent-setup`은 설정 값을 출력한 뒤 표준 입력으로 정확히 `acknowledge`를 받아야 기록합니다. 어떤 환경 변수로도 대체할 수 없습니다. 값이 하나라도 다르면 그대로 중단하고 위 단계로 돌아가세요. +저장소 연결, 지식 원본, incident platform, 응답 계획은 읽을 수 있는 공식 안정 API가 없으므로 위 표를 보고 포털에서 직접 확인하는 것이 유일한 방법입니다. ## 실패했을 때 | 증상 | 조치 | |---|---| -| `doctor`의 `Python environment` 검사가 `FAIL` | `app/.venv`가 없거나 Pillow가 안 잡힙니다. 로컬 문제이며 클라우드 배포와는 무관하니 바로 실행: `./scripts/setup-venv.sh` | -| `doctor`의 Reader 검사가 `FAIL` | 두 principal ID가 근거 파일과 같은지 확인하고 리소스 그룹에 Reader를 다시 부여합니다 | -| `baseline`이 telemetry 없음으로 종료 | 10분 더 기다린 뒤 다시 실행합니다. 계속 실패하면 `azd env get-value AZURE_CONTAINER_APP_FQDN`으로 앱을 직접 호출해 봅니다 | -| `acknowledge`가 기록되지 않음 | 입력한 단어가 정확한지, `azd env select`로 올바른 환경을 골랐는지 확인합니다 | +| `app/.venv`가 없음 (다음 문서의 캡처 단계가 Pillow를 씁니다) | `azd up`의 postprovision 단계가 만들어 둡니다. 없거나 깨졌다면 로컬 문제이므로 바로 실행: `./scripts/setup-venv.sh` | +| 두 principal ID로 Reader 권한이 확인되지 않음 | 두 ID가 근거 파일과 같은지 확인하고 리소스 그룹에 Reader를 다시 부여합니다 | +| 쿼리에 요청이 보이지 않음 | 10분 더 기다린 뒤 다시 조회합니다. 계속 비어 있으면 `curl -sS -o /dev/null -w '%{http_code}\n' "https://${APP_FQDN}/healthz"`로 앱을 직접 호출해 봅니다 | +| `acknowledge-agent`가 기록되지 않음 | 입력한 단어가 정확한지, `azd env select`로 올바른 환경을 골랐는지 확인합니다 | | 경고가 Agent에 도착하지 않음 | Monitoring Contributor 범위가 구독인지, 응답 계획이 `On`인지 확인합니다 | ## 다음 단계 diff --git a/monitor/sre-agent-event-lab/guides/02-scenario-s1.md b/monitor/sre-agent-event-lab/guides/02-scenario-s1.md index 54501c7..a3eeea9 100644 --- a/monitor/sre-agent-event-lab/guides/02-scenario-s1.md +++ b/monitor/sre-agent-event-lab/guides/02-scenario-s1.md @@ -7,7 +7,9 @@ - [01-agent-setup.md](01-agent-setup.md)를 마쳤고 `evidence/state.json`에 `baseline_passed`와 `agent_setup_acknowledged`가 기록되어 있습니다. - 현재 활성 구독이 azd 환경의 구독과 같습니다. -이 두 가지는 `evidence/state.json`을 통해 강제됩니다. 아래 "수동 실행"도 첫 단계에서 `lab_state.py begin-run`을 호출하므로 같은 순서·중복 실행 게이트를 그대로 적용받습니다. 차이는 실패했을 때입니다. 지름길은 종료 트랩이 장애를 자동으로 되돌리지만, 수동 실행에서는 복구 명령을 운영자가 직접 실행해야 합니다. +이 두 가지는 `evidence/state.json`을 통해 강제됩니다. 아래 "수동 실행"의 첫 단계인 `lab_state.py begin-run`이 순서와 중복 실행을 함께 확인합니다. + +어떤 시나리오든 실행이 `running`이나 `failed`로 남아 있으면 세 시나리오 모두 새 실행이 거부됩니다. 세 시나리오는 같은 Container App 하나를 쓰고, 끝나지 않은 실행은 장애가 아직 살아 있을 수 있는 상태이기 때문입니다. `failed`는 그 시나리오를 다시 실행하면 풀리고, `running`은 실행이 끝나기를 기다리거나 `python3 scripts/lab_state.py mark-failed s1`처럼 끝난 방식을 기록해야 풀립니다. 장애를 되돌리는 일은 운영자가 직접 해야 하므로, 중간에 멈췄다면 복구 명령까지 마치고 결과를 기록하세요. 조건이 하나라도 없으면 실행이 시작 전에 거부되고 무엇을 먼저 하라는 안내가 출력됩니다. @@ -236,12 +238,14 @@ fi 복구는 워크로드 정상화(`RECOVERY_OK`)와 경고 해제가 모두 확인될 때만 인정합니다. 둘 중 하나라도 어긋나면 실패로 기록되므로, 장애가 남아 있는 실행이 성공으로 채점되지 않습니다. -`timeline.json`은 캡처 도구가 대상 경고를 찾는 파일이며, 이 파일이 없으면 지름길의 `lab.sh capture`도 진행하지 못합니다. +`timeline.json`은 캡처 도구가 대상 경고를 찾는 파일이므로, 이 파일이 없으면 다음 캡처 단계가 진행되지 못합니다. ### 6. 조사 근거 수집 Agent 스레드를 내려받아 정규화하고, 관측 결과를 먼저 기록한 뒤 그림을 만듭니다. +이 단계부터는 `app/.venv`의 인터프리터를 씁니다. 렌더링에 Pillow가 필요하기 때문입니다. `azd up`의 postprovision 단계가 만들어 두므로 보통 그대로 있고, 없다면 `./scripts/setup-venv.sh`를 실행한 뒤 이어서 진행합니다. + ```bash AGENT_ENDPOINT="$(jq -r '.agent_endpoint // empty' evidence/agent-setup.json)" case "${AGENT_ENDPOINT}" in @@ -287,19 +291,6 @@ if (( LAB_READY )) && [[ -n "${INJECTED_AT}" ]]; then fi ``` -## 지름길 - -위 여섯 단계를 그대로 자동화한 것이 아래 두 명령입니다. 같은 Azure 호출을 같은 순서로 실행하고, 추가로 실행 상태 기록과 종료 시 자동 복구를 처리합니다. 무엇이 실행되는지 이미 알고 반복할 때만 쓰세요. - -```bash -./scripts/lab.sh run s1 -./scripts/lab.sh capture s1 -``` - -`run`은 장애 주입 → 부하 → 경고 대기 → 복구 → 타임라인 저장까지 진행하고, 경고가 발생하지 않으면 12분 뒤 실패로 기록합니다. 중간에 끊어도 종료 트랩이 복구를 시도합니다. `capture`는 대상 디렉터리를 `evidence/state.json`에서 찾으므로 경로를 입력하지 않습니다. - -지름길은 순서 게이트도 적용합니다. 어떤 시나리오든 실행이 `running`이나 `failed`로 남아 있으면 세 시나리오 모두 새 실행이 거부됩니다. 세 시나리오는 같은 Container App 하나를 쓰고, 끝나지 않은 실행은 장애가 아직 살아 있을 수 있는 상태이기 때문입니다. `failed`는 그 시나리오를 다시 실행하면 풀리고, `running`은 실행이 끝나기를 기다리거나 `python3 scripts/lab_state.py mark-failed s1`처럼 끝난 방식을 기록해야 풀립니다. 수동 실행도 `begin-run`으로 같은 게이트를 적용받지만, 장애를 되돌리는 일은 운영자가 직접 해야 합니다. - ## Azure에서 발생하는 변화 | 순서 | 변화 | @@ -344,7 +335,7 @@ fi 1. Container App의 활성 revision이 다시 정상입니다. 2. 이 실행이 발생시킨 경고가 Azure Monitor에서 `Resolved`가 됩니다. -경고 해제는 최대 25분, 워크로드 정상화는 최대 10분까지 기다립니다. 1분 주기의 stateful log alert는 실패 요청이 5분 조회 창에서 빠진 뒤에도 조건이 10분간 불충족이어야 `Resolved`가 되므로 여유 시간을 포함합니다. 둘 중 하나라도 시간 안에 확인되지 않으면 실행은 실패로 기록되고 S2는 계속 막힙니다. 실패한 실행은 원인을 고친 뒤 `./scripts/lab.sh run s1`을 다시 실행하면 새 시도로 이어집니다. 다시 실행하는 순간 이전 시도의 `s1_recovered`와 `s1_captured` 기록은 장애를 주입하기 전에 지워지므로, 새 시도가 복구되고 `capture`까지 끝날 때까지 S2는 다시 막힙니다. 이미 성공한 시나리오를 한 번 더 돌릴 때도 같습니다. +경고 해제는 최대 25분, 워크로드 정상화는 최대 10분까지 기다립니다. 1분 주기의 stateful log alert는 실패 요청이 5분 조회 창에서 빠진 뒤에도 조건이 10분간 불충족이어야 `Resolved`가 되므로 여유 시간을 포함합니다. 둘 중 하나라도 시간 안에 확인되지 않으면 실행은 실패로 기록되고 S2는 계속 막힙니다. 실패한 실행은 원인을 고친 뒤 이 절을 `begin-run`부터 다시 실행하면 새 시도로 이어집니다. `begin-run`을 다시 호출하는 순간 이전 시도의 `s1_recovered`와 `s1_captured` 기록은 장애를 주입하기 전에 지워지므로, 새 시도가 복구되고 `record-capture`까지 끝날 때까지 S2는 다시 막힙니다. 이미 성공한 시나리오를 한 번 더 돌릴 때도 같습니다. 되돌리기 자체가 실패하면(예: `az containerapp update` 거부, 새 revision이 준비되지 않음) 스크립트는 `CRITICAL:` 두 줄을 출력하고 0이 아닌 코드로 끝냅니다. 주입한 장애가 그대로 남아 있다는 뜻이므로, 다음 시나리오를 실행하기 전에 `FAILURE_MODE=none`을 수동으로 되돌리고 revision이 정상인지 확인하세요. diff --git a/monitor/sre-agent-event-lab/guides/03-scenario-s2.md b/monitor/sre-agent-event-lab/guides/03-scenario-s2.md index b30045c..f240444 100644 --- a/monitor/sre-agent-event-lab/guides/03-scenario-s2.md +++ b/monitor/sre-agent-event-lab/guides/03-scenario-s2.md @@ -261,15 +261,6 @@ fi `agent_endpoint`는 `https://`로 시작하고 자리표시자 괄호가 없어야 합니다. `http://`를 쓰면 데이터 평면 토큰이 평문으로 나갑니다. `capture_agent.py`는 제한 시간까지 결론을 받지 못하면 종료 코드 3으로 끝나며, 이는 "결론 없음"을 그대로 기록하는 정상 경로입니다. 결과 기록을 렌더링보다 먼저 하는 이유는 이미지 생성이 실패해도 관측한 결과를 잃지 않기 위해서입니다. -## 지름길 - -```bash -./scripts/lab.sh run s2 -./scripts/lab.sh capture s2 -``` - -같은 명령을 같은 순서로 실행하면서 종료 시 자동 복구까지 처리합니다. 순서 게이트는 수동 실행도 `begin-run`을 통해 동일하게 적용받으며, 차이는 실패 시 복구를 누가 하느냐입니다. - ## Azure에서 발생하는 변화 | 순서 | 변화 | @@ -303,7 +294,7 @@ fi 1. 활성 revision이 정상이고 `/api/orders`가 다시 빠르게 응답합니다. 2. `alert-sre-lab-s2-latency`가 `Resolved`입니다. -경고가 해제되지 않으면 부하가 남아 있는지, 새 revision으로 트래픽이 100% 넘어갔는지 확인합니다. 실패로 기록된 실행은 `./scripts/lab.sh run s2`를 다시 실행해 새 시도로 이어 갑니다. 다시 실행하면 이전 시도의 `s2_recovered`와 `s2_captured` 기록이 주입 전에 지워지고, 새 시도가 복구되고 `capture`될 때까지 S3는 다시 막힙니다. +경고가 해제되지 않으면 부하가 남아 있는지, 새 revision으로 트래픽이 100% 넘어갔는지 확인합니다. 실패로 기록된 실행은 이 절을 `begin-run`부터 다시 실행해 새 시도로 이어 갑니다. 다시 실행하면 이전 시도의 `s2_recovered`와 `s2_captured` 기록이 주입 전에 지워지고, 새 시도가 복구되고 `record-capture`될 때까지 S3는 다시 막힙니다. 되돌리기 자체가 실패하면 스크립트는 `CRITICAL:` 두 줄을 출력하고 0이 아닌 코드로 끝냅니다. 지연이 그대로 남아 있다는 뜻이므로, 다음 시나리오를 실행하기 전에 `ORDER_DELAY_MS=0`을 수동으로 되돌리고 새 revision이 정상인지 확인하세요. diff --git a/monitor/sre-agent-event-lab/guides/04-scenario-s3.md b/monitor/sre-agent-event-lab/guides/04-scenario-s3.md index 45ee6b7..b71264d 100644 --- a/monitor/sre-agent-event-lab/guides/04-scenario-s3.md +++ b/monitor/sre-agent-event-lab/guides/04-scenario-s3.md @@ -234,15 +234,6 @@ fi `agent_endpoint`는 `https://`로 시작하고 자리표시자 괄호가 없어야 합니다. `http://`를 쓰면 데이터 평면 토큰이 평문으로 나갑니다. `capture_agent.py`는 제한 시간까지 결론을 받지 못하면 종료 코드 3으로 끝나며, 이는 "결론 없음"을 그대로 기록하는 정상 경로입니다. 결과 기록을 렌더링보다 먼저 하는 이유는 이미지 생성이 실패해도 관측한 결과를 잃지 않기 위해서입니다. -## 지름길 - -```bash -./scripts/lab.sh run s3 -./scripts/lab.sh capture s3 -``` - -`run`은 위 순서에 더해 필요한 배포 출력이 비어 있지 않은지 삭제 전에 확인하고, 전파 대기를 최대 5분으로 제한하며, 종료 시 역할 복구를 보장합니다. 수동 실행에서는 이 확인들이 없으므로 삭제 후 복구 명령까지 반드시 직접 완료해야 합니다. - ## Azure에서 발생하는 변화 | 순서 | 변화 | @@ -283,7 +274,7 @@ fi 1. `Storage Blob Data Reader` 할당이 원래 Blob 컨테이너 범위에 다시 존재합니다. 2. `alert-sre-lab-s3-storage-rbac`가 `Resolved`입니다. -역할 전파에는 몇 분이 걸릴 수 있습니다. `/api/documents`가 200을 돌려주는지 직접 호출해 확인하고, 실패로 기록되었다면 `./scripts/lab.sh run s3`으로 새 시도를 시작합니다. 다시 실행하면 이전 시도의 `s3_recovered`와 `s3_captured` 기록이 주입 전에 지워지므로, 새 시도가 복구되고 `capture`될 때까지 채점은 막힙니다. +역할 전파에는 몇 분이 걸릴 수 있습니다. `/api/documents`가 200을 돌려주는지 직접 호출해 확인하고, 실패로 기록되었다면 이 절을 `begin-run`부터 다시 실행해 새 시도를 시작합니다. 다시 실행하면 이전 시도의 `s3_recovered`와 `s3_captured` 기록이 주입 전에 지워지므로, 새 시도가 복구되고 `record-capture`될 때까지 채점은 막힙니다. 역할 복구 자체가 실패하면 스크립트는 `CRITICAL:` 두 줄을 출력하고 0이 아닌 코드로 끝냅니다. 워크로드에 Blob 권한이 없는 상태가 그대로 남으므로, 같은 이름·같은 범위의 `Storage Blob Data Reader` 할당을 수동으로 다시 만든 뒤 다음 단계로 넘어가세요. diff --git a/monitor/sre-agent-event-lab/guides/05-results.md b/monitor/sre-agent-event-lab/guides/05-results.md index 743dae5..9acb66c 100644 --- a/monitor/sre-agent-event-lab/guides/05-results.md +++ b/monitor/sre-agent-event-lab/guides/05-results.md @@ -12,7 +12,7 @@ ```bash cd monitor/sre-agent-event-lab -./scripts/lab.sh score +app/.venv/bin/python scripts/score.py --evidence-root evidence ``` `evidence/scorecard.json`과 `SCENARIOCRITERIONSTATUSPOINTSDETAIL` 표를 출력합니다. 종합 판정이 `FAIL`일 때만 종료 코드 1을 반환합니다. diff --git a/monitor/sre-agent-event-lab/infra/main.bicep b/monitor/sre-agent-event-lab/infra/main.bicep index 27a99d6..7e3f204 100644 --- a/monitor/sre-agent-event-lab/infra/main.bicep +++ b/monitor/sre-agent-event-lab/infra/main.bicep @@ -83,6 +83,7 @@ output AZURE_APP_INSIGHTS_NAME string = lab.outputs.appInsightsName output AZURE_STORAGE_CONTAINER_SCOPE string = lab.outputs.storageContainerScope output AZURE_BLOB_ROLE_ASSIGNMENT_NAME string = lab.outputs.blobRoleAssignmentName output AZURE_TELEMETRY_SERVICE_NAME string = lab.outputs.telemetryServiceName +output AZURE_WORKSPACE_CUSTOMER_ID string = lab.outputs.workspaceCustomerId // Deployment outputs the lab scripts (common.sh `deployment_output`, // run-scenario.sh, query-evidence.sh) still read by their original names. diff --git a/monitor/sre-agent-event-lab/scripts/baseline.sh b/monitor/sre-agent-event-lab/scripts/baseline.sh deleted file mode 100755 index a5c4956..0000000 --- a/monitor/sre-agent-event-lab/scripts/baseline.sh +++ /dev/null @@ -1,103 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P)" -source "${SCRIPT_DIR}/common.sh" - -# Overridable only for tests: production runs always use the 600s/20s -# defaults. Bounding the poll (rather than querying once) tolerates -# Application Insights ingestion lag without hanging forever. -readonly TELEMETRY_TIMEOUT_SECONDS="${LAB_BASELINE_TELEMETRY_TIMEOUT_SECONDS:-600}" -readonly TELEMETRY_POLL_INTERVAL_SECONDS="${LAB_BASELINE_TELEMETRY_POLL_INTERVAL_SECONDS:-20}" - -require_lab_config -verify_subscription -verify_lab_resource_group - -APP_FQDN="$(deployment_output containerAppFqdn)" -WORKSPACE_CUSTOMER_ID="$(deployment_output workspaceCustomerId)" -TELEMETRY_SERVICE_NAME="$(deployment_output telemetryServiceName)" -readonly APP_FQDN WORKSPACE_CUSTOMER_ID TELEMETRY_SERVICE_NAME - -if [[ -z "${APP_FQDN}" || -z "${WORKSPACE_CUSTOMER_ID}" || -z "${TELEMETRY_SERVICE_NAME}" ]]; then - echo "Deployment outputs are missing. Run: azd provision" >&2 - exit 1 -fi - -EVIDENCE_DIR="$(create_evidence_dir baseline)" -readonly EVIDENCE_DIR - -if ! python3 "${SCRIPT_DIR}/loadgen.py" \ - "https://${APP_FQDN}/api/orders" \ - --requests 30 \ - --concurrency 4 \ - --expect-status 200 \ - --output "${EVIDENCE_DIR}/orders.json"; then - echo "Baseline /api/orders requests did not all succeed: ${EVIDENCE_DIR}/orders.json" >&2 - exit 1 -fi - -if ! python3 "${SCRIPT_DIR}/loadgen.py" \ - "https://${APP_FQDN}/api/documents" \ - --requests 10 \ - --concurrency 2 \ - --expect-status 200 \ - --output "${EVIDENCE_DIR}/documents.json"; then - echo "Baseline /api/documents requests did not all succeed: ${EVIDENCE_DIR}/documents.json" >&2 - exit 1 -fi - -# The workspace answers with a flat JSON array of row objects (see -# `log_analytics_row_count` in common.sh), so the poll asks for real rows -- -# projected and bounded with `take 1` -- rather than a `| count`, whose -# single row would look like data even for an empty workspace. `contains` -# (substring) is used rather than `has` (term match) so a path like -# `/api/orders` matches inside `GET /api/orders` regardless of tokenization. -telemetry_seen() { - local path_fragment="$1" - local rows - rows="$(log_analytics_row_count "${WORKSPACE_CUSTOMER_ID}" PT30M \ - "AppRequests | where AppRoleName == '${TELEMETRY_SERVICE_NAME}' | where Name contains '${path_fragment}' | project TimeGenerated, Name | take 1")" - [[ "${rows:-0}" -gt 0 ]] -} - -ORDERS_SEEN=0 -DOCUMENTS_SEEN=0 -# Bounded poll: always at least one honest attempt (even with a zero -# timeout), and never a sleep that would run past the deadline. -TELEMETRY_DEADLINE=$(( SECONDS + TELEMETRY_TIMEOUT_SECONDS )) -while :; do - if [[ "${ORDERS_SEEN}" -eq 0 ]] && telemetry_seen "/api/orders"; then - ORDERS_SEEN=1 - fi - if [[ "${DOCUMENTS_SEEN}" -eq 0 ]] && telemetry_seen "/api/documents"; then - DOCUMENTS_SEEN=1 - fi - if [[ "${ORDERS_SEEN}" -eq 1 && "${DOCUMENTS_SEEN}" -eq 1 ]]; then - break - fi - if (( SECONDS + TELEMETRY_POLL_INTERVAL_SECONDS > TELEMETRY_DEADLINE )); then - break - fi - sleep "${TELEMETRY_POLL_INTERVAL_SECONDS}" -done - -jq -n \ - --argjson ordersSeen "${ORDERS_SEEN}" \ - --argjson documentsSeen "${DOCUMENTS_SEEN}" \ - --arg checkedAt "$(utc_now)" \ - '{orders_telemetry_seen: ($ordersSeen == 1), documents_telemetry_seen: ($documentsSeen == 1), checked_at: $checkedAt}' \ - >"${EVIDENCE_DIR}/telemetry-check.json" - -if [[ "${ORDERS_SEEN}" -ne 1 || "${DOCUMENTS_SEEN}" -ne 1 ]]; then - echo "Application Insights did not show both request types within ${TELEMETRY_TIMEOUT_SECONDS}s (orders=${ORDERS_SEEN} documents=${DOCUMENTS_SEEN})." >&2 - echo "Evidence directory: ${EVIDENCE_DIR}" >&2 - exit 1 -fi - -# Only a baseline that really produced both request types unlocks S1: a -# scenario run against a workload whose telemetry never arrived cannot be -# told apart from the failure it is supposed to inject. -lab_state mark baseline_passed --evidence-dir "${EVIDENCE_DIR}" - -echo "Evidence directory: ${EVIDENCE_DIR}" diff --git a/monitor/sre-agent-event-lab/scripts/capture-scenario.sh b/monitor/sre-agent-event-lab/scripts/capture-scenario.sh deleted file mode 100755 index 29c6faa..0000000 --- a/monitor/sre-agent-event-lab/scripts/capture-scenario.sh +++ /dev/null @@ -1,137 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P)" -source "${SCRIPT_DIR}/common.sh" - -if [[ "$#" -ne 1 ]] || [[ ! "$1" =~ ^s[123]$ ]]; then - echo "Usage: $0 s1|s2|s3" >&2 - exit 2 -fi - -readonly SCENARIO="$1" -readonly ASSET_DIR="${LAB_ROOT}/assets/captures/${SCENARIO}" -readonly PYTHON="${LAB_ROOT}/app/.venv/bin/python" - -require_lab_config -verify_subscription -verify_lab_resource_group - -# The public command is `lab.sh capture s1`, with no timestamped path: the -# directory this scenario's run recorded in `evidence/state.json` is the -# only one whose timeline belongs to the alert being captured. No override -# is accepted -- capturing an arbitrary directory would record its outcome -# as this environment's current capture status and could unblock the next -# scenario on evidence from another run. Re-rendering an archived run is a -# read-only job for the lower-level tools instead, which touch no state: -# -# app/.venv/bin/python scripts/capture_agent.py --output-dir ... -# app/.venv/bin/python scripts/render_capture.py /normalized-timeline.json --scenario s1 -if ! EVIDENCE_DIR="$(lab_state evidence-dir "${SCENARIO}")"; then - exit 1 -fi -readonly EVIDENCE_DIR -readonly TIMELINE_FILE="${EVIDENCE_DIR}/timeline.json" -readonly NORMALIZED_FILE="${EVIDENCE_DIR}/normalized-timeline.json" - -CURRENT_CAPTURE_STATE="$(lab_state capture-status "${SCENARIO}")" -if [[ "${CURRENT_CAPTURE_STATE}" == "conclusion" ]]; then - echo "${SCENARIO} already has a conclusion for this run." >&2 - echo "To collect another capture, start a new attempt first: ./scripts/lab.sh run ${SCENARIO}" >&2 - echo "To regenerate only the rendered assets, run: ${PYTHON} ${SCRIPT_DIR}/render_capture.py ${NORMALIZED_FILE} ${ASSET_DIR} --scenario ${SCENARIO}" >&2 - exit 1 -fi - -if [[ ! -f "${AGENT_SETUP_FILE}" ]]; then - echo "Missing Agent setup evidence: ${AGENT_SETUP_FILE}" >&2 - echo "Create evidence/agent-setup.json as shown in guides/01-agent-setup.md." >&2 - exit 1 -fi -if [[ ! -f "${TIMELINE_FILE}" ]]; then - echo "Missing scenario timeline: ${TIMELINE_FILE}" >&2 - exit 1 -fi -if [[ ! -x "${PYTHON}" ]]; then - echo "Missing Python environment: ${PYTHON}" >&2 - echo "Cloud resources are already deployed; only this local step needs to be retried. Re-run: ./scripts/setup-venv.sh" >&2 - exit 1 -fi -if ! "${PYTHON}" -c "import PIL" >/dev/null 2>&1; then - echo "Python environment at ${PYTHON} is missing Pillow (PIL), which render_capture.py needs." >&2 - echo "Cloud resources are already deployed; only this local step needs to be retried. Re-run: ./scripts/setup-venv.sh" >&2 - exit 1 -fi - -AGENT_ENDPOINT="$(jq -r '.agent_endpoint // empty' "${AGENT_SETUP_FILE}")" -ALERT_ID="$(jq -r '.alert_id // empty' "${TIMELINE_FILE}")" -if [[ -z "${AGENT_ENDPOINT}" || "${AGENT_ENDPOINT}" != https://* || "${AGENT_ENDPOINT}" == *"<"* || "${AGENT_ENDPOINT}" == *">"* ]]; then - echo "agent_endpoint is missing or is not a valid HTTPS endpoint in ${AGENT_SETUP_FILE}." >&2 - echo "Update evidence/agent-setup.json as shown in guides/01-agent-setup.md." >&2 - exit 1 -fi -if [[ -z "${ALERT_ID}" ]]; then - echo "alert_id is missing from ${TIMELINE_FILE}." >&2 - echo "Create a new scenario timeline by running: ./scripts/lab.sh run ${SCENARIO}" >&2 - exit 1 -fi -readonly AGENT_ENDPOINT ALERT_ID - -set +e -"${PYTHON}" "${SCRIPT_DIR}/capture_agent.py" \ - --scenario "${SCENARIO}" \ - --alert-id "${ALERT_ID}" \ - --endpoint "${AGENT_ENDPOINT}" \ - --output-dir "${EVIDENCE_DIR}" \ - --timeout 1200 \ - --interval 15 -capture_status="$?" -set -e -if [[ "${capture_status}" -ne 0 && "${capture_status}" -ne 3 ]]; then - echo "Evidence collection failed with status ${capture_status}." >&2 - exit "${capture_status}" -fi - -# Recorded before any further check can abort the script: the terminal -# state of the normalized timeline is the honest outcome of this capture, -# including `thread-not-created`, `investigation-missing` and -# `conclusion-missing`. Only a real `conclusion` counts as a successful -# capture and unblocks the next scenario -- and `lab_state.py` refuses even -# that when the run it belongs to did not recover, because a conclusion -# collected against an unresolved incident is indistinguishable from a real -# one once it is on disk. That refusal must not look like a crash: the -# capture pipeline has already written real files, and they stay. -if ! CAPTURE_STATE="$(lab_state record-capture "${SCENARIO}" \ - --timeline "${NORMALIZED_FILE}" \ - --evidence-dir "${EVIDENCE_DIR}")"; then - echo "The capture was collected but not recorded." >&2 - echo "Raw evidence is on disk and unchanged: ${EVIDENCE_DIR}" >&2 - exit 1 -fi -readonly CAPTURE_STATE - -event_count="$(jq 'length' "${NORMALIZED_FILE}")" -if (( event_count < 4 )); then - echo "Capture has fewer than four explicit states: ${event_count}" >&2 - exit 1 -fi - -mkdir -p "${ASSET_DIR}" -find "${ASSET_DIR}" -maxdepth 1 -type f \ - \( -name '*.png' -o -name 'investigation.gif' -o -name 'timeline.mmd' \ - -o -name 'timeline.md' \) -delete - -"${PYTHON}" "${SCRIPT_DIR}/render_capture.py" \ - "${NORMALIZED_FILE}" \ - "${ASSET_DIR}" \ - --scenario "${SCENARIO}" - -echo "Raw evidence: ${EVIDENCE_DIR}" -echo "Rendered capture: ${ASSET_DIR}/investigation.gif" -echo "Capture status: ${CAPTURE_STATE}" -if [[ "${CAPTURE_STATE}" != "conclusion" ]]; then - echo "The Agent produced no conclusion for ${SCENARIO} (${CAPTURE_STATE})." - echo "Recorded as-is; the next scenario stays blocked until a capture ends in a conclusion." -fi -if [[ "${capture_status}" -eq 3 ]]; then - echo "Capture ended at the deadline; missing states are explicit in the output." -fi diff --git a/monitor/sre-agent-event-lab/scripts/cleanup-external.sh b/monitor/sre-agent-event-lab/scripts/cleanup-external.sh index 8e5ff64..389898a 100755 --- a/monitor/sre-agent-event-lab/scripts/cleanup-external.sh +++ b/monitor/sre-agent-event-lab/scripts/cleanup-external.sh @@ -160,7 +160,7 @@ if ! RECORDED_ASSIGNMENTS="$(jq -r ' | join("|") ' "${CLEANUP_SETUP_FILE}" 2>/dev/null)"; then echo "Agent setup evidence is not valid JSON: ${CLEANUP_SETUP_FILE}" >&2 - echo "Rewrite evidence/agent-setup.json as shown in guides/01-agent-setup.md, then run: lab.sh acknowledge agent-setup" >&2 + echo "Rewrite evidence/agent-setup.json as shown in guides/01-agent-setup.md, then run: python3 scripts/lab_state.py acknowledge-agent" >&2 exit 1 fi readonly RECORDED_ASSIGNMENTS @@ -260,12 +260,12 @@ while IFS='|' read -r assignment_key assignment_id principal_key expected_princi echo "${assignment_key} is empty for principal ${expected_principal_id:-unknown}." >&2 echo "Find the live ID: az role assignment list --assignee-object-id ${expected_principal_id:-} --role \"Monitoring Contributor\" --scope ${SUBSCRIPTION_SCOPE} --query \"[0].id\" -o tsv" >&2 echo "If no assignment is returned and teardown should continue, clear both ${assignment_key} and ${principal_key} in the evidence file." >&2 - echo "Update evidence/agent-setup.json as shown in guides/01-agent-setup.md, then run: lab.sh acknowledge agent-setup" >&2 + echo "Update evidence/agent-setup.json as shown in guides/01-agent-setup.md, then run: python3 scripts/lab_state.py acknowledge-agent" >&2 exit 1 fi if [[ -z "${expected_principal_id}" ]]; then echo "Incomplete Agent setup evidence: ${assignment_id} was recorded without ${principal_key}." >&2 - echo "Update evidence/agent-setup.json as shown in guides/01-agent-setup.md, then run: lab.sh acknowledge agent-setup" >&2 + echo "Update evidence/agent-setup.json as shown in guides/01-agent-setup.md, then run: python3 scripts/lab_state.py acknowledge-agent" >&2 exit 1 fi case "${VERIFIED_ASSIGNMENT_IDS}" in diff --git a/monitor/sre-agent-event-lab/scripts/cleanup.sh b/monitor/sre-agent-event-lab/scripts/cleanup.sh deleted file mode 100755 index 60e3ee5..0000000 --- a/monitor/sre-agent-event-lab/scripts/cleanup.sh +++ /dev/null @@ -1,83 +0,0 @@ -#!/usr/bin/env bash -# Compatibility wrapper for the documented cleanup command. -# -# The lab is provisioned with azd, so `azd down --purge` is what tears it -# down: azd deletes the resource group it created, and its predown/postdown -# hooks run `cleanup-external.sh` for the two things azd cannot see (the -# recorded subscription-scoped role assignments, and the image values the -# postdeploy hook stored). This script therefore only forwards to -# `cleanup-external.sh` -- it never deletes a resource group of its own, -# because a broad deletion here would also take resources azd did not -# create and knows nothing about. -# -# `--legacy-delete-resource-group` keeps the pre-azd recovery path -# available: a lab whose azd environment was lost still has to be -# deletable by hand. It runs the same subscription and tag checks the old -# script ran, so it can only ever delete a resource group tagged -# `purpose=sre-agent-event-lab` for the current azd environment. -set -euo pipefail - -SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P)" -source "${SCRIPT_DIR}/common.sh" - -CONFIRMED=0 -LEGACY_RESOURCE_GROUP=0 -EXTERNAL_ARGS=() -while [[ "$#" -gt 0 ]]; do - case "$1" in - --yes) - CONFIRMED=1 - EXTERNAL_ARGS+=("--yes") - ;; - --legacy-delete-resource-group) LEGACY_RESOURCE_GROUP=1 ;; - *) - echo "Usage: $0 [--yes] [--legacy-delete-resource-group]" >&2 - exit 2 - ;; - esac - shift -done -readonly CONFIRMED LEGACY_RESOURCE_GROUP - -echo "Use 'azd down --purge' for complete cleanup." - -external_cleanup() { - # Bash 3.2 aborts on "${ARRAY[@]}" for an empty array under `set -u`. - if (( ${#EXTERNAL_ARGS[@]} > 0 )); then - "${SCRIPT_DIR}/cleanup-external.sh" "${EXTERNAL_ARGS[@]}" - else - "${SCRIPT_DIR}/cleanup-external.sh" - fi -} - -if [[ "${LEGACY_RESOURCE_GROUP}" -ne 1 ]]; then - external_cleanup - exit 0 -fi - -require_lab_config -verify_subscription - -if ! resource_group_exists; then - echo "Resource group ${RESOURCE_GROUP} is already absent." - external_cleanup - exit 0 -fi -verify_lab_resource_group - -external_cleanup - -echo "Planned legacy cleanup:" -echo " Delete tagged resource group: ${RESOURCE_GROUP}" - -if [[ "${CONFIRMED}" -ne 1 ]]; then - echo "Dry run only. Re-run with --yes to execute." - exit 0 -fi - -az group delete \ - --name "${RESOURCE_GROUP}" \ - --yes \ - --no-wait - -echo "Deletion started for ${RESOURCE_GROUP}." diff --git a/monitor/sre-agent-event-lab/scripts/common.sh b/monitor/sre-agent-event-lab/scripts/common.sh index 4f68bb2..46c8b21 100755 --- a/monitor/sre-agent-event-lab/scripts/common.sh +++ b/monitor/sre-agent-event-lab/scripts/common.sh @@ -103,7 +103,7 @@ load_lab_config() { # Set by the deploy phase (`postdeploy` hook, scripts/azd-deploy-app.sh) # once it has built the lab image and switched the Container App onto it; # empty until then, which is the reliable "has azd deploy run yet?" signal - # doctor.sh needs to tell the intermediate placeholder state (expected) + # Callers need to tell the intermediate placeholder state (expected) # apart from a real post-deploy health regression (not expected). SRE_CONTAINER_IMAGE="$(setting SRE_CONTAINER_IMAGE "${SRE_CONTAINER_IMAGE:-}" "")" # Deployment outputs without an AZURE_-prefixed duplicate (see @@ -273,7 +273,7 @@ evidence_dir_path() { # create_evidence_dir SCENARIO -- name it and create it in one step, for # callers that write into it immediately and have nothing left to refuse -# them (`baseline.sh`). +# them (the baseline step in `guides/01-agent-setup.md`). create_evidence_dir() { local directory directory="$(evidence_dir_path "$1")" @@ -393,7 +393,7 @@ wait_for_new_revision_ready() { # alert_monitor_condition ALERT_ID -- Azure Monitor's own word for the # alert's current state: `Fired`, `Resolved`, or empty when the read failed. # ALERT_ID is the alert's full ARM resource ID, exactly as -# `run-scenario.sh` recorded it from the Alerts Management list. +# the scenario's timeline recorded it from the Alerts Management list. alert_monitor_condition() { local alert_id="$1" az rest \ diff --git a/monitor/sre-agent-event-lab/scripts/deploy.sh b/monitor/sre-agent-event-lab/scripts/deploy.sh deleted file mode 100755 index 01f6a39..0000000 --- a/monitor/sre-agent-event-lab/scripts/deploy.sh +++ /dev/null @@ -1,22 +0,0 @@ -#!/usr/bin/env bash -# Compatibility wrapper. The lab is an azd project now: azure.yaml owns the -# Bicep entry point, the preprovision/postprovision hooks register providers -# and prepare the local environment, and the postdeploy hook builds the image -# in ACR and moves the Container App onto it once AcrPull has propagated. This -# script only forwards to `azd up` -- which runs both phases -- so the -# previously documented command keeps working. -set -euo pipefail - -SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P)" -readonly SCRIPT_DIR -LAB_ROOT="$(cd "${SCRIPT_DIR}/.." && pwd -P)" -readonly LAB_ROOT - -command -v azd >/dev/null 2>&1 || { - echo "Required command not found: azd (https://aka.ms/azd-install)" >&2 - exit 1 -} - -echo "deploy.sh now runs 'azd up' in ${LAB_ROOT}." >&2 -cd "${LAB_ROOT}" -exec azd up "$@" diff --git a/monitor/sre-agent-event-lab/scripts/doctor.sh b/monitor/sre-agent-event-lab/scripts/doctor.sh deleted file mode 100755 index c6fd482..0000000 --- a/monitor/sre-agent-event-lab/scripts/doctor.sh +++ /dev/null @@ -1,354 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P)" -source "${SCRIPT_DIR}/common.sh" - -# `require_lab_config` may return before every value below is assigned (a -# missing required setting makes it `return 1` immediately). Seeding these -# names from the process environment keeps every later `${NAME}`/`-n` check -# well-defined under `set -u` no matter how far configuration loading got. -# The process environment is exactly where `common.sh`'s `setting` takes its -# highest-precedence value from, so this reads the same source it would -- -# it does not invent a fallback, and any name that survives here unset stays -# empty rather than becoming a configuration value on its own. -SUBSCRIPTION_ID="${SUBSCRIPTION_ID:-}" -RESOURCE_GROUP="${RESOURCE_GROUP:-}" -AZURE_ENV_NAME="${AZURE_ENV_NAME:-}" -SRE_AGENT_RESOURCE_ID="${SRE_AGENT_RESOURCE_ID:-}" - -ANY_FAIL=0 -AZURE_SAFE=1 - -# report CHECK STATUS DETAIL -- the doctor output contract: a single -# tab-separated line per check. STATUS is PASS, FAIL, or MANUAL; only FAIL -# marks the overall run unhealthy (`acknowledge`d MANUAL items are a human -# decision, not this script's to make). -report() { - local check_name="$1" status="$2" detail="$3" - printf '%s\t%s\t%s\n' "${check_name}" "${status}" "${detail}" - if [[ "${status}" == "FAIL" ]]; then - ANY_FAIL=1 - fi -} - -# 1. Required commands ------------------------------------------------------ -MISSING_COMMANDS=() -for command_name in az azd jq curl python3; do - command -v "${command_name}" >/dev/null 2>&1 || MISSING_COMMANDS+=("${command_name}") -done -if [[ "${#MISSING_COMMANDS[@]}" -eq 0 ]]; then - report "Required commands" PASS "az, azd, jq, curl, python3 found on PATH." -else - AZURE_SAFE=0 - report "Required commands" FAIL "Install missing commands: ${MISSING_COMMANDS[*]}." -fi - -# Python environment (app/.venv + Pillow) ------------------------------------- -# The documented azd-first flow prepares `app/.venv` from the `postprovision` -# hook (`scripts/azd-postprovision-local.sh` -> `scripts/setup-venv.sh`, its -# whole job), before any scenario is captured. This check -# reports whether that step actually finished -- Pillow importable, not just -# a venv directory present -- so a partially-run or pre-uv-created venv is -# caught here rather than surfacing later as `capture-scenario.sh`'s "Missing -# Python environment" failure. This is a local precondition, independent of -# Azure reachability, so it neither reads AZURE_SAFE nor sets it. -VENV_PYTHON="${SCRIPT_DIR}/../app/.venv/bin/python" -if [[ -x "${VENV_PYTHON}" ]] && "${VENV_PYTHON}" -c "import PIL" >/dev/null 2>&1; then - report "Python environment" PASS "app/.venv is ready (Pillow importable) for capture-scenario.sh." -else - report "Python environment" FAIL "app/.venv is missing or incomplete (Pillow not importable). Run: ./scripts/setup-venv.sh" -fi - -# Log Analytics CLI extension ------------------------------------------------ -# `az monitor log-analytics query` -- the only read behind the telemetry -# check and behind `baseline.sh`/`query-evidence.sh` -- ships in an extension -# that is not installed with the core CLI, so its absence is a prerequisite -# failure with its own row rather than a mysterious empty query result. -if ! command -v az >/dev/null 2>&1; then - AZURE_SAFE=0 - report "Log Analytics CLI extension" FAIL "Blocked: install the Azure CLI first, then run: az extension add --name ${LOG_ANALYTICS_EXTENSION_NAME}" -elif log_analytics_extension_installed; then - report "Log Analytics CLI extension" PASS "az monitor log-analytics query is available (extension ${LOG_ANALYTICS_EXTENSION_NAME})." -else - AZURE_SAFE=0 - report "Log Analytics CLI extension" FAIL "az monitor log-analytics query is unavailable. Install it: az extension add --name ${LOG_ANALYTICS_EXTENSION_NAME}" -fi - -# Azure CLI login ------------------------------------------------------------ -if login_error="$(az account show --query id -o tsv 2>&1 1>/dev/null)"; then - report "Azure CLI login" PASS "Signed in to Azure CLI." -else - AZURE_SAFE=0 - report "Azure CLI login" FAIL "Run: az login -- ${login_error:-not signed in}" -fi - -# azd authentication --------------------------------------------------------- -# azd keeps its own credential store: `az login` alone does not make -# `azd env`/`azd provision` work. `azd auth login --check-status` is the only -# non-interactive read of that state and always exits 0, so the row is -# decided by the status it prints, never by its exit code. -AZD_AUTH_STATUS="" -if command -v azd >/dev/null 2>&1; then - AZD_AUTH_STATUS="$(azd_auth_status)" -fi -case "${AZD_AUTH_STATUS}" in - success) - report "azd authentication" PASS "azd reports an authenticated session (azd auth login --check-status)." - ;; - unauthenticated) - AZURE_SAFE=0 - report "azd authentication" FAIL "azd is not signed in. Run: azd auth login" - ;; - *) - AZURE_SAFE=0 - report "azd authentication" FAIL "azd did not report a login status. Check it by hand: azd auth login --check-status, then run: azd auth login" - ;; -esac - -# azd configuration ----------------------------------------------------------- -# `require_lab_config` must run in *this* shell (not a subshell) so the -# resolved SUBSCRIPTION_ID/RESOURCE_GROUP/etc. it makes readonly survive for -# every check below; only its stderr is captured to a scratch file. -mkdir -p "${EVIDENCE_ROOT}" -CONFIG_ERROR_FILE="${EVIDENCE_ROOT}/.doctor-config-error" -: >"${CONFIG_ERROR_FILE}" -if require_lab_config 2>"${CONFIG_ERROR_FILE}"; then - report "azd configuration" PASS "Resolved subscription ${SUBSCRIPTION_ID}, resource group ${RESOURCE_GROUP}, environment ${AZURE_ENV_NAME}." -else - AZURE_SAFE=0 - CONFIG_DETAIL="$(tr '\n' ' ' <"${CONFIG_ERROR_FILE}" | sed 's/[[:space:]]*$//')" - report "azd configuration" FAIL "${CONFIG_DETAIL:-Missing required azd configuration.}" - # `load_lab_config` returns before its `readonly` line whenever a - # required setting is missing, so none of these names are readonly yet - # on this path -- safe (and necessary, under `set -u`) to re-seed them. - SUBSCRIPTION_ID="${SUBSCRIPTION_ID:-}" - RESOURCE_GROUP="${RESOURCE_GROUP:-}" - AZURE_ENV_NAME="${AZURE_ENV_NAME:-}" - SRE_AGENT_RESOURCE_ID="${SRE_AGENT_RESOURCE_ID:-}" -fi -rm -f "${CONFIG_ERROR_FILE}" - -# 2. Subscription equality --------------------------------------------------- -if [[ "${AZURE_SAFE}" -eq 1 ]]; then - if subscription_error="$(verify_subscription 2>&1 1>/dev/null)"; then - report "Subscription match" PASS "Active subscription matches ${SUBSCRIPTION_ID}." - else - AZURE_SAFE=0 - report "Subscription match" FAIL "${subscription_error}" - fi -else - report "Subscription match" FAIL "Blocked: resolve the failing check above first." -fi - -# 3. Resource group tags ----------------------------------------------------- -if [[ "${AZURE_SAFE}" -eq 1 ]]; then - if rg_error="$(verify_lab_resource_group 2>&1 1>/dev/null)"; then - report "Resource group tags" PASS "Resource group ${RESOURCE_GROUP} is tagged purpose=sre-agent-event-lab, azd-env-name=${AZURE_ENV_NAME}." - else - AZURE_SAFE=0 - report "Resource group tags" FAIL "${rg_error}" - fi -else - report "Resource group tags" FAIL "Blocked: resolve the failing check above first." -fi - -if [[ "${AZURE_SAFE}" -eq 1 ]]; then - APP_NAME="$(deployment_output containerAppName)" - APP_FQDN="$(deployment_output containerAppFqdn)" - WORKSPACE_CUSTOMER_ID="$(deployment_output workspaceCustomerId)" - TELEMETRY_SERVICE_NAME="$(deployment_output telemetryServiceName)" -else - APP_NAME="" - APP_FQDN="" - WORKSPACE_CUSTOMER_ID="" - TELEMETRY_SERVICE_NAME="" -fi - -# 4. Container App health and /healthz --------------------------------------- -if [[ "${AZURE_SAFE}" -eq 1 && -n "${APP_NAME}" ]]; then - health_state="$(az containerapp revision list \ - --resource-group "${RESOURCE_GROUP}" \ - --name "${APP_NAME}" \ - --query "[?properties.active].properties.healthState | [0]" \ - -o tsv 2>/dev/null || true)" - if [[ "${health_state}" == "Healthy" ]]; then - report "Container App health" PASS "Active revision of ${APP_NAME} is Healthy." - else - report "Container App health" FAIL "Active revision health is '${health_state:-unknown}'. Investigate: az containerapp revision list --resource-group ${RESOURCE_GROUP} --name ${APP_NAME}" - fi -elif [[ "${AZURE_SAFE}" -eq 1 ]]; then - report "Container App health" FAIL "Deployment output containerAppName is empty. Run: azd provision" -else - report "Container App health" FAIL "Blocked: resolve the failing check above first." -fi - -if [[ "${AZURE_SAFE}" -eq 1 && -n "${APP_FQDN}" ]]; then - http_status="$(curl --max-time 10 --silent --output /dev/null --write-out '%{http_code}' "https://${APP_FQDN}/healthz" 2>/dev/null || echo 000)" - if [[ "${http_status}" == "200" ]]; then - report "Health endpoint" PASS "https://${APP_FQDN}/healthz returned HTTP 200." - elif [[ -z "${SRE_CONTAINER_IMAGE}" ]]; then - # SRE_CONTAINER_IMAGE is only set once the deploy phase (`postdeploy` - # hook, scripts/azd-deploy-app.sh) has built the lab image and switched - # the Container App onto it. Empty here means `azd provision` ran but - # `azd deploy` has not (yet): the app is still the public placeholder - # image (port 80, no /healthz -- see infra/main.bicep), which is the - # documented, expected intermediate state, not a broken deployment. The - # generic "investigate with curl -v" wording below would misclassify - # that state, so this branch names the actual remedy instead. - report "Health endpoint" FAIL "https://${APP_FQDN}/healthz returned HTTP ${http_status}, and SRE_CONTAINER_IMAGE is not recorded -- the deploy phase has not run yet (this is the expected placeholder state after 'azd provision' alone, not a broken deployment). Run: azd deploy --no-prompt" - else - report "Health endpoint" FAIL "https://${APP_FQDN}/healthz returned HTTP ${http_status}. Investigate: curl -v https://${APP_FQDN}/healthz" - fi -elif [[ "${AZURE_SAFE}" -eq 1 ]]; then - report "Health endpoint" FAIL "Deployment output containerAppFqdn is empty. Run: azd provision" -else - report "Health endpoint" FAIL "Blocked: resolve the failing check above first." -fi - -# 5. Application Insights request telemetry in the last 30 minutes ---------- -# The query projects and `take`s real rows instead of `| count`: KQL's -# `count` always returns exactly one row (`Count: 0` for an empty table), so -# a row count taken from it can never distinguish data from no data. The -# extension prints a flat JSON array, and `log_analytics_row_count` parses -# that shape (see common.sh). -if [[ "${AZURE_SAFE}" -eq 1 && -n "${WORKSPACE_CUSTOMER_ID}" && -n "${TELEMETRY_SERVICE_NAME}" ]]; then - request_rows="$(log_analytics_row_count "${WORKSPACE_CUSTOMER_ID}" PT30M \ - "AppRequests | where AppRoleName == '${TELEMETRY_SERVICE_NAME}' | project TimeGenerated, Name | take 1")" - if [[ "${request_rows:-0}" -gt 0 ]]; then - report "Application Insights telemetry" PASS "AppRequests present for ${TELEMETRY_SERVICE_NAME} in the last 30 minutes." - else - report "Application Insights telemetry" FAIL "No AppRequests telemetry in the last 30 minutes for role ${TELEMETRY_SERVICE_NAME}. Run: lab.sh baseline" - fi -elif [[ "${AZURE_SAFE}" -eq 1 ]]; then - report "Application Insights telemetry" FAIL "Deployment outputs workspaceCustomerId/telemetryServiceName are empty. Run: azd provision" -else - report "Application Insights telemetry" FAIL "Blocked: resolve the failing check above first." -fi - -# 6. Three alert rules enabled ------------------------------------------------ -if [[ "${AZURE_SAFE}" -eq 1 ]]; then - DISABLED_RULES=() - for rule_name in alert-sre-lab-s1-http500 alert-sre-lab-s2-latency alert-sre-lab-s3-storage-rbac; do - rule_json="$(az rest --method get \ - --url "https://management.azure.com/subscriptions/${SUBSCRIPTION_ID}/resourceGroups/${RESOURCE_GROUP}/providers/microsoft.insights/scheduledqueryrules/${rule_name}?api-version=2023-12-01" \ - -o json 2>/dev/null || true)" - if [[ -z "${rule_json}" ]]; then - rule_json='{}' - fi - rule_enabled="$(jq -r '.properties.enabled // false' <<<"${rule_json}" 2>/dev/null || echo false)" - if [[ "${rule_enabled}" != "true" ]]; then - DISABLED_RULES+=("${rule_name}") - fi - done - if [[ "${#DISABLED_RULES[@]}" -eq 0 ]]; then - report "Alert rules enabled" PASS "alert-sre-lab-s1-http500, alert-sre-lab-s2-latency, alert-sre-lab-s3-storage-rbac are all enabled." - else - report "Alert rules enabled" FAIL "Not enabled or missing: ${DISABLED_RULES[*]}. Re-enable: az resource update --ids --set properties.enabled=true (or re-run azd provision)." - fi -else - report "Alert rules enabled" FAIL "Blocked: resolve the failing check above first." -fi - -# 7. SRE Agent resource, only when SRE_AGENT_RESOURCE_ID is configured ------- -if [[ -n "${SRE_AGENT_RESOURCE_ID}" ]]; then - if [[ "${AZURE_SAFE}" -eq 1 ]]; then - if az resource show --ids "${SRE_AGENT_RESOURCE_ID}" -o none 2>/dev/null; then - report "SRE Agent resource" PASS "Resource exists: ${SRE_AGENT_RESOURCE_ID}." - else - report "SRE Agent resource" FAIL "Resource not found: ${SRE_AGENT_RESOURCE_ID}. Verify: az resource show --ids ${SRE_AGENT_RESOURCE_ID}" - fi - else - report "SRE Agent resource" FAIL "Blocked: resolve the failing check above first." - fi -fi - -# 8. Reader role assignment on the lab resource group ------------------------ -# `--include-inherited` is deliberate: Reader granted at the subscription (or -# management group) gives the Agent exactly the effective read access it -# needs on this resource group, and omitting the flag hides those grants -# entirely -- `az role assignment list` returns only assignments made at the -# queried scope without it -- which would report a working setup as broken. -# The detail still distinguishes the two, because an operator who requires an -# explicit resource-group-scoped assignment has to be able to see that the -# access is only inherited. -if [[ "${AZURE_SAFE}" -eq 1 ]]; then - RESOURCE_GROUP_SCOPE="/subscriptions/${SUBSCRIPTION_ID}/resourceGroups/${RESOURCE_GROUP}" - if [[ ! -f "${AGENT_SETUP_FILE}" ]]; then - report "Reader role assignment" FAIL "Agent setup evidence missing: ${AGENT_SETUP_FILE}. Create evidence/agent-setup.json as shown in guides/01-agent-setup.md, then run: lab.sh acknowledge agent-setup." - elif ! AGENT_SETUP_JSON="$(jq '.' "${AGENT_SETUP_FILE}" 2>/dev/null)"; then - # A hand-edited or truncated evidence file is one FAIL row, not a raw - # `jq` abort: under `set -e` an unguarded parse would kill the run and - # swallow every remaining check, including the MANUAL rows an operator - # still needs. - report "Reader role assignment" FAIL "Agent setup evidence is not valid JSON: ${AGENT_SETUP_FILE}. Rewrite evidence/agent-setup.json as shown in guides/01-agent-setup.md, then run: lab.sh acknowledge agent-setup." - else - agent_principal_id="$(jq -r '.agent_principal_id // empty' <<<"${AGENT_SETUP_JSON}" 2>/dev/null || true)" - agent_uami_principal_id="$(jq -r '.agent_user_assigned_principal_id // empty' <<<"${AGENT_SETUP_JSON}" 2>/dev/null || true)" - monitoring_assignment_id="$(jq -r '.monitoring_contributor_assignment_id // empty' <<<"${AGENT_SETUP_JSON}" 2>/dev/null || true)" - uami_monitoring_assignment_id="$(jq -r '.uami_monitoring_contributor_assignment_id // empty' <<<"${AGENT_SETUP_JSON}" 2>/dev/null || true)" - EXPECTED_ASSIGNMENT_PREFIX="/subscriptions/${SUBSCRIPTION_ID}/providers/Microsoft.Authorization/roleAssignments/" - MISSING_CLEANUP_KEYS=() - [[ "$(lowercase "${monitoring_assignment_id}")" == "$(lowercase "${EXPECTED_ASSIGNMENT_PREFIX}")"* ]] || MISSING_CLEANUP_KEYS+=("monitoring_contributor_assignment_id") - [[ "$(lowercase "${uami_monitoring_assignment_id}")" == "$(lowercase "${EXPECTED_ASSIGNMENT_PREFIX}")"* ]] || MISSING_CLEANUP_KEYS+=("uami_monitoring_contributor_assignment_id") - if [[ "${#MISSING_CLEANUP_KEYS[@]}" -gt 0 ]]; then - report "Monitoring Contributor cleanup evidence" FAIL "Missing or invalid: ${MISSING_CLEANUP_KEYS[*]}. Find each live ID with: az role assignment list --assignee-object-id --role \"Monitoring Contributor\" --scope /subscriptions/${SUBSCRIPTION_ID} --query \"[0].id\" -o tsv; then update ${AGENT_SETUP_FILE}." - else - report "Monitoring Contributor cleanup evidence" PASS "Both subscription-scoped role assignment IDs are recorded." - fi - if [[ -z "${agent_principal_id}" || -z "${agent_uami_principal_id}" ]]; then - report "Reader role assignment" FAIL "Agent setup evidence is missing agent_principal_id/agent_user_assigned_principal_id: ${AGENT_SETUP_FILE}." - else - MISSING_READER=() - INHERITED_READER=() - for principal_id in "${agent_principal_id}" "${agent_uami_principal_id}"; do - assignments_json="$(az role assignment list \ - --resource-group "${RESOURCE_GROUP}" \ - --assignee-object-id "${principal_id}" \ - --include-inherited \ - -o json 2>/dev/null || true)" - if [[ -z "${assignments_json}" ]]; then - assignments_json='[]' - fi - direct_count="$(jq --arg scope "${RESOURCE_GROUP_SCOPE}" ' - [.[]? | select((.roleDefinitionName // "") == "Reader") - | select(((.scope // "") | ascii_downcase) == ($scope | ascii_downcase))] - | length' <<<"${assignments_json}" 2>/dev/null || echo 0)" - inherited_scope="$(jq -r --arg scope "${RESOURCE_GROUP_SCOPE}" ' - [.[]? | select((.roleDefinitionName // "") == "Reader") - | select(((.scope // "") | ascii_downcase) != ($scope | ascii_downcase)) - | .scope] - | first // empty' <<<"${assignments_json}" 2>/dev/null || true)" - if [[ "${direct_count:-0}" -gt 0 ]]; then - continue - elif [[ -n "${inherited_scope}" ]]; then - INHERITED_READER+=("${principal_id} (inherited from ${inherited_scope})") - else - MISSING_READER+=("${principal_id}") - fi - done - if [[ "${#MISSING_READER[@]}" -gt 0 ]]; then - report "Reader role assignment" FAIL "Missing Reader on ${RESOURCE_GROUP} (direct or inherited) for: ${MISSING_READER[*]}. Grant: az role assignment create --assignee-object-id --assignee-principal-type ServicePrincipal --role Reader --resource-group ${RESOURCE_GROUP}" - elif [[ "${#INHERITED_READER[@]}" -gt 0 ]]; then - report "Reader role assignment" PASS "Reader is effective on ${RESOURCE_GROUP} for both recorded Agent identities; not assigned directly for: ${INHERITED_READER[*]}. Inherited access is sufficient to read the lab; assign it on ${RESOURCE_GROUP} if the setup must be scoped to this lab only." - else - report "Reader role assignment" PASS "Reader is assigned directly on ${RESOURCE_GROUP} for both recorded Agent identities." - fi - fi - fi -else - report "Reader role assignment" FAIL "Blocked: resolve the failing check above first." -fi - -# 9. Portal-only settings: never inferred, always MANUAL --------------------- -# No official stable Azure SRE Agent API currently reads back the -# repository connection, knowledge sources, incident platform, or response -# plan mode, so these are never reported as PASS/FAIL -- doing so would -# mean guessing. They stay MANUAL until an official stable API can prove -# them, matching the portal path an operator needs to check by hand. -report "Repository connection" MANUAL "No official stable API exposes the Agent's repository connection state. Verify in the portal: https://sre.azure.com > Agent > Settings > Repository." -report "Knowledge source" MANUAL "No official stable API exposes configured knowledge sources. Verify in the portal: https://sre.azure.com > Agent > Settings > Knowledge." -report "Incident platform" MANUAL "No official stable API exposes the incident platform connection. Verify in the portal: https://sre.azure.com > Agent > Settings > Incident platform." -report "Response plan" MANUAL "No official stable API confirms the response plan mode. Verify in the portal: https://sre.azure.com > Agent > Response plans (must be Review)." - -exit "${ANY_FAIL}" diff --git a/monitor/sre-agent-event-lab/scripts/lab-env.sh b/monitor/sre-agent-event-lab/scripts/lab-env.sh index 8fd942c..d23d054 100755 --- a/monitor/sre-agent-event-lab/scripts/lab-env.sh +++ b/monitor/sre-agent-event-lab/scripts/lab-env.sh @@ -118,6 +118,8 @@ lab_env_bind APP_FQDN AZURE_CONTAINER_APP_FQDN || true lab_env_bind WORKLOAD_PRINCIPAL_ID AZURE_CONTAINER_APP_PRINCIPAL_ID || true lab_env_bind STORAGE_CONTAINER_SCOPE AZURE_STORAGE_CONTAINER_SCOPE || true lab_env_bind BLOB_ROLE_ASSIGNMENT_NAME AZURE_BLOB_ROLE_ASSIGNMENT_NAME || true +lab_env_bind WORKSPACE_CUSTOMER_ID AZURE_WORKSPACE_CUSTOMER_ID || true +lab_env_bind TELEMETRY_SERVICE_NAME AZURE_TELEMETRY_SERVICE_NAME || true # The manual walkthrough issues `az` commands that inherit the CLI's active # subscription, so a mismatch would inject the failure into a same-named diff --git a/monitor/sre-agent-event-lab/scripts/lab.sh b/monitor/sre-agent-event-lab/scripts/lab.sh deleted file mode 100755 index cf91ea9..0000000 --- a/monitor/sre-agent-event-lab/scripts/lab.sh +++ /dev/null @@ -1,60 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P)" -source "${SCRIPT_DIR}/common.sh" - -usage() { - cat <<'USAGE' -Usage: lab.sh [args] - -Commands: - doctor Diagnose the lab environment - baseline Run baseline load and verify telemetry - acknowledge agent-setup Record manually-verified Agent setup evidence - run s1|s2|s3 Run a failure scenario - capture s1|s2|s3 Capture Azure SRE Agent evidence for a scenario - score Score the collected evidence - -Commands run in this order: doctor, baseline, acknowledge agent-setup, -then run/capture for s1, s2 and s3 in turn, then score. Each step refuses -to start until `evidence/state.json` records the previous one. -USAGE -} - -# Sub-scripts are run through `bash` explicitly (not a bare `exec path`) so -# dispatch never depends on that file's executable bit. -case "${1:-}" in - doctor) - exec bash "${SCRIPT_DIR}/doctor.sh" - ;; - baseline) - exec bash "${SCRIPT_DIR}/baseline.sh" - ;; - acknowledge) - [[ "${2:-}" == "agent-setup" ]] || { usage >&2; exit 2; } - # Reads the operator's typed answer from this process's stdin; the - # configuration is loaded first so the acknowledgement is recorded - # against the azd environment it was given for. - require_lab_config - lab_state acknowledge-agent - ;; - run) - [[ "${2:-}" =~ ^s[123]$ ]] || { usage >&2; exit 2; } - exec bash "${SCRIPT_DIR}/run-scenario.sh" "${2}" - ;; - capture) - [[ "${2:-}" =~ ^s[123]$ ]] || { usage >&2; exit 2; } - # capture-scenario.sh resolves the evidence directory this scenario's - # recorded run wrote, so the public command needs no timestamp. - exec bash "${SCRIPT_DIR}/capture-scenario.sh" "${2}" - ;; - score) - require_lab_config - lab_tool score.py --evidence-root "${EVIDENCE_ROOT}" - ;; - *) - usage >&2 - exit 2 - ;; -esac diff --git a/monitor/sre-agent-event-lab/scripts/lab_state.py b/monitor/sre-agent-event-lab/scripts/lab_state.py index 9f2cead..6b3b63b 100755 --- a/monitor/sre-agent-event-lab/scripts/lab_state.py +++ b/monitor/sre-agent-event-lab/scripts/lab_state.py @@ -174,6 +174,22 @@ def terminal_state(events: Iterable[Dict[str, Any]]) -> str: return "thread-not-created" +# The document that walks an operator through each scenario. Refusal +# messages name it instead of a command, because the lab is run by hand: +# there is no script that performs a scenario. +SCENARIO_GUIDES = { + "s1": "guides/02-scenario-s1.md", + "s2": "guides/03-scenario-s2.md", + "s3": "guides/04-scenario-s3.md", +} + + +def scenario_guide(scenario: str) -> str: + # A remedy that cannot be acted on is not a remedy: an unmapped + # scenario still has to say where to look. + return SCENARIO_GUIDES.get(scenario, "the matching guide under guides/") + + def _scenario_stage(stage: str) -> Optional[Sequence[str]]: """('s1', 'recovered') for `s1_recovered`, else None.""" for scenario in SCENARIOS: @@ -458,7 +474,7 @@ def _unfinished_remedy(scenario: str, status: str) -> str: "wait for it to finish, or record how it ended with " "lab_state.py mark-failed {0}".format(scenario) ) - return "Run: lab.sh run {0}".format(scenario) + return "follow {0}".format(scenario_guide(scenario)) def _remedy(self, missing: Sequence[str]) -> str: """The next command that is actually reachable for a missing stage. @@ -471,8 +487,11 @@ def _remedy(self, missing: Sequence[str]) -> str: gate would demand first. """ remedies = { - "baseline_passed": "Run: lab.sh baseline", - "agent_setup_acknowledged": "Run: lab.sh acknowledge agent-setup", + "baseline_passed": ( + "Run the baseline steps in guides/01-agent-setup.md, then: " + "lab_state.py mark baseline_passed" + ), + "agent_setup_acknowledged": "Run: lab_state.py acknowledge-agent", } for stage in missing: if stage in remedies: @@ -487,8 +506,12 @@ def _remedy(self, missing: Sequence[str]) -> str: blocked, status, self._unfinished_remedy(blocked, status) ) if suffix == "recovered": - return "Run: lab.sh run {0}".format(scenario) - return "Run: lab.sh capture {0}".format(scenario) + return "Run the {0} steps in {1}".format( + scenario, scenario_guide(scenario) + ) + return "Capture the {0} evidence as {1} describes".format( + scenario, scenario_guide(scenario) + ) return "" @staticmethod @@ -694,7 +717,7 @@ def acknowledge_agent(state: LabState, stream=None, output=None) -> int: """Print the configured Agent wiring and require a typed acknowledgement. None of these settings has an official, stable API to read back (see - `doctor.sh`'s four permanent `MANUAL` rows), so the only honest evidence + the four settings no stable API exposes), so the only honest evidence that the portal side is done is a human who looked at it. A configured environment variable proves intent, never completion -- which is why this command reads the answer from stdin and accepts nothing but the @@ -832,7 +855,9 @@ def main(argv: Optional[Sequence[str]] = None) -> int: if not directory: raise LabStateError( "No evidence directory recorded for {0}. " - "Run: lab.sh run {0}".format(args.scenario) + "Run the {0} steps in {1}".format( + args.scenario, scenario_guide(args.scenario) + ) ) print(directory) return 0 diff --git a/monitor/sre-agent-event-lab/scripts/run-scenario.sh b/monitor/sre-agent-event-lab/scripts/run-scenario.sh deleted file mode 100755 index 808d2e0..0000000 --- a/monitor/sre-agent-event-lab/scripts/run-scenario.sh +++ /dev/null @@ -1,432 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P)" -source "${SCRIPT_DIR}/common.sh" - -if [[ "$#" -ne 1 ]] || [[ ! "$1" =~ ^s[123]$ ]]; then - echo "Usage: $0 s1|s2|s3" >&2 - exit 2 -fi - -readonly SCENARIO="$1" -require_lab_config -verify_subscription -verify_lab_resource_group - -# The run order is a safety boundary, not a convenience: a scenario started -# before the previous one recovered and was captured overlaps two incidents -# in one workload, and neither capture can then be read. The same applies -# to any scenario left `running` or `failed` -- its fault may still be live -# -- so an unfinished run anywhere refuses every scenario, not just the -# next one. Checked before the first Azure call that breaks anything. -lab_state require-run "${SCENARIO}" - -# Overridable only for tests; production runs use the defaults. -readonly ALERT_RESOLVE_TIMEOUT_SECONDS="${LAB_ALERT_RESOLVE_TIMEOUT_SECONDS:-1500}" -readonly ALERT_RESOLVE_POLL_INTERVAL_SECONDS="${LAB_ALERT_RESOLVE_POLL_INTERVAL_SECONDS:-20}" -readonly RECOVERY_HEALTH_TIMEOUT_SECONDS="${LAB_RECOVERY_HEALTH_TIMEOUT_SECONDS:-600}" -readonly ALERT_FIRE_TIMEOUT_SECONDS="${LAB_ALERT_FIRE_TIMEOUT_SECONDS:-720}" -readonly ALERT_FIRE_POLL_INTERVAL_SECONDS="${LAB_ALERT_FIRE_POLL_INTERVAL_SECONDS:-20}" -readonly REVISION_READY_TIMEOUT_SECONDS="${LAB_REVISION_READY_TIMEOUT_SECONDS:-600}" -readonly REVISION_READY_POLL_INTERVAL_SECONDS="${LAB_REVISION_READY_POLL_INTERVAL_SECONDS:-10}" -readonly S3_PROPAGATION_TIMEOUT_SECONDS="${LAB_S3_PROPAGATION_TIMEOUT_SECONDS:-300}" -readonly S3_PROPAGATION_POLL_INTERVAL_SECONDS="${LAB_S3_PROPAGATION_POLL_INTERVAL_SECONDS:-10}" - -APP_NAME="$(deployment_output containerAppName)" -APP_FQDN="$(deployment_output containerAppFqdn)" -WORKLOAD_PRINCIPAL_ID="$(deployment_output containerAppPrincipalId)" -STORAGE_CONTAINER_SCOPE="$(deployment_output storageContainerScope)" -BLOB_ROLE_ASSIGNMENT_NAME="$(deployment_output blobRoleAssignmentName)" -readonly APP_NAME APP_FQDN WORKLOAD_PRINCIPAL_ID STORAGE_CONTAINER_SCOPE -readonly BLOB_ROLE_ASSIGNMENT_NAME - -MISSING_OUTPUTS="" -[[ -n "${APP_NAME}" ]] || MISSING_OUTPUTS="${MISSING_OUTPUTS} containerAppName" -[[ -n "${APP_FQDN}" ]] || MISSING_OUTPUTS="${MISSING_OUTPUTS} containerAppFqdn" -if [[ "${SCENARIO}" == "s3" ]]; then - [[ -n "${WORKLOAD_PRINCIPAL_ID}" ]] || MISSING_OUTPUTS="${MISSING_OUTPUTS} containerAppPrincipalId" - [[ -n "${STORAGE_CONTAINER_SCOPE}" ]] || MISSING_OUTPUTS="${MISSING_OUTPUTS} storageContainerScope" - [[ -n "${BLOB_ROLE_ASSIGNMENT_NAME}" ]] || MISSING_OUTPUTS="${MISSING_OUTPUTS} blobRoleAssignmentName" -fi -if [[ -n "${MISSING_OUTPUTS}" ]]; then - echo "Missing deployment outputs:${MISSING_OUTPUTS}. Run: azd provision" >&2 - exit 1 -fi - -EVIDENCE_DIR="$(evidence_dir_path "${SCENARIO}")" -readonly EVIDENCE_DIR - -# The attempt is recorded before anything can break, and clears whatever -# the previous attempt left behind. Everything below this line can exit -# without reaching `mark-recovered`/`mark-failed` -- a rejected injection, -# a recovery the EXIT trap cannot complete, a Ctrl-C -- and the state file -# must never keep describing the run this one replaces: a re-run of an -# already-captured scenario would otherwise leave `recovered` + -# `conclusion` in place and admit the next scenario on evidence from a run -# that no longer exists. -# -# `begin-run` is also the last gate: it re-reads `state.json` and refuses -# while any scenario is still `running` or `failed`, which covers the -# window between the `require-run` above and here. The directory is only -# created afterwards, so a refusal leaves nothing behind in `evidence/`. -lab_state begin-run "${SCENARIO}" "${EVIDENCE_DIR}" -mkdir -p "${EVIDENCE_DIR}" -readonly ALERT_LIST_RESPONSE_FILE="${EVIDENCE_DIR}/alert-list-response.json" -readonly ALERT_LIST_ERROR_FILE="${EVIDENCE_DIR}/alert-list-error.log" -readonly ALERT_LIST_ATTEMPT_ERROR_FILE="${EVIDENCE_DIR}/.alert-list-error.tmp" - -RECOVERED=0 -OUTCOME_RECORDED=0 -PENDING_FAILURE_REASON="" -ALERT_LIST_FAILURES=0 -ALERT_LIST_VALID_RESPONSES=0 -ALERT_LIST_POLLS=0 -INJECTED_AT="" -REVISION_READY_AT="" -ROLE_DELETED_AT="" -ALERT_RULE_NAME="" -ALERT_ID="" -ALERT_FIRED_AT="" - -record_failed_run() { - local reason="$1" - PENDING_FAILURE_REASON="${reason}" - if ! lab_state mark-failed "${SCENARIO}" "${EVIDENCE_DIR}" --reason "${reason}"; then - return 1 - fi - OUTCOME_RECORDED=1 - PENDING_FAILURE_REASON="" -} - -wait_for_s3_fault() { - local started="${SECONDS}" - local probe_file="${EVIDENCE_DIR}/rbac-probe.json" - - while (( SECONDS - started < S3_PROPAGATION_TIMEOUT_SECONDS )); do - if python3 "${SCRIPT_DIR}/loadgen.py" \ - "https://${APP_FQDN}/api/documents" \ - --requests 1 \ - --concurrency 1 \ - --expect-status 503 \ - --output "${probe_file}"; then - return 0 - fi - echo "Waiting for the Storage RBAC deletion to reach the data plane..." >&2 - sleep "${S3_PROPAGATION_POLL_INTERVAL_SECONDS}" - done - - echo "Storage RBAC deletion did not produce HTTP 503 within ${S3_PROPAGATION_TIMEOUT_SECONDS}s." >&2 - return 1 -} - -# restore_container_app_env SETTING -- reverts one injected Container App -# setting and waits for the revision that carries it to become active. -# -# Every step is checked explicitly. `recover` is also called as -# `if ! recover` from the EXIT trap, and bash disables `set -e` inside a -# function invoked in a condition: an unchecked `az` failure or a timed-out -# wait would fall through to `RECOVERED=1` and report a recovery that never -# happened, leaving the fault live in the Container App. -restore_container_app_env() { - local setting="$1" - local old_revision - - if ! old_revision="$(latest_revision_name "${APP_NAME}")"; then - echo "Recovery failed: could not read the current revision of ${APP_NAME}." >&2 - return 1 - fi - if ! az containerapp update \ - --resource-group "${RESOURCE_GROUP}" \ - --name "${APP_NAME}" \ - --set-env-vars "${setting}" \ - --output none; then - echo "Recovery failed: az containerapp update ${setting} was rejected." >&2 - return 1 - fi - if ! wait_for_new_revision_ready \ - "${APP_NAME}" \ - "${old_revision}" \ - "${REVISION_READY_TIMEOUT_SECONDS}" \ - "${REVISION_READY_POLL_INTERVAL_SECONDS}" >/dev/null; then - echo "Recovery failed: no new healthy revision carrying ${setting}." >&2 - return 1 - fi -} - -# Restores S3's deleted `Storage Blob Data Reader` assignment. A read that -# fails is not "the assignment is missing": it is an unknown state, and -# creating on top of an unknown state is not a recovery either, so both -# propagate. -restore_blob_role() { - local existing_assignment - - if ! existing_assignment="$(az role assignment list \ - --scope "${STORAGE_CONTAINER_SCOPE}" \ - --assignee-object-id "${WORKLOAD_PRINCIPAL_ID}" \ - --query "[?roleDefinitionName=='Storage Blob Data Reader'].id | [0]" \ - -o tsv)"; then - echo "Recovery failed: could not read the blob role assignments of ${STORAGE_CONTAINER_SCOPE}." >&2 - return 1 - fi - if [[ -n "${existing_assignment}" ]]; then - return 0 - fi - if ! az role assignment create \ - --name "${BLOB_ROLE_ASSIGNMENT_NAME}" \ - --assignee-object-id "${WORKLOAD_PRINCIPAL_ID}" \ - --assignee-principal-type ServicePrincipal \ - --role "Storage Blob Data Reader" \ - --scope "${STORAGE_CONTAINER_SCOPE}" \ - --output none; then - echo "Recovery failed: could not restore Storage Blob Data Reader for ${WORKLOAD_PRINCIPAL_ID}." >&2 - return 1 - fi -} - -# `RECOVERED=1` is reached only when the whole branch succeeded, so a failed -# attempt is retried by the EXIT trap instead of being remembered as done. -recover() { - if [[ "${RECOVERED}" -eq 1 ]]; then - return 0 - fi - - case "${SCENARIO}" in - s1) restore_container_app_env FAILURE_MODE=none || return 1 ;; - s2) restore_container_app_env ORDER_DELAY_MS=0 || return 1 ;; - s3) restore_blob_role || return 1 ;; - *) - echo "Recovery failed: no recovery is defined for ${SCENARIO}." >&2 - return 1 - ;; - esac - - RECOVERED=1 - return 0 -} - -recover_on_exit() { - local original_status="$?" - local recovery_status=0 - local failure_reason - trap - EXIT - if ! recover; then - recovery_status=1 - echo "CRITICAL: scenario recovery failed for ${SCENARIO}." >&2 - echo "CRITICAL: the injected fault is still active. Revert it by hand before running any other scenario." >&2 - fi - if [[ "${OUTCOME_RECORDED}" -eq 0 ]]; then - failure_reason="${PENDING_FAILURE_REASON:-run aborted with status ${original_status}}" - if [[ "${recovery_status}" -ne 0 ]]; then - failure_reason="${failure_reason}; automatic recovery failed" - fi - if ! record_failed_run "${failure_reason}"; then - echo "CRITICAL: could not record the aborted ${SCENARIO} run as failed." >&2 - fi - fi - if [[ "${recovery_status}" -ne 0 ]]; then - exit 1 - fi - exit "${original_status}" -} -trap recover_on_exit EXIT - -case "${SCENARIO}" in - s1) - ALERT_RULE_NAME="alert-sre-lab-s1-http500" - OLD_REVISION="$(latest_revision_name "${APP_NAME}")" - INJECTED_AT="$(utc_now)" - az containerapp update \ - --resource-group "${RESOURCE_GROUP}" \ - --name "${APP_NAME}" \ - --set-env-vars FAILURE_MODE=http500 \ - --output none - wait_for_new_revision_ready \ - "${APP_NAME}" \ - "${OLD_REVISION}" \ - "${REVISION_READY_TIMEOUT_SECONDS}" \ - "${REVISION_READY_POLL_INTERVAL_SECONDS}" >/dev/null - REVISION_READY_AT="$(utc_now)" - python3 "${SCRIPT_DIR}/loadgen.py" \ - "https://${APP_FQDN}/api/orders" \ - --requests 120 \ - --concurrency 4 \ - --expect-status 500 \ - --output "${EVIDENCE_DIR}/load.json" - ;; - s2) - ALERT_RULE_NAME="alert-sre-lab-s2-latency" - OLD_REVISION="$(latest_revision_name "${APP_NAME}")" - INJECTED_AT="$(utc_now)" - az containerapp update \ - --resource-group "${RESOURCE_GROUP}" \ - --name "${APP_NAME}" \ - --set-env-vars ORDER_DELAY_MS=4000 \ - --output none - wait_for_new_revision_ready \ - "${APP_NAME}" \ - "${OLD_REVISION}" \ - "${REVISION_READY_TIMEOUT_SECONDS}" \ - "${REVISION_READY_POLL_INTERVAL_SECONDS}" >/dev/null - REVISION_READY_AT="$(utc_now)" - python3 "${SCRIPT_DIR}/loadgen.py" \ - "https://${APP_FQDN}/api/orders" \ - --requests 90 \ - --concurrency 8 \ - --expect-status 200 \ - --timeout 15 \ - --output "${EVIDENCE_DIR}/load.json" - ;; - s3) - ALERT_RULE_NAME="alert-sre-lab-s3-storage-rbac" - ROLE_ASSIGNMENT_ID="${STORAGE_CONTAINER_SCOPE}/providers/Microsoft.Authorization/roleAssignments/${BLOB_ROLE_ASSIGNMENT_NAME}" - readonly ROLE_ASSIGNMENT_ID - INJECTED_AT="$(utc_now)" - az role assignment delete --ids "${ROLE_ASSIGNMENT_ID}" - ROLE_DELETED_AT="$(utc_now)" - if ! wait_for_s3_fault; then - S3_PROPAGATION_FAILURE="Storage RBAC deletion did not produce HTTP 503 within ${S3_PROPAGATION_TIMEOUT_SECONDS}s." - record_failed_run "${S3_PROPAGATION_FAILURE}" - exit 1 - fi - python3 "${SCRIPT_DIR}/loadgen.py" \ - "https://${APP_FQDN}/api/documents" \ - --requests 60 \ - --concurrency 4 \ - --expect-status 503 \ - --output "${EVIDENCE_DIR}/load.json" - ;; -esac -readonly ALERT_RULE_NAME - -started="${SECONDS}" -while (( SECONDS - started < ALERT_FIRE_TIMEOUT_SECONDS )); do - ALERT_LIST_POLLS=$((ALERT_LIST_POLLS + 1)) - if ! az rest \ - --method get \ - --url "https://management.azure.com/subscriptions/${SUBSCRIPTION_ID}/providers/Microsoft.AlertsManagement/alerts?api-version=2019-03-01&targetResourceGroup=${RESOURCE_GROUP}&monitorCondition=Fired" \ - >"${ALERT_LIST_RESPONSE_FILE}" 2>"${ALERT_LIST_ATTEMPT_ERROR_FILE}"; then - ALERT_LIST_FAILURES=$((ALERT_LIST_FAILURES + 1)) - mv "${ALERT_LIST_ATTEMPT_ERROR_FILE}" "${ALERT_LIST_ERROR_FILE}" - echo "Azure Alerts list request failed; the last error is in ${ALERT_LIST_ERROR_FILE}. Retrying within the ${ALERT_FIRE_TIMEOUT_SECONDS}s budget." >&2 - sleep "${ALERT_FIRE_POLL_INTERVAL_SECONDS}" - continue - fi - rm -f "${ALERT_LIST_ATTEMPT_ERROR_FILE}" - if [[ ! -s "${ALERT_LIST_RESPONSE_FILE}" ]]; then - ALERT_LIST_FAILURES=$((ALERT_LIST_FAILURES + 1)) - ALERT_ID="" - echo "Azure Alerts list returned an empty response." >"${ALERT_LIST_ERROR_FILE}" - echo "Azure Alerts list returned an empty response; details are in ${ALERT_LIST_ERROR_FILE}. Retrying within the ${ALERT_FIRE_TIMEOUT_SECONDS}s budget." >&2 - sleep "${ALERT_FIRE_POLL_INTERVAL_SECONDS}" - continue - fi - alerts_json="$(cat "${ALERT_LIST_RESPONSE_FILE}")" - candidate_alert_id="" - if ! candidate_alert_id="$(jq -r \ - --arg rule "${ALERT_RULE_NAME}" \ - 'first(.value[] | select(.properties.essentials.alertRule | endswith($rule)) | .id) // ""' \ - <<<"${alerts_json}" 2>/dev/null)"; then - ALERT_LIST_FAILURES=$((ALERT_LIST_FAILURES + 1)) - ALERT_ID="" - printf 'Azure Alerts list returned invalid JSON. Response: %s\n' "${alerts_json}" >"${ALERT_LIST_ERROR_FILE}" - echo "Azure Alerts list returned invalid JSON; details are in ${ALERT_LIST_ERROR_FILE}. Retrying within the ${ALERT_FIRE_TIMEOUT_SECONDS}s budget." >&2 - sleep "${ALERT_FIRE_POLL_INTERVAL_SECONDS}" - continue - fi - ALERT_LIST_VALID_RESPONSES=$((ALERT_LIST_VALID_RESPONSES + 1)) - ALERT_ID="${candidate_alert_id}" - if [[ -n "${ALERT_ID}" ]]; then - ALERT_FIRED_AT="$(jq -r \ - --arg id "${ALERT_ID}" \ - '.value[] | select(.id == $id) | .properties.essentials.startDateTime' \ - <<<"${alerts_json}")" - break - fi - sleep "${ALERT_FIRE_POLL_INTERVAL_SECONDS}" -done - -# No alert ever firing is a distinct failure from one that fires and never -# resolves, but it is just as unusable as evidence: recorded as a failed -# run -- with the evidence directory and a reason -- before this exits, so -# a later `lab.sh score` or the next scenario's gate can never read this -# attempt as anything but failed. The EXIT trap (still armed here) still -# reverts whatever was injected above. -if [[ -z "${ALERT_ID}" ]]; then - if [[ "${ALERT_LIST_VALID_RESPONSES}" -eq 0 && "${ALERT_LIST_FAILURES}" -gt 0 ]]; then - ALERT_NEVER_FIRED_REASON="could not query Azure Alerts after ${ALERT_LIST_FAILURES} failed attempts within ${ALERT_FIRE_TIMEOUT_SECONDS}s; last error: ${ALERT_LIST_ERROR_FILE}" - else - ALERT_NEVER_FIRED_REASON="alert ${ALERT_RULE_NAME} did not fire within ${ALERT_FIRE_TIMEOUT_SECONDS}s." - if [[ "${ALERT_LIST_FAILURES}" -gt 0 ]]; then - ALERT_NEVER_FIRED_REASON="${ALERT_NEVER_FIRED_REASON} ${ALERT_LIST_FAILURES} of ${ALERT_LIST_POLLS} polls failed; last error: ${ALERT_LIST_ERROR_FILE}" - fi - fi - echo "${ALERT_NEVER_FIRED_REASON}" >&2 - echo "Evidence directory: ${EVIDENCE_DIR}" >&2 - record_failed_run "${ALERT_NEVER_FIRED_REASON}" - exit 1 -fi - -recover -RECOVERED_AT="$(utc_now)" -readonly RECOVERED_AT - -# Recovery is only real when the workload is healthy again *and* Azure -# Monitor closed the alert this run fired. Both are waited for before any -# state transition, so a timeout leaves the scenario failed and the next -# scenario blocked instead of recording a recovery nobody confirmed. -RECOVERY_CONFIRMED=1 -RECOVERY_FAILURE="" -if ! wait_for_app_ready "${APP_NAME}" "${RECOVERY_HEALTH_TIMEOUT_SECONDS}"; then - RECOVERY_CONFIRMED=0 - RECOVERY_FAILURE="workload did not become healthy within ${RECOVERY_HEALTH_TIMEOUT_SECONDS}s" -fi - -ALERT_RESOLVED_AT="" -if [[ "${RECOVERY_CONFIRMED}" -eq 1 ]]; then - if ALERT_RESOLVED_AT="$(wait_for_alert_resolved \ - "${ALERT_ID}" \ - "${ALERT_RESOLVE_TIMEOUT_SECONDS}" \ - "${ALERT_RESOLVE_POLL_INTERVAL_SECONDS}")"; then - : - else - ALERT_RESOLVED_AT="" - RECOVERY_CONFIRMED=0 - RECOVERY_FAILURE="alert ${ALERT_RULE_NAME} was not Resolved within ${ALERT_RESOLVE_TIMEOUT_SECONDS}s" - fi -fi -readonly ALERT_RESOLVED_AT RECOVERY_CONFIRMED RECOVERY_FAILURE - -jq -n \ - --arg scenario "${SCENARIO}" \ - --arg injectedAt "${INJECTED_AT}" \ - --arg revisionReadyAt "${REVISION_READY_AT}" \ - --arg roleDeletedAt "${ROLE_DELETED_AT}" \ - --arg alertRule "${ALERT_RULE_NAME}" \ - --arg alertId "${ALERT_ID}" \ - --arg alertFiredAt "${ALERT_FIRED_AT}" \ - --arg recoveredAt "${RECOVERED_AT}" \ - --arg alertResolvedAt "${ALERT_RESOLVED_AT}" \ - '{ - scenario: $scenario, - injected_at: $injectedAt, - revision_ready_at: (if $revisionReadyAt == "" then null else $revisionReadyAt end), - role_deleted_at: (if $roleDeletedAt == "" then null else $roleDeletedAt end), - alert_rule: $alertRule, - alert_id: $alertId, - alert_fired_at: $alertFiredAt, - recovered_at: $recoveredAt, - alert_resolved_at: (if $alertResolvedAt == "" then null else $alertResolvedAt end) - }' | tee "${EVIDENCE_DIR}/timeline.json" - -if [[ "${RECOVERY_CONFIRMED}" -ne 1 ]]; then - echo "Scenario ${SCENARIO} is recorded as failed: ${RECOVERY_FAILURE}." >&2 - echo "Evidence directory: ${EVIDENCE_DIR}" >&2 - record_failed_run "${RECOVERY_FAILURE}" - exit 1 -fi - -lab_state mark-recovered "${SCENARIO}" "${EVIDENCE_DIR}" -OUTCOME_RECORDED=1 - -printf 'Evidence directory: %s\n' "${EVIDENCE_DIR}" diff --git a/monitor/sre-agent-event-lab/scripts/score.py b/monitor/sre-agent-event-lab/scripts/score.py index 01fef96..edf0b5c 100755 --- a/monitor/sre-agent-event-lab/scripts/score.py +++ b/monitor/sre-agent-event-lab/scripts/score.py @@ -25,7 +25,7 @@ Output is `evidence/scorecard.json` plus a tab-separated table (`SCENARIOCRITERIONSTATUSPOINTSDETAIL`), the same -machine-readable shape `doctor.sh` prints. +machine-readable shape the evidence files carry. Python 3.9 compatible: no PEP 604 unions and no third-party imports. """ @@ -332,8 +332,8 @@ def main(argv: Optional[Sequence[str]] = None) -> int: state = state_from_environment(state_path) if not any(state.capture_status(scenario) for scenario in SCENARIOS): print( - "No captured scenario evidence in {0}. Run: lab.sh run s1, " - "then lab.sh capture s1.".format(evidence_root), + "No captured scenario evidence in {0}. Run the s1 steps in " + "guides/02-scenario-s1.md, capture included.".format(evidence_root), file=sys.stderr, ) return 1 diff --git a/monitor/sre-agent-event-lab/scripts/setup-venv.sh b/monitor/sre-agent-event-lab/scripts/setup-venv.sh index fed9df3..b0489ba 100755 --- a/monitor/sre-agent-event-lab/scripts/setup-venv.sh +++ b/monitor/sre-agent-event-lab/scripts/setup-venv.sh @@ -5,7 +5,8 @@ set -euo pipefail # needs: the app's own runtime dependencies (`requirements.txt`, pulled in # by `-r requirements.txt` at the top of `requirements-dev.txt`), Pillow for # `render_capture.py`'s PNG/GIF rendering, and pytest/httpx for `app/tests`. -# `capture-scenario.sh` and `guides/05-results.md`'s notification step both +# The capture step in each scenario guide and `guides/05-results.md`'s +# notification step both # run under this interpreter. # # `uv` is mandatory here, not merely preferred: this lab runs behind a diff --git a/monitor/sre-agent-event-lab/scripts/tests/azd_fake.py b/monitor/sre-agent-event-lab/scripts/tests/azd_fake.py index 72d1283..3809283 100644 --- a/monitor/sre-agent-event-lab/scripts/tests/azd_fake.py +++ b/monitor/sre-agent-event-lab/scripts/tests/azd_fake.py @@ -19,7 +19,7 @@ rc=0, stdout='{"status": "success", "expiresOn": "2026-08-14T07:57:15Z"}' ``` -Three properties matter for `common.sh`/`doctor.sh` and are therefore +Three properties matter for `common.sh` and its callers, and are therefore modelled here: 1. azd writes its `ERROR:` diagnostics to **stdout**, not stderr, and signals diff --git a/monitor/sre-agent-event-lab/scripts/tests/cleanup_harness.py b/monitor/sre-agent-event-lab/scripts/tests/cleanup_harness.py index 99f54de..9af25b2 100644 --- a/monitor/sre-agent-event-lab/scripts/tests/cleanup_harness.py +++ b/monitor/sre-agent-event-lab/scripts/tests/cleanup_harness.py @@ -99,7 +99,7 @@ def agent_setup( uami_assignment_id=None, uami_principal_id=AGENT_UAMI_PRINCIPAL_ID, ): - """The evidence file `lab.sh acknowledge agent-setup` leaves behind.""" + """The evidence file `lab_state.py acknowledge-agent` leaves behind.""" if monitoring_assignment_id is None: monitoring_assignment_id = assignment_id(AGENT_ASSIGNMENT_NAME) if uami_assignment_id is None: diff --git a/monitor/sre-agent-event-lab/scripts/tests/doctor_harness.py b/monitor/sre-agent-event-lab/scripts/tests/doctor_harness.py deleted file mode 100644 index 9aadd67..0000000 --- a/monitor/sre-agent-event-lab/scripts/tests/doctor_harness.py +++ /dev/null @@ -1,550 +0,0 @@ -"""Fake-CLI harness for `doctor.sh`, `baseline.sh`, and `lab.sh`. - -`lab_script_harness.py`'s generic `az` fake models `run-scenario.sh` / -`query-evidence.sh` / `capture-scenario.sh` / `cleanup.sh`'s call surface. It -does not model the surfaces `doctor.sh` and `baseline.sh` add: a container -app's *current* health state (no polling loop), a `curl` probe of -`/healthz`, `az extension show` for the `log-analytics` extension, -`az resource show` for the SRE Agent resource, a per-rule `az rest` read of -`Microsoft.Insights/scheduledQueryRules`, and `az role assignment list` -keyed by a specific `--assignee-object-id`. This module gives each test -full, mutable control over that state through a single `FakeAz` object so -`doctor.sh`/`baseline.sh`/`lab.sh` are driven as real programs -- not -grepped as text -- exactly like the other lab scripts. - -Three observable contracts here were re-verified against the *real* CLIs on -2026-08-14 (azure-cli 2.86.0 / log-analytics 1.0.0b1 / azd 1.29.0) rather -than assumed, because each one had been modelled incorrectly before: - -1. `az monitor log-analytics query -o json` prints a flat JSON array of row - objects, not the `{"tables": [...]}` REST envelope (see `_rows`). -2. `az role assignment list` hides parent-scope assignments unless - `--include-inherited` is passed. -3. `azd auth login --check-status` always exits 0 and reports the real - answer only in its output (see `azd_fake.py`). -""" -import json -import os -import shutil -import subprocess -from dataclasses import dataclass, field -from pathlib import Path -from typing import Dict - -from azd_fake import write_azd_stub, write_executable -from lab_script_harness import ( - ENV_NAME, - REAL_PYTHON, - RESOURCE_GROUP, - SCRIPTS_DIR, - SUBSCRIPTION_ID, -) - - -BASH = shutil.which("bash") or "/bin/bash" - -APP_NAME = "ca-sre-lab" -APP_FQDN = "ca-sre-lab.example.azurecontainerapps.io" -WORKSPACE_CUSTOMER_ID = "9d1a0b2c-3d4e-5f60-7182-93a4b5c6d7e8" -TELEMETRY_SERVICE_NAME = "sre-lab-order-api" -AGENT_PRINCIPAL_ID = "8c8a4f0e-0000-4000-8000-2b1f9a0c1234" -AGENT_UAMI_PRINCIPAL_ID = "9c8a4f0e-1111-4000-8000-2b1f9a0c5678" - -ALERT_RULE_NAMES = ( - "alert-sre-lab-s1-http500", - "alert-sre-lab-s2-latency", - "alert-sre-lab-s3-storage-rbac", -) - -RESOURCE_GROUP_SCOPE = f"/subscriptions/{SUBSCRIPTION_ID}/resourceGroups/{RESOURCE_GROUP}" -SUBSCRIPTION_SCOPE = f"/subscriptions/{SUBSCRIPTION_ID}" - -# `az monitor log-analytics query -o json` does NOT print the REST envelope -# (`{"tables": [...]}`). The `log-analytics` extension's `Query._output` -# flattens every table into one JSON array with a single object per row -- -# `TableName` plus one *stringified* value per projected column -- so an -# empty result set prints exactly `[]`. Verified against the installed -# extension source (log-analytics 1.0.0b1, azure-cli 2.86.0): -# `~/.azure/cliextensions/log-analytics/azext_loganalytics/custom.py`. -def _rows(*rows) -> str: - return json.dumps(list(rows)) - - -APP_REQUESTS_ROW = { - "TableName": "PrimaryResult", - "TimeGenerated": "2026-08-14T00:05:00Z", - "Name": "GET /api/orders", -} -DOCUMENT_REQUESTS_ROW = { - "TableName": "PrimaryResult", - "TimeGenerated": "2026-08-14T00:06:00Z", - "Name": "GET /api/documents", -} -EMPTY_RESULT = "[]" - -# What KQL's `count` operator really returns: exactly one row, whatever the -# data looks like. A fake that answers `| count` with an empty array would -# hide the very defect the doctor/baseline telemetry checks must not have. -COUNT_ZERO_RESULT = _rows({"TableName": "PrimaryResult", "Count": "0"}) - - -def _default_alert_rules_enabled() -> Dict[str, bool]: - return {name: True for name in ALERT_RULE_NAMES} - - -def _no_inherited_reader() -> Dict[str, bool]: - return {AGENT_PRINCIPAL_ID: False, AGENT_UAMI_PRINCIPAL_ID: False} - - -@dataclass -class FakeAz: - """Mutable state for the fake `az`/`azd`/`curl`/`python3` a run uses. - - Every field defaults to a fully healthy lab so a test only has to set - the one attribute it wants to exercise. `workdir` is normally the - fixture's `tmp_path`. State is (re)materialized each time - `run_doctor`/`run_baseline`/`run_lab_cli` is called, so mutating a - field *after* the fixture is created (as the brief's examples do) is - honoured; a lab directory that already exists for `workdir` is reused - (not wiped), so evidence a test or a prior call wrote survives. - """ - - workdir: Path - logged_in: bool = True - azd_logged_in: bool = True - log_analytics_extension_installed: bool = True - active_subscription_id: str = SUBSCRIPTION_ID - resource_group_exists: bool = True - resource_group_purpose: str = "sre-agent-event-lab" - resource_group_env_tag: str = ENV_NAME - container_app_health: str = "Healthy" - healthz_status: int = 200 - app_insights_has_recent_requests: bool = True - app_insights_orders_seen: bool = True - app_insights_documents_seen: bool = True - alert_rules_present: Dict[str, bool] = field(default_factory=_default_alert_rules_enabled) - alert_rules_enabled: Dict[str, bool] = field(default_factory=_default_alert_rules_enabled) - sre_agent_resource_exists: bool = True - agent_setup_present: bool = True - agent_setup_body: "str | None" = None - agent_principal_id: str = AGENT_PRINCIPAL_ID - agent_uami_principal_id: str = AGENT_UAMI_PRINCIPAL_ID - # Reader assigned *directly* on the lab resource group. - reader_role_assigned: Dict[str, bool] = field( - default_factory=lambda: {AGENT_PRINCIPAL_ID: True, AGENT_UAMI_PRINCIPAL_ID: True} - ) - # Reader assigned on the subscription and therefore only visible to a - # lookup that asks for inherited assignments. - reader_role_inherited: Dict[str, bool] = field(default_factory=_no_inherited_reader) - baseline_orders_succeed: bool = True - baseline_documents_succeed: bool = True - # Whether `app/.venv/bin/python` exists at all, and whether Pillow is - # importable from it -- the two facts `scripts/setup-venv.sh` (run from - # `postprovision`) is responsible for making true, and doctor's "Python - # environment" check reports on. Both default to a fully set-up venv so - # only the tests exercising this check need to touch either field. - venv_present: bool = True - pillow_importable: bool = True - # None means "use the module default AZD_VALUES"; a test passes {} (or - # a partial dict) to exercise the missing/partial-configuration paths. - azd_values: "Dict[str, str] | None" = None - - -def _bool_json(value: bool) -> str: - return "true" if value else "false" - - -def _az_stub_source(fake_az: FakeAz, log_path: Path) -> str: - rule_branches = [] - for rule_name in ALERT_RULE_NAMES: - present = fake_az.alert_rules_present.get(rule_name, True) - enabled = fake_az.alert_rules_enabled.get(rule_name, True) - if not present: - rule_branches.append(f' *"/scheduledqueryrules/{rule_name}?"*) exit 1 ;;') - else: - rule_branches.append( - f' *"/scheduledqueryrules/{rule_name}?"*) ' - f'printf \'{{"properties": {{"enabled": {_bool_json(enabled)}}}}}\\n\' ;;' - ) - rule_case = "\n".join(rule_branches) - - # `az role assignment list` only returns assignments made at *parent* - # scopes when `--include-inherited` is passed; without it, a Reader - # granted on the subscription is invisible to a resource-group scoped - # lookup. The fake reproduces that, so a doctor that drops the flag - # cannot pass the inherited-Reader test by accident. - principal_ids = set(fake_az.reader_role_assigned) | set(fake_az.reader_role_inherited) - reader_branches = [] - for principal_id in sorted(principal_ids): - direct = [] - inherited = [] - if fake_az.reader_role_assigned.get(principal_id, False): - direct.append( - { - "principalId": principal_id, - "roleDefinitionName": "Reader", - "scope": RESOURCE_GROUP_SCOPE, - } - ) - if fake_az.reader_role_inherited.get(principal_id, False): - inherited.append( - { - "principalId": principal_id, - "roleDefinitionName": "Reader", - "scope": SUBSCRIPTION_SCOPE, - } - ) - reader_branches.append( - f' "{principal_id}")\n' - f" if [[ \"${{all_args}}\" == *--include-inherited* ]]; then\n" - f" printf '%s\\n' '{json.dumps(direct + inherited)}'\n" - f" else\n" - f" printf '%s\\n' '{json.dumps(direct)}'\n" - f" fi ;;" - ) - reader_branches.append(" *) printf '[]\\n' ;;") - reader_case = "\n".join(reader_branches) - - orders_rows = _rows(APP_REQUESTS_ROW) if fake_az.app_insights_orders_seen else EMPTY_RESULT - documents_rows = ( - _rows(DOCUMENT_REQUESTS_ROW) if fake_az.app_insights_documents_seen else EMPTY_RESULT - ) - any_rows = _rows(APP_REQUESTS_ROW) if fake_az.app_insights_has_recent_requests else EMPTY_RESULT - account_show = ( - f"printf '%s\\n' '{fake_az.active_subscription_id}'" if fake_az.logged_in else "exit 1" - ) - group_exists = "printf 'true\\n'" if fake_az.resource_group_exists else "printf 'false\\n'" - resource_show = "exit 0" if fake_az.sre_agent_resource_exists else "exit 1" - log_analytics_extension = "exit 0" if fake_az.log_analytics_extension_installed else ( - "printf 'ERROR: The extension log-analytics is not installed.\\n' >&2\n exit 1" - ) - - return f"""#!/usr/bin/env bash -printf '%s\\n' "$*" >> "{log_path}" -all_args="$*" -case "${{1:-}} ${{2:-}}" in - "account show") - {account_show} - ;; - "extension show") - if [[ "${{all_args}}" == *"--name log-analytics"* ]]; then - {log_analytics_extension} - fi - ;; - "group exists") - {group_exists} - ;; - "group show") - if [[ "$*" == *"azd-env-name"* ]]; then - printf '%s\\n' "{fake_az.resource_group_env_tag}" - else - printf '%s\\n' "{fake_az.resource_group_purpose}" - fi - ;; - "containerapp revision") - printf '%s\\n' "{fake_az.container_app_health}" - ;; - "monitor log-analytics") - # KQL's `count` operator always returns exactly one row, even for an - # empty table -- reproduced here so any caller that infers "data - # exists" from a `| count` result's row count fails loudly. - if [[ "${{all_args}}" == *"| count"* ]]; then - printf '%s\\n' '{COUNT_ZERO_RESULT}' - elif [[ "${{all_args}}" == *"/api/orders"* ]]; then - printf '%s\\n' '{orders_rows}' - elif [[ "${{all_args}}" == *"/api/documents"* ]]; then - printf '%s\\n' '{documents_rows}' - else - printf '%s\\n' '{any_rows}' - fi - ;; - "rest --method") - case "$*" in -{rule_case} - *) printf '{{}}\\n' ;; - esac - ;; - "resource show") - {resource_show} - ;; - "role assignment") - principal="${{all_args##*--assignee-object-id }}" - principal="${{principal%% *}}" - case "${{principal}}" in -{reader_case} - esac - ;; - *) - : ;; -esac -exit 0 -""" - - -def _curl_stub_source(fake_az: FakeAz) -> str: - return f"""#!/usr/bin/env bash -printf '%s' "{fake_az.healthz_status}" -""" - - -def _python3_stub_source(fake_az: FakeAz, log_path: Path, is_venv_python: bool = False) -> str: - """Fake `python3`/`.venv/bin/python` for `loadgen.py`: writes a minimal, - valid summary and exits with loadgen's real contract (0 success, 2 a - request mismatch), keyed by the target URL so orders/documents can be - made to succeed or fail independently. - - Every other script -- notably `lab_state.py` and `score.py`, whose - behaviour these tests are checking -- runs under the real interpreter, - so a lab script that records or reads state is exercised, not faked. - - When `is_venv_python` is set, this stub also answers the - `-c "import PIL"` probe that `scripts/setup-venv.sh` and doctor's - "Python environment" check use to decide whether Pillow is importable - from `app/.venv`, per `fake_az.pillow_importable` -- loadgen never sends - that probe, so it cannot collide with the loadgen branch below.""" - orders_ok = 1 if fake_az.baseline_orders_succeed else 0 - documents_ok = 1 if fake_az.baseline_documents_succeed else 0 - pil_probe = "" - if is_venv_python: - pil_exit = 0 if fake_az.pillow_importable else 1 - pil_probe = f"""if [[ "${{1:-}}" == "-c" && "${{2:-}}" == *PIL* ]]; then - exit {pil_exit} -fi -""" - return f"""#!/usr/bin/env bash -printf '%s\\n' "$*" >> "{log_path}" -{pil_probe}case "${{1:-}}" in - *loadgen.py) ;; - *) exec "{REAL_PYTHON}" "$@" ;; -esac -output="" -requests=1 -args=("$@") -for ((i=0; i<${{#args[@]}}; i++)); do - case "${{args[$i]}}" in - --output) output="${{args[$((i+1))]}}" ;; - --requests) requests="${{args[$((i+1))]}}" ;; - esac -done -succeed=1 -if [[ "$*" == *"/api/orders"* ]]; then - succeed={orders_ok} -elif [[ "$*" == *"/api/documents"* ]]; then - succeed={documents_ok} -fi -if [[ -n "${{output}}" ]]; then - mkdir -p "$(dirname "${{output}}")" - printf '{{"total": %s, "errors": 0}}\\n' "${{requests}}" > "${{output}}" -fi -if [[ "${{succeed}}" -eq 1 ]]; then - exit 0 -else - exit 2 -fi -""" - - -AZD_VALUES = { - "AZURE_SUBSCRIPTION_ID": SUBSCRIPTION_ID, - "AZURE_RESOURCE_GROUP": RESOURCE_GROUP, - "AZURE_ENV_NAME": ENV_NAME, - "AZURE_LOCATION": "koreacentral", - "AZURE_CONTAINER_APP_NAME": APP_NAME, - "AZURE_CONTAINER_APP_FQDN": APP_FQDN, - "AZURE_STORAGE_CONTAINER_SCOPE": ( - f"/subscriptions/{SUBSCRIPTION_ID}/resourceGroups/{RESOURCE_GROUP}" - "/providers/Microsoft.Storage/storageAccounts/stsrelab/blobServices/default" - "/containers/documents" - ), - "AZURE_BLOB_ROLE_ASSIGNMENT_NAME": "3f2504e0-4f89-11d3-9a0c-0305e82c3301", - "AZURE_WORKSPACE_ID": ( - f"/subscriptions/{SUBSCRIPTION_ID}/resourceGroups/{RESOURCE_GROUP}" - "/providers/Microsoft.OperationalInsights/workspaces/log-sre-lab" - ), - "AZURE_APP_INSIGHTS_NAME": "appi-sre-lab", - "AZURE_TELEMETRY_SERVICE_NAME": TELEMETRY_SERVICE_NAME, - "containerAppPrincipalId": "8c8a4f0e-aaaa-4000-8000-2b1f9a0c1234", - "workspaceCustomerId": WORKSPACE_CUSTOMER_ID, -} - - -class LabRun: - def __init__( - self, - lab: Path, - bin_dir: Path, - workdir: Path, - az_log: Path, - python_log: Path, - azd_log: Path, - ): - self.lab = lab - self.bin_dir = bin_dir - self.workdir = workdir - self.az_log = az_log - self.python_log = python_log - self.azd_log = azd_log - - def run(self, script_name, args=(), env=None, stdin=None): - process_env = { - "PATH": f"{self.bin_dir}{os.pathsep}{os.environ.get('PATH', '')}", - "HOME": os.environ.get("HOME", str(self.lab)), - } - process_env.update(env or {}) - return subprocess.run( - [BASH, str(self.lab / "scripts" / script_name), *args], - capture_output=True, - text=True, - input=stdin, - env=process_env, - cwd=str(self.workdir), - ) - - def az_calls(self): - return self.az_log.read_text() if self.az_log.exists() else "" - - def azd_calls(self): - return self.azd_log.read_text() if self.azd_log.exists() else "" - - def evidence_dir(self): - return self.lab / "evidence" - - -def _write_agent_setup(lab: Path, fake_az: FakeAz) -> None: - agent_setup_path = lab / "evidence" / "agent-setup.json" - if not fake_az.agent_setup_present: - agent_setup_path.unlink(missing_ok=True) - return - if fake_az.agent_setup_body is not None: - agent_setup_path.write_text(fake_az.agent_setup_body) - return - setup = { - "agent_endpoint": "https://sre-agent.example.com/api/incidents", - "monitoring_contributor_assignment_id": ( - f"/subscriptions/{SUBSCRIPTION_ID}/providers" - "/Microsoft.Authorization/roleAssignments/principal-one" - ), - "agent_principal_id": fake_az.agent_principal_id, - "uami_monitoring_contributor_assignment_id": ( - f"/subscriptions/{SUBSCRIPTION_ID}/providers" - "/Microsoft.Authorization/roleAssignments/principal-two" - ), - "agent_user_assigned_principal_id": fake_az.agent_uami_principal_id, - } - agent_setup_path.write_text(json.dumps(setup)) - - -def _materialize(fake_az: FakeAz) -> LabRun: - """Create (once) or refresh the throwaway lab + fake CLIs for `fake_az`. - - The lab's `scripts/` and `evidence/` directories are only created the - first time a given `fake_az.workdir` is used, so evidence written by an - earlier call (or by the test itself) survives across repeated - `run_doctor`/`run_lab_cli` calls with the same `fake_az`. The fake - `az`/`curl`/`python3` executables are always rewritten so the latest - mutations to `fake_az` take effect immediately. - """ - tmp_path = fake_az.workdir - lab = tmp_path / "lab" - if not lab.exists(): - shutil.copytree( - SCRIPTS_DIR, - lab / "scripts", - ignore=shutil.ignore_patterns("tests", "__pycache__"), - ) - (lab / "azure.yaml").write_text("name: sre-agent-event-lab\n") - (lab / "evidence").mkdir() - - _write_agent_setup(lab, fake_az) - - bin_dir = tmp_path / "bin" - bin_dir.mkdir(exist_ok=True) - az_log = tmp_path / "az-calls.log" - python_log = tmp_path / "python-calls.log" - azd_log = tmp_path / "azd-calls.log" - - write_executable(bin_dir / "az", _az_stub_source(fake_az, az_log)) - azd_values = fake_az.azd_values if fake_az.azd_values is not None else AZD_VALUES - write_azd_stub(bin_dir, azd_values, "azd_1_29", azd_log, logged_in=fake_az.azd_logged_in) - write_executable(bin_dir / "curl", _curl_stub_source(fake_az)) - write_executable(bin_dir / "python3", _python3_stub_source(fake_az, python_log)) - - # `app/.venv` is always removed and recreated (rather than only - # `mkdir -p`'d once) so a test that flips `venv_present`/ - # `pillow_importable` between two `run_doctor` calls on the same - # `fake_az` -- exactly like every other mutable field here -- is - # honoured on the second call too. - venv_dir = lab / "app" / ".venv" - shutil.rmtree(venv_dir, ignore_errors=True) - if fake_az.venv_present: - venv_bin = venv_dir / "bin" - venv_bin.mkdir(parents=True, exist_ok=True) - write_executable( - venv_bin / "python", - _python3_stub_source(fake_az, python_log, is_venv_python=True), - ) - - workdir = tmp_path / "elsewhere" - workdir.mkdir(exist_ok=True) - - return LabRun(lab, bin_dir, workdir, az_log, python_log, azd_log) - - -def _split_env(env_overrides): - env = {key.upper(): value for key, value in env_overrides.items() if key != "env"} - if "env" in env_overrides: - env.update(env_overrides["env"]) - return env - - -def run_doctor(fake_az: FakeAz, **env_overrides) -> subprocess.CompletedProcess: - """Run `doctor.sh` against `fake_az`'s current state. - - `env_overrides` keys use the process-environment names `common.sh` - reads, lower-cased for readability (e.g. `sre_agent_resource_id="..."` - becomes `SRE_AGENT_RESOURCE_ID`). - """ - run = _materialize(fake_az) - return run.run("doctor.sh", env=_split_env(env_overrides)) - - -def run_baseline(fake_az: FakeAz, **env_overrides) -> subprocess.CompletedProcess: - run = _materialize(fake_az) - return run.run("baseline.sh", env=_split_env(env_overrides)) - - -def run_lab_cli(fake_az: FakeAz, args, stdin=None, **env_overrides) -> subprocess.CompletedProcess: - """Run `lab.sh` against `fake_az`'s current state. - - `stdin` feeds the interactive commands (`acknowledge agent-setup` reads - the operator's typed answer), so the acknowledgement is driven exactly - as a human drives it -- through the process's standard input.""" - run = _materialize(fake_az) - return run.run("lab.sh", args=args, stdin=stdin, env=_split_env(env_overrides)) - - -def state_path_for(fake_az: FakeAz) -> Path: - """Where `lab_state.py` keeps this lab's ordered-run state.""" - return lab_dir_for(fake_az) / "evidence" / "state.json" - - -def lab_dir_for(fake_az: FakeAz) -> Path: - """The lab directory `run_doctor`/`run_baseline`/`run_lab_cli` will use - for `fake_az`. A test can call this *before* the first run to pre-seed - evidence (e.g. a scenario's `timeline.json`) once the directory exists, - or after a run to inspect what the script produced.""" - return fake_az.workdir / "lab" - - -def az_calls_for(fake_az: FakeAz) -> str: - """Every `az` invocation logged for `fake_az.workdir`'s most recent - run, in order, one `argv` per line.""" - log_path = fake_az.workdir / "az-calls.log" - return log_path.read_text() if log_path.exists() else "" - - -def azd_calls_for(fake_az: FakeAz) -> str: - """Every `azd` invocation logged for `fake_az.workdir`'s most recent - run (argv, then `cwd=...`, one per line -- see `azd_fake.py`).""" - log_path = fake_az.workdir / "azd-calls.log" - return log_path.read_text() if log_path.exists() else "" diff --git a/monitor/sre-agent-event-lab/scripts/tests/lab_script_harness.py b/monitor/sre-agent-event-lab/scripts/tests/lab_script_harness.py index bea2a24..9a7689a 100644 --- a/monitor/sre-agent-event-lab/scripts/tests/lab_script_harness.py +++ b/monitor/sre-agent-event-lab/scripts/tests/lab_script_harness.py @@ -59,7 +59,7 @@ def _alert(scenario, alert_uuid): """One entry of the Alerts Management list response. - All three lab rules are listed, because `run-scenario.sh` picks the + All three lab rules are listed, because a scenario picks the alert whose rule matches the scenario it just ran: a list containing only S1's alert would let an S2 run pass on S1's evidence. """ @@ -125,7 +125,7 @@ def _az_stub_source(log_path, state_dir): call (clearing the failure mode/delay, or restoring the blob role) flips it to `Resolved` -- unless `${state}/alert_stays_fired` exists, which reproduces an alert that never closes. That is the only way to - exercise the recovery gate honestly: `run-scenario.sh` must not record + exercise the recovery gate honestly: a scenario must not record a recovery Azure Monitor never confirmed. Recovery itself can fail, which is the other half of that gate. Three @@ -290,7 +290,7 @@ def _lab_python_stub_source( `score.py`, which are the behaviour under test) runs under the real interpreter. - Also answers the `-c "import PIL"` probe `capture-scenario.sh` and + Also answers the `-c "import PIL"` probe the capture step and doctor's "Python environment" check use to verify Pillow is importable, per `pillow_importable` -- explicitly, rather than delegating to whichever real interpreter happens to run the test suite, so this fake diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_baseline.py b/monitor/sre-agent-event-lab/scripts/tests/test_baseline.py deleted file mode 100644 index 9908a1b..0000000 --- a/monitor/sre-agent-event-lab/scripts/tests/test_baseline.py +++ /dev/null @@ -1,121 +0,0 @@ -"""Behavioural tests for `baseline.sh`. - -`baseline.sh` is the one script that has to decide "did my traffic actually -reach Application Insights?" from a Log Analytics answer, so its parsing of -the real `az monitor log-analytics query` output shape -- a flat JSON array -of row objects, never the `{"tables": [...]}` REST envelope -- and the bound -on its polling loop are the two properties worth pinning. Both are exercised -by running the script as a real program against the fake CLIs in -`doctor_harness.py`. -""" -import json -import time - -import pytest - -from doctor_harness import FakeAz, az_calls_for, lab_dir_for, run_baseline - - -# Long enough to allow several poll rounds, short enough that a genuinely -# unbounded loop fails the test instead of hanging the suite. -TELEMETRY_TIMEOUT_SECONDS = "5" -POLL_INTERVAL_SECONDS = "1" - - -@pytest.fixture -def fake_az(tmp_path): - return FakeAz(workdir=tmp_path) - - -def run_bounded_baseline(fake_az, **env_overrides): - return run_baseline( - fake_az, - lab_baseline_telemetry_timeout_seconds=TELEMETRY_TIMEOUT_SECONDS, - lab_baseline_telemetry_poll_interval_seconds=POLL_INTERVAL_SECONDS, - **env_overrides, - ) - - -def telemetry_check(fake_az): - evidence_dirs = sorted((lab_dir_for(fake_az) / "evidence").glob("baseline-*")) - assert evidence_dirs, "baseline.sh wrote no evidence directory" - return json.loads((evidence_dirs[-1] / "telemetry-check.json").read_text()) - - -def analytics_queries(fake_az): - return [line for line in az_calls_for(fake_az).splitlines() if "monitor log-analytics" in line] - - -def test_baseline_succeeds_when_both_request_types_appear(fake_az): - """The healthy case: the workspace answers with a non-empty flat array - for both request types.""" - result = run_bounded_baseline(fake_az) - - assert result.returncode == 0, result.stdout + result.stderr - assert telemetry_check(fake_az) == { - "orders_telemetry_seen": True, - "documents_telemetry_seen": True, - "checked_at": telemetry_check(fake_az)["checked_at"], - } - - -def test_baseline_fails_when_the_workspace_returns_no_rows(fake_az): - """The empty case: `[]` for `/api/orders` must not be mistaken for data, - and the failure has to name what was and was not seen.""" - fake_az.app_insights_orders_seen = False - - result = run_bounded_baseline(fake_az) - - assert result.returncode != 0 - assert telemetry_check(fake_az)["orders_telemetry_seen"] is False - assert telemetry_check(fake_az)["documents_telemetry_seen"] is True - assert "orders=0" in result.stderr - - -def test_baseline_query_does_not_count_rows_of_a_count(fake_az): - """KQL `count` always returns exactly one row, so a row count taken from - a `| count` query reports "data exists" for an empty workspace.""" - run_bounded_baseline(fake_az) - - queries = analytics_queries(fake_az) - assert queries - for query in queries: - assert "| count" not in query, f"baseline query relies on `| count`: {query}" - - -def test_baseline_polling_is_bounded_by_the_timeout(fake_az): - """Telemetry that never arrives must end the run near the timeout, not - hang and not exit after a single try.""" - fake_az.app_insights_orders_seen = False - fake_az.app_insights_documents_seen = False - - started = time.monotonic() - result = run_bounded_baseline(fake_az) - elapsed = time.monotonic() - started - - assert result.returncode != 0 - assert elapsed < int(TELEMETRY_TIMEOUT_SECONDS) + 20, f"poll overran its bound: {elapsed}s" - assert len(analytics_queries(fake_az)) > 2, "baseline gave up without polling" - - -def test_baseline_polls_at_least_once_with_a_zero_timeout(fake_az): - """A degenerate timeout must still produce one honest attempt and a - telemetry-check record, never an unexplained silent pass.""" - result = run_baseline( - fake_az, - lab_baseline_telemetry_timeout_seconds="0", - lab_baseline_telemetry_poll_interval_seconds="1", - ) - - assert analytics_queries(fake_az), "baseline never queried the workspace" - assert result.returncode == 0, result.stdout + result.stderr - assert telemetry_check(fake_az)["orders_telemetry_seen"] is True - - -def test_baseline_records_evidence_for_both_load_phases(fake_az): - result = run_bounded_baseline(fake_az) - - assert result.returncode == 0, result.stdout + result.stderr - evidence_dir = sorted((lab_dir_for(fake_az) / "evidence").glob("baseline-*"))[-1] - assert (evidence_dir / "orders.json").is_file() - assert (evidence_dir / "documents.json").is_file() diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_cleanup_external.py b/monitor/sre-agent-event-lab/scripts/tests/test_cleanup_external.py index 4f4fa08..2993ba0 100644 --- a/monitor/sre-agent-event-lab/scripts/tests/test_cleanup_external.py +++ b/monitor/sre-agent-event-lab/scripts/tests/test_cleanup_external.py @@ -597,7 +597,7 @@ def test_readme_documents_the_teardown_hooks_and_manual_recovery(): "the README must not claim canceling leaves a fully working environment " "-- predown already removed the recorded roles before the prompt" ) - assert "acknowledge agent-setup" in section, ( + assert "acknowledge-agent" in section, ( "the README must name the recovery step (re-running Agent setup/role " "assignment) an operator needs after canceling" ) diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_common.py b/monitor/sre-agent-event-lab/scripts/tests/test_common.py index 15ede57..9b144ab 100644 --- a/monitor/sre-agent-event-lab/scripts/tests/test_common.py +++ b/monitor/sre-agent-event-lab/scripts/tests/test_common.py @@ -7,19 +7,13 @@ COMMON_SH = Path(__file__).parents[1] / "common.sh" -DEPLOY_SH = Path(__file__).parents[1] / "deploy.sh" -CLEANUP_SH = Path(__file__).parents[1] / "cleanup.sh" CLEANUP_EXTERNAL_SH = Path(__file__).parents[1] / "cleanup-external.sh" QUERY_EVIDENCE_SH = Path(__file__).parents[1] / "query-evidence.sh" -RUN_SCENARIO_SH = Path(__file__).parents[1] / "run-scenario.sh" -CAPTURE_SCENARIO_SH = Path(__file__).parents[1] / "capture-scenario.sh" -BASELINE_SH = Path(__file__).parents[1] / "baseline.sh" BASH = shutil.which("bash") or "/bin/bash" UUID_PATTERN = re.compile(r"\b[0-9a-fA-F]{8}-(?:[0-9a-fA-F]{4}-){3}[0-9a-fA-F]{12}\b") -LEGACY_RESOURCE_GROUP_FLAG = "--legacy-delete-resource-group" REQUIRED_ENV = { "AZURE_SUBSCRIPTION_ID": "11111111-2222-3333-4444-555555555555", @@ -125,22 +119,6 @@ def test_verify_subscription_reports_only_subscription_id_on_mismatch(tmp_path): assert REQUIRED_ENV["AZURE_SUBSCRIPTION_ID"] in result.stderr -def test_deploy_delegates_to_azd_up(): - """The subscription-scope templates deploy.sh used to deploy were removed - when the lab moved to azd, so deploy.sh must not reference them any more. - It stays as a thin compatibility wrapper so the documented command keeps - working. - """ - script = DEPLOY_SH.read_text() - - assert "azd up" in script - assert "subscription" + ".bicep" not in script - assert "az deployment sub validate" not in script - assert "az deployment sub create" not in script - assert "az deployment group" not in script - assert 'IMAGE_TAG="20260812.4"' not in script - - def test_no_tracked_lab_file_references_the_deleted_subscription_templates(): lab_root = Path(__file__).parents[2] tracked = subprocess.run( @@ -172,13 +150,13 @@ def test_readme_documents_a_working_deployment_command(): assert "azd env get-value AZURE_CONTAINER_APP_FQDN" in readme -def test_readme_documents_scenario_scripts_read_the_current_azd_environment(): - """`run-scenario.sh` and `query-evidence.sh` now resolve deployment - outputs through `common.sh`'s `load_lab_config` (explicit env > current - `azd env get-value` > allowed default), so they work against whatever - azd environment is currently selected -- not a fixed pre-azd resource - group. The README's scenario-execution section must describe that - mechanism instead of the old "legacy, not yet rewritten" caveat. +def test_readme_documents_that_evidence_collection_reads_the_current_azd_environment(): + """`query-evidence.sh` resolves deployment outputs through `common.sh`'s + `load_lab_config` (explicit env > current `azd env get-value` > allowed + default), so it works against whatever azd environment is currently + selected -- not a fixed pre-azd resource group. The README's + scenario-execution section must describe that mechanism instead of the + old "legacy, not yet rewritten" caveat. """ readme = (Path(__file__).parents[2] / "README.md").read_text() @@ -186,35 +164,22 @@ def test_readme_documents_scenario_scripts_read_the_current_azd_environment(): assert scenario_heading in readme section = readme.split(scenario_heading, 1)[1].split("##", 1)[0] - assert "run-scenario.sh" in section + assert "query-evidence.sh" in section assert "load_lab_config" in section assert "레거시" not in section, ( "README's scenario-execution section must no longer describe " - "run-scenario.sh/query-evidence.sh as reading a legacy, pre-azd " - "deployment lookup -- load_lab_config now reads the current azd " - "environment." + "query-evidence.sh as reading a legacy, pre-azd deployment lookup " + "-- load_lab_config now reads the current azd environment." ) -def test_scenario_waits_for_new_revision_before_load(): - common = COMMON_SH.read_text() - scenario = (Path(__file__).parents[1] / "run-scenario.sh").read_text() - # Line continuations are formatting, not behaviour: the call is checked - # with its own wrapping collapsed. - collapsed = " ".join(scenario.replace("\\\n", " ").split()) - - assert "wait_for_new_revision_ready()" in common - assert 'OLD_REVISION="$(latest_revision_name "${APP_NAME}")"' in scenario - assert 'wait_for_new_revision_ready "${APP_NAME}" "${OLD_REVISION}"' in collapsed - - def test_cleanup_removes_both_subscription_monitoring_assignments(): - """The verified deletion of the two subscription-scoped assignments now + """The verified deletion of the two subscription-scoped assignments lives in `cleanup-external.sh` -- the script `azd down`'s `predown` hook - runs and the one `cleanup.sh` forwards to -- so both recorded records, - their principals, the Monitoring Contributor role definition and the - read-back that verifies them must be there. `test_cleanup_external.py` - exercises the resulting behaviour against a staged subscription. + runs -- so both recorded records, their principals, the Monitoring + Contributor role definition and the read-back that verifies them must be + there. `test_cleanup_external.py` exercises the resulting behaviour + against a staged subscription. """ script = CLEANUP_EXTERNAL_SH.read_text() @@ -233,29 +198,6 @@ def test_cleanup_removes_both_subscription_monitoring_assignments(): assert "Nothing outside the azd resource group to clean up." in script -def test_s1_and_s2_record_injection_before_container_app_update(): - script = (Path(__file__).parents[1] / "run-scenario.sh").read_text() - main_case = script.rsplit('case "${SCENARIO}" in', 1)[1] - - for branch, next_branch in ((" s1)", " s2)"), (" s2)", " s3)")): - section = main_case.split(branch, 1)[1].split(next_branch, 1)[0] - assert section.index('INJECTED_AT="$(utc_now)"') < section.index( - "az containerapp update" - ) - assert 'REVISION_READY_AT="$(utc_now)"' in section - - -def test_s3_records_injection_before_role_deletion(): - script = (Path(__file__).parents[1] / "run-scenario.sh").read_text() - main_case = script.rsplit('case "${SCENARIO}" in', 1)[1] - section = main_case.split(" s3)", 1)[1].split("esac", 1)[0] - - assert section.index('INJECTED_AT="$(utc_now)"') < section.index( - "az role assignment delete" - ) - assert 'ROLE_DELETED_AT="$(utc_now)"' in section - - def test_lab_state_runs_bound_to_the_resolved_configuration(): """`lab_state.py` must never resolve the lab's identity itself. @@ -276,57 +218,6 @@ def test_lab_state_runs_bound_to_the_resolved_configuration(): assert 'AZURE_RESOURCE_GROUP="${RESOURCE_GROUP}"' in tool_helper -def test_run_scenario_checks_the_run_order_before_injecting_a_failure(): - """The gate is worthless after the fact: `require-run` has to run before - the first `az` call that breaks the workload.""" - script = RUN_SCENARIO_SH.read_text() - - assert script.index('lab_state require-run "${SCENARIO}"') < script.index( - "az containerapp update" - ) - assert script.index('lab_state require-run "${SCENARIO}"') < script.index( - "az role assignment delete" - ) - - -def test_run_scenario_records_recovery_only_after_health_and_alert_checks(): - script = RUN_SCENARIO_SH.read_text() - - assert "wait_for_app_ready" in script - assert "wait_for_alert_resolved" in script - assert script.index("wait_for_alert_resolved") < script.index( - 'lab_state mark-recovered' - ) - assert "lab_state mark-failed" in script - - -def test_run_scenario_default_alert_resolution_budget_covers_stateful_log_alerts(): - script = RUN_SCENARIO_SH.read_text() - - assert 'LAB_ALERT_RESOLVE_TIMEOUT_SECONDS:-1500' in script - assert 'local timeout_seconds="${2:-1500}"' in COMMON_SH.read_text() - - -def test_capture_scenario_records_the_terminal_state_from_the_timeline(): - """The capture status is derived from the normalized timeline, so a - missing thread/investigation/conclusion is recorded as itself and can - never be reported as a successful capture.""" - script = CAPTURE_SCENARIO_SH.read_text() - - assert 'lab_state record-capture "${SCENARIO}"' in script - assert '--timeline "${NORMALIZED_FILE}"' in script - assert 'lab_state evidence-dir "${SCENARIO}"' in script - - -def test_baseline_records_the_passing_baseline_stage(): - script = BASELINE_SH.read_text() - - assert "lab_state mark baseline_passed" in script - assert script.index("lab_state mark baseline_passed") > script.index( - "did not show both request types" - ) - - def test_lab_state_and_score_are_exercised_as_programs(): """`lab_state.py` and `score.py` decide whether a scenario may run and what the evidence is worth, so both are driven through their real API @@ -349,45 +240,68 @@ def test_activity_log_export_projects_only_incident_fields(): assert "claims:" not in script -def test_cleanup_delegates_to_external_cleanup_and_keeps_recovery_deletion_behind_a_flag(): - """`cleanup.sh` used to delete the whole resource group itself, which is - now `azd down`'s job. It stays as a compatibility wrapper: it names the - supported command, forwards to `cleanup-external.sh`, and only deletes a - resource group when an operator explicitly asks for the documented - recovery path. `test_lab_scripts.py` runs both paths as programs. - """ - script = CLEANUP_SH.read_text() - readme = (Path(__file__).parents[2] / "README.md").read_text() +def test_the_lab_keeps_only_the_shell_scripts_it_cannot_do_without(): + """The walkthrough is manual: the guides run `az` and the Python tools + directly, and no script runs a scenario on the operator's behalf. What + survives is the set nothing else can cover -- the four hooks + `azure.yaml` invokes, the environment `uv` needs, the values exported + once per shell, the shared library those read, and the evidence + collection that is eight queries in a row. - assert "azd down --purge" in script - assert "cleanup-external.sh" in script - assert LEGACY_RESOURCE_GROUP_FLAG in script - assert LEGACY_RESOURCE_GROUP_FLAG in readme, ( - "the legacy resource-group deletion must be documented for recovery" + Pinning the set keeps a new wrapper from arriving unnoticed: adding one + means changing this list, in review. + """ + scripts = sorted( + path.name for path in (Path(__file__).parents[1]).glob("*.sh") ) - # Nothing may delete a resource group before the legacy flag is parsed. - assert script.index(LEGACY_RESOURCE_GROUP_FLAG) < script.index("az group delete") + + assert scripts == [ + "azd-configure.sh", + "azd-deploy-app.sh", + "azd-postprovision-local.sh", + "cleanup-external.sh", + "common.sh", + "lab-env.sh", + "query-evidence.sh", + "setup-venv.sh", + ], scripts -def test_scenario_query_capture_cleanup_scripts_are_exercised_as_programs(): - """The four entry points are covered by execution tests, not by reading - their text: `test_lab_scripts.py` runs each one against fake - `az`/`azd`/`python` executables from a working directory outside the - lab, which is the only way to catch a caller that reassigns a name - `common.sh` already made readonly. +def test_nothing_points_an_operator_at_a_script_that_is_gone(): + """A refusal message or an instruction that names a deleted script is + worse than no message: it sends the operator to `No such file or + directory` at the moment they are already stuck. Scripts, tools and + documents are all checked, because all three talk to the operator. """ - lab_script_tests = (Path(__file__).parent / "test_lab_scripts.py").read_text() + lab_root = Path(__file__).parents[2] + scripts_dir = lab_root / "scripts" + sources = ( + list(scripts_dir.glob("*.sh")) + + list(scripts_dir.glob("*.py")) + + [lab_root / "README.md"] + + sorted((lab_root / "guides").glob("*.md")) + ) + + for path in sources: + # A URL is not a script reference: `https://astral.sh/uv/...` ends in + # a country-code domain that looks exactly like a shell script. + text = re.sub(r"https?://\S+", " ", path.read_text()) + for name in set(re.findall(r"\b([a-z0-9_-]+\.sh)\b", text)): + assert (scripts_dir / name).is_file(), (path.name, name) - for script_name in ( - "run-scenario.sh", - "query-evidence.sh", - "capture-scenario.sh", - "cleanup.sh", - ): - assert f'"{script_name}"' in lab_script_tests, ( - f"{script_name} has no execution test" - ) +def test_every_shell_script_is_exercised_by_a_test(): + """Reading a script's text cannot catch a caller that reassigns a name + `common.sh` already made readonly; only running it can.""" + tests_dir = Path(__file__).parent + corpus = "\n".join( + path.read_text() for path in tests_dir.glob("*.py") if path != Path(__file__) + ) + + for path in sorted((Path(__file__).parents[1]).glob("*.sh")): + assert f'"{path.name}"' in corpus or path.name in corpus, ( + f"{path.name} has no test that runs or sources it" + ) # --- The evidence directory is a name first, a directory second ------------- @@ -413,7 +327,7 @@ def run_in_throwaway_lab(tmp_path, command): def test_evidence_dir_path_names_a_directory_without_creating_it(tmp_path): - """`run-scenario.sh` needs the evidence path *before* it asks + """A scenario needs the evidence path *before* it asks `lab_state.py` to admit the run, because the path is what it registers. Creating the directory at that point left an empty `sN-/` behind whenever the run was then refused -- litter an operator has to @@ -437,7 +351,7 @@ def test_evidence_dir_path_names_a_directory_without_creating_it(tmp_path): def test_create_evidence_dir_still_creates_what_it_names(tmp_path): - """`baseline.sh` writes into the directory immediately, so the eager + """Some callers write into the directory immediately, so the eager helper must keep working -- the split adds a step, it does not move the responsibility.""" result, evidence_root = run_in_throwaway_lab( diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_doctor.py b/monitor/sre-agent-event-lab/scripts/tests/test_doctor.py deleted file mode 100644 index 2da687b..0000000 --- a/monitor/sre-agent-event-lab/scripts/tests/test_doctor.py +++ /dev/null @@ -1,556 +0,0 @@ -"""Behavioural tests for `doctor.sh`. - -Every test drives `doctor.sh` as a real program against a fake `az`/`azd`/ -`curl` on PATH (see `doctor_harness.py`), not by grepping its text: the -output-contract tests below prove PASS/FAIL/MANUAL rows are produced from -actual (faked) CLI responses, that a single unhealthy signal both flips the -exit code and is distinguishable from a portal-only MANUAL check, and that -the script never queries Azure once its own safety gate (commands, login, -azd configuration, subscription equality, resource-group tags) has failed. -""" -import json - -import pytest - -from doctor_harness import ( - AGENT_PRINCIPAL_ID, - AGENT_UAMI_PRINCIPAL_ID, - RESOURCE_GROUP, - SUBSCRIPTION_SCOPE, - FakeAz, - az_calls_for, - azd_calls_for, - lab_dir_for, - run_doctor, -) - - -MANUAL_CHECKS = ( - "Repository connection", - "Knowledge source", - "Incident platform", - "Response plan", -) - - -def _rows_of(result): - return dict(line.split("\t", 2)[0:2] for line in result.stdout.splitlines() if "\t" in line) - - -def _detail_of(result, check_name): - row = next( - line for line in result.stdout.splitlines() if line.startswith(f"{check_name}\t") - ) - return row.split("\t", 2)[2] - - -def _analytics_queries(fake_az): - return [ - line for line in az_calls_for(fake_az).splitlines() if "monitor log-analytics" in line - ] - - -@pytest.fixture -def fake_az(tmp_path): - return FakeAz(workdir=tmp_path) - - -def test_doctor_reports_manual_for_unverifiable_portal_settings(fake_az): - result = run_doctor(fake_az, sre_agent_resource_id="/subscriptions/sub/...") - - assert "Repository connection\tMANUAL" in result.stdout - assert "Response plan\tMANUAL" in result.stdout - - -def test_doctor_fails_when_workload_is_unhealthy(fake_az): - fake_az.container_app_health = "Unhealthy" - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Container App health\tFAIL" in result.stdout - - -def test_doctor_passes_fully_healthy_environment(fake_az): - """Every gating and diagnostic check reaches PASS; only the four - portal-only settings are MANUAL, and the overall exit code is 0.""" - result = run_doctor(fake_az, sre_agent_resource_id="/subscriptions/sub/.../sreAgents/a") - - assert result.returncode == 0, result.stdout + result.stderr - rows = _rows_of(result) - assert rows["Required commands"] == "PASS" - assert rows["Python environment"] == "PASS" - assert rows["Log Analytics CLI extension"] == "PASS" - assert rows["Azure CLI login"] == "PASS" - assert rows["azd authentication"] == "PASS" - assert rows["azd configuration"] == "PASS" - assert rows["Subscription match"] == "PASS" - assert rows["Resource group tags"] == "PASS" - assert rows["Container App health"] == "PASS" - assert rows["Health endpoint"] == "PASS" - assert rows["Application Insights telemetry"] == "PASS" - assert rows["Alert rules enabled"] == "PASS" - assert rows["SRE Agent resource"] == "PASS" - assert rows["Reader role assignment"] == "PASS" - for manual_check in MANUAL_CHECKS: - assert rows[manual_check] == "MANUAL" - - -def test_doctor_omits_sre_agent_resource_row_when_not_configured(fake_az): - """The check is only meaningful -- and only printed -- once an operator - has recorded SRE_AGENT_RESOURCE_ID; otherwise nothing has been created - yet, and there is nothing to verify.""" - result = run_doctor(fake_az) - - assert result.returncode == 0, result.stdout + result.stderr - assert "SRE Agent resource" not in result.stdout - - -def test_doctor_fails_when_sre_agent_resource_is_missing(fake_az): - fake_az.sre_agent_resource_exists = False - - result = run_doctor(fake_az, sre_agent_resource_id="/subscriptions/sub/.../sreAgents/missing") - - assert result.returncode == 1 - assert "SRE Agent resource\tFAIL" in result.stdout - - -def test_doctor_fails_when_healthz_does_not_return_200(fake_az): - fake_az.healthz_status = 503 - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Health endpoint\tFAIL" in result.stdout - assert "503" in result.stdout - - -def test_doctor_explains_the_pre_deploy_placeholder_state_instead_of_misclassifying_it(fake_az): - """`azd provision` alone leaves the public placeholder image running - (port 80, no `/healthz`) until `azd deploy` builds and switches to the - lab image (see infra/main.bicep, scripts/azd-deploy-app.sh). Without - SRE_CONTAINER_IMAGE recorded, a failing `/healthz` is that documented, - expected intermediate state -- not a broken deployment -- so doctor - must name the exact remedy (`azd deploy --no-prompt`) instead of the - generic "investigate with curl -v" wording that implies something is - actually broken.""" - fake_az.healthz_status = 404 - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Health endpoint\tFAIL" in result.stdout - detail = _detail_of(result, "Health endpoint") - assert "azd deploy --no-prompt" in detail - assert "curl -v" not in detail - - -def test_doctor_reports_a_genuine_health_failure_once_the_lab_image_is_deployed(fake_az): - """Once SRE_CONTAINER_IMAGE is recorded, the deploy phase already - replaced the placeholder with the lab image and its /healthz probes; - a failure at that point is a real regression and must not be softened - into the pre-deploy placeholder message.""" - fake_az.healthz_status = 503 - - result = run_doctor( - fake_az, - sre_container_image="acrsrelabtest.azurecr.io/sre-event-lab:abc123", - ) - - assert result.returncode == 1 - assert "Health endpoint\tFAIL" in result.stdout - detail = _detail_of(result, "Health endpoint") - assert "azd deploy --no-prompt" not in detail - assert "curl -v" in detail - - -def test_doctor_fails_when_venv_is_missing(fake_az): - """Finding #1: doctor must report on the venv `setup-venv.sh` (run from - `postprovision`) is responsible for creating, with a remedy pointing at - that exact script -- not just a generic "python3 missing" message, - since `python3` itself is still on PATH.""" - fake_az.venv_present = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Python environment\tFAIL" in result.stdout - assert "setup-venv.sh" in _detail_of(result, "Python environment") - - -def test_doctor_fails_when_pillow_is_not_importable_from_the_venv(fake_az): - """A venv that exists but never finished installing (or was created by - something other than `setup-venv.sh`) must fail this check too, not - just an absent venv directory.""" - fake_az.pillow_importable = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Python environment\tFAIL" in result.stdout - assert "setup-venv.sh" in _detail_of(result, "Python environment") - - -def test_doctor_passes_venv_check_independently_of_azure_reachability(fake_az): - """The venv/Pillow readiness check is a local precondition, not an - Azure fact: it must still report accurately (and PASS when the venv is - fine) even when every Azure-dependent check is blocked.""" - fake_az.logged_in = False - - result = run_doctor(fake_az) - - assert "Python environment\tPASS" in result.stdout - assert "Azure CLI login\tFAIL" in result.stdout - - -def test_doctor_fails_when_app_insights_has_no_recent_requests(fake_az): - fake_az.app_insights_has_recent_requests = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Application Insights telemetry\tFAIL" in result.stdout - - -def test_doctor_fails_when_an_alert_rule_is_disabled(fake_az): - fake_az.alert_rules_enabled["alert-sre-lab-s2-latency"] = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - row = next(line for line in result.stdout.splitlines() if line.startswith("Alert rules enabled\t")) - assert "FAIL" in row - assert "alert-sre-lab-s2-latency" in row - - -def test_doctor_fails_when_an_alert_rule_is_missing(fake_az): - fake_az.alert_rules_present["alert-sre-lab-s3-storage-rbac"] = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - row = next(line for line in result.stdout.splitlines() if line.startswith("Alert rules enabled\t")) - assert "FAIL" in row - assert "alert-sre-lab-s3-storage-rbac" in row - - -def test_doctor_fails_when_agent_setup_evidence_is_missing(fake_az): - """Unlike the repository/knowledge/response-plan settings, Reader is a - real RBAC role assignment a stable API can check -- so a missing - prerequisite is FAIL, never a shrug-and-guess MANUAL.""" - fake_az.agent_setup_present = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Reader role assignment\tFAIL" in result.stdout - assert "agent-setup.json" in result.stdout - assert "guides/01-agent-setup.md" in result.stdout - assert "lab.sh acknowledge agent-setup" in result.stdout - - -def test_doctor_fails_when_subscription_assignment_ids_are_empty(fake_az): - fake_az.agent_setup_body = json.dumps( - { - "agent_endpoint": "https://sre-agent.example.com/api/incidents", - "monitoring_contributor_assignment_id": "", - "agent_principal_id": AGENT_PRINCIPAL_ID, - "uami_monitoring_contributor_assignment_id": ( - f"{SUBSCRIPTION_SCOPE}/providers/Microsoft.Authorization/" - "roleAssignments/principal-two" - ), - "agent_user_assigned_principal_id": AGENT_UAMI_PRINCIPAL_ID, - } - ) - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Monitoring Contributor cleanup evidence\tFAIL" in result.stdout - assert "monitoring_contributor_assignment_id" in result.stdout - assert "az role assignment list" in result.stdout - - -def test_doctor_accepts_case_insensitive_arm_assignment_ids(fake_az): - assignment_prefix = ( - f"{SUBSCRIPTION_SCOPE}/providers/microsoft.authorization/roleassignments/" - ) - fake_az.agent_setup_body = json.dumps( - { - "agent_endpoint": "https://sre-agent.example.com/api/incidents", - "monitoring_contributor_assignment_id": assignment_prefix + "principal-one", - "agent_principal_id": AGENT_PRINCIPAL_ID, - "uami_monitoring_contributor_assignment_id": assignment_prefix + "principal-two", - "agent_user_assigned_principal_id": AGENT_UAMI_PRINCIPAL_ID, - } - ) - - result = run_doctor(fake_az) - - assert result.returncode == 0, result.stdout + result.stderr - assert "Monitoring Contributor cleanup evidence\tPASS" in result.stdout - - -def test_doctor_fails_when_one_recorded_identity_lacks_reader(fake_az): - fake_az.reader_role_assigned[AGENT_UAMI_PRINCIPAL_ID] = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - row = next( - line for line in result.stdout.splitlines() if line.startswith("Reader role assignment\t") - ) - assert "FAIL" in row - assert AGENT_UAMI_PRINCIPAL_ID in row - - -def test_doctor_never_calls_azure_once_subscription_mismatches(fake_az): - """Fail closed: once the pinned-subscription check fails, no further - Azure resource lookups may happen, even to produce diagnostics.""" - result = run_doctor(fake_az, azure_subscription_id="99999999-9999-9999-9999-999999999999") - - assert result.returncode == 1 - assert "Subscription match\tFAIL" in result.stdout - assert "Blocked: resolve the failing check above first." in result.stdout - assert "containerapp revision" not in az_calls_for(fake_az) - assert "role assignment" not in az_calls_for(fake_az) - - -def test_doctor_never_calls_azure_when_azd_configuration_is_missing(fake_az): - """No azd value and no explicit environment: doctor must stop with the - actionable `azd env set` message before touching Azure, exactly like - every other entry point.""" - fake_az.azd_values = {} - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "azd configuration\tFAIL" in result.stdout - assert "azd env set AZURE_SUBSCRIPTION_ID" in result.stdout - assert "containerapp revision" not in az_calls_for(fake_az) - - -def test_doctor_pins_azd_lookups_to_the_lab_project_root(fake_az): - """`azd env get-value` must be pinned with `--cwd` to this lab's own - project root, exactly as `common.sh`'s other callers already are.""" - result = run_doctor(fake_az) - - assert result.returncode == 0, result.stdout + result.stderr - assert f"cwd={lab_dir_for(fake_az)}" in azd_calls_for(fake_az) - - - -# --- Application Insights telemetry: the real `az monitor log-analytics -# query` output contract ------------------------------------------------ -# -# The extension flattens the REST envelope into a JSON array of row objects -# (`{"tables": [...]}` is never printed), and KQL's `count` operator always -# returns exactly one row -- so "the result had rows" only distinguishes -# data from no data when the query itself is not a `| count`. - - -def test_doctor_reports_telemetry_pass_from_flat_query_output(fake_az): - """A workspace that has data answers with a non-empty flat array.""" - fake_az.app_insights_has_recent_requests = True - - result = run_doctor(fake_az) - - assert result.returncode == 0, result.stdout + result.stderr - assert "Application Insights telemetry\tPASS" in result.stdout - assert _analytics_queries(fake_az), "doctor never queried the workspace" - - -def test_doctor_reports_telemetry_fail_from_empty_flat_query_output(fake_az): - """A workspace with no matching rows answers with exactly `[]`.""" - fake_az.app_insights_has_recent_requests = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Application Insights telemetry\tFAIL" in result.stdout - assert "lab.sh baseline" in _detail_of(result, "Application Insights telemetry") - - -def test_doctor_telemetry_query_does_not_count_rows_of_a_count(fake_az): - """`| count` returns one row even for an empty table, so counting its - rows can never distinguish data from no data.""" - run_doctor(fake_az) - - queries = _analytics_queries(fake_az) - assert queries - for query in queries: - assert "| count" not in query, f"telemetry query relies on `| count`: {query}" - - -def test_doctor_telemetry_does_not_parse_the_rest_tables_envelope(fake_az): - """Regression guard for the shape defect itself: with the real flat - output faked, a `.tables[0].rows` parse always yields zero rows and can - only ever report FAIL, so a healthy workspace must still PASS.""" - fake_az.app_insights_has_recent_requests = True - fake_az.app_insights_orders_seen = True - - result = run_doctor(fake_az) - - assert "Application Insights telemetry\tPASS" in result.stdout - - -# --- Prerequisite: the `log-analytics` CLI extension --------------------- - - -def test_doctor_fails_when_log_analytics_extension_is_missing(fake_az): - """`az monitor log-analytics query` lives in an extension that is not - installed by default; without it every telemetry check is meaningless.""" - fake_az.log_analytics_extension_installed = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Log Analytics CLI extension\tFAIL" in result.stdout - assert "az extension add --name log-analytics" in _detail_of( - result, "Log Analytics CLI extension" - ) - - -def test_doctor_does_not_query_the_workspace_without_the_extension(fake_az): - """Fail closed: no point issuing a query the CLI cannot run.""" - fake_az.log_analytics_extension_installed = False - - run_doctor(fake_az) - - assert "monitor log-analytics" not in az_calls_for(fake_az) - - -def test_doctor_checks_the_extension_with_a_stable_command(fake_az): - run_doctor(fake_az) - - assert "extension show --name log-analytics" in az_calls_for(fake_az) - - -# --- Prerequisite: azd authentication ------------------------------------ - - -def test_doctor_reports_azd_authentication_from_check_status(fake_az): - """`azd auth login --check-status` is azd's only non-interactive login - read, and it is queried with `--output json` so the machine-readable - status -- not a human sentence -- decides the row.""" - run_doctor(fake_az) - - azd_calls = azd_calls_for(fake_az) - assert "auth login --check-status" in azd_calls - assert "--output json" in azd_calls - - -def test_doctor_fails_when_azd_is_not_authenticated(fake_az): - fake_az.azd_logged_in = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "azd authentication\tFAIL" in result.stdout - assert "azd auth login" in _detail_of(result, "azd authentication") - - -def test_doctor_does_not_trust_azd_check_status_exit_code(fake_az): - """`azd auth login --check-status` always exits 0. A doctor that reads - the exit status would report a signed-out operator as authenticated and - then charge on into Azure calls.""" - fake_az.azd_logged_in = False - - result = run_doctor(fake_az) - - assert "azd authentication\tPASS" not in result.stdout - assert "containerapp revision" not in az_calls_for(fake_az) - - -# --- Reader role assignment: direct vs inherited ------------------------- - - -def test_doctor_reports_reader_assigned_directly_on_the_resource_group(fake_az): - result = run_doctor(fake_az) - - assert result.returncode == 0, result.stdout + result.stderr - detail = _detail_of(result, "Reader role assignment") - assert "directly" in detail - assert RESOURCE_GROUP in detail - - -def test_doctor_accepts_reader_inherited_from_the_subscription(fake_az): - """Reader granted on the subscription gives the Agent the same effective - read access on the lab resource group, so it must not be a false FAIL -- - but the detail has to say it is inherited, not a direct assignment.""" - fake_az.reader_role_assigned[AGENT_UAMI_PRINCIPAL_ID] = False - fake_az.reader_role_inherited[AGENT_UAMI_PRINCIPAL_ID] = True - - result = run_doctor(fake_az) - - assert result.returncode == 0, result.stdout + result.stderr - detail = _detail_of(result, "Reader role assignment") - assert "inherited" in detail - assert AGENT_UAMI_PRINCIPAL_ID in detail - assert SUBSCRIPTION_SCOPE in detail - - -def test_doctor_asks_azure_for_inherited_role_assignments(fake_az): - """Without `--include-inherited`, `az role assignment list` hides - parent-scope grants entirely.""" - run_doctor(fake_az) - - role_calls = [ - line for line in az_calls_for(fake_az).splitlines() if line.startswith("role assignment") - ] - assert role_calls - for call in role_calls: - assert "--include-inherited" in call - - -def test_doctor_still_fails_when_no_reader_exists_at_any_scope(fake_az): - fake_az.reader_role_assigned[AGENT_PRINCIPAL_ID] = False - fake_az.reader_role_inherited[AGENT_PRINCIPAL_ID] = False - - result = run_doctor(fake_az) - - assert result.returncode == 1 - detail = _detail_of(result, "Reader role assignment") - assert AGENT_PRINCIPAL_ID in detail - assert "az role assignment create" in detail - - -# --- Malformed agent-setup.json ------------------------------------------ - - -def test_doctor_fails_gracefully_on_malformed_agent_setup_evidence(fake_az): - """A truncated/hand-edited evidence file must be one FAIL row, not a raw - `jq` abort that kills the run mid-report.""" - fake_az.agent_setup_body = '{"agent_principal_id": "8c8a4f0e"' - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Reader role assignment\tFAIL" in result.stdout - detail = _detail_of(result, "Reader role assignment") - assert "valid JSON" in detail - assert "agent-setup.json" in detail - assert "guides/01-agent-setup.md" in detail - assert "lab.sh acknowledge agent-setup" in detail - - -def test_doctor_finishes_the_report_after_malformed_agent_setup_evidence(fake_az): - """The rows after the failing check still have to be printed, and the - raw parser error must not leak to stderr.""" - fake_az.agent_setup_body = "not json at all" - - result = run_doctor(fake_az) - - assert "Response plan\tMANUAL" in result.stdout - assert "parse error" not in result.stderr - assert "jq:" not in result.stderr - - -def test_doctor_fails_when_agent_setup_evidence_is_valid_json_but_empty(fake_az): - fake_az.agent_setup_body = "{}" - - result = run_doctor(fake_az) - - assert result.returncode == 1 - assert "Reader role assignment\tFAIL" in result.stdout - assert "agent_principal_id" in _detail_of(result, "Reader role assignment") diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_lab_cli.py b/monitor/sre-agent-event-lab/scripts/tests/test_lab_cli.py deleted file mode 100644 index b3b7532..0000000 --- a/monitor/sre-agent-event-lab/scripts/tests/test_lab_cli.py +++ /dev/null @@ -1,258 +0,0 @@ -"""Behavioural tests for `lab.sh`, the guided single-command entry point. - -`lab.sh` itself contains no Azure logic -- it only dispatches to -`doctor.sh`, `baseline.sh`, `run-scenario.sh`, `capture-scenario.sh`, and -(once a later task adds them) `lab_state.py`/`score.py`. These tests drive -it as a real program: dispatch to the scripts that already exist is proven -by observing their actual (faked) side effects, not by grepping `lab.sh`'s -source for a case label. The one text-based exception is -`test_lab_cli_dispatches_known_commands`, which is a brief-mandated -regression guard for the dispatcher's own case statement. -""" -import json -from pathlib import Path - -import pytest - -from doctor_harness import FakeAz, lab_dir_for, run_lab_cli, state_path_for -from lab_script_harness import make_lab - - -COMMANDS = ("doctor", "baseline", "acknowledge", "run", "capture", "score") - -# Bounded recovery waits: `run-scenario.sh` polls the workload health and the -# fired alert's condition before it records a recovery. -BOUNDED_WAITS = { - "LAB_ALERT_RESOLVE_TIMEOUT_SECONDS": "5", - "LAB_ALERT_RESOLVE_POLL_INTERVAL_SECONDS": "1", - "LAB_RECOVERY_HEALTH_TIMEOUT_SECONDS": "5", -} - -CONCLUSION_TIMELINE = [ - {"state": "alert-fired"}, - {"state": "thread-created"}, - {"state": "investigating"}, - {"state": "conclusion"}, -] -FULL_REVIEW = { - "impact_scope": {"met": True, "detail": "Named both routes."}, - "direct_cause": {"met": True, "detail": "Named the injected failure mode."}, - "actual_evidence": {"met": True, "detail": "Quoted AppRequests rows."}, - "safe_minimum_mitigation": {"met": True, "detail": "Proposed the revert."}, - "uncertainty": {"met": True, "detail": "Flagged what it could not verify."}, -} - - -@pytest.fixture -def fake_az(tmp_path): - return FakeAz(workdir=tmp_path) - - -def test_lab_cli_dispatches_known_commands(): - lab_cli = Path(__file__).parents[1].joinpath("lab.sh").read_text() - for command in COMMANDS: - assert f"{command})" in lab_cli, f"lab.sh has no dispatch case for {command}" - - -def test_lab_cli_doctor_dispatches_to_doctor_sh(fake_az): - result = run_lab_cli(fake_az, ["doctor"]) - - assert result.returncode == 0, result.stdout + result.stderr - assert "Required commands\tPASS" in result.stdout - assert "Repository connection\tMANUAL" in result.stdout - - -def test_lab_cli_doctor_surfaces_doctor_sh_failure(fake_az): - fake_az.container_app_health = "Unhealthy" - - result = run_lab_cli(fake_az, ["doctor"]) - - assert result.returncode == 1 - assert "Container App health\tFAIL" in result.stdout - - -def test_lab_cli_baseline_dispatches_to_baseline_sh(fake_az): - result = run_lab_cli( - fake_az, - ["baseline"], - lab_baseline_telemetry_timeout_seconds="5", - lab_baseline_telemetry_poll_interval_seconds="1", - ) - - assert result.returncode == 0, result.stdout + result.stderr - evidence_dirs = sorted((lab_dir_for(fake_az) / "evidence").glob("baseline-*")) - assert evidence_dirs, "lab.sh baseline never invoked baseline.sh" - assert (evidence_dirs[-1] / "telemetry-check.json").is_file() - assert (evidence_dirs[-1] / "orders.json").is_file() - assert (evidence_dirs[-1] / "documents.json").is_file() - - -def test_lab_cli_baseline_surfaces_baseline_sh_failure(fake_az): - fake_az.baseline_orders_succeed = False - - result = run_lab_cli( - fake_az, - ["baseline"], - lab_baseline_telemetry_timeout_seconds="5", - lab_baseline_telemetry_poll_interval_seconds="1", - ) - - assert result.returncode != 0 - - -def test_lab_cli_run_dispatches_to_run_scenario_sh(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.seed_state() - - result = lab_run.run("lab.sh", ["run", "s1"], env=BOUNDED_WAITS) - - assert result.returncode == 0, result.stderr - evidence_dirs = sorted((lab_run.lab / "evidence").glob("s1-*")) - assert evidence_dirs, "lab.sh run never invoked run-scenario.sh" - timeline = json.loads((evidence_dirs[-1] / "timeline.json").read_text()) - assert timeline["scenario"] == "s1" - - -def test_lab_cli_run_rejects_an_unknown_scenario(tmp_path): - lab_run = make_lab(tmp_path) - - result = lab_run.run("lab.sh", ["run", "s9"]) - - assert result.returncode == 2 - assert "Usage" in result.stderr - - -def test_lab_cli_capture_resolves_the_evidence_directory_from_the_state(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - run_result = lab_run.run("lab.sh", ["run", "s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - evidence_dir = sorted((lab_run.lab / "evidence").glob("s1-*"))[-1] - - result = lab_run.run("lab.sh", ["capture", "s1"]) - - assert result.returncode == 0, result.stdout + result.stderr - assert (evidence_dir / "normalized-timeline.json").is_file() - assert (lab_run.lab / "assets" / "captures" / "s1" / "investigation.gif").is_file() - assert lab_run.scenario_state("s1")["capture_status"] == "conclusion" - - -def test_lab_cli_capture_fails_clearly_when_no_evidence_exists(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - - result = lab_run.run("lab.sh", ["capture", "s1"]) - - assert result.returncode != 0 - assert "No such file or directory" not in result.stderr - assert "s1" in result.stderr - assert "lab.sh run s1" in result.stderr - - -def test_lab_cli_capture_works_even_when_capture_scenario_is_not_executable(tmp_path): - """Dispatch must not depend on a sub-script's executable bit -- `lab.sh` - always invokes it through `bash`, never a bare `exec path`.""" - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - run_result = lab_run.run("lab.sh", ["run", "s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - (lab_run.lab / "scripts" / "capture-scenario.sh").chmod(0o644) - - result = lab_run.run("lab.sh", ["capture", "s1"]) - - assert result.returncode == 0, result.stdout + result.stderr - - -def test_lab_cli_acknowledge_prints_the_settings_and_records_the_answer(fake_az): - result = run_lab_cli( - fake_az, - ["acknowledge", "agent-setup"], - stdin="acknowledge\n", - sre_agent_name="sre-agent-lab", - sre_repository_url="https://github.com/example/devguidesample", - sre_repository_branch="feature/sre-agent-azd-lab", - sre_knowledge_path="runbooks/incident-response.md", - ) - - assert result.returncode == 0, result.stdout + result.stderr - assert "https://github.com/example/devguidesample" in result.stdout - assert "feature/sre-agent-azd-lab" in result.stdout - assert "runbooks/incident-response.md" in result.stdout - assert "Review" in result.stdout - assert "alert-sre-lab-s1-http500" in result.stdout - state = json.loads(state_path_for(fake_az).read_text()) - assert "agent_setup_acknowledged" in state["stages"] - assert state["environment"] == "sre-lab-exec" - - -def test_lab_cli_acknowledge_records_nothing_without_the_exact_word(fake_az): - result = run_lab_cli(fake_az, ["acknowledge", "agent-setup"], stdin="yes\n") - - assert result.returncode != 0 - assert not state_path_for(fake_az).exists() - - -def test_lab_cli_score_without_evidence_explains_what_to_run(fake_az): - result = run_lab_cli(fake_az, ["score"]) - - assert result.returncode == 1 - assert "lab.sh run" in result.stderr - assert "No such file or directory" not in result.stderr - - -def test_lab_cli_score_scores_the_collected_evidence(fake_az): - run_lab_cli(fake_az, ["score"]) # materializes the lab - evidence_root = lab_dir_for(fake_az) / "evidence" - scenarios = {} - for scenario in ("s1", "s2", "s3"): - evidence_dir = evidence_root / f"{scenario}-20260814T000000Z" - evidence_dir.mkdir(parents=True) - (evidence_dir / "normalized-timeline.json").write_text(json.dumps(CONCLUSION_TIMELINE)) - (evidence_dir / "conclusion-review.json").write_text(json.dumps(FULL_REVIEW)) - scenarios[scenario] = { - "run_status": "recovered", - "capture_status": "conclusion", - "evidence_dir": str(evidence_dir), - } - state_path_for(fake_az).write_text( - json.dumps( - { - "environment": "sre-lab-exec", - "subscription_id": "11111111-2222-3333-4444-555555555555", - "resource_group": "rg-sre-lab-exec", - "stages": {"baseline_passed": {"at": "2026-08-14T00:00:00Z"}}, - "scenarios": scenarios, - } - ) - ) - - result = run_lab_cli(fake_az, ["score"]) - - assert result.returncode == 0, result.stdout + result.stderr - assert "OVERALL\tTOTAL\tPASS\t30/30" in result.stdout - scorecard = json.loads((evidence_root / "scorecard.json").read_text()) - assert scorecard["overall"]["verdict"] == "PASS" - - -def test_lab_cli_acknowledge_rejects_an_unknown_subcommand(fake_az): - result = run_lab_cli(fake_az, ["acknowledge", "not-a-setup"]) - - assert result.returncode == 2 - assert "Usage" in result.stderr - - -def test_lab_cli_rejects_an_unknown_command(fake_az): - result = run_lab_cli(fake_az, ["bogus"]) - - assert result.returncode == 2 - assert "Usage" in result.stderr - - -def test_lab_cli_with_no_arguments_prints_usage(fake_az): - result = run_lab_cli(fake_az, []) - - assert result.returncode == 2 - assert "Usage" in result.stderr diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_lab_env.py b/monitor/sre-agent-event-lab/scripts/tests/test_lab_env.py index 216ed21..c73297f 100644 --- a/monitor/sre-agent-event-lab/scripts/tests/test_lab_env.py +++ b/monitor/sre-agent-event-lab/scripts/tests/test_lab_env.py @@ -1,12 +1,12 @@ -"""Contract tests for the one-shot lab environment and its Codespaces entry. - -An operator who opens the lab in Codespaces runs `az login` once and then -one `source`; every later command reads exported values instead of binding -each one by hand. These tests pin the pieces that promise makes: the -devcontainer Codespaces actually offers, the sourced script that resolves -and exports the values, and the guides that stopped re-binding them. +"""Contract tests for the one-shot lab environment. + +An operator runs `az login` once and then one `source`; every later command +reads exported values instead of binding each one by hand. These tests pin +the pieces that promise makes: the sourced script that resolves and exports +the values, and the guides that stopped re-binding them. The container that +supplies the tools is the whole repository's, and is pinned separately in +`test_repo_devcontainer.py`. """ -import json import os import re import subprocess @@ -20,11 +20,6 @@ GUIDES = LAB_ROOT / "guides" README = LAB_ROOT / "README.md" -# GitHub only offers configurations stored at the repository root, either as -# `.devcontainer/devcontainer.json` or one level deep under `.devcontainer/`. -# A file anywhere else is invisible in the Codespaces creation UI. -DEVCONTAINER = REPO_ROOT / ".devcontainer" / "sre-agent-event-lab" / "devcontainer.json" - SCENARIO_GUIDES = { "02-scenario-s1.md": "s1", "03-scenario-s2.md": "s2", @@ -41,49 +36,14 @@ "WORKLOAD_PRINCIPAL_ID": "AZURE_CONTAINER_APP_PRINCIPAL_ID", "STORAGE_CONTAINER_SCOPE": "AZURE_STORAGE_CONTAINER_SCOPE", "BLOB_ROLE_ASSIGNMENT_NAME": "AZURE_BLOB_ROLE_ASSIGNMENT_NAME", + "WORKSPACE_CUSTOMER_ID": "AZURE_WORKSPACE_CUSTOMER_ID", + "TELEMETRY_SERVICE_NAME": "AZURE_TELEMETRY_SERVICE_NAME", } def manual_section(name): text = (GUIDES / name).read_text() - return text[text.index("## 수동 실행"):text.index("## 지름길")] - - -# --- Codespaces entry point ------------------------------------------------ - - -def test_codespaces_offers_a_configuration_for_this_lab(): - assert DEVCONTAINER.is_file(), ( - "Codespaces only lists configurations under the repository's own " - f".devcontainer directory; expected {DEVCONTAINER}" - ) - config = json.loads(DEVCONTAINER.read_text()) - assert config["name"] - features = " ".join(config.get("features", {})) - for required in ("azure-cli", "azure-dev/azd", "github-cli", "python"): - assert required in features, required - - -def test_devcontainer_records_no_secret_and_no_subscription(): - raw = DEVCONTAINER.read_text() - assert not re.search( - r"\b[0-9a-fA-F]{8}-(?:[0-9a-fA-F]{4}-){3}[0-9a-fA-F]{12}\b", raw - ), "a GUID in the devcontainer would pin every user to one subscription" - for forbidden in ("PASSWORD", "SECRET", "TOKEN", "_KEY", "CONNECTION_STRING"): - assert forbidden not in raw.upper(), forbidden - - -def test_devcontainer_prepares_the_lab_without_logging_in_for_the_user(): - """`az login` is interactive and account-specific: the container may - prepare the workspace, but it must never attempt a login or bake in - credentials of whoever authored the file.""" - config = json.loads(DEVCONTAINER.read_text()) - lifecycle = " ".join( - str(config.get(hook, "")) - for hook in ("onCreateCommand", "postCreateCommand", "postStartCommand") - ) - assert "setup-venv.sh" in lifecycle - assert "az login" not in lifecycle + return text[text.index("## 수동 실행"):].split("\n## ", 1)[0] # --- One-shot environment -------------------------------------------------- @@ -128,6 +88,8 @@ def run_sourced(tmp_path, azd_mode="ok", az_mode="ok", extra=""): AZURE_CONTAINER_APP_PRINCIPAL_ID) echo 8c8a4f0e-0000-4000-8000-2b1f9a0c1234 ;; AZURE_STORAGE_CONTAINER_SCOPE) echo /subscriptions/s/rg/containers/documents ;; AZURE_BLOB_ROLE_ASSIGNMENT_NAME) echo 3f2504e0-4f89-11d3-9a0c-0305e82c3301 ;; + AZURE_WORKSPACE_CUSTOMER_ID) echo 9d1a0b2c-3d4e-5f60-7182-93a4b5c6d7e8 ;; + AZURE_TELEMETRY_SERVICE_NAME) echo sre-lab-app ;; *) echo "" ;; esac """ @@ -186,6 +148,7 @@ def test_sourcing_exports_the_resolved_values(tmp_path): assert values["RESOURCE_GROUP"] == "rg-sre-lab" assert values["APP_FQDN"] == "ca-sre-lab.example.io" assert values["BLOB_ROLE_ASSIGNMENT_NAME"] == "3f2504e0-4f89-11d3-9a0c-0305e82c3301" + assert values["WORKSPACE_CUSTOMER_ID"] == "9d1a0b2c-3d4e-5f60-7182-93a4b5c6d7e8" def test_a_failed_lookup_reports_zero_and_exports_nothing(tmp_path): @@ -243,26 +206,6 @@ def test_scenario_guides_still_refuse_to_run_when_the_environment_is_missing(): assert "begin-run {0}".format(scenario) in section, name -# --- The container can actually run what it promises ------------------------ - - -def test_devcontainer_provides_uv_because_setup_venv_requires_it(): - """`setup-venv.sh` is uv-only by design and exits 1 without it, so a - container that runs it in postCreate must install uv or every Codespace - starts with a failed lifecycle hook and no `app/.venv`.""" - config = json.loads(DEVCONTAINER.read_text()) - lifecycle = " ".join( - str(config.get(hook, "")) - for hook in ("onCreateCommand", "postCreateCommand", "postStartCommand") - ) - if "setup-venv.sh" not in lifecycle: - return - raw = DEVCONTAINER.read_text() - assert "uv" in raw, ( - "setup-venv.sh refuses to run without uv; the container must supply it" - ) - - # --- The repository URL never carries a credential -------------------------- diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_lab_guides.py b/monitor/sre-agent-event-lab/scripts/tests/test_lab_guides.py index 4af44f6..b4a2421 100644 --- a/monitor/sre-agent-event-lab/scripts/tests/test_lab_guides.py +++ b/monitor/sre-agent-event-lab/scripts/tests/test_lab_guides.py @@ -2,7 +2,7 @@ The README is the quickstart an operator reads first, and `guides/` holds the step-by-step documents it hands off to. These tests check behaviour a -reader depends on -- that the commands are the ones `lab.sh` really +reader depends on -- that the commands are the ones the lab really accepts, in the order `lab_state.py` really enforces; that every path, link and screenshot resolves; and that nothing here asks anyone to paste a credential into a file or an environment variable. @@ -17,7 +17,6 @@ GUIDES = LAB_ROOT / "guides" OFFICIAL_ASSETS = LAB_ROOT / "assets" / "official" RUNBOOK = LAB_ROOT / "runbooks" / "incident-response.md" -LAB_SH = LAB_ROOT / "scripts" / "lab.sh" VALIDATION_RESULTS = LAB_ROOT / "validation-results.md" DYNAMIC_THRESHOLDS = LAB_ROOT / "dynamic-thresholds.md" RESULTS_GUIDE = LAB_ROOT / "guides" / "05-results.md" @@ -179,24 +178,33 @@ def blocks(text: str): def test_readme_is_azd_first_and_ordered(): text = README.read_text() - commands = [ - "azd env new", - "azd up", - "./scripts/lab.sh doctor", - "./scripts/lab.sh baseline", - "./scripts/lab.sh acknowledge agent-setup", - "./scripts/lab.sh run s1", - "./scripts/lab.sh capture s1", - "./scripts/lab.sh run s2", - "./scripts/lab.sh capture s2", - "./scripts/lab.sh run s3", - "./scripts/lab.sh capture s3", - "./scripts/lab.sh score", - "azd down --purge", + headings = [ + "## 시작하기: fork와 Codespaces", + "## 사전 조건", + "## azd 환경 만들기", + "## 배포", + "## Azure SRE Agent 설정", + "## 정상 상태 확인과 승인", + "## 시나리오 실행", + "## 결과 확인", + "## 정리", ] - positions = [text.index(command) for command in commands] + positions = [text.index(heading) for heading in headings] assert positions == sorted(positions) + # The commands that carry each step, in the section that owns them. + def section(heading): + return text.split(heading, 1)[1].split("\n## ", 1)[0] + + assert "azd env new" in section("## azd 환경 만들기") + assert "azd up" in section("## 배포") + assert "lab_state.py acknowledge-agent" in section("## 정상 상태 확인과 승인") + assert "scripts/score.py" in section("## 결과 확인") + assert "azd down --purge" in section("## 정리") + scenarios = section("## 시나리오 실행") + for name in ("02-scenario-s1.md", "03-scenario-s2.md", "04-scenario-s3.md"): + assert name in scenarios, name + def test_readme_warns_about_cost_and_teardown_before_the_first_azure_command(): text = README.read_text() @@ -279,8 +287,8 @@ def test_deployment_plan_status_matches_what_was_actually_deployed(): def test_deployment_plan_does_not_claim_the_manual_scenario_sequence_ran(): """The operator said they would connect the Agent and run S1/S2/S3 one - by one themselves. Only the pre-acknowledgement refusal (`lab.sh run s1` - rejected before `acknowledge agent-setup`) was exercised, so the plan + by one themselves. Only the pre-acknowledgement refusal (a scenario + rejected before `acknowledge-agent`) was exercised, so the plan must record the scenario sequence as pending -- in the Live Deployment Proof table, where a reader looks for what the deployment proved. """ @@ -324,7 +332,7 @@ def test_deployment_plan_records_one_current_test_and_bicep_count(): def test_scenario_guides_document_the_critical_recovery_failure_path(): - """`run-scenario.sh` reverts the injected fault from an EXIT trap, and + """The operator reverts the injected fault by hand, and that revert can itself fail (a rejected `az containerapp update`, a revision that never becomes ready, a refused role restore). It then prints `CRITICAL:` and exits non-zero with the fault still live, so each @@ -415,7 +423,7 @@ def test_autonomy_screenshots_warn_that_the_lab_must_choose_review(): def test_scenario_guides_say_a_rerun_retires_the_previous_attempt(): - """`run-scenario.sh` records the new attempt before it injects + """The guide records the new attempt before injecting anything, which clears whatever the previous attempt recorded -- including a `conclusion` that was already unblocking the next scenario. An operator who re-runs a captured scenario to collect a @@ -487,26 +495,24 @@ def test_s3_guide_documents_propagation_wait_and_manual_restore_command(): def test_scenario_guides_lead_with_the_real_azure_commands(): - """`lab.sh run sX` hides the operation it performs, which is the part - worth learning. Each scenario guide states the injection and the revert + """A script that runs the scenario hides the operation it performs, + which is the part worth learning. Each scenario guide states the injection and the revert as commands the operator runs, so the lab teaches the Azure change instead of a wrapper.""" for name, commands in MANUAL_SCENARIO_COMMANDS.items(): text = (GUIDES / name).read_text() manual_index = text.index("## 수동 실행") - shortcut_index = text.index("## 지름길") - assert manual_index < shortcut_index, name - manual_section = text[manual_index:shortcut_index] + manual_section = text[manual_index:].split("\n## ", 1)[0] for command in commands: assert command in manual_section, (name, command) -def test_scenario_guides_keep_every_manual_step_runnable_without_lab_sh(): +def test_scenario_guides_keep_every_manual_step_runnable(): """A manual run still has to produce the same evidence and recorded state the scorer reads, or the manual path dead-ends at scoring.""" for name, scenario in SCENARIO_GUIDES.items(): text = (GUIDES / name).read_text() - manual_section = text[text.index("## 수동 실행"):text.index("## 지름길")] + manual_section = text[text.index("## 수동 실행"):].split("\n## ", 1)[0] assert "scripts/loadgen.py" in manual_section, name assert "az rest" in manual_section, name assert ( @@ -520,12 +526,15 @@ def test_scenario_guides_keep_every_manual_step_runnable_without_lab_sh(): ), name -def test_scenario_guides_present_lab_sh_as_an_optional_shortcut(): - for name, scenario in SCENARIO_GUIDES.items(): - text = (GUIDES / name).read_text() - shortcut_section = text[text.index("## 지름길"):] - assert "./scripts/lab.sh run {0}".format(scenario) in shortcut_section, name - assert "./scripts/lab.sh capture {0}".format(scenario) in shortcut_section, name +def test_no_guide_offers_a_script_that_runs_the_scenario_for_the_operator(): + """A one-command shortcut is the fastest way to finish the lab having + learned nothing: the operator watches a script scroll instead of making + the Azure change and seeing what the Agent does with it. There is no + shortcut section, and nothing dispatches the scenarios.""" + for path in guide_paths() + [README]: + text = path.read_text() + assert "## 지름길" not in text, path.name + assert "lab.sh" not in text, path.name def test_validation_results_keeps_the_one_minute_static_run_and_explains_it(): @@ -631,13 +640,13 @@ def test_readme_links_every_numbered_guide_in_order(): assert positions == sorted(positions) -def test_readme_troubleshooting_index_routes_to_doctor_and_guides(): +def test_readme_troubleshooting_index_routes_to_the_step_documents(): text = README.read_text() heading = "## 문제 해결" assert heading in text section = text.split(heading, 1)[1] - assert "lab.sh doctor" in section + assert "azd env get-value" in section assert "guides/" in section @@ -746,7 +755,6 @@ def test_scenario_guides_use_the_required_section_order(): required = [ "## 시작 조건", "## 수동 실행", - "## 지름길", "## Azure에서 발생하는 변화", "## SRE Agent에서 확인할 항목", "## 성공·부분 성공·실패 판정", @@ -794,8 +802,8 @@ def test_scenario_guides_judge_success_partial_and_failure_with_recovery(): recovery = text.split("## 복구 확인", 1)[1].split("\n## ", 1)[0] assert "Resolved" in recovery, name assert "state.json" in text, name - assert "./scripts/lab.sh run {0}".format(scenario) in text, name - assert "./scripts/lab.sh capture {0}".format(scenario) in text, name + assert "lab_state.py begin-run {0}".format(scenario) in text, name + assert "lab_state.py record-capture {0}".format(scenario) in text, name def test_results_guide_documents_the_scoring_thresholds_and_manual_gap(): @@ -816,18 +824,20 @@ def test_results_guide_documents_the_scoring_thresholds_and_manual_gap(): # --- commands match the scripts ----------------------------------------- -def test_documented_lab_commands_exist_in_lab_sh(): - documented = set(re.findall(r"lab\.sh\s+([a-z-]+)", joined_docs())) - supported = set(re.findall(r"^\s{2}([a-z-]+)\)", LAB_SH.read_text(), re.MULTILINE)) +def test_every_script_the_docs_name_actually_exists(): + """Deleting a script has to delete the instructions that call it, or the + operator meets `No such file or directory` mid-lab.""" + referenced = set(re.findall(r"scripts/([A-Za-z0-9_.-]+\.(?:sh|py))", joined_docs())) - assert documented, "no lab.sh commands are documented" - assert documented <= supported, sorted(documented - supported) + assert referenced, "the documents name no script at all" + for name in sorted(referenced): + assert (LAB_ROOT / "scripts" / name).is_file(), name def test_agent_setup_guide_matches_the_interactive_acknowledge_contract(): text = (GUIDES / "01-agent-setup.md").read_text() - assert "./scripts/lab.sh acknowledge agent-setup" in text + assert "scripts/lab_state.py acknowledge-agent" in text assert "acknowledge" in text assert "표준 입력" in text or "stdin" in text assert re.search(r"환경 변수[^.]{0,60}(대체할 수 없|불가)", text), ( @@ -851,10 +861,10 @@ def test_agent_setup_guide_offers_azd_env_set_without_storing_secrets(): def test_agent_setup_guide_offers_a_python_environment_remedy(): - """Finding #4 (Task 6 follow-up): `doctor.sh` gained a `Python - environment` check, but the guide's failure table never told an - operator what to do about it. The remedy is local-only and independent - of the postprovision-hook ordering contract (see + """The first commands an operator runs need `app/.venv`, and a missing + or half-built one is a local failure with a local fix. The guide's + failure table has to say so: the remedy is independent of the + postprovision-hook ordering contract (see `test_setup_venv_orders_local_recovery_correctly` below), so it may -- and should -- name `./scripts/setup-venv.sh` directly. """ @@ -863,7 +873,7 @@ def test_agent_setup_guide_offers_a_python_environment_remedy(): assert heading in text section = text.split(heading, 1)[1] - assert "Python environment" in section + assert "app/.venv" in section assert "./scripts/setup-venv.sh" in section diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_lab_scripts.py b/monitor/sre-agent-event-lab/scripts/tests/test_lab_scripts.py deleted file mode 100644 index 94d508d..0000000 --- a/monitor/sre-agent-event-lab/scripts/tests/test_lab_scripts.py +++ /dev/null @@ -1,1221 +0,0 @@ -"""Execution tests for the lab's four shell entry points. - -Each script is run as a program against fake `az`/`azd`/`python` -executables, from a working directory that is not the lab. Running them -proves what reading their text cannot: that configuration actually loads, -that no variable a script assigns collides with a `readonly` name -`common.sh` already declared, and that the safety checks run before any -Azure operation. -""" -import json -from pathlib import Path - -import pytest - -from lab_script_harness import ( - AZD_VALUES, - ENV_NAME, - MISSING_CONCLUSION_TIMELINE, - RESOURCE_GROUP, - SUBSCRIPTION_ID, - make_lab, -) - - -CALLERS = ("run-scenario.sh", "query-evidence.sh", "capture-scenario.sh", "cleanup.sh") - -# Short enough that a genuinely unbounded wait fails the test instead of -# hanging the suite, long enough for several poll rounds. -BOUNDED_WAITS = { - "LAB_ALERT_RESOLVE_TIMEOUT_SECONDS": "5", - "LAB_ALERT_RESOLVE_POLL_INTERVAL_SECONDS": "1", - "LAB_RECOVERY_HEALTH_TIMEOUT_SECONDS": "5", - "LAB_REVISION_READY_TIMEOUT_SECONDS": "5", - "LAB_REVISION_READY_POLL_INTERVAL_SECONDS": "1", - "LAB_S3_PROPAGATION_TIMEOUT_SECONDS": "5", - "LAB_S3_PROPAGATION_POLL_INTERVAL_SECONDS": "1", -} - -NO_ALERT_WAITS = dict( - BOUNDED_WAITS, - LAB_ALERT_FIRE_TIMEOUT_SECONDS="3", - LAB_ALERT_FIRE_POLL_INTERVAL_SECONDS="1", -) - - -def captured(scenario, evidence_dir): - """The state entry of a scenario that already ran and captured cleanly.""" - return { - scenario: { - "run_status": "recovered", - "capture_status": "conclusion", - "evidence_dir": str(evidence_dir), - } - } - - -def _assert_loaded_config(result, lab_run): - """Every caller must get past `require_lab_config` + the safety checks.""" - assert "readonly variable" not in result.stderr, ( - "a script assigned a name common.sh already made readonly: " - f"{result.stderr!r}" - ) - assert "azd env set" not in result.stderr, ( - f"configuration failed to load: {result.stderr!r}" - ) - az_calls = lab_run.az_calls() - assert "account show" in az_calls, ( - f"verify_subscription never ran: {az_calls!r} / {result.stderr!r}" - ) - assert 'tags."azd-env-name"' in az_calls, ( - f"verify_lab_resource_group never ran: {az_calls!r}" - ) - assert f"cwd={lab_run.lab}" in lab_run.azd_calls(), ( - "azd lookups must be pinned to the lab project root: " - f"{lab_run.azd_calls()!r}" - ) - - -def test_run_scenario_s1_runs_to_completion_from_another_directory(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - _assert_loaded_config(result, lab_run) - assert result.returncode == 0, result.stderr - evidence_dirs = sorted((lab_run.lab / "evidence").glob("s1-*")) - assert evidence_dirs, "no evidence directory was created" - timeline = json.loads((evidence_dirs[-1] / "timeline.json").read_text()) - assert timeline["scenario"] == "s1" - assert timeline["injected_at"] - assert timeline["alert_id"] - assert timeline["recovered_at"] - assert "FAILURE_MODE=none" in lab_run.az_calls(), "the scenario never recovered" - - # The evidence is only usable if the recorded moments really bracket the - # Azure operations they claim to describe. - update_times = [ - line.split("\t")[1] - for line in lab_run.az_calls().splitlines() - if "containerapp update" in line - ] - assert update_times, "the failure was never injected" - assert timeline["injected_at"] <= update_times[0], ( - "injected_at must be recorded before the Container App is updated" - ) - assert update_times[0] <= timeline["revision_ready_at"] - assert timeline["revision_ready_at"] <= timeline["recovered_at"] - - -def test_query_evidence_collects_every_artifact_from_another_directory(tmp_path): - lab_run = make_lab(tmp_path) - evidence_dir = tmp_path / "evidence-out" - - result = lab_run.run( - "query-evidence.sh", - ["s1", str(evidence_dir), "2026-08-14T00:00:00Z", "2026-08-14T01:00:00Z"], - ) - - _assert_loaded_config(result, lab_run) - assert result.returncode == 0, result.stderr - for artifact in ( - "app-requests.json", - "app-dependencies.json", - "app-exceptions.json", - "activity-log.json", - "alerts.json", - "revisions-redacted.json", - "storage-role-assignments.json", - "query-window.json", - ): - assert (evidence_dir / artifact).is_file(), f"missing {artifact}" - window = json.loads((evidence_dir / "query-window.json").read_text()) - assert window["scenario"] == "s1" - - -def test_query_evidence_queries_the_resolved_workspace_and_principal(tmp_path): - """The values that used to collide with `common.sh`'s readonly names - (the workspace customer ID and the workload principal ID) must reach the - Azure CLI calls that consume them.""" - lab_run = make_lab(tmp_path) - evidence_dir = tmp_path / "evidence-out" - - result = lab_run.run( - "query-evidence.sh", - ["s1", str(evidence_dir), "2026-08-14T00:00:00Z", "2026-08-14T01:00:00Z"], - ) - - assert result.returncode == 0, result.stderr - az_calls = lab_run.az_calls() - assert "--workspace 9d1a0b2c-3d4e-5f60-7182-93a4b5c6d7e8" in az_calls - assert "--assignee-object-id 8c8a4f0e-0000-4000-8000-2b1f9a0c1234" in az_calls - - -def test_run_scenario_refuses_a_scenario_the_state_does_not_allow(tmp_path): - """No baseline and no acknowledgement recorded: the failure must be - injected into nothing at all.""" - lab_run = make_lab(tmp_path) - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - assert "baseline_passed" in result.stderr - assert "agent_setup_acknowledged" in result.stderr - assert "containerapp update" not in lab_run.az_calls(), ( - "run-scenario.sh injected a failure before checking the run order" - ) - assert not sorted((lab_run.lab / "evidence").glob("s1-*")) - - -def test_run_scenario_s2_refuses_to_start_before_s1_was_captured(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.seed_state( - scenarios={"s1": {"run_status": "recovered", "evidence_dir": str(tmp_path / "s1")}} - ) - - result = lab_run.run("run-scenario.sh", ["s2"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - assert "s1_captured" in result.stderr - assert "containerapp update" not in lab_run.az_calls() - - -def test_run_scenario_s2_starts_once_s1_recovered_and_was_captured(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.seed_state(scenarios=captured("s1", tmp_path / "s1")) - - result = lab_run.run("run-scenario.sh", ["s2"], env=BOUNDED_WAITS) - - assert result.returncode == 0, result.stdout + result.stderr - assert "ORDER_DELAY_MS=4000" in lab_run.az_calls() - assert lab_run.scenario_state("s2")["run_status"] == "recovered" - - -def test_run_scenario_records_recovery_only_after_the_alert_resolved(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode == 0, result.stdout + result.stderr - evidence_dir = sorted((lab_run.lab / "evidence").glob("s1-*"))[-1] - scenario_state = lab_run.scenario_state("s1") - assert scenario_state["run_status"] == "recovered" - assert scenario_state["evidence_dir"] == str(evidence_dir) - timeline = json.loads((evidence_dir / "timeline.json").read_text()) - assert timeline["alert_resolved_at"], "the resolved moment was never recorded" - assert timeline["recovered_at"] <= timeline["alert_resolved_at"] - - -def test_run_scenario_fails_when_the_alert_never_resolves(tmp_path): - """An alert Azure Monitor never closed means the workload is not proven - healthy again: the run stays failed, so the next scenario cannot start - on top of an unresolved incident.""" - lab_run = make_lab(tmp_path, alert_resolves=False) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - assert "Resolved" in result.stderr - assert "FAILURE_MODE=none" in lab_run.az_calls(), ( - "the injected failure must still be reverted before the run gives up" - ) - assert lab_run.scenario_state("s1")["run_status"] == "failed" - evidence_dir = sorted((lab_run.lab / "evidence").glob("s1-*"))[-1] - timeline = json.loads((evidence_dir / "timeline.json").read_text()) - assert timeline["alert_resolved_at"] is None - - -def test_run_scenario_marks_failed_when_the_alert_never_fires(tmp_path): - """No alert ever firing is a different failure than one that fires and - never resolves: nothing to recover from Azure Monitor's point of view, - but the run is still unusable evidence. It must be recorded as failed - -- with the evidence directory and a reason -- exactly like every other - failed run, and the fault that was injected must still be reverted by - the same recovery trap that protects every other exit path.""" - lab_run = make_lab(tmp_path, alert_fires=False) - lab_run.seed_state() - - result = lab_run.run( - "run-scenario.sh", - ["s1"], - env=NO_ALERT_WAITS, - ) - - assert result.returncode != 0 - assert "did not fire" in result.stderr - assert "FAILURE_MODE=none" in lab_run.az_calls(), ( - "the injected failure must still be reverted even though no alert ever fired" - ) - scenario_state = lab_run.scenario_state("s1") - assert scenario_state["run_status"] == "failed" - assert "did not fire" in scenario_state.get("failure_reason", "") - evidence_dir = sorted((lab_run.lab / "evidence").glob("s1-*"))[-1] - assert scenario_state["evidence_dir"] == str(evidence_dir) - - -def test_run_scenario_retries_a_transient_alert_list_failure(tmp_path): - lab_run = make_lab(tmp_path, alert_list_failures=1) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode == 0, result.stdout + result.stderr - alert_reads = [ - line - for line in lab_run.az_calls().splitlines() - if "monitorCondition=Fired" in line - ] - assert len(alert_reads) >= 2 - assert lab_run.scenario_state("s1")["run_status"] == "recovered" - - -def test_run_scenario_records_an_unexpected_abort_as_failed(tmp_path): - lab_run = make_lab(tmp_path, loadgen_fails=True) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - assert "FAILURE_MODE=none" in lab_run.az_calls() - scenario = lab_run.scenario_state("s1") - assert scenario["run_status"] == "failed" - assert "aborted" in scenario["failure_reason"] - - -def test_run_scenario_reports_when_aborted_state_cannot_be_recorded(tmp_path): - lab_run = make_lab(tmp_path, loadgen_fails=True, mark_failed_fails=True) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - assert "CRITICAL: could not record the aborted s1 run as failed" in result.stderr - assert lab_run.scenario_state("s1")["run_status"] == "running" - - -def test_run_scenario_records_a_post_recovery_timeline_failure(tmp_path): - lab_run = make_lab(tmp_path, timeline_jq_fails=True) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - scenario = lab_run.scenario_state("s1") - assert scenario["run_status"] == "failed" - assert "aborted" in scenario["failure_reason"] - - -def test_run_scenario_discards_partial_jq_output_from_invalid_alert_json(tmp_path): - lab_run = make_lab(tmp_path, alert_list_invalid_json=True) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=NO_ALERT_WAITS) - - assert result.returncode != 0 - scenario = lab_run.scenario_state("s1") - assert scenario["run_status"] == "failed" - assert "could not query Azure Alerts" in scenario["failure_reason"] - - -def test_run_scenario_records_persistent_alert_api_errors_separately(tmp_path): - lab_run = make_lab(tmp_path, alert_list_failures=100) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=NO_ALERT_WAITS) - - assert result.returncode != 0 - scenario = lab_run.scenario_state("s1") - assert "could not query Azure Alerts" in scenario["failure_reason"] - assert "did not fire" not in scenario["failure_reason"] - error_file = Path(scenario["evidence_dir"]) / "alert-list-error.log" - assert error_file.is_file() - assert "TooManyRequests" in error_file.read_text() - - -def test_run_scenario_reports_alert_errors_after_an_earlier_valid_poll(tmp_path): - lab_run = make_lab( - tmp_path, - alert_fires=False, - alert_failure_after_success=True, - ) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=NO_ALERT_WAITS) - - assert result.returncode != 0 - reason = lab_run.scenario_state("s1")["failure_reason"] - assert "did not fire" in reason - assert "polls failed" in reason - assert "alert-list-error.log" in reason - error_file = Path(lab_run.scenario_state("s1")["evidence_dir"]) / "alert-list-error.log" - assert "Forbidden" in error_file.read_text() - - -def test_run_scenario_rejects_an_empty_successful_alert_response(tmp_path): - lab_run = make_lab(tmp_path, alert_list_empty_body=True) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=NO_ALERT_WAITS) - - assert result.returncode != 0 - reason = lab_run.scenario_state("s1")["failure_reason"] - assert "could not query Azure Alerts" in reason - assert "empty response" in ( - Path(lab_run.scenario_state("s1")["evidence_dir"]) / "alert-list-error.log" - ).read_text() - - -def test_run_scenario_preserves_reason_when_failure_state_write_fails(tmp_path): - lab_run = make_lab( - tmp_path, - alert_fires=False, - mark_failed_fails=True, - ) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=NO_ALERT_WAITS) - - assert result.returncode != 0 - assert "did not fire" in result.stderr - assert "Evidence directory:" in result.stderr - assert "CRITICAL: could not record" in result.stderr - - -def test_run_scenario_retries_the_original_failure_reason_from_the_exit_trap(tmp_path): - lab_run = make_lab( - tmp_path, - alert_fires=False, - mark_failed_failures=1, - ) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=NO_ALERT_WAITS) - - assert result.returncode != 0 - scenario = lab_run.scenario_state("s1") - assert scenario["run_status"] == "failed" - assert "did not fire" in scenario["failure_reason"] - assert "run aborted" not in scenario["failure_reason"] - - -def test_run_scenario_s1_reports_a_rejected_recovery_update_as_critical(tmp_path): - """`recover` runs both directly and from the EXIT trap, and the trap - calls it as `if ! recover`, which turns `set -e` off for the whole - function body. A recovery whose `az containerapp update` was rejected - must therefore report the failure by return value: otherwise the fault - is still injected in a live Container App while the script exits - claiming it recovered.""" - lab_run = make_lab(tmp_path, recovery_update_fails=True) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - assert "CRITICAL" in result.stderr, ( - "a failed recovery must be reported as CRITICAL, not swallowed: " - f"{result.stderr!r}" - ) - assert lab_run.scenario_state("s1").get("run_status") != "recovered", ( - "a run whose recovery failed must never be recorded as recovered" - ) - attempts = [ - line - for line in lab_run.az_calls().splitlines() - if "containerapp update" in line and "FAILURE_MODE=none" in line - ] - assert len(attempts) >= 2, ( - "a failed recovery must not mark itself recovered, so the EXIT trap " - f"has to try again: {attempts!r}" - ) - - -def test_run_scenario_s2_reports_a_stalled_recovery_revision_as_critical(tmp_path): - """The recovery update is accepted but no new healthy revision ever - becomes active: the workload is still slow, so the wait timing out has - to fail the recovery instead of falling through to success.""" - lab_run = make_lab(tmp_path, recovery_revision_stalls=True) - lab_run.seed_state(scenarios=captured("s1", tmp_path / "s1")) - - result = lab_run.run("run-scenario.sh", ["s2"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - assert "CRITICAL" in result.stderr, ( - f"a timed-out recovery wait must be reported as CRITICAL: {result.stderr!r}" - ) - assert lab_run.scenario_state("s2").get("run_status") != "recovered" - assert "ORDER_DELAY_MS=0" in lab_run.az_calls(), "recovery was never attempted" - - -def test_run_scenario_s3_reports_a_failed_role_restore_from_the_exit_trap(tmp_path): - """The S3 fault is a deleted role assignment and the alert never fires, - so recovery only ever runs from the EXIT trap -- the exact path where - `if ! recover` disables `set -e`. A refused `az role assignment create` - must still surface: the workload is left without its blob permission - until an operator restores it.""" - lab_run = make_lab(tmp_path, alert_fires=False, role_create_fails=True) - lab_run.seed_state( - scenarios=dict( - captured("s1", tmp_path / "s1"), **captured("s2", tmp_path / "s2") - ) - ) - - result = lab_run.run("run-scenario.sh", ["s3"], env=NO_ALERT_WAITS) - - assert result.returncode != 0 - assert "role assignment create" in lab_run.az_calls(), ( - "the exit trap never tried to restore the deleted role assignment" - ) - assert "CRITICAL" in result.stderr, ( - "a refused role restore must be reported as CRITICAL, not swallowed: " - f"{result.stderr!r}" - ) - assert lab_run.scenario_state("s3").get("run_status") != "recovered" - - -def test_run_scenario_s3_recovers_and_records_a_successful_run(tmp_path): - """The unchanged happy path: the blob role is restored, the alert - resolves, and the run is recorded as recovered.""" - lab_run = make_lab(tmp_path) - lab_run.seed_state( - scenarios=dict( - captured("s1", tmp_path / "s1"), **captured("s2", tmp_path / "s2") - ) - ) - - result = lab_run.run("run-scenario.sh", ["s3"], env=BOUNDED_WAITS) - - assert result.returncode == 0, result.stdout + result.stderr - assert "CRITICAL" not in result.stderr - assert "role assignment delete" in lab_run.az_calls() - assert "role assignment create" in lab_run.az_calls() - assert lab_run.scenario_state("s3")["run_status"] == "recovered" - - -def test_run_scenario_s3_waits_for_rbac_revocation_before_the_final_load(tmp_path): - lab_run = make_lab(tmp_path, s3_probe_failures=2) - lab_run.seed_state( - scenarios=dict( - captured("s1", tmp_path / "s1"), **captured("s2", tmp_path / "s2") - ) - ) - - result = lab_run.run("run-scenario.sh", ["s3"], env=BOUNDED_WAITS) - - assert result.returncode == 0, result.stdout + result.stderr - calls = lab_run.python_calls().splitlines() - probe_indexes = [ - index - for index, call in enumerate(calls) - if "/api/documents" in call and "--requests 1" in call - ] - final_index = next( - index - for index, call in enumerate(calls) - if "/api/documents" in call and "--requests 60" in call - ) - assert len(probe_indexes) >= 3 - assert max(probe_indexes) < final_index - - -def test_run_scenario_s3_records_the_propagation_timeout_reason(tmp_path): - lab_run = make_lab(tmp_path, s3_probe_failures=1000) - lab_run.seed_state( - scenarios=dict( - captured("s1", tmp_path / "s1"), **captured("s2", tmp_path / "s2") - ) - ) - waits = dict(BOUNDED_WAITS, LAB_S3_PROPAGATION_TIMEOUT_SECONDS="2") - - result = lab_run.run("run-scenario.sh", ["s3"], env=waits) - - assert result.returncode != 0 - scenario = lab_run.scenario_state("s3") - assert scenario["run_status"] == "failed" - assert "did not produce HTTP 503 within 2s" in scenario["failure_reason"] - - -def test_run_scenario_s3_refuses_missing_recovery_outputs_before_deletion(tmp_path): - values = dict(AZD_VALUES) - values["containerAppPrincipalId"] = "" - values["AZURE_STORAGE_CONTAINER_SCOPE"] = "" - values["AZURE_BLOB_ROLE_ASSIGNMENT_NAME"] = "" - lab_run = make_lab(tmp_path, azd_values=values) - lab_run.seed_state( - scenarios=dict( - captured("s1", tmp_path / "s1"), **captured("s2", tmp_path / "s2") - ) - ) - - result = lab_run.run("run-scenario.sh", ["s3"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - assert "containerAppPrincipalId" in result.stderr - assert "storageContainerScope" in result.stderr - assert "blobRoleAssignmentName" in result.stderr - assert "azd provision" in result.stderr - assert "role assignment delete" not in lab_run.az_calls() - assert not sorted((lab_run.lab / "evidence").glob("s3-*")) - - -def test_a_failed_run_blocks_the_next_scenario(tmp_path): - lab_run = make_lab(tmp_path, alert_resolves=False) - lab_run.seed_state() - first = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert first.returncode != 0 - - result = lab_run.run("run-scenario.sh", ["s2"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - assert "s1_recovered" in result.stderr - - -# --- A re-run that dies early never leaves the previous success standing --- - - -@pytest.mark.parametrize("break_the_rerun", ("injection", "recovery")) -def test_a_rerun_that_dies_early_retires_the_previous_success(tmp_path, break_the_rerun): - """S1 recovers and captures a real conclusion, then is re-run and the - re-run fails *before* it can record any outcome of its own -- the - injecting `az containerapp update` is rejected, or the recovery is and - the EXIT trap gives up. - - Without an attempt recorded before the first destructive call, the - scenario entry still read `recovered` + `conclusion` from the run that - was just superseded, so `run-scenario.sh s2` was admitted and injected - a second fault into a workload whose first incident had not been - reproduced. The started attempt has to clear that, so every later gate - -- the next scenario and the scorer -- refuses. - """ - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - first_run = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert first_run.returncode == 0, first_run.stderr - first_capture = lab_run.run("capture-scenario.sh", ["s1"]) - assert first_capture.returncode == 0, first_capture.stderr - finished = lab_run.scenario_state("s1") - assert set(finished) == { - "run_status", - "started_at", - "capture_status", - "evidence_dir", - }, finished - assert finished["run_status"] == "recovered" - assert finished["capture_status"] == "conclusion" - first_evidence_dir = finished["evidence_dir"] - - if break_the_rerun == "injection": - lab_run.break_injection() - else: - lab_run.break_recovery() - rerun = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert rerun.returncode != 0, rerun.stdout - entry = lab_run.scenario_state("s1") - assert entry.get("run_status") in ("running", "failed"), entry - assert "capture_status" not in entry, ( - "a conclusion captured against the superseded run must not survive " - f"the re-run: {entry!r}" - ) - assert entry.get("evidence_dir") != first_evidence_dir, ( - "the re-run must not keep pointing at the previous attempt's evidence" - ) - - blocked = lab_run.run("run-scenario.sh", ["s2"], env=BOUNDED_WAITS) - - assert blocked.returncode != 0, blocked.stdout - assert "s1_recovered" in blocked.stderr - assert "ORDER_DELAY_MS=4000" not in lab_run.az_calls(), ( - "S2 injected its fault although S1's re-run never recovered" - ) - scored = lab_run.run("lab.sh", ["score"]) - assert scored.returncode != 0, scored.stdout - assert "lab.sh run" in scored.stderr - - -# --- One unfinished run stops the whole lab, not just the next scenario ---- - - -def finish_scenario(lab_run, scenario): - """Run and capture one scenario end to end, the way an operator does.""" - run_result = lab_run.run("run-scenario.sh", [scenario], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stdout + run_result.stderr - capture = lab_run.run("capture-scenario.sh", [scenario]) - assert capture.returncode == 0, capture.stdout + capture.stderr - assert lab_run.scenario_state(scenario)["capture_status"] == "conclusion" - - -def test_a_broken_s1_rerun_stops_s3_although_s2_is_still_captured(tmp_path): - """The gap this closes, end to end. - - All three scenarios run and capture cleanly, then S1 is re-run and the - re-run dies before it can record an outcome. S1 is now `running` or - `failed` -- its fault may still be live in the shared Container App -- - but S2's entry is untouched, still `recovered` + `conclusion`. The - ordered rules only look one scenario back, so `run-scenario.sh s3` read - S2's stale success and was admitted: a third fault injected on top of an - incident nobody had resolved, and two captures that can no longer be - told apart. - """ - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - for scenario in ("s1", "s2", "s3"): - finish_scenario(lab_run, scenario) - lab_run.break_injection() - rerun = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert rerun.returncode != 0, rerun.stdout - assert lab_run.scenario_state("s1")["run_status"] in ("running", "failed") - assert lab_run.scenario_state("s2")["capture_status"] == "conclusion" - az_before = lab_run.az_calls() - - blocked = lab_run.run("run-scenario.sh", ["s3"], env=BOUNDED_WAITS) - - assert blocked.returncode != 0, blocked.stdout - assert "s1" in blocked.stderr - new_calls = lab_run.az_calls()[len(az_before):] - assert "containerapp update" not in new_calls, new_calls - assert "role assignment delete" not in new_calls, ( - f"a refused run injected S3's fault anyway: {new_calls!r}" - ) - assert not sorted((lab_run.lab / "evidence").glob("s3-*"))[1:], ( - "a refused run must not leave a second S3 evidence directory behind" - ) - - -def test_a_refused_run_leaves_no_evidence_directory_behind(tmp_path): - """The evidence directory is registered with the run, so its path has to - exist as a string before `begin-run` -- but the directory itself must - only be created once the run was admitted. Otherwise every refusal - litters `evidence/` with an empty `sN-/` that reads exactly - like an attempt that ran and produced nothing. - """ - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - finish_scenario(lab_run, "s1") - lab_run.break_injection() - assert lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS).returncode != 0 - before = sorted(path.name for path in (lab_run.lab / "evidence").glob("s2-*")) - - blocked = lab_run.run("run-scenario.sh", ["s2"], env=BOUNDED_WAITS) - - assert blocked.returncode != 0, blocked.stdout - after = sorted(path.name for path in (lab_run.lab / "evidence").glob("s2-*")) - assert after == before, f"a refused run created {set(after) - set(before)}" - - -def test_the_evidence_directory_is_created_only_after_the_run_is_admitted(tmp_path): - """Ordering, observed at the moment it matters: when `begin-run` is - called the directory must not exist yet, and by the time the run does - its work it must.""" - lab_run = make_lab(tmp_path) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode == 0, result.stdout + result.stderr - probes = lab_run.begin_run_probes() - assert probes, "begin-run was never called" - assert [existed for existed, _ in probes] == ["absent"], probes - registered = lab_run.scenario_state("s1")["evidence_dir"] - assert probes[0][1] == registered, ( - "the path registered with the run must be the one that was created" - ) - assert (lab_run.lab / registered).is_dir() or Path(registered).is_dir() - - -def test_a_running_scenario_cannot_be_started_a_second_time(tmp_path): - """A run left `running` -- a Ctrl-C, a crashed terminal -- must not be - restarted blindly: two live injections of the same fault leave neither - capture readable. The operator has to record how the first one ended.""" - lab_run = make_lab(tmp_path) - lab_run.seed_state( - scenarios={"s1": {"run_status": "running", "started_at": "2026-08-14T00:00:00Z"}} - ) - az_before = lab_run.az_calls() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode != 0, result.stdout - assert "running" in result.stderr - assert "mark-failed s1" in result.stderr - new_calls = lab_run.az_calls()[len(az_before):] - assert "containerapp update" not in new_calls, new_calls - - -def test_capture_scenario_reports_a_refused_record_without_losing_evidence(tmp_path): - """`record-capture` now refuses a conclusion for a run that did not - recover. The capture pipeline has already written real files by then, so - the failure must say so and name where they are -- not exit on an - unexplained non-zero from a command substitution.""" - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - document = json.loads(lab_run.state_path.read_text()) - document["scenarios"]["s1"]["run_status"] = "failed" - lab_run.state_path.write_text(json.dumps(document)) - - result = lab_run.run("capture-scenario.sh", ["s1"]) - - assert result.returncode != 0, result.stdout - assert "recovered" in result.stderr - evidence_dir = lab_run.scenario_state("s1")["evidence_dir"] - assert evidence_dir in result.stderr, ( - f"the operator must be told the raw evidence survived: {result.stderr!r}" - ) - assert (Path(evidence_dir) / "normalized-timeline.json").is_file() - assert "capture_status" not in lab_run.scenario_state("s1") - - -def test_a_started_run_is_recorded_before_the_fault_is_injected(tmp_path): - """Ordering is the whole point: the attempt must be persisted *before* - the first destructive Azure call, because that call is what can fail - and leave nothing else to write the state.""" - lab_run = make_lab(tmp_path, injection_update_fails=True) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - entry = lab_run.scenario_state("s1") - assert entry.get("run_status") in ("running", "failed"), entry - assert entry.get("started_at", "").endswith("Z"), entry - evidence_dirs = sorted((lab_run.lab / "evidence").glob("s1-*")) - assert entry.get("evidence_dir") == str(evidence_dirs[-1]) - - -def test_a_started_run_is_completed_by_a_healthy_run(tmp_path): - """The started attempt is a transition, not a terminal state: a run - that recovers must end as `recovered`, with no `running` left behind.""" - lab_run = make_lab(tmp_path) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode == 0, result.stdout + result.stderr - assert lab_run.scenario_state("s1")["run_status"] == "recovered" - - -def test_run_scenario_refuses_a_state_file_from_another_environment(tmp_path): - """A `state.json` left behind by another lab must never unlock a run - here: the file records the environment, subscription and resource group - it belongs to, and every command checks them.""" - lab_run = make_lab(tmp_path) - lab_run.seed_state(environment="sre-lab-somewhere-else") - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode != 0 - assert "sre-lab-somewhere-else" in result.stderr - assert "containerapp update" not in lab_run.az_calls() - - -def test_run_scenario_binds_new_state_to_the_current_environment(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.state_path.unlink(missing_ok=True) - lab_run.seed_state() - - result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - - assert result.returncode == 0, result.stdout + result.stderr - state = lab_run.state() - assert state["environment"] == ENV_NAME - assert state["subscription_id"] == SUBSCRIPTION_ID - assert state["resource_group"] == RESOURCE_GROUP - - -def test_capture_scenario_resolves_the_evidence_directory_from_the_state(tmp_path): - """The public command is `lab.sh capture s1` -- no timestamped path -- - so `capture-scenario.sh` has to find the directory the recorded run - wrote, not the newest directory that happens to be on disk.""" - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - evidence_dir = sorted((lab_run.lab / "evidence").glob("s1-*"))[-1] - - result = lab_run.run("capture-scenario.sh", ["s1"]) - - _assert_loaded_config(result, lab_run) - assert result.returncode == 0, result.stdout + result.stderr - assert (evidence_dir / "normalized-timeline.json").is_file() - assert (lab_run.lab / "assets" / "captures" / "s1" / "investigation.gif").is_file() - assert lab_run.scenario_state("s1")["capture_status"] == "conclusion" - - -def test_capture_scenario_refuses_an_explicit_evidence_directory(tmp_path): - """The legacy second argument let a capture of *any* directory be - recorded as this environment's current capture status -- re-rendering an - old run would unblock the next scenario on evidence that does not belong - to the alert being captured. The public script takes the scenario only; - regenerating artifacts from an archived directory is a `capture_agent.py` - / `render_capture.py` job, which records no state.""" - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - stale_dir = tmp_path / "evidence-out" - stale_dir.mkdir() - (stale_dir / "timeline.json").write_text( - json.dumps({"scenario": "s1", "alert_id": "/alerts/aaaa0000"}) - ) - - result = lab_run.run("capture-scenario.sh", ["s1", str(stale_dir)]) - - assert result.returncode == 2, result.stdout + result.stderr - assert "Usage:" in result.stderr - assert str(stale_dir) not in result.stderr - assert not (stale_dir / "normalized-timeline.json").exists(), ( - "an explicit directory must never be captured" - ) - assert not lab_run.scenario_state("s1"), ( - "a rejected invocation must not record a capture status" - ) - - -def test_capture_scenario_usage_documents_only_the_scenario_argument(tmp_path): - lab_run = make_lab(tmp_path) - - result = lab_run.run("capture-scenario.sh", []) - - assert result.returncode == 2 - assert "Usage:" in result.stderr - assert "EVIDENCE_DIR" not in result.stderr - - -def test_capture_scenario_records_a_missing_conclusion_as_itself(tmp_path): - lab_run = make_lab(tmp_path, capture_timeline=MISSING_CONCLUSION_TIMELINE) - lab_run.write_agent_setup() - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - - result = lab_run.run("capture-scenario.sh", ["s1"]) - - assert result.returncode == 0, result.stdout + result.stderr - assert lab_run.scenario_state("s1")["capture_status"] == "conclusion-missing" - assert "conclusion-missing" in result.stdout - blocked = lab_run.run("run-scenario.sh", ["s2"], env=BOUNDED_WAITS) - assert blocked.returncode != 0 - assert "s1_captured" in blocked.stderr - - -def test_capture_scenario_refuses_to_replace_a_successful_capture(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - first = lab_run.run("capture-scenario.sh", ["s1"]) - assert first.returncode == 0, first.stdout + first.stderr - calls_before = lab_run.lab_python_log.read_text().count("capture_agent.py") - - second = lab_run.run("capture-scenario.sh", ["s1"]) - - assert second.returncode != 0 - assert "already has a conclusion" in second.stderr - assert "lab.sh run s1" in second.stderr - assert lab_run.scenario_state("s1")["capture_status"] == "conclusion" - assert lab_run.lab_python_log.read_text().count("capture_agent.py") == calls_before - - -def test_capture_scenario_render_failure_names_the_direct_retry(tmp_path): - lab_run = make_lab(tmp_path, render_fails=True) - lab_run.write_agent_setup() - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - first = lab_run.run("capture-scenario.sh", ["s1"]) - assert first.returncode != 0 - assert lab_run.scenario_state("s1")["capture_status"] == "conclusion" - - second = lab_run.run("capture-scenario.sh", ["s1"]) - - assert second.returncode != 0 - assert "render_capture.py" in second.stderr - assert "normalized-timeline.json" in second.stderr - - -def test_query_evidence_refuses_missing_outputs_before_writing_artifacts(tmp_path): - values = dict(AZD_VALUES) - values["containerAppPrincipalId"] = "" - values["AZURE_STORAGE_CONTAINER_SCOPE"] = "" - lab_run = make_lab(tmp_path, azd_values=values) - evidence_dir = tmp_path / "evidence-out" - - result = lab_run.run( - "query-evidence.sh", - ["s1", str(evidence_dir), "2026-08-14T00:00:00Z", "2026-08-14T01:00:00Z"], - ) - - assert result.returncode != 0 - assert "containerAppPrincipalId" in result.stderr - assert "storageContainerScope" in result.stderr - assert "azd provision" in result.stderr - assert not evidence_dir.exists() - - -def test_capture_scenario_missing_setup_names_the_file_and_guide(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - - result = lab_run.run("capture-scenario.sh", ["s1"]) - - assert result.returncode != 0 - assert "evidence/agent-setup.json" in result.stderr - assert "guides/01-agent-setup.md" in result.stderr - - -def test_capture_scenario_missing_endpoint_names_the_setup_file(tmp_path): - lab_run = make_lab(tmp_path) - setup_path = lab_run.write_agent_setup() - setup = json.loads(setup_path.read_text()) - setup["agent_endpoint"] = "" - setup_path.write_text(json.dumps(setup)) - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - - result = lab_run.run("capture-scenario.sh", ["s1"]) - - assert result.returncode != 0 - assert "agent_endpoint" in result.stderr - assert "evidence/agent-setup.json" in result.stderr - assert "guides/01-agent-setup.md" in result.stderr - - -def test_capture_scenario_rejects_a_placeholder_endpoint(tmp_path): - lab_run = make_lab(tmp_path) - setup_path = lab_run.write_agent_setup() - setup = json.loads(setup_path.read_text()) - setup["agent_endpoint"] = "https://..azuresre.ai" - setup_path.write_text(json.dumps(setup)) - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - - result = lab_run.run("capture-scenario.sh", ["s1"]) - - assert result.returncode != 0 - assert "valid HTTPS" in result.stderr - assert "agent-setup.json" in result.stderr - - -def test_capture_scenario_missing_alert_id_names_the_timeline_and_rerun(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - evidence_dir = Path(lab_run.scenario_state("s1")["evidence_dir"]) - timeline_path = evidence_dir / "timeline.json" - timeline = json.loads(timeline_path.read_text()) - timeline["alert_id"] = "" - timeline_path.write_text(json.dumps(timeline)) - - result = lab_run.run("capture-scenario.sh", ["s1"]) - - assert result.returncode != 0 - assert str(timeline_path) in result.stderr - assert "lab.sh run s1" in result.stderr - - -def test_capture_scenario_without_a_recorded_run_names_the_command_to_run(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - lab_run.seed_state() - - result = lab_run.run("capture-scenario.sh", ["s1"]) - - assert result.returncode != 0 - assert "No such file or directory" not in result.stderr - assert "lab.sh run s1" in result.stderr - - -def test_capture_scenario_fails_actionably_when_venv_is_missing(tmp_path): - """Finding #1: cloud resources (the alert rules, the app, etc.) may - already be deployed by the time this local-only precondition fails, so - the message must name the exact rerun command, not just what's wrong.""" - lab_run = make_lab(tmp_path, venv_present=False) - lab_run.write_agent_setup() - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - - result = lab_run.run("capture-scenario.sh", ["s1"]) - - assert result.returncode != 0 - assert "Missing Python environment" in result.stderr - assert "setup-venv.sh" in result.stderr - - -def test_capture_scenario_fails_actionably_when_pillow_is_not_importable(tmp_path): - lab_run = make_lab(tmp_path, pillow_importable=False) - lab_run.write_agent_setup() - lab_run.seed_state() - run_result = lab_run.run("run-scenario.sh", ["s1"], env=BOUNDED_WAITS) - assert run_result.returncode == 0, run_result.stderr - - result = lab_run.run("capture-scenario.sh", ["s1"]) - - assert result.returncode != 0 - assert "Pillow" in result.stderr - assert "setup-venv.sh" in result.stderr - - -def test_cleanup_dry_run_plans_without_deleting_from_another_directory(tmp_path): - """`cleanup.sh` is a compatibility wrapper around the external cleanup - `azd down` runs: it plans the recorded role-assignment removal and never - proposes a resource-group deletion of its own.""" - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - - result = lab_run.run("cleanup.sh") - - assert result.returncode == 0, result.stderr - assert "azd down --purge" in result.stdout - assert "Planned external cleanup" in result.stdout - assert f"Delete tagged resource group: {RESOURCE_GROUP}" not in result.stdout - assert "Dry run only" in result.stdout - az_calls = lab_run.az_calls() - assert "group delete" not in az_calls, "a dry run must delete nothing" - assert "role assignment delete" not in az_calls - - -def test_cleanup_deletes_only_the_recorded_external_assignments(tmp_path): - """Even with --yes, the wrapper must not delete a resource group: that - is `azd down`'s job, and a broad deletion here would take resources azd - never created with it.""" - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - - result = lab_run.run("cleanup.sh", ["--yes"]) - - assert result.returncode == 0, result.stderr - az_calls = lab_run.az_calls() - assert "role assignment delete --ids /subscriptions/" in az_calls - assert "group delete" not in az_calls - - -def test_cleanup_legacy_flag_deletes_the_tagged_resource_group(tmp_path): - """The pre-azd resource groups still have to be recoverable by hand, so - the broad deletion stays available behind an explicit flag -- after the - same tag and subscription checks it always ran.""" - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - - result = lab_run.run( - "cleanup.sh", ["--legacy-delete-resource-group", "--yes"] - ) - - _assert_loaded_config(result, lab_run) - assert result.returncode == 0, result.stderr - az_calls = lab_run.az_calls() - assert f"group delete --name {RESOURCE_GROUP} --yes --no-wait" in az_calls - assert "role assignment delete --ids /subscriptions/" in az_calls - - -def test_cleanup_legacy_dry_run_plans_the_resource_group_deletion(tmp_path): - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - - result = lab_run.run("cleanup.sh", ["--legacy-delete-resource-group"]) - - assert result.returncode == 0, result.stderr - assert f"Delete tagged resource group: {RESOURCE_GROUP}" in result.stdout - assert "group delete" not in lab_run.az_calls() - - -@pytest.mark.parametrize("script_name", CALLERS) -def test_every_caller_fails_closed_when_configuration_is_missing(script_name, tmp_path): - """No azd value and no explicit environment: every entry point must stop - with the actionable `azd env set` message before touching Azure.""" - lab_run = make_lab(tmp_path, azd_values={}) - lab_run.write_agent_setup() - arguments = { - "run-scenario.sh": ["s1"], - "query-evidence.sh": [ - "s1", - str(tmp_path / "out"), - "2026-08-14T00:00:00Z", - "2026-08-14T01:00:00Z", - ], - "capture-scenario.sh": ["s1"], - "cleanup.sh": [], - }[script_name] - - result = lab_run.run(script_name, arguments) - - assert result.returncode != 0 - assert "azd env set AZURE_SUBSCRIPTION_ID" in result.stderr - assert "group delete" not in lab_run.az_calls() - assert "containerapp update" not in lab_run.az_calls() - - -@pytest.mark.parametrize("script_name", CALLERS) -def test_every_caller_refuses_a_foreign_subscription(script_name, tmp_path): - """The subscription-equality boundary must hold for every entry point.""" - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - arguments = { - "run-scenario.sh": ["s1"], - "query-evidence.sh": [ - "s1", - str(tmp_path / "out"), - "2026-08-14T00:00:00Z", - "2026-08-14T01:00:00Z", - ], - "capture-scenario.sh": ["s1"], - "cleanup.sh": ["--yes"], - }[script_name] - - result = lab_run.run( - script_name, - arguments, - env={"AZURE_SUBSCRIPTION_ID": "99999999-9999-9999-9999-999999999999"}, - ) - - assert result.returncode != 0 - assert "Refusing to continue in subscription" in result.stderr - assert "Expected 99999999" in result.stderr - az_calls = lab_run.az_calls() - assert "group delete" not in az_calls - assert "containerapp update" not in az_calls - assert "role assignment delete" not in az_calls - - -def test_environment_name_tag_mismatch_stops_every_caller(tmp_path): - """A resource group tagged for another azd environment is refused even - when its purpose tag matches -- including on the legacy recovery path, - the only one that still deletes a resource group.""" - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - - result = lab_run.run( - "cleanup.sh", - ["--legacy-delete-resource-group", "--yes"], - env={"AZURE_ENV_NAME": f"{ENV_NAME}-other"}, - ) - - assert result.returncode != 0 - assert "Refusing to operate on untagged resource group" in result.stderr - assert "group delete" not in lab_run.az_calls() - - -def test_subscription_id_is_read_from_the_azd_environment(tmp_path): - """Nothing in the callers may pin a subscription of its own.""" - lab_run = make_lab(tmp_path) - lab_run.write_agent_setup() - - result = lab_run.run("cleanup.sh") - - assert result.returncode == 0, result.stderr - assert SUBSCRIPTION_ID in lab_run.az_calls() diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_lab_state.py b/monitor/sre-agent-event-lab/scripts/tests/test_lab_state.py index bc5bafc..97c3257 100644 --- a/monitor/sre-agent-event-lab/scripts/tests/test_lab_state.py +++ b/monitor/sre-agent-event-lab/scripts/tests/test_lab_state.py @@ -11,6 +11,7 @@ import importlib.util import json import os +import pathlib import subprocess import sys from pathlib import Path @@ -42,6 +43,19 @@ def load_module(): } +def test_every_scenario_has_a_guide_to_send_an_operator_to(): + """A refusal names the document that walks the step. If the two lists + drift, the message degrades to "the matching guide under guides/" -- + which is a direction, not an answer -- so they are pinned together. + """ + import lab_state + + assert set(lab_state.SCENARIO_GUIDES) == set(lab_state.SCENARIOS) + lab_root = pathlib.Path(__file__).parents[2] + for scenario, guide in lab_state.SCENARIO_GUIDES.items(): + assert (lab_root / guide).is_file(), (scenario, guide) + + def run_cli(state_path, args, stdin="", env=None): process_env = dict(os.environ) process_env.update(ENVIRONMENT) @@ -314,7 +328,7 @@ def test_begin_run_blocks_the_next_scenario_and_the_capture_gate(tmp_path): def test_begin_run_without_an_evidence_directory_drops_the_previous_one(tmp_path): - """`capture-scenario.sh` captures whatever directory the state names. + """The capture step uses whatever directory the state names. Keeping the finished attempt's directory across a new attempt would let a capture of the *old* timeline be recorded as this attempt's outcome.""" path = tmp_path / "state.json" @@ -563,7 +577,7 @@ def test_the_refusal_lists_every_blocker_earliest_first(tmp_path): assert message.index("s2") < message.index("s3"), message assert "running" in message and "failed" in message assert "mark-failed s2" in message - assert "lab.sh run s3" in message + assert "guides/04-scenario-s3.md" in message def test_a_repair_is_refused_while_an_earlier_run_is_still_running(tmp_path): @@ -595,7 +609,7 @@ def test_the_refusal_names_the_blocking_scenario_its_status_and_a_remedy(tmp_pat message = str(refusal.value) assert "s2" in message assert "failed" in message - assert "lab.sh run s2" in message, message + assert "guides/03-scenario-s2.md" in message, message def test_the_refusal_for_a_running_scenario_names_how_to_end_it(tmp_path): @@ -627,7 +641,7 @@ def test_the_ordered_remedy_never_tells_an_operator_to_restart_a_running_run(tmp message = str(refusal.value) assert "s1_recovered" in message - assert "lab.sh run s1" not in message, message + assert "guides/02-scenario-s1.md" not in message, message assert "mark-failed s1" in message, message @@ -768,9 +782,9 @@ def test_a_conclusion_cannot_be_recorded_against_a_run_that_never_recovered( @pytest.mark.parametrize( "run_status, expected, forbidden", ( - ("running", "mark-failed s1", "lab.sh run s1"), - ("failed", "lab.sh run s1", "mark-failed s1"), - (None, "lab.sh run s1", "mark-failed s1"), + ("running", "mark-failed s1", "guides/02-scenario-s1.md"), + ("failed", "guides/02-scenario-s1.md", "mark-failed s1"), + (None, "guides/02-scenario-s1.md", "mark-failed s1"), ), ids=("running", "failed", "none"), ) @@ -1202,12 +1216,12 @@ def test_cli_evidence_dir_without_a_run_names_the_command_to_run(tmp_path): result = run_cli(tmp_path / "state.json", ["evidence-dir", "s1"]) assert result.returncode == 1 - assert "lab.sh run s1" in result.stderr + assert "guides/02-scenario-s1.md" in result.stderr assert "Traceback" not in result.stderr def test_cli_begin_run_starts_a_new_attempt_and_blocks_the_next_scenario(tmp_path): - """The command `run-scenario.sh` runs between `require-run` and the + """The command an operator runs between `require-run` and the first destructive Azure call: it must clear the finished attempt and leave the scenario `running`, which satisfies no gate.""" path = tmp_path / "state.json" @@ -1276,7 +1290,7 @@ def test_cli_require_run_refuses_every_scenario_while_one_run_is_unfinished(tmp_ assert refused.returncode == 1, refused.stdout assert "s3" in refused.stderr assert "failed" in refused.stderr - assert "lab.sh run s3" in refused.stderr + assert "guides/04-scenario-s3.md" in refused.stderr assert "Traceback" not in refused.stderr assert run_cli(path, ["require-run", "s3"]).returncode == 0 diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_query_evidence.py b/monitor/sre-agent-event-lab/scripts/tests/test_query_evidence.py new file mode 100644 index 0000000..2304f04 --- /dev/null +++ b/monitor/sre-agent-event-lab/scripts/tests/test_query_evidence.py @@ -0,0 +1,162 @@ +"""Execution tests for `query-evidence.sh`. + +The lab is walked by hand: the guides run `az` and the Python tools +directly. `query-evidence.sh` is the one shell entry point an operator +still invokes, because collecting a scenario's evidence means running eight +queries whose results have to land in files `score.py` can read. + +It is run here as a program against fake `az`/`azd` executables, from a +working directory that is not the lab. Running it proves what reading its +text cannot: that configuration actually loads, that no variable it assigns +collides with a `readonly` name `common.sh` already declared, and that the +subscription and resource-group checks run before any Azure call. +""" +import json + +from lab_script_harness import ( + AZD_VALUES, + ENV_NAME, + SUBSCRIPTION_ID, + make_lab, +) + + +WINDOW = ["2026-08-14T00:00:00Z", "2026-08-14T01:00:00Z"] + + +def query_arguments(evidence_dir): + return ["s1", str(evidence_dir), *WINDOW] + + +def _assert_loaded_config(result, lab_run): + """The script must get past `require_lab_config` + the safety checks.""" + assert "readonly variable" not in result.stderr, ( + "a script assigned a name common.sh already made readonly: " + f"{result.stderr!r}" + ) + assert "azd env set" not in result.stderr, ( + f"configuration failed to load: {result.stderr!r}" + ) + az_calls = lab_run.az_calls() + assert "account show" in az_calls, ( + f"verify_subscription never ran: {az_calls!r} / {result.stderr!r}" + ) + assert 'tags."azd-env-name"' in az_calls, ( + f"verify_lab_resource_group never ran: {az_calls!r}" + ) + assert f"cwd={lab_run.lab}" in lab_run.azd_calls(), ( + "azd lookups must be pinned to the lab project root: " + f"{lab_run.azd_calls()!r}" + ) + + +def test_query_evidence_collects_every_artifact_from_another_directory(tmp_path): + lab_run = make_lab(tmp_path) + evidence_dir = tmp_path / "evidence-out" + + result = lab_run.run("query-evidence.sh", query_arguments(evidence_dir)) + + _assert_loaded_config(result, lab_run) + assert result.returncode == 0, result.stderr + for artifact in ( + "app-requests.json", + "app-dependencies.json", + "app-exceptions.json", + "activity-log.json", + "alerts.json", + "revisions-redacted.json", + "storage-role-assignments.json", + "query-window.json", + ): + assert (evidence_dir / artifact).is_file(), f"missing {artifact}" + window = json.loads((evidence_dir / "query-window.json").read_text()) + assert window["scenario"] == "s1" + + +def test_query_evidence_queries_the_resolved_workspace_and_principal(tmp_path): + """The values that used to collide with `common.sh`'s readonly names + (the workspace customer ID and the workload principal ID) must reach the + Azure CLI calls that consume them.""" + lab_run = make_lab(tmp_path) + evidence_dir = tmp_path / "evidence-out" + + result = lab_run.run("query-evidence.sh", query_arguments(evidence_dir)) + + assert result.returncode == 0, result.stderr + az_calls = lab_run.az_calls() + assert "--workspace 9d1a0b2c-3d4e-5f60-7182-93a4b5c6d7e8" in az_calls + assert "--assignee-object-id 8c8a4f0e-0000-4000-8000-2b1f9a0c1234" in az_calls + + +def test_query_evidence_refuses_missing_outputs_before_writing_artifacts(tmp_path): + values = dict(AZD_VALUES) + values["containerAppPrincipalId"] = "" + values["AZURE_STORAGE_CONTAINER_SCOPE"] = "" + lab_run = make_lab(tmp_path, azd_values=values) + evidence_dir = tmp_path / "evidence-out" + + result = lab_run.run("query-evidence.sh", query_arguments(evidence_dir)) + + assert result.returncode != 0 + assert "containerAppPrincipalId" in result.stderr + assert "storageContainerScope" in result.stderr + assert "azd provision" in result.stderr + assert not evidence_dir.exists() + + +def test_query_evidence_fails_closed_when_configuration_is_missing(tmp_path): + """No azd value and no explicit environment: the entry point must stop + with the actionable `azd env set` message before touching Azure.""" + lab_run = make_lab(tmp_path, azd_values={}) + evidence_dir = tmp_path / "out" + + result = lab_run.run("query-evidence.sh", query_arguments(evidence_dir)) + + assert result.returncode != 0 + assert "azd env set AZURE_SUBSCRIPTION_ID" in result.stderr + assert "monitor log-analytics" not in lab_run.az_calls() + + +def test_query_evidence_refuses_a_foreign_subscription(tmp_path): + """The subscription-equality boundary must hold before any query runs.""" + lab_run = make_lab(tmp_path) + evidence_dir = tmp_path / "out" + + result = lab_run.run( + "query-evidence.sh", + query_arguments(evidence_dir), + env={"AZURE_SUBSCRIPTION_ID": "99999999-9999-9999-9999-999999999999"}, + ) + + assert result.returncode != 0 + assert "Refusing to continue in subscription" in result.stderr + assert "Expected 99999999" in result.stderr + assert "monitor log-analytics" not in lab_run.az_calls() + + +def test_a_resource_group_tagged_for_another_environment_is_refused(tmp_path): + """The purpose tag alone is not enough: a group belonging to a different + azd environment is someone else's lab.""" + lab_run = make_lab(tmp_path) + evidence_dir = tmp_path / "out" + + result = lab_run.run( + "query-evidence.sh", + query_arguments(evidence_dir), + env={"AZURE_ENV_NAME": f"{ENV_NAME}-other"}, + ) + + assert result.returncode != 0 + assert "Refusing to operate on untagged resource group" in result.stderr + assert "monitor log-analytics" not in lab_run.az_calls() + + +def test_subscription_id_is_read_from_the_azd_environment(tmp_path): + """Nothing in the lab may pin a subscription of its own.""" + lab_run = make_lab(tmp_path) + evidence_dir = tmp_path / "out" + + result = lab_run.run("query-evidence.sh", query_arguments(evidence_dir)) + + assert result.returncode == 0, result.stderr + assert SUBSCRIPTION_ID in lab_run.az_calls() diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_repo_devcontainer.py b/monitor/sre-agent-event-lab/scripts/tests/test_repo_devcontainer.py new file mode 100644 index 0000000..27ac609 --- /dev/null +++ b/monitor/sre-agent-event-lab/scripts/tests/test_repo_devcontainer.py @@ -0,0 +1,117 @@ +"""Contract tests for the repository-wide dev container. + +One container serves the whole repository. VS Code's "Reopen in Container" +and Codespaces both read `.devcontainer/devcontainer.json` by default, so a +configuration stored there works from any directory without the operator +picking anything. A configuration filed under a lab's own name does not: +Codespaces offers it as one choice among several, and VS Code ignores it. + +These tests keep the container general as labs are added: it supplies the +shared Azure toolchain and nothing that names a particular lab, and the +tools it supplies stay in step with the install instructions someone +working outside a container follows. +""" +import json +import re +from pathlib import Path + + +REPO_ROOT = Path(__file__).parents[4] +DEVCONTAINER_DIR = REPO_ROOT / ".devcontainer" +DEVCONTAINER = DEVCONTAINER_DIR / "devcontainer.json" +TOOLCHAIN_DOC = DEVCONTAINER_DIR / "README.md" + +# The command names every lab in this repository expects on PATH. The +# container supplies them and the toolchain document explains how to get +# each one without a container. +REQUIRED_TOOLS = ("az", "azd", "gh", "python3", "uv", "jq", "curl") + +# The dev container features that install the tools not already in the base +# image. +REQUIRED_FEATURES = ("azure-cli", "azure-dev/azd", "github-cli", "python") + + +def config(): + assert DEVCONTAINER.is_file(), ( + "VS Code and Codespaces read .devcontainer/devcontainer.json by " + f"default; expected {DEVCONTAINER}" + ) + return json.loads(DEVCONTAINER.read_text()) + + +def test_the_container_is_the_repository_default(): + """Anywhere in the repository, with no configuration picker.""" + assert DEVCONTAINER.is_file(), ( + "VS Code and Codespaces read .devcontainer/devcontainer.json by " + f"default; expected {DEVCONTAINER}" + ) + assert config()["name"] + + +def test_there_is_exactly_one_container_configuration(): + """A second configuration brings the Codespaces picker back, and the + default stops being the only answer to 'which container am I in'.""" + found = sorted( + path.relative_to(REPO_ROOT) + for path in DEVCONTAINER_DIR.rglob("devcontainer.json") + ) + assert found == [Path(".devcontainer/devcontainer.json")], found + + +def test_the_container_names_no_individual_lab(): + """The moment the container knows one lab's path, it stops being the + repository's container: every new lab has to edit it, and one lab's + setup failure breaks container creation for everyone.""" + raw = DEVCONTAINER.read_text() + labs = [ + azure_yaml.relative_to(REPO_ROOT).parent + for azure_yaml in REPO_ROOT.rglob("azure.yaml") + if not {".git", ".worktrees", ".venv", "node_modules"} + & set(azure_yaml.relative_to(REPO_ROOT).parts) + ] + assert labs, "no lab found; the discovery rule is wrong" + for lab in labs: + assert str(lab) not in raw, lab + assert lab.name not in raw, lab.name + + +def test_the_container_supplies_the_shared_toolchain(): + features = " ".join(config().get("features", {})) + for required in REQUIRED_FEATURES: + assert required in features, required + assert "uv" in DEVCONTAINER.read_text(), ( + "labs install Python dependencies through uv only" + ) + + +def test_the_container_carries_no_credential_and_no_subscription(): + raw = DEVCONTAINER.read_text() + assert not re.search( + r"\b[0-9a-fA-F]{8}-(?:[0-9a-fA-F]{4}-){3}[0-9a-fA-F]{12}\b", raw + ), "a GUID here would pin every user to whoever authored the file" + for forbidden in ("PASSWORD", "SECRET", "TOKEN", "_KEY", "CONNECTION_STRING"): + assert forbidden not in raw.upper(), forbidden + + +def test_the_container_never_logs_in_for_the_user(): + """`az login` is interactive and account-specific.""" + raw = DEVCONTAINER.read_text() + assert "az login" not in raw + assert "azd auth login" not in raw + + +def test_one_document_owns_the_install_instructions(): + """Someone working outside a container needs the same tools. Keeping + that list in one place is what stops each lab from carrying its own + copy and drifting.""" + assert TOOLCHAIN_DOC.is_file(), TOOLCHAIN_DOC + text = TOOLCHAIN_DOC.read_text() + for tool in REQUIRED_TOOLS: + assert tool in text, tool + + +def test_the_lab_points_at_that_document_instead_of_repeating_it(): + lab_readme = REPO_ROOT / "monitor" / "sre-agent-event-lab" / "README.md" + assert ".devcontainer/README.md" in lab_readme.read_text(), ( + "the lab must link the shared toolchain document, not restate it" + ) diff --git a/monitor/sre-agent-event-lab/scripts/tests/test_score.py b/monitor/sre-agent-event-lab/scripts/tests/test_score.py index 48366a0..1117315 100644 --- a/monitor/sre-agent-event-lab/scripts/tests/test_score.py +++ b/monitor/sre-agent-event-lab/scripts/tests/test_score.py @@ -412,5 +412,5 @@ def test_cli_without_any_state_explains_what_to_run_first(tmp_path): result = run_cli(tmp_path) assert result.returncode == 1 - assert "lab.sh run" in result.stderr + assert "guides/02-scenario-s1.md" in result.stderr assert "Traceback" not in result.stderr