diff --git a/ptu_lb/README.md b/ptu_lb/README.md new file mode 100644 index 0000000..1303429 --- /dev/null +++ b/ptu_lb/README.md @@ -0,0 +1,3 @@ +# PTU LB Test Docs + +메인 문서는 [ptu-lb-test.md](ptu-lb-test.md) 입니다. diff --git a/ptu_lb/ptu-lb-test.md b/ptu_lb/ptu-lb-test.md new file mode 100644 index 0000000..7d235bd --- /dev/null +++ b/ptu_lb/ptu-lb-test.md @@ -0,0 +1,319 @@ +# Adaptive PTU TPM Quota 조정 가이드 (with Sample Test) + +## 문제 상황 + +동일 워크로드를 처리하는 2개 모델 배포(Korea/UK)가 있으며, 각 리전의 사용률(utilization)에 따라 모델별 TPM quota(= deployment SKU capacity)를 안정적으로 자동 조정하는 adaptive 자동화가 필요하다. + +- 현재 테스트 환경: Global Standard deployment 2개 +- 목표 운영 환경: PTU deployment 기반 자동화 +- 핵심 목표: TPM quota 조정 자동화, 성능 안정화, 플래핑 방지 + +> **핵심 설계 결정 — Reserved quota 공유 상한 하 비례 재분배** +> +> 두 리전(Korea/UK)은 사용자가 **Reservation한 총 TPM quota**(예: 20,000 TPM)를 나눠 쓰며, 두 배포 capacity의 **합은 이 Reserved 총량을 어떤 시점에도 초과할 수 없다.** 초과하는 증설 요청은 거부된다. +> 따라서 각 리전을 독립으로 올리는 방식이 아니라, **고정된 Reserved 총량을 트래픽 비율에 맞춰 두 리전에 재분배**한다. +> - **감축 먼저, 증설 나중**: 총량이 꽉 찬 상태에서 증설을 먼저 하면 실패하므로, 여유를 내주는 리전을 **먼저 감축**해 headroom을 확보한 뒤 증설한다. +> - **loss 최소화**: 증설은 우선 미할당 버퍼로 충당(무손실)하고, 버퍼가 부족할 때만 잉여 리전에서 회수하되 감축·증설을 같은 사이클에서 연속 처리해 총 가용량이 줄어드는 구간을 최소화한다. +> - **리전 최소 보장**: 각 리전은 Reserved 총량의 25% 이상을 항상 유지한다. +> +> ※ 위 수치(총량 20,000 TPM, 리전 최소 보장 25% 등)는 **샘플 값**이며, 런북 상단 파라미터로 조정 가능한 **변수**다. Reserved 총량(`RESERVED_TOTAL_TPM`), capacity 환산(`TPM_PER_CAPACITY_UNIT`), 최소 보장 비율(`MIN_CAPACITY_RATIO`), 목표/상하한 사용률(`AIM_UTIL`/`TARGET_UTIL_HIGH`/`TARGET_UTIL_LOW`), 스텝(`MAX_STEP_UNITS`), 쿨다운(`COOLDOWN_MINUTES`) 등을 운영 환경에 맞게 바꿔 쓰면 된다. + +## 아키텍처 + +Azure Automation은 **스케줄러 + 실행/인증 호스트** 역할이며, 실제 제어 로직은 Python 런북(Azure SDK)이 담당한다. 제어 루프는 **① 메트릭 적재 → ② 윈도우 조회 → ③ 목표 재분배·capacity 갱신(감축→증설, 쿨다운 가드) → ④ 이력 기록** 순으로 동작한다. + +- 데이터 플레인(파랑): Client → App/Router → Foundry(KR/UK) → Azure Monitor +- 제어 플레인(초록): 스케줄이 Hybrid Runbook Worker를 트리거하고, Worker가 VNet 안에서 Python 런북을 실행(관리 ID 인증) +- 상태 저장소(보라): Table Storage(TrafficWindow/ControlDecision/ScaleAction)에 **Private Endpoint**로 접근 (Storage `publicNetworkAccess=Disabled`) +- capacity 조정(빨강): 런북이 Reserved 총량 제약 하에 **감축→증설 순서**로 Foundry 배포 SKU capacity를 ARM(공용 경로)으로 갱신 + + +![PTU TPM Quota Adaptive Controller 아키텍처](ptu_architecture.png) + +## Azure 리소스 구성 + +### 필수 + +1. Azure AI Foundry model deployment (Korea) +2. Azure AI Foundry model deployment (UK) +3. Azure Storage Account (Table Storage) +4. Azure Automation Account +5. Managed Identity (Runbook 실행 주체) +6. Azure Monitor 또는 Log Analytics + +### 권장 + +1. Application Insights +2. Azure Key Vault +3. Action Group + +## 데이터 모델 (Table Storage) + +### TrafficWindow + +- 목적: n분 단위 메트릭 집계 저장 +- PartitionKey: region (korea | uk) +- RowKey: windowStartUtc (ISO8601) +- 컬럼: requestCount, inputTokens, outputTokens, p95LatencyMs, errorRate, assignedTpm + +### ControlDecision + +- 목적: 리전별 제어 의사결정 저장 +- PartitionKey: date (yyyyMMdd) +- RowKey: `{timestamp}-{region}-{guid}` +- 컬럼: region, utilization, avgConsumedTpm, capacityBefore, capacityTarget, decision(up|down|hold), reason, cooldownBlocked, executed + +### ScaleAction + +- 목적: 리전별 실제 capacity 조정 실행 결과 저장 +- PartitionKey: date (yyyyMMdd) +- RowKey: `{timestamp}-{region}-{guid}` +- 컬럼: region, requestPayload, capacityBefore, capacityTarget, changed, responseStatus, success, errorMessage, durationMs + +## 안정적 자동화 운영 원칙 + +짧은 주기로 빈번하게 조정하면 오히려 불안정해질 수 있으므로, 제어 루프는 안정성을 우선한다. + +1. 관측 주기와 조정 주기를 분리 +2. 단일 포인트가 아닌 이동평균으로 판단 +3. 쿨다운 동안 재조정 금지 +4. 1회 최대 변경폭 제한 +5. 리전별 최소/최대 capacity 경계값 강제(최소 = Reserved 총량의 25%) +6. 두 리전 capacity 합은 Reserved 총량을 항상 준수(초과 증설 불가) +7. 재분배는 감축을 먼저 실행해 headroom 확보 후 증설 + +## 운영 파라미터 (변수화) + +제어 knob은 리전별 **capacity unit**(배포 SKU capacity)이며, `TPM_PER_CAPACITY_UNIT`로 TPM과 매핑한다. 두 리전 capacity의 합은 `RESERVED_TOTAL_TPM`(Reserved quota)을 넘을 수 없다. + +| 파라미터 | 의미 | 권장 시작값 | +|---|---|---| +| METRIC_WINDOW_MINUTES | 메트릭 집계 윈도우(분) | 5 | +| SMOOTHING_WINDOWS | 이동평균 윈도우 개수 | 3 | +| TPM_PER_CAPACITY_UNIT | capacity 1 unit ≈ TPM | 1000 | +| RESERVED_TOTAL_TPM | Reservation한 총 TPM 상한(합산 불변) | 20000 | +| TOTAL_CAPACITY_UNITS | Reserved 총량의 capacity 환산(= RESERVED_TOTAL_TPM / TPM_PER_CAPACITY_UNIT) | 20 | +| TARGET_UTIL_HIGH | 이 사용률 초과 → 증설 압력 | 0.75 | +| TARGET_UTIL_LOW | 이 사용률 미만 → 감축(잉여 회수) | 0.30 | +| AIM_UTIL | 조정 후 목표 사용률 | 0.60 | +| MAX_STEP_UNITS | 1회 최대 capacity 변경폭 | 5 | +| MIN_CAPACITY_RATIO | 리전별 최소 보장 비율(Reserved 총량 대비) | 0.25 | +| MIN_CAPACITY | 리전별 최소 보장 capacity(= ceil(TOTAL × 0.25)) | 5 | +| MAX_CAPACITY | 리전별 최대 capacity. 2리전 기본값 = TOTAL − MIN_CAPACITY, N리전은 `TOTAL − (N−1)·MIN_CAPACITY` | 15 | +| COOLDOWN_MINUTES | 조정 후 재조정 금지 시간(분) | 20 | + +> 불변식: 항상 `Σ cap_r ≤ TOTAL_CAPACITY_UNITS`. 리전 최대는 `TOTAL − (N−1)·MIN_CAPACITY`로 일반화되며(2리전이면 `TOTAL − MIN_CAPACITY`), 최소는 총량의 25%를 보장한다. +> +> **리전/모델 확대 시**: 코드 변경은 `REGIONS`(PS `$Regions`) 매핑에 배포 항목(리전 또는 모델별 `account`+`deployment` 쌍)을 추가하는 것으로 끝난다. 모델을 늘리는 경우도 동일하게 각 모델 배포를 하나의 항목으로 등록하면 같은 Reserved 총량을 공유한다. 항목별 최대는 `region_max(n)`(PS `Get-RegionMax`)이 항목 수 n으로 자동 일반화하고, 총량 제약·감축→증설 순서·MIN 보장은 그대로 유지된다. +> +> **PayGo fallback 병행 시**: 상위 트래픽이 PTU 초과분을 PayGo(표준) 배포로 흘리는 구성에서는 PTU util을 100%에 근접(TARGET_UTIL_HIGH·AIM_UTIL ≈ 1.0)시켜도 무방하다. 초과분은 fallback LLM이 흡수하므로, 비용 최적화 관점에서 적정 PTU/PayGo 비중을 도출하는 것이 목표다. + +## 샘플 테스트 Sizing (권장) + +샘플 테스트 목적이므로 비용과 변동성을 낮추는 보수적 크기로 시작한다. + +### 권장 시작값 + +1. 모델 배포 수: 2개 (Korea, UK) +2. Reserved 총량: 20,000 TPM (= 20 capacity unit) +3. 초기 capacity: 10 / 10 (합 20 = Reserved 총량 만재, 버퍼 0) +4. 리전 최소 capacity: 5 (Reserved 총량의 25%) +5. 리전 최대 capacity: 15 (= 총량 − 상대 리전 최소) +6. 1회 조정폭: 최대 5 unit +7. 사용률 밴드: LOW 30% ~ HIGH 75% (AIM 60%) +8. 쿨다운: 20분 + +### 부하 프로파일 (샘플) + +Reserved 총량이 만재(10/10)인 상태에서 트래픽을 기울여 리전 간 재분배를 유도한다. + +1. Baseline 20분: 양 리전 util 40~60% (유지 구간, 재분배 없음) +2. 기울기 30분: Korea util 85%+(PayGo fallback 병행 시 ≈100%까지 허용) / UK util 20% → UK 먼저 감축 후 Korea 증설(감축→증설) +3. 역전 30분: UK util 85%+ / Korea util 20% → 반대 방향 재분배(쿨다운 경과 후) +4. 진동 30분: 10분 단위 고/저 교차 → 쿨다운/스텝 제한으로 왕복 억제 + +### 샘플 성공 기준 + +1. 두 리전 capacity 합이 Reserved 총량을 항상 준수(초과 0회) +2. 재분배 시 감축이 증설보다 먼저 실행됨(증설 선행 실패 0회) +3. 고사용률 리전 증설분이 잉여 리전 감축분으로 충당됨(loss 최소) +4. 쿨다운 중 불필요한 재조정이 발생하지 않음 +5. ControlDecision, ScaleAction 이력이 누락 없이 기록됨 +6. 메트릭 열화(조회 실패) 시 RI 균등분배 fallback으로 전환되어 미할당 버퍼 없이 총량을 소진함 + +### 파라미터 적용 규칙 + +1. 스케줄 주기는 METRIC_WINDOW_MINUTES의 배수 +2. COOLDOWN_MINUTES는 스케줄 주기보다 크게 설정 +3. 초기에는 MAX_STEP_UNITS를 낮게 시작해 플래핑 여부 검증 + +## Adaptive 제어 로직 (Reserved 총량 공유) + +두 리전을 **하나의 Reserved 총량 예산**으로 함께 평가하고, 목표를 먼저 산출한 뒤 총량 제약을 지키는 순서로 실행한다. + +1. 각 배포의 현재 capacity를 SDK로 조회(폐루프 기준값) +2. 최근 SMOOTHING_WINDOWS 구간의 분당 소비 TPM 이동평균 산출 + +$$ +consumedTpm_r = \frac{1}{N}\sum_{i=1}^{N} \frac{inputTokens_i + outputTokens_i}{METRIC\_WINDOW\_MINUTES} +$$ + +3. 리전별 사용률과 AIM_UTIL 기준 필요 capacity(need) 산출 + +$$ +utilization_r = \frac{consumedTpm_r}{capacity_r \times TPM\_PER\_CAPACITY\_UNIT} +\qquad +need_r = \left\lceil \frac{consumedTpm_r}{AIM\_UTIL \times TPM\_PER\_CAPACITY\_UNIT} \right\rceil +$$ + +4. 목표 capacity 산출 (Reserved 총량 제약) +- **여유 있음** ($\sum_r need_r \le TOTAL$): 각 리전에 need만 배정, 나머지는 미할당 버퍼로 둔다 +- **경합** ($\sum_r need_r > TOTAL$): 각 리전에 MIN을 먼저 보장하고, 남은 용량을 소비 TPM 비율로 배분 +- **fallback (메트릭 열화)**: 메트릭 조회가 실패해 판단 근거가 없을 때는 소비 기반 배분 대신 **RI 균등분배**(총량을 리전 수로 균등, $target_r \approx TOTAL / n$)로 전환해 미할당 버퍼 없이 RI에 맞는 PTU를 최대 소진한다 + +$$ +target_r = MIN\_CAPACITY + \left\lfloor (TOTAL - n \cdot MIN\_CAPACITY)\times\frac{consumedTpm_r}{\sum_j consumedTpm_j} \right\rfloor +$$ + +5. dead-band·스텝 완충 +- 경합이 아니고 사용률이 [LOW, HIGH] 안이면 유지(hold)로 churn 방지 +- 목표로의 이동을 ±MAX_STEP_UNITS로 제한 +- 리전 최소/최대(MIN_CAPACITY ~ `region_max(n)` = TOTAL − (n−1)·MIN) 경계 적용. 2리전이면 MAX_CAPACITY = TOTAL − MIN 과 동일 + +6. 쿨다운 가드 +- 해당 리전의 마지막 '실제 변경(changed=true)' ScaleAction이 COOLDOWN_MINUTES 이내면 조정 차단 + +7. 총량 제약 실행 (감축 먼저 → 증설 나중) +- 현재 headroom = TOTAL − Σ(현재 capacity) +- planned < 현재인 리전(잉여 회수)을 **먼저 감축** → headroom 증가 +- planned > 현재인 리전을 **사용률 높은 순**으로, 남은 headroom 한도 안에서만 증설 +- headroom이 없으면 이번 사이클 증설은 보류(다음 사이클 수렴), 어떤 시점에도 합 ≤ TOTAL 보장 + +8. 실행 및 기록 +- capacity 갱신 API 호출(azure-mgmt-cognitiveservices) +- ControlDecision, ScaleAction 기록(경합 여부·headroom 제한 사유 포함) +- 다음 주기에서 개선 여부 재평가 + +## Runbook Workflow + +한 사이클은 두 리전을 함께 처리한다(공유 총량 재분배): + +1. (수집) 각 배포의 현재 capacity 조회 + Table 최근 윈도우로 이동평균 사용률 계산 +2. (산출) 리전별 need 계산 → Reserved 총량 제약으로 목표(target) 배분(여유/경합 분기) +3. (완충) dead-band·스텝·쿨다운 가드로 이번 사이클 planned 결정 +4. (실행 1) planned < 현재인 리전을 먼저 감축 → headroom 확보 +5. (실행 2) planned > 현재인 리전을 사용률 높은 순으로 headroom 한도 내 증설 +6. (기록) 리전별 ControlDecision / ScaleAction 기록 +7. (안전) 개별 API 실패 시 기록 후 계속, 합이 총량을 넘지 않도록 보장 +8. (fallback) 메트릭 조회 열화 시 소비 기반 재분배 대신 **RI 균등분배**(총량을 리전 수로 균등)로 전환 → 판단 근거가 없을 때도 미할당 버퍼 없이 RI에 맞는 PTU를 최대 소진(감축→증설 순서·총량 불변식 유지) + +## 테스트 시나리오 + +### Test-1 Baseline + +- 양 리전 util 40~60% 유지, 총량 만재(10/10) +- 불필요 조정 0회 확인 + +### Test-2 재분배 (감축→증설) + +- Korea util 85%+ / UK util 20% 유도 (PayGo fallback 병행 시 Korea ≈100%도 정상으로 판정) +- UK 감축이 Korea 증설보다 먼저 실행되고, 합이 총량을 유지하는지 확인 + +### Test-3 증설 선행 방지 + +- 총량 만재 상태에서 Korea 증설 요청 발생 +- 감축 없이 증설을 시도하지 않음(증설 선행 실패 0회) 확인 + +### Test-4 경합 비례 배분 + +- 양 리전 동시 고부하(합 need > 총량) +- 소비 TPM 비율대로 target이 배분되고 합 ≤ 총량 확인 + +### Test-5 플래핑 억제 + +- 10분 단위로 사용률 고/저 교차 +- 쿨다운/스텝 제한으로 왕복 조정 억제 확인 + +### Test-6 장애 복원력 + +- capacity 변경 API 실패 강제 +- 재시도/기록/안전 종료 및 총량 불변식 유지 확인 +- 메트릭(Table) 조회 실패 강제 → RI 균등분배 fallback 전환 및 총량 불변식 유지 확인 + +## 테스트 결과 + +실제 Azure 리소스(Foundry 배포 2개 + Storage Table)에 대해 라이브 검증한 결과다. Reserved 총량 = 20 unit(20000 TPM), baseline 10/10. + +| # | 시나리오 | 입력(소비 TPM/min) | 판단 | capacity 변화 | 합 | 결과 | +|---|----------|--------------------|------|----------------|-----|------| +| A | 경합 비례 배분(dry-run) | KR 9000 / UK 7000 | contention=True | 10/10 → 11/9 | 20 | 소비 비율대로 재분배, 합 ≤ 총량 | +| B | 재분배 실제 조정 | KR 9000 / UK 저부하 | contention=True | 10/10 → 13/7 | 20 | **UK 감축이 먼저 실행되어 headroom 확보 후 KR 증설**, SDK로 실 capacity 13/7 확인 | +| C | 수렴·안정 | B 직후 동일 부하 | hold/hold | 13/7 유지 | 20 | 조정 후 util이 dead-band(69%/52%)에 안착 → 불필요 조정 0회 | +| D | 쿨다운 가드 | 큰 부하 변화 재유도 | hold/hold | 13/7 유지 | 20 | target(14/6) 산출됐으나 쿨다운(20분)으로 차단 → 플래핑 억제 | + +- **총량 불변식**: 모든 사이클에서 KR+UK 합 = 20 ≤ Reserved 총량(초과 0회). +- **감축 선행**: Test-B에서 증설 선행 실패 없이 감축→증설 순서로 폐루프 조정 완료. +- 오프라인 시뮬레이션(Azure SDK 스텁)에서도 baseline·경합·유휴 등 케이스 전부 합 ≤ 20, 감축 선행 순서 준수 확인. + +## 완료 기준 + +1. 두 리전 capacity 합이 Reserved 총량을 항상 준수(초과 0회) +2. 재분배가 감축→증설 순서로 수행되고, 증설 선행 실패가 0회 +3. 조정 후 고부하 리전의 p95 latency 또는 error rate 개선 관찰 +4. 플래핑 없이 안정적 제어(쿨다운/스텝 준수) +5. 모든 판단/실행 이력 저장 + +## 파일 구성 + +- 메인 문서: [ptu-lb-test.md](ptu-lb-test.md) +- 아키텍처 다이어그램(원본): [ptu_architecture.excalidraw](ptu_architecture.excalidraw) +- 아키텍처 다이어그램(이미지): [ptu_architecture.png](ptu_architecture.png) +- 다이어그램 렌더러(유틸): [render_excalidraw.py](render_excalidraw.py) +- **Runbook(기본, Python): [runbooks/adaptive_tpm_quota_controller.py](runbooks/adaptive_tpm_quota_controller.py)** +- Runbook 의존성: [runbooks/requirements.txt](runbooks/requirements.txt) +- Runbook(선택, PowerShell 동등 이식본): [runbooks/optional_adaptive_tpm_quota_controller.ps1](runbooks/optional_adaptive_tpm_quota_controller.ps1) + +> 기본 구현은 Python 런북이다. PowerShell 버전은 동일한 Reserved 총량 공유 재분배 로직(감축→증설, 실제 actuation·쿨다운 포함)을 Az 모듈(Az.Accounts / Az.CognitiveServices / AzTable)로 이식한 **선택적 대체 구현**으로, PowerShell 선호 환경을 위한 참조본이다. Storage가 Private Endpoint 구성이면 이 런북도 프라이빗 경로가 닿는 워커(예: Windows Hybrid Worker)에서 실행해야 한다. + +## 배포 구성 (프로덕션 아키텍처) + +무인(unattended) 자동 운영을 위한 권장 구성이다. 핵심 제약은 Storage `publicNetworkAccess=Disabled` 환경에서 Automation 클라우드 샌드박스(공용 인프라)가 Table에 접근할 수 없다는 점이며, **VNet 내 Hybrid Runbook Worker + Table Private Endpoint**로 해결한다. Foundry capacity 갱신은 ARM(`management.azure.com`) 관리 평면이라 항상 공용 경로로 동작하므로 Foundry용 Private Endpoint는 필요 없다. + +### 제약과 해결 + +| 제약 | 원인 | 해결 | +|---|---|---| +| 샌드박스가 Storage에 접근 불가 | Storage public 접근 차단 + 샌드박스는 공용 인프라에서 실행 | VNet 내 Hybrid Runbook Worker + Table Private Endpoint | +| Foundry 접근용 PE 필요 여부 | capacity 갱신은 ARM 관리 평면(항상 공용) | Foundry PE 불필요, Table PE만 구성 | +| 10분 주기 스케줄 | Automation 네이티브 스케줄 최소 간격이 1시간 | 10분 오프셋 시간 스케줄 6개(`sched-min-00/10/20/30/40/50`) | + +### 네트워크 · 워커 + +1. VNet(예: `10.20.0.0/16`) — 워커 서브넷과 PE 서브넷(PE 네트워크 정책 비활성)을 분리 +2. Private DNS Zone `privatelink.table.core.windows.net` 생성 후 VNet에 링크 +3. Storage Table Private Endpoint(group-id=`table`) → 사설 IP로 해석 +4. Hybrid Worker VM(Ubuntu, 공용 IP 없음, System-assigned MI) + `HybridWorkerForLinux` 확장 → Hybrid Worker Group에 등록 +5. VM에 `python3` + `azure-identity` / `azure-data-tables` / `azure-mgmt-cognitiveservices` 설치 (인증은 IMDS로 VM MI 사용) + +### 권한 (RBAC) + +- Hybrid Worker VM MI: **Storage Table Data Contributor**(Storage) + **Cognitive Services Contributor**(양 Foundry) +- 인증 경로: `DefaultAzureCredential` → IMDS → VM System-assigned MI + +### 런북 · 스케줄 + +1. `runbooks/adaptive_tpm_quota_controller.py`를 Automation Python 3 런북으로 게시하고 Hybrid Worker Group에서 실행 +2. 6개 시간 스케줄을 `runOn=`으로 연결해 10분 주기를 구현 + - `az automation`에는 `job-schedule` 서브커맨드가 없으므로 REST로 연결: `az rest --method put --url ".../jobSchedules/{guid}?api-version=2023-11-01"` +3. 최초 배포는 `--dry-run`으로 판단 로직을 검증한 뒤 실제 실행으로 전환 + +### 배포 검증 체크리스트 + +- 런북 잡이 워커에서 **Completed**로 종료 +- 워커가 사설 경로로 Table을 읽고 Foundry capacity를 실제 변경 → 테스트 후 baseline(10/10) 복원 +- 고사용률 리전 단독 증설 · 저사용률 리전 단독 감축 · 쿨다운 중 재조정 차단 확인 +- ControlDecision · ScaleAction 이력이 누락 없이 기록 + +### 비용 참고 + +- Hybrid Worker VM은 상시 과금된다. 미사용 시 중지: `az vm deallocate -g -n ` diff --git a/ptu_lb/ptu_architecture.excalidraw b/ptu_lb/ptu_architecture.excalidraw new file mode 100644 index 0000000..7f4e033 --- /dev/null +++ b/ptu_lb/ptu_architecture.excalidraw @@ -0,0 +1,1801 @@ +{ + "type": "excalidraw", + "version": 2, + "source": "https://excalidraw.com", + "elements": [ + { + "id": "title", + "type": "text", + "x": 60, + "y": 24, + "width": 1140, + "height": 34, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 101, + "version": 1, + "versionNonce": 101, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "PTU TPM Quota Adaptive Controller — 리전 독립 오토스케일", + "fontSize": 26, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "PTU TPM Quota Adaptive Controller — 리전 독립 오토스케일", + "lineHeight": 1.2 + }, + { + "id": "subtitle", + "type": "text", + "x": 60, + "y": 64, + "width": 1000, + "height": 20, + "angle": 0, + "strokeColor": "#868e96", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 102, + "version": 1, + "versionNonce": 102, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "Hybrid Runbook Worker + Storage Private Endpoint (unattended · 10분 주기)", + "fontSize": 15, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "Hybrid Runbook Worker + Storage Private Endpoint (unattended · 10분 주기)", + "lineHeight": 1.2 + }, + { + "id": "cFoundry", + "type": "rectangle", + "x": 520, + "y": 110, + "width": 250, + "height": 180, + "angle": 0, + "strokeColor": "#1971c2", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1.5, + "strokeStyle": "dashed", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 110, + "version": 1, + "versionNonce": 110, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "cFoundryT", + "type": "text", + "x": 532, + "y": 116, + "width": 230, + "height": 18, + "angle": 0, + "strokeColor": "#1971c2", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 111, + "version": 1, + "versionNonce": 111, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "모델 배포 (리전 독립 quota)", + "fontSize": 13, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "모델 배포 (리전 독립 quota)", + "lineHeight": 1.2 + }, + { + "id": "cAuto", + "type": "rectangle", + "x": 60, + "y": 380, + "width": 700, + "height": 310, + "angle": 0, + "strokeColor": "#f08c00", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1.5, + "strokeStyle": "dashed", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 112, + "version": 1, + "versionNonce": 112, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "cAutoT", + "type": "text", + "x": 76, + "y": 388, + "width": 600, + "height": 20, + "angle": 0, + "strokeColor": "#f08c00", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 113, + "version": 1, + "versionNonce": 113, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "제어 플레인 — Azure Automation", + "fontSize": 15, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "제어 플레인 — Azure Automation", + "lineHeight": 1.2 + }, + { + "id": "cStore", + "type": "rectangle", + "x": 820, + "y": 400, + "width": 400, + "height": 290, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1.5, + "strokeStyle": "dashed", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 114, + "version": 1, + "versionNonce": 114, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "cStoreT", + "type": "text", + "x": 836, + "y": 408, + "width": 380, + "height": 34, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 115, + "version": 1, + "versionNonce": 115, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "상태 저장소 — Table Storage\n(Private Endpoint · public 접근 차단)", + "fontSize": 13, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "상태 저장소 — Table Storage\n(Private Endpoint · public 접근 차단)", + "lineHeight": 1.2 + }, + { + "id": "bClient", + "type": "rectangle", + "x": 60, + "y": 140, + "width": 150, + "height": 70, + "angle": 0, + "strokeColor": "#1971c2", + "backgroundColor": "#a5d8ff", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 120, + "version": 1, + "versionNonce": 120, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bClientT", + "type": "text", + "x": 60, + "y": 153, + "width": 150, + "height": 44, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 121, + "version": 1, + "versionNonce": 121, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "Client\n(부하·실사용)", + "fontSize": 16, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "Client\n(부하·실사용)", + "lineHeight": 1.2 + }, + { + "id": "bRouter", + "type": "rectangle", + "x": 280, + "y": 140, + "width": 170, + "height": 70, + "angle": 0, + "strokeColor": "#1971c2", + "backgroundColor": "#a5d8ff", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 122, + "version": 1, + "versionNonce": 122, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bRouterT", + "type": "text", + "x": 280, + "y": 153, + "width": 170, + "height": 44, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 123, + "version": 1, + "versionNonce": 123, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "App / Router\n(요청 분배)", + "fontSize": 16, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "App / Router\n(요청 분배)", + "lineHeight": 1.2 + }, + { + "id": "bKR", + "type": "rectangle", + "x": 540, + "y": 150, + "width": 210, + "height": 54, + "angle": 0, + "strokeColor": "#f08c00", + "backgroundColor": "#ffe066", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 124, + "version": 1, + "versionNonce": 124, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bKRT", + "type": "text", + "x": 540, + "y": 155, + "width": 210, + "height": 44, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 125, + "version": 1, + "versionNonce": 125, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "Foundry (Korea)\ntpmtest-kr", + "fontSize": 15, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "Foundry (Korea)\ntpmtest-kr", + "lineHeight": 1.2 + }, + { + "id": "bUK", + "type": "rectangle", + "x": 540, + "y": 222, + "width": 210, + "height": 54, + "angle": 0, + "strokeColor": "#f08c00", + "backgroundColor": "#ffe066", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 126, + "version": 1, + "versionNonce": 126, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bUKT", + "type": "text", + "x": 540, + "y": 227, + "width": 210, + "height": 44, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 127, + "version": 1, + "versionNonce": 127, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "Foundry (UK)\ntpmtest-uk", + "fontSize": 15, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "Foundry (UK)\ntpmtest-uk", + "lineHeight": 1.2 + }, + { + "id": "bMon", + "type": "rectangle", + "x": 880, + "y": 170, + "width": 180, + "height": 70, + "angle": 0, + "strokeColor": "#495057", + "backgroundColor": "#e9ecef", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 128, + "version": 1, + "versionNonce": 128, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bMonT", + "type": "text", + "x": 880, + "y": 183, + "width": 180, + "height": 44, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 129, + "version": 1, + "versionNonce": 129, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "Azure Monitor\n(토큰·지연 메트릭)", + "fontSize": 15, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "Azure Monitor\n(토큰·지연 메트릭)", + "lineHeight": 1.2 + }, + { + "id": "bSch", + "type": "rectangle", + "x": 90, + "y": 440, + "width": 170, + "height": 95, + "angle": 0, + "strokeColor": "#f08c00", + "backgroundColor": "#ffec99", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 130, + "version": 1, + "versionNonce": 130, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bSchT", + "type": "text", + "x": 90, + "y": 454, + "width": 170, + "height": 66, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 131, + "version": 1, + "versionNonce": 131, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "스케줄 10분 주기\n(6× 시간 스케줄,\n10분 오프셋)", + "fontSize": 14, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "스케줄 10분 주기\n(6× 시간 스케줄,\n10분 오프셋)", + "lineHeight": 1.2 + }, + { + "id": "bMI", + "type": "rectangle", + "x": 90, + "y": 565, + "width": 170, + "height": 90, + "angle": 0, + "strokeColor": "#f08c00", + "backgroundColor": "#ffec99", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 132, + "version": 1, + "versionNonce": 132, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bMIT", + "type": "text", + "x": 90, + "y": 588, + "width": 170, + "height": 44, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 133, + "version": 1, + "versionNonce": 133, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "Managed Identity\n(System-assigned)", + "fontSize": 14, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "Managed Identity\n(System-assigned)", + "lineHeight": 1.2 + }, + { + "id": "bWk", + "type": "rectangle", + "x": 320, + "y": 455, + "width": 200, + "height": 110, + "angle": 0, + "strokeColor": "#2f9e44", + "backgroundColor": "#b2f2bb", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 134, + "version": 1, + "versionNonce": 134, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bWkT", + "type": "text", + "x": 320, + "y": 477, + "width": 200, + "height": 66, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 135, + "version": 1, + "versionNonce": 135, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "Hybrid Runbook Worker\nhrw-vm (VNet)\nUbuntu · B2s", + "fontSize": 15, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "Hybrid Runbook Worker\nhrw-vm (VNet)\nUbuntu · B2s", + "lineHeight": 1.2 + }, + { + "id": "bRb", + "type": "rectangle", + "x": 560, + "y": 455, + "width": 170, + "height": 110, + "angle": 0, + "strokeColor": "#2f9e44", + "backgroundColor": "#b2f2bb", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 136, + "version": 1, + "versionNonce": 136, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bRbT", + "type": "text", + "x": 560, + "y": 477, + "width": 170, + "height": 66, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 137, + "version": 1, + "versionNonce": 137, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "Python 런북\nAdaptiveTpmQuota\nControllerPy", + "fontSize": 14, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "Python 런북\nAdaptiveTpmQuota\nControllerPy", + "lineHeight": 1.2 + }, + { + "id": "bTW", + "type": "rectangle", + "x": 850, + "y": 460, + "width": 340, + "height": 54, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "#d0bfff", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 138, + "version": 1, + "versionNonce": 138, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bTWT", + "type": "text", + "x": 850, + "y": 477, + "width": 340, + "height": 22, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 139, + "version": 1, + "versionNonce": 139, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "TrafficWindow · 집계 메트릭", + "fontSize": 15, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "TrafficWindow · 집계 메트릭", + "lineHeight": 1.2 + }, + { + "id": "bCD", + "type": "rectangle", + "x": 850, + "y": 535, + "width": 340, + "height": 54, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "#d0bfff", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 140, + "version": 1, + "versionNonce": 140, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bCDT", + "type": "text", + "x": 850, + "y": 552, + "width": 340, + "height": 22, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 141, + "version": 1, + "versionNonce": 141, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "ControlDecision · 의사결정", + "fontSize": 15, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "ControlDecision · 의사결정", + "lineHeight": 1.2 + }, + { + "id": "bSA", + "type": "rectangle", + "x": 850, + "y": 610, + "width": 340, + "height": 54, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "#d0bfff", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 3 + }, + "seed": 142, + "version": 1, + "versionNonce": 142, + "isDeleted": false, + "boundElements": [], + "updated": 1762128000000, + "link": null, + "locked": false + }, + { + "id": "bSAT", + "type": "text", + "x": 850, + "y": 627, + "width": 340, + "height": 22, + "angle": 0, + "strokeColor": "#1e1e1e", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 143, + "version": 1, + "versionNonce": 143, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "ScaleAction · 실행 이력 · 쿨다운 기준", + "fontSize": 14, + "fontFamily": 2, + "textAlign": "center", + "verticalAlign": "top", + "containerId": null, + "originalText": "ScaleAction · 실행 이력 · 쿨다운 기준", + "lineHeight": 1.2 + }, + { + "id": "a1", + "type": "arrow", + "x": 210, + "y": 175, + "width": 70, + "height": 0, + "angle": 0, + "strokeColor": "#1971c2", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 150, + "version": 1, + "versionNonce": 150, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 70, + 0 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "a2", + "type": "arrow", + "x": 450, + "y": 170, + "width": 90, + "height": 7, + "angle": 0, + "strokeColor": "#1971c2", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 2 + }, + "seed": 151, + "version": 1, + "versionNonce": 151, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 47, + 0 + ], + [ + 47, + 7 + ], + [ + 90, + 7 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "a3", + "type": "arrow", + "x": 450, + "y": 185, + "width": 90, + "height": 64, + "angle": 0, + "strokeColor": "#1971c2", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 2 + }, + "seed": 152, + "version": 1, + "versionNonce": 152, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 47, + 0 + ], + [ + 47, + 64 + ], + [ + 90, + 64 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "a4", + "type": "arrow", + "x": 750, + "y": 177, + "width": 130, + "height": 13, + "angle": 0, + "strokeColor": "#1971c2", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 2 + }, + "seed": 153, + "version": 1, + "versionNonce": 153, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 65, + 0 + ], + [ + 65, + 13 + ], + [ + 130, + 13 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "a5", + "type": "arrow", + "x": 750, + "y": 249, + "width": 130, + "height": 29, + "angle": 0, + "strokeColor": "#1971c2", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 2 + }, + "seed": 154, + "version": 1, + "versionNonce": 154, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 95, + 0 + ], + [ + 95, + -29 + ], + [ + 130, + -29 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "a6", + "type": "arrow", + "x": 970, + "y": 240, + "width": 0, + "height": 220, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 155, + "version": 1, + "versionNonce": 155, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 0, + 220 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "a7", + "type": "arrow", + "x": 730, + "y": 487, + "width": 120, + "height": 0, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 156, + "version": 1, + "versionNonce": 156, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 120, + 0 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "a8", + "type": "arrow", + "x": 645, + "y": 565, + "width": 205, + "height": 72, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 2 + }, + "seed": 157, + "version": 1, + "versionNonce": 157, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 0, + 72 + ], + [ + 205, + 72 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "a9", + "type": "arrow", + "x": 645, + "y": 455, + "width": 0, + "height": 179, + "angle": 0, + "strokeColor": "#e03131", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 3, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 158, + "version": 1, + "versionNonce": 158, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 0, + -179 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "a10", + "type": "arrow", + "x": 260, + "y": 487, + "width": 60, + "height": 0, + "angle": 0, + "strokeColor": "#2f9e44", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 159, + "version": 1, + "versionNonce": 159, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 60, + 0 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "a11", + "type": "arrow", + "x": 260, + "y": 600, + "width": 160, + "height": 35, + "angle": 0, + "strokeColor": "#2f9e44", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 2, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": { + "type": 2 + }, + "seed": 160, + "version": 1, + "versionNonce": 160, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "points": [ + [ + 0, + 0 + ], + [ + 160, + 0 + ], + [ + 160, + -35 + ] + ], + "lastCommittedPoint": null, + "startBinding": null, + "endBinding": null, + "startArrowhead": null, + "endArrowhead": "arrow" + }, + { + "id": "L1", + "type": "text", + "x": 985, + "y": 330, + "width": 150, + "height": 20, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 161, + "version": 1, + "versionNonce": 161, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "① 메트릭 적재", + "fontSize": 14, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "① 메트릭 적재", + "lineHeight": 1.2 + }, + { + "id": "L2", + "type": "text", + "x": 735, + "y": 462, + "width": 120, + "height": 20, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 162, + "version": 1, + "versionNonce": 162, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "② 윈도우 조회", + "fontSize": 14, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "② 윈도우 조회", + "lineHeight": 1.2 + }, + { + "id": "L3", + "type": "text", + "x": 658, + "y": 300, + "width": 160, + "height": 36, + "angle": 0, + "strokeColor": "#e03131", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 163, + "version": 1, + "versionNonce": 163, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "③ capacity 갱신\n(ARM · 공용 경로)", + "fontSize": 14, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "③ capacity 갱신\n(ARM · 공용 경로)", + "lineHeight": 1.2 + }, + { + "id": "L4", + "type": "text", + "x": 470, + "y": 598, + "width": 160, + "height": 36, + "angle": 0, + "strokeColor": "#6741d9", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 164, + "version": 1, + "versionNonce": 164, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "④ 이력 기록\n(Decision·Action)", + "fontSize": 13, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "④ 이력 기록\n(Decision·Action)", + "lineHeight": 1.2 + }, + { + "id": "L5", + "type": "text", + "x": 262, + "y": 463, + "width": 60, + "height": 20, + "angle": 0, + "strokeColor": "#2f9e44", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 165, + "version": 1, + "versionNonce": 165, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "트리거", + "fontSize": 13, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "트리거", + "lineHeight": 1.2 + }, + { + "id": "L6", + "type": "text", + "x": 300, + "y": 576, + "width": 90, + "height": 20, + "angle": 0, + "strokeColor": "#2f9e44", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 166, + "version": 1, + "versionNonce": 166, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "인증(MI)", + "fontSize": 13, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "인증(MI)", + "lineHeight": 1.2 + }, + { + "id": "legend", + "type": "text", + "x": 60, + "y": 712, + "width": 1140, + "height": 20, + "angle": 0, + "strokeColor": "#495057", + "backgroundColor": "transparent", + "fillStyle": "solid", + "strokeWidth": 1, + "strokeStyle": "solid", + "roughness": 0, + "opacity": 100, + "groupIds": [], + "frameId": null, + "roundness": null, + "seed": 167, + "version": 1, + "versionNonce": 167, + "isDeleted": false, + "boundElements": null, + "updated": 1762128000000, + "link": null, + "locked": false, + "text": "색상: 파랑=데이터 흐름 · 보라=제어 루프(①②④) · 빨강=capacity 조정(③) · 초록=Automation 실행·인증", + "fontSize": 13, + "fontFamily": 2, + "textAlign": "left", + "verticalAlign": "top", + "containerId": null, + "originalText": "색상: 파랑=데이터 흐름 · 보라=제어 루프(①②④) · 빨강=capacity 조정(③) · 초록=Automation 실행·인증", + "lineHeight": 1.2 + } + ], + "appState": { + "gridSize": 20, + "viewBackgroundColor": "#ffffff" + }, + "files": {} +} \ No newline at end of file diff --git a/ptu_lb/ptu_architecture.png b/ptu_lb/ptu_architecture.png new file mode 100644 index 0000000..c79ced6 Binary files /dev/null and b/ptu_lb/ptu_architecture.png differ diff --git a/ptu_lb/render_excalidraw.py b/ptu_lb/render_excalidraw.py new file mode 100644 index 0000000..271ad91 --- /dev/null +++ b/ptu_lb/render_excalidraw.py @@ -0,0 +1,106 @@ +"""Render an .excalidraw file (rectangles, text, multi-point arrows) to PNG with Pillow. +Lightweight renderer tailored for the ptu_lb workflow diagrams.""" +import json +import sys +from PIL import Image, ImageDraw, ImageFont + +SCALE = 2 # supersample for crisp output + + +def load_font(size): + for name in [ + "C:/Windows/Fonts/malgun.ttf", # Korean support + "C:/Windows/Fonts/segoeui.ttf", + "C:/Windows/Fonts/arial.ttf", + ]: + try: + return ImageFont.truetype(name, size) + except OSError: + continue + return ImageFont.load_default() + + +def main(src, dst): + with open(src, encoding="utf-8") as f: + data = json.load(f) + elements = [e for e in data["elements"] if not e.get("isDeleted")] + + # bounds + max_x = max_y = 0 + for e in elements: + ex = e["x"] + (e.get("width") or 0) + ey = e["y"] + (e.get("height") or 0) + if e["type"] == "arrow": + for px, py in e["points"]: + ex = max(ex, e["x"] + px) + ey = max(ey, e["y"] + py) + max_x = max(max_x, ex) + max_y = max(max_y, ey) + W = int((max_x + 40) * SCALE) + H = int((max_y + 40) * SCALE) + + img = Image.new("RGB", (W, H), "#ffffff") + d = ImageDraw.Draw(img) + + def sc(v): + return int(v * SCALE) + + # rectangles first + for e in elements: + if e["type"] != "rectangle": + continue + x0, y0 = sc(e["x"]), sc(e["y"]) + x1, y1 = sc(e["x"] + e["width"]), sc(e["y"] + e["height"]) + fill = e.get("backgroundColor") + if fill == "transparent": + fill = None + r = 10 * SCALE + d.rounded_rectangle([x0, y0, x1, y1], radius=r, fill=fill, + outline=e.get("strokeColor", "#1e1e1e"), + width=int(max(1, e.get("strokeWidth", 2)) * SCALE)) + + # arrows + def draw_arrow(pts, color, width): + for i in range(len(pts) - 1): + d.line([pts[i], pts[i + 1]], fill=color, width=width) + # arrowhead on last segment + (x0, y0), (x1, y1) = pts[-2], pts[-1] + import math + ang = math.atan2(y1 - y0, x1 - x0) + L = 12 * SCALE + for da in (math.radians(150), math.radians(-150)): + hx = x1 + L * math.cos(ang + da) + hy = y1 + L * math.sin(ang + da) + d.line([(x1, y1), (hx, hy)], fill=color, width=width) + + for e in elements: + if e["type"] != "arrow": + continue + pts = [(sc(e["x"] + px), sc(e["y"] + py)) for px, py in e["points"]] + draw_arrow(pts, e.get("strokeColor", "#1e1e1e"), + int(max(1, e.get("strokeWidth", 2)) * SCALE)) + + # text on top + for e in elements: + if e["type"] != "text": + continue + font = load_font(int(e.get("fontSize", 18) * SCALE)) + color = e.get("strokeColor", "#1e1e1e") + lines = e["text"].split("\n") + lh = int(e.get("fontSize", 18) * SCALE * e.get("lineHeight", 1.2)) + align = e.get("textAlign", "left") + box_w = sc(e.get("width") or 0) + for i, line in enumerate(lines): + tw = d.textlength(line, font=font) + tx = sc(e["x"]) + if align == "center": + tx = sc(e["x"]) + (box_w - tw) / 2 + ty = sc(e["y"]) + i * lh + d.text((tx, ty), line, fill=color, font=font) + + img.save(dst) + print(f"saved {dst} ({img.width}x{img.height})") + + +if __name__ == "__main__": + main(sys.argv[1], sys.argv[2]) diff --git a/ptu_lb/runbooks/adaptive_tpm_quota_controller.py b/ptu_lb/runbooks/adaptive_tpm_quota_controller.py new file mode 100644 index 0000000..e3a1573 --- /dev/null +++ b/ptu_lb/runbooks/adaptive_tpm_quota_controller.py @@ -0,0 +1,490 @@ +#!/usr/bin/env python3 +""" +Adaptive TPM Quota Controller (Python / Azure Automation Python 3 runbook) +========================================================================== + +목적 +---- +Korea / UK 2개 Foundry(gpt) 배포가 나눠 쓰는 **Reserved TPM quota 총량**을, +각 리전의 실제 사용률(utilization)에 맞게 **재분배**하도록 자동 조정한다. + +핵심 설계 결정 (중요) +-------------------- +* 두 배포 capacity 의 합은 사용자가 Reservation 한 총량(RESERVED_TOTAL_TPM)을 + **어떤 시점에도 초과할 수 없다.** 초과하는 증설 요청은 거부된다. +* 따라서 리전을 독립으로 올리지 않고, 고정된 Reserved 총량을 트래픽 비율에 + 맞춰 재분배한다: + - 필요량 합이 총량 이하면 각자 need 만 배정(나머지는 미할당 버퍼) + - 필요량 합이 총량을 초과하면(경합) 소비 TPM 비율로 비례 배분 +* **감축 먼저, 증설 나중**: 총량이 꽉 찬 상태에서 증설을 먼저 하면 실패하므로, + 잉여 리전을 먼저 감축해 headroom 을 확보한 뒤 증설한다. +* **loss 최소화**: 증설은 우선 미할당 버퍼로 충당(무손실)하고, 부족할 때만 + 잉여를 회수하되 감축·증설을 같은 사이클에서 연속 처리한다. +* 리전 최소 보장(MIN_CAPACITY)은 Reserved 총량의 25% 로 유지한다. +* capacity 단위와 TPM 의 매핑은 TPM_PER_CAPACITY_UNIT 로 명시한다. + +채운 갭 +------- +1. 실제 actuation: azure-mgmt-cognitiveservices 로 배포 sku.capacity 를 실제 갱신 +2. 상태 추적(폐루프): 실행 시 배포의 현재 capacity 를 SDK 로 읽어 기준값으로 사용 +3. 쿨다운: ScaleAction 이력에서 리전별 마지막 실제 변경 시각을 확인해 재조정 차단 +4. 실 메트릭 경로: TrafficWindow(Table) 에서 리전별 최근 N 윈도우를 읽어 이동평균 +5. 총량 제약: 두 리전 합 ≤ Reserved 총량, 감축→증설 순서로 불변식 보장 +6. az CLI 제거: Azure SDK(Table/mgmt) 사용 → Automation Python 샌드박스에서 동작 + +인증 +---- +* DefaultAzureCredential 사용. + - 로컬: `az login` 세션(AzureCliCredential)을 자동 사용 + - Azure Automation: 시스템 할당 관리 ID(ManagedIdentityCredential)를 자동 사용 +* 필요한 RBAC: + - Storage: "Storage Table Data Contributor" (TrafficWindow 읽기 / 이력 쓰기) + - Foundry: "Cognitive Services Contributor" (배포 capacity 갱신) + +로컬 실행 예시 +------------- + pip install -r requirements.txt + python adaptive_tpm_quota_controller.py --dry-run # 계산만 + python adaptive_tpm_quota_controller.py # 실제 조정 +""" + +from __future__ import annotations + +import argparse +import math +import sys +import uuid +from dataclasses import dataclass, field +from datetime import datetime, timezone + +from azure.identity import DefaultAzureCredential +from azure.data.tables import TableServiceClient +from azure.mgmt.cognitiveservices import CognitiveServicesManagementClient + + +# --------------------------------------------------------------------------- +# 설정 (운영 시 환경변수/파라미터로 외부화 권장) +# --------------------------------------------------------------------------- + +SUBSCRIPTION_ID = "2cf925b6-80cb-4567-abda-5ccd3010aab5" +RESOURCE_GROUP = "ptu-lb" +STORAGE_ACCOUNT = "ptulb59018sa" + +TRAFFIC_TABLE = "TrafficWindow" +DECISION_TABLE = "ControlDecision" +ACTION_TABLE = "ScaleAction" + +# 리전 → Foundry 계정/배포 매핑 (Reserved 총량을 두 리전이 공유) +REGIONS: dict[str, dict[str, str]] = { + "korea": {"account": "ptufoundrykr001", "deployment": "tpmtest-kr"}, + "uk": {"account": "ptufoundryuk001", "deployment": "tpmtest-uk"}, +} + +# --- 제어 파라미터 ----------------------------------------------------------- +METRIC_WINDOW_MINUTES = 5 # TrafficWindow 1건이 대표하는 시간(분) +SMOOTHING_WINDOWS = 3 # 이동평균에 사용할 최근 윈도우 개수 +TPM_PER_CAPACITY_UNIT = 1000 # capacity 1 unit ≈ 1000 TPM (운영값으로 조정) + +# Reserved quota 총량(합산 불변) — 두 리전 capacity 합은 이 값을 넘을 수 없다. +RESERVED_TOTAL_TPM = 20000 # Reservation 한 총 TPM 상한 +TOTAL_CAPACITY_UNITS = RESERVED_TOTAL_TPM // TPM_PER_CAPACITY_UNIT # capacity 환산(=20) + +TARGET_UTIL_HIGH = 0.75 # 이 사용률 초과 → 증설 압력 +TARGET_UTIL_LOW = 0.30 # 이 사용률 미만 → 감축(잉여 회수) +AIM_UTIL = 0.60 # 조정 후 목표로 삼는 사용률(과도한 왕복 방지용 중간값) + +MAX_STEP_UNITS = 5 # 1회 조정 최대 capacity 변경폭(단위) +MIN_CAPACITY_RATIO = 0.25 # 리전별 최소 보장 = Reserved 총량의 25% +MIN_CAPACITY = max(1, math.ceil(TOTAL_CAPACITY_UNITS * MIN_CAPACITY_RATIO)) # =5 +# 리전 최대 = 총량 − 나머지 리전 최소 보장 합. 2리전 기본값(=TOTAL−MIN)이며, +# N리전 확장 시에는 region_max(n) 으로 (n-1)*MIN 을 차감해 일반화한다. +MAX_CAPACITY = TOTAL_CAPACITY_UNITS - MIN_CAPACITY # 2리전 기본값 =15 +COOLDOWN_MINUTES = 20 # 실제 조정 후 재조정 금지 시간(분) + + +def clamp(value: int, low: int, high: int) -> int: + return max(low, min(value, high)) + + +def region_max(n: int) -> int: + """N리전 일반화된 리전별 최대 capacity(= 총량 − 나머지 (n-1) 리전 최소 보장). + + 한 리전이 최대를 가져가도 나머지 리전이 각각 MIN_CAPACITY 를 확보할 수 있어야 + 하므로 상한은 TOTAL − (n-1)*MIN 이다. n=2 이면 TOTAL − MIN 로 기존과 동일. + """ + return max(MIN_CAPACITY, TOTAL_CAPACITY_UNITS - (max(n, 1) - 1) * MIN_CAPACITY) + + +def even_split_targets(regions: list[str]) -> dict[str, int]: + """RI(Reserved) 총량을 리전 수로 균등 분배(내림 + 나머지 앞 리전부터). + + 메트릭 조회 실패 등 판단 근거가 열화됐을 때, 미할당 버퍼로 PTU 를 낭비하지 않고 + RI 에 맞는 PTU 를 최대 소진하기 위한 **fallback 목표**다. 합은 항상 TOTAL 이다. + """ + n = len(regions) or 1 + base = TOTAL_CAPACITY_UNITS // n + rem = TOTAL_CAPACITY_UNITS - base * n + return {r: base + (1 if i < rem else 0) for i, r in enumerate(regions)} + + +@dataclass +class RegionResult: + region: str + avg_consumed_tpm: float + current_capacity: int + utilization: float + decision: str # "up" | "down" | "hold" + target_capacity: int + reason: str + cooldown_blocked: bool + executed: bool + contention: bool = False + error: str = "" + + +# --------------------------------------------------------------------------- +# 메트릭 조회 +# --------------------------------------------------------------------------- + +def get_recent_windows(table_svc: TableServiceClient, region: str, limit: int) -> list[dict]: + """리전 파티션의 최근(RowKey 내림차순) limit 개 윈도우를 반환.""" + tc = table_svc.get_table_client(TRAFFIC_TABLE) + entities = tc.query_entities( + query_filter="PartitionKey eq @pk", + parameters={"pk": region}, + ) + rows = list(entities) + # RowKey 는 ISO8601 타임스탬프 문자열이므로 문자열 정렬 = 시간 정렬 + rows.sort(key=lambda e: e["RowKey"], reverse=True) + return rows[:limit] + + +def avg_consumed_tpm(rows: list[dict]) -> float: + """윈도우별 (inputTokens+outputTokens)/window_minutes 의 평균(분당 소비 TPM).""" + if not rows: + return 0.0 + total = 0.0 + for r in rows: + tokens = float(r.get("inputTokens", 0)) + float(r.get("outputTokens", 0)) + total += tokens / METRIC_WINDOW_MINUTES + return round(total / len(rows), 2) + + +# --------------------------------------------------------------------------- +# 배포 capacity 조회/갱신 (실제 actuation) +# --------------------------------------------------------------------------- + +def get_deployment_capacity(mgmt: CognitiveServicesManagementClient, account: str, deployment: str) -> int: + dep = mgmt.deployments.get(RESOURCE_GROUP, account, deployment) + return int(dep.sku.capacity) + + +def set_deployment_capacity( + mgmt: CognitiveServicesManagementClient, account: str, deployment: str, capacity: int +) -> None: + """배포 SKU capacity 를 실제로 갱신한다(비동기 → 완료 대기).""" + dep = mgmt.deployments.get(RESOURCE_GROUP, account, deployment) + dep.sku.capacity = capacity + poller = mgmt.deployments.begin_create_or_update( + RESOURCE_GROUP, account, deployment, dep + ) + poller.result() + + +# --------------------------------------------------------------------------- +# 쿨다운 체크 +# --------------------------------------------------------------------------- + +def is_in_cooldown(table_svc: TableServiceClient, region: str, now: datetime) -> bool: + """해당 리전의 마지막 '실제 변경' ScaleAction 이 쿨다운 이내면 True.""" + tc = table_svc.get_table_client(ACTION_TABLE) + actions = list( + tc.query_entities( + query_filter="region eq @r", + parameters={"r": region}, + ) + ) + # 실제 capacity 가 바뀐(=changed=True) 액션만 대상으로 최신 것을 찾는다 + changed = [a for a in actions if str(a.get("changed", "")).lower() == "true"] + if not changed: + return False + changed.sort(key=lambda a: a.metadata["timestamp"], reverse=True) + last_ts: datetime = changed[0].metadata["timestamp"] + if last_ts.tzinfo is None: + last_ts = last_ts.replace(tzinfo=timezone.utc) + elapsed_min = (now - last_ts).total_seconds() / 60.0 + return elapsed_min < COOLDOWN_MINUTES + + +# --------------------------------------------------------------------------- +# 공유 상한(Reserved quota) 재분배 로직 +# --------------------------------------------------------------------------- + +@dataclass +class RegionState: + region: str + account: str + deployment: str + current: int + consumed: float + utilization: float + + +def need_capacity(consumed_tpm: float, cap_max: int = MAX_CAPACITY) -> int: + """AIM_UTIL 에서 돌릴 수 있는 capacity 를 역산(리전 최소/최대로 clamp). + + cap_max 는 리전 수에 따른 상한(region_max(n))으로, 기본값은 2리전 기준이다. + """ + if consumed_tpm <= 0: + return MIN_CAPACITY + raw = math.ceil(consumed_tpm / (AIM_UTIL * TPM_PER_CAPACITY_UNIT)) + return clamp(raw, MIN_CAPACITY, cap_max) + + +def compute_targets(states: list[RegionState]) -> tuple[dict[str, int], dict[str, int], bool]: + """ + Reserved 총량 제약(sum <= TOTAL_CAPACITY_UNITS) 하에서 리전별 목표 capacity 를 산출. + + 반환: (targets, needs, contention) + - needs: 각 리전이 AIM_UTIL 로 돌기 위해 필요한 capacity + - contention: 총 필요량이 Reserved 총량을 초과(True) → 소비 TPM 비율로 비례 배분 + """ + n = len(states) + cap_max = region_max(n) # 리전 수에 따른 상한(2리전이면 TOTAL−MIN 로 기존과 동일) + needs = {s.region: need_capacity(s.consumed, cap_max) for s in states} + total_need = sum(needs.values()) + + if total_need <= TOTAL_CAPACITY_UNITS: + # 여유 있음: 필요량만 배정하고 남는 건 미할당 버퍼로 둔다(증설은 무손실). + return dict(needs), needs, False + + # 경합: MIN 을 먼저 보장하고 남은 용량을 소비 TPM 비율로 배분한다. + rem = TOTAL_CAPACITY_UNITS - n * MIN_CAPACITY + total_consumed = sum(max(s.consumed, 0.0) for s in states) or 1.0 + targets: dict[str, int] = {} + for s in states: + extra = int(rem * max(s.consumed, 0.0) / total_consumed) + targets[s.region] = clamp(MIN_CAPACITY + extra, MIN_CAPACITY, cap_max) + # 내림으로 남은 잔여 unit 을 소비가 큰 리전부터 채운다(합이 TOTAL 을 넘지 않게). + leftover = TOTAL_CAPACITY_UNITS - sum(targets.values()) + for s in sorted(states, key=lambda x: x.consumed, reverse=True): + while leftover > 0 and targets[s.region] < cap_max: + targets[s.region] += 1 + leftover -= 1 + return targets, needs, True + + +def plan_region(state: RegionState, target: int, contention: bool) -> tuple[str, int, str]: + """ + 목표치를 dead-band/스텝 제한으로 완충해 이번 사이클의 planned capacity 를 만든다. + 반환: (decision, planned, reason) + """ + within_band = TARGET_UTIL_LOW <= state.utilization <= TARGET_UTIL_HIGH + if within_band and not contention: + # 편안한 구간이고 경합도 없음 → 불필요한 왕복(churn) 방지 위해 유지 + return "hold", state.current, f"util {state.utilization:.2%} in dead-band" + + delta = clamp(target - state.current, -MAX_STEP_UNITS, MAX_STEP_UNITS) + planned = state.current + delta + tag = f"util {state.utilization:.2%} -> target {target} (contention={contention})" + if planned > state.current: + return "up", planned, tag + if planned < state.current: + return "down", planned, tag + return "hold", state.current, f"util {state.utilization:.2%} bounded, no change" + + +# --------------------------------------------------------------------------- +# 이력 기록 +# --------------------------------------------------------------------------- + +def _date_pk(now: datetime) -> str: + return now.strftime("%Y%m%d") + + +def _row_key(now: datetime, region: str) -> str: + return f"{now.strftime('%Y%m%dT%H%M%S%f')}-{region}-{uuid.uuid4().hex[:8]}" + + +def write_decision(table_svc: TableServiceClient, now: datetime, res: RegionResult) -> None: + tc = table_svc.get_table_client(DECISION_TABLE) + tc.create_entity({ + "PartitionKey": _date_pk(now), + "RowKey": _row_key(now, res.region), + "region": res.region, + "utilization": res.utilization, + "avgConsumedTpm": res.avg_consumed_tpm, + "capacityBefore": res.current_capacity, + "capacityTarget": res.target_capacity, + "decision": res.decision, + "reason": res.reason, + "cooldownBlocked": res.cooldown_blocked, + "executed": res.executed, + }) + + +def write_action(table_svc: TableServiceClient, now: datetime, res: RegionResult, duration_ms: int) -> None: + tc = table_svc.get_table_client(ACTION_TABLE) + changed = res.executed and res.target_capacity != res.current_capacity + tc.create_entity({ + "PartitionKey": _date_pk(now), + "RowKey": _row_key(now, res.region), + "region": res.region, + "requestPayload": f"{res.region} capacity {res.current_capacity} to {res.target_capacity}", + "capacityBefore": res.current_capacity, + "capacityTarget": res.target_capacity, + "changed": changed, + "responseStatus": 200 if not res.error else 500, + "success": res.error == "", + "errorMessage": res.error, + "durationMs": duration_ms, + }) + + +# --------------------------------------------------------------------------- +# 메인 제어 사이클 +# --------------------------------------------------------------------------- + +def run_cycle(dry_run: bool) -> list[RegionResult]: + now = datetime.now(timezone.utc) + cred = DefaultAzureCredential() + + table_svc = TableServiceClient( + endpoint=f"https://{STORAGE_ACCOUNT}.table.core.windows.net", + credential=cred, + ) + mgmt = CognitiveServicesManagementClient(cred, SUBSCRIPTION_ID) + + # --- Phase 1: 상태 수집 (리전별 현재 capacity + 이동평균 사용률) -------------- + # 메트릭 조회 실패는 리전을 degraded 로 표시한다(판단 근거 열화 → RI 균등분배 fallback). + states: list[RegionState] = [] + degraded: list[str] = [] + for region, cfg in REGIONS.items(): + account, deployment = cfg["account"], cfg["deployment"] + current = get_deployment_capacity(mgmt, account, deployment) + try: + rows = get_recent_windows(table_svc, region, SMOOTHING_WINDOWS) + consumed = avg_consumed_tpm(rows) + except Exception: # noqa: BLE001 - 메트릭 열화 시 fallback 으로 진행 + consumed = 0.0 + degraded.append(region) + provisioned = current * TPM_PER_CAPACITY_UNIT + util = round(consumed / provisioned, 4) if provisioned > 0 else 0.0 + states.append(RegionState(region, account, deployment, current, consumed, util)) + + # --- Phase 2: 목표 산출(총량 제약) + dead-band/스텝/쿨다운 완충 --------------- + # 정상: 소비 비율 재분배. Fallback: 메트릭 열화 시 RI(Reserved) 총량을 균등분배해 + # 미할당 버퍼 없이 RI 에 맞는 PTU 를 최대 소진한다(감축→증설 순서는 그대로 유지). + fallback = bool(degraded) + if fallback: + targets = even_split_targets([s.region for s in states]) + _needs = dict(targets) + contention = True # 균등분배는 총량 만재 배분(버퍼 0) + else: + targets, _needs, contention = compute_targets(states) + + results: dict[str, RegionResult] = {} + planned_map: dict[str, int] = {} + for s in states: + decision, planned, reason = plan_region(s, targets[s.region], contention) + if fallback: + reason += " | FALLBACK RI 균등분배(메트릭 열화)" + + cooldown_blocked = False + if decision != "hold" and is_in_cooldown(table_svc, s.region, now): + cooldown_blocked = True + planned = s.current + decision = "hold" + reason += " | blocked by cooldown" + + planned_map[s.region] = planned + results[s.region] = RegionResult( + region=s.region, + avg_consumed_tpm=s.consumed, + current_capacity=s.current, + utilization=s.utilization, + decision=decision, + target_capacity=planned, # 실제 적용 후 갱신될 수 있음 + reason=reason, + cooldown_blocked=cooldown_blocked, + executed=False, + contention=contention, + ) + + # --- Phase 3: 총량 제약 실행 (감축 먼저 → 증설은 남은 headroom 한도) ---------- + free = TOTAL_CAPACITY_UNITS - sum(s.current for s in states) + durations: dict[str, int] = {s.region: 0 for s in states} + + def _apply(st: RegionState, new_cap: int, res: RegionResult) -> int: + start = datetime.now(timezone.utc) + try: + if not dry_run: + set_deployment_capacity(mgmt, st.account, st.deployment, new_cap) + res.target_capacity = new_cap + res.executed = (not dry_run) and (new_cap != st.current) + except Exception as exc: # noqa: BLE001 - 기록 후 다음 리전 진행 + res.error = str(exc)[:500] + res.target_capacity = st.current + return int((datetime.now(timezone.utc) - start).total_seconds() * 1000) + + # 3-1) 감축(donor) 먼저 → headroom 확보 (loss 없이 여유 리전에서 회수) + for s in states: + res = results[s.region] + planned = planned_map[s.region] + if not res.cooldown_blocked and planned < s.current: + durations[s.region] = _apply(s, planned, res) + if not res.error: + free += s.current - res.target_capacity + + # 3-2) 증설(recipient): 사용률 높은 리전부터, 남은 headroom 한도로만 + for s in sorted(states, key=lambda x: x.utilization, reverse=True): + res = results[s.region] + planned = planned_map[s.region] + if res.cooldown_blocked or planned <= s.current: + continue + inc = min(planned - s.current, free) + if inc <= 0: + res.decision = "hold" + res.reason += " | no headroom (Reserved 총량 소진)" + continue + new_cap = s.current + inc + durations[s.region] = _apply(s, new_cap, res) + if not res.error: + free -= (res.target_capacity - s.current) + if res.target_capacity < planned: + res.reason += f" | headroom-limited to {res.target_capacity}" + + # --- Phase 4: 이력 기록 (판단/액션 모두 남겨 관측 가능하게) ------------------- + for s in states: + res = results[s.region] + write_decision(table_svc, now, res) + write_action(table_svc, now, res, durations[s.region]) + + return [results[s.region] for s in states] + + +def main(argv: list[str]) -> int: + parser = argparse.ArgumentParser(description="Adaptive TPM quota controller (Reserved 총량 공유 재분배)") + parser.add_argument("--dry-run", action="store_true", help="계산/기록만 하고 실제 capacity 는 바꾸지 않음") + args = parser.parse_args(argv) + + results = run_cycle(dry_run=args.dry_run) + + print(f"=== Adaptive TPM Quota Controller (dryRun={args.dry_run}) ===") + total_after = 0 + for r in results: + total_after += r.target_capacity + print( + f"[{r.region}] consumedTpm={r.avg_consumed_tpm} " + f"capacity={r.current_capacity}->{r.target_capacity} util={r.utilization:.2%} " + f"decision={r.decision} contention={r.contention} " + f"cooldown={r.cooldown_blocked} executed={r.executed} " + f"reason='{r.reason}'" + + (f" error='{r.error}'" if r.error else "") + ) + print(f"total capacity after={total_after}/{TOTAL_CAPACITY_UNITS} (Reserved cap)") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main(sys.argv[1:])) diff --git a/ptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1 b/ptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1 new file mode 100644 index 0000000..5758cdf --- /dev/null +++ b/ptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1 @@ -0,0 +1,459 @@ +<# +.SYNOPSIS + Adaptive TPM Quota Controller (PowerShell / Azure Automation) — OPTIONAL 대체 구현 + +.DESCRIPTION + Python 런북(adaptive_tpm_quota_controller.py)과 **동일한 제어 로직**을 PowerShell(Az 모듈)로 + 이식한 "선택적(optional) 대체 구현"이다. 기본 구현은 Python 런북이며, 본 스크립트는 + PowerShell 선호 환경을 위한 동등 참조본이다. + + [Python 버전과의 동등성 — 핵심 설계 결정] + * 두 배포 capacity 의 합은 Reservation 한 총량(ReservedTotalTpm)을 어떤 시점에도 + 초과할 수 없다. 고정된 Reserved 총량을 트래픽 비율에 맞춰 재분배한다: + - 필요량 합 ≤ 총량 → 각자 need 만 배정(나머지는 미할당 버퍼) + - 필요량 합 > 총량 → 소비 TPM 비율로 비례 배분(경합) + * 감축 먼저, 증설 나중: 총량이 꽉 찬 상태에서 증설을 먼저 하면 실패하므로, + 잉여 리전을 먼저 감축해 headroom 을 확보한 뒤 증설한다(loss 최소화). + * 리전 최소 보장(MinCapacity)은 Reserved 총량의 25% 로 유지한다. + * capacity(배포 SKU capacity) 를 조정 knob 으로 직접 사용. 1 unit ≈ TpmPerCapacityUnit. + * 실제 actuation: Az.CognitiveServices 로 배포 SKU capacity 를 실제 갱신(폐루프). + * 쿨다운: ScaleAction 이력에서 리전별 마지막 실제 변경 시각을 확인해 재조정 차단. + +.NOTES + [인증 / 의존 모듈] + * 인증: 관리 ID -> Connect-AzAccount -Identity (로컬은 Connect-AzAccount 로 로그인) + * 필요 모듈: Az.Accounts, Az.CognitiveServices, AzTable + * 필요 RBAC: + - Storage: "Storage Table Data Contributor" (TrafficWindow 읽기 / 이력 쓰기) + - Foundry: "Cognitive Services Contributor" (배포 capacity 갱신) + + [네트워크 주의] + * Storage 가 publicNetworkAccess=Disabled + Private Endpoint 구성인 경우, 이 런북도 + 프라이빗 경로가 닿는 곳(예: Hybrid Runbook Worker)에서 실행되어야 테이블에 접근된다. + (Linux Hybrid Worker 에는 PowerShell 이 기본 설치되어 있지 않을 수 있으므로, + PowerShell 로 운영하려면 Windows Hybrid Worker 를 사용) + +.EXAMPLE + .\optional_adaptive_tpm_quota_controller.ps1 -DryRun # 계산/기록만 + +.EXAMPLE + .\optional_adaptive_tpm_quota_controller.ps1 # 실제 조정 +#> + +param( + # --- 대상 리소스 ------------------------------------------------------------- + [string]$SubscriptionId = "2cf925b6-80cb-4567-abda-5ccd3010aab5", + [string]$ResourceGroup = "ptu-lb", + [string]$StorageAccountName = "ptulb59018sa", + + [string]$TrafficTable = "TrafficWindow", + [string]$DecisionTable = "ControlDecision", + [string]$ActionTable = "ScaleAction", + + # --- 제어 파라미터 (Python 버전 상수와 동일) -------------------------------- + [int]$MetricWindowMinutes = 5, # TrafficWindow 1건이 대표하는 시간(분) + [int]$SmoothingWindows = 3, # 이동평균에 사용할 최근 윈도우 개수 + [int]$TpmPerCapacityUnit = 1000, # capacity 1 unit ≈ 1000 TPM (운영값으로 조정) + + [double]$TargetUtilHigh = 0.75, # 초과 시 증설 압력 + [double]$TargetUtilLow = 0.30, # 미만 시 감축(잉여 회수) + [double]$AimUtil = 0.60, # 조정 후 목표 사용률(중간값) + + [int]$MaxStepUnits = 5, # 1회 조정 최대 capacity 변경폭(단위) + [int]$ReservedTotalTpm = 20000, # Reservation 한 총 TPM 상한(합산 불변) + [double]$MinCapacityRatio = 0.25, # 리전별 최소 보장 = Reserved 총량의 25% + [int]$CooldownMinutes = 20, # 실제 조정 후 재조정 금지 시간(분) + + [switch]$DryRun +) + +$ErrorActionPreference = "Stop" + +# --- Reserved 총량 파생값 (합산 불변식의 기준) --------------------------------- +$TotalCapacityUnits = [int]($ReservedTotalTpm / $TpmPerCapacityUnit) # =20 +$MinCapacity = [int][math]::Max(1, [math]::Ceiling($TotalCapacityUnits * $MinCapacityRatio)) # =5 +# 리전 최대 = 총량 − 나머지 리전 최소 보장 합. 2리전 기본값(=TOTAL−MIN)이며 +# N리전 확장 시에는 Get-RegionMax 로 (n-1)*MIN 을 차감해 일반화한다. +$MaxCapacity = $TotalCapacityUnits - $MinCapacity # 2리전 기본값 =15 + +# 리전 -> Foundry 계정/배포 매핑 (Reserved 총량을 두 리전이 공유) +$Regions = [ordered]@{ + korea = @{ Account = "ptufoundrykr001"; Deployment = "tpmtest-kr" } + uk = @{ Account = "ptufoundryuk001"; Deployment = "tpmtest-uk" } +} + +# --------------------------------------------------------------------------- +# 인증 + 테이블 컨텍스트 +# --------------------------------------------------------------------------- + +function Connect-Context { + # Azure Automation: 시스템 할당 관리 ID. 로컬: 이미 로그인된 컨텍스트가 있으면 재사용. + if (-not (Get-AzContext)) { + try { + Connect-AzAccount -Identity -Subscription $SubscriptionId | Out-Null + } + catch { + Connect-AzAccount -Subscription $SubscriptionId | Out-Null + } + } + Set-AzContext -Subscription $SubscriptionId | Out-Null +} + +function Get-CloudTable { + param([string]$TableName) + # 공유키 없이 Entra(RBAC) 자격으로 접근 + $ctx = New-AzStorageContext -StorageAccountName $StorageAccountName -UseConnectedAccount + return (Get-AzStorageTable -Name $TableName -Context $ctx).CloudTable +} + +function Clamp { + param([int]$Value, [int]$Min, [int]$Max) + return [math]::Min([math]::Max($Value, $Min), $Max) +} + +function Get-RegionMax { + # N리전 일반화된 리전별 최대 capacity(= 총량 − (n-1)*MIN). n=2 이면 TOTAL−MIN 로 기존과 동일. + param([int]$N) + $nn = [math]::Max($N, 1) + return [math]::Max($MinCapacity, $TotalCapacityUnits - (($nn - 1) * $MinCapacity)) +} + +function Get-EvenSplitTargets { + # RI(Reserved) 총량을 리전 수로 균등 분배(내림 + 나머지 앞 리전부터). 합은 항상 TOTAL. + # 메트릭 열화 등 판단 근거 상실 시, 미할당 버퍼 없이 RI 에 맞는 PTU 를 최대 소진하는 fallback 목표. + param([array]$Regions) + $n = [math]::Max($Regions.Count, 1) + $base = [int][math]::Floor($TotalCapacityUnits / $n) + $rem = $TotalCapacityUnits - ($base * $n) + $targets = @{} + for ($i = 0; $i -lt $Regions.Count; $i++) { + $targets[$Regions[$i]] = $base + (&{ if ($i -lt $rem) { 1 } else { 0 } }) + } + return $targets +} + +# --------------------------------------------------------------------------- +# 메트릭 조회 +# --------------------------------------------------------------------------- + +function Get-RecentWindows { + param([string]$Region, [int]$Limit) + + $table = Get-CloudTable -TableName $TrafficTable + $rows = Get-AzTableRow -Table $table -CustomFilter "PartitionKey eq '$Region'" + if (-not $rows) { return @() } + + # RowKey 는 ISO8601 타임스탬프 문자열이므로 문자열 정렬 = 시간 정렬 + return @($rows | Sort-Object -Property RowKey -Descending | Select-Object -First $Limit) +} + +function Get-AvgConsumedTpm { + param([array]$Rows) + # 윈도우별 (inputTokens+outputTokens)/window_minutes 의 평균(분당 소비 TPM) + if ($Rows.Count -eq 0) { return 0.0 } + + $total = 0.0 + foreach ($r in $Rows) { + $tokens = [double]($r.inputTokens) + [double]($r.outputTokens) + $total += $tokens / $MetricWindowMinutes + } + return [math]::Round($total / $Rows.Count, 2) +} + +# --------------------------------------------------------------------------- +# 배포 capacity 조회/갱신 (실제 actuation) +# --------------------------------------------------------------------------- + +function Get-DeploymentCapacity { + param([string]$Account, [string]$Deployment) + $dep = Get-AzCognitiveServicesAccountDeployment ` + -ResourceGroupName $ResourceGroup -AccountName $Account -Name $Deployment + return [int]$dep.Sku.Capacity +} + +function Set-DeploymentCapacity { + param([string]$Account, [string]$Deployment, [int]$Capacity) + # 기존 배포의 model/properties 는 유지하고 SKU capacity 만 갱신(create-or-update) + $existing = Get-AzCognitiveServicesAccountDeployment ` + -ResourceGroupName $ResourceGroup -AccountName $Account -Name $Deployment + $sku = New-Object 'Microsoft.Azure.Management.CognitiveServices.Models.Sku' ` + -ArgumentList $existing.Sku.Name, $null, $null, $Capacity, $null + New-AzCognitiveServicesAccountDeployment ` + -ResourceGroupName $ResourceGroup -AccountName $Account -Name $Deployment ` + -Properties $existing.Properties -Sku $sku | Out-Null +} + +# --------------------------------------------------------------------------- +# 쿨다운 체크 +# --------------------------------------------------------------------------- + +function Test-InCooldown { + param([string]$Region, [datetime]$Now) + + $table = Get-CloudTable -TableName $ActionTable + $actions = Get-AzTableRow -Table $table -CustomFilter "region eq '$Region'" + if (-not $actions) { return $false } + + # 실제 capacity 가 바뀐(changed=true) 액션만 대상으로 최신 것을 찾는다 + $changed = @($actions | Where-Object { "$($_.changed)".ToLower() -eq "true" }) + if ($changed.Count -eq 0) { return $false } + + $last = $changed | Sort-Object -Property TableTimestamp -Descending | Select-Object -First 1 + $elapsedMin = ($Now.ToUniversalTime() - $last.TableTimestamp.UtcDateTime).TotalMinutes + return ($elapsedMin -lt $CooldownMinutes) +} + +# --------------------------------------------------------------------------- +# 공유 상한(Reserved quota) 재분배 로직 +# --------------------------------------------------------------------------- + +function Get-NeedCapacity { + param([double]$ConsumedTpm, [int]$CapMax = $MaxCapacity) + if ($ConsumedTpm -le 0) { return $MinCapacity } + $raw = [int][math]::Ceiling($ConsumedTpm / ($AimUtil * $TpmPerCapacityUnit)) + return (Clamp -Value $raw -Min $MinCapacity -Max $CapMax) +} + +function Get-Targets { + # $States: 각 원소 = @{ Region; Account; Deployment; Current; Consumed; Utilization } + param([array]$States) + + $n = $States.Count + $capMax = Get-RegionMax -N $n # 리전 수에 따른 상한(2리전이면 TOTAL−MIN 로 기존과 동일) + $needs = @{} + foreach ($s in $States) { $needs[$s.Region] = Get-NeedCapacity -ConsumedTpm $s.Consumed -CapMax $capMax } + $totalNeed = ($needs.Values | Measure-Object -Sum).Sum + + if ($totalNeed -le $TotalCapacityUnits) { + # 여유 있음: 필요량만 배정, 나머지는 미할당 버퍼(증설은 무손실) + return @{ Targets = $needs; Contention = $false } + } + + # 경합: MIN 을 먼저 보장하고 남은 용량을 소비 TPM 비율로 배분 + $rem = $TotalCapacityUnits - ($n * $MinCapacity) + $totalConsumed = ($States | ForEach-Object { [math]::Max($_.Consumed, 0.0) } | Measure-Object -Sum).Sum + if ($totalConsumed -le 0) { $totalConsumed = 1.0 } + + $targets = @{} + foreach ($s in $States) { + $extra = [int][math]::Floor($rem * [math]::Max($s.Consumed, 0.0) / $totalConsumed) + $targets[$s.Region] = Clamp -Value ($MinCapacity + $extra) -Min $MinCapacity -Max $capMax + } + # 내림으로 남은 잔여 unit 을 소비가 큰 리전부터 채운다(합이 TOTAL 을 넘지 않게) + $leftover = $TotalCapacityUnits - ($targets.Values | Measure-Object -Sum).Sum + foreach ($s in ($States | Sort-Object -Property Consumed -Descending)) { + while ($leftover -gt 0 -and $targets[$s.Region] -lt $capMax) { + $targets[$s.Region] += 1 + $leftover -= 1 + } + } + return @{ Targets = $targets; Contention = $true } +} + +function Get-RegionPlan { + param([hashtable]$State, [int]$Target, [bool]$Contention) + + $util = $State.Utilization + $withinBand = ($util -ge $TargetUtilLow -and $util -le $TargetUtilHigh) + if ($withinBand -and -not $Contention) { + # 편안한 구간이고 경합도 없음 → 불필요한 왕복(churn) 방지 위해 유지 + return @{ Decision = "hold"; Planned = $State.Current; Reason = "util $([math]::Round($util*100,2))% in dead-band" } + } + + $delta = Clamp -Value ([int]($Target - $State.Current)) -Min (-$MaxStepUnits) -Max $MaxStepUnits + $planned = $State.Current + $delta + $tag = "util $([math]::Round($util*100,2))% -> target $Target (contention=$Contention)" + if ($planned -gt $State.Current) { return @{ Decision = "up"; Planned = $planned; Reason = $tag } } + if ($planned -lt $State.Current) { return @{ Decision = "down"; Planned = $planned; Reason = $tag } } + return @{ Decision = "hold"; Planned = $State.Current; Reason = "util $([math]::Round($util*100,2))% bounded, no change" } +} + +# --------------------------------------------------------------------------- +# 이력 기록 +# --------------------------------------------------------------------------- + +function New-RowKey { + param([datetime]$Now, [string]$Region) + return "$($Now.ToString('yyyyMMddTHHmmssffffff'))-$Region-$([guid]::NewGuid().ToString('N').Substring(0,8))" +} + +function Write-Decision { + param([datetime]$Now, [hashtable]$Res) + $table = Get-CloudTable -TableName $DecisionTable + Add-AzTableRow -Table $table ` + -PartitionKey $Now.ToString('yyyyMMdd') -RowKey (New-RowKey -Now $Now -Region $Res.Region) ` + -Property @{ + region = $Res.Region + utilization = $Res.Utilization + avgConsumedTpm = $Res.AvgConsumedTpm + capacityBefore = $Res.CurrentCapacity + capacityTarget = $Res.Target + decision = $Res.Decision + reason = $Res.Reason + cooldownBlocked = $Res.CooldownBlocked + executed = $Res.Executed + } | Out-Null +} + +function Write-Action { + param([datetime]$Now, [hashtable]$Res, [int]$DurationMs) + $table = Get-CloudTable -TableName $ActionTable + $changed = ($Res.Executed -and ($Res.Target -ne $Res.CurrentCapacity)) + Add-AzTableRow -Table $table ` + -PartitionKey $Now.ToString('yyyyMMdd') -RowKey (New-RowKey -Now $Now -Region $Res.Region) ` + -Property @{ + region = $Res.Region + requestPayload = "$($Res.Region) capacity $($Res.CurrentCapacity) to $($Res.Target)" + capacityBefore = $Res.CurrentCapacity + capacityTarget = $Res.Target + changed = $changed + responseStatus = if ($Res.Error) { 500 } else { 200 } + success = [string]::IsNullOrEmpty($Res.Error) + errorMessage = $Res.Error + durationMs = $DurationMs + } | Out-Null +} + +# --------------------------------------------------------------------------- +# 메인 제어 사이클 +# --------------------------------------------------------------------------- + +Connect-Context +$now = (Get-Date).ToUniversalTime() + +# --- Phase 1: 상태 수집 (리전별 현재 capacity + 이동평균 사용률) --------------- +# 메트릭 조회 실패는 리전을 degraded 로 표시한다(판단 근거 열화 → RI 균등분배 fallback). +$states = @() +$degraded = @() +foreach ($region in $Regions.Keys) { + $cfg = $Regions[$region] + $current = Get-DeploymentCapacity -Account $cfg.Account -Deployment $cfg.Deployment + try { + $rows = Get-RecentWindows -Region $region -Limit $SmoothingWindows + $consumed = Get-AvgConsumedTpm -Rows $rows + } + catch { + $consumed = 0.0 + $degraded += $region + } + $provisioned = $current * $TpmPerCapacityUnit + $util = if ($provisioned -gt 0) { [math]::Round($consumed / $provisioned, 4) } else { 0.0 } + $states += @{ + Region = $region; Account = $cfg.Account; Deployment = $cfg.Deployment + Current = $current; Consumed = $consumed; Utilization = $util + } +} + +# --- Phase 2: 목표 산출(총량 제약) + dead-band/스텝/쿨다운 완충 ---------------- +# 정상: 소비 비율 재분배. Fallback: 메트릭 열화 시 RI 총량을 균등분배해 RI 에 맞는 PTU 를 최대 소진. +$fallback = ($degraded.Count -gt 0) +if ($fallback) { + $targets = Get-EvenSplitTargets -Regions @($states | ForEach-Object { $_.Region }) + $contention = $true # 균등분배는 총량 만재 배분(버퍼 0) +} +else { + $alloc = Get-Targets -States $states + $targets = $alloc.Targets + $contention = $alloc.Contention +} + +$results = @{} +$plannedMap = @{} +$durations = @{} +foreach ($s in $states) { + $p = Get-RegionPlan -State $s -Target ([int]$targets[$s.Region]) -Contention $contention + $decision = $p.Decision + $planned = $p.Planned + $reason = $p.Reason + if ($fallback) { $reason = "$reason | FALLBACK RI 균등분배(메트릭 열화)" } + + $cooldownBlocked = $false + if ($decision -ne "hold" -and (Test-InCooldown -Region $s.Region -Now $now)) { + $cooldownBlocked = $true + $planned = $s.Current + $decision = "hold" + $reason = "$reason | blocked by cooldown" + } + + $plannedMap[$s.Region] = $planned + $durations[$s.Region] = 0 + $results[$s.Region] = @{ + Region = $s.Region; Account = $s.Account; Deployment = $s.Deployment + AvgConsumedTpm = $s.Consumed; CurrentCapacity = $s.Current; Utilization = $s.Utilization + Decision = $decision; Target = $planned; Reason = $reason + CooldownBlocked = $cooldownBlocked; Executed = $false; Contention = $contention; Error = "" + } +} + +# --- Phase 3: 총량 제약 실행 (감축 먼저 → 증설은 남은 headroom 한도) ----------- +$free = $TotalCapacityUnits - (($states | ForEach-Object { $_.Current } | Measure-Object -Sum).Sum) + +function Invoke-Apply { + param([hashtable]$St, [hashtable]$Res, [int]$NewCap) + $sw = [System.Diagnostics.Stopwatch]::StartNew() + try { + if (-not $DryRun) { + Set-DeploymentCapacity -Account $St.Account -Deployment $St.Deployment -Capacity $NewCap + } + $Res.Target = $NewCap + $Res.Executed = ((-not $DryRun) -and ($NewCap -ne $St.Current)) + } + catch { + $msg = "$($_.Exception.Message)" + $Res.Error = $msg.Substring(0, [math]::Min(500, $msg.Length)) + $Res.Target = $St.Current + } + $sw.Stop() + return [int]$sw.ElapsedMilliseconds +} + +# 3-1) 감축(donor) 먼저 → headroom 확보 (여유 리전에서 회수, loss 없음) +foreach ($s in $states) { + $res = $results[$s.Region] + $planned = $plannedMap[$s.Region] + if (-not $res.CooldownBlocked -and $planned -lt $s.Current) { + $durations[$s.Region] = Invoke-Apply -St $s -Res $res -NewCap $planned + if (-not $res.Error) { $free += ($s.Current - $res.Target) } + } +} + +# 3-2) 증설(recipient): 사용률 높은 리전부터, 남은 headroom 한도로만 +foreach ($s in ($states | Sort-Object -Property Utilization -Descending)) { + $res = $results[$s.Region] + $planned = $plannedMap[$s.Region] + if ($res.CooldownBlocked -or $planned -le $s.Current) { continue } + $inc = [math]::Min($planned - $s.Current, $free) + if ($inc -le 0) { + $res.Decision = "hold" + $res.Reason = "$($res.Reason) | no headroom (Reserved 총량 소진)" + continue + } + $newCap = $s.Current + $inc + $durations[$s.Region] = Invoke-Apply -St $s -Res $res -NewCap $newCap + if (-not $res.Error) { + $free -= ($res.Target - $s.Current) + if ($res.Target -lt $planned) { $res.Reason = "$($res.Reason) | headroom-limited to $($res.Target)" } + } +} + +# --- Phase 4: 이력 기록 ------------------------------------------------------ +$results_ordered = @() +foreach ($s in $states) { + $res = $results[$s.Region] + Write-Decision -Now $now -Res $res + Write-Action -Now $now -Res $res -DurationMs $durations[$s.Region] + $results_ordered += $res +} +$results = $results_ordered + +Write-Output "=== Adaptive TPM Quota Controller (PowerShell, dryRun=$DryRun) ===" +$totalAfter = 0 +foreach ($r in $results) { + $totalAfter += $r.Target + $line = "[$($r.Region)] consumedTpm=$($r.AvgConsumedTpm) capacity=$($r.CurrentCapacity)->$($r.Target) " + + "util=$([math]::Round($r.Utilization*100,2))% decision=$($r.Decision) contention=$($r.Contention) " + + "cooldown=$($r.CooldownBlocked) executed=$($r.Executed) reason='$($r.Reason)'" + if ($r.Error) { $line += " error='$($r.Error)'" } + Write-Output $line +} +Write-Output "total capacity after=$totalAfter/$TotalCapacityUnits (Reserved cap)" diff --git a/ptu_lb/runbooks/requirements.txt b/ptu_lb/runbooks/requirements.txt new file mode 100644 index 0000000..db746c4 --- /dev/null +++ b/ptu_lb/runbooks/requirements.txt @@ -0,0 +1,3 @@ +azure-identity==1.19.0 +azure-data-tables==12.5.0 +azure-mgmt-cognitiveservices==13.5.0