Repository navigation
feat(ptu_lb): Adaptive TPM Quota Controller - Reserved-total 공유 캡 재분배 - #31
Conversation
- 다중 리전(KR/UK) Foundry 배포 용량을 고정 예약 총량 내에서 재분배하는 컨트롤러 추가 - Python 런북(adaptive_tpm_quota_controller.py) 실제 actuation + --dry-run 지원 - PowerShell 옵션 런북, requirements.txt - 아키텍처 다이어그램(excalidraw/png), 테스트 문서(ptu-lb-test.md), README
|
Warning Review limit reached
Next review available in: 6 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (3)
📝 WalkthroughWalkthroughThis PR adds a new ChangesDocumentation and Architecture Diagram
Adaptive TPM Quota Controller Runbooks
Estimated code review effort: 4 (Complex) | ~60 minutes Sequence Diagram(s)sequenceDiagram
participant Scheduler as Automation Schedule
participant Controller as Adaptive Controller
participant Table as Azure Table Storage
participant CogSvc as Cognitive Services Mgmt
Scheduler->>Controller: trigger run_cycle
Controller->>Table: query TrafficWindow (per region)
Table-->>Controller: recent traffic rows
Controller->>CogSvc: get current deployment capacity
CogSvc-->>Controller: current capacity
Controller->>Table: query ScaleAction (cooldown check)
Table-->>Controller: last change timestamp
Controller->>Controller: compute targets, contention, plan (down/up/hold)
Controller->>CogSvc: set_deployment_capacity (reductions first)
Controller->>CogSvc: set_deployment_capacity (increases, within headroom)
Controller->>Table: write ControlDecision
Controller->>Table: write ScaleAction
Controller-->>Scheduler: summary + total capacity
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 5
🧹 Nitpick comments (4)
ptu_lb/ptu_architecture.excalidraw (1)
32-32: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winRename the diagram to match the shared-quota design.
리전 독립 오토스케일conflicts with the controller described in the docs, which rebalances a shared reserved total across KR/UK. Please rename the title and the Foundry label so the diagram matches the actual control model.♻️ Suggested text update
- PTU TPM Quota Adaptive Controller — 리전 독립 오토스케일 + PTU TPM Quota Adaptive Controller — 공유 Reserved quota 재분배 ... - 모델 배포 (리전 독립 quota) + 모델 배포 (공유 Reserved quota)Also applies to: 131-131
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@ptu_lb/ptu_architecture.excalidraw` at line 32, The diagram title text still says “리전 독립 오토스케일,” which conflicts with the shared-quota control model used in the docs. Update the title text in ptu_architecture.excalidraw and the corresponding Foundry label text (the matching text entry referenced in the diff) so both clearly describe the shared reserved total rebalance across KR/UK.ptu_lb/runbooks/adaptive_tpm_quota_controller.py (1)
178-193: 🚀 Performance & Scalability | 🔵 Trivial | ⚡ Quick winCooldown query is unbounded and cross-partition.
Rows are written to
ScaleActionwith aPartitionKeyof the date (_date_pk), but the cooldown lookup filters only onregion eq@r``. As history accumulates this scans every action row across all date partitions on each cycle, and materializes them into a list before filtering. Consider constraining the scan (e.g., add a recent date-partition bound or a server-sideTimestamp/`changed` filter) and TTL/prune the history tables.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@ptu_lb/runbooks/adaptive_tpm_quota_controller.py` around lines 178 - 193, The cooldown lookup in the ScaleAction query is unbounded and scans across all date partitions because it only filters by region, then materializes every row before filtering changed records. Update the cooldown check in the adaptive_tpm_quota_controller logic to add a tighter server-side bound, such as a recent date-partition constraint or Timestamp/changed filter, and avoid loading the full result set into memory. Also review the ScaleAction history retention strategy so old rows are pruned or expire via TTL.ptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1 (2)
73-73: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win
[int]cast uses banker's rounding, diverging from Python's floor division.Python computes
TOTAL_CAPACITY_UNITS = RESERVED_TOTAL_TPM // TPM_PER_CAPACITY_UNIT(floor). PowerShell's[int]cast rounds half-to-even, so for a non-multipleReservedTotalTpm(e.g.21500) this yields22vs Python's21, breaking the stated equivalence and the reserved-total invariant. Defaults (20000/1000=20) hide it. Use[math]::Floorto match.♻️ Proposed fix
-$TotalCapacityUnits = [int]($ReservedTotalTpm / $TpmPerCapacityUnit) # =20 +$TotalCapacityUnits = [int][math]::Floor($ReservedTotalTpm / $TpmPerCapacityUnit) # =20🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@ptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1` at line 73, The $TotalCapacityUnits calculation in optional_adaptive_tpm_quota_controller.ps1 uses an [int] cast, which can round differently from Python’s floor division and break the intended equivalence with TOTAL_CAPACITY_UNITS. Update the computation in the capacity unit calculation block to use floor semantics instead of casting so it matches the Python logic for ReservedTotalTpm and TpmPerCapacityUnit, especially when the total is not an exact multiple.Source: Linters/SAST tools
100-105: 🚀 Performance & Scalability | 🔵 Trivial | 💤 Low valueOptional: cache the storage context / table clients instead of recreating per call.
Get-CloudTablerunsNew-AzStorageContext+Get-AzStorageTableon every invocation, and it is called for each region inGet-RecentWindows,Test-InCooldown,Write-Decision, andWrite-Action. Caching the context and per-nameCloudTable(e.g. in a script-scoped hashtable) removes repeated resolution round-trips.Also applies to: 116-181
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@ptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1` around lines 100 - 105, `Get-CloudTable` is recreating the storage context and resolving the table on every call, which is repeated by `Get-RecentWindows`, `Test-InCooldown`, `Write-Decision`, and `Write-Action`. Update the shared helper to cache the `New-AzStorageContext` result and memoize `Get-AzStorageTable` lookups by table name in a script-scoped hashtable, then have those callers reuse the cached `CloudTable` objects instead of rebuilding them each time.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@ptu_lb/render_excalidraw.py`:
- Around line 23-106: The CLI entry point in main() is called directly with
sys.argv[1] and sys.argv[2], which can fail with IndexError when the script is
launched without both paths. Add an argument count check in the __main__ block
before calling main(), and print a short usage message or exit cleanly when the
required src and dst arguments are missing.
In `@ptu_lb/runbooks/adaptive_tpm_quota_controller.py`:
- Around line 69-71: Hardcoded Azure account identifiers are committed in the
adaptive TPM quota controller setup and should be externalized. Update the
module-level configuration in adaptive_tpm_quota_controller.py so
SUBSCRIPTION_ID, RESOURCE_GROUP, and STORAGE_ACCOUNT are read from environment
or Automation variables instead of fixed literals, and keep the runbook
entrypoint using those symbols to resolve values at runtime. Align the existing
configuration loading with the docstring guidance already present near the top
of the file.
- Around line 165-168: The capacity update flow can hang because
`poller.result()` waits indefinitely after
`mgmt.deployments.begin_create_or_update(...)`; update the `poller.result` call
in the deployment path to use a bounded timeout so stalled ARM operations fail
fast. Keep the change localized around the `poller` returned from
`begin_create_or_update`, and let `_apply` receive and surface any timeout or
operation failure instead of blocking the runbook mid-cycle.
In `@ptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1`:
- Around line 1-41: The script contains persisted Korean literals in the
adaptive TPM quota controller flow that can be misread under Windows PowerShell
5.1 when the file is saved without a UTF-8 BOM. Update the
optional_adaptive_tpm_quota_controller.ps1 script encoding to UTF-8 with BOM so
the `reason` strings used by `ControlDecision` and `ScaleAction` (including the
`"... | no headroom (Reserved 총량 소진)"` branch) are preserved correctly when
written to stdout and Table storage.
In `@ptu_lb/runbooks/requirements.txt`:
- Around line 1-3: Add an explicit pyjwt dependency floor in requirements.txt
alongside azure-identity so the resolver cannot pick an affected transitive
version; update the dependency list to include pyjwt>=2.13.0 near
azure-identity, azure-data-tables, and azure-mgmt-cognitiveservices. Keep the
change minimal and ensure the pinned floor is visible in this requirements file
rather than relying on msal’s transitive PyJWT constraint.
---
Nitpick comments:
In `@ptu_lb/ptu_architecture.excalidraw`:
- Line 32: The diagram title text still says “리전 독립 오토스케일,” which conflicts with
the shared-quota control model used in the docs. Update the title text in
ptu_architecture.excalidraw and the corresponding Foundry label text (the
matching text entry referenced in the diff) so both clearly describe the shared
reserved total rebalance across KR/UK.
In `@ptu_lb/runbooks/adaptive_tpm_quota_controller.py`:
- Around line 178-193: The cooldown lookup in the ScaleAction query is unbounded
and scans across all date partitions because it only filters by region, then
materializes every row before filtering changed records. Update the cooldown
check in the adaptive_tpm_quota_controller logic to add a tighter server-side
bound, such as a recent date-partition constraint or Timestamp/changed filter,
and avoid loading the full result set into memory. Also review the ScaleAction
history retention strategy so old rows are pruned or expire via TTL.
In `@ptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1`:
- Line 73: The $TotalCapacityUnits calculation in
optional_adaptive_tpm_quota_controller.ps1 uses an [int] cast, which can round
differently from Python’s floor division and break the intended equivalence with
TOTAL_CAPACITY_UNITS. Update the computation in the capacity unit calculation
block to use floor semantics instead of casting so it matches the Python logic
for ReservedTotalTpm and TpmPerCapacityUnit, especially when the total is not an
exact multiple.
- Around line 100-105: `Get-CloudTable` is recreating the storage context and
resolving the table on every call, which is repeated by `Get-RecentWindows`,
`Test-InCooldown`, `Write-Decision`, and `Write-Action`. Update the shared
helper to cache the `New-AzStorageContext` result and memoize
`Get-AzStorageTable` lookups by table name in a script-scoped hashtable, then
have those callers reuse the cached `CloudTable` objects instead of rebuilding
them each time.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: f03cbd19-ed2e-43a2-a9ac-3c1a32cdf4db
⛔ Files ignored due to path filters (1)
ptu_lb/ptu_architecture.pngis excluded by!**/*.png
📒 Files selected for processing (7)
ptu_lb/README.mdptu_lb/ptu-lb-test.mdptu_lb/ptu_architecture.excalidrawptu_lb/render_excalidraw.pyptu_lb/runbooks/adaptive_tpm_quota_controller.pyptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1ptu_lb/runbooks/requirements.txt
| def main(src, dst): | ||
| with open(src, encoding="utf-8") as f: | ||
| data = json.load(f) | ||
| elements = [e for e in data["elements"] if not e.get("isDeleted")] | ||
|
|
||
| # bounds | ||
| max_x = max_y = 0 | ||
| for e in elements: | ||
| ex = e["x"] + (e.get("width") or 0) | ||
| ey = e["y"] + (e.get("height") or 0) | ||
| if e["type"] == "arrow": | ||
| for px, py in e["points"]: | ||
| ex = max(ex, e["x"] + px) | ||
| ey = max(ey, e["y"] + py) | ||
| max_x = max(max_x, ex) | ||
| max_y = max(max_y, ey) | ||
| W = int((max_x + 40) * SCALE) | ||
| H = int((max_y + 40) * SCALE) | ||
|
|
||
| img = Image.new("RGB", (W, H), "#ffffff") | ||
| d = ImageDraw.Draw(img) | ||
|
|
||
| def sc(v): | ||
| return int(v * SCALE) | ||
|
|
||
| # rectangles first | ||
| for e in elements: | ||
| if e["type"] != "rectangle": | ||
| continue | ||
| x0, y0 = sc(e["x"]), sc(e["y"]) | ||
| x1, y1 = sc(e["x"] + e["width"]), sc(e["y"] + e["height"]) | ||
| fill = e.get("backgroundColor") | ||
| if fill == "transparent": | ||
| fill = None | ||
| r = 10 * SCALE | ||
| d.rounded_rectangle([x0, y0, x1, y1], radius=r, fill=fill, | ||
| outline=e.get("strokeColor", "#1e1e1e"), | ||
| width=int(max(1, e.get("strokeWidth", 2)) * SCALE)) | ||
|
|
||
| # arrows | ||
| def draw_arrow(pts, color, width): | ||
| for i in range(len(pts) - 1): | ||
| d.line([pts[i], pts[i + 1]], fill=color, width=width) | ||
| # arrowhead on last segment | ||
| (x0, y0), (x1, y1) = pts[-2], pts[-1] | ||
| import math | ||
| ang = math.atan2(y1 - y0, x1 - x0) | ||
| L = 12 * SCALE | ||
| for da in (math.radians(150), math.radians(-150)): | ||
| hx = x1 + L * math.cos(ang + da) | ||
| hy = y1 + L * math.sin(ang + da) | ||
| d.line([(x1, y1), (hx, hy)], fill=color, width=width) | ||
|
|
||
| for e in elements: | ||
| if e["type"] != "arrow": | ||
| continue | ||
| pts = [(sc(e["x"] + px), sc(e["y"] + py)) for px, py in e["points"]] | ||
| draw_arrow(pts, e.get("strokeColor", "#1e1e1e"), | ||
| int(max(1, e.get("strokeWidth", 2)) * SCALE)) | ||
|
|
||
| # text on top | ||
| for e in elements: | ||
| if e["type"] != "text": | ||
| continue | ||
| font = load_font(int(e.get("fontSize", 18) * SCALE)) | ||
| color = e.get("strokeColor", "#1e1e1e") | ||
| lines = e["text"].split("\n") | ||
| lh = int(e.get("fontSize", 18) * SCALE * e.get("lineHeight", 1.2)) | ||
| align = e.get("textAlign", "left") | ||
| box_w = sc(e.get("width") or 0) | ||
| for i, line in enumerate(lines): | ||
| tw = d.textlength(line, font=font) | ||
| tx = sc(e["x"]) | ||
| if align == "center": | ||
| tx = sc(e["x"]) + (box_w - tw) / 2 | ||
| ty = sc(e["y"]) + i * lh | ||
| d.text((tx, ty), line, fill=color, font=font) | ||
|
|
||
| img.save(dst) | ||
| print(f"saved {dst} ({img.width}x{img.height})") | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| main(sys.argv[1], sys.argv[2]) |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
sed -n '1,220p' ptu_lb/render_excalidraw.pyRepository: hellices/devguidesample
Length of output: 3672
🏁 Script executed:
rg -n "render_excalidraw\.py|excalidraw" -S .Repository: hellices/devguidesample
Length of output: 2207
🏁 Script executed:
sed -n '240,280p' ptu_lb/ptu-lb-test.mdRepository: hellices/devguidesample
Length of output: 2617
Guard the CLI arguments. sys.argv[1] / sys.argv[2] will raise IndexError if the script is run without both paths; add a small usage check before calling main().
🧰 Tools
🪛 ast-grep (0.44.0)
[warning] 23-23: File path is request-/variable-derived; validate and normalize to prevent path traversal.
Context: open(src, encoding="utf-8")
Note: [CWE-22] Improper Limitation of a Pathname to a Restricted Directory ('Path Traversal').
(open-filename-from-request)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@ptu_lb/render_excalidraw.py` around lines 23 - 106, The CLI entry point in
main() is called directly with sys.argv[1] and sys.argv[2], which can fail with
IndexError when the script is launched without both paths. Add an argument count
check in the __main__ block before calling main(), and print a short usage
message or exit cleanly when the required src and dst arguments are missing.
Source: Linters/SAST tools
| SUBSCRIPTION_ID = "2cf925b6-80cb-4567-abda-5ccd3010aab5" | ||
| RESOURCE_GROUP = "ptu-lb" | ||
| STORAGE_ACCOUNT = "ptulb59018sa" |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win
Externalize account identifiers instead of hardcoding them.
The SUBSCRIPTION_ID, RESOURCE_GROUP, and STORAGE_ACCOUNT are committed into source (the docstring at Line 66 already recommends externalizing). Prefer reading from environment variables / Automation variables so the same runbook works across environments and infra identifiers aren't baked into VCS history.
🔧 Suggested change
-SUBSCRIPTION_ID = "2cf925b6-80cb-4567-abda-5ccd3010aab5"
-RESOURCE_GROUP = "ptu-lb"
-STORAGE_ACCOUNT = "ptulb59018sa"
+import os
+SUBSCRIPTION_ID = os.environ["PTU_SUBSCRIPTION_ID"]
+RESOURCE_GROUP = os.environ.get("PTU_RESOURCE_GROUP", "ptu-lb")
+STORAGE_ACCOUNT = os.environ["PTU_STORAGE_ACCOUNT"]📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| SUBSCRIPTION_ID = "2cf925b6-80cb-4567-abda-5ccd3010aab5" | |
| RESOURCE_GROUP = "ptu-lb" | |
| STORAGE_ACCOUNT = "ptulb59018sa" | |
| import os | |
| SUBSCRIPTION_ID = os.environ["PTU_SUBSCRIPTION_ID"] | |
| RESOURCE_GROUP = os.environ.get("PTU_RESOURCE_GROUP", "ptu-lb") | |
| STORAGE_ACCOUNT = os.environ["PTU_STORAGE_ACCOUNT"] |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@ptu_lb/runbooks/adaptive_tpm_quota_controller.py` around lines 69 - 71,
Hardcoded Azure account identifiers are committed in the adaptive TPM quota
controller setup and should be externalized. Update the module-level
configuration in adaptive_tpm_quota_controller.py so SUBSCRIPTION_ID,
RESOURCE_GROUP, and STORAGE_ACCOUNT are read from environment or Automation
variables instead of fixed literals, and keep the runbook entrypoint using those
symbols to resolve values at runtime. Align the existing configuration loading
with the docstring guidance already present near the top of the file.
| poller = mgmt.deployments.begin_create_or_update( | ||
| RESOURCE_GROUP, account, deployment, dep | ||
| ) | ||
| poller.result() |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Map the target file and inspect the relevant region.
git ls-files ptu_lb/runbooks/adaptive_tpm_quota_controller.py
wc -l ptu_lb/runbooks/adaptive_tpm_quota_controller.py
sed -n '130,210p' ptu_lb/runbooks/adaptive_tpm_quota_controller.py
# Find other poller.result usage patterns in the repo.
rg -n "poller\.result\(" -S .Repository: hellices/devguidesample
Length of output: 3137
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Inspect the target file around the commented lines.
nl -ba ptu_lb/runbooks/adaptive_tpm_quota_controller.py | sed -n '145,185p'
# Inspect nearby helper logic that might already handle timeouts or poller errors.
rg -n "def _apply|begin_create_or_update|poller\.result|timeout=" ptu_lb/runbooks/adaptive_tpm_quota_controller.py -n -S
# Check Azure LROPoller result semantics from installed docs/source if present.
python3 - <<'PY'
import inspect
mods = [
("azure.core.polling", "LROPoller"),
("azure.mgmt.resource.resources.v2022_09_01.operations._deployments_operations", "DeploymentsOperations"),
]
for mod_name, attr in mods:
try:
mod = __import__(mod_name, fromlist=[attr])
obj = getattr(mod, attr)
print(f"## {mod_name}.{attr}")
print(inspect.signature(obj.result))
except Exception as e:
print(f"## {mod_name}.{attr} unavailable: {e}")
PYRepository: hellices/devguidesample
Length of output: 201
🌐 Web query:
Azure Python SDK LROPoller.result timeout parameter behavior begin_create_or_update timeout
💡 Result:
In the Azure Python SDK, the LROPoller.result(timeout=...) method and the begin_create_or_update method handle timeouts in distinct ways. 1. LROPoller.result(timeout=...) Behavior The timeout parameter in LROPoller.result() specifies a duration in seconds to wait for the long-running operation to complete before the method returns control to your code [1][2][3]. - If the operation completes within the specified timeout, result() returns the final deserialized resource [1][2]. - If the operation does not complete within the timeout, the method returns control, but it does not automatically cancel the underlying cloud operation [4][1]. You can verify whether the operation has finished by calling the done() method on the poller object [4][3]. 2. begin_create_or_update Timeout The begin_create_or_update method itself typically does not accept a timeout parameter to control the duration of the long-running operation. Instead, it initiates the request and returns an LROPoller instance [5]. If you need to influence the timing of the polling process, you generally have two approaches: - Poller Timeout: As noted above, you can use .result(timeout=...) on the returned poller to limit how long your local code waits for the final result [1][2]. - Client-Level Configuration: To adjust the polling interval (the frequency at which the SDK checks the status of the operation), you may need to configure the client's long-running operation settings or, in some cases, provide a custom polling strategy [6][7]. Important Note: The timeout parameter in result() is distinct from connection or HTTP request timeouts, which are generally managed via the client's configuration or transport layer settings [8][9]. If an operation is still running after the result() timeout, the cloud resource creation or update continues in the background on the Azure server [4][1].
Citations:
- 1: https://azuresdkdocs.z19.web.core.windows.net/python/azure-core/latest/azure.core.polling.html
- 2: https://github.com/MicrosoftDocs/azure-docs-sdk-python/blob/master/docs-ref-autogen/azure-core/azure.core.polling.LROPoller.yml
- 3: https://azuresdkdocs.blob.core.windows.net/$web/python/azure-core/1.0.0/azure.core.polling.html
- 4: https://learn.microsoft.com/en-us/python/api/azure-core/azure.core.polling.lropoller?view=azure-python
- 5: https://github.com/Azure/azure-sdk-for-python/wiki/Common-issues-about-Python-SDK
- 6: Doc of how to adjust op status polling interval Azure/azure-sdk-for-python#1120
- 7: Are there default values of "timeout" and "retries"? Azure/azure-sdk-for-python#4599
- 8: https://azure.github.io/azure-sdk/python_design.html
- 9: Connection timeout doesn't work. Documentation is unclear in regards to how to use that. Azure/azure-sdk-for-python#39580
🏁 Script executed:
#!/bin/bash
set -euo pipefail
python3 - <<'PY'
from pathlib import Path
path = Path("ptu_lb/runbooks/adaptive_tpm_quota_controller.py")
text = path.read_text()
for i, line in enumerate(text.splitlines(), 1):
if 145 <= i <= 185:
print(f"{i:4d}: {line}")
PYRepository: hellices/devguidesample
Length of output: 2034
Add a timeout to the poller result call. poller.result() can block the runbook indefinitely if the ARM operation stalls; a bounded wait keeps the capacity update from hanging mid-cycle and lets _apply surface the failure.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@ptu_lb/runbooks/adaptive_tpm_quota_controller.py` around lines 165 - 168, The
capacity update flow can hang because `poller.result()` waits indefinitely after
`mgmt.deployments.begin_create_or_update(...)`; update the `poller.result` call
in the deployment path to use a bounded timeout so stalled ARM operations fail
fast. Keep the change localized around the `poller` returned from
`begin_create_or_update`, and let `_apply` receive and surface any timeout or
operation failure instead of blocking the runbook mid-cycle.
| <# | ||
| .SYNOPSIS | ||
| Adaptive TPM Quota Controller (PowerShell / Azure Automation) — OPTIONAL 대체 구현 | ||
|
|
||
| .DESCRIPTION | ||
| Python 런북(adaptive_tpm_quota_controller.py)과 **동일한 제어 로직**을 PowerShell(Az 모듈)로 | ||
| 이식한 "선택적(optional) 대체 구현"이다. 기본 구현은 Python 런북이며, 본 스크립트는 | ||
| PowerShell 선호 환경을 위한 동등 참조본이다. | ||
|
|
||
| [Python 버전과의 동등성 — 핵심 설계 결정] | ||
| * 두 배포 capacity 의 합은 Reservation 한 총량(ReservedTotalTpm)을 어떤 시점에도 | ||
| 초과할 수 없다. 고정된 Reserved 총량을 트래픽 비율에 맞춰 재분배한다: | ||
| - 필요량 합 ≤ 총량 → 각자 need 만 배정(나머지는 미할당 버퍼) | ||
| - 필요량 합 > 총량 → 소비 TPM 비율로 비례 배분(경합) | ||
| * 감축 먼저, 증설 나중: 총량이 꽉 찬 상태에서 증설을 먼저 하면 실패하므로, | ||
| 잉여 리전을 먼저 감축해 headroom 을 확보한 뒤 증설한다(loss 최소화). | ||
| * 리전 최소 보장(MinCapacity)은 Reserved 총량의 25% 로 유지한다. | ||
| * capacity(배포 SKU capacity) 를 조정 knob 으로 직접 사용. 1 unit ≈ TpmPerCapacityUnit. | ||
| * 실제 actuation: Az.CognitiveServices 로 배포 SKU capacity 를 실제 갱신(폐루프). | ||
| * 쿨다운: ScaleAction 이력에서 리전별 마지막 실제 변경 시각을 확인해 재조정 차단. | ||
|
|
||
| .NOTES | ||
| [인증 / 의존 모듈] | ||
| * 인증: 관리 ID -> Connect-AzAccount -Identity (로컬은 Connect-AzAccount 로 로그인) | ||
| * 필요 모듈: Az.Accounts, Az.CognitiveServices, AzTable | ||
| * 필요 RBAC: | ||
| - Storage: "Storage Table Data Contributor" (TrafficWindow 읽기 / 이력 쓰기) | ||
| - Foundry: "Cognitive Services Contributor" (배포 capacity 갱신) | ||
|
|
||
| [네트워크 주의] | ||
| * Storage 가 publicNetworkAccess=Disabled + Private Endpoint 구성인 경우, 이 런북도 | ||
| 프라이빗 경로가 닿는 곳(예: Hybrid Runbook Worker)에서 실행되어야 테이블에 접근된다. | ||
| (Linux Hybrid Worker 에는 PowerShell 이 기본 설치되어 있지 않을 수 있으므로, | ||
| PowerShell 로 운영하려면 Windows Hybrid Worker 를 사용) | ||
|
|
||
| .EXAMPLE | ||
| .\optional_adaptive_tpm_quota_controller.ps1 -DryRun # 계산/기록만 | ||
|
|
||
| .EXAMPLE | ||
| .\optional_adaptive_tpm_quota_controller.ps1 # 실제 조정 | ||
| #> |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Add a UTF-8 BOM — persisted Korean reason strings can be corrupted under Windows PowerShell 5.1.
Line 387 ("... | no headroom (Reserved 총량 소진)") writes non-ASCII text into the reason column of ControlDecision/ScaleAction and to stdout. Since the NOTES recommend a Windows Hybrid Worker (PS 5.1), a BOM-less UTF-8 file is decoded with the system codepage, garbling the Korean literals and the data written to Table storage. Save the file as UTF-8 with BOM.
Also applies to: 387-387
🧰 Tools
🪛 PSScriptAnalyzer (1.25.0)
[warning] Missing BOM encoding for non-ASCII encoded file 'optional_adaptive_tpm_quota_controller.ps1'
(PSUseBOMForUnicodeEncodedFile)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@ptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1` around lines 1 -
41, The script contains persisted Korean literals in the adaptive TPM quota
controller flow that can be misread under Windows PowerShell 5.1 when the file
is saved without a UTF-8 BOM. Update the
optional_adaptive_tpm_quota_controller.ps1 script encoding to UTF-8 with BOM so
the `reason` strings used by `ControlDecision` and `ScaleAction` (including the
`"... | no headroom (Reserved 총량 소진)"` branch) are preserved correctly when
written to stdout and Table storage.
Source: Linters/SAST tools
| azure-identity==1.19.0 | ||
| azure-data-tables==12.5.0 | ||
| azure-mgmt-cognitiveservices==13.5.0 |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
curl -s https://pypi.org/pypi/msal/json | jq -r '.info.requires_dist[]? | select(test("pyjwt";"i"))'
curl -s https://pypi.org/pypi/pyjwt/json | jq -r '.info.version'Repository: hellices/devguidesample
Length of output: 192
🏁 Script executed:
#!/bin/bash
set -euo pipefail
echo '--- ptu_lb/runbooks/requirements.txt ---'
cat -n ptu_lb/runbooks/requirements.txt
echo
echo '--- azure-identity 1.19.0 requires_dist ---'
curl -s https://pypi.org/pypi/azure-identity/1.19.0/json \
| jq -r '.info.requires_dist[]?'
echo
echo '--- msal latest requires_dist (filtered pyjwt-related) ---'
curl -s https://pypi.org/pypi/msal/json \
| jq -r '.info.requires_dist[]? | select(test("pyjwt";"i"))'Repository: hellices/devguidesample
Length of output: 536
Pin pyjwt>=2.13.0 here. azure-identity pulls msal, and msal allows PyJWT[crypto]<3,>=1.0.0, so an explicit floor avoids resolving to an affected transitive release.
🔧 Suggested change
azure-identity==1.19.0
azure-data-tables==12.5.0
azure-mgmt-cognitiveservices==13.5.0
+pyjwt>=2.13.0📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| azure-identity==1.19.0 | |
| azure-data-tables==12.5.0 | |
| azure-mgmt-cognitiveservices==13.5.0 | |
| azure-identity==1.19.0 | |
| azure-data-tables==12.5.0 | |
| azure-mgmt-cognitiveservices==13.5.0 | |
| pyjwt>=2.13.0 |
🧰 Tools
🪛 OSV Scanner (2.4.0)
[HIGH] 1-1: pyjwt 2.9.0: undefined
(PYSEC-2025-183)
[HIGH] 1-1: pyjwt 2.9.0: undefined
(PYSEC-2026-120)
[HIGH] 1-1: pyjwt 2.9.0: undefined
(PYSEC-2026-175)
[HIGH] 1-1: pyjwt 2.9.0: undefined
(PYSEC-2026-176)
[HIGH] 1-1: pyjwt 2.9.0: undefined
(PYSEC-2026-177)
[HIGH] 1-1: pyjwt 2.9.0: undefined
(PYSEC-2026-178)
[HIGH] 1-1: pyjwt 2.9.0: undefined
(PYSEC-2026-179)
[HIGH] 1-1: pyjwt 2.9.0: PyJWT accepts unknown crit header extensions
[HIGH] 1-1: pyjwt 2.9.0: PyJWKClient: missing scheme allowlist enables CVE-2024-21643-class SSRF + token forgery via file://, ftp://, data: schemes
[HIGH] 1-1: pyjwt 2.9.0: PyJWKClient unbounded JWKS endpoint requests via attacker-controlled kid values (DoS)
[HIGH] 1-1: pyjwt 2.9.0: PyJWT: Algorithm allow-list bypass when decoding with PyJWK / PyJWKClient keys
[HIGH] 1-1: pyjwt 2.9.0: PyJWT: Unauthenticated DoS via unbounded Base64URL decoding of unused payload segment in b64=false detached JWS
[HIGH] 1-1: pyjwt 2.9.0: PyJWT: Public-key JWK accepted as HMAC secret enables forged HS256 tokens when mixed families are allowed
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@ptu_lb/runbooks/requirements.txt` around lines 1 - 3, Add an explicit pyjwt
dependency floor in requirements.txt alongside azure-identity so the resolver
cannot pick an affected transitive version; update the dependency list to
include pyjwt>=2.13.0 near azure-identity, azure-data-tables, and
azure-mgmt-cognitiveservices. Keep the change minimal and ensure the pinned
floor is visible in this requirements file rather than relying on msal’s
transitive PyJWT constraint.
Source: Linters/SAST tools
| @@ -0,0 +1,311 @@ | |||
| # Adaptive PTU TPM Quota 조정 가이드 (with Sample Test) | |||
|
|
||
| ### Test-2 재분배 (감축→증설) | ||
|
|
||
| - Korea util 85%+ / UK util 20% 유도 |
There was a problem hiding this comment.
실제로는 100%를 사용해도 fallback llm이 있기 때문에 100% 도달로 판단해도 될듯.
비용 최적화를 위해서 적정 ptu / paygo 를 도출하고 있음.
There was a problem hiding this comment.
Pull request overview
Adds an Adaptive TPM Quota Controller for a 2-region (KR/UK) Azure AI Foundry deployment that rebalances shared Reserved TPM quota without ever exceeding the reserved total, using a “decrease-first then increase” actuation sequence plus cooldown safeguards. This also introduces accompanying documentation and an architecture diagram for operating the controller via Azure Automation + Hybrid Runbook Worker.
Changes:
- Added Python Azure Automation runbook implementing shared Reserved-total capacity rebalancing (with cooldown + decrease-first actuation).
- Added optional equivalent PowerShell (Az modules) runbook implementation.
- Added operational documentation, Excalidraw diagram source + rendered PNG, and a small Excalidraw-to-PNG renderer utility.
Reviewed changes
Copilot reviewed 7 out of 8 changed files in this pull request and generated 11 comments.
Show a summary per file
| File | Description |
|---|---|
| ptu_lb/runbooks/requirements.txt | Pins Azure SDK dependencies for the Python runbook. |
| ptu_lb/runbooks/adaptive_tpm_quota_controller.py | Main controller logic: metrics read → target allocation → decrease-first apply → history writeback. |
| ptu_lb/runbooks/optional_adaptive_tpm_quota_controller.ps1 | Optional PowerShell port of the same control loop using Az + AzTable. |
| ptu_lb/render_excalidraw.py | Utility to render .excalidraw diagrams to PNG for docs. |
| ptu_lb/README.md | Minimal entry-point pointing to the main guide. |
| ptu_lb/ptu-lb-test.md | Full guide: design rationale, runbook workflow, test scenarios, and deployment notes. |
| ptu_lb/ptu_architecture.excalidraw | Architecture diagram source for the PTU LB workflow. |
| ptu_lb/ptu_architecture.png | Rendered architecture diagram image referenced by docs. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| entities = tc.query_entities( | ||
| query_filter="PartitionKey eq @pk", | ||
| parameters={"pk": region}, | ||
| ) |
| def is_in_cooldown(table_svc: TableServiceClient, region: str, now: datetime) -> bool: | ||
| """해당 리전의 마지막 '실제 변경' ScaleAction 이 쿨다운 이내면 True.""" | ||
| tc = table_svc.get_table_client(ACTION_TABLE) | ||
| actions = list( | ||
| tc.query_entities( | ||
| query_filter="region eq @r", | ||
| parameters={"r": region}, | ||
| ) | ||
| ) | ||
| # 실제 capacity 가 바뀐(=changed=True) 액션만 대상으로 최신 것을 찾는다 | ||
| changed = [a for a in actions if str(a.get("changed", "")).lower() == "true"] | ||
| if not changed: | ||
| return False | ||
| changed.sort(key=lambda a: a.metadata["timestamp"], reverse=True) | ||
| last_ts: datetime = changed[0].metadata["timestamp"] | ||
| if last_ts.tzinfo is None: | ||
| last_ts = last_ts.replace(tzinfo=timezone.utc) | ||
| elapsed_min = (now - last_ts).total_seconds() / 60.0 | ||
| return elapsed_min < COOLDOWN_MINUTES |
| [string]$SubscriptionId = "2cf925b6-80cb-4567-abda-5ccd3010aab5", | ||
| [string]$ResourceGroup = "ptu-lb", | ||
| [string]$StorageAccountName = "ptulb59018sa", |
| $table = Get-CloudTable -TableName $ActionTable | ||
| $actions = Get-AzTableRow -Table $table -CustomFilter "region eq '$Region'" | ||
| if (-not $actions) { return $false } |
| "updated": 1762128000000, | ||
| "link": null, | ||
| "locked": false, | ||
| "text": "PTU TPM Quota Adaptive Controller — 리전 독립 오토스케일", |
| "updated": 1762128000000, | ||
| "link": null, | ||
| "locked": false, | ||
| "text": "모델 배포 (리전 독립 quota)", |
| "textAlign": "left", | ||
| "verticalAlign": "top", | ||
| "containerId": null, | ||
| "originalText": "모델 배포 (리전 독립 quota)", |
| import json | ||
| import sys | ||
| from PIL import Image, ImageDraw, ImageFont | ||
|
|
| "C:/Windows/Fonts/malgun.ttf", # Korean support | ||
| "C:/Windows/Fonts/segoeui.ttf", | ||
| "C:/Windows/Fonts/arial.ttf", | ||
| ]: | ||
| try: |
| import argparse | ||
| import math | ||
| import sys | ||
| import uuid | ||
| from dataclasses import dataclass, field | ||
| from datetime import datetime, timezone | ||
|
|
||
| from azure.identity import DefaultAzureCredential | ||
| from azure.data.tables import TableServiceClient | ||
| from azure.mgmt.cognitiveservices import CognitiveServicesManagementClient | ||
|
|
||
|
|
||
| # --------------------------------------------------------------------------- | ||
| # 설정 (운영 시 환경변수/파라미터로 외부화 권장) | ||
| # --------------------------------------------------------------------------- | ||
|
|
||
| SUBSCRIPTION_ID = "2cf925b6-80cb-4567-abda-5ccd3010aab5" | ||
| RESOURCE_GROUP = "ptu-lb" | ||
| STORAGE_ACCOUNT = "ptulb59018sa" |
- region_max/Get-RegionMax 로 리전별 최대 capacity 를 N리전으로 일반화 - 메트릭 열화 시 RI 균등분배 fallback(even_split_targets/Get-EvenSplitTargets) 추가 - PayGo fallback 시 PTU util 100% 근접 허용 및 리전 확대 방안 문서화
| import argparse | ||
| import math | ||
| import sys | ||
| import uuid | ||
| from dataclasses import dataclass, field | ||
| from datetime import datetime, timezone | ||
|
|
||
| from azure.identity import DefaultAzureCredential | ||
| from azure.data.tables import TableServiceClient | ||
| from azure.mgmt.cognitiveservices import CognitiveServicesManagementClient | ||
|
|
||
|
|
||
| # --------------------------------------------------------------------------- | ||
| # 설정 (운영 시 환경변수/파라미터로 외부화 권장) | ||
| # --------------------------------------------------------------------------- | ||
|
|
||
| SUBSCRIPTION_ID = "2cf925b6-80cb-4567-abda-5ccd3010aab5" | ||
| RESOURCE_GROUP = "ptu-lb" | ||
| STORAGE_ACCOUNT = "ptulb59018sa" | ||
|
|
||
| TRAFFIC_TABLE = "TrafficWindow" | ||
| DECISION_TABLE = "ControlDecision" | ||
| ACTION_TABLE = "ScaleAction" | ||
|
|
| n = len(states) | ||
| cap_max = region_max(n) # 리전 수에 따른 상한(2리전이면 TOTAL−MIN 로 기존과 동일) | ||
| needs = {s.region: need_capacity(s.consumed, cap_max) for s in states} | ||
| total_need = sum(needs.values()) | ||
|
|
||
| if total_need <= TOTAL_CAPACITY_UNITS: | ||
| # 여유 있음: 필요량만 배정하고 남는 건 미할당 버퍼로 둔다(증설은 무손실). | ||
| return dict(needs), needs, False | ||
|
|
||
| # 경합: MIN 을 먼저 보장하고 남은 용량을 소비 TPM 비율로 배분한다. | ||
| rem = TOTAL_CAPACITY_UNITS - n * MIN_CAPACITY | ||
| total_consumed = sum(max(s.consumed, 0.0) for s in states) or 1.0 |
| entities = tc.query_entities( | ||
| query_filter="PartitionKey eq @pk", | ||
| parameters={"pk": region}, | ||
| ) | ||
| rows = list(entities) |
| tc = table_svc.get_table_client(ACTION_TABLE) | ||
| actions = list( | ||
| tc.query_entities( | ||
| query_filter="region eq @r", | ||
| parameters={"r": region}, |
| $n = $States.Count | ||
| $capMax = Get-RegionMax -N $n # 리전 수에 따른 상한(2리전이면 TOTAL−MIN 로 기존과 동일) | ||
| $needs = @{} | ||
| foreach ($s in $States) { $needs[$s.Region] = Get-NeedCapacity -ConsumedTpm $s.Consumed -CapMax $capMax } | ||
| $totalNeed = ($needs.Values | Measure-Object -Sum).Sum | ||
|
|
||
| if ($totalNeed -le $TotalCapacityUnits) { | ||
| # 여유 있음: 필요량만 배정, 나머지는 미할당 버퍼(증설은 무손실) | ||
| return @{ Targets = $needs; Contention = $false } | ||
| } | ||
|
|
||
| # 경합: MIN 을 먼저 보장하고 남은 용량을 소비 TPM 비율로 배분 | ||
| $rem = $TotalCapacityUnits - ($n * $MinCapacity) | ||
| $totalConsumed = ($States | ForEach-Object { [math]::Max($_.Consumed, 0.0) } | Measure-Object -Sum).Sum | ||
| if ($totalConsumed -le 0) { $totalConsumed = 1.0 } |
| $table = Get-CloudTable -TableName $ActionTable | ||
| $actions = Get-AzTableRow -Table $table -CustomFilter "region eq '$Region'" | ||
| if (-not $actions) { return $false } | ||
|
|
||
| # 실제 capacity 가 바뀐(changed=true) 액션만 대상으로 최신 것을 찾는다 |
| "text": "PTU TPM Quota Adaptive Controller — 리전 독립 오토스케일", | ||
| "fontSize": 26, | ||
| "fontFamily": 2, | ||
| "textAlign": "left", | ||
| "verticalAlign": "top", | ||
| "containerId": null, | ||
| "originalText": "PTU TPM Quota Adaptive Controller — 리전 독립 오토스케일", |
| "text": "모델 배포 (리전 독립 quota)", | ||
| "fontSize": 13, | ||
| "fontFamily": 2, | ||
| "textAlign": "left", | ||
| "verticalAlign": "top", | ||
| "containerId": null, | ||
| "originalText": "모델 배포 (리전 독립 quota)", |
| def load_font(size): | ||
| for name in [ | ||
| "C:/Windows/Fonts/malgun.ttf", # Korean support | ||
| "C:/Windows/Fonts/segoeui.ttf", | ||
| "C:/Windows/Fonts/arial.ttf", | ||
| ]: |
| - 메인 문서: [ptu-lb-test.md](ptu-lb-test.md) | ||
| - 아키텍처 다이어그램(원본): [ptu_architecture.excalidraw](ptu_architecture.excalidraw) | ||
| - 아키텍처 다이어그램(이미지): [ptu_architecture.png](ptu_architecture.png) | ||
| - 다이어그램 렌더러(유틸): [render_excalidraw.py](render_excalidraw.py) | ||
| - **Runbook(기본, Python): [runbooks/adaptive_tpm_quota_controller.py](runbooks/adaptive_tpm_quota_controller.py)** |
목적
테스트
Summary by CodeRabbit