Skip to content

feat(monitoring): health check, alerts, backend-requested logging; v0.1.0 verification - #28

Merged
ghchinoy merged 1 commit into
mainfrom
feat/monitoring-alerts-probe
Sep 28, 2026
Merged

ghchinoy merged 1 commit into
mainfrom
feat/monitoring-alerts-probe

Conversation

@ghchinoy

Copy link
Copy Markdown
Owner

Monitoring

  • scripts/probe.py: health check per target (API contract cases with expected
    statuses, latency sample, gateway decisions must be answered by Vertex); one
    JSON log line per target (probe_event=dgem.probe). Works with ADC locally and
    the metadata server on Cloud Run; IAP gateways via PROBE_GATEWAY_AUDIENCE.
  • scripts/deploy_probe.sh + deploy/probe/Dockerfile: least-privilege service
    account, two Cloud Run jobs (hourly Vertex + gateway, daily Cloud Run) and
    Cloud Scheduler triggers.
  • scripts/setup_alerts.py: log-based metrics (decisions by backend, decision
    wall time, probe results and latency), email channel, seven 'dgem:' alert
    policies (PromQL), every threshold overridable; idempotent, --dry-run.
  • docs/operate/monitoring.md: what each alert means and what to do, how to change
    thresholds, deploy and remove; linked from the runbook and journey.

Gateway logging

  • dgem_backend_requested logged next to dgem_backend so a vertex_first failover
    is distinguishable from an explicit Cloud Run request; failed requests log
    backend 'none' instead of 'cloudrun'.

Verification (benchmarks/runs/20260928-v010-verification)

  • All sample suites against production v0.1.0 pass; two pre-existing harness
    issues documented (bench-intents banking77 > 26 options; harness-only
    templates). Rollback targets exercised, including a traffic drill on the
    redeployed Vertex rollback model. scripts/template_sweep.py added.

CHANGELOG v0.1.2 (unreleased); deployment log updated.

Already applied to the hosting project (alerts, metrics, probe jobs). The logging change needs a gateway redeploy (v0.1.2) for the failover and p95 alerts to see data.

…uested logging; v0.1.0 verification run

Monitoring
- scripts/probe.py: health check per target (API contract cases with expected
  statuses, latency sample, gateway decisions must be answered by Vertex); one
  JSON log line per target (probe_event=dgem.probe). Works with ADC locally and
  the metadata server on Cloud Run; IAP gateways via PROBE_GATEWAY_AUDIENCE.
- scripts/deploy_probe.sh + deploy/probe/Dockerfile: least-privilege service
  account, two Cloud Run jobs (hourly Vertex + gateway, daily Cloud Run) and
  Cloud Scheduler triggers.
- scripts/setup_alerts.py: log-based metrics (decisions by backend, decision
  wall time, probe results and latency), email channel, seven 'dgem:' alert
  policies (PromQL), every threshold overridable; idempotent, --dry-run.
- docs/operate/monitoring.md: what each alert means and what to do, how to change
  thresholds, deploy and remove; linked from the runbook and journey.

Gateway logging
- dgem_backend_requested logged next to dgem_backend so a vertex_first failover
  is distinguishable from an explicit Cloud Run request; failed requests log
  backend 'none' instead of 'cloudrun'.

Verification (benchmarks/runs/20260928-v010-verification)
- All sample suites against production v0.1.0 pass; two pre-existing harness
  issues documented (bench-intents banking77 > 26 options; harness-only
  templates). Rollback targets exercised, including a traffic drill on the
  redeployed Vertex rollback model. scripts/template_sweep.py added.

CHANGELOG v0.1.2 (unreleased); deployment log updated.
@ghchinoy
ghchinoy merged commit 7110342 into main Sep 28, 2026
3 checks passed
@ghchinoy
ghchinoy deleted the feat/monitoring-alerts-probe branch October 3, 2026 19:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant