Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
69 changes: 69 additions & 0 deletions .github/workflows/postgres-profile.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
name: postgres-profile

on:
pull_request:
paths:
- "adapters/postgres/**"
- "evaluation/postgres_profile.py"
- "tests/test_postgres_profile.py"
- ".github/workflows/postgres-profile.yml"
push:
branches:
- main
paths:
- "adapters/postgres/**"
- "evaluation/postgres_profile.py"
- "tests/test_postgres_profile.py"
- ".github/workflows/postgres-profile.yml"
workflow_dispatch:

permissions:
contents: read

jobs:
profile:
runs-on: ubuntu-latest
timeout-minutes: 15
services:
postgres:
image: postgres:16-alpine
env:
POSTGRES_USER: postgres
POSTGRES_PASSWORD: postgres
POSTGRES_DB: fabops_bench
ports:
- 5432:5432
options: >-
--health-cmd "pg_isready -U postgres -d fabops_bench"
--health-interval 5s
--health-timeout 5s
--health-retries 12

steps:
- name: Checkout
uses: actions/checkout@v4

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"

- name: Install uv
run: python -m pip install --no-cache-dir uv

- name: Install locked dependencies
run: uv sync --locked --dev

- name: Lint and unit tests
run: |
uv run ruff check evaluation/postgres_profile.py tests/test_postgres_profile.py
uv run pytest -q tests/test_postgres_profile.py

- name: Verify PostgreSQL plans and latency direction
run: >-
uv run python -m evaluation.postgres_profile
--dsn postgresql://postgres:postgres@127.0.0.1:5432/fabops_bench
--event-rows 50000
--case-rows 20000
--repeats 3
--strict
8 changes: 8 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -380,6 +380,14 @@ These numbers are **not** PostgreSQL/Neo4j/Redpanda production capacity claims.

이 수치는 PostgreSQL/Neo4j/Redpanda의 **production capacity 주장**이 아닙니다. 자세한 정의와 제한은 `docs/operations/SLO.md`에 있습니다.

### PostgreSQL query-plan evidence / PostgreSQL 실행계획 근거

Two real repository query shapes were profiled against an isolated PostgreSQL 16 benchmark database using `EXPLAIN (ANALYZE, BUFFERS)` over 200,000 event rows and 80,000 case rows. A composite `(event_type, sequence DESC)` index reduced the recent-measurement query p50 from **10.230 ms to 0.033 ms** and buffer hits from **12,078 to 14**. A `(classification, lot_id DESC, updated_at DESC)` index reduced the related-case query p50 from **0.162 ms to 0.018 ms**, while removing the previous incremental-sort step.

실제 repository query shape 두 개를 격리 PostgreSQL 16 환경에서 20만 event / 8만 case fixture와 `EXPLAIN (ANALYZE, BUFFERS)`로 측정했습니다. `(event_type, sequence DESC)` composite index는 recent-measurement query p50을 **10.230 ms → 0.033 ms**, buffer hit을 **12,078 → 14**로 줄였고, `(classification, lot_id DESC, updated_at DESC)` index는 related-case query p50을 **0.162 ms → 0.018 ms**로 줄이면서 기존 incremental-sort 단계를 제거했습니다.

Full experiment, plan names, variance and limitations: `docs/POSTGRESQL_PROFILING.md` and `evidence/postgres/operational-index-profile.json`.

### Executed incident exercise / 실제 장애 훈련

The M6 Neo4j dependency outage exercise actually stopped the project Neo4j container, observed degraded readiness, restarted it, and observed recovery:
Expand Down
12 changes: 12 additions & 0 deletions adapters/postgres/migrations/008_operational_query_indexes.sql
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
-- Indexes selected from measured repository query shapes.
-- Keep these narrow: JSONB payload columns stay in the heap to avoid bloating
-- the write path for the authoritative event/case tables.
BEGIN;

CREATE INDEX IF NOT EXISTS fabops_event_log_event_type_sequence_idx
ON fabops_event_log(event_type, sequence DESC);

CREATE INDEX IF NOT EXISTS fabops_cases_classification_lot_updated_idx
ON fabops_cases(classification, lot_id DESC, updated_at DESC);

COMMIT;
62 changes: 62 additions & 0 deletions docs/POSTGRESQL_PROFILING.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
# PostgreSQL operational query profiling

## Problem

Two repository reads are order-sensitive and can become expensive as the event/case tables grow:

- `recent_measurement_events()` filters one event type and asks for the newest sequences.
- `related_cases()` filters one classification and orders by lot/update recency.

The existing single-column indexes did not fully match those filter + order shapes.

## Measurement

`evaluation/postgres_profile.py` creates an isolated benchmark database, applies the real FabOps migrations, seeds **200,000 events** and **80,000 cases**, and runs the repository-equivalent SQL with `EXPLAIN (ANALYZE, BUFFERS, FORMAT JSON)`.

Each before/after phase is repeated 15 times. The profiler refuses to seed or drop indexes unless the database name contains `bench` or `test`.

Command used for the committed result:

```bash
uv run python -m evaluation.postgres_profile \
--dsn postgresql://postgres:postgres@127.0.0.1:55432/fabops_bench \
--event-rows 200000 \
--case-rows 80000 \
--repeats 15
```

The DSN above is a disposable local benchmark database, not a production connection string.

## Change

Migration `008_operational_query_indexes.sql` adds only the two indexes that match measured repository access patterns:

```sql
CREATE INDEX fabops_event_log_event_type_sequence_idx
ON fabops_event_log(event_type, sequence DESC);

CREATE INDEX fabops_cases_classification_lot_updated_idx
ON fabops_cases(classification, lot_id DESC, updated_at DESC);
```

No GIN/GiST/partitioning feature was added because these reads do not justify them.

## Result

| Query | Before | After | p50 | p95 | Buffer hits |
|---|---|---|---:|---:|---:|
| recent measurements | `fabops_event_log_pkey` | `fabops_event_log_event_type_sequence_idx` | **10.230 → 0.033 ms** | **10.627 → 0.0413 ms** | **12,078 → 14** |
| related cases | `idx_fabops_cases_lot_id` + incremental sort | `fabops_cases_classification_lot_updated_idx` | **0.162 → 0.018 ms** | **0.1799 → 0.0324 ms** | **2,844 → 46** |

Measured p50 reductions were **99.68%** and **88.89%** respectively on this fixture.

The complete machine-readable result is committed at `evidence/postgres/operational-index-profile.json`.

The CI regression gate runs a smaller isolated fixture with `--strict`; it fails if the intended index is not selected, p50 does not improve, or shared buffer hits do not decrease. The smaller CI fixture is a regression direction check, not the source of the README benchmark numbers above.

## Limitation

- This is a synthetic portfolio fixture using the real schema/query shape, not a production-fab capacity benchmark.
- Measurements are warm-cache runs on one local PostgreSQL 16 instance.
- The experiment isolates read plans; it does not quantify index write amplification, vacuum behavior, or production concurrency.
- The profiler records latency variance, but it is not a substitute for workload-level contention testing.
Loading
Loading