Skip to content

Latest commit

 

History

History
264 lines (220 loc) · 13.5 KB

File metadata and controls

264 lines (220 loc) · 13.5 KB

Gateway performance benchmark

The benchmark has two modes. The default suite provisions invocation-owned gateways, databases, Redis coordination, and loopback fixtures; --suite=false is a manual mode for an already-running gateway. Neither mode builds or reconfigures an operator service. Every measured comparison uses the same deterministic fixture contract for a direct target and a gateway target.

Reports and paths beginning with .artifacts/ are private operator-local outputs. They are not hosted proof and are not shipped with the repository.

Complete benchmark suite (default)

First complete the source-build sequence, including the generated console assets. Then build the benchmark runner, cache the suite images, and run the complete suite:

mkdir -p .artifacts
go build -trimpath -ldflags='-s -w' -o .artifacts/hoorific ./cmd/hoorific
go build -trimpath -ldflags='-s -w' -o .artifacts/bench-runner ./tools/bench

podman pull docker.io/library/postgres:17
podman pull docker.io/library/redis:7

.artifacts/bench-runner \
  --suite \
  --upstream-transport http \
  --binary .artifacts/hoorific \
  --output .artifacts/bench

The suite is Linux-only and requires /proc, util-linux taskset, local rootless Podman (the suite does not fall back to Docker), and permission to use logical CPU IDs 0,2,4,6. The PostgreSQL 17 and Redis 7 images must already be available locally; the suite never pulls them. A fresh suite-<profile>-* directory is created under the output path, with a suite.json manifest and one child directory per deployment/route pair: standalone/native, standalone/translated, cluster/native, and cluster/translated.

Each child runs the complete nine-workload matrix against both direct fixture and gateway targets, with three measured runs and a separate warmup for every case. A child therefore emits 54 measured records; the four-entry suite emits the corresponding matrix independently for each deployment/route pair. The suite's fixed measurement window is 30 seconds, warmup is 5 seconds, per-request timeout is 15 seconds, and the suite child uses an 85-minute whole-session deadline (the manual-mode default is 2 hours). Use a separate output directory when collecting the tls profile:

.artifacts/bench-runner \
  --suite \
  --upstream-transport tls \
  --binary .artifacts/hoorific \
  --output .artifacts/bench-tls

Manual provisioned-gateway mode

Use manual mode only with a running gateway that is already configured to route the bench-fixture alias to the runner's loopback fixture at http://127.0.0.1:18089/v1. The runner starts that fixture, performs unary and SSE identity preflight against the direct fixture and gateway endpoint, and then measures all nine workloads. It never starts, migrates, or reconfigures the gateway binary.

The manual mode requires --suite=false and a private transport-proof JSON in addition to the URL and key files:

.artifacts/bench-runner \
  --suite=false \
  --upstream-transport http \
  --binary .artifacts/hoorific \
  --output .artifacts/bench-manual-run \
  --url-file /run/user/$UID/hoorific-bench-url \
  --key-file /run/user/$UID/hoorific-bench-key \
  --transport-proof-file /run/user/$UID/hoorific-bench-transport-proof.json \
  --fixture-listen 127.0.0.1:18089 \
  --deployment standalone \
  --route native

The URL file contains the complete absolute loopback HTTP /v1/chat/completions endpoint with no userinfo, query, or fragment. The key file contains only the raw API key (not a Bearer prefix); it must be a regular file no larger than 64 KiB with no group/other permission bits and no embedded CR/LF. Both files should remain private. Use a fresh output directory: the runner writes a permanent profile marker and refuses to reuse a directory containing benchmark artifacts.

The proof file must be a regular private JSON object with exactly these nine fields, and every endpoint must match the running gateway and the --fixture-listen address:

{
  "profile": "http",
  "direct_endpoint": "http://127.0.0.1:<fixture-port>/v1/chat/completions",
  "gateway_endpoint": "http://127.0.0.1:<gateway-port>/v1/chat/completions",
  "gateway_upstream": "http://127.0.0.1:<fixture-port>/v1",
  "fixture_origin": "http://127.0.0.1:<fixture-port>",
  "ca_file": "",
  "ca_sha256": "",
  "bridge_present": false,
  "source": "manual configuration assertion"
}

For tls, use a matching HTTPS direct endpoint, private CA file and digest, and bridge_present: true; do not reuse an HTTP proof. A proof generated by a suite-owned entry is acceptable only while its endpoints and profile still describe the current run. Missing, extra, null, or mismatched proof fields are setup failures. The client-to-gateway endpoint remains HTTP even in the tls profile; that profile qualifies the direct/upstream TLS path, not end-to-end gateway TLS.

--gateway-pid is optional; when supplied it must be the actual gateway process whose /proc/PID/status is sampled, not a supervisor. A bounded manual measurement may set --duration, --warmup, and --request-timeout to 1s (the minimum accepted value), but it still runs every workload, target, and measured repetition; do not present a shortened run as the complete suite.

Deployment and protocol prerequisites

--deployment standalone|cluster and --route native|translated are required annotations in manual mode and are recorded as metadata; the runner does not infer topology or protocol translation. Standalone and cluster runs, and native and translated routes, must be collected separately. A native route exercises the gateway's native request/stream codec. A translated route deliberately exercises the configured adapter. Do not merge these modes or present them as universal protocol comparisons.

Fixed workload matrix

Every case has three independent measured runs and a separate warmup. Warmup observations are retained but never included in measured percentiles. The runner uses bounded request deadlines and a whole-session deadline; interruption or timeout leaves an incomplete artifact.

  • Unary (unary_1k_1000rps_c64, unary_64k_1000rps_c64): 1 KiB or 64 KiB input, 1 KiB output, offered 1000 requests/s, concurrency 64.
  • Offered SSE (sse_1k_1000rps_c256, sse_64k_1000rps_c256): 1 KiB or 64 KiB input, 128 non-empty events, 1 ms event spacing, offered 1000 requests/s, concurrency 256.
  • Closed-loop SSE (closed_loop_sse_1k_c1, closed_loop_sse_1k_c64, closed_loop_sse_64k_c1, closed_loop_sse_64k_c64): 1 KiB or 64 KiB input, 128 events, concurrency 1 or 64. A new request is admitted only after the previous stream in that slot completes.
  • Slow-reader SSE (slow_sse_1000): 1000 simultaneous streams using 1 KiB input. The runner reads the first event, deliberately leaves response bodies unread during the bounded hold, and then cancels all streams. This case is not replaced by a smaller load when resources are insufficient.

Input contains a benchmark-invocation challenge nonce. Successful output must retain that nonce and the exact fixture shape. An incomplete stream, malformed SSE, HTTP rejection, timeout, or content mismatch is an error. Offered arrivals rejected because all configured concurrency slots are occupied are counted as dropped arrivals; the runner never lowers the rate to make a case pass.

Metrics and artifacts

results.json is written incrementally and atomically with mode 0600. summary.txt is a human-readable index. Each measured record identifies target, workload, run, offered and achieved rates, attempted/completed/error/dropped counts, total latency p50/p95/p99, streaming TTFT p50/p95/p99, warmup counts, and RSS samples where available. Percentiles are computed from real completed requests. Queueing/dispatch lateness and HTTP rejection are retained as separate counts when available.

Gateway RSS is sampled from /proc/<gateway-pid>/status and is unavailable when no PID is supplied or the process exits. Allocation metrics are unavailable: this runner does not substitute its own Go runtime allocation counters for gateway allocations. Environment metadata includes OS/architecture, Go version, CPU count, hostname when available, binary sizes, deployment/route annotations, and timestamp. Secrets and the key-bearing URL are excluded.

A result state of complete_with_failures is still an honest record of the configured workload, not a pass. unmet_prerequisite_or_incomplete means setup, preflight, interruption, or deadline prevented a complete matrix. No native/translated, standalone/cluster, or throughput superiority claim may be made from incomplete or unequal rows.

Container build time and footprint

tools/bench/image_build.py compares a captured baseline Dockerfile and ignore file with the current candidate. It builds separate temporary source contexts, does not edit the checkout, and never prunes shared builder state. Both variants receive identical source mutations. The result images have unique tags and are retained for inspection unless --cleanup-images is explicitly supplied.

python tools/bench/image_build.py --engine podman \
  --baseline-dockerfile .artifacts/image-review-baseline/Dockerfile \
  --baseline-dockerignore .artifacts/image-review-baseline/.dockerignore \
  --output .artifacts/image-build-comparison.json --timeout 600

Capture baseline files before editing them; their hashes are recorded in the report. The runner also accepts --engine docker. The serial matrix covers a cold build, an unchanged warm build, a backend-source comment, a frontend CSS asset change, and a verification-tool-only comment. The frontend probe adds an unused custom property so that it changes the generated asset rather than being discarded as a comment.

“Cold” means image-layer reuse is disabled and the candidate receives a fresh HOORIFIC_CACHE_NAMESPACE. Base images must already be local; registry/proxy and OS caches are not cleared. Later candidate builds reuse that same namespace. The original baseline has no cache mounts. Build duration excludes subsequent image inspection and export, and builds run serially to avoid contention from the comparison itself. Do not run other builds or load tests during measurement.

Timed-out builds and interrupted measurements terminate only the owned process group, with at most ten seconds of bounded cleanup. The report stays incomplete and retains earlier command evidence; cancellation is never reported as a successful build. Optional image cleanup removes only the run's tags without forcing the removal of containers or images still in use.

Image Size/VirtualSize are reported as engine-provided values, not inferred from human-readable history. Podman OCI exports additionally record layer blob bytes and media types; a compressed-layer total is labelled as such only when compression is verified. Docker's archive size is a different metric and is not presented as OCI compressed bytes. These are image-size measurements, not application RSS or throughput benchmarks.

Runtime qualification must accompany a footprint change: migration, readiness, authenticated console operation, and persisted state after restart, under the existing non-root/read-only contract. Certificate trust, timezone loading, identity lookup, and /tmp writability must remain available even when the runtime has no shell.

Historical local reference measurement

The table below records one Linux/amd64 comparison made with Podman 6.1.0 and the same pinned Go/Bun base images. It is the snapshot summarized in the private .artifacts/image-review-proof.json and final report .artifacts/image-build-comparison-final.json (run 20260907t193707z-1478adb3e7, started 2026-09-07T19:37:07.869337Z, finished 2026-09-07T19:41:37.881519Z). Medians are from two serial runs on one machine. They are workload-specific observations, not CI timing guarantees, capacity commitments, or a promise about another builder.

Build case Previous Dockerfile Optimized Dockerfile
Cold, base images already present 54.29 s 46.91 s
Unchanged warm build 10.92 s 10.03 s
Backend-only source change 32.59 s 10.10 s
Frontend CSS asset change 42.38 s 13.37 s
Verification-tool-only change 33.89 s 7.23 s

That local comparison measured 125,752,509 bytes for the previous image and 40,746,582 bytes for the optimized image (67.6% smaller). Verified gzip-compressed OCI layers measured 47,299,273 versus 13,946,682 bytes (70.5% smaller). The delivered Go binaries were byte-identical; the difference came from the runtime filesystem and build graph, not from removing application features or compressing the executable.

Reproduce the comparison with the private output path below after capturing the baseline Dockerfile and ignore file. The command creates temporary source contexts, records hashes, and does not edit the checkout:

python tools/bench/image_build.py \
  --engine podman \
  --baseline-dockerfile .artifacts/image-review-baseline/Dockerfile \
  --baseline-dockerignore .artifacts/image-review-baseline/.dockerignore \
  --output .artifacts/image-build-comparison.json \
  --timeout 600

The baseline files and the output directory in that example are operator inputs; do not treat them as repository-provided evidence. The comparison should be run serially with no other builds or load tests. Runtime qualification must accompany any footprint change: migration, readiness, authenticated console operation, and persisted state after restart under the existing non-root/read-only contract. Certificate trust, timezone loading, identity lookup, and /tmp writability must remain available even when the runtime has no shell.