OpenAI-compatible LLM gateway for multi-provider routing, reliability, and observability — built for AI Infra / SRE / DevOps portfolio work.
- OpenAI-compatible API —
POST /v1/chat/completions(JSON + SSE streaming) - Multi-provider routing — route by model name with optional fallback
- Providers — OpenAI API, Ollama (any OpenAI-compatible upstream)
- Health probes —
GET /health,GET /ready - Structured logging — JSON logs with request ID
- Optional gateway auth —
Authorization: Bearer <GATEWAY_API_KEY>
- Redis rate limiting — token-bucket per API key / IP (
429+Retry-After) - Circuit breaker — per-provider failure isolation with half-open recovery
- Mock provider — fast local responses for load testing
- k6 load test — see
docs/load-test.md
- Prometheus metrics — request rate, latency, tokens, errors, circuit breaker state, rate limit stats
- Grafana dashboard — pre-provisioned with 11 panels (QPS, latency, tokens, errors, CB state, etc.)
- Metrics endpoint —
GET /metricsfor Prometheus scraping
- CI/CD — GitHub Actions: lint, test, build, Docker push to GHCR
- K8s manifests — Deployment, Service, ConfigMap, Secret, HPA, PDB, Ingress
- Redis StatefulSet — Persistent Redis for rate limiting
- Monitoring stack — Prometheus + Grafana deployed to K8s
- Prometheus alerts — HighErrorRate, HighLatency, CircuitBreakerOpen, HighRateLimitRejection
- Runbooks — deployment, troubleshooting, scaling, monitoring
cp config.example.yaml config.yaml
export OPENAI_API_KEY=sk-... # optional if using Ollama onlymake build
./bin/gateway -config config.yaml# Health
curl http://localhost:8080/health
curl http://localhost:8080/ready
# Chat (Ollama — start Ollama first)
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.2",
"messages": [{"role": "user", "content": "Hello"}]
}'
# Streaming
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.2",
"stream": true,
"messages": [{"role": "user", "content": "Hello"}]
}'Use any OpenAI SDK by pointing base_url to http://localhost:8080/v1.
See config.example.yaml.
| Section | Purpose |
|---|---|
server.addr |
Listen address (default :8080) |
gateway.api_key |
Optional client auth for the gateway |
metrics |
Prometheus endpoint (enabled, path) |
rate_limit |
Redis token-bucket (enabled, redis_url, rpm, burst) |
circuit_breaker |
Per-provider breaker thresholds |
providers |
Upstream LLM backends |
routing |
Model → provider (+ optional fallback) |
Environment variables in config use ${VAR} syntax.
cmd/gateway/ Entrypoint
internal/config/ YAML config loading
internal/model/ OpenAI-compatible types
internal/provider/ Upstream adapters
internal/router/ Model routing + fallback
internal/handler/ HTTP handlers
internal/middleware/ Logging, request ID, rate limit
internal/circuitbreaker/ Provider circuit breaker
internal/ratelimit/ Redis token bucket
internal/metrics/ Prometheus metrics
internal/gateway/ HTTP server wiring
scripts/ k6 load test
deploy/ Docker Compose + Prometheus + Grafana
k8s/ Kubernetes manifests
docs/ Architecture & runbooks
docker run --rm -p 6379:6379 redis:7-alpine
make run-loadtest
k6 run scripts/loadtest.jsDetails: docs/load-test.md
cp config.example.yaml config.yaml
docker compose -f deploy/docker-compose.yml up --buildServices: gateway (:8080), redis (:6379), prometheus (:9090), grafana (:3000).
# View metrics
curl http://localhost:8080/metrics
# Grafana dashboard (Docker stack)
open http://localhost:3000 # admin/adminThe gateway's own single-service scrape config is deploy/prometheus/prometheus.yml (used by the deploy/docker-compose.yml stack). For the full AI inference platform observability rollout, deploy/prometheus/prometheus.integration.yml scrapes all three services:
| Job | Target | Metrics path |
|---|---|---|
ai-gateway |
gateway:8080 |
/metrics |
enterprise |
enterprise-backend:8000 |
/metrics |
vllm-adapter |
vllm-adapter:8000 |
/api/system/metrics |
The integration config requires the three service DNS names to be reachable from Prometheus (same Docker network / Kubernetes namespace). The companion deploy/grafana/dashboards/ai-platform-overview.json dashboard shows the six minimum panels: three-service HTTP QPS, 5xx rate, p95 latency, and Gateway provider latency / fallback rate / embedding rates.
Note: Prometheus/Grafana are not currently running in the WSL development environment (no Docker/GPU). The integration YAML and dashboards are validated statically by
go test ./...(internal/monitoring); the real stack is pending the Docker environment.
| Phase | Focus |
|---|---|
| 1 ✅ | Core gateway, routing, streaming, health |
| 2 ✅ | Redis rate limit, circuit breaker, load test |
| 3 ✅ | Prometheus metrics + Grafana dashboard |
| 4 ✅ | CI/CD, K8s manifests, runbooks |
# 1. Create the secrets from your real values (see k8s/secret.example.yml)
kubectl create secret generic ai-gateway-secrets \
--namespace ai-gateway \
--from-literal=OPENAI_API_KEY='sk-...' \
--from-literal=GATEWAY_API_KEY='...' \
--dry-run=client -o yaml | kubectl apply -f -
# 2. Deploy to K8s
kubectl apply -f k8s/namespace.yml
kubectl apply -f k8s/configmap.yml
kubectl apply -f k8s/redis/
kubectl apply -f k8s/monitoring/
kubectl apply -f k8s/See docs/runbooks/deployment.md for detailed instructions.
MIT