Observability stack for local Ollama inference: Prometheus, Grafana, LiteLLM, PostgreSQL, and a custom Python exporter.
| Mode | Prerequisites | Start |
|---|---|---|
| Podman (recommended) | Podman only | LINUX.md / PODMAN.md |
| Conda + binaries (local dev) | Conda, Ollama, portable binaries | LINUX.md or SETUP.MD (Windows) |
| Platform | Setup guide |
|---|---|
| Linux | LINUX.md — Podman (portable) or native Conda |
| Windows | SETUP.MD or PODMAN.md |
| Architecture | architecture.md |
Full setup guide: see SETUP.MD for Windows step-by-step commands, troubleshooting, and notes from implementation.
Client
-> LiteLLM :4000 (API keys, usage logging)
-> Ollama :11434
-> PostgreSQL (request_logs, usage_events)
-> Exporter :8000 (metrics + direct generate)
-> Ollama :11434
-> Prometheus metrics + JSONL/CSV logs
Prometheus :9090 -> Grafana :3000
PostgreSQL -> NSE portal :3080 (dashboards, API keys)
No Conda, Ollama, or PostgreSQL on the host. Full steps: LINUX.md.
cd llm-monitoring-poc
chmod +x scripts/podman/*.sh
./scripts/podman/build.sh
./scripts/podman/up.sh
./scripts/podman/pull-models.sh
./scripts/podman/status.sh.\scripts\podman\build.ps1
.\scripts\podman\up.ps1
.\scripts\podman\pull-models.ps1See PODMAN.md for offline image transfer to another machine.
| Tool | Purpose |
|---|---|
| Conda | Python env for the exporter |
| Ollama | Local LLM inference |
After installing Miniconda, initialize PowerShell (once), then open a new terminal:
conda init powershellcd llm-monitoring-poc
chmod +x scripts/*.sh scripts/lib/stack.sh
./scripts/setup.sh
ollama pull tinyllama && ollama pull phi3:mini
./scripts/start-stack.sh
./scripts/status.sh
./scripts/generate-traffic.shSee architecture.md for full Linux commands, diagrams, and component roles.
cd llm-monitoring-poc
# If the broken partial env exists, remove it first (optional but clean):
conda env remove -n llm-monitoring-poc -y
.\scripts\setup.ps1
.\scripts\start-stack.ps1
# One-time setup (conda env + Grafana download + dashboards)
.\scripts\setup.ps1
# Pull both monitored models
ollama pull tinyllama
ollama pull phi3:mini
# Start everything: Prometheus, Grafana, exporter (+ LiteLLM via start-litellm.ps1)
.\scripts\setup.ps1 # first time only
.\scripts\start-stack.ps1
.\scripts\status.ps1 # verifies Grafana, Prometheus, exporter, Ollama
#.\scripts\start-stack.ps1
# Clean failed Grafana extract (if present)
Remove-Item -Recurse -Force .tools\_extract-grafana -ErrorAction SilentlyContinue
Remove-Item -Recurse -Force "$env:TEMP\lm-graf" -ErrorAction SilentlyContinue
# Reinstall binaries (Grafana may take a few minutes)
.\scripts\install-tools.ps1
# Start stack
.\scripts\start-stack.ps1
# Generate sample traffic for both models
.\scripts\generate-traffic.ps1
# Health check
.\scripts\status.ps1Open http://localhost:3000 (login admin / admin).
Dashboards: Dashboards → LLM Observability
| Dashboard | Description |
|---|---|
| LLM — TinyLlama | Metrics for tinyllama only |
| LLM — Phi-3 Mini | Metrics for phi3:mini only |
| LLM — All Models | Comparison charts + per-model rows + model filter |
| LLM Observability POC | Same as All Models (legacy title) |
Stop the stack:
.\scripts\stop-stack.ps1If Grafana install failed with long Remove-Item / Expand-Archive errors, clean up and reinstall tools:
Remove-Item -Recurse -Force .tools\_extract-grafana -ErrorAction SilentlyContinue
.\scripts\install-tools.ps1The NSE portal at http://localhost:3080 shows the API Keys & Users table with token usage, request counts, latency, and host GPU/CPU metrics. Use the Light mode / Dark mode button in the header to switch themes (Models tab Grafana iframe follows the same theme). Data is stored in PostgreSQL (or SQLite fallback at .data/llm_monitoring.db when USE_SQLITE=true).
After changing Grafana dashboards, regenerate and restart:
py scripts/build-dashboards.py
.\scripts\stop-stack.ps1
.\scripts\start-stack.ps1| Test project | Model | Virtual API key |
|---|---|---|
| Project A | tinyllama | sk-tiny-A-1 |
| Project B | tinyllama | sk-tiny-B-2 |
| Project C | phi3:mini | sk-phi3-C-1 |
| Project D | phi3:mini | sk-phi3-D-2 |
# One-time DB + dashboards (after setup.ps1)
.\scripts\init-db.ps1
# Real traffic for all 4 projects (updates dashboard live)
.\scripts\generate-project-traffic.ps1
# Many parallel requests (default 5 per key; skips disabled keys like Project B)
.\scripts\generate-bulk-traffic.ps1 -CountPerKey 10
# Or single project via curl / PowerShell:
curl -X POST http://localhost:8000/generate -H "X-API-Key: sk-tiny-A-1" -H "Content-Type: application/json" -d "{\"model\":\"tinyllama\",\"prompt\":\"Hello\"}"
Invoke-RestMethod "http://localhost:8000/test?model=tinyllama&api_key=sk-tiny-A-1"
# LiteLLM proxy on :4000 (logs usage to PostgreSQL)
.\scripts\start-litellm.ps1Grafana: Dashboards → LLM — API Keys & Usage (llm-api-keys)
PostgreSQL (optional):
$env:DATABASE_URL = "postgresql://llm:llm@localhost:5432/llm_monitoring"
$env:USE_SQLITE = ""
.\scripts\init-db.ps1| Service | URL |
|---|---|
| NSE Portal (recommended) | http://localhost:3080 |
| NSE API keys | http://localhost:3080/api/keys |
| NSE Prometheus metrics | http://localhost:3080/metrics |
| Grafana | http://localhost:3000 |
| Prometheus | http://localhost:9090 |
| LiteLLM | http://localhost:4000 |
| Exporter | http://localhost:8000 |
| Ollama | http://localhost:11434 |
Invoke-RestMethod "http://localhost:8000/test?model=tinyllama"
Invoke-RestMethod "http://localhost:8000/test?model=phi3:mini"
Invoke-RestMethod -Method Post -Uri "http://localhost:8000/generate" `
-ContentType "application/json" `
-Body '{"model":"phi3:mini","prompt":"Explain LLM observability briefly."}'
Invoke-RestMethod "http://localhost:8000/requests?limit=20"
Invoke-RestMethod "http://localhost:8000/slow-prompts?limit=10"Per-request logs: logs/requests.jsonl, logs/requests.csv
llm-monitoring-poc/
environment.yml # Conda: Python + prometheus-client
monitoring/ # DB store, infra metrics, Prometheus helpers
db/
schema.sql
seed.sql
litellm/
config.yaml
db_usage_callback.py
nse-portal/
server.py # API + /metrics + Grafana proxy
dashboard.js
ops-dashboard.css
ollama-exporter/
exporter.py # Gateway, metrics, logs, API-key attribution
requirements.txt
prometheus/
prometheus.yml # Scrape localhost:8000
alert-rules.yml
grafana/
provisioning/ # Datasources + dashboard provider
dashboards/ # Generated JSON dashboards
compose.yaml # Podman Compose stack
Containerfile # Python services image (exporter + portal)
LINUX.md # Linux: Podman (portable) + native Conda
PODMAN.md # Container quick start + offline image export
scripts/
podman/ # build, up, down, export-images, pull-models
setup.ps1 / setup.sh # Full one-time setup (Windows / Linux)
install-tools.ps1|.sh # Download Prometheus, Grafana to .tools/
start-stack.ps1|.sh # Start all services
stop-stack.ps1|.sh
status.ps1|.sh
generate-traffic.ps1|.sh
lib/stack.ps1|stack.sh # Shared stack helpers
build-dashboards.py
architecture.md # Diagrams, Prometheus/Grafana/LiteLLM roles, file graph
logs/ # Request history (gitignored)
.tools/ # Grafana binaries (gitignored)
.data/ # TSDB + Grafana data (gitignored)
After editing scripts/build-dashboards.py or adding models in MODELS:
py -3 scripts/build-dashboards.py
# or: conda run -n llm-monitoring-poc python scripts/build-dashboards.pyRestart the stack so Grafana reloads JSON files.
Loaded from prometheus/alert-rules.yml. View at http://localhost:9090/alerts
LLMExporterDownOllamaDownLLMHighFailureRateLLMHighP95LatencyLLMNoRecentSuccess
Set LOG_PROMPT_TEXT=false before starting the exporter (in scripts/start-stack.ps1 or your shell) to log only prompt_hash and prompt_preview.
Replace this PoC exporter with metrics from LiteLLM, vLLM, or TGI on real servers. Keep Prometheus, Grafana, PostgreSQL, and GPU telemetry (DCGM) as the platform layer.
See OBSERVABILITY_REVIEW.md for gap analysis and production roadmap.