Closed-loop story in one command: signals → stream → curated dataset → LoRA → gated promote → canary → rollback.
# Python 3.12 + Docker Desktop
cp .env.example .env
make install # once
make demo # full loop (~2–4 min with smoke LoRA)
make interview # 30s pitch + rollback proof on existing runDefaults:
- Uses tiny HF model (
--smoke) — no GPU / no 3B download INJECT_FAULT=1so canary fails on purpose and triggers auto-rollback (interview punchline)- Healthy path:
make demo-healthy
Talk track: INTERVIEW.md
Optional gateway proof (LLMOps running + migrated through 0005_adapter_canary_percent):
export LLMOPS_GATEWAY_URL=http://localhost:8000
export LLMOPS_ADMIN_API_KEY=llmops_dev_default_key
SKIP_GATEWAY=0 make demo| Step | Proof |
|---|---|
| Spine | training_examples rows; Redis adaptloop:agg:* |
| Curation | .artifacts/datasets/<run>/train.parquet + holdout + review_queue |
| Train | .artifacts/runs/<run>/adapter + MinIO checkpoints |
| Promote | MLflow UI http://localhost:5001 — model in Staging; holdout includes judge_score |
| Canary (fault) | adaptloop-canary exits non-zero; rollback archives + disables gateway route |
| Gateway (optional) | X-Adapter-Id + sticky canary_percent split when calling model: "epc-qa" |
- Training/serving skew — shared
adaptloop/normalization+ golden tests; same chat template in train and synthetic serve path. - Feedback poisoning — schema validation;
needs_reviewrows go toreview_queue.parquet, never intotrain.parquet. - Consumer lag / stale promotion — promote gate refuses when Redis
produced - processedlag is high. - Canary regression — Staging promote watched via error_rate + latency_p95; breach → MLflow Archived + gateway route disable.
- Bad generations — holdout LLM-judge rubric (optional API judge) feeds
judge_scoreinto promote gates. - Risky full cutover — Staging
canary_percentsticky-hashes a % of alias traffic before full Production.
make spine
make curate
make train
make promote
make canary
make gateway # needs LLMOpsAdaptLoop — Event-driven ML adaptation platform: Redpanda ingest of production LLM traces, Bytewax stream aggregation, PostgreSQL/pgvector curation, LoRA fine-tuning with MLflow registry, and lag/canary-gated adapter promotion into a multi-tenant inference gateway.