Skip to content

Latest commit

 

History

History
63 lines (48 loc) · 2.58 KB

File metadata and controls

63 lines (48 loc) · 2.58 KB

AdaptLoop Demo Guide

Closed-loop story in one command: signals → stream → curated dataset → LoRA → gated promote → canary → rollback.

One-command demo

# Python 3.12 + Docker Desktop
cp .env.example .env
make install          # once
make demo             # full loop (~2–4 min with smoke LoRA)
make interview        # 30s pitch + rollback proof on existing run

Defaults:

  • Uses tiny HF model (--smoke) — no GPU / no 3B download
  • INJECT_FAULT=1 so canary fails on purpose and triggers auto-rollback (interview punchline)
  • Healthy path: make demo-healthy

Talk track: INTERVIEW.md

Optional gateway proof (LLMOps running + migrated through 0005_adapter_canary_percent):

export LLMOPS_GATEWAY_URL=http://localhost:8000
export LLMOPS_ADMIN_API_KEY=llmops_dev_default_key
SKIP_GATEWAY=0 make demo

What you should see

Step Proof
Spine training_examples rows; Redis adaptloop:agg:*
Curation .artifacts/datasets/<run>/train.parquet + holdout + review_queue
Train .artifacts/runs/<run>/adapter + MinIO checkpoints
Promote MLflow UI http://localhost:5001 — model in Staging; holdout includes judge_score
Canary (fault) adaptloop-canary exits non-zero; rollback archives + disables gateway route
Gateway (optional) X-Adapter-Id + sticky canary_percent split when calling model: "epc-qa"

Interview failure modes (talk track)

  1. Training/serving skew — shared adaptloop/normalization + golden tests; same chat template in train and synthetic serve path.
  2. Feedback poisoning — schema validation; needs_review rows go to review_queue.parquet, never into train.parquet.
  3. Consumer lag / stale promotion — promote gate refuses when Redis produced - processed lag is high.
  4. Canary regression — Staging promote watched via error_rate + latency_p95; breach → MLflow Archived + gateway route disable.
  5. Bad generations — holdout LLM-judge rubric (optional API judge) feeds judge_score into promote gates.
  6. Risky full cutover — Staging canary_percent sticky-hashes a % of alias traffic before full Production.

Manual slice commands

make spine
make curate
make train
make promote
make canary
make gateway   # needs LLMOps

Resume one-liner

AdaptLoop — Event-driven ML adaptation platform: Redpanda ingest of production LLM traces, Bytewax stream aggregation, PostgreSQL/pgvector curation, LoRA fine-tuning with MLflow registry, and lag/canary-gated adapter promotion into a multi-tenant inference gateway.