- Training pipeline lives in the
src/distillpackage (python -m distill.<module>); root-level scripts are legacy and being retired. - Services are standalone under
services/<name>/— the api service must not import fromsrc/distill(keeps torch out of its Docker image). - Python: snake_case modules. Treat 200 LOC as a review signal; split only when it produces a real responsibility boundary.
- Follow ecosystem conventions for non-Python files (
App.tsx,Dockerfile,README.md); use kebab-case for new project-owned files when no stronger convention exists. - Imports at top, grouped: stdlib → third-party → local.
- Training tunables in
src/distill/config.py, all env-overridable; secrets only via.env(gitignored) — never hardcoded, never printed. - Service config in
services/api/app/config.py, env-only (container-friendly).
- Teacher API calls: classify retryable vs fatal, exponential backoff + jitter,
validate outputs (empty/short/U+FFFD rejected) — see
distill/teacher_client.py. - Dataset writes are atomic (temp file + rename), resumable across runs.
- File I/O: explicit
encoding="utf-8"; console viadistill/logging_utils.py. - Training: let it crash for visibility; fix root cause (see phase-03 incident log).
python -m pytest tests/ -q(package) andservices/api: python -m pytest tests/ -q(fake runtime — no GPU/model needed); web:pnpm test+pnpm build.ruff check src/ tests/ services/api/must be clean (CI gate).docs/openapi.yamlis the manually maintained frontend contract. The web client is generated from it (pnpm run generate-client); CI checks YAML → TypeScript drift, not FastAPI → YAML drift.
- Conventional commits:
feat:,fix:,ci:,build:,docs:,refactor:,test:,chore: - No AI refs in commit messages.
- Focused commits; no secrets, no private notes, no model artifacts.
- bf16 end-to-end — the base checkpoint is bf16 and fp16 conversion at load crashes the current torch nightly on Windows (see phase-03 incident log).
- Model loads go CPU-first, then
.to("cuda:0")— safetensors direct-to-GPU is broken against the current torch nightly. - LoRA on full bf16 (rank env-overridable via
LORA_R; v0.5/v0.6 ran r=16, v0.7 ran r=32 to test capacity), targets q/k/v/o, with gradient checkpointing; the QLoRA 4-bit path is kept behindLOAD_IN_4BITbut bnb on-the-fly quantization crashes in this environment. - MAX_SEQ_LENGTH 512 — vocab-sized logits (152K) OOM 6GB at 1024.
- adamw_8bit optimizer; batch 1 × grad-accum 8; validation split with early stopping and best-checkpoint restore.
- 570 prompts / 10 categories (v0.5: 530; v0.6/v0.7: +40 for weak categories); every sample validated at generation time.
- Exact Qwen2.5 chat template with
<|im_start|>/<|im_end|>special tokens (a plain-text approximation trains a format no inference stack uses). - Quality gate before splitting: mojibake, min-length, duplicate instructions.
- Stratified 80/10/10 train/validation/test, seed 42, splits disjoint by id.
- Perplexity on the held-out test split (never in-sample), per category.
- ROUGE-L + token-F1 vs teacher reference; optional LLM-as-judge via 9Router.
- Reports versioned under
plans/reports/evaluation-<label>.mdwith baseline comparison and explicit caveats.