From 8a83046fb7bf45b68b6572848e38aeb7d55029f5 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Bj=C3=B6rn=20Rochel?= Date: Sat, 19 Sep 2026 15:58:25 +0200 Subject: [PATCH] docs(decisions): record eval runner extraction boundary --- docs/decisions/ADR-0001-eval-runner.md | 3 +- ...xtract-the-evaluation-runner-into-sevro.md | 68 +++++++++++++++++++ docs/decisions/README.md | 6 +- 3 files changed, 74 insertions(+), 3 deletions(-) create mode 100644 docs/decisions/ADR-0010-extract-the-evaluation-runner-into-sevro.md diff --git a/docs/decisions/ADR-0001-eval-runner.md b/docs/decisions/ADR-0001-eval-runner.md index 1a57cf60..e85924b1 100644 --- a/docs/decisions/ADR-0001-eval-runner.md +++ b/docs/decisions/ADR-0001-eval-runner.md @@ -1,8 +1,9 @@ # ADR-0001: Custom eval runner (TS/Bun) with harness adapters -Status: Accepted +Status: Superseded Date: 2026-07-05 Summary: Use a thin TypeScript-on-Bun eval runner with declarative YAML cases, real harness adapters, pinned model matrices, repeated trials, and outcome-based checks. +Superseded by: ADR-0010 ## Context diff --git a/docs/decisions/ADR-0010-extract-the-evaluation-runner-into-sevro.md b/docs/decisions/ADR-0010-extract-the-evaluation-runner-into-sevro.md new file mode 100644 index 00000000..7e9e833c --- /dev/null +++ b/docs/decisions/ADR-0010-extract-the-evaluation-runner-into-sevro.md @@ -0,0 +1,68 @@ +# ADR-0010: Extract the evaluation runner into Sevro + +Status: Accepted +Date: 2026-09-19 +Summary: Move generic evaluation infrastructure into the versioned Sevro package while Darrow retains its evaluation policy, cases, and product integration. +Supersedes: ADR-0001 + +## Context + +Darrow's shared evaluation runner has grown beyond the thin repository helper +selected by ADR-0001. It now contains reusable execution, isolation, host +adapters, grading, evidence, and reporting infrastructure alongside policy and +assets that are specific to Darrow. + +Keeping both responsibilities in `evals/runner/` couples runner development and +dependencies to marketplace changes. It also prevents other projects from +using the generic runner without adopting Darrow's source tree and evaluation +policy. + +The standalone runner has a separate durable owner in +[`bjro/sevro`](https://github.com/BjRo/sevro). Darrow still needs to own the +evaluation cases and policy that define evidence for this repository. + +## Decision + +Extract generic evaluation infrastructure into Sevro and consume it through a +versioned public package. + +- Sevro owns the execution engine, fixture and check mechanics, isolation, + cancellation, run ownership, evidence persistence, cleanup, generic + reporting, host adapters, public schemas, and their generic unit tests. +- Darrow owns its evaluation extension, cases, suites, conditions, corpus + definitions, snapshots, normative specifications, plugin-local assets, and + product integration tests. +- Darrow integrates through Sevro's public CLI and extension protocol rather + than importing unpublished runner internals or copying runner source. +- Normal Darrow use pins an exact Sevro release and does not require a Sevro + source checkout or Git metadata. +- Coordinated development may select an explicit local Sevro checkout. Evidence + from such a run records the local development identity separately from a + packaged release identity. +- Evaluation tooling remains repository development infrastructure. No + marketplace plugin runtime gains a Bun, TypeScript, or Sevro dependency. +- The migration is staged. Generic implementation remains in Darrow until + compatibility is proven through public interfaces and the pinned package can + replace it without losing required behavior or evidence. + +This decision establishes ownership and distribution boundaries. Package +naming and registry details, extension transport and framing, exact public JSON +schemas, and the CLI exit-code mapping remain open for the applicable +specifications and implementation planning. + +## Consequences + +- Sevro can evolve and release generic evaluation infrastructure independently + of the Darrow marketplace. +- Darrow can evolve its policy, cases, and extension with the capabilities they + assess while consuming a reproducible runner release. +- Coordinated changes require compatibility checks across both repositories and + evidence that distinguishes packaged releases from local development builds. +- Extraction work must introduce explicit project and configuration roots plus + configurable result and active-run storage; package installation paths cannot + stand in for project state. +- Existing commands, evidence, isolation boundaries, and historical results + need behavior-parity coverage before Darrow switches to the package. Any + deliberate incompatibility must be documented as a migration. +- ADR-0001 remains the historical record of the original in-repository runner + choice but no longer governs its long-term ownership. diff --git a/docs/decisions/README.md b/docs/decisions/README.md index 0be7786a..a2697478 100644 --- a/docs/decisions/README.md +++ b/docs/decisions/README.md @@ -16,7 +16,6 @@ Rebuild it with `decision catalog rebuild` and verify it with `decision catalog | Decision | Status | Date | Summary | Relationships | | --- | --- | --- | --- | --- | -| [ADR-0001: Custom eval runner (TS/Bun) with harness adapters](ADR-0001-eval-runner.md) | Accepted | 2026-07-05 | Use a thin TypeScript-on-Bun eval runner with declarative YAML cases, real harness adapters, pinned model matrices, repeated trials, and outcome-based checks. | None | | [ADR-0002: Separate capabilities from orchestration](ADR-0002-separate-capabilities-from-orchestration.md) | Accepted | 2026-08-13 | Keep capabilities intent-matched and independently selectable while starting continuation-owning orchestration only through explicit user invocation. | None | | [ADR-0003: Treat plugins as optionality boundaries](ADR-0003-treat-plugins-as-optionality-boundaries.md) | Accepted | 2026-08-13 | Treat each plugin as a self-contained unit of adoption, compatibility, and ownership that composes through host-visible skill intent rather than sibling dependencies or a separate capability registry. | None | | [ADR-0004: Use native goal ownership for core orchestration](ADR-0004-use-native-goal-ownership-for-core-orchestration.md) | Accepted | 2026-08-13 | Use Darrow to compile and launch one host-native goal owner instead of operating a second execution controller or general workflow runtime. | Revisit when: Matched multi-trial evidence on supported hosts shows that a Darrow-owned execution controller materially improves task outcomes over native goal ownership after accounting for wall time, model usage, child invocations, and human interruptions. | @@ -24,8 +23,8 @@ Rebuild it with `decision catalog rebuild` and verify it with `decision catalog | [ADR-0007: Separate skill evaluation evidence dimensions](ADR-0007-separate-skill-evaluation-evidence-dimensions.md) | Accepted | 2026-08-13 | Represent invariant coverage, task outcomes, matched skill ablation, and skill activation as separate evaluation evidence dimensions. | None | | [ADR-0008: Allow Python and UV for Langfuse observability](ADR-0008-allow-python-and-uv-for-langfuse-observability.md) | Accepted | 2026-08-24 | Permit the independently installable Langfuse observability plugin to use a locked Python backend managed by UV while retaining a portable Bash hook launcher and keeping the exception scoped to that plugin. | Revisit when: Codex exposes equivalent native Langfuse export, the Langfuse SDK no longer requires Python, or the plugin can meet its rollout-reconstruction and export contract with the portable Bash baseline alone. | | [ADR-0009: Adopt Python and UV for substantial plugin mechanics](ADR-0009-adopt-python-and-uv-for-substantial-plugin-mechanics.md) | Accepted | 2026-09-17 | Adopt contained Python packages managed by UV incrementally for substantial cross-platform plugin mechanics, invoking them directly from installed skill-relative paths. | Supersedes: ADR-0005; Revisit when: UV and supported Python cannot provide independently installable helpers across every supported native host, or a lighter common runtime offers materially better portability and containment. | +| [ADR-0010: Extract the evaluation runner into Sevro](ADR-0010-extract-the-evaluation-runner-into-sevro.md) | Accepted | 2026-09-19 | Move generic evaluation infrastructure into the versioned Sevro package while Darrow retains its evaluation policy, cases, and product integration. | Supersedes: ADR-0001 | - @@ -33,6 +32,7 @@ Rebuild it with `decision catalog rebuild` and verify it with `decision catalog + ## Rejected @@ -51,6 +51,8 @@ Rebuild it with `decision catalog rebuild` and verify it with `decision catalog | Decision | Status | Date | Summary | Relationships | | --- | --- | --- | --- | --- | +| [ADR-0001: Custom eval runner (TS/Bun) with harness adapters](ADR-0001-eval-runner.md) | Superseded | 2026-07-05 | Use a thin TypeScript-on-Bun eval runner with declarative YAML cases, real harness adapters, pinned model matrices, repeated trials, and outcome-based checks. | Superseded by: ADR-0010 | | [ADR-0005: Use portable Bash facades for plugin mechanics](ADR-0005-use-portable-bash-facades-for-plugin-mechanics.md) | Superseded | 2026-08-13 | Put deterministic plugin mechanics behind narrow portable Bash facades while keeping authority and contextual judgment in skills. | Superseded by: ADR-0009 | +