Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion docs/decisions/ADR-0001-eval-runner.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,9 @@
# ADR-0001: Custom eval runner (TS/Bun) with harness adapters

Status: Accepted
Status: Superseded
Date: 2026-07-05
Summary: Use a thin TypeScript-on-Bun eval runner with declarative YAML cases, real harness adapters, pinned model matrices, repeated trials, and outcome-based checks.
Superseded by: ADR-0010

## Context

Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
# ADR-0010: Extract the evaluation runner into Sevro

Status: Accepted
Date: 2026-09-19
Summary: Move generic evaluation infrastructure into the versioned Sevro package while Darrow retains its evaluation policy, cases, and product integration.
Supersedes: ADR-0001

## Context

Darrow's shared evaluation runner has grown beyond the thin repository helper
selected by ADR-0001. It now contains reusable execution, isolation, host
adapters, grading, evidence, and reporting infrastructure alongside policy and
assets that are specific to Darrow.

Keeping both responsibilities in `evals/runner/` couples runner development and
dependencies to marketplace changes. It also prevents other projects from
using the generic runner without adopting Darrow's source tree and evaluation
policy.

The standalone runner has a separate durable owner in
[`bjro/sevro`](https://github.com/BjRo/sevro). Darrow still needs to own the
evaluation cases and policy that define evidence for this repository.

## Decision

Extract generic evaluation infrastructure into Sevro and consume it through a
versioned public package.

- Sevro owns the execution engine, fixture and check mechanics, isolation,
cancellation, run ownership, evidence persistence, cleanup, generic
reporting, host adapters, public schemas, and their generic unit tests.
- Darrow owns its evaluation extension, cases, suites, conditions, corpus
definitions, snapshots, normative specifications, plugin-local assets, and
product integration tests.
- Darrow integrates through Sevro's public CLI and extension protocol rather
than importing unpublished runner internals or copying runner source.
- Normal Darrow use pins an exact Sevro release and does not require a Sevro
source checkout or Git metadata.
- Coordinated development may select an explicit local Sevro checkout. Evidence
from such a run records the local development identity separately from a
packaged release identity.
- Evaluation tooling remains repository development infrastructure. No
marketplace plugin runtime gains a Bun, TypeScript, or Sevro dependency.
- The migration is staged. Generic implementation remains in Darrow until
compatibility is proven through public interfaces and the pinned package can
replace it without losing required behavior or evidence.

This decision establishes ownership and distribution boundaries. Package
naming and registry details, extension transport and framing, exact public JSON
schemas, and the CLI exit-code mapping remain open for the applicable
specifications and implementation planning.

## Consequences

- Sevro can evolve and release generic evaluation infrastructure independently
of the Darrow marketplace.
- Darrow can evolve its policy, cases, and extension with the capabilities they
assess while consuming a reproducible runner release.
- Coordinated changes require compatibility checks across both repositories and
evidence that distinguishes packaged releases from local development builds.
- Extraction work must introduce explicit project and configuration roots plus
configurable result and active-run storage; package installation paths cannot
stand in for project state.
- Existing commands, evidence, isolation boundaries, and historical results
need behavior-parity coverage before Darrow switches to the package. Any
deliberate incompatibility must be documented as a migration.
- ADR-0001 remains the historical record of the original in-repository runner
choice but no longer governs its long-term ownership.
6 changes: 4 additions & 2 deletions docs/decisions/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,23 +16,23 @@ Rebuild it with `decision catalog rebuild` and verify it with `decision catalog
<!-- prettier-ignore -->
| Decision | Status | Date | Summary | Relationships |
| --- | --- | --- | --- | --- |
| [ADR-0001: Custom eval runner (TS/Bun) with harness adapters](ADR-0001-eval-runner.md) | Accepted | 2026-07-05 | Use a thin TypeScript-on-Bun eval runner with declarative YAML cases, real harness adapters, pinned model matrices, repeated trials, and outcome-based checks. | None |
| [ADR-0002: Separate capabilities from orchestration](ADR-0002-separate-capabilities-from-orchestration.md) | Accepted | 2026-08-13 | Keep capabilities intent-matched and independently selectable while starting continuation-owning orchestration only through explicit user invocation. | None |
| [ADR-0003: Treat plugins as optionality boundaries](ADR-0003-treat-plugins-as-optionality-boundaries.md) | Accepted | 2026-08-13 | Treat each plugin as a self-contained unit of adoption, compatibility, and ownership that composes through host-visible skill intent rather than sibling dependencies or a separate capability registry. | None |
| [ADR-0004: Use native goal ownership for core orchestration](ADR-0004-use-native-goal-ownership-for-core-orchestration.md) | Accepted | 2026-08-13 | Use Darrow to compile and launch one host-native goal owner instead of operating a second execution controller or general workflow runtime. | Revisit when: Matched multi-trial evidence on supported hosts shows that a Darrow-owned execution controller materially improves task outcomes over native goal ownership after accounting for wall time, model usage, child invocations, and human interruptions. |
| [ADR-0006: Keep decisions with their authoritative owners](ADR-0006-keep-decisions-with-their-authoritative-owners.md) | Accepted | 2026-08-13 | Keep each decision at the narrowest durable authoritative owner its consumers obey, with one canonical sink per effect and honest gaps for inaccessible owners. | None |
| [ADR-0007: Separate skill evaluation evidence dimensions](ADR-0007-separate-skill-evaluation-evidence-dimensions.md) | Accepted | 2026-08-13 | Represent invariant coverage, task outcomes, matched skill ablation, and skill activation as separate evaluation evidence dimensions. | None |
| [ADR-0008: Allow Python and UV for Langfuse observability](ADR-0008-allow-python-and-uv-for-langfuse-observability.md) | Accepted | 2026-08-24 | Permit the independently installable Langfuse observability plugin to use a locked Python backend managed by UV while retaining a portable Bash hook launcher and keeping the exception scoped to that plugin. | Revisit when: Codex exposes equivalent native Langfuse export, the Langfuse SDK no longer requires Python, or the plugin can meet its rollout-reconstruction and export contract with the portable Bash baseline alone. |
| [ADR-0009: Adopt Python and UV for substantial plugin mechanics](ADR-0009-adopt-python-and-uv-for-substantial-plugin-mechanics.md) | Accepted | 2026-09-17 | Adopt contained Python packages managed by UV incrementally for substantial cross-platform plugin mechanics, invoking them directly from installed skill-relative paths. | Supersedes: ADR-0005; Revisit when: UV and supported Python cannot provide independently installable helpers across every supported native host, or a lighter common runtime offers materially better portability and containment. |
| [ADR-0010: Extract the evaluation runner into Sevro](ADR-0010-extract-the-evaluation-runner-into-sevro.md) | Accepted | 2026-09-19 | Move generic evaluation infrastructure into the versioned Sevro package while Darrow retains its evaluation policy, cases, and product integration. | Supersedes: ADR-0001 |

<!-- darrow-source: 1a57cf608fb4cfa4b770abebf34ca450e79c0160 2711262332 330 ADR-0001-eval-runner.md -->
<!-- darrow-source: cd9996efb2f66b1602ead9a405a48686e14f67a9 1820997609 370 ADR-0002-separate-capabilities-from-orchestration.md -->
<!-- darrow-source: 452327e39a9dff558c6897356d0e1e85cf65906f 3673420192 418 ADR-0003-treat-plugins-as-optionality-boundaries.md -->
<!-- darrow-source: 42bc1095831a4d097a6bda6462d29315bc0f0529 2382776868 628 ADR-0004-use-native-goal-ownership-for-core-orchestration.md -->
<!-- darrow-source: 739b2fa46ad8b0ed220c4d341317bf670d3d7a41 243106577 398 ADR-0006-keep-decisions-with-their-authoritative-owners.md -->
<!-- darrow-source: ae5948d48fa4b0b0ce2c80ea76e23a71d582aa71 4074249863 369 ADR-0007-separate-skill-evaluation-evidence-dimensions.md -->
<!-- darrow-source: aab600b869cf357f81f4424f1813a75596d0a64c 2479862079 646 ADR-0008-allow-python-and-uv-for-langfuse-observability.md -->
<!-- darrow-source: 152ebad0159783e7f2299a3abb6df5b3d3ae6176 3731273062 623 ADR-0009-adopt-python-and-uv-for-substantial-plugin-mechanics.md -->
<!-- darrow-source: 7e9e833caf7a83145daacfa932b43da384f1ef93 4046602581 376 ADR-0010-extract-the-evaluation-runner-into-sevro.md -->

## Rejected

Expand All @@ -51,6 +51,8 @@ Rebuild it with `decision catalog rebuild` and verify it with `decision catalog
<!-- prettier-ignore -->
| Decision | Status | Date | Summary | Relationships |
| --- | --- | --- | --- | --- |
| [ADR-0001: Custom eval runner (TS/Bun) with harness adapters](ADR-0001-eval-runner.md) | Superseded | 2026-07-05 | Use a thin TypeScript-on-Bun eval runner with declarative YAML cases, real harness adapters, pinned model matrices, repeated trials, and outcome-based checks. | Superseded by: ADR-0010 |
| [ADR-0005: Use portable Bash facades for plugin mechanics](ADR-0005-use-portable-bash-facades-for-plugin-mechanics.md) | Superseded | 2026-08-13 | Put deterministic plugin mechanics behind narrow portable Bash facades while keeping authority and contextual judgment in skills. | Superseded by: ADR-0009 |

<!-- darrow-source: e85924b1a6bbbb6f57670af8b500c6227f55bc42 1322776374 340 ADR-0001-eval-runner.md -->
<!-- darrow-source: 4121fc1ae3de0d07c230af478cf687fca0ff8766 3926343988 378 ADR-0005-use-portable-bash-facades-for-plugin-mechanics.md -->
Loading