Goal
Produce a reproducible Claude Code OFF vs MARGINAL Shadow experiment where task success is independently verified.
Negative results are welcome.
Contribution wanted
Build at least one matched Claude Code workload and publish the evidence needed to reproduce it.
Required controls
Freeze as much as practical across OFF and Shadow runs:
- task;
- starting repository snapshot;
- Claude Code version;
- model/version where available;
- agent configuration;
- tool access;
- permissions;
- verifier;
- time limits.
Required measurements
Where genuinely available, report:
- task success;
- governance decisions;
- repeated-work observations;
- false-stop recommendations;
- governance overhead;
- tokens;
- latency;
- tool calls.
Unavailable metrics must remain unavailable rather than being reported as zero.
Scientific discipline
Task quality comes before compute savings.
- If both lanes fail, the experiment must not be presented as evidence of savings.
- If Shadow applies no intervention, any observed token/cost difference must not be attributed causally to MARGINAL.
- Keep infrastructure failures in the evidence bundle.
- Report regressions and negative results.
- Preserve the exact starting state or provide a reproducible fixture.
- Do not cherry-pick only favorable runs.
Privacy
Do not publish prompts, credentials, private repository names, proprietary source code or raw tool output.
Use a public/synthetic workload or a fixture specifically created for evaluation.
Acceptance criteria
- At least one paired OFF/Shadow workload with verified task outcome.
- Same starting state demonstrated for both lanes.
- Exact Claude Code/MARGINAL versions recorded.
- Verifier is deterministic or its limitations are documented.
- Raw aggregate evidence needed for reproduction is committed.
- Governance overhead is reported separately from agent workload cost.
- Result wording matches what the experiment actually proves.
- Existing benchmark claims are not broadened without evidence.
Related work
Do not duplicate the existing cross-platform governance latency issue. This issue is specifically about paired Claude Code task-level OFF vs MARGINAL Shadow evidence.
Goal
Produce a reproducible Claude Code OFF vs MARGINAL Shadow experiment where task success is independently verified.
Negative results are welcome.
Contribution wanted
Build at least one matched Claude Code workload and publish the evidence needed to reproduce it.
Required controls
Freeze as much as practical across OFF and Shadow runs:
Required measurements
Where genuinely available, report:
Unavailable metrics must remain unavailable rather than being reported as zero.
Scientific discipline
Task quality comes before compute savings.
Privacy
Do not publish prompts, credentials, private repository names, proprietary source code or raw tool output.
Use a public/synthetic workload or a fixture specifically created for evaluation.
Acceptance criteria
Related work
Do not duplicate the existing cross-platform governance latency issue. This issue is specifically about paired Claude Code task-level OFF vs MARGINAL Shadow evidence.