Skip to content

benchmark: contribute a verified Claude Code OFF vs MARGINAL Shadow workload #62

Description

@SignalLayerLabs

Goal

Produce a reproducible Claude Code OFF vs MARGINAL Shadow experiment where task success is independently verified.

Negative results are welcome.

Contribution wanted

Build at least one matched Claude Code workload and publish the evidence needed to reproduce it.

Required controls

Freeze as much as practical across OFF and Shadow runs:

  • task;
  • starting repository snapshot;
  • Claude Code version;
  • model/version where available;
  • agent configuration;
  • tool access;
  • permissions;
  • verifier;
  • time limits.

Required measurements

Where genuinely available, report:

  • task success;
  • governance decisions;
  • repeated-work observations;
  • false-stop recommendations;
  • governance overhead;
  • tokens;
  • latency;
  • tool calls.

Unavailable metrics must remain unavailable rather than being reported as zero.

Scientific discipline

Task quality comes before compute savings.

  • If both lanes fail, the experiment must not be presented as evidence of savings.
  • If Shadow applies no intervention, any observed token/cost difference must not be attributed causally to MARGINAL.
  • Keep infrastructure failures in the evidence bundle.
  • Report regressions and negative results.
  • Preserve the exact starting state or provide a reproducible fixture.
  • Do not cherry-pick only favorable runs.

Privacy

Do not publish prompts, credentials, private repository names, proprietary source code or raw tool output.

Use a public/synthetic workload or a fixture specifically created for evaluation.

Acceptance criteria

  • At least one paired OFF/Shadow workload with verified task outcome.
  • Same starting state demonstrated for both lanes.
  • Exact Claude Code/MARGINAL versions recorded.
  • Verifier is deterministic or its limitations are documented.
  • Raw aggregate evidence needed for reproduction is committed.
  • Governance overhead is reported separately from agent workload cost.
  • Result wording matches what the experiment actually proves.
  • Existing benchmark claims are not broadened without evidence.

Related work

Do not duplicate the existing cross-platform governance latency issue. This issue is specifically about paired Claude Code task-level OFF vs MARGINAL Shadow evidence.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarkReproducible evaluation workclaude-codeClaude Code integrationevidenceEvidence quality, attribution and validationhelp wantedExtra attention is needed

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions